The actual documents the agents read and work from, shown exactly as they are on disk — not a summary. See the progress view instead · All projects
# PLAN.md — THE COHERENCE FACTORY: every feature always improving, bugs fixed on their own, a person planning only what recurs
Owner: Nick. Purpose: plan and how-it-works for every agent, on every machine. Status as of 2026-09-27.
This file is the one record. It lives in the brain repo (`nick-deck/deck-brain-2`) under `projects/business/coherence-factory/`; the Hub's own guide (`deck-business/CLAUDE.md` §7) points here. From any machine: `gh api repos/nick-deck/deck-brain-2/contents/projects/business/coherence-factory/PLAN.md -H "Accept: application/vnd.github.raw"`.
**🔴🔴 THIS IS THE ONLY PLANNING DOCUMENT FOR THIS PROJECT. Do not create a second plan, tracker, summary, or scratch state file — extend THIS file. Any status view is GENERATED from this plan; if a view disagrees with the plan, the plan wins.**
**NORTH STAR:** the requester, 2026-09-27: *"i basically want each of our features always improving - plus looking for new ideas to make the whole UX more valuable and streamlined based on whats going on in the ecosystem both for the target audience and the evolution of software ai tech etc - user issues and bug get logged and fixed automatically and recurring issues get a human involved to plan a fix"* and *"assume we dont have anyone that reads code yet so reviews need to be set up for non developers to review - the whole process needs to work for a team that knows the product and user very well but not the tech - tech problems probably get solved by multiple agent reviews like triads across astra fable opus for hard problems and lower models or single or dual reviews for simpler and medium difficult problems."* Finished looks like: Mae is on the Candidates screen on her phone, something is off, she tells Screen Buddy "the tag button does nothing" and goes back to work. An hour later her Hub shows a card in her own words with a before-and-after picture, one sentence on what changed, a note that two independent AI reviewers checked it, and the fix is already live because it only put back what was meant to work; her tap closes the card. When Dean asks for the huddle list to sort differently, the same card comes back with **Try it** switched on for him alone, and his tap on **Looks good** is what makes it live for everyone; his **Send back** with a sentence reopens it. When the same thing breaks a second time, or a fix needs three tries, the screen's owner gets one card that says what was tried, offers two or three numbered options with their effort, and asks for a number. Every Monday each screen's owner finds at most three improvement ideas for their screen waiting in their Inbox, and Nick and Chantelle find one eight-line note on what changed out in the world for the people Coherence is for and in the AI and software it runs on, each line pointing at a Coherence screen. Nobody on the team reads a line of code, and the numbers on one tile say how many issues came in, how many fixed themselves, how many needed a person, how long from report to live, and what it cost.
**FINISH LINE:**
1. From any Hub screen, on a phone and on a computer, a team member says one sentence to Screen Buddy ("something's wrong: …" or "I wish this …") and within a minute a card exists on the **Coherence Improvement** board carrying the screen, the open record, their words and the screen summary, owned by the agent `factory`; proven as Mae on a phone and as Dean on a computer, read back from the card by id.
2. A page error nobody reports is caught by PostHog (live in the Hub since 2026-08-09), grouped into one issue per root cause, and lands as one factory row within one runner tick, with the issue's count, first-seen and last-seen (and the replay of the session that raised it once replay is switched on by Nick's or Chantelle's dated word); the same error twice makes one row with count 2, never two rows; proven with an injected test error on the live Hub on an `agent-test` record; the exception that leaves the page keeps its stack frames, carries a blank message value and nothing matching the floor shapes, and while replay is off no recording request leaves the page (when on, a page carrying a marked-secret element replays as masked text only).
3. A seeded easy bug on a test card (one screen, one file) goes report → fix → a red-first harness that failed before the fix and passes after → checked by a model that did not build it → driven in a browser as the reporter → merged and published → the reporter's card in plain words with the before-and-after pictures, with no human step in between, and the run's clock from card to live is written on the card; proven once end to end on an `agent-test` issue and the clock recorded as the baseline (the target is under two hours; the first run sets the number).
4. A seeded medium **change** (new or altered behaviour) stops at **Ready to look at** with the plain-words Look page and its flag on for the reviewer alone; the reviewer's **Looks good** switches it on for everyone; **Send back** with a sentence reopens it and the sentence reaches the builder's next brief; proven as Dean on the live Hub with an `agent-test` change, both branches.
5. The classifier puts any change touching a TRUNK path, a money path or client-facing words (the hard-coded lists and rules in §3, contract C3) in the **hard** tier, and the runner merges a hard-tier change only with a recorded three-seat review (Fable proposes, Astra attacks the premise, Opus reads cold) bound to the reviewed commit; proven red-first by a change to `app/css/one.css` that tries to skip the review and is refused with the reason named, and by a stale record for another commit that is refused; the tile counts any commit on `main` that bypassed the runner.
6. The same fingerprint reopened twice, a fix reverted, three issues on one screen in fourteen days, or a fix that hit its bounds each produce exactly one **Plan with a person** card to that screen's owner on Coherence Improvement, and a review failing its third loop produces exactly one card to Nick on the **Ecosystem Fixes** board (BUILD §1a): what was tried, two or three numbered options with how long until live, one recommendation, answerable with a number; proven with a seeded recurrence, one card and not two.
7. Every day the runner (not a hand run) files at most three improvement ideas per owner as **suggestions** in that owner's Inbox when there is new evidence, and up to three **What's new out there** lines to Nick and Chantelle, each naming a Coherence screen or a gap and a recommendation; every night it files a correction to its own sizing rules or briefs for anything sent back or undone; a dismissed suggestion is remembered by the suggestion store and never refiled; proven by runner-written heartbeat rows and the suggestions read back by id. Daily until Nick says otherwise.
8. A **Factory** tile on Home for Nick and Chantelle only shows issues in, fixed on their own, needing a person, median time from report to live, send-back rate and cost, every number derived from the issue rows and the run records, none typed; proven by a run whose numbers a checker recomputes from the rows, and by Mae and Rizza not seeing it.
9. Every §2 row is driven on a real signed-in session by a checker that did not build it; the manifest reads 100%; the final sign-off is Fable or Opus reading the closed steps' proofs.
**Owner:** the Fable session that wrote this plan (overseer; opens the card and drives) · **Overseer:** ONE — Fable; never builds · **Design authority:** none — nothing new is drawn: every surface reuses the Hub's task panel, the shared Look · Test · Approve card and the Home tile wall; the one new page (the Look page, STEP 7) is text, pictures and the Hub's own tokens
> **STEP 0 — ARM THE LOOP, BEFORE ANYTHING ELSE.** Set a 5-minute loop. Every time it fires, answer
> these five in order and CORRECT any failure before doing anything else:
> 1. **NORTH STAR** — is what I am doing this minute moving this plan's North Star? If not, drop it.
> 2. **FAN-OUT** — declare the whole actual roster, dispatch useful ready work, and shed your own unnecessary processes. Coordinate through peers or the launching dispatcher; no numeric cap or load-wait rule applies.
> 3. **CHEAP** — are cheap models doing the building AND the per-step checking? If anything on
> Anthropic or OpenAI is building or checking a step, move it down now (§M).
> 4. **STUCK** — for anything I have called blocked: name the input that does not exist yet, or the
> three concrete things I tried. If I cannot, it is not blocked — drive through it now.
> 5. **NEXT** — did something just finish? Then the next step whose inputs exist starts THIS minute.
> A finished step is never a place to stop, a report is never a reason to wait, and Nick being
> away or asleep is the reason to keep going, not to pause.
> Then keep building. The loop never stops until the FINISH LINE is proven.
**Rule: a step starts the moment its named inputs exist, whatever its number. A step closes on ONE independent check by a different model. Nothing waits on Nick to test.**
**Working directory for every DONE-PROOF:** Hub steps run in `projects/business/business-app/` (the Hub's own repository, `nick-deck/deck-business`); brain steps in the brain repo (`nick-deck/skippy-code`, the running copy on the Mac mini, landed with `projects/ops/bin/brain-land.sh`; the workspace's copy under `projects/personal/skippy-app/skippy-code/` is stale and is not the surface); scheduler steps from the workspace root (`nick-deck/deck-brain-2`, `$WORKSPACE`). Browser proofs use the local server `node harness/pm/serve.mjs` (serves `app/dist` at `http://127.0.0.1:8788`) and are serialised through `/tmp/bzvisual-chrome.lock`; anything that may outlive ten minutes is launched through `projects/ops/skippy-jobs/lib/run-detached.sh start … -- <command>` and polled with its `wait`.
## 🔴 STEP 19 — IMPORT READINESS for the regrouped feature lanes (opened Fri 2 Oct 2026, ~9:30 pm Central)
**Why:** Nick, 2 Oct ~9:20 pm, to this lane: *"id like you to stay fanned out as much as possible until this dev facoty project is ready to import the lanes we are regrouping right now - always act as overseer not builder"*. The regroup (projects/ops/project-files/PLAN.md STEP 9; design `projects/ops/project-files/evidence/regroup-2026-10-02/design/design-locked.txt` §6-§8) gives each of 29 features an owner agent that runs improvement rounds THROUGH this factory. Feature ids and their cards/channels: `projects/ops/project-files/evidence/regroup-2026-10-02/pm-lane/feature-cards.json` (30 rows; `audience-plus` is a sub-manual of `content-production` by Nick's word ~8:55 pm, "audience+ can be subdoc", so 29 lanes).
**Overseer:** this session (Opus 5.5) writes briefs, checks, lands; it never types code. **Builders:** the cheap lane (route-build / cheap-task, launched through run-detached.sh). **Checker:** an independent verifier agent per piece, plus the piece's red-first proof.
**Done when:** I1-I7 are live and proven, and one owner agent can run one round (load its packet, file a row with its lane, see its lane's numbers) on a test row.
| Piece | What | Files | Red-first proof |
|---|---|---|---|
| I1 lane field | Factory rows carry `lane` (one of the 29 ids, or "" = unmapped); unknown values refused | Hub: new `app/functions/api/_factory-lanes.js` (LANES list), `_factory-issues.js` (PATCHABLE + row default + validation) | `_factory-issues.selftest.mjs`: patch lane "crm-inbox" accepted; "nope" refused 400; a new row has lane "" |
| I4 per-lane numbers | `GET /api/factory-metrics?lane=<id>` returns the tile's numbers for that lane only; no lane = whole factory as today; unknown lane 400 | Hub: `factory-metrics.js` + its harness | a seeded test row in lane A is not counted under lane B |
| I2 owner packet | a new script, owner-packet (in the factory folder, to be created), given a lane id prints one JSON packet: manual + open steps, open/dismissed factory rows of the lane, page errors / Screen Buddy reports / failed checks for its screens, benchmark list, "What we tried", test identities, per-lane numbers; anything missing is `"unknown"`, never omitted | brain: the new owner-packet script and its red-first test (both to be created in projects/ops/factory) | a lane with no manual yet yields `manual: "unknown"`, not a crash and not `{}` |
| I3 briefing pack | a versioned `briefing-pack-<lane>@<sha>` text (manual sections + "don't retry" lines) attached to every dispatch route: factory build/check briefs, route-build/cheap-task, Codex reviews | brain: `factory/briefs/fix-build.md` placeholder + the dispatch helpers (named per route when briefed) | a factory build brief for a lane row contains the pack's version line |
| I5 reconciliation events | every merge to main touching a feature's paths (and every rollback) queues one durable event per lane: dedupe key, retries, ack; a manual-only change queues none | brain: publish stage `lib/factory-tool-publish.mjs` + a small queue in the factory state | the same merge twice → one event; a manual-only commit → zero |
| I6 claims | Suite card `claims` (acquire / renew / release / expiry) read by the merge gate as a wait | Hub card field + brain `lib/factory-tool-merge.mjs` | an unexpired claim on a fix's file → wait, not attempt; expired → proceeds |
| I7 observer + scout | `factory-observe` / `factory-scout` scheduled entries file real suggestions per lane (STEP 15) | brain jobs | a seeded lane with new evidence → ≤3 suggestions, re-run → no duplicates |
**STATUS (2 Oct 2026, ~11:00 pm Central, overseer):**
- I1 lane field — LIVE on the Hub (#4457, f53b27ad31): rows carry `lane` (29 ids in `_factory-lanes.js`).
- I4 per-lane numbers — LIVE (#4457): `/api/factory-metrics?lane=<id>`.
- I2 owner packet — on main (d35e509b18): `owner-packet.mjs <lane>`, tested on the real shop, jev and content-production manuals (jev's manual has no SCREENS AND DOORS block; regroup told).
- I3 briefing pack — on main (cc9f4cf4c7, 2b4bf5737d, c6b4725f77): versioned pack from the manual, appended to a lane row's build brief; wired into the runner. Not yet on the check (reader) briefs.
- I5 change notices — on main (f3b53e1fbe, 8496c7309d, c6b4725f77): durable deduplicated queue; the publish stage queues a notice when a lane's fix goes live. Not yet: rollbacks, and other lanes' merges (needs the screen-to-lane map).
- I6 claims — factory side on main (e8be0cd6c7): claims its fix files before merging, 409/unreadable = wait; dormant until the PM lane's claims door (D1, branch pm-card-claims) is live and the door client gains claimAcquire/claimRelease.
- I7 observer/scout — not started. Screen-to-lane map — test building.
- STEP 11: first fix merged by the factory (#4454) and went live, then wrongly auto-reverted (#4456) by a live check comparing against a newer local build; check fixed (afbb5dc4c5); row reopened, fresh round at build.
- Runner restart owed once STEP 11 passes publish (loads the briefing pack / queue wiring).
**Order:** wave 1 now, in parallel: I1, I4, I2 (skeleton that marks missing sources unknown). Wave 2 after I1 is live: I5, I6, I3. Wave 3: I7, then the screen-to-lane map from the manuals' SCREENS AND DOORS blocks (agreed with the regroup session in #agents thread msg-86c6c8a3, 2 Oct ~8:05 pm). STEP 11 keeps running underneath (test row fx-4e8dead05fef at merge, waiting on a green Hub publish).
## 🔴 HANDED BACK TO A LOCAL SESSION — Fri 2 Oct 2026, ~5:05 pm (Central)
**Pick-up:** read this block, then the STEPS list. The first action is STEP 11 on the Mac Studio, item 1 of "Next" below.
- **Where it stands:**
- Live on the Hub: #4236 (the review door's `recur_screen`, plus the beacon harness listed as a by-hand proof).
- On brain main:
- the drive machine-fault wait (`factory-fix.mjs`, b88a461);
- the monitor's PostHog recurrence count (`factory-monitor.mjs`, 1ec79a2);
- the daily recurrence job wired live (`factory-recur.mjs`, 4a8cbe9): it asks `recur_screen` once per screen with three issues in 14 days, and sends a problem reopened twice to a person through the plan action. Test-marked rows only until OPENED.
- Test row `fx-4e8dead05fef` is still `plan` (since 4:57 am).
- No PR of this lane is open. The cloud session left nothing uncommitted; its scratch files go with the session.
- **Done this run:**
- STEP 12's third-screen trigger, live and reviewed.
- STEP 12's reopened-twice trigger, reviewed and landed.
- STEP 5's beacon harness, reviewed. Its live run is owed.
- The drive's machine-fault wait and the monitor's PostHog count.
- **Next, in order:**
1. **Mac Studio: finish STEP 11.**
- Confirm the job runner is on brain main at 4a8cbe9 or later, and restart it if not.
- Set row `fx-4e8dead05fef` from `plan` to `drive` through the factory door, with the factory key.
- Let the scheduled fix job take it through drive, merge, publish and the hand-off to Mae.
- Then run `node projects/ops/factory/prove-e2e.mjs --issue fx-4e8dead05fef`.
2. **Mac: the beacon harness, live.** In the Hub repo, run `FACTORY_AGENT_KEY=<from the vault> node _selfchecks/harness-factory-beacon-20260927.mjs`, then the same with `--mutant=one`, which must FAIL case 4.
3. **Nick's word on OPENED** once STEP 11 passes.
4. **STEP 2:** the Screen Buddy core lane's draft door.
5. **STEP 12's last piece:** the third-failed-loop card on Ecosystem Fixes. It needs a Hub door that makes a card on that board.
- **Tried and failed in the cloud** (each a reason this work was handed back):
- The browser does not trust the session proxy, and widening that trust was refused by the safety check.
- Brain files fetched loose are not in a git repository, so the cheap lane refuses them; each edit was a recorded override.
- The safety check refused changing the Studio-owned test row.
- The safety check refused reading a build result while a PR edited CLAUDE.md.
- **Needs a person:**
- Nick: open the factory to real reports after STEP 11, yes or not yet.
- Nick: the draft cloud allow rules (1, a narrow test-row reset tool; 2, browser trust for the beacon harness; 3, restoring build side-effect files), if cloud sessions are to stop less.
## 🔴 HOW THIS LANE MERGES NOW — from 2 Oct 2026, ~2:10 pm (Central)
The merge desk is closed (Nick, 2 Oct 2026). This lane merges its own Hub pull requests under RULE 64 and posts no more @merge-desk hand-offs. Before a merge, all measured:
1. **Proof and review.** The proof is red-first and green, and an independent reviewer's MERGE is written ON THE PR: a body line or comment naming the reviewer, the head it read, and MERGE.
2. **Checks.** In a private worktree at the PR head, with origin/main merged in, run `hub-api-selftest-diff.mjs` and `hub-tier1-diff.mjs` (clean main copy vs this copy, `--jobs 6`), launched through `run-detached.sh`. In a cloud session the TIER-1 side-by-side is asked of a Mac session in #agents.
3. **Read the "red on both" list**, and re-run a newly red gate alone on both copies.
4. **Merge only the tested head:** `node projects/ops/toolkit/land/hub-pr-land.mjs --pr <n>`, after `gh pr view <n> --json headRefOid` equals the tested commit.
5. **After the publish,** the live `/api/build-integrity` `source_revision` must carry the merge commit. A red live check is reverted at once by a new PR.
No PR of this lane is open: #4236 merged and is live. The next Hub PR, if any, follows this process.
**Ownership note for the next step.** `jobs/factory-recur.mjs` (calling the live `recur_screen`) is not in this cloud session's named brain files (`factory-fix.mjs`, `_test-factory-fix.mjs`, `factory-monitor.mjs`, this plan). The next session that is handed it builds the call with a red-first case in `_test-factory-recur.mjs`.
## 🔴 LIVE — Fri 2 Oct 2026, 12:23 pm (Central)
- **#4236 is merged and live.** The merge desk merged it as 267bbb00f7 at 11:52 am.
- **Live proof, measured 12:22 pm:**
- the live `/api/build-integrity` `source_revision` 78920e3ce carries the merge;
- `/api/health` answers 200 with JSON;
- the live review door answers `recur_screen` from a person with 403 "the factory agent only", where an unknown action answers 400.
- **One check could not run here:** `verify-live.mjs` compares the live site with a local build of the same commit, and this session's local build is older, so it reads stale. That is a limit of this session, not a live fault.
- **The merge desk was told** in its thread.
- **Next, in order:**
1. Brain: `jobs/factory-recur.mjs` calls the live door's `recur_screen` for each new row, so the third-screen card fires on its own. This is this lane's work, with a red-first test.
2. A Mac session runs the beacon harness live with `FACTORY_AGENT_KEY`, then `--mutant=one` (must FAIL case 4).
3. The Mac Studio finishes STEP 11 (row `fx-4e8dead05fef`).
## 🔴 HANDED TO THE MERGE DESK — Fri 2 Oct 2026, ~11:30 am (Central)
- **#4236 is ready and handed to the merge desk** in Hub Chat #agents. The thread is `msg-3ce6af7f-377d-452c-b047-927be8ef6ec0`; post every later word on #4236 there with `--thread`. Chat name: `coherence-factory-lane`.
- **Head 7bb470ae4, current with main.** The diff is four files: factory-review.js, its selftest, the beacon harness, and the one gates.js `NOT_A_CHECK` line.
- **Proofs:** selftest 64/64, red 10 on main's door; gate-coverage PASS; an independent reader said MERGE.
- **Triad of 2 Oct, done:**
- The 12 fixtures and the skip log the cloud build rewrote were restored from git.
- The 2 new payroll-adjustment fixtures it wrote were moved to the session scratch folder. They go when the session ends and are regenerated by the build.
- The Hub guide edit was taken out of #4236.
- **Proposed Hub guide wording, for HUB-DATA STEP 4's documentation step.** These are the lines that went out with commit 6ca37f34b on this branch; read them with `git show 6ca37f34b -- CLAUDE.md`. They cover §3 line 53 and §7 lines 73–74: building in the cloud through `cloud-build-host.sh`, publishing on a `hub-publisher` runner, and the side-by-side check in the cloud. Since 2 Oct merges go through the merge desk, so that wording must say "hand off to the merge desk", not "merges itself".
- **Next after the merge:**
1. Check that the live `/api/build-integrity` `source_revision` carries the merge commit.
2. Wire `jobs/factory-recur.mjs` to call `recur_screen`.
3. Run the beacon harness live, then with `--mutant=one`.
## 🔴 OFF-THE-MAC TRIAD — Fri 2 Oct 2026, ~10:55 am (Central)
Nick, 2 Oct 2026: "i dont want things to be dependant on local machines - triad the best solution and act".
The proposal was attacked (a Sonnet seat) and cold-read (an Opus seat): PROCEED WITH CHANGES. The decisions:
1. **Cloud lanes build, run the side-by-side check and merge in the cloud** under RULE 64, then run `verify-live.mjs` after the edge window and revert a failed live check at once.
- The build host was measured in this session: `Hub build host: ready (data 52 min old)`.
- The first cloud fast build of #4236 BLOCKED on gate-coverage: the beacon harness was in no tier. That is a real fault, caught in the cloud. It is fixed: the harness joins `NOT_A_CHECK` beside `cc-live-walk.mjs`, an edit made by the cheap lane (zai, attempt 1).
- The Hub guide's stale lines (§3 line 53, §7 lines 73–74) are corrected in the same pull request (6ca37f34b).
- **Held:** the session's safety check refused to let this session read the second build's result once the pull request edits the Hub guide (CLAUDE.md), classing it as self-modification. #4236 is NOT merged; the decision is with Nick.
2. **Publishing** stays on any `hub-publisher` runner (the Studio or Chantelle's Mac mini; deploy.yml lines 100–107) until HUB-DATA STEP 4. This lane builds no second path.
3. **STRUCK: moving the factory's jobs to a cloud Routine.** The measured blockers, handed to HUB-DATA STEP 7, which owns that move:
- the jobs-role gate (`lib/factory-wire.mjs:43`, `machine-roles.json` "jobs" = Nicks-Mac-Studio);
- the gitignored state folder (`factory-wire.mjs:54`);
- the hard-coded `/Users/nickdeck/actions-runner/...` checkout (`factory-wire.mjs:55`);
- the drive's Mac arm64 headless-shell Chrome (`factory/factory-drive.mjs:57`);
- no Codex in a cloud session (asserted, not measured);
- KV has no compare-and-swap, so a host lock would not hold.
4. **The beacon harness stays a by-hand live proof**, for HUB-DATA STEP 7's measure.
5. **Cheap edits of brain files from the cloud.** The gap is exact: `cloud-brain.sh` keeps a shallow, blob-less, sparse git cache and copies files out of it, so the copies in the workspace are not inside a git repository and the cheap lane refuses them. The owner of `cloud-brain.sh` is unconfirmed; it is not edited here.
## 🔴 CLOUD STOP — Fri 2 Oct 2026, 9:30 am (Central)
- **Where it stands:** STEP 12's third-screen trigger and the drive's machine-fault wait are built, proven and reviewed. The monitor now counts PostHog recurrences. STEP 5's beacon harness is written and reviewed but has not run live yet. The Hub half waits in draft pull request nick-deck/deck-business#4236 (branch `factory-recur-third-screen`) for a Mac session's build, side-by-side check and merge.
- **Done this run:**
- The Hub review door gains `recur_screen`. Three issues on one screen in 14 days (same test mark, `never` rows left out) make ONE card to the screen's owner, listing every issue. The screen is claimed for 14 days, so a repeat or a fourth issue makes none. A failed card write releases the claim. Selftest: 10 new cases, `FAIL 10` exit 1 on the old door, 64/64 after.
- The beacon harness, rewritten to F6's scope. It throws the same digit-free `factory-test` error twice on a live page and asserts: one test row from PostHog with its count up by two; frames kept; the exception floor-clean; no recording request; the fallback beacon counts. It refuses without `FACTORY_AGENT_KEY`. Cleanup sets every touched row to `never`. `--mutant=one` must fail.
- Brain: in `jobs/factory-fix.mjs` a drive machine fault is a wait, not an attempt, a hand-off or a slot. `_test-factory-fix.mjs` 15/15, red 5 on the old job. Landed b88a461 and ad79e8c.
- Brain: `jobs/factory-monitor.mjs` reads `posthog-personal-api-key` by name and counts occurrences since publish through `occurrencesSince`. A PostHog error counts 0 and is said in the beat. Landed 1ec79a2.
- **Next step:** on a Mac:
1. Run the beacon harness with `FACTORY_AGENT_KEY` set, then `--mutant=one` (must FAIL case 4).
2. Run the Hub TIER-1 side-by-side on #4236 and merge it.
3. Wire `jobs/factory-recur.mjs` to call the door's `recur_screen` for each new row, so the trigger fires on its own.
- **Tried and failed:**
- The low-cost build lane cannot run in a cloud session: cheap-build runs git in the fetched brain folder, which is not a git repository (exit 128). Each edit was recorded as a cheap-vendor-failed override.
- The beacon harness cannot run in a cloud session: the browser does not trust the session's proxy, and widening browser trust was refused by the session's safety check.
- **Needs a person:** none. Merging #4236 and running the harness are Mac-session work, not a decision.
- **Cheap lane, 9:55 am Central (this cloud session):** the governance refresh was run; `grep -c statusSnapshot` in cheap-build reads 3. A trial cheap edit on the brain file `projects/ops/skippy-jobs/jobs/factory-monitor.mjs` (to remove an unused `phNote` variable) was refused by design before any vendor was called. The last lines were:
```
because : cleared by the data wall
handing off…
job : rb-mur401rl-rkdj4q
🔴 REFUSED — no git repository holds projects/ops/skippy-jobs/jobs/factory-monitor.mjs (looked from its folder and from the workspace root), so a stray write could not be detected — refusing before any vendor is called
🔴 THE CHEAP LANE TRIED AND FAILED. It has already reverted the file to its snapshot.
```
So in a cloud session, brain files fetched by `cloud-brain.sh get` cannot be cheap-edited, while Hub files can. The `phNote` cleanup was not done and no override was recorded for it; it is harmless and can ride the next Mac edit of that file. Pull request #4236 carries the one-line note that the overrides are no longer needed.
- **Waiting on a Mac (unchanged):** finishing STEP 11 (row fx-4e8dead05fef back through drive, merge, publish, hand-off to Mae, prove-e2e); the Hub build and side-by-side check; the hard-tier readers; the job runner restart; Slack as Skippy; any of the four acts.
- **Reviewer's optional notes, for the next builder:**
- On the Mac rig, the harness rebuilds request bodies from puppeteer's text `postData()`, which can damage gzip. If case 1 reads "0 seen" there, look at that first, and use CDP `postDataEntries` or turn compression off.
- `recur_screen` keeps its claim when the card is made but the hand-off fails, and nothing retries that hand-off.
- A PostHog error in the monitor beats OK, not WARN.
## 🔴 HANDOFF TO CLOUD — Fri 2 Oct 2026, 8:55 am
**1 · Pick-up:** `projects/business/coherence-factory/PLAN.md`, this block ("HANDOFF TO CLOUD — Fri 2 Oct 2026, 8:55 am"); first thing to do: on a Hub branch `factory-recur-third-screen`, build STEP 12's missing trigger (a third screen report in 14 days makes one plan card) with its selftest case red first.
**2 · State, measured 8:52 am local:**
- **Live:** the Hub serves revision 17cf0cfffc. It carries every Hub change of this project: #4029 (seeded test bug), #4050 and #4123 and #4151 (Factory tile on Home), #4077 (plain-words plan exit), #4091 (guide pointer), #4142 (PostHog test mark). Nothing of this project is merged and still publishing.
- **Brain main:** all factory code and this plan; last factory code commit 051dc1cac3 (the drive's picture check). The Studio's shared copy has it, and the Studio restarted at about 7:50 am, so its job runner is on that code.
- **Open branch, no pull request:** Hub `factory/fx-4e8dead05fef-a2` (806b82ed4) — the factory's own fix for test row `fx-4e8dead05fef` plus its harness. It is NOT for a person or a cloud session to merge: the point of STEP 11 is that the factory merges it itself.
- **Test row `fx-4e8dead05fef`:** status `plan` since 4:40 am, attempts 0, reviewer Mae, a test row. Its two working folders on the Studio (`fx-4e8dead05fef-a2`, `-a2-base`) are the factory's own.
- **Running:** nothing of the handing-off session. The Studio's scheduled factory jobs (triage, fix, monitor) run as built, on test rows only.
- **Switches:** `ARMED = true`, `OPENED = null` in `projects/ops/skippy-jobs/lib/factory-first-live.mjs` — test rows only, not opened to real reports. Unchanged since 1 Oct night.
- **Latest tests:** full factory suite (25 kit sections, 5 standalone tests) 0 fails at 6:48 am; the third seed's drive run by hand at 6:58 am reads "nothing outside the band changed" at 375 and 1280.
**3 · Done this session:** see the block below, "LANDED FOR MACHINE RESTART — Fri 2 Oct 2026, 7:05 am", part 3 (unchanged since): 13 live faults found by three test seeds and fixed red-first; before, no seed left the inbox; after, seed three passed harness, build and both readers on its own and its screen check grades clean.
**4 · Postmortem:** the same block, part 4. One addition: **the landing left the first real factory merge undone at the last stage.** Cause: each picture-check guess cost a 90-second live run and the restart came before the row could be sent round again. Rule: when a live seed is one stage from the finish, send the row round before polishing anything else.
**5 · Next steps, ranked:**
1. **STAYS ON A MAC — finish STEP 11:** reset row `fx-4e8dead05fef` from `plan` to its drive stage through the factory door, let the Studio's scheduled fix job take it through drive, merge, publish and hand-off to Mae, then `node projects/ops/factory/prove-e2e.mjs --issue fx-4e8dead05fef`.
2. **Cloud — STEP 12:** the Hub review door's third-screen-in-14-days trigger (`app/functions/api/factory-review.js`, selftest `_factory-review.selftest.mjs`): exactly one plan card to the screen owner, a repeat makes none.
3. **Cloud — STEP 5:** the beacon harness `_selfchecks/harness-factory-beacon-20260927.mjs` in the Hub repo (two identical injected errors make one row with count 2); then the monitor's PostHog occurrence count in brain `projects/ops/skippy-jobs/jobs/factory-monitor.mjs` using `lib/factory-posthog.mjs` `occurrencesSince`.
4. **Cloud — brain, one file and its test:** make a drive machine fault a wait, not an attempt and not a plan hand-off (`projects/ops/skippy-jobs/jobs/factory-fix.mjs`, drive stage, "unmeasurable"; proof in `projects/ops/skippy-jobs/_test-factory-fix.mjs`, red first).
5. **STAYS ON A MAC:** STEP 9's one live hard-tier run through the three seats; STEP 10's independent VERIFIED line and heartbeats; STEP 16's checker re-run; STEP 17 sign-off and the audit rounds.
**STAYS ON A MAC (the Coherence Factory lane on the Mac Studio does these; a cloud session must not try):**
- Everything in STEP 11: the factory's fix job, its drive (a local Hub plus Chrome), its merge and the Hub publish all run on the Studio.
- The Hub build and the pre-merge side-by-side check: the cloud build script (cloud-build-host, under the Hub's scripts folder) is not on the Hub's main (HUB-DATA STEP 1 not landed, measured 8:52 am), so every cloud pull request waits for a Mac session to run that check and merge.
- The hard-tier readers (Astra through the Studio's Codex accounts, Claude from the Studio's command line), the job runner restart, and any Slack message as Skippy.
**6 · Held for a person:**
- **Opening the factory to real reports** (`OPENED`): Nick's; default: stays on test rows only.
- **Mae's Done tap on the first live test fix:** Mae's, once STEP 11 hands it to her; default: the machine proof stands.
- The test plan card for `fx-4e8dead05fef` handed at 4:40 am is a test card; the Mac resume carries it forward or archives it with the toolkit's archive-test-cards tool.
## 🔴 LANDED FOR MACHINE RESTART — Fri 2 Oct 2026, 7:05 am
**1 · Pick-up prompt:** `Read projects/business/coherence-factory/PLAN.md, the block "LANDED FOR MACHINE RESTART — Fri 2 Oct 2026, 7:05 am", then restart the job runner on the shared copy and send test row fx-4e8dead05fef back through the drive so the first factory fix reaches the live Hub (STEP 11); /finish with a 5-minute loop, no updates until done.`
**2 · State, measured at landing (7:05 am local):**
- **Running:** nothing of this session. The 5-minute loop is cancelled; the local test Hub this session left on port 8790 is stopped; the drive runs and the proof suite have all exited. The Studio's scheduled factory jobs (triage, fix, monitor) keep running as built; they act on test rows only.
- **Switches:** the factory is ARMED on test rows only and NOT opened to real reports (`lib/factory-first-live.mjs`: `ARMED = true`, `OPENED = null`). Opening is a person's commit, never an agent's.
- **In the cloud:** workspace main carries everything, last code commit 051dc1cac3 (the drive's picture check). All six Hub pull requests of this run are merged: #4029 (seeded bug), #4050 and #4123 (Factory tile), #4077 (plain-words plan exit), #4091 (guide pointer), #4142 (PostHog test mark). The last publish seen green was build 4ced5c6cc1.
- **The third test fix** (row `fx-4e8dead05fef`, the seeded bug "What you said shows nothing"): status `plan` (handed at 4:40 am after one failed drive), reviewer Mae, a test row. Its fix is correct and proven by hand: harness red then green, the cheap and Sonnet readers both said MERGE, and the drive run by hand at 6:58 am with the landed code reads "nothing outside the band changed" on phone and computer. Its two commits (harness bd26006ca, fix 806b82ed4) are on the Hub branch `factory/fx-4e8dead05fef-a2`. Its two working folders on the Studio (`fx-4e8dead05fef-a2`, `-a2-base` under the Hub's `.claude/worktrees`) are the factory's own and were left for the resume.
- **Not yet picked up:** the Studio's shared copy was at d44cfb2aeb, behind 051dc1cac3. It pulls on its own; the runner caches code, so it needs `bash projects/ops/skippy-jobs/restart-daemon.sh` after the pull.
- **Removed:** every Hub working folder this session made (factory-hub, -pointer, -posthog, -seed, -tile, -tile-tidy, -main-ref), the retired seeds' folders, and the workspace folder factory-finish. The proof suite script pointed `FACTORY_HUB_API` at factory-hub; point it at any current Hub checkout's `app/functions/api`.
- **Last tests:** the full factory suite (25 kit sections and 5 standalone tests) 0 fails at 6:48 am on the landed code.
**3 · Done this run (1 Oct 6:50 pm → 2 Oct 7:05 am):**
- Four independent reviews (Fable, Astra, senior engineer, cold reader) rewrote the order of work (the FINISH RUN block below).
- The factory ran live on a test bug for the first time. Three seeds exposed 13 live faults, each fixed with a proof that was red first: the model door's data class and reply shape, the account gauge's shape, a missing state folder, a build brief that named forbidden files, a harness red for the wrong reason, no cause file, false "hard" ratings, merges refused on a fast-moving main, world refusals burning attempts, the drive unable to open a test card, and the picture check failing on side panels. Before: no seed left the inbox. After: seed three passed harness, build and both readers on its own, and its drive now grades clean.
- STEP 14 (Factory tile on Home) verified 100% in a browser on the live Hub. STEP 16 (pointers) 95% with a working proof. STEP 13 (monitor) 75% with daily tripwires. STEP 9: the three reader seats (Fable, Astra in an isolated run, Opus) are built and the plain-words plan exit is live. STEP 5: the PostHog reader is built, live read-only, and triage syncs it first.
- The drive's picture check (051dc1cac3): only the scrolling box around the change is treated as moved, pinned boxes are masked by their own shadows, a two-level colour wobble is excused, and each picture's band and masks are saved beside it.
**4 · Postmortem:**
- **Went well:** running a real seed found more in one night than every stand-in proof had; each fault was fixed and proven within the hour; outside verifiers caught real gaps each round.
- **Wrong: stand-ins matched my code, not the real door** (data class, reply shape, gauge shape), so the suite was green while nothing worked live. Rule: build every stand-in from a captured real reply, and run one live seed before calling a stage built.
- **Wrong: I guessed at the picture failure three times** (row slack, column edges, shadow margin) before looking. The cause was one pixel one colour level apart, found in two minutes once the band and masks were saved and the pixels printed. Rule: when a picture check says "changed", print the differing pixels and the masks first; never fix from a hypothesis.
- **Wrong: a failed drive spent the seed's place** (one transient "region missing" sent the row to plan and a test card to Mae's side at 4:40 am). Rule: a machine fault in the drive is a wait, not an attempt and not a hand-off; treat "unmeasurable" like the merge stage's world refusals.
- **Wrong: measuring in one window size and photographing in another** (the full-page picture stretches the window). Rule: measure in the same layout the picture is taken in.
- **Wrong: in-page code kept inside a template string hid a broken pattern until a 90-second live run.** Rule: after editing in-page code, evaluate the template and parse it before any live run.
**5 · Ideas for next steps, ranked:**
1. **Finish STEP 11 (the finish line's heart):** restart the runner on 051dc1cac3 or later; set row `fx-4e8dead05fef` from `plan` back to its drive stage (a status reset through the factory door); let the scheduled fix job take it through drive, merge, publish and hand-off to Mae; then `node projects/ops/factory/prove-e2e.mjs --issue fx-4e8dead05fef`.
2. **Make a drive machine fault a wait** (`jobs/factory-fix.mjs`, the drive stage: "unmeasurable" returns wait, no attempt, no plan), with a red-first case.
3. **STEP 5:** wire the monitor's PostHog occurrence count and build the beacon harness.
4. **STEP 9:** one live hard-tier run through the three seats.
5. **STEP 10:** the independent VERIFIED line and the eight heartbeats; **STEP 12:** the third-screen-in-14-days trigger; **STEP 16:** the checker's re-run.
6. **STEP 17 and the audits:** sign-off, then fresh-eyes audit rounds until a round's best finding is minor. STEP 15's live wiring and STEP 18 wait until the factory is opened.
**6 · Held for a person:**
- **Opening the factory to real reports** (setting `OPENED`): Nick's, by design; default if nobody answers: stays on test rows only.
- **Mae's Done tap on the first live test fix** (FINISH LINE 9 evidence): Mae's, once step 1 above hands it to her; default: the machine proof stands and the tap is recorded when it comes.
- The test plan card for `fx-4e8dead05fef` handed at 4:40 am is a test card; the resume either carries it forward or archives it with the toolkit's archive-test-cards tool.
## 🔴 FINISH RUN — 1 Oct 2026, night: the order, the safety set and the six ideas (READ THIS BEFORE ANY STEP BLOCK)
**Why:** Nick, 1 Oct 2026 ~6:50 pm: *"get fable astra boris and cold reader to eval plan and ideas then write a proper update for the plan and then you get it done … /finish … dont turn it off until this project is done no stopping for updates or approals just get it built"*. The ideas are six lessons from Anthropic's "The AI-native SDLC playbook" (claude.com/blog/the-ai-native-sdlc-playbook, 21 Aug 2026). Four independent read-only reviews were run the same night: Fable (design/strategy), Astra (Codex gpt-6-astra, engineering attack; it reproduced three gate gaps with probes), Boris (`senior-engineer`, code reality) and a cold reader (`spec-breaker`). All four said: right goal, wrong order, unsafe to arm as it stands. Where this section disagrees with a STEP block below, THIS SECTION WINS; the STEP blocks keep their history.
**Facts corrected (measured 1 Oct 2026, 9:15 pm, Mac Studio):**
- The Codex CLI IS on the Studio (`~/.local/bin/codex`). The hard tier is still closed, for a different reason: STEP 9 has the gate and the seat loop only — no seat adapters, no `bind` command, no triad briefs, nothing calls `run()`. STEP 9 is 60%, not 95%.
- R12 makes every factory fix at least **medium** (it carries its own harness). STEP 11 and FINISH LINE 3 mean a one-file bug that the classifier rates medium; the card says "medium", never "easy".
- **One revert rule (R13), replacing STEP 10's "the runner never merges a revert" and settling STEP 13:** a fix whose live check fails, or whose own harness goes red after publish, is reverted at once through the revert path (F2-S6), as REPO §1 requires; the row then goes to `plan` with a card.
- Nothing reads `spec` or `plan` rows yet (STEP 12 is 0%, `go_ahead` has no owner): armed today, those rows would go silent. PostHog error rows never leave `publish` (`jobs/factory-fix.mjs:126-127`, `lib/factory-run.mjs:100-101`).
- The Studio's load average was 604 at 9:08 pm (other lanes' browser runs). The drive's two real-run failures were pages photographed mid-load under load; time-based waiting cannot fix that, readiness must be asserted.
- `prove-e2e.mjs` waits up to 24 h for Mae's own Done tap inside one command (breaks the ten-minute rule). The machine proof of STEP 11 ends at `finished` with the card in Mae's hands; Mae's tap is asked of her on that card and recorded as FINISH LINE 9 evidence.
- OPENED (the factory acting on real team reports) is a product decision held by Nick or Chantelle, not a safety rule; ARMED (acting on `agent-test` rows only) is this lane's to commit once F1 and F2 are green.
**The order (F1 → F7). A step starts the moment its inputs exist; F2's eight items run in parallel.**
- **F1 · Drive readiness (closes STEP 10's remainder).** Each local Hub is warmed once (a throwaway visit to the route and the sentinel page). A page counts as ready only when no loading placeholder is visible on two consecutive checks, the drive's region is present, and the network has been idle for 500 ms; the placeholder wait rises to 25 s; a page that never becomes ready answers `unmeasurable`, never a picture. PROOF: `--only real` passes twice in a row, launched detached on the Studio with `FACTORY_HUB_BUNDLE` set; a checker that did not build it writes STEP 10's VERIFIED line.
- **F2 · The safety set, before ARMED (added to STEP 10's definition of done; each red-first in `_test-factory-tools.mjs`).**
- S1 · the gate and the merge both refuse unless `origin/main` is an ancestor of the fix's head (main moved → rebase and re-prove).
- S2 · the hard gate refuses a record without three distinct seat verdicts, all proceed/agree, bound to the head commit.
- S3 · the gate refuses when main's latest publish run is red, or when an open pull request or a non-factory commit in the last 24 h touches a path the fix touches (RULE 64: no regression, lanes respected). This is this lane's substitute for STEP 18, which the Pages lane's thread owns.
- S4 · idea C hardened: one exported `fence()` used by the build tool, the merge gate and the reader; "the harness is unchanged" checked by `git diff --quiet <harness_commit> HEAD -- <own>`, with a case where the harness is edited AND committed.
- S5 · the harness and the fix's checks run under `sandbox-exec`: reads limited to the worktree and node, writes to scratch only, network to 127.0.0.1 only; red-first canaries (a harness that reads the vault folder, one that calls an outside host) must fail.
- S6 · idea E, the rehearsed undo: at publish the revert branch is built and the fix's own harness is run on it and must go RED (`revert_rehearsed_sha` recorded); publish refuses `live` without it; on a failed live check the rehearsed revert merges through the revert path (its diff must be the exact inverse of the merged fix, its harness red on the revert head) and the row goes to `plan`.
- S7 · no silent rows: every row that reaches `spec` or `plan` gets one card to the screen owner within two ticks (STEP 12's minimal slice: what was tried, in plain words); the heartbeat warns on any `spec`/`plan` row older than 24 h with no card; PostHog rows leave `publish`.
- S8 · `opened()` checks its shape (quoted words, a date, a proof-file hash); a change to `OPENED`, `ARMED` or `factory-classify.mjs` without them is refused.
- **F3 · STEP 11, rewritten.** Start when: STEPS 3, 4 (the renderer), 7, 8, 9's gate and 10 are live, and F1 and F2 are green. The seed is the one-line defect in `app/js/task-panel-factory.js` that shows only on cards whose name begins `agent-test`, with NO flag (a flagged seed reads as already fixed on the local Hub). The report is filed through the issue door as `factory` riding Mae, scripted (STEP 3's live proof path); STEP 2 and STEP 5 are not needed. This lane commits `ARMED=true`; OPENED stays null. `prove-e2e.mjs` is launched detached and ends at `finished`. Then idea E live: the seed's rehearsed revert is merged, the "before" picture and the red harness come back, and the factory fixes it again — the first rehearsed undo and a second clock sample.
- **F4 · STEP 13 with idea B and a narrow idea D.** Every harness under `_selfchecks/factory/` runs daily; a harness retires only into a cheaper permanent check, never into nothing; a weekly revert drill on the test copy writes a "last undo proven" date. Fixed-count tripwires, no σ bands and no model in detection: rows stuck in `spec`/`plan` over 24 h, publish failures, reverts, spend per day, a factory job that has gone quiet; a doubling against the 7-day median, or a quiet job, sets `factory:pause` (honoured by the merge stage) and files one Plan card.
- **F5 · STEP 12 in full, and the spec exit:** the three-seat runner as a job, `bind`, `go_ahead`, seat adapters (Fable and Opus by `claude -p`, Astra by `astra-run.sh`), so wishes and hard rows can move. STEP 9 to 100%; FINISH LINE 4 is owned here.
- **F5 build spec, second version (2 Oct 2026, ~1:05 am; the first version, a separate review job around the three-seat loop, read FIX-SPEC from Astra with eleven findings and is replaced by Astra's simpler design).** The hard tier moves only through three more independent readers of the exact commit, inside the check stage that already runs readers as detached jobs; nothing else changes.
- **Three seats in the check stage's reader program** (beside its `cheap` and `sonnet`): **propose** = Fable (`claude -p --model fable`), **attack** = Astra (Codex `gpt-6-astra` through `astra-run.sh`, read-only sandbox, an empty folder as its directory, homes `~/.codex`, `~/.codex3`, `~/.codex4` in turn, never `~/.codex2`), **cold** = Opus (`claude -p --model opus`). For a hard card the check stage starts all five readers at once on ONE frozen input file (`sha`, `diff`, `words`); no seat ever receives another seat's answer, so the cold read is independent by construction. Every seat uses the reader's existing restrictions (no tools, no settings, no MCP, no session, its own empty folder), fences the words AND the diff between nonce-marked DATA lines, and scans the words and the diff with the floor check and the plan writer's secret shapes before anything is sent (a hit = CHANGES, never sent); its answer is scanned the same way before it is stored.
- **Verdicts normalise to the stage's own words:** propose `proceed` → MERGE, `stop` → CHANGES; attack `proceed` → MERGE, `proceed with changes` or `stop` → CHANGES; cold `agree` → MERGE, `disagree` → CHANGES. CHANGES from any seat sends the card back to build with the notes (an attempt; three attempts → plan), so a request for changes is never approval of the same commit. A seat that cannot run answers "later": its finished neighbours are kept for the same commit, only it restarts, and not sooner than 30 minutes; the 48-hour and cost bounds still end the card.
- **Evidence the coordinator writes, not the models:** the seat program stamps `seat`, `model` (from the command it ran and the model the command line reports, never from the model's own words; a mismatch = "later"), `sha` and the input file's sha256; the check stage accepts an answer only when its sha is the head and its digest is the input it wrote. The readers' outputs live in the factory's state folder, which no harness (sandboxed) or builder (confined to its worktree) can write.
- **The gate (S2 completed):** a hard change needs, among the head's checks, MERGE from all five seats `cheap`, `sonnet`, `propose`, `attack` and `cold`, each stamped with its expected model family. The old `record` path is not used for hard any more. A review belongs to one commit (that is `bind`): a new head after a rebase is read again from scratch. Known gap, kept: between the gate's last ancestor check and GitHub's squash there is a window of seconds in which main can move; branch protection's "up to date" rule closes it and needs a paid GitHub plan (a money item after STEP 11).
- **Routing:** triage sends a hard bug to `building` like any other bug (it stops at `ready` for its reviewer, R10); a hard wish gets the plain-words plan and Go ahead like the others; the hard review happens in its check stage.
- **Proof:** each seat's parse and normalisation, floor refusal of words and diff, a model mismatch → later; the check stage starting five seats for hard and two for medium, keeping finished seats when one says later, refusing an answer whose digest or sha differs; the gate refusing a hard change missing any seat, a seat on another commit, a seat with the wrong model family, a proposer's `stop`; triage and the round routing hard rows. A live hard run waits for an agent-test hard row after STEP 11.
- **F6 · STEP 2 wiring** from the Studio through `brain-land.sh`; STEP 11 re-run through the Screen Buddy drawer as Mae on a phone. **STEP 5:** the PostHog issue reader only; source maps and replay deferred.
- **F7 · STEP 14** (the tile, with "last undo proven"), **STEP 15** with idea A, **STEP 16, STEP 17.** Idea A here means: the nightly correction lands a brief or classifier change only if the replay set passes — every classifier fixture and every past run's recorded diff gets a tier no lower than before — and each send-back or mis-tier adds a fixture. Steering files outside the factory (CLAUDE.md, CORE, skills) belong to ZION-13, not this plan.
**The six ideas, decided:** A ADAPT (F7, STEP 15) · B ADOPT (F4, STEP 13) · C ADOPT, mostly built, hardened (F2-S4) · D ADAPT narrow (F4: fixed counts and a pause switch, not σ bands — six users make bands noise) · E ADOPT (F2-S6, F3, F4) · F ADOPT (F3).
**Still on other lanes or people, none blocking F1–F4:** STEP 4's two board-scan changes (PM lane) · STEP 2's draft door, needed only for changes (Screen Buddy core lane) · STEP 18 (the Pages lane's thread) · OPENED (Nick or Chantelle, a dated word) · Mae's one Done tap on the seeded card · branch protection on the Hub's main (a paid GitHub plan: money, after STEP 11). Housekeeping owed: `FLOWS-order.json` (a leftover cheap-lane work order) and `FLOWS.html` sit beside this plan without a RULE 51 ticket.
## The factory floor, in one picture (read this before the steps)
```
INPUTS ──────────────► TRIAGE ──────► SPEC? ──────────► BUILD ─────► MACHINE REVIEW ─────► VERIFY ────────► RELEASE ──────────► MONITOR ──┐
Screen Buddy dedupe by hard tier: cheap easy: 1 checker biz-app-qa merge → publish beacon 7 days │
"something's wrong" fingerprint 3-seat review builder in medium: checker drives the bug: live for all; reopen on │
"I wish this…" tier: easy / writes a plain- a worktree, + Sonnet review real screen as the reporter closes recurrence │
error beacon medium / hard words spec → one file or hard: 3 seats before, the reporter; change: flag on for │
hub-checks FAIL rows severity screen owner one feature verifier + Sonnet pictures to the reviewer → Looks │
deploy-watch, sweep owner by map reads it first red-first after; bound to the the Look page good = everyone; │
Monday audit findings (never code) harness reviewed commit Send back = reopen │
NPS / feedback words (flag off · revert PR) │
▲ │
│ RECURRING (same fingerprint twice · fix reverted · 3 on a screen in 14 days) ──► ONE "Plan with a person" card │
│ to the screen owner on Coherence Improvement; a review that failed 3 loops ──► ONE card on Ecosystem Fixes (BUILD §1a) │
│ │
│ ALWAYS IMPROVING (daily): per-owner observer → ≤3 suggestions to the owner's Inbox · ecosystem scout → ≤3 lines │
└────────────────────────────────────── approved suggestions re-enter as changes ◄─────────────────────────────────────────────────────┘
SELF-IMPROVEMENT (nightly): send-backs and reverts on easy/medium changes, grouped by Jev's reason · cost → the classifier lists and the briefs are corrected; one Factory tile shows the numbers.
JEV (the Hub's cheap judge, every 30 minutes, watch → suggest → act per question): kind · severity · same-thing · tier hint (raise only) · beacon noise · send-back reason · idea rank; a Jev answer never closes a step or lowers a tier.
```
**The one rule a non-developer needs:** they never see code. A card carries their own words, one sentence on what changed for them, before-and-after pictures, **Try it** with three lines on how to try it and what it touches, what was checked in plain words, the risk in plain words backed by a real mechanism, who gets it after Looks good, and two buttons. Correctness is settled by machines and by other models; taste, priority and "is this what I meant" are settled by the person who knows the product.
**Review depth by difficulty (the requester's rule, made checkable):**
| Tier | What puts a change here (the classifier, contract C3) | Before building | Machine review after building | A person |
|---|---|---|---|---|
| **easy** | one file, not on the TRUNK or SHARED lists, no data shape change, no door signature change, no send and no money path | nothing | one cheap checker (a different vendor) re-runs the red-first harness and the build gates | a bug: told, with pictures, after it is live, and closes the card; a change: Looks good / Send back |
| **medium** | two to five files, or one screen plus its door, or a shared stylesheet token, still no TRUNK or SHARED path and no data shape change | the plain-words spec is written and sits on the card | one cheap checker re-runs the harness **and** one Sonnet review of the diff against the spec | reads the spec on the card first if it is a change; Looks good / Send back at the end |
| **hard** | any TRUNK or SHARED path, any data shape or store change, any door signature change, anything on a money, credential, sending or destruction path, or more than five files | a three-seat review: Fable proposes with the revert command named, Astra attacks the premise, Opus reads the corrected plan cold; the plain-words spec goes to the screen owner before a line is built | the review's verifier that did not build it, plus one Sonnet review; RULE 60's loop up to three times; the record is bound to the reviewed commit | reads the spec first; Looks good / Send back at the end; a third failed loop is a card on Ecosystem Fixes |
A builder may raise a tier and never lower one; the tier is stored on the run record at first classification and only ever rises. The TRUNK and SHARED lists are code, not prose (C3).
**Where this plan's model choices come from, and the one conflict it resolves:** `projects/ops/walkaway/MODEL-MATRIX.md` (2026-08-22, Chantelle's 2026-08-18 ruling) puts worker verification on Sonnet "never cheap" and test-authoring review on Opus; the plan skill's §M carries Nick's later word (2026-09-09): *"that independent verifier can still be Qwen or DeepSeek or GLM"* and *"the final QA stays with OpenAI or Anthropic, but only the final."* His later word wins: per-step checkers here are cheap, Sonnet reviews diffs on the medium and hard tiers and authors the tests, Opus and Fable hold the three-seat review and the final sign-off. The matrix's weekly correction pass is where that row is brought into line; this plan records the choice rather than claiming the matrix already says it.
## Already true (facts, not story)
- Coherence is the working name of the software product — the Hub sold as a product — and its own brand; the business is HyperProfits; the Hub is `nick-deck/deck-business`, live at hub.heroesandsidekicks.io, used today by Nick, Chantelle, Mae, Dean, DinDin and Rizza — evidence: `projects/business/savings-service/AI-Operating-Layer-Business-Master-Plan.md` (Nick, 2026-09-19: "coherence is working product name"); commit `2fa263c80c` (Nick, 2026-09-22: "1 its own brand"); `projects/ops/skippy-jobs/lib/hub-session.mjs:15` (`IDENTITIES`).
- **The Hub already has a Coherence Builds board**, one of seven department boards with the six agent stages and no name pattern, made on Nick's word (2026-09-22/23: "Coherence Builds will be where everything related to the hub builds into coherene lives") — evidence: `app/functions/api/tasks.js:267–274` (`DEPARTMENT_BOARDS`, `coherence-builds`); `board-report.mjs:600` (`EXTRA_WRITABLE_LANES`). It keeps large builds and human updates; the factory's cards live on the new Coherence Improvement board (the requester, 2026-09-28; C11). This plan's own card was opened on Coherence Builds on 2026-09-27 as `nt-20260928-003139-2211`.
- The task door validates `group` against its registered lanes and refuses unknown ones; a project card needs `project_spec`, `north_star`, `finish_line` (each ≤ 500 characters), `project_thread` and `project_thread_engine`, and is born in `Backlog`; the editable keys are fixed (`EDITABLE_KEYS`, `tasks.js:434`) and do not include a free `factory` object; a test card must carry `agent-test` in its name; a real card is completed only on a person's dated words, never by an agent on its own (§12.13) — evidence: `tasks.js:1952–1954, 2242–2243, 3204–3207`, and this plan's own card creation on 2026-09-27, refused six times until every rule above was met.
- The task door has a suggestion path: `action:"suggest"` files a proposal (name, why, assignee, group, due_date, what, tags) into its own store with a stable id, dedupes, remembers a dismissal, and makes a card only when a person presses Approve in their Inbox (Nick, 2026-09-22: suggestions "get approved before becomeing tasks") — evidence: `app/functions/api/_suggestions.js:1–12, 40–62, 77–102`; `tasks.js:1416, 3719–3732`; `inbox-feed.js:1433`.
- Screen Buddy already sits on every Hub screen (drawer, ⌘J, type or talk), sends `{screen, record_label, summary}` with every message with nine-plus-digit runs masked, signs in to the Hub as the person, and has one shared approval card: one sentence, **Look** and **Test** links, **Approve** and **Discard**, which sends only `{nonce, decision}` to `/api/skippy-approvals` from the brain's approval queue; a lane plugs in by returning the §5.1 draft shape, adding its approve action to `APPROVAL_ONLY` in `server.js` and its tool to `HUB_HELPER_TOOLS` — evidence: `projects/business/screen-buddy/PLAN.md` §3.1, §3.6, §5.1, §6.1 item 1; `app/js/workflow-voice.js:14–52`. The card has no picture layout, no free-text field and no numbered choice, and this plan does not change it.
- The Hub publishes on a **push to main** through the one Studio runner, never on a pull request; a newer push cancels an older run's sweep (`cancel-in-progress: true`); a docs-only push is free — evidence: `.github/workflows/deploy.yml:38–40, 67–69`; `projects/business/business-app/CLAUDE.md` §1, §3. **`main` is NOT protected**: the cold reader's live read on 2026-09-27 returned `protected: false` and the protection settings answer 403 (a paid GitHub plan is needed for a private repository); this Mac has no `gh` login to re-read it, so the plan treats `main` as unprotected. Agents merge proven work themselves (REPO §1, RULE 64).
- The task door lets an agent own a card only under its OWN declared name (`tasks.js:1978–1986`); a hand-off `kind:"review"` goes only to gracie or neeko, every other kind goes only to a person, and no kind hands a card to an agent (`tasks.js:1606–1611`); an agent starts running a person's card with `action:"take"` (`tasks.js:3641`). The Hub's working day is Asia/Manila (`_suggestions.js:31`, `manilaToday`). A "finished" hand-off lands in the column named **Nick's Review** whoever the owner is — a Hub-wide name the PM lane owns.
- `HS_DB` is hs-database, the client and Sidekick master database (rates, tax IDs, placements) with one choke point `_hsdb.js` and its schema in a sibling repo; every later feature got its own D1 database (`POOL_DB`, `CC_DB`), created on Nick's Slack approval — evidence: `app/wrangler.toml:66–75, 129–143`; `_hsdb.js:1–5`.
- The suggestion door's Approve makes an ordinary card for the suggestion's assignee (`_suggestions.js:150–176`); leadership sees every open suggestion in their Inbox (`inbox-feed.js:1454`); the local server starts with an empty local store (`harness/pm/serve.mjs`).
- The Ecosystem Fixes board is reserved: only a three-seat review that failed three loops routes there (Nick, 2026-09-19) — evidence: `projects/ops/blocks/BUILD.md` §1a.
- Checking already exists in five places and none of them takes a report from a person or fixes anything: 474 self-checks and the drift gate run on every publish (`app/build-dist.js`); the visual sweep posts one Hub inbox item on failure (`.github/workflows/deploy.yml:271`); `hub-checks-fast` and `hub-checks-daily` run two hard-coded lists in `projects/ops/life-os/REGROUP-2026-09-08/plans/HUB/harness/run-hub-checks.mjs:37–57` and record FAIL rows for the lane; `bizapp-deploy-watch` retries a failed deploy once and cards the second failure; `biz-app-qa` walks every live app on Mondays and folds findings into `projects/ops/ecosystem-audit/findings.json` — evidence: those files, opened 2026-09-27.
- **PostHog is already live in the Hub** (the cold reader of 2026-09-28 found what two earlier reads and the plan author missed): `app/index.html:193–276` has loaded PostHog since 2026-08-09 after a three-seat review — the project key in the page (US cloud), `capture_exceptions: true`, `autocapture: false`, `disable_session_recording: true` with the recorded reason "replay would ship payroll and client money figures … that is Nick's call, not an agent's", a `before_send` that strips `#` and `?` from addresses because the hash carries client and person names, and an early return when `navigator.webdriver` is set because the build's visual sweep hung twice with PostHog loading. So the project and its keys exist, errors are already being caught, and the two things the factory adds are a masked replay (a recorded ruling to supersede, on Nick's or Chantelle's dated word) and our own reading of PostHog's issues. The Hub has no "report a problem" control; NPS and survey feedback exist for clients, not for the product — evidence: `app/index.html:193–276`; `command grep -rn -i "report a bug\|report an issue\|report a problem" app/js/*.js` returns nothing, 2026-09-27.
- The build minifies with esbuild and produces no source maps (`build-dist.js:1008–1030`); PostHog's source-map inject rewrites built files, which must happen inside the build before its content hashes are taken, never as a later workflow step; `/api/build-integrity` already serves `source_revision`, the commit a build came from — evidence: `app/build-dist.js:1008–1030`; `app/functions/api/build-integrity.js`; posthog.com/docs/error-tracking/upload-source-maps/cli (read 2026-09-28).
- The Hub's Functions have no npm dependencies (no `package.json`, every import relative or `node:`), Pages has no clock of its own, and a PostHog flag-definitions fetch counts as ten flag requests; the department boards are one list (`DEPARTMENT_BOARDS`, `tasks.js:273`) mirrored in ten places that "change all four together" (`app/js/tasks.js:327–333`, `priority-boards.js:38`, `recurrences.js:57`, `pm-task-panel.js:302`, `inbox-feed.js:1461`, `projects.js:139`, `_jev-batches.js:939`, `_department-boards.selftest.mjs:19`, `board-report.mjs:600`); a board field is set only by the robot bearer, Nick or Chantelle (`tasks.js:1796`), `create()` takes no fields, and the task board groups only by due horizon (`app/js/tasks.js:4639`); the hand-off door has five kinds; Neeko's scan runs checks 2, 5, 10, 12 and 14 by code and pings a card owned by an agent and quiet for 24 hours (check 12) or an `agent-test` card older than two hours (check 14), and three pings put it in Nick's end-of-day report — evidence: the cold read of 2026-09-28, each line re-opened.
- `BZ.knowsFinance` is true for Nick, Chantelle, Mae and Rizza (`BUSINESS_FINANCE`), so it cannot mean leadership — evidence: `app/js/util.js:464–466`; `app/functions/api/_session.js:220–224`. Leadership is `LEADERSHIP` in `tasks.js` (Nick, Chantelle).
- Per-workspace feature toggles are specified and not built, and they belong to the workspace-setup lane; the Hub's `bizapp:` KV has no compare-and-swap and can read stale for a minute, and the Hub already binds D1 (`HS_DB`) — evidence: `app/_design/workspace-setup/SPEC.md` §3.3, §3.10; `app/_design/model-pool/PLAN.md` Already true; `app/wrangler.toml`.
- The model ladder is written (see the reconciliation above); Astra (`gpt-6-astra`) is the Codex reviewer, launched only detached through `astra-run.sh`; review-then-act is RULE 60 — evidence: `projects/ops/walkaway/MODEL-MATRIX.md`; `projects/ops/blocks/BUILD.md` §1a, §2; `~/.codex/models_cache.json` carries `gpt-6-astra` (2026-09-27).
- Failures already feed a Sunday review: `failure-weekly-review` posts one Hub card per ISO week for the week just ended and stores that week's card id in its own state file (`failure-weekly-review.json` under the jobs' `state/` folder) — evidence: `jobs/failure-weekly-review.mjs:143–176`.
- Spend is metered per call with vendor, model, tokens and cost — evidence: `projects/ops/spend-meter.cjs:208-361`.
- **Jev** is the Hub's own cheap decision model (Workers AI, `jev-1.13.0`, about $0.04 per million tokens plus a 5% fee, under a $10-a-month ceiling Nick approved on 2026-09-25): one door `app/functions/api/_jev.js` (`judge(env, use, subjectKey, state)`), a registry of typed questions in `_jev-uses.js` (question types `choice`, `score`, `noul`; effects `annotate`, `prefill`, `order`, `route`, `tag`; never money, keys, deletion or sending), per-use stages `off → watch → suggest → act` held as data in `bizapp:jev-config` and switched in Settings → Judgment, a 7-day cache, a floor check on every input, an outside-text fence, a log of every decision (never the text), and a 30-minute clock (`jev-tick.js`) that judges what changed; on any failure Jev gives no judgment and the caller does exactly what it did before; a Jev "yes" is never the final word (plan skill §BOARD) — evidence: `_jev.js:1–30`, `_jev-uses.js:1–20, 1543–1565`, `jev-tick.js:1–12`, `.github/workflows/jev-tick.yml`, `app/_design/jev/PLAN.md` (Nick, 2026-09-25: "look at every feature … determine how we can make JEV a part of it … require less human intervention"). Uses already registered that the factory reuses: `suggestion.wanted` (would a person want this suggested card), `task.prefill`.
- The plan skill's doctrine already names the human review points as taste, priority and outcome, never code: "Nick is never the tester"; "the plans themselves need to be clear as to what the outcome is for the human user" — evidence: `.claude/skills/plan/SKILL.md` §AUTH, §U.
- **The boards, on the requester's word (2026-09-28):** *"continuous improvement cycles probably deserve their own board as we'll have one for each feature suite/lane and that'll mean several lanes active at any given time on that board - it'll likely have its own PM rules and neeko will treat that differently to how we treat other builds"* and *"the coherence builds is for large new builds and human updates etc"*. So: a new group-backed board **Coherence Improvement** holds the factory's cycles, with one lane per feature suite (a board field, not seven boards — C11); **Coherence Builds** keeps large new builds and human-made updates; where bugs and issues live was left open and is decided in C11.
- **PostHog, read 2026-09-28 from its own docs:** error tracking auto-captures `window.onerror` and unhandled rejections in the browser (`capture_exceptions`), groups exceptions into issues by type, message and stack, supports custom grouping rules, assignment and alerts, needs source maps for minified code, and has an API to list issues, read fingerprints and link an issue to an outside tracker; session replay masks every input by default and can mask all text and unmask only marked elements, with masking done in the browser so masked text never leaves it; feature flags target a person by their distinct id and are evaluated from Cloudflare Workers with the `posthog-node` `workerd` build (a blocking remote call per request unless definitions are cached in KV); the free tier resets monthly at 100,000 exceptions, 5,000 replays, 1,000,000 flag requests and 1,000,000 events; autocapture (pageviews, clicks) is on by default and can be turned off or filtered with `before_send`; US and EU clouds exist — evidence: posthog.com/docs/error-tracking/capture, /docs/error-tracking/integrations, /docs/api/error-tracking-2, /docs/session-replay/privacy, /docs/libraries/cloudflare-workers, /docs/privacy/data-collection, and the 2026 pricing summaries (flexprice.io, dev.to/beton), all read 2026-09-28. Not found in its docs: whether an exception event carries its replay's link automatically (the replay is found by session id), and whether PostHog uses customer data for model training — both to be read from the account's own terms before the first real event (STEP 5).
## PostHog: what it takes over, what stays ours (decided 2026-09-28 on the reading above; the owed cold read covers it)
**Verdict: a hybrid, not either/or.** PostHog is the factory's **sensor layer**; the factory's **ledger, judgment and release** stay ours. The requester asked whether to lean on PostHog or build our own "given the platform complexity and where we're aiming to take this app"; the honest answer is that PostHog already does four things well that we would otherwise build badly, and does none of the things that make this a factory rather than a dashboard.
| Job | Who does it | Why |
|---|---|---|
| Catch page errors nobody reports, group them into one issue per root cause, keep counts and first/last seen | **PostHog error tracking** (replaces our beacon and fingerprinting) | its grouping by type, message and stack is better than ours and improves without us; 100,000 exceptions a month free; an API the runner reads |
| Show what the person actually did before an error or a report | **PostHog session replay**, mask-everything by default, unmask only marked safe elements | the replay link is the best "how to reproduce" a harness author can get; 5,000 replays a month free; masking happens in the browser so nothing on the floor leaves it |
| Turn a change on for one person, then everyone, then off in seconds | **PostHog feature flags** (replace our D1 flag store; our door stays as a thin wrapper) | per-person targeting, instant off from an API call, 1,000,000 requests a month free; server-side checks read a KV-cached copy so a door never waits on PostHog |
| Know which screens are used, by whom, how often, and where people give up | **PostHog product analytics** (pageviews and clicks, allow-listed) | the daily observer gets real usage evidence instead of guessing; nothing else in the Hub measures use |
| Take a person's words ("something's wrong", "I wish") from any screen | **ours** (Screen Buddy) | PostHog has surveys, not a conversation on the screen with the record open |
| Decide bug or wish, size it, set the review depth, run the three seats, gate the merge, publish, roll back | **ours** | this is the factory; PostHog has no opinion on any of it |
| The ledger: one row per issue with kind, tier, status, attempts, cost, the run record, the card | **ours** (the Hub's KV under `factory:` keys, one choke point), each row carrying `posthog_issue_id` and `replay_url` | the team's record lives in the Hub; PostHog issues are one input, linked never copied |
| The plain-words Look pages, the cards, the board, the recurrence options, the tile | **ours** | the team lives in the Hub; PostHog's Linear and GitHub issue links are for developers and are not used |
**What it costs:** nothing at the team's volume (six people; every free ceiling is orders of magnitude above it, and the flag-definition refresh is budgeted at about 432,000 of the 1,000,000 monthly flag requests with one writer refreshing once a minute); a paid tier only matters when Coherence has customers. **What it needs from a person:** nothing to start — the project and its key have been in the Hub since 2026-08-09; ONE decision: turning session replay on, masked, supersedes the recorded ruling that replay stays off because it "would ship payroll and client money figures … that is Nick's call" — so replay is a switch in C2 that stays OFF until Nick or Chantelle say the word, and everything else in this plan works without it (the harness author gets the person's words, the screen and the record instead of a replay). Agents handle the keys (CORE §2); the runner's personal API key and the source-map CLI key are minted by an agent from the existing account and go to the vault. **What it must never receive:** anything on the floor — replays (when on) mask all inputs and all text and unmask only elements marked safe, with console-log and network recording off in the project; exceptions keep their frames but the message value is blanked in the browser and a browser copy of the floor shapes drops anything that matches; addresses keep today's `#`/`?` scrub; autocapture stays off (the ruling) and screen use is measured by one explicit page-view event carrying only the route's first segment; IP capture is off. **What it changes in this plan:** STEP 5 replaces the inline PostHog block with the factory's own script (same key, same guards), reads issues, and puts source maps inside the build; STEP 6's flags move to PostHog behind our door with no vendor library in the Functions; STEP 13's monitor reads PostHog issues; STEP 15's observer reads screen use; C2 and C6 are rewritten below; the build-our-own beacon stays in the plan only as the fallback if PostHog ever refuses us.
## 0 · Gate Zero receipts (the plan may not exist without these)
- Failure Mode Registry loaded: 2026-09-27, `.claude/skills/plan/references/failure-registry.md`; the entries this build is exposed to are named in §4 (ten, plus five novel).
- Canonical specs loaded: code → `projects/ops/agents/CODE-STANDARD.md`; QA → `DEV-QA-SPEC.md`, `projects/ops/HANDBACK-GATE-SPEC.md`, `projects/ops/agents/HONEST-OUTPUT-STANDARD.md`; Hub → `projects/business/business-app/CLAUDE.md`, `KANBAN-AND-AGENT-BOARDS-SPEC.md` §10–§12 (§12.13 read in full); Screen Buddy → `projects/business/screen-buddy/PLAN.md` §3.4, §5, §5.1, §6.1, §7; build doctrine → `.claude/skills/plan/SKILL.md`; the three-seat review → `.claude/skills/triad/SKILL.md`; the model ladder → `projects/ops/walkaway/MODEL-MATRIX.md`.
- Ownership check: `projects/ops/artifacts/project-status/registry.json` has no row for a factory, an issue intake, an error beacon, a review tier or an ecosystem scout (searched `coheren|factory|improve|bug|qa` on 2026-09-27: the hits are the Model Pool, the Access Pool, ZION-17 board-PM enforcement, ZION-13 Claude upgrades, and the school's Lesson Factory); the neighbours are composed, not rebuilt: the Coherence Builds board, the suggestion door, the shared approval card, board-scan, biz-app-qa, failure-weekly-review, the Screen Buddy core.
- Expected inputs confirmed to exist: the Coherence Builds board (`tasks.js:274`; this plan's card sits on it); the suggestion door (`_suggestions.js`); the Screen Buddy §5.1 card (live); the hand-off door (`tasks.js`); D1 binding `HS_DB` (`wrangler.toml`); `harness/pm/serve.mjs`; `projects/ops/cheap-task.mjs` and `route-build.mjs`; `astra-run.sh` and the `gpt-6-astra` slug (used for this plan's own attack seat, 2026-09-27); `biz-app-qa`; `spend-meter.cjs`; `hub-session.mjs` `mintAgentSession` (used to open this plan's own card, 2026-09-27); the Studio runner's deploy workflow. Named as inputs to verify at their step, not assumed: `gh auth status` on the Mac Studio (STEP 11; this Mac's shell has no `gh` login; the Studio pushes with a keychain PAT); the brain's live check on the Mac mini (`_live-screen-buddy.mjs`, STEP 2); a job-callable way to store a draft in the brain's approval queue for a named person (STEP 2 proves it or names the core ask).
- PLAN AUTHOR: Fable (Claude Code desktop session on Chantelle's Mac, 2026-09-27).
- COLD READER: two seats, in sequence. Astra (Codex `gpt-6-astra`, read-only, told to disagree) returned **do not proceed** on the first draft with seven premise attacks, eight contract faults and seven rule conflicts, every one checked against the code and folded in. Opus (a fresh session, without sight of the attack) returned **do not proceed** on the second draft with eight premise findings, eight loop faults and seven rule conflicts, checked against the code and folded in ("What the reviewers changed", below); one finding is recorded as not re-verifiable from this Mac (branch protection). 🔴 This third draft has not yet been read cold: RULE 60's loop allows three, and the first act of the drive is one more cold read of this draft before any lane opens.
- PROMPT-SPEC scan (P1–P7): P1 "the Coherence app" → the Hub as a product (Nick, 2026-09-19), recorded in Already true. P1 "reviews for non developers" → §1 WHAT IT LOOKS LIKE and the card shape (C5). P1 "hard / medium / simple" → the classifier, C3, made checkable. P4 scope → NOT in scope, eight items with owners. P5 destination → the Coherence Builds board, the Inbox and the Home tile; the brain repo for this file. P6 "cohernece" read as Coherence. No P7.
## 1 · Goal and definition of done
- **What we're building, one paragraph.** A continuous-improvement loop for Coherence that runs on machinery the Hub already has: Screen Buddy takes a one-sentence report or wish from any screen; a beacon logs page errors nobody reports; every input becomes one issue row (deduplicated by fingerprint) and one card on the existing **Coherence Builds** board; a script classifies each change as easy, medium or hard by its blast radius and stores the tier so it can only rise; cheap models fix easy and medium bugs in a worktree with a red-first harness, a different model checks, the tester drives the real screen as the reporter, the runner merges and the Studio publishes, and the reporter gets a plain-words Look page with pictures and closes the card. Bugs (restoring what was meant to work) go live for everyone once checked; changes (new or altered behaviour) go live for the reviewer alone behind a per-change flag until their **Looks good**. Hard changes run a three-seat review (Fable, Astra, Opus) before a line is built, and its record is bound to the reviewed commit at merge. Recurrence and reverts produce one **Plan with a person** card with numbered options to the screen owner; a third failed review loop produces one card on Ecosystem Fixes. Weekly, an observer files at most three improvement suggestions per screen to its owner's Inbox and a scout files one eight-line ecosystem note to Nick and Chantelle, through the Hub's own suggestion door. Every Sunday the factory reads its own outcomes and corrects its classifier and briefs. One Home tile shows the numbers.
- **HOW IT'S USED:** a team member on any screen tells Screen Buddy what is wrong or what they wish for, in one sentence, and carries on; later they get a card in their own words with pictures and, for a change, a **Try it** and two buttons; a screen owner reads a plain-words spec before a hard or medium change is built and a **Plan with a person** card when something recurs; owners approve or dismiss ideas in their Inbox; Nick and Chantelle read one weekly ecosystem note and one tile. · HOW WE KNOW: the requester's words, 2026-09-27; Screen Buddy's rulings (phone = Screen Buddy only; draft then Look · Test · Approve; one-tap create) in `projects/business/screen-buddy/PLAN.md` §1.1; the suggestion ruling (Nick, 2026-09-22).
- **WHAT IT LOOKS LIKE:** nothing new is drawn. The report is a sentence in the existing drawer; the cards sit on a new six-stage Coherence Improvement board, one Suite field per feature suite (C11), and open in the existing task panel; the review is the existing shared approval card whose **Look** opens a plain **Look page** (the one new page: the C5 sections and the before-and-after pictures, in the Hub's own tokens), whose **Test** is **Try it**, whose **Approve** is **Looks good** and whose **Discard** is **Send back** (the drawer then asks for the sentence); ideas and the weekly note are suggestions in the Inbox; the tile is one more tile on the Home tile wall. A generated card section never shows code, a file path or a model name (C5). · HOW WE KNOW: `app/js/workflow-voice.js:14–52`; Screen Buddy §5.1; `inbox-feed.js:1433`; Home tile wall (`home.js` FUNCTIONS); CORE §1.
- **WHERE IT LIVES:** hub.heroesandsidekicks.io — the drawer on every screen; the **Coherence Improvement** board (`#tasks`, lens `coherence-improvement`; Coherence Builds keeps large builds and human updates); the task panel `#task/<id>`; the Look page `#factory/<issue id>`; suggestions in each person's Inbox; the **Factory** tile on Home for Nick and Chantelle. Issue rows, flag bookkeeping and run summaries in the Hub's KV (`BIZ_KV`, keys `factory:issue:*`, `factory:flag:*`, `factory:run:*`, one choke point); the runner's full run records under `projects/ops/skippy-jobs/state/factory/`. Opened by Nick, Chantelle, Mae, Dean, DinDin and Rizza under their own logins. · HOW WE KNOW: `tasks.js:274`; `app/wrangler.toml` D1 bindings; `hub-session.mjs` `IDENTITIES`.
- **WHAT IT MUST DO:** nine verifiable capabilities, each an eval in §6:
1. Take a one-sentence report or wish from Screen Buddy on any screen and make one issue row and one card carrying the screen, record, words and summary.
2. Log a page error nobody reports as one fingerprinted row with count, first-seen and last-seen, sending only an allow-listed payload.
3. Classify a change as easy, medium or hard from its diff and the TRUNK and SHARED lists, store the tier, and never lower it.
4. Fix an easy or medium bug on its own: worktree, red-first harness, cheap build, different-model check, browser drive as the reporter, merge, publish, Look page with pictures, card handed to the reporter to close.
5. Hold a change (not a bug) at Ready to look at behind a per-person flag; Looks good = everyone; Send back with words = reopened with the words in the next brief; flag off (server-authoritative, re-read by the client on every route change and within a minute) = rollback; a bug's rollback is a revert merged by the runner, named before publish.
6. Run a three-seat review before building any hard change and refuse a merge without its record bound to the merged commit.
7. Detect recurrence (same-source fingerprint reopened twice, a revert, three on a screen in fourteen days) and post exactly one Plan with a person card; route a third failed review loop to Ecosystem Fixes.
8. File daily, through the suggestion door: at most three ideas per owner when there is new evidence, up to three ecosystem lines to Nick and Chantelle; a dismissal is remembered by that door; Jev ranks the candidates and pre-screens the filed suggestions at its stage.
9. Show the factory's numbers on one tile for leadership only, derived, never typed, and correct its own classifier and briefs every night from send-backs and reverts, with Jev's reason hints grouping the send-backs.
10. Give Jev seven typed questions (kind, severity, same-thing, tier hint, beacon noise, send-back reason, idea rank), each starting at watch and promoted only in Settings on measured agreement, each with a fallback that is exactly what the step did before.
- **NOT in scope:** eight things a reasonable person would assume are in, each with its owner:
- Security or privacy work of any kind — owner: `projects/ops/sp-sec/PLAN.md` (one line there, then back to the step; plan skill §S). The beacon's allow-list is a data-floor routing rule, not security work.
- Per-workspace feature on/off (a customer turning the time tracker off) — owner: the workspace-setup lane (`app/_design/workspace-setup/SPEC.md` §3.3). The factory's flags are per-change and temporary: promoted to everyone or switched off within fourteen days.
- Reports and wishes from customers of Coherence, their support queue and SLAs — owner: the Daily workspace plan and Nick's go-to-market ruling ("get it to a proper MVP, then introduce it to our clients"). The same door will take them; this plan proves it for the six team members.
- Which cheap vendors customers may use, rebilling, and live web data inside features — owner: the Model Pool and Access Pool plans.
- Redrawing or rebuilding any Hub screen, the PM screens included, and any change to the shared approval card's renderer — owner: the PM plan (`app/_design/pm/plan/PLAN.md`), each screen's lane, and the Screen Buddy core lane. The factory changes screens only through issues and approved suggestions.
- Adoption of new Claude Code tooling for our own agents — owner: ZION-13. The scout reads its output and writes about the product, not our tools.
- A public roadmap, release notes for customers, marketing of what shipped — owner: Nick ("I'll work on marketing stuff later"). The Friday what-shipped digest already exists (`hub-ops-manager`) and is fed, not replaced.
- Anything on Nick's four acts (money leaving, a credential, irreversible destruction, a message sent as him) — the factory builds around them and files the act through the door; RULE 46 stands: no live client-facing card is ever a test. No option on a Plan with a person card ever involves money, by construction.
- Turning session replay on — a recorded ruling of 2026-08-09 keeps it off ("that is Nick's call"); the factory carries the masked replay as a switch and does not flip it; per-customer PostHog projects, the EU cloud and any paid tier are decisions for the day Coherence has customers, owned by the Daily workspace plan.
- **Trip-over protocol:** a lane that finds something outside the fence writes one handover line to its named owner (a security- or privacy-shaped thing: one line in `projects/ops/sp-sec/PLAN.md`), then back to building — never investigates, never fixes.
## 1a · Critical variables — the confirmation sheet is GENERATED from this table
> 🔴 **No lane opens while any row reads `UNCONFIRMED`.** The sheet Nick signs is this table rendered by `render_sheet.py` — never a separate summary. A variable belongs here only if the named person alone can settle it (V1) AND two plausible answers produce two different builds. V2 facts (what a thing actually is) never go on the sheet — open the thing and record what was seen, dated. Cap: seven rows.
| # | The variable, in plain words | Value chosen | Alternatives rejected | Class | HOW WE KNOW | Cost if wrong | CONFIRMED |
|---|---|---|---|---|---|---|---|
| 1 | **SURFACE — which screen this lands on, and who opens it** | The Hub itself: Screen Buddy's drawer on every screen for reporting, the new Coherence Improvement board (row 5) and the existing task panel for the cards, the Inbox for ideas, one Home tile for the numbers; opened by the six team members under their own logins | A separate issue tracker (GitHub Issues, Linear); a Slack channel as the intake; a form page; the cards on Coherence Builds | V1 | the requester, 2026-09-27, "i want to design a factory for our cohernece app - continuous improvement bug fixes etc" and "the whole process needs to work for a team that knows the product and user very well but not the tech"; Nick, 2026-09-27 (Screen Buddy): "we need it to be available on all users on the hub"; Nick, 2026-09-27: "Phone = Screen Buddy only"; Nick, 2026-09-22/23: "Coherence Builds will be where everything related to the hub builds into coherene lives" | the team reports somewhere they do not already work, and nothing gets reported | the requester, 2026-09-27, "user issues and bug get logged and fixed automatically and recurring issues get a human involved to plan a fix" |
| 2 | Who is the person who reads a spec and taps Looks good — the product reviewer for a screen | The reporter reviews their own report; for a beacon row or an approved idea the screen's owner reviews, from a hard-coded map (C4): Rizza finance, Mae workforce and hours, Dean SLA, NPS, huddles, Heroes and Sidekicks, DinDin the RO tracker and waiting-on-client, Chantelle look-and-feel, onboarding, pages and broadcasts, Nick leadership, CRM and Captus content; Chantelle when no owner is mapped. Never routed to Nick for permission | Nick reviews everything; a developer reviews | V1 | Nick, 2026-09-25 (RULE 63): "i dont want team to ask me for permission on our work they act and it gets done and approved period"; Nick, 2026-09-19: "each team member will need a dept specific dashboard of some sort" (the department map already exists in `home.js` FUNCTIONS) | cards pile up in front of the wrong person and nothing ships | Nick, 2026-09-25, "i dont want team to ask me for permission on our work they act and it gets done and approved period" |
| 3 | What goes live on its own and what waits for a person's tap | A **bug** (restores what was meant to work; the harness proves before-broken, after-working; nothing new appears) goes live for everyone once checked; the reporter is told with pictures and their tap closes the card (§12.13: an agent never closes a real card on its own). A **change** (new or altered behaviour, a wish, an approved idea) goes live for the reviewer alone and waits for Looks good | everything waits for a tap (slow; every typo fix needs a person); everything ships (the team wakes up to screens they did not ask for) | V1 | the requester, 2026-09-27, "user issues and bug get logged and fixed automatically"; Nick, 2026-09-18 (RULE 49): "the bar is verified and better, never complete or signed off"; Nick, 2026-09-27 (Screen Buddy): "Draft, then Look · Test · Approve. nothing goes live … until the person approves" — read together: a fix is verified-and-better, a new behaviour is a draft | either the team is flooded with approvals, or a screen changes under them | the requester, 2026-09-27, "user issues and bug get logged and fixed automatically and recurring issues get a human involved to plan a fix" |
| 4 | How deep the machine review goes, by difficulty | Three tiers by blast radius: easy = one cheap checker; medium = cheap checker plus one Sonnet review; hard = a three-seat review (Fable proposes, Astra attacks, Opus cold) before building, a non-builder verifier and Sonnet after, RULE 60's three loops, the record bound to the merged commit | one review depth for everything; a human code review | V1 | the requester, 2026-09-27, "triads across astra fable opus for hard problems and lower models or single or dual reviews for simpler and medium difficult problems"; Nick, 2026-09-09: "One independent review per task step … not a check at every micro stage" | either tokens burn on trivia or a trunk change ships on one cheap opinion | the requester, 2026-09-27, "tech problems probably get solved by multiple agent reviews like triads across astra fable opus for hard problems and lower models or single or dual reviews for simpler and medium difficult problems" |
| 5 | Where bugs and issues live: their own board, or inside one of the two | Inside **Coherence Improvement** (the factory's board, one Suite per feature suite): every issue is a row in the factory's ledger, and a card is made only when a person is or will be involved (a report or wish, a change waiting on a person, a plan card, a silent error whose fix failed or recurred); a silent error that fixes itself gets no card and is told to the owner in the next morning's evidence and counted on the tile | a third board for bugs; a card for every issue including silent fixes (each needing a tap an agent cannot give); bugs on Coherence Builds | V1 | the requester, 2026-09-28, asked for the call: "not sure if bugs and issues gets its own board or system or if it integrates into either of those", and "continuous improvement cycles probably deserve their own board … several lanes active at any given time on that board … its own PM rules"; §12.13 (an agent never closes a real card) decides the no-card-for-silent-fixes half | either a bug board nobody watches, or hundreds of cards waiting for taps nobody asked for | the requester, 2026-09-28, "not sure if bugs and issues gets its own board or system or if it integrates into either of those" — the call was delegated by that sentence and this row is the answer, reversible in C11 alone |
| 6 | Whether PostHog is the sensor layer (errors, replays, per-person switches, usage) with the ledger and judgment ours, or everything built in-house | PostHog for the four sensor jobs, ours for intake, ledger, sizing, review, release and the cards; the plain beacon kept as a proven fallback | build everything ourselves (a worse error grouper, no replays, no usage data); PostHog as the tracker too (its issue links are for developers, the team lives in the Hub) | V1 | the requester, 2026-09-28, asked for the verdict: "look into how we can leverage posthog for issues tracking or if we are better off building our own custom version given the platform complexity and where we're aiming to take this app"; the reading in "PostHog: what it takes over, what stays ours" | months spent rebuilding what is free, or a vendor that sees more than it should | the requester, 2026-09-28, "look into how we can leverage posthog for issues tracking or if we are better off building our own custom version" — the verdict was asked for by that sentence and this row is it, reversible in C2 and C6 alone |
- V1 confirmation reads `<name>, <date>, "<their own words>"` — the date is required.
- V2 confirmation reads `opened <what>, <date>, saw: <what was actually there>`.
**Considered and ruled NOT critical:**
- Where the issue rows live — KV (`BIZ_KV`, `factory:` keys) behind one choke point whose `revision` check refuses a stale write with 409, decided 2026-09-28 (the 2026-09-27 D1 choice is superseded: rows are one per issue, every write reads back, and the selftest runs in-process like every other Hub selftest).
- Whether the weekly ideas go to owners on Monday or Friday — Monday; changes on the sheet's word in one line of the job.
- The beacon's batching (every ten seconds, per-fingerprint counts, deduplicated on the server) — a fact about volume, not a preference.
- The per-issue bounds (three build attempts, 48 hours of wall clock, five dollars of metered spend, one three-seat review, then a person) — bounds, each changed in one constant.
## 1b · Subproject decomposition — could a piece of this ship on its own?
| Subproject | End goal (one sentence — what's TRUE when done) | Depends on (named artefact) | Owner | Own PLAN.md path | Confirmation-sheet status |
|---|---|---|---|---|---|
| INTAKE (report, beacon, card fields, issue rows) | Any team member's sentence or any page error is one row and, where a person is involved, one card on Coherence Improvement within a minute | Screen Buddy §5.1 card (exists); KV `BIZ_KV` (exists); the harness kit (STEP 1) | HUB + BRAIN lanes | this file (§3b STEPS 2–5) | rows 1–2 confirmed |
| RELEASE (flags, the Look page, Looks good / Send back) | A change is live for one person, then everyone on their tap, and off within a minute on a flag | STEP 4's card fields; the §5.1 approve card | HUB + BRAIN lanes | this file (STEPS 6–7) | row 3 confirmed |
| FIX LOOP (classifier, tiers, three-seat review, fix runner, merge gate, monitor) | An easy or medium bug fixes itself end to end with the right depth of review; a hard change cannot be merged without its review | STEPS 2–7 artefacts; `cheap-task.mjs`; `biz-app-qa`; `astra-run.sh`; `gh` on the Studio | PROCESS + JOBS lanes | this file (STEPS 8–11, 13) | row 4 confirmed |
| RECUR + IMPROVE (recurrence card, observer, scout, tile, Sunday correction) | A person plans only what recurs; owners get ideas weekly in their Inbox; the factory corrects itself and shows its numbers | the issue rows (STEP 3) and the run records (STEP 10) | JOBS lane | this file (STEPS 12, 14–16) | none needed |
**Carve-out rule:** anything left out of every subproject's scope is named with a real owner in the same edit, or it may not be left out. Carved out and owned: per-workspace toggles (workspace-setup lane); customer intake (Daily workspace plan); the Friday what-shipped digest (hub-ops-manager, fed by STEP 10's run summaries); the shared approval card's renderer (Screen Buddy core lane; STEP 2 names the one ask if the brain has no job-callable draft door).
## 2 · The complete UX map (this becomes the test manifest verbatim)
| Id | Screen / entry point | State (default·empty·error·loading) | Element / interaction | Expected behavior | Navigation from → to |
|---|---|---|---|---|---|
| U01 | Any Hub screen, Screen Buddy drawer (phone 375×812) | default | say or type "something's wrong: the tag button does nothing" | Screen Buddy answers in one sentence that it has logged it and names the card; one issue row and one card exist within 60 s carrying screen, record, words, summary; the card is owned by `factory` on Coherence Improvement in the screen's Suite, stage Backlog, tagged `factory`, `stage:inbox`, `kind:bug` | screen → same screen (drawer closes on Done) |
| U02 | Any Hub screen, drawer (desktop 1280×800) | default | type "I wish the huddle list sorted by overdue first" | same as U01 with `kind:wish` and the tier left for triage; the reply names it a wish, not a bug | same |
| U03 | Drawer | error | the issue door answers 5xx or times out | Screen Buddy says plainly it could not log it and to try again; the brain retries the same `request_key` once; the door's idempotency on `request_key` means at most one row and one card exist afterwards, never two | same |
| U04 | Drawer | edge | the sentence contains a run of nine or more digits | the row and the card carry the masked form, never the digits | same |
| U05 | Any screen (background) | error | a JavaScript error or an unhandled promise rejection happens | PostHog captures it and groups it into one issue per root cause; within one runner tick one factory row exists for that issue (fingerprint = the PostHog issue id, `source: posthog`), count 1; a second identical error updates the same row's count and last-seen, no second row; the exception that leaves the page keeps its frames, carries a blank message value and nothing matching the floor shapes | none (invisible) |
| U06 | Any screen (background) | edge | PostHog does not answer | the page keeps working; PostHog's own script queues and retries on its side; nothing is shown to the person; the plain beacon to our own door stays proven in the fixtures so it can be switched in if PostHog ever refuses us | none |
| U07 | Board chooser (`#tasks`) | default | tick **Coherence Improvement** | the board shows the factory cards, and the Suite filter chip narrows them to one feature suite; a factory card shows the person's words as its title and the tag `factory` | tasks → board |
| U08 | Coherence Improvement board | empty | no factory issues | the board renders its six columns and the Hub's usual empty words; every suite reads empty | — |
| U09 | Card (task panel), stage Backlog | tag `stage:inbox` | open it | the panel's Details section for a `factory`-tagged card is fetched from the issue door by card id and shows: the words, the screen, the record, the summary, kind (bug / wish), owner `factory`; no code, no path, no model name; the door down → "Details are not available right now" and nothing else | board → panel → board (scroll restored) |
| U10 | Card, In Progress | tag `stage:triage` | read | Details add: tier in plain words ("small: one screen" / "a few screens" / "a shared part of the Hub"), severity in plain words, the reviewer's name; Conversation carries the triage note | — |
| U11 | Card, In Progress | tag `stage:spec` (medium or hard change) | the reviewer reads | Conversation carries the plain-words spec: what will change for you, what will not change, how you will know; the approval card in the drawer reads "Go ahead with: <one sentence>" with **Look** (the spec on the Look page), **Approve** (= Go ahead) and **Discard** (= Not this; the drawer then asks "what should be different?" and the words go on the card and reopen it as Inbox) | panel → drawer |
| U12 | Card, In Progress | tag `stage:building` | read | Conversation shows one line per attempt ("attempt 1 of 3: building"), never code; the attempt clock; when every cheap vendor is down, "waiting for a builder" | — |
| U13 | Card, Internal Review | tag `stage:checking` | read | Conversation shows in plain words who checked ("one independent AI reviewer re-ran the test" / "two independent AI reviewers" / "three: proposed, attacked, read cold") and the browser drive result as the reporter | — |
| U14 | Card, Nick's Review → the reviewer | tag `stage:ready` (a change) | the reviewer opens the drawer | the shared approval card: "<one sentence on what changed for you> — Look · Try it · Looks good · Send back"; **Look** opens the Look page (`#factory/<id>`): What you said · What changed for you · Before / After pictures (phone and computer) · **How to try it** (three lines: open which screen, do what, expect what; "trying it changes nothing real: it uses an agent-test record" or "it changes your own <thing> and you can undo it") · What we checked (plain words) · Risk (backed by the mechanism: "on for you alone; can be switched off for everyone within a minute") · **Who gets it after Looks good: everyone on the Hub** · a **Something's not right** line that opens the drawer with the card attached; **Try it** opens the screen with the flag on for them alone | panel → drawer → Look page → the screen → back |
| U15 | U14 | action | tap **Looks good** (Approve) | the flag turns on for everyone; the card moves to Done through the door with the approval in their words; a second tap is refused (410) | — |
| U16 | U14 | action | tap **Send back** (Discard) | the drawer asks for one sentence; without words nothing changes; with words the card returns to Building as attempt n+1, the words in Conversation and in the next brief; the flag stays on for the reviewer only | — |
| U17 | Card, the review column → the reporter (the column is named "Nick's Review" for everyone; the Look page's first line reads "Waiting for your look, Mae") | tag `stage:live` (a bug) | the reporter opens | the approval card reads "This is live for everyone: <what changed for you> — Look · Done"; the Look page shows the pictures and what was checked, and its Risk line reads "already live; say 'put it back' and it is undone"; **Done** (Approve) closes the card with the reporter's dated words (§12.13, the request shape); **Something's still wrong** (Discard) opens the drawer, which offers two choices: "put it back the way it was" (merges the prepared revert and reopens the row) or "it's still broken" (reopens the row as Inbox with `reopened: 1`); a card the reporter never closes stays open and Neeko's hourly scan pings it as any other | panel → drawer |
| U18 | Card, the review column → the screen owner | tag `stage:plan` | the owner opens | the approval card reads "<the problem in one sentence> — Look · Go with the recommendation"; the Look page shows: the problem in the reporter's words; what was tried (one line per attempt); two or three numbered options, each with **what it means for you**, **how long until it is live** ("about a day" / "about a week"), **who acts next**, whether it touches anyone else's screen, and "this changes the software only; nothing is bought"; one recommendation; a due date; the owner taps Approve for the recommendation, says "option 2" in the drawer, or Discards with words ("none of these"); the chosen option becomes a spec row and a Go ahead card (U11) | — |
| U19 | Board | edge | the same fingerprint arrives while a card is open (a second report of the same thing, or a new occurrence on the same PostHog issue) | no second card; the open card's count rises and Conversation gets one line "seen again by <person> on <screen>" (for a PostHog occurrence: "seen again on <screen>, <count> times"); a new reporter is added to the row's reporters | — |
| U20 | Board | edge | the same fingerprint arrives after its card was Done | the row's `reopened` becomes 1 and a NEW card is born, linked to the Done one in its first comment (an agent never reopens a person's closed card); a second recurrence makes the Plan with a person card (U18) and not a third fix | — |
| U21 | Any hard-tier change | edge | the runner tries to merge without a review record, or with a record made for a different commit | refused with the reason named on the card; the change stays in Building | — |
| U22 | Home (Nick, Chantelle) | default | the **Factory** tile | issues in (7 d / 30 d), fixed on their own, needing a person, median report-to-live, send-back rate, cost, stale flags; tapping opens the Coherence Improvement board | home → board |
| U23 | Home (Mae, Dean, DinDin, Rizza) | default | no Factory tile | the metrics door answers 403 for them and the tile is not drawn; their own screen's ideas arrive as suggestions in their Inbox | — |
| U24 | Owner's Inbox, every morning | default | up to three suggested cards for the screens they own (none on a day with no new evidence; Jev's `suggestion.wanted` folds the ones it judges unwanted, never hides them) | each: a name in plain words and a **why** (what it means for you, what was seen, how big — the suggestion door's two fields); the Inbox's own **Approve** makes the card, which triage turns into a `wish` row and takes (so the owner never has to build it), and **Dismiss** is remembered by the suggestion store so the same idea is never refiled | inbox → card |
| U25 | Nick's and Chantelle's Inbox, every morning | default | the **What's new out there** set: at most three suggested cards a day, one per line, none on a quiet day | each names a Coherence screen or a gap, a recommendation, and the source in words; Approve makes a `wish` row owned by that screen's owner; Dismiss is remembered | inbox → card |
| U26 | Any card | error | the cheap vendors are all down | the card's Conversation says "waiting for a builder" and the attempt clock pauses; nothing escalates to a person; the runner retries next tick; the 48-hour wall clock still runs | — |
| U27 | Any card | edge | the run passes three attempts, 48 hours, five dollars of metered spend, or three review loops | attempts, hours or dollars → Plan with a person to the screen owner (U18); three review loops → one card on Ecosystem Fixes in the same shape; the reporter is told in one line | — |
| U28 | Any card | edge | a flag reaches fourteen days neither promoted nor off | the flag switches itself off, the card returns to the reviewer with "this switched itself off because nobody decided", and the tile counts it under stale | — |
## 2d · DESIGN FIDELITY GATE
DESIGN FIDELITY GATE: N/A — nothing rendered that is new in kind: the Look page (STEP 7) is text, four pictures and the Hub's own tokens with no new component, the approval card is the shared one (Screen Buddy §5.1, already graded), the tile joins the Home tile wall (PM plan, already graded). If the Look page acquires any new component, a dated PLAN-CHANGES.md delta fills this block before it is built; the phone layout of the Look page is checked at 375×812 by STEP 7's harness.
## 3 · Lanes and frozen contracts
| Lane | Scope (in / out) | Owner | Definition of done | Builder (cheap, named) | Backup builder | Checker (different model) | Backup checker |
|---|---|---|---|---|---|---|---|
| KIT | in: every harness, selftest and fixture this plan's proofs name, written red-first before the thing they test / out: any product code | this plan's overseer | STEP 1 closed | Sonnet (`TEST-AUTHORING:` line; the matrix's test-authoring row) | Opus | Qwen 3.8 (runs every file; each must be red or refuse by name) | DeepSeek V4 Pro |
| HUB | in: doors `factory-issues`, `factory-flags`, `factory-review`, `factory-metrics`, the beacon, the Look page, the card Details renderer, the tile, the migration / out: any screen's own behaviour, the PM screens, the approval card's renderer, workspace toggles, `tasks.js` | this plan's overseer | STEPS 3–7, 14 closed | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 (harness re-run) or Sonnet (diff review) | DeepSeek V4 Pro |
| BRAIN | in: Screen Buddy helper `hub-factory.mjs` with tools `report_issue`, `suggest_idea`, `explain_change`, `send_back`, `choose_option`, `factory_ready` (the job-callable draft) and their `APPROVAL_ONLY` / `HUB_HELPER_TOOLS` registration / out: the drawer, voice, the approval card itself (core) | this plan's overseer | STEP 2 closed | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 | Sonnet |
| PROCESS | in: `factory-classify.mjs`, `factory-review.mjs` (tiers), `factory-triad.mjs` (the three seats and the merge gate), `factory-card.mjs` (the plain-words text), the briefs, all under `projects/ops/factory/` / out: the cheap tools themselves, `astra-run.sh`, the matrix | this plan's overseer | STEPS 8–9 closed | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8; Sonnet for the tests | DeepSeek V4 Pro |
| JOBS | in: scheduler jobs `factory-triage`, `factory-fix`, `factory-monitor`, `factory-recur`, `factory-observe`, `factory-scout`, `factory-metrics`, `factory-correct` under `projects/ops/skippy-jobs/jobs/`, `factory-state.mjs` under `projects/ops/skippy-jobs/lib/`, their schedule rows / out: the runner itself, other jobs | this plan's overseer | STEPS 10–13, 15–16 closed | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 | Sonnet |
**Contracts between lanes (FROZEN at plan time — change = dated PLAN-CHANGES.md delta):**
- **C1 · The issue row** (KV `BIZ_KV`, key `factory:issue:<id>` with the index `factory:issues`, one choke point `_factory-issues.js`; HUB owns the door, JOBS reads it): `id` (`fx-<12 hex>`) · `revision` (int, +1 on every write; a write carrying a stale `expected_revision` is refused 409) · `request_key` (the caller's idempotency key; the same key returns the same row) · `fingerprint` (sha256 of `beacon|screen|error_name|file|line` for a beacon row; sha256 of `report|screen|record_kind|normalised words` for a report; the source word is part of the hash so the two kinds never collide by accident) · `kind` (`bug|wish|error`) · `source` (`screen-buddy|posthog|beacon|hub-checks|deploy-watch|visual-sweep|monday-audit|nps`) · `posthog_issue_id` · `replay_url` · `repo` (`hub|brain`) · `screen` · `record_kind` · `record_id` · `record_label` · `summary` · `words` (masked) · `reporters[]` (identities, first is the original; `beacon` for beacon rows) · `reviewer` (from C4) · `owner` (from C4) · `tier` (`easy|medium|hard|null`, only ever rises) · `severity` (`blocks|wrong|cosmetic|null`) · `count` · `first_seen` · `last_seen` · `status` (`inbox|triage|spec|building|checking|ready|live|done|plan|never`) · `card_id` · `reopened` (int) · `attempts` (int) · `flag_id` · `related[]` (issue ids the triage model linked as probably the same thing, never merged) · `created_at` · `updated_at` · `details` (AMENDED 2026-09-29 on STEP 4's independent check: a card's finished C5 body, `{shape: intake|spec|ready|plan, sections: [{label, lines[], own_words?}]}`, at most twelve sections of at most twenty lines of at most 500 characters, written by the factory runner through `update` and drawn by the card's Details renderer; each line is generated and passes the C5 gate unless its section says `own_words: true`, which the person's own words always do). Door: `POST /api/factory-issues` actions `report` (from the brain as the person) · `report_posthog` (agent key `factory` only: creates or updates one row from a PostHog issue id, count, first and last seen, replay; no card) · `beacon` (the fallback only; from the page, the session cookie is the identity) · `list` · `get` (by id or `card_id`) · `update` (agent key `factory` only, with `expected_revision`) · `reopen`. Every write reads the row back and returns it. A `report` creates the card through the tasks door with `id` = the issue id (the door's own caller-id idempotency, `tasks.js:2188`), so a retried report finds the existing card and never makes a second; the row stores `card_id` only after the card reads back. A card is made only in the cases C11 names (a person's report or wish, a change waiting on a person, a plan card, a silent error whose fix failed or recurred); it lives on **`coherence-improvement`** in its Suite (C11), carries `due_date` = today + 2 (an agent cannot open an undated card) and is created by an agent session named **`factory`** riding the reporter's identity (`mintAgentSession("factory", person)`; the door lets an agent own a card only under its own declared name, `tasks.js:1978–1986`, so the brain's helper mints this session, not `voiceagent`), on `coherence-improvement` with its Suite (C11), assignee `factory`, `session_id`/`session_title` from the run, `what` = the words, `tags` = `factory`, `stage:<status>`, `kind:<kind>`, and `test_card: true` with `agent-test` in the name when the words begin `agent-test`; nothing else is written onto the card — the Details renderer fetches the row (U09). Card ownership through the loop: `factory` owns it until review; `ready` and `live` hand it to the reviewer with `kind:"finished"` (the only kind that reaches a person; `tasks.js:1610`), so the reviewer owns it from then on; a **Send back** never hands it back (no kind hands a card to an agent): the reviewer keeps the card and `factory` runs it with `action:"take"` (`tasks.js:3641`; `take` records the agent on the card and moves nothing) until the next `finished`. Progress while a card is being worked (attempt lines, "waiting for a builder", the checker's verdict in plain words) is written as **comments** through the comments door signed `factory`, never as hand-offs: a hand-off always changes the owner and the column and is used only at a stop. A row that recurs after its card is Done never reopens that card (an agent cannot undo a completion without a person's dated request, §12.13): the row's `reopened` count rises and a NEW card is born, linked to the old one in `related[]` and in its first comment. The column a finished card lands in is named "Nick's Review" for everyone; the Look page's first line therefore reads "Waiting for YOUR look, <name>", and the column's name is on the NEXT list for the PM lane.
- **C2 · Errors, replays and usage through PostHog** (HUB, `app/js/factory-posthog.js`, which REPLACES the inline PostHog block at `app/index.html:193–276` — one `posthog.init`, never two — and keeps that block's key, its `navigator.webdriver` and `file:` early returns and its `#`/`?` scrub): `capture_exceptions` on for unhandled errors and rejections (console errors off), `autocapture: false` (the recorded ruling stands), one explicit `$pageview` per route change carrying only the route's first segment, IP capture off in the project, `identify(<hub identity>)` with a person property `hub_identity` so a flag and a replay belong to a person; **session replay behind a switch that is OFF until Nick or Chantelle's dated word supersedes the 2026-08-09 ruling**, and when on: `maskAllInputs: true`, `maskTextSelector: "*"`, a `maskTextFn` that unmasks only elements carrying `data-record="true"` (the C5 Look page marks its own generated text safe; nothing on the floor is ever marked), console-log and network recording off in the project; a `before_send` that keeps `$exception_list` with its type and stack frames (grouping and source maps need them) but blanks every `value` and every breadcrumb's data, keeps `$session_id`, and drops the whole event when any property matches the browser copy of the floor shapes (`app/js/factory-floor.js`, the same regular expressions as the server's floor scanner, kept in step by a selftest that runs both over one fixture). PostHog groups the exceptions into issues; the runner reads the issues API with a personal API key an agent mints from the existing account and puts in the vault — STEP 5 first probes the real response and, if the list carries no occurrence count, last-seen or session id, reads them through PostHog's query API instead — and makes one factory row per PostHog issue through the door's `report_posthog` action with `posthog_issue_id`, `source: posthog`, `fingerprint` = PostHog's issue id (one definition: a report's fingerprint is ours, a PostHog row's is PostHog's), `count`, `first_seen`, `last_seen` and `replay_url` when replay is on; a resolved PostHog issue that recurs goes through the U20 path. Source maps are generated by esbuild and injected by PostHog's CLI inside `build-dist.js` before the content hashes are taken (a TRUNK change, reviewed once), with the CLI key as a Studio runner secret from the vault. **Fallback, only if PostHog ever refuses us:** the plain beacon this contract replaced (an allow-listed batch to our own door), kept in STEP 5's fixtures. Red-first test: the harness spoofs `navigator.webdriver` to false so the script loads under automation, asserts that a captured exception reaching the network carries its frames and a blank value and nothing matching the floor shapes, and, when replay is on, that a replay frame of a page with a marked-secret element contains only masked text.
- **C3 · The classifier** (PROCESS, `factory-classify.mjs` under `projects/ops/factory/`): input a unified diff (never only a path list; a path list is refused with `refused: give me the diff`); output `{tier, reasons[], files[]}`. Two hard-coded arrays in the file (never env, never config): **TRUNK** — `app/css/one.css`, `app/js/tasks.js`, `app/js/app.js`, `app/js/panel.js` and its symlink target `projects/personal/family-app/js/panel.js`, `app/index.html`, `app/build-dist.js`, `app/gates.js`, `app/functions/api/_session.js`, `app/functions/api/_authgate.js`, `app/functions/api/_permissions.js`, `app/functions/api/tasks.js`, `app/functions/api/comments.js`, `app/functions/api/_taskrun.js`, `app/functions/api/_suggestions.js`, `app/functions/api/_kv.js`, `app/wrangler.toml`, every file under `.github/workflows/`, every `*.selftest.mjs`, everything under `_selfchecks/` EXCEPT the one harness the run itself owns (`_selfchecks/factory/harness-factory-<this run's issue id>.mjs`, added by the harness stage; any other harness in the diff is TRUNK), everything under `projects/ops/factory/` (the factory's own gate and code, whichever repo the diff is in), any `app/functions/api/_*db*.js`; **SHARED** — `app/js/util.js`, `app/js/workflow-voice.js`, `app/js/page-context.js`, `app/js/voice.js`, `app/js/factory-flags.js`, `app/js/factory-posthog.js`, `app/js/factory-floor.js`, `app/functions/api/_factory-flags.js`, `app/functions/api/_factory-issues.js`, `app/functions/api/_sha256.js` (an explicit list only: the earlier "referenced by more than three files" rule made 43 of 170 screen files hard and is dropped; the Sunday job adds to SHARED from evidence). Rules, in order: any TRUNK or SHARED hit → `hard`; more than five files → `hard`; **the money rule** — any file whose route prefix maps to `rizza` in C4 (finance, payroll, receipts, invoices, loans) or any file under `app/functions/api/` whose name contains `loan`, `payroll`, `invoice`, `receipt`, `finance`, `pnl` → `hard`; **the client-words rule** (RULE 48) — a diff on a screen whose prefix maps to `chantelle` in C4 (website, pages, broadcasts, forms, surveys, esign) that adds or changes a string literal → `hard`, and the words are written by the Fable seat, never by the cheap builder; a door's exported signature changes (`onRequest*` params, a new or removed `action` string, a changed `EDITABLE_KEYS`-style constant) → `hard`; a diff adding or removing a D1 `CREATE TABLE` / `ALTER TABLE` / `DROP` → `hard`; a diff touching `sendBeacon`, a `fetch(` to a host other than the Hub's own, `stripe`, `payroll`, `esign` together with `send`, `vault`, or `x-skippy-token` → `hard`; two to five files, or one `app/js/*.js` plus its `app/functions/api/*.js`, or any `--` token in a stylesheet → `medium`; else `easy`. A caller may pass `--raise <tier>` and never `--lower`; the stored tier on the run record is compared and a lower result is discarded. Red-first test: `_test-factory-classify.mjs` (STEP 1) proves each rule from a fixture diff, proves a path list is refused, proves `--lower` is refused, and proves a one-file change to `app/js/util.js` is `hard`.
- **C4 · Reviewer and owner** (HUB, `app/functions/api/_factory-owners.js`, exported const, hard-coded): `ownerFor({screen, severity, reporters})` returns `{owner, reviewer}`. `owner` is the longest matching route prefix: `#finance`, `#payroll`, `#receipts`, `#invoices` → `rizza` · `#workforce`, `#leave`, `#standups`, `#hours`, `#loops` → `mae` · `#sla`, `#nps`, `#huddles`, `#heroes`, `#sidekicks`, `#roster` → `dean` · `#ro`, `#waiting`, `#clients` → `dindin` · `#onboarding`, `#pages`, `#broadcasts`, `#website`, `#forms`, `#surveys`, `#esign` → `chantelle` · `#home`, `#crm`, `#content`, `#captus`, `#projects`, `#tasks` → `nick` · no match → `chantelle`; `#roster` beats `#ro` because the longest prefix wins; `severity: cosmetic` sets `owner` to `chantelle` only when no prefix matched. `#loans` → `rizza`. `reviewer` is `reporters[0]` for a `bug` reported by a person (they can judge "fixed"); for a `wish`, an approved idea or a beacon row it is `owner` (the reporter is copied on the card), so a wish about someone else's screen never goes live on the wisher's tap. The selftest proves every prefix, the longest-prefix rule, the cosmetic rule and both reviewer rules.
- **C5 · The card text** (PROCESS, `factory-card.mjs`, used by every job that writes a Look page or a card comment): every generated body is built by one function from the row and the run record, never from a model's free prose, and has one of four shapes — **intake** (What you said · Where · What happens next), **spec** (What you said · What will change for you · What will not change · How you will know), **ready / live** (What you said · What changed for you (≤ 25 words, from the builder's brief, checked by the reviewer step) · Before / After (two `/api/files/<id>` picture links per viewport) · How to try it (three lines; ready only) · What we checked · Risk · Who gets it after Looks good (ready only) · Something's not right), **plan** (The problem · What was tried, one line per attempt · Options 1–3, each with what it means for you, effort in hours or days, who acts next, "no money involved" · Recommendation). **What we checked** draws only on evidence present in the run record: "the test that failed before now passes" (needs `harness_red_at` and `harness_green_at`) · "one independent AI reviewer re-ran it" / "two independent AI reviewers" / "three: proposed, attacked, read cold" (needs the matching `review_kind`) · "driven on the real screen as you, on a phone and on a computer" (needs four evidence ids) · "nothing else on the screen changed" (needs the tester's outside-the-change screenshot diff at zero). **Risk** likewise: "on for you alone; can be switched off for everyone within a minute" (needs `flag_id`) · "already live; can be undone with one revert if you say so" (needs `revert_pr` prepared) · "a shared part of the Hub; three reviewers checked it; the undo command was written before it went live" (needs `triad` and `revert_command`). A gate refuses a generated section containing a path (`/` followed by a letter and a dot-extension), a model name (`glm|deepseek|qwen|sonnet|opus|fable|astra|codex`), the words `commit`, `PR`, `diff`, `function`, `selector`, or a code fence; **the person's own words are never gated**. The red-first test `_test-factory-card.mjs` (STEP 1) proves each refusal, each evidence-gated phrase absent when its evidence is absent, and a clean body passing.
- **C6 · Flags — AMENDED 2026-09-28 on STEP 6's three-seat review: our own KV record is the ONE source of truth and PostHog's flag feature is not used.** The adversary measured that a switch kept in two places disagrees: racing an off against a promote left PostHog on for everyone while our record read off (78 of 300 trials), an off was impossible while PostHog was down because PostHog was written first, and a real browser reading PostHog never saw a per-person switch because nothing identifies the person to PostHog. So: every door reads `factory:flag:<id>`; the page helper `window.FactoryFlag` reads the flags door, never PostHog; an off is written to our record first and wins over any later set or promote until an explicit re-enable; every change to the flag record carries its own marker, so a write lost to a simultaneous one answers 409, never 200 — AS FAR AS ANY STORE WITHOUT COMPARE-AND-SWAP CAN SEE (stated on the third review, 2026-09-28: a write that lands after the caller's own read-back cannot be seen by that caller, so two simultaneous sets can both answer 200 and the later one's list of people wins; the next read shows the truth; a set may name the revision it read, and is then refused when the flag has moved on; the direction that matters for safety, OFF, is immune because it lives on its own key that set and promote never write, and a re-enable names the one off it cancelled rather than comparing clocks — the off standing when the re-enable runs, not one the caller saw earlier; an off that lands before the re-enable's own read-back makes the re-enable answer 409); a promoted flag does not expire (C6's rule: only a flag neither promoted nor off switches itself off at 14 days, STEP 13); the flags and issues doors accept the factory agent only with the dedicated factory key as well as the agent name, because any teammate's agent pass can claim the name. PostHog keeps the jobs it does well — errors, replays when switched on, usage. The text below is the original contract, kept for the record; where it says PostHog flags, read our record. Original: **C6 · Flags** (PostHog feature flags behind our own thin door, `/api/factory-flags`; a KV record `factory:flag:<id>` keeps only the bookkeeping PostHog does not: `flag_id` (= the issue id, also the PostHog flag key) · `created_at` · `promoted_at` · `expires_at` (created + 14 d) · `off` (bool) · `revision`). The runner creates the flag through PostHog's API targeting the person property `hub_identity` = the reviewer (the property `identify` sets, passed explicitly on every evaluation), promotes it to everyone, or turns it off — three API calls, each read back. Server: `factoryFlagOn(env, id, identity)` evaluates PostHog's flag-definitions JSON itself in `_factory-flags.js` (the Functions carry no npm dependency, so no `posthog-node`; the conditions the factory uses are only `hub_identity` equals and rollout 100 or 0), from a KV copy that ONE writer — the door itself, on a request, with `waitUntil` — refreshes when it is older than 60 seconds (Pages has no clock), keeping the last copy when PostHog does not answer; it is authoritative for every door that carries a change. Client: `posthog-js` evaluates the same flags for the signed-in person; `app/js/factory-flags.js` exposes `window.FactoryFlag.on(id)` over it and, when PostHog is not loaded (an automated browser, the webdriver guard), over the door's `GET`, so drives and harnesses see flags (false until loaded, false for unknown ids, never throws; re-read on every route change and within a minute); a change is written as `if (FactoryFlag.on("fx-…")) { new } else { old }` in the client and behind `factoryFlagOn` in a door. **Looks good** promotes the flag and sets `promoted_at`; **Send back** leaves it; `off` is the rollback and takes effect for every open page within about two minutes (a KV read can be a minute stale on top of the minute's refresh). At `expires_at` neither promoted nor off, the daily job sets `off: true` and returns the card to the reviewer (U28). A flag never covers a data-shape or store change: those are `hard`, and the three-seat review's named revert command is their rollback. A bug (C1 `kind: bug`) carries no flag: its rollback is a prepared revert pull request the runner can merge on the reporter's "something's still wrong" (`revert_pr` on the run record, made at publish).
- **C7 · The run record** (JOBS, one JSON per issue under `projects/ops/skippy-jobs/state/factory/`, read and written only through `factory-state.mjs`, which takes a lease file per issue and increments `revision` on every write; a writer with a stale revision loses and re-reads): `issue_id` · `revision` · `tier_predicted` (at triage) · `tier_actual` (from the final diff; only ever ≥ predicted) · `attempts[]` each `{n, started, ended, builder_model, checker_model, review_kind, harness, harness_red_at, harness_green_at, browser_drive {as, viewports, evidence[], outside_diff}, outcome}` · `triad` (`{proposal, attack, cold, verdicts[], reviewed_sha}` or null) · `pr` · `merged_sha` · `published_at` · `revert_command` · `revert_pr` · `flag_id` · `clock_ms` (card created → live) · `cost_usd` (summed from the spend meter by run id) · `spend_cap_hit` · `wall_clock_hit` · `sent_back[]` (`{by, words, at}`) · `reverted_at`. A daily job (`factory-metrics`) pushes each record's summary (the scalar fields above, no evidence bodies) into D1 `factory_runs` through `factory-issues` `update` so the Pages tile reads D1, never a Mac's disk.
- **C8 · Approvals and hand-offs use the doors, never free comments**: every stop is `action:"handoff"` with `kind` and the five-field update; `factory` is the agent name on every card it makes (`x-skippy-agent: factory`, its own key). A person's decisions travel one way only: the shared approval card's **Approve** / **Discard** (`{nonce, decision}` to `/api/skippy-approvals`), where the brain runs the stored draft as that person — `looks-good`, `done`, `go-ahead`, `choose-recommendation` — carrying `asked_by {from, words}`, the issue `revision` and, for a ready card, the `merged_sha` it was shown for; free text (Send back words, "option 2", Not this words) is taken by the drawer conversation through the brain tools `send_back {issue_id, words}`, `choose_option {issue_id, n}`, and written through the door as the person. The brain never runs an approve on its own. The runner asks the brain to store a draft for a named person through the brain's helper door with `factory_ready {issue_id, person}` (STEP 2 proves the door exists or names the one ask of the core lane); the approve action is registered in `APPROVAL_ONLY` and the tools in `HUB_HELPER_TOOLS`.
- **C9 · The merge gate** (PROCESS, `factory-triad.mjs gate`): the runner may merge only when the run record's `tier` matches the classifier's result on the pull request's head diff (the tier guessed at triage, `tier_predicted`, is never built on: the diff decides), `harness_green_at` is set, the review evidence for that tier is present — **every tier includes one independent diff read** (easy: the cheap checker reads the diff against the reporter's words in the same call that re-runs the harness; medium and hard: Sonnet; RULE 64's "a reviewer who is not the builder, given the diff … returned MERGE") — and for `hard` the `triad.reviewed_sha` equals the pull request's head; a record for another commit is refused (`refused: review is for <sha>, head is <sha>`); a `proceed with changes` verdict is applied, re-attacked once, and must end `proceed`. **The gate is runner-side and honest about it:** `main` is not protected and cannot be on the repository's current GitHub plan (Already true), so the only path that merges is `factory-fix`'s merge stage, which calls `gate` first, and `factory-monitor`'s daily pass reads every commit on `main` since its last run and flags any commit not made by the runner or by a person's own login as `unreviewed` on the tile. Turning on branch protection is money leaving (a paid plan) and is filed through the money door when the loop is proven (STEP 11), never assumed.
- **C10 · Jev's seats in the factory — hints, never decisions** (HUB: seven entries in the registry `_jev-uses.js`, seven batch adapters in `_jev-batches.js`, seven outcome joins in `_jev-outcomes.js`, all SHARED files, so the change is one `hard` change reviewed once by the three seats and landed by STEP 3's lane; the Jev lane keeps the door and Settings. JOBS reads hints only through `GET /api/jev?hints=…`, never `env.AI`). Facts that shape this (Jev plan C2, `_jev.js:224–231`): a use whose input carries the person's words (`untrusted`) is capped at **suggest** in code — only effect `order`, which reorders visible things and hides nothing, may reach act; **no Jev judgment is ever written onto a task, card or record** — a hint is read by the step and the step's own logic decides; an unconfigured use is **off** (not watch) until Nick or Chantelle switch it on in Settings → Judgment, and promotion beyond watch happens only there on the Jev plan's day-14 numbers; on no judgment (off, budget, cache miss, unsure, refused) the step does exactly what it did before. The seven, each with its real fallback: `factory.kind` (choice bug · wish · question; untrusted words → suggest at most: prefills triage's DeepSeek call, which still answers; outcome join = the harness author's label) · `factory.severity` (choice blocks · wrong · cosmetic; prefills; outcome join = the reviewer's severity on the Look page) · `factory.same` (noul over ONE candidate pair per report: the newest open row on the same screen, never a scan; suggest = a `related[]` link on the row, never a merge and never a count change; outcome join = a person's "same thing" tap on the card) · `factory.tier-hint` (choice easy · medium · hard from the words; suggest = triage RAISES `tier_predicted` when Jev says higher, never lowers; outcome join = the classifier's tier on the built diff) · `factory.beacon.noise` (noul over structured fields only — screen, error name, file, line, count; no words — so it may reach act; suggest = the row is tagged "probably noise" and triage takes it last; act = the row is FOLDED (never fixed, tile counts it as noise, kept for 30 days and shown to the screen owner in the daily observer's evidence so a false negative is seen); a silent error never has a card at any stage (C11); outcome join = the owner's Dismiss in the morning evidence or a real fix on that fingerprint) · `factory.sendback.reason` (choice wrong behaviour · looks wrong · not what I meant · broke something else; suggest = the reason is written beside the words in `sent_back[]` marked "Jev suggests"; with no hint the nightly job groups by "unlabelled"; outcome join = the owner's pick when the Look page asks "which of these was it?" on the next round) · `factory.idea.rank` (score with five levels over the THREE candidates Sonnet already chose for that owner; effect `order`: it only orders those three in the Inbox and hides none; outcome join = Approve or Dismiss). Bounds: each use is limited to 200 subjects a day (the clock's default) and one candidate per subject, the hint carries its own effective stage and the runner refuses to treat any hint as more than suggest unless the door says `act` for that hint (an ask of the Jev lane, recorded in its NEXT list: preserve the judgment's stage on the hints door); the factory's share of the $10 monthly ceiling is a constant, $3, checked by the door's own budget before every call. Nothing here touches money, keys, deletion or sending, which the registry forbids by construction.
- **C11 · The boards, the suites and Neeko's rules for the factory.** **Coherence Improvement** is a new department board (`coherence-improvement`) added to `DEPARTMENT_BOARDS` in `app/functions/api/tasks.js:273` and to every mirror that names the seven boards (verified by `git grep coherence-builds` in the Hub on 2026-09-28: `app/js/tasks.js:332`, `app/js/priority-boards.js:38`, `app/js/pm-task-panel.js:302`, `app/functions/api/recurrences.js:57`, `app/functions/api/inbox-feed.js:1461`, `app/functions/api/_department-boards.selftest.mjs:19`, the two harness board lists `_selfchecks/harness-allviewAZ-20260730.mjs:403` and `_selfchecks/harness-selectBL-20260730.mjs:208`, and in the workspace the brain's `projects/ops/skippy-jobs/lib/board-report.mjs:600` with its test `projects/ops/skippy-jobs/_test-ai-builds-lane.mjs:50`, which asserts the exact list) — eleven files, one entry each, plus a re-baseline of `_selfchecks/harness-recurringAN-20260730.mjs` (the Add bar gains one destination, so its byte-frozen tree moves, the way that file prescribes: a byte copy `app/js/tasks.js.baseline-20260928-coherence-improvement`, md5 measured). Two lists that name the boards are deliberately NOT changed, on the skeptic seat's evidence (three-seat review of 2026-09-28): `app/functions/api/projects.js:139` (`GROUP_TYPE` maps a board to a project TYPE, and the type list, the Projects screen's chips and its picker are a frozen contract of the projects-overview plan — a factory project reads "Other" there, as it does today) and `app/functions/api/_jev-batches.js:939` (`W3C_FUNCTION_BOARDS` enrols every open card on the board in Jev's SOP-for-task batch, a behaviour C10 decides, not a list entry) — TRUNK, one three-seat review, owned by STEP 4 — with the six agent stages and a board field **Suite** (a select of seven, set through `board_fields`; the runner writes it with the robot bearer in a second write right after each create, because a board field is set only by the robot bearer, Nick or Chantelle and `create()` takes no fields) whose values are the feature suites — Finance (Rizza) · Workforce & hours (Mae) · Sales, SLA & people (Dean) · Client delivery (DinDin) · Content, sites & forms (Chantelle) · Leadership, CRM & boards (Nick) · Platform (rows whose `repo` is `brain`, the factory's own files) — mapped from C4's prefixes; the board is FILTERED by suite with the filter chips the board already has (grouping by a board field is not built and is not built here), so several suites are active at once on one board; the requester's "one lane per feature suite" is this field, not seven boards. **Coherence Builds** keeps large new builds and human-made updates and the factory never writes there except the plan's own card. **Where bugs and issues live:** every issue is a row in the factory's ledger (D1, linked to PostHog); a row gets a **card on Coherence Improvement, in its suite, only when a person is or will be involved** — a person's report or wish (they want to see it), a change waiting for a spec or a Looks good, a Plan with a person, and any silently caught error whose fix fails or recurs; a silently caught error that fixes itself gets no card at all: it is closed in the ledger, counted on the tile, and listed in the owner's next morning evidence ("3 errors fixed on your screens yesterday"), because a card an agent cannot close would otherwise sit waiting for a tap nobody asked for. Factory cards obey Neeko's checks as they are today, and the two changes that need code are made by STEP 4 on the PM lane's dated word: **as today** — a card waiting on a person (Go ahead / Not this, Looks good / Send back, a plan choice) is handed to that person with a `decision` or `finished` hand-off, so it is theirs and check 12 (an agent's card quiet for 24 hours) never fires; `factory` runs it again with `take` after their answer; progress lines are comments and hand-offs happen only at the five kinds the door has; an `agent-test` card is closed or handed off `finished` within two hours (check 14), so STEP 11's card is handed to Mae before its second hour; "Nick's Review" on this board means "waiting for the person named on the card" and the Look page's first line says who; the daily heartbeat is the Factory tile, not a card update. **Two code changes in `board-scan.mjs` and `_test-board-scan.mjs`, STEP 4, on the PM lane's word:** a factory card without a Suite is pinged, and a card on `coherence-improvement` is never put in Nick's end-of-day report after three pings (its escalation is the Plan with a person card). **One rule the Kanban spec's owner is asked to adopt** (a dated ask on the PM lane's NEXT list): an exception to §11.2 for a silently caught error the factory fixes with no person involved — it opens no card; the ledger row, the tile and the owner's morning evidence are its record.
**Buckets that share a goal message each other:** a dated line into `projects/business/business-app/KANBAN-AND-AGENT-BOARDS-SPEC.md`'s owner (the PM lane) when C11 lands, carrying the seven rules above; a dated line into `app/_design/jev/PLAN.md`'s NEXT list when C10's seven uses land (the Jev lane owns the door and Settings; the factory owns only its entries); a dated line into `projects/business/screen-buddy/PLAN.md` §6.1 when STEP 2's helper lands (the core lane owns the drawer) or when STEP 2 must ask it for the job-callable draft; into `projects/ops/agent-fleet/failure-learning/PLAN.md` when STEP 15 adds the factory section to the Sunday card; into `app/_design/workspace-setup/SPEC.md` §3.3 when C6 lands (so the toggles lane knows the per-change flag exists and is not theirs).
## 3b · Execution map — the Step map, then one STEP block per row
> 🔴 The table is the INDEX; the `### STEP` blocks are the INSTRUCTIONS — 1:1. EXECUTOR comes by name from `projects/ops/walkaway/MODEL-MATRIX.md`; the CHECKER never shares a model (glm≡zai) or a session with its row's EXECUTOR; DONE-PROOF is a runnable command in backticks. Keep the next line verbatim — the gate looks for it:
A task is DONE only when its review-ledger row is CLOSED by a reviewer that is not the builder.
**Step map (read this first):**
| Stage | # | Task (step name) | Needs (named artefact, or `none — start now`) | EXECUTOR (cheap model) | EXECUTOR BACKUP | CHECKER (different model) | CHECKER BACKUP | DONE-PROOF (runnable command) |
|---|---|---|---|---|---|---|---|---|
| Plan | 1 | The harness kit: every proof written red-first (KIT) | none — start now | Sonnet (`TEST-AUTHORING:`) | Opus | Qwen 3.8 | DeepSeek V4 Pro | `cd "$WORKSPACE" && ls projects/ops/factory/_test-factory-*.mjs projects/ops/skippy-jobs/_test-factory-*.mjs projects/business/business-app/_selfchecks/harness-factory-*.mjs projects/business/business-app/app/functions/api/_factory-*.selftest.mjs | wc -l` → `15`, plus `prove-e2e.mjs` under `projects/ops/factory/` and `_test-hub-factory.mjs` in the brain repo (17 in all); each run against the unbuilt surface exits non-zero naming what is missing |
| Framing | 2 | Screen Buddy takes a report, a wish, a send-back, a choice; the job-callable draft (BRAIN) | STEP 3's door live on the local server; `_test-hub-factory.mjs` CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 | Sonnet | `node _test-hub-factory.mjs && node _test-hub-factory.mjs --sabotage; echo $?` CREATED BY STEP 1 → `PASS 6 tools` then non-zero |
| Framing | 3 | The issue door, the owners map and the KV store (HUB) | the selftest CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 | DeepSeek V4 Pro | `node app/functions/api/_factory-issues.selftest.mjs` → `PASS 43/43` |
| Framing | 4 | The Coherence Improvement board, the Suite field, the card and its Details renderer (HUB) | STEP 3's door; the harness CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 (browser) | Sonnet | `node _selfchecks/harness-factory-board-20260927.mjs` → `PASS board present, suite field, 3 fixture cards on coherence-improvement in their suites, details fetched, no code strings, door-down message` |
| Framing | 5 | PostHog: the existing sensor moved into the factory's script, issues read, source maps in the build (HUB + JOBS) | the two API keys minted into the vault by an agent; STEP 3's door; the harness CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 (browser) | DeepSeek V4 Pro | `node _selfchecks/harness-factory-beacon-20260927.mjs` → `PASS injected 2 errors → 1 row count 2 linked to 1 posthog issue; exception frames kept, value blank, floor clean; replay off-silent|on-masked; fallback beacon green` |
| Elements | 6 | Per-change flags (HUB) | the selftest CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 | Sonnet | `node app/functions/api/_factory-flags.selftest.mjs` → `PASS 13/13` |
| Elements | 7 | The Look page and the review door: Try it · Looks good · Send back (HUB + PROCESS) | STEP 4's renderer; STEP 6's flags; STEP 2's tools; C5; the harness CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Sonnet (browser, as Dean, phone and computer) | Qwen 3.8 | `node _selfchecks/harness-factory-review-20260927.mjs` → `PASS look page 8 sections at 375 and 1280; try-it dean:true mae:false; looks-good → all; second tap 410; send-back words → reopened; body clean` |
| Elements | 8 | The classifier and the tier runner (PROCESS) | the tests CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 | Sonnet | `node projects/ops/factory/_test-factory-classify.mjs && node projects/ops/factory/_test-factory-review.mjs` CREATED BY STEP 1 → `PASS 12 rules, path-list refused, --lower refused, util.js → hard` and `PASS easy=1 check, medium=2, hard-without-record refused` |
| Elements | 9 | The three-seat review runner, the merge gate and the headless drive (PROCESS) | STEP 8; `astra-run.sh`; the Claude CLI on the Studio (verified in the step); the test CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Sonnet | Qwen 3.8 | `node projects/ops/factory/_test-factory-triad.mjs` CREATED BY STEP 1 → `PASS: no-record merge refused; stale-sha refused; record → allowed; changes → re-attacked once; 3rd loop → ecosystem-fixes card; drive → 4 pictures` |
| Details | 10 | Triage and the fix runner as a state machine, registered and beating (JOBS) | STEPS 3, 6, 7, 8, 9; the test CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 | Sonnet | `node projects/ops/skippy-jobs/_test-factory-fix.mjs` CREATED BY STEP 1 → `PASS 14 transitions; 3 cards per tick; green-harness rejected; 3 attempts → plan; 48h → plan; $5 → plan; vendors-down → wait; tier-rise → spec; merge → publish → handoff`, and eight runner-written heartbeat rows |
| Details | 11 | One seeded easy bug, end to end, live (JOBS + all) | STEPS 2–10 landed and deployed; `prove-e2e.mjs` CREATED BY STEP 1; `gh auth status` on the Studio | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Sonnet (`biz-app-qa`, browser, as Mae) | Qwen 3.8 | `node projects/ops/factory/prove-e2e.mjs --issue <id>` CREATED BY STEP 1 → `PASS live; clock <ms>; look page clean; pictures 4; revert_pr prepared; card with reporter` |
| Details | 12 | Recurrence → one Plan with a person card; third loop → Ecosystem Fixes (JOBS) | STEP 10's state; C7; the test CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 | Sonnet | `node projects/ops/skippy-jobs/_test-factory-recur.mjs` CREATED BY STEP 1 → `PASS 4 triggers → 1 card each on coherence-improvement; 3rd loop → 1 card on ecosystem-fixes to nick; repeat → 0; choose 2 → spec row; none → reopened` |
| Details | 13 | Monitor after publish: recurrence, revert, stale flags, the factory's own daily harness list (JOBS) | STEPS 5, 10; the test CREATED BY STEP 1 | DeepSeek V4 Pro | Qwen 3.8 | GLM 5.3 (`zai`) | Sonnet | `node projects/ops/skippy-jobs/_test-factory-monitor.mjs` CREATED BY STEP 1 → `PASS beacon recurrence reopens; harness red → revert merged + recur; stale flag → off + card back; quiet day → 0 writes` |
| Output | 14 | The Factory tile on Home, leadership only (HUB) | STEP 3's rows; C7's run summaries in KV; the harness CREATED BY STEP 1 | GLM 5.3 (`zai`) | DeepSeek V4 Pro | Qwen 3.8 (browser) | Sonnet | `node _selfchecks/harness-factory-tile-20260927.mjs` → `PASS 7 numbers = recomputed; 403 and no tile for mae and rizza` |
| Output | 15 | The daily observer, the daily scout, the nightly correction, through the suggestion door (JOBS) | STEP 3's rows; C7; C10 at watch; `failure-weekly-review.mjs`; the test CREATED BY STEP 1 | Qwen 3.8 (gather) | DeepSeek V4 Pro | Sonnet | GLM 5.3 (`zai`) | `node projects/ops/skippy-jobs/_test-factory-weekly.mjs` CREATED BY STEP 1 → `PASS ≤3 per owner per day; quiet day → 0; dismissed not refiled; scout ≤3 each with screen|gap; nightly: 1 classifier change per evidenced mis-tier, 1 brief change per repeated reason; sunday comment on the week's card, 7 numbers` |
| Proof | 16 | Pointers, registry, toolkit (POLISH) | STEPS 2–15; the test CREATED BY STEP 1 | DeepSeek V4 Pro (registry, toolkit, manual bullet) with the Hub `CLAUDE.md` §7 written by the overseer (Fable; CORE §3) | Qwen 3.8 | GLM 5.3 (`zai`) | Sonnet | `node projects/ops/skippy-jobs/_test-factory-wiring.mjs` CREATED BY STEP 1 → `PASS CLAUDE.md §7 present; registry row present; toolkit 3/3; manual bullet present` |
| Proof | 17 | Final sign-off of the FINISH LINE | STEPS 1–16 closed | Fable (reads proofs; never builds) | Opus | Opus (reads the same proofs cold; the one check that stays on Anthropic) | Fable | `python3 projects/ops/agents/check_plan.py <this file>` → `PASS` and the nine FINISH LINE lines each cite a closed step |
**Then one block per step:**
### STEP 1 — The harness kit: every proof written red-first
**FOR NICK:** nothing he notices; this is what makes every later "it works" a claim a machine can refuse, written by a reader who did not build the thing. · **Tier:** POLISH
**Start when:** none — start now.
**Builder:** Sonnet (the matrix's test-authoring row; the brief carries `TEST-AUTHORING:` naming only these files) · **Builder backup:** Opus · **Checker:** Qwen 3.8 · **Checker backup:** DeepSeek V4 Pro
**Files you may touch (all new, and only these seventeen):** under `projects/ops/factory/`: `_test-factory-card.mjs`, `_test-factory-classify.mjs`, `_test-factory-review.mjs`, `_test-factory-triad.mjs`, `prove-e2e.mjs`, and the folder `fixtures/` (twelve diff fixtures, one per C3 rule, the C5 refusal and evidence fixtures, the fake reviewer and fake tool replies); under `projects/ops/skippy-jobs/`: `_test-factory-fix.mjs`, `_test-factory-recur.mjs`, `_test-factory-monitor.mjs`, `_test-factory-weekly.mjs`, `_test-factory-wiring.mjs`; in the Hub repo under `_selfchecks/`: `harness-factory-board-20260927.mjs`, `harness-factory-beacon-20260927.mjs`, `harness-factory-review-20260927.mjs`, `harness-factory-tile-20260927.mjs`; in the Hub repo under `app/functions/api/`: `_factory-issues.selftest.mjs`, `_factory-flags.selftest.mjs`; in the brain repo (the Mac mini's running copy): `_test-hub-factory.mjs`. **Never** any product file; never a file another step names as its own.
**Do exactly this:**
1. For each STEP 2–16 block below, write the harness or selftest its PROOF names so that it asserts exactly that block's DEFINITION OF DONE and prints exactly its expected PASS line, and exits non-zero naming the first missing thing when the surface it tests does not exist yet ("refused: no door at /api/factory-issues", "refused: factory-classify.mjs not found").
2. Every harness that drives a browser signs in through `projects/ops/skippy-jobs/lib/hub-session.mjs` on the local server, serialises through `/tmp/bzvisual-chrome.lock`, marks every card it makes `test_card: true` with `agent-test` in the name, drives phone (375×812) and computer (1280×800) where the step says so, and ends with a cleanup line `cleanup LEFT <n>`.
3. Every test that exercises a job or a review uses fakes behind `FACTORY_FAKE_TOOLS=1` / `FACTORY_FAKE_REVIEWERS=1` so no vendor is called; every fake is a file under `fixtures/`, never inline prose.
4. Run each file once against the unbuilt surface and record its non-zero exit and its refusal line in the brief's hand-back.
**DEFINITION OF DONE:** all seventeen kit files exist, and each one run against the unbuilt surface exits non-zero naming what is missing (a kit file that passes before the build is a check that cannot fail and is rewritten).
**PROOF:** `cd "$WORKSPACE" && ls projects/ops/factory/_test-factory-*.mjs projects/ops/skippy-jobs/_test-factory-*.mjs projects/business/business-app/_selfchecks/harness-factory-*.mjs projects/business/business-app/app/functions/api/_factory-*.selftest.mjs | wc -l` → `15`, and `ls projects/ops/factory/prove-e2e.mjs` and the brain repo's `_test-hub-factory.mjs` both present (17); then each file run in its working directory → exit non-zero with a `refused:` line · **FAILS IF:** fewer than seventeen, or any file exits 0 before its step is built.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
**VERIFIED:** 2026-09-28 (100%, 17 of 17 kit files exit 1 with a `refused:` line against the unbuilt surface, proven by a red-first checker script run once per file in its working directory; the Hub six on branch `factory-kit` 0f6a915f1, the workspace eleven on origin/main 9ed242d0f6; the cheap lane was tried six times on the first file and never wrote it — every vendor timed out or wandered the repo for 40 steps — so the files were written on Anthropic: six workspace tests under recorded `cheap-vendor-failed` overrides after route-build's own attempt failed its proof twice on each, five job-runner tests as control-plane files the router keeps inside, and the six Hub files the same way; a one-page recipe per file lives under `projects/ops/factory/briefs/kit` so the step builders and a cheap checker can read what each test asserts; the checker's own re-run of the seventeen is owed and named in the SUMMARY.)
### STEP 2 — Screen Buddy takes a report, a wish, a send-back, a choice; the job-callable draft
**Measured 2026-09-28 (binding on the brain's tool):** a session minted through the Hub's session door for a person carries the agent mark of the minter (Skippy), and the issue door accepts a `report` only from a plain person session or the `factory` agent. The brain's six tools therefore sign in as the agent `factory` riding the person they act for (`x-skippy-agent: factory` on the agent-session call, the person's identity in the body), never as the person alone and never as Skippy.
**FOR NICK:** anyone on the team, on any screen, on a phone, says one sentence and it is logged where it will be fixed; the same drawer takes their "send it back because…" and their "option 2"; nobody fills in a form. · **Tier:** FRONT
**Start when:** STEP 3's door answers `POST /api/factory-issues {action:"report"}` on the local server (`curl -s http://127.0.0.1:8788/api/factory-issues -d '{"action":"list"}'` returns JSON); `_test-hub-factory.mjs` exists (CREATED BY STEP 1).
**Builder:** GLM 5.3 (`zai`) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Qwen 3.8 · **Checker backup:** Sonnet
**Files you may touch:** in the brain repo: `hub-factory.mjs` (new) under its `lib/` folder; `server.js` (the import, the dispatch case, `TOOLS.push`, `CHANGE_TOOLS`, `HUB_HELPER_TOOLS`, `APPROVAL_ONLY` lines only); `projects/ops/bin/brain-land.sh` (one test line). **Never** `voice.js`, `workflow-voice.js`, `page-context.js` (the Screen Buddy core lane); never `_test-hub-factory.mjs` (STEP 1's).
**Do exactly this:**
1. Copy the helper recipe in `projects/business/screen-buddy/PLAN.md` §3.4 and §7: one file `lib/hub-factory.mjs` exporting six tools. `report_issue {words}` and `suggest_idea {words}`: sign in as the person (the brain's `hub-person` helper), read the screen context the brain already holds (`screen`, `record_label`, `summary`, record kind and id from `SkpScreenRecord`), post `POST /api/factory-issues {action:"report", kind:"bug"|"wish", words, screen, record_kind, record_id, record_label, summary, request_key}` (the `request_key` is the brain turn's id, so a retry after a 5xx finds the same row), read the row back by id, and answer only after the read-back with one sentence naming the card. `explain_change {issue_id}`: reads the row and the run summary's plain-words fields (C5) and answers in plain words; never quotes a diff. `send_back {issue_id, words}` and `choose_option {issue_id, n}`: write through the review door as the person with `asked_by {from, words}` and the row's `revision`. `factory_ready {issue_id, person}`: called by the runner through the brain's helper door, stores a draft for that person whose §5.1 shape is `{summary, look_url: #factory/<id>, test_url, approve:{action:"looks-good"|"done"|"go-ahead"|"choose-recommendation", input:{issue_id, revision, merged_sha}}, discard:{action:"send-back-start"}}`; the approve actions go into `APPROVAL_ONLY`; `send-back-start` makes the drawer ask for the sentence and then calls `send_back`.
2. If the brain has no job-callable way to store a draft for a named person, write the one ask into `projects/business/screen-buddy/PLAN.md` §6.1 in one dated line and build the five other tools; the step stays OPEN at five tools until the core delivers the draft door (STEP 7 and everything after it depend on `factory_ready`; a review card that cannot be stored is a loop that stalls in silence).
3. Add the tools to `server.js` (import, dispatch, `TOOLS.push`, `CHANGE_TOOLS` for the four that write, `HUB_HELPER_TOOLS` for receipts, `APPROVAL_ONLY` for the approve actions) and one line to `brain-land.sh`.
4. Land with `projects/ops/bin/brain-land.sh`; then drive it live as Mae on the Mac mini with the brain's live check (`_live-screen-buddy.mjs --case factory-report`, the live check gaining one case in this step).
**DEFINITION OF DONE:** as Mae on a phone, the sentence "something's wrong: the tag button does nothing" on Candidates makes one issue row and one card within 60 s, read back by id, and the drawer's reply names the card; a stored draft for Dean shows in his drawer as the shared card with Look, Test, Approve and Discard.
**PROOF:** in the brain repo: `node _test-hub-factory.mjs && node _test-hub-factory.mjs --sabotage; echo $?` (CREATED BY STEP 1) → `PASS 6 tools: report_issue row fx-… read back; card ac-… read back; retry same request_key → same row; factory_ready draft nonce …` then a non-zero exit on sabotage · **FAILS IF:** a tool answers ok before the read-back, a retry makes a second row, or sabotage exits 0.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** the moment this step closes, post one dated line into `projects/business/screen-buddy/PLAN.md` §6.1: `Coherence Factory STEP 2 closed <date> — report_issue, suggest_idea, explain_change, send_back, choose_option, factory_ready live in the drawer; the core owns nothing new` (or the one ask, if step 2 applied).
### STEP 3 — The issue door, the owners map and the KV store
**FOR NICK:** every report and every error has one home and one count, so the same thing is never chased twice. · **Tier:** POLISH (nothing he notices; STEP 2 and STEP 5 are what he sees)
**Start when:** `app/functions/api/_factory-issues.selftest.mjs` exists (CREATED BY STEP 1); STEP 4's board entry (`coherence-improvement` in the thirteen files) is on `main`, because `report` makes its card on that board and the door's create is refused for a board the task door does not know.
**Builder:** GLM 5.3 (`zai`) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Qwen 3.8 · **Checker backup:** DeepSeek V4 Pro
**Files you may touch:** `app/functions/api/factory-issues.js` (new), `app/functions/api/_factory-issues.js` (new, the store, the fingerprint, the revision — the ONE choke point for the factory's database, the `_hsdb.js` convention), `app/functions/api/_factory-owners.js` (new, C4). No new binding, no migration: the store is the Hub's own KV (`BIZ_KV`, the store every task, suggestion and profile already lives in) under `factory:` keys, so nothing is created on Cloudflare and the selftest runs in-process over a mock KV like every other Hub selftest (decided 2026-09-28 by the overseer when the kit was written: the earlier D1 choice bought a compare-and-swap the choke point's `revision` check gives anyway, at the price of a database creation, a migration and a selftest that could not run without a server; the rows are one per issue, never per occurrence — PostHog holds those — so KV's read-after-write lag of about a minute is covered by every write reading its row back before answering). **Never** `HS_DB` (hs-database is the clients' and Sidekicks' master database with one choke point and a schema owned elsewhere; nothing of the factory's lives in it), `wrangler.toml`, `tasks.js`, `comments.js`, `_suggestions.js` (the board owner), the selftest (STEP 1's).
**Do exactly this:**
1. Write `_factory-issues.js` as the ONE choke point over KV: keys `factory:issue:<id>` (the C1 row), `factory:issues` (the index: ids with `fingerprint`, `status`, `request_key`, `card_id`, so `list` and the fingerprint and request-key lookups read one key), `factory:flag:<id>` (C6's bookkeeping), `factory:run:<id>` (C7's scalar summary); `fingerprintFor(row)`, `upsertReport()`, `upsertBeacon()` (merges counts), `upsertPosthog()`, `reopen()`, `list()`, `get()`, `update()` with `expected_revision` (a stale one → 409, never a silent overwrite); every write reads the row back before answering; a row's `words` are masked with the same nine-plus-digit rule Screen Buddy uses before they are stored. **Exact shapes for the builder (added 2026-09-28 when the kit was written; the selftest asserts them):** the store has NO shared index key (redesigned 2026-09-28 after two three-seat review rounds measured lost and dangling rows): a row's id is `fx-` + the first 12 hex of the hash of its fingerprint, so every report of one thing lands on one row key; the row is written first through the Hub's `rmw` helper by a pure mutator that merges a reporter, counts once per `request_key` (a double-tap changes nothing) and reopens a closed row; rows are listed by the store's key listing (prefix `factory:issue:`), and the `request_key` and `card_id` lookups are single-value keys `factory:rk:<key>` and `factory:card:<id>` written after their row; a damaged row is refused, never replaced; every merge carries one occurrence key minted before the write (a report's is its caller-scoped `request_key`, so one key from two people is two things), a PostHog call sets the issue's total and reopens a closed row only for an occurrence later than the close, update/reopen/setCard answer 409 when their change did not survive, and a count can come out low under simultaneous writes (KV has no compare-and-swap), never high, never the row; a report's fingerprint is its screen, the record it was about (`record_kind`, `record_id`) and its words, so one sentence about two records is two things (decided 2026-09-28 on the review); a row's id never changes, so the step that owns U20 (a recurrence after Done gets a NEW card) gives each recurrence its own card id — the issue id plus the `reopened` count — never the bare issue id; `fingerprintFor` prefixes a report `r:` (sha1 of screen + `|` + the words lowercased, digit runs collapsed to `#`, whitespace collapsed), a beacon `b:` (sha1 of screen `|` name `|` file `|` line) and a PostHog row `p:` + the issue id, so the three kinds never collide; `maskWords` turns every run of nine or more digits into nine `#`; `upsertReport` with a known `request_key` returns the existing row without a write, and a second report of the same thing adds the reporter and `count + 1` on the same row; a beacon or a report on a `done` row reopens it (`reopened + 1`); `update` applies only known fields, never `id`, `revision`, `created_at` or `fingerprint`, and throws an error carrying `status: 409` on a stale `expected_revision`; `setCard` writes `card_id` and its lookup key; `getRecord`/`putRecord` serve the flags door and the runner for `factory:flag:*` and `factory:run:*`. The door `factory-issues.js` reads the session the way the other doors do (`_session.js`, the `db_ident` cookie; an agent session's `agent` field equals `factory`), answers `{ ok, data }` / `{ ok: false, error }`, 401 with no session, 405 off POST, 400 on an unknown action or a missing `screen`/`words`, 403 to a person on `report_posthog`/`update`, 404 on `get` with no row; `report` makes the card in-process through `tasks.js` `onRequest` with the same cookie and the body `{ action: "create", board: "coherence-improvement", name: <masked words, "agent-test: " kept when test is true>, assignee: "factory", due_date: <today + 2 days>, test_card: <test>, tags: ["factory", "stage:inbox", "kind:<kind>"], monday_status: "Backlog" }` and stores the returned id; `look` answers `{ body }` (the row's `run.look_body` or an empty string until STEP 7 fills it). **Four facts measured while building this step (2026-09-28), binding on every later Hub step:** (1) the Hub's session secret setting is `APP_GATE_SECRET`, and an agent session is a known identity carrying the agent's name (`signSession(secret, "nick", undefined, "factory")`), never an identity called `factory`; (2) only the `factory` agent may own a card under its own name, so the door makes the card with an agent session riding the caller's identity; (3) the tasks door allows one open card per `session_id` and only the tag `business`, so each issue is its own thread (`session_id` = the row's id) and a factory card is recognized by its board and the row's `card_id`, never by a tag — wherever this plan says a card "carries the tag factory", read "sits on Coherence Improvement and is named by a row's `card_id`"; (4) the tasks list hides test cards, so a check that counts cards reads the stored task record. **And one rule for landing the kit:** every `app/functions/api/*.selftest.mjs` is a TIER-1 build gate found by folder scan, and a `_selfchecks/harness-*` file in no tier is refused by the gate-coverage ratchet, so a Hub kit file reaches `main` only in the same change as the step it proves, green, and (for a harness) registered in its tier; until then the Hub kit waits on the branch `factory-kit`.
3. Write `factory-issues.js` with actions `report` (a signed-in person or an agent session as the person), `beacon` (any signed-in session; body must carry exactly the C2 keys or it is refused 400), `list`, `get`, `update` and `reopen` (agent key `factory` only; 403 otherwise; 409 on a stale revision); `report` opens the card through the tasks door exactly as C1 says (id = the issue id, `coherence-improvement`, the Suite from C11, `factory`, the tags, `what`, `test_card` when the words begin `agent-test`) and stores `card_id` after the read-back.
4. Write `_factory-owners.js` (C4) with `ownerFor()`.
**DEFINITION OF DONE:** the selftest passes 43/43 (plan change, 2026-09-28, third: eight cases added after the adversary's third round — racing beacons never count more than their occurrences, a PostHog total is set not added and an unchanged total leaves a done row closed, one request key from two people makes two rows, a factory test session's card is named agent-test, a failed card write answers 502 with no card id stored, the beacon and words caps, a damaged row named in the list; every merge now carries an occurrence key minted before the write, and the stated limit reads: a count can come out low under simultaneous writes, never high, and no row is ever lost. Plan change, second: six cases added after the adversary's re-check — ten racing reports of one thing make one row, six racing double-taps count once, eight racing different reports are all listed, a damaged row is refused, an unknown status or field is refused, a tier only rises — all over a store answering with random delays; the store was redesigned so a row's id is its fingerprint's hash and no shared index key exists, and a count may undercount under simultaneous identical reports, never the row. Plan change, first: eleven cases added after the three-seat review showed the first eighteen could not see a real report without a card, lost rows under simultaneous reports, or the authorization answers — a real report makes a card without the test mark; 401 with no session; 405 off POST; 400 on an unknown action, missing words, floor words and a beacon event missing a field; 403 to a person on `report_posthog` and `reopen`; 404 on an unknown id; ten simultaneous reports all listed; and before them the original eighteen: two identical beacons → one row count 2; a beacon body with an extra key → 400; a done row + same fingerprint → `reopened: 1`; a report and a beacon on the same screen never share a fingerprint; the same `request_key` twice → the same row and one card; `update` without the agent key → 403; a stale `expected_revision` → 409; a nine-digit run in words → masked; every C4 prefix, the longest-prefix rule, the cosmetic rule and the reviewer rule) and a `report` on the local server produces one row and one card whose `card_id` reads back from `/api/tasks`.
**PROOF:** `cd projects/business/business-app && node app/functions/api/_factory-issues.selftest.mjs` → `PASS 43/43` · **FAILS IF:** any case fails, or two rows share a fingerprint while neither is done.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 4 — The Coherence Improvement board, the Suite field, the card and its Details renderer
**FOR NICK:** the factory has its own board, one filter chip per feature suite, in the team's own words and the stages the Hub already uses; Coherence Builds stays his board for big builds. · **Tier:** FRONT
**Start when:** `_selfchecks/harness-factory-board-20260927.mjs` exists (CREATED BY STEP 1); the PM lane's dated word on the two `board-scan.mjs` changes (asked at lane-open; until it arrives, the board and the field land and the two scan changes wait).
**Builder:** Sonnet for the board list and the scan changes (TRUNK and a control-plane job; the brief carries a `CHEAP-FIRST-OVERRIDE: FLOOR` line naming `tasks.js` only if the gate's scan finds a floor value there, else `NICK-ASKED` is not available and the work goes to GLM with the three-seat review before merge as C9 requires) · **Builder backup:** GLM 5.3 (`zai`) · **Checker:** Qwen 3.8 (browser) · **Checker backup:** DeepSeek V4 Pro
**Files you may touch:** the thirteen board-list files C11 names, one added entry each (`coherence-improvement`, "Coherence Improvement"; the two harness lists and the brain's test gain the entry so they stay green); the Suite field definition written through `board_fields` (a select of the seven suites); `app/js/task-panel-factory.js` (new: when the open card carries the tag `factory`, fetches `/api/factory-issues {action:"get", card_id}` and renders the C5 intake or spec shape as a read-only Details section; the door down → one line "Details are not available right now"); `app/index.html` (the script tag for the new file); `projects/ops/skippy-jobs/jobs/board-scan.mjs` and its `_test-board-scan.mjs` (the two C11 changes, red-first, only with the PM lane's word). **Never** any other line of either `tasks.js`, `panel.js`, the harness (STEP 1's).
**Do exactly this:**
1. Add the board to the thirteen files in one change (its three-seat review runs once, C9; the Hub files land as one pull request, the two workspace files as one commit), define the Suite field, and write `task-panel-factory.js` as above; it refuses to render any generated string that fails the C5 gate (render "—" and log one console warning); the person's own words render untouched.
1a. With the PM lane's word: the two `board-scan.mjs` changes, each proven red-first in `_test-board-scan.mjs`.
**Progress, 2026-09-28 (part 1 of STEP 4 landed and live):** the board entry went through its three-seat review (the cold verifier agreed; the adversary's five required changes were applied — two lists deliberately left alone, the recurring harness re-baselined — and its re-check's one comment fix too), merged as nick-deck/deck-business#2397, and is in the live build (source revision 7dc2c8bd0, which contains the merge). Proven on the live Hub: a test card created on `coherence-improvement` answered 201 and read back on that board, then was archived through the Hub's own archive action; the board reads empty again. The workspace mirror (`board-report.mjs` and its lane test) landed on main the same hour. Found on the way: main's publish had been blocked for every lane by the source-parity gate (the inbox rule gained a helper the gate did not supply); fixed by handing the gate the real helper, merged as nick-deck/deck-business#2400. Still open in STEP 4: the Suite field, the Details renderer and the two scan changes.
**Progress, 2026-09-28 night (part 2 built, in check):** the Details renderer (`app/js/task-panel-factory.js`) was written whole by a cheap model on the new whole-file mode and made loop-proof in-house after a real browser measured the page freezing with it; it is on the Hub branch factory-step4b (864a8a5cc) with its script tag. The board harness had never run its browser half (a page API the rig lacks, and a crash that printed PASS); it was rebuilt to drive the Hub's own browser rig and the real doors (kit branch e9a211e2e) and now passes 8/8 on the branch, with the Details cases red without the renderer. An independent check is running; the two board-scan changes still wait on the PM lane's word.
2. Run the harness: on the local server it seeds three `agent-test` cards (inbox, spec, plan) through the door, signs in as Nick, opens `#tasks`, ticks Coherence Improvement, opens each card and asserts the Details show the right C5 shape, that the generated sections contain no `/`-path, no model name and no code fence, that the door being down (503 through the harness's interceptor) shows the one line, and cleans the cards up (they carry `test_card: true`).
**DEFINITION OF DONE:** the board and its Suite field exist, the three fixture cards on Coherence Improvement sit in their suites and render their factory Details from the door, the door-down line shows, and the cards are removed at the end of the run.
**PROOF:** `cd projects/business/business-app && node _selfchecks/harness-factory-board-20260927.mjs` → `PASS board present, suite field, 3 fixture cards on coherence-improvement in their suites, details fetched, no code strings, door-down message, cleanup LEFT 0` · **FAILS IF:** a generated section carries a path or a model name, the door-down state renders anything else, or `LEFT` is not 0.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 5 — PostHog: errors caught and grouped, replays masked, source maps on every publish
**FOR NICK:** things that break on a screen get logged and grouped even when nobody says anything, the same error twice is one thing, and the day he or Chantelle says the word, a masked replay of what the person actually did comes with every error. · **Tier:** FRONT (he sees the fixes on the tile and, when one fails, the card)
**Start when:** the runner's personal API key and the source-map CLI key are minted by an agent from the existing PostHog account and put in the vault (no person touches a key, CORE §2); STEP 3's door answers `report_posthog`; `_selfchecks/harness-factory-beacon-20260927.mjs` exists (CREATED BY STEP 1). Replay stays off until Nick or Chantelle's dated word, and this step does not wait for it.
**Builder:** GLM 5.3 (`zai`) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Qwen 3.8 (browser) · **Checker backup:** DeepSeek V4 Pro
**Files you may touch:** `app/js/factory-posthog.js` (new, C2) and `app/js/factory-floor.js` (new, the browser copy of the floor shapes); `app/index.html` (the inline PostHog block at lines 193–276 is REPLACED by one script tag for the new file, first in the body so boot errors count — TRUNK, reviewed once); `app/build-dist.js` (esbuild source maps plus PostHog's inject before the content hashes — TRUNK, reviewed once in the same review); the runner's issue reader in `projects/ops/skippy-jobs/lib/` (`factory-posthog.mjs`, new: probes the issues API's real response once, reads count, last-seen and session through the issues list or PostHog's query API, makes rows through the door's `report_posthog`). **Never** `page-context.js` (Screen Buddy core), any screen's own file, the harness (STEP 1's). The plain beacon (the pre-PostHog C2) lives only in the harness fixtures as the fallback.
**Do exactly this:**
1. Write `factory-posthog.js` and `factory-floor.js` to C2 exactly, moving the existing block's key, guards and scrub into the file; the replay switch reads a constant `REPLAY_ON` that stays `false` until the dated word is recorded in this plan; turn off console-log and network recording in the project settings.
2. Write the runner's reader: probe the real issues response first and record its fields in the run; every tick, read what changed and for each issue make or update one factory row through `report_posthog` (`posthog_issue_id`, `source: posthog`, count, first and last seen, `replay_url` when replay is on); a resolved issue that recurs → the U20 path.
3. Put source-map generation and inject inside `build-dist.js` before hashing, with the CLI key as a Studio runner secret from the vault; this and the `index.html` replacement are one TRUNK change with one three-seat review.
4. Run the harness: on the live Hub signed in as Nick on an `agent-test` record, with `navigator.webdriver` spoofed to false so the script loads under automation, it injects `setTimeout(()=>{throw new Error("factory-test 123456789012")})` twice, waits one runner tick, asserts one factory row with count 2 carrying a `posthog_issue_id` and no digits anywhere, asserts the network body of the captured exception carried its frames, a blank value and nothing matching the floor shapes; when `REPLAY_ON` is true it also opens a page with a marked-secret element under replay and asserts the replay frame holds only masked text, else it asserts no recording request left the page; cleans the `agent-test` row and resolves the PostHog issue. The fallback fixture (the plain beacon against our own door) is run once too, so the fallback is proven live.
**DEFINITION OF DONE:** two injected identical errors make one factory row with count 2 linked to one PostHog issue, the exception that left the page kept its frames with a blank value and nothing on the floor, replay is either proven masked (when on) or proven silent (when off), and the fallback beacon still proves green.
**PROOF:** `cd projects/business/business-app && node _selfchecks/harness-factory-beacon-20260927.mjs` → `PASS injected 2 errors → 1 row count 2 linked to 1 posthog issue; exception frames kept, value blank, floor clean; replay off-silent|on-masked; fallback beacon green; cleanup LEFT 0` · **FAILS IF:** two rows, a message value or a floor shape in the exception, unmasked secret text in a replay, a recording request while replay is off, or the fallback red.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 6 — Per-change flags
**Review, 2026-09-28:** built (13/13) and REFUTED by the adversary seat — a racing off could leave PostHog on, off was impossible while PostHog was down, the page helper was never loaded and never saw per-person switches, expiry disagreed between the two copies, and the selftest could not see these (two mutants stayed green); the cold seat agreed with a named change (the same two blind spots). Rework under the amended C6 above, with every regression given a red-first case, then the three seats again.
**Second rework, 2026-09-28 (built, pushed on the Hub branch factory-step6 at f1184e336, the three seats running again):** the rework under the amended C6 was refuted a second time — the adversary built a fixed ordering (a promote that read the record before an off and wrote after it) that turned the change back on with both calls answering 200, and the cold seat lost the same race one time in six on the code's own selftest; KV has no compare-and-swap, so no marker on a shared record can see a write that lands after the off returned. The design changed: an off is written to its own key (`factory:flag-off:<id>`) that set and promote never write, a switch is on only when no off stands against it, and a re-enable stamps the flag record so an older off no longer counts while a fresh off wins again. The dedicated factory key is now real on BOTH doors (flags and issues): header `x-factory-key` with the factory name, the Cloudflare setting `FACTORY_KEY` or, until it is set, a hash stored once in KV through `set_key` by Nick, Chantelle or the robot; with neither, every factory write answers 503, so nothing is open by default. The key itself was minted 2026-09-28 into the family vault as `factory-agent-key` (ours, no outside account; RULE 32). The page helper has a real sixty-second timer and fails closed; a set on an expired switch restarts its fourteen days. Proof on the branch: the flags selftest 24/24 in ten consecutive runs (fixed-order race cases, a lost write answering 409, the timer, fail-closed, expiry reset, the key, 503 with no key, set_key once, re-enable), the issues selftest 44/44 (the name without the key is 403), the endpoint gate 7/7, the API selftest folder 383 pass with the four pre-existing fixture reds. Consequences for the other steps: every harness or job that mints a factory session must send the key (the STEP 4 board harness and the STEP 5 beacon harness on branch factory-kit do not yet — fixed when those steps are built); after the merge the live issue door refuses the factory agent until the key's hash is set on the live Hub through `set_key` (one call by Chantelle's or Nick's session, or the robot), which is part of landing this step; a person's own report never needs the key. Read switches only through `factoryFlagOn`, `getFlag` and `flagsFor` — STEP 7, 10 and 13 never parse the records.
**Third round, 2026-09-28:** the adversary re-checked the second rework and said proceed with changes — seven of its nine findings closed; two schedules still ended on after an off answered 200 (two offs and a re-enable in flight with the older off landing last; and a re-enable stamped by a fast clock), because "does this off still count" was decided by comparing clocks and the off key could be overwritten by an older off's retry. Changed: a re-enable now records the marker of the one off it cancelled, any other off stands whatever order the writes land in or what the clocks say, and an off is one write with a read-back and never a second write; the key compare is constant-time; GET reads the off key only for a switch that would otherwise be on for the caller (the adversary's cost finding); set_key is reachable by the robot bearer the tasks door already trusts (the adversary found the robot path dead — a bearer carries no session), so no person handles the key. The cold seat agreed with everything but one contract line: "a lost write answers 409, never 200" cannot hold for two simultaneous sets in a store without compare-and-swap (339 of 500 forced two-way races answered 200 for a write that was then overtaken), a truth no marker can change — C6 now states that limit exactly, a set may name the revision it read, and the safe direction (off) is where the guarantee is absolute. Disagreement recorded: the cold seat called this a fail of the promise as written; the overseer resolved it by rewriting the promise to what the store can keep and proving the part that matters, rather than adding a Durable Object (a new binding and a TRUNK change, out of this step's fence). Proof: flags selftest 28/28 in eight runs (new fixed-order cases: two offs and a re-enable, the fast clock, the robot bearer on set_key, the revision on set — each red on the second rework first), issues 44/44.
**FOR NICK:** a change can be on for one person, then everyone, and off within a couple of minutes without a deploy; that is the safety net under everything the factory ships. · **Tier:** POLISH (STEP 7 is what he sees)
**Start when:** `app/functions/api/_factory-flags.selftest.mjs` exists (CREATED BY STEP 1).
**Builder:** GLM 5.3 (`zai`) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Qwen 3.8 · **Checker backup:** Sonnet
**Files you may touch:** `app/functions/api/factory-flags.js` (new), `app/functions/api/_factory-flags.js` (new), `app/js/factory-flags.js` (new), `app/index.html` (one script tag); AMENDED on the third review 2026-09-28: `app/functions/api/factory-issues.js` and its selftest (the same factory-key check on every factory-agent path, C6 as amended), and this step's own selftest gains red-first cases for each refuted claim (the review's reason for each case is in its header). **Never** `_permissions.js`, `_session.js` (TRUNK).
**Do exactly this:**
1. Write `_factory-flags.js` to C6: `factoryFlagOn(env, flagId, identity)` evaluates PostHog's flag-definitions JSON itself (the `hub_identity` person property; no vendor library in the Functions, which carry no npm dependencies); the JSON lives in a KV copy that ONE writer refreshes through `waitUntil` when a door finds it older than 60 seconds (Pages has no clock; a door never waits on PostHog and keeps the last copy when it does not answer); `setFlag`, `promote`, `turnOff` call PostHog's flags API with the vault's personal key, read the flag back, and write the bookkeeping row (`expected_revision`, `expires_at`).
2. Write `factory-flags.js`: `GET` returns the signed-in person's flags `{flags: {id: true}, at}` from the same KV copy; `POST` actions `promote`, `off`, `set` accept only the `factory` agent key or a §5.1 approve carrying `asked_by` and the revision.
3. Write `app/js/factory-flags.js`: `window.FactoryFlag.on(id)` over `posthog-js`'s flag evaluation for the signed-in person when PostHog's script is loaded, and over the door's `GET` when it is not (under automation, or a file: page), false until loaded, false for unknown ids, never throws; re-read on every `hashchange` and at most every 60 seconds.
4. The selftest runs against a PostHog stub in its fixtures (the API's request and response shapes recorded once from the real project), so the check needs no network and cannot pass by luck.
**DEFINITION OF DONE:** the selftest passes 13/13 (on for one identity is false for another; `all` is true for everyone; `off` beats `all`; `promote` without the key → 403; a stale revision → 409; an unknown flag is false; `expires_at` = created + 14 d; a flag past `expires_at` is reported by `GET` as `expired` and never as on; the client helper returns false before load; the client falls back to the door's GET without PostHog's script; the client re-fetches on a route change and after 60 s) and `FactoryFlag.on("fx-demo")` reads true as Dean and false as Mae on the local server after `set on_for:["dean"]`, and false for Dean within about two minutes of `off` (one definitions refresh plus one client re-read).
**PROOF:** `cd projects/business/business-app && node app/functions/api/_factory-flags.selftest.mjs` → `PASS 13/13` · **FAILS IF:** any case fails, or a flag past `expires_at` reads as on.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** the moment this step closes, post one dated line into `app/_design/workspace-setup/SPEC.md` §3.3: `Coherence Factory STEP 6 closed <date> — per-change flags exist (`factory:flag:*` records, 14-day self-off); they are not the per-workspace feature toggles and never replace them.`
### STEP 7 — The Look page and the review door: Try it · Looks good · Send back
**AMENDED 2026-09-29 before the build (the overseer, recorded here so the builder and checker read one contract):** the door's actions are named with underscores — `ready`, `live`, `looks_good`, `done`, `go_ahead`, `choose_recommendation`, `send_back`, `choose` — matching the brain's six tools already on main (`projects/ops/factory/brain/hub-factory.mjs`); the card body travels in the issue row's `details` (C1, amended 2026-09-29), built by `factory-card.mjs`'s `buildDetails` in the runner and passed to `ready`, so the card's Details section and the Look page draw the same sections; `ready` also takes `reviewer`, `before[]` and `after[]` (picture ids) and sets the flag on for the reviewer; `send_back` sets `status: building`. The kit's review harness was rewritten the same day to drive the Hub's own browser rig and the real doors (kit branch b51e90c05); the Look page is `app/js/factory-look.js`, the door `app/functions/api/factory-review.js`.
**Progress, 2026-09-29 (part 1 landed, the door in build):** the card writer `projects/ops/factory/factory-card.mjs` is on main (72700beddd) — written by a cheap model on the whole-file mode, then four independent checks drove its fixes, each red-first in `_test-factory-card.mjs`: a person's own words mentioning money no longer stop the card; the live shape is one the store accepts; a plan offers two or three options; and the money rule moved to the source. **REQUIREMENT FOR STEPS 10 AND 12 (the code that generates plan options), set by the fourth check:** every plan option object carries `involves_money: false`, written as a literal in code where the option is built — never copied from, inferred from or filled in by a model's drafted text — and an option that would need real spend is never offered at all (C5; money leaving is a four-acts item). Three checks proved a word list cannot tell "pay dates" from "pay $50", so the card writer's text check is only a narrow backstop (currency signs and amounts). The review door (`factory-review.js`) is on the Hub branch factory-step7 with its door-level proof (every refusal and status code right on the model's first whole-file attempt; three faults being fixed — one of them the overseer's own brief, whose hand-off sentences were shorter than the tasks door allows); the store there accepts `attempts` and `chosen` (selftest 47/47).
**FOR NICK:** Dean sees what changed for him in one sentence with pictures, reads three lines on how to try it and what it touches, tries it himself, and his tap is what makes it live for everyone; his Send back with a sentence is heard. · **Tier:** FRONT
**Start when:** STEP 4's renderer runs; STEP 6's flags door answers; STEP 2's `send_back`, `choose_option` and `factory_ready` exist (or the core ask is recorded); `_selfchecks/harness-factory-review-20260927.mjs` and `_test-factory-card.mjs` exist (CREATED BY STEP 1).
**Builder:** GLM 5.3 (`zai`) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Sonnet (browser, as Dean, phone and computer) · **Checker backup:** Qwen 3.8
**Files you may touch:** `factory-card.mjs` (new, C5 — the one function that writes every generated body) under `projects/ops/factory/`; `app/js/factory-look.js` (new: the route `#factory/<issue id>`, fetching the row and its run summary and rendering the C5 shape with the four pictures, in the Hub's tokens, 375 and 1280); `app/index.html` (one script tag); `app/functions/api/factory-review.js` (new: actions `ready` (agent key: requires `before[]`, `after[]` (four `/api/files/<id>` ids), `what_changed` (≤ 25 words), `how_to_try[]` (three lines), the evidence keys C5 needs; sets the flag `on_for:[reviewer]`, hands the card to the reviewer through the hand-off door with `kind:"finished"` — the only kind that reaches a person, `tasks.js:1610` — and calls the brain's `factory_ready`), `live` (agent key: the bug shape; hands the card to the reporter `kind:"finished"`), `looks-good` / `done` / `go-ahead` / `choose-recommendation` (a person's approve carrying `asked_by`, `revision`, `merged_sha`; `looks-good` promotes the flag and completes the card with the person's dated words through the task door's own `complete`; `done` completes a bug card the same way; a second call → 410), `send-back` (a person; requires `asked_by.words`; sets `status: building`, `attempts+1`, appends `sent_back[]` with Jev's `factory.sendback.reason` hint beside the words (C10), and `factory` takes the card), `choose {n}`). **Never** the shared approval card's renderer (Screen Buddy core), `tasks.js`, the harness and the test (STEP 1's).
**Do exactly this:**
1. Write `factory-card.mjs` (C5): `body({issue, run, shape})` → the sectioned text; `assertClean(text)` throws naming the offending token; evidence-gated phrases appear only with their evidence. `_test-factory-card.mjs` must now pass.
2. Write `factory-look.js` and `factory-review.js` as above.
3. Run the harness: on the local server it seeds an `agent-test` change with a flag; calls `ready` as `factory`; signs in as Dean at 375 and at 1280; opens the drawer card and the Look page; asserts the eight sections, four pictures, the three how-to-try lines, the who-gets-it line; Try it opens the screen with `FactoryFlag.on(id)` true for Dean and false for Mae; Approve → flag `all`, card Done with Dean's words, second Approve 410; reseeds; Discard → the drawer asks; no words → nothing changes; "wrong order" → card back to In Progress, words in Conversation, flag still Dean-only; cleans up.
**DEFINITION OF DONE:** as Dean on the local server, at phone and computer width, a Ready to look at card opens a Look page with only the C5 sections, Try it is on for him alone, Looks good makes the flag `all` and closes the card with his words and the second tap is refused, and Send back needs words and reopens the card with them.
**PROOF:** `cd projects/business/business-app && node _selfchecks/harness-factory-review-20260927.mjs` → `PASS look page 8 sections at 375 and 1280; try-it dean:true mae:false; looks-good → all; second tap 410; send-back words → reopened; body clean; cleanup LEFT 0`; and `cd "$WORKSPACE" && node projects/ops/factory/_test-factory-card.mjs` (CREATED BY STEP 1) → `PASS 8 refusals; 6 evidence gates; clean body passes` · **FAILS IF:** any section is missing at either width, the body fails `assertClean`, Try it is on for Mae, or a second Looks good succeeds.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 8 — The classifier and the tier runner
**FOR NICK:** nothing he notices; this is what makes "one review for a small thing, three for a shared thing" a rule a machine keeps rather than a habit. · **Tier:** POLISH
**Start when:** `_test-factory-classify.mjs`, `_test-factory-review.mjs` and the twelve fixture diffs exist under `projects/ops/factory/` (CREATED BY STEP 1).
**Builder:** Sonnet (the classifier and the tier runner are a load-bearing safety wall, and CORE §3 keeps such walls on Anthropic; the brief carries a `CHEAP-FIRST-OVERRIDE: FLOOR` line naming the classifier file) · **Builder backup:** Opus · **Checker:** Qwen 3.8 · **Checker backup:** DeepSeek V4 Pro
**Files you may touch:** under `projects/ops/factory/`: `factory-classify.mjs` (new, C3) and `factory-review.mjs` (new: given a tier and a run record, runs the review depth — easy: one cheap checker re-runs the harness; medium: that plus one Sonnet diff review against the spec through the Anthropic-tier path of `projects/ops/cheap-task.mjs`; hard: requires the three-seat record for this commit (STEP 9) then the same as medium with the review's verifier). **Never** `cheap-task.mjs`, `route-build.mjs`, `MODEL-MATRIX.md`, the tests and fixtures (STEP 1's).
**Do exactly this:**
1. Write `factory-classify.mjs` to C3 exactly: the two hard-coded arrays, the import-count rule computed from the repo, the diff-only input; `--diff <file>` prints `{tier, reasons, files}`; a path list is refused; `--raise` accepted; `--lower` exits 2 with `refused: a tier is never lowered`.
2. Write `factory-review.mjs`: `review({tier, run, worktree, sha})` → writes `review_kind`, the checker model (never the builder's vendor; `glm≡zai`) and the verdicts into the run record (C7); a `hard` tier with no `run.triad` or with `triad.reviewed_sha !== sha` → exit 3 `refused: hard tier needs a review record for this commit`.
**DEFINITION OF DONE:** every fixture diff lands in its tier with its reason, a path list is refused, lowering is refused, a one-file change to `app/js/util.js` is `hard`, a hard review without a record for this commit is refused, a medium review invokes exactly two checks and an easy review exactly one.
**PROOF:** `cd "$WORKSPACE" && node projects/ops/factory/_test-factory-classify.mjs && node projects/ops/factory/_test-factory-review.mjs` (CREATED BY STEP 1) → `PASS 12 rules, path-list refused, --lower refused, util.js → hard` and `PASS easy=1 check, medium=2, hard-without-record refused` · **FAILS IF:** any fixture mis-tiers, or `--lower` exits 0, or a hard review runs without a record for the commit.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 9 — The three-seat review runner and the merge gate
**FOR NICK:** a change to a shared part of the Hub is argued over by three different minds before anyone builds it, nothing can be merged without that record for that exact change, and a change that fails three times reaches the Ecosystem Fixes board as options, never as a mess. · **Tier:** POLISH
**Start when:** STEP 8's `factory-review.mjs` exists; `projects/ops/skippy-jobs/lib/astra-run.sh` runs (`zsh projects/ops/skippy-jobs/lib/astra-run.sh 2>&1 | head -1` prints usage); `_test-factory-triad.mjs` exists (CREATED BY STEP 1).
**Builder:** Sonnet for the review runner and its merge gate (a safety wall, CORE §3; the brief carries a `CHEAP-FIRST-OVERRIDE: FLOOR` line naming that file), GLM 5.3 (`zai`) for the headless drive and the briefs · **Builder backup:** Opus / DeepSeek V4 Pro · **Checker:** Qwen 3.8 · **Checker backup:** DeepSeek V4 Pro
**Files you may touch:** under `projects/ops/factory/`: `factory-triad.mjs` (new), `factory-drive.mjs` (new: the headless drive STEP 10 uses) and the folder `briefs/` with `triad-propose.md`, `triad-attack.md`, `triad-cold.md` (new; each carries CORE and BUILD verbatim per the dispatch header). **Never** `astra-run.sh`, `astra-review.sh`, the review skill, the Hub's workflow files, the test (STEP 1's).
**Do exactly this:**
0. Verify the seats are callable from a scheduled job on the Studio and record it: Astra through `astra-run.sh` (the Codex CLI, proven 2026-09-27 for this plan's own attack seat); Fable and Opus through the Claude Code CLI in print mode (`claude -p --model <slug>`), launched detached through `run-detached.sh` exactly as Astra is — a job cannot dispatch a Claude Code agent, so the CLI is the route; if the CLI is absent on the Studio or refuses the model, the step records `NOT MEASURABLE — no Claude CLI on the Studio` and names the install as its one missing thing, and the HARD tier stays closed in STEP 10 (a hard row waits at `spec` with "waiting for the reviewers") until the seats are proven; easy and medium run.
1. `factory-triad.mjs run --issue <id>`: (a) PROPOSE on Fable through the CLI (`ROLE: RECONCILER` for the proposal writer), producing the plain-words spec (C5 spec shape) and the technical proposal with the revert command named; (b) ATTACK on Astra through `astra-run.sh start … --sandbox read-only --effort high` with `triad-attack.md`, polled with `wait --seconds 480` until DONE, requiring a verdict `proceed | proceed with changes | do not proceed` with `file:line` behind every claim; `proceed with changes` → the changes are applied to the proposal and the attack seat runs once more on the corrected proposal, which must then read `proceed`; (c) COLD on Opus through the CLI with only the corrected proposal (never the attack); (d) write `run.triad = {proposal, attack, cold, verdicts[], reviewed_sha: null}` (C7) and post the spec onto the card as `stage:spec` with Go ahead / Not this through `factory-review.js` for the reviewer; `do not proceed` from either seat → loop counter +1 and back to (a) with the objection; the third failed loop → `factory-recur.mjs plan --issue <id> --board ecosystem-fixes` (STEP 12) and stop.
2. After the build, `factory-triad.mjs bind --issue <id> --sha <head>`: the cold seat re-reads the built diff against the proposal once and sets `reviewed_sha` when it matches the proposal's scope; a diff outside the scope → `do not proceed` and a loop.
3. `factory-triad.mjs gate --issue <id> --sha <head>`: C9 exactly; exit 0 only when the record matches the tier of the head diff, the harness went red then green, the diff-read verdict for the tier is present, and for `hard` `reviewed_sha` equals the head; the only caller is the fix runner's merge stage (`main` is unprotected; C9 says so and the monitor flags any commit that bypassed the runner).
4. Write `factory-drive.mjs`: signs in as a named identity through `hub-session.mjs` on the local server, opens a route at 375×812 and 1280×800, performs a scripted click path, saves before/after screenshots through `/api/files`, and diffs the screenshot outside the changed element; it is the drive STEP 10 uses and the kit's browser harnesses may reuse it.
5. Tests with a fake CLI and a fake `astra-run.sh` (env `FACTORY_FAKE_REVIEWERS=1`): no record → `gate` exits 3; a record for another sha → exits 3 naming both; a record with both proceeds and the head sha → 0; `proceed with changes` → applied and re-attacked once; three `do not proceed` loops → one plan call with `--board ecosystem-fixes` and stop; `factory-drive.mjs` on the local server → four picture ids and an `outside_diff` number.
**DEFINITION OF DONE:** a hard change cannot pass the gate without a record bound to its head commit, a stale record is refused by name, `proceed with changes` is re-attacked exactly once, three failed loops produce exactly one Ecosystem Fixes call, the three seats are proven callable from a job (or recorded NOT MEASURABLE with the one missing thing), and the headless drive returns four pictures.
**PROOF:** `cd "$WORKSPACE" && node projects/ops/factory/_test-factory-triad.mjs` (CREATED BY STEP 1) → `PASS: no-record merge refused; stale-sha refused; record → allowed; changes → re-attacked once; 3rd loop → ecosystem-fixes card; drive → 4 pictures` · **FAILS IF:** the gate passes without a record, a stale record passes, the third loop retries instead of planning, or the drive returns fewer than four pictures.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 10 — Triage and the fix runner as a state machine, registered and beating
**FOR NICK:** a reported bug moves itself from words to a fix to a check to a merged, published, live screen while the team sleeps, a few at a time, one bounded step per tick, and a stranger can read where each one is from its card. · **Tier:** FRONT
**Progress, 2026-09-29 (the machine and its record store built):** `jobs/factory-fix.mjs` exports `tick(state, tools)` and passes the kit test; `lib/factory-state.mjs` is C7's store (one JSON per issue, a lease per issue, revision +1 per write, a stale writer 409, C7's fields only). Three independent checks shaped it: waiting (for the merge slot, for the publish, a tool that throws) is never a failed attempt; a green-before-fix harness is re-briefed once, then plan; a merged card is never sent back by the attempts, age or cost bounds, and instead goes to a person with its revert prepared if it is not live within six hours of its merge; waiting cards never use one of a tick's three slots, so stuck cards cannot starve new ones. Step 3 landed (e2804ba437): the eight jobs are rows in the runner with distinct minutes (Manila times converted to the Studio's UTC-5); seven beat "not built yet" until their step lands and `factory-fix` beats "tools not wired yet" and acts on nothing. Design slices S1–S5 landed 2026-09-29, checked by an independent checker that did not build them (SHIP after two rounds; every proof re-run twice and each shown to fail on a mutant): S1 C7's `runner` field; S2 the tick's pending/started and relabel answers, R3 no second merge of a merged card, R10 a hard bug stops at ready; S3 the Hub door client `lib/factory-door.mjs` (list, get, update, ready, live, comment on a task card only, upload; never sets done, live or ready itself; the key and session scrubbed from every error — the data wall keeps this file on Anthropic); S4 `lib/factory-jobs.mjs` (detached steps polled across ticks; overdue killed; names fenced); S5 the two briefs `fix-harness.md` and `fix-build.md` (the harness proves both switch states; a fix may not touch its proof or anything R4 names). The first check found two gaps — comment and upload missing from the door client, and the harness brief silent on the switch — both closed red-first. Then, 2026-09-29, each checked by an independent checker that did not build it (SHIP; every proof re-run and shown to fail on a mutant): S12 `lib/factory-run.mjs` `runOnce(ctx)` — the jobs machine only, one round at a time (a lock broken after eleven minutes), the run record saved before the row on every round, a 409 re-read (retried when only a new occurrence moved the revision, dropped when a person changed the row), a send-back starting a clean round, spend metered, live and ready never written through update, the merge slot kept between rounds, one card's error never stopping the others; S13 `lib/factory-triage-core.mjs` `triageRow` — an easy or medium bug builds, a wish or anything hard goes to spec, a PostHog error that cannot be built goes to plan with a card for its owner, a tier is never lowered, posts to a person only 08:00–20:00 Manila; and R11 (the fix's own harness in `_selfchecks/factory/`, the classifier's exemption moved and the old place made hard). The checks also found and closed: C7's store crashing on a new machine's first record (it now makes its folder before its lease), a stranded hard PostHog error, and stale per-fix paths in the plan and kit. S10 landed 2026-09-29, checked by an independent checker that did not build it (SHIP after two rounds, confirmed with the checker's own probes): `lib/factory-tool-merge.mjs` — gate reads every changed name with renames split (R4), refuses a changed harness, evidence for another commit and a mis-tiered diff, and turns every failure into a refusal (R5); merge waits on a running publish, refuses a HEAD that moved since the gate, saves merge_started_at before any pull request, opens once, squash-merges the gated commit and finds an earlier merge after a restart (R3). The merge gate learned which harness is the fix's own (`own`), and R12 was added after a second checker defeated a word-list content check. Then S7 `lib/factory-tool-build.mjs` and S11 `lib/factory-tool-publish.mjs`, each checked by an independent checker that did not build it (SHIP after two rounds, with the checkers' own probes). The build starts the cheap lane detached, confined to the fix's worktree, with the filled brief (the switch rule for a wish or a hard bug), a proof that runs the fix's own harness and requires it unchanged — every path quoted for bash, proven against a real folder named with a space — and the run id for spend; on finish it commits, fails the attempt on any R4 path read with renames split, and re-tiers with the fix's own harness; every vendor down is a wait. The publish is live only when the merged commit is served, no later commit reverts it, the revert is prepared as a draft pull request, the served build's tier-1 gates ran and the live check passes — anything else is not live yet, never a throw, so the six-hour limit hands it to a person; the review hand-off carries the card's risk and checked facts. The checks caught, and the rebuilds closed: an unquoted path that would have failed every real build, R4 left to the gate, a revert that still read as live, a failed live check that bypassed the six-hour limit, and a review card that always said nothing to flag; a first rebuild that only listed the proof's own examples was discarded. S0's key probe passed from this machine (2026-09-29: the live Hub accepts the factory key) and found two faults in the door client, fixed red-first: it read the session helper's cookie under the wrong name, and an older vault reader's labeled answer would have been sent whole as the key. CORRECTION to S7's wording: the Hub squash-merges, so the commit to revert exists only after the merge; the publish tool records revert_command and revert_pr. BUILD §1a's five beneficial tests (the checker corrected the overseer, who had called them undefined) are measured from the run before a fix is live: the goal and the before-and-after are the fix's own test failing before and passing after; nothing a person depends on got worse is the served build's tier-1 gates, the live check and the drive's outside diff showing nothing else on the screen changed; no work created is a reader's judgment recorded by the check stage; anything unmeasured is unproven and not live, so the six-hour limit hands it to a person with the revert prepared. This makes the drive's outside diff (S9) and the check stage's creates-work judgment (S8) required, not deferred. KNOWN LIMIT, named by the publish tool's checker: the revert check finds a commit that names the merged sha (git revert and GitHub's one-click revert both do); a hand-made rollback that never names it is not seen here, and is caught instead by STEP 13's monitor, which re-runs every fix's own harness daily and treats a red one as a recurrence. REQUIREMENT FOR S6, named by the build tool's checker: a refused or failed attempt leaves its commit in its worktree, so every new attempt gets a fresh worktree of its own (`<id>-a<n>` with n raised, from origin/main), never the last attempt's. S6 landed 2026-09-29, `lib/factory-tool-harness.mjs`, checked by an independent checker that did not build it (SHIP after two rounds; every proof re-run twice and five mutants each shown to turn their check red): a fresh worktree per attempt from origin/main, the seed file outside it, the test author run detached with the filled brief, the author allowed to write its one harness file and nothing else (a stray file is refused before anything is committed and the attempt starts fresh), the harness committed and run detached, and red only when it printed a failing case and did not crash. The same check found that no build ever recorded its harness going green, so no fix could ever have merged or gone live; the rebuilt build tool now records it on the committed head, resets a failed or refused attempt to the harness commit, and marks a tier rise for a fresh worktree (a plain failure reuses the worktree reset to the harness commit, which the checker judged a sound reading of the fresh-worktree requirement). The drive planner found that the round saved each record with the revision it read before the tools wrote it, so every real round's save was refused and the row never moved; the round now re-reads before saving and keeps what the tools wrote (a red-first case in the round proof). S0 on the Studio, measured by the Hub's probe workflow 2026-09-29 (two runs): the Claude command line is installed (2.1.283) and its sign-in is present, but a Sonnet call exits 1 with nothing on its error stream; gh is signed in; the workspace is current; the eight factory rows and job files are present and three factory heartbeat rows are already beating, so no restart was needed. S8 landed 2026-09-29, `lib/factory-tool-check.mjs`, checked by an independent checker that did not build it (SHIP; three mutants each turned the proof red): one reader for an easy fix and two for medium and hard, each started detached; MERGE only when every reader says so; each verdict tied to the commit it read and a moved commit refused; the creates-work judgment recorded as given, never assumed; any failed or unreadable reader is a wait. S5′ landed the same day: the harness brief's DRIVE is now a JSON literal (route, as, one region selector never html or body, at most twenty click, type or wait steps, fixtures through a local door carrying agent-test), read by the drive as text and never run. S9 is built as three files under `projects/ops/factory/`: `factory-pixels.mjs` (the PNG reader and the outside-the-band row diff), `factory-drive-read.mjs` (reading DRIVE and grading outside_diff) and `factory-drive.mjs` (the command line that runs the local Hub and takes the pictures); each is written whole, so no cheap build ever inserts into a file. Landed 2026-09-29, every piece checked by independent checkers that did not build it (each proof re-run, and mutants shown to turn it red): S9, the drive — `factory-pixels.mjs` (PNG reading, proven byte for byte against two independent decoders on 60 real screenshots; the outside-the-band row diff, fuzzed 500 ways without ever under-counting), `factory-drive-read.mjs` (DRIVE read as text and never run; outside_diff graded none, changed or unmeasurable; a deep-nesting crash the checker found, fixed), `factory-drive-hub.mjs` (the Hub's own local server on a dedicated port, a local session, seeded fixtures, a factory flag, snapshot and restore; the port taken back only from a Hub left by a hard-killed drive in the factory's own folder; no key, token or secret reaches the local Hub, proven by a planted secret absent from its running processes; every web call time-limited), `factory-drive.mjs` (the command: seven real cases on a real local Hub and a real headless browser — a change inside the region is none with four pictures at 375 and 1280, a change outside it or visible with the switch off is changed, a clean switched change is none, page noise is unmeasurable, a stranger and a busy port are refused), `lib/factory-tool-drive.mjs` (the drive tool; its job never inherits a secret), the runner's two answers (unproven goes to a person with no attempt, a change outside is an attempt) and the merge gate's requirement (the drive ran on exactly the merged commit and saw nothing else change). What the real runs taught, each fixed red-first: a locked Mac never finishes a Chrome picture, so the drive uses the headless shell (on this Mac and on the Studio, whose screen is locked); a card that grows by a fraction of a pixel redraws the text below it, so the region's height is rounded up to a whole pixel in every picture; a fix's first page load is slower, so a page counts as settled only once its loading placeholders are gone; the page is measured before and after its picture and shot again if it moved. S8b, `factory-read.mjs`, the check stage's reader program: exact verdicts only; the R4 fence, a red harness, an oversized diff and floor-shaped content answered in code before any model; the low-cost reader is a vendor the builder never used; the Sonnet reader signs in with an allowed account by environment (the Studio's `claude -p` said "Not logged in" under its runner because the keychain is unreadable there); it and the author choose the first Claude command line whose help lists `--restricted` (the Studio's first copy on the path is 2.1.237, which refuses the flags; `~/.local/bin/claude` is 2.1.284). S15 `lib/factory-first-live.mjs`: until STEP 11 has passed and Nick or Chantelle gives a dated word (a commit setting OPENED), the runner and triage see and reach test rows only; and ARMED ships false, so on the Studio the jobs wire, sign in, report their gaps and act on nothing until a commit arms them for STEP 11. S16: worktrees under the Hub's own `.claude/worktrees` (the Hub's shared links resolve only there); the model-written harness runs with every key removed; the stray-file check sees ignored places. S17: the build names its attempt with the cheap lane's declared `--job` and records the spend key the lane meters under, so the five-dollar bound can fire. S19 `factory-author.mjs`: Sonnet writes the one harness file, in restricted mode with file tools only, writes allowed to that path, no settings, MCP or session, an allowed account's sign-in by environment, twenty minutes. S20 `lib/factory-wire.mjs` and the runner's entry point: one overlay per stage (the two briefs never mixed), the limited door, the bundle from the deploy checkout, spend summed from recorded keys, the GitHub key read from the vault by name and handed to gh and git only (the Studio's gh cannot use its keychain under the scheduler); every missing piece is a named gap; a dry run acts on nothing; moved rows are read back. S21 `lib/factory-guess.mjs` and `jobs/factory-triage.mjs`: one low-cost guess per report with only real Hub files kept, the owners map read from the deployed checkout (the current main), run records made before a row moves. The grunt lane caps the two new labels (factory-read 150, factory-triage 300 a day). Measured on the Studio 2026-09-29 (read-only): it holds the jobs role; the generated bundle is in the deploy checkout; the headless shell is installed; the Hub's worktree folder is ignored; all five allowed accounts have a sign-in. KNOWN LIMIT, named by a checker: the Claude command-line probe is bounded at 20 seconds, but a candidate whose `--help` itself starts a child could leave that child running; candidates are only the operator's own install paths. On the Studio 2026-09-29 after its job runner restarted onto this code: the factory key was copied into the Studio's own vault (it held only this Mac's copy; same key, not a rotation, value never shown), and the runner-written rows then read `factory-fix OK — wired; not armed, so acting on nothing` (no stage waiting: bundle, headless shell and Hub checkout all found) and `factory-triage OK — not armed: 0 inbox rows seen, none routed` (signed in to the live Hub through the limited door). R9 landed 2026-09-29, `lib/factory-record.mjs`: the record the reporter had open is read through the Hub's own single-record doors (a task card, a project, a contact, a survey, a page) with the factory's session and masked before it is written into the harness author's seed — emails, phone numbers, long numbers, money by symbol or currency code, values after secret-looking names, cloud key ids, bare hex runs, random-looking blobs and uppercase codes, and every value under a sensitive field name; anything floor-shaped left after masking, or a floor check that fails, withholds the whole record; every other kind gets its id only, with a note. An independent checker attacked the masking over four rounds (a one-line key pair, currency codes, over-masking of paths and references, hex and uppercase secrets) until SHIP; the named cost is that a bare commit sha or uppercase order number in a report is hidden too, since neither can be told from a secret. OPEN: arming the factory and STEP 11's seeded run (needs STEPS 2–9 landed and beating, per STEP 11's start rule); the runner-written heartbeat rows from the new code appear after the Studio's job runner picks it up.
**Start when:** STEP 3's door, STEP 6's flags, STEP 7's `factory-review.js`, STEP 8's classifier and reviewer, STEP 9's gate exist on the local server / in the workspace; `_test-factory-fix.mjs` exists (CREATED BY STEP 1).
**Builder:** GLM 5.3 (`zai`) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Qwen 3.8 · **Checker backup:** Sonnet
**Files you may touch:** under `projects/ops/skippy-jobs/jobs/`: `factory-triage.mjs` (new), `factory-fix.mjs` (new); under `projects/ops/skippy-jobs/lib/`: `factory-state.mjs` (new, C7's one module, lease and revision); the schedule rows for the eight factory jobs in `projects/ops/skippy-jobs/runner.mjs` (rows only; the six later jobs are registered here so they beat from the start and STEPS 12–15 fill them in); `projects/ops/skippy-jobs/lib/long-runners.json` (the build stage); under `projects/ops/factory/briefs/`: `fix-build.md`, `fix-harness.md` (new); added by the reviewed tools design (2026-09-29): `lib/factory-jobs.mjs`, `lib/factory-door.mjs`, `lib/factory-tools.mjs` (built as `lib/factory-run.mjs`, `lib/factory-triage-core.mjs` and one `lib/factory-tool-<stage>.mjs` per stage — harness, build, check, drive, merge, publish) under `projects/ops/skippy-jobs/`; under `projects/ops/factory/`: `factory-drive.mjs` (the drive's command line) with its parts `factory-pixels.mjs`, `factory-drive-read.mjs` and `factory-drive-hub.mjs`, and `factory-read.mjs` (the check stage's reader program); one entry, `factory-read`, in the grunt lane's daily cap table (`grunt-lane.mjs` under `projects/personal/skippy-app/lib/`, the entry only), the kit file `_test-factory-tools.mjs` under `projects/ops/skippy-jobs/` with its proofs in `factory-proofs/` beside it (one red-first proof per module, each shown by an independent checker to fail on a mutant; `--only <section>`; a section that cannot run is NOT MEASURABLE and fails), and the `runner` field of C7. **Never** the runner's code beyond its rows, other jobs, `cheap-task.mjs`, the test (STEP 1's).
**AMENDED 2026-09-29 — the tools design, reviewed three ways before the build (a planner proposed it; an adversary refuted it with eight findings, all taken; a cold third reviewer agreed with one named change, R10). Binding on every STEP 10 slice; where it disagrees with the text above, it wins.**
<details><summary>The design: slices S0–S14 and binding rules R1–R10</summary>
W = /Users/chantellelamoreaux/Documents/Claude 2.0 ; H = W/projects/business/business-app (Hub; cite origin/main).
#### Plan assumptions measured
- Build-integrity commit field: already exists as `source_revision` (H/app/build-dist.js:1232,1239; H/app/functions/api/build-integrity.js:59). No slice.
- `factory-drive.mjs` missing; drive loop in W/projects/ops/factory/factory-triad.mjs:59 needs a `tools.shot`. Separate slice S9; blocks live only.
- `harness/tools/drive.mjs` is a CLI; reuse `app/_design/pm/shot.mjs` (newPage), `harness/pm/local-keys.mjs` (LOCAL_SECRET), `harness/pm/serve.mjs`.
- Worktrees "under the runner's _work root" (/Users/nickdeck/actions-runner/_work) are outside the workspace; cheap-task.mjs:1489-1492 refuses them. Use W/projects/ops/skippy-jobs/state/factory/work/<id>-a<n> (gitignored .gitignore:172; writable per vendor-fence.mjs:272).
- Sonnet authoring the harness from a job via the Claude CLI on the Studio: never measured (PLAN.md:464). Test-authoring never goes cheap. Blocker for every fix until measured (S0).
- C7 has no field for the runner's working state (factory-state.mjs:8). Amendment S1.
- verify-live.mjs compares live against a local app/dist (header 9-13): needs a local build of that commit (S11).
- Jev hints: `_jev-uses.js` has no factory.* uses; triage runs hints off (C10 fallback).
- Factory key on live Hub: unverified; doors 503 until set (factory-issues.js:24-26). Probe in S0; robot bearer can set it.
- Row status already accepts harness/build/check/drive/merge/publish (_factory-issues.js:37); `run` and `attempts` patchable (:38); stale revision 409 (:181).
#### Keys on the Studio (no secret printed)
Hub session: `mintAgentSession("factory","nick")` (W/projects/ops/skippy-jobs/lib/hub-session.mjs:51) reads the Hub token from the vault in memory. Factory key: `vault.py get factory-agent-key --caller=skippy --value-only`, held in memory, sent as `x-factory-key`. Child processes (gh, claude CLI): wrap with `run-with-key.mjs --caller=skippy --key=github-pat-classic --env=GH_TOKEN -- gh …`. The job itself runs in-process in the runner.
#### Slices (tests written first by Sonnet under TEST-AUTHORING in a new kit file _test-factory-tools.mjs --only <fn>; builder GLM unless noted; checker Qwen)
- S0 Studio measurements: `claude -p --model sonnet "reply ok"` detached; `gh auth status`; key probe (update with expected_revision -1 → 409 accepted / 503 unset / 403 wrong); the review harness inside a worktree under the new root (space in "Claude 2.0").
- S1 C7 amendment: one field `runner {started_at, merged_at, slot_at, harness_green_once, drove, note, job, worktree, head_sha, gated_sha, n}`; add 'runner' to factory-state FIELDS.
- S2 tick learns two answers: in harness/build/check/drive, `{pending:true, started?:bool}` → no attempt, slot freed unless started; `{relabel:"wish"}` → spec. New kit cases (pending harness today reads as green-before-fix; 4 pending + 3 fresh → the fresh advance; pending at 49h → plan).
- S3 door client lib/factory-door.mjs: loadCredentials(); createDoor({origin, fetchImpl, creds}) → {list, get, update(id,rev,patch), ready, live, comment, upload}; allow-list issues/review/upload/comments; refuses complete, looks_good, done, go_ahead, choose_recommendation. Proof: fetchImpl routed into the Hub's real doors over mock KV; key header sent; stale 409; refused action never fetches; key never in an error or log. Builder Sonnet (FLOOR: handles a key).
- S4 lib/factory-jobs.mjs: startJob({name,cwd,argv}) over run-detached.sh; pollJob(job,{now,maxMs}) → running|done|gone|overdue (overdue kills).
- S5 briefs fix-harness.md and fix-build.md under projects/ops/factory/briefs/ (harness exports DRIVE = {route, as, clicks[]}; build brief forbids editing the harness).
- S6-S11 lib/factory-tools.mjs makeTools(ctx), ctx = {door, store, jobs, sh, fetchImpl, now, paths:{hubMain, workRoot}}:
- S6 harness: first call worktree from origin/main + detached Sonnet CLI → {pending, started}; when done run the harness inline, non-zero exit = red → attempts[n].harness_red_at, {red:true}; missing file or relabel → {relabel:"wish"}.
- S7 build: detached cheap-task --do fix-build --dir <worktree folder> --prove "harness && harness sha unchanged"; on finish classify(git diff origin/main...HEAD), revert_command, {ok, tier}. A harness edit is a failure.
- S8 check: factory-review.mjs review(tier, ctx) (:11); cheap seat inline (harness re-run + one gruntFetch on a different vendor); Sonnet seat detached only after S0; until then triage sends medium to spec.
- S9 factory-drive.mjs (PROCESS) + drive: shot({issue, as, when, viewport}) with two local servers (main and the fix worktree), a local session signed with LOCAL_SECRET, uploads through the door to the live Hub; outside_diff REQUIRED (it is BUILD §1a's nothing-else-changed test; a fix without it is unproven and never live — corrected 2026-09-29, it was deferred).
- S10 gate + merge: gate calls factory-triad.mjs gate() (:19), saves gated_sha; merge refuses unless HEAD == gated_sha, throws (a wait) while a deploy.yml run is in progress or queued, then gh pr create + gh pr merge --squash --match-head-commit <gated_sha>; prepares the revert PR, never merges it. Builder Sonnet (FLOOR).
- S11 publish/handoff/vendorsDown: live when merged_sha is an ancestor of the live source_revision; verify-live.mjs run in the Actions runner's deploy checkout whose dist is the served build (path measured in S0); beneficial tests from the record + gates.tier1. handoff "review" → review door ready with details, look_body, 2+2 pictures; "finished" → live, reported bugs only. vendorsDown = an attempt ended "vendors down" in the last 15 minutes.
- S12 run(): loadState → tick → persist; checks this machine holds the jobs role (role-owner.mjs:74,186) else beats a warning; writes C7 first then update {status, attempts} with expected_revision; 409 → the person's change wins, the transition dropped; tools read results from job output files so repeating a stage is harmless; stops starting work after 6 minutes.
- S13 factory-triage: triageRow(row,{hints, guess, classify, ownerFor}) → {patch, shape}, pure; header-only diff from the named files, raise medium when unsure; ownerFor imported from the runner's own main worktree; human-facing posts only 08:00-20:00 Manila.
- S14 live proof: restart on the Studio, ≥8 runner heartbeats, STEP 11's seeded agent-test bug through prove-e2e.mjs.
#### Risks (planner's ranking)
1 the Claude CLI on the Studio unproven; 2 worktree location, the cheap fence and Hub harnesses with the space in the path; 3 status vs C7 races when a person sends back mid-tick; 4 verify-live against a dist that is not the served build; 5 auto-merging a revert (recommend never; card to plan with the revert prepared); 6 a tick overrunning 9 minutes.
#### Two choices kept
(a) teaching tick the pending/started answers (S2) rather than starting and collecting jobs outside it — kept inside so the 48h and $5 bounds work on hung jobs;
(b) the issue row's status as the single source of truth for the stage, with one `runner` field added to C7.
#### Binding rules (they override anything above that disagrees)
- R1 One tick at a time. run() takes an exclusive lock `state/factory/tick.lock` (pid + start time, stale after 11 minutes) around the whole tick; while it is held the tick beats "previous tick still running" and exits. The job runner has no per-job lock (runner.mjs origin/main:1900-1907, the two-slot semaphore :1252-1256, an abandoned timed-out tick keeps running :1378-1380), and the row is every 5 minutes with a 9-minute timeout. The harness run, the cheap seat and the drive are detached jobs, never inline.
- R2 A 409 is re-read, not assumed to be a person: every new occurrence raises the row's revision (_factory-issues.js:116-131). On 409 re-read the row; if its status still equals the tick's `from` and sent_back has not grown, retry once at the fresh revision; drop the transition only when the status changed.
- R3 Merge is idempotent. Before calling gh, write C7 runner.merge_started_at and use a branch `factory/<id>-a<n>`; look the pull request up with `gh pr list --head <branch> --state all --json number,state,mergeCommit`: MERGED → take its merge oid; OPEN → merge it; else create it. tick's merge case reconciles whenever merged_sha or merge_started_at is set (never a second merge).
- R4 A fix may not touch what proves it: the build step fails any diff that touches harness/**, _selfchecks/**, app/_design/**, **/fixtures/**, *.selftest.mjs, package*.json or .github/**. The gate re-runs the proof itself with the harness and its support files from origin/main and the app from gated_sha: red on main, green on HEAD.
- R5 Evidence is tied to a commit: harness_green and every check store the sha they ran on; tools.gate fetches, recomputes `git diff origin/main...HEAD`, requires every stored sha to equal HEAD, and returns `{ok:false, reason}` instead of throwing (factory-triad.mjs:23-33 trusts evidence for easy and medium and throws on refusal).
- R6 A send-back starts a clean round: when sent_back grows, the previous merged_sha moves into attempts[n] and runner.{drove, harness_green_once, gated_sha, head_sha, worktree, merge_started_at, merged_at} are cleared, so no check or bound is skipped (factory-fix.mjs:40-52, :94).
- R7 A change is only on for its reviewer: for kind ≠ bug the build brief puts the change behind `flag(<issue id>)`, and the harness proves it off with the flag off and on with the flag on; a change that cannot be flagged stops at spec (factory-review.js:46,82-83 promise "on for you alone").
- R8 Spend is metered: run() passes the run id to cheap-task and gruntFetch and sums cost_usd from the spend meter before tick, so the five-dollar bound can fire.
- R9 The seeded live record goes through `floorHits` (_model-pool-floor.js) and any hit sends the card to plan; the seed file lives under state/factory/seed/<id>, outside the worktree; the harness loads it by path and never writes record values into itself.
- R10 A hard bug is treated as a change: a `kind: "bug"` card whose tier is `hard` (C3: client words, money, trunk, schema, a door's signature) is built behind `flag(<issue id>)` like R7 and its publish ends at `ready` with the hand-off for review, never at `live`; only the reviewer's Looks good reaches everyone (tick's publish case, factory-fix.mjs:123-131, checks the tier before the bug branch).
- R11 A fix's harness lives in `_selfchecks/factory/harness-factory-<id>.mjs`, never at the top of `_selfchecks/` (found 2026-09-29 while wiring the gate, measured on the Hub's main: the tier ratchet `harness-gate-coverage` scans only the top of `_selfchecks/` and its ceiling of 9 untiered checks had 9 — no headroom — so the first merged fix would have blocked every lane's publish, and registering each harness would touch `app/gates.js`, a TRUNK file, making every fix hard). The classifier's own-harness exemption points at that subfolder and the same file at the top of `_selfchecks/` is hard (`_test-factory-classify.mjs`); the harness imports the shared browser rig from the folder above it; STEP 13's monitor runs `_selfchecks/factory/` daily in place of a tier, and retires a file by `git mv` into `_selfchecks/retired/` thirty days after its fix stayed green. Every per-fix harness path in this plan and its kit recipes names that subfolder; the kit's own named harnesses (harness-factory-board, -review, -beacon, -tile) are tiered checks and stay at the top.
- R12 A fix is never easier than medium: every fix carries its own executable harness, which runs unattended on the Studio, and no word list can prove a harness safe (an independent check, 2026-09-29, slipped `process['env']`, a split `import()` name, indirect eval and the Function-constructor escape past a content check). So a change carrying `_selfchecks/factory/harness-factory-<id>.mjs` is classified at least medium — two independent reads, one of them a strong reader of the harness itself — and a harness that plainly launches programs, opens raw sockets, evaluates code, reads the environment or imports beyond the shared rig and safe built-ins is hard. Triage predicts with the harness included, so the build never reads it as a rise. The easy tier stays for changes that carry no factory harness. OPEN, named by the checker: how well the strong reader catches a disguised harness is not provable with fakes (the Sonnet seat is stubbed in every test and not yet measured on the Studio, S0), and a crafted comment in a diff could try to talk the reader round; the live measurement is owed before the factory merges unattended.
- Corrections: vendor-fence.mjs:272 is an outbound-send allow entry, not a write permission — the worktree root must be proven writable by the cheap lane in S0. The factory key's vault entry `factory-agent-key` is not in the vault index; S0 probes it. deploy.yml has cancel-in-progress false (:89-91); the top risk is a double side effect or a stuck merge on the live merge path.
#### Needs a person: nothing (the factory key is ours; the robot bearer can set it; branch protection needs a paid plan, out of scope).
</details>
**Do exactly this:**
1. `factory-triage` (at :05 and :35, five minutes after each Jev tick, around the clock; a row younger than one tick waits one more so its hints exist): reads `inbox` rows (kind `bug`, `wish` and `error` — a PostHog row (`source: posthog`) is routed exactly as a bug with the screen owner as reviewer and NO card, since nobody reported it and a silent fix has no tap to wait for (C11); it gets a card only when it goes to `plan`); for each, it first reads Jev's hints for the row through the door (C10: `factory.kind`, `factory.severity`, `factory.same`, `factory.tier-hint`, `factory.beacon.noise`; at watch they are logged only, at suggest they prefill the next call, at act they set the field and the DeepSeek call is skipped for that field), then one DeepSeek call names the likely file(s) and fills whatever Jev did not, the kind (`bug|wish`), the severity, and any open row on the same screen it is probably the same thing as (`related[]`, linked never merged); a beacon row Jev calls noise at act goes `never` without a card; the classifier tiers the named files' likely diff (`--raise` when unsure, and raised again by `factory.tier-hint`, never lowered) and the tier is stored; `ownerFor()` sets `owner` and `reviewer`; the card gets its C5 intake Details; a `bug` in `easy|medium` → `building`; a `wish` or any `hard` → `spec` (STEP 9 for hard; for a medium wish the same call drafts the C5 spec shape and the reviewer gets Go ahead / Not this); the human-facing posts (`spec`, `ready`, `plan`) are made between 08:00 and 20:00 **Asia/Manila** (the Hub's working day, `manilaToday`) and queued otherwise; bug fixing never waits for the clock. An approved factory suggestion (a card tagged `factory` with no row yet, made by the suggestion door's Approve for its owner) is picked up here: the row is created as a `wish` with the card's owner as reviewer, and `factory` runs the card with `action:"take"`.
2. `factory-fix` (every 5 minutes, ON THE MAC STUDIO — the only machine that can build the Hub, in the clean-copy layout the Hub block §3 names, `projects/business/biz` beside the symlinks, with `HUB_ARTIFACT_WORKSPACE` set; worktrees under the runner's `_work` root, never `/tmp`; the runner's own scheduler row says `host: studio`): a state machine over `building → harness → build → check → drive → merge → publish → ready|live` advancing up to THREE cards per tick, each card as many stages as fit in eight minutes of that tick, through `factory-state.mjs` (C7): **harness** — a Sonnet dispatch with `TEST-AUTHORING: _selfchecks/factory/harness-factory-<id>.mjs` writes a harness that reproduces the reported behaviour on the local server **seeded** with the PM harness fixtures plus the record the reporter had open (read from the live Hub by `record_kind`/`record_id` as `factory`, nine-plus-digit runs masked, written into the local store only), and must be RED on the unfixed worktree (`harness_red_at`); the harness author also confirms the label: a `bug` needs a pre-existing intended behaviour it can name from the screen, else it relabels the row `wish` and the card goes to `spec`; a harness green before the fix is rejected and re-briefed once, then the card goes to `plan`; **build** — `git worktree add` of the target repo's `main` (`repo: hub` for Hub files, `repo: brain` for `projects/ops/factory/**`), then `node projects/ops/cheap-task.mjs --do "<fix-build.md filled>" --dir <the named file's folder> --prove "node _selfchecks/factory/harness-factory-<id>.mjs"` through `run-detached.sh`, polled next tick; the diff re-tiered (a rise re-routes to `spec` and the worktree is discarded, never reused); the revert command written to the record; **check** — `factory-review.mjs review` at the tier (every tier reads the diff, C9); **drive** — `factory-drive.mjs` (PROCESS: a headless drive built on the Hub's own harness tools, `harness/tools/drive.mjs` and the PM harness's `shot.mjs`; a scheduled job cannot dispatch `biz-app-qa` or any Claude Code agent), on the local server built from the worktree, signed in as the reporter through `hub-session.mjs`, phone and computer, before (from `main`) and after (from the worktree) screenshots saved through `/api/files` (four ids) and an outside-the-change screenshot diff; **merge** — `factory-triad.mjs gate`, then `gh pr create` and `gh pr merge --squash` from the worktree (agents merge proven work, REPO §1), at most ONE merge per thirty minutes across all cards and only when the Studio runner reports no publish or sweep running (read through the workflow-runs API), so a merge never cancels another's sweep; `merged_sha` recorded, and a revert pull request prepared (`revert_pr`) for a bug; **publish** — wait until the live Hub's `/api/build-integrity` reports a build whose commit is `merged_sha` or a descendant of it (the build's integrity JSON must carry the commit it was built from; if it does not today, STEP 10 adds that one field to `build-dist.js`'s JSON — a TRUNK change, reviewed once by the three seats before this stage is trusted), then `node app/verify-live.mjs --origin https://hub.heroesandsidekicks.io`, then BUILD §1a's five beneficial tests from the record (measured against the harness and the screen; unmeasurable → the card goes to `plan` with the revert prepared as a draft pull request and never merged by the runner; a person decides — reconciled 2026-09-29 with the design's risk 5, never auto-merge a revert); a reported `bug` → `live` (`factory-review.js live`: the card is handed to the reporter as `finished` with the Look page; the reporter's Done closes it, §12.13); a PostHog `error` row → `live` closes in the ledger with no card and no hand-off (there is no reporter), counted on the tile and told to the screen owner in the next morning's evidence (STEP 15); a change → `ready` with the flag on for the reviewer. Bounds: `attempts` caps at 3, wall clock at 48 hours from `building`, metered spend at five dollars per issue and at most one three-seat review (Anthropic and Codex seats run on subscription tokens and are counted, not priced) → `plan` (STEP 12). All cheap vendors failing → the card says "waiting for a builder" and the tick ends (U26). Every card update through the hand-off door with the five-field update (C8). A fix's harness leaves the daily browser list thirty days after its fix went live and stayed green (it is moved by `git mv` into `_selfchecks/retired/`, never deleted — deletion is its own approval, CORE §2 — and a browser-free assertion replaces it where one exists), so the Chrome lock is never the factory's ceiling. A change to files in the brain repo (`repo: brain`) lands through `projects/ops/bin/brain-land.sh` and is confirmed by its own restart receipt, never by the Hub's publish.
3. Register the eight jobs with distinct minutes; `bash projects/ops/skippy-jobs/restart-daemon.sh`; wait for a runner-written heartbeat row for each (JOBS §1: a hand run proves nothing); a job whose code STEPS 12–15 have not yet written beats `not built yet` and does nothing.
**DEFINITION OF DONE:** with fakes (`FACTORY_FAKE_TOOLS=1`), the machine advances three fixture bugs one stage each per tick through fourteen transitions, rejects a harness that is green before the fix, sends the third attempt, the 48th hour and the fifth dollar to plan, re-routes a tier rise to spec, waits when every vendor is down, merges only through the gate, and hands a bug card to the reporter rather than closing it; and eight runner-written heartbeat rows exist after the restart.
**PROOF:** `cd "$WORKSPACE" && node projects/ops/skippy-jobs/_test-factory-fix.mjs` (CREATED BY STEP 1) → `PASS 14 transitions; 3 cards per tick; green-harness rejected; 3 attempts → plan; 48h → plan; $5 → plan; vendors-down → wait; tier-rise → spec; merge → publish → handoff`, and `command grep -c "factory-" HEARTBEAT.md` after the restart ≥ 8 with the runner as the writer · **FAILS IF:** two stages advance in one tick for one card, a green-before-fix harness is accepted, a fourth attempt runs, a merge happens without the gate, an agent completes a real card, or a heartbeat is from a hand run.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** the moment this step closes, post one dated line into `projects/ops/agent-fleet/failure-learning/PLAN.md`: `Coherence Factory STEP 10 closed <date> — every fix attempt writes a C7 run record under skippy-jobs/state/factory/ and a summary row in D1 factory_runs; the Sunday review may read them.`
### STEP 11 — One seeded easy bug, end to end, live
**FOR NICK:** the first real proof that a reported bug fixes itself: Mae says it on her phone, and later her card shows it live with pictures, with nobody in between. · **Tier:** FRONT
**Start when:** STEPS 2–10 are landed, deployed (`node app/verify-live.mjs --origin https://hub.heroesandsidekicks.io` exits 0) and beating; `prove-e2e.mjs` exists under `projects/ops/factory/` (CREATED BY STEP 1); `gh auth status` succeeds on the Mac Studio (checked in this step, recorded).
**Builder:** GLM 5.3 (`zai`) (the builder of the seeded bug's fix is whatever the machine dispatches) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Sonnet (`biz-app-qa`, browser, as Mae) · **Checker backup:** Qwen 3.8
**Files you may touch:** under `projects/ops/factory/`: the folder `seeds/` with `easy-bug.md` (new: one leaf-file defect in the factory's own Details renderer, `app/js/task-panel-factory.js`, that shows only on a card whose name begins `agent-test` — the Hub renders nothing else for test identities alone, so the seed lives in the factory's own leaf file and is visible only on a test card). **Never** any live screen's real behaviour (RULE 46 spirit: the seeded defect is visible on an `agent-test` card only), `prove-e2e.mjs` (STEP 1's).
**Do exactly this:**
1. Pick the seed: in `task-panel-factory.js`, make the "What you said" line render blank when the card's name begins `agent-test` (one attribute, one leaf file); land it as an `agent-test` change behind a flag on for Mae only, so nobody else's card is touched.
2. As Mae on a phone (the brain's live check on the Mac mini, `--as mae --viewport phone`), open an `agent-test` card the seed made and say "something's wrong: my words are missing from this card".
3. Run `prove-e2e.mjs --issue <id>`: it polls the row every minute, prints each transition with its clock, and at `live` asserts `assertClean` on the Look page's generated sections, four picture ids resolve, `clock_ms` and `merged_sha` written, `revert_pr` prepared, the card sits with Mae as `finished`, the harness file exists and is in the factory's own daily list (STEP 13); then, as Mae, taps Done; then cleans the `agent-test` seed and confirms the card is closed.
4. Write the measured clock into this file's SUMMARY as the baseline.
**DEFINITION OF DONE:** one `agent-test` bug goes from Mae's sentence to live with no human step, the Look page is clean with four pictures, the revert is prepared, Mae's tap closes the card, and the clock is recorded.
**PROOF:** `cd "$WORKSPACE" && node projects/ops/factory/prove-e2e.mjs --issue <id>` (CREATED BY STEP 1) → `PASS live; clock <ms>; look page clean; pictures 4; revert_pr prepared; card with reporter; closed by mae; cleanup LEFT 0` · **FAILS IF:** any stage needed a hand, the generated sections fail `assertClean`, fewer than four pictures, or the card was closed by anyone but Mae.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 12 — Recurrence → one Plan with a person card; a third failed loop → Ecosystem Fixes
**FOR NICK:** a person is pulled in only when the machine has genuinely failed or the same thing keeps coming back, and then with options and a number to answer, never a mess to untangle. · **Tier:** FRONT
**Start when:** `factory-state.mjs` (STEP 10) exists; `_test-factory-recur.mjs` exists (CREATED BY STEP 1).
**Builder:** GLM 5.3 (`zai`) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Qwen 3.8 · **Checker backup:** Sonnet
**Files you may touch:** under `projects/ops/skippy-jobs/jobs/`: `factory-recur.mjs` (new, replacing the placeholder STEP 10 registered); under `projects/ops/factory/briefs/`: `plan-options.md` (new). **Never** `factory-fix.mjs`, the test (STEP 1's).
**Do exactly this:**
1. `factory-recur` (daily 09:00 Manila, and callable `plan --issue <id> [--board ecosystem-fixes]` by STEPS 9 and 10): four triggers make ONE card on Coherence Improvement, in the screen's Suite, to the screen owner — `reopened >= 2`; `reverted_at` set; three issues with the same `screen` in fourteen days; `attempts >= 3` or the wall-clock or spend bound; ONLY a third failed three-seat review loop makes ONE card on **Ecosystem Fixes**, assigned to `nick` (BUILD §1a: "after the THIRD failed loop, go to him … One card on the Hub's Ecosystem Fixes board … Nothing else routes to that board"). Every card is idempotent on `issue_id + trigger` and carries the C5 plan shape: the problem in the reporter's words; what was tried, one line per attempt from C7 with what it measured; two or three options drafted by Sonnet from the run record and the diff, numbered, each with what it means for you, **how long until it is live** ("about a day" / "about a week"), who acts next, and whether it touches anyone else's screen; the line "this changes the software only; nothing is bought" is written because the factory can only change code (a real spend is a four-acts item and never an option here); one recommendation; a due date on the card as the Hub requires; the approval card reads "Go with the recommendation", the drawer takes "option n", and Discard with words means "none of these" and reopens the row as Inbox with the words. The chosen option becomes a `spec` row with the option's text as its C5 spec and re-enters STEP 10 as a change.
2. Test with fixtures: each of the four recurrence triggers → exactly one card on `coherence-improvement`; the third-loop trigger → exactly one card on `ecosystem-fixes` assigned to nick; the same trigger twice → zero new cards; `choose 2` → one new spec row carrying option 2's text; "none of these" with words → the row reopened as Inbox; every option carries a time-to-live and who acts next.
**DEFINITION OF DONE:** on fixtures, each of the four recurrence triggers makes exactly one card on Coherence Improvement to the screen owner, the third-loop trigger makes exactly one on Ecosystem Fixes assigned to Nick, a repeat makes none, `choose 2` makes one spec row carrying option 2's text, and "none of these" reopens the row.
**PROOF:** `cd "$WORKSPACE" && node projects/ops/skippy-jobs/_test-factory-recur.mjs` (CREATED BY STEP 1) → `PASS 4 triggers → 1 card each on coherence-improvement; 3rd loop → 1 card on ecosystem-fixes to nick; repeat → 0; choose 2 → spec row; none → reopened` · **FAILS IF:** two cards for one trigger, a bound or a recurrence landing on Ecosystem Fixes, or an option without a number, a time-to-live and who acts next.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 13 — Monitor after publish: recurrence, revert, stale flags, the factory's own daily harness list
**FOR NICK:** a fix is watched for a week after it goes live; if it comes back, it reopens on its own; if its own test goes red, the undo is merged and a person is told; a flag left half-on switches itself off. · **Tier:** POLISH
**AMENDED 2026-09-29 (STEP 10 design R11):** the factory's own daily harness list is the folder `_selfchecks/factory/` in the Hub, run in full once a day by this job (not registered in a Hub tier, which has no headroom); a red file there is a recurrence of its fix.
**Start when:** STEP 5's PostHog rows and STEP 10's run records exist; `_test-factory-monitor.mjs` exists (CREATED BY STEP 1).
**Builder:** DeepSeek V4 Pro · **Builder backup:** Qwen 3.8 · **Checker:** GLM 5.3 (`zai`) · **Checker backup:** Sonnet
**Files you may touch:** under `projects/ops/skippy-jobs/jobs/`: `factory-monitor.mjs` (new, replacing the placeholder). **Never** `hub-checks-daily.mjs` or `run-hub-checks.mjs` (another lane's hard-coded lists; the factory runs its own), the test (STEP 1's).
**Do exactly this:**
1. `factory-monitor` (hourly, with a daily pass at 06:00 Manila): for every `live` row under seven days that carries a `posthog_issue_id`, a new occurrence on that PostHog issue since `published_at` (the issue reader's cursor) → the U20 path; for every `live` row, the daily pass runs `_selfchecks/factory/harness-factory-<id>.mjs` on the local server built from `main` and a red result → `reverted_at`, the prepared `revert_pr` merged through the gate, the reporter told in one line, and `recur plan`; every flag past `expires_at` neither `all` nor `off` → `off: true`, the card back to the reviewer with "this switched itself off because nobody decided", a `stale` mark for the tile (U28).
2. Test with fixtures for each of the three outcomes and a no-op day.
**DEFINITION OF DONE:** on fixtures, a new PostHog occurrence within seven days of the fix reopens the row, a red harness merges the revert and calls recur, a stale flag switches off and returns the card, and a quiet day writes nothing.
**PROOF:** `cd "$WORKSPACE" && node projects/ops/skippy-jobs/_test-factory-monitor.mjs` (CREATED BY STEP 1) → `PASS posthog recurrence reopens; harness red → revert merged + recur; stale flag → off + card back; quiet day → 0 writes` · **FAILS IF:** a quiet day writes anything, a recurrence is missed, or a red harness leaves the change live.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 14 — The Factory tile on Home, leadership only
**FOR NICK:** one tile tells him whether the factory is earning its keep: how much came in, how much fixed itself, how much needed a person, how fast, what it cost, and how many flags are stale. · **Tier:** FRONT
**Start when:** STEP 3's store holds run summaries (`factory:run:*`) and STEP 10's `factory-metrics` pushes into it; the Home tile wall is live (PM plan, department block); `_selfchecks/harness-factory-tile-20260927.mjs` exists (CREATED BY STEP 1).
**Builder:** GLM 5.3 (`zai`) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Qwen 3.8 (browser) · **Checker backup:** Sonnet
**Files you may touch:** `app/functions/api/factory-metrics.js` (new: 403 unless the signed-in identity is in the task door's `LEADERSHIP` set (Nick, Chantelle); derives the seven numbers from the KV rows and run summaries through `_factory-issues.js`), `app/js/home-factory-tile.js` (new; registers one tile with the Home tile wall's existing registry and draws it only when the door answers 200), `app/index.html` (one script tag); under `projects/ops/skippy-jobs/jobs/`: `factory-metrics.mjs` (new, replacing the placeholder: daily, pushes C7 summaries into `factory:run:*` through the door). **Never** `home.js` beyond the registry call the tile wall already exposes, `util.js`, the harness (STEP 1's).
**Do exactly this:**
1. The door: `GET /api/factory-metrics` → `{in_7d, in_30d, auto_fixed, needed_person, median_report_to_live_ms, send_back_rate, cost_usd_30d, stale_flags}`, every number a query, none stored; 403 for anyone outside `LEADERSHIP`.
2. The tile: seven figures with plain labels; tap → `#tasks` with Coherence Improvement ticked; drawn only on a 200.
3. Run the harness: on the local server it seeds rows and runs, reads the door as Nick, recomputes the seven numbers from the raw rows itself, asserts equality; signs in as Nick and asserts the tile shows those figures; signs in as Mae and as Rizza and asserts 403 and no tile; cleans up.
**DEFINITION OF DONE:** the seven numbers on the tile equal a recomputation from the raw rows, and Mae and Rizza get 403 and no tile.
**PROOF:** `cd projects/business/business-app && node _selfchecks/harness-factory-tile-20260927.mjs` → `PASS 7 numbers = recomputed; 403 and no tile for mae and rizza; cleanup LEFT 0` · **FAILS IF:** any number differs, or Mae or Rizza sees the tile.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 15 — The daily observer, the daily scout, the nightly correction, through the suggestion door
**FOR NICK:** every morning each screen's owner finds up to three ideas for their screens in their Inbox when there is new evidence, he and Chantelle find up to three lines on what changed in the world Coherence lives in, and the factory fixes its own judgement every night from what got sent back or undone — daily, not weekly, until he says otherwise (the requester, 2026-09-27: "the cycle of continuous improvement needs to be daily not weekly until further notice"). · **Tier:** FRONT
**Start when:** STEP 3's rows and STEP 10's run records exist; C10's uses are registered (at watch is enough); `projects/ops/skippy-jobs/jobs/failure-weekly-review.mjs` runs; `_test-factory-weekly.mjs` exists (CREATED BY STEP 1).
**Builder:** Qwen 3.8 (gathering, long context) · **Builder backup:** DeepSeek V4 Pro · **Checker:** Sonnet · **Checker backup:** GLM 5.3 (`zai`)
**Files you may touch:** under `projects/ops/skippy-jobs/jobs/`: `factory-observe.mjs`, `factory-scout.mjs`, `factory-correct.mjs` (new, replacing the placeholders; `factory-correct` is the nightly job STEP 10 registers); under `projects/ops/factory/briefs/`: `observe.md`, `scout.md` (new). **Never** `failure-weekly-review.mjs` itself, `_suggestions.js` (the ideas go through its door, not around it), `_jev-uses.js` (C10 lands through STEP 3's hard change), the test (STEP 1's).
**Do exactly this:**
1. `factory-observe` (every day 07:00 Manila): per OWNER (the six people; C4 groups the screens), gather with Qwen only what is NEW since the last run: that owner's screens' issue rows and reopens, the silent errors fixed on their screens (no card was made; this is where they are told), the Screen Buddy turns on those screens where the buddy could not do what was asked (the brain's refusal log, by screen), the send-back words with Jev's reason hints, the harness failures, and PostHog's usage for those screens (views, clicks, where people leave — allow-listed events only, C2); an owner with no new evidence that day gets nothing; draft candidates with Sonnet and let Sonnet choose at most THREE per owner (Jev's `factory.idea.rank`, C10, only orders those three in the Inbox and hides none); before filing, ask Jev's `factory.same` about each candidate against that owner's open and dismissed factory suggestions of the last sixty days (a paraphrase of a dismissed idea is dropped at suggest); **backpressure**: file nothing for an owner who already has twelve open factory suggestions, and nothing at all when thirty are open across everyone (the store refuses a writer at 200 open and prunes an open item only thirty days after its due date); each suggestion is `{name (≤ 200, plain words), why (what it means for you, what was seen, how big, ≤ 500), assignee: owner, group: "coherence-improvement", due_date: today + 7 (an item due tomorrow leaves the Inbox on day three), tags: ["factory","idea","screen:<prefix>"], what: the C5 spec shape, id: "fx-idea-" + sha16(screen + normalised name)}` through `POST /api/tasks {action:"suggest"}` as `factory`; the suggestion store's own dedupe, its remembered dismissal and the existing `suggestion.wanted` fold do the rest (a `dismissed` reply is counted and never refiled). Native research only; DeepAPI only where a native read fails (RULE 61).
2. `factory-scout` (every day 07:30 Manila): gather with Qwen through `WebSearch`/`WebFetch` (and `gh` for tooling releases) what is new in the last day across two lists frozen in `scout.md`: the audience's world (AI agency owners and marketers who would rather do it in-house — what they ask for, what competitors shipped: GoHighLevel, Monday, Notion, HubSpot, Slack AI, Linear, Cursor-style assistants inside business tools) and the runtime's world (Anthropic and OpenAI model and agent releases, MCP, computer-use, Cloudflare Workers and Pages changes that touch what the Hub can do); judge on Fable (design tier): at most THREE lines a day, each `{what changed, in words · which Coherence screen or gap · recommendation · source named in words}`, and nothing on a day with nothing new; ZION-13's latest output is read first and its items are not repeated; each line is filed as one suggestion to `nick` (Chantelle sees leadership suggestions in her Inbox as well) with tags `["factory","scout","day:<date>"]`, `assignee` for the eventual card = the screen's owner from C4 in `what`.
3. `factory-correct` (every night 23:00 Manila): reads the day's C7 records and files at most ONE change of each kind a night, each a `hard` change on the factory's own files that runs its three-seat review and then ACTS (RULE 60: the review is the gate; the factory's own code has no screen owner and needs no person's tap, so this creates no work for a person — BUILD §1a's fifth test): (a) **a sizing miss**, only when the evidence says the review depth missed a defect — an easy or medium change reverted by a red harness, or sent back with the reason "broke something else", or whose `tier_actual > tier_predicted` — files a change to the classifier's lists or rules with the diff attached, deduplicated by the rule it names against the last thirty days; a send-back that reads "not what I meant" is a spec miss and files a change to the spec brief, "looks wrong" is a drive-evidence miss and files a change to the drive's screenshot rules, and neither touches the classifier; (b) **a repeated reason** — one Jev reason (or "unlabelled") three times in seven days on one tier — files a change to that tier's builder brief (`fix-build.md` under `projects/ops/factory/briefs/`); (c) it writes the day's seven numbers into the run summary. On Sundays it also writes ONE comment onto that ISO week's failure-review card (its id read from the review's own state file, `failure-weekly-review.json` under the jobs' `state/` folder, for the week just ended) with the week's factory section (in / auto-fixed / needed a person / mis-tiered / sent back and why / reverted / cost).
4. Test: fixtures produce ≤ 3 suggestions per owner per day and none on a day with no new evidence; a `dismissed` reply is not refiled; the scout files ≤ 3 suggestions each naming a screen or `gap` and none on a quiet day; the nightly job files one classifier change per evidenced mis-tier, never two for one rule in thirty days, one brief change when a reason repeats three times, and the Sunday comment lands on the right week's card id with the seven numbers.
**DEFINITION OF DONE:** on fixtures, the observer files at most three suggestions per owner per day only when there is new evidence and honours a dismissal, the scout files at most three suggestions a day each naming a screen or a gap, the nightly job files one classifier change per evidenced mis-tier and one brief change per repeated reason, and the Sunday comment lands on the right week's card with the seven numbers.
**PROOF:** `cd "$WORKSPACE" && node projects/ops/skippy-jobs/_test-factory-weekly.mjs` (CREATED BY STEP 1) → `PASS ≤3 per owner per day; quiet day → 0; dismissed not refiled; scout ≤3 each with screen|gap; nightly: 1 classifier change per evidenced mis-tier, 1 brief change per repeated reason; sunday comment on the week's card, 7 numbers` · **FAILS IF:** four ideas for one owner in a day, a suggestion on a day with no evidence, a dismissed idea returns, a line names no screen, two changes for one rule in thirty days, or the comment lands on the wrong card.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** the moment this step closes, post one dated line into `projects/ops/agent-fleet/failure-learning/PLAN.md`: `Coherence Factory STEP 15 closed <date> — the Sunday card carries a factory section as one comment; nothing in failure-weekly-review.mjs changed.`
### STEP 16 — Pointers, registry, toolkit
**FOR NICK:** nothing he notices; this makes the factory something a stranger can find. · **Tier:** POLISH
**Start when:** STEPS 2–15 closed; `_test-factory-wiring.mjs` exists (CREATED BY STEP 1).
**Builder:** DeepSeek V4 Pro for the registry row, the toolkit rows and the manual bullet; the Hub `CLAUDE.md` §7 is written by the overseer (Fable) — CORE §3: editing CLAUDE.md stays on Anthropic · **Builder backup:** Qwen 3.8 · **Checker:** GLM 5.3 (`zai`) · **Checker backup:** Sonnet
**Files you may touch:** `projects/business/business-app/CLAUDE.md` (a new §7 of five lines pointing here; overseer only), `projects/ops/agents/neeko/HUB-MANUAL.md` WHERE THINGS STAND (one bullet, rewritten in place), `projects/ops/artifacts/project-status/registry.json` (one row: `coherence-factory`), `projects/ops/toolkit/TOOLKIT.md` (rows for the classifier, the card writer and the state writer). **Never** the runner's rows (STEP 10's), the test (STEP 1's).
**Do exactly this:**
1. Write the five-line §7 in the Hub's `CLAUDE.md` (overseer), the manual bullet, the registry row (`publicOk` after the same secrets grep the other rows use), the toolkit rows.
**DEFINITION OF DONE:** every pointer resolves.
**PROOF:** `cd "$WORKSPACE" && node projects/ops/skippy-jobs/_test-factory-wiring.mjs` (CREATED BY STEP 1) → `PASS CLAUDE.md §7 present; registry row present; toolkit 3/3; manual bullet present` · **FAILS IF:** a pointer is dead.
**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.
**Handoff (if any):** none.
### STEP 17 — Final sign-off of the FINISH LINE
**FOR NICK:** the factory is declared working by a reader who built none of it, against the nine lines at the top, and the numbers on the tile are the ongoing report. · **Tier:** POLISH
**Start when:** STEPS 1–16 closed with `VERIFIED:` lines.
**Builder:** Fable (reads; never builds) · **Builder backup:** Opus · **Checker:** Opus (reads the same proofs cold; the one check that stays on Anthropic, §1C rule 6) · **Checker backup:** Fable
**Files you may touch:** this file's SUMMARY and STEPS sections; `HISTORY.md` beside it for the postmortem line. **Never** any code.
**Do exactly this:**
1. For each of the nine FINISH LINE lines, name the closed step and its `VERIFIED:` line that proves it; a line with no proof keeps the plan open.
2. Run the plan gate and write the postmortem (what failed, what was confused, what to keep) into `HISTORY.md`, and the failure registry gets its four-column rows the same turn.
**DEFINITION OF DONE:** every FINISH LINE line cites a closed step's dated `VERIFIED:` line and the plan gate passes.
**PROOF:** `python3 projects/ops/agents/check_plan.py <this file>` → `PASS`, and the SUMMARY names the nine proofs · **FAILS IF:** a FINISH LINE line has no closed step behind it.
**If the check fails:** the overseer names the open line and the step that must close; no new bar is added.
**Checker's job:** re-read the nine proofs cold; PASS is the sign-off.
**Handoff (if any):** none.
**Step-writing rules:** every step names the literal command and the literal expected output — "verify it works" is a defect · as many steps as the North Star needs, no more · red-first for any fix step · builds and per-step checks on the cheap tier by name; the overseer never builds; the plan is written and the FINISH LINE signed off on Anthropic or OpenAI.
### STEP 18 — Every change passes the publish checks before it merges, so one lane's break never stops another's publish
**Raised 2026-09-28 by the Pages lane, with Chantelle's "yes write it up"; her question, verbatim: "why are we waiting for other lanes why cant each lane publich as needed".** Measured that evening: the Hub publishes only on a push to main, through one Studio runner, and runs every TIER-1 gate there; between 20:27 and 23:32 six different lanes' failing gates each blocked every lane's publish in turn (the finance tabs harness, the CRM quick-add selftest, the Pages "coming soon" check, the pages-words check, the endpoint gate on a finance robot flag, and the offline-cache check on a stylesheet version); merges land every minute or two and cancel the run under way, so a good change can sit unpublished for an hour; and a lane cannot run the gates before merging on Chantelle's Mac (the Hub's tier-1 helper refuses there, and the fast tier needs the Studio layout and the business database). This is the factory's own need as well: STEP 10's fix runner merges and publishes unattended and cannot do so into a main that any lane can break.
**FOR NICK:** each lane's change is checked on its own before it can reach the live Hub, so a lane that breaks something is stopped at its own door and everyone else keeps publishing. · **Tier:** FRONT
**Start when:** Chantelle's or Nick's dated word on the option below, in their own thread (a teammate lane's report of it counts as the report, not the word); STEP 9's merge gate exists (it does, 90%). **HANDED OFF 2026-09-29 on Nick's word, in Chantelle's thread ("lets handthis off then i want to keep you focused on the bjuild and ill fix thisin another lane"): this step now lives in its own thread; its three review rounds are saved on branch factory-scratch-2026-09-28 under projects/ops/factory/scratch-2026-09-28/step18-review/.** Nick, 2026-09-29: "you can do whatever you need to do to fix it as long as you triuad and find a solution". **MET 2026-09-28: Chantelle, in her own thread: "fix the root problem - triad astra fable to get it done".**
**Review, round 1 (2026-09-28):** the first draft (a check on every pull request, enforced in the landing tool and a hook) was refuted by the attack seat (Astra) and amended by the cold seat (Fable) on the same three points — enforcement was advisory (the Pages lane merges by Python HTTP, the Threads lane by curl, the web UI stays open on this free-plan private repository, which offers no branch protection); a green check on a pull request goes stale the moment main moves; and one runner at 40 pushes in 75 minutes with a 6 min 37 s build cannot also check every push. The attack seat added that a pull_request trigger would run unmerged branch code on the Studio with access to the business database and personal files. The second draft is a serialized, batched landing train (dispatched only from main's own workflow): it builds the ready pull requests merged onto current main with the publish job's own steps and pushes to main only a green candidate, bisecting a red batch; deploy.yml stays as the safety net and names any bypass. Both seats are reading it.
**Builder:** Opus (a Studio-runner workflow and the lane merge tools — control-plane) · **Builder backup:** Fable · **Checker:** a fresh verifier on the real runner · **Checker backup:** Opus
**Files you may touch:** in the Hub repository, the publish workflow file and one new pull-request check workflow; the lane tools that merge by API (`lane/pr.py` and its siblings) — each to wait for the check and refuse on red. **Never** any gate's own assertions (a gate is fixed in the app, never loosened), `app/gates.js`'s tiering, or another lane's files.
**Options (the three-seat review picks, then builds):**
1. **A required check on each pull request (recommended):** the Studio runner builds the pull request's merge commit with the fast tier and every TIER-1 gate; main refuses a red pull request; the lane merge tools wait for the check. Cost: each merge waits one fast build (about ninety seconds) and the single runner queues them.
2. **A merge queue:** pull requests are batched, the batch is built, a breaking one is ejected. Cost: GitHub's merge queue on the repository's plan, and longer waits at busy times.
3. **Revert on red:** a merge whose publish goes red on a gate its own diff touched is reverted automatically and its lane told. Cost: main is briefly broken each time; attribution is by diff and can be wrong.
**Do exactly this (after the word):** run the three-seat review on the chosen option against this block; build it; prove it with a deliberate red pull request on an `agent-test` branch that must be refused, and a green one that merges and publishes; then leave it on for a day.
**DEFINITION OF DONE:** a pull request that breaks a TIER-1 gate cannot reach main; a lane sees its own check's result before it merges; main's publish run is green on consecutive merges for one day.
**PROOF:** the deliberate red `agent-test` pull request is refused with the failing gate named, the green one merges and its publish succeeds, and `gh run list --repo nick-deck/deck-business --branch main` shows no gate-blocked publish for twenty-four hours after.
**If the check fails:** the builder fixes and re-checks the named failure until it passes.
**Checker's job:** re-run the PROOF yourself, once. PASS closes the step.
**Handoff (if any):** the moment this step closes, one dated line in the Hub's own CLAUDE.md §3 replacing "A push to main runs two jobs" with how a change now reaches main.
## 4 · Regret Check (the registry failures this build is actually exposed to)
| Failure mode (registry entry) | The measure in THIS plan that prevents it | Where it lives (section / artifact / gate) |
|---|---|---|
| A second system was built because the first was invisible | The first draft missed the Coherence Builds board, the suggestion door and the approval card's real shape; the attack seat found all three, and the plan now composes them: one new board on the requester's word (2026-09-28) added as one entry in the thirteen files the existing boards live in, no new idea store, no new approval renderer, no second PostHog init | Already true; §0 Ownership check; C1, C8; "What the reviewers changed" |
| A check existed that could not fail | Every proof in this plan is written before the thing it tests and must exit non-zero against the unbuilt surface (STEP 1); every fix harness must be RED on the unfixed worktree before the fix; a green-before-fix harness is rejected | STEP 1; STEP 10 step 2; C7 |
| Work was written to a queue no reader ever visits | Every store names its reader and the arrival is proven: issue rows → `factory-triage` (heartbeat) → a card the reviewer sees → the Look page; the Sunday comment lands on the week's real card id, read from the review's own state | C1; STEPS 10, 12, 15 |
| A new failure state was detected but reached no human | Recurrence, reverts and bounds each produce one card to a named owner, proven to arrive; a red harness after publish merges the revert and tells the reporter; vendors-down never escalates but the card says so | STEPS 12, 13; U26; U27 |
| The builder graded its own work and passed it | Checker ≠ builder on every row (glm≡zai); the tier runner picks the checker from a different vendor; the drive is `biz-app-qa`, which never fixes; the harnesses are authored by Sonnet, not by the builder; the merge gate binds the review to the merged commit | §3b map; C7; C9; STEP 1; STEP 8 |
| A whole screen handed to a cheap model failed six of six times | Fix briefs name ONE file or one feature with named hooks; more than five files is `hard` and goes to the three-seat review, never to one cheap brief | C3; STEP 10 step 2 |
| Three learning jobs were built and never scheduled | The eight jobs are registered and beat from STEP 10, before their code lands; a job not yet built beats "not built yet"; a hand run is not a proof | STEP 10 step 3; JOBS §1 |
| A detector's death was invisible because only its target read it | The Factory tile shows stale flags and the jobs' silence; `factory-monitor` is read by the Sunday job, and the Sunday card is read by people | STEP 14; STEP 15 step 3 |
| A report used names/shorthand only the writer understood | C5: one function writes every generated body from four fixed shapes; a gate refuses paths, model names, `commit`, `PR`, `diff`, `function`, `selector`, code fences; evidence-gated phrases cannot appear without their evidence; the person's own words are never touched | C5; STEP 7 |
| Done was declared before the live surface was checked | The publish stage runs `verify-live.mjs` against the real origin and BUILD §1a's five beneficial tests; the drive is on the real screen as the reporter; STEP 11 is a live end-to-end on the real Hub | STEP 10 publish; STEP 11 |
| **Novel:** the classifier under-tiers a shared change and one cheap opinion ships it | TRUNK and SHARED are hard-coded lists plus an import-count rule proven by fixture (`util.js` → hard); the tier is stored and only rises; the built diff is re-tiered; the merge gate re-runs the classifier on the head; the Sunday job files a classifier change for every evidenced mis-tier | C3; C9; STEP 10 build; STEP 15 step 3 |
| **Novel:** the fix loop thrashes and burns tokens | Three attempts, 48 hours, five dollars, one three-seat review, then a person; three cards per tick, one stage each; cost summed per run and shown on the tile; vendors-down pauses the attempts but not the wall clock | STEP 10; C7; U26; U27 |
| **Novel:** flags pile up half-on and nobody knows what is live for whom | Every flag switches itself off at fourteen days and the card returns to the reviewer; the server is authoritative and open pages re-read within a minute; the tile counts stale flags; a bug carries no flag and has a prepared revert instead | C6; STEP 13; STEP 14 |
| **Novel:** the error and replay vendor receives something on the floor | Every exception keeps its frames but leaves the page with a blank message value and blank breadcrumbs, and the whole event is dropped in the browser when anything matches the floor shapes (the same expressions as the server's scanner, kept in step by a selftest); replay is off until a dated word, and when on masks all inputs and all text and unmasks only elements marked safe; autocapture is off; IP capture is off; the harness proves the exception body and the replay state live; a floor-shaped observation goes to sp-sec in one line | C2; STEP 5; NOT in scope |
| **Novel:** the vendor is down or refuses us and the factory goes blind | Errors: the plain beacon to our own door stays in the fixtures and is proven green at STEP 5, so it can be switched in; flags: every door evaluates the definitions from its own KV copy, refreshed on request by one writer, and keeps the last copy when PostHog does not answer, and the page's helper falls back to the door's GET when PostHog's script is not loaded; replays and usage are evidence, never a gate, so their absence slows nothing | C2; C6; STEP 5; STEP 6 |
| **Novel:** an agent closes a real card on its own (§12.13) | No factory code path calls `complete` as an agent on a real card: a bug card is handed to the reporter as `finished` and their Done closes it; `agent-test` cards close through the test-card path | C1; STEP 7; STEP 10; STEP 11 |
| **Novel:** a Jev "yes" quietly becomes the decision | Every C10 use starts at watch, is promoted only in Settings on measured agreement, never lowers a tier, never closes a step, never sends, and "unsure" goes to whoever the step already names; the fallback on no judgment is the step's old behaviour, proven by a selftest case | C10; STEP 3; STEP 10 |
| **Novel:** the factory's own nightly corrections become a daily stream of human review work | A correction is filed only when the evidence names the miss, at most one of each kind a night, on the factory's own files, reviewed by the three seats and then acted on with no person's tap (RULE 60) | STEP 15 step 3 |
| **Novel:** daily ideas flood the Inboxes | At most three per owner per day and only with new evidence; twelve open per owner and thirty overall as backpressure; due in seven days; at most three scout lines a day and none on a quiet day; the suggestion door's dedupe, remembered dismissal, Jev's paraphrase check and the `suggestion.wanted` fold; the cadence is Nick's word (daily until further notice) and one constant per job | STEP 15; U24; U25 |
| **Novel:** every fix adds a permanent browser harness and the one Chrome lock becomes the ceiling | A fix's harness retires thirty days after its last green into a browser-free assertion or is deleted with a note on the run record; one merge per tick; the drive is headless | STEP 10; STEP 13 |
| **Novel:** a merge bypasses the runner because `main` is unprotected | The monitor reads every commit on `main` since its last run and counts any not made by the runner or a person's own login on the tile as `unreviewed`; branch protection is filed through the money door once STEP 11 proves the loop | C9; STEP 13; STEP 14 |
## 5 · Topology and roles
- **OVERSEER-AUTHORITY:** none named (`projects/ops/OVERSEER-AUTHORITY.md` CURRENT HOLDER: "NO SEAT IS NAMED. THIS GRANT IS DORMANT.", 2026-08-28). **The four approval classes (money leaving · credential rotation · irreversible destruction · a message sent as Nick) and the floor (logins · credentials, tokens and keys · government IDs · card, bank and routing numbers) never move on the overseer's word.**
- Thread layout: one overseer thread (this plan's Fable session) for all five lanes; builders and checkers as cheap dispatches through `cheap-task.mjs` / `route-build.mjs`; Sonnet for test authoring (`TEST-AUTHORING:` line) and diff review; `biz-app-qa` for browser drives; Astra through `astra-run.sh` detached; Opus through the Agent route for cold reads.
- Overseer: Fable · Workers: GLM 5.3 (`zai`), DeepSeek V4 Pro, Qwen 3.8, Sonnet (tests, reviews) · Cap: no numeric cap (BUILD §2: "no numeric agent cap applies"; the template's 8 and 40 do not bind); the fix runner's own bound is one merge per tick and three cards in flight per tick
- State files location: this file is its own state (RULE 51: beside it only `HISTORY.md` for the postmortem and `STEPS.json` if the update tool writes one); per-issue run records under `projects/ops/skippy-jobs/state/factory/` (C7); questions and assumptions are lines in this file's SUMMARY until answered.
- **Board card id:** nt-20260928-003139-2211
- The card sits on Coherence Builds under the agent name `factory-planner` (opened 2026-09-27 as a project card: spec, north star and finish line each under 500 characters, thread and engine named, born in Backlog).
- **Artefact consumers:** issue rows → `factory-triage`, the Details renderer, the Look page and the tile; run records → `factory-fix`, `factory-recur`, `factory-monitor`, `factory-metrics`, `factory-correct`; cards → the six team members (proven by read-back by id in every step that writes one); suggestions → each owner's Inbox; the Sunday comment → the failure-review card people already read.
- **Write-contention (parallel lanes in a shared checkout):** KIT lane writes only the seventeen files STEP 1 names; HUB lane writes only under `app/functions/api/factory-*`, `app/functions/api/_factory-*`, `app/js/factory-*`, `app/js/task-panel-factory.js`, `app/js/home-factory-tile.js`, `app/migrations/factory.sql`, `.github/workflows/factory-gate.yml`, and the script tags in `app/index.html`; BRAIN lane writes only `lib/hub-factory.mjs` and the named `server.js` lines; PROCESS lane writes only under `projects/ops/factory/` (never STEP 1's tests); JOBS lane writes only `factory-*.mjs` under `projects/ops/skippy-jobs/jobs/`, `factory-state.mjs` under `lib/`, the runner's factory rows, `long-runners.json`'s factory row, and `state/factory/`. Each lane declares its files at the top of its `LANE-*.md` in the Hub repo (the collision gate). Checkouts proven writable at lane open by `git status --porcelain` and a one-byte write test.
**Per-stage topology — counts DECLARED at plan time (machine-gated: a number in every row):**
| Stage | Overseer | Sub-overseers | Workers |
|---|---|---|---|
| Plan (STEP 1) | 1 | 0 | 2 |
| Framing (STEPS 2–5) | 1 | 0 | 4 |
| Elements (STEPS 6–9) | 1 | 0 | 4 |
| Details (STEPS 10–13) | 1 | 0 | 4 |
| Output (STEPS 14–15) | 1 | 0 | 3 |
| Proof (STEPS 16–17) | 1 | 0 | 1 |
**The walk-away contract — a stranger resumes the drive from files alone:**
- **STATE FILE:** `projects/business/coherence-factory/PLAN.md` (this file: SUMMARY and STEPS are the state) and `projects/ops/skippy-jobs/state/factory/*.json` for per-issue runs
- **HEARTBEAT ROW:** `coherence-factory` in `projects/personal/skippy-app/ala-state/work-threads.json` (written when the drive opens)
- **MORNING-REPORT LINE:** `Coherence Factory — <n> of 17 steps closed; e2e clock <ms or not yet>; issues in / auto-fixed / needing a person this week` in `projects/ops/walkaway/REPORT.md`
## 6 · Evals — what "working" means, decided now
| Capability | Check (exact command or procedure) | Pass looks like |
|---|---|---|
| 1 · a sentence in the drawer makes one row and one card | `node _test-hub-factory.mjs` (CREATED BY STEP 1) in the brain repo, then the live case on the Mac mini as Mae at phone width | `PASS 6 tools`; a row and a card read back by id within 60 s |
| 2 · an unreported page error is one fingerprinted row | `node _selfchecks/harness-factory-beacon-20260927.mjs` | `1 row count 2; payload keys exact; queue drained` |
| 3 · classify easy / medium / hard, never lower | `node projects/ops/factory/_test-factory-classify.mjs` (CREATED BY STEP 1) | `12 rules, path-list refused, --lower refused, util.js → hard` |
| 4 · an easy bug fixes itself end to end | `node projects/ops/factory/prove-e2e.mjs --issue <id>` (CREATED BY STEP 1) on the live Hub | `live; clock <ms>; look page clean; pictures 4; revert_pr prepared; card with reporter; closed by mae` |
| 5 · a change waits behind a flag; Looks good / Send back behave | `node _selfchecks/harness-factory-review-20260927.mjs` | `look page 8 sections at 375 and 1280; try-it dean:true mae:false; looks-good → all; second tap 410; send-back words → reopened` |
| 6 · a hard change cannot skip its review or reuse a stale one | `node projects/ops/factory/_test-factory-triad.mjs` (CREATED BY STEP 1); the branch protection listing | `no-record merge refused; stale-sha refused; record → allowed; changes → re-attacked once; 3rd loop → ecosystem-fixes card`; `factory-gate` required |
| 7 · recurrence makes exactly one card, on the right board | `node projects/ops/skippy-jobs/_test-factory-recur.mjs` (CREATED BY STEP 1) | `4 triggers → 1 card each on coherence-improvement; 3rd loop → 1 card on ecosystem-fixes to nick; repeat → 0` |
| 8 · daily ideas and the ecosystem note through the suggestion door | `node projects/ops/skippy-jobs/_test-factory-weekly.mjs` (CREATED BY STEP 1) and the daily heartbeat rows | `≤3 per owner per day; quiet day → 0; dismissed not refiled; scout ≤3 each with screen|gap` and two runner-written rows a day |
| 9 · the tile's numbers are derived, leadership only, and the factory corrects itself nightly | `node _selfchecks/harness-factory-tile-20260927.mjs`; the nightly run's filed changes; the Sunday comment on the week's failure-review card | `7 numbers = recomputed; 403 and no tile for mae and rizza`; one classifier change per evidenced mis-tier and one brief change per repeated reason; a Sunday comment with seven numbers |
| 10 · Jev's seven questions run at watch with fallbacks | `node app/functions/api/_factory-issues.selftest.mjs` (the C10 cases: every use answers or refuses through the one door; a refused judgment leaves the step's result exactly as before) | `PASS 43/43` including `jev: 7 uses registered at watch; no-judgment → unchanged` |
## If you get stuck (all steps)
Before writing "blocked": (1) re-read the step's START WHEN line — most "stuck" is a misread gate, (2) try a concrete workaround, (3) write one line to the overseer naming the ONE missing artefact. Then keep working every other step whose inputs exist. Never idle on a blocker; never end a turn waiting on a background result.
## Your loop
Every pass: every step whose START WHEN inputs exist and which is not yet CLOSED is running, up to the cap → each builder runs its own PROOF, hands to its checker → PASS closes it, FAIL loops it → repeat until the FINISH LINE is proven.
## What the reviewers changed (the cold reads, 2026-09-27)
**Astra (attack seat, Codex `gpt-6-astra`, read-only, high effort, 2026-09-27): verdict on the first draft "do not proceed".** Every finding was re-checked against the file it cited before being folded in; none was overruled.
- A Coherence Builds board already exists on Nick's word (`tasks.js:267–274`) and stays his board for large builds and human updates; the factory's cards live on the new Coherence Improvement board, one entry in the same ten files (the requester, 2026-09-28; C11, STEP 4).
- The task door validates groups, has fixed editable keys, needs a name with `agent-test` for a test card, and never lets an agent complete a real card (§12.13) → the Details renderer fetches the row by card id instead of new card fields; test cards are named `agent-test`; a bug card is handed to the reporter and their Done closes it (C1, STEP 7, STEP 10, STEP 11).
- The shared approval card is one sentence, Look, Test, Approve, Discard, `{nonce, decision}` from the brain's queue, with no picture layout or free text → a Look page carries the pictures and sections; Send back words and "option n" go through the drawer conversation and brain tools; the runner stores the draft through a `factory_ready` tool registered in `APPROVAL_ONLY`, or STEP 2 records the one ask of the core (C5, C8, STEP 2, STEP 7).
- The Hub publishes on a push to `main`, not on a pull request → the runner merges (REPO §1) and a `factory-gate` status check binds the review record to the merged commit (C9, STEP 9, STEP 10).
- Monitoring named no rollback; report and beacon fingerprints could never match; the daily-check list is another lane's hard-coded array → flags switch off (server-authoritative, re-read within a minute); bugs get a prepared revert pull request; recurrence joins only within a source and `related[]` links across; the factory runs its own daily harness list (C6, C7, STEP 13).
- `BZ.knowsFinance` includes Mae and Rizza → the metrics door is 403 outside the task door's `LEADERSHIP` set (STEP 14).
- Ideas already have a home: the suggestion door with dedupe and remembered dismissal, and Nick's ruling that suggestions are approved before they become tasks → the observer and the scout file suggestions; the separate refusal file is gone (STEP 15).
- C1 gained `revision` and `request_key`; C2 became an allow-list with per-fingerprint counts; C3 gained SHARED, the import-count rule, diff-only input and a stored tier that only rises; C4 returns `{owner, reviewer}` with longest-prefix; C5 became four evidence-gated shapes that never touch the person's words; C7 gained a lease, revisions and `factory_runs`; the third failed loop routes to Ecosystem Fixes (BUILD §1a); the five beneficial tests run at publish; the model-matrix conflict is recorded and resolved by Nick's later word; the Hub `CLAUDE.md` edit stays with the overseer (CORE §3); the STEP 10/16 cycle is broken by registering the jobs in STEP 10; the kit is seventeen files counted honestly; three cards advance per tick; triage runs around the clock; `proceed with changes` is re-attacked once; per-issue bounds of 48 hours and five dollars; the Sunday mis-tier is evidenced by a send-back or a revert; the weekly card id is read from the review's state file; U14 carries how-to-try and who-gets-it lines; U18 options carry effort, who acts next and "no money involved".
**Opus (cold seat, a fresh session, 2026-09-27, without sight of the attack): verdict on the second draft "do not proceed".** Every finding was re-opened in the code before being folded in; one could not be re-verified from this Mac and is recorded as the reader's claim.
- `main` is not protected and cannot be on the current GitHub plan (reader's live read; not re-verifiable here, no `gh` login) → the merge gate is runner-side and says so; the monitor flags commits that bypassed the runner; branch protection is a money-door item after STEP 11 (C9).
- An agent owns a card only under its own name; `review` hand-offs go only to gracie or neeko; nothing hands a card to an agent; `take` is how an agent runs a person's card (`tasks.js:1978–1986, 1606–1611, 3641`) → the brain mints a `factory` agent session for reports; `ready`/`live` hand off `finished` to the reviewer; Send back keeps the reviewer as owner and `factory` takes the card (C1). The "Nick's Review" column name is the PM lane's and is on the NEXT list; the Look page's first line names the actual reviewer.
- `HS_DB` is the clients' master database with one choke point; every later feature got its own D1 → the factory was going to get `FACTORY_DB`; superseded 2026-09-28: the ledger lives in the Hub's KV under `factory:` keys behind one choke point, so no database is created and the selftest runs in-process (STEP 3, C1).
- The Hub builds only on the Studio in the clean-copy layout; a scheduled job cannot dispatch `biz-app-qa` or a Claude Code agent; the one runner cancels older publishes; the local store starts empty; the team's day is Manila → the fix runner lives on the Studio, drives with a headless `factory-drive.mjs`, calls Fable and Opus through the Claude CLI (verified in STEP 9 or recorded NOT MEASURABLE), merges once per tick, waits for any publish containing `merged_sha`, seeds the local store with the reporter's record, and posts to people in Manila hours (STEP 9, STEP 10).
- No element renders only for test identities → the STEP 11 seed lives in the factory's own renderer and shows only on an `agent-test` card.
- Only a third failed three-seat loop routes to Ecosystem Fixes, and that card goes to Nick (BUILD §1a); bounds and recurrence go to the screen owner (STEP 12). A money rule (Rizza's screens, `loans`) and a client-words rule (RULE 48) join the classifier; the import-count rule that made 43 of 170 files hard is dropped; every tier reads the diff (RULE 64); gate code and self-tests are TRUNK (RULE 12) (C3, C9).
- Reviewer = the reporter for a bug, the screen owner for a wish (C4); U17 offers "put it back the way it was" (the prepared revert) as well as "still wrong"; U18 says how long until live instead of agent-hours and offers "none of these"; the observer files at most three ideas per owner, only with evidence; an approved suggestion is picked up by triage and taken; harnesses retire after thirty days green; the numeric cap line follows BUILD §2; Anthropic and Codex seats are counted on the tile, not priced.
- Overruled, with the reason: "cut the scout, the Sunday self-correction, the tile and the flags until STEP 11 proves the loop" — the requester asked for all four outcomes on 2026-09-27 (always improving, ideas from the ecosystem, automatic fixes, a person for what recurs); the plan keeps the scope and orders it so STEP 11 proves the loop before the weekly and Sunday jobs are filled in (they are registered and beating "not built yet" from STEP 10).
**Opus (cold seat, second round, 2026-09-28, on the sixth draft — PostHog and the boards): verdict "proceed with changes".** Every finding re-opened and folded in; none overruled. PostHog has been live in the Hub since 2026-08-09 with errors on and replay off by a recorded ruling → no sign-up, keys minted by an agent, the factory's script replaces the inline block and keeps its guards and scrub, replay is a switch off until Nick's or Chantelle's dated word, the exception keeps its frames and blanks the value, a browser floor check is ported, source maps go inside the build; flags evaluate in our own door from the definitions JSON (no vendor library in the Functions), one writer refreshes on request, `FactoryFlag` falls back to the door under automation, "off within about two minutes"; the board is one entry in thirteen mirrored files, verified by grep (TRUNK, one review), the Suite is written with the robot bearer after each create, the board is filtered not grouped, factory cards obey Neeko's checks as they are (decision hand-offs while a person is owed, `agent-test` cards closed within two hours) and two scan changes wait on the PM lane's word; the door gains `report_posthog` and `source: posthog`; one fingerprint definition; the noise card is gone; STEP 13 reads PostHog occurrences.
**Astra (attack seat, second round, 2026-09-27, on the fourth draft — Jev and the daily cadence): verdict "do not proceed".** Every finding re-opened in the code and folded in; none overruled.
- Jev's door caps any use fed the person's words at suggest, only an `order` use may reach act, no judgment is ever written onto a record, and an unconfigured use is off → C10 is rewritten as hints only: prefill and annotate, `same` over one candidate, `beacon.noise` on structured fields with a fold instead of a silent `never`, `idea.rank` ordering the three Sonnet chose and hiding none; registry, batch adapters, outcome joins and Settings switches named as the deliverable; a $3 share of the ceiling; 200 subjects a day per use; the hints door's stage-stripping named as an ask of the Jev lane.
- The task door needs a due date on an agent's card; `review` never reaches a person; a hand-off always moves the card; an agent cannot undo a completion → due date on every card, `finished` only, comments for progress, a NEW linked card on recurrence after Done (C1, STEP 7, U20).
- Every fix adds a harness under a folder the sizing rules call TRUNK, which made every easy fix hard → the run's own harness file is exempt, nothing else (C3).
- Publish evidence: `verify-live.mjs` compares a served token to a local bundle, not a commit → the live build-integrity door must carry the commit (one TRUNK change, reviewed once); brain-repo changes land through the brain's own landing script (STEP 10).
- One merge per tick still cancels sweeps → one merge per thirty minutes, only when the runner is idle (STEP 10). Deleting retired harnesses is a deletion → `git mv` into a retired folder (CORE §2).
- The classifier and the merge gate are a safety wall → built by Sonnet, not GLM (CORE §3; STEP 8, STEP 9). A missing draft door or missing review seats no longer let their steps close (STEP 2, STEP 9).
- Daily suggestions would fill the store and the leadership Inbox → three per owner chosen by Sonnet, twelve open per owner and thirty overall as backpressure, due in seven days, paraphrases caught by Jev against dismissed ones (STEP 15). Triage runs five minutes after each Jev tick so hints exist (STEP 10).
- The nightly correction filed hard work for every send-back → it files at most one change of each kind a night, only when the evidence names the miss (a revert or "broke something else" → the sizing rules; "not what I meant" → the spec brief; "looks wrong" → the drive rules), and the factory's own changes act after the three seats with no person's tap (STEP 15).
**Sixth draft (2026-09-28, the requester's two asks): unreviewed.** PostHog as the sensor layer and the Coherence Improvement board with its Suite field and Neeko rules (C11) went in after the third read; the owed cold read covers them.
**Next cold read: owed on this sixth draft before any lane opens** (RULE 60 allows three loops; this plan has had three reads: attack, cold, attack).
## SUMMARY — a few plain-English lines, read by the status generator
**Handed to a cloud session, 2 Oct 2026, 8:55 am:** the state below is unchanged; the cloud session takes the Hub-side and single-file work, and finishing the first factory fix live stays on the Mac Studio. The pick-up is the HANDOFF TO CLOUD block at the top of this plan.
**Where it stands, 2 Oct 2026, 7:05 am (landed for a machine restart):** the factory ran live on test bugs overnight and 13 real faults were found and fixed; the third test fix passed its test, its build and both readers on its own and its screen check now reads clean, but it has not yet been merged and published by the factory itself, which is the next step. The Factory tile is live on Home. The factory is armed on test reports only and is not opened to real ones. Nothing is running from the finishing session; the pick-up is the LANDED FOR MACHINE RESTART block at the top of this plan.
**Where it stands, 2026-09-29 (lane moving from Chantelle's Mac mini to the Mac Studio):** STEPS 1, 3 and 8 are done; STEP 6 is 95% (live on the Hub), STEP 7 95%, STEP 9 95%, STEP 4 90%, STEP 2 50%, STEP 10 78% by an independent checker, STEPS 5 and 11–17 untouched. STEP 10's whole code path is on the cloud main (commits d7bed41271, 2697d51fe8, 99ea0fdc7c): the fix runner's state machine and its round; one tool per stage (harness, build, check, drive, merge, publish); the check stage's reader program `projects/ops/factory/factory-read.mjs`; the harness author `projects/ops/factory/factory-author.mjs`; the drive `projects/ops/factory/factory-drive.mjs` with `factory-pixels.mjs`, `factory-drive-read.mjs` and `factory-drive-hub.mjs`; the wiring `projects/ops/skippy-jobs/lib/factory-wire.mjs`; the triage job with its guess; the reporter-record reader `lib/factory-record.mjs` (R9); the first-live limiter `lib/factory-first-live.mjs`. The kit `node projects/ops/skippy-jobs/_test-factory-tools.mjs` has 26 sections (`--only <name>`; set FACTORY_HUB_API to a Hub checkout's app/functions/api, and FACTORY_HUB_BUNDLE to a generated app/functions/_artifacts-bundle.js for `hub` and `real`). On the Studio the factory runs UNARMED: after its job runner restarted onto this code and the factory key was copied into the Studio's own vault (it held only the Mac mini's copy; same key, not a rotation), jobs.log reads `factory-fix OK — wired; not armed, so acting on nothing` and `factory-triage OK — not armed: 0 inbox rows seen, none routed`. ARMED (in `lib/factory-first-live.mjs`) is false until a commit arms it for STEP 11; OPENED (real rows) stays null until STEP 11 passes and Nick or Chantelle gives a dated word. The independent sign-off of STEP 10 (2026-09-29) passed STEP 10's PROOF (`_test-factory-fix.mjs`: `PASS 14 transitions; 3 cards per tick; …`), 24 of the 26 kit sections and R1–R12, and held STEP 10 at 78% for two reasons: (a) only 3 of the 8 factory jobs had beaten on the Studio (factory-triage, factory-fix, factory-monitor); the other five run once a day at 10:07, 18:03, 18:33, 20:11 and 21:48 Studio time and their first windows had not come yet — recheck jobs.log on the Studio; (b) the real drive proof (`--only real`) failed twice on its first case, once as `the sentinel page changed` (565 rows) and once as `the control was not clean on identical code` (1028 rows). Root cause, measured the same day: a freshly started local Hub is slow on its first page loads (its functions compile per tree), so under machine load a picture is taken while the page still shows its grey loading placeholders; the drive already waits up to 15 s for placeholders to clear (a page on Home keeps one that never clears, so every settle waits the full 15 s and a real case takes 3–6 minutes). The fix that was about to be built: warm each local Hub before photographing (one throwaway visit to the route and to the second page, then the real pictures), and raise the placeholder wait to 25 s; prove it by `--only real` passing twice in a row with nothing else using the drive's port. What the drive's real runs taught, each already fixed: a locked Mac never finishes a Chrome for Testing picture, so the drive uses `~/.cache/puppeteer/chrome-headless-shell` (installed on the Studio, whose screen is locked too); a region that grows by a fraction of a pixel redraws the text below it, so its height is rounded up in every picture; a leftover local Hub from a killed run holds the port, so the helper takes back a port held only from the factory's own work folder. The Studio facts: it holds the jobs role; the generated bundle is at `~/actions-runner/_work/deck-business/deck-business/app/functions/_artifacts-bundle.js`; `/usr/local/bin/claude` is 2.1.237 and refuses `--restricted`, `~/.local/bin/claude` is 2.1.284 and is the one the reader and author pick; `claude -p` fails under the runner with "Not logged in" unless an account's sign-in is passed as CLAUDE_CODE_OAUTH_TOKEN, which is how the reader and author call it; Sonnet, Opus and Fable all answered that way on the Studio; Codex is not installed there, so the hard tier stays closed (hard rows wait at spec); `gh` cannot use its keychain under the scheduler, so the wiring hands the vault key `github-pat-classic` to gh and git by environment. Next, in order: (1) the drive warm-up above, then `--only real` twice; (2) recheck the five once-a-day factory heartbeats on the Studio; (3) ask a checker that did not build it for STEP 10's VERIFIED line; (4) correct R9's text, which names `_model-pool-floor.js` — the floor check actually wired is `hardFloorHits` in `projects/personal/skippy-app/lib/grunt-egress-scan.mjs`; (5) STEP 11 needs STEPS 2–9 landed and beating first: STEP 2 waits on the Screen Buddy core lane's draft door, STEP 4's last part on the PM lane's board-scan changes, STEP 5 (PostHog) is untouched, STEP 9 needs only Codex on the Studio for the hard tier. Other sessions change files in this workspace all day; the cloud main moves every few minutes, so land with fetch, rebase, push, retried. On the Mac mini this lane left nothing uncommitted; its scratch folder is gone.
Seventh draft, 2026-09-28: the requester said the plan looks good as a starting place and asked two things, both now in it, and a fourth independent read (Opus, cold, on the sixth draft) found that PostHog has been live in the Hub since 2026-08-09 with errors on and replay off by a recorded ruling, so nobody has to sign up for anything: the factory's script replaces the existing PostHog block, an agent mints the two keys, and replay stays off until Nick or Chantelle say the word. PostHog is the sensor layer — it catches and groups page errors, runs the per-person switches and measures which screens are used, and keeps a masked replay once switched on — while the intake, the ledger, the sizing, the reviews, the release and the cards stay ours (the section "PostHog: what it takes over, what stays ours"). The factory gets its own board, Coherence Improvement, one entry in the thirteen files the existing boards live in, with a Suite filter per feature suite and Neeko's checks obeyed as they are (C11); Coherence Builds keeps large builds and human updates; a bug or issue is a row in the ledger and becomes a card only when a person is or will be involved, and a silent error the factory fixes on its own never has a card. The plan passes the plan gate and has had four independent reads. STEP 1 is done (2026-09-28): all seventeen red-first test files exist and each refuses with a named reason against the unbuilt surface; the Hub six sit on the branch factory-kit and the workspace eleven are on main, with a one-page recipe per file under projects/ops/factory/briefs/kit. The cheap lane could not write one of them in six tries (every vendor timed out or wandered), so they were written on Anthropic under recorded overrides, and the cheap checker's own re-run of the seventeen is still owed. The ledger now lives in the Hub's own KV behind one choke point, not a new database. STEP 4's board entry is edited in all thirteen files on the branch factory-board (the Hub's board selftest 9/9 green) and waits on its three-seat review before it merges. Nothing is live for the team yet. Jev sits in seven places as an adviser only (C10); the observer, the scout and the self-correction run daily (STEP 15). The six cycles are drawn in FLOWS.html beside this file and shared in the app as a PDF. Its card is nt-20260928-003139-2211 on Coherence Builds, with Chantelle for a go. NEXT list (not worked until the plan is approved): the review column is named "Nick's Review" for everyone (the PM lane's to rename); the Jev hints door strips a judgment's stage (an ask of the Jev lane); branch protection on the Hub's main needs a paid GitHub plan (a money item after STEP 11); the requester is assumed to be Chantelle from the machine and the sign-in, and the card goes to her. When it is built, the team will report a problem or a wish by saying one sentence to the helper that already sits on every Hub screen, page errors will log themselves, small and medium bugs will fix themselves and come back as a picture page the reporter closes with one tap, changes will wait for the screen owner's tap behind a switch that turns itself off if nobody decides, shared parts of the Hub will be argued over by three different AI minds before anyone builds and nothing can be merged without that record, anything that keeps coming back will reach a person as numbered options, every screen owner will find up to three ideas a week in their Inbox, Nick and Chantelle will find one short weekly note on the outside world, and one tile will show the numbers. The first live end-to-end run will set the report-to-live baseline.
## STEPS
<!-- The live status checklist, read by `status-regen.mjs` / `project-status-page.py`. The heading
above must be exactly "## STEPS" with nothing else on the line. -->
> One numbered line per STEP block above, same numbers. CURRENT STATE, rewritten in place — the STEP blocks say what to do; this section records what has been proven.
```
1. The harness kit: every proof written red-first — 100%
DEFINITION OF DONE: all seventeen kit files exist and each exits non-zero against the unbuilt surface naming what is missing
PROOF: `ls` of the four kit folders counts 15, plus prove-e2e.mjs and the brain's _test-hub-factory.mjs, then each file run once → 17 red-first, 0 failing, 0 missing (the Hub half on branch factory-kit 0f6a915f1, the workspace half on origin/main 9ed242d0f6)
VERIFIED: 2026-09-28 (100%, 17 of 17 kit files exit 1 with a refused: line against the unbuilt surface; one recipe per file under projects/ops/factory/briefs/kit; the cheap lane could not write one file in six tries, so the eleven workspace files were written on Anthropic — six under recorded cheap-vendor-failed overrides, five as control-plane files the router itself keeps inside — and the six Hub files the same way; checker's re-run owed)
2. [UI] Screen Buddy takes a report, a wish, a send-back, a choice; the job-callable draft — 50%
DEFINITION OF DONE: as Mae on a phone, one sentence makes one row and one card within 60 s, read back by id; a stored draft for Dean shows as the shared card
PROOF: `node _test-hub-factory.mjs && node _test-hub-factory.mjs --sabotage; echo $?` (CREATED BY STEP 1)
VERIFIED: 2026-09-29 (50%, the tool module only: projects/ops/factory/brain/hub-factory.mjs on main (f83b3648da), written by a cheap model and its five robustness fixes by a second cheap-model work order; by an independent checker that did not build it — the kit test and its sabotage run twice each, the builder's robust proof twice, and its own probe of all six tools against a throwing door, a rowless 200, a 503, a non-JSON body, a missing id and identity forwarding, 42/42; the real door accepts the forwarded record fields. OPEN, the other 50%: wiring the six tools into the brain's server.js (nick-deck/skippy-code, which this Mac's GitHub sign-in cannot see; the brain lands from the Mac Studio through brain-land.sh) and the live proof as Mae on a phone. The tools must send the factory key with the factory session.)
3. The issue door, the owners map and the KV store — 100%
DEFINITION OF DONE: the selftest passes 43/43 and a report on the local server produces one row and one card on Coherence Improvement
PROOF: `node app/functions/api/_factory-issues.selftest.mjs`
VERIFIED: 2026-09-28 (90%, merged as nick-deck/deck-business#2448 (6699b7c19); by an independent checker that did not build it: 43/43 three times under a second, the endpoint gate 7/7 on its own commit, and each of seven mutations — a merge without its occurrence key, PostHog totals added, request keys unscoped, the test-session card prefix dropped, 200 with no card, the words cap dropped, a damaged row served — turned the selftest red; three adversary rounds, the last accepting the design; the whole Hub selftest folder 383 pass with the four failures that predate it. LIVE PROOF 2026-09-28 on hub.heroesandsidekicks.io (build with the merge): one report as the factory agent riding Mae → row fx-dee018ecc727 with a card of the same id on coherence-improvement, read back through the tasks door and by card id through the issue door; the same words again → the same row, count 2; the row marked never and the card archived, nothing left behind. Measured on the way: a session a script mints for a person carries the agent mark Skippy, so the door (rightly) refused it as a report — STEP 2's brain tool must mint the factory session riding the person. Tracked: the name `factory` can be claimed by any teammate's agent pass, fixed in STEP 6's rework)
4. [UI] The Coherence Improvement board, the Suite field, the card and its Details renderer — 90%
DEFINITION OF DONE: three fixture cards render their Details from the door, the door-down line shows, cleanup LEFT 0
PROOF: `node _selfchecks/harness-factory-board-20260927.mjs`
VERIFIED: 2026-09-29 (90%, merged as nick-deck/deck-business#2651 (ec5926c79): the board and Suite field (part 1, live since #2397), and the card's Details section — by two independent checks that did not build it: the first found the C5 filter untested, a lookup for every card, a reporter's own words filtered and three of four shapes unreachable (the store refused their bodies); all fixed — C1 gained `details`, the renderer draws any shape from it and never filters own words, factory ids only — and the re-check reproduced every fix (filter off and own-words-filtered mutants caught by case 8, selftest 45/45, endpoint gate 7/7, board harness 9/9 in a real browser, only the five named files changed) and named one gap, now closed and proven red-first (case 9: an ordinary card is never looked up; fails with the guard removed). The board harness passes 10/10 in a real browser with the Hub's own rig and the real doors; it had never run its browser half before this step (a page API the rig lacks, a crash that printed PASS). OPEN, the other 10%: the two board-scan changes, waiting on the PM lane's word; and the Details section seen on the live Hub by a person, which STEP 11's end-to-end run covers.)
5. [UI] PostHog: errors caught and grouped, replays masked, source maps on every publish — 0%
DEFINITION OF DONE: two identical injected errors make one factory row with count 2 linked to one PostHog issue; the exception left the page with its frames, a blank value and nothing on the floor; replay proven silent while off (masked when on); the fallback beacon proves green
PROOF: `node _selfchecks/harness-factory-beacon-20260927.mjs`
PROBED 2026-10-02, ~4:20 am (F6 narrows this step to the reader; source maps and replay deferred): the vault holds `posthog-personal-api-key`; it sees one project, "Default project" (id 549957, us.posthog.com); `/api/environments/549957/error_tracking/issues/` answers 200 with 12 issues whose fields are id, status, severity, name, description, first_seen, assignee, external_issues, cohort — no count or last-seen, which come from the query API (`$exception` events grouped by `$exception_issue_id`). The Hub door `report_posthog` makes rows WITHOUT a test mark, and before OPENED the limiter lists only test rows, so the reader needs one Hub change: an issue whose name or description carries `factory-test` arrives marked test (the beacon harness injects exactly that). Design: `lib/factory-posthog.mjs` (read issues + counts; before OPENED pass only `factory-test` issues) called from the triage job each pass; the door client and limiter gain `reportPosthog`; the monitor's `posthogOccurrences` reads the same counts.
PROGRESS 2026-10-02, ~4:30 am: the reader is built and on main (e5a5a66b2c; 5 proof cases, run live read-only: before OPENED 0 issues pass, opened all 12 with real counts); triage syncs it first each pass and keeps routing through a PostHog outage; the Hub half (`report_posthog` marks the factory's own injected errors test at birth, selftest 48/48) is shipping. Open: the monitor's posthogOccurrences, and the beacon harness (an injected `factory-test` error → one row with count 2).
PROGRESS 2026-10-02, 9:30 am (cloud): the monitor's `posthogOccurrences` now reads `occurrencesSince` with the vault key (brain 1ec79a2). The beacon harness is rewritten to F6's scope on Hub branch `factory-recur-third-screen` (nick-deck/deck-business#4236, 051dadc30). Owed: its first live run on a Mac with `FACTORY_AGENT_KEY`, then `--mutant=one` red.
VERIFIED: 2026-10-02 (the monitor's PostHog count and the harness's code, read by an independent reviewer that did not build them: MERGE, after one FIX round on the harness — the floor check read PostHog's own timestamps, gzip and `data=` bodies were unreadable, and a run without the key could leave a row open; all three fixed and re-read MERGE. The live run is not yet proven.)
6. Per-change flags — 95%
DEFINITION OF DONE: the selftest passes 24/24 (off on its own key wins every fixed-order race, a lost write answers 409, the factory key gates both doors, the helper re-reads within sixty seconds and fails closed); a flag on for Dean reads true for Dean and false for Mae, and false for Dean within about two minutes of off; the key's hash set on the live Hub
PROOF: `node app/functions/api/_factory-flags.selftest.mjs`
VERIFIED: 2026-09-28 (95%, merged as nick-deck/deck-business#2506 (3a546b0e0) and live: by the cold checker that did not build it — 28/28 five times on 6454d8d46, its own race script 0 losses in 200-trial batches, the endpoint gate 7/7, the fence held — and by the adversary, whose third round accepted the design after 3,000 random interleavings with 0 off-undone cases and asked only for test pins; those four pins (the one-write off, the bearer gate, two re-enable cycles, an overtaken re-enable) were added and proven red on their mutants, 31/31 six times, re-run by the builder only — the checker's re-run of the 31-case version is the 5% owed. LIVE PROOF 2026-09-28 on hub.heroesandsidekicks.io: the factory key's hash set once as the robot bearer (200); a switch on for Dean read true for Dean and false for Mae, promote reached Mae, off reached both at once, a set after off answered 409, the factory name without the key 403, the switch left off; STEP 3's live proof re-run with the key: one report → one row and one card, the same words → count 2, cleaned up)
7. [UI] The Look page and the review door: Try it · Looks good · Send back — 95%
DEFINITION OF DONE: as Dean at phone and computer width, the Look page shows only the C5 sections, Try it is on for him alone, Looks good promotes and closes with his words and a second tap is refused, Send back needs words
PROOF: `node _selfchecks/harness-factory-review-20260927.mjs`
VERIFIED: 2026-09-29 (90%, merged as nick-deck/deck-business#2710 (9f61686fb): the review door `factory-review.js` and the Look page `factory-look.js`, with the door's own TIER-1 selftest `_factory-review.selftest.mjs`; by an independent checker that did not build it, over two rounds — the first found three door faults (a bug marked live could never be closed by its reporter, ready could re-target a change already with its reviewer, live hid why a hand-off failed) and the second a real gap (choose had no reviewer check: any signed-in person could set another's choice) plus an unasked leadership bypass on every review action, two-picture sides unenforced and Done able to promote a wish; all fixed red-first by cheap-model work orders and the re-grade reads SHIP, the checker reproducing each fix with its own scripts. The browser proof passed in the overseer's run at 375 and 1280 (eight sections and four pictures, Try it on for Dean alone, Looks good to everyone with a second tap refused, Send back needs words, cleanup LEFT 0); the checker's own browser run then passed as well (a third round, 2026-09-29: every case at 375 and 1280, exit 0). LIVE 2026-09-29: the live Hub serves the Look page and the door refuses an unsigned call 401 and a GET 405. OWED for 100%: a walk on the live Hub as the reviewer, which STEP 11's seeded bug does end to end. DEFERRED by design: go_ahead and choose_recommendation belong to STEPs 9 and 12)
8. The classifier and the tier runner — 100%
DEFINITION OF DONE: every fixture diff lands in its tier, a path list and lowering are refused, util.js is hard, a hard review without a record for the commit is refused
PROOF: `node projects/ops/factory/_test-factory-classify.mjs && node projects/ops/factory/_test-factory-review.mjs` (CREATED BY STEP 1)
VERIFIED: 2026-09-28 (100%, by an independent checker that did not build it: both tests passed twice, and mutating a rule at a time — the commit comparison, easy running two checks, the --lower refusal — turned each test red; the first check found three gaps (no commit comparison in a hard review, two C3 rules missing, a --lower assertion that could not fail), all closed and re-proven; one named gap left, over-broad not unsafe: the money and client-words rules match a file's name anywhere rather than C4's route prefix, which can only raise a tier, never lower it)
9. The three-seat review runner and the merge gate — 60%
DEFINITION OF DONE: no record or a stale record → merge refused; record for the head → allowed; proceed-with-changes re-attacked once; third loop → one Ecosystem Fixes card; the three seats proven callable from a job; the headless drive returns four pictures
PROOF: `node projects/ops/factory/_test-factory-triad.mjs` (CREATED BY STEP 1)
VERIFIED: 2026-09-28 (90%, by an independent checker that did not build it: 13 cases pass and each of seven mutations — skipping the re-classification, an easy change with no review, one read for medium, a harness with no red time, a fourth loop, skipping the re-attack, three pictures — turns the test red; the first check failed the gate for trusting the declared tier and asking for evidence only on hard, fixed and re-proven; the re-check's one gap, a harness green before red, now has its case. NOT MEASURABLE here, and the one missing thing: the three seats called from a scheduled job — this Mac's Claude command line is signed out, and the seats are wired on the Mac Studio by STEP 10)
PROGRESS 2026-09-29: the seats are callable from a job on the Mac Studio, measured there over a non-interactive shell exactly the way the check stage's reader calls them (the first Claude command line whose help lists --restricted, `~/.local/bin/claude` 2.1.284, an allowed account's sign-in passed by environment, no keychain): Sonnet, Opus and Fable each answered. Astra is NOT MEASURABLE — the Codex command line is not installed on the Studio — so, per step 0, the hard tier stays closed (a hard row waits at spec) while easy and medium fixes run. The headless drive returns four pictures (STEP 10's drive, seven real cases passing).
CORRECTED 2026-10-01 (four reviews, see FINISH RUN): the Codex command line IS on the Studio now, but only the gate and the seat loop exist — no seat adapters, no `bind`, no triad briefs, nothing calls `run()`; and Astra's probe showed the hard gate accepts a record holding only the matching commit (F2-S2). 60% until F5.
PROGRESS 2026-10-02, ~12:30 am (F5's spec exit, FINISH LINE 4): a change below the hard tier now gets a plain-words plan (`lib/factory-spec.mjs`: what you said, what will change, what it touches, how you will try it; one low-cost model call; file names, folder paths, model and vendor names, builder words and code dropped; secret shapes refused in the request, the screen, the send-back words and the answer; no floor check, no writer), handed through the Hub's new `spec_ready` door to the person who asked, who taps Go ahead (row to building) or Not this with words (plan rewritten); a hard change, an unknown tier, or a plan the door refuses goes to a person the same tick. Workspace on main (8fd8db3965, 7fcc05917f); Hub on branch `factory-spec-exit` (3615af8b8, 42babb279): door actions and card buttons shown only to the asker. The independent verifier read FIX with nine findings; all fixed red-first (spec proof 22 cases, 15 of 15 guard mutants killed; round proof 4 new cases; door selftest 3 new cases; the board's browser check gains case 10, red on the old buttons); its re-check closed every leak and named two test gaps, both closed; Hub side merged as nick-deck/deck-business#4077 (TIER-1 side by side clean, API selftests PASS).
PROGRESS 2026-10-02, ~1:00 am (F5 seats, second spec after Astra's FIX-SPEC): built and on main (80c840c52e). The check stage's reader program has three more seats — propose (Fable), attack (Astra through astra-run.sh in an empty folder, one 38-minute deadline across attempts, never `~/.codex2`), cold (Opus) — each with the reader's restrictions, words and diff fenced as data and scanned for secret shapes first, its notes scanned before they are stored, and its model stamped from the command line (Claude's reported model, Codex's log), never from the answer. For a hard card the check stage starts all five readers on one frozen input, accepts an answer only when its seat, commit and input digest match, lets a "later" seat wait 30 minutes while keeping the others, and sends any CHANGES back to build. The gate needs all five MERGE reads of the head with the three seats by their own models (the old record no longer counts). Triage builds a hard bug; a hard wish gets the plain-words plan. Proofs: reader 19 new cases, check stage 9, gate 11, merge tool 3, all red on the old code; full suite FAILS 0. Independent verification: FIX with six findings, then a re-check: fixed — Astra now runs isolated (`astra-run.sh --isolated`: no user config or connectors, no shell, code host, web, browser, computer use or plugins; probed live twice by this lane and once by the verifier with the production flags), the secret check no longer refuses ordinary code, the check stage waits for running readers, caps "later" at three and restart-alls at three and hands an unknown tier to a person (ef9d7953b4, 7b6a51f121). Known and kept: Codex's agent-collaboration tools stay listed but run under the same isolation; a job killed by the helper's 45-minute stop can leave its Codex run until Astra's own 38-minute deadline; reader spend is not in the card's cost bound (the three seats run on subscription accounts). Still open for 100%: a live hard run on an agent-test row.
10. Triage and the fix runner as a state machine, registered and beating — 85%
PROGRESS 2026-10-02, ~12:30 am: the real drive proof (`--only real`, a real local Hub and Chrome) passed twice in a row on the Studio under heavy load (load 54–91; run 2 ended 11:38 pm) — reason (b) below is closed. The whole kit (27 sections plus the five STEP 1 tests) reads FAILS 0. First armed live tick, 12:10 am: the fix round failed because the factory's state folder (not in git) did not exist on the jobs machine; fixed red-first (the round makes its own folder and only a held lock is aged; 7fcc05917f) and the folder made on the Studio. Owed: the independent VERIFIED line and the five once-a-day heartbeats.
DEFINITION OF DONE: fourteen transitions, three cards per tick, green-before-fix rejected, the bounds send to plan, vendors-down waits, merge only through the gate, a bug card handed to the reporter; eight runner-written heartbeats
PROOF: `node projects/ops/skippy-jobs/_test-factory-fix.mjs` (CREATED BY STEP 1)
PROGRESS 2026-10-01 night (FINISH RUN, builder's own proofs; the independent check is running): F1 readiness built (the drive asserts no placeholder, a complete document and the region, waits for network idle, warms each local Hub, and answers unmeasurable instead of photographing a half-drawn page) — `--only real` twice still owed on a quiet machine. F2 built, each red-first: S1 and S3 in the merge gate (7 new cases); S2 in the hard gate (3 cases); S4 already covered by R4 and the harness hash; S5 the sandbox (`lib/factory-sandbox.mjs`, 11 cases run for real under sandbox-exec, red on an unsandboxed mutant; Chrome runs inside it with its own sandbox off); S6/R13 the rehearsed undo in publish (8 cases E1-E8, red on the old tool); S7 the sweep in the round (2 cases) plus the Hub's new `plan` action; S8 `ARMED_BY` and `selfCheck` enforced in factory-fix and factory-triage. The tick sends an undone fix to plan at once and hands errors off through the door. Branch `factory-finish-run` (5716c7c713, d7ffbed9f7).
INDEPENDENT CHECK 2026-10-01, ~10:20 pm (verifier that did not build it): MERGE on all three change sets (workspace, Hub doors, the seed), every new rule shown red on a mutant; seven findings, fixed the same night before arming: (1) the undo rehearsal's sandbox now reads the seed folder and the Hub bundle like the harness run; (2) the sandbox now denies the whole home folder (a new canary, red on the old rules); (3) a stuck row counts as handed only once its card reached the owner, the round reports any it could not reach and the fix job's heartbeat turns WARN naming them (this replaces S7's 24-hour clause with an immediate warning); (4) the merge gate fails closed when git or gh cannot be read, reads up to 300 open pull requests and refuses at the limit, and recognises the factory's own commits only by their exact title shape; (5) the live card's "what failed" line follows the real number of tries. Left open on purpose: the hard-gate record does not name its seats (the hard tier stays closed until F5's seat adapters exist), and the undo's inverse check compares file names, not the full diff. F4's tripwire (`lib/factory-tripwire.mjs`, 11 cases) pauses every merge while its pause file exists.
11. [UI] One seeded bug (rated medium by R12), end to end, live — 0%
REWRITTEN 2026-10-01: start conditions, report path, seed and revert rehearsal per FINISH RUN F3; machine proof ends at `finished`; Mae's own Done tap is FINISH LINE 9 evidence
DEFINITION OF DONE: an agent-test bug goes from Mae's sentence to live with no human step; the Look page is clean with four pictures; the revert is prepared; Mae's tap closes the card; the clock is recorded
PROOF: `node projects/ops/factory/prove-e2e.mjs --issue <id>` (CREATED BY STEP 1)
PROGRESS 2026-10-02, ~12:30 am: armed on test rows only (ARMED_BY set, OPENED null) and in the Studio's shared copy; its job runner restarted onto that code at ~11:45 pm. The seeded bug was filed at 11:41 pm as the factory riding Mae, marked test: row `fx-9d18cd38a5cd` ("agent-test: seeded bug — my own words are missing…", the blank-own-words seed from #4029). Triage holds a report for 30 minutes by design, then routes it; prove-e2e reads PENDING inbox. Next: watch it to `finished` with prove-e2e, then idea E's live undo.
FOUND LIVE 2026-10-02, ~12:50 am: the seed never left the inbox because triage's guess (and the plan writer, and the check stage's cheap reader) named a data class the low-cost model door does not have ("personal"; the door's name for team, family and client words is "health"), so every guess threw and every row "waited". The proofs had copied the same wrong name. Fixed on main (80c840c52e); the proofs now check against the door's own list. Next: the Studio's shared copy pulls it, the job runner restarts, triage routes the seed.
FOUND LIVE, second fault, 2 Oct ~1:20 am: with the class fixed the guess still read "unreadable twice" — the door answers `{ json: { content: [thinking, text] } }` and the factory read `r.content` (always empty) with DeepSeek's thinking left on; the guess, the plan writer and the check stage's cheap reader all had it, and their test stand-ins had copied the wrong shape. Fixed (7b6a51f121: `r.json.content` text blocks, thinking off as the door's other callers do; stand-ins return the real shape, a thinking block first); the fixed guess answered live. 1:36 am: triage routed the seed ("1 routed"), the fix job moved it to `harness`, and the harness author's job is running in a fresh Hub worktree. KNOWN BLOCKER AHEAD, not this lane's: no Hub publish has gone live since 11:19 pm; every deploy since fails or times out on the finance lane's TIER-1 gate `harness-laneB1triage-20260802.mjs` (15-minute limit on the runner), so the seed's merge will wait at `publish` until that lane's gate is healthy (the Forms lane raised it; this lane is not touching it).
FOUND LIVE, faults 3 to 5, 2 Oct ~1:40–2:20 am, each fixed red-first and landed the same hour: (3) the test author read the account gauge's retired shape and threw before any call, so no harness had ever been written (538e8e2302; the author then wrote the seed's harness, which diagnosed the seed exactly and went red at 2:00 am); (4) the cheap lane refused the build brief — it named forbidden files, which the lane reads as files to change, and named no file to change when the guess found none (956300bf73: the fence names folders, the build takes the app file the harness names as the cause, the proof names every file); (5) the brief's short paths went through the workspace's `app/` and `_selfchecks/` links to the MAIN Hub checkout and were refused as symbolic links (ddb42ab7d2: every path is the workspace path inside the fix's worktree). Also landed: a reopened row starts a clean round and a row reopened twice goes to a person (8388ddcf72). The seed's attempts may run out on the old code before the runner restarts; if it lands in plan, a fresh seed is filed on the fixed code.
2:38–3:05 am: the first seed used its three attempts on the old build code and went to plan (its card archived, its row set to never — it found five live faults). A second seed was filed (`fx-b7cb3c729206`, same seeded defect, new words). An independent verifier read tonight's fixes: FIX, eight findings, all fixed the same hour (c82600b6de) — chiefly S1: the Hub's main takes about twenty commits an hour, so "main moved" refused nearly every fix; an easy or medium fix whose files main has not changed since its base now merges onto the moved main (hard ones and overlapping ones are rebuilt). Also the gauge's real account names, file names in reporters' words, the harness-file fallback's filters, a slash-bearing key, a reopen's attempts and spend, merge-gate world refusals as waits (3b0630d14e), and proofs for every surviving mutant. STEP 13: the monitor's daily pass now counts the tripwires and pauses or clears every merge.
3:10–3:40 am, second seed: routed at 3:10, harness red at 3:20, the builder (now reading the right folder) made the correct one-line fix, but the proof failed all ten cases — FAULT 6: the author's harness looked for the card on the Tasks list, where test cards are hidden by design, so it was red for the wrong reason and nothing could turn it green. Fixed (6aaff06777): red needs a passing setup case beside the failing ones; a harness whose every case fails is rewritten with the reason in the author's next brief, and the brief demands passing setup cases and opens test cards directly at #task/<id>. The second seed was sent back to its harness stage on a fresh worktree (attempt 2, attempts reset; a test row) and its author is running with the rejection in its brief.
3:39–4:06 am: the rewritten harness was a proper red (seven setup cases ok, only "What you said shows the sentence" failing, at 375 and 1280); FAULT 7: it named no app file and the guess named none, so the build had nothing to change (76d16d0ea1: the build takes the files that draw the harness's region; the author writes a CAUSE line); the cheap model then made the one-line fix and all nine cases passed (3:55 am). FAULT 8: the build stage re-tiered it HARD — the classifier's eval rule matched the rig's browser.eval and the payroll rule read the harness's own sign-in header and cdp.send (20798c9ad9: those rules now read the change, not its proof). A tier only ever rises, so the second seed (now hard) was retired and its spec card, which had reached Mae at 3:55 am, archived. Third seed filed 4:06 am: `fx-4e8dead05fef`.
LANDED FOR RESTART 2026-10-02, 7:05 am: the third seed (fx-4e8dead05fef) passed harness, build and both readers on its own; its one scheduled drive failed on a transient and sent it to plan at 4:40 am; six more faults in the picture check were then fixed (051dc1cac3) and the drive run by hand reads clean on phone and computer. Not yet done: restart the runner, reset the row to its drive stage, merge, publish, hand-off, prove-e2e. The pick-up is the LANDED FOR MACHINE RESTART block at the top.
12. Recurrence → one Plan with a person card; a third failed loop → Ecosystem Fixes — 40%
DEFINITION OF DONE: each of the four recurrence triggers makes exactly one card on Coherence Improvement to the screen owner, the third-loop trigger exactly one on Ecosystem Fixes to Nick, a repeat makes none, choose 2 makes one spec row, none-of-these reopens the row
PROOF: `node projects/ops/skippy-jobs/_test-factory-recur.mjs` (CREATED BY STEP 1)
PROGRESS 2026-10-01 night: the minimal slice F2-S7 is built — the Hub review door's `plan` action (one card to the screen owner saying why and what was tried, a second call refused 409, six selftest cases) and the round's sweep that calls it for every plan or spec row nobody was told about. The recurrence core is built too (`jobs/factory-recur.mjs`: four triggers make one card each to the screen owner, a third failed review loop one card to Nick on Ecosystem Fixes, a repeat makes none, choose records the owner's answer); STEP 1's `_test-factory-recur.mjs` passes all five cases (it was red: no job). Its live tools (the card door, the key store) are not wired, and its job beats that honestly.
PROGRESS 2026-10-02, 9:30 am (cloud): the third-screen trigger is a Hub door action, `recur_screen` in `factory-review.js` (nick-deck/deck-business#4236, 6102df3). Three issues on one screen in 14 days make one card to the owner; a repeat makes none; test and real rows never mix. Selftest 64/64, red 10 on the old door. Open: `jobs/factory-recur.mjs` calls it for each new row (its live tools are still not wired); merge #4236 on a Mac.
LIVE 2026-10-02, 12:22 pm: merged as nick-deck/deck-business#4236 (267bbb00f7) by the merge desk; live build 78920e3ce carries it; the live door answers recur_screen 403 to a person. Open: the brain job that calls it.
PROGRESS 2026-10-02, 4:30 pm (cloud): the daily job (`jobs/factory-recur.mjs`, Studio daily at 20:11 UTC) now asks the live door's `recur_screen` once per screen whose rows reach three in 14 days (`screenRun`; test-marked rows only until OPENED). Proof: a scratch proof red first (no `screenRun`), then PASS; `_test-factory-recur.mjs` PASS. Independent reader: MERGE; its one finding (unmarked rows before opening) is fixed. Open: the other triggers' live tools (reopened twice, reverted, bounds, third loop); the Studio's job runner picks this up on its next pull and restart.
PROGRESS 2026-10-02, ~4:55 pm (cloud): reopened twice is live too. `reopenedRun` sends a row reopened two or more times, not yet handed and not ready, live, done or never, to a person once through the review door's `plan` action; the daily job calls it. Proof red first (no `reopenedRun`), then PASS; `_test-factory-recur.mjs` PASS; independent reader MERGE. Bounds and undone fixes already reach a person (the fix round's plan sweep, the monitor). Left in STEP 12: the third failed review loop's Ecosystem Fixes card (no Hub door makes a card on that board yet), and choose/none-of-these from the owner's reply.
VERIFIED: 2026-10-02 (the third-screen trigger: an independent reviewer that did not build it re-ran the selftest green and the red-first copy red, `FAIL 10` exit 1, and read the claim, the counting and the failure path; MERGE)
13. Monitor after publish — 75%
PROGRESS 2026-10-02, ~2:50 am: the tripwires now have their daily count — the monitor's daily pass counts rows stuck a day in spec or plan with nobody told, fixes merged but not live within six hours, automatic undos and the day's spend, reads the factory jobs' heartbeats (quiet = no beat for an hour on the fix job, two on triage), keeps fourteen days of counts, and sets or clears the pause every merge reads (`dailyTripwire`, 7 proof cases, red on the old monitor). Open: the PostHog recurrence (STEP 5) and the weekly undo drill on the test copy.
DEFINITION OF DONE: a beacon recurrence within seven days reopens; a red harness merges the revert and calls recur; a stale flag switches off and returns the card; a quiet day writes nothing
PROOF: `node projects/ops/skippy-jobs/_test-factory-monitor.mjs` (CREATED BY STEP 1)
PROGRESS 2026-10-01 night: the core `run(rows, flags, tools, now)` is built and its STEP 1 test passes all four cases (it was red: no run exported). The scheduled job is wired (landed 5c4f392190): once armed, each live fix's own harness re-runs on the daily pass (06:00 Manila) inside the sandbox in a fresh Hub worktree; a red one merges its REHEARSED undo and goes to a person; a review switch past its expiry is switched off and the change goes back to its owner; outside the pass nothing runs (`prove-monitor-job.mjs`, 6 cases, kit section `monitorjob`). The F4 tripwire and pause are built (`lib/factory-tripwire.mjs`; the gate honours the pause file). Open: the PostHog recurrence (needs STEP 5's reader), computing the tripwire's daily counts in the monitor, and the weekly undo drill on the test copy.
14. [UI] The Factory tile on Home, leadership only — 100%
VERIFIED: 2026-10-02, ~4:25 am (100%, by an independent checker that did not build it, in a real browser on the live Hub: as Nick the Factory tile sits under the greeting at 1440 and 390 with every number readable; /api/factory-metrics answers 403 "leadership only" to Mae and Rizza and their Home has no tile at either width; the 7-day and 30-day counts (5 and 5) match a recount from the five rows read through the factory door; limit stated: with no row live or in plan yet, fixed-alone, the median, the send-back rate and cost are all zero, so only the two counts tell a right formula from a wrong one)
PROGRESS 2026-10-02, ~4:15 am: the placement fix is live (nick-deck/deck-business#4123 in build 5147c548e4, 4:03 am); photographed on the live Hub as Nick at 1440 and 390: the tile sits under the greeting and above the Leadership card, numbers readable on a phone; Mae and Rizza 403. Owed for 100%: an independent checker's VERIFIED line in a real browser.
PROGRESS 2026-10-02, ~3:20 am: LIVE since the 2:49 am publish (build 9718d9d899, which carries #4050). Measured on the live Hub: as Nick, Home shows the tile with its numbers (4 this week, 4 this month — tonight's test rows, the only rows until OPENED); /api/factory-metrics answers 403 "leadership only" to Mae and Rizza, and the tile draws only on a 200. One defect seen live and fixed: the tile drew above the greeting (Home draws after it) — it now anchors below the greeting and the leadership dashboard and re-places on every redraw; browser check case 5, red on the old tile; Hub checks clean; shipping. Owed: the live photo after that publish, and the independent VERIFIED line ("browser").
DEFINITION OF DONE: seven numbers equal a recomputation from raw rows; Mae and Rizza get 403 and no tile
PROOF: `node _selfchecks/harness-factory-tile-20260927.mjs`
PROGRESS 2026-10-01 night: on Hub branch `factory-tile` (not merged yet): the door `app/functions/api/factory-metrics.js` (every number computed at request time from the rows and the switch records, 403 outside leadership), the self-contained tile `app/js/home-factory-tile.js` (adds itself to Home only on a 200; touches no other Home code), the issue store accepting `clock_ms`, `cost_usd`, `fixed_alone` and a dated `created_at` (47/47), and the harness brought over from `factory-kit`. The factory records those three facts on the row when a fix goes live (landed 2656ebb736). Owed: the page's script tag, the harness run, the gate comparison and the merge.
PROGRESS 2026-10-02, ~12:30 am: merged as nick-deck/deck-business#4050 (script tag, harness in TIER-2, tile placed on the PM Home). NOT LIVE YET, measured: the live Hub still serves d1a01f1beb; every publish since 9:21 pm failed or was cancelled (the latest on the onboarding lane's ZO11 gate, which that lane pushed a fix for at 12:13 am), and a signed-in photo of the live Home as Nick shows no tile. Next: once a publish lands, photograph Home as Nick (tile) and as Mae (no tile).
15. The daily observer, the daily scout, the nightly correction, through the suggestion door — 30%
DEFINITION OF DONE: ≤3 suggestions per owner per day only with new evidence, honouring a dismissal; ≤3 scout lines a day each naming a screen or a gap; one nightly correction per evidenced mis-tier and per repeated send-back reason; the Sunday comment on the week's card with seven numbers
PROOF: `node projects/ops/skippy-jobs/_test-factory-weekly.mjs` (CREATED BY STEP 1)
PROGRESS 2026-10-01 night: the core (`observe`, `scout`, `nightly`, `sunday` in `jobs/factory-observe.mjs`) is built and STEP 1's test passes all six cases (it was red: no job). Not wired: the evidence reader, the suggestion door, the outside-world source for the scout, and idea A's replay set that must pass before a nightly correction lands; the scout and correct jobs are still placeholders.
16. Pointers, registry, toolkit — 95%
CHECKED 2026-10-02, ~4:25 am by an independent checker: registry, toolkit and manual pointers resolve; the proof hard-coded the guide's section number (§7) while the Coherence Factory section is §8 — the proof now looks for the section by name and its plan path, reading the Hub's main when the working copy lags (e5a5a66b2c); PASS. Owed: the checker's own re-run for the VERIFIED line.
PROGRESS 2026-10-02, ~1:20 am: the proof itself was broken (it looked one folder above the workspace, so no pointer could ever be found) — fixed; registry row, toolkit rows (classifier, card writer, run-record store) and manual bullet resolve; the Hub guide gained §8 on the Coherence Factory (nick-deck/deck-business#4091, merged; documents only). The proof reads the Hub guide from the workspace's Hub checkout, which a person owns and auto-pull skips, so §8 shows there once that checkout is on main. Owed: the independent checker's run.
DEFINITION OF DONE: every pointer resolves
PROOF: `node projects/ops/skippy-jobs/_test-factory-wiring.mjs` (CREATED BY STEP 1)
17. Final sign-off of the FINISH LINE — 0%
DEFINITION OF DONE: every FINISH LINE line cites a closed step's dated VERIFIED line and the plan gate passes
PROOF: `python3 projects/ops/agents/check_plan.py <this file>`
18. Every change passes the publish checks before it merges — 0%
DEFINITION OF DONE: a pull request that breaks a TIER-1 gate cannot reach main; a lane sees its own check before it merges; main publishes green on consecutive merges for a day
PROOF: a deliberate red agent-test pull request is refused with the gate named, a green one merges and publishes, and no gate-blocked publish for twenty-four hours after (CREATED BY STEP 18)
```
- **DONE (green ✅) requires ONE dated `VERIFIED:` line from the independent checker** — the builder's own proof run is not a `VERIFIED:` line. A `[UI]` step's `VERIFIED:` line must contain the word "browser" (a real logged-in drive-through). A `VERIFIED:` line naming a failing count keeps the step open at that percentage.
- **IN PROGRESS (◐) shows a percentage**; **NOT STARTED (⬜)** = no `VERIFIED:` line.