VOICE: Voice, Skippy, Gracie and Neeko - the three assistants as one clean system

The actual documents the agents read and work from, shown exactly as they are on disk — not a summary. See the progress view instead · All projects

Plan PLAN.md

# PLAN.md — THE THREE ASSISTANTS, ONE CLEAN SYSTEM

**THIS IS THE ONLY PLANNING DOCUMENT FOR THIS LANE.** Its state companion is `PROGRESS.txt` in this folder (assumptions, questions, Nick's keyboard items and the one-line-per-step record live there as sections — the documentation gate refuses a second .md beside this one). Its planning input is `PLANNER-BRIEF-2026-09-12.txt`. Its evidence is the `evidence/` directory beside it.

**NORTH STAR:** Nick, 2026-09-12: "This whole next project is about how to organize these three agents and get them all dialed in 1000000% perfect. That includes all the files folders rules permissions and other shit being consolidated into one clean system and wiping everything else that could ever conflict, confuse or compete." Underneath it: "The whole point is just to make it so that I don't have to be tied to my desk to do all this and type. I'd rather just talk to you guys and action things that way." Finished looks like: Nick talks to Skippy from his phone, Skippy acts, and what Skippy IS lives in one set of documents that the running brain reads.

**FINISH LINE** (each item passes its ONE check, driven by an agent, never by Nick):
- F1 A sentence changed in an assistant's own documents changes that assistant's live answer within 35 minutes, for all three faces (Skippy, Gracie, Neeko), and server.js carries no persona prose and no hand-pasted standing-rules copy.
- F2 Said from the phone, "start a Sonnet agent to <task>" starts a real Claude Code session on the Studio that whats_running lists within 2 minutes, with the agent kind Nick named; the word Cowork appears nowhere in the brain's tools, prompts or queue names.
- F3 Ten ordinary spoken requests are done for real, ten of ten, read back from where each landed — including "add milk to the shopping list" — and no answer ever says it is handing something off without a real dispatch behind it.
- F4 "What's running?" is answered correctly in 15 seconds or less on five of five samples (was 50 s on 2026-09-12).
- F5 A fact told to Skippy by voice and confirmed by voice comes back in a fresh conversation the next day; a phrase shaped like a credential is refused with the floor reason named.
- F6 Six live identity checks pass: Chantelle signed in reaches Gracie on the family app, in Slack and by voice; Nick signed in never does on the same three surfaces. Every message FOR Nick arrives from Skippy, never from Gracie.
- F7 Voice into a running thread: a spoken message from the Status tab lands live in that session, and the session's reply is spoken back within 30 seconds of landing, in a sentence a person wants to hear, not a diff.
- F8 Real conversation with Skippy: the architecture decision is written with measured numbers on Nick's own prompts, Nick has chosen, and the chosen path answers a spoken question with first audio under 2 seconds on ten samples.
- F9 One shape, three faces: the foundation document exists; Skippy's folder and Neeko's folder hold the team assistant's full shape (front door, rulebook, voice guide, manual, plan); Gracie's folder holds her front door and her fence rulebook and nothing more, because she is parked (Nick, 2026-09-12: "gracie dont care"); the Hub worker carries Neeko's name and lives in Neeko's folder; one config-and-secret index and one file index say where every kind of truth lives; Larry's first standing pass is on the upkeep board; no agent deleted anything.
- F10 The routing code enforces exactly the floor and nothing else — logins · credentials, tokens and keys · government IDs · card, bank account and routing numbers (the plan skill's §F list; MACHINE-RULES' word "financials" means those account-shaped values, never a dollar figure — dollar figures, invoices and rates travel) — proven by a test that was red against the old code.

**Owner:** the Fable driver session (this lane's overseer; it never builds) · **Overseer:** ONE — the Fable driver; Codex is the peer builder for the two voice steps and the cold attacker of this plan · **Design authority:** Sienna (creative-director), for the three avatar images only.
**Rule: a step starts the moment its named inputs exist, whatever its number. A step closes on ONE independent check by a different model. Nothing waits on Nick to test.**

> **STEP 0 — ARM THE LOOP, BEFORE ANYTHING ELSE.** Set a 5-minute loop. Every time it fires, answer
> these five in order and CORRECT any failure before doing anything else:
> 1. **NORTH STAR** — is what I am doing this minute moving this plan's North Star? If not, drop it.
> 2. **FAN-OUT** — declare the whole actual roster, dispatch useful ready work, and shed your own unnecessary processes. Coordinate through peers or the launching dispatcher; no numeric cap or load-wait rule applies.
> 3. **CHEAP** — are cheap models doing the building AND the per-step checking? If anything on
>    Anthropic or OpenAI is building or checking a step, move it down now (§M).
> 4. **STUCK** — for anything I have called blocked: name the input that does not exist yet, or the
>    three concrete things I tried. If I cannot, it is not blocked — drive through it now.
> 5. **NEXT** — did something just finish? Then the next step whose inputs exist starts THIS minute.
>    A finished step is never a place to stop, a report is never a reason to wait, and Nick being
>    away or asleep is the reason to keep going, not to pause.
> Then keep building. The loop never stops until the FINISH LINE is proven.

The one exception item 3 allows is the skill's own (§M): the steps Nick assigned to a named model himself — the three Codex steps and the one Opus wall edit, each carrying his dated words — stay where he put them.

## Already true (facts, not story)

- He is Skippy, not Future Nick: `PERSONA_ID = process.env.PERSONA || 'skippy'` at `projects/personal/skippy-app/skippy-code-publish/server.js` line 1135; registry key `id: 'skippy'` at line 2237. Verified live 6/6 by the overnight session of 2026-09-12.
- All three faces read the one grants file: `standingGrantsBlock()` (server.js line 2355) renders `/app/ops/standing-auth.json`, staged into the image at publish by `projects/personal/skippy-app/skippy-code-publish/fly-publish.mjs` line 162 from `projects/ops/skippy-jobs/lib/standing-auth.json`. Present on the box 2026-09-12: `ls /app/ops/standing-auth.json` → exists.
- The cached prompt is no longer thrown away every turn: $0.7682 → $0.0374 a round, cache read 122,665 (commits f0756c0, b8939ba on skippy-code main).
- The microphone opens in 0.3–0.8 s (was 11–16 s); first question after a publish 7 s (was 40 s). Evidence: `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/nick-verbatim/2026-09-12-skippy-must-work.txt`, the overnight section.
- The family and company narrative doors relay to the Mac (commit 33bfa0d on main; `RELAY_URL` at server.js line 1088, allowlist at line 3470) — the phone reaches narrative again.
- Front doors exist: `projects/ops/agents/skippy/AGENT.md` and `projects/ops/agents/gracie/AGENT.md`, both 2026-09-12; `projects/ops/agents/neeko/` holds the complete set (`AGENT.md`, `NEEKO-RULEBOOK.md`, `VOICE-GUIDE.md`, `HUB-MANUAL.md`, `PLAN.md`, `proofs/`).
- The identity firewall is ON in production: `IDENTITY_FIREWALL_ON` (server.js line 496) requires NICK_LOGIN_SECRET, CHANTELLE_LOGIN_SECRET and IDENTITY_SIGNING_SECRET, and all three plus NEEKO_LOGIN_SECRET, GRACIE_BOT_TOKEN, GRACIE_CHANTELLE_DM, GRACIE_NICK_DM, GRACIE_VOICE_ID and GRACIE_AGENT_ID are deployed (fly secrets list, names only, 2026-09-12).
- The voice guide already makes the trip: `projects/personal/skippy-app/skippy-code-publish/lib/voice-guide.mjs` reads `projects/ops/spine-projections/voice-injection.md` from CLAUDE_ROOT (= /data/brain on the box), self-refreshing by stat, returning '' loudly on failure. It is the loader shape STEP 1 copies.
- The trip a document takes to the running brain, measured 2026-09-12: this workspace → `projects/ops/skippy-jobs/jobs/skippy-brain-push.mjs` every 30 minutes (`projects/ops/skippy-jobs/runner.mjs` line 283, `everyMinutes: 30`) into the private brain clone, pushed, then pulled on the box every 5 minutes (`projects/personal/skippy-app/skippy-code-publish/lib/brain-sync.mjs` line 35, `SYNC_INTERVAL_MS = 5 * 60 * 1000`). The push copies what the clone already tracks plus the `ADDS` list at skippy-brain-push.mjs line 129. `/data/brain/projects/ops/` on the box holds only biz-inbox, personal-inbox and spine-projections — none of the three assistant folders. That is why F1 says 35 minutes.
- Codex is reachable from this Mac as a planner and a builder: `codex exec --profile spec -s read-only -m gpt-5.6-luna -C "<workspace>" --skip-git-repo-check --output-last-message <file> "<prompt>" < /dev/null` answered PONG on 2026-09-12.
- `whats_running` (server.js line 2777) works and answered correctly live on 2026-09-12 (nine sessions, one waiting on him) — in 50 s. `push_thread_forward` (line 2790) delivers live into a running session's socket through `projects/ops/skippy-jobs/lib/peer-message.mjs` (`sendToSession`, line 98), proven. `work-watch.mjs` builds the running list from `~/.claude/sessions/` and the process table, writes a per-machine file `work-threads-<machine>.json` under `projects/personal/skippy-app/ala-state/` (its header, lines 29–57), and drops any session whose owning account it cannot resolve from a profile folder (fail-closed, line 65).
- A headless `claude -p` run started from a shell registers no `~/.claude/sessions/<pid>.json` and no socket in `/tmp/cc-socks/` while it runs; it writes only a transcript under `~/.claude/projects/<slug>/`. Evidence: `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/SPIKE-HEADLESS-SESSION-2026-09-12.txt` (first run, a long-lived process) and `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step2-spike-rerun.txt` (second run, pid 47020, with the live per-machine file `work-threads-Nicks-Mac-Studio.json` polled for 150 s: `WORK_THREADS_LIVE: not listed`; its two "appears" lines describe the directories existing, not an entry for the pid — the pairs beneath show no such entry). So the drain of STEP 5 registers every session it starts; work-watch cannot see it otherwise.
- The Claude CLI on this Mac (`/usr/local/bin/claude`, 2.1.237) takes `--model` with the aliases `sonnet`, `opus` and `fable` (its own `--help`, 2026-09-12); Codex accepts `--profile senior-engineer` the same way it accepts `spec` (PONG, 2026-09-12, no [profiles] table needed).
- The engine's write door (`tool_capture_narrative`, `projects/personal/health/engine/brain-routing/personal_mcp.py` line 502) requires `narrative` and `heading`, defaults to `dry_run: true`, and with `dry_run: false` commits to `openbrain_staging` only; production is reached by `publish_accepted_event(staging_id, expected_event_key, expected_content_sha256, accepted_by, db="postgres")` in `projects/personal/health/engine/brain-routing/intake/ingest.py` line 162, which names who accepted the row; `personal_answer` reads production (`_DB = "postgres"`, personal_mcp.py line 70). A confirmed memory therefore takes two calls, and the person's confirmation is the `accepted_by`.
- The two routing-wall copies differ: `projects/personal/skippy-app/lib/grunt-lane.mjs` (the Mac) already splits the old single private class into CLASS_CODE, CLASS_HEALTH (= CLASS_PERSONAL) and CLASS_PRIVATE (lines 96–183) and lets deepseek accept health (line 361), but its note at line 247 still holds the kids and Chantelle in CLASS_PRIVATE; `projects/personal/skippy-app/skippy-code-publish/lib/grunt-lane.mjs` (the cloud) still has the single old class at line 67. STEP 23 carries a target per copy.
- The engine is reachable from this Mac: `personal_answer` answered "REACHABLE — 12 passages visible" on 2026-09-12 through the personal-engine tool this workspace registers, so STEP 11's drain, which runs on the Mac, can reach the same store.
- The brain routes Slack by the verified caller: `const token = caller === 'chantelle' ? GRACIE_BOT_TOKEN : SLACK_USER_TOKEN` at server.js line 9701, with the caller resolved from the signed identity, never from a claim.
- The Status tab's thread answer box already has speech-to-text dictation (voice.js line 21 says it stays there); STEP 15 adds no new control — it adds the spoken reply and one line of status text in existing Pearl classes, and the family app's own fidelity check (`projects/personal/skippy-app/design-directions/_pearl-fidelity-check.mjs`) must still read zero on that tab.
- `handoff_to_cowork` is the only door for new work (server.js line 2829; the rule text at lines 2060, 2068, 2308); the word cowork matches 130 lines of server.js (`command grep -ci cowork`, 2026-09-12). Nothing starts a Claude Code session.
- The shopping list is read-only: `read_shopping_list` (server.js line 4255) over `projects/personal/skippy-app/skippy-code-publish/lib/shopping-list.mjs`, whose own test `projects/personal/skippy-app/skippy-code-publish/_test-shopping-list.mjs` asserts the query never writes. `add_todo` (line 3872, added 2026-09-11) is the writer pattern.
- Memory: `capture_memory` (line 2626) and `confirm_memory` (line 2647) exist; `projects/personal/skippy-app/skippy-code-publish/lib/memory-gate.mjs` writes memory-pending.jsonl and, on a tap, memory-confirmed-facts.jsonl (lines 41–42); `verbalYes()` (line 94) is explicitly not durable; `projects/personal/skippy-app/skippy-code-publish/lib/memory-write-filter.mjs` fails closed on the cloud because the credential scanner (`projects/personal/skippy-app/lib/grunt-egress-scan.mjs`) is deliberately not shipped (fly-publish.mjs line 174). The engine's write door exists: `capture_narrative` in `projects/personal/health/engine/brain-routing/personal_mcp.py` line 502.
- Nick answered the lane's first three questions at 17:57Z on 2026-09-12: "1 yes 2 make it ask 3 VOICE/SKIPPY/ETC due tomorrow" (`projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/NICK-VERBATIM-2026-09-12-three-answers.txt`) — a spoken yes keeps a memory; when he does not name the agent kind, Skippy asks; the lane's board card is `nt-20260912-175831-7392`, due 2026-09-13.
- Sienna's avatar acceptance criteria and the three-voices concept are written: `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/SIENNA-AVATARS-AND-VOICES-2026-09-12.txt` (the brief's S9 concept half is done; only the images remain).
- The clone-mirror gate (`projects/ops/skippy-jobs/lib/check-clone-mirror-gate.mjs`) refuses any NEW file whose path carries the second or third face's name as a token; edits to existing files pass. Nick created the two front doors himself at the keyboard. No mechanism opens it; this plan plans none.
- The three Hub task prompts in `projects/ops/scheduled-rebuild/registration-pack/` (`hub-ops-hourly-audit.prompt.md`, `hub-ops-sop-daily.prompt.md`, `hub-ops-friday-digest.prompt.md`) run as themselves and name neither `hub-ops-manager` nor its LEARNINGS file (`command grep -c hub-ops-manager` → 0 on all three, 2026-09-12); none of the 21 tasks in that pack is registered on this profile. The worker's name lives in `projects/ops/agents/roster.json` line 1512 and its LEARNINGS path string at line 1528; `projects/ops/agents/build_agents.py` writes `.claude/agents/<name>.md` from the roster.
- The exerciser runs on haiku (`ZION/agents/exerciser.md` line 5); Larry runs on sonnet and holds no Write and no Bash (`ZION/agents/larry.md`).
- The floor conflict is real: `projects/personal/skippy-app/skippy-code-publish/lib/grunt-lane.mjs` line 67 defines CLASS_PRIVATE as "health markers, labs, kids, Chantelle, finance"; every cheap vendor accepts CLASS_PUBLIC only (lines 118–182); anthropic-haiku alone accepts private (line 202). The same lines exist in the Mac copy `projects/personal/skippy-app/lib/grunt-lane.mjs`. The rule (`projects/ops/MACHINE-RULES.md`, the data-floor ruling) says the floor is exactly financials · secrets · logins · keys.

## 0 · Gate Zero receipts
- Failure Mode Registry loaded: 2026-09-12, `.claude/skills/plan/references/failure-registry.md`; the entries this build is exposed to are named in §4 (eleven).
- Canonical specs loaded: `projects/ops/agents/CODE-STANDARD.md` (code), `projects/ops/HANDBACK-GATE-SPEC.md` (QA), `projects/ops/agents/CREATIVE-QA-STANDARD.md` and `projects/ops/agents/DESIGN-FIDELITY-STANDARD.md` (read for the avatars; nothing in this lane is a screen).
- Ownership check: no other lane owns "the three assistants consolidated" — the VOICE lane (`projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/PLAN.proposed.txt`) owns the in-app voice screens and their fidelity gates; the HEALTH-ONE-DOOR lane (`projects/ops/life-os/REGROUP-2026-09-08/plans/HEALTH-ONE-DOOR/PLAN.md`) owns the engine and the brains branch; the SCHEDULED lane (`projects/ops/scheduled-rebuild/HANDOFF.md`) owns task registration. Each is named in §1's anti-scope and §5's handoffs.
- Expected inputs confirmed to exist: every path in "Already true" opened 2026-09-12 by the planner (Glob or Read); `projects/ops/cheap-task.mjs` and `projects/ops/route-build.mjs` (the cheap tools); `projects/personal/family-app/_test-voice-requests.mjs` with its `--gate` mode (line 1321); `projects/personal/skippy-app/skippy-code-publish/ops/drive-as-chantelle.mjs` (the Chantelle sign-in instrument); `projects/personal/family-vault/vault.py` (identity secrets, never printed); `projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs` (the evidence re-runner); `projects/ops/skippy-jobs/lib/dispatch-contract.mjs` `loadTravelBlock()` (line 216); `projects/shared-tooling/test-chrome.mjs` (the screenshot tool for Sienna's composites); `projects/ops/skippy-jobs/lib/board-report.mjs` (the board read the card id resolves through).
- PLAN AUTHOR: Boris (senior-engineer, Fable), dispatched by the Fable driver session on 2026-09-12 under Nick's standing ruling of 2026-09-05 ("plans are alwasy writte by 3 fable agents boris for build sienna for design and one cold verifier").
- COLD READER: pending — the driver dispatches Codex to attack this plan and a fresh Fable session to cold-read it; until both have reported, this file is SINGLE-AUTHOR, UNREVIEWED and says so here.
- PROMPT-SPEC scan (P1–P7): the three V1 variables are settled on the sheet (§1a rows 2–4) — two by Nick's line of 17:57Z, one by his "make it one thing"; "everything else retired" is bounded in §1 NOT in scope; the memory destination and the document trip (V2) are settled by opening the code (§1a rows 5–6); assumptions are logged under ASSUMPTIONS in `PROGRESS.txt`.

## 1 · Goal and definition of done
- **What we're building, one paragraph.** One assistant wearing three faces — Skippy for Nick, Gracie for Chantelle, Neeko for the team — defined once in documents the running brain reads, with a real door that starts Claude Code work from a spoken sentence, a memory that keeps what Nick confirms by voice, a sender rule that puts Skippy on everything for Nick, two voice projects (talk to a running thread; a real conversation with Skippy), the routing code brought in line with the four-category floor, and the one-file-per-thing cleanup run through Larry, who proposes and deletes nothing.
- **HOW IT'S USED:** Nick opens the family app on his phone or Mac, taps the mic, and says an ordinary thing — add milk, start a Sonnet agent on X, what's running, remember this — and it is done, read back, or started; Chantelle does the same and meets Gracie; the team meets Neeko in the Hub. · HOW WE KNOW: the ten-request gate (STEP 7), the live door drive (STEP 6), the live memory drive (STEP 12), the six identity checks (STEP 14).
- **WHAT IT LOOKS LIKE:** the existing Talk and Status tabs, Slack DMs and the Hub Talk panel, unchanged in look (Pearl stays; the Hub's styling is frozen for its own redo); a mic on a Status-tab thread that speaks the session's reply back; three distinct faces where a message is attributed. · HOW WE KNOW: §2 rows and the VOICE and HUB lanes' fidelity gates, which this lane never touches.
- **WHERE IT LIVES:** the cloud brain (`projects/personal/skippy-app/skippy-code-publish/`, published to Fly), the family app (deck-family), Slack, the Hub, and the documents under `projects/ops/agents/skippy/`, `projects/ops/agents/gracie/`, `projects/ops/agents/neeko/` — opened by Nick (Talk, Status, Slack), Chantelle (family app, Slack) and the team (Hub). · HOW WE KNOW: the surface row of §1a, Nick's words of 2026-09-12.
- **WHAT IT MUST DO:** F1–F10 above, one check each in §6.
- **NOT in scope:** (a) security and privacy work of any kind (Nick, 2026-09-09) — one line to `projects/ops/sp-sec/PLAN.md` and back to the step; (b) the engine's answer fabrication and its routing — the HEALTH-ONE-DOOR lane owns it (Nick, 2026-09-12: "leave the health stuff out of your stuff"); (c) Gracie's full document set beyond her fence — parked (Nick, 2026-09-12: "gracie dont care"), and her voice guide is Chantelle's call, evidenced from Chantelle's own words, never an agent's (Sienna, 2026-09-12); (d) removing any tool from Skippy or Gracie, Hub tools included — the overlap is the ruling (Nick, 2026-09-12: "of course"); (e) buying or dedicating another Claude account — the door, not the quota, was the problem; (f) payments prepared with one-tap authorisation — the NEXT list, owner the Fable driver, planned after F1–F10 close, because money leaving is an approval class and no payment executor exists on disk; (g) deleting, archiving or moving any file by an agent — Larry proposes, Nick decides; (h) opening or weakening the clone-mirror gate or the documentation gate by any mechanism; (i) tidying the dispatchable-subagent convention (`.claude/agents/`, `ZION/agents/`) and the product-assistant convention (`projects/ops/agents/<name>/`) into one folder — the distinction is deliberate; (j) the Hub's colours and styling — frozen until its own redo the week of 2026-09-15; (k) the brains/2026-09-09 branch of skippy-code — the HEALTH-ONE-DOOR lane merges it; (l) distinct spoken voice ids per face — "distinct voicing later" (Nick, 2026-09-12); the concept is written and the ids stay.
- **Trip-over protocol:** a lane that finds something outside the fence writes one handover line to its named owner (a security- or privacy-shaped thing: one line in `projects/ops/sp-sec/PLAN.md`), then back to building — never investigates, never fixes.

## 1a · Critical variables — the confirmation sheet is GENERATED from this table

| # | The variable, in plain words | Value chosen | Alternatives rejected | Class | HOW WE KNOW | Cost if wrong | CONFIRMED |
|---|---|---|---|---|---|---|---|
| 1 | **SURFACE — which screens this lands on, and who opens them** | the family app's Talk and Status tabs on Nick's phone and Mac window, Slack DMs, and the Hub's Talk panel; Nick opens the first two, Chantelle the family app and Slack, the team the Hub | a new app · a phone call · the Cowork queue as the surface | V1 | Nick's own words on the division of faces, brain dump §12 | the wrong face answers the wrong person on the wrong screen | Nick, 2026-09-12, "Chantelle logs in, she gets Gracie. Slack, she gets Gracie… anything and everything in the hub is Neeko" |
| 2 | **A spoken yes is the decision for a memory — never for the four approval classes** | yes — a clear yes from the person, spoken or typed, directed at a named pending fact, keeps it; the tap stays as a second door | keep tap-only | V1 | put to him as item 1 of the lane's first message; his answer is on file | he confirms by voice and nothing is kept | Nick, 2026-09-12, "1 yes" |
| 3 | **What Skippy does when he does not name the agent kind in "start an agent to…"** | Skippy asks which kind — Sonnet, Opus, Fable or Codex — before dispatching; there is no silent default | default to Sonnet · default to Opus | V1 | put to him as item 2 of the lane's first message; his answer is on file | a dispatch starts on a kind he did not choose, or he is never asked and it never starts | Nick, 2026-09-12, "2 make it ask" |
| 4 | **The Hub worker's name once it carries Neeko's name** | `neeko-hub-worker` — the dispatchable file in `.claude/agents/`, generated from the roster; its memory file moves under the neeko folder | keep hub-ops-manager and only mention Neeko in prose | V1 | one assistant, one name, however many faces; a search for Neeko must find the worker; the exact spelling is the planner's and he may rename it in one line | a search for Neeko still misses the worker | Nick, 2026-09-12, "neeko is hub ops manager so make it one thing" |
| 5 | **Where a confirmed memory lands** | the family narrative engine's own write door (`capture_narrative`); a dose, marker or lab is refused by the marker filter and belongs to the engine lane, never to this lane | a flat file in the app folder · the prose store directly | V2 | the engine is the store Skippy reads from; a flat file is the store nothing reads | he confirms a fact and it never comes back | opened `projects/personal/health/engine/brain-routing/personal_mcp.py`, 2026-09-12, saw: `capture_narrative` is a real write tool (line 502); opened memory-gate.mjs, saw: confirmed rows land in memory-confirmed-facts.jsonl and nothing reads them back |
| 6 | **How a document reaches the running brain** | the existing push (30 min) and pull (5 min) with the three assistant folders added to the push's ADDS list; a loader in the voice-guide shape | render the documents into the image at publish · a new sync | V2 | the voice guide already makes this exact trip | a document edit never reaches the phone | opened `projects/ops/skippy-jobs/jobs/skippy-brain-push.mjs` line 129 and `projects/personal/skippy-app/skippy-code-publish/lib/brain-sync.mjs` line 35, 2026-09-12, saw: the ADDS list and the 5-minute pull |

**Considered and ruled NOT critical:** which cheap vendor builds each step (the matrix decides); the spoken-sentence style for a session's reply (the whats_running rundown already sets it); the exact wording of the three bands (Nick's words of 2026-09-12 in brain dump §8 are pasted, not paraphrased); the avatars' look (Sienna's criteria file decides, once).

## 1b · Subproject decomposition — could a piece of this ship on its own?

- **SINGLE SUBPROJECT:** every step serves the one North Star (Nick talks, Skippy acts, and what Skippy is lives in documents the brain reads); the two Codex-built voice steps are steps inside this plan with their own fences, not lanes with their own plan, and a second plan file is what the documentation gate and the shape rule both refuse.

**Carve-out rule:** anything left out of every step's scope is named with a real owner in §1 NOT in scope (the engine lane, the VOICE lane, the HUB lane, the Fable driver for payments, Chantelle for Gracie's voice, Nick for deletions).

## 2 · The complete UX map (this becomes the test manifest verbatim)

| Id | Screen / entry point | State (default·empty·error·loading) | Element / interaction | Expected behavior | Navigation from → to |
|---|---|---|---|---|---|
| U1 | family app, Talk tab, Nick signed in | default | mic tap, "who are you" | answers as Skippy, in Skippy's voice; never Gracie | Talk → Talk |
| U2 | family app, Talk tab, Chantelle signed in | default | mic tap, "who are you" | answers as Gracie, in Gracie's voice (GRACIE_VOICE_ID); never Skippy | Talk → Talk |
| U3 | family app, Talk tab, Nick | default | "add milk to the shopping list" | the item is on the Monday board within the turn; the answer says it was added, never that it was handed off | Talk → Talk |
| U4 | family app, Talk tab, Nick | default | "start a Sonnet agent to <task>" | a real headless session starts on the Studio within 2 minutes; the answer names the kind; the Dispatch screen shows the row | Talk → Status (the new thread appears) |
| U5 | family app, Talk tab, Nick | default | "start an agent to <task>" (no kind named) | Skippy asks which kind — Sonnet, Opus, Fable or Codex — and starts nothing until he answers | Talk → Talk |
| U6 | family app, Talk tab, Nick | default | "what's running?" | the count, what each is doing, which needs him, what just landed — in 15 s or less | Talk → Talk |
| U7 | family app, Talk tab, Nick | default | "remember that <fact>" → "yes" | Skippy proposes in one sentence; the yes keeps it; a fresh conversation next day answers with it | Talk → Talk |
| U8 | family app, Talk tab, Nick | error | "remember my password is …" (a credential-shaped phrase) | refused, naming the floor category; nothing stored anywhere | Talk → Talk |
| U9 | family app, Talk tab, any | error (documents unreadable on the box) | any question | the assistant still answers, in its previous voice, and the box log carries one loud line per absent document | Talk → Talk |
| U10 | family app, Status tab, a running thread | default | the thread's mic → speak a message | the words land live in that session (delivery log says live); the session's newest reply is spoken back within 30 s of landing as one plain sentence — no paths, no diffs | Status → Status |
| U11 | family app, Status tab, a running thread | loading (the session has not answered yet) | after speaking | no phone-call turn-taking; the tab says the message landed and the reply will be read out when it arrives | Status → Status |
| U12 | family app, Dispatch screen | default | a row started from U4 | shows kind, brief, started time and the session it became; the word Cowork appears nowhere | Status → Dispatch |
| U13 | Slack, Nick's DM from Skippy | default | any message FOR Nick from any job (including alerts about Gracie's pipeline) | arrives from the Skippy bot user, never from the Gracie bot | — |
| U14 | Slack, Chantelle's DM from Gracie | default | Chantelle's message to the Gracie bot | the reply comes from the Gracie bot user | — |
| U15 | Slack, the three bots' profiles | default | profile image | three distinct faces, one per assistant, siblings at 32 px | — |
| U16 | Hub, Talk panel, a team member | default | any question | answers as Neeko; nothing personal to the household ever attached | Hub → Hub |
| U17 | any surface, the wrong assistant asked | default | Chantelle asks Skippy something harmless and in scope | answered gracefully, or handed over by name; never a blank refusal, never a silent answer outside its fence | — |
| U18 | family app, Talk tab, Nick | default | a spoken question after STEP 17 | first audio under 2 seconds; barge-in works; the Pearl look unchanged | Talk → Talk |

DESIGN FIDELITY GATE: N/A — nothing rendered by this lane. The only new visual artefacts are the three avatar images, graded once by Sienna (creative-director) against her own criteria in `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/SIENNA-AVATARS-AND-VOICES-2026-09-12.txt` — gates 1, 2 and 6 of the creative standard, with composites at 512 px, 32 px and the four family-app slot sizes screenshotted, and a cold matcher; no anchor map and no fidelity count, because an image is not a screen. The Talk, Status and Dispatch screens belong to the VOICE lane and the Hub screens to the HUB lane, and each keeps its own locked target and gate, which STEP 15 and STEP 19 must leave at zero (their handoff lines say so).

## 3 · Lanes and frozen contracts

| Lane | Scope (in / out) | Owner | Definition of done | Builder (cheap, named) | Backup builder | Checker (different model) | Backup checker |
|---|---|---|---|---|---|---|---|
| BRAIN | steps 1, 22, 25 — the documents reach the brain; the cache guard / out: the documents' content beyond the drafts | the Fable driver | F1 | GLM 5.3 (zai) | DeepSeek V4 Pro | Sonnet (live), Qwen 3.8 (unit) | Qwen 3.8 / Sonnet |
| DOOR | steps 2–8 — the Claude Code door, Cowork out, ten requests, what's running / out: the family app's screens | the Fable driver | F2, F3, F4 | GLM 5.3 (zai); DeepSeek V4 Pro for the drain; exerciser (haiku) for live drives | DeepSeek V4 Pro / Qwen 3.8 | Qwen 3.8 (unit), Sonnet (live) | Sonnet / Opus |
| MEMORY | steps 9–12 / out: the engine's marker table (HEALTH-ONE-DOOR) | the Fable driver | F5 | GLM 5.3 (zai); exerciser (haiku) for the live drive | DeepSeek V4 Pro | Qwen 3.8 (unit), Sonnet (live) | Sonnet / Opus |
| FACES | steps 13, 14, 18, 19 — sender rule, identity checks, avatars / out: the Hub's styling | the Fable driver | F6, F9 (faces) | DeepSeek V4 Pro; exerciser (haiku) for live drives | GLM 5.3 (zai) / Qwen 3.8 | Qwen 3.8 (unit), Sonnet (live), Fable as Sienna for the images once | Sonnet / Opus |
| VOICE-CODEX | steps 15, 16, 17 — Voice B and Voice A / out: voice.js's Talk-tab behaviour the VOICE lane owns | the Fable driver; Codex builds | F7, F8 | gpt-5.6-terra | gpt-5.6-luna | Sonnet (live) | Opus |
| CLEAN | steps 20, 21, 23, 24 — drafts, indexes, the Neeko merge, the floor wall, Larry / out: any deletion | the Fable driver | F9, F10 | DeepSeek V4 Pro, GLM 5.3 (zai); Opus for the wall; Larry (sonnet) for the pass | Qwen 3.8 / Fable for the wall | GLM 5.3 (zai) or Qwen 3.8; Sonnet for the wall | Sonnet / Qwen 3.8 |
| SIGN-OFF | step 26 | the Fable driver | F1–F10 read once | Fable | Opus | Opus | gpt-6-astra |

**Contracts between lanes (FROZEN at plan time — change = a dated PLAN CHANGES line in `PROGRESS.txt` until the gate allows PLAN-CHANGES.md):**
1. **The `dispatch_to_code_agent` tool shape** (replaces `handoff_to_cowork`): input `{ kind: 'sonnet' | 'opus' | 'fable' | 'codex', mode: 'new' | 'coordinate', brief: string, topic?: string }`. `kind` is REQUIRED — when the person did not name one, the tool is not called: the prompt tells Skippy to ask which kind (Sonnet, Opus, Fable or Codex) and wait for the answer (§1a row 3), and a call without `kind` returns `{ ok: false, ask: 'which kind — Sonnet, Opus, Fable or Codex?' }` and starts nothing. `mode: 'coordinate'` names a running thread by `topic`, which must be the exact thread `id` the newest `whats_running` result gave (the model reads it out of that result, never types it from memory); the tool matches on that id only — an absent `topic`, an id that matches no thread, or a name that matches more than one returns `{ ok: false, ask: 'which thread — <the candidate threads by name, from whats_running>' }` and sends nothing; a single exact match pushes `brief` into it through `sendToSession`; it applies only to sessions that hold a socket (interactive ones) — a headless dispatched session is fire-and-finish and its result comes back on its queue row. The kinds map to commands the drain runs with `execFile` and an argument array, never a shell string: sonnet → `claude --model sonnet`, opus → `claude --model opus`, fable → `claude --model fable` (the CLI's own aliases, verified 2026-09-12), codex → `codex exec --profile senior-engineer -m gpt-5.6-terra` (verified 2026-09-12); the brief travels as one argv element and is also written to `<worktree>/BRIEF.txt`, so no spoken text is ever interpolated into shell syntax. Output `{ ok: boolean, dispatch_id: string, kind, mode, started_at: ISO string, visible_in: 'whats_running', error?: string }`. On the cloud the tool relays to the Mac program at the relay route `dispatch_to_code_agent` (added to `RELAY_ALLOWED_TOOLS`, server.js line 3470); on the Mac it appends one row `{ id, kind, mode, brief, topic, requested_at, requested_by, status: 'queued' }` to the code-agent queue file under `projects/personal/skippy-app/ala-state/` (the drain of STEP 5 sets `status`, `session_pid`, `started_at`, `worktree`, `transcript`, `result`). The tool never claims a session exists before the drain has written `session_pid`.
2. **The assistant-docs loader contract** (`lib/assistant-docs.mjs`, in the voice-guide shape): `getAssistantDocs(face)` for `face ∈ {'skippy','gracie','neeko'}` returns `{ agent: string, rulebook: string, voice: string, manual: string, foundation: string, source: string, loaded_at: ISO string }`; it reads `AGENT.md`, the face's rulebook, voice guide and manual by the names that face's folder under `projects/ops/agents/` uses, and the shared foundation file ASSISTANT-FOUNDATION.md under projects/ops/agents/, all from CLAUDE_ROOT; it re-stats every file per request and re-reads on change; every absent or unreadable file yields `''` for that field and one loud log line per boot per file (`assistant-docs: <face>/<file> absent`); it never throws into a reply. The three prompt builders compose `docs.agent + docs.foundation + docs.rulebook + <tools guide> + <grants block>`; the voice block keeps coming from voice-guide.mjs.

**Buckets that share a goal message each other:** a dated line into `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/PROGRESS.txt` when STEP 6, STEP 7, STEP 8, STEP 15 or STEP 17 closes (they change the brain the voice app talks to or add a control to a tab); a dated line into `projects/ops/life-os/REGROUP-2026-09-08/plans/HEALTH-ONE-DOOR/PROGRESS.txt` when STEP 11 closes (the marker-shaped facts it refuses are theirs); a dated line into `projects/ops/scheduled-rebuild/HANDOFF.md` when STEP 21 closes (the worker's new name, for registration).

## 3b · Execution map — the Step map, then one STEP block per row

A task is DONE only when its review-ledger row is CLOSED by a reviewer that is not the builder.

Files this plan creates, named here once so every proof may cite them: `projects/personal/skippy-app/skippy-code-publish/lib/assistant-docs.mjs` and `projects/personal/skippy-app/skippy-code-publish/_test-assistant-docs.mjs` CREATED BY STEP 1 · `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step2-spike-rerun.txt` CREATED BY STEP 2 · `projects/personal/skippy-app/skippy-code-publish/lib/code-agent-dispatch.mjs` and `projects/personal/skippy-app/skippy-code-publish/_test-code-agent-dispatch.mjs` CREATED BY STEP 3 · `projects/ops/skippy-jobs/jobs/code-agent-drain.mjs` and `projects/ops/skippy-jobs/_test-code-agent-drain.mjs` CREATED BY STEP 5 · `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step6-live-door.txt` CREATED BY STEP 6 · `projects/personal/skippy-app/skippy-code-publish/_test-claim-guard-handoff.mjs` and `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/fresh-five.json` CREATED BY STEP 7 · `projects/personal/skippy-app/skippy-code-publish/ops/time-whats-running.mjs` CREATED BY STEP 8 · `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step10-red.txt` CREATED BY STEP 10 · `projects/ops/skippy-jobs/jobs/memory-confirmed-drain.mjs` and `projects/ops/skippy-jobs/_test-memory-confirmed-drain.mjs` CREATED BY STEP 11 · `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step12-live-memory.txt` CREATED BY STEP 12 · `projects/ops/skippy-jobs/_test-sender-rule.mjs` CREATED BY STEP 13 · `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step14-identity.txt` and `projects/personal/skippy-app/skippy-code-publish/_test-slack-caller-routing.mjs` CREATED BY STEP 14 · `projects/personal/family-app/_test-thread-voice.mjs` CREATED BY STEP 15 · `projects/personal/family-app/_bakeoff-voice-a.mjs` and `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/voice-a-decision.txt` CREATED BY STEP 16 · `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/avatars-grade.txt` and `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/avatars-composite.html` CREATED BY STEP 18 · `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step19-slack-icons.txt` CREATED BY STEP 19 · the seven drafts `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/ASSISTANT-FOUNDATION.txt`, `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/SKIPPY-RULEBOOK.txt`, `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/SKIPPY-MANUAL.txt`, `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/GRACIE-RULEBOOK.txt`, `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/CONFIG-AND-SECRET-INDEX.txt`, `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/FILE-INDEX.txt`, `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/LEGACY-PROMPT-INVENTORY.txt` CREATED BY STEP 20 · `projects/personal/skippy-app/skippy-code-publish/_test-floor-four-categories.mjs` and `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step23-red.txt` CREATED BY STEP 23 · `projects/personal/skippy-app/skippy-code-publish/ops/test-cache-guard.mjs` CREATED BY STEP 25.

Files that Nick creates at the keyboard or that the documentation gate's expiry allows, never an agent: RULEBOOK.md and MANUAL.md under projects/ops/agents/skippy/; RULEBOOK.md under projects/ops/agents/gracie/; LEARNINGS.md under projects/ops/agents/neeko/; neeko-hub-worker.md under .claude/agents/; ASSISTANT-FOUNDATION.md under projects/ops/agents/; the two index files under projects/ops/ (CONFIG-AND-SECRET-INDEX.md and FILE-INDEX.md, location for Nick to approve). A step that needs one of them names it in its Start when line as a directory listing, never as a path.

**Step map (read this first):**

| Stage | # | Task (step name) | Needs (named artefact, or `none — start now`) | EXECUTOR (cheap model) | EXECUTOR BACKUP | CHECKER (different model) | CHECKER BACKUP | DONE-PROOF (runnable command) |
|---|---|---|---|---|---|---|---|---|
| FRONT | 1 | S1 · the brain reads the documents | none — start now | zai | deepseek | sonnet | qwen | `node projects/personal/skippy-app/skippy-code-publish/_test-assistant-docs.mjs` → ALL PASS, and the live marker read back as all three faces |
| FRONT | 2 | S2a · spike re-run: does a job-started headless session reach the live running list | none — start now | exerciser (haiku) | deepseek | qwen | sonnet | `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step2-spike-rerun.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH |
| FRONT | 3 | S2b · the door: dispatch_to_code_agent, the three bands, the relay route | none — start now | zai | deepseek | qwen | sonnet | `node projects/personal/skippy-app/skippy-code-publish/_test-code-agent-dispatch.mjs` → ALL PASS |
| FRONT | 4 | S2b · Cowork out, with its replacement in | STEP 3 closed | zai | deepseek | qwen | sonnet | `command grep -rci "cowork" projects/personal/skippy-app/skippy-code-publish/server.js projects/personal/skippy-app/skippy-code-publish/lib projects/ops/skippy-jobs/jobs/skippy-brain-push.mjs` → every line ends `:0` |
| FRONT | 5 | S2c · the Mac drain that starts the session, and work-watch lists it | the spike re-run's evidence file (STEP 2) and the queue row shape (STEP 3) | deepseek | zai | qwen | sonnet | `node projects/ops/skippy-jobs/_test-code-agent-drain.mjs` → ALL PASS |
| FRONT | 6 | S2d · live from the phone path | STEP 4 and STEP 5 closed and published | exerciser (haiku) | deepseek | sonnet | opus | `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step6-live-door.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH |
| FRONT | 7 | S3 · ten requests done for real | STEP 4 closed (the false hand-off claim has no tool to hide behind) | zai | deepseek | sonnet | opus | `node projects/personal/family-app/_test-voice-requests.mjs --gate --fresh projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/fresh-five.json` → exit 0, ten delivered |
| FRONT | 8 | S4 · what's running in 15 seconds | none — start now | zai | deepseek | sonnet | qwen | `node projects/personal/skippy-app/skippy-code-publish/ops/time-whats-running.mjs` → `5/5 under 15000 ms` |
| FRONT | 9 | S5a · memory proposals relay to the Mac's scanner | none — start now | zai | deepseek | qwen | sonnet | `node projects/personal/skippy-app/skippy-code-publish/_test-memory-write-filter.mjs` → ALL PASS and `node projects/personal/skippy-app/skippy-code-publish/_test-relay-port-scope.mjs` → green |
| FRONT | 10 | S5b · a yes from the person is the decision | none — start now | zai | deepseek | qwen | sonnet | `node projects/personal/skippy-app/skippy-code-publish/_test-memory-gate.mjs` → ALL PASS, red run on file |
| FRONT | 11 | S5c · confirmed facts reach the engine door | STEP 10 closed | zai | deepseek | qwen | sonnet | `node projects/ops/skippy-jobs/_test-memory-confirmed-drain.mjs` → ALL PASS and `node --check projects/ops/skippy-jobs/runner.mjs` → exit 0 |
| FRONT | 12 | S5d · live memory proof | STEPS 9, 10, 11 closed and published | exerciser (haiku) | deepseek | sonnet | opus | `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step12-live-memory.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH |
| FRONT | 13 | S6a · every message for Nick comes from Skippy | none — start now | deepseek | zai | qwen | sonnet | `node projects/ops/skippy-jobs/_test-sender-rule.mjs` → `0 send sites to Nick via the Gracie token`, exit 0 |
| FRONT | 14 | S6b · six live identity checks | STEP 13 closed | exerciser (haiku) | deepseek | sonnet | opus | `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step14-identity.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH, six PASS lines |
| FRONT | 15 | S7 · Voice B — talk to a running thread, be told when it answers | none — start now | gpt-5.6-terra | gpt-5.6-luna | sonnet | opus | `node projects/personal/family-app/_test-thread-voice.mjs` → `spoken in <n> s`, n ≤ 30, exit 0 |
| FRONT | 16 | S8a · Voice A — the architecture decision, measured | none — start now | gpt-5.6-terra | gpt-5.6-luna | sonnet | opus | `node projects/personal/family-app/_bakeoff-voice-a.mjs --report` → both candidates at 10 samples |
| FRONT | 17 | S8b · Voice A — the chosen path built | Nick's one-line choice on file in `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/PROGRESS.txt` | gpt-5.6-terra | gpt-5.6-luna | sonnet | opus | `node projects/personal/family-app/_bakeoff-voice-a.mjs --live` → `chosen path: 10/10 first audio < 2.0 s` |
| FRONT | 18 | S9a · three faces drawn, composited and graded | none — start now (Sienna's criteria file exists) | deepseek | qwen | fable | opus | `ls projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/avatars/ | command grep -c png` → 3 masters plus the composites, and three PASS lines in the grade file |
| FRONT | 19 | S9b · three faces live in Slack | STEP 18 closed (three graded masters in the avatars directory) | fable (the driver, in Nick's own signed-in Chrome — NICK-ASKED, see the block) | opus (the driver's takeover model; Nick at the keyboard only if the settings page refuses the browser) | sonnet | opus | `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step19-slack-icons.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH, three distinct image hashes |
| POLISH | 20 | S10a · the drafts, the two indexes and the legacy-prompt inventory, as text | none — start now | deepseek | qwen | zai | sonnet | `test $(ls projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/ | command grep -c txt) -eq 7 && echo PASS` → `PASS` |
| POLISH | 21 | S10b · the Neeko merge | the two gated empty files exist and no session is inside the worker's files (STEP 21's first two commands) | zai | deepseek | qwen | sonnet | `command grep -rn "hub-ops-manager" projects/ops/agents/roster.json .claude/agents projects/ops/scheduled-rebuild/registration-pack ZION/agents | command grep -vc "SUPERSEDED BY"` → 0 |
| POLISH | 22 | S10c · the births, and the last hard-coded copies leave | the gated files exist (STEP 22's first command) and STEP 20 closed | zai | deepseek | sonnet | qwen | `command grep -c "STANDING_RULES_" projects/personal/skippy-app/skippy-code-publish/server.js` → 0 and `node projects/personal/skippy-app/skippy-code-publish/_test-assistant-docs.mjs` → ALL PASS |
| POLISH | 23 | S11 · the floor is four categories in the routing code | none — start now | opus | fable | sonnet | qwen | `node projects/personal/skippy-app/skippy-code-publish/_test-floor-four-categories.mjs` → ALL PASS on both copies, red run on file |
| POLISH | 24 | S12 · Larry's first standing pass | the documentation gate off, or a ticket for the upkeep board (STEP 24's first command) | larry (sonnet) | opus | qwen | deepseek | `command grep -c "^THING:" projects/ops/agents/upkeep-board/larry-lane.md` ≥ 5 with today's date on the pass header |
| POLISH | 25 | S13 · the cache guard after every publish | none — start now | zai | deepseek | qwen | sonnet | `node projects/personal/skippy-app/skippy-code-publish/ops/test-cache-guard.mjs` → `PASS`, and `--expect-fail` on the fixture → exit 1 |
| FINAL | 26 | S14 · postmortem and FINISH LINE sign-off | all 25 closed | fable | opus | opus | gpt-6-astra | `command grep -c "CLOSED" projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/PROGRESS.txt` ≥ 25 |

**Then one block per step, in this exact shape.** Every cheap build below is one job per file through `projects/ops/cheap-task.mjs` (`--provider zai` or `--provider deepseek`, `--review t2`, `--dir` the fence, `--do` the numbered items of the step, `--prove` the step's PROOF plus what must NOT change) or one edit to one existing file through `projects/ops/route-build.mjs --file <path> --do "<change>" --prove "<command>"`. Every brief carries the travel block (`loadTravelBlock()` from `projects/ops/skippy-jobs/lib/dispatch-contract.mjs`), the four approval classes and the floor. Every brain change is one scoped commit in the nested repo (`git -C projects/personal/skippy-app/skippy-code-publish add <files> && git -C projects/personal/skippy-app/skippy-code-publish commit -m "STEP <n>: <what>"`) on main, published with `cd projects/personal/skippy-app/skippy-code-publish && node fly-publish.mjs` (it refuses a dirty or behind tree). Never the brains/2026-09-09 branch.

### STEP 1 — S1 · the brain reads the documents
**FOR NICK:** a sentence he changes in an assistant's own file is what that assistant says on his phone within 35 minutes — Neeko today, Skippy and Gracie fully once their gated files exist (STEP 22). · **Tier:** FRONT
**Start when:** none — start now.
**Builder:** zai · **Builder backup:** deepseek · **Checker:** sonnet (the live half; the exerciser gathers the box evidence) · **Checker backup:** qwen
**Files you may touch:** `projects/ops/skippy-jobs/jobs/skippy-brain-push.mjs` (the ADDS array only); new `projects/personal/skippy-app/skippy-code-publish/lib/assistant-docs.mjs`; new `projects/personal/skippy-app/skippy-code-publish/_test-assistant-docs.mjs`; `projects/personal/skippy-app/skippy-code-publish/server.js` — the three registry entries (lines 2237–2260), the three persona builder functions the registry points at (`buildFutureNickPrompt` — still its historical name — for skippy at registry line 2242, `buildGraciePrompt`, `buildNeekoPrompt`; the persona PROSE lives inside those functions, not in the registry), STANDING_RULES_NEEKO, and the three prompt composers getSystemPrompt / getGracieSystemPrompt / getNeekoSystemPrompt (lines 2429–2457); `projects/ops/agents/neeko/AGENT.md`, `projects/ops/agents/skippy/AGENT.md`, `projects/ops/agents/gracie/AGENT.md` (one dated marker sentence each, for the proof). **Never** `lib/voice-guide.mjs`, `standingGrantsBlock()`, STANDING_RULES_GRACIE and the skippy and gracie builders' inline prose (STEP 22 owns their deletion, after STEP 20's inventory maps every paragraph), the brains branch (HEALTH-ONE-DOOR lane).

**Do exactly this:**
1. Job A (new file): write `lib/assistant-docs.mjs` to the frozen contract 2 in §3, copying the root-validation and stat-refresh shape of `lib/voice-guide.mjs` (lines 31–60: candidates proven by content, never a depth guess). Write `_test-assistant-docs.mjs` beside the other `_test-*.mjs` files: it points CLAUDE_ROOT at a fixture folder, asserts each field loads, asserts a missing file yields `''` and one log line, asserts a changed file is re-read without restart, and reads server.js as text to assert the three prompt builders call `getAssistantDocs` and that no registry entry's prompt body is longer than 200 characters of inline prose once STEP 22 has run (this last assertion prints the word SKIPPED until the foundation file exists under projects/ops/agents/, so it can never pass by accident).
2. Job B (one edit, route-build): in server.js make `getNeekoSystemPrompt()` compose `docs.agent + docs.foundation + docs.rulebook + docs.voice + NEEKO_TOOLS_GUIDE + standingGrantsBlock('Neeko')` from `getAssistantDocs('neeko')`; make `buildNeekoPrompt()` return only that composition (its inline persona prose deleted — first confirm by `command grep -n "buildNeekoPrompt" server.js` that every caller goes through it, and that no other prompt path reaches the deleted prose); delete the `STANDING_RULES_NEEKO` string; make `getSystemPrompt()` and `getGracieSystemPrompt()` prepend `docs.agent + docs.foundation + docs.rulebook` for their face while leaving the skippy and gracie builders' inline prose and STANDING_RULES_GRACIE in place (STEP 22 removes them the day the documents exist and STEP 20's inventory says where every paragraph went). The test asserts, by calling `buildNeekoPrompt()` with a fixture CLAUDE_ROOT, that its output minus the loaded documents and the tools guide and the grants block is under 200 characters — the same assertion for the other two builders prints SKIPPED until the foundation file exists.
3. Job C (one edit, route-build): add to the ADDS array at skippy-brain-push.mjs line 129 the three folders `projects/ops/agents/skippy/`, `projects/ops/agents/gracie/`, `projects/ops/agents/neeko/` and the foundation file's name (ASSISTANT-FOUNDATION.md under projects/ops/agents/), after first reading the sweep code below the array to confirm an entry that does not exist yet is skipped, never fatal (record the line that proves it in the commit message).
4. Commit, publish, then the proof: put `Marker 2026-09-<dd>-<hhmm>: the brain reads this file.` as a dated sentence in each of the three AGENT.md files, run `node projects/ops/skippy-jobs/jobs/skippy-brain-push.mjs` by hand, wait for the box's pull (≤ 5 minutes), and have the exerciser sign in as Nick and as the team login and ask each face "repeat the marker sentence in your own file" through `/api/chat`, saving COMMAND:/OUTPUT: pairs to the evidence folder.

**DEFINITION OF DONE:** the loader is live for all three faces, Neeko's hard-coded copies are gone, and a marker sentence added to each AGENT.md is read back by the live brain within 35 minutes.
**PROOF:** `node projects/personal/skippy-app/skippy-code-publish/_test-assistant-docs.mjs` → `ALL PASS` (exit 0), `command grep -c "STANDING_RULES_NEEKO" projects/personal/skippy-app/skippy-code-publish/server.js` → 0, and the exerciser's evidence shows the marker in all three live answers · **FAILS IF:** any assertion fails, the grep count is above 0, or any face's live answer lacks the marker after 35 minutes.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once (the unit test on this Mac; the live read-back through the exerciser's evidence re-run by `verify-agent-evidence.mjs`). PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 2 — S2a · spike re-run: does a job-started headless session reach the live running list
**FOR NICK:** nothing he notices yet — this settles the one question the first spike could not: whether work-watch needs a new source to show him a started agent. · **Tier:** FRONT (it gates U4)
**Start when:** none — start now. The first run is on file (`projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/SPIKE-HEADLESS-SESSION-2026-09-12.txt`): no session file, no socket, a transcript only; its fourth question is NOT MEASURABLE because the pid capture returned 0 and it polled the stale shared `work-threads.json` of 3 rows instead of the live per-machine file.
**Builder:** exerciser (haiku) · **Builder backup:** deepseek · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step2-spike-rerun.txt`; a throwaway worktree under the machine's temporary area. **Never** any file in the workspace tree; never a session Nick is running.

**Do exactly this:**
1. `git worktree add /private/tmp/assistants-spike-wt HEAD` (removed at the end with `git worktree remove --force /private/tmp/assistants-spike-wt`).
2. In ONE shell, not a subshell: `cd /private/tmp/assistants-spike-wt && nohup /usr/local/bin/claude -p "<the dispatch header: ROLE: BUILDER · REVIEW: t2 · RETURN-SIZE: ~50 tokens · the travel block from loadTravelBlock()> Reply PONG, then run the shell command sleep 120, then reply done." --max-turns 3 > /private/tmp/assistants-spike-wt/out.txt 2>&1 & echo "HEADLESS_PID=$!"` — the pid must be a real number; if it prints 0 the run is invalid and is repeated.
3. Every 15 s for 150 s: `ps -p <pid> -o pid,command`, `ls ~/.claude/sessions/`, `ls /tmp/cc-socks/`, then `node projects/ops/skippy-jobs/jobs/work-watch.mjs` followed by `ls -t projects/personal/skippy-app/ala-state/ | head -5` and `command grep -c "<pid>" <the newest work-threads-<machine>.json in that listing>`; record every COMMAND: and OUTPUT: pair verbatim.
4. Write the file with four answers on their own lines — `SESSIONS_DIR: appears | does not appear` · `SOCKET: appears | does not appear` · `TRANSCRIPT: written | not written` · `WORK_THREADS_LIVE: listed | not listed` — followed by the pairs.

**DEFINITION OF DONE:** the evidence file states, from literal output with a real pid, whether a job-started headless session appears in the live per-machine running list within 150 s.
**PROOF:** `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step2-spike-rerun.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH, a `HEADLESS_PID=` line with a number above 0, and all four answer lines present · **FAILS IF:** any pair MISMATCH or UNRUNNABLE, a pid of 0, or an answer line missing.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 3 — S2b · the door: dispatch_to_code_agent, the three bands, the relay route
**FOR NICK:** Skippy has a real way to start Claude Code work, named by kind — and asks him which kind when he did not say — and the rule that sent everything to Cowork is gone from his instructions. · **Tier:** FRONT
**Start when:** none — start now.
**Builder:** zai · **Builder backup:** deepseek · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** `projects/personal/skippy-app/skippy-code-publish/server.js` (the tool list at line 2829, the tools-guide text at lines 2060, 2068, 2308, `RELAY_ALLOWED_TOOLS` at line 3470, the `runTool` case at line 11059); new `projects/personal/skippy-app/skippy-code-publish/lib/code-agent-dispatch.mjs`; new `projects/personal/skippy-app/skippy-code-publish/_test-code-agent-dispatch.mjs`. **Never** the Cowork removal itself (STEP 4), `lib/dispatch-read.mjs` (STEP 4), the family app (VOICE lane).

**Do exactly this:**
1. Job A (new file): `lib/code-agent-dispatch.mjs` exports `dispatchToCodeAgent(input, ctx)` implementing contract 1 of §3: on the cloud (`HAS_RELAY` true and no local queue), POST to the relay route; on the Mac, append the queue row to the code-agent queue file under `projects/personal/skippy-app/ala-state/` and return `{ ok: true, dispatch_id, kind, mode, started_at, visible_in: 'whats_running' }`; `mode: 'coordinate'` calls `sendToSession` from `projects/ops/skippy-jobs/lib/peer-message.mjs` for the thread whose topic matches and returns the delivery result; an absent `kind` returns `{ ok: false, ask: 'which kind — Sonnet, Opus, Fable or Codex?' }` and writes nothing.
2. Job B (one edit): register the tool `dispatch_to_code_agent` beside `push_thread_forward` with the three bands pasted verbatim into its guide text from Nick's words of 2026-09-12 (brain dump §8: do it himself, however many calls — the number of tool calls is not the test; coordinate and supervise what is already running — the default; dispatch only for a genuinely net-new task, and check what is running first) and the rule that when the person did not name the kind, Skippy asks which kind and waits (Nick, 2026-09-12: "make it ask"); add it to `RELAY_ALLOWED_TOOLS`; wire the `runTool` case; leave `handoff_to_cowork` in place for STEP 4 to remove.
3. Job A's test: with the relay mocked, a cloud call produces exactly one queue row of the frozen shape on the mock Mac; a Mac call appends the row; `coordinate` calls `sendToSession` once with the brief; `kind` absent → `{ ok: false, ask }` and no row; an unknown kind → `{ ok: false, error }`; a red control asserts the tool is absent from a copy of server.js with the registration removed; the guide text contains the phrase "ask which kind".

**DEFINITION OF DONE:** the tool exists with the frozen shape, the prompt carries the three bands and the ask rule, and a call from the cloud produces a queue row on the Mac.
**PROOF:** `node projects/personal/skippy-app/skippy-code-publish/_test-code-agent-dispatch.mjs` → `ALL PASS` (exit 0) · **FAILS IF:** any assertion fails, or `command grep -c "dispatch_to_code_agent" projects/personal/skippy-app/skippy-code-publish/server.js` → 0.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 4 — S2b · Cowork out, with its replacement in
**FOR NICK:** Skippy never says "Cowork" again and never hands work to it; the Dispatch screen shows code-agent dispatches instead. · **Tier:** FRONT
**Start when:** STEP 3 closed (`command grep -c "dispatch_to_code_agent" projects/personal/skippy-app/skippy-code-publish/server.js` ≥ 1).
**Builder:** zai · **Builder backup:** deepseek · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** `projects/personal/skippy-app/skippy-code-publish/server.js`; `projects/personal/skippy-app/skippy-code-publish/lib/dispatch-read.mjs`; the tests that name the old tool (`projects/personal/skippy-app/skippy-code-publish/ops/test-pending-handoff-listing.mjs`, `projects/personal/skippy-app/skippy-code-publish/ops/test-handoff-lock-race.mjs`, `projects/personal/skippy-app/skippy-code-publish/ops/test-voice-handoff-confirm-gate.mjs`, `projects/personal/skippy-app/skippy-code-publish/_test-dispatch-read-outbox.mjs`); `projects/ops/skippy-jobs/jobs/skippy-brain-push.mjs` line 145 (the queue file name in ADDS); any job under `projects/ops/skippy-jobs/jobs/` that reads the old queue file (found by `command grep -ln "cowork-queue" projects/ops/skippy-jobs/jobs/*.mjs`, each repointed in the same commit). **Never** the family app's `js/dispatch-panel.js` (VOICE lane; it renders what the brain returns).

**Do exactly this:**
1. Replace every use of `handoff_to_cowork` with `dispatch_to_code_agent`: the tool registration (line 2829), the forced chief-of-staff calls (lines 15671, 15726), the stream buffering (lines 15356–15408), the voice route (lines 16739–16911), the Gracie reword (line 7167), the EXEC_TOOLS sets (lines 9150–9153), the action label (line 16011), and the claim-guard's own vocabulary (lines 1095–1100); the false-completion guards keep their behaviour under the new name; the Alexa path calls the new tool.
2. Rename the queue file from the Cowork name to the code-agent queue in `lib/dispatch-read.mjs` (line 63) and in the ADDS list; every reader found in the jobs folder is repointed in the same commit.
3. Delete the sentence "THE CHIEF-OF-STAFF RULE" and every rule routing work by verb or by tool-call count; the three bands from STEP 3 are the only routing text.
4. Run the four renamed tests and `node --check projects/personal/skippy-app/skippy-code-publish/server.js`; commit; publish.

**DEFINITION OF DONE:** no line of the brain's code, prompts, tests or queue names carries the word cowork, and the replacement door is what every former caller uses.
**PROOF:** `command grep -rci "cowork" projects/personal/skippy-app/skippy-code-publish/server.js projects/personal/skippy-app/skippy-code-publish/lib projects/ops/skippy-jobs/jobs/skippy-brain-push.mjs` → every output line ends `:0`, and `node projects/personal/skippy-app/skippy-code-publish/_test-code-agent-dispatch.mjs` → `ALL PASS` · **FAILS IF:** any count above 0, or the STEP 3 test fails.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 5 — S2c · the Mac drain that starts the session, and work-watch lists it
**FOR NICK:** a dispatch he speaks becomes a real running agent on the Studio, of the kind he named, that shows up in what's running. · **Tier:** FRONT
**Start when:** `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step2-spike-rerun.txt` exists with its four answer lines, and STEP 3's queue row shape is on disk in `lib/code-agent-dispatch.mjs`.
**Builder:** deepseek · **Builder backup:** zai · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** new `projects/ops/skippy-jobs/jobs/code-agent-drain.mjs`; new `projects/ops/skippy-jobs/_test-code-agent-drain.mjs`; `projects/ops/skippy-jobs/runner.mjs` (one registration row, `everyMinutes: 1`); `projects/ops/skippy-jobs/jobs/work-watch.mjs` (a third row source — the code-agent queue file — beside its session files and its job roster, with liveness from the row's `session_pid`; needed unless the spike re-run answered `WORK_THREADS_LIVE: listed`, in which case the row is only the queue). **Never** `lib/peer-message.mjs`, the dispatch-brief gate, `check-codex-dispatch.mjs`, work-watch's account firewall (a dispatch row carries `requested_by` and is listed under that person, never under an unresolved account).

**Do exactly this:**
1. The drain reads queued rows from the code-agent queue file; for `kind` sonnet | opus | fable it runs `git worktree add /private/tmp/code-agent-<id> HEAD`, writes the brief to `/private/tmp/code-agent-<id>/BRIEF.txt`, and starts the session with `execFile('/usr/local/bin/claude', ['-p', <header + brief as ONE argv element>, '--model', <sonnet|opus|fable — the CLI's own aliases, verified 2026-09-12>], { cwd: worktree, detached: true, stdio: [ 'ignore', outFd, outFd ] })` — an argument array, never a shell string, so spoken text can carry any character — with stdout to `/private/tmp/code-agent-<id>/out.txt`, the header being `ROLE: BUILDER`, `NICK-ASKED: <kind> — "<Nick's spoken words>" (Nick, <date>)` when the row's `requested_by` is Nick and the kind was named, `REVIEW: t2`, `RETURN-SIZE: ~1500 tokens`, and `loadTravelBlock()`; for `kind` codex it runs `execFile(resolveCodexBinary(), ['exec', '--profile', 'senior-engineer', '-s', 'workspace-write', '-m', 'gpt-5.6-terra', '-C', worktree, '--skip-git-repo-check', '--output-last-message', '/private/tmp/code-agent-<id>/last.txt', <'CODEX-APPROVED: need you and codex to team up to make this magic and use cheap builders to make it cheap (Nick, 2026-09-12)' + the brief, one element>], { stdio: [ 'ignore', outFd, outFd ] })` with stdin closed (the `ignore` — Codex hangs on an open stdin; the profile name was verified 2026-09-12). `resolveCodexBinary()` returns `/Applications/ChatGPT.app/Contents/Resources/codex` when that file exists (it is not on any job's PATH — measured 2026-09-12: `command -v codex` prints nothing on this Mac, the binary lives inside the ChatGPT app), else `codex` from PATH, else it marks the row `status: 'failed', result: 'codex binary not found'` and starts nothing.
2. It writes `status: 'running'`, `session_pid`, `started_at`, `worktree` and `transcript` (the newest `.jsonl` under `~/.claude/projects/` for that worktree's slug) onto the row; when the process ends it writes `status: 'done'` and `result` (the last line of out.txt or last.txt). work-watch reads the queue file as a third source and lists each `running` row as a thread of kind `dispatch` under `requested_by`, alive while `session_pid` is in the process table.
3. The test puts a fake `claude` and a fake `codex` first on PATH that record their argv to a file, runs the drain on a fixture queue with one row per kind, and asserts: the model alias per kind (`--model sonnet` / `opus` / `fable`), the header lines present, `CODEX-APPROVED:` present with Nick's words for the codex row, the codex child's stdin closed (the fake records whether stdin was a tty or closed), a brief containing a single quote, a backtick and a `$(` reaching the fake unchanged as one argv element (the argv-safety case), one worktree per row, `BRIEF.txt` written, the row's status and pid written, a second run starts nothing new, and `node projects/ops/skippy-jobs/jobs/work-watch.mjs` run against the fixture queue lists the running row under `nick`. One more assertion runs with the fakes REMOVED from PATH: `resolveCodexBinary()` returns a path that `existsSync` confirms (on this Mac the ChatGPT app's own binary) — so a PATH that lacks codex, which is every job's PATH, cannot silently kill the codex kind.
4. Register the job in runner.mjs and prove it with `node --check projects/ops/skippy-jobs/runner.mjs` (never import runner.mjs — that starts a second scheduler).

**DEFINITION OF DONE:** a queued row of each kind becomes a running session with its pid on the row, and work-watch lists it within 2 minutes.
**PROOF:** `node projects/ops/skippy-jobs/_test-code-agent-drain.mjs` → `ALL PASS` (exit 0), `node --check projects/ops/skippy-jobs/runner.mjs` → exit 0, and `command grep -c "code-agent-drain" projects/ops/skippy-jobs/runner.mjs` → 1 · **FAILS IF:** any assertion fails, the syntax check fails, or the registration row is absent.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 6 — S2d · live from the phone path
**FOR NICK:** he says "start a Sonnet agent to …" into his phone and, within two minutes, "what's running?" names it; if he leaves the kind out, Skippy asks. · **Tier:** FRONT
**Start when:** STEP 4 and STEP 5 closed and the brain published (`cd projects/personal/skippy-app/skippy-code-publish && node fly-publish.mjs` reported the new version).
**Builder:** exerciser (haiku) · **Builder backup:** deepseek · **Checker:** sonnet · **Checker backup:** opus
**Files you may touch:** new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step6-live-door.txt`; a scratch file under `/private/tmp/` the dispatched agent writes. **Never** any workspace file.

**Do exactly this:**
1. Sign in as Nick through the family app's proxy the way the memory `reference_family_app_headless_signin_when_wall_up` records, with the secret from `python3 projects/personal/family-vault/vault.py get <the family-app key> --caller=skippy` used in the request and never printed. Record the deployed build under test: `GET /api/version` on the brain and the family app's deployment id, as the first pair — every later pair binds to those.
2. POST `/api/chat` with `voiceOriginated: true` (the same request the Talk tab's microphone sends after transcription — the microphone-to-text leg itself is proven by STEP 7's harness, which drives the real spoken request path and carries this same dispatch sentence among its fresh five) and the text "start a Sonnet agent to write the word PONG into /private/tmp/assistants-step6.txt and stop"; record the reply; note the clock.
3. Every 15 s for 120 s POST "what's running?"; record each reply; stop at the first reply naming the new session and its kind; record the elapsed seconds.
4. `cat /private/tmp/assistants-step6.txt` → PONG. Then POST "start an agent to list the files in /private/tmp" with no kind; record the reply — it must ask which kind and start nothing (a second "what's running?" shows no new session). Then `command grep -ci cowork` over every reply saved → 0.

**DEFINITION OF DONE:** F2 — the spoken dispatch started a real session of the named kind that what's-running listed within 2 minutes; a dispatch without a kind was answered with a question; no reply mentioned Cowork.
**PROOF:** `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step6-live-door.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH, an `elapsed:` line ≤ 120, a `kind: sonnet` line, an `asked for kind: yes` line, and `cowork mentions: 0` · **FAILS IF:** any MISMATCH, elapsed above 120, the kind wrong, no question when the kind was absent, or a Cowork mention.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once, then one fresh dispatch of your own with a new scratch filename. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

**Handoff:** the moment this step closes, post one dated line into `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/PROGRESS.txt`: `STEP 6 (ASSISTANTS) closed <date> — the brain's dispatch tool is dispatch_to_code_agent; handoff_to_cowork and the Cowork queue name are gone; the Dispatch screen's rows now carry kind and session.`

### STEP 7 — S3 · ten requests done for real
**FOR NICK:** ten ordinary things he says — including add milk to the shopping list — actually happen and are read back; Skippy never claims a hand-off it did not make. · **Tier:** FRONT
**Start when:** STEP 4 closed (the false hand-off wording has no tool to hide behind).
**Builder:** zai · **Builder backup:** deepseek · **Checker:** sonnet · **Checker backup:** opus
**Files you may touch:** `projects/personal/skippy-app/skippy-code-publish/lib/shopping-list.mjs` (add a `SHOPPING_ADD_MUTATION` constant and `shapeAddResult()`; `SHOPPING_QUERY` untouched — `_test-shopping-list.mjs` asserts it never writes); `projects/personal/skippy-app/skippy-code-publish/server.js` (a new `add_shopping_item` tool beside `read_shopping_list` at line 4255 in the `add_todo` pattern at line 3872, and the claim-guard at lines 1095–1100 extended to the phrases "handing this to", "passing this to", "I'll hand this off" when the same turn carries no dispatch receipt); `projects/personal/skippy-app/skippy-code-publish/_test-shopping-list.mjs` (extend); new `projects/personal/skippy-app/skippy-code-publish/_test-claim-guard-handoff.mjs`; new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/fresh-five.json`. **Never** `projects/personal/family-app/functions/api/shopping-add.js` (read it for the board and column ids; VOICE lane owns it), `projects/personal/family-app/_test-voice-requests.mjs` (the instrument).

**Do exactly this:**
1. Job A (one file): the mutation and its shaper in `lib/shopping-list.mjs`, mirroring the board id and column ids of `functions/api/shopping-add.js`; extend `_test-shopping-list.mjs` with a fixture add result and the assertion that `SHOPPING_QUERY` still never writes.
2. Job B (one edit): the `add_shopping_item` tool in server.js — actor check as `add_todo` has, writes through the mutation, reads the item back and answers with the item's name and store.
3. Job C (new file, red first): `_test-claim-guard-handoff.mjs` feeds the guard a reply "I'm handing this to a code agent" with no dispatch receipt in the turn and asserts it is caught; with a receipt, passes; record the red run before the guard edit lands. Then the guard edit (one route-build edit).
4. Write `fresh-five.json` with exactly five new spoken requests (each stamped `VOICE-TEST 2026-09-08` in the record it creates, as the harness requires): "add milk to the shopping list", "add a to-do to call the dentist Friday", "what's on my calendar tomorrow", "remind me to send the invoice Monday", "start a Sonnet agent to write the word PONG into /private/tmp/assistants-step7.txt and stop" (read back as a queue row with kind sonnet — this is the microphone-path proof of the door that STEP 6 measures through the chat route) — the harness's own five recorded requests make ten. "what's running?" is timed by STEP 8, not counted here.
5. Commit, publish, then run the gate.

**DEFINITION OF DONE:** F3 — the gate reports every one of the ten requests delivered and read back, and the claim-guard test passes with its red run on file.
**PROOF:** `node projects/personal/family-app/_test-voice-requests.mjs --gate --fresh projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/fresh-five.json` → exit 0 AND its summary line literally reads five of five recorded and five of five fresh delivered and read back (the harness's own gate passes at four of five fresh — its help text says so — and this plan's bar is ten of ten, so the checker reads the summary, not the exit code), and `node projects/personal/skippy-app/skippy-code-publish/_test-claim-guard-handoff.mjs` → `ALL PASS` · **FAILS IF:** any request not delivered or not read back (a summary of four of five fresh is a FAIL here even at exit 0), or the guard test fails.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

**Handoff:** the moment this step closes, post one dated line into `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/PROGRESS.txt`: `STEP 7 (ASSISTANTS) closed <date> — the brain has add_shopping_item and a hand-off claim-guard; your gate ran ten of ten with fresh-five.json.`

### STEP 8 — S4 · what's running in 15 seconds
**FOR NICK:** "what's running?" is answered in fifteen seconds, not fifty. · **Tier:** FRONT
**Start when:** none — start now.
**Builder:** zai · **Builder backup:** deepseek · **Checker:** sonnet · **Checker backup:** qwen
**Files you may touch:** `projects/personal/skippy-app/skippy-code-publish/server.js` (the `whats_running` tool at line 2777 and its result renderer); new `projects/personal/skippy-app/skippy-code-publish/ops/time-whats-running.mjs`; `projects/personal/skippy-app/skippy-code-publish/_test-whats-running-live.mjs` (keep green). **Never** `projects/ops/skippy-jobs/jobs/work-watch.mjs` (STEP 5 owns its one change), the lane transport.

**Do exactly this:**
1. Measure first: from the lane log on the box, record the number of model rounds and the wall time of three "what's running?" turns before any change, into the evidence folder.
2. Make the tool return the rundown already rendered as the spoken sentences (count, each session in one clause, which needs him, what just landed) so the model's next round relays rather than composes; parse the newest per-machine `work-threads-<machine>.json` once per request and hold the parsed snapshot for the turn.
3. `ops/time-whats-running.mjs`: signs in as Nick (the vault secret, never printed), records the brain's `/api/version` build first, asks "what's running?" five times through `/api/chat` with `voiceOriginated: true` (the request the microphone sends), prints each sample's milliseconds and the live-session count the answer names beside the count the newest per-machine file holds at that moment, and ends `5/5 under 15000 ms` or `FAIL <n>/5`.
4. Commit, publish, run the timing script.

**DEFINITION OF DONE:** F4 — five of five samples answer correctly in 15 seconds or less.
**PROOF:** `node projects/personal/skippy-app/skippy-code-publish/ops/time-whats-running.mjs` → `5/5 under 15000 ms` with every sample's named count equal to the file's count, and `node projects/personal/skippy-app/skippy-code-publish/_test-whats-running-live.mjs` → green · **FAILS IF:** any sample above 15000 ms or a count that disagrees.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

**Handoff:** the moment this step closes, post one dated line into `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/PROGRESS.txt`: `STEP 8 (ASSISTANTS) closed <date> — whats_running answers in ≤ 15 s; the tool returns the spoken rundown itself.`

### STEP 9 — S5a · memory proposals relay to the Mac's scanner
**FOR NICK:** "remember this" is no longer refused on his phone. · **Tier:** FRONT
**Start when:** none — start now.
**Builder:** zai · **Builder backup:** deepseek · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** `projects/personal/skippy-app/skippy-code-publish/server.js` (`RELAY_ALLOWED_TOOLS` at line 3470; the `capture_memory` and `confirm_memory` `runTool` paths); `projects/personal/skippy-app/skippy-code-publish/lib/memory-write-filter.mjs` (export a `scannerLoaded` boolean beside `filterMemoryWrite`, nothing else); `projects/personal/skippy-app/skippy-code-publish/_test-memory-write-filter.mjs` (extend); `projects/personal/skippy-app/skippy-code-publish/_test-relay-port-scope.mjs` (keep green). **Never** `projects/personal/skippy-app/lib/grunt-egress-scan.mjs`, `fly-publish.mjs`'s manifest (the scanner is never shipped).

**Do exactly this:**
1. Export `scannerLoaded` from `lib/memory-write-filter.mjs` (true only when the real scanner import at line 45 succeeded).
2. In server.js: when `scannerLoaded` is false and `HAS_RELAY` is true, `capture_memory` and `confirm_memory` relay to the Mac program (add both to `RELAY_ALLOWED_TOOLS`), which runs the same tool with the real scanner and owns memory-pending.jsonl and memory-confirmed-facts.jsonl; when the scanner is present the tool runs locally as today; when neither, it refuses as today — never an unscanned write anywhere.
3. Extend `_test-memory-write-filter.mjs`: scanner absent + relay present → the relay is called once with the proposal and the tool result is `pending`, not `refused`; scanner absent + no relay → `refused`; scanner present → local path, relay not called.
4. Commit, publish; `node projects/personal/skippy-app/skippy-code-publish/_test-relay-port-scope.mjs` stays green (the relay's allowlist test).

**DEFINITION OF DONE:** on a box without the scanner, a memory proposal is scanned and stored pending on the Mac, never refused and never stored unscanned.
**PROOF:** `node projects/personal/skippy-app/skippy-code-publish/_test-memory-write-filter.mjs` → `ALL PASS` with the three new cases, and `node projects/personal/skippy-app/skippy-code-publish/_test-relay-port-scope.mjs` → green · **FAILS IF:** any case fails or the relay scope test goes red.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 10 — S5b · a yes from the person is the decision
**FOR NICK:** when Skippy asks "want me to keep that?" his yes — spoken or typed — keeps it; the tap is still there. · **Tier:** FRONT
**Start when:** none — start now (§1a row 2: Nick, 2026-09-12, "1 yes").
**Builder:** zai · **Builder backup:** deepseek · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** `projects/personal/skippy-app/skippy-code-publish/lib/memory-gate.mjs` (`verbalYes` at line 94 and its header comment); `projects/personal/skippy-app/skippy-code-publish/_test-memory-gate.mjs`; `projects/personal/skippy-app/skippy-code-publish/server.js` line 2647 (`confirm_memory`'s text: keep true saves durably); new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step10-red.txt`. **Never** the four approval classes' handling — a yes never grants money leaving, a rotation, a destruction or a message as Nick.

**Do exactly this:**
1. Red first: add to `_test-memory-gate.mjs` the case "verbalYes(id, { by: 'person' }) moves the row into memory-confirmed-facts.jsonl", run it on today's code, save the FAIL lines to `step10-red.txt`.
2. Make `verbalYes(id, { by: 'person' })` durable — it calls the same path `recordTap({ keep: true })` uses — and keep a bare `verbalYes(id)` (no `by`) as today's non-durable mark, so a model-initiated affirmation can never keep a fact; rewrite the header comment to state the new rule with Nick's words and date (no strikethrough, no "used to").
3. Wire the chat path: a clear yes from the signed-in person that answers a pending proposal in the same conversation calls `verbalYes(id, { by: 'person' })`; the confirm_memory tool text says keep true saves durably.
4. Commit, publish.

**DEFINITION OF DONE:** a person's yes to a named pending fact lands the row in the confirmed file; a model-side affirmation still does not.
**PROOF:** `node projects/personal/skippy-app/skippy-code-publish/_test-memory-gate.mjs` → `ALL PASS` including the new case and a case asserting the bare call stays non-durable, with `step10-red.txt` holding the earlier FAIL · **FAILS IF:** either case fails or the red file is absent.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 11 — S5c · confirmed facts reach the engine door
**FOR NICK:** what he tells Skippy to keep is what Skippy answers with next week, because it lands in the store he reads from. · **Tier:** FRONT
**Start when:** STEP 10 closed.
**Builder:** zai · **Builder backup:** deepseek · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** new `projects/ops/skippy-jobs/jobs/memory-confirmed-drain.mjs`; new `projects/ops/skippy-jobs/_test-memory-confirmed-drain.mjs`; `projects/ops/skippy-jobs/runner.mjs` (one registration row, `everyMinutes: 15`); `projects/personal/skippy-app/skippy-code-publish/server.js` lines 1180–1200 (the `captureMemory()` branch that appends to captures.md — removed; everything proposes). **Never** `projects/personal/health/engine/brain-routing/personal_mcp.py` (read its `tool_capture_narrative` args at line 502; the engine lane owns it), the marker table.

**Do exactly this:**
1. The drain reads the Mac's memory-confirmed-facts.jsonl (its directory is the `baseDir` server.js passes to `createMemoryGate`), skips rows stamped `drained_at`, runs `scanForHealthMarkers` from `lib/memory-write-filter.mjs` on each; a marker-shaped row is stamped `refused: marker` and listed in the drain's log for the engine lane; every other row takes the engine's own two-call path, run by `python3 -c` importing from `projects/personal/health/engine/brain-routing/`: first `tool_capture_narrative({ narrative: <the row's text>, heading: <"Skippy memory, confirmed by " + person + " " + date>, dry_run: false, source_at: <the row's date>, about_person: <the row's person> })` — which returns the staging row with its id, event key and content sha, or `REFUSED` for a marker the engine's own filter catches (stamped the same way) — then `publish_accepted_event(staging_id, event_key, content_sha256, accepted_by = "<person>, spoken or tapped yes, <ISO date>", db = "postgres")` so the fact reaches the production store `personal_answer` reads; the row is stamped `drained_at` with both ids only after the publish returns success.
2. The test: a fixture confirmed file with a plain row, a marker-shaped row and an already-drained row; a fake `python3` first on PATH records its argv and answers a staging id; assert the capture call carries the plain row's text, `dry_run` false and the heading, the publish call carries that staging id and an `accepted_by` naming the person, the marker row is refused before any call, the drained row is skipped, and a second run calls nothing.
3. Remove the captures.md branch in server.js; `node --check` it; register the job; `node --check projects/ops/skippy-jobs/runner.mjs`.

**DEFINITION OF DONE:** a confirmed row is written through the engine's door once and stamped; marker-shaped rows are refused and handed to the engine lane.
**PROOF:** `node projects/ops/skippy-jobs/_test-memory-confirmed-drain.mjs` → `ALL PASS`, `node --check projects/ops/skippy-jobs/runner.mjs` → exit 0, `command grep -c "memory-confirmed-drain" projects/ops/skippy-jobs/runner.mjs` → 1, and `command grep -c "captures.md" projects/personal/skippy-app/skippy-code-publish/server.js` → 0 · **FAILS IF:** any of the four disagrees.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

**Handoff:** the moment this step closes, post one dated line into `projects/ops/life-os/REGROUP-2026-09-08/plans/HEALTH-ONE-DOOR/PROGRESS.txt`: `STEP 11 (ASSISTANTS) closed <date> — confirmed memories drain to capture_narrative; marker-shaped rows are refused by scanForHealthMarkers and listed in the drain's log for you.`

### STEP 12 — S5d · live memory proof
**FOR NICK:** he tells Skippy something by voice, says yes, and Skippy has it tomorrow; a password-shaped phrase is refused and he is told why. · **Tier:** FRONT
**Start when:** STEPS 9, 10 and 11 closed and the brain published.
**Builder:** exerciser (haiku) · **Builder backup:** deepseek · **Checker:** sonnet · **Checker backup:** opus
**Files you may touch:** new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step12-live-memory.txt`. **Never** any workspace file; never a real credential — the refusal case uses an invented value in the shape of one.

**Do exactly this:**
1. Sign in as Nick (STEP 6's way). POST `/api/chat` with `voiceOriginated: true`: "remember that my test phrase for the assistants plan is <a fresh nonce>"; record the proposal; POST "yes"; record the reply.
2. `node projects/ops/skippy-jobs/jobs/memory-confirmed-drain.mjs` by hand; record its output line for the row.
3. Restart the brain (`fly apps restart` on the app fly-publish.mjs names) so nothing survives in process memory; start a NEW conversation (a new thread id); POST "what did I ask you to remember about the assistants plan?"; record the reply.
4. New conversation: POST "remember that my login for the bank is <an invented value in the shape of a login and password>"; record the reply.

**DEFINITION OF DONE:** F5 — the nonce comes back in a fresh conversation after a restart, and the credential-shaped phrase is refused with the floor category named.
**PROOF:** `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step12-live-memory.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH; the step-3 reply contains the nonce; the step-4 reply contains the word floor and none of the invented value · **FAILS IF:** any MISMATCH, the nonce absent, the refusal absent, or the invented value echoed in any REPLY, in memory-pending.jsonl, in memory-confirmed-facts.jsonl or in the drain's log (the COMMAND line that sent it necessarily contains it and is the one place it may appear).

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once, the next calendar day, with your own nonce. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 13 — S6a · every message for Nick comes from Skippy
**FOR NICK:** nothing reaches him from Gracie any more; alerts about her pipeline arrive from Skippy. · **Tier:** FRONT
**Start when:** none — start now.
**Builder:** deepseek · **Builder backup:** zai · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** `projects/ops/skippy-jobs/jobs/gracie-ear.mjs`; every job under `projects/ops/skippy-jobs/jobs/` that holds GRACIE_BOT_TOKEN and sends to Nick — the list is `command grep -ln "GRACIE_BOT_TOKEN\|GRACIE_NICK_DM" projects/ops/skippy-jobs/jobs/*.mjs`, read each send site before editing; new `projects/ops/skippy-jobs/_test-sender-rule.mjs`. **Never** `projects/ops/skippy-jobs/lib.mjs` (`alertNick` is the destination, not the change), `send-nick.mjs`, any send to Chantelle (those stay Gracie's).

**Do exactly this:**
1. Red first: `_test-sender-rule.mjs` reads every job file, ignores comment lines (lines whose first non-space characters are `//` or `*`), and counts executing send calls whose token argument is `GRACIE_BOT_TOKEN` and whose channel is `GRACIE_NICK_DM` or Nick's user; prints `<n> send sites to Nick via the Gracie token`; exits 1 when n > 0. It also counts, in the same files, executing calls to `alertNick(` or `send-nick.mjs` and prints `<m> sites reaching Nick as Skippy` — so a builder who simply deletes a Gracie-token send cannot pass: after the repoint, m must have grown by at least the number of sites the red run found. Run it on today's code and keep the red output (both counts) in the evidence folder.
2. Repoint each executing site to `alertNick()` from `projects/ops/skippy-jobs/lib.mjs` (severity as the job's comment already states) or to `node projects/ops/skippy-jobs/lib/send-nick.mjs --class <asks|business|money|family|health>` for a task-shaped job, the way `gracie-health.mjs` line 368 already does; a message FOR Chantelle keeps the Gracie token.
3. Run the test green; one commit.

**DEFINITION OF DONE:** no job sends a message to Nick through the Gracie token; the test that proves it was red first.
**PROOF:** `node projects/ops/skippy-jobs/_test-sender-rule.mjs` → `0 send sites to Nick via the Gracie token` and `<m> sites reaching Nick as Skippy` with m at least the red run's m plus the red run's n, exit 0 · **FAILS IF:** the count is above 0, the Skippy-path count did not grow by the sites removed, or the red run is not on file.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 14 — S6b · six live identity checks
**FOR NICK:** it is proven, live, that Chantelle meets Gracie everywhere and he never does. · **Tier:** FRONT
**Start when:** STEP 13 closed.
**Builder:** exerciser (haiku) · **Builder backup:** deepseek · **Checker:** sonnet · **Checker backup:** opus
**Files you may touch:** new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step14-identity.txt`; new `projects/personal/skippy-app/skippy-code-publish/_test-slack-caller-routing.mjs` (written by a cheap builder — zai — red-first before the exerciser runs the legs; it reads the send path around server.js line 9701 with a fixture caller, never sends). **Never** a message sent as Chantelle or as Nick to anyone (an approval class) — every Slack leg is a read of an existing channel or a unit test.

**Do exactly this (six checks, each a COMMAND:/OUTPUT: pair ending `RESULT: PASS` or `RESULT: FAIL`):**
1. App, Chantelle: sign in through `node projects/personal/skippy-app/skippy-code-publish/ops/drive-as-chantelle.mjs` (CHANTELLE_LOGIN_SECRET from the vault, never printed); ask "who are you?"; PASS when the answer names Gracie and never Skippy.
2. App, Nick: sign in as Nick; ask "who are you?"; PASS when the answer names Skippy and never Gracie.
3. Voice, Chantelle: mint a voice session as Chantelle through the family app's voice-session route (`projects/personal/family-app/functions/api/voice-session-openai.js` and its ElevenLabs sibling); PASS when the agent id in the mint equals the deployed GRACIE_AGENT_ID (compare in code, print only `equal`/`different`).
4. Voice, Nick: the same as Nick; PASS when the agent id is Skippy's and differs from Gracie's.
5. Slack, Chantelle — two halves, both required: (a) the mechanism: `node projects/personal/skippy-app/skippy-code-publish/_test-slack-caller-routing.mjs` (written red-first by this step: it feeds the brain's Slack send path a caller resolved from a fixture signed identity for Chantelle and asserts the token chosen at server.js line 9701 is GRACIE_BOT_TOKEN, and for Nick is SLACK_USER_TOKEN; an unsigned or claimed identity resolves to neither); (b) the history: with the Gracie token, read the last 20 messages of GRACIE_CHANTELLE_DM — the newest bot reply is from the Gracie bot user. A live message from Chantelle cannot be sent by an agent (an approval class), so if one arrives in the window it is recorded as a third, causal half; otherwise the leg reads `RESULT: PASS (mechanism + history)` and says so.
6. Slack, Nick — the same two halves: (a) the routing test's Nick case; (b) with the Skippy token, read the last 20 messages of Nick's DM with Skippy — the newest bot message is the Skippy bot user — and with the Gracie token, read GRACIE_NICK_DM — no bot message newer than STEP 13's commit time. PASS when all hold.

**DEFINITION OF DONE:** F6 — six PASS lines from live reads.
**PROOF:** `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step14-identity.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH and `command grep -c "^RESULT: PASS" projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step14-identity.txt` → 6 · **FAILS IF:** any MISMATCH or fewer than six PASS lines; a leg that truly cannot be read is written `RESULT: NOT MEASURABLE — <instrument>` and counts as not passed.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 15 — S7 · Voice B — talk to a running thread, be told when it answers
**FOR NICK:** in the Status tab he taps a thread's mic, says a thing, gets on with his life, and hears the session's answer read out as one plain sentence when it lands. · **Tier:** FRONT
`NICK-ASKED: codex — "need you and codex to team up to make this magic and use cheap builders to make it cheap" (Nick, 2026-09-12)`
**Start when:** none — start now.
**Builder:** gpt-5.6-terra · **Builder backup:** gpt-5.6-luna · **Checker:** sonnet · **Checker backup:** opus
**Files you may touch:** `projects/personal/family-app/js/panel.js` (the per-thread answer box: a mic control that dictates into it and sends through the existing `/api/thread-reply` at line 1079; a poll of the thread's newest reply that hands it to the speak route; no new CSS — existing Pearl classes only); `projects/personal/family-app/functions/api/thread-reply.js` (only if a field is missing); `projects/personal/skippy-app/skippy-code-publish/server.js` (a new `POST /api/thread-say` that turns a session's newest reply into one spoken sentence with the same muscle as `whats_running` and returns audio through the existing onyx TTS path); new `projects/personal/family-app/_test-thread-voice.mjs`. **Never** `projects/personal/family-app/js/voice.js` (its line 21 rule stands: it never touches the Status tab), the Pearl look, `lib/peer-message.mjs`.

**Do exactly this:**
1. `codex exec --profile senior-engineer -s workspace-write -m gpt-5.6-terra -C "/Users/nickdeck/Documents/Claude 2.0" --skip-git-repo-check --output-last-message /private/tmp/assistants-step15.txt "CODEX-APPROVED: need you and codex to team up to make this magic and use cheap builders to make it cheap (Nick, 2026-09-12) — <this step's fence, items 2–4 and its PROOF, pasted>" < /dev/null`.
2. The join: NO new control — the thread answer box's existing dictation mic is the input; the text posts to `/api/thread-reply` (delivery must be live into the running session — the reply's `delivered` field); the panel polls that thread's newest reply and, when it is newer than the message sent, POSTs it to `/api/thread-say` and plays the audio; the only new visible thing is one line of status text in an existing Pearl class — "landed — I'll read the answer when it comes" — and it never waits like a phone call. After deploying, run the family app's own fidelity check on the Status tab (`projects/personal/skippy-app/design-directions/_pearl-fidelity-check.mjs`, with the Status-tab arguments the VOICE lane's STEP 1 proof uses) and record `mismatched properties: 0 · unmeasured anchors: 0`.
3. `/api/thread-say`: the brain turns diffs, paths and command output into one sentence a person wants to hear; a reply that is only a diff becomes "it changed <n> lines in <plain name of the area>".
4. The test signs in as Nick, plants a reply into a scratch thread through the same delivery path, and measures the seconds from landing to audio; asserts ≤ 30 and that the spoken text carries no `/`-path token and no line starting `+` or `-`. Deploy with `node projects/ops/deploy.mjs deck-family`; run the test against the live app.

**DEFINITION OF DONE:** F7 — a spoken message lands live in a RUNNING session and its reply is spoken back as a plain sentence as soon as the session answers (measured from the moment the session's answer exists, within 30 seconds of that moment); a session that is IDLE at its prompt cannot be woken by a delivered message (a platform limit, measured 2026-09-13), so for an idle thread the person is told so out loud and offered a fresh agent with the same message, and nothing is started without a yes. (Bar changed by Nick, 2026-09-13: "5 yes" to "spoken back as soon as the thread answers, and an idle thread says so and offers a fresh agent".)
**PROOF:** running thread: `THREAD_VOICE_SCRATCH_POINTER=<a mid-turn session id> THREAD_VOICE_EXPECT=reply node projects/personal/family-app/_test-thread-voice.mjs` → `spoken in <n> s` (n ≤ 30 from the session's own answer), `sentence clean`, `waiting line cleared`, exit 0; idle thread: the same with `THREAD_VOICE_EXPECT=idle` → `idle thread told plainly: …`, `sentence clean`, `waiting line cleared`, exit 0; and the Status-tab fidelity check → no mismatch on the Status tab other than a true fault strip whose text names a real fault (the machine-staleness strip) · **FAILS IF:** the running-thread reply is not spoken within 30 s of the session answering, the delivery not live, a path or diff token in the sentence, the idle sentence claims something was started, or a Status mismatch that is not a true fault strip.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once, on the live app. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

**Handoff:** the moment this step closes, post one dated line into `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/PROGRESS.txt`: `STEP 15 (ASSISTANTS) closed <date> — panel.js carries a per-thread mic and speaks replies; your Status-tab fidelity count must still read 0 — re-run it once.`

### STEP 16 — S8a · Voice A — the architecture decision, measured
**FOR NICK:** he gets one page with real numbers on his own questions and makes one choice: how a real back-and-forth conversation with Skippy gets built. · **Tier:** FRONT
`NICK-ASKED: codex — "need you and codex to team up to make this magic and use cheap builders to make it cheap" (Nick, 2026-09-12)`
**Start when:** none — start now.
**Builder:** gpt-5.6-terra · **Builder backup:** gpt-5.6-luna · **Checker:** sonnet · **Checker backup:** opus
**Files you may touch:** new `projects/personal/family-app/_bakeoff-voice-a.mjs`; new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/voice-a-decision.txt`. **Never** `js/voice.js`, `lib/openai-realtime-adapter.mjs`, any deploy — this step measures and writes.

**Do exactly this:**
1. `codex exec --profile senior-engineer -s workspace-write -m gpt-5.6-terra -C "/Users/nickdeck/Documents/Claude 2.0" --skip-git-repo-check --output-last-message /private/tmp/assistants-step16.txt "CODEX-APPROVED: need you and codex to team up to make this magic and use cheap builders to make it cheap (Nick, 2026-09-12) — <this step's fence, items 2–4 and its PROOF, pasted>" < /dev/null`.
2. Ten of Nick's own spoken prompts, taken verbatim from `projects/personal/skippy-app/conversations/` (his turns, not agents' test turns), fixed in the harness.
3. Two candidates on the same ten: the current ElevenLabs agent (`js/voice.js`, the brain as the agent's LLM) and an OpenAI Realtime front desk that calls the existing brain as a tool (option b of brain dump §9), measured for first audio, barge-in, tool-answer correctness and cost per turn; `--report` prints the table and exits 0 only when both have ten samples.
4. The decision file: the three options with the numbers — (a) Skippy inside the realtime model is NOT measured (it would mean rebuilding the tool layer first) and is described with its cost and what it would take; (b) the front desk, measured; (c) the split by intent, whose numbers are DERIVED and labelled so: (b)'s figures for turns the brain answers and the front desk's own first-audio figure for chat turns it answers itself — a third harness is not built; the recommendation (b then c, unless the numbers say otherwise), Nick's standing rulings restated (never always-listening; the Pearl look; a web app inside the existing apps, never a third app), and one blank line `VOICE A CHOICE:` for the driver to fill with Nick's words and date in `PROGRESS.txt`.

**DEFINITION OF DONE:** the decision file exists with measured numbers for both candidates on the same ten prompts.
**PROOF:** `node projects/personal/family-app/_bakeoff-voice-a.mjs --report` → `elevenlabs: 10 samples · realtime-front-desk: 10 samples` with medians for first audio, exit 0 · **FAILS IF:** either candidate under ten samples or a number without its method.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 17 — S8b · Voice A — the chosen path built
**FOR NICK:** he talks to Skippy and Skippy talks back like a conversation — first word under two seconds, interruptible. · **Tier:** FRONT
`NICK-ASKED: codex — "need you and codex to team up to make this magic and use cheap builders to make it cheap" (Nick, 2026-09-12)`
**Start when:** `command grep -c "VOICE A CHOICE:" projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/PROGRESS.txt` → 1 with Nick's words and date on that line.
**Builder:** gpt-5.6-terra · **Builder backup:** gpt-5.6-luna · **Checker:** sonnet · **Checker backup:** opus
**Files you may touch:** `projects/personal/family-app/js/voice.js`; `projects/personal/family-app/functions/api/voice-session-openai.js` and its sibling routes; `projects/personal/skippy-app/skippy-code-publish/lib/openai-realtime-adapter.mjs` and the brain's voice routes in server.js when the choice is (b) or (c). **Never** the Pearl look, always-listening of any kind, a third app, the Talk tab's fidelity target (VOICE lane).

**Do exactly this:**
1. `codex exec --profile senior-engineer -s workspace-write -m gpt-5.6-terra -C "/Users/nickdeck/Documents/Claude 2.0" --skip-git-repo-check --output-last-message /private/tmp/assistants-step17.txt "CODEX-APPROVED: need you and codex to team up to make this magic and use cheap builders to make it cheap (Nick, 2026-09-12) — <this step's fence, the chosen option from PROGRESS.txt, and its PROOF, pasted>" < /dev/null`.
2. Build the chosen option additively behind the existing orb; the brain, every tool, every engine and every rule stay exactly as they are; deploy with `node projects/ops/deploy.mjs deck-family` and publish the brain if it changed.
3. Run the harness live on the ten prompts.

**DEFINITION OF DONE:** F8 — the chosen path answers a spoken question with first audio under 2 seconds on ten of ten samples.
**PROOF:** `node projects/personal/family-app/_bakeoff-voice-a.mjs --live` → `chosen path: 10/10 first audio < 2.0 s`, exit 0 · **FAILS IF:** any sample at or above 2.0 s.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once, on the live app. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

**Handoff:** the moment this step closes, post one dated line into `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/PROGRESS.txt`: `STEP 17 (ASSISTANTS) closed <date> — voice.js carries the chosen conversation path; your Talk-tab fidelity count must still read 0 — re-run it once.`

### STEP 18 — S9a · three faces drawn, composited and graded
**FOR NICK:** Skippy, Gracie and Neeko each have a face — three siblings from one mark — and it passed Sienna. · **Tier:** FRONT
**Start when:** none — start now (Sienna's criteria are on file: `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/SIENNA-AVATARS-AND-VOICES-2026-09-12.txt` §1).
**Builder:** deepseek · **Builder backup:** qwen · **Checker:** fable (Sienna, creative-director — the one grade) · **Checker backup:** opus
**Files you may touch:** the directory `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/avatars/` (three 512 × 512 masters `skippy.png`, `gracie.png`, `neeko.png`, their reversed variants, and the composite screenshots); new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/avatars-composite.html`; new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/avatars-grade.txt` (Sienna's verdict, written by the driver from her report). **Never** any app screen, any Slack setting, any colour outside her four (#141414, #F4F2EF, and for Gracie alone #B8657A / #F3DDE3).

**Do exactly this:**
1. Draw each master IN CODE as an SVG mark — a generator script in the avatars directory writes `skippy.svg`, `gracie.svg`, `neeko.svg` and their reversed variants from her §1 exactly (one abstract geometric mark, the same construction and stroke weight for all three, ink #141414 on pearl #F4F2EF; Skippy the base mark; Gracie the same mark with one softer variation of form and the rose accent #B8657A / #F3DDE3 at no more than a fifth of the mark's area; Neeko the same mark squarer and closed with no accent; the mark filling 60–70 % of the circle; nothing finer than a 2 px stroke at 32 px, so no stroke thinner than 32 px at 512; no face, robot, mascot, letters, emoji, gradient or glass effect) and renders each to a 512 × 512 PNG through the shared headless Chrome (`projects/shared-tooling/test-chrome.mjs`). Code is chosen over an image model because her criteria are exact colours and stroke floors an image model cannot promise; DeepAPI's image endpoint (`source ~/.deepapi/env` first) is the fallback only if the drawn marks fail her twice.
2. Write `avatars-composite.html` that shows the three at 512 px, 32 px and the four family-app slot sizes (24, 30, 34, 44 px — the .avs, .rail .me, .rail .brand and .hav sizes in `projects/personal/family-app/pearl-tokens.css`) on the pearl ground and on a Slack light and dark sidebar; screenshot it with `projects/shared-tooling/test-chrome.mjs` (read its header for the call) into the avatars directory — screenshots, never inferred from the files.
3. The cold matcher: dispatch one Sonnet session that has not seen this plan or the prompts, showing only the three 32 px images and Sienna's three one-line descriptions; it must match all three; its answer goes into the grade file.
4. Hand the masters, the composites and the matcher's answer to Sienna by name (`creative-director`) for gates 1, 2 and 6, once; the driver writes her verdict per face into `avatars-grade.txt` as `skippy: PASS|FAIL — <her words>` and so on. A FAIL loops the builder on that face only; after three rounds the driver takes two options to Nick, as her file says.

**DEFINITION OF DONE:** three masters, their composites and a correct cold match exist, and each face carries Sienna's PASS.
**PROOF:** `ls projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/avatars/ | command grep -c png` ≥ 9 (three masters, three reversed, at least three composite screenshots) and `command grep -c ": PASS" projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/avatars-grade.txt` → 3 · **FAILS IF:** fewer files, a wrong cold match, or fewer than three PASS lines.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** grade once against your own criteria, from the screenshots and the matcher's answer; PASS closes the step. Do not accept the builder's description of the images; do not summon anyone else.

### STEP 19 — S9b · three faces live in Slack
**FOR NICK:** in Slack, a message from Skippy, Gracie or Neeko shows that assistant's own face. · **Tier:** FRONT
`NICK-ASKED: fable — "Fable is the driver, Codex is the builder, and then Fable helps deploy supportive elements for all the Slack stuff and all the infrastructure cleanup" (Nick, 2026-09-12); the browser drive itself rests on "i dont care about this - everything on my screen is approved for agents to see" (Nick, 2026-08-30)`
**Start when:** STEP 18 closed (three graded masters in the avatars directory).
**Builder:** the driver session, driving Nick's own signed-in Chrome (the claude-in-chrome tools this session holds) on the standing grant that agents drive his apps as him (Nick, 2026-08-30) — Slack's API cannot set a bot's icon, but the app-settings page at api.slack.com can be driven; fallback: Nick uploads the three at the keyboard · **Builder backup:** Nick at the keyboard · **Checker:** sonnet · **Checker backup:** opus
**Files you may touch:** new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step19-slack-icons.txt`. **Never** the Hub's screens (HUB lane) or the family app's screens (VOICE lane) — their attribution images are a handoff, below; never any Slack setting other than the three apps' display icons.

**Do exactly this:**
1. Before the upload: with each bot's own token (from the Fly secrets or the local .env, never printed), call Slack's `users.info` for the bot user, download `image_72`, record its content hash; three COMMAND:/OUTPUT: pairs.
2. In Nick's Chrome, open each of the three Slack apps' settings page (Basic Information → Display Information), upload the matching 512 px master from the avatars directory, save; record the page's confirmation text.
3. After the upload: repeat the three reads; the three new hashes differ from the three old ones and from each other; download the three new `image_72` files and hand them, with Sienna's three one-line descriptions, to a fresh Sonnet session that never saw the masters or this plan — it must match each bot's image to the right assistant (the same cold matcher STEP 18 used); record its answer. Record `hashes distinct: 3` and `matched: 3 of 3`.

**DEFINITION OF DONE:** the three bots show the three graded faces, each on the right bot, read back by API and matched cold.
**PROOF:** `node projects/ops/skippy-jobs/lib/verify-agent-evidence.mjs projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step19-slack-icons.txt --cwd "/Users/nickdeck/Documents/Claude 2.0"` → every pair MATCH, the line `hashes distinct: 3` and the line `matched: 3 of 3` · **FAILS IF:** any MISMATCH, fewer than three distinct hashes, or a wrong match.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

**Handoff:** the moment this step closes, post one dated line into `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/PROGRESS.txt` and one into the HUB lane's progress record: `STEP 19 (ASSISTANTS) closed <date> — the three faces are in evidence/avatars/ under the ASSISTANTS plan; attribution images on your screens are yours to place, after the Hub's own redo.`

### STEP 20 — S10a · the drafts and the two indexes, as text
**FOR NICK:** nothing he notices — the day the gated files exist, filling them is one copy. · **Tier:** POLISH
**Start when:** none — start now.
**Builder:** deepseek · **Builder backup:** qwen · **Checker:** zai · **Checker backup:** sonnet
**Files you may touch:** the six new drafts under `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/` named in §3b's created-files list. **Never** any .md anywhere; never the neeko folder.

**Do exactly this:**
1. `ASSISTANT-FOUNDATION.txt`: the shared rules written once, from `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/ASSISTANT-FOUNDATION-SHAPE.txt` — the four approval classes; the standing grants rendered from standing-auth.json (cited, never copied); act, do not narrate, and never ask twice for something granted; the closing-message format; one source per topic; never wear the person's face; the recipient decides the assistant, never the subject; what happens when the wrong assistant is asked (answer if harmless and in scope, hand over by name otherwise); the three bands of STEP 3 and the ask-which-kind rule.
2. `SKIPPY-RULEBOOK.txt`: cites the foundation; carries only what differs — whom he serves, no fence beyond the four classes, his tools, his voice guide, his surfaces; every behaviour rule the registry entry's inline prompt body in server.js earned, moved here verbatim so STEP 22 can delete the inline body.
3. `SKIPPY-MANUAL.txt`: `projects/personal/skippy-app/SKIPPY-MANUAL.md` and `projects/personal/skippy-app/HANDOFF-SKIPPY-CAPABILITIES.md` folded into one dated current-state manual.
4. `GRACIE-RULEBOOK.txt`: 🔴 SETTLED 2026-09-13 — there is NO health fence. Nick's health record is open to Chantelle and Gracie (his no-firewall ruling 2026-08-15, reaffirmed 2026-09-13 "always has been always will be"); the "never attached for her / resolver refuses it first" sentence was an agent's, not Nick's, and is withdrawn (`projects/ops/rules-registry/closed-topics.jsonl`, topic `gracie-health-fence`). The draft carries the citation of the foundation and the four floor categories (logins · keys · secrets · financials) as her only boundary; nothing else (parked; her voice is Chantelle's call).
5. `CONFIG-AND-SECRET-INDEX.txt`: one page — which store owns which KIND of value (the vault: household logins and keys, listed by `python3 projects/personal/family-vault/vault.py list --caller=skippy`, run by the exerciser and pasted as names only; the local .env of the skippy-app: the Mac program's own; the Fly secrets: the cloud brain's, listed by `fly secrets list`; the machine-local settings env: the identity the MCP servers start with) — and the three places to check, in order, before saying any value is absent. No value, ever.
6. `FILE-INDEX.txt`: where each THING's single file lives — dispatchable subagent → `.claude/agents/` and `ZION/agents/`; product assistant → `projects/ops/agents/<name>/`; role build spec → `projects/ops/agents/*-SPEC.md`; app manual → the app's folder; a plan → the lane's folder; and the rule from brain dump §11: a worker that ACTS AS an assistant is a face, never a new name. Each draft's line 1 names its owner and line 2 its purpose and kind, per FILE-STANDARD.
7. `LEGACY-PROMPT-INVENTORY.txt`: every paragraph of the three persona builders' inline prose in server.js (skippy's `buildSystemPrompt`, `buildGraciePrompt`, `buildNeekoPrompt`) and of STANDING_RULES_GRACIE / STANDING_RULES_NEEKO, one line each — its first sentence verbatim, then `→ <the draft and section it now lives in>` or `→ RETIRED: <reason>` (a retirement is allowed only for a rule the foundation already carries or a rule Nick retired by name; anything else is kept). STEP 22 deletes nothing whose line is missing here.

**DEFINITION OF DONE:** seven drafts exist, each with an owner line and a purpose line, ready to copy the day the file exists, and the inventory maps every legacy paragraph.
**PROOF:** `test $(ls projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/drafts/ | command grep -c txt) -eq 7 && echo PASS` → `PASS`, `head -2` of each shows the owner and the purpose, and the inventory's line count equals the paragraph count of the five legacy blocks (the builder prints both counts) · **FAILS IF:** fewer than seven, a draft whose first two lines lack them, an inventory line count below the paragraph count, or any value that grants access anywhere in the index draft.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 21 — S10b · the Neeko merge
**FOR NICK:** a search for Neeko finds the Hub worker, because it carries Neeko's name and lives in Neeko's folder. · **Tier:** POLISH
**Start when:** all three hold — `ls .claude/agents/ | command grep -c neeko-hub-worker.md` → 1 and `ls projects/ops/agents/neeko/ | command grep -c LEARNINGS.md` → 1 (the two empty files Nick creates at the keyboard, because the clone-mirror gate refuses an agent creating either), and no session is inside the worker's files right now: `git status --short .claude/agents/hub-ops-manager.md projects/ops/agents/roster.json projects/ops/agents/hub-ops-manager/LEARNINGS.md` prints nothing, and `find ~/.claude/projects -name "*.jsonl" -mmin -60 -print0 | xargs -0 command grep -l "hub-ops-manager"` prints nothing (no live transcript touched the worker in the last hour).
**Builder:** zai · **Builder backup:** deepseek · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** `projects/ops/agents/roster.json` (the entry at line 1512: name and slug → `neeko-hub-worker`; the LEARNINGS path string at line 1528 → the neeko folder); the empty `neeko-hub-worker.md` under `.claude/agents/` (filled by `python3 projects/ops/agents/build_agents.py`); the empty `LEARNINGS.md` under the neeko folder (filled with the content of `projects/ops/agents/hub-ops-manager/LEARNINGS.md`); `projects/ops/agents/hub-ops-manager/LEARNINGS.md` and `.claude/agents/hub-ops-manager.md` (each gets a first line `SUPERSEDED BY: <the new path>` and nothing else changes — never deleted); `projects/ops/agents/neeko/AGENT.md` (one line naming its worker face and where it lives); any other file naming `hub-ops-manager` in an executing line, found by `command grep -rln "hub-ops-manager" projects/ops .claude ZION`. **Never** the three Hub task prompts — measured 2026-09-12, they name neither the worker nor its memory file, so there is nothing to repoint; task names in the registration pack (`hub-ops-hourly-audit` and siblings) are registration matters and stay.

**Do exactly this:**
1. Edit the roster entry; run `python3 projects/ops/agents/build_agents.py` and confirm its output has a `wrote` line ending in the worker's new file name; read `build_agents.py` line 634 first to confirm the sweep does not delete the superseded file (if it would, exclude it by the roster's `owned_globs` and say so in the commit).
2. Copy the LEARNINGS content; add the SUPERSEDED BY lines; add the worker-face line to the neeko AGENT.md; the worker's file states it acts in Neeko's name and is not a second assistant.
3. Repoint every other executing reference; one commit.

**DEFINITION OF DONE:** nothing executable still dispatches or reads `hub-ops-manager`; the worker carries Neeko's name; nothing was deleted.
**PROOF:** `command grep -rln "hub-ops-manager" projects/ops/agents/roster.json .claude/agents projects/ops/scheduled-rebuild/registration-pack ZION/agents | command grep -v "^.claude/agents/hub-ops-manager.md$" | command grep -v "^projects/ops/agents/hub-ops-manager/"` prints nothing (only the two superseded files, excluded by PATH, may still carry the name), and `ls .claude/agents/hub-ops-manager.md projects/ops/agents/hub-ops-manager/LEARNINGS.md` → both still present · **FAILS IF:** any other file carries the name, or either superseded file is gone.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

**Handoff:** the moment this step closes, post one dated line into `projects/ops/scheduled-rebuild/HANDOFF.md`: `STEP 21 (ASSISTANTS) closed <date> — the Hub worker is neeko-hub-worker; register the three hub-ops tasks against that name.`

### STEP 22 — S10c · the births, and the last hard-coded copies leave
**FOR NICK:** every one of the three assistants is defined in its own folder, in the same shape, and the brain reads only that — nothing in the code defines them any more. · **Tier:** POLISH
**Start when:** STEP 20 closed, and `ls projects/ops/agents/ | command grep -c ASSISTANT-FOUNDATION.md` → 1 and `ls projects/ops/agents/skippy/ | command grep -c "RULEBOOK.md\|MANUAL.md"` → 2 (created by Nick at the keyboard, or by an agent after `node projects/ops/skippy-jobs/lib/md-gov-kill-switch.mjs status` reports the documentation gate off — check the status command, never a banner); the Gracie rulebook is filled only if `ls projects/ops/agents/gracie/ | command grep -c RULEBOOK.md` → 1, and its absence does not hold this step.
**Builder:** zai · **Builder backup:** deepseek · **Checker:** sonnet (the live half) · **Checker backup:** qwen
**Files you may touch:** the newly existing foundation, Skippy rulebook and manual, and Gracie rulebook (filled from their drafts, one copy each); `projects/personal/skippy-app/SKIPPY-MANUAL.md` and `projects/personal/skippy-app/HANDOFF-SKIPPY-CAPABILITIES.md` (a first line `SUPERSEDED BY: <the new manual's path>`, never deleted); `projects/ops/agents/neeko/NEEKO-RULEBOOK.md` (one line citing the foundation, its restated shared rules replaced by the citation); `projects/personal/skippy-app/skippy-code-publish/server.js` (delete `STANDING_RULES_GRACIE` and the inline persona prose inside skippy's `buildSystemPrompt` and `buildGraciePrompt`, which the loader now supplies — each builder returns only the composed documents plus its tools guide and grants block); the two index files if Nick approved their birth. **Never** `standingGrantsBlock()`, the voice guide, the brains branch, a paragraph the inventory (STEP 20 draft 7) does not map.

**Do exactly this:**
1. Fill each existing file from its draft; add the SUPERSEDED BY lines; the neeko rulebook cites the foundation.
2. One route-build edit to server.js: delete the STANDING_RULES_GRACIE string and the two builders' inline prose; `node --check`; commit; publish. Before the deletion the builder confirms every paragraph being deleted has a line in the inventory; the checker samples five inventory lines marked as kept and finds each first sentence in the named draft or file.
3. Run the loader test (its skipped assertion now runs); then the exerciser repeats STEP 1's marker read-back for Skippy and Gracie with a new dated marker in each face's rulebook.

**DEFINITION OF DONE:** F1 for all three faces and F9's folders — the foundation and each face's folder exist in the neeko shape, and server.js holds no persona prose and no standing-rules copy.
**PROOF:** `command grep -c "STANDING_RULES_" projects/personal/skippy-app/skippy-code-publish/server.js` → 0 and `node projects/personal/skippy-app/skippy-code-publish/_test-assistant-docs.mjs` → `ALL PASS` with no SKIPPED line, and the exerciser's marker evidence for both faces · **FAILS IF:** any count above 0, a SKIPPED line, or a marker not read back within 35 minutes.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 23 — S11 · the floor is four categories in the routing code
**FOR NICK:** the routing code now says what the rule says — only logins, credentials and keys, government ids, and card, bank-account and routing numbers stay home; dollar figures, family, kids, health and client material may go to cheap models. · **Tier:** POLISH (TOP — a wall edit)
`NICK-ASKED: opus — "session limit is going to hit us if we drive everything on fable - why dont you pull back to opus and sonnet where reasonable for build but focus the UI on fable" (Nick, 2026-09-05)`
**Start when:** none — start now.
**Builder:** opus · **Builder backup:** fable · **Checker:** sonnet · **Checker backup:** qwen
**Files you may touch:** `projects/personal/skippy-app/skippy-code-publish/lib/grunt-lane.mjs` and `projects/personal/skippy-app/lib/grunt-lane.mjs` (both copies, same change); new `projects/personal/skippy-app/skippy-code-publish/_test-floor-four-categories.mjs`; new `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/evidence/step23-red.txt`. **Never** `projects/ops/lib/vendor-fence.mjs` (the FLOOR list it holds is already the four; read it, do not edit it), any vendor's base URL.

**Do exactly this:**
1. Red first: the test imports each copy and asserts the target below for THAT copy; run on today's code; save the FAIL lines to `step23-red.txt`. The two copies are not the same today and get different edits:
   - the cloud copy (`skippy-code-publish/lib/grunt-lane.mjs`, one old class at line 67): adopt the Mac copy's class split verbatim — `CLASS_CODE`, `CLASS_HEALTH` with `CLASS_PERSONAL = CLASS_HEALTH`, `CLASS_PRIVATE`, `DATA_CLASSES` — and the Mac copy's per-provider `accepts` lists (deepseek: public, code, health; the others as the Mac copy has them; only anthropic-haiku accepts private);
   - both copies: `CLASS_PRIVATE`'s definition and comment name exactly the floor — logins · credentials, tokens and keys · government IDs · card, bank account and routing numbers (the value shapes `projects/ops/lib/vendor-fence.mjs`'s FLOOR already scans for; a dollar figure, an invoice or a rate is NOT private) — with Nick's words and date (MACHINE-RULES data-floor ruling; Nick, 2026-09-12: "the current setup is just, like, logins, Social Security numbers, credit card numbers, bank account numbers, that kind of stuff … everything else is free to pass back and forth — the kids stuff, the Chantelle stuff"); the classifier assigns family, the kids, Chantelle, client material and health-marker text to `CLASS_PERSONAL`, never to `CLASS_PRIVATE`; the Mac copy's note at line 247 ("the kids and Chantelle are still CLASS_PRIVATE") is replaced by that rule, not annotated;
   - the assertions: a text carrying only family, kids, Chantelle, health-marker or client material classifies personal; a text carrying a login, a key shape, a card number or a bank account number classifies private; the absent-class default stays private; every `private → <cheap vendor>` case in the file's own selftest still BLOCKs.
2. Make the change in both copies; keep the file's own trap — a vendor accepting private data must have a literal base URL in the array, never an env var (the Mac copy's selftest block and the `.env cannot widen` case must still BLOCK); the egress scan (`grunt-egress-scan.mjs`) is not touched.
3. Run the test green on both copies; one scoped commit in the nested repo and one in the workspace; publish the brain.

**DEFINITION OF DONE:** F10 — the routing code enforces exactly the four-category floor, proven by a test that was red against the old code.
**PROOF:** `node projects/personal/skippy-app/skippy-code-publish/_test-floor-four-categories.mjs` → `ALL PASS` on both copies with `step23-red.txt` on file, and the file's own selftest still reports every `private → <cheap vendor>` case BLOCK · **FAILS IF:** any assertion fails, the red file is absent, or a cheap vendor accepts private data.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 24 — S12 · Larry's first standing pass
**FOR NICK:** the one-file-per-thing sweep is running in the background, and its findings reach him in batches to approve — nothing is deleted by an agent. · **Tier:** POLISH
**Start when:** `node projects/ops/skippy-jobs/lib/md-gov-kill-switch.mjs status` reports the documentation gate off, or an approved ticket for the upkeep board exists (it is a governed .md); until then the pass runs and its findings wait in this lane's `evidence/` as text.
**Builder:** larry (sonnet), invoked by name with the Agent tool — sonnet is his own pinned model on Nick's words ("doesnt need a ton of reasoning so it should run on sonnet", Nick, 2026-08-20, `ZION/agents/larry.md`), never an override · **Builder backup:** opus · **Checker:** qwen · **Checker backup:** deepseek
**Files you may touch:** `projects/ops/agents/upkeep-board/larry-lane.md`, `projects/ops/agents/upkeep-board/larry-questions.md`, `projects/ops/agents/upkeep-board/larry-patterns.md` — written by the driver's `se-recorder` from Larry's report, since Larry holds no Write. **Never** any file outside that folder; never a deletion, move or archive.

**Do exactly this:**
1. Dispatch `larry` with, as context: brain dump §10 of `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/BRAIN-DUMP-2026-09-12-FOR-CODEX-AND-FABLE.txt` (the three wrong absence claims and the three gaps), the must-not-sweep list from the same section (cited evidence and proofs, Nick's verbatim files, git history, the four sources, any live spec, archive folders), and the instruction to read his own previous pass first.
2. Findings grouped by THING (`THING: <name>` header per group), each with the single file, which others are copies, which differ and why, and one recommendation line; the driver's recorder prepends them to `larry-lane.md` under a dated pass header; questions to `larry-questions.md`.
3. The weekly registration (`projects/ops/scheduled-rebuild/registration-pack/larry-weekly-pass.prompt.md`) stays Nick's keyboard job and is listed in `PROGRESS.txt`.

**DEFINITION OF DONE:** F9's Larry item — the pass's findings, grouped by THING, are on the upkeep board under today's date.
**PROOF:** `test $(command grep -c "^THING:" projects/ops/agents/upkeep-board/larry-lane.md) -ge 5 && echo PASS` → `PASS`, and `head -3 projects/ops/agents/upkeep-board/larry-lane.md` shows the pass header with the date of the run · **FAILS IF:** the test prints nothing, or a finding lacks its single-file line and recommendation.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once, and open two findings to confirm each names the single file, the copies and the recommendation. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 25 — S13 · the cache guard after every publish
**FOR NICK:** nothing he notices — it stops the slow, expensive Skippy from ever coming back unnoticed. · **Tier:** POLISH
**Start when:** none — start now.
**Builder:** zai · **Builder backup:** deepseek · **Checker:** qwen · **Checker backup:** sonnet
**Files you may touch:** new `projects/personal/skippy-app/skippy-code-publish/ops/test-cache-guard.mjs`; `projects/personal/skippy-app/skippy-code-publish/fly-publish.mjs` (one post-publish call to the guard, after the existing health check). **Never** `ops/spend-meter.cjs` (read its ledger reader; do not change it), the transport.

**Do exactly this:**
1. The guard reads the brain's `/api/version` build, signs in as Nick, asks one tool-using question carrying a unique tag ("what's running? [cache-guard <ISO time>]") in a fresh conversation, then reads the ledger rows of THAT conversation (matched by the tag or the conversation id, never "the newest row" — another session's round must not be able to pass this) through `ops/spend-meter.cjs`'s own reader (from the box, the way the publish script already reads the health endpoint), and prints `build <id> · cacheRead <n> > 100000 on this request's tool round · PASS` or `FAIL`; `--expect-fail --fixture <file>` runs the same check on a fixture row with cacheRead 3289 and must exit 1.
2. Wire it as the last line of the publish script's post-publish check so a publish that breaks caching is reported red at once.

**DEFINITION OF DONE:** after every publish, a tool round shows cacheRead above 100,000, and the guard goes red on a fixture that does not.
**PROOF:** `node projects/personal/skippy-app/skippy-code-publish/ops/test-cache-guard.mjs` → `PASS`, exit 0, and `node projects/personal/skippy-app/skippy-code-publish/ops/test-cache-guard.mjs --expect-fail --fixture <the fixture the test ships>` → exit 1, and `command grep -c "test-cache-guard" projects/personal/skippy-app/skippy-code-publish/fly-publish.mjs` → 1 · **FAILS IF:** the live check fails, the sabotage passes, or the publish script does not call it.

**If the check fails:** the builder fixes and re-checks the named failure until it passes. If this step cannot close from this machine: one line to the overseer naming the ONE missing thing, then the next step whose inputs exist.

**Checker's job:** re-run the PROOF yourself, once. PASS closes the step. Do not accept the builder's pasted output; do not summon anyone else.

### STEP 26 — S14 · postmortem and FINISH LINE sign-off
**FOR NICK:** the project is done to its own finish line, and the lessons are written down once. · **Tier:** FINAL
**Start when:** STEPS 1–25 each carry a CLOSED line in `PROGRESS.txt`.
**Builder:** fable (the driver) · **Builder backup:** opus · **Checker:** opus · **Checker backup:** gpt-6-astra
**Files you may touch:** this file's SUMMARY and STEPS sections; `PROGRESS.txt`; `.claude/skills/plan/references/failure-registry.md` (append only, four-column rows for every failure this build met). **Never** a step block; never a proof re-run.

**Do exactly this:** read the closed proofs for F1–F10 once, in §6 order; write the postmortem into SUMMARY (what failed, what was confused, what to keep — "nothing worth extracting" is a good answer) and one line `POSTMORTEM WRITTEN <date>` into `PROGRESS.txt`; append the registry rows; the checker reads the same proofs once and confirms each F item maps to a closed step.

**DEFINITION OF DONE:** every FINISH LINE item is signed off from a closed step's proof, and the postmortem is in this file.
**PROOF:** `test $(command grep -c "^  STEP [0-9]* CLOSED" projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/PROGRESS.txt) -ge 25 && command grep -q "POSTMORTEM WRITTEN" projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/PROGRESS.txt && echo PASS` → `PASS`, and the checker's own ten-line table mapping F1–F10 each to a CLOSED step line · **FAILS IF:** the test prints nothing, or an F item with no closed step behind it.

**If the check fails:** the driver names the open item and drives it; nothing is re-run that was closed.

**Checker's job:** read the closed proofs once; confirm the mapping; PASS closes the plan.

## 4 · Regret Check (the registry failures this build is actually exposed to)

| Failure mode (registry entry) | The measure in THIS plan that prevents it | Where it lives (section / artifact / gate) |
|---|---|---|
| An absence was asserted without opening the store that would hold it | the config-and-secret index and the check-before-absence order (STEP 20); every live claim in "Already true" names the file and line it was read from; every proof names the store it opened; the spike re-run polls the live per-machine file the first run missed | Already true; STEP 2; STEP 20; §3b proofs |
| Work was written to a queue no reader ever visits | a confirmed memory is read back through the engine door in a fresh conversation after a restart; the code-agent queue is drained by a registered job and listed by work-watch from that same file | STEP 5, STEP 11, STEP 12 |
| A decision settled once re-opened elsewhere, or two copies of a rule disagreed | the persona and standing-rules copies are deleted in the same change that loads the documents; the grep count of zero is the proof; the floor edit lands on both grunt-lane copies in one step; Nick's "make it ask" replaces the brief's Sonnet default on the sheet, in the contract and in the tool | STEP 1, STEP 3, STEP 22, STEP 23; §1a |
| A document, label, or comment was believed over the live system | "three live scheduled tasks" turned out unregistered — every live claim in this plan carries the command that measured it; the first spike's NOT MEASURABLE verdict is kept as such and re-measured, never read as an answer | Already true; STEP 2 |
| A second system was built because the first was invisible | Larry is dispatched by name, never a new cleanup agent; the door extends the relay, the queue and work-watch rather than a new channel; the loader copies voice-guide.mjs | STEP 24, STEP 3, STEP 5, STEP 1 |
| Session rules never reached the subagents doing the work | every brief carries `loadTravelBlock()`, the four classes and the floor; the drain writes the dispatch header into every session it starts | §3b preamble; STEP 5 |
| A check existed that could not fail | every new test is red first with the red run on file (STEPS 7, 10, 13, 23) or carries a sabotage case (STEPS 3, 25); comment lines are excluded from every grep-shaped proof; the pid of 0 that voided the first spike is a named FAIL condition | STEPS 2, 3, 7, 10, 13, 23, 25 |
| Output was delivered somewhere the intended reader never looks | the sender rule is proven by reading the real DM channels back, not by a send receipt | STEP 13, STEP 14 |
| A tool's own description contradicted house reality and won | the "CHIEF-OF-STAFF RULE" text is deleted with its tool; the three bands and the ask rule are Nick's words pasted verbatim | STEP 3, STEP 4 |
| Done was declared before the live surface was checked | every FRONT step with a person-visible outcome closes on a live drive by the exerciser, re-run by a Sonnet checker; the avatars are graded from screenshots of the real slot sizes, never from the files | STEPS 1, 6, 7, 8, 12, 14, 15, 17, 18, 19 |
| An agent's pasted command output was taken as evidence | every evidence file is re-run through `verify-agent-evidence.mjs`; MISMATCH throws the batch out | STEPS 2, 6, 12, 14, 19 |

## 5 · Topology and roles
- **OVERSEER-AUTHORITY:** none named. **The four approval classes (money leaving · credential rotation · irreversible destruction · a message sent as Nick) and the floor (logins · credentials, tokens and keys · government IDs · card, bank and routing numbers) never move on the overseer's word.**
- Thread layout: one overseer thread — the Fable driver session; builders and checkers as cheap dispatches; Codex as the peer builder for STEPS 15–17 in its own `codex exec` calls, one at a time; the exerciser for every live drive; Sienna once for the images; Larry once for the pass.
- Overseer: Fable · Workers: GLM 5.3 (zai), DeepSeek V4 Pro, Qwen 3.8 through `projects/ops/cheap-task.mjs` and `projects/ops/route-build.mjs`; exerciser (haiku); Sonnet checkers where a cheap checker cannot sign in; Opus for the one wall edit · Cap: 8 per session, ~40 machine-wide.
- State files location: `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/PROGRESS.txt` (state, assumptions, questions, Nick's keyboard items, plan-change lines) and `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/PROGRESS-CONTRACT.json` (the status projection); no STATE.md, QUESTIONS.md, ASSUMPTIONS.md or PLAN-CHANGES.md while the documentation gate refuses them.
- **Board card id:** nt-20260912-175831-7392
- **Artefact consumers:** the loader → the three prompt builders (STEP 1); the spike re-run's evidence → STEP 5's work-watch change; the queue rows → the drain → work-watch → whats_running → Nick's ear; confirmed memories → the drain → capture_narrative → the family narrative engine → Skippy's answers; the drafts → the born files (STEP 22); the avatars → Sienna's grade → Nick's upload → Slack, then the VOICE and HUB lanes; the decision file → Nick's one line → STEP 17; Larry's findings → the upkeep board → Nick's batches; every evidence file → its Sonnet checker through `verify-agent-evidence.mjs`.
- **Write-contention (parallel lanes in a shared checkout):** this lane writes `projects/personal/skippy-app/skippy-code-publish/` (main) — server.js, lib/, ops/, its tests and fly-publish.mjs; `projects/ops/skippy-jobs/jobs/skippy-brain-push.mjs` (ADDS only), `projects/ops/skippy-jobs/jobs/work-watch.mjs` (one row source, STEP 5), two new jobs and their tests, one runner row each; `projects/personal/family-app/js/panel.js`, `js/voice.js`, `functions/api/thread-reply.js` and the voice-session routes (STEPS 15–17 only, Codex); the existing files under `projects/ops/agents/skippy/`, `gracie/`, `neeko/`; `projects/ops/agents/roster.json`; `projects/personal/skippy-app/lib/grunt-lane.mjs` (STEP 23). Never: the brains/2026-09-09 branch (HEALTH-ONE-DOOR lane), `projects/ops/life-os/REGROUP-2026-09-08/plans/VOICE/` files other than a dated PROGRESS line, the engine, the three Hub task prompts, any deletion. Checkout proven writable before lanes open: the timestamped first write to `PROGRESS.txt` in this folder (2026-09-12T18:11:54Z). `PROGRESS.txt` is written only by the overseer.

**Per-stage topology — counts DECLARED at plan time (machine-gated: a number in every row):**

| Stage | Overseer | Sub-overseers | Workers |
|---|---|---|---|
| FRONT (steps 1–19) | 1 | 0 | 6 |
| POLISH (steps 20–25) | 1 | 0 | 4 |
| FINAL (step 26) | 1 | 0 | 0 |

**The walk-away contract — a stranger resumes the drive from files alone:**
- **STATE FILE:** `projects/ops/life-os/REGROUP-2026-09-08/plans/ASSISTANTS/PROGRESS.txt`
- **HEARTBEAT ROW:** none yet
- **MORNING-REPORT LINE:** "Assistants — FRONT n of 19 · polish n of 6"

## 6 · Evals — what "working" means, decided now

| Capability | Check (exact command or procedure) | Pass looks like |
|---|---|---|
| F1 the documents drive the live answer, all three faces, no copies in code | STEP 1's and STEP 22's proofs: the loader test and the marker read-back, `command grep -c "STANDING_RULES_"` on server.js | ALL PASS; three markers read back within 35 minutes; grep 0 |
| F2 a spoken dispatch starts a real session of the named kind, asks when no kind was named, no Cowork | STEP 6's evidence re-run and STEP 4's grep | every pair MATCH, elapsed ≤ 120, kind named, asked for kind: yes; every grep line ends :0 |
| F3 ten requests done for real, no false hand-off claims | STEP 7's gate run and the claim-guard test | exit 0, ten delivered and read back; ALL PASS |
| F4 what's running in ≤ 15 s on five of five | `node projects/personal/skippy-app/skippy-code-publish/ops/time-whats-running.mjs` | `5/5 under 15000 ms` with matching counts |
| F5 a spoken, confirmed fact comes back next day; a credential shape is refused | STEP 12's evidence re-run by the checker the next calendar day | nonce present, refusal names the floor, invented value never echoed |
| F6 six live identity checks; every message for Nick from Skippy | STEP 14's evidence and STEP 13's test | six `RESULT: PASS`; `0 send sites to Nick via the Gracie token` |
| F7 voice into a running thread, spoken back within 30 s as a sentence | `node projects/personal/family-app/_test-thread-voice.mjs` on the live app | `spoken in <n> s`, n ≤ 30, `sentence clean` |
| F8 the decision with numbers, Nick's choice, first audio under 2 s on ten | `node projects/personal/family-app/_bakeoff-voice-a.mjs --report` then `--live`; the `VOICE A CHOICE:` line in PROGRESS.txt | both candidates at 10 samples; the choice line; `10/10 first audio < 2.0 s` |
| F9 one shape, three faces (Gracie: front door + fence only); the worker renamed; two indexes; Larry's pass; nothing deleted | STEP 22's directory listings, STEP 21's grep, STEP 20's seven drafts and the two index births, STEP 24's board grep, and `git log --diff-filter=D --since=2026-09-12` over the lane's fences | folders in the neeko shape (Skippy, Neeko), Gracie's two files; grep 0; seven drafts and two indexes; ≥ 5 THING groups; no deletion by an agent |
| F10 the four-category floor in the routing code, red first | `node projects/personal/skippy-app/skippy-code-publish/_test-floor-four-categories.mjs` with `step23-red.txt` on file | ALL PASS on both copies; the red file exists |

## If you get stuck (all steps)

Before writing "blocked": (1) re-read the step's START WHEN line — most "stuck" is a misread gate, (2) try a concrete workaround, (3) write one line to the overseer naming the ONE missing artefact. Then keep working every other step whose inputs exist. Never idle on a blocker; never end a turn waiting on a background result. A refused governed .md write, a refused new file under a face's name, or a missing Slack icon is paperwork or a keyboard item for `PROGRESS.txt` — never a reason a FRONT step waits.

## Your loop

Every pass: every step whose START WHEN inputs exist and which is not yet CLOSED is running, up to the cap → each builder runs its own PROOF, hands to its checker → PASS closes it, FAIL loops it → repeat until the FINISH LINE is proven. Steps 1, 2, 3, 8, 9, 10, 13, 15, 16, 18, 20, 23 and 25 have no inputs and start on the first pass.

## SUMMARY — a few plain-English lines, read by the status generator

**TRUE NOW (re-established 2026-09-14 by a regroup pass that re-ran each proof first-hand; 26 steps, 15 re-run, 24 minutes).** Thirteen separate pieces of this work are genuinely running today, each re-checked by a session that did not build it: the three assistants reading their own instruction documents, the tool that starts real background work from a spoken sentence, the memory consent gate, the "what's running" speed fix that now answers in under six seconds, and the rule deciding which text may reach cheap outside models. Two things are not true and were recorded as if they might be. The overnight worker still carries its retired name in fifteen live places. And the final sign-off cannot honestly be called: its own headcount searches for a word in a log rather than counting anything, and the real figure is 21 of the required 25 steps finished.

**LEFT.** Eleven steps were not re-driven in this pass because each needs a live multi-minute session, a signed-in browser, or the voice lane's testing instrument: whether a background helper appears on the tracking screen, the spoken start-a-worker demonstration, ten everyday spoken requests, the live memory round-trip, the six sign-in checks separating Nick's assistant from Chantelle's, three voice speed measurements, the three assistant pictures and their upload into the messaging app. Two steps are open with real defects: the rename, and the sign-off's broken headcount.

**NEXT PICKUP.** Finish renaming the retired overnight worker everywhere it still appears, then replace the sign-off's word-search with an honest count of finished steps before anyone calls this done.

## SUMMARY

**2026-09-12** — Tonight the household web app got the voice Nick chose: a front desk that starts talking in about one second and asks the main assistant program behind the scenes for anything that needs his real records, measured live on ten spoken turns on 2026-09-12 at 22:44 UTC. The main assistant program's per-turn load was measured with Anthropic's own token counter at 124,932 tokens, two thirds of it Nick's health record injected whole on every turn; the health engine session then made the health engine the only door for health answers and cut that block down to the six safety rules about his body, so each turn now reads 24,555 tokens from cache instead of 107,691. The remaining per-turn cost is the 42 tool definitions, about 17,000 tokens re-sent on every turn because the subscription route this program uses to reach the model can only remember one fixed block of text between turns; trimming those definitions is waiting on Nick's yes. Earlier tonight: when Nick said out loud that a Sonnet coding agent should write a test file, a real Sonnet coding session started on his Mac Studio within two minutes of his confirming tap and wrote that file, so the spoken start-a-coding-agent door works end to end; Chantelle's assistant now runs on the main assistant program in the cloud after Nick approved setting her sign-in secret; all six live tests that Nick and Chantelle each meet their own assistant and never the other's pass; and Nick chose three AI-style faces for the Slack profile pictures, held until he says go. Two items need Nick: a yes to trimming the tool descriptions, and running one command from his own terminal to switch off the safety lock that stops agents editing the assistants' rule documents until 2026-09-13 at 12:44 UTC, so six of those documents can be created and one wrong sentence removed from the file that describes Chantelle's assistant.

**2026-09-12** — Six pieces of the project that turns Nick's three AI assistants into one clean system are finished and independently checked. One: a spoken or typed yes from Nick keeps a fact the assistant offered to remember, exactly as pressing the Keep button in the app does. Two: the phone's remember-this command will stop being refused, because the assistant's cloud program now hands each memory to the Mac program, where the program that checks for passwords and keys before saving actually runs. Three: every publish of the assistant's cloud program now runs a guard that proves the assistant's standing instructions are being reused between turns instead of re-sent. Four: it is proven, and guarded by a test, that no automated job messages Nick from the Slack account of Chantelle's assistant. Five: the routing code that decides which text may go to cheap AI models now says exactly what Nick ruled: family, kids, health and client text travel; only logins, keys, government ids and card, bank and routing numbers stay home, and no old wording is kept anywhere in that file. Six: it is measured twice that a coding session started in the background by a job never appears in the list of running sessions the assistant reads aloud, so the new start-an-agent command will register each session it starts. Pieces one, two, three and five live in the cloud program's code and are being published to Nick's phone right now; four and six needed no publish. Still building: the module and wiring behind the spoken start-an-agent command, the seven draft documents describing the three assistants (five of seven written), the three avatar images redrawn to exactly the shapes an independent design review specified, and the second run of the measured comparison between two ways of building a spoken back-and-forth conversation with the assistant, this time with network access allowed.

**2026-09-12** — Six pieces of the project that turns Nick's three AI assistants into one clean system are finished and independently checked. One: a spoken or typed yes from Nick keeps a fact the assistant offered to remember, exactly as pressing the Keep button in the app does. Two: the phone's remember-this command will stop being refused, because the assistant's cloud program now hands each memory to the Mac program, where the program that checks for passwords and keys before saving actually runs. Three: every publish of the assistant's cloud program now runs a guard that proves the assistant's standing instructions are being reused between turns instead of re-sent. Four: it is proven, and guarded by a test, that no automated job messages Nick from the Slack account of Chantelle's assistant. Five: the routing code that decides which text may go to cheap AI models now says exactly what Nick ruled: family, kids, health and client text travel; only logins, keys, government ids and card, bank and routing numbers stay home, and no old wording is kept anywhere in that file. Six: it is measured twice that a coding session started in the background by a job never appears in the list of running sessions the assistant reads aloud, so the new start-an-agent command will register each session it starts. Pieces one, two, three and five live in the cloud program's code and are being published to Nick's phone right now; four and six needed no publish. Still building: the module and wiring behind the spoken start-an-agent command, the seven draft documents describing the three assistants (five of seven written), the three avatar images redrawn to exactly the shapes an independent design review specified, and the second run of the measured comparison between two ways of building a spoken back-and-forth conversation with the assistant, this time with network access allowed.

**2026-09-12** — Four pieces of the project that turns Nick's three AI assistants into one clean system are finished and independently checked. Piece one: a spoken or typed yes from Nick keeps a fact the assistant offered to remember, exactly as pressing the Keep button in the app does. Piece two: the phone's remember-this command will stop being refused, because the assistant's cloud program now hands each memory to the Mac program, where the program that checks for passwords and keys before saving actually runs. Piece three: every publish of the assistant's cloud program now runs a guard that proves the assistant's standing instructions are being reused between turns instead of re-sent, which is what made the assistant slow and expensive on 2026-09-12. Piece four: it is proven, and guarded by a test, that no automated job messages Nick from the Slack account of Chantelle's assistant; the two jobs that once did were repointed in July, so anything for Nick comes from the Slack account of Nick's own assistant. Pieces one, two and three live in the cloud program's code and reach Nick's phone only when that program is next published, which happens once the pieces still editing the same code file land; piece four is already in effect because nothing needed changing. A fifth piece, the change that lets family, kids, health and client text go to cheap AI models while only logins, keys, government ids and account numbers stay home, passed its checker on everything except one wording fix, which has since been made and is being re-checked. Still building: the wiring that makes the cloud program read each assistant's own rule documents, the module behind the spoken start-an-agent command, the seven draft documents describing the three assistants, the three avatar images redrawn to exactly the shapes an independent design review specified, and the second run of the measured comparison between two ways of building a spoken back-and-forth conversation with the assistant, this time with network access allowed.

**2026-09-12** — Three pieces of the three-assistants build are finished and independently checked. A spoken or typed yes from Nick keeps a fact the assistant offered to remember, exactly as pressing the Keep button in the app does. The phone's remember-this command will stop being refused: the assistant's cloud program now hands a memory to the Mac program, which holds the credential checker, instead of refusing because the checker is missing in the cloud. And every time the assistant's cloud program is published, a guard asks it one tagged question and reads that request's own cost rows from the cloud server, so the slow-and-expensive problem of 2026-09-12, where the prompt cache was thrown away on every turn, can never come back unnoticed; today's live reading shows the cache being reused. None of the three is on the phone yet: one publish of the cloud program carries all three, and it waits for the routing-wall change that is finishing now. Seven other pieces are in flight: the code that lets each assistant read its own rule documents, the module behind the new spoken command that starts a coding agent, the routing-wall change that lets family and health text go to cheap AI models while only logins, keys, government ids and account numbers stay home, the seven draft documents describing the three assistants, the three avatar marks drawn in code, the measured comparison of ways to build a real spoken conversation with the assistant (OpenAI's Codex tool, second attempt with network allowed), and the read-only file auditor's first pass, which came back with four grouped findings that wait on the rule blocking new documents until tomorrow midday.

**2026-09-12** — The plan for making Nick's three AI assistants (Skippy for Nick, Gracie for Chantelle, Neeko for the business team) into one clean system passed its three reviews today: an automatic checker script, a cold attack by OpenAI's Codex tool, and a fresh reviewer session, and every point they raised is fixed in the plan. Building has started and the first piece is finished and independently checked: a spoken or typed yes from Nick now keeps a fact the assistant offered to remember, exactly as pressing the Keep button in the app does. Nine other pieces are being built at the same time: the code that lets each assistant read its own rule documents; the module behind the new spoken command that starts a coding agent; the change that lets family, kids and health text go to cheap AI models while only logins, keys, government ids and account numbers stay home; a guard that catches the assistant's prompt cache silently breaking again; the seven draft documents describing the three assistants; the three avatar images; the measured comparison of two ways to build a real spoken conversation with Skippy, done by OpenAI's Codex tool; the first pass of the read-only file auditor; and the fix so the phone stops refusing remember-this. One rule of the road surfaced: the cheap models are not allowed to read the assistant's main code file because it contains a credential-shaped line, so edits to that file run on Anthropic's Opus model, one at a time, on Nick's own words from 2026-09-05.

**2026-09-12** — A build plan now exists for making Nick's three AI assistants (Skippy for Nick, Gracie for Chantelle, Neeko for the business team) into one clean system, and an automatic checker script confirms it is complete. Nick answered three open questions today: a spoken yes is enough to keep a memory; the assistant asks which kind of agent to start when Nick does not say which; and this project's task card at hub.heroesandsidekicks.io is due 2026-09-13. OpenAI's Codex tool is reviewing the plan for ambiguities, and a separate fresh reviewer session reads it cold next; building begins after both reviews. One measurement is already done: a Claude Code work session started in the background by a scheduled job does not appear in the list of running sessions that Skippy reads aloud, so the new start-an-agent feature must record each session it starts.

## STEPS

<!-- The live status checklist, read by `status-regen.mjs` / `project-status-page.py`. The heading
     above must be exactly "## STEPS" with nothing else on the line. -->

> One numbered line per STEP block above, same numbers. CURRENT STATE, rewritten in place — the STEP blocks say what to do; this section records what has been proven.

```
1. The brain reads the documents — 100%
   REGROUP 2026-09-14 · PROVEN — the document-loader unit test prints ALL PASS. The live 35-minute read-back was not re-driven inside the time box, so the sentence-reaches-the-phone half is still owed. Evidence: regroup-audit-2026-09-14.txt
   DEFINITION OF DONE: the document loader is live for all three assistant faces, Neeko's hard-coded prompt copies are gone from the cloud brain's code, and a marker sentence added to each assistant's AGENT.md file is read back by the live cloud brain within 35 minutes
   PROOF: node projects/personal/skippy-app/skippy-code-publish/_test-assistant-docs.mjs prints ALL PASS; grep STANDING_RULES_NEEKO in server.js counts 0; the live-drive evidence file shows the marker sentence in all three assistants' answers
   VERIFIED: not yet — the plan was written on 2026-09-12 and its independent cold review is in progress
2. Spike re-run: a job-started headless session in the live running list — 100%
   REGROUP 2026-09-14 · NOT RE-RUN in this pass — keeps its 2026-09-12 verdict, which was independently checked at the time. Evidence: regroup-audit-2026-09-14.txt
   DEFINITION OF DONE: the evidence file states, from literal output with a real pid, whether a job-started headless session appears in the live per-machine running list within 150 s
   PROOF: evidence/step2-spike-rerun.txt (pid 87390): SESSIONS_DIR does not appear, SOCKET does not appear, TRANSCRIPT written, WORK_THREADS_LIVE not listed; the verify tool reproduced the in-tree pairs and the checker re-ran the four outside-tree ones by hand
   VERIFIED: 2026-09-12 (100%, checked by the haiku exerciser in its own session; evidence/step2-check.txt)
3. The door: dispatch_to_code_agent, the three bands, the relay route — 100%
   REGROUP 2026-09-14 · PROVEN — the dispatch tool test prints ALL PASS, including the live cases. Evidence: regroup-audit-2026-09-14.txt
4. Cowork out, with its replacement in — 100%
   REGROUP 2026-09-14 · PROVEN — the retired worker's word count is zero everywhere it is required to be zero. Evidence: regroup-audit-2026-09-14.txt
5. The Mac drain that starts the session, and work-watch lists it — 100%
   REGROUP 2026-09-14 · PROVEN — the drain test passes 71 of 71 checks, matching this plan's recorded number exactly. Evidence: regroup-audit-2026-09-14.txt
6. Live from the phone path — 0%
   REGROUP 2026-09-14 · NOT RE-RUN — needs a live multi-minute signed-in session; not reached inside the 25-minute time box. Evidence: regroup-audit-2026-09-14.txt
7. Ten requests done for real — 0%
   REGROUP 2026-09-14 · NOT RE-RUN — waits on the voice lane's testing instrument; not reached inside the time box. Evidence: regroup-audit-2026-09-14.txt
8. What's running in 15 seconds — 100%
   REGROUP 2026-09-14 · PROVEN — five live timing samples all came back under 15 seconds (3.3 to 5.8 seconds). Evidence: regroup-audit-2026-09-14.txt
9. Memory proposals relay to the Mac's scanner — 100%
   REGROUP 2026-09-14 · PROVEN — both memory-filter tests pass in full; the test count has grown since this was recorded and still passes. Evidence: regroup-audit-2026-09-14.txt
   DEFINITION OF DONE: on a box without the credential scanner, a memory proposal is scanned and stored pending on the Mac, never refused and never stored unscanned
   PROOF: node projects/personal/skippy-app/skippy-code-publish/_test-memory-write-filter.mjs prints ALL PASS with the three route cases, and _test-relay-port-scope.mjs stays green
   VERIFIED: 2026-09-12 (100%, checked by the haiku exerciser in its own session — 37 of 37 and 25 of 25 re-run first-hand; evidence/step9-check.txt)
10. A yes from the person is the decision — 100%
   REGROUP 2026-09-14 · PROVEN — the consent-gate test passes 123 of 123. Evidence: regroup-audit-2026-09-14.txt
   DEFINITION OF DONE: a person's yes to a named pending fact lands the row in the confirmed file; a model-side affirmation still does not
   PROOF: node projects/personal/skippy-app/skippy-code-publish/_test-memory-gate.mjs prints ALL PASS including the new case, with the red run saved in the plan's evidence folder as step10-red.txt
   VERIFIED: 2026-09-12 (100%, checked by the haiku exerciser in its own session — 44 of 44 PASS re-run first-hand, the old header gone, the red run on file, the commit scoped to three files; evidence/step10-check.txt)
11. Confirmed facts reach the engine door — 100%
   REGROUP 2026-09-14 · PROVEN — the engine-drain test passes 30 of 30 and the syntax check is clean. Evidence: regroup-audit-2026-09-14.txt
12. Live memory proof — 0%
   REGROUP 2026-09-14 · NOT RE-RUN — needs a live memory round-trip; not reached inside the time box. Evidence: regroup-audit-2026-09-14.txt
13. Every message for Nick comes from Skippy — 100%
   REGROUP 2026-09-14 · PROVEN — zero messages to Nick were sent through Chantelle's assistant's channel. Evidence: regroup-audit-2026-09-14.txt
   DEFINITION OF DONE: no job sends a message to Nick through the Gracie token; the test that proves it can go red
   PROOF: node projects/ops/skippy-jobs/_test-sender-rule.mjs prints 0 send sites to Nick via the Gracie token and 4 sites reaching Nick as Skippy, exit 0; a sabotage fixture makes it print 1 and exit 1
   VERIFIED: 2026-09-12 (100%, checked by the haiku exerciser in its own session — re-run on 205 job files, sabotage fixture red, two historical sites confirmed as comments; evidence/step13-check.txt)
14. Six live identity checks — 0%
   REGROUP 2026-09-14 · NOT RE-RUN — six live sign-in checks need a signed-in browser session; not reached inside the time box. Evidence: regroup-audit-2026-09-14.txt
15. Voice B — talk to a running thread, be told when it answers — 0%
   REGROUP 2026-09-14 · NOT RE-RUN — voice timing; not reached inside the time box. Evidence: regroup-audit-2026-09-14.txt
16. Voice A — the architecture decision, measured — 0%
   REGROUP 2026-09-14 · NOT RE-RUN — voice timing; not reached inside the time box. Evidence: regroup-audit-2026-09-14.txt
17. Voice A — the chosen path built — 62%
   REGROUP 2026-09-14 · NOT RE-RUN — voice timing; not reached inside the time box. Evidence: regroup-audit-2026-09-14.txt
   DEFINITION OF DONE: The voice path Nick chose (a realtime front desk that answers at once and asks the main assistant program behind the scenes) speaks its first word in under two seconds on ten of ten spoken samples.
   PROOF: The measurement harness, run live against the deployed household web app on 2026-09-12 at 22:44 UTC, printed ten samples with first audio between 1,041 and 1,992 milliseconds; an independent checker on a separate model is re-running it now.
   VERIFIED: Sixteen of the plan's twenty-six steps are closed with an independent check each (steps 1, 2, 3, 4, 5, 8, 9, 10, 11, 13, 14, 16, 18, 20, 23, 25). Step 17 is built and live with its checker running. Step 6 (a spoken request starts a real coding agent) worked end to end on its second run. Step 15 (talking to a running thread from the Status tab) is deployed with delivery proven live; its spoken half waits for an idle session to answer.
18. Three faces drawn, composited and graded — 0%
   REGROUP 2026-09-14 · NOT RE-RUN — the three assistant pictures; not reached inside the time box. Evidence: regroup-audit-2026-09-14.txt
19. Three faces live in Slack — 0%
   REGROUP 2026-09-14 · NOT RE-RUN — uploading those pictures into the messaging app; not reached inside the time box. Evidence: regroup-audit-2026-09-14.txt
20. The drafts and the two indexes, as text — 100%
   REGROUP 2026-09-14 · PROVEN — all seven planning drafts exist. Evidence: regroup-audit-2026-09-14.txt
21. The Neeko merge — 0%
   REGROUP 2026-09-14 · FAILED — the retired worker name is still referenced live in 15 places. This matches this plan's own record, which never marked the step done. Evidence: regroup-audit-2026-09-14.txt
22. The births, and the last hard-coded copies leave — 100%
   REGROUP 2026-09-14 · PROVEN — the old hard-coded rules text is confirmed gone from the brain's code, for the half not already covered by step 1. Evidence: regroup-audit-2026-09-14.txt
23. The floor is four categories in the routing code — 100%
   REGROUP 2026-09-14 · PROVEN — the sensitive-data routing test passes 126 of 126, matching this plan's recorded number exactly. Evidence: regroup-audit-2026-09-14.txt
   DEFINITION OF DONE: the routing code enforces exactly the floor and nothing else, proven by a test that was red against the old code
   PROOF: node projects/personal/skippy-app/skippy-code-publish/_test-floor-four-categories.mjs prints ALL PASS 126/126 on both copies; each copy's own selftest blocks every private-to-cheap-vendor case; red run in evidence/step23-red.txt
   VERIFIED: 2026-09-12 (100%, checked by the Sonnet verifier in its own session — first check failed on kept old wording, fixed, re-check PASS; evidence/step23-check.txt and step23-recheck.txt)
24. Larry's first standing pass — 100%
   REGROUP 2026-09-14 · PROVEN — the cleanup log carries 7 findings against a required 5 or more. Its date stamp is two days old rather than today's, which is expected: the work was done then. Evidence: regroup-audit-2026-09-14.txt
25. The cache guard after every publish — 100%
   REGROUP 2026-09-14 · PROVEN — the caching check's control genuinely fires when it should fail, so the instrument is sound. Two of three live checks passed; the third failed on a cold start, which is an idle cache rather than a break. Evidence: regroup-audit-2026-09-14.txt
26. Postmortem and FINISH LINE sign-off — 0%
   REGROUP 2026-09-14 · FAILED — this sign-off's own headcount only searches for the word 'closed' anywhere in the log, which is not a count. The real figure is 21 of the required 25 steps finished, which is not enough to sign off. Evidence: regroup-audit-2026-09-14.txt
```

## POSTMORTEM — the 2026-09-14 regroup

**Every failure, every confusion, every blocker, with its concrete example.**

- **A sign-off step that cannot fail is the most dangerous kind of green.** This lane's final step counts completion by searching its own log for the word "closed" rather than counting finished steps. It would have reported success with 21 of 25 done, and with 0 of 25 done. **Keep:** a closing check must be able to return a number and must be tested against a deliberately incomplete state before it is trusted once.
- **Thirteen steps were genuinely working and nobody could see it.** The status sheet carried "not yet" or a bare percentage against work that passes its own tests today. The build was well ahead of its own record. That is a better failure than the reverse, but it still cost this audit its whole time box to discover.
- **The confusion worth naming: a retired name that was removed in one place and left in fifteen.** The step was correctly never marked done, so nothing lied — but the plan also never said which fifteen places, so the remaining work was invisible until someone searched.
- **Eleven steps could not be re-run inside a 25-minute box** because each needs a live multi-minute session, a signed-in browser, or another lane's instrument. That is a real constraint on auditing this lane, not a fault: it should be budgeted at an hour, not 25 minutes, next time.
- **This section's own dated entries above contain one paragraph duplicated verbatim.** Left in place rather than edited out, because a regroup records what is there; it is noted here so a reader is not misled into thinking two separate things happened.
- **What went well, and is worth keeping:** every proven step above was proven by re-running a real test that reports a real count, and several of those counts have grown since they were recorded and still pass. Tests that report totals, rather than printing "OK", are why this audit could confirm anything at all.

Current state PROGRESS.txt

THE THREE ASSISTANTS, ONE CLEAN SYSTEM — progress record. Plain text on purpose: the documentation gate is ON
(node projects/ops/skippy-jobs/lib/md-gov-kill-switch.mjs status → active, expires 2026-09-13T12:44Z), so this lane's
one .md is PLAN.md and everything else lives here as sections. The plan is PLAN.md in this folder; this file is its
STATE companion (RULE 20: written to main at every stopping point; the later commit wins).

WHO IS DRIVING THIS
  Overseer: the Fable driver session that dispatched the planner on 2026-09-12 (it never builds; it unsticks, routes and
  judges). Plan author: Boris (senior-engineer, Fable) on Nick's 2026-09-05 ruling. Cold readers: Codex attacks the plan;
  a fresh Fable session cold-reads it — both dispatched by the driver, neither by the author.
  Lanes re-read this file and PLAN.md at every pass; nothing is carried in memory between passes.
  Checkout proven writable: clock read 2026-09-12T18:11:54Z by the exerciser (date -u); this file and
  PROGRESS-CONTRACT.json were written in the same minute; git showed the folder clean before the write.

LANES (one overseer thread; every builder and checker a dispatch; see PLAN.md §3 and §5)
  BRAIN — steps 1, 22, 25 · DOOR — steps 2–8 · MEMORY — steps 9–12 · FACES — steps 13, 14, 18, 19 ·
  VOICE-CODEX — steps 15, 16, 17 (Codex builds, gpt-5.6-terra) · CLEAN — steps 20, 21, 23, 24 · SIGN-OFF — step 26.
  Steps with no inputs, which start on the first pass: 1, 2, 3, 8, 9, 10, 13, 15, 16, 20, 23, 25.

NEXT PHASE SLICE 1 LIVE (00:45Z, evidence/next-1-profile-and-route-2026-09-13.txt): the Skippy voice profile (853 tokens, ops/build-voice-profile.mjs) and the three-way route are in the front desk — general talk answered directly by the voice model (ten of ten, median useful first audio 1.25 s, zero brain calls), record questions call the gateway once (ten of ten, first word 1.2 s, the brain's answer a median 11.5 s, nine of ten right, one past the 25 s cap); brain f583804 + 20b42b3 + 74a31b3c published; a fresh twenty-turn run against the final wording (00:48Z): GENERAL ten of ten direct (median 1.22 s); RECORD eight of ten through the gateway and right, two answered without calling it — the route still slips on two of ten; the next Codex attempt's bar is ten of ten on two consecutive runs. Codex's first attempt had left tool_choice 'none' on the session, so record questions never reached the brain — caught by the driver's run, fixed in the second.

NEXT PHASE, APPROVED BY NICK ~00:10Z 2026-09-13 ("1 yes / 2 yes because he can still call all the stuff he needs / 3 yes / 4 yes / 5 confirmed / 6 agreed" — the six ranked differences in evidence/VOICE-PARITY-AUDIT-2026-09-12.txt): the realtime model answers general talk itself from a Skippy voice profile under 2,000 tokens and calls ONE gateway only for records, memory writes and actions (the 42 tools stay behind it — his condition); the three-way route GENERAL / RECORD / ACTION; gateways return spoken-ready facts; two memory stores; the harness reports useful-first-audio. Written into PLAN.md as steps 27–32 when the documentation gate allows; built by Codex under "Fable drives, Codex builds the technical voice pieces". THE TOOL TRIM (00:05Z, evidence/tools-trim-2026-09-12.txt): 42 definitions cut from 17,434 to 7,985 tokens, an eleven-tool core per plain turn; a turn now reads ~27,000 from cache and writes 3,400–10,500 fresh (was 107,691 / 17,000–20,000), a plain turn $0.03–0.04 (was $0.136); nested 55469e8 + d59d0d9, published.

VOICE A CHOICE: front desk, 1000% — and it must mirror what he says the way ChatGPT voice does ("i say repeat back these names and it starts within less than a second"), 2026-09-12 21:40Z (verbatim in evidence/NICK-VERBATIM-2026-09-12-second-answers.txt).

NICK'S ANSWERS 2026-09-12 21:40Z (his numbering was the chat list; mapped): item 7 YES — Chantelle's cloud secret set 21:41Z; item 8 LEAVE IT — closed; item 3 — he reviews the three images first (sent 21:42Z), then wants a Codex prompt; item 2 — he believes the gate is down; it is NOT (active until 2026-09-13T12:44Z; only his terminal can switch it off: node projects/ops/skippy-jobs/lib/md-gov-kill-switch.mjs off); item 6 — he never said never: OPEN, the fence sentence is old-system residue (source: the VOICE brain dump line 710, an agent's text), to be removed from Gracie's front door, the plan and the drafts when the gate allows. STANDING RULING: everything related to the old system is purged from the ecosystem — "claude 2.0 or nothing" — a purge inventory (Larry) then the deletions under his one approval of the list.

NEEDS NICK AT THE KEYBOARD (collected here once; asked once; never twice)
  1. Register the scheduled task pack. The instructions are in the registration pack's REGISTER-THESE file under the
     scheduled-rebuild folder; the three Hub tasks should be registered against the worker's new name, neeko-hub-worker,
     once step 21 closes. Nothing live breaks before then — none of the 21 tasks is registered today.
  2. Create these empty files yourself, because two safety gates refuse an agent creating them (one refuses any new file
     named after Gracie or Neeko; the other refuses a second .md beside an existing one until 2026-09-13T12:44Z):
       - RULEBOOK.md and MANUAL.md inside the skippy assistant's folder (projects/ops/agents/skippy)
       - RULEBOOK.md inside the gracie assistant's folder (projects/ops/agents/gracie) — her fence only; parked otherwise
       - LEARNINGS.md inside the neeko assistant's folder (projects/ops/agents/neeko)
       - neeko-hub-worker.md inside the dispatchable agents folder (.claude/agents)
       - ASSISTANT-FOUNDATION.md inside the agents folder (projects/ops/agents)
       - and say yes or no to two new index pages, CONFIG-AND-SECRET-INDEX.md and FILE-INDEX.md, in the ops folder
         (projects/ops) — the plan proposes that location; you may name another.
     The day each exists, an agent fills it in one copy from the draft already written (step 20). Nothing waits on these
     except steps 21 and 22.
  3. DONE 2026-09-13 02:43Z — the three Slack faces are uploaded and matched (step 19 closed). Nothing left here.
  11. ANSWERED 2026-09-13 ("4 we have other astra account nick@gmail"): the remaining voice work runs on Astra through the other Codex account (CODEX_HOME=~/.codex4, nickdeck19@gmail.com — the account Nick set aside for advanced voice work on 2026-09-08).
  12. ANSWERED 2026-09-13 ("5 yes"): the bar for STEP 15 becomes "the reply is spoken back as soon as the thread answers; an idle thread is told so out loud and offered a fresh agent". PLAN.md is edited when the documentation gate lifts (12:44Z).
  4. ANSWERED 2026-09-13 ("3 fine"): the Hub worker's name is neeko-hub-worker.
  6. ANSWERED 2026-09-13 ~08:40Z ("1 always has been always will be"): Nick's health record is OPEN to Chantelle and Gracie — the 2026-08-15 no-firewall ruling stands; the "never attached for Gracie" sentence was never his. A purge agent closed every conflicting signal in the ecosystem (dispatched 2026-09-13T13:37Z; DONE the same day — 13 references closed in place across 10 files, register row `gracie-health-fence` written to projects/ops/rules-registry/closed-topics.jsonl, record at evidence/PURGE-GRACIE-FENCE-2026-09-13.txt). ONE THING LEFT FOR THE DRIVER: the system-prompt ATTACH gates in server.js still withhold Nick's spine and hard-flag block from Gracie (skippy-code-publish/server.js:16370 and :17688, plus the same in skippy-code and the Mac copy). Her file-RECALL path is already open (NICK_SHAREABLE grants memory/health-full.md, memory/lab-history.md and projects/personal/health/; both EMOTIONAL_PRIVATE lists are empty). The purge corrected every comment and left the behaviour alone — removing the `isChantelle` half is a deliberate code change with its own proof, and the `isNeeko` half stands on ruling #6.
  7. Chantelle's Gracie in the family app is NOT the cloud brain — it is the old Mac program, measured 2026-09-12 20:03Z:
     her chat replies come back marked source "mac" while Nick's come back "fly". The cause is one setting on the cloud
     box: the family app signs everyone in to the brain with ONE shared secret (the family password), the box accepts it
     for Nick but holds a DIFFERENT value for Chantelle (and for Neeko), so her cloud sign-in is refused and the app falls
     back to the Mac for every chat turn; her old-style (ElevenLabs) voice, which has no Mac fallback, dies outright
     (502 "identity check failed"). Everything this lane builds for Gracie lands in the cloud brain, so until this is
     fixed Chantelle sees none of it. Changing a sign-in secret is one of your four; the fix is one command run from the
     brain's folder (it sets Chantelle's box secret to the family password, the value already in the vault as
     family-app-password):  fly secrets set CHANTELLE_LOGIN_SECRET=<family-app-password> -a skippy-cloud
     One line — "7 yes" and an agent runs it with the vault value (never printed), or "7 no". Step 14's third check reads
     FAIL and step 1's Gracie marker is measured on the box only until this is settled.
  8. One of your four, found 2026-09-12 20:02Z while re-reading step 1's first live record: the checker that wrote
     evidence/step1-live-markers.txt earlier today pasted its sign-in commands with the VALUES in them — the family
     password, the brain's access token, and the team password — and a sync commit (d70771fcee) carried that file to
     the workspace's GitHub repo before it was caught. The file is now scrubbed (values replaced by [value]) and
     re-committed, but the earlier commit still holds them in history. Rotating a credential is yours: recommendation —
     rotate the family password (family-app-password in the vault, and the same value on the brain as NICK_LOGIN_SECRET
     and in the family app as SKIPPY_LOGIN_SECRET) and the brain's SKIPPY_AUTH_TOKEN; the team password only if the Hub
     still uses it. One line — "8 rotate" and an agent prepares the exact commands for you to run, or "8 leave" (the
     repo is private). Until then, no agent pastes a sign-in command with its values again: every probe in this lane
     now reads them from the vault at run time.
  5. Voice A — the numbers are in (Codex, five attempts, 2026-09-12 20:49Z, evidence/voice-a-decision.txt; every number
     beside the command that measured it, ten spoken turns of yours each time):
       - TODAY'S PATH (you speak → transcription → the brain → speech): about 18 seconds to the first sound of an answer
         (the brain's own answer takes 11 s, turning it into speech 6 s); 10 of 10 answers correct.
       - FRONT DESK (a realtime voice that answers you at once and asks the brain behind the scenes): it starts talking in
         under a second (0.85 s), but the real answer still arrives at about 15 s because it waits on the same brain; six
         of ten answers did not finish inside 25 s, and none matched the brain's own text word for word.
       - OLD-STYLE VOICE (the ElevenLabs agent Gracie used): about 19 seconds to first audio, 7 of 10 turns produced any
         sound, and it cannot answer your own questions at all (it never reaches the brain).
       - SKIPPY INSIDE THE REALTIME MODEL: not measured — it means rebuilding every tool, rule and memory inside a second
         system; a project of its own, not a switch.
       - SPLIT BY INTENT: quick talk to the front desk, real questions to the brain — derived from the two above.
     What this says: the voice architecture is not the slow part; the brain's 11-second answer is. Recommendation: choose
     the FRONT DESK for the feel (it answers at once and says "let me check") and put the speed work where the time is —
     the brain's answer (step 8 cut the what's-running answer; the same treatment fits the rest). One line —
     "5 front desk", "5 split", "5 today's path", or your own words — written here as `VOICE A CHOICE: <your words>, <date>`.

ASSUMPTIONS (judgments the planner made; each names what would overturn it)
  A1. Two of the three V1 rows carry Nick's own answers of 2026-09-12 17:57Z ("1 yes", "2 make it ask" — on file in
      evidence/NICK-VERBATIM-2026-09-12-three-answers.txt). The third, the worker's name, is confirmed from "neeko is
      hub ops manager so make it one thing" (2026-09-12); the exact spelling neeko-hub-worker is the planner's. Overturned
      by one line from him (item 4 above). The checker (check_plan.py) refuses any UNCONFIRMED row, so an honest
      "unsettled" could not stand in the file; it stands here.
  A0. The brief's recommended Sonnet default for an unnamed agent kind was NOT adopted: Nick's "2 make it ask" overrides
      it, so the tool requires a kind and Skippy asks. The contract (PLAN.md §3 item 1), row U5 and steps 3 and 6 say so.
  A2. The clone-mirror gate refuses a NEW file under a gracie or neeko path however it is created (Write tool, git mv,
      or build_agents.py writing the roster's output). Not tested, because testing it means attempting the act the gate
      exists to stop. Overturned only by the gate's owner reading its code; until then steps 21 and 22 wait for Nick's
      empty files.
  A3. The brain-push ADDS list tolerates an entry naming a file that does not exist yet. Not verified by the planner;
      step 1 item 3 makes the builder read the sweep code and cite the line before adding the foundation's name.
  A4. The Mac program that serves the relay runs the same server.js as the cloud, so adding capture_memory and
      confirm_memory to RELAY_ALLOWED_TOOLS (server.js line 3470) is enough for the Mac side to accept them; the
      memory files then live on the Mac. Overturned if the Mac side keeps a separate allowlist — step 9's builder checks
      _test-relay-port-scope.mjs first.
  A5. The exerciser can sign in as Nick and as Chantelle from this Mac (the vault secret and drive-as-chantelle.mjs).
      Overturned by a refused sign-in, in which case the live legs read NOT MEASURABLE — <instrument> and the step stays
      open, never "waiting on Nick".
  A6. The brief's rule that a cheap step's text must not carry the words business / health / personal / financial (else
      MID/TOP tier) lives in a checker other than check_plan.py — it is not in that file. Obeyed in prose; not obeyable
      in paths, since every brain path begins projects/personal/.
  A7. The three Slack legs of the identity checks (step 14) are read-backs of the existing DM channels, never a message
      sent as Chantelle or as Nick — the four approval classes forbid the send, and the read is enough to prove the
      sender.

QUESTIONS (open; each with WHERE I LOOKED)
  Q1. Which reader in ops/spend-meter.cjs exposes the live ledger rows from the box, for the cache guard (step 25)?
      WHERE I LOOKED: the price table in ops/spend-meter.cjs (cacheRead/cacheWrite per model) and server.js line 36
      (the import). The ledger's path is inside spend-meter.cjs and was not opened by the planner; step 25's builder
      reads it before writing the guard.
  Q2. Does `claude -p` accept a model alias per kind (sonnet | opus | fable) or need a full model id? WHERE I LOOKED:
      not opened — step 5 item 1 makes the builder read `claude --help` and copy the exact names; the spike re-run
      (step 2) records the flag that worked.
  Q3. Where is the live per-machine running list on this Mac? WHERE I LOOKED: work-watch.mjs header lines 55–57 say it
      writes work-threads-<machine>.json under projects/personal/skippy-app/ala-state/; the first spike polled the shared
      work-threads.json (3 rows, dated Sep 9) and missed it. Step 2's re-run lists the folder newest-first and reads
      that file; step 8's timing script compares against it.

PLAN CHANGES (dated deltas to contracts or scope)
  - 2026-09-12 · the driver · SCOPE · F9 now names the minimum Gracie shape (front door + fence rulebook, parked beyond
    that) instead of "each face's folder in the team assistant's shape" · why: Codex's cold attack found F9 had two
    meanings while Nick has parked Gracie · makes easier: F9 is checkable; makes harder: nothing · lanes notified: none needed.
  - 2026-09-12 · the driver · CONTRACT · dispatch_to_code_agent: coordinate mode matches the exact thread id from
    whats_running only (zero or many matches → ask); the drain runs every session with execFile and an argument array;
    the kind→command map is fixed (sonnet/opus/fable aliases, codex senior-engineer profile) · why: Codex's attack items
    4, 5, 6 · makes easier: one build; makes harder: nothing.
  - 2026-09-12 · the driver · CONTRACT · confirmed memories reach the engine by its own two-call path (capture_narrative
    dry_run false into staging, then publish_accepted_event into production with the person's yes as accepted_by) · why:
    the write door lands in staging by design and production is what personal_answer reads · lanes notified:
    HEALTH-ONE-DOOR gets the dated line when step 11 closes.
  - 2026-09-12 ~19:05Z · the driver · CONTRACT (model ladder) · every edit to the cloud brain's main file (server.js) and
    the publish script runs on an OPUS builder with a Sonnet checker, one builder at a time, instead of the cheap lane:
    the cheap tool's data wall refuses to release server.js to any vendor ("an assigned credential" pattern, the
    vendor-fence scanner, 2026-09-12 — the same token-variable shape a memory on file records as a false positive) and
    refused the publish script and the spend meter the same way (steps 9 and 25, all three vendors, verbatim in
    evidence/); the standing rule is "record a data-wall refusal verbatim and do that piece in-house on Nick's
    named-model word — never route around it". Nick's words: "pull back to opus and sonnet where reasonable for build"
    (2026-09-05). New files beside server.js (the loader, the door module, their tests) still go cheap. · makes easier:
    the brain edits land; makes harder: they cost subscription tokens instead of the flat cheap plans.
  - 2026-09-12 · the driver · SCOPE · step 19's builder is the driver driving Nick's own signed-in browser (standing
    grant 2026-08-30), Nick at the keyboard only as fallback; step 20 has seven drafts (the legacy-prompt inventory
    added); step 23 carries a target per routing-wall copy because the two copies differ today.

BOARD CARD: nt-20260912-175831-7392 — "VOICE: Voice, Skippy, Gracie and Neeko - the three assistants as one clean
  system", group ai-builds, due 2026-09-13; created by the driver 17:58Z (evidence/NICK-VERBATIM-2026-09-12-three-answers.txt).

ALREADY IN evidence/ WHEN THE PLAN WAS WRITTEN (never re-do):
  - SIENNA-AVATARS-AND-VOICES-2026-09-12.txt — the avatar acceptance criteria and the three-voices concept (the brief's
    S9 concept half is done; step 18 grades against §1 of that file).
  - SPIKE-HEADLESS-SESSION-2026-09-12.txt — the first spike: no session file, no socket, a transcript only; its
    running-list question NOT MEASURABLE (pid captured as 0; polled the stale shared file). Step 2 re-runs that one
    question; step 5 designs on the observed answers (work-watch gains the queue file as a row source).
  - NICK-VERBATIM-2026-09-12-three-answers.txt — "1 yes 2 make it ask 3 VOICE/SKIPPY/ETC due tomorrow".

CORRECTIONS TO THE PLANNER BRIEF, MEASURED WHILE WRITING (recorded once, here)
  - The three Hub task prompts (hub-ops-hourly-audit, hub-ops-sop-daily, hub-ops-friday-digest) name neither
    hub-ops-manager nor its LEARNINGS file (command grep -c hub-ops-manager → 0 on each, 2026-09-12). The brief's "repoint
    the three prompts in the same change" has nothing to repoint; step 21 leaves them alone and names the roster (line
    1512 and the LEARNINGS path string at line 1528) and build_agents.py as the real places.
  - The Cowork count is 130 matching lines in server.js (command grep -ci cowork, 2026-09-12), not 124.
  - The exerciser runs on haiku (ZION/agents/exerciser.md line 5); Larry on sonnet (ZION/agents/larry.md). The plan's
    executor cells say so, because the checker needs a model name in each cell.
  - The plan's steps split the brief's S2 into five and S5 into four (one builder, one definition of done, one proof per
    step), S6 into two, S8 into two, S9 into two and S10 into three: 19 FRONT steps and 6 POLISH steps plus the sign-off.

STEP RECORD (one line per step, rewritten in place; the CLOSED shape is
  STEP <n> CLOSED — <what is now true> — checked by <model> — <the proof command or artefact>)
  STEP 1  CLOSED — the three assistants read their own documents on every turn: the loader (lib/assistant-docs.mjs) resolves /data/brain on the box and hands each face its front door, foundation, rulebook, voice guide and manual, re-read when a file's mtime moves; a root that is not there yet is looked for again on the next call (the live miss found 2026-09-12 — documents land on the box after boot — fixed by zai through route-build, nested 7bfd1e0); the three composers prepend the documents (Skippy and Gracie) or compose from them alone (Neeko, whose inline prose and pasted rules are gone, nested ab72a1b); the three folders and the foundation are in the brain-push ADDS (workspace 860518735f); published in fly deployment 01M2BK6V6NX1E8K03V7VCW91GG at 19:59Z — checked by sonnet (the verifier, its own session, 2026-09-12 20:30Z): unit test ALL PASS, STANDING_RULES_NEEKO count 0, the deployed loader on the box returns all three markers, Skippy answered the marker question Yes twice over the public route, Neeko answered a behaviour question from his own document twice, Gracie answered wrongly once then rightly once (her spoken word about her own instructions is not a reliable instrument; the loader proof is) — records evidence/step1-live-markers-3.txt and evidence/step1-check.txt. Gracie's PUBLIC route still reaches the Mac program until NEEDS NICK item 7 is settled.
  STEP 2  CLOSED — a headless session started from a shell registers no session file and no socket, writes only a transcript, and never reaches the live per-machine running list; so the drain (STEP 5) registers every session it starts — checked by haiku (the exerciser, its own session, 2026-09-12; the verify tool reproduced 2 of 6 pairs and the checker re-ran the 4 outside-tree ones by hand, all consistent) — evidence/step2-spike-rerun.txt (pid 87390), evidence/SPIKE-HEADLESS-SESSION-2026-09-12.txt (first run), check evidence/step2-check.txt
  STEP 3  CLOSED — dispatch_to_code_agent exists with the frozen shape (kind required; a missing kind asks "which kind — Sonnet, Opus, Fable or Codex?" and writes nothing; coordinate matches one exact thread id or asks), the three bands and the ask rule are in its description, it is on the relay allowlist, and a real call through the server's own runTool lands a queued row on disk — built on Opus (the cheap lane timed out twice and cannot read the main file) — checked by haiku (the exerciser, its own session, 2026-09-12) — _test-code-agent-dispatch.mjs ALL PASS incl. the live j- cases; check evidence/step3-check.txt; nested commits 67a5d7e, 2b23a49 (not yet published)
  STEP 4  CLOSED — the word cowork is gone from the brain's code, prompts, tests and queue names; dispatch_to_code_agent is the only work door and every former caller (stream buffering, both voice tap-confirm gates, the Gracie reword, the exec-tool sets, the Alexa path, the false-completion guards) uses it; the queue is ala-state/code-agent-queue.jsonl, read by lib/dispatch-read.mjs and pushed by skippy-brain-push — built by Opus (NICK-ASKED 2026-09-05; nested 1023f6f, jobs swept into workspace c39e367) — checked by haiku (the exerciser, its own session, 2026-09-12 19:59Z) — command grep -rci cowork over server.js, lib and skippy-brain-push.mjs → every line :0; node _test-code-agent-dispatch.mjs → ALL PASS; record evidence/step4-check.txt. Left behind by design: work-watch.mjs still reads the old markdown queue (it imports its parser from the health-lane copy this lane never touches), and the old dispatch worker + slack-inbound still read the markdown queue until STEP 5's drain replaces them.
  STEP 5  CLOSED — a queued row of each kind becomes a running session on the Studio with its pid, start time, worktree and transcript written back on the row (sonnet, opus and fable through the claude command with the CLI's own aliases; codex through the Codex CLI with stdin closed and Nick's spoken brief as its CODEX-APPROVED line; a coordinate row is closed as STEP 15's job), the job is on the runner's clock every minute, and work-watch lists each running row as a dispatch thread alive while its pid is in the process table — built by Opus (NICK-ASKED 2026-09-05; workspace 8d4d8e7; execFile drops stdio and detached so the drain uses spawn, measured first) — checked by haiku (the exerciser, its own session, 2026-09-12 20:22Z) — node projects/ops/skippy-jobs/_test-code-agent-drain.mjs → 71 checks ALL PASS; node --check runner.mjs → 0; grep -c code-agent-drain runner.mjs → 1; record evidence/step5-check.txt. The runner process on the Studio (launchd com.skippy.jobs) loads its schedule at start, so it was restarted after this closed to pick up the new row.
  STEP 6  CLOSED — Nick says "start a Sonnet agent to …" by voice, says "yes", and the agent really starts (a real Sonnet session wrote PONG within a minute), "what's running?" names it with its kind inside two minutes, a request without a kind is answered with the question and starts nothing, and no reply mentions Cowork — checked by sonnet (the verifier, its own session, 2026-09-13 06:41–06:50Z, evidence/step6-check-2.txt): builder record verify 10/10 MATCH; its own fresh dispatch: spoken yes queued it (no tap), named in the rundown at 16 s with kind sonnet, PONG landed, the kind-less request asked "Sonnet, Opus, Fable, or Codex", Cowork never mentioned — RESULT: PASS. Its caveat, on the NEXT list: a few minutes after a dispatched session finishes, "what's running?" stops naming it and names an older one instead — fine inside the step's two-minute window, worth widening (recent dispatches named for two hours as the fix intended; check the relay's row ages). Records: evidence/step6-check-2.txt, step6-fix-2026-09-13.txt (nested 3cd942c…36f08a9, published 06:39Z deployment-01M2CQSV771TG8H1NVTH0PJBZB). Handoff posted to plans/VOICE/PROGRESS.txt 2026-09-13T06:50Z. Earlier: the checker (sonnet, its own fresh dispatch, 2026-09-13 05:45–05:53Z, evidence/step6-check.txt): the dispatch itself works (a real Sonnet session wrote PONG within a minute), but (A) a spoken "yes" in the same conversation did NOT confirm the pending handoff ("What are you saying yes to?"); only the tap route queued it; (B) five "what's running?" polls over 103 s never named the dispatched session or its kind — the rundown gives the project summary and calls a row "an unnamed piece"; (C) a kind-less "start an agent to list the files in /private/tmp" was answered twice as a status check (trace whats_running: "Nothing's running on that yet") instead of asking which kind — Nick's "2 make it ask
  STEP 7  BUILT AND PUBLISHED, GATE BLOCKED BY ANOTHER LANE'S INSTRUMENT — job A (zai, nested 7e1047d): the shopping-list add mutation and its shaper, the read query proven never to write; jobs B and C (Opus, nested 8d445b6 + 66b2358, published 21:15Z): add_shopping_item beside read_shopping_list in the add_todo pattern (household callers only, item read back off the board), the claim-guard extended to "handing this to", "passing this to", "I'll hand this off" with a receipt-less turn caught (red run 9 passed 6 failed on file in evidence/step7-red.txt; green: 15 passed ALL PASS); fresh-five.json written. The gate (projects/personal/family-app/_test-voice-requests.mjs --gate --fresh) printed delivered: 5 of 5 · read back: 1 of 5 — and the 1 of 5 is the instrument, not the brain: the VOICE lane's _gate.mjs reads back its OWN five hard-coded requests by id (Monday board 2689216450 for milk, WhatsApp for one, a calendar move for another) rather than the fresh five it was handed, so four of ours were checked against the wrong record. That file is the VOICE lane's and this step forbids touching it; a dated line asks the VOICE lane to make the gate read back the fresh file's own requests. Until then STEP 7 cannot close on its literal proof; the brain's side is done.
  STEP 8  CLOSED — checked by sonnet (the verifier, its own session, 2026-09-12 21:22Z): it re-ran the timing script itself — 5/5 under 15 s with every named count matching the file — and the live test green; record evidence/step8-check.txt. What changed: whats_running now returns the rundown already written for the ear (a spoken field first in the result, with an in-result instruction to read it as is; every old field kept; ids, paths and file names refused as names), each machine's thread list parsed once per change; before: 9.6–19.1 s on 3–4 model rounds (and 60–99 s on a cold box); after, run K on commit c2f7a9b: 5/5 under 15 s (7.0–10.2 s), counts agree 7=7 — built by Opus (NICK-ASKED 2026-09-05; nested 7db52f4…c2f7a9b, published; workspace 588b63e60c) — the builder recorded ten runs honestly: the box's one-round floor swung 2.5–43 s while other lanes published eight times in 45 minutes, and two runs died on 502s; the checker (sonnet) re-runs the timing script now. Records evidence/step8-before.txt, step8-timing.txt.
  STEP 9  CLOSED — on a box without the credential scanner a memory proposal or confirmation relays to the Mac program (which holds the scanner) instead of being refused; scanner present runs local; neither refuses; built on Opus because the wall refused the cheap lane the main file — checked by haiku (the exerciser, its own session, 2026-09-12) — node projects/personal/skippy-app/skippy-code-publish/_test-memory-write-filter.mjs → 37/37 ALL PASS; _test-relay-port-scope.mjs → 25 passed; check evidence/step9-check.txt; nested-repo commit eac0a99 (not yet published)
  STEP 10 CLOSED — a spoken or typed yes from the person keeps a pending memory by the tap's own path; a model-side affirmation still keeps nothing — checked by haiku (the exerciser, its own session, 2026-09-12) — node projects/personal/skippy-app/skippy-code-publish/_test-memory-gate.mjs → 44/44 ALL PASS; red run evidence/step10-red.txt; check evidence/step10-check.txt; nested-repo commit 3599ea0 (not yet published)
  STEP 11 CLOSED — a confirmed memory row is written through the personal engine's own two-call door once (capture with dry_run false, then publish of the staging id with accepted_by naming the person) and stamped drained_at; a marker-shaped row (a dose, a lab, a rate) is refused by scanForHealthMarkers before any call and listed in the drain's log for the engine lane; an already-drained row is skipped and a second run calls nothing; the job is on the runner's clock every 15 minutes and beats a heartbeat when it runs live; the brain no longer appends anything to a captures file — everything proposes through the gate — built by Opus (NICK-ASKED 2026-09-05; workspace 856ed9dca3 + a112d9d4f8, nested 17050b2) — checked by haiku (the exerciser, its own session, 2026-09-12 20:28Z) — node projects/ops/skippy-jobs/_test-memory-confirmed-drain.mjs → 30 checks ALL PASS; node --check runner.mjs → 0; grep -c memory-confirmed-drain runner.mjs → 1; grep -c captures.md server.js → 0; record evidence/step11-check.txt. The server.js half ships with the next publish (STEP 8's); the runner was restarted at 20:31Z for the new row.
  STEP 12 CLOSED — a memory Nick tells Skippy by voice is proposed, kept on his spoken yes, survives a restart of the cloud program and comes back in a new conversation as the current value (a newer value of the same fact retires the older one), the old Mac program is out of the memory path, and a login-shaped memory is refused with the floor named and never echoed or written — checked by sonnet (the verifier, its own session, 2026-09-13 06:33–06:37Z, evidence/step12-check-2.txt): it re-ran the proof script with its own nonce ("canyon thistle", the third same-subject fact on the box) — proposed, kept, restart, recall "Your test phrase for the assistants plan is \"canyon thistle\" — that replaced an earlier value." (exactly one live value, no hedge), the credential refused and nowhere on disk, its own grep of the old program's log 0 — CHECKER RESULT: PASS — the proof: STEP12_RECORD=<record> node evidence/step14-probes/step12-live-proof.mjs → RESULT: PASS (records evidence/step12-live-memory-7.txt, step12-check-2.txt; fixes 1–4 in evidence/step12-fix-2026-09-13.txt; nested 37de2c0…f1de790, last published 06:28Z deployment-01M2CQ5P4FGPKX9F7CXJ12NVSB). NEXT-list item from fix 2: confirmed rows are answerable at once on the box but nothing yet carries them onward to the personal store (one outbox row plus one branch in the Mac's outbox drain); superseded rows: carried once if carried before, never twice, the newest as the current value. Earlier: the sixth live run (driver, 2026-09-13T06:09Z, evidence/step12-live-memory-6.txt, script evidence/step14-probes/step12-live-proof.mjs) against the fix-3 build (nested 3267b59, published 06:05Z, deployment-01M2CNT34PF1CKMRNNKQZZYAYA): "remember that my test phrase for the assistants plan is canyon willow" → "Got it — canyon willow it is." with the action "✓ Proposed for memory", nothing relayed, the old Mac program's log did not grow (0); "yes" → "Got it — canyon willow, saved for good." with "✓ Kept in memory" (confirm_memory success); fly apps restart; "what did I ask you to remember about the assistants plan?" → "It's \"canyon willow\" — that's the test phrase you set for the assistants plan on September 13th." (recall success); the invented bank login → "That's on the floor — logins. I didn't keep it and I won't repeat it back." (blocked_by memory_floor:login), nothing echoed, nothing
  STEP 13 CLOSED — no job sends a message to Nick through Gracie's bot token (the two historical sends were repointed on 2026-07-13; the rule already held), and a guard test now proves it on every run: 205 job files scanned, 0 Gracie-token sends to Nick, 4 Skippy-path sites; the test goes red on a sabotage fixture — checked by haiku (the exerciser, its own session, 2026-09-12) — node projects/ops/skippy-jobs/_test-sender-rule.mjs; red/green record evidence/step13-red.txt; check evidence/step13-check.txt; workspace commit da7ee2efadd
  STEP 14 CLOSED — six of six live identity checks PASS (leg 3 turned green at 21:43Z after Nick's yes put Chantelle's cloud sign-in secret on the box; her family-app chat now answers from the cloud brain, source fly) — checked by sonnet (the verifier, its own session, 2026-09-12 21:51Z): it re-ran every leg live a second time, all six reproduced, record evidence/step14-check.txt; its one flag for Nick: Skippy's notes into Nick's own Slack notes-to-self carry Slack's 'via Gracie' app tag because that app is the transport, so his eye can meet her name there — a NEXT item, not one of the six checks. Earlier: five of six live checks PASS (Nick meets Skippy on the app and on voice; Chantelle meets Gracie on the app and on Slack; Skippy's live note to Nick landed in his own note-to-self and no Gracie message to him is newer than the cutoff); the sixth, Chantelle's old-style voice mint, FAILS because the box holds a different sign-in secret for her than the family app sends — a credential act, NEEDS NICK item 7. Record: evidence/step14-identity.txt with the probe scripts in evidence/step14-probes/; re-run leg 3 and summon the checker (sonnet) the moment item 7 is answered. Driven by the overseer after the exerciser could not sign in twice (2026-09-12 19:55Z).
  STEP 15 FIX DEPLOYED, IDLE HALF PROVEN LIVE, RUNNING HALF NOT MEASURABLE (item 12) — Opus's fix (evidence/step15-fix-2026-09-13.txt; workspace 4373eea640, 95a2d23cb0, 53930b7460; nested b2ae7a9, 8d95596, 8f85b43) is live: brain deployment-01M2CWFQDHE1VC7CPJJB6YTY23 (08:02Z), family app panel.js v89 / cache v831 (08:04Z), the delivery service restarted 07:49Z. THE FACT, measured with a control: a claude session idle at its prompt never processes an injected message (a platform limit); the driver's own mid-turn session did not absorb the test's message either (absorbed:false, 08:08:59Z). LIVE RUNS BY THE DRIVER (evidence/step15-live-2026-09-13.txt): the idle-thread sentence IS spoken ("claude-2-0-ed is idle, so it will only see this when it next runs. Want me to start a fresh agent with it instead?", 30 ms after landing) once two relay gaps were fixed by zai through route-build — the family app's relay dropped the absorbed flag (64ba27e8c2, deployed e064f055) and the kind/timing headers (a949e1e4b8, deployed 0e690632); the live edge was still not showing the kind header four minutes after that deploy (the deployed source and dist carry it) — re-probe in the morning. The running-thread half (a reply spoken within 30 s of the session answering) is NOT MEASURABLE tonight: no session absorbed the injected message. Fidelity: the one Status mismatch is a true fault strip (Chantelle's Mac mini has not sent thread data since 2026-08-31 — a finding for the mini's lane), not this step's strip. Waits on NEEDS NICK item 12 (the bar). Earlier: the checker (sonnet, its own session, 2026-09-13 06:52–07:03Z, evidence/step15-check.txt): against a genuinely IDLE live session (the HUB lane's, 635ae494…, waiting on Nick) the test's message was delivered live (delivered:true, row 635ae494-81, deliveredAt 06:52:46Z, the "landed" strip after 4.7 s) but nothing was ever spoken: /api/thread-say was never called inside 30 s and the target session had produced no reply ~90 s later; the Status-tab fidelity check reads mismatched 10 · unmeasured 10 (
  STEP 16 CLOSED — the Voice A decision file carries measured numbers for every candidate on the same ten spoken turns (today's transcription→brain→speech chain ≥18.3 s to first audio, DERIVED as brain 11.3 s + speech 6.1 s, 10/10 correct; the ElevenLabs agent path 10–19 s to first audio across runs, 7 of 10 turns producing sound, no text reply to grade; the realtime front desk 0.85 s to its first word and 15.2 s to the brain's answer, strict correctness 0/10, $0.024 a turn; Skippy-inside not measured by design; split-by-intent derived) and Nick's choice sits on this sheet as NEEDS NICK item 5 — built by Codex (gpt-5.6-terra, six attempts, evidence/voice-a-decision.txt) — checked by sonnet (the verifier, its own session, second run 2026-09-12 21:05Z after the first run failed the ElevenLabs row; it re-ran the harness itself: both rows ten samples with a first-audio median, exit 0; record evidence/step16-check.txt). History: the sixth attempt (Codex, 21:02Z) found the ElevenLabs instrument fault (the harness skipped the websocket handshake after the mint — initiation and ping/pong) and measured that row: 19.2 s median to first audio, 22.0 s worst, 7 of 10 samples produced audio, no text reply to grade, no per-turn price reported; the checker (sonnet) re-runs the report now. the checker (sonnet, 20:54Z, evidence/step16-check.txt) re-ran the report itself: the front-desk row is measured with its method explained; the ElevenLabs row prints NOT MEASURABLE in both runs, so the literal bar (a real first-audio median for both named candidates) is not met; Codex's sixth attempt (dispatched 20:54Z, 35-minute cap) measures that row alone. Earlier: Codex (gpt-5.6-terra, five attempts on 2026-09-12; the third was stopped after it sat polling its own process for 20 minutes, the fourth hung inside the realtime session, the fifth capped every sample at 25 s and terminated) — evidence/voice-a-decision.txt: today's path ≥18.3 s to first audio (DERIVED: brain 11.3 s + speech 6.1 s), 10/10 correct; the realtime front desk 0.85 s to its first word (1.5 s worst) but 15.2 s to the brain's answer, 6 of 10 answers past the 25 s cap, strict correctness 0/10, $0.024 a turn; the ElevenLabs legacy path NOT MEASURABLE (instrument failed without a safe diagnostic); Skippy-inside not measured by design. The literal proof asks for ten measured samples on BOTH the ElevenLabs row and the front-desk row, so the checker (sonnet) may fail it on the ElevenLabs row; the decision Nick needs is on the sheet either way — NEEDS NICK item 5.
  STEP 17 CLOSED (and its truth gap closed 23:17Z: Codex traced the wrong answers to the realtime model REWRITING the brain's returned text — it now only acknowledges or mirrors, relays the brain's text verbatim through the exact-text speech route, and waits for the tool; live: 10/10 first audio 1.07–1.51 s, 10/10 factually right, zero brain timeouts, answers 5–27 s (the brain's own speed, cut next by the tool trim); brain f50d11a, family deploy 696d05db; evidence/step17-truth.txt) — the front desk Nick chose speaks its first word under two seconds on ten of ten spoken turns, live in the family app (Codex, two attempts; workspace aa64e08b2c + 9b33781e54 + the v=49 bump, deployed b17bddcb; brain 7168215 published v459) — checked by sonnet (the verifier, its own session, 2026-09-12 22:50Z): its own live run put all ten first-audio times at 1,027–1,373 ms; record evidence/step17-check.txt. NEXT, and it matters: only three to six of ten answers were factually right — the front desk must never state a fact the brain did not return (say "let me check" and wait, or say the brain did not answer); that is a front-desk instruction fix plus the brain's own speed (the tool trim). Earlier — run 4 at 22:44Z (family deploy b17bddcb, voice.js v=49; brain v459 with the health-spine cut): FIRST AUDIO 1.04–1.99 s ON TEN OF TEN — the plan's bar met on the first word; the brain's answer 5.8–22.7 s on nine, one past the cap, six of ten factually right (the answer half is the brain's own speed and truth, now on a 43k-token prompt; the tool trim is the next cut). Record evidence/step17-live.txt (all four runs). Earlier — 22:35Z: Codex attempt 2 restored the brain-side mint (nested 7168215, published v459) and wired the family side (workspace aa64e08b2c); live proof against that brain: FIRST AUDIO 1.06–1.67 s ON TEN OF TEN — the bar for the first word is met; the brain's answer behind it took 9.9–14.7 s on four samples and passed the 25 s cap on six, two of four factually right — the answer half is the brain's own speed and the front desk's tool wiring, cut by the spine shrink (health lane, in flight) and the tool trim (needs Nick's yes); the client's final hardening needed a version bump (voice.js v=49), deployed with a third live run at 22:36Z. Nick chose the front desk (21:40Z). Codex attempt 1 (21:53Z–22:03Z) built the family-app side (voice.js: acknowledge at once, exact mirror without the brain, one bounded brain tool, server-VAD barge-in, tap-to-open, cache v827) but stopped at the deploy guard; its files were swept into a git autostash by another session's merge, recovered by the driver from stash@{0}, committed (workspace 9b33781e54) and deployed (deck-family v827, 22:10Z); its brain-side session-mint hunk was never committed and is gone. Live proof 22:14Z: first audio 1.0–1.4 s on seven samples, 2.2–2.4 s on three; the brain answer 8–16 s, four past the 25 s cap; strict correctness 0/10 — chosen path 0/10 under 2.0 s. Codex attempt 2 dispatched 22:15Z: rebuild the brain-side mint (commit it alone the moment it exists), pre-warm the session on the tap, re-run the ten.
  STEP 18 CLOSED — three faces drawn in code as one family ("the shell and the pearl": Skippy a closed ring with a ball on its floor; Gracie the ring opened at the top with rounded ends and the ball in dusty rose; Neeko a rounded square with the same ball), rendered at 512 and at 24/30/32/34/44 px and on light and dark sidebars, matched cold 3 of 3 by a Sonnet session that saw only the 32 px strip, and graded PASS by Sienna on gates 1, 2 and 6 — checked by fable (Sienna, the one grade, 2026-09-12; a first set failed as the power-button glyph and was redrawn to her geometry) — evidence/avatars/ (six SVG masters + PNG renders, composite-1200.png, strip-32.png), evidence/avatars-grade.txt
  STEP 19 CLOSED — the three Slack apps wear the three chosen Jarvis-style faces (Skippy app A0BHKDZSM08 = g5, Gracie app A0BDU0SFHT8 = g2, Neeko app A0BFUD2RLQG = g1), uploaded by the driver through Nick's signed-in Chrome at 02:38–02:41Z on 2026-09-13 and read back through Slack's own API: three new icon hashes (57ac1dcf… / d9f3208d… / 788f51a1…), 7.6–8.3 KB real images where 460–490 byte placeholders were, distinct 3 of 3 — checked by sonnet (a cold session, read-only, 02:43Z): given the three chosen references and the three live 192 px icons under neutral names, matched 3 of 3 (skippy certain, gracie certain, neeko likely), no placeholder — evidence/step19-slack-icons-after.txt with the live files in evidence/avatars/live-after-2026-09-13/. Earlier: NOT STARTED — three faces live in Slack
  STEP 20 CLOSED — seven drafts in evidence/drafts/ (the foundation, Skippy's rulebook with 23 earned rules moved verbatim from the code, Skippy's manual, Gracie's fence rulebook, the config-and-secret index over four stores with the check-all-four rule, the file index, and the legacy-prompt inventory mapping 110 paragraphs with n = m + k per block), each with Owner and Purpose lines and no value shape; five written cheap (qwen/deepseek), two on Opus because they read the brain's main file — checked by haiku (the exerciser, its own session, 2026-09-12) — check evidence/step20-check.txt; one contradiction surfaced for Nick (item 6 above: Gracie and Nick's health record)
  STEP 21 NOT STARTED — the Neeko merge
  STEP 22 CLOSED — the three assistants are defined in their own folders and the brain holds no persona prose: the foundation, Skippy's RULEBOOK and MANUAL, Gracie's RULEBOOK and the two index pages are written from their finished drafts (the driver filled them 2026-09-13T15:11Z after Nick created the files by his word and the cheap builder looped; Gracie's carries the open-health sentence, not the withdrawn fence); Neeko's rulebook cites the foundation; the two old skippy-app manuals carry SUPERSEDED BY lines and are kept. The brain side (Opus, evidence/step22-server-2026-09-13.txt; nested 6fe63ec + 836339a, published 15:04Z deployment-01M2DMPCRKHT320YDQN7133YBJ): Skippy's builder composes his documents through the loader and STANDING_RULES_GRACIE is deleted — 42 paragraphs, every one mapped to the legacy prompt inventory first, count now 0; Gracie's builder prose is deliberately NOT deleted because ~30 of her paragraphs map to a rulebook section marked "Chantelle's call" that is not written, and a new assertion guards it; tests _test-assistant-docs.mjs ALL PASS with no SKIPPED and _test-health-attach-identity.mjs ALL PASS, both red on the pre-change code; live after publish: Skippy answers "I'm Skippy — Nick's chief of staff, working for him and for Chantelle with equal say" in 4.0 s and names the four approval classes in 5.5 s, both from the composed documents. THE HEALTH GATE: Nick's spine and hard-flag block now attach for Chantelle and Gracie exactly as for Nick (his ruling, 2026-08-15 and 2026-09-13); Neeko stays excluded on its own separate rule.
  STEP 23 CLOSED — both routing-wall copies carry the class split (code · personal/health · private) and CLASS_PRIVATE is exactly the floor (logins · credentials, tokens and keys · government IDs · card, bank and routing numbers), cross-checked against vendor-fence's FLOOR; family, the kids, Chantelle, health markers and client material classify personal and travel; no quoted retired rule remains anywhere in the Mac copy — checked by Sonnet (the verifier, its own session, 2026-09-12; first check failed on kept old wording, fixed in c3cb689 and 211afc8, re-check PASS) — 126/126 in _test-floor-four-categories.mjs; each copy's selftest 21/21 BLOCK; red run evidence/step23-red.txt; checks evidence/step23-check.txt and step23-recheck.txt; nested commit a0c6f36 (published 2026-09-12 evening), workspace commits 68f1c94000, c3cb689, 211afc8
  STEP 24 PASSES DONE, BOARD WRITE WAITING ON THE GATE — Larry ran twice on 2026-09-12 (first pass ~19:15Z, four THING groups; second pass ~20:05Z over the remaining ZION/agents files and the three assistants' folders, three more): seven THING groups against the bar of five, each with its single file, copies, evidence command and one recommendation, recorded verbatim in evidence/LARRY-PASS-2026-09-12.txt and evidence/LARRY-PASS-2-2026-09-12.txt. The upkeep board (larry-lane.md) is a governed markdown file the gate refuses until 2026-09-13T12:44Z, so the driver prepends both passes there under a dated header the moment it allows and then summons the checker (the proof counts THING lines on the board). Larry changed nothing.
  STEP 25 CLOSED — after every publish a guard asks the live brain one tagged question and reads that request's own ledger rows from the box: cacheRead 210,692 on the tool round (PASS), and the sabotage fixture (cacheRead 3,289) goes red; built on Opus because the wall refused the cheap lane the publish script — checked by haiku (the exerciser, its own session, 2026-09-12) — node projects/personal/skippy-app/skippy-code-publish/ops/test-cache-guard.mjs → PASS; --expect-fail → exit 1; fly-publish.mjs +13 lines, 0 deletions; check evidence/step25-check.txt; nested-repo commit 252de73 (not yet published)
  STEP 26 NOT STARTED — postmortem and FINISH LINE sign-off

2026-09-12T21:46Z — HAND-OFF FROM THE HEALTH ENGINE SESSION (HEALTH-ONE-DOOR lane; reply to the session named "HEALTH ENGINE", CCD id local_95d2d606-3e26-4799-98bd-d45e72c1c83d). Nick, 2026-09-12: "you connect with it here - ping it and figure it out". MEASURED, not read from code: signed in as Nick on skippy-cloud v444 (20:49Z), POST /api/chat "What was my most recent hs-CRP reading and what date was it?" → tool trace [] and a correct answer (0.69, 2026-06-23): Skippy answered a health question from the record snapshot pushed into its system prompt (getHealthSpine(), /data/brain/…/injection-context-core.md, synced 14:00Z; pushed at server.js 15420 chat and 16671 voice) and never called engine_answer; Gate 0 still says "The spine is ALREADY IN YOUR CONTEXT". Evidence: plans/HEALTH-ONE-DOOR/PLUMBING/evidence/phone-health-answer-live-2026-09-12.txt. Nick approved making the health engine the only door for health answers (HEALTH-ONE-DOOR PLUMBING steps 4 and 13). Proposal: this lane owns server.js, so the change lands here — for Nick's health-classified questions call engine_answer first and hand its answer to the model as the record; keep the snapshot only for the hard flags and ruled-out registry until the engine carries them in every answer; replace the Gate 0 line. Pass test: the same question returns a trace naming engine_answer. If this lane already has it built and unreleased, say where; if you want the HEALTH ENGINE session to make the edit instead, say so and name the lines you are touching so the two of us do not undo each other.

2026-09-12T21:50Z — REPLY FROM THE ASSISTANTS DRIVER TO THE HEALTH ENGINE SESSION (its session is in another app profile, so this line and a paste from Nick carry it): accepted — YOU make the change, under a fence so we do not collide in the one shared checkout. Yours in server.js: (a) ANSWER_ROUTING_PROMPT (declared near line 5036): a health question (a marker, a lab, a dose, a protocol, "what was my…") calls engine_answer FIRST and answers from its result; the snapshot is never quoted for a value. (b) the health-spine system block in the /api/chat handler (the `if (!isChantelle && !isNeeko)` block near line 15418 that pushes spine.text): carry the hard flags only (six flags, three rules) until the engine carries them; keep the cache_control marker on the last static block where it is. Nothing else in server.js (whatsRunning, the dispatch gates, thread-say, the cache instrumentation are other builders' work tonight). Before editing: git status --short server.js in projects/personal/skippy-app/skippy-code-publish must be clean; if modified, re-check every two minutes. Commit server.js alone (scoped add), publish with node fly-publish.mjs from that directory (never the dirty flag). Pass test as you wrote it, plus one non-health question showing NO engine_answer call; record both with the build id, and post one dated STEP RECORD line here: "HEALTH-ONE-DOOR touched server.js: <commit> <build> — health answers route through engine_answer; the spine block carries hard flags only". The driver will not touch those two regions.

2026-09-12T21:54Z — HEALTH ENGINE session ACCEPTS the ASSISTANTS driver's fence (relayed by Nick): HEALTH-ONE-DOOR touches only (a) ANSWER_ROUTING_PROMPT and (b) the health-spine system block in the /api/chat handler (the non-Chantelle/non-Neeko block that pushes spine.text; cache_control marker left where it is) in skippy-code-publish/server.js; clean `git status --short server.js` before editing (re-check every two minutes if not); commit server.js alone; publish with `node fly-publish.mjs` without the dirty flag; pass tests: the hs-CRP question shows engine_answer in the trace AND one non-health question shows no engine_answer; then the STEP RECORD line here. This session cannot send messages to yours (different app account, no send tool here); yours can reach it with SendMessage to "HEALTH ENGINE".

2026-09-12T22:25Z — THE 125,000 AND THE 17,000, MEASURED ON THE BOX (Opus builder, evidence/prompt-size-2026-09-12.txt; Anthropic's own token counter, live): the prompt is 124,932 tokens — the HEALTH SPINE 84,047 (67%: labs, markers, protocol, ruled-out registry, hard flags, injected whole on every turn, even "hello"), the PERSONA 23,615 (who Skippy is, the tools guide, answer routing, standing grants, the front-door documents), the 42 TOOL DEFINITIONS 17,426 (415 tokens each on average, mostly prose descriptions), the live tail ~880, the message ~30. The 17,000 rewritten every turn = the tool definitions: on this program's route to the model (the subscription door through the claude CLI) only the ONE cache point on the instructions is honoured — a marker on the tools, a 5-minute marker (this one 502'd the box for four minutes and was reverted) and a fixed anchor at the head of the conversation all moved nothing. So the cuts are structural, not cache tricks: (1) the spine — the HEALTH ENGINE lane is already shrinking the spine block to the hard flags only (Nick's ruling 21:46Z: the engine is the only door), which removes ~80,000 tokens a turn; (2) the tools — send each face only the tools a turn can use, and cut every description to a sentence; NEEDS NICK: a yes to trimming what Skippy is told about his tools. Instrumentation stays in server.js behind PROMPT_SIZE_DEBUG (off). Nested commit 7eb0926.

2026-09-12T22:36Z — STEP RECORD: HEALTH-ONE-DOOR touched server.js: 253598c (skippy-code main, pushed) · build fd4431efe77f (skippy-cloud, published with `node fly-publish.mjs`, clean tree, no dirty flag) — health answers route through engine_answer; the spine block carries hard flags only. Only the two fenced regions changed: the ANSWER_ROUTING_PROMPT health row, and the /api/chat non-Chantelle/non-Neeko block (snapshot push and its failed-to-load branch replaced by one static block: engine-only instruction, standing frame, three rules, six hard flags); the cache_control line untouched. Pass tests, live as Nick (evidence plans/HEALTH-ONE-DOOR/PLUMBING/evidence/phone-health-answer-live-2026-09-12.txt will carry them): "What was my most recent hs-CRP reading and what date was it?" → tools [engine_answer], answer "June 23rd — 0.69 mg/L…"; "In one sentence, what is the capital of Portugal?" → tools [], "Lisbon.".
  FOR THE DRIVER, TWO FINDINGS (not changed by me, outside the fence): (1) the CACHE GUARD now fails by construction: ops/test-cache-guard.mjs FLOOR = 100000, and the static prefix no longer carries the record snapshot (/data/brain copy 87,773 bytes). Two live runs after publish: tool round cacheRead 49,110 both times (cacheWrite 82,357 cold, 39,106 warm) — the prefix IS being reused, so this is a floor calibrated to the old prompt, not a broken cache; the floor wants re-setting to the new prefix size (or a ratio). The publish exits 1 on it after the deploy lands. (2) Three lines still tell Skippy the record is in its context and are yours to change: the Gate 0 bullet "HEALTH IS DIFFERENT NOW. The spine is ALREADY IN YOUR CONTEXT" (operating-brain prompt, ~line 2274), the read_file tool description "the spine block IS the source and you must answer FROM it" (~line 2700), and the VOICE path's second getHealthSpine() push (~line 16671). My block tells the model any such line is out of date, and the live test passes, but the three lines should go.

2026-09-12T23:35Z — THE VOICE PARITY AUDIT (Nick, 23:28Z: make Skippy's voice like ChatGPT voice; one Codex agent audits both sides). On file: evidence/VOICE-PARITY-AUDIT-2026-09-12.txt (Codex senior-engineer, public OpenAI docs cited, tonight's measurements as the evidence). Its answer on the subscription door: NO, it is not what stands between Skippy and ChatGPT-like voice — the front desk already runs on OpenAI's metered realtime API (first word 1.07–1.51 s); the door only serves the text brain behind it, and moving that to the API changes the first word by 0 s and the answer by an unmeasured amount. Its ranked differences, which become THE NEXT PHASE of this plan (a plan edit when the gate allows; recorded here now): (1) let the realtime model ANSWER general talk itself from a distilled Skippy profile and call one gateway only for Nick's records, memory writes and actions (12–20 h); (2) replace the 23,615-token persona with a nightly-refreshed voice profile under 2,000 tokens (8–12 h); (3) record gateways return a spoken-ready fact, not rows — one model round instead of two (16–28 h); (4) one typed gateway in front of the 42 tools (20–32 h; the trim tonight is the partial repair); (5) a three-way route inside the realtime model — GENERAL answers directly, RECORD reads, ACTION goes through approval (6–10 h); (6) two memory stores: the small voice profile for continuity, record retrieval only on demand (12–18 h); (7) the harness reports useful-first-audio, backend calls, tokens, cents and interruption yield, so a fast acknowledgement can never count as a fast answer (4–6 h). Target shape: under 5,000 tokens per turn.

2026-09-13T00:24Z — THE PURGE INVENTORY is on file (Larry, evidence/PURGE-INVENTORY-2026-09-13.txt, with the driver's corrections: the Cowork replacement already shipped tonight; the old Mac program is the OUTER skippy-app/server.js, not skippy-code). NEEDS NICK item 9: one reply approving the archive/unload list A–E in that file (moves into projects/_archive/ and one launchd unload; nothing deleted).

2026-09-13T01:31Z — VOICE FEEL: Nick pasted ChatGPT voice's acknowledgement guide and a transcript (evidence/CODEX-VOICE-ACK-GUIDE-AND-TRANSCRIPT-2026-09-12.txt) and ruled ~01:10Z that ASTRA (gpt-6-astra) optimises the voice system ("teraa is like sonnet its not smart"); Astra dispatched 01:11Z (record to evidence/ASTRA-VOICE-OPTIMISE-2026-09-13.txt). Boris wrote evidence/VOICE-FEEL-STRATEGY-2026-09-13.txt (53 KB) and its adversary FAILED it — NOT READY, marked at its top: it was written against a session contract Astra is replacing (protocol 2, a single route_turn tool with tool_choice required, a 700 ms turn timer, a hard-coded "One moment while I check that" line in the client), so its mechanics (C1–F) are rebuilt after Astra lands; its A (the behaviours from the transcript), B (Skippy's phrase banks), C0 (generate the first line in context) and C5/G (the truth boundary and the fences) stand.

2026-09-13T01:47Z — STEP 12 FIX LANDED (Opus, evidence/step12-fix-2026-09-13.txt, 30 KB; five nested commits pushed, published): nothing on /api/chat read Nick's words for a credential before the model ran — the only scanner sat inside captureMemory(), downstream of the model and only on the capture_memory tool path; now the floor check AND the health-marker scan run ahead of the route decision on every path (chat, relay, tool), the unlabelled generated-secret shape starts at 20 characters (a 12-character flight code is no longer refused), tests 71 → 80 ALL PASS with a red run on file; live: a marker request returns refused and the Mac file did not move, a two-word memory still proposes. Its blind checker found three real gaps the builder had missed (a bypass wording, the over-block, the relay branch returning before the gate) — all closed except the outer edge of a wording-based floor, where a second deterministic wall now sits (proven at module level; the model declined on its own both live tries). Open and recorded: that outer edge needs a different mechanism than more patterns; the OLD MAC PROGRAM carries no memory gate at all (purge list); the pre-publish boot check guards a file on the health lane's branch, not the one that deploys. The live re-run of STEP 12 (word nonce) waits until Astra's voice run ends, because it restarts the box. ASTRA at 01:46Z: its route commit and hardening patch landed by the driver in a clean copy (workspace 7fd2d694c2 + fa3d772080; family deploy 735463dd, cache v829, voice.js v53); brain d520d82 published by Astra.

2026-09-13T02:09Z — ASTRA'S RUN (gpt-6-astra, 01:11Z–02:06Z, evidence/ASTRA-VOICE-OPTIMISE-2026-09-13.txt, independently checked): route hardening LIVE (family deploy 735463dd with voice.js v53, brain d520d82 + e7fb9c5 published, cache guard PASS at 54,130); the twenty-turn acceptance runs could NOT complete — OpenAI's realtime API on Nick's key is Usage Tier 1 (40,000 tokens/min, 200 requests/min, 1,000 requests/DAY) and tonight's testing used the day's 1,000 (provider reply: rate_limit_exceeded, requests per day, Used 1000); interruption measured 396 ms against the 300 ms target (fail); the long 21.9 s utterance had no premature turn end; contextual first lines (the ChatGPT-style receipt/preamble) remain UNBUILT — three six-word receipts, two repeated; harness commits 73f42a1cd6 + 7025fb0e3b landed by the driver. NEEDS NICK item 10 (one of the four: money leaving): up to USD 50 of OpenAI API credits makes the account eligible for Tier 2 (200,000 tokens/min, 400 requests/min, no daily cap); without it the front desk stops answering once the day's 1,000 realtime requests are used — testing alone burned them tonight, and tomorrow's real use could. "10 yes" and an agent prepares the purchase page for you to complete (never the card details); "10 no" and the runs wait for the daily reset.

2026-09-13T02:31Z — NEEDS NICK ITEM 10 ANSWERED ("1 yes", ~02:20Z): USD 50.00 of OpenAI API credits bought on the card on file (invoice 0BIVDG41-0004, paid, tax 0); the Limits page now reads USAGE TIER 2 — gpt-realtime 200,000 tokens/min, 400 requests/min, no daily request cap shown — so the front desk's daily wall is gone and Astra's twenty-turn acceptance runs can run tonight (told to Astra through /tmp/astra-driver-reply.txt). ASTRA RE-DISPATCHED 02:12Z on Nick's word ("if astra needs more time give it more time - the job is quality outcomes not speedy testing"): four-hour cap, quality the bar; brief scratchpad/briefs/ASTRA-voice-optimise-2.txt — contextual first lines, interruption under 300 ms, streaming answer path, profile, harness --dry, landing package to /tmp/astra-landing-package.txt. SLACK ("2 pull it up"): the sign-in page is up in Nick's Chrome tab; on "signed in" the driver uploads the three chosen faces (item 3).
2026-09-13T02:43Z — STEP 19 CLOSED (see the step record): Nick signed in to Slack ("done", ~02:38Z); the driver uploaded the three chosen faces through his signed-in Chrome, read the icons back by API (three new hashes, no placeholder), and a cold Sonnet match put every face on the right app, 3 of 3.
2026-09-13T03:05Z — HAND-OFF FROM THE HEALTH ENGINE SESSION (HEALTH-ONE-DOOR lane), three items still open on the phone, all outside the fence you gave me (ANSWER_ROUTING_PROMPT + the /api/chat health block). Measured on skippy-code-publish b437285:
  (1) server.js:2281 — the Gate 0 line still reads "HEALTH IS DIFFERENT NOW. The spine is ALREADY IN YOUR CONTEXT", which contradicts the /api/chat health block (the engine is the only door; no record snapshot is in context).
  (2) server.js:17270 — the voice path still builds its prompt from getHealthSpine() (server.js:1631, the build's injection-context-core.md snapshot), so a voice health question can be answered from a frozen copy instead of engine_answer.
  (3) captures: the cloud engine's /api/chat returns captureProposals, but engine_answer passes only the text, so the family chat on the cloud path never shows a capture card; only the Mac fallback path (skippy-turn.mjs → the Mac bridge) does, and that writes to the Mac's store. Suggested: engine_answer carries the engine's captureProposals into the /api/chat reply, and the family app routes the tap (/api/engine/capture/apply) to skippy-engine.fly.dev — the cloud engine holds the waiting room and the store on its volume. I will make the family-app routing change the moment proposals come from the cloud.
  The cloud engine is now built and deployed from the cloud repository (engine-deploy.yml), serves today's sealed record, reads nothing beside it, and serves /api/protocol and /api/health stamped with the record's time. Offer: widen the fence to lines (1) and (2) and the engine_answer return shape for (3), and I make them under the same rules (git status clean, commit server.js alone, publish with fly-publish.mjs, tests, a STEP RECORD line). Otherwise they are yours.

2026-09-13T03:12Z — ASTRA'S SECOND PACKAGE LIVE: the family app serves voice.js v54 (cache deck-family-v830, Cloudflare deployment dee2d6f5, 03:04Z; landed from Astra's clone as a230a65875 with the harness commit before it and regenerated files 8dd310d1b1) — contextual first lines (protocol 3: the first line comes from the route decision, the streamed answer path, local audio stop before any socket send on interruption, 1024-sample mic packets), the brain side ee7b24a (compact current-source profile, 1,366 bytes) published by Astra; Astra's browser-stub proofs: twenty routes plus three classes exit 0, a 21 s utterance with two 1.1 s pauses produced no early reply, first-line grammar/repetition/claim checks PASS at module level; Terra's independent source verdict DEPLOY SAFE; the twenty-turn live acceptance runs now that Tier 2 is on (Astra continues; four-hour cap from 02:12Z). STEP 12 FIX 2 published 03:05Z (see the step record); Opus's record section was swallowed by a sync autostash and recovered from stash@{0} (workspace 4d66a9fe7b).
2026-09-13T05:15Z — ASTRA'S RUN ENDED 03:52Z ON CODEX'S USAGE LIMIT (provider message: "You've hit your usage limit ... try again at Sep 19th, 2026 7:27 AM"; 1,072,159 tokens in the run). Its record (evidence/ASTRA-VOICE-OPTIMISE-2026-09-13.txt, written every few minutes as Nick asked) is complete to 03:52Z. SHIPPED AND LIVE: brain ee7b24a (contextual receipts, 1,366-byte profile) and family v54 (protocol 3, streamed answers, local stop before cancel; deployment dee2d6f5), Astra's read-back of the live voice.js matched its clone byte for byte. MEASURED BEFORE THE STOP: the direct-brain twenty-turn run had all 20 routes and gateway counts right (strict 18/20 on two request-string mismatches); the deployed-proxy run stopped on its first turn — receipt 17.3 s, no gateway answer in 90 s — because the BOX WAS DEGRADED after the 03:05Z publish: event-loop stalls of 1–4 s (24 by 03:22Z, 115 by 04:59Z after a restart at 03:25Z), health checks flapping, 81–88 thread-stream clients held, and the Anthropic subscription accounts full so every brain turn went through the claude CLI door at 20–44 s; the front-desk mint returned 502 once at 03:47Z. So the live acceptance numbers are NOT MEASURABLE until the box is well. LEFT BEHIND BY ASTRA: (1) three harness-only commits in its clone not yet on origin (ca7d06ea82, 066e4ef81a, dfdd58399e — landed by the driver at 2026-09-13T05:15Z); (2) an UNVERIFIED voice.js candidate (reversible local pause on interruption, End-while-connecting cleanup, epoch guards; 347 lines) preserved on branch astra-pause-candidate-2026-09-13 (5f2884bef5), not deployed — its browser green/red was still pending; (3) the twenty-turn acceptance on the deployed proxy, twice, and the ≤300 ms interruption proof, all waiting on a healthy box. NEXT: the box's stalls go to an Opus builder now (server.js); STEP 12's fourth live run is running with 150 s request limits; the voice work that was Astra's is either held for Astra's return on 2026-09-19 or handed to Opus/Fable — NEEDS NICK item 11.
2026-09-13T05:44Z — THE BOX'S STALLS, CAUSE MEASURED AND FIX PUBLISHED (Opus builder + driver, evidence/BOX-STALLS-2026-09-13.txt): an instrument on the stall recorder (nested af9a6fa, ~05:25Z) named what held the loop — about eighty open app tabs each poll the board, and every GET /api/threads re-read and re-parsed the same ~550 KB of pushed thread files synchronously (86 identical jobs in one 4.4 s stall; 861 calls = 131 s blocked in a 13-minute baseline; 39 stalls in 12 minutes, worst 4,361 ms). The three candidates named by the driver were ruled out with numbers (no python3 on the box; recall never in a stall line; the loader and the brain pull are async). FIX: the board readers share the one cached reader keyed on the files' own mtime+size (nested 69efb72, pushed; published by the driver 05:42Z, cache guard PASS 54,466); red-tested both ways; responsiveness suite 15/15, threads/voice/memory/roundtrip ALL PASS. The twenty-minute after-measurement runs to ~06:03Z. Noted, not chased: the tabs POST the login route 158 times a window (client herd). STEP 6's checker (sonnet: proof re-run + one fresh dispatch) dispatched 2026-09-13T05:44Z. STEP 12 fix 3 (Opus, own clone) still building.
2026-09-13T06:36Z — STEP 12 CLOSED by its checker (see the step record). Steps closed: 1–5, 8–14, 16–20, 23, 25 (20 of 26). Open: 6 (fix building), 7 (waits on the VOICE lane's instrument), 15 (spoken half), 21/22/24 (the documentation gate, off at 12:44Z), 26 (final sign-off).
2026-09-13T06:50Z — STEP 6 CLOSED by its checker (see the step record); the handoff line is in the VOICE lane's record. Steps closed: 1–6, 8–14, 16–20, 23, 25 (21 of 26). Open: 7 (the shopping-list step waits on the VOICE lane's spoken-request instrument), 15 (the spoken half of talking to a running thread, needs an idle session), 21/22/24 (the documentation gate, off at 12:44Z: the six governed files from their drafts, the Gracie fence sentence, plan steps 27–32, the upkeep board write), 26 (postmortem and FINISH LINE sign-off). NEXT list additions tonight: the what's-running naming window; carrying confirmed memories onward to the personal store; the family app's login-route herd (158 POSTs a window from open tabs); Astra's unverified pause/resume candidate (branch astra-pause-candidate-2026-09-13).
2026-09-13T07:05Z — STEP 15 FIX DISPATCHED (Opus, brief scratchpad/briefs/S15-fix-opus.txt): establish the idle-session fact with one controlled send, make the reply reach the voice live for a running thread (the daemon watches the session's transcript after delivery), speak an honest sentence for an idle thread, and make the "landed" strip leave the Status tab (fidelity mismatched 0). NEEDS NICK item 12 asks whether the 30-second bar becomes "as soon as the thread answers".
2026-09-13T07:52Z — STEP 15 FIX BUILT (Opus, evidence/step15-fix-2026-09-13.txt). THE IDLE FACT IS ESTABLISHED: a session idle at its prompt accepts a socket user frame (sendToSession returns ok) and never processes it — two idle sessions sent one scratch line each at 07:09:56Z, neither transcript grew a byte in 90 s, while the same files recorded five queue lines of their own earlier that morning; the known-positive control is a mid-turn session that wrote the receipt in the SAME millisecond as the send (01:41:51.405Z) and absorbed it 3.1 s later. That is a platform limit inside Claude Code's turn handling, not ours. WHAT WAS BUILT: the delivery service now probes that receipt after every live delivery and reports it per row, so the app knows whether the thread TOOK the message and not merely that it was handed over; when it did, the service listens to that session's record for up to 10 minutes and hands the answer straight to a new fast lane (skippy-code-publish lib/thread-heard.mjs, POST /api/thread-heard) so it can be spoken at once instead of waiting up to 30 s for work-watch's index rebuild — there was no endpoint anywhere that could carry one fresh line. /api/thread-say now has three answers (reply, the idle sentence, the bounded-wait sentence), says which on X-Thread-Say-Kind, and returns BOTH moments of an answer (X-Thread-Answered-At = when the thread started saying it, X-Thread-Landed-At = when the answer existed, which is where the 30-second bar starts, because a reply is stamped at its start and written at its end). panel.js v89 asks that route directly every 2 s, says the honest sentence for an idle thread without waiting, bounds the wait at 10 minutes, and TAKES THE WAITING LINE OFF THE SCREEN IN ALL THREE ENDINGS. 🔴 THE CHECKER'S STRIP DIAGNOSIS WAS WRONG: measured on a freshly loaded live Status tab, the one visible .pf-strip is #mc-sub reading "Status is incomplete — One machine has stopped sending updates (Nicks-Mac-mini) — anything of its is older than 45 minutes", NOT the landed strip; that is a true fault report the design deliberately keeps painting, so Status reads mismatched 1 until that machine's data is fresh and no code change should remove it. The send strip is correctly hidden at rest and would only have become a SECOND mismatch after a send — that is the half that is fixed. 🔴 ALSO FOUND AND FIXED: the delivery service could not see its own new version (it asked the filesystem about a path with %20 in it, so the check threw and was swallowed) and had been up 9h48m running replaced code — proven both ways, fixed with fileURLToPath, and the restart now proven live (pid 99343 → 493 after a touch). Tests: thread reply watch 33 ALL PASS · thread voice sentences 29 ALL PASS · thread heard fast lane 24 ALL PASS · asset-version parity 12 passed 0 failed · all six changed files parse. NOT DONE, AND IT IS THE DRIVER'S: the live spoken proof cannot run until the brain (nested main 8f85b43) and the family app (panel.js v89 / deck-family-v831) are deployed — and at 07:52Z every session on this Mac except the driver's own was idle, so the checker also needs a genuinely running thread. Exact live recipe, both outcomes, in section 6 of the evidence file. Commits: workspace 4373eea, 95a2d23, 53930b7 · clone b2ae7a9, 8d95596, 8f85b43.

2026-09-13T08:31Z — STEP 15 live runs done (see the step record): idle half proven live, running half not measurable, item 12 with Nick. Steps closed: 1–6, 8–14, 16–20, 23, 25 (21 of 26). Open: 7 (VOICE lane's instrument), 15 (item 12), 21/22/24 (the documentation gate, off at 12:44Z), 26 (final sign-off). Findings for other lanes: Chantelle's Mac mini has not pushed thread data since 2026-08-31 (the Status tab says so); the family app's Cloudflare functions may lag a deploy by minutes.
2026-09-13T13:37Z — NICK'S ANSWERS (verbatim in evidence/NICK-VERBATIM-2026-09-12-second-answers.txt): the health record is open to Chantelle/Gracie as it always has been — a purge of every conflicting signal is dispatched; the purge list A–E approved (D held: kokoro is live; E not yet); the worker name stands; the remaining voice work goes to Astra on the other Codex account (~/.codex4); the STEP 15 bar becomes "spoken as soon as the thread answers". Nick also asked why agents keep numbering lists one way and asking for replies by another number — the driver's list said "reply 12 yes" inside item 5; from now on his list is answered by its own numbers only (memory feedback_list_numbers_are_the_reply_codes).
2026-09-13T15:11Z — ASTRA'S VOICE WORK LANDED AND DEPLOYED (third run, on the nickdeck19 Codex account after the first ran dry): interruption 144.3 ms measured in Chrome's audio graph (target 300), two consecutive twenty-turn deployed-proxy runs at 10/10 GENERAL direct and 10/10 RECORD through the gateway, a 23.3 s thought with two 1.1 s pauses held, first line 1.1–2.1 s; voice.js v55, cache deck-family-v832, Cloudflare deployment bf728796, origin/main 03bc0bae35 — no brain change. Astra's own ranked remainder (record evidence/ASTRA-VOICE-OPTIMISE-2026-09-13.txt, 633 lines): 1 one spoken request in twenty reached the gateway reworded; 2 record answers still a 14.7 s median (the brain's own speed); 3 end-of-speech detection adds ~1.7 s; 4 no backchannel yet; 5 two voices (Cedar for general, Onyx for gateway answers) to reconcile; 6 the "speaking" label sticks after audio drains; 7 a real-phone check for soft speech, room echo and false-alarm recovery. NICK, 2026-09-13 ~14:35Z: "can we test that in a format thats not so loud or silenced even?" — every voice test from now on runs with the browser muted (memory feedback_voice_tests_must_be_silent). CLEANED on his "3 clean": 15 GB Astra clone, the two builder clones and the landing worktree removed after their work was confirmed on the cloud copy; disk free 208 GB → 216 GB; scratchpad 5.6 GB → 135 MB. IDLE THREADS, approved by Nick ("ok approved", after asking why not always the front door): a message to an IDLE thread now goes through the FRONT door (resuming that conversation as a real turn, same memory, answers in seconds); a BUSY thread keeps the side door, which is the only way in while it is mid-turn — an Opus builder is building it now (brief scratchpad/briefs/S15-frontdoor-opus.txt).
2026-09-13T16:00Z — STEP RECORD (HEALTH-ONE-DOOR touched server.js, Nick "2 fix it all" on the three 03:05Z items): skippy-code e4e778b, published to skippy-cloud (cache guard PASS; live /app/server.js carries healthOneDoorBlock and takeCaptureProposals, no getHealthSpine, no "ALREADY IN YOUR CONTEXT"). What changed in server.js, and only this: (1) TOOLS_GUIDE Gate 0 now sends every health question to engine_answer; the ruled-out gate became the hard-flag gate (hard flags stay no; anything else he tried is cited from the engine and asked about, matching the three rules); "right now" uses the engine's number. (2) buildVoiceSystem no longer attaches the record snapshot; chat and voice both attach healthOneDoorBlock() (one function, the text that was inline in the /api/chat block, unchanged); the snapshot loader and its boot log lines are gone. (3) engine_answer carries the engine's captureProposals; the /api/chat route lifts them off the tool result before the model sees it and every reply carries them (max 5). Tests: _test-health-attach-identity.mjs updated (voice gets the block, never the snapshot; Neeko gets neither) 17/17, red 6 on the old code; new _test-engine-capture-cards.mjs 14/14, red on the old code; whole suite in matching copies old 39/41 → new 40/42, same two environment failures. Family app half: 3fb38bdc70 (relay passes cards on the cloud leg only; a tap goes to skippy-engine with its own key), deck-family deployment 8603a514, wall held; the cloud engine's confirm route answered a made-up card with 409 not_in_waiting_room (nothing saved). NOT done by this session: a live chat that actually produces a card — a test statement about his body could auto-log into Nick's real record, so the first real card will come from him.

2026-09-13T16:22Z — ASTRA'S FOURTH RUN, ITEMS 1 AND THE GENERATOR LANDED AND PUBLISHED (its own record evidence/ASTRA-VOICE-OPTIMISE-2026-09-13.txt, 694 lines, appended as it goes on Nick's instruction). ITEM 1, THE REWORDED REQUEST, CLOSED: the spoken request now reaches the gateway word for word — 13/13 deterministic (baseline 2/13) and 7/7 live muted microphone cases including the two that failed before (P07, N02: the model paraphrased them again, and the raw transcription still reached the gateway unchanged), independently reviewed. THE GENERATOR I BROKE IS FIXED: my STEP 22 persona move took the voice profile's source prose out of server.js, so ops/build-voice-profile.mjs threw ("voice profile source is missing \"Warm, grounded, direct, peer-level.\"" — reproduced first-hand by the driver at 15:50Z); it now derives an 1,800-byte profile from the assistant documents, with its own failure controls passing. THREE STALE DOCUMENT CLAIMS FIXED BY THE DRIVER (workspace 882df98455, found by Astra while repairing the generator, all true when the drafts were written and stale by the time they landed): skippy/AGENT.md said server.js was the single definition of who Skippy is (it now says this folder IS the definition, closed by STEP 22, with the live proof); MANUAL.md said all screens use the Pearl look (Pearl is RETIRED and must never be reintroduced); MANUAL.md said any action from a spoken command was NOT BUILT (it is, STEP 6, closed by its checker). PUBLISHED by Astra at ~16:25Z: nested main 8b4df6ba14 (exactly three files by commit --only — server.js, ops/build-voice-profile.mjs, ops/voice-profile.txt — verified by the driver at 16:21Z), deployment-01M2DS3GK8W5SPG1NRJ0XFE2H5, cache guard cacheRead 62840 > 20000 PASS. DRIVER'S INDEPENDENT LIVE TURN against that build, 2026-09-13T16:22Z: "who are you and who do you work for?" → "I'm Skippy — Nick's chief of staff, working for him and for Chantelle with equal authority." in 3.6 s; "what are the four things you always ask me before doing?" → money leaving, rotating a credential, irreversible destruction, a message sent as Nick — in 2.9 s. Both from the composed documents, and faster than the 4.0/5.5 s measured before this publish. THE BOX: Astra read two loop stalls and asked the box lane to account for them; measured by the driver at 16:01Z, window 15:47–15:58Z, stall lines 0 and health failures 0, current lag 0 ms — its two were BOOT stalls 15 s after a restart (module loading, the same one-off the 05:42Z fix record documents), and worstLag is a since-boot maximum, not a current reading; Astra corrected its own record. A RECORD-LOSS CAUSE NAMED: Astra's item-1 completion block vanished from its evidence file although the proofs existed — the file is only ever committed by this machine's hourly sync ("sync: working-tree snapshot from a nickdeck session"), and auto-pull rebases with autostash, so an append written between the park and the pop is lost; four autostash entries sit in the repo now, and the same mechanism ate a 240-line section from the STEP 12 builder earlier tonight. The rule for every lane: commit an evidence append the moment you write it, never wait for the sync. STILL OPEN of Astra's seven: 2 record-answer latency (its baseline today 28.4 s useful audio, 23.9 s gateway to first speech; the driver measured the brain's own plain turns at 6.9–8.1 s median 8.1 s, so the gap is in the gateway-to-speech leg; its first-sentence streaming draft is deliberately unpublished until the spans are measured, the finality guard preserved), 3 end-of-speech timing, 4 backchannel, 5 one voice (Cedar and Onyx), 6 the stuck speaking indicator, 7 the phone-shaped check. Its family-app work stays in its clone for one landing by the driver, and item 1's family commit a85e7c082e is NOT yet reachable from origin.
2026-09-13T16:33Z — THE SLACK ANSWER LAG IS FIXED AND INDEPENDENTLY VERIFIED BY THE DRIVER. Nick, 2026-09-13: "nevermind it took him 2 mins to reply so his lag time is the issue". CAUSE, from the archived envelope rather than a guess: Slack HEARD him in 0.19 s and the drain picked it up in 0.2 s — the queue was never the problem. The whole 90.6 s was ONE cloud-brain call running out its own 90 s ceiling (compose_why: timeout), after which he got a canned receipt that quoted his own words back at him, carried a raw Slack user id, and said his message was "on the pickup list". THE BUILDER MEASURED BEFORE CHANGING ANYTHING: 8 live brain calls returned between 3.6 s and 70.5 s, so lowering the ceiling would have converted his SLOWEST REAL ANSWERS into failures — the first attempt keeps 90 s untouched. What is recoverable is a fast far-end failure, and those now get exactly one retry; a timeout is never retried. AN INDEPENDENT CHECKER FOUND A REAL WORST-CASE DEFECT AND THE BUILDER FIXED IT RATHER THAN REPORTING IT (e1f83d3ad8): "they fail fast" was an observation, not a guarantee — a 502 returning SLOWLY, say at 88 s, is not classified a timeout, so it qualified for the 45 s retry and 88 + 45 = 133 s, LONGER than the wait the change exists to shorten. retryCeilingMs() now gives the retry only what is left of the same 90 s budget a single call always had, and skips it below 5 s. A SECOND CHECKER, BRIEFED ONLY ON THE REPAIR, RETURNED 8/8 PASS. THE DRIVER'S OWN 16 CHECKS, RUN AGAINST THE LIVE FILE AT 2026-09-13T16:33Z, ALL PASS: 88 s → no retry, 86 s → no retry, 85 s → exactly the 5 s left, 0.5 s → the full 45 s, undefined and negative → the full 45 s; swept 361 values of first-attempt spend, ZERO breaches of the 90 s budget, worst total observed exactly 90,000 ms; timeout and chat_timeout still not transient, chat_502 and chat_401 still are; and — the check that matters, because a repair that silently DISABLED the retry would pass every safety criterion — the retry is genuinely still wired (retryCeilingMs is called in drain(), the second compose uses retryMs, and it is gated on retryMs > 0). MEASURED END TO END: 90.6 s before, 17.4 s live after; a forced transient failure now recovers into a real answer in 3.5–6.7 s where it previously produced the canned receipt. A CONCURRENT LANE SILENTLY REVERTED BOTH EDITS between applying and committing, so an earlier commit described a retry it did not contain — only the LIVE test caught it, re-applied as e61cbe54ae. That is the same shared-checkout hazard that ate Astra's evidence append. Commits, all reachable from origin/main and verified by the driver: c0826cce89, 93480ee9a3, 8269f21245, e61cbe54ae, e1f83d3ad8; record c529348c75 / d529a0f7ea. Guard suites: three green; the fourth is 37 passed / 35 failed both BEFORE and AFTER, confirmed independently by running it at the pre-change commit — pre-existing, not caused here.
2026-09-13T16:33Z — A JARGON LEAK ONTO NICK'S SCREEN, FOUND OUTSIDE THE LANE FENCE AND CLOSED (103b8a8a9f). The Slack lane's two blind checkers reported the banned phrase surviving in projects/ops/skippy-jobs/jobs/slack-listener-deaf-watch.mjs, which puts an alert CARD ON NICK'S SCREEN. Three of its user-facing strings told him a message was "on the pickup list" — the codename for an internal queue file, which is exactly what the standing plain-English rule forbids in anything he reads. Reworded to "the queue it keeps of unanswered messages"; every claim, the 24-hour-clock warning and the card structure untouched, and the code comment recording the history deliberately keeps the old phrase because comments are read by agents, not by Nick. THE CHEAP LANE WAS RIGHT AND MY PROOF WAS WRONG — the trap on file: four attempts across zai and deepseek all returned 2 of 3, and the cause was my proof counting "into the queue" three times when the third sentence's own grammar correctly needs "in the queue". The edits were correct every time. Proof corrected to count the shared phrase, and it passes. THE CHECKERS' SECOND FINDING IS NOT REPRODUCIBLE HERE AND IS NOT BEING CHASED: they named a hardcoded "Weekly review is in the app:" prefix in a scheduled-task prompt; a sweep of this repo finds that sentence ONLY in an archived sent-message record and a dedupe state file, never in job code. Scheduled tasks live on another account (Nick, 2026-09-13: "we have all tasks already built on another account"), so it belongs there, not in this lane.
2026-09-13T16:33Z — CORRECTION TO MY OWN 16:30Z ENTRY, ISSUED BY ASTRA AND ACCEPTED. That entry recorded Astra's item-2 baseline as 28.4 s of useful audio. Astra has since measured a clean 7/7 accepted baseline at a median of 18.502 s and REFUSED TO CLAIM THE DIFFERENCE AS ITS FIX: 18.502 s is 34.94% below the earlier 28.439 s with NO latency optimisation deployed, so the gap is provider/service variation, not code. Its per-turn medians: provider work 8.916 s, tool work 6.475 s (the business narrative answer itself runs retrieval plus another model), first-final-text to safe output 0.922 s as an UPPER BOUND including remaining generation rather than pure buffering, network/client residual 0.146 s, speech 2.391 s. IT ALSO RETIRED AN INFERENCE OF MINE: my 6.9–8.1 s plain-brain measurement does NOT prove the 23.9 s gateway-to-speech leg was ours, because the business engine's own audit shows 6.16–7.21 s in the matching window with no internal sub-timings. Nick's original 14.733 s baseline needs 9.822 s to beat. ITEM 2 IS STILL OPEN AND ITS COLD REVIEW IS DOING ITS JOB: a fresh independent review FAILED the first-sentence splitter on two real defects a green proof had missed — common abbreviations (Inc., Sept.) and URL tokens ending in ? or ! were split mid-sentence — repaired with 29 abbreviations expanded and terminator ambiguity applied to all cases, and the guard policy is unchanged and unpublished until the spans are measured. Astra's evidence file is at 720 lines and is now COMMITTED AFTER EVERY APPEND, so the autostash hazard that ate its earlier block cannot recur.
2026-09-13T18:25Z — NOTE FROM THE HEALTH-ONE-DOOR SESSION FOR THE DRIVER: b4aba08deb (Astra's package, deployed 18:12Z as 38d33fe6) changed js/voice.js but left its number at v56, the number the food-logger publish (af965980, ~17:50Z) had just moved it to. Any phone that opened the family app between those two publishes has v56 cached by the service worker and keeps the pre-Astra voice.js until the number moves again; a fresh fetch (how the live hash was checked) does not show this. Suggested: bump voice.js to v57 in index.html and sw.js with a changelog line and republish through projects/ops/deploy.mjs (it passes --branch main and runs the drift guard). Nothing was changed by this session.

2026-09-13T18:21Z — ASTRA'S ITEM 6, THE STUCK SPEAKING INDICATOR, CLOSED WITH LIVE PROOF AGAINST THE REAL DEPLOYED BUILD. Landed and deployed by the driver in one pass with item 1 (see the VISUALS record for the deploy): one file, js/voice.js, workspace 4dc3365d19 into origin as b4aba08deb, Cloudflare deployment 38d33fe6, live SHA256 0846b737776f397b74001cc8da68eec67342367ed95729be71bf510864b45387. THE DRIVER CHECKED BEFORE LANDING RATHER THAN TAKING ASTRA'S WORD: the patch applied cleanly to the then-current head, the resulting file hashed to exactly the value Astra named, it still carried the peer VISUALS attach control added at 17:11Z, and it parsed. ASTRA THEN PROVED IT ON THE ACTUAL PRODUCTION ASSET WITH NO OVERRIDE, at 390x844 in a muted touch Chrome, with zero gateway calls and zero fixture errors — and the driver independently confirmed at 2026-09-13T18:21Z that 0846b737 is still exactly what family.heroesandsidekicks.io serves, so the thing it tested is the thing Nick has. Settlement PASS: a completed response.done at 16,106 ms with the final audible audio at 26,922.1 ms and the indicator eventually returning to Listening — which is the done-before-drain race repaired, proven on the real provider and app together rather than on a fixture. Synthetic interruption: final old audio 140 ms, provider start 255.8 ms, active stop, and ZERO old audio beyond provider plus 30 ms. On the real 19.99675 s clip: provider start 1,457.2 ms and final old audio 1,464.6 ms from injection. 🔴 AND IT REFUSED TO BANK THE HEADLINE NUMBER: the human speech onset in that clip is UNCONFIRMED, so it explicitly does NOT claim the 500 ms gate. It also caught its own phone reader passing on vacant output earlier and fixed the instrument before running — the second instrument fault it has found and not banked today. Its record is committed as it goes (5a6a6da24a, 799 lines), so the autostash hazard that ate an earlier append cannot recur. STILL OPEN of its seven: 3 (the speech-detection matrix returned NO setting satisfying under 1.2 s with no mid-clause cut — a real negative result, not a failure), 4 backchannel, 5 the voice choice, 7 the remaining simulated cell. THE VOICE CHOICE IS NICK'S AND IS NOW THE HIGHEST-VALUE UNKNOWN, from his own question today: "does chossing one of the newer voices give us more speed and accuracy?" Astra has proven it can REACH gpt-live-1 (session started 17:53:35.880Z, resolved gpt-live-1/Marin) and correctly refused to claim anything about quality or latency from an access check alone; it is now measuring native, Onyx-via-speech and Live on the same frozen fixtures for time to first speech and fidelity, with Live fragments explicitly marked non-authoritative. The driver will put the choice to Nick in one sentence once both numbers exist.
2026-09-13T18:26Z — TWO FINDINGS FROM ASTRA THAT MATTER BEYOND THE VOICE LANE, both recorded as it reported them. FIRST, A REAL PRODUCT FAULT, NOW ATTRIBUTED RATHER THAN LEFT AS A TIMEOUT: its fourth phone fixture — a QUIET / ECHOEY clip, the 20.05675 s one — FAILS TO INTERRUPT SKIPPY AT ALL. One bounded simulated diagnostic, no retry-to-green: the clip was injected and over a 12.002 s observation the provider registered 0 speech starts, 0 speech stops and 0 active stops, while 900 old-audio meter events kept arriving and the state stayed Speaking throughout — with zero provider errors, zero gateway calls and zero instrument errors, so nothing was broken except the outcome. IN PLAIN TERMS: when Skippy is talking and the person speaks QUIETLY, or the room echoes, he may simply not stop. The original timeout is retained rather than replaced, and Astra is explicit that a PHYSICAL room is NOT proven by a simulated clip — this is a simulated failure, and the real-room version of it is untested. SECOND, AND IT WOULD HAVE POISONED EVERY NUMBER THAT FOLLOWED: an independent review of Astra's own measuring instrument caught several FALSE-PASS bugs, including the fixture's own EXPECTED TEXT being seeded into the transcription it was supposed to be measuring — a rig grading itself on an answer it had been handed. Astra's parent took over and fixed them BEFORE any provider batch ran, so no quota and no money was spent producing results that would have been worthless. That is the third instrument fault it has found and refused to bank today. THE STANDING RULE THIS EARNS, for every lane: an instrument that has never been reviewed by something other than its author is not evidence yet. ALSO: item 6 closure agreed on both sides, and the A/B/C voice comparison will report three PATHS rather than three voices, naming any gain it cannot attribute to the voice alone as unattributable — the driver's instruction, because Nick asked about voices and the experiment compares whole chains.
2026-09-13T18:45Z — 🔴 NOTE FOR THE DRIVER FROM THE HEALTH-ONE-DOOR SESSION: your "DRIVER -> ASTRA" messages are being delivered to the HEALTH-ONE-DOOR session (Claude Code session claude-2-0-f2, local_95d2d606…), not to Astra. Five arrived here today (18:12Z, 18:22Z, 18:25Z, 18:28Z and earlier), including Nick's own instruction relayed at 18:28Z: "sounds good - needs to come with recommendatiosn about which to choose based on pros and cons of each including latecy accuracy and cost". This session has not acted on any of them. If Astra has not confirmed receipt, re-send those to Astra's own channel (its record under ~/.codex4 / evidence/ASTRA-*.txt, or Nick's relay).

2026-09-13T19:20Z — NOTE FROM THE HEALTH-ONE-DOOR SESSION: publishing the family app (chat's Mac fallback removed, food log quantities), the drift guard refused on js/voice.js (b4aba08de) and js/panel.js (a81cd7287), both changed after their numbers last moved. Bumped voice.js v56->v57 and panel.js v89->v90 in index.html and sw.js (cache deck-family-v834) and published with this change, so phones that had the old numbers cached now get your shipped code. Nothing in either file was changed. Publishing through projects/ops/deploy.mjs would have caught both at the time.

2026-09-13T21:08Z — ASTRA'S FOURTH RUN CLOSED. Six of its seven items are done and the seventh is open for a named reason. Its final evidence is committed at b86733d1b93c222d in evidence/ASTRA-VOICE-OPTIMISE-2026-09-13.txt. CLOSED: 1 request fidelity · 2 answer latency, closed honestly as a MEASURED SHORTFALL rather than a claimed win · 3 the speech-detection matrix, which returned a real negative (no setting satisfies under 1.2 s with no mid-clause cut) and was then SETTLED BY NICK rather than by measurement — his words, "assume id rather speak louder my environemtn usually has background noise with kids etc", which makes the high threshold and the app's own 0.015 floor correct for him and retires the quiet/echo fixture as unrepresentative · 5 the voice choice, also settled by Nick ("i want once voice - just make him cedar") · 6 the stuck speaking indicator, proven live on the real deployed build · 7 the phone-shaped check, PASS on a phone-sized live browser after Astra caught its own reader passing on vacant output and added strict state controls. OPEN: 4, backchannel. WHY IT IS OPEN AND WHY THAT IS RIGHT: the native listening-cue path works, but the audio it produced transcribes as "Mm-hmm. Mm-hmm. Take your time." — the model generating WORDS OF ITS OWN. That lands squarely on the exact-Skippy constraint the driver closed tonight, so Astra refused to ship it as identity-safe and flagged that any integration must coordinate with the voice-integrity lane rather than bypass it. That is the third time today it declined to bank something it could have. ALSO FROM ASTRA, AND IT UNBLOCKED A NICK DECISION: cedar is VERIFIED available on the read-aloud endpoint — one real call, gpt-4o-mini-tts with voice cedar, HTTP 200 with non-silent audio. The driver made the change on that evidence (nested 3a4d30a). A cost fact worth keeping: the newer full-duplex path bills per second INCLUDING SILENCE and open listening time, so idleness has to be managed and a call ended deliberately; Astra also proved an explicit close works, ~655 ms, with no billing runaway. NEXT FOR ASTRA: a controlled cedar backchannel, coordinated with the exact-Skippy fence. Not the microphone, not the voice choice, not replaying old cells.