Knowledge & Grading

The actual documents the agents read and work from, shown exactly as they are on disk β€” not a summary. See the progress view instead Β· All projects

Plan PLAN-7-KNOWLEDGE-GRADING.md

# SMP-7 Β· Knowledge grading and kids-app transcript closure

## πŸ”΄ MISSING-PROOF REGISTER β€” SMP-7 Knowledge & Grading β€” 2026-08-30 β€” FOR THE PLANNER, NOT FOR A BUILDER

**This is a RECORD, not a work order. Do not create any file listed below.** They are named so the
plan rework can decide what each should be. Nick's instruction, 2026-08-30: *"i told you to have the
planner do that - note they are missing and thats all."*

This list is generated by `check_plan.py`'s own `_dead_proof_paths()` β€” the same function the gate
uses β€” not by a second hand-rolled scan. An earlier attempt at a parallel scanner disagreed with the
checker in BOTH directions and was discarded; **one instrument, never two.** Re-derive with
`python3 projects/ops/agents/check_plan.py projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md`.

**TWO DIFFERENT PROBLEMS ARE LISTED HERE AND THEY NEED DIFFERENT ANSWERS.**

### B Β· EVIDENCE ARTEFACTS THAT DO NOT EXIST (1)

πŸ”΄ **These must NEVER be hand-written.** They are OUTPUTS β€” screenshots, reports, logs a step
produces when it runs. Creating them by hand is fabricating proof, the precise failure this
register exists to prevent.

    projects/personal/skippy-app/ala-state/agent-signals/smp7-garble-probe-real-event.json

For the planner: the defect is the PROOF SHAPE, not the missing file. A proof of the form
`test -s <artefact>` passes on any text at all and so cannot fail β€” and *a proof that cannot fail is
not a proof*. Each needs replacing with a command that RUNS the step, produces the artefact, and
has a stated result that would make it FAIL.



## πŸ”΄πŸ”΄ OVERSEER FINDINGS, 2026-08-30 β€” READ EVERY WORD BEFORE DRIVING ONE STEP OF THIS PLAN

Nothing below removes a stage, a capability or a goal. It records what was measured on 2026-08-30
and why the verdicts in this lane's STATE file cannot be taken at face value.

### FINDING 1 β€” THE ENTIRE 2026-08-30 RUN WAS MEASURED THROUGH A DEAD INSTRUMENT

Every lane in the programme reported it could not reach the things its proofs require. The overseer
session ran the SAME probes, on the SAME Mac, within minutes of reading those reports:

| Probe (run it yourself) | Every lane reported | Overseer measured, same machine, same hour |
|---|---|---|
| `curl -s -o /dev/null -w "%{http_code}" https://github.com/` | `Could not resolve host: github.com` | **200** |
| `curl -s -o /dev/null -w "%{http_code}" https://skippy-cloud.fly.dev/` | `Could not resolve host` | **401** β€” alive, wants auth |
| `node -e 'require("net").createServer().listen(31877,"127.0.0.1")'` | `listen EPERM: operation not permitted` | **bound successfully** |
| `ioreg -n Root -d1 -a \| awk '/CGSSessionScreenIsLocked/{getline; print}'` | `ioreg: error: can't open file.` | **exit 0, screen unlocked** |
| `screencapture -x /tmp/probe.png` | `could not create image from display` | **wrote 983,519 bytes** |

**The workers were sandboxed. The Mac was healthy throughout.** Roughly 182 UNPROVEN and 58 FAILED
verdicts across nine lanes describe the cage, not the product. Reproduce all five rows before acting
on any verdict in this lane. If they do not reproduce for you, YOU are in the sandbox too, and that
is the first thing to report β€” not a product finding.

**How this should have been caught, and by whom.** SMP-3 proved it rather than asserting it: its
known-positive control (`github.com`) failed identically to its real target. This workspace's own
rule says *pair every empty result with a known-positive control; a failed control is instrument
failure, never proof of absence.* That pairing already invalidated the run. The rule was applied
correctly at the step level in every lane and nobody applied it at the programme level. **That gap β€”
step-level discipline with no programme-level check β€” is the root process failure of the night.**

### FINDING 2 β€” EVERY BROWSER STEP FAILED ON AN UNANSWERED PERMISSION PROMPT, NOT ON CHROME

Verbatim, from the lanes' own evidence files:

    Browser Use rejected this action due to browser security policy.
    Reason: The user declined permission for this action.

and on a localhost target:

    Browser use cannot access http://localhost:3000 because the user denied permission for this request.

**Nick was asleep. Nobody declined anything.** The permission gate asks a human, receives no answer
because there is no human at the keyboard, and records that silence as a DECLINE.

Nick holds a STANDING, WRITTEN authorisation covering exactly this β€” agents drive his own apps as
him, click through his own sign-in gates without re-asking, and may look at anything on his screen
(`~/.claude/CLAUDE.md`, section "EVERYTHING ON NICK'S SCREEN IS APPROVED FOR AGENTS TO SEE"). The
gate never sees that ruling.

Confirmed instances in at least five lanes. Citations:
  Β· `projects/ops/smp-2-dispatch/STATE-2-DISPATCH.md:1148` β€” an INDEPENDENT CHECKER, not the builder:
    *"I connected to the available Chrome surface and made one read-only navigation attempt …
    browser policy returned: The user declined permission for this action. I did not use a workaround."*
  Β· `projects/ops/smp5-live-reliability/evidence/step-2-builder-20260830T0933Z.txt:7`
  Β· `projects/ops/skippy-master-plan/smp5b-reach-presence/evidence/step-14-leading-turn.txt:24`

**CREDIT WHERE IT IS DUE, AND DO NOT "FIX" THIS THE WRONG WAY.** Not one agent routed around the
gate. They recorded the refusal, refused to claim a live result, marked the step UNPROVEN and moved
on. That is correct. The remedy is a narrow, explicit pre-authorisation β€” never teaching agents to
bypass a permission gate, and never blanket browser approval: these signed-in surfaces reach the
business hub and messaging paths that can text real people as Nick, i.e. all four of the things that
require him personally.

### FINDING 3 β€” SOME PLANS NAME PROOF FILES THAT WERE NEVER WRITTEN

Repo-wide, **127 of 439 proof files named by live plans do not exist (29%)**. A worker sent to run
one gets `MODULE_NOT_FOUND`, which reads as a product failure and is recorded as one. Two gates now
refuse this at authoring time: `projects/ops/skippy-jobs/_test-plan-proofs-exist.mjs` and a
document-wide check inside `projects/ops/agents/check_plan.py`.

Per-lane counts measured 2026-08-30 (scan controls passed both ways β€” a known-present file resolved,
a known-absent one did not):

πŸ”΄ THE OVERSEER'S FIRST COUNT OF THIS WAS TOO LOW AND `check_plan.py` CAUGHT IT. The overseer
scanned only for `.mjs/.js/.py/.sh` paths starting with `projects/`. The real gate also catches
EVIDENCE artefacts a DONE-PROOF names (`.png`, `.md`, `.jsonl`, `.json`) and bare filenames with no
directory prefix. Authoritative counts, straight from `python3 projects/ops/agents/check_plan.py`:

    SMP-1 Talk          12 missing       SMP-5A Reliability   2 missing
    SMP-2 Dispatch       0 missing       SMP-6 Reach          9 missing
    SMP-3 Status         2 missing       SMP-7 Knowledge      2 missing
    SMP-4 Desktop       15 missing       SMP-8 Watchers       1 missing

**SMP-2 is the ONLY lane at zero.** An earlier overseer message called SMP-7 and SMP-8 "clean and
drivable with no plan repair" β€” that was wrong, and the correction is recorded here rather than
quietly fixed. Both carry missing proof artefacts and both need those written before their affected
steps can close. Re-derive any figure above with the command named; do not trust this table on its
own β€” that is the exact mistake it exists to record.

### MANDATORY PREFLIGHT β€” ADDED 2026-08-30, APPLIES TO EVERY STEP IN THIS PLAN

Before executing ANY step, the worker runs, as its first action: reach a known-good host Β· bind a
throwaway port on 127.0.0.1 Β· read the screen-lock state. Any failure means record
`NOT MEASURABLE FROM HERE β€” <which probe failed>` and stop. **A worker may never record a product
verdict through an instrument it has not shown to be working.** This preflight matters more than any
individual step below.

### FINDING 4 β€” HALF OF THIS LANE IS GENUINELY FINISHED; THE REST IS NEARLY CLEAR BUT NOT AT ZERO

Measured 2026-08-30: all 8 named SCRIPT files exist. πŸ”΄ BUT `check_plan.py` finds **2 missing proof
artefacts** this plan names β€” `projects/personal/family-app/specs/REBRAND-AND-NATIVE-SPEC.md` and
`projects/personal/skippy-app/ala-state/agent-signals/smp7-garble-probe-real-event.json`. An earlier
overseer message called this lane "clean and drivable with no plan repair"; that was WRONG and is
corrected here rather than quietly fixed. Those two must exist before their steps can close. This
is still among the least-blocked lanes β€” but it is not at zero, and only SMP-2 is.

**PIECE (a), GRADING β€” DONE AND REAL, NOT A CLAIM.** A live, dated grading pass across 3 of 3
domains, $2.04 spent, evidence recorded in the lane's own artefacts. Nick confirmed the
five-question rubric wording stays as-is, directly, on 2026-08-26. Do not re-open or re-litigate it.

**PIECE (b), THE KIDS APP β€” OPEN.** The garbled and repeated transcripts, Steps 6–11. Ownership was
settled by Nick on 2026-08-27: it belongs to this lane, under the same worker as piece (a). No other
chunk owns it.

### FINDING 5 β€” ONLY ONE STEP IS TOUCHED BY THE TWO ROOT CAUSES

Step 5 (browser-native capture) is blocked by Finding 2, the permission refusal. Nothing else in
this lane depends on the browser, and its non-live steps are unaffected by Finding 1. **This is the
lane to drive first once a working session exists**, because the least of it is blocked.

### CURRENT POSITION

Piece (a) closed. Steps 6–11 not started.


**Owner:** Boris, SMP-7 single writer. **Overseer:** current SMP coordinator, re-verified at every handoff. **Design authority:** Sienna for any screen claim. No step starts before its enter gate is PROVEN and independently closed.

## Already true

- Current state is 17 PROVEN / 235 UNPROVEN / 18 REFUTED; no old completion prose is proof β€” evidence: `STATE-7-KNOWLEDGE-GRADING.md`.
- Stage 7 is explicitly not started; its Stage-4 probe is not a substitute β€” evidence: `projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft`.
- Kids-app deployment, destination delivery, real-device capture, and unsuppressed watcher alert delivery remain UNPROVEN β€” evidence: `STATE-7-KNOWLEDGE-GRADING.md`.
- The kids app is deliberately not Skippy: no engine, conversation, health, financial, or family-context wiring β€” evidence: `STATE-7-KNOWLEDGE-GRADING.md`.

## 0 Β· Gate Zero receipts

- Failure Mode Registry loaded: `.claude/skills/plan/references/failure-registry.md`, 2026-08-29; Β§4 maps every entry mechanically.
- Canonical specs loaded: `.claude/skills/plan/SKILL.md`, `projects/ops/agents/check_plan.py`, this chunk's state record, canonical Stage-7 draft, learning-app source/test, watcher and deploy surfaces.
- Ownership check: this plan extends the existing Stage-7 record, `projects/personal/learning-app/`, and `projects/ops/skippy-jobs/`; it creates no rubric, app, store, watcher, scheduler, or plan rival.
- Expected inputs confirmed to exist: Steps 1 and 4–8 reopen the draft, response destination, learning-app files, test, watcher job/test/state, runner, KV path, and governed deploy path before relying on any one.
- PROMPT-SPEC scan (P1–P7): no unresolved build-changing ambiguity. Nick alone supplies grading answers and normal real-device use; both are V1 acts, neither authorizes agents to invent an answer or substitute a synthetic call.

## 1 Β· Goal and definition of done

- **HOW IT'S USED:** Nick gives the existing five answers for each grading lane; Noah or Willow uses the current kids app normally and Nick receives one clean transcript/one safety alert. **HOW WE KNOW:** Steps 3, 8, and 9 independently reopen the response/destination records.
- **WHAT IT LOOKS LIKE:** a complete Grade-7 record with every β€œno” routed to its owner, and no family-visible repeated-prefix transcript. **HOW WE KNOW:** Step 3 counts real answers; Steps 7–9 run red/green/sabotage and real-path proof.
- **WHERE IT LIVES:** the existing Stage-7 response destination found in Step 1; `https://skippytutor.pages.dev/`, opened by Noah/Willow and read by Nick in the existing family destination. **HOW WE KNOW:** source/destination evidence is saved in Steps 1, 8 and 9.
- **WHAT IT MUST DO:** make grading runnable and recorded; prove watcher alert delivery; prove browser-native capture; fix only evidence-supported transcript behavior; prove real-shape RED/GREEN/sabotage and production alert delivery; deploy exact bytes; prove a real-device clean family outcome; retire temporary diagnostic through owner; hand off unchanged state. **HOW WE KNOW:** Β§Β§2, 3b and 6.
- **NOT in scope:** rewriting/answering the rubric; child-facing Skippy/personal/health/family wiring; reviving session cron or `wrangler tail`; Gracie, Neeko, or unrelated learning-app changes; disabling a safety guard. Discoveries go as one dated line to owner (security: `projects/ops/sp-sec/PLAN.md`), never a local fix.
- **REPLACING / RETIRING:** temporary KV watch stays alive until real-path proof. Watchers infra alone retires its schedule in Step 10; session cron and tail remain retired.

## 1a Β· Critical variables β€” the confirmation sheet is GENERATED from this table

| # | The variable, in plain words | Value chosen | Alternatives rejected | Class | HOW WE KNOW | Cost if wrong | CONFIRMED |
|---|---|---|---|---|---|---|---|
| 1 | SURFACE β€” which screen this lands on, and who opens it | existing kids app; Noah/Willow use it, Nick reads its existing destination | test-only page | V1 | real bug screenshots and assignment | fix wrong surface | Nick, 2026-08-27 |
| 2 | Grade-7 answers | Nick's exact five answers per lane | agents infer/average answers | V1 | canonical Stage-7 close condition | fabricated approval | Nick, 2026-08-27 |

**Considered and ruled NOT critical:** paths, worker URL, namespace, deployed revision, and coordinator identity are V2: reopen the thing, do not ask a person.

## 1b Β· Subproject decomposition

- **SINGLE SUBPROJECT:** Stage-7 grading and the transcript closure share this chunk's final outcome and coordinator handoff; neither is a complete SMP-7 delivery by itself.

## 2 Β· The complete UX map (this becomes the test manifest verbatim)

| Id | Screen / entry point | State (defaultΒ·emptyΒ·errorΒ·loading) | Element / interaction | Expected behavior | Navigation from β†’ to |
|---|---|---|---|---|---|
| UX-1 | Stage-7 rubric | empty = no record; ready = questions/routes known | open canonical stage | five questions per lane; no answer invented | draft β†’ response record |
| UX-2 | Grade-7 sitting | valid / missing / no route | Nick answers | actual answer or exact owner-stage route | question β†’ owner route |
| UX-3 | KV watch | baseline / novel / timeout | controlled novel event | separate reader sees alert inside bound | KV β†’ signal store |
| UX-4 | real app capture | unavailable / device / error | microphone control and normal use | browser-native path writes distinguishable diagnostic event | device β†’ KV |
| UX-5 | transcript + safety regression | old red / new green / sabotage red | real `r[0].transcript`, safety utterance | clean join and one production alert | source β†’ destination |
| UX-6 | production + family result | loading / HTTP error / 200 / new event | governed deploy and normal use | exact bytes and clean family-visible result | source β†’ Pages β†’ KV β†’ family |
| UX-7 | closeout | watch active / retired / mismatch | owner cleanup + byte comparison | only post-proof retirement, containment, exact handoff | proof β†’ owner β†’ coordinator |

## 3 Β· Lanes and frozen contracts

| Lane | Scope (in / out) | Owner | Definition of done | Model (explicit) |
|---|---|---|---|---|
| Grading | existing rubric/record only; never answer authoring | Boris | real per-lane responses independently reopened | Qwen |
| Watch/capture | existing job, KV, browser routes; no new scheduler/tail | Boris | durable independent alert and real-device evidence | GLM 5.3 / Qwen |
| Transcript closure | existing children-app/deploy/containment only | Boris | real clean result then owner-controlled cleanup | Qwen |

**Frozen contracts:** `dbg:garble:` is the diagnostic prefix; existing family destination is the delivery surface; `runner.mjs` belongs to Watchers infra; no child surface receives Skippy/personal/health/family data.

## 3b Β· Execution map

A task is DONE only when its review-ledger row is CLOSED by a reviewer that is not the builder.

| Stage | # | Task | Gate to enter | EXECUTOR | CHECKER | DONE-PROOF | Ends when |
|---|---:|---|---|---|---|---|---|
| Grading | 1 | Re-derive Stage-7 contract | none β€” start here | Qwen | Sonnet | `command grep -nE '^### Stage 7|five questions|not started' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft` | scope/destination saved |
| Grading | 2 | Prepare, do not invent, the sitting | 1 | GLM 5.3 | Sonnet | `command grep -n 'STEP 3 INPUT' projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md` | exact record contract exists |
| Grading | 3 | Record/reopen Stage-7 responses | 2 | Qwen | Sonnet | `test -f projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft` | every lane answered/routed |
| Watch | 4 | Prove unsuppressed watcher delivery | 3 answered | GLM 5.3 | Sonnet | `node projects/ops/skippy-jobs/_test-watch-garble-probe.mjs` | novel key reaches reader |
| Capture | 5 | Prove browser-native capture or exhaust routes | 4 | Qwen | Sonnet | watcher state JSON has `checkedAt`, `knownBaseline`, and `hasUnexplained`, plus Step-5 browser/control content | live status honestly stated |

### Execution evidence and named handoffs β€” applies to Steps 1–11

`PLAN-7-KNOWLEDGE-GRADING.md` is the existing, approved durable record for this chunk.  It is the **only** execution-evidence writer: Boris appends a dated `STEP N EVIDENCE` block below `## STEPS` after receiving literal stdout/stderr, screenshots, or destination readbacks from an executor.  Builders and checkers never write this plan and never create an evidence file.  A block contains the command or UI route, UTC time, exit/status, redacted content result, known-positive control (where an empty result is possible), reviewer verdict, and the precise next action on failure.  A block that omits one field is UNPROVEN.

The pre-existing `evidence/` directory is historical material only; no step may create a new file there.  The only persistent machine-owned state in this plan is `projects/ops/skippy-jobs/state/watch-garble-probe.json`, written solely by `node projects/ops/skippy-jobs/jobs/watch-garble-probe.mjs --live`; every step may read it but may not write it.  This resolves the writer fence: Boris writes the plan, the watcher writes its state, and no other step writes either.

**Named external routes:** the controlled diagnostic writer is `POST https://skippytutor.pages.dev/debug/garble-probe` with non-private control text; the reader is the same Worker’s authenticated `GET /debug/garble-probe?k=<RESULTS_KEY>` using the key loaded inside the existing deploy runner, never printed.  The watcher’s separately readable signal route is `projects/personal/skippy-app/ala-state/agent-signals/smp7-garble-probe-real-event.json`, created by the existing `raise-signal.mjs` path.  The safety destination is the app’s own `escalateToParents()` route: its Slack webhook posts as the application to `#skippy-school`, its relay is `PARENT_ALERT_URL`, and its durable reader is the parent page’s `alert:` record in the existing RESULTS KV.  It never sends a message as Nick.

**Dependency rule:** each `RUNNABLE WHEN` names the exact access it needs.  A satisfied enter gate does not make a missing dependency look like β€œnot started”: record `BLOCKED β€” <named dependency>; verified by <command/control>; next action <literal action>` in the plan evidence block, then continue independent eligible steps.  The authoritative coordinator lookup is the live-agent list: `ListAgents` must return the live SP-G overseer session; send to that returned session with `SendMessage`.  `projects/ops/REBUILD-2026-08-21/STATE.md` is the positive-control instruction that requires this lookup.  No matching live session means Step 10 is BLOCKED and no handoff is attempted.

### STEP 1 β€” Re-derive the Stage-7 rubric, lanes, routes, and record location

**Enter this step when:** none β€” start here. **RUNNABLE WHEN:** `command test -f projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft` exits 0 and `command grep -nE '^### Stage 7|five questions|not started|six lanes'` returns at least one matching line.  The grep is the known-positive control; an empty result blocks Step 1 rather than claiming the draft lacks Stage 7.

**Builder:** Qwen. **Checker:** Sonnet, a different read-only session, instructed to REFUTE.

**Files you may touch:** canonical draft read-only; Boris may write this plan’s existing `STEP 1 EVIDENCE` block only. **Never:** canonical draft (owner: grading-spec owner), learning-app tree (owner: transcript lane), a new record, or `evidence/`.

**Do exactly this:**
1. Run `command grep -nE '^### Stage 7|five questions|not started|six lanes' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft`.
2. Run `sed -n '702,747p' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft`. **Range confirmation:** fresh heading search found `### Stage 7` at line 702 and the next peer heading, `### Stage 8`, at line 748; 702–747 is therefore the complete Stage-7 section, not the unrelated old 816–865 excerpt.
3. Give the complete literal stdout, stderr, and exit status to Boris; Boris writes it into the existing `STEP 1 EVIDENCE` block, including the matching-line count and the `test -f` control.

**PROOF:** saved output names Stage 7, five questions, lane count, route rules and current unchecked condition; artefact opens. **Instrument:** source `grep` and `sed`, executing the rubric-read path. **FAILS IF β€” Nick's words:** β€œwe are calling it ready without knowing what it asks or where answers go.” **Evidence:** ARTIFACT SAVED at the named path.

**If it fails:** preserve exact error, post one line to grading-spec owner, leave Step 2 blocked; proceed only with an independent ready step.

**Checker's job:** rerun both reads, open artefact, and refute any claim Stage 7 is complete.

**Handoff:** Boris writes `STEP 1 closed <date> β€” Stage-7 contract re-derived; Step 2 may prepare it.` in this plan's `## STEPS` section.

### STEP 2 β€” Prepare the existing Grade-7 sitting without inventing a question or answer

**Enter this step when:** Step 1 PROVEN. **RUNNABLE WHEN:** Step 1’s evidence block names the exact existing response-record path and the canonical draft remains readable; the literal `STEP 3 INPUT` text is supplied to Boris from that block.  No new credential, app, ledger, file, or writer is needed.

**Builder:** GLM 5.3. **Checker:** Sonnet, a different read-only session, instructed to REFUTE.

**Files you may touch:** this plan only. **Never:** canonical draft (owner: grading-spec owner), durable response record before Nick answers, or kids-app files.

**Do exactly this:**
1. Add a `STEP 3 INPUT` block below this step quoting exact questions, lane identifiers, β€œno” route, and existing response destination from Step 1.
2. Run `command grep -nE 'Stage 7|five questions|no.*lane|not started' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft`; give literal output and exit to Boris for the existing `STEP 2 EVIDENCE` block.
3. Run `command grep -n 'STEP 3 INPUT' projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md`.

**PROOF:** source excerpt and plan block agree on questions/lanes/routes/destination and contain no response. **Instrument:** source/plan `grep`, executing preparation only. **FAILS IF β€” Nick's words:** β€œthe agents filled in my grade or invented a rubric.” **Evidence:** ARTIFACT SAVED.

**If it fails:** restore plan block from source, do not open the sitting, notify grading-spec owner; Step 3 remains blocked.

**Checker's job:** compare source and block line-by-line; refute any added item or answer.

**Handoff:** post `STEP 2 closed <date> β€” existing Stage-7 contract is runnable and contains no invented answer.`

### STEP 3 β€” Conduct and independently reopen the real Stage-7 sitting

**Enter this step when:** Step 2 PROVEN. **RUNNABLE WHEN:** (a) Step 2 identifies one existing durable response-record path and its normal writer, (b) Qwen, GLM 5.3, and DeepSeek can each perform its named read route without exposing response contents, and (c) Nick is available for the V1 answers.  The three agents are dispatched by Boris through their existing Codex sessions; their returned literals go only to Boris’s plan block.  A missing route is BLOCKED, not β€œnot started.”

**Builder:** Qwen. **Checker:** Sonnet, a different read-only session, instructed to REFUTE.

**Files you may touch:** only the existing durable response record named by Step 2, through its named normal writer; Boris may write this plan’s evidence block. **Never:** a new ledger, rubric draft, children-app files, or `evidence/`.

**Do exactly this:**
1. Before involving Nick, Qwen, GLM 5.3, and DeepSeek each independently open/search the Step-2-named response route; each returns command/tool, exit, and a redacted result to Boris for one `STEP 3 EVIDENCE` block.
2. If all three find the route and only answers are missing, present exactly the existing five questions per lane in one sitting and store answers only in that existing destination.
3. For every β€œno”, write the exact owner-stage route; do not average/reword it.
4. Sonnet returns a redacted readback to Boris; Boris records it in the same evidence block.

**PROOF:** three distinct preflight artefacts; each lane has five actual answers or an owner route; Sonnet reopens the durable record. **Instrument:** three independent route reads plus durable record reader. **FAILS IF β€” Nick's words:** β€œyou said it was graded when I never gave grades, or turned no into yes.” **Evidence:** ARTIFACT SAVED; signed-in capture may be DESCRIBED, NOT PRESERVED only to protect private responses.

**If it fails:** preserve completed actual answers only, mark remaining lanes UNPROVEN, post coordinator line, and let independent Step 4 continue.

**Checker's job:** reopen record; count five responses per lane; find altered/missing no-route; verify all preflight artefacts.

**Handoff:** post to verified coordinator: `STEP 3 closed <date> β€” Stage-7 record independently reopened; unresolved routes: <list/none>.`

### STEP 4 β€” Prove the existing watcher alerts a separate reader

**Enter this step when:** Step 3 PROVEN. **RUNNABLE WHEN:** (a) the watcher script and state path exist, (b) `projects/personal/skippy-app/.env` is readable only to the existing watcher/deploy-runner process that loads `CLOUDFLARE_API_TOKEN`, (c) `curl -fsS https://skippytutor.pages.dev/debug/garble-probe` reaches the deployed writer, and (d) the existing agent-signal directory is writable by `raise-signal.mjs`.  The first successful baseline list is the positive control for the remote KV reader.

**Builder:** GLM 5.3. **Checker:** Sonnet, a different read-only session, instructed to REFUTE.

**Files you may touch:** no source. The existing watcher alone may write its state and existing signal route; Boris may write this plan evidence. **Never:** `runner.mjs`, kids-app source, a scheduler, direct KV credentials, or `evidence/`.

**Do exactly this:**
1. Run the watcher’s own remote list path: `node projects/ops/skippy-jobs/jobs/watch-garble-probe.mjs --live`.  Record its complete result and the state JSON readback in `STEP 4 EVIDENCE`; it must establish a non-empty known-baseline control before any empty result can count.
2. Write one controlled novel event through the real deployed writer, not a fixture: `curl -fsS -X POST https://skippytutor.pages.dev/debug/garble-probe -H 'content-type: application/json' --data '{"marker":"SMP7-control","point":"control","raw":"smp7 controlled probe"}'`; then rerun the watcher with a 60-second wall-clock bound.
3. Sonnet reads the existing `smp7-garble-probe-real-event` signal through its named signal path and confirms the control marker; Boris records the redacted result.  The watcher owner clears that same signal with `raise-signal.mjs --clear smp7-garble-probe-real-event --reason 'SMP-7 controlled probe complete'`, then reruns the watcher.  No raw KV deletion is permitted.

**PROOF:** known baseline control non-empty; novel key is new; watcher completes ≀60 seconds; separate reader sees alert; cleanup succeeds. **Instrument:** real Wrangler KV, watcher, existing signal reader β€” actual alert code path. **FAILS IF β€” Nick's words:** β€œthe watch went quiet and nobody knew.” **Evidence:** ARTIFACT SAVED at named repo paths.

**If it fails:** record exact error and whether writer, watcher, signal, or cleanup failed; leave watch active; send the named owner one executable request: `Watchers infra: run the Step-4 command, return watcher exit, state JSON, and signal-path readback; repair only the failed named component.` Steps 5/9 cannot treat silence as evidence.

**Checker's job:** rerun baseline, inspect novel output/readback, challenge timeout/cleanup, and reject classification-only evidence.

**Handoff:** `STEP 4 closed <date> β€” controlled new KV event reached independent reader; later real event may count.`

### STEP 5 β€” Prove browser-native capture or exhaust the real machine routes

**Enter this step when:** Step 4 PROVEN. **RUNNABLE WHEN:** one of three named pre-existing instruments is usable without changing permissions: the in-app Browser logged-in session with microphone permission, local Chrome GUI with microphone permission, or macOS GUI/AppleScript with an existing Chrome profile.  Before an β€œunavailable” result, that same instrument must open `https://skippytutor.pages.dev/` and report the page title as its known-positive control.

**Builder:** Qwen. **Checker:** Sonnet, a different read-only session, instructed to REFUTE.

**Files you may touch:** no source, browser permission, watcher state, or new evidence file; Boris may write this plan evidence. **Never:** app source, browser/mic permissions, a new diagnostic, a safety guard, or `evidence/`.

**Do exactly this:**
1. Drive the existing logged-in browser session at `https://skippytutor.pages.dev/`: focus before each click, use the existing microphone control, and return a redacted screenshot plus coordinate/action log to Boris for `STEP 5 EVIDENCE`.
2. If unavailable, Qwen uses the in-app Browser control, GLM 5.3 uses local Chrome GUI, and DeepSeek uses AppleScript GUI.  Each first opens the public URL (positive control), then reports literal availability/permission state to Boris; none changes a permission.
3. Reopen the watcher state.  If a real-device `dbg:garble:` event exists, Boris records only timestamp/marker classification and the comparison to the Step-4 control window.

**PROOF:** either a real browser/device interaction makes a distinguishable destination event, or three named controlled attempts show unavailable instruments and live event is honestly UNPROVEN. **Instrument:** browser/microphone and destination; direct function invocation never counts. **FAILS IF β€” Nick's words:** β€œyou called a fake scripted call the kids using the app.” **Evidence:** ARTIFACT SAVED; signed-in child capture may be DESCRIBED, NOT PRESERVED to protect private content.

**If it fails:** do not ask a human before all three attempts; then one ordinary-use request, with attempts/results, is permitted. Continue source diagnosis only if it makes no live claim.

**Checker's job:** retry real browser route; verify three controls if blocked; reject source/direct-invocation substitutes and empty destination output without baseline.

**Handoff:** `STEP 5 closed <date> β€” browser-native capture <proven/unproven>; evidence <path>; real-event <status>.`

## 4 Β· Regret Check

| Failure mode (registry entry) | The measure in THIS plan that prevents it | Where it lives |
|---|---|---|
| Every Failure Mode Registry entry | **Registry coverage is written in the existing plan’s Step-11 evidence block:** each registry row is mapped to an applicable numbered step or `N/A: this plan does not perform the act`. A blank row fails review. | Step 11 evidence block |
| Fabricated proof text | Builder evidence cannot close: Sonnet reruns instrument/code path and cited artefacts must open. | every step |
| Empty result read as answer | Every empty KV/search result has a known-positive control on the same instrument. | Steps 1, 4, 5, 8, 10 |
| Watcher dies silently | Controlled novel event, bounded run, separate signal reader before real silence is read. | Step 4 |
| Deployment accepted as delivery | Live bytes plus destination-side family read and real event. | Steps 8–9 |
| Test-local alert counted as production | Production alert destination is exercised; local counter is not accepted. | Step 7 |
+| A second system was built because the first was invisible | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A capability was declared impossible from a stale or unverified claim | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An absence was asserted without opening the store that would hold it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A known constraint's reason was lost, and it silently capped the product | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An instruction assumed capacity the executor doesn't have | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Expectations/manifest rows carried no grounding | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Work was written to a queue no reader ever visits | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A detector's death was invisible because only its target read it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A decision settled once re-opened elsewhere, or two copies of a rule disagreed | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A rule constraining the user turned out to be an agent's invention | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Remediation was ordered with diagnosis last | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A document, label, or comment was believed over the live system | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A proposal was sold on a capability never opened and read | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A cause was named and acted on without eliminating alternatives | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The human was asked a question the record already answers | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A spec and its guard were authored by the same hand and ratified the same defect | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Session rules never reached the subagents doing the work | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| One rule was blanket-applied across items needing per-item answers | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Pattern-matching scoped too loosely produced false connections | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Rules existed but were psychologically dormant at answer-time | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A run exceeded its cost/time ceiling or hung unbounded | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A helper was dispatched on a brief with a wrong or missing constraint | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A claim about the user/system was made without its source | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A conclusion was drawn from a partial read | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A fact was quoted as current without its date | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A computed value never reached the persistent record | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A missing lookup key fell back silently to a wrong default | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A hardcoded identifier broke when the referent was recreated | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A placeholder or wrong-level path shipped as a literal instruction | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A UI reported success while the backend silently failed | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Mid-session state was assumed unchanged | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Uncertainty was silently absorbed instead of marked | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A serial multi-step operation blew its time budget | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An external action went unlogged and became unrecoverable | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A tool's own description contradicted house reality and won | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Personal/identifying data exposed, or a record written to the wrong subject | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| One instance of a defect class was fixed while its siblings stayed broken | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A read operation mutated state | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The three biggest absence-claims variants: empty result, broken probe, discarded stderr | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A generated mirror was hand-edited, or its generator never re-ran | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Deployed config silently diverged from source config | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A delivery path was reordered and its notification behavior changed | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A critical boundary was config-editable and could be silently widened | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A "growing" archive had actually frozen | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Files were archived but their citations kept pointing at them | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A pipeline broke silently and looked identical to a working one | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Output was delivered somewhere the intended reader never looks | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Concurrent sessions clobbered each other's work in a shared file | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An enforcement gate covered fewer paths than its rule, or failed open | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Identity or authority was read from a value the caller supplies | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A new failure state was detected but reached no human | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The builder graded its own work and passed it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A check existed that could not fail | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The review didn't cover the shipped artifact | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A narrowing/refactoring change broke the cases that were already correct | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A check's verdict depended on wall-clock, machine load, or a concurrent writer | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A test existed but nothing ran it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An interactive element or view shipped untested / unseen | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Coverage was reported optimistically | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A staleness/freshness check used the wrong proxy | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A quantitative claim shipped without its method | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Done was declared before the live surface was checked | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A biometric/metric overrode the human's stated reality | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A correlation was asserted as a cause | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A nuanced reality was collapsed into a clean binary | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A recommendation repeated something already tried, uncited | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A wrong record was disclaimed instead of corrected | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Open items were re-typed from memory and drifted | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A deliverable was referenced instead of delivered | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A report used names/shorthand only the writer understood | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Commands were sent to a surface that can't run them | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A number was published without the population it was counted over | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A finding existed only in the session's output and died with it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The plan named a target with total precision, and the target was wrong | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The human approved a summary, and the summary was silent on the deciding variable | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A project stated its scope and never its anti-scope, and lanes leaked into adjacent work | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A new rule was written as prose inside its own fix, with nothing enforcing it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A confirmation was satisfied by checking the wrong kind of fact | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A blocker common to every lane was carved out of all of them and given to nobody | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Lanes were built to stop: one pass, land, idle β€” while fixed ceremony ate the context | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A caveat nobody measured travelled as fact through multiple independent lanes | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The environment destroyed work silently, and the lane wrote a wrong lesson from it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A specification described ONE lifecycle in several places, and the copies drifted independently β€” four consecutive cold reviews each found ~5-8 blocking ambiguities, because every patch added another partial description of the same state machine | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A task brief on an existing project was treated as the plan, and a generated status checklist was treated as the task list | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A regression test's "red-proof" failed for a reason unrelated to the thing it claimed to prove, twice in one session, two different mechanisms | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A standing instruction to route work to an outside/cheap engine eroded over a long session into doing the work directly | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A plan's own second line named a different document as the authority, and the reader proceeded without opening it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A live bug got three consecutive confident wrong-or-unproven diagnoses, two claiming live verification | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Fourteen guards stayed green all day while the live screen showed the wrong thing | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An agent was accused of fabricating its report because a narrow search failed to find the file it cited | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A tool's failure verdict was believed without checking the disk β€” and separately, a success verdict shipped a syntax error | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A build with several independently-shippable pieces was planned and run as one monolithic project, too large for one agent to hold | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A rule written only in prose, with no template slot and no machine gate, behaved as if it didn't exist | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A row-quality check counted TOTAL filled cells instead of checking the specific columns it claimed to require | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Three independent readers reported wildly different "% complete" for the exact same objective state β€” twice, on two different subprojects | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A V2 "opened it, here's what I saw" confirmation was wrong three separate times because it opened the WRONG PATH β€” the plan's own stated location, never independently rediscovered | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A shared coordination file used by several subprojects at once had no per-subproject write fence, and one subproject's list silently filled with rows belonging to the others | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The single cheapest, most decisive test of a build's core hypothesis was defined at planning time (correctly) but not RUN until after most of the build effort was already spent | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A dispatched build agent reported an interim status ("build is in progress, will resume once a Monitor delivers the completion notification") as its FINAL answer and returned, instead of waiting for the real result | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A sandbox restriction produced the EXACT error text this same repo's own CLAUDE.md already documents as a sign of a genuinely broken machine ("chrome exited early, code null" / Chrome preflight failure), and it was initially read as that known problem rather than investigated as a new one | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A paid external tool (Codex CLI) ran out of its own usage quota mid-build, and the agent that hit the limit chose to switch to running the command directly via its own Bash tool instead of the mandated Codex path β€” correctly, but this is a real, recurring risk that needs a standing rule, not a one-off judgment call | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A card-creation script reported success ("card opened... read back and confirmed") and its own internal counter incremented, but the card did not actually exist on live re-query β€” twice, for two different cards, requiring full manual re-creation | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Three separate, independently-fatal wiring gaps each made the same feature (the ai-builds board) non-functional in a different way, and NONE of them were caught by a passing `build-dist.js` run, any TIER-1 or TIER-2 gate, or any API-level check | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A real, deployed code fix (the three fixes directly above) did not reach a real user's already-open browser tab, even after that user hard-refreshed multiple times | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A correct, intentional, previously-ruled-on design decision (the task screen's default view narrows to "my own tasks" even for leadership identities) was mistaken for a bug because it was checked from only ONE identity's login | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The Updates panel β€” the actual surface a person opens to read what an agent posted about a card β€” is wired to Monday.com sync data ONLY, and an app-native card (this entire board) has no Monday board behind it, so it will read "No Monday updates on record for this item" FOREVER, regardless of how many real, correctly-formatted updates were posted server-side | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A pure oversight/QA dispatch (re-run four questions, grade the answers, write nothing) was refused twice in a row by the WORK-TYPE gate as "unclear," burning two full agent-spawn round-trips before the actual task began | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The same brief, past the work-type gate, was then refused by a SEPARATE gate for missing the ~6,000-word MACHINE-RULES travel block β€” a requirement with no automatic injection and no template a brief author can copy from without hitting the refusal first | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A fix (new SYSTEM-prompt grounding rules) was drafted, partially applied to disk, and left in a syntactically-valid but COMPLETELY UNVERIFIED state when the tool writing it (Codex CLI) hit its own account-wide usage cap mid-task | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The above fix's failure was found ONLY because a second, genuinely fresh-context pass re-ran the real test live β€” the first pass's own self-check (syntax valid, code present) had already been satisfied and would have been reported "done" without it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A confirmed, applied data fix was verified as working because it had only been applied to ONE of two live copies of the same data (production) β€” the copy actually being tested against (staging) still held the old, wrong text | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A 16-question regression suite meant to catch exactly this bug class had been silently crashing on question 1 and reporting nothing useful for a full day, because a dependency it called gained a new required argument and nobody re-ran the suite after that change landed | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Two entire bodies of real, load-bearing work β€” a 34-file answer pipeline and this drive's own PLAN.md/STATE.md tracking pair β€” had never been committed to git, on any machine, the whole time they were being built, found only by accident while fixing something else | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A request to deepen an existing artifact was answered by re-polishing the context already in hand, while named, existing sources were never opened | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A gate protecting one specific, highly sensitive file covered some tool surfaces (Write/Edit/MultiEdit) but not others (Bash), and the gap sat honestly documented in the file's own header for a day before being closed | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A function parameter's DEFAULT value silently made an entire decision branch unreachable, under a fully green test suite, since the day the branch was written | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A write-then-rename ("atomic write") pattern was used to update one row in a file that has a SECOND, independent writer appending new rows β€” the pattern is genuinely atomic against a torn read, and genuinely loses any row the other writer appended during the read-modify-write window | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A test suite's own "red-proof" claimed a safety property held ("removing the fix would fail the test") without ever actually removing the fix and running the suite | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Test files that exercised a shared module's logging path wrote real output into the REAL production log file, even though every other piece of test state (queue, tickets, journal) was correctly scoped to scratch directories | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An identity verified once, in memory, from a live authenticated source, was designed to be re-derived later from a file any process could write β€” which would have made the file, not the live authentication, the actual source of trust | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A background daemon process registered a global crash-and-exit handler for unhandled promise rejections; a later feature fired a promise without a `.catch()` in that same process, meaning any transient failure in that one feature (a network timeout) would have crashed the ENTIRE daemon, including everything unrelated it was doing | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A build's supersession of one design ("a standalone daemon" β†’ "extend the existing listener") correctly re-scoped every task around the new mechanism's natural shape, and in doing so quietly dropped a piece of functionality that had no obvious home in the new shape | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| `fs.watch()` on a shared state directory was assumed to be a sufficient delivery trigger, and was not β€” under real concurrent load from ~235 other sessions writing to sibling files in the same directory, two real queued requests sat with zero fs.watch event ever firing | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A plan asserted facts about the repo it never checked β€” one step named a symbol that travels under a different name; another's file fence named a file that does not exist (merges log items A3, A4, D2) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The program fixed what was BROKEN instead of building what was ASKED FOR β€” a day's good work landed on a component its own plan retires (log item J1) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A plan passed every gate β€” well-formed steps, real proofs β€” and still could not deliver what the user asked for (log item J2) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An assistant's first-person account of its own failure was taken as the root cause by every reader, and it was false (log item J3) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Three verifications were real and all three had the wrong SCOPE: verifying a quote is not verifying the claim; verifying a file once is not verifying it now; verifying the code path is not verifying the thing (merges H1, H2, H3 β€” one defect, three extents) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An orchestrator's confident relay propagated a wrong conclusion to five sessions faster than any plan could β€” a real acceptance criterion was deleted on it β€” and the builder that refused the relay with evidence was right (merges F1, J4) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| One writer in three read the same handoff as a gate and serialized nine of fourteen steps behind another chunk's tenth step (log item F3) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Every failure mode of the file-approval machinery was silent: an approved-once path became permanently un-requestable; a legitimate handoff into a shared governed file consumed another chunk's pending approval; approval never notified the requester; one approval unlocked exactly one edit operation, losing a two-part edit's second half; and a plan tracker named STATE.md missed the PLAN-shaped free-edit carve-out, costing ~10 approval taps in one evening (merges B1, B2, B3, B4, I2) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A governance CLI silently dropped unrecognized flags (exit 0), let a two-token flag value overwrite the file path, let --reason swallow the next flag as its value, and its own written spec documented the broken form in two copies (merges C1, C2, C3, C4) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Plan shape existed as convention, not enforcement: plans degenerated into 1,000-line session logs; the plan template itself failed the machine gate; the checker validates a plan's parts, never its shape (merges A1, A2, D1) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A punchlist item condensed to six words pointed its reader at exactly the wrong action β€” implementing it literally would have silently rerouted every assistant reply into manual approval (log item I3) | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A production secret read as SET when its value was EMPTY, and every check agreed with the wrong answer for 90 minutes across three sessions | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The SAME claim, on the SAME evidence, was CONFIRMED by a checker asked to verify it and REFUTED by a checker asked to break it β€” and the refuting one was right | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Reasoning ABOUT a system instead of ASKING it β€” the single most repeated failure of the 2026-08-27/28 night, four times across three different sessions, every time producing a confident and wrong claim from real evidence | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A hard prerequisite discovered AFTER a decision, with no owner assigned, silently converts a made decision into an unimplementable one | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A relayed instruction is acted on, or held, by whether the RELAY ITSELF could be the attack β€” and sessions had no test for that, so they either obeyed every relay or refused every relay | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Two independent programs audited themselves on the same night and found the same disease β€” every instrument reported a state that was not the system's state β€” while both had been reading the reports as ground truth | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A PROOF block read as complete while still containing its own template placeholders β€” four times in one plan, and the shape is mechanically detectable | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Real evidence, deliberately destroyed for a good reason, is indistinguishable from evidence that never existed | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A capability was ruled impossible on the strength of a query that structurally could not see the answer β€” the same shape as an earlier logged incident, on a different tool, and it was not recognised | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| The instruments used to verify a UI lie in four distinct ways, and a "drive the real surface" standard that does not name them produces confident false results | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A step's entry gate was satisfied and the step still could not run, and the format had nowhere to say so | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An automated proof's own internal check detected failure and the surrounding pipeline logged success anyway β€” the checking logic and the reporting logic disagreed, and reporting won | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A dispatch gate blocked the exact defensive pattern its own preceding line prescribed, for the exact reason that pattern exists | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A fallback held in place to make a cutover safe was itself the reason the cutover could never succeed β€” every retry failed, and each failure made the fallback look more necessary | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An approved instruction was correct when it was approved and harmful by the time it could be delivered β€” and every existing rule for handling relayed instructions asked only whether it was AUTHENTIC, never whether it was still TRUE | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| "I fixed the file" Β· "I deployed it" Β· "that is what the user sees" are THREE different claims, and a chunk can be right about the first two and wrong about the third β€” the gap is a client cache that no repo read, no deploy log and no server-side fetch can see | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| In a multi-session build, code read from the working tree is not the state of the system β€” it may be another session's half-finished fix, and reading it as established behaviour produces a confident diagnosis of a bug that does not exist | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Three successive rounds of fixes each produced an honest, passing proof, and the user's original complaint was untouched by all three β€” because every proof measured the mechanism the fixer had chosen to fix, never the sentence the user actually said | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A correct local caution was escalated into a fleet-wide halt across eight sessions on a crisis that did not exist β€” and the escalation priced only one side of the decision | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An overseer reported two pieces of work as missing because no message about them had reached its inbox β€” both had landed, were logged with dates and real terms, and one had already passed a full triad | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An acknowledgement from the system under test was read as evidence of the outcome β€” the same word, `queued`, covered a genuine pass and a silent 40-minute failure on the same endpoint the same night | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An overseer authorized an action by bridging a DIFFERENT ruling of the user's onto the question β€” reasoning correctly from a real quote that was about something else, three relay hops from where it was said | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An agent, blocked by a safety guard mid-test, offered the user a choice between loosening the guard and accepting weaker proof β€” presenting a load-bearing protection as one of two equal options | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A fault that repairs itself faster than anyone reports it is invisible to every alarm in the system β€” two family-facing surfaces cut out roughly twice a day for a MONTH and nobody escalated once | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A relayed approval was acted on as if the work were still outstanding β€” and the same file had already been written, by the session doing the relaying | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An investigator noticed that a metric could not possibly detect what it was being asked to detect, WROTE THAT DOWN, and then built a headline claim on it anyway β€” because the number it produced agreed with the conclusion | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An investigation's own searches and relays contaminated the evidence it was searching for β€” 80 of 84 occurrences of the string were manufactured by the act of investigating it | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| Three unrelated lanes in one night each ran an honest check against an intermittent fault and each got a clean answer, because a point-in-time probe is mathematically almost certain to miss a fault that heals itself | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| An overseer holding the user's GENUINE first-hand instructions relayed them as authority to four sessions β€” and one correctly refused, because accuracy and standing are different things and only one of them travels | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |
| A file that documents its own version history in prose ABOVE its code turns every unanchored search into a lie β€” three sessions in one hour read the changelog and believed it was the declaration | Steps 1–11: named model, literal instrument, adversarial Sonnet rerun, durable artefact, and a fail branch; N/A only where this chunk does not perform the registry act. | Β§3b and matching STEP block |

## 5 Β· Topology and roles

- **OVERSEER-AUTHORITY:** none named; coordinator identity is opened before handoff. Four approval classes remain fixed.
- Thread layout: one writing thread; every Sonnet checker is a fresh read-only session with no plan-edit instruction.
- **STATE FILE:** this plan's `## STEPS` section.
- **HEARTBEAT ROW:** `smp-7-knowledge-grading` in `projects/personal/skippy-app/ala-state/work-threads.json`.
- **MORNING-REPORT LINE:** `SMP-7 (Knowledge grading): <step> β€” <verified>/<total> steps proven β€” blocked: <list or none>` in `projects/ops/walkaway/REPORT.md`.

| Stage | Overseer | Sub-overseers | Workers |
|---|---:|---:|---:|
| 1–3 grading | 1 | 0 | 2 |
| 4–5 watch/capture | 1 | 0 | 2 |

## 6 Β· Evals

| Capability | Check | Pass looks like |
|---|---|---|
| grading rubric | Steps 1–3 exact source/record procedure | five actual answers or owner route per lane |
| E-1 probe/watcher | Step 4 novel-key procedure | independent durable alert read |
| E-2 real capture | Step 5 browser/device procedure | distinguishable non-synthetic diagnostic event |

## If you get stuck (all steps)

Try a concrete workaround, reread proof requirements, post one line to coordinator and blocker owner, then log `STEP N BLOCKED β€” tried: a,b,c. Need: one sentence.` Continue any independent eligible work.

## Your loop

Find the lowest eligible step, build, save durable proof, have Sonnet refute it, update its matching `## STEPS` line, repeat.

## SUMMARY β€” a few plain-English lines, read by the status generator

This is a clean rebuild from the verified demotions. It makes the rubric genuinely gradeable and proves the full kids-app path, including delivery, before removing the temporary watch.

## STEPS β€” the live status checklist, read by `status-regen.mjs` / `project-status-page.py`

1. Re-derive Grade-7 contract β€” UNPROVEN (2026-08-29: the mandated `sed -n '816,865p'` excerpt does not contain the Stage-7 contract; see STEP 1 EVIDENCE)
   DEFINITION OF DONE: source/destination evidence saved and independently reread.
   PROOF: `command grep -nE '^### Stage 7|five questions|not started' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft`
2. Prepare Grade-7 sitting β€” 0%
   DEFINITION OF DONE: exact source-derived contract contains no answers.
   PROOF: `command grep -n 'STEP 3 INPUT' projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md`
3. Conduct/reopen Grade-7 sitting β€” 0%
   DEFINITION OF DONE: every lane has five actual answers or owner route.
   PROOF: durable record readback named in Step 2.
4. Prove watcher alert delivery β€” 0%
   DEFINITION OF DONE: novel event reaches separate reader inside bound.
   PROOF: `node projects/ops/skippy-jobs/_test-watch-garble-probe.mjs`
5. [UI] Prove browser-native capture or exhaust routes β€” 0%
   DEFINITION OF DONE: browser event or three controlled agent attempts.
   PROOF: Step 5 named evidence.

### STEP 1 EVIDENCE β€” 2026-08-29 β€” UNPROVEN

**Evidence state:** UNPROVEN. The durable receipt is this existing plan at `projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md`; it preserves the literal reads below. No new evidence file was created.

**Run context:** `/Users/nickdeck/Documents/Claude 2.0`; 2026-08-29 America/Cancun. The canonical draft was read only.

**Exact command run:**

```sh
command test -f projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft
step1_test_status=$?
command grep -nE '^### Stage 7|five questions|not started|six lanes' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft
step1_grep_status=$?
sed -n '816,865p' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft
step1_sed_status=$?
echo "test -f exit: $step1_test_status"
echo "grep exit: $step1_grep_status"
echo "sed exit: $step1_sed_status"
```

**Literal stdout/stderr and exit status:**

```text
55:- **WHAT IT LOOKS LIKE:** six lanes each with a green heartbeat row that has once been made to go red on purpose; a real reply in a real thread; the desktop corner widget visible over other apps; a dated Stage-7 record of five yes answers per lane. HOW WE KNOW: Β§1's five measurable done bullets + Stage 2's corrected real-surface test (Nick's 2026-08-21 verbatim, quoted there).
91:| U10 | The real app, Stage-7 sitting | default | Nick answers five questions per lane, live, one sitting | Six rows, five yes/no answers each, dated, in his own words where he adds any β€” πŸ”΄ **CORRECTED 2026-08-23 (third pass): wording alone does not prove NICK answered** (no authorship/provenance check exists today, and the family app/hub share one login with Chantelle) β€” see Stage 7's own checklist for the honest gap and the strongest available mitigation (captured live, inside a real-time conversation turn on one of his three named surfaces, session/thread named alongside the date) | Stage 7 β†’ REBUILD-RECORD.md |
115:  the six lanes as polish, and the phone-transport re-verification in this spec (Stage 0) is
136:| 1 | What "phone transport" means for this spec's re-verification | Nick-reachability on his phone, via the native-push plan | Full PSTN telephony (place/receive real calls) | V2 | `skippy.md` corrections #2 and #4 scope everything past the six lanes as deferred polish (πŸ”΄ pre-2026-08-22 content β€” the live file is now a pointer; recoverable at `projects/ops/rules-registry/backups/20260822-strip-agents-skills/agents/skippy.md` or `git show 5b7295e685:.claude/agents/skippy.md`, verified present 2026-08-23); `projects/personal/family-app/specs/REBRAND-AND-NATIVE-SPEC.md` Β§4 item 3 is the live push-build record | CONFIRMED |
702:### Stage 7 β€” Nick's grade: dialed in and approved for action
704:**What this is:** the finish line is Nick personally answering five questions, once per lane, in
706:lane that has never been graded by him is not done. These five questions are constructed for this
711:**The five questions, asked per lane:**
728:- [ ] Nick answers all five questions, per lane, in one sitting, on the real app.
745:- [ ] Once all six lanes are six-for-six, record "dialed in and approved for action" with the
939:  0/3 β€” not started, Nick-gated. Stage 8: 0/4 β€” correctly still gated closed.
`skippy-acts-health.mjs`'s entry from `runner.mjs`'s schedule for one run, without telling anyone
building the meta-watcher in advance. `skippy-watchers-health.mjs` must name the missing watcher
by file name in its next run β€” not report a generic "5 of 5" that silently became "4 of 5 plus a
rounding pass." Restore the entry; confirm the report clears on the following run. This is the
proof that the sixth lane watches for real, not the proof that a job exists.

---

## 6 Β· DEPENDENCIES

| Waits on | What must be true before this spec proceeds | Why |
|---|---|---|
| **SP-4** (family app strip) | The "explicit updates only" gate is landed in `functions/api/_classify-ask.js` (or its replacement) | Stage 3's integration test needs a real gate to test against β€” this spec does not build one |
| **SP-7** (the voice app) | The transport bake-off has named a winner | Stage 2 cannot pick or test a transport SP-7 has not yet chosen β€” πŸ”΄ **DE FACTO CLEARED 2026-08-24, see Stage 2's own opening note and Β§3 row 4**: SP-7's formal bake-off still has not run, but live evidence already settles the winner (the existing Talk pipeline, Fly leg), so this row no longer blocks Stage 2 in practice |

Not listed in the master plan's own dependency row for SP-8, but true of the whole programme:
**SP-0** (routing/budgets/secrets) and **SP-1** (the sign-off queue) run first for every
subproject, this one included β€” Stage 7's sign-off is recorded through SP-1's queue mechanics,
and every model call in this spec's own build work routes through SP-0's router.

**Never touched by this spec, regardless of dependency state:** the credential floor named in
Section 6 of `MACHINE-RULES.md`'s travel block β€” no watcher built here rotates, requests, or
prints a credential to prove itself; a broken-credential test uses a disposable test path, never a
real one.

---

## 7 Β· OWNER + WRITE FENCE

From the master plan's Β§4 table, SP-8 row, restated exactly:

| Owner | Writes only inside | Waits on |
|---|---|---|
| **PS** β€” a Personal-account session | Agent and watcher files: `.claude/agents/skippy.md` itself, and new watcher scripts under `projects/ops/skippy-jobs/jobs/` | SP-4, SP-7 |

This spec never writes inside `projects/personal/family-app/`, `projects/business/business-app/`,
or any Gracie/Neeko-named path β€” those belong to other rows in the same table. Where this spec
needs something from one of those, it tests against what that row has already landed; it does not
build a second copy.

---

**Left UNCONFIRMED for Nick:** whether the five-question rubric in Stage 7 should be his own
literal words from a document I could not locate on disk β€” I searched for a standalone
"five-question rubric" or "dialed in and approved for action" file and found neither; the
questions above are built from three things he has actually said, cited inline, and should be
corrected in his own words if the construction is wrong. Also left open: which of Chantelle's four
drafted opt-in choices (Stage 8) she actually wants worded that way β€” that draft is research, not
her answer.
test -f exit: 0
grep exit: 0
sed exit: 0
```

**Known-positive control and match count:**

```sh
command grep -nE '^### Stage 7|five questions|not started|six lanes' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft | /usr/bin/wc -l
step1_count_status=$?
echo "grep-to-wc exit: $step1_count_status"
```

```text
      11
grep-to-wc exit: 0
```

The earlier standalone `grep` exited 0, so the pipe's last-command exit is not being used as evidence of the grep's status. Its 11 matching lines prove the reader can find the named Stage-7 terms. This is the control beside the failed contract excerpt.

**Result:** prerequisite file test passed; the known-positive search found 11 lines and the Stage-7 heading at line 702. The required `sed -n '816,865p'` command also exited 0, but its literal output begins in a watcher test description and contains only dependency/ownership material, not the Stage-7 questions, route rules, or unchecked Stage-7 condition. The mandated excerpt therefore cannot prove the specified contract. Step 1 is UNPROVEN; Step 2 remains blocked.

**What could not be obtained:** a prescribed line-range read that names Stage 7, its five questions, lane count, route rules, and current unchecked condition. Reason: those requirements are at/near lines 702–745, while the required range is 816–865. No source was altered to make the receipt appear to pass.

**Reviewer verdict:** not obtained in this execution. The step requires Sonnet, a different read-only session, to rerun both reads and refute any completion claim. Until that occurs and the range discrepancy is resolved in the plan by its owner, this remains UNPROVEN.

**Owner lookup command and literal output:**

```sh
git log -1 --format='author=%an <%ae>%ncommit=%H%nsubject=%s' -- projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft
```

```text
author=Nick Deck <nickdeck@Nicks-Mac-Studio.local>
commit=61b957bd7dd48e5c38c8218cc41e5094b7c7c6ec
subject=SP-8: agent-driven mic test PASSED for real (speaker->Yeti->transcript, no human); records what is genuinely still unproven and why
```

**Live-session lookup:** `ListAgents` returned only `/root` (running). No live grading-spec-owner agent/session was available to receive the failure post, and Nick is asleep; no message was sent.

**Precise next action on failure:** grading-spec owner must correct Step 1's prescribed source range to the actual Stage-7 section, then rerun Step 1 with a fresh independent checker.

**Receipt write verification command and literal output:**

```sh
git diff --check -- projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md projects/ops/skippy-master-plan/smp7-knowledge-grading/STATE-7-KNOWLEDGE-GRADING.md
receipt_diff_check_status=$?
command grep -n '^### STEP 1 EVIDENCE' projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md
receipt_plan_control_status=$?
command grep -n '^## Step execution receipt' projects/ops/skippy-master-plan/smp7-knowledge-grading/STATE-7-KNOWLEDGE-GRADING.md
receipt_state_control_status=$?
echo "git diff --check exit: $receipt_diff_check_status"
echo "plan receipt control exit: $receipt_plan_control_status"
echo "state receipt control exit: $receipt_state_control_status"
```

```text
418:### STEP 1 EVIDENCE β€” 2026-08-29 β€” UNPROVEN
74:## Step execution receipt β€” 2026-08-29
git diff --check exit: 0
plan receipt control exit: 0
state receipt control exit: 0
```

#### 2026-08-30 corrected-range re-execution β€” UNPROVEN

**Evidence state:** UNPROVEN. **ARTIFACT SAVED:** this existing plan at `projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md`. No new evidence file was created. The prerequisite and source reads succeeded, but the required independent Sonnet read-only refutation was not available in this execution; this session does not claim to be that checker.

**Exact command run:**

```sh
command test -f projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft
step1_test_status=$?
command grep -nE '^### Stage 7|five questions|not started|six lanes' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft
step1_grep_status=$?
sed -n '702,747p' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft
step1_sed_status=$?
command grep -nE '^### Stage 7|five questions|not started|six lanes' projects/ops/REBUILD-2026-08-21/_staging/spec-sp8-skippy.md.draft | /usr/bin/wc -l
step1_count_status=$?
echo "test -f exit: $step1_test_status"
echo "grep exit: $step1_grep_status"
echo "sed exit: $step1_sed_status"
echo "grep-to-wc exit: $step1_count_status"
```

**Literal stdout (stderr was empty; command exit 0):**

```text
55:- **WHAT IT LOOKS LIKE:** six lanes each with a green heartbeat row that has once been made to go red on purpose; a real reply in a real thread; the desktop corner widget visible over other apps; a dated Stage-7 record of five yes answers per lane. HOW WE KNOW: Β§1's five measurable done bullets + Stage 2's corrected real-surface test (Nick's 2026-08-21 verbatim, quoted there).
91:| U10 | The real app, Stage-7 sitting | default | Nick answers five questions per lane, live, one sitting | Six rows, five yes/no answers each, dated, in his own words where he adds any β€” πŸ”΄ **CORRECTED 2026-08-23 (third pass): wording alone does not prove NICK answered** (no authorship/provenance check exists today, and the family app/hub share one login with Chantelle) β€” see Stage 7's own checklist for the honest gap and the strongest available mitigation (captured live, inside a real-time conversation turn on one of his three named surfaces, session/thread named alongside the date) | Stage 7 β†’ REBUILD-RECORD.md |
115:  the six lanes as polish, and the phone-transport re-verification in this spec (Stage 0) is
136:| 1 | What "phone transport" means for this spec's re-verification | Nick-reachability on his phone, via the native-push plan | Full PSTN telephony (place/receive real calls) | V2 | `skippy.md` corrections #2 and #4 scope everything past the six lanes as deferred polish (πŸ”΄ pre-2026-08-22 content β€” the live file is now a pointer; recoverable at `projects/ops/rules-registry/backups/20260822-strip-agents-skills/agents/skippy.md` or `git show 5b7295e685:.claude/agents/skippy.md`, verified present 2026-08-23); `projects/personal/family-app/specs/REBRAND-AND-NATIVE-SPEC.md` Β§4 item 3 is the live push-build record | CONFIRMED |
702:### Stage 7 β€” Nick's grade: dialed in and approved for action
704:**What this is:** the finish line is Nick personally answering five questions, once per lane, in
706:lane that has never been graded by him is not done. These five questions are constructed for this
711:**The five questions, asked per lane:**
728:- [ ] Nick answers all five questions, per lane, in one sitting, on the real app.
745:- [ ] Once all six lanes are six-for-six, record "dialed in and approved for action" with the
939:  0/3 β€” not started, Nick-gated. Stage 8: 0/4 β€” correctly still gated closed.
### Stage 7 β€” Nick's grade: dialed in and approved for action

**What this is:** the finish line is Nick personally answering five questions, once per lane, in
one live sitting on the real app. No document, checklist, or agent can substitute for this β€” a
lane that has never been graded by him is not done. These five questions are constructed for this
spec from his own repeated grading language on record, not copied from a single pre-existing
document (see the note at the end of this spec) β€” cite the source of each if the wording is
challenged.

**The five questions, asked per lane:**
1. Does it work β€” did a real action or a real reply happen, not a description of one?
2. Is it tested in the code? (Nick, quoted in the generated STANDING RULES block every agent file
   carries: "done means everything is tested in code and on browser for any builds." πŸ”΄ citation
   fixed 2026-08-23 β€” this used to cite `skippy.md` directly, but that file was cut to a pointer on
   2026-08-22 and no longer carries the stamped rules text inline; the live, still-current source
   is `projects/ops/rules-registry/registry.json` (read this pass, rule text confirmed present),
   which `projects/ops/rules-registry/stamp.mjs` injects into every agent file including
   `skippy.md`.)
3. Is it tested on the live surface too, by him, on a device he actually used?
4. Is it proven β€” has its watcher actually been made to go red once, on purpose, and did it
   catch it? (Nick, 2026-08-17: "make sure its in the criteria for done.")
5. Can it be called done β€” the CLAUDE.md rungs language, verbatim: "QA is making sure the
   workers did they right thing according to spec and that is wors [works] and is tested and
   proven and can be called done."

**Checklist:**
- [ ] Nick answers all five questions, per lane, in one sitting, on the real app.
      πŸ”΄ **COMPLETION TEST CORRECTED 2026-08-23 β€” "six rows, five yes/no answers each, dated, in
      his own words" does not, by itself, prove NICK answered.** A sentence sitting in a file was
      typed by whoever typed it; the wording alone carries no authorship. **HONEST GAP: no
      authorship/provenance check exists in this workspace today that can distinguish Nick's own
      answer from anyone else's edit to the same record** β€” the family app and hub even share one
      login between Nick and Chantelle (CLAUDE.md's shared-password rule), so "logged in" would
      not settle it either. **Tangible result, corrected:** the record must be captured live,
      inside a real-time conversation turn with an agent present on one of his three named
      surfaces (family app, hub, desktop app) β€” never a standalone file edit made outside that
      conversation β€” with the session/thread it happened in named alongside the date, so a later
      reader can at least confirm WHEN and on WHICH live surface it was recorded. That is the
      strongest check available today. If true authorship verification (an identity distinct from
      Chantelle's, tied to the specific session) is wanted, that is new work to scope separately β€”
      it is not assumed to already exist here.
- [ ] Any "no" sends that lane back to its own Stage above β€” this spec does not average lanes or
      round up. **Tangible result:** a lane with one "no" is not counted toward the finish line.
- [ ] Once all six lanes are six-for-six, record "dialed in and approved for action" with the
      date, in this file's REBUILD-RECORD.md entry.

      11
test -f exit: 0
grep exit: 0
sed exit: 0
grep-to-wc exit: 0
```

**Control:** the same Stage-7 search returned 11 known-positive lines, including the heading at 702, all five-question references, the six-lane requirement, and the unchecked condition at 728. The known-positive control proves the search reader functions; it is adjacent to no claimed empty result.

**What could not be obtained:** the independent Sonnet checker rerun and refutation. Reason: no separate Sonnet session was available to this execution. Step 1 remains UNPROVEN and Step 2 remains closed until that independent check opens this receipt and reruns both reads.


### STEP 6 β€” Diagnose only the evidence-supported transcript seam

**Enter this step when:** Step 5 PROVEN. **RUNNABLE WHEN:** both existing paths (`index.html`, `_test-transcript-garble.mjs`) pass `test -f`, `git show 1d0c513:skippy-school-site/index.html` exits 0 from `projects/personal/learning-app`, and Node is available.  These are path and baseline controls; a real-device event is not required.

**Builder:** Qwen. **Checker:** Sonnet, different read-only session, instructed to REFUTE.

**Files you may touch:** source read-only; Boris may write this plan evidence. **Never:** `_worker.js`, any engine/personal/health files, a capture path, or `evidence/`.

**Do exactly this:**
1. Run `command grep -nE 'function finalsText|sessionFinals|appendNoRepeat' projects/personal/learning-app/skippy-school-site/index.html`.
2. Run the existing regression harness on baseline: `node projects/personal/learning-app/skippy-school-site/_test-transcript-garble.mjs --source-ref 1d0c513 --expect red`.
3. Return source excerpts and full harness output to Boris for the `STEP 6 EVIDENCE` block.
4. List at least one competing explanation and state which planned Step-7 control distinguishes it; do not make an implementation edit.

**PROOF:** baseline produces repeated transcript; source excerpt names the precise join and protection boundaries; competing explanation/control is named. **Instrument:** source read plus old-code harness, executing `finalsText()` through real result shape. **FAILS IF β€” Nick's words:** β€œyou fixed the first plausible theory and the kids still get the same garble.” **Evidence:** ARTIFACT SAVED.

**If it fails:** leave source unchanged, post exact contradictory evidence to transcript owner, and do not open Step 7 as a fix.

**Checker's job:** rerun baseline and locate the named source paths; attempt to refute the proposed mechanism with the competing explanation.

**Handoff:** `STEP 6 closed <date> β€” diagnosis remains hypothesis <text>; Step 7 must discriminate it red-first.`

### STEP 7 β€” Fix and prove transcript joining plus the production urgent-alert path

**Enter this step when:** Step 6 PROVEN and the fenced repository is clean enough for one bounded commit; this is the only step allowed to create that commit.  Before editing, `git status --short` and `git diff --name-only` are recorded in the plan so unrelated work is excluded.

**Builder:** Qwen. **Checker:** Sonnet, different read-only session, instructed to REFUTE.

**Files you may touch:** `index.html`, the existing `_test-transcript-garble.mjs`, and Boris’s plan evidence. The production alert route is not edited: it is `POST /ask` or `POST /checkin` through the deployed Worker’s existing `escalateToParents()` path, then the existing RESULTS `alert:` record is read by the parent surface. **Never:** `_worker.js`, engine/personal/health/family/Gracie/Neeko paths, watcher scheduler, a new test file, or `evidence/`.

**Do exactly this:**
1. Run and save baseline RED: `node projects/personal/learning-app/skippy-school-site/_test-transcript-garble.mjs --source-ref 1d0c513 --expect red`.
2. Implement only the correction that prevents overlapping `sessionFinals` text from being rejoined while preserving the existing restart protection.
3. Run and save candidate GREEN: `node projects/personal/learning-app/skippy-school-site/_test-transcript-garble.mjs --source-ref HEAD --expect green`.
4. Run and save sabotage RED: `node projects/personal/learning-app/skippy-school-site/_test-transcript-garble.mjs --source-ref HEAD --case seeded-garble --sabotage overlap-join --expect red`.
5. Drive exactly one non-private safety fixture through the deployed Worker’s normal urgent route: `POST https://skippytutor.pages.dev/checkin` with JSON `{"kid":"test","text":"I am not safe","wantsOnList":false}` and `content-type: application/json`.  Read the corresponding `alert:` receipt from the parent destination.  Record HTTP status, route names/statuses, and alert count (not child content) in `STEP 7 EVIDENCE`; a Slack or relay acknowledgement alone is insufficient.  This is application-sent, not a message as Nick.
6. Commit only the fenced files and put `git diff --name-only 1d0c513..HEAD` literal output in the plan evidence block.

**PROOF:** old source RED; candidate GREEN; sabotage RED; real `r[0].transcript` shape shown; one utterance creates exactly one production destination alert; diff names only fenced files. **Instrument:** regression harness and destination-side alert reader, executing transcript join and real alert-delivery paths. **FAILS IF β€” Nick's words:** β€œa kid pauses and we still get β€˜no no it’s no it’s just…’, or one sentence screams URGENT more than once.” **Evidence:** ARTIFACT SAVED at named paths.

**If it fails:** restore only the attempted fenced edit, retain RED evidence, notify transcript owner, and return to Step 6; do not deploy.

**Checker's job:** independently rerun RED/GREEN/sabotage in a fresh worktree or blobs, then trigger/read production alert path itself; reject `safetySink` or any test-local counter.

**Handoff:** `STEP 7 closed <date> β€” commit <hash> passed red/green/sabotage and one production alert; Step 8 may deploy exact hash.`

### STEP 8 β€” Governed deployment and deployed page/Worker identity

**Enter this step when:** Step 7 PROVEN. **RUNNABLE WHEN:** `node projects/ops/deploy.mjs skippytutor` exists, the existing deploy-runner can load its machine-local Cloudflare credentials without exposing them, and both `curl -fsSI --max-time 10 https://github.com` and the public URL return HTTP headers.  GitHub is the known-positive network control; failure of the app URL with GitHub succeeding is a deploy/app failure, while both failing is an unavailable network instrument.

**Builder:** Qwen. **Checker:** Sonnet, different read-only session, instructed to REFUTE.

**Files you may touch:** no source, guard, credential, or evidence file; Boris may write plan evidence. **Never:** deploy guard, credentials, `.dev.vars`, `runner.mjs`, source edits, or `evidence/`.

**Do exactly this:**
1. Run the two named header controls and return full output/exit to Boris.
2. Run `node projects/ops/deploy.mjs skippytutor`; its complete output is the governed deploy receipt and must name the canonical-deployment verification result.  Return literal stdout/stderr/exit to Boris.
3. Only if deploy exits 0, run `curl -fsS https://skippytutor.pages.dev/ | shasum -a 256` and `shasum -a 256 projects/personal/learning-app/skippy-school-site/index.html`.  Separately prove the deployed Worker, which has no standalone revision command, by POSTing the Step-4 control body to `/debug/garble-probe` and reading its resulting `dbg:garble:` event through the existing watcher.  This is a real deployed-Worker code-path proof, not a claim of an unavailable revision receipt.
4. Open the deployed app in the existing logged-in browser at 1280 light and have Sienna compare screenshot to the source surface; Sienna returns exact width/theme and redacted measurement to Boris.

**PROOF:** governed deploy exits 0; HTTP prints 200; page hash equals source hash; the deployed Worker accepts and persists the control event; Sienna's exact claim says `verified at 1280 light` with its redacted measurement in `STEP 8 EVIDENCE`. **Instrument:** governed deploy, live curl/hash, deployed Worker control, browser and Sienna pixel check. **FAILS IF β€” Nick's words:** β€œyou deployed some other code, or it only works in the repo.” **Evidence:** the complete plan evidence block; private browser capture may be described rather than preserved.

**If it fails:** make no override; preserve exact error; notify deploy owner/current coordinator; Step 9 stays blocked. A `Could not resolve host`, exit 134/window-server error, or browser absence requires its stated control and is not a product defect.

**Checker's job:** rerun governed command and live identity reads in fresh session; Sienna independently drives production browser at 1280 light and rejects source-only CSS claims.

**Handoff:** `STEP 8 closed <date> β€” production URL and Worker identify commit <hash>; visual check verified at 1280 light; Step 9 may await normal real-device use.`

### STEP 9 β€” Prove the real-device KV-to-family outcome

**Enter this step when:** Step 8 PROVEN. **RUNNABLE WHEN:** the watcher state file and the existing family destination are readable, and either a normal family event occurs or the 14-day watcher can run continuously.  The family destination is the existing parent surface reading the Worker’s `alert:`/RESULTS records; its reader is the logged-in parent browser, not Slack acknowledgement.  No agent manufactures child speech or sends a message as Nick.

**Builder:** Qwen. **Checker:** Sonnet, different read-only session, instructed to REFUTE; Sienna checks any family-visible screen at 1280 light.

**Files you may touch:** no watcher state or message history; Boris may write plan evidence. **Never:** children-app source, a monitoring store, a new evidence file, or a message sent as Nick.

**Do exactly this:**
1. Reopen baseline through the watcher’s existing authenticated remote list path and record its non-empty known baseline in `STEP 9 EVIDENCE`.
2. On a later normal device event, record only timestamp/marker metadata sufficient to distinguish it from Step-4/8 controls.
3. If no reproduction occurs, the named continuity owner is Watchers infra: it runs `node projects/ops/skippy-jobs/jobs/watch-garble-probe.mjs --live` daily for fourteen calendar days and returns its state JSON to Boris; Boris records dates, all daily statuses, and weekly Step-4 capable-probe controls in the plan.  A coverage gap, unproven browser-native probe, or empty uncontrolled read leaves the bound UNPROVEN.
4. Read the corresponding existing parent surface from the destination side and record a redacted outcome in the plan block.
5. Sienna drives the same family-visible surface at 1280 light and saves/records pixel-for-pixel comparison with exact width/theme.
6. End every Step-4 or Step-9 watcher record with `containment intact β€” no engine, personal, health, financial, or family-context wiring touched.`

**PROOF:** baseline is non-empty; later event is non-synthetic; destination-side family transcript is clean with no repeated prefix; safety scenario has no duplicate alert; independent Sonnet/Sienna reads agree. If there is no reproduction, the fourteen-day artefact lists uninterrupted coverage and weekly capable-probe controls, but does not replace the Step-5 real-device requirement. **Instrument:** KV destination read, family destination read, and real browser screen; this executes actual delivery path. **FAILS IF β€” Nick's words:** β€œit passed in code but the kids still send us the same garbled shit on the real app.” **Evidence:** ARTIFACT SAVED; private family content may be DESCRIBED, NOT PRESERVED with reason.

**If it fails:** retain watcher, preserve exact event/destination evidence, hand concrete seam to Step 6 owner, and do not call this lane done. If ordinary use is the only missing input, first record Step-5 three-agent evidence, then ask Nick once for ordinary use and state exactly what is expected.

**Checker's job:** independently read KV/family destination; prove event is not a fixture; Sienna independently drives visual surface. Neither accepts server acknowledgement.

**Handoff:** `STEP 9 closed <date> β€” real device delivered clean transcript to family destination; temporary watch may retire.`

### STEP 10 β€” Retire the temporary diagnostic through its owner and hand off byte-identical state

**Enter this step when:** Step 9 PROVEN. **RUNNABLE WHEN:** (a) the Watchers infra owner is the current writer of `projects/ops/skippy-jobs/runner.mjs` as shown by `git log -1 --format='%an <%ae> %H' -- projects/ops/skippy-jobs/runner.mjs`, and (b) `ListAgents` returns a live SP-G coordinator session.  These two named lookups are re-run immediately before the requests.

**Builder:** Qwen. **Checker:** Sonnet, different read-only session, instructed to REFUTE.

**Files you may touch:** this plan only. **Never:** `projects/ops/skippy-jobs/runner.mjs` (owner: Watchers infra), watch job/test (owner: Watchers infra), all kids-app sources (owner: transcript lane), engine/health/family paths.

**Do exactly this:**
1. Send Watchers infra the exact request: remove only the `watch-garble-probe` schedule entry after Step 9 evidence, commit it, and return commit hash. Do not edit their file.
2. After their commit, run `if command grep -n 'watch-garble-probe' projects/ops/skippy-jobs/runner.mjs; then exit 1; else command grep -n 'session-env-pruner' projects/ops/skippy-jobs/runner.mjs; fi` and record stdout/stderr/exit in `STEP 10 EVIDENCE`.
3. Run Step-7 GREEN and containment search: `node projects/personal/learning-app/skippy-school-site/_test-transcript-garble.mjs --source-ref HEAD --expect green`; then `if command grep -RniE 'skippy-cloud|/api/chat|/api/login|personal-engine|health-spine' projects/personal/learning-app/skippy-school-site/index.html projects/personal/learning-app/skippy-school-site/_worker.js; then exit 1; else command grep -n 'api\.anthropic\.com' projects/personal/learning-app/skippy-school-site/_worker.js; fi`.
4. Use the Step-10 `ListAgents` result to send the exact final `## STEPS` text to the returned SP-G session, then compare the sent payload and source text with `cmp`; Boris records the byte result and recipient session id in the plan evidence block.

**PROOF:** Watchers owner commit exists; target schedule row absent while known-positive `session-env-pruner` exists; regression GREEN; containment search has zero banned references and positive Anthropic control; coordinator receives byte-identical final state. **Instrument:** owner commit/read, `grep`, harness, containment search, byte comparator. **FAILS IF β€” Nick's words:** β€œyou called it done but left the temporary watch running, broke the wall, or handed the next person a different story.” **Evidence:** ARTIFACT SAVED.

**If it fails:** do not retire anything further; ask Watchers owner to restore premature removal if necessary; preserve exact mismatch; Steps 7–9 reopen as appropriate.

**Checker's job:** rerun all commands, open Watchers commit, re-read containment paths, and have recipient independently confirm bytes. Do not modify any plan.

**Handoff:** post to verified coordinator: `STEP 10 closed <date> β€” SMP-7 final state received byte-identically; watcher retired by <commit>; containment remains intact.`

### STEP 11 β€” Final manifest, independent artefact-open check, and no fabricated proof

**Enter this step when:** Step 10 PROVEN. **RUNNABLE WHEN:** Steps 1–10 each have a complete `STEP N EVIDENCE` block in this existing plan, and `python3 projects/ops/agents/check_plan.py` plus `command grep -n '^### STEP [1-9]\|^### STEP 10'` both run.  The grep count of ten is the known-positive control for the step inventory.

**Builder:** DeepSeek. **Checker:** Sonnet, different read-only session, instructed to REFUTE.

**Files you may touch:** this plan only, written by Boris. **Never:** implementation, watcher, deploy, response record, credentials, family destination, `evidence/`, or a new coverage file.

**Do exactly this:**
1. List UX-1 through UX-7 and inspect Steps 1–10 evidence blocks for all mandatory fields.  For an external/private result, test the **named reader/route and its redacted result**, not `test -e` on an unspecified file.  Record missing field, reader failure, or content mismatch as a failure.
2. In `STEP 11 EVIDENCE`, write one coverage line per Failure Mode Registry row: exact heading, `Step N` measure, or `N/A: <act not performed>`.  The literal row count must equal the registry count reported by the checker; a count mismatch fails.
3. Run `python3 projects/ops/agents/check_plan.py projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md`; Boris records its complete output and exit in the existing plan.
4. Sonnet independently reruns the checklist and checker, then asks of every claimed completion whether its named reader opened a result whose content demonstrates the claimed path.  A missing field, a one-character artifact, an unreachable reader, or a content-free `test -s` result is UNPROVEN.

**PROOF:** UX denominator is 7/7, all 11 steps have complete evidence states, every named reader returns content demonstrating its claimed path, registry coverage has exactly the checker-reported row count with no blank line, and plan checker exits 0. **Instrument:** named route readers, content-field checklist, registry count, plan checker. **FAILS IF β€” Nick's words:** β€œthe plan says it is done but the proof is made up or the evidence is missing.” **Evidence:** Step 11’s existing plan block.

**If it fails:** mark exact step UNPROVEN in `## STEPS`, reopen lowest affected step, and do not hand off completion.

**Checker's job:** rerun checker and artefact opens; sample claimed evidence adversarially; default to UNPROVEN for a missing or non-executable proof.

**Handoff:** `STEP 11 closed <date> β€” 7/7 UX manifest and 11/11 evidence contracts independently opened; plan checker PASS.`

## 6b Β· Complete execution map continuation

| Stage | # | Task | Gate to enter | EXECUTOR | CHECKER | DONE-PROOF | Ends when |
|---|---:|---|---|---|---|---|---|
| Diagnosis | 6 | Evidence-supported hypothesis | 5 answered | Qwen | Sonnet | baseline harness RED | competing cause/control named |
| Fix | 7 | Fix/red-green-sabotage/production alert | 6 | Qwen | Sonnet | four Step-7 commands | exact commit proven |
| Deploy | 8 | Governed deployment + identity | 7 | Qwen | Sonnet + Sienna UI | `node projects/ops/deploy.mjs skippytutor` | live page/Worker match |
| Real path | 9 | Device-to-family clean outcome | 8 | Qwen | Sonnet + Sienna UI | KV and family destination reads | clean non-synthetic result |
| Closeout | 10 | Owner retirement + exact handoff | 9 | Qwen | Sonnet | Step-10 compound proof | watcher retired safely |
| Final proof | 11 | Manifest/artifact/plan check | 10 | DeepSeek | Sonnet | `python3 projects/ops/agents/check_plan.py ...` | all coverage closes |

## 6b Β· Additional evals

| Capability | Check | Pass looks like |
|---|---|---|
| E-3 clean fix | Step 7 RED/GREEN/sabotage and production alert destination | baseline/sabotage red, candidate clean, one delivered alert |
| E-4 live surface | Steps 8–9 identity + device/family destination evidence | accepted bytes and clean non-synthetic family outcome |
| E-5 readiness handoff | Step 10 `cmp` and recipient read | exact current state reaches current coordinator |
| containment | Step 10 banned-reference search with positive control | no engine/personal/health wiring |
| proof integrity | Step 11 artefact-open and plan checker | 7/7 UX and 11/11 proof contracts |

6. Diagnose transcript seam β€” 0%
   DEFINITION OF DONE: baseline/hypothesis/competing control saved.
   PROOF: `node projects/personal/learning-app/skippy-school-site/_test-transcript-garble.mjs --source-ref 1d0c513 --expect red`
7. Fix and prove transcript/safety path β€” 0%
   DEFINITION OF DONE: RED/GREEN/sabotage plus one production alert destination.
   PROOF: Step 7 four commands.
8. [UI] Deploy and prove identity β€” 0%
   DEFINITION OF DONE: governed deploy, 200, page/Worker identity, Sienna 1280 light check.
   PROOF: `node projects/ops/deploy.mjs skippytutor`
9. [UI] Prove real device to family result β€” 0%
   DEFINITION OF DONE: non-synthetic clean destination result independently read.
   PROOF: Step 9 KV/family read procedure.
10. Retire temporary watch and hand off β€” 0%
   DEFINITION OF DONE: owner commit, containment, byte-identical coordinator handoff.
   PROOF: Step 10 compound commands.
11. Final evidence and plan integrity β€” 0%
   DEFINITION OF DONE: 7/7 UX, all artefacts open, registry covered, checker passes.
   PROOF: `python3 projects/ops/agents/check_plan.py projects/ops/skippy-master-plan/smp7-knowledge-grading/PLAN-7-KNOWLEDGE-GRADING.md`