Copilot — Milestone 1b: live granularity, card lifetime and the adaptive call plan for CurrencyTransfer replays
1. Overview
Problem statement. Stevan reviewed a replay of call 3339895706 on 2026-09-27 (
docs/design/FEEDBACK-2026-09-27.md). His session1d160d96matches run49e3e48fdecision for decision, on policy v4. The review found four problems:- Nothing feels live. The transcript and Jev both run once per stitched turn:
LiveTranscript.updatedraws a line at itst_end, andUtteranceDriver.ticksendsutterance{i}at that moment. This call's turns run to 44.5 s and 112 words. The word-weighted wait from spoken to visible is 9.1 s, and the longest gap between decision points is 45.6 s. Jev itself answers in 304 ms at p50. - The plan reads as a score sheet. It has 10 rows, 6 of them never ticked, and a live tally of 2 of 7.
- Facts the copilot should simply know nag as steps ("Check we can serve them").
- The purpose was never captured: "I get paid on the fifteenth" misses the trigger.
- One tax-residency worry became two jurisdiction rows. A 3-word fragment closed it, and a re-raise 1.7 s later inserted a second row.
- A complaint about the client's own bank at 02:22 opened "Client thinks our rate is too high".
- Suggestions are rare and vanish unread. Only 12 of 96 decisions showed an open point, on 3 distinct cards. The amount point was tinted for 1.0 s, and the first-transfer-date point was withdrawn uncovered after 23 s.
- The player jumps. "catching up (N queued)" moves the scrubber.
The engineering analysis is
docs/architecture/proposal-m1b-live-plan.md: measures in §2, rules in §3-§6, the story sketch in §7, risks in §8, disagreements in §9 and review changes in §10. Stevan's rulings are inpolicy/approvals/2026-09-27-m1b-decisions.json.- Nothing feels live. The transcript and Jev both run once per stitched turn:
Proposed solution. Four strands inside design G's one-card live view, plus verification:
- (a) Live granularity in replay.
- Words appear as they are spoken, in the browser only.
STITCH_V2cuts long turns into sentence-sized parts (decision 1a).- Feedback moves at sentence ends; the card switches only at the end of the client's run (decision 1b).
- The thresholds counted in decisions are re-tuned for the new cadence.
- The per-call cost gate becomes $0.06 (decision 16).
- (b) Card lifetime.
- In the presenter: a 12 s floor, stacking up to 3, a sticky last suggestion and a rep-turn grace.
- In the engine: hysteresis in the [0.3, 0.6) band, and pre-judged fits for the item about to open (decisions 3 and 17).
- (c) The plan as a guide.
- Every item gets a kind: fact, must-say or judgement.
- Pre-call facts are seeded at evidence i=-1 from a closed-vocabulary
precall_json(decision 2a). - The live list becomes known chips, Now, Coming up and Covered, with no counts (decision 2b).
- A coaching review is shown to the rep (decision 2g).
- The must-say rules follow decisions 2c, 2d and 2e.
- Stevan's approved wording is applied (decision 2f).
- The purpose trigger gains new words, and a salary-home prior ranks situations (decision 2h).
- (d) The duplicate tick and the rate false positive.
- A concern clears only on a client line of 4 or more words, and a re-raise within 60 s reopens the same row.
- The objection criteria stop typing a complaint about the client's own bank as a CurrencyTransfer rate concern.
This needs no new Cloudflare service and one D1 migration (
0004_precall.sql).- (a) Live granularity in replay.
Target users. Stevan (founder; the only Access identity) reviewing replays now. CurrencyTransfer reps later, each reading their own review (decision 2g).
Success metrics (M1b exit).
scripts/m1b-acceptance-check.mjs(COPILOT-118b) asserts these, and COPILOT-118 runs it on the deployed Worker. The baseline is the 2026-09-27 run49e3e48fandeval/reports/final-20260927*.json.- Cadence on 3339895706 (on estimated word times):
- Decision points go from 96 to 125-145.
- The longest gap between decision points is ≤ 15 s (was 45.6 s), and the p90 gap is ≤ 11 s (was 16.0 s).
- Transcript text never appears before its estimated spoken time.
- Sentence-end-to-first-feedback p50 and p95 at 1x are reported, measured on wall-clock pacing.
- Suggestions on 3339895706:
- More than 12 decisions show an uncovered point on the current item (today 12).
- More than 3 distinct items ever show a point (today 3); 20 and 5 are expected.
- No revealed point leaves before 12 s, except for the exemptions defined in COPILOT-107.
- At most 3 points show at once. The amount point never renders tinted.
- Plan on 3339895706:
- At the last decision the live list has ≤ 4 open rows (10 today), no tally and no "Not needed" group.
- At most one jurisdiction row (two today).
- No rate or fees concern opens before 04:00.
- With the demo seed, "What they're sending, and why" and "Where they live and bank" are satisfied at i=0.
- The review reads "Must-says: 2 of 3 that applied".
- Gates on every promoted policy candidate:
- L1 ≥ 90 % per scenario (baseline v4: onboarding 91/91, cs 79/81).
- L4 and L6 pass. Every
input_tokensis ≤ 12,000 (decision 15), and cost per call is ≤L6_LIMITS.cost_per_call_usd= 0.06 (decision 16). - e2e p95 at 1x is ≤ 3 s.
- Curation survives the re-import:
- 3 marks on 3339895706 and 1 label on 3303259297, before and after.
- 3485591407 keeps
pd_ct_id109829.
- Cadence on 3339895706 (on estimated word times):
2. Requirements
Functional Requirements
Core
- FR-1 Screenshot tooling for the deployed dashboard (COPILOT-102a). Transcript words are revealed between each line's start and end (COPILOT-102).
- FR-2 Re-stitching and curation:
STITCH_V2parts (103);- the curation export (104a);
- the
iremap (104f); - the re-import CLI with an atomic curation precondition (104b);
- the validated curation-restore route (104c);
- Worker ingest on v2 (104d);
- the guarded re-import of the four samples, with the curation restored (104e).
- FR-3 A presenter simulator (105a). The presenter's change moment becomes the end of the client's run, with new points at sentence ends;
LIVE_CARD_AT_SENTENCE_END = false(105). - FR-4 Decision-counted thresholds are re-tuned, the window becomes 16, the cost gate becomes $0.06, and a policy candidate carries
stitch_version2 (106). - FR-5 Card lifetime:
- a design board for the live-view changes (121);
- presenter point lifetime R1-R4 (107);
- engine hysteresis R5 (108a) and pre-judged fits R6 (108b);
- the player's reserved status slot (120).
- FR-6 The plan as a guide:
- Grammar and policy data: plan grammar, text identity and registry (110a); compatible engine readers and contract types (110f); generator and eval tooling (110g); the opt-in derived checklist (110h); the approved wording (110c); typed templates (110d); the
bank_countryslot with the served list (110b); an L1 fact-value check (110e). - Engine: fact and judgement semantics (111a); must-say
applies/due(111b); amount corrections against second legs (111c). - Pre-call facts: schema and validation (112); write paths (112b); pure seeding and conflict provenance (113); the session propagation (113c); the served-list jurisdiction gate (113b).
- Views: the guide pane (114); chips with their source (114b); the review model (115) and its view (115b).
- Capture: purpose trigger words and the prior (116).
- Grammar and policy data: plan grammar, text identity and registry (110a); compatible engine readers and contract types (110f); generator and eval tooling (110g); the opt-in derived checklist (110h); the approved wording (110c); typed templates (110d); the
- FR-7 The concern clearing guard and reopen (117). The objection criteria exclude complaints about the client's own provider (119).
- FR-8 Verification: wall-clock replay pacing and arrival logs (118a); measurement and acceptance tooling (118b); final verification, register and runbook (118).
Secondary
- FR-9 The views render every new label only from an approved text in the pinned bundle; an older bundle keeps today's words (INV-COPILOT-015).
- FR-10 The served-country list ships empty. The jurisdiction gate and the "outside the served list" surfacing do nothing until a list is supplied (§7 Q3).
Out of scope
- COPILOT-109: decision 3 is option A. Its id stays reserved.
- Live audio, nova-3 streaming, the Aircall transport spike and the paid Workers plan (M3, decision 14). True partial decisions.
- The real sign-up join through
pd_ct_id(M2; M1b fixes the shape it must produce). Multi-leg amount and pair slots. A tax-residency concern type (M2). A Saudi overlay template (decision 2h). - The following, each excluded by the decision named:
- raising the 12,000 cap (decision 15);
- live-list variants B and C (decision 2b);
- a team leaderboard (decision 2g);
- a per-leg minimum and any batching line (decision 2e).
- Also out of scope:
- a dismiss or defer control on suggestions;
- a
jointvalue for the funding-account list (§7 Q5); - nova-3 word times for uploads (§7 Q4).
Non-Functional Requirements
- Performance.
- e2e p95 (utterance sent to decision received) ≤ 3 s at 1x, the M1 gate; the 1.5 s target is reported.
- Provider p95 ≤ 1.5 s. At most 2 subrequests per decision (COPILOT-058b).
- A 125-165-decision replay at 20x completes through
reconnect{after_i}with 0 errors. - The word reveal rewrites only the rows still being spoken.
- Cost and tokens.
- Cost per replayed strong call ≤
L6_LIMITS.cost_per_call_usd(0.06 from 106). Stories read the constant, never a literal. - Every request ≤ 12,000 input tokens; p95 ≤ 11,000 is reported.
prejudgeis the first drop under budget. - Caps are measured on session requests only, never on a whole-bank probe; M1's COPILOT-023 failed on exactly that.
- Cost per replayed strong call ≤
- Security and privacy.
- Access fronts every route.
PUT /api/calls/:id/precallandPOST /api/calls/:id/curationaccept the service token only, validate every row (including the ownership of deterministic ids), and writeaudit_log.- A forced transcript replacement is refused atomically when the call's curation changed since its export.
precall_jsonholds only closed-list ids, ISO codes, digits and a source, on every write path (INV-COPILOT-004, INV-COPILOT-020).- Exports and fixtures are pseudonymised envelopes.
docs/runbook/is never published.node scripts/check-secrets.mjsexits 0.
- Scalability. D1 batch limits are unchanged (INV-COPILOT-008). The same
step()serves 96 or 165 decision points. The served list and theprecallshape are policy data, ready for the M2 join. - Accessibility.
- Both themes on every UI change. Colour is never the only signal: a rep-turn point arrives grey with its tick.
- The "Covered" line is a keyboard-operable button with
aria-expanded. prefers-reduced-motioncollapses the point fades. The word reveal follows the clock and is not an animation.- The status slot never moves the scrubber.
Decisions (Stevan, 2026-09-27; policy/approvals/2026-09-27-m1b-decisions.json)
| # | Decision | Lands in |
|---|---|---|
| 1a | Cut at a sentence end after ≥ 6 words and ≥ 5 s since the last cut; forced at 12 s | 103 |
| 1b | Feedback (ticks, chips, caution, new points on the open card) moves at sentence ends; the card switches only at the end of the client's run; one flag, defaulting to that | 105 |
| 2a | Pre-call facts from the CT sign-up wizard via pd_ct_id, closed-list values only, chips labelled with their source; the join is M2; the sample call is seeded by hand as "Seeded for the demo" |
112, 112b, 113, 114b |
| 2b | Known chips, Now (one card), Coming up (≤ 3 applying must-says), Covered (collapsed); no counts live | 114 |
| 2c | Who holds the money is a must-say that is due when the client raises safety, and otherwise belongs to the first-transfer walkthrough. It is never listed on a first call where the client does not raise safety | 110d, 111b |
| 2d | Fund from your own account is a must-say only while the payer might be someone else (third party, joint, company or unclear). A salary into the client's own account never triggers it | 110d, 111b |
| 2e | The GBP 5,000 minimum applies per client relationship or booking total, not per currency leg. A small second leg passes once the main amount clears the minimum: the item is not needed, and no batching line is written. Multiple amounts are tracked only so that a mid-call correction is not mistaken for a second leg | 111c |
| 2f | The reworded labels are approved as drafted (policy/proposals/2026-09-27.json, approved_by stevan@currencytransfer.com). Its one [VERIFY] line waits for decision 13 |
110c |
| 2g | The rep sees their own review | 115, 115b |
| 2h | Ranked situations; no Saudi overlay | 116 |
| 3 | Keep the like-for-like gate and add the six lifetime rules (12 s floor, stack to 3, sticky last suggestion 12 s, rep-turn grace, hysteresis per 17, pre-judge the opening item) | 107, 108a, 108b |
| 14 | Live audio is M3 | out of scope |
| 15 | The hard token cap stays 12,000 | 106, 108b, 118 |
| 16 | L6_LIMITS.cost_per_call_usd = 0.06 |
106 |
| 17 | A shown situation survives one answer in [0.3, 0.6) or one dropped question; it leaves on the second, at once under 0.3, or when another candidate qualifies | 108a |
| 12, 13 | Still open. Stories carry [DECISION 12] / [DECISION 13] |
§7 Q1, Q2 |
3. Technical Architecture
flowchart LR
subgraph devbox["devbox / otto worktree (no data/)"]
SMP[("tests/fixtures/samples/*.json<br/>pseudonymised fragments")]
ST2["stitch() STITCH_V2<br/>src/ingest/stitch.ts"]
EXP["export-call-curation.mjs<br/>remap-i.mjs"]
IMP["import-fixture.mjs --stitch 2"]
PRE["precall-payload.mjs"]
POL["policy/src → build-policy --next-version<br/>→ publish-policy (draft) → eval gate → publish"]
EVAL["eval.mjs l1,l4,l6 · replay-ws · ws-probe<br/>presenter-sim · m1b-acceptance-check"]
SMP --> ST2 --> IMP
end
subgraph CF["Cloudflare (Access on every hostname)"]
W["Worker jev-copilot<br/>+ PUT /api/calls/:id/precall<br/>+ POST /api/calls/:id/curation"]
D1[("D1 copilot<br/>+ 0004_precall.sql")]
DO[("DO CallSession<br/>initialState(pin, seed) · step()")]
AI[["typesafe/jev via AI Gateway jev-copilot"]]
end
subgraph WEB["Browser (owns the clock)"]
TR["transcript.ts: word reveal"]
PR["presenter.ts: run end, R1-R4"]
PV["plan.ts / chips.ts: guide<br/>review-plan.ts: coaching"]
end
IMP -->|"POST /api/calls/import (force)"| W
EXP -->|"D1 reads · POST curation"| W
PRE -->|"PUT precall (service token)"| W
POL -->|"POST /api/policy · eval-runs · publish"| W
EVAL --> W
W --> D1
DO --> D1
DO -->|"≤ 2 subrequests per decision"| AI
TR --- PR --- PV
WEB <-->|"WS decision{plan: items, facts[source], known, must_says, dropped_hard}"| DO
Key components. Line numbers are at 51dbb7e (code unchanged since a83e53e). The symbol is the anchor: executors re-locate by symbol if a line moves.
| Area | Existing (symbol, file:line) | M1b change | Stories |
|---|---|---|---|
| Stitch | STITCH_V1 src/ingest/stitch.ts:53; maxWords: 120 :59; DECISION_MIN_WORDS = 4 :70; merge step :160-171; parts loop :186-213 (every part keeps t: turn.t); output order :230-233 |
STITCH_V2; each part's t is the previous part's t_end |
103 |
| Import | DERIVED_TABLES src/routes/calls-import.ts:24 (a forced import deletes decisions, marks, moments, labels, answers, move_answers, rewrites); callRowValues :133 writes pd_ct_id from the payload; contentHash src/ingest/import-payload.ts:168 covers stitch_version; buildFixturePayload scripts/lib/fixture-payload.mjs:24 (no pd_ct_id); momentId src/engine/moments.ts:114 embeds start_i; labelsOf src/session/checkpoint.ts:279 reads gold_json, which embeds i |
curation export, i remap, an atomic curation precondition, a validated restore route |
104a, 104b, 104c, 104e |
| Ingest | stitch(roled) src/routes/ingest.ts:381; stitch(applyRoles(…)) src/routes/speaker-map.ts:308; INGEST src/policy/build.ts:119, bound into policy_hash by hashPolicy (src/policy/canonical.ts) |
uploads stitched under v2 (104d); INGEST.stitch_version 2 in the 106 candidate |
104d, 106 |
| Transcript | LiveTranscript.update web/src/transcript.ts:207; UtteranceDriver.tick web/src/clock.ts:427 |
word reveal from u.t |
102 |
| Presenter | isInsert web/src/presenter.ts:125; nextClientEnd :140; dueFor :153; releaseAt :164; apply :176; LIVE_MIN_DWELL :23 |
run end, sentence-end points, R1-R4, client-due must-says treated as inserted | 105a, 105, 107 |
| Situation gate | fitsToAsk src/engine/situation.ts:152; recordFits :172 (idempotent, called from step() and again from updatePlan); retargetFit :205; shownSituationId :228; planRequest src/engine/budget.ts:97 (fact_values drop :121) |
hysteresis, pre-judge, a prejudge drop step |
108a, 108b |
| Plan | COVERED_DWELL = 4 src/engine/plan.ts:69; coveredToAsk :302; updatePlan :354; def.done_when.length :422; episode insert :443-459; asks :490-498; CallPlan src/engine/types.ts:341; FactValue :160 (value: string); FactChip :176 (at: Evidence, non-null); factChips src/engine/fact-values.ts:381 |
kinds, known, must_says, chip source, seeded evidence at i=-1, reopen, dropped_hard |
106, 108a, 110f, 111a, 111b, 113, 117 |
| Concern | updateConcern src/engine/concern.ts:91: opens on client_objecting ≥ concern_open (:102), clears on any client line (:120) |
≥ 4-word clear, reopen window, served-list gate | 113b, 117 |
| Amounts | clientAmountsGbp src/engine/amount.ts:385; recordAmounts :362; when_amount_below_gbp in src/engine/hero.ts:40 (applies unless some client amount ≥ threshold) |
a correction replaces the corrected amount; a second leg is kept (decision 2e) | 111c |
| State | initialState(p: Pick<Policy,'scenario'|'version'>) src/engine/state.ts:17. Callers: src/session/handlers/load.ts:131, seek.ts:55, src/session/checkpoint.ts:254 (rehydrate, with a PolicyPin), src/engine/recompute.ts:73 (pure; called by src/session/handlers/set_weights.ts:61). buildJevState :63 |
initialState(pin, seed?), with the seed carried in DO meta and sessions.precall_json |
113 |
| Policy | policy/src/plan-onboarding.json: templates (first_call_after_signup ten items; the per-template text ids plan.template.<template>.<item>.*), fact_slots, purpose trigger :176, purpose_max_decisions 8. rules-*.json: checklist :20-104, thresholds :251-277, token_budget.window 12. Bank onboarding.json: onboarding_objection_type :221, fact_lists.countries (35 ids incl. SA). customer-success.json cs_objection_type :116. shared.json client_objecting. policy/proposals/2026-09-27.json: Stevan's approved texts |
kinds, bank_country, served_countries, thresholds, criteria, the approved wording |
106, 108a, 110b, 110c, 110d, 116, 117, 119 |
| Eval | L6_LIMITS scripts/lib/eval-l6.ts:21 (cost_per_call_usd: 0.05); the L4 partition in scripts/lib/eval-l4.ts; engine checks in scripts/lib/eval-l1.ts (checklist_unchanged :174); eval/labelled/{onboarding,cs,exemplar-moments,must-say}.json; scripts/demo-check.mjs:53 reads stitched/<call>.v1.json |
gate 0.06; retired ids in L4; the fact_value check; demo check on v2 |
104b, 104e, 106, 110d, 110e |
| Replay tooling | paceDelays scripts/lib/replay-client.ts:41 (caps each gap at MAX_GAP_S even at 1x); scripts/replay-ws.mjs (--policy-version, --session; its summary prints reconnects and max_input_tokens); scripts/ws-probe.mjs (--resume; flags live in scripts/lib/probe-flags/); scripts/probe-player.mjs; scripts/design/shoot-dashboard.mjs (--deployed, --play) |
--wall-clock and --log; presenter-sim.mjs; --count-points and --known; --at, --until, --expect and --expect-none; --check status-slot |
102a, 105a, 108b, 110f, 118a, 118b, 120 |
| Player | catchingUpLabel web/src/player.ts:71 |
reserved status slot | 120 |
Data flow for one replayed part. Under v2, a 44.5 s client turn arrives as 4-6 kind:'turn' parts, each with its own t and t_end, so the protocol, the DO and the decision table are unchanged.
- The browser clock passes a part's
u.t: its row appears and grows word by word. - At the part's
t_end, the browser sendsutterance{i}. - The DO builds the state: seeded facts at i=-1, a 16-unit window, and the pre-judge questions on client parts.
- The DO answers from the cache or from Jev and runs
step(): kinds, known facts, must-sayapplies/due, hysteresis, and the concern guard and reopen. - It commits in ≤ 2 subrequests and sends
decision. - The presenter applies ticks, chips, the caution and any new point on the open card as the decision arrives. It applies a card switch at the client's run end, after the 20 s dwell, and never during a rep turn.
- R1-R4 hold what is already on screen.
Integration points.
- Jev: unchanged. The DO calls
env.AI.run('typesafe/jev', …, {gateway:{id:'jev-copilot'}}). - D1: migration
0004_precall.sql, applied remotely by 112b. - Policy gate:
POST /api/eval-runs, thenPOST /api/policy/:scenario/publish(architecture §9). - Deploys and live commands:
node scripts/deploy.mjsfor deploys; live commands run throughnode scripts/with-cf-env.mjs, using thejev-scriptsAccess service token.
Conventions. The M1 conventions in prd.md §3 still apply: test commands, with-cf-env for live commands, pseudonymised fixture envelopes, replacement semantics, and null for unknown. M1b adds:
Deploy and probe.
node scripts/deploy.mjsrefuses a dirty tree or a branch behindmain, and deploys with--var GIT_SHA:<HEAD>. Every live script assertsGET /api/healthversion == git rev-parse HEADfirst:replay-ws.mjs,ws-probe.mjs,eval.mjs,probe-player.mjs,presenter-sim.mjs,m1b-acceptance-check.mjs, andshoot-dashboard.mjs --deployed.Commit sequence, per story:
- Commit the code. That commit is the deployed revision.
node scripts/deploy.mjs.- Run the live steps.
- Make at most one derived-artefact commit. It holds only what the live run produced or pinned: screenshots,
eval/reports/*,docs/reports/*,docs/verification/*, and test pins, remap fixtures and label remaps derived from the run. It touches no runtime file undersrc/,web/,scripts/,policy/ormigrations/. - Run
npm run typecheckandnpx vitest runon the final commit; both pass.
This is M1's practice, and it is what makes
health.versionmeaningful.
Headless criteria only. Every criterion must be satisfiable by the executor under the
jev-scriptsservice token. None needs a signed-in human, a Cloudflare dashboard figure, or a reply or sign-off from Stevan; those are in §5.6. In M1, COPILOT-023, 054, 055, 060 and 072 had to be amended mid-run because they did.Token caps on session requests only. The 12,000 cap is asserted on replay and eval
max_input_tokens, and on the worst case (tests/unit/worst-case.ts, vianpx vitest run tests/unit/engine-groups.test.ts tests/unit/policy-build.test.ts). Whole-bank figures andprobe-questions.mjsfigures are information only.Policy candidates. A story that changes
policy/src(the policy chain in HOT FILES) follows these steps:- Regenerate what the change makes stale: the policy fixtures (the
scripts/build-policy.mjsheader commands), the stub answers (node scripts/gen-stub-answers.mjs),policy/REDPEN.md(node scripts/redpen.mjs) andpolicy/LABELS.md(node scripts/labels-sheet.mjs). - Build both scenarios with
node scripts/build-policy.mjs --scenario <s> --next-version --out policy/build/<s>.v<N>.json. Never use a literal version.--next-versionreads the trackedpolicy/build/, so no server is needed. - Commit the code, the sources and the bundles together, then deploy, so the Worker's loader accepts the new shapes before any upload.
- Upload each draft with
node scripts/publish-policy.mjs --scenario <s> --version <N>. It answers 409not_evaluated, which is expected. - Run the story's live probes pinned to the candidate (
--policy-version <N>). - Promote only through the gate: run
node scripts/eval.mjs --scenario both --layers l1,l4,l6 --policy-version <N> --fresh --run <story>-<yyyymmdd>, thenpublish-policy.mjs --scenario <s> --version <N>for both scenarios. The gate is L1 ≥ 90 % per scenario, plus L4 and L6. - List every case that passed in
eval/reports/final-20260927-{onboarding,cs}.l1.jsonand now fails, with its score.
A failed candidate stays a draft, and the story stays
passes:falsewith its failing cases. No threshold is lowered and no label is edited to make a gate pass.- Regenerate what the change makes stale: the policy fixtures (the
New policy fields are optional. Every field M1b adds to
src/policy/types.tsis optional; absent means M1 behaviour. This coverskind,applies/due/satisfied_by/surface_when,texts_from,known_note,view_labels,situation_fit_hard,concern_reopen_window_s,served_countries,priorsand theprejudgedrop step. New code deployed before its candidate is published therefore never breaks the published bundle. Each story tests its field absent on a clone of the committed fixture.Times, not line numbers, in live probes. Line numbers in this PRD are
transcript_rev1 (the v1 stitch); after 104e the four samples are v2, with newi. Unit tests ontests/fixtures/stitched/<call>.v1.jsonkeep the v1 numbers. Live probes select by call time (utterances.t), or throughtests/fixtures/remap/<call>.json(104e), never by a literali.UI evidence from the deployed dashboard. Use
node scripts/design/shoot-dashboard.mjs docs/design/evidence '<shots>' --deployed --at <mm:ss> --expect '<selector>', where each shot is[name, query, scheme, 1440, 900]and--at/--until/--expect/--expect-nonecome from 102a. Take a light shot and a dark shot; the review view addsview=reviewto the query. Before shooting, run a cache-warmingreplay-wswithout--eval-uid, so the dashboard's session hits the same answers. A state that depends on Jev is shot with--until '<selector>', which pauses the page's own session the first time that state renders. A candidate story shoots after it publishes, because the page's session uses the published policy.Fixtures. The new kinds
curation,decisionsandremapjoinFIXTURE_KINDS(scripts/lib/fixtures.ts) and the schema-driventests/unit/fixtures-clean.test.ts. They hold pseudonymised text only.tests/fixtures/exports/keeps its call-coach exports. Test-only inputs that are not fixtures (a sample approvals reply, synthetic transcripts) are built inside the test and written to a temporary directory.Status notes. Block reasons and status notes use plain words, with no apostrophes or shell metacharacters. In M1 an apostrophe crashed otto (fixed in otto
ab20ebb); nothing relies on that fix.Rep-facing words.
- A policy text (item titles, hints and notes,
plan.*labels, chips, situation texts) renders only when the pinned bundle holds it approved (INV-COPILOT-003, INV-COPILOT-015). - UI chrome is the fixed strings in
web/srcnamed bydocs/design/DESIGN.md: pane and section names ("Now", "Coming up", "Covered", "Plan for this call" today), review-panel labels, and the rule reasons listed in DESIGN.md §Review view. Chrome is never built from call data, so it needs no red pen, as today.
- A policy text (item titles, hints and notes,
HOT FILES. Stories that edit the same file are ordered by dependencies. Every edit is additive: append fields, rules and registrations; never reorder or rename. The run is sequential: otto implement, one story per iteration, as in M1.
| File or group | Stories, in order | Rule |
|---|---|---|
Policy chain: policy/src/*.json, tests/fixtures/policy/*, tests/fixtures/answers/*, policy/REDPEN.md, policy/LABELS.md |
110c → 110d → 117 → 119 → 106 → 108a → 108b → 110b → 116 | one candidate per story; rebase, regenerate, then --next-version |
src/policy/{types,build,loader,text,canonical}.ts, src/policy/fact-slots.ts |
110a (grammar, text identity and registry) → 110h (derived checklist, hash ownership) → 110c (the apply tool) → 110d → 117 and 108a (thresholds) → 110b (slot) → 116 (priors) | additive types; the builder validates the new shapes |
scripts/redpen.mjs, scripts/labels-sheet.mjs |
110g → 110c | 110g resolves texts_from and the derived checklist through the builder's helpers |
web/src/review-plan.ts |
110a (mustSayOf reader) → 115b |
|
web/src/presenter.ts |
105 → 107 | 105 owns the change moment; 107 owns the lifetime state and the client-due insert clause |
web/src/plan.ts, web/src/shell.ts |
105 → 107 → 117 → 114 | 105 adds data-midrun; 107 the last-suggestion line; 117 the reopened state word; 114 restructures the pane |
web/src/chips.ts, web/src/render.ts |
110b (data-slot, the country template) → 114b (data-source, no link when at.i < 0, suffix) |
|
src/engine/plan.ts, src/engine/types.ts |
110a (readers: done_when, applies) → 110f (types) → 117 → 111a → 111b → 113 (templateFor, the seed and its conflict); 108a adds only dropped_hard, and 108b only the pre-judge wiring |
typed fields appended |
src/engine/concern.ts |
117 → 113b | |
src/engine/amount.ts, and src/engine/fact-values.ts for the amount reconcile |
111c | |
src/session/protocol.ts |
110f | re-exports the additive CallPlan fields only |
src/routes/calls-import.ts |
104b (curation precondition, pd_ct_id) → 112b (precall provenance and preservation) |
|
src/index.ts |
104c (curation route), 112b (precall route) | one dispatch line each |
scripts/lib/probe-flags/index.ts |
110f (--known), 108b (--count-points) |
one handler file and one registration line each |
scripts/design/shoot-dashboard.mjs |
102a adds --at, --until, --expect and --expect-none; later stories only call it |
|
scripts/lib/replay-client.ts, scripts/replay-ws.mjs |
118a | summary fields and flags appended |
scripts/presenter-sim.mjs |
105a → 107 → 118b | output fields appended |
docs/design/DESIGN.md |
102 §Transcript pane; 105 → 107 → 111b §Cadence rules; 113 → 114b §Client facts chips; 117 §Plan item rows → 114 §Plan pane; 115b §Review view; 120 §Player and §States; 121 §Variants | disjoint sections |
eval/labelled/*.json |
104e (must-say.json i remap), 110e → 116 (onboarding.json), 119 (exemplar-moments.json) |
append cases; labels are ground truth |
scripts/lib/eval-l4.ts, scripts/lib/eval-l6.ts, scripts/lib/eval-l1.ts |
118a (wall-clock 1x sample) → 110g (L4 retired) → 106 (L6 constant) → 110e (fact_value) |
118a precedes the first gate (110c) |
docs/reports/cost-latency.md |
104e, 106, 108b, 118 | append a section each |
migrations/ |
112b applies 0004_precall.sql, which 112 adds (0001-0003 exist). prd-m2-corpus.md also plans a 0004 and must renumber (§7 Q6) |
4. Product Invariants
INV-COPILOT-001 to INV-COPILOT-014 carry forward unchanged; the full text is in prd.md §4, and stories cite them by id.
- INV-COPILOT-001
src/engine/**is pure: nocloudflare:workersimport, and noDate.now,Math.random,fetchorcrypto.getRandomValues. Identical input gives identical output, so every decision is replayable. - INV-COPILOT-002
said_*items persist only on rep turns.*_knownfacts persist on any turn. Nothing un-persists within a session. - INV-COPILOT-003 Every rep-facing policy string is approved text or a typed approved template. The loader hard-rejects
[VERIFYand any line without an allowlistedapproved_by, with no bypass. - INV-COPILOT-004 Transcript text is pseudonymised before it reaches D1, the DO, Jev, R2
eval/or the browser. Coach prose never enters D1. - INV-COPILOT-005 Every answer row stores its hashes, its model and the exact post-budget state. Every paid request writes a
jev_requestsrow. A cached answer is never served to a different state. - INV-COPILOT-006 No request above 12,000 estimated tokens leaves the DO. Score levels are strings.
chars_per_tokenis policy data. - INV-COPILOT-007 Access fronts every hostname, and the Worker verifies the JWT. No credential appears in the repo or in a published page.
- INV-COPILOT-008 D1 statements bind ≤ 100 parameters (
chunkForD1). - INV-COPILOT-009 The browser owns the clock. The DO sets no timers or alarms and processes messages on one serial chain.
- INV-COPILOT-010 HTTP 402 code 2021 stops the session and is never retried. Errors hold the previous snapshot and are marked
unknown. - INV-COPILOT-011
REWRITE_ENABLEDdefaults to false. - INV-COPILOT-012 No closing or activation probability appears anywhere.
- INV-COPILOT-013 A policy is promoted only through a passing eval run; there is no bootstrap.
- INV-COPILOT-014 Every transcript line, policy text, note or tag reaches the DOM only through
textContent/createTextNode.
New in M1b:
- INV-COPILOT-015 No unapproved text is shown (it extends INV-003).
- A policy text enters
policy/srconly approved. The approval comes from an approver-supplied file (Stevan'spolicy/proposals/2026-09-27.jsoncarriesapproved_by), applied byscripts/apply-label-approvals.mjs. No story setsapproved_byitself. - The views render a new text id only when the pinned bundle holds it approved; otherwise they keep today's approved text or omit the line.
- The situation
address_ties_me_to_tax_residencystays out ofpolicy/srcuntil decision 13. Its[VERIFY]point is a draft; its approved headline and hint wait with it, because the builder requires a situation'score_text_idto be one of its points (src/policy/build.ts:587). - Sources: FEEDBACK §2d; decision 2f.
- A policy text enters
- INV-COPILOT-016 A fact item never nags.
- A
factitem has no live row. When satisfied it moves toplan.knownwith its source, and it returns as the card only while itssurface_whenfires. - Fact and judgement items are never counted in any score, live or in the review.
- Sources: FEEDBACK §2c-§2d; proposal §5.1.
- A
- INV-COPILOT-017 One concern row per episode window.
- A concern clears only on a client utterance of ≥ 4 words that carries the Jev clearing signal.
- A raise of the same type within
concern_reopen_window_s(60 s of call time) of the last close keeps the previousstart_i, so the same row reopens. - Sources: FEEDBACK §2b; proposal §6.
- INV-COPILOT-018 Nothing renders before it is said.
- A transcript row appears only once the clock passes its
u.t, and word k only at its estimated time. Nothing beyond the clock is drawn. - A decision's feedback appears only after its utterance's
t_endand the decision's arrival. - A seeded fact names its real source: "From sign-up" only for
signup. A seeded chip has no time link. - Sources: FEEDBACK §1, §2a; proposal §3.2.
- A transcript row appears only once the clock passes its
- INV-COPILOT-019 Curation survives re-stitching.
- A forced re-import of a call that holds marks or labels runs only with the curation counts and latest timestamps of its export (
tests/fixtures/curation/<call>.rev<N>.json, currenttranscript_rev). The replacement's claim statement checks them atomically, so a mark written after the export makes the replacement change nothing and answer 409. - Marked moments, marks and labels (including
gold_json.i) are re-inserted at the remappediby an idempotent, audited batch. Mark and label counts are equal before and after. - Sources: proposal §8 risk 1, §10 item 1.
- A forced re-import of a call that holds marks or labels runs only with the curation counts and latest timestamps of its export (
- INV-COPILOT-020 Pre-call facts are closed-vocabulary and snapshotted.
- On every write path,
precall_jsonholds only policy slot label ids, ISO 4217 codes from the bank,{digits, code}amounts and asourcein {signup, crm, demo}.signupis reserved for the M2 join. - The value is copied, with its converted seed, into
sessions.precall_jsononce at session start, and never changes for that session. - A forced transcript replacement without
precallin its payload keeps the call's value (112b). - Sources: proposal §5.2; INV-001, INV-004.
- On every write path,
5. User Stories
Every story is one code commit of one concern, plus at most one derived-artefact commit (§3 convention 1). "Estimate" counts otto iterations. A story with an eval gate or a Jev-dependent probe also carries a one-iteration retry budget, reported separately in the sizing table. Line numbers are v1 (transcript_rev 1) unless a criterion says otherwise. Ordering reasons for dependencies are in each story's Notes. Completed M1 stories on main (COPILOT-046, 058e, 065, 070 and the rest) are prerequisites, never dependency ids.
Phase A — Live granularity in replay
ID: COPILOT-102a
Title: Screenshot tooling:
--at,--until,--expectand--expect-nonefor the deployed dashboard [INTEGRATION-CRITICAL]Description: As the builder, I want the screenshot script to pause on a time or on a rendered state and to assert what the page shows, so that every M1b UI story's evidence is taken in the page's own session and fails when the state is absent.
Acceptance Criteria:
scripts/design/shoot-dashboard.mjs(HOT: this story owns the flags) gains four flags, documented in its header:--at <mm:ss>: after load, press 20× and Play, and pause once the player clock reaches the time. Wait until the session queue is empty, then shoot.--until '<css selector>': play at 20× and pause the first time the selector matches, polling every frame; exit 1 if it never matches before the call ends.--expect '<css selector>': after the pause, poll up to 10 s for a match; print the match count; exit 1 when there is none. It never prints page text.--expect-none '<css selector>': exit 1 when there is any match.
Each flag's core is a pure function in
scripts/design/lib/shoot-flags.mjs(new).Tests.
tests/unit/shoot-dashboard.test.ts, over happy-dom fixture pages, covers flag parsing;--untilpausing on the first match;--expectfailing with no match; and--expect-nonefailing on a match.npx vitest run tests/unit/shoot-dashboard.test.ts→ all passed;npm run typecheckexits 0.SAFETY: without the new flags, the script behaves exactly as today (
--deployed,--play,--wait).[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD, asserted by the script), against real Playwright playback of 3339895706. Each case prints its exit code:- Positive.
node scripts/design/shoot-dashboard.mjs /tmp/m1b-102a '[["pos","?call=3339895706","light",1440,900]]' --deployed --until '.item.current' --expect '.item.current'→ 0, and the PNG exists. --at.… '[["at","?call=3339895706","dark",1440,900]]' --deployed --at 00:30 --expect '.item'→ 0; the printed pause time is ≥ 30 s, and the queue is empty at the shot.- Negative,
--expect-none.… --deployed --at 00:30 --expect-none '.item'→ 1. - Negative,
--until.… --deployed --until '[data-never-present]'→ 1, after the call ends at 20×.
Quote the four exit codes.
- Positive.
Dependencies: none
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-102
Title: Transcript text appears word by word between a line's start and end [INTEGRATION-CRITICAL]
Description: As a rep watching a replay, I want each line's words to appear as they are spoken rather than when the line ends, so that a 45 s client turn reads as it happens and nothing appears before its estimated spoken moment.
Acceptance Criteria:
web/src/transcript.tsLiveTranscript.update(todaypassedIndexovert_end,:207-212) changes as follows:- The row for line
iis inserted once the clockt ≥ u.t. - Its text is the first
kwords, wherek = floor(u.words × clamp((t − u.t) / (u.t_end − u.t), 0, 1)). A zero-span line shows whole. - The row carries
data-partialuntilt ≥ u.t_end. - The latest-line tint sits on the row being spoken.
- Rows stay in
iorder, so a backchannel whosetfalls inside a partial line sits above it (v1 lines 6 and 7 inside line 8). - A seek back removes rows with
u.t > tand truncates partial ones. - Covered tags attach only to complete rows.
- Nothing beyond the clock is drawn.
- Words are spread evenly over the line's own span, because
GET /api/calls/:idhas no fragment times. The v2 parts keep each span to about 12 s. - Text reaches the DOM only through
textContent. - The reveal rule is a pure exported function,
revealedWords(u, t), which 118b reuses.
- The row for line
docs/design/DESIGN.md§Transcript pane gains: "Text appears word by word between the line's start and end. Inside a line the times are estimated by spreading its words evenly until live transcription supplies real ones (M3)."Tests.
tests/unit/web-transcript.test.tsontests/fixtures/stitched/3339895706.v1.json. Line 8 hast44.99,t_end89.45 and 112 words.- At 44.9 s, row 8 is absent.
- At 60.0 s, it shows 37 ± 1 words and has
data-partial. - At 89.5 s, it is complete, without
data-partial. - At 70.0 s, row 7 (
t69.35) sits above row 8. - A seek from 89.5 s to 50.0 s leaves row 8 at 12 ± 1 words and removes row 7.
npx vitest run tests/unit/web-transcript.test.ts tests/unit/web-dom-safety.test.ts→ all passed.npm run typecheckexits 0.grep -rn "innerHTML\|insertAdjacentHTML" web/srcprints nothing.[UI] Design artifact: DESIGN.md §Transcript pane (G) and
docs/design/shotgun/G/live.html. No new token or colour.[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/deploy.mjs.node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-102-light","?call=3339895706","light",1440,900],["COPILOT-102-dark","?call=3339895706","dark",1440,900]]' --deployed --at 01:00 --expect '[data-partial]'exits 0 and writes both PNGs. Quote the match counts.
INVARIANT: INV-COPILOT-018; INV-COPILOT-014; INV-COPILOT-009.
Dependencies: COPILOT-102a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-103
Title:
STITCH_V2: sentence-and-time split of long turns, with true part start timesDescription: As the engine, I want long turns split into sentence-sized parts under decision 1a's rule, so that Jev is asked every 5-12 s during a monologue, while v1 output stays byte-identical.
Acceptance Criteria:
- Parameters (
src/ingest/stitch.ts):StitchParamsgains optionalsplitMinWords,splitMinS,splitForceSandsplitLookaheadS.export const STITCH_V2 = Object.freeze({ ...STITCH_V1, version: 2, splitMinWords: 6, splitMinS: 5, splitForceS: 12, splitLookaheadS: 3 }).
- Where it runs: inside each merged turn, after the merge step (
:160-171) and before the parts loop (:186-213). - Cut rules:
- A part is cut at a sentence end (a word ending
.,?or!) once it holds ≥splitMinWordswords and its estimated span is ≥splitMinS. - When the span passes
splitForceSwithout such a cut, the part is cut at the first sentence or clause boundary (. ? ! ;or,) within the nextsplitLookaheadSseconds. With no boundary in that window, it is cut at the word boundary where the lookahead ends. The lookahead is bounded, so no part exceedssplitForceS + splitLookaheadS, with one exception: a single word longer than that. - A trailing piece under 4 words that is not the last of its turn joins the part before it.
- A part is cut at a sentence end (a word ending
- Spans come from the member fragments' real start and end, and from word share inside a fragment.
- Part times: the first part keeps
turn.t; each later part'stis the previous part'st_end; the last part'st_endisturn.t_end. - Unchanged from v1:
- The v1
maxWordssplit still applies inside a part. kind,member_ids(the turn's) anddecision_point, as today.- The output order (
t_end, thent,:230-233). - No merge across speakers.
- Without the new params,
stitch()is exactly v1.
- The v1
- Tests (
tests/unit/stitch.test.ts):- Golden.
stitch(fragments, STITCH_V1)on the four samples' fragments (tests/fixtures/samples/<call>.json) deep-equals the utterances oftests/fixtures/stitched/<call>.v1.json. - 3339895706 under
STITCH_V2:- 125-145 decision points.
- The longest gap between consecutive decision points (by
t_end) is ≤ 15 s, and the p90 is ≤ 11 s. - No part is under 4 words unless it is the last of its turn.
- Within each turn,
tis non-decreasing and everyt_end ≥ t. - The test prints the three numbers, plus the other three samples' decision-point counts (about 165, 68 and 44 expected).
- Sparse punctuation. A synthetic 40 s fragment with no punctuation cuts every ≤ 15 s. A comma 2 s after 12 s is taken, and one 5 s after is not.
- Golden.
- v2 fixtures.
node scripts/make-fixtures.mjs --stitch-only(new; it needs nodata/, so it runs in an otto worktree) rebuildstests/fixtures/stitched/<call>.v1.jsonand writes<call>.v2.json(kindstitched,stitch_version: 2) from the committedtests/fixtures/samples/<call>.jsonfor the four samples. After the run:git diff --exit-code tests/fixtures/stitched/*.v1.jsonpasses;git status --porcelain tests/fixtures/stitchedlists exactly the four.v2.jsonfiles.
npx vitest run tests/unit/stitch.test.ts tests/unit/fixtures-clean.test.ts→ all passed;npm run typecheckexits 0.- SAFETY:
STITCH_V1, the default parameter ofstitch(),INGESTand every Worker caller stay unchanged here. 104d moves the Worker and 106 the policy. - INVARIANT: INV-COPILOT-004 (fixtures are built from pseudonymised samples only).
- Parameters (
Dependencies: none
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-104a
Title: Export the review sessions, the curation and the CRM key before any re-import, from one consistent snapshot [INTEGRATION-CRITICAL]
Description: As Stevan, I want my replay session, the run it matched, and every mark, label and CRM key saved as fixtures before a transcript is replaced, so that a forced re-import, which deletes every derived row of the call (
DERIVED_TABLES,src/routes/calls-import.ts:24), loses nothing.Acceptance Criteria:
- Curation export.
scripts/export-call-curation.mjs --call <id>runs read-onlySELECTs throughnode scripts/with-cf-env.mjs npx wrangler d1 execute copilot --remote --json. It writestests/fixtures/curation/<call>.rev<N>.json(kindcuration), where N is the call's currenttranscript_rev: 1 for 3339895706, 3303259297 and 3347356034, and 2 for 3485591407 on 2026-09-27.- Contents:
- the call row's
transcript_rev,stitch_versionandpd_ct_id; - the moments that carry a mark, with their
marks; - the call's
labels, with theirlabel_idandgold_json; i, t, t_end, speaker, kind, words, member_idsfor every utterance, without text;expected_curation: {marks, labels, last_mark_at, last_label_at};policy_at_export: {version, checklist_ids, risk_ids}: the call scenario's published policy at export time, the evidence for a label whose target id a later policy retires.
- the call row's
- Consistency.
(transcript_rev, content_hash, expected_curation, sha256 of the exported rows)is read before and after the rows. The export is written only when both reads are equal and match the exported rows. Otherwise it retries, up to 3 times, then exits 1. So a replacement, a restore or a new mark during the export is detected.tests/unit/export-curation.test.tssimulates a replacement between the two reads. - Quoted counts: 3339895706 has 3 marks; 3303259297 has 1 label (
i75,gold_json.i75); 3485591407 haspd_ct_id109829; the rest have none.
- Contents:
- Decisions export.
--sessions 49e3e48f-c169-458d-85a6-ae6c5c4e745d,1d160d96-9b60-480c-bb67-c146f8739b13writestests/fixtures/decisions/3339895706.v4.<first 8 characters of the session id>.json(kinddecisions).- Per decision:
i;decision_jsonwithoutjevtelemetry; and the Jev answers it consumed (theanswersrow atanswers_state_hash/answers_qset_hash, themove_answersrow atmove_state_hash/move_qset_hash, ornull). - Quoted counts: 96 decisions up to i=121, and 86 up to i=107.
- These answers are the evidence that 105a, 107 and 117 replay and 119 quotes (v1 i=30). A re-import deletes the D1 rows.
- Per decision:
- Fixture kinds.
FIXTURE_KINDS(scripts/lib/fixtures.ts) gainscuration,decisionsandremap.tests/unit/fixtures-clean.test.tscovers them, including a test of the consistency retry on changing counts.npx vitest run tests/unit/fixtures-clean.test.ts tests/unit/export-curation.test.ts→ all passed. - Docs.
docs/demo-script.mdanddocs/runbook/m1-acceptance.mdeach gain one line: "Utterance numbers here aretranscript_rev1 (the v1 stitch); COPILOT-104e regenerates them." - [INTEGRATION-CRITICAL] Live probe: both export commands, run against remote D1, print the counts and
expected_curationabove. Quote them. - SAFETY:
SELECTonly; no write. - INVARIANT: INV-COPILOT-019; INV-COPILOT-004.
- Curation export.
Dependencies: none
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-104f
Title: The v1→v2
iremap on effective spans, speakers and fragmentsDescription: As the operator, I want every exported moment, mark and label mapped from the v1 lines to the v2 parts that hold the same words, so that curation lands on the same words after the re-import.
Acceptance Criteria:
scripts/remap-i.mjs --export <curation file> --to tests/fixtures/stitched/<call>.v2.jsonprints the map as JSON. Its pure core isscripts/lib/remap-i.ts, and--applyis 104c's.- Effective spans. v1 split parts of one turn share
tandmember_ids(v1 lines 89 and 90:t464.435,member_ids[119],t_end500.045 and 505.81). So each later part's effective span starts at the previous same-turn part'st_end: line 90's span is [500.045, 505.81]. - Line to parts. An old line maps to the v2 parts of the same speaker whose
member_idsintersect its own and whose[t, t_end)overlaps its effective span. A backchannel maps to the v2 backchannel with the samemember_ids. - Tie-breaks. A time belongs to the part whose
[t, t_end)contains it; the last part owns itst_end. - Moments. A moment's
start_i/end_imap to the first and last matched part. - Labels. A label maps to the part containing the old line's effective end. Both
labels.iandgold_json.iare rewritten, andlabel_idis kept. - Moment ids.
moment_idis rebuilt bymomentId()(src/engine/moments.ts:114), keeping the policy-hash suffix. Each mark'smoment_idfollows; mark ids are kept. - Unmatched lines. It lists every old line with no match; none are expected.
- Effective spans. v1 split parts of one turn share
Tests.
tests/unit/remap-i.test.tscovers:- a split turn (v1 i=8);
- v1 i=89 and i=90, each to the parts inside its own effective span (a moment on i=90 must not map back to 464.435);
- a label at the end of a long rep turn;
- overlapping speakers;
- a boundary tie;
gold_json.irewritten withlabel_idkept;- a mark following its moment;
- a restored false-positive label still dismissing its item when
applyDurableMutations(src/session/checkpoint.ts:290) re-applies it on a v2 state.
npx vitest run tests/unit/remap-i.test.ts→ all passed;npm run typecheckexits 0.
Dependencies: COPILOT-103, COPILOT-104a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-104b
Title: Re-import CLI and route:
--stitch 2,pd_ct_idkept, and an atomic curation precondition on forced replacement [INTEGRATION-CRITICAL]Description: As the operator, I want a forced transcript replacement to refuse atomically when the call's marks or labels changed since the export, and to keep the call's CRM key, so that 104e's destructive step cannot delete curation made after the export.
Acceptance Criteria:
Route (
src/routes/calls-import.ts; HOT, first):- A forced replacement of a call that holds marks or labels requires
expected_curation: {marks, labels, last_mark_at, last_label_at}in the request; without it the answer is 409curation_required. - The replacement's claim statement (the
UPDATE calls … WHERE …of the onedb.batch()) addsANDconditions comparing the current counts and the latestcreated_atof the call's marks and labels with those values. A mark or label written after the export therefore makes the claim change 0 rows, and the route answers 409curation_changedwith the database untouched. - A first import (no call row) needs no
expected_curation. - A forced replacement that omits it has its claim assert zero marks and zero labels, atomically. A mark written after a preliminary empty check therefore makes the claim change 0 rows (409
curation_changed). - A payload without
pd_ct_idkeeps the stored one (COALESCE); one with it replaces it.
- A forced replacement of a call that holds marks or labels requires
CLI.
scripts/import-fixture.mjs --stitch 1|2(default 2) wrapstests/fixtures/stitched/<id>.v<stitch>.json, and sendspd_ct_idandexpected_curationfromtests/fixtures/curation/<id>.rev<N>.json.- On a forced replacement of a call with curation, it exits 1 when that export is missing for the call's current
transcript_rev. --print-payloadprints the request without sending it.scripts/import-call-coach.mjsgains the same--stitchflag.
- On a forced replacement of a call with curation, it exits 1 when that export is missing for the call's current
Demo check.
scripts/lib/demo-check.tsandscripts/demo-check.mjs:53readtests/fixtures/stitched/<call>.v<N>.jsonfor--stitch N(default 2).tests/unit/demo-check.test.tscovers both versions, and its existing integration test pins--stitch 1with the v1final-20260927report until 104e re-pins it to v2.Tests (
tests/workers/calls-import.test.ts):- a mark written between reading
expected_curationand the claim gives 409curation_changed, and every row is unchanged; - a missing
expected_curationon a call with curation gives 409curation_required; - with
expected_curationomitted, a mark written after an empty check gives 409curation_changed; - a first import succeeds with no export (regression);
pd_ct_idis kept on a forced v2 replacement.
npx vitest run -c vitest.workers.config.ts tests/workers/calls-import.test.tsandnpx vitest run tests/unit/demo-check.test.ts→ all passed;npm run typecheckexits 0.- a mark written between reading
[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD), negative only:node scripts/import-fixture.mjs --call 3339895706 --stitch 2 --force --expected-rev 1 --print-payload, withexpected_curation.marksset to 0 byjq, is sent withcurl -s -X POSTand the service-token headers to/api/calls/import?force=1&expected_rev=1. It answers 409curation_changed.node scripts/with-cf-env.mjs npx wrangler d1 execute copilot --remote --command "SELECT transcript_rev, stitch_version FROM calls WHERE call_id='3339895706'"still shows 1 and 1. Quote both.
INVARIANT: INV-COPILOT-019; INV-COPILOT-008.
Dependencies: COPILOT-103, COPILOT-104a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-104c
Title: A validated, idempotent curation-restore route and
remap-i --apply[INTEGRATION-CRITICAL]Description: As the operator, I want the remapped marks and labels restored through one validated route, which refuses rows that belong to another call and ignores a repeated post, so that restoration after a re-import is safe to repeat.
Acceptance Criteria:
src/routes/curation.tsaddsPOST /api/calls/:id/curation, registered by one dispatch line insrc/index.ts(HOT). Service token only; an email actor gets 403.- Body:
{expected_rev, moments[], marks[], labels[], policy_at_export?}. - Field-level schema. Every row is validated against it before any write; a failure answers 422, naming the row and field.
- Columns (the schema's real ones,
migrations/0001_init.sql):moments:moment_id, call_id, scenario, kind, start_i, end_i, topic, jev_json, policy_hash, priority, review_state, created_at;marks:mark_id, moment_id, verdict, note, use_as_example, admin, created_at;labels:label_id, call_id, i, question_id, gold_json, admin, created_at.
- Identifiers:
call_idequals the URL's;moment_idequalsmomentId(call_id, kind, topic, start_i, policy_hash)recomputed from the row;mark_idandlabel_idare strings of the existing id formats;- each mark's
moment_idis a moment in the same body; question_idis a policy id;kindandverdictare the schema'sCHECKvalues.
- Integers:
0 ≤ start_i ≤ end_i < n_utterancesatexpected_rev, and0 ≤ labels.i < n_utterances. - Audit metadata:
created_atis ISO 8601, andadminis an email or a service-token common name, as the writers store it. gold_json: it parses to exactly the durable-label shapeFalsePositiveLabel(src/session/checkpoint.ts:207):{value: false, kind: 'false_positive', target, id, i}, with no other property.targetis one ofDismissTarget(src/engine/dismiss.ts:8):risk,must_sayorconcern.idfits the target: a risk-flag id or a checklist id of the call's scenario policy, or forconcernthe id form the dismiss path writes.- Retired ids. A
must_sayorrisklabel whoseidthe current policy no longer checks (110d retires six onboarding ids) is accepted when the body carriespolicy_at_export(104a) and the id is in itschecklist_idsorrisk_ids. It is restored unchanged (the closedgold_jsonshape) and reported underretiredin the response and in the audit row, with the export's policy version as its evidence. A dismissal of an id the engine no longer checks has no effect, so restoring it is harmless and keeps the record. An id in neither the current nor the exported policy is still 422. iequals the label'si, andidequalsquestion_id.- Any other shape, kind or property answers 422.
- Text:
marks.noteandmoments.topicpassassertRedactedText/isPseudonymised;- every string inside
moments.jev_jsonpasses them too, andjev_jsonparses; - identifiers, timestamps and
adminare not prose and skip the guards.
- Input duplicates: a duplicate id within the body answers 422.
- Columns (the schema's real ones,
- Collisions. The primary keys are the identity (
moment_id,mark_id,label_id). A key that already exists answers 409conflict, naming the row, unless the stored row belongs to this call and is identical in every column. An identical stored row is a no-op. - Writes. One
db.batch()guarded byrevGuard(call_id, expected_rev), plusaudit_log{action:'call.remap_i'}with the counts. A staleexpected_revanswers 409 with nothing written. - Apply.
scripts/remap-i.mjs --applyposts the remapped export. - Tests (
tests/workers/curation.test.ts): a real-row round trip (rows exported from a seeded test database, remapped by 104f and posted back are identical except the remapped fields); insert; identical re-post (zero counts); another call'smoment_idorlabel_id→ 409; a non-deterministicmoment_id→ 422; a duplicate in the body → 422; each field rule; agold_jsonwith an extra property, anotherkind, an unknowntargetor a foreignid→ 422; a label on a retired id withpolicy_at_exportevidence → restored, and reported asretired; stale rev → 409; email → 403; the audit row.npx vitest run -c vitest.workers.config.ts tests/workers/curation.test.ts→ all passed. - [INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD), on a disposable synthetic call:node scripts/make-synthetic-call.mjs --from 3347356034 --decision-points 20 --id m1b-curation-probeimports it.- POST one moment, one mark and one label built for it → 200, counts 1/1/1.
- POST the identical body again → 200, zero counts.
SELECTthe three rows and compare them with the body; quote both.- POST a label reusing the probe's
label_idwith a differentgold_json→ 409conflict. - Quote the audit rows.
- INVARIANT: INV-COPILOT-019; INV-COPILOT-004; INV-COPILOT-007; INV-COPILOT-008.
Dependencies: COPILOT-104f
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-104d
Title: Worker ingest on
STITCH_V2: uploads and speaker-map re-stitching produce sentence-sized parts [INTEGRATION-CRITICAL]Description: As Stevan, I want uploaded calls stitched under the same v2 rule as the imported samples, so that an upload never replays with 45 s turns.
Acceptance Criteria:
src/routes/ingest.ts:381andsrc/routes/speaker-map.ts:308callstitch(…, STITCH_V2)and writestitch_version = 2(ingest.ts:407,speaker-map.ts:354).INGEST(src/policy/build.ts) is unchanged here; 106 changes it with its candidate.Forced paths. Both forced replacement paths (
ingest.ts's claim at:400-412andspeaker-map.ts's at:350-356) gain 104b's atomic curation condition. The claim asserts theexpected_curationsent, or zero marks and labels when none is sent, so a mark written concurrently makes the claim change 0 rows (409curation_changed).Tests.
tests/workers/ingest.test.tsassertsstitch_version = 2and v2 decision points on a new upload.tests/workers/speaker-map-apply.test.tsasserts the same after a forced speaker-map re-stitch.- For both paths, a concurrent mark written before the claim gives 409 and leaves every row unchanged.
npx vitest run -c vitest.workers.config.ts tests/workers/ingest.test.ts tests/workers/speaker-map-apply.test.ts→ all passed.[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):- Pick one of the four uploaded copies of 3339895706 (
up_06337679…,up_0873645c…,up_869446e2…,up_ef077387…) with 0 marks and 0 labels. - Re-stitch it through
POST /api/calls/<upload id>/speaker-mapwithforceand its stored role map (the COPILOT-070 path; no transcription cost). SELECT stitch_version, n_decision_points FROM calls WHERE call_id = '<upload id>'shows 2 and a count above the v1 96. Quote it.
The other copies stay v1, and 118's register says so.
- Pick one of the four uploaded copies of 3339895706 (
INVARIANT: INV-COPILOT-004 (the re-stitch works on stored pseudonymised fragments).
Dependencies: COPILOT-103, COPILOT-104b
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-104e
Title: Re-import the four samples under
STITCH_V2, restore the curation, remap the labels and the demo script, measure cost [INTEGRATION-CRITICAL]Description: As Stevan, I want the four sample calls re-imported at the new granularity, with my marks, labels and the CRM key carried over and the cost measured, so that the deployed replay shows sentence-level cadence without losing any curation.
Acceptance Criteria:
[INTEGRATION-CRITICAL] Live run on the deployed Worker (
health.version== HEAD), in this order:Before.
node scripts/with-cf-env.mjs npx wrangler d1 execute copilot --remote --command "SELECT COUNT(*) AS n FROM marks k JOIN moments m ON m.moment_id = k.moment_id WHERE m.call_id = '3339895706'"→ 3.… "SELECT COUNT(*) AS n FROM labels WHERE call_id = '3303259297'"→ 1.- Quote the published policy versions (
GET /api/policy/onboarding,…/customer_success). This story's evaluation pins them.
For each sample:
node scripts/import-fixture.mjs --call <id> --stitch 2 --force --expected-rev <N>printsimported <id> rev <N+1>; 104b's precondition guards it.node scripts/remap-i.mjs --export tests/fixtures/curation/<id>.rev<N>.json --to tests/fixtures/stitched/<id>.v2.json --applyprints the inserted counts.- After an interruption, re-running
--applycompletes the restore (104c is idempotent). Ifimport-fixtureanswers 409curation_changed, re-export (104a) and repeat.
After.
… "SELECT c.call_id, c.transcript_rev, c.stitch_version, c.pd_ct_id, c.audio_r2_key IS NOT NULL AS has_audio, SUM(u.decision_point) AS dp FROM calls c JOIN utterances u ON u.call_id = c.call_id WHERE c.call_id IN ('3339895706','3303259297','3485591407','3347356034') GROUP BY c.call_id"must show:stitch_version2 on all four;dp125-145 for 3339895706 (the others quoted);has_audio1 for 3339895706;pd_ct_id109829 for 3485591407.
The two counts of step 1 again give 3 and 1.
Replay.
node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --eval-uid m1b-stitch2-<yyyymmdd>-r→errors == 0,requests == decision_points,max_input_tokens ≤ 12000. Quotereconnects(about 5).Eval.
node scripts/eval.mjs --scenario both --layers l4,l6 --policy-version <the version of step 1> --run m1b-stitch2-<yyyymmdd>.- L6 passes: every input ≤ 12,000, and
cost_by_call≤L6_LIMITS.cost_per_call_usdfor 3339895706 and 3303259297 (about $0.047 and $0.048 expected). - Quote L4, pass or fail; the L4 gate on v2 is 106's.
- Append "M1b: STITCH_V2 (104e)" to
docs/reports/cost-latency.md.
- L6 passes: every input ≤ 12,000, and
Derived-artefact commit (§3 convention 1):
tests/fixtures/remap/<call>.json(kindremap), the four v1→v2 maps.eval/labelled/must-say.jsonwith everyirewritten through the map.tand label values stay unchanged, with aremapped: {from_rev, date}note per call.docs/demo-script.mdrows (i,[mm:ss]) fromeval/reports/m1b-stitch2-<yyyymmdd>.l4l6.json, withtests/unit/demo-check.test.tsre-pinned to it and--stitch 2(a test-only change).- The
ireferences indocs/runbook/m1-acceptance.md, rewritten.
Then
node scripts/demo-check.mjs eval/reports/m1b-stitch2-<yyyymmdd>.l4l6.json docs/demo-script.mdexits 0, and on the final commitnpx vitest runandnpm run typecheckpass (includingtests/unit/eval-cases.test.ts,eval-like-for-like-cases.test.ts,eval-l4l6.test.tsanddemo-check.test.ts).INVARIANT: INV-COPILOT-019; INV-COPILOT-004; INV-COPILOT-005.
Dependencies: COPILOT-104b, COPILOT-104c, COPILOT-118a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, eval)
Notes: The destructive step. 110c and 110d may have promoted candidates before this story runs, so the evaluated versions are the ones published at step 1, not an assumed v4.
ID: COPILOT-105a
Title: A presenter simulator over a recorded session [INTEGRATION-CRITICAL]
Description: As the builder, I want to run the page's presenter over a session's recorded decisions and see every change, point and exit on the call clock, so that the cadence and lifetime rules can be measured and their screenshots located without guessing.
Acceptance Criteria:
scripts/presenter-sim.mjs --session <session_id> [--log <replay log>](new; assertshealth.version== HEAD):- It reads the session's decisions from D1 (
with-cf-env, read-only) and the utterances fromGET /api/calls/:id. - It runs
present()(web/src/presenter.ts) at 0.1 s steps. With--log(118a's format), each decision uses its recorded arrival; otherwise arrival is projected ast_end+ 0.3 s and the output is labelledprojected. - It prints JSON:
- every card switch, with whether a client run was in progress;
- the first point added to the open card mid-run;
- the first time ≥ 2 points show;
- every
lastSuggestioninterval; - every point's on-screen interval, with its exit reason;
- the first time each item id opened.
Later stories append fields.
- It reads the session's decisions from D1 (
Tests.
tests/unit/presenter-sim.test.tsruns the simulator's core overtests/fixtures/decisions/3339895706.v4.49e3e48f.jsonand the v1 utterances. It reproduces the FEEDBACK §3 figures under today's presenter: the amount point on screen for 1.0 s, and the first-transfer-date point for 23 s.npx vitest run tests/unit/presenter-sim.test.ts→ all passed.[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-105a-<yyyymmdd>→errors == 0.node scripts/presenter-sim.mjs --session <session_id>prints its JSON. Quote the switch count and the point intervals.
INVARIANT: INV-COPILOT-009 (it reads; it sets no timer anywhere).
Dependencies: COPILOT-104a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-105
Title: Presenter: the change moment is the end of the client's run; new points on the open card move at sentence ends [INTEGRATION-CRITICAL]
Description: As a rep, I want feedback (ticks, chips, caution, a new point on the open card) within a sentence of what was said, while the card itself waits for the client to finish (decision 1b), so that sentence-level decisions feel live without the headline jumping mid-turn.
Acceptance Criteria:
web/src/presenter.ts(HOT, first):- Run end.
isClientRunEnd(us, u, upto = us.length)is a clientturnwhose nextturn(skipping backchannels and acks) is not the client's, or does not exist beforeupto.uptomeans that in live audio (M3) only what has arrived is read. - Change moments.
dueFor(:153),nextClientEnd(:140) andreleaseAt(:164) use run ends only.- A plan that changes the current item applies at the first boundary at or after L = max(arrival, dwell end, the end of the client run its decision belongs to). A boundary is L itself when no utterance is in progress at L; otherwise it is the end of the client run or the rep turn in progress. A client part that ends mid-run is never a boundary.
- A decision on a run's last part therefore applies on arrival, once the dwell has passed.
- An inserted item (
isInsert,:125) uses the same rule, without the dwell. - After the call's last
t_end, anything pending applies (as today).
- On arrival. Ticks, chips, the caution and the purpose still apply on arrival. A plan whose current item id equals the open card's applies on arrival for any client-part decision: its new point shows at once, at most one per applied plan (
apply,:176). - Flag.
export const LIVE_CARD_AT_SENTENCE_END = false. When true, every client part end is a change moment. - Marker. A point revealed while a client run is still in progress carries
data-midrun="1"(set by the presenter, rendered byweb/src/plan.ts; HOT, first) until that run ends.
- Run end.
docs/design/DESIGN.md§Cadence rules:- Rule 1 is rewritten: items change state at the end of the client's run (the last client part before the rep speaks), after the dwell; or at every client part end while the flag is true.
- Rule 4 gains "a new point on the open card may appear at any client sentence end".
- The constants table gains the flag.
Tests.
tests/unit/web-presenter.test.tsontests/fixtures/stitched/3339895706.v2.json, with synthetic plans:- a card switch decided on a middle part of v1 line 8's parts applies at the last part's
t_end; - a switch whose dwell ends mid-run applies at that run's end, not at the dwell's end;
- a decision on a run's last part, arriving after the run ended, applies on arrival once the dwell has passed;
- a point added to the open card on a middle part shows at that part's
t_endplus arrival, withdata-midrun="1", and the card id is unchanged until the run ends; - with the flag true, the switch applies at the middle part;
- the existing cases pass, updated only where the change moment moved (named in the progress note).
npx vitest run tests/unit/web-presenter.test.ts tests/unit/presenter-sim.test.ts tests/unit/web-plan.test.ts→ all passed;npm run typecheckexits 0.- a card switch decided on a middle part of v1 line 8's parts applies at the last part's
[UI] Design artifact: DESIGN.md §Cadence rules (rules 1 and 4) and
docs/design/shotgun/G/NOTES.md(c).[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-105-<yyyymmdd>→errors == 0.node scripts/presenter-sim.mjs --session <session_id>prints 0 switches while a client run is in progress, and the first mid-run point time. Quote both. LIVE-PROBE [INTEGRATION-CRITICAL] on the deployed Worker (health.version == HEAD): (1) node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-105-→ errors == 0; (2) node scripts/presenter-sim.mjs --session <session_id> prints 0 switches while a client run is in progress, and the first mid-run point time with the run's state at that instant (quote both); (3) node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-105-light","?call=3339895706","light",1440,900],["COPILOT-105-dark","?call=3339895706","dark",1440,900]]' --deployed --until '.item.current [data-midrun="1"]' exits 0: it pauses on a point the presenter added while the client's run was still open (data-midrun="1" on the current item), and the presenter-sim line for that instant shows client_run true and the same card id open from that point until the run's end; whether the client's text was still being revealed at that instant ([data-partial]) is reported, not required, because on this call the only mid-run point lands after the client's part ended while a trailing fragment keeps the run open (amended 2026-09-27: the rule under test is that points move at sentence ends while the card waits for the run's end, not a coincidence of reveal and point on one recording). Quote it.
INVARIANT: INV-COPILOT-009; INV-COPILOT-018.
Dependencies: COPILOT-102, COPILOT-102a, COPILOT-103, COPILOT-104e, COPILOT-105a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
Notes: 104e puts the v2 transcript in D1 for the live probe.
ID: COPILOT-106
Title: Re-tune the decision-counted thresholds and the window for sentence-level cadence; cost gate $0.06; a policy candidate on
stitch_version2 [INTEGRATION-CRITICAL]Description: As the policy owner, I want the counters tuned for one decision per turn re-tuned for the new cadence and gated by the eval, and the cost gate set as Stevan ruled (decision 16), so that concerns, purposes and cooldowns keep their real-time meaning and the gate says what it means.
Acceptance Criteria:
Policy source and engine:
Setting Where Change thresholds.concern_max_agepolicy/src/rules-onboarding.json,rules-customer-success.json8 → 12 thresholds.card_cooldown_uttssame 3 → 4 token_budget.windowsame 12 → 16 purpose_max_decisionsplan-onboarding.json,plan-customer-success.json8 → 12 COVERED_DWELLsrc/engine/plan.ts:694 → 6 INGEST.stitch_versionsrc/policy/build.ts:120reads STITCH_V2.version;tests/unit/policy-build.test.tsexpects{stitch_version: 2, redact_version: 1}Kept, and named in the progress note:
confirm_updates2,stage_confirm2,ema_alpha0.4, soft 11,000 and hard 12,000 (decision 15).Cost gate.
scripts/lib/eval-l6.tsL6_LIMITS.cost_per_call_usdgoes 0.05 → 0.06 (decision 16).tests/unit/eval-l4l6.test.tsasserts that the check reads the constant: $0.055 passes and $0.065 fails.docs/architecture.md§9 (the L6 row) anddocs/reports/cost-latency.mdstate $0.06.Regeneration and worst case. Fixtures and stubs are regenerated.
npx vitest run tests/unit/engine-groups.test.ts tests/unit/policy-build.test.ts: the worst case at window 16 is ≤ 12,000 for both scenarios (quoted).npx vitest run tests/unitandnpm run typecheckexit 0.[INTEGRATION-CRITICAL] Candidate and gate (§3 convention 4):
node scripts/eval.mjs --scenario both --layers l1,l4,l6 --policy-version <N> --fresh --run m1b-tune-<yyyymmdd>. It must show:- L1 onboarding and cs ≥ 90 %;
L4: PASSon the v2 transcripts;- L6 PASS, with cost per call ≤ 0.06 quoted per call;
evaluated=true.
window_to_8fallbacks ≤ 11 % of the run's decisions (baseline 10.5 %). Quotenode scripts/with-cf-env.mjs npx wrangler d1 execute copilot --remote --command "SELECT COUNT(*) AS n, SUM(json_extract(d.decision_json,'$.jev.budget.dropped') LIKE '%window_to_8%') AS k FROM decisions d JOIN sessions s ON s.session_id = d.session_id WHERE s.eval_uid LIKE 'm1b-tune-<yyyymmdd>%'".nmust be > 0, andk / nmust be numeric and ≤ 0.11. The path is$.jev.budget.dropped, because$.budgetis a number.- Publish both. Append "M1b: cadence re-tune (106)" to
docs/reports/cost-latency.md.
INVARIANT: INV-COPILOT-006; INV-COPILOT-013.
Dependencies: COPILOT-104e, COPILOT-105, COPILOT-119
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, gate)
Notes: 119 is the policy-chain order. The L6 constant and
INGESTchange with this candidate because its gate evaluates them (round-1 judgement, §7).
Phase B — Card lifetime
ID: COPILOT-121
Title: Design board for the M1b live view: stacked points with the last suggestion, the guide pane, the player status slot [INTEGRATION-CRITICAL]
Description: As Stevan, I want M1b's three visual changes drawn in G's own markup, in both themes, with a recommended variant each, so that 107, 114 and 120 build one design and my sign-off is a single review. This is the shotgun round the proposal put before 107 and 114; FEEDBACK §4 adds the status slot.
Acceptance Criteria:
docs/design/shotgun/M1b/(new) holds three things:board.html;- variant pages built from copies of
docs/design/shotgun/G/live.htmlandlive.css(G's files stay unchanged:git diff --exit-code docs/design/shotgun/G); - a
data.jsof states drawn from 3339895706 (pseudonymised text only, and the approved wording ofpolicy/proposals/2026-09-27.json).
- The board poses three questions, with two variants each, in the
/design-shotgunboard format:- Stacked points. Two stacked points against three (decision 3 set 3; the board shows the reading load). Both variants show a point that arrives grey with its tick after a rep turn (R4), and the grey "Last suggestion: …" line under a new card (R3).
- The guide pane (decision 2b). Chips with a source suffix, the "Guide for this call" head with its subtitle, Now, "Coming up" (≤ 3 grey rows, no state word) and a collapsed "Covered: …" line. Two densities of the Covered line.
- The player status. "catching up (N queued)" in a reserved slot after the speed control, or as an overlay above the bar.
node docs/design/shotgun/M1b/render.mjs(adapted fromdocs/design/shotgun/G/render.mjs) exits 0. It writes every variant at 1440x900, in light and dark, todocs/design/evidence/COPILOT-121-<question>-<variant>-{light,dark}.png.docs/design/evidence/COPILOT-121-board.mdrecords, per question, the variants, the recommendation and the reason, plus "Sign-off: pending (PRD §5.6 M-1)".docs/design/DESIGN.md§Variants gains "M1b board (COPILOT-121): recommended, pending sign-off", with a link. Only §Tokens "Live view tokens (G)" are used, with no new colour; the accent stays on points, ticks, the current dot and the Live pill.- [UI] Design artifact: DESIGN.md §Live view (G) and
docs/design/shotgun/G/. Both themes, as above. - [INTEGRATION-CRITICAL] Live probe (Playwright, a real browser):
node docs/design/shotgun/M1b/render.mjsexits 0.ls docs/design/evidence/COPILOT-121-*-light.png docs/design/evidence/COPILOT-121-*-dark.png | wc -lprints 12: 3 questions × 2 variants × 2 themes. Quote both.
Dependencies: none
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-107
Title: Presenter point lifetime: 12 s floor, stacking to three, sticky last suggestion, rep-turn grace [INTEGRATION-CRITICAL]
Description: As a rep, I want a suggestion to stay long enough to read, later suggestions to stack under it, and never to see a point I have just made flash on, so that "say this" is usable while listening (decision 3, rules R1-R4).
Acceptance Criteria:
web/src/presenter.ts(HOT, after 105):- Constants:
LIVE_POINT_MIN = 12,LIVE_POINTS_MAX = 3(decision 3),LIVE_LAST_SUGGESTION = 12andLIVE_REP_POINT_GRACE_S = 3. - State:
PresentedgainsrevealedAt: Record<string, number>andlastSuggestion: { point, until } | null. - The rules below are pure functions of the clock, and
seekPresentedresets them. They apply in this order of precedence:- Covered. A covered point folds, as today.
- Hard drop. A point whose situation is in the plan's
dropped_hard(108a; absent means none) leaves within 1 s. - Cap (R2). When a new point arrives while
LIVE_POINTS_MAXare shown, the oldest uncovered point leaves at once, even inside its floor. This is the one exception to R1, and decision 3 sets it. - Card switch (R3). On a card switch, the old card's points leave with it. Its newest uncovered point continues as
lastSuggestionforLIVE_LAST_SUGGESTION. - Floor (R1). A point dropped by an uncertain miss, or by a new plan on the same card, keeps its slot, tinted or grey, until its 12 s floor ends.
- The sticky line counts toward
LIVE_POINTS_MAX. While it shows, the card shows at most 2 points. A new point that would exceed the cap removes the sticky line first, because it is the oldest suggestion. A switch onto a card with 3 points drops it at once. - Rep-turn grace (R4). A point first introduced by a plan decided on a rep turn is held until the next rep-turn decision arrives, or until
LIVE_REP_POINT_GRACE_Spasses, whichever comes first.- A client decision in between (v1 line 10) does not release it, because only rep turns carry
covered::(coveredToAsk,src/engine/plan.ts:302). - If the releasing decision carries the point's
covered_at, the point renders grey without the tint.
- A client decision in between (v1 line 10) does not release it, because only rep turns carry
- Client-due insert.
isInsert(:125) also treats an item withdue_by: 'client'(110f's contract field) as inserted.
- Constants:
web/src/plan.ts(HOT, first):lastSuggestionrenders as one grey line under the new card, labelledplan.last_suggestionfrom the pinned bundle's approvedplan.view_labels("Last suggestion", 110c).- Without that label, the line is the approved point text alone, grey, with no prefix (INV-COPILOT-015).
- The card carries
data-points="<n>".
docs/design/DESIGN.md§Cadence rules:- rule 10: the floor, with the precedence above;
- rule 11: stacking, capped at 3;
- rule 12: the last suggestion and the rep-turn grace;
- rule 2 gains "a must-say made due by the client is inserted like a concern";
- the constants table gains the four values.
scripts/presenter-sim.mjs(105a) additionally prints the shortest on-screen time of any point, excluding exits 1-4, and the points that left within 12 s for any other reason.Tests.
tests/unit/web-presenter.test.tsreplaystests/fixtures/decisions/3339895706.v4.49e3e48f.jsonover the v1 utterances, with arrival =t_end+ 0.3 s on a 0.1 s clock:- The amount point, first introduced on v1 i=9 (rep,
t_end99.2), is never rendered tinted. - Every point that exits for a reason other than 1-4 has been on screen ≥ 12 s.
- After each switch the presenter computes, the switched-away card's newest uncovered point is
lastSuggestionfor 12 s, unless a cap eviction removes it first. On this fixture that is the first-transfer-date point, after the switch following v1 i=17. - Never more than 3 suggestions show, counting points and the sticky line together.
- A switch onto a three-point card drops the sticky line.
- A synthetic item with
due_by: 'client'opens without waiting out the dwell.
- The amount point, first introduced on v1 i=9 (rep,
tests/unit/web-plan.test.tscovers the line with and without the label.
npx vitest run tests/unit/web-presenter.test.ts tests/unit/web-plan.test.ts tests/unit/presenter-sim.test.ts tests/unit/web-dom-safety.test.ts→ all passed;npm run typecheckexits 0.[UI] Design artifact: DESIGN.md §Cadence rules 10-12 and §The current item; the recommended card variant on the 121 board.
[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-107-<yyyymmdd>(published policy; no--eval-uid) →errors == 0.node scripts/presenter-sim.mjs --session <session_id>. Quote the first ≥ 2-point time and the shortest on-screen time.node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-107-light","?call=3339895706","light",1440,900],["COPILOT-107-dark","?call=3339895706","dark",1440,900]]' --deployed --until '[data-points="2"],[data-points="3"]'exits 0.
If 3339895706 never shows two points, steps 1-3 run on 3303259297, and the note says so.
INVARIANT: INV-COPILOT-009; INV-COPILOT-015; INV-COPILOT-018.
Dependencies: COPILOT-102a, COPILOT-104a, COPILOT-105, COPILOT-105a, COPILOT-110c, COPILOT-110f, COPILOT-121
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-108a
Title: Engine: situation hysteresis in the uncertain band, counted once per utterance [INTEGRATION-CRITICAL]
Description: As the engine, I want a shown situation to survive one uncertain answer but leave at once on a clear miss (decision 17), so that cards do not flicker on one borderline answer.
Acceptance Criteria:
src/engine/situation.ts(R5):- New fields.
SituationFitgainsshown: string | null,misses: numberandmiss_i: number | null, additively insrc/engine/types.ts. - New threshold. An optional
situation_fit_hard: 0.3goes inrules-*.json. When it is absent, the M1 rule applies (§3 convention 5). - Counting a miss.
recordFits(:172) is called twice per decision, bystep()and byupdatePlan, and stays idempotent. It changesmissesonly whenmiss_i !== u.i, so one decision counts at most one miss.- It increments
misseswhen the shown situation's answer is in[situation_fit_hard, situation_fit), or when its question was dropped. It setsmiss_i = u.i. - It resets
missesto 0 at ≥situation_fit. - Below
situation_fit_hard, it clearsshownat once and adds the situation id to the plan's newdropped_hard: string[](one additive field insrc/engine/plan.ts).
- It increments
- Keeping the shown situation.
shownSituationId(:228) keepsshownwhilemisses < 2, unless another candidate qualifies under like-for-like §3.3 rules 1-2.
- New fields.
docs/policy/like-for-like.md§3.3 gains the amendment: "rules 1-2 decide entering; a shown situation leaves on a second answer in [0.3, 0.6) or a second dropped question, at once under 0.3, or when another candidate qualifies (decision 17, 2026-09-27)."Tests (
tests/unit/engine-situations.test.ts,tests/unit/engine-fit.test.ts):- Through the full
step()path, one answer of 0.45 counts one miss and the situation survives it. - A second such answer on the next decision removes it.
- One dropped question is survived.
- An answer of 0.05 removes it at once, with its id in
dropped_hard. - It leaves when another candidate qualifies (fit ≥ 0.6, with the Choice agreeing).
- With
situation_fit_hardabsent, M1 behaviour holds. - The delivery-date MUST-NOT case (like-for-like §4.1) still shows nothing.
npx vitest run tests/unit/engine-*.test.ts→ all passed;npm run typecheckexits 0.- Through the full
[INTEGRATION-CRITICAL] Candidate and gate (§3 convention 4). After the build and the deploy (
health.version== HEAD):node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --policy-version <N> --session m1b-108a-<yyyymmdd> --eval-uid m1b-108a-<yyyymmdd>→errors == 0,max_input_tokens ≤ 12000.- Quote the count of decisions whose
plan.dropped_hardis non-empty. - The gate run
--run m1b-hysteresis-<yyyymmdd>passes. Publish.
INVARIANT: INV-COPILOT-001; INV-COPILOT-003.
Dependencies: COPILOT-106
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, gate)
ID: COPILOT-108b
Title: Engine: pre-judged fits and situation choice for the item about to open, with a
prejudgedrop step [INTEGRATION-CRITICAL]Description: As the engine, I want the next item's situations judged before it opens, so that a new card opens with points instead of "Listening" (decision 3, R6).
Acceptance Criteria:
Pre-judge questions. On client parts,
fitsToAsk(src/engine/situation.ts:152) also asksfits::for up to 3 situations of the item that would become current next: the ranking's leading challenger, or else the next open item. When that item has ≥ 2 candidates, it also asks the item'ssituation::<move>Choice, becauseshownSituationIdneeds both.Seeding.
retargetFit(:205) seeds the new item'sfitsandchoicefrom those answers, instead of{}/null, when they were asked on the previous decision.Budget. The pre-judge questions are a list of their own in
planRequest, with a code-rule drop stepprejudgedropped beforefact_values(src/engine/budget.ts:121). The drop is recorded inbudget.dropped.docs/policy/like-for-like.md§3.4 gains the pre-judge budget line.Probe flag.
scripts/lib/probe-flags/count-points.ts, plus one registration line inscripts/lib/probe-flags/index.ts(HOT), adds--count-points. It reconnects withload{resume:true}and prints, overhello.history:decisions_with_point: decisions whose current item has ≥ 1 uncovered point;items_with_point: distinct such items;withdrawn_within_12s: points that leave the current item uncovered and not hard-dropped within 12 s of call time.
Tests.
tests/unit/engine-fit.test.ts:ask_about_the_transfer(situationsamount_not_saidandfirst_transfer_date_not_said) opens after a pre-judgedfits::amount_not_said≥ 0.6 and a Choice ofamount_not_saidat ≥ 0.4, and shows a point on its first decision.tests/unit/engine-budget.test.ts:prejudgedrops beforefact_values.npx vitest run tests/unit/engine-groups.test.ts tests/unit/policy-build.test.ts: the worst case with the pre-judge questions is ≤ 12,000 (quoted).
npx vitest run tests/unit/engine-*.test.ts→ all passed;npm run typecheckexits 0.[INTEGRATION-CRITICAL] Candidate and gate. The drop step is policy data (
token_budget.drop_order). After the build and the deploy:node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --policy-version <N> --session m1b-108b-<yyyymmdd> --eval-uid m1b-108b-<yyyymmdd>→errors == 0,max_input_tokens ≤ 12000.node scripts/ws-probe.mjs --call 3339895706 --session m1b-108b-<yyyymmdd> --resume --count-pointsprintsdecisions_with_point> 12 anditems_with_point> 3 (today's 12 and 3; 20 and 5 expected), andwithdrawn_within_12s. Quote all three.- The gate run
--run m1b-prejudge-<yyyymmdd>passes, with cost per call ≤L6_LIMITS.cost_per_call_usd. Quote the delta against 106's run (about $0.003-0.005 expected). - Publish. Append the values to
docs/reports/cost-latency.md.
INVARIANT: INV-COPILOT-001; INV-COPILOT-006.
Dependencies: COPILOT-108a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, gate)
ID: COPILOT-120
Title: The player status has a reserved slot: "catching up (N queued)" never moves the scrubber [INTEGRATION-CRITICAL]
Description: As a rep, I want the player bar to stay still when the session falls behind and catches up, so that the scrubber does not jump left and right while I watch (FEEDBACK §4).
Acceptance Criteria:
web/src/player.tsand its stylesheet:- The status text (
catchingUpLabel,:71) renders in an element that is always present (data-status,aria-live="polite"). - It sits where the 121 board recommends.
- Its width is reserved for the longest status, so nothing else in the bar changes size or position.
- The status text (
tests/unit/web-player-status.test.ts(happy-dom): the element exists with and without a queue; only its text changes; its reserved-width class is present both ways.npx vitest run tests/unit/web-player-status.test.ts→ all passed.scripts/probe-player.mjsgains--check status-slotand--scheme light|dark:- It plays at 20× (a queue forms) and samples the scrubber's
getBoundingClientRect()every 100 ms while the status changes. - It exits 1 if the scrubber's x or width ever differs by more than 0.5 px.
- It prints the sample count, how many samples had the status visible (≥ 1 required), and the largest deviation.
--shot <file>shoots while the status is visible.
- It plays at 20× (a queue forms) and samples the scrubber's
docs/design/DESIGN.md§Player and §States "Catching-up badge" gain the slot rule.- [UI] Design artifact: DESIGN.md §Top bar (the replay player) and §States; the 121 board's status question; both themes.
- [INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/probe-player.mjs --call 3303259297 --view live --check status-slot --scheme light --shot docs/design/evidence/COPILOT-120-light.png, and the same with--scheme dark --shot docs/design/evidence/COPILOT-120-dark.png. Both exit 0. This is the live view's compact bar;probe-playerdefaults to the review view. Quote the largest deviation.
Dependencies: COPILOT-121
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration
Notes: The M1 player (COPILOT-046) is a completed prerequisite on
main; it is not a dependency id in this plan.
Phase C — Plan semantics and pre-call facts
Phase C — Plan semantics and pre-call facts
ID: COPILOT-110a
Title: Plan grammar: policy types, builder validation, text identity and registry, and the typed readers; no policy data change
Description: As the policy owner, I want the item-kind grammar, the new
appliesforms and the two new text locations (view labels and known notes) defined, validated, registered for approval and read safely, before any policy data uses them, so that a bundle with or without them builds, loads and typechecks.Acceptance Criteria:
Policy types (
src/policy/types.ts; HOT, first). Every field is optional (§3 convention 5).- Template items gain:
kind: 'fact' | 'must_say' | 'judgement';satisfied_by: { slots?, facts?, served_check?, rule?: 'served' };surface_when: { any: [ {unknown_after_stage}, {slot_outside_served}, {source_conflict}, {signal}, {slot_in}, {stage}, {always: true} ] };due: {signal} | {any: [...]};texts_from: <template id>;known_note: ApprovedText.
done_whenbecomes optional; it is absent on a judgement item.appliesgains{slot_in: {<slot>: [ids]}, or_unknown?: true, unknown_unless?: {<slot>: [ids]}}and{signal: <id>}.plan.view_labels?: ApprovedText[].plan.initial_template?: <template id>: the template before a purpose locks (113).
- Template items gain:
Builder (
src/policy/build.ts):- It validates each new shape when present: every named signal, slot, stage, template and must-say id exists, else it throws naming the template and the item.
- The rule at
:480-481is relaxed:not_needed_reasonis required whenappliesis notalways, and allowed when it is.
Text identity (
assertTexts,:242):- An item's own texts keep the scheme
plan.template.<purpose>.<item>.{title,hint,not_needed}, plus.known_notewhen present. - A field filled by
texts_from: Tkeeps its canonical id,plan.template.T.<item>.<field>.assertTextsaccepts that id for exactly the inherited fields, and only whenT's item of the same id holds the identical text. plan.view_labelsids must be one ofplan.head.title,plan.head.subtitle,plan.known.{signup,call,demo},plan.last_suggestionandplan.concern.raised_again.- The duplicate check counts an inherited text once. The same id with different content is still rejected.
- An item's own texts keep the scheme
Text registry (
src/policy/text.ts).textEntriesyieldsknown_noteandplan.view_labelsentries (kindplan), inherited texts once. The loader's approval check (unapproved_text) andlookupTextthen cover them.Typed readers, so the typecheck stays green in this commit:
src/engine/plan.ts:422treats a missingdone_whenas neverdone.web/src/review-plan.ts:423mustSayOfskips an item withoutdone_when.- The
appliesevaluator:planApplies(src/engine/plan.ts:136) today narrowswhen_slot, then passes the rest tochecklistApplies. It handlesslot_in(withor_unknownandunknown_unless) andsignalbefore that call, and the hero rule (src/engine/hero.ts:40) accepts the widened union. SessionState.signals_seen: Record<string, number>records the firstiat which each signal was ≥signal_onon a client turn. A checkpoint without it reads as{}.
Tests.
tests/unit/policy-plan.test.ts: each invalid shape is rejected, naming the item;texts_fromfills only the missing fields and keeps canonical ids; the real mapping (the walkthrough'sexplain_who_holds_the_moneyinheriting the first-call title and hint) builds.tests/unit/policy-build.test.ts: a.known_noteon a fact item builds; the committed fixtures build byte-identically.tests/unit/policy-loader.test.ts: a draftview_labelsentry and a draftknown_noteare rejected (unapproved_text).tests/unit/rep-facing-text.test.ts: approved ones resolve throughlookupText.tests/unit/web-review-plan.test.ts: a judgement item withoutdone_whenrenders without error.tests/unit/engine-hero.test.ts: each newappliesform holds and fails as specified.tests/unit/engine-step.test.ts:signals_seenis recorded, and an old checkpoint rehydrates.
npx vitest run tests/unit/policy-plan.test.ts tests/unit/policy-build.test.ts tests/unit/policy-loader.test.ts tests/unit/rep-facing-text.test.ts tests/unit/engine-plan.test.ts tests/unit/engine-hero.test.ts tests/unit/engine-step.test.ts tests/unit/web-review-plan.test.ts→ all passed;npm run typecheckexits 0.SAFETY:
policy/srcis unchanged (git diff --exit-code policy/src); no candidate, no deploy.INVARIANT: INV-COPILOT-003; INV-COPILOT-015; INV-COPILOT-001.
Dependencies: none
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-110f
Title: The M1b wire contract types and the
--knownprobe [INTEGRATION-CRITICAL]Description: As the engine and the probes, I want the new grammar read safely and the new wire fields defined before any policy data or view uses them, so that every later story can deploy without breaking the published bundle.
Acceptance Criteria:
Contract types (
src/engine/types.ts,src/session/protocol.ts; HOT, additive and optional):CallPlan.known,CallPlan.must_saysandCallPlan.dropped_hard;PlanItem.due_by: 'client';FactValue.sourceandFactChip.source, over'signup' | 'crm' | 'demo' | 'call';FactValue.replaced: {value, source}.
FactChip.atstays non-null. A seeded chip'satis{i: -1, t: 0}(113), and 114b renders no link for it.Probe flag.
scripts/lib/probe-flags/known.ts, plus one registration line inscripts/lib/probe-flags/index.ts(HOT), adds--known. It reconnects withload{resume:true}and prints, per decision ofhello.history: the chips' slots and sources, theplan.knownids, and any known id also inplan.items.Tests.
tests/unit/session-protocol.test.ts: the new optional fields round-trip, and a decision without them still parses.- The committed fixtures give plans deep-equal to today's.
npx vitest run tests/unit/session-protocol.test.ts tests/unit/engine-plan.test.ts tests/unit/engine-purity.test.ts→ all passed;npm run typecheckexits 0.[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/replay-ws.mjs --call 3303259297 --scenario customer_success --speed 20 --session m1b-110f-<yyyymmdd>→errors == 0, with decisions unchanged in shape.node scripts/ws-probe.mjs --call 3303259297 --session m1b-110f-<yyyymmdd> --resume --knownprints chips with sourcecalland an emptyplan.known. Quote it.
INVARIANT: INV-COPILOT-001; INV-COPILOT-005.
Dependencies: COPILOT-110a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-110g
Title: Generators and eval tooling read the new grammar: the red-pen sheet, the labels sheet, and L4's retired ids
Description: As the policy owner, I want the red-pen and labels sheets to describe kinds, the new
appliesforms, items withoutdone_when, known notes and view labels, and L4 to tolerate checklist ids the policy retires, so that regeneration and the gate keep working once 110d converts the data.Acceptance Criteria:
Red-pen sheet (
scripts/redpen.mjs):describeApplies(:61) describesslot_in(withor_unknownandunknown_unless) andsignal.describeDone(:80) handles a missingdone_when("a judgement call: not ticked"), and the sheet gains a kind column.- Before rendering, it resolves the source through the builder's own helpers, exported from
src/policy/build.ts:resolveTextsFromfills inherited titles and hints, since:108readsitem.title.text;deriveChecklist(110h) replaceschecklist: "derived", since:285callschecklist.map.
- It still throws on a truly unknown form.
Labels sheet (
scripts/labels-sheet.mjs) listsknown_noteandplan.view_labelstexts with their approval state.L4.
scripts/lib/eval-l4.tspartitionseval/labelled/must-say.jsonagainst the evaluated checklist. A labelled id the policy no longer checks is reported underretired, left out of recall and precision, and never exits 2.must-say.jsonis not edited.Tests.
tests/unit/redpen.test.ts(new) renders a fixture clone converted to 110d's real shapes: every newappliesform, a judgement item withoutdone_when, the walkthrough item withtexts_fromand no own title, andchecklist: "derived". It asserts the inherited title and the derived checklist ids appear. An unknown form throws.tests/unit/eval-l4l6.test.ts: a retired id.node scripts/redpen.mjs --checkandnode scripts/labels-sheet.mjs --checkstill exit 0 on the committed sources.
npx vitest run tests/unit/redpen.test.ts tests/unit/eval-l4l6.test.ts→ all passed.
Dependencies: COPILOT-110a, COPILOT-110h
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-110h
Title: The derived must-say checklist per template, opt-in, and the hero reading it
Description: As the policy owner, I want the must-say checklist derived from each template's must-say items instead of a hand list, so that the checklist and the plan cannot drift. The derivation is opt-in per bundle, so no policy data changes here.
Acceptance Criteria:
Opt-in.
src/policy/build.ts: withchecklist: "derived"inrules-*.json, the builder deriveschecklist_by_template[template]: that template's items withkind: must_say(or, without kinds, adone_whenmust-say), with theirapplies, plusrate_pressure_handled. The bundle'schecklistis the union by id. Without"derived", the hand list is used, as today.Hero.
src/engine/hero.tsreadschecklist_by_templatefor the plan's current template (templateFor, including 113's initial template) when present, elsechecklist.Hash ownership and validation.
src/policy/canonical.tsHASH_FIELDS.weightsgainschecklist_by_template.- The loader (
src/policy/loader.ts) validates it: every key is a template id; every entry is a must-say of that template, or israte_pressure_handled, the one scenario-wide entry the derivation adds to every template with its applicability (when_fact: rate_pressure_raised) and completion rules unchanged; and the union equalschecklist. deriveChecklistis exported for 110g.
Tests.
tests/unit/policy-build.test.tsuses a fixture clone with"derived": it yields per-template entries and their union, andbank_hashandplaybook_hashare unchanged.tests/unit/engine-hero.test.ts: the hero reads the current template's entries; an old bundle readschecklist.tests/unit/policy-loader.test.tsandtests/unit/policy-canonical.test.ts: mutatingchecklist_by_templatechangesweights_hash; a bundle whose stored hash was not recomputed is refused (hash_mismatch); an entry that is not a must-say of its template (other thanrate_pressure_handled), or a union that differs fromchecklist, is refused. A real derived bundle, built from a converted fixture clone withrate_pressure_handled, passes both the builder and the loader.
npx vitest run tests/unit/policy-build.test.ts tests/unit/engine-hero.test.ts tests/unit/policy-loader.test.ts tests/unit/policy-canonical.test.ts→ all passed;npm run typecheckexits 0.SAFETY:
policy/srcunchanged.INVARIANT: INV-COPILOT-001; INV-COPILOT-002.
Dependencies: COPILOT-110a, COPILOT-110f
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-110c
Title: Apply Stevan's approved 2026-09-27 wording to
policy/src; drafts and homeless entries are skipped and reported [INTEGRATION-CRITICAL]Description: As the policy owner, I want the titles, hints, notes and new labels that Stevan approved as drafted (decision 2f;
policy/proposals/2026-09-27.json) applied in one step, so that the guide reads in his words and no unapproved text enters a build.Acceptance Criteria:
The tool.
scripts/apply-label-approvals.mjs <file>reads the proposals format:{source, decision, applies_to, texts[]}, where each text hastext_id,text,status,approved_byandapproved_at.- It first classifies every entry:
- invalid: malformed, an approver not in
policy/approvers.json, or an id with no known home. Any invalid entry makes it exit 1, write nothing, and name them. - draft:
statusother thanapproved. Reported asskipped (draft)and never applied. - deferred: approved, but its home is created by a later step. The known deferred homes are
plan.fact.bank_country.templateuntil 110b adds the slot, andsituation.address_ties_me_to_tax_residency.*until its point is approved (decision 13). Reported asdeferred; not an error. - applicable: everything else.
- invalid: malformed, an approver not in
- It then writes every applicable entry. A re-run is idempotent, and applies deferred entries whose home now exists.
--dry-runprints the classification and the diff.
- It first classifies every entry:
Homes.
- An existing title or hint id is reworded and stamped from the entry.
…<item>.known_noteon a fact item becomes itsknown_note, under the same id (110a registers it).- On a must-say item it becomes the
not_needed_reason, under the id the builder requires,plan.template.<purpose>.<item>.not_needed: an existing reason is reworded, and a missing one is created. The text,status,approved_byandapproved_atare copied from the entry. The id mapping (known_note→not_needed) is printed, and recorded onpolicy/LABELS.md. - 110a allows a reason while
appliesisalways. plan.head.title,plan.head.subtitle,plan.known.{signup,call,demo},plan.last_suggestionandplan.concern.raised_againjoinplan.view_labelsin both plan files.plan.fact.bank_country.templatehas no home until 110b adds the slot.situation.address_ties_me_to_tax_residency.{headline,hint}has no home while the situation's point is a draft:assertSituationsrequirescore_text_idamong its points (src/policy/build.ts:587). Its draft.pointis skipped [DECISION 13].
Test.
tests/unit/label-approvals.test.tsbuilds files inline on a temporary copy ofpolicy/src, and covers:- each home, including both
known_notemappings (a fact's own id; a must-say's.not_neededid); - the real mix of approved and draft entries (the draft reported, never applied);
- a deferred entry reported, then applied on a re-run once its home exists;
- an unknown id → invalid → exit 1 with nothing written;
- an approver outside the allowlist, which exits 1 and writes nothing.
- each home, including both
Apply the real file. Run
node scripts/apply-label-approvals.mjs policy/proposals/2026-09-27.json. It exits 0 and prints the[VERIFY]point as a draft, and the two situation texts and the bank-country template as deferred. Then runnode scripts/labels-sheet.mjsandnode scripts/redpen.mjs.node scripts/labels-sheet.mjs --checkexits 0, andgrep -c "Known from sign-up" policy/LABELS.mdcounts the two applied fact notes.tests/unit/label-approvals.test.tsasserts that the regenerated sheet lists every appliedknown_noteand view label.grep -rn "address_ties_me_to_tax_residency" policy/srcprints nothing.- Fixtures and stubs are regenerated.
npx vitest run tests/unit/label-approvals.test.ts tests/unit/policy-approved.test.ts tests/unit/policy-drafts.test.ts tests/unit/rep-facing-text.test.ts tests/unit/policy-build.test.ts→ all passed.[INTEGRATION-CRITICAL] Candidate and gate
--run m1b-wording-<yyyymmdd>(the weights and playbook hashes change): L1 ≥ 90 % for both, L4 PASS, L6 PASS,evaluated=true. Publish.SAFETY: the story never edits
policy/proposals/2026-09-27.jsonorpolicy/approvals/*, and never setsapproved_by; every stamp comes from Stevan's file.INVARIANT: INV-COPILOT-015; INV-COPILOT-003.
Dependencies: COPILOT-110a, COPILOT-110g, COPILOT-118a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, gate)
Notes: First in the policy chain; 118a's wall-clock pacing lands before any gate. The coordinator asked for everything but the
[VERIFY]line to be applied. The situation's approved headline and hint cannot enter without a point, so they wait with it (§7 Q2).ID: COPILOT-110d
Title: Typed templates: every item gets its kind and its rules (decisions 2c, 2d, 2e), and the checklist switches to derived [INTEGRATION-CRITICAL]
Description: As the policy owner, I want every plan item typed as a fact, a must-say or a judgement call, with Stevan's rules, and the must-say checklist derived from the must-say items, so that the plan can be a guide and the two lists cannot drift. No wording changes here.
Acceptance Criteria:
The ten
first_call_after_signupitems (policy/src/plan-onboarding.json; kinds from FEEDBACK §2d):Item Kind Rule open_with_agendajudgement surface_when {stage: opening}understand_the_transferfact slots [transfer_purpose, amount, frequency];surface_whenany of{unknown_after_stage: discovery},{source_conflict: transfer_purpose},{source_conflict: amount},{source_conflict: frequency}check_serviceability_firstfact slots [residency],rule: served;surface_whenany of{slot_outside_served: residency},{unknown_after_stage: discovery},{source_conflict: residency}confirm_funding_accountmust_say applies {slot_in: {funding_account: [third_party, company_account]}, or_unknown: true, unknown_unless: {transfer_purpose: [salary_or_pension]}}(decision 2d; §7 Q5);due {signal: client_ready_to_book}explain_who_holds_the_moneymust_say applies {signal: client_asked_about_safety},due {signal: client_asked_about_safety}: never listed on a first call where safety is not raised (decision 2c)explain_settlement_and_cutoffjudgement surface_whenany of{slot_in: {timing: [this_week, this_month]}},{signal: client_asked_about_timing}; nodone_whenexplain_booking_is_bindingmust_say applies always,due {signal: client_ready_to_book}close_the_document_gapmust_say appliesunchangedminimum_transfermust_say applies {when_amount_below_gbp: 5000}, unchanged: per relationship, not per leg (decision 2e)agree_next_stepfact facts [next_step_agreed]The
not_needed_reasonof the must-says comes from 110c's approved notes.The same rules in every template.
confirm_funding_accountcarries decision 2d's rule wherever it appears. The walkthrough item's missingnot_needed_reasoncomes throughtexts_from: first_call_after_signup, under 110a's identity rule.first_transfer_walkthroughgainsexplain_who_holds_the_money(decision 2c):texts_from: first_call_after_signup,applies always, anddueonclient_asked_about_safetyorclient_ready_to_book.explain_settlement_and_cutoffis judgement in every template.Every other item, across both plan files, takes the kind its
done_whenimplies:- a
must_saygives must_say, withappliesunchanged; - slots or facts only give fact, with
satisfied_byderived from those conditions; deliveredonly gives judgement withsurface_when {always: true}. Itsappliesis kept (customer successadd_user_accessandexplain_document_requeststay conditional), and 111a requires both.
The progress note lists every item with its kind and predicates.
- a
Initial template.
plan-onboarding.jsongainsinitial_template: first_call_after_signup. It takes effect once 113's reader lands, and until then it is ignored (§3 convention 5).The derived checklist.
rules-onboarding.jsonandrules-customer-success.jsonswitch tochecklist: "derived"(110h).- The onboarding union becomes
said_who_holds_funds,said_fund_from_own_account,said_booking_is_binding,said_minimum,said_docs_neededandrate_pressure_handled. - These leave it:
said_rate_transparency,said_quote_lifetime,said_settlement_timing,said_payment_reason,said_rm_contactandsaid_recording_disclosure. Their questions stay in the bank. - The customer-success union is quoted.
tests/unit/policy-build.test.tsasserts thatbank_hashandplaybook_hashare unchanged.
- The onboarding union becomes
Tests. Fixtures, stubs,
policy/REDPEN.mdandpolicy/LABELS.mdare regenerated with 110g's generators.tests/unit/engine-hero.test.ts:- who-holds is not applicable on a first call before the safety signal, and applicable on the walkthrough;
confirm_funding_accountdoes not apply with a salary purpose and an unknown funding account, on both templates;- it applies with
third_party; minimum_transferdoes not apply with NZD 25,000 and CAD 5,000 held.
tests/unit/policy-plan.test.ts: every mapped fact has a non-emptysatisfied_by, and every mapped judgement keeps its sourceapplies.
npx vitest run tests/unit/policy-plan.test.ts tests/unit/policy-build.test.ts tests/unit/policy-src.test.ts tests/unit/policy-approved.test.ts tests/unit/engine-hero.test.ts tests/unit/eval-l4l6.test.ts→ all passed;node scripts/redpen.mjs --check,npx vitest run tests/unitandnpm run typecheckexit 0.[INTEGRATION-CRITICAL] Candidate and gate
--run m1b-kinds-<yyyymmdd>: L1 ≥ 90 % for both,L4: PASSwithretiredquoted, L6 PASS,evaluated=true. Publish.SAFETY: no text changes; 110c applied the wording.
INVARIANT: INV-COPILOT-016; INV-COPILOT-013; INV-COPILOT-003.
Dependencies: COPILOT-110c, COPILOT-110f, COPILOT-110g, COPILOT-110h
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, gate)
Notes: 110f's readers deploy before this data (§3 convention 5).
ID: COPILOT-110b
Title: The
bank_countryslot, its approved chip template, and the served-country list (a bank change) [INTEGRATION-CRITICAL]Description: As the engine, I want to know the country of the bank the client names, and which countries the partners serve, so that "Where they live and bank" is satisfied from the call and surfaced only when a country is off the served list.
Acceptance Criteria:
Slot grammar (code):
src/policy/fact-slots.tsaddsbank_countryto the slot ids as an optional slot; a bundle without it is valid.- The builder accepts a country span slot with a
templatewhose{country}param is a country id rendered throughfact_lists.countries[].display(as situation points render{country}), instead of per-country labels. - Its candidate trigger is
\b(bank|banks|banking)\bover the gazetteer. parsePrecall(112) acceptsbank_countryas a country id when the policy has the slot.web/src/render.tstemplateText(:71) renders a chip template's{country}param throughfact_lists.countries[].display, as its point branch does (:87).web/src/chips.ts(HOT, first) gives each chipdata-slot.docs/design/DESIGN.md§Client facts chips adds the slot ("Banks in Saudi Arabia", in the transfer group after residency).
Data.
policy/src/plan-onboarding.jsonfact_slotsgainsbank_country:source {kind: span, question: span_bank_country, over: countries}, with noknown_fact, asresidencyhas none.policy/src/onboarding.jsongainsspan_bank_country: "Which place is the country of the bank the client uses? …". It is Jev-facing, so it needs no red-pen.fact_lists.served_countries: []is added (underbank_hash, empty until supplied; §7 Q3).check_serviceability_first.satisfied_by.served_checkbecomes[residency, bank_country], and itssurface_whengains{slot_outside_served: bank_country}.- Re-run
node scripts/apply-label-approvals.mjs policy/proposals/2026-09-27.json. The approvedplan.fact.bank_country.template("Banks in {country}") now finds its home, and the other entries are no-ops.
Tests.
tests/unit/policy-build.test.ts: an old bundle (no slot) builds and loads; the customer-success bundle is unaffected.tests/unit/engine-plan.test.ts: an empty list, and['GB']with residency or bank countrySA(the card surfaces).tests/unit/engine-fact-values.test.ts: the slot, and its chip as a typed template.tests/unit/web-chips.test.ts: the chip renders "Banks in Saudi Arabia".tests/unit/precall.test.ts: a round trip withbank_country: 'SA'.
npx vitest run tests/unit/policy-build.test.ts tests/unit/engine-plan.test.ts tests/unit/engine-fact-values.test.ts tests/unit/precall.test.ts tests/unit/web-chips.test.ts→ all passed.[UI] Design artifact: DESIGN.md §Client facts chips.
[INTEGRATION-CRITICAL] Candidate and gate. Deploy before the upload (§3 convention 4):
node scripts/probe-questions.mjs --call 3339895706 --at 38 --planned(the draft bank) plansspan_bank_countryand answers the Saudi candidate at ≥ 0.6. Quote it.- The gate run
--run m1b-bank-country-<yyyymmdd>: L1 onboarding and cs ≥ 90 %. Publish. - On the published policy,
node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-110b-light","?call=3339895706","light",1440,900],["COPILOT-110b-dark","?call=3339895706","dark",1440,900]]' --deployed --until '[data-slot="bank_country"]'exits 0.
INVARIANT: INV-COPILOT-015; INV-COPILOT-016; INV-COPILOT-020.
Dependencies: COPILOT-102a, COPILOT-108b, COPILOT-110d, COPILOT-111a, COPILOT-112
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, gate)
Notes: 111a implements the served rule that the tests exercise; 108b is the policy-chain order.
ID: COPILOT-110e
Title: L1 can check an engine fact value from a span question; the bank-country case [INTEGRATION-CRITICAL]
Description: As the policy owner, I want an L1 case to pin a fact the engine derives from a span Choice, so that "Saudi National Bank" giving a bank country of
SAis part of the gate.Acceptance Criteria:
- The check.
scripts/lib/eval-l1.tsgains the engine checkfact_value: {slot, value}. It runsstep()on the case's answers and compares the slot's value. When a case names a span slot, the L1 request includes that slot's span Choice, built from the case's window candidates asprobe-questions --questions span_<slot>builds it.eval/schema.jsonandscripts/lib/eval-cases.tsaccept the check.tests/unit/eval-l1.test.tscovers a pass and a fail. - The case.
eval/labelled/onboarding.jsongainsob_3339895706_i38_bank_country: context v1 lines 34-37 and latest line 38 ("Which is Saudi National Bank."), verbatim from the v1 fixture, checkingfact_value {slot: bank_country, value: SA}.npx vitest run tests/unit/eval-l1.test.ts tests/unit/eval-cases.test.ts→ all passed. - [INTEGRATION-CRITICAL] On the deployed Worker (
health.version== HEAD), with the published version from 110b:node scripts/eval.mjs --scenario onboarding --layers l1 --run m1b-fact-value-<yyyymmdd>passes the new case (its answer quoted), with L1 onboarding ≥ 90 %. - INVARIANT: INV-COPILOT-004 (the case text passes the production redactor).
- The check.
Dependencies: COPILOT-110b
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, eval)
ID: COPILOT-111a
Title: Engine: fact and judgement semantics, silent facts,
plan.known, chips with their source [INTEGRATION-CRITICAL]Description: As the engine, I want facts to tick silently from any source into a known list with their source, and judgement items to exist only while triggered, so that the live list holds only what helps right now.
Acceptance Criteria:
src/engine/plan.ts(HOT, after 117), for items with akind; others behave as in M1.- fact.
- Satisfied when every
satisfied_by.slotsslot is known (a call value at ≥fact_value_min_conf, or a seeded value at i=-1), everysatisfied_by.factsfact is held, and, forrule: served, every knownserved_checkslot is on the served list. An empty or absent list passes. - A satisfied fact item is not in
items. It goes intoplan.known: {id, source, at: {i, t} | null}[], wheresourcecomes from the slot that completed it andatisnullat i=-1. - An unsatisfied fact item has no row.
- A fact item is the current card only while a
surface_whenclause holds. While surfaced it leavesplan.knownand is initemsas current. When the clause stops holding, its satisfaction is evaluated again: it returns toknownonly ifsatisfied_byholds, else it has no row. The probe below therefore holds even for a surfaced fact.unknown_after_stage:stage.currentis known (non-null, and notnone) and is neitheropeningnor the named stage, and a slot is still unknown;slot_outside_served: a known slot is not on a non-empty served list;source_conflict: the slot'sFactValue.conflictis true (set and cleared by 113).
- Satisfied when every
- judgement. No row until both its
appliesand itssurface_whenhold. It then competes for the current card under the existing challenger rules. It is neverdoneornot_needed, and it leaves when either stops holding. - One card. A surfaced fact or judgement is a candidate in the existing ranking, never current by fiat. The existing rules (a challenger leads for two decisions; at most one change of the current item per decision) keep exactly one current item. Among several surfaced candidates the earliest in template order ranks first; the others stay hidden until it resolves.
- Neither kind counts in the hero.
- fact.
factChips(src/engine/fact-values.ts:381) passesFactValue.sourcethrough (absent meanscall).Tests.
tests/unit/engine-plan.test.tson the v1 fixture of 3339895706, with its stub answers and 110d's policy fixture:open_with_agendashows no row once the stage has leftopening;check_serviceability_firstis inknown(sourcecall) from the decision where residency becomes known, and never initemsunless surfaced;explain_settlement_and_cutoffhas no row until its trigger fires;- a synthetic state with
conflictset surfacescheck_serviceability_first, which is then not inknown; - facts never change the hero's completeness;
- in the initial state (
stage.currentnull), and after anonestage answer, no fact surfaces onunknown_after_stage; - with
understand_the_transferandcheck_serviceability_firstboth surfaced, exactly one item is current on every decision; - customer success
add_user_access(a judgement withappliesondecision_maker) has no row whiledecision_makeriscaller, andexplain_document_requestnone while documents arecomplete; - an unsatisfied mapped fact (customer success
review_upcoming_exposure, from 110d's derivedsatisfied_by) has no row and is not inknownuntil its conditions hold; - a composite fact (
understand_the_transfer) that surfaced on a conflict, and whose conflict then clears whilefrequencyis still unknown, returns to no row, not toknown.
npx vitest run tests/unit/engine-plan.test.ts tests/unit/session-protocol.test.ts tests/unit/engine-purity.test.ts tests/unit/engine-recompute.test.ts→ all passed;npm run typecheckexits 0.[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD, with 110d's published kinds):node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-111a-<yyyymmdd>→errors == 0.node scripts/ws-probe.mjs --call 3339895706 --session m1b-111a-<yyyymmdd> --resume --knownshows no decision with a known id that is also inplan.items. Quote it.
INVARIANT: INV-COPILOT-016; INV-COPILOT-001.
Dependencies: COPILOT-110d, COPILOT-117
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
Notes: 117 is the hot-file order for
src/engine/plan.ts.ID: COPILOT-111b
Title: Engine: must-say
appliesanddue, no not-needed rows in the live plan, the must-say record for the review [INTEGRATION-CRITICAL]Description: As the engine, I want must-says listed only while they apply and surfaced when due, and the live plan free of not-needed rows, so that the live list is a guide and the review keeps the score.
Acceptance Criteria:
src/engine/plan.ts:- A
must_sayitem is initemsonly while itsappliesholds. It islateruntil itsduefires (with nodue, it is due once it applies), and it may then become current. - When
duefired on a client signal, the item carriesdue_by: 'client', which 107's presenter inserts without the dwell. - The live
decision.plan.itemsholds nonot_neededitem. plan.must_says: {id, applies, why, applied_at, due_at, due_by, said_at}[]covers every must-say item of the plan's template, with{i, t}times.whyis the rule kind:always,signal,slot_in,when_amount_below_gbporwhen_slot.appliesis the value at this decision.due_atisnullwhile the item'sduehas not fired. An applying must-say whose due never fires (for example booking-is-binding on a call with noclient_ready_to_book) ends the call withdue_at: null.
- A
Tests.
tests/unit/engine-plan.test.tson the v1 fixture of 3339895706, with the stub answers and these overrides (listed in the test):said_docs_needed0.9 at i=46,said_booking_is_binding0.9 at i=89,client_asked_about_safety0.9 at i=107 (08:50), andfact_transfer_purposesalary_or_pensionat i=71.- At the last decision the live
itemshas ≤ 4 open rows and nonot_needed. explain_who_holds_the_moneyfirst applies, and becomes due, at i=107 withdue_by: 'client'(decision 2c).confirm_funding_accountdoes not apply at the end (decision 2d).minimum_transferdoes not apply (decision 2e).must_saysends with 3 applying, 2 of them said.- On a synthetic call with no
client_ready_to_bookand nothing said,explain_booking_is_bindingends withapplies: true,due_at: null.
tests/unit/web-presenter.test.ts: the who-holds item opens without waiting out the dwell.npx vitest run tests/unit/engine-plan.test.ts tests/unit/engine-hero.test.ts tests/unit/web-presenter.test.ts tests/unit/session-protocol.test.ts→ all passed;npm run typecheckexits 0.- At the last decision the live
[UI] Design artifact: DESIGN.md §Cadence rules rule 2 (drawn by 107) and §The current item.
[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-111b-<yyyymmdd>→errors == 0.node scripts/with-cf-env.mjs npx wrangler d1 execute copilot --remote --command "SELECT COUNT(*) AS n FROM decisions d, json_each(json_extract(d.decision_json,'$.plan.items')) e WHERE d.session_id = '<session_id>' AND json_extract(e.value,'$.state') = 'not_needed'"→n: 0. Quote the last decision'splan.must_says.node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-111b-light","?call=3339895706","light",1440,900],["COPILOT-111b-dark","?call=3339895706","dark",1440,900]]' --deployed --until '[data-id="explain_who_holds_the_money"][aria-current="step"]'exits 0.
If the item never opened, the story quotes
client_asked_about_safetyat the parts around 08:50 and stayspasses:false: decision 2c is then not met live.INVARIANT: INV-COPILOT-016; INV-COPILOT-001.
Dependencies: COPILOT-107, COPILOT-111a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, Jev-dependent probe)
ID: COPILOT-111c
Title: Amounts: a mid-call correction replaces the amount, a second currency leg is kept, and the minimum stays per relationship (decision 2e) [INTEGRATION-CRITICAL]
Description: As the engine, I want "fifty thousand, actually fifteen hundred" read as one corrected amount, and a second currency leg read as a second amount, so that the minimum rule (per client relationship, not per leg) is neither cleared by a figure the client took back nor re-opened by a small second leg.
Acceptance Criteria:
Amount history (
src/engine/amount.ts,src/engine/types.ts; HOT).SessionStategainsamount_history: {gbp, code, digits, i, superseded_by?}[]next to today'samounts: number[], which is kept and derived as the GBP of the non-superseded entries, so old checkpoints read as before.recordAmounts(:362) appends each client amount (and each client-confirmed rep amount) with its currency.- An amount without a currency takes the currency of the amount it corrects.
- A client amount in the same currency as the latest held one supersedes it when stated in the same client run, or in the next client run, with a correction marker ("actually", "I mean", "sorry", "no,", "rather", "correction"). A client run is the conversational unit: consecutive client
turnparts, skipping backchannels and acks, as 105'sisClientRunEnddefines it. It is never a single v2 part. - With a correction marker, a new client amount supersedes the latest held amount whatever its currency, in the same client run or the next one. "Fifty thousand euros, sorry, I meant fifteen hundred dollars" leaves USD 1,500 only.
- Without a marker, a different currency is a second leg and is kept, and a same-currency amount in a later run is kept (a second transfer).
- The rule is written at the top of
src/engine/amount.ts: a marker means a correction; no marker means an addition. - A seeded entry (113, i=-1) is superseded by the first call amount in its currency, at any distance and without a marker: the call's own figure replaces sign-up's. Call-to-call supersession follows the correction rule above.
Reconcile with the fact.
clientAmountsGbp(:385) reads the non-superseded history, plus theamountfact value only when it matches a non-superseded entry or no history exists. When the amount span question was dropped by the budget, the history still holds the text amounts.Unchanged.
when_amount_below_gbp(src/engine/hero.ts:40): any client amount at or above the threshold clears the minimum. There is no per-leg rule and no batching line.Tests.
tests/unit/engine-features.test.ts(theextractAmountsblock) andtests/unit/engine-hero.test.ts:- "fifty thousand euros, actually fifteen hundred" holds EUR 1,500 only, so the minimum applies;
- "fifty thousand euros, sorry, I meant fifteen hundred dollars" holds USD 1,500 only, so the minimum applies;
- "fifty thousand euros and about five thousand Canadian a month" (no marker) holds both;
- v1 i=13 (NZD 30,000) then i=33 ("about 5,000 Canadian every month") holds both, so the minimum is not needed;
- a later same-currency amount without a marker is kept;
- a seeded NZD 25,000, followed several turns later by the call's "about 2,000 NZD", supersedes the seed, so the minimum applies;
- a dropped span answer is covered;
- an old checkpoint with
amountsbut noamount_historymigrates on first read: each legacy GBP value becomes an entry withcode: null, kept and never superseded. Rehydrate it, append another transfer, and the minimum is still evaluated over both; - a correction spanning three v2 parts of one client run supersedes;
- the 058d cases pass.
npx vitest run tests/unit/engine-features.test.ts tests/unit/engine-hero.test.ts tests/unit/engine-fact-values.test.ts tests/unit/engine-state.test.ts→ all passed.[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-111c-<yyyymmdd>→errors == 0. The last decision'splan.must_saysentry forminimum_transferreadsapplies: falseafter the 02:41 CAD leg. Quote it.INVARIANT: INV-COPILOT-001; INV-COPILOT-002.
Dependencies: COPILOT-111b
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-112
Title: Pre-call facts: the
0004_precall.sqlmigration file, closed-vocabulary validation and the demo payload toolDescription: As Stevan, I want the shape of pre-call facts fixed and validated in one place, so that every write path accepts only closed-list values with their source, and the demo seed is generated rather than hand-typed (decision 2a).
Acceptance Criteria:
Migration file.
migrations/0004_precall.sql(0001-0003 exist) isALTER TABLE calls ADD COLUMN precall_json TEXT; ALTER TABLE sessions ADD COLUMN precall_json TEXT;.tests/workers/migrations.test.tsapplies it in miniflare and lists both columns.Validation.
src/ingest/precall.tsparsePrecall(json, policy): Precall:- A key is accepted only when it is a fact slot of the pinned onboarding policy with a closed label list:
client_type,residency,transfer_purpose,frequency,documents_status,current_provider,funding_account,timing, andbank_countryas a country id fromfact_lists.countrieswhen the pinned policy has that slot (110b). - Its value must be one of that slot's label ids.
- Also accepted:
currency_pairas two ISO 4217 codes fromfact_lists.currencies, andamountas{digits, code}. sourcein {signup, crm, demo} is required.- Anything else raises
invalid_precallnaming the field. There is no free-text field.
- A key is accepted only when it is a fact slot of the pinned onboarding policy with a closed label list:
Import payload.
scripts/lib/import-payload.schema.jsonandsrc/ingest/import-payload.tsgain an optionalprecall, validated byparsePrecall.Demo payload tool.
scripts/precall-payload.mjs --call <id> --source demoprints the demo payload from a table keyed by call id, validated againstpolicy/src. For 3339895706 it is{"source":"demo","client_type":"personal","residency":"SA","currency_pair":["SAR","NZD"],"transfer_purpose":"salary_or_pension","frequency":"monthly","documents_status":"proof_of_address_outstanding","amount":{"digits":"25000","code":"NZD"}}: the call's own values, with the amount at the low end of line 13's 25-30k NZD. Itspd_ct_idis NULL, hencedemo.Tests.
tests/unit/precall.test.ts: every key, and"transfer_purpose":"salary"→invalid_precall.tests/unit/pseudonymised.test.ts: a name-shaped string in any field is rejected.
npx vitest run tests/unit/precall.test.ts tests/unit/pseudonymised.test.tsandnpx vitest run -c vitest.workers.config.ts tests/workers/migrations.test.ts→ all passed;npm run typecheckexits 0.INVARIANT: INV-COPILOT-020; INV-COPILOT-004.
Dependencies: none
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-112b
Title: Pre-call facts: the write paths (a service-token PUT and the import field), provenance on every path, preservation on replacement, the migration applied, the demo seed written [INTEGRATION-CRITICAL]
Description: As Stevan, I want pre-call facts written only by trusted callers with an honest source, and kept when a transcript is replaced, so that the demo seed and the future sign-up join cannot be spoofed or silently lost.
Acceptance Criteria:
PUT route.
src/routes/precall.tsaddsPUT /api/calls/:id/precall, registered by one dispatch line insrc/index.ts(HOT). It takes the service token only (an email actor gets 403), acceptssourcedemoorcrmonly (signupgets 422, reserved for the M2 join), and is auditedcall.precall_set.Import path (
src/routes/calls-import.ts; HOT, after 104b), same rules:- A payload
precallis accepted only from the service token (an email actor gets 403), and only withsourcedemoorcrm. - A forced replacement without
precallkeeps the stored value (COALESCE); one with it replaces it.
- A payload
Tests.
tests/workers/precall.test.tscovers:- a PUT, then a forced transcript replacement without
precall, leavingcalls.precall_jsonunchanged; signuprefused on both paths;- an email actor refused on both paths;
"residency":"Narnia"→ 400 namingresidency;- the audit row.
npx vitest run -c vitest.workers.config.ts tests/workers/precall.test.ts tests/workers/calls-import.test.ts→ all passed.- a PUT, then a forced transcript replacement without
[INTEGRATION-CRITICAL] Live probe:
node scripts/with-cf-env.mjs npx wrangler d1 migrations apply copilot --remoteapplies0004_precall.sql.- Deploy (
health.version== HEAD). - In one shell,
set -a; . ~/.config/jev/cloudflare.env; set +a, thencurl -s -X PUT "$WORKER_URL/api/calls/3339895706/precall" -H "CF-Access-Client-Id: $CF_ACCESS_CLIENT_ID" -H "CF-Access-Client-Secret: $CF_ACCESS_CLIENT_SECRET" --data "$(node scripts/precall-payload.mjs --call 3339895706 --source demo)"→ 200. - The same with
"residency":"Narnia"→ 400, and with"source":"signup"→ 422. - Quote
SELECT precall_json FROM calls WHERE call_id='3339895706'and thecall.precall_setaudit row.
INVARIANT: INV-COPILOT-020; INV-COPILOT-007; INV-COPILOT-004.
Dependencies: COPILOT-104b, COPILOT-104e, COPILOT-112
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
Notes: It runs after 104e so the live seed is written to the v2 transcript's call row. This story owns the
precall_jsoncolumn's preservation, because 104b runs before the column exists.ID: COPILOT-113
Title: Seeding from pre-call facts at evidence i=-1 (pure): conversion,
initialStateandrecomputeparameters, replacement provenance and an active conflict that resolvesDescription: As the engine, I want a pure conversion from the stored pre-call facts to engine state, and a record of when the call contradicts them, so that items satisfied before the call never nag and a contradicted fact surfaces once, not forever.
Acceptance Criteria:
Converting the precall.
src/ingest/precall.tsseedFromPrecall(precall, policy): Seed(pure) converts each value into the form the call path produces:- a slot label id as
FactValue.value; currency_pairas the stringpairLegsreads (for exampleSAR/NZD);amountas the stringamountValuereads (for example25000 NZD), plus a seededamount_historyentry (111c).
Each value gets
conf: 1, i: -1, t: 0, source, alongsideclient_typeandpersisted[known_fact] = {i: -1, value: 1}for slots that have aknown_fact.- a slot label id as
Initial template.
templateFor(src/engine/plan.ts:72) returnsplan.initial_template(110a) before a purpose locks, instead ofother. When the field is absent, it returnsother, as today. 110d setsplan.initial_template: first_call_after_signupfor onboarding, and customer success leaves it absent. So the first decision's plan, before the purpose Choice locks (the greeting is low-confidence), already holds the first-call fact items.Applying it.
src/engine/state.tsinitialState(pin, seed?)applies the seed.src/engine/recompute.tsrecompute(…, seed?)takes it as a parameter and stays pure. A seeded chip'satis{i: -1, t: 0}, soFactChip.atstays non-null; 114b renders no link fori < 0.Provenance and conflict.
- A confident different call value replaces a seeded one: the chip reads "updated" and the
sourcebecomescall. FactValue.replaced: {value, source}keeps the seeded value as history only.- An active
FactValue.conflict: trueis set at the same time. It clears on engine time: once the fact's item has been current continuously for at leastthresholds.conflict_hold_sof call time (optional, 20 s, not less thanLIVE_MIN_DWELL), and a client run has ended and a rep-turn decision has arrived since it became current. - Limitation: the hold is measured on engine time, so when the presenter defers the card (mid-run, or during the dwell), the time the rep actually sees it may be shorter than 20 s. §6 risk 21 records this.
- A new confident value that agrees with the seed also clears it.
source_conflict(111a) readsconflict, notreplaced, so a resolved fact stops resurfacing.
- A confident different call value replaces a seeded one: the chip reads "updated" and the
Tests.
tests/unit/precall.test.tsround-trips each slot throughfactChipsandclientAmountsGbp.tests/unit/engine-state.test.tsandtests/unit/engine-plan.test.ts, with 3339895706's demo payload:- at i=0
plan.knownholdscheck_serviceability_firstandunderstand_the_transfer(sourcedemo); plan.factshas 7 chips with sourcedemoandat.i-1;- a later confident
residencyofGBsetsreplacedandconflict, and surfacescheck_serviceability_firstonce; - over consecutive mid-run client decisions,
conflictstays true; - after 20 s current, a client run end and a rep-turn decision, it clears;
- once cleared, it does not surface again;
- with a synthetic
plan.initial_templateand a low-confidence greeting, the first decision'splan.knownholds both fact items; with the field absent, the template isother, as today.
- at i=0
tests/unit/engine-recompute.test.ts: zero drift with the same seed.
npx vitest run tests/unit/engine-*.test.ts tests/unit/precall.test.ts→ all passed;npm run typecheckexits 0.INVARIANT: INV-COPILOT-020; INV-COPILOT-001; INV-COPILOT-016.
Dependencies: COPILOT-111c, COPILOT-112
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-113c
Title: Pre-call seed through every session path: the snapshot at load, seek, rehydrate and
set_weights[INTEGRATION-CRITICAL]Description: As the Durable Object, I want every session to take its pre-call seed once and carry it through every path that rebuilds state, so that a replay, a seek, a rehydrate and a recompute all see the same start.
Acceptance Criteria:
Snapshot at load.
src/session/handlers/load.tscopiescalls.precall_jsonand its seed (113) intosessions.precall_json({precall, seed}) and into the DO meta, once, at session start. A later PUT never changes an existing session.Every rebuild path passes
meta.seed:load.ts:131;seek.ts:55;src/session/checkpoint.ts:254(the rehydrate with no checkpoint);src/session/handlers/set_weights.ts:61(intorecompute).
Tests (
tests/workers/session-load.test.ts,tests/workers/session-seek.test.ts,tests/workers/session-recompute.test.ts,tests/workers/session-checkpoint.test.ts):- the snapshot is copied once;
- a later PUT leaves it unchanged;
- a rehydrate after an eviction, a seek and a
set_weightsrecompute all keep the seed (the first decision's 7 demo chips).
npx vitest run -c vitest.workers.config.ts tests/workers/session-load.test.ts tests/workers/session-seek.test.ts tests/workers/session-recompute.test.ts tests/workers/session-checkpoint.test.ts→ all passed;npm run typecheckexits 0.[INTEGRATION-CRITICAL] Live probe, after 112b's PUT, on the deployed Worker (
health.version== HEAD):node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-113c-<yyyymmdd>→errors == 0.node scripts/ws-probe.mjs --call 3339895706 --session m1b-113c-<yyyymmdd> --resume --known: the first decision shows 7 chips sourceddemo, andcheck_serviceability_firstandunderstand_the_transferknown and not inplan.items. Quote it.
INVARIANT: INV-COPILOT-020; INV-COPILOT-009; INV-COPILOT-004.
Dependencies: COPILOT-112b, COPILOT-113
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-113b
Title: The served-list gate on the jurisdiction concern
Description: As the engine, I want the jurisdiction concern to stay closed while the client's country is known and on the partners' served list, and never to be silenced by "known" alone, so that the "we can't serve you" exit is kept.
Acceptance Criteria:
src/engine/concern.tsupdateConcern(HOT, after 117): a raise typedjurisdictionis ignored whileresidencyis known and infact_lists.served_countries. The gate does nothing while the list is absent or empty, as it is today.Tests (
tests/unit/engine-concern.test.ts, with fixture banks):served_countries: ['SA']with residencySA: a jurisdiction raise opens nothing.['GB']with residencySA: a raise atclient_objecting0.9 typedjurisdictionopens an episode.- An empty list: it opens.
- Residency unknown with
['SA']: it opens.
npx vitest run tests/unit/engine-concern.test.ts tests/unit/engine-purity.test.ts→ all passed.SAFETY: the gate is never keyed on "known" alone;
served_countriesinpolicy/srcstays empty (§7 Q3).INVARIANT: INV-COPILOT-001; INV-COPILOT-017.
Dependencies: COPILOT-110b, COPILOT-117
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration
Notes: Inert on the live call until a served list exists, so it needs no live probe; 118's register records it as such.
ID: COPILOT-114
Title: Live view as a guide: Now, Coming up, collapsed Covered, no tallies, the approved head [INTEGRATION-CRITICAL]
Description: As a rep, I want the plan to read as a guide (one card, up to three things coming up, and what I have covered in one line) so that it never looks like a score sheet (decision 2b).
Acceptance Criteria:
web/src/plan.tsandweb/src/shell.ts(HOT, after 107 and 117) lay the pane out in this order (decision 2b):- Head.
plan.head.title("Guide for this call") andplan.head.subtitle("Adapts as you talk. Use what fits.") from the pinned bundle'splan.view_labelswhen approved; otherwise "Plan for this call" and no subtitle. There is no tally. - Chips. The chips row, where it is today.
- Now (
data-section="now"): the current card, with 107's points. - Coming up (
data-section="coming-up"): at most 3 applying must-say rows in plan order (statesnext/later), grey, with no state word and no time. - Covered (
data-section="covered"): the done rows collapsed into one button line, "Covered:mm:ss · …", with aria-expanded, which expands them.
There is no "Not needed on this call" group. Fact and judgement items have no rows. "Now", "Coming up" and "Covered" are UI chrome from decision 2b (§3 convention 10).
- Head.
docs/design/DESIGN.md: §Plan pane and §Plan item rows are rewritten to match. §Not needed on this call moves to the review view. §Removed from the live view gains the tally and the not-needed group.Tests (
tests/unit/web-plan.test.ts):- the head with and without
plan.head.titleapproved; - no tally element;
- no not-needed group;
- at most 3 Coming up rows;
- Covered collapsed, and expandable by keyboard.
npx vitest run tests/unit/web-plan.test.ts tests/unit/web-dom-safety.test.ts→ all passed;npm run typecheckexits 0.- the head with and without
[UI] Design artifact: DESIGN.md §Plan pane, as rewritten; the recommended guide variant on the 121 board.
[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD), with the published policy carrying 110c's labels:- A cache-warming
replay-wsof 3339895706 →errors == 0. node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-114-0245-light","?call=3339895706","light",1440,900],["COPILOT-114-0245-dark","?call=3339895706","dark",1440,900]]' --deployed --at 02:45 --expect '[data-section="coming-up"]'exits 0.- The same at
--at 08:55, with namesCOPILOT-114-0855-{light,dark}and--expect '[data-section="covered"]', exits 0.
- A cache-warming
INVARIANT: INV-COPILOT-015; INV-COPILOT-016; INV-COPILOT-014.
Dependencies: COPILOT-102a, COPILOT-107, COPILOT-110c, COPILOT-111b, COPILOT-113c, COPILOT-121
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-114b
Title: Chips name their source, carry no link when seeded, and render countries by name [INTEGRATION-CRITICAL]
Description: As a rep, I want each fact chip to say where the fact came from, with no link for facts known before the call, so that I know what sign-up gave us, what the call established, and what was typed in for a demo (decision 2a).
Acceptance Criteria:
web/src/chips.ts(HOT, after 110b):- Each chip carries
data-sourceanddata-slot. - It links to
updated_at ?? at(src/engine/fact-values.ts:314keeps the original evidence and setsupdated_aton a replacement). It renders no link only when that evidence hasi < 0: a seeded value never corrected during the call. - It renders the source suffix only for the three approved sources, from the pinned bundle's
plan.view_labels:plan.known.signup"From sign-up",plan.known.call"From the call",plan.known.demo"Seeded for the demo". A missing label renders no suffix. - M1b has no "From CRM" label: a
crmchip renders without a suffix, and CRM provenance arrives with the M2 join. - A
demoorcrmchip never reads "From sign-up".
- Each chip carries
docs/design/DESIGN.md§Client facts chips gains:- "a chip seeded before the call has no time link and carries its source";
- the suffix style: 12px
--ink-3, after the value, within the chip.
Tests (
tests/unit/web-chips.test.ts):- signup, call and demo each get their suffix, and
crmgets none; - a missing label gives no suffix;
demonever shows "From sign-up";at.iof -1 with noupdated_atgives no link;- a seeded value later corrected by the call links to the correction;
- the 8-chip cap holds.
npx vitest run tests/unit/web-chips.test.ts tests/unit/web-dom-safety.test.ts→ all passed.- signup, call and demo each get their suffix, and
[UI] Design artifact: DESIGN.md §Client facts chips; the chips on the 121 board's guide variant.
[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD), with 112b's demo seed:node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-114b-light","?call=3339895706","light",1440,900],["COPILOT-114b-dark","?call=3339895706","dark",1440,900]]' --deployed --until '[data-source="demo"]' --expect '[data-source="demo"]'exits 0.INVARIANT: INV-COPILOT-018; INV-COPILOT-015.
Dependencies: COPILOT-102a, COPILOT-110b, COPILOT-110c, COPILOT-113c, COPILOT-114
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-115
Title: Review model: must-says against what applied, known facts by source, "covered before the guide suggested it", with a fallback for pre-M1b decisions
Description: As a rep reading my own call afterwards (decision 2g), I want the review's numbers computed from the decisions and the transcript by one pure function, so that the score is coaching, is reproducible, and still works for calls recorded before M1b.
Acceptance Criteria:
web/src/review-model.ts(new, pure)reviewModel(decisions, utterances, policy), whereutterancesare theGET /api/calls/:idrows (i, t, t_end, speaker, kind). It returns:- Must-says:
must_says: {counted, said, rows, not_due}. A must-say is counted when it was said, or when it applies at the call's end and itsdue_atis set. An applying must-say that never became due and was not said is not counted: it is listed innot_due("Not due on this call"), with no due link. An unsaid must-say is exempt only when it became due after the start of the rep's last run (the rep's final consecutiveturnparts, skipping backchannels), when the rep had no later opportunity to say it. A said must-say always counts. Each row carrieswhy,due_at, and eithersaid_atormissedwith the item's suggested approved point and its due moment. An item without points, such asminimum_transfer, uses its approved hint instead. - Known facts:
known: {n, m, by_source, notes}. Known fact items are counted against the template's fact items, with theirknown_notewhen present. - Ahead of the guide: must-says said before any decision showed them as the card.
- Legacy:
legacy: truewhen the decisions carry noplan.must_says(recorded before M1b).
- Must-says:
Tests (
tests/unit/review-model.test.ts):- On the v1 fixture of 3339895706 with 111b's overrides:
must_saysis{counted: 3, said: 2}, with who-holds missed and due at v1 i=107 (08:50);- ahead of the guide is 2 of 2 (documents at i=46, booking at i=89);
- judgement items appear in no count.
- A synthetic call where a must-say becomes due during the rep's final run, which spans several v2 parts, does not count it unsaid.
- One due before that run counts, as does one said in it.
- Trailing client speech after the rep's final run changes nothing.
- A missed
minimum_transferrow carries its hint, not a point. - On a call with no
client_ready_to_book, an unsaid booking-is-binding is innot_due, not counted, and has no due link. tests/fixtures/decisions/3339895706.v4.49e3e48f.json(pre-M1b) giveslegacy: true.
npx vitest run tests/unit/review-model.test.ts tests/unit/engine-purity.test.ts→ all passed.- On the v1 fixture of 3339895706 with 111b's overrides:
INVARIANT: INV-COPILOT-016.
Dependencies: COPILOT-111b, COPILOT-113
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-115b
Title: Review view: completeness as coaching, rendered for the rep [INTEGRATION-CRITICAL]
Description: As a rep reading my own call (decision 2g), I want the review to show what applied, what I covered, what I missed with the suggested line and a link to the moment, and where I was ahead of the guide, so that it builds trust rather than policing.
Acceptance Criteria:
web/src/review-plan.tsrendersreviewModel:- "Must-says: N of M that applied", one row each. A row gives its
whyin words, when it became due (a time link), and either when it was said (a time link) or "Missed" with the suggested line (rendered as the live view renders points) and a link to the due moment. - "Known at the end: N of M" with counts by source. Each item shows its
known_noteonly when the fact's source issignup, because the approved notes read "Known from sign-up"; other sources show theirplan.known.*label. - "Covered before the guide suggested it: N of M".
- A quiet "Not due on this call" line listing
not_dueitems, with no link and no score. - "Reconstructed from the decisions" where the live presentation is inferred (COPILOT-049's caveat).
- No ranking or comparison with other reps or calls.
- With
legacy: true, today's plan panel, labelled "Recorded before M1b".
The reasons ("always on a first call", "the client asked whether it is safe", "someone else might pay", "an amount under the minimum") and the panel labels are UI chrome listed in DESIGN.md (§3 convention 10).
- "Must-says: N of M that applied", one row each. A row gives its
docs/design/DESIGN.md§Review view, §Panels "The plan (was Must-say checklist)" is rewritten to match, with the list of chrome strings.Tests (
tests/unit/web-review-plan.test.ts): the 115 model renders "Must-says: 2 of 3 that applied"; the who-holds row reads "Missed" with its link to 08:50 and the suggested line; a legacy model renders the old panel.npx vitest run tests/unit/web-review-plan.test.ts tests/unit/web-dom-safety.test.ts→ all passed;npm run typecheckexits 0.[UI] Design artifact: DESIGN.md §Review view (variant D), as rewritten, and §Tokens "Review view tokens (C and D)".
[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD). After a cache-warmingreplay-wsof 3339895706 (errors == 0),node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-115b-light","?view=review&call=3339895706","light",1440,900],["COPILOT-115b-dark","?view=review&call=3339895706","dark",1440,900]]' --deployed --at 09:10 --expect '[data-section="must-says"]'exits 0.INVARIANT: INV-COPILOT-016; INV-COPILOT-014; INV-COPILOT-015.
Dependencies: COPILOT-102a, COPILOT-110c, COPILOT-115
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-116b
Title: A known transfer purpose establishes purpose_known: a seeded or confident purpose makes agree_next_step eligible
Description: As the engine, I want the transfer_purpose slot to establish purpose_known the way frequency, pair and documents already establish theirs, so that a call whose purpose the pre-call seed (or a confident call answer) gives can offer agree_next_step and pin_down_timing_and_amount, and onboarding_strong's card share no longer rides on a false-positive purpose_known lock at i 72 (0.67-0.78 on byte-identical requests, persist_fact 0.70). INTEGRATION-CRITICAL. Policy candidate (next version = current + this slot change).
Acceptance Criteria:
- Data, the only policy/src change: in policy/src/plan-onboarding.json the transfer_purpose fact slot gains "known_fact": "purpose_known" and "known_at_conf": 0.9 (the documents_status precedent). No engine code changes: by the existing rules seedFromPrecall then persists purpose_known {i: -1, value: 1} when the precall carries transfer_purpose; a call value ≥ 0.9 persists it at the value's evidence i; a purpose_known lock while the slot is empty makes transfer_purpose pending. bank_hash and playbook_hash equal the current published version; only the onboarding weights_hash moves; customer_success is unchanged.
- Tests: tests/unit/engine-fact-values.test.ts: the slot pins {known_fact: 'purpose_known', known_at_conf: 0.9}; seedFromPrecall on the 3339895706 demo precall persists purpose_known {i: -1, value: 1}; a client-turn fact_transfer_purpose salary_or_pension at 0.95 persists purpose_known at that i, at 0.7 sets the value and not purpose_known, none sets neither. tests/unit/engine-state.test.ts: the seeded call_facts.known_facts is ['docs_status_known', 'frequency_known', 'pair_known', 'purpose_known']. tests/unit/engine-card.test.ts: on initialState(onboarding, demo seed) with no concern open, allowedMoves allows agree_next_step and pin_down_timing_and_amount and suppresses open_with_agenda and ask_about_the_transfer with rule 'blocked_by:purpose_known'; the unseeded requires:purpose_known pins are unchanged. tests/unit/policy-build.test.ts: onboarding differs from the previous version in weights_hash only; cs unchanged. Regenerate the four policy fixtures; stubs re-pinned by policy_hash only (as in 106/116); any replay or plan digest that moves is re-pinned and named in the progress entry with its reason. npx vitest run tests/unit/engine-fact-values.test.ts tests/unit/engine-state.test.ts tests/unit/engine-card.test.ts tests/unit/policy-build.test.ts tests/unit/policy-src.test.ts → all passed; npm run typecheck → 0; npx vitest run → all passed.
- Docs: docs/design/DESIGN.md §Client fact values and the pre-call seed paragraph of docs/architecture.md: a seeded purpose, or a call purpose at ≥ 0.9, establishes purpose_known like frequency and pair; on a call seeded with a purpose, open_with_agenda and ask_about_the_transfer are off the card from the first decision, and pin_down_timing_and_amount and agree_next_step are eligible.
- LIVE-PROBE [INTEGRATION-CRITICAL] candidate and gate; deploy before the upload (§3 convention 4): (1) node scripts/build-policy.mjs --next-version → onboarding v
(weights_hash only), customer_success v (hashes unchanged); node scripts/deploy.mjs; node scripts/publish-policy.mjs --scenario onboarding --version and --scenario customer_success --version → uploaded as drafts, publish refused 409 not_evaluated (expected). (2) node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --policy-version --session m1b-116b- --eval-uid m1b-116b- → REPLAY_EXIT=0, summary quoted; D1: SELECT count(*) n, sum(json_extract(session_state_json,'$.persisted.purpose_known.i') = -1) seeded, sum(json_extract(decision_json,'$.card.state') = 'shown') shown, sum(json_extract(decision_json,'$.card.move_id') = 'agree_next_step') agree FROM decisions WHERE session_id = (SELECT session_id FROM sessions WHERE eval_uid = 'm1b-116b- ') → n 128, seeded 128, shown ≥ 70, agree ≥ 15; quote that session's i 72 purpose_known noul (the card no longer depends on it). (3) node scripts/eval.mjs --scenario both --layers l1,l4,l6 --policy-version --fresh --run m1b-purpose-known- → EVAL_EXIT=0, 'L4: PASS', onboarding_strong card_shown_share ≥ 70/128 quoted, onboarding_weak's share quoted (no seed, no transfer_purpose value: expected in its 61-66 % band); then a second draw node scripts/eval.mjs --scenario onboarding --layers l4 --policy-version --fresh --run m1b-purpose-known- b → L4 PASS, onboarding_strong ≥ 70/128. A '2018: Invalid User Credentials' burst in jev_requests fails zero_errors, not this story: rerun that run once, as in 106 and 108a. (4) node scripts/publish-policy.mjs --scenario onboarding --version and --scenario customer_success --version → 'published … v , previous retired'. - INVARIANT: INV-COPILOT-002; INV-COPILOT-013; INV-COPILOT-020.
Dependencies: COPILOT-113c, COPILOT-106
Priority: MEDIUM
Executor: claude:opus
ID: COPILOT-116
Title: Transfer-purpose trigger words and the salary-home prior [INTEGRATION-CRITICAL]
Description: As the engine, I want "I get paid on the fifteenth" to trigger the purpose question, and a salary sender's likely situations ranked first (decision 2h), so that the purpose is captured on calls like 3339895706.
Acceptance Criteria:
Trigger words. In
policy/src/plan-onboarding.json, thetransfer_purposetrigger (:176) gainspaid|pay ?day|pay home|earn(ings)?|income. The question and its labels are unchanged, and so isbank_hash; the trigger is plan data underweights_hash.The prior.
plan.priorsis optional (absent means M1):[{id: 'salary_home', when: {sell_in: ['SAR', 'AED'], residency_in_sell_country: true}, ranks: {ask_about_the_transfer: ['first_transfer_date_not_said', 'amount_not_said'], explain_who_holds_the_money: ['is_my_money_safe', 'where_do_i_send_funds']}}].- While it holds,
fitsToAskand 108b's pre-judge order that move's candidates byranksbefore the cap of 3. - A prior never writes a slot, a chip or an answer.
- With every move at ≤ 3 candidates, the prior changes no question until M2 adds the segment's situations.
docs/policy/like-for-like.mdlists the three Saudi situations (test transfer first, first-payment delays, address worries) as M2 harvest drafts; they are not encoded.
- While it holds,
Tests.
tests/unit/engine-fact-values.test.ts: v1 line 71 ("… I get paid on the fifteenth …") and v1 line 88 ("pay home. Yeah.") triggerfact_transfer_purpose; no chip comes from a prior alone.tests/unit/engine-situations.test.ts: on a fixture move with 4 candidates, the ranked ones are asked first.eval/labelled/onboarding.jsongainsob_3339895706_i71_purpose_salary(context v1 lines 67-70, latest line 71, verbatim) with{q: 'fact_transfer_purpose', op: 'in', v: ['salary_or_pension']}.
npx vitest run tests/unit/engine-fact-values.test.ts tests/unit/engine-situations.test.ts tests/unit/eval-cases.test.ts→ all passed.[INTEGRATION-CRITICAL] Live probe:
node scripts/probe-questions.mjs --call 3339895706 --at 71 --planned(the draft bank) plansfact_transfer_purposeand answerssalary_or_pension. Quote the confidence.- Candidate and gate
--run m1b-purpose-<yyyymmdd>: the new case passes, and L1 onboarding is ≥ 90 %. Then publish.
INVARIANT: INV-COPILOT-003; INV-COPILOT-001.
Dependencies: COPILOT-110c, COPILOT-110e
Priority: MEDIUM
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, gate)
Notes: The last story of the policy chain; 110e also writes
eval/labelled/onboarding.json.
Phase D — The duplicate tick and the rate false positive
Phase D — The duplicate tick and the rate false positive
ID: COPILOT-117
Title: Concern episodes: never clear on a line under 4 words; reopen the same row within 60 s [INTEGRATION-CRITICAL]
Description: As a rep, I want a three-word fragment never to close a concern, and a concern that comes back seconds after it closed to reopen the same row, so that a single worry is one row with one tick (FEEDBACK §2b).
Acceptance Criteria:
Clearing guard. In
src/engine/concern.tsupdateConcern(HOT, first), clearing (:120) additionally requires a client utterance withu.words ≥ CONCERN_CLEAR_MIN_WORDS(4). The Jev clearing signal ≥concern_clearis still required. There is no punctuation or run-end test: the engine cannot see the next turn.Reopen.
- A raise of the same
typewithinthresholds.concern_reopen_window_sof the last close keeps the previousstart_i, and setsreopened_at: u.iandraises + 1. - The threshold is new and optional in
rules-*.json: 60 s of call time, compared onu.t. Absent means M1. SessionState.concerngainslast_closed: {type, start_i, t}.
- A raise of the same
The row. In
src/engine/plan.ts, at the episode rows (:443-459), a row with thatstart_ireturns tocurrentwithdone_at: nullandreopened_at.web/src/plan.ts(HOT, after 107) renders its state word as "Raised by the client 05:08, raised again 05:27", using the approvedplan.concern.raised_againfromplan.view_labels, and without the second part when that label is absent.docs/design/DESIGN.md§Plan item rows gains that state word.
Tests (
tests/unit/engine-concern.test.ts,tests/unit/engine-plan.test.ts,tests/unit/web-plan.test.ts):- Re-stepping
step()over the v1 utterances of 3339895706 with the recorded answers of run49e3e48f(tests/fixtures/decisions/3339895706.v4.49e3e48f.json):- exactly one
concern:jurisdiction:54row, raised at 308.5 s; raisesis 2 (v1 line 58, 7 words);- done at 364.3 s on v1 line 64 (4 words);
- v1 line 56 ("Sure. Well, I", 3 words,
client_accepts0.71-0.77) does not clear it.
- exactly one
- On a synthetic transcript: a re-raise 20 s after a clear on a 5-word accepting line reopens the same row, and a re-raise 90 s after the next close opens a new one.
- With the threshold absent, M1 behaviour holds.
- The reopened state word renders, with and without the label.
npx vitest run tests/unit/engine-concern.test.ts tests/unit/engine-plan.test.ts tests/unit/engine-recompute.test.ts tests/unit/web-plan.test.ts→ all passed;npm run typecheckexits 0.- Re-stepping
[UI] Design artifact: DESIGN.md §Plan item rows (the client-item state words).
[INTEGRATION-CRITICAL] Candidate and gate (§3 convention 4):
node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --policy-version <N> --session m1b-117-<yyyymmdd> --eval-uid m1b-117-<yyyymmdd>→errors == 0.node scripts/with-cf-env.mjs npx wrangler d1 execute copilot --remote --command "SELECT COUNT(DISTINCT json_extract(e.value,'$.id')) AS n FROM decisions d, json_each(json_extract(d.decision_json,'$.plan.items')) e WHERE d.session_id = '<session_id>' AND json_extract(e.value,'$.id') LIKE 'concern:jurisdiction:%'"→n ≤ 1. Quote it.- The gate run
--run m1b-concern-<yyyymmdd>passes. Publish. - On the published policy,
node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-117-light","?call=3339895706","light",1440,900],["COPILOT-117-dark","?call=3339895706","dark",1440,900]]' --deployed --until '[data-id^="concern:jurisdiction"]'exits 0. This is the plan rows'data-id(web/src/plan.ts:163).
If no jurisdiction concern opens in the page's session (Jev variance), step 4 exits 1: repeat it, at most twice, and quote each outcome. Recompute is not used, because
concern_openfeeds the state Jev saw.INVARIANT: INV-COPILOT-017; INV-COPILOT-001.
Dependencies: COPILOT-104a, COPILOT-105a, COPILOT-107, COPILOT-110d
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, gate)
Notes: 110d is the policy-chain order, and 107 the hot-file order for
web/src/plan.ts.ID: COPILOT-119
Title: A rate concern needs a CurrencyTransfer price: the objection criteria exclude complaints about the client's current provider [INTEGRATION-CRITICAL]
Description: As a rep, I want "Client thinks our rate is too high" to open only when the client pushes back on a CurrencyTransfer price, so that a complaint about their own bank's rate (call 3339895706, v1 line 30, 02:22) never becomes a plan row.
Acceptance Criteria:
Root cause first, quoted in the progress note from 104a's exported answers for v1 i=30 (
tests/fixtures/decisions/3339895706.v4.49e3e48f.json). All six stored rows, measured in D1 on 2026-09-27, agree:client_objecting0.82-0.84;onboarding_objection_typerateat confidence 0.44-0.58, withfeessecond;rate_pressure_raisednot asked (the rate group is gated on a rate, fees or alternative-provider concern, COPILOT-081).
So the concern opened through
updateConcern(src/engine/concern.ts:102) on a real objection signal, and the type Choice mis-typed a complaint about her own bank. The proposal's engine clause would change nothing here and is not built (§7 Q7).Bank text (Jev-facing, so no red-pen; COPILOT-058e precedent):
- In
policy/src/onboarding.jsononboarding_objection_typeandpolicy/src/customer-success.jsoncs_objection_type, therateandfeesoptions name a CurrencyTransfer rate or fee explicitly. - The instructions add: "A complaint about what the client's current bank or provider charges, given as the reason to look at CurrencyTransfer, is not a concern about CurrencyTransfer's rate or fees: choose
alternative_providerwhen they prefer or were quoted by that provider, elsenone." client_objectingcriteria.false(policy/src/shared.json) gains the same case.- The before and after texts go in
docs/runbook/rate-concern.md(new, never published).node scripts/redpen.mjsis re-run.
- In
L1 cases.
eval/labelled/exemplar-moments.jsongainsob_3339895706_i30_objection_type, with the same context and latest as the existingob_3339895706_i30_31_rate_pressure_raised, checking{q: 'onboarding_objection_type', op: 'not_in', v: ['rate', 'fees']}.- It also gains one positive case per scenario where the client pushes back on a CurrencyTransfer quote, checking
client_objecting >= 0.6and the scenario's type Choicein ['rate']. - The existing
client_mentions_wise_quotecase must still pass.
npx vitest run tests/unit/eval-cases.test.ts tests/unit/engine-concern.test.ts→ all passed.Blast radius, before the build.
node scripts/probe-questions.mjs --cases eval/labelled/<file>.jsonruns on the draft bank foronboarding.json,cs.json,rate.json,like-for-like.jsonandexemplar-moments.json. Quote each new check, and every check namingclient_objecting,onboarding_objection_typeorcs_objection_type, before → after.[INTEGRATION-CRITICAL] Candidate and gate (§3 convention 4):
--run m1b-rate-concern-<yyyymmdd>→ L1 onboarding and cs ≥ 90 %,L4: PASS,evaluated=true. Publish.node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-119-<yyyymmdd> --eval-uid m1b-119-<yyyymmdd>→errors == 0.node scripts/with-cf-env.mjs npx wrangler d1 execute copilot --remote --command "SELECT COUNT(*) AS n FROM decisions d JOIN utterances u ON u.call_id = d.call_id AND u.i = d.i, json_each(json_extract(d.decision_json,'$.plan.items')) e WHERE d.session_id = '<session_id>' AND u.t < 240 AND (json_extract(e.value,'$.id') LIKE 'concern:rate:%' OR json_extract(e.value,'$.id') LIKE 'concern:fees:%' OR json_extract(e.value,'$.id') LIKE 'rate:%')"→n: 0. Also… "SELECT COUNT(*) AS n FROM decisions d JOIN utterances u ON u.call_id = d.call_id AND u.i = d.i WHERE d.session_id = '<session_id>' AND u.t < 240 AND json_extract(d.session_state_json,'$.concern.type') IN ('rate','fees')"→n: 0: the concern itself, not only its row.tests/unit/engine-concern.test.tshas a failing regression fixture: the old criteria's answers at v1 i=30 open a rate episode (rate:*). Quote whatever concern opened near 02:22.
INVARIANT: INV-COPILOT-003 (only Jev-facing texts change); INV-COPILOT-006.
Dependencies: COPILOT-104a, COPILOT-117
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, gate)
Notes: COPILOT-058e (the
rate_pressure_raisedexclusion) is a completed M1 prerequisite onmain, not a dependency id here. 117 is the policy-chain order.
Phase E — Verification
Phase E — Verification
ID: COPILOT-122
Title: Documents-not-shared caution and the approved line: a rep who says documents are not shared gets the caution and the line to say instead [INTEGRATION-CRITICAL]
Description: As a rep, I want a caution when I tell a client their documents are not shared with anyone (the activating partner does receive KYC data) and the approved line to say instead, so that the wrong statement is corrected on the call.
Acceptance Criteria:
policy/src/rules-onboarding.jsonandrules-customer-success.json: a new risk flagrep_said_documents_not_shared(noul on the rep's turn: the rep said the client's documents or data are not shared with anyone or stay with CurrencyTransfer only; false when the rep names the partner) with the approved cautioncaution.rep_said_documents_not_sharedas itslive_note, thresholds as the other rep_* flags;policy/src/playbook-onboarding.json: the approved lineclose_the_document_gap.h4added as a point of the documents situation (documents_outstanding_next_step) withapplies: {when_flag: 'rep_said_documents_not_shared'}so it surfaces when the flag fires; both texts read frompolicy/proposals/2026-09-27.jsonwith their approvals; bank and policy hashes bump;node scripts/build-policy.mjsboth scenarios exit 0.eval/labelled/onboarding.jsongains two cases from call 3339895706 line 55 ('we don't share it with anyone else' → flag ≥ 0.6) and a negative case where the rep names the partner (flag < 0.4);npx vitest run tests/unit/eval-cases.test.ts tests/unit/engine-risk.test.ts→ all passed, the risk test asserting the caution note and the surfaced line on the flag.- [INTEGRATION-CRITICAL] Live probe: after
build-policy --next-version,publish-policy(draft),node scripts/eval.mjs --scenario both --layers l1,l4,l6 --policy-version <N> --fresh --run docs-shared-<date>→L1 onboarding ≥ 90 %,L1 cs ≥ 90 %,L4: PASS,evaluated=true(quoted), publish → 200; a freshnode scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20whose decision at line 55 carriesriskscontainingrep_said_documents_not_sharedand whose next card shows pointclose_the_document_gap.h4(quoted from D1).
Dependencies: COPILOT-110j, COPILOT-119
Priority: HIGH
Executor: claude:opus
ID: COPILOT-122b
Title: Tax-residency situation under the documents step: headline, hint and the approved proof-of-address line [INTEGRATION-CRITICAL]
Description: As a rep, I want the worry "will this address tie me to that country" recognised as its own situation under the documents step, with the approved line naming the accepted proofs of address, so that the copilot no longer files it under jurisdiction.
Acceptance Criteria:
policy/src/playbook-onboarding.json: a new situationaddress_ties_me_to_tax_residencyunderclose_the_document_gapwithwhat(the client worries that giving an address in a country ties them to tax residency or privacy exposure there),not_for(a genuine question about where the client lives or banks, which stays with serviceability), threeclient_examplesfrom call 3339895706 lines 54-58 (pseudonymised, verbatim) plus one drafted,headlineandhintfrompolicy/proposals/2026-09-27.json(approved), andpoints= the approvedsituation.address_ties_me_to_tax_residency.pointascore_text_id;{country}renders from the residency slot as the other country lines do;node scripts/build-policy.mjsboth scenarios exit 0 andnode scripts/redpen.mjsregenerates.eval/labelled/like-for-like.jsongains a positive case (line 54: 'a little bit worried about that just in terms of being a tax resident' →fits::address_ties_me_to_tax_residency ≥ 0.6,concern_type != jurisdiction) and a negative case (a client asking where to send funds → fits < 0.4);npx vitest run tests/unit/eval-cases.test.ts tests/unit/policy-src.test.ts→ all passed.- [INTEGRATION-CRITICAL] Live probe: the same gated candidate flow as COPILOT-122 (one policy version may carry both),
L1 onboarding ≥ 90 %,L4: PASS,evaluated=truequoted, publish → 200; a fresh replay of 3339895706 whose decision at line 54 shows the situation's headline with its point and noconcern:jurisdiction:*row opened by that line (quoted from D1).
Dependencies: COPILOT-122, COPILOT-117
Priority: HIGH
Executor: claude:opus
ID: COPILOT-118a
Title: Wall-clock replay pacing and a per-decision arrival log [INTEGRATION-CRITICAL]
Description: As Stevan, I want the 1x replay to pace utterances exactly as the call did and to log when each decision arrived, so that "sentence end to first feedback" and point lifetimes are measured, not projected.
Acceptance Criteria:
- Pacing.
scripts/lib/replay-client.tsandscripts/replay-ws.mjsgain--wall-clock. At--speed 1, each utterance is sent at an absolute deadlinet0 + t_end_i:t0is the wall time of call time 0, which is the log's origin. So the first utterance goes at its ownt_end(19.44 s on 3339895706), not afterMIN_GAP_S. Gaps are neither capped (MAX_GAP_S) nor floored (paceDelays,:41). After areconnect{after_i}, sends resume against the same deadlines. - Log.
--log <file>writes a header line{session_id, call_id, transcript_rev, policy_version, git_sha, speed, wall_clock, decision_points}, then JSON lines{i, call_t, due_ms, sent_ms, arrived_ms}, with all times in ms fromt0. - Summary. The summary gains
transport_after_part_end_p50_msandtransport_after_part_end_p95_ms:arrived_ms − call_t_end × 1000over decision points. This is receipt, not display; the displayed figure is 118b's. - The gate's 1x sample. The 1x replay of the L4/L6 gate (
scripts/lib/eval-l4.ts/eval-l6.ts, the customer-success weak call 3347356034) uses the same wall-clock pacing, so the gate's e2e p95 is measured on true pacing. - Tests (
tests/unit/replay-client.test.ts, fake timers): the first send att_end_0; no cap and no floor; alignment after a reconnect; the log fields; the summary figures.npx vitest run tests/unit/replay-client.test.ts→ all passed;npm run typecheckexits 0. - [INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/replay-ws.mjs --call 3347356034 --scenario customer_success --speed 1 --wall-clock --log /tmp/m1b-118a.jsonl(329 s) →errors == 0. The header matches the session. The first data line'ssent_msis within 200 ms ofdue_ms. Quote both transport figures.
- Pacing.
Dependencies: none
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-118b
Title: The measured table and a failing M1b acceptance check [INTEGRATION-CRITICAL]
Description: As Stevan, I want one command that fails when any M1b exit condition does not hold, and one that prints the measured table, so that "M1b is done" is a test rather than a set of printouts.
Acceptance Criteria:
Recorded arrivals.
scripts/presenter-sim.mjs --log <file>(105a, 107) uses the recorded arrivals of 118a's log. Its output also gives:- the amount point's first render state;
- the maximum suggestions shown at once;
presented_after_part_end_p50_msand_p95_ms: for each decision point, the first presenter change caused by that decision (a tick, chip, caution or point it carries), minus its part'st_end. A change caused by an earlier decision, or by the clock (fades, dwell releases), is not counted.no_change: the count of decisions whose plan changed nothing on screen, reported separately and left out of the percentiles.
This is the displayed feedback, distinct from 118a's transport figures.
tests/unit/presenter-sim.test.tscovers an unrelated clock-driven update and two overlapping arrivals.The measured table.
scripts/m1b-measure.mjs --call <id> --session <session_id> [--log <file>]prints the rows:- from the fixtures: decision points, and the longest and p90 gap, for v1 and v2;
- the word-weighted spoken-to-visible wait under 102's rule, against v1's end-of-line rule (9.1 s);
- from
presenter-sim: decisions with an uncovered point, distinct items with a point, the shortest point display (excluding 107's exempt exits), and the points withdrawn within 12 s for any other reason.
The acceptance check.
scripts/m1b-acceptance-check.mjs --call 3339895706 --session <session_id> --log <file> --expect-seed demoassertshealth.version== HEAD, then exits 1 naming each failed condition, else 0.- A pass requires
--log, an arrival log from a--wall-clockreplay of that session. The log is rejected unless its header'ssession_id,call_id,transcript_rev,policy_versionandgit_shamatch the session and the deployment, and it holds one line per decision point. - With
--projectedinstead, arrivals are projected, and the run only reports, exiting 2. - The conditions:
- The call's D1 utterances have
stitch_version2, 125-145 decision points, a longest gap between decision points ≤ 15 s, and a p90 ≤ 11 s. - More than 12 decisions with an uncovered point, and more than 3 items with a point.
- At most 3 points shown at once, and no non-exempt point withdrawn within 12 s.
- The amount point, from
ask_about_the_transfer/amount_not_said, never renders tinted after a rep-turn decision. - The pure transcript reveal (
revealedWords, 102) draws no word before its estimated time, on this call's utterances. - At the last decision: ≤ 4 open live rows and no
not_neededitem. - No
concern:rate:%,concern:fees:%orrate:%item, and no concern typedrateorfeesinsession_state_json, before 240 s; and ≤ 1concern:jurisdiction:id. - At the first decision: 7 chips sourced
demo, andplan.knownholdingcheck_serviceability_firstandunderstand_the_transfer.--expect-seed demomakes this required, and a missing seed fails. reviewModel(115) over the session gives{counted: 3, said: 2}.- 3 marks on 3339895706, 1 label on 3303259297, and
pd_ct_id109829 on 3485591407.
The UI tally is checked in the browser by 118.
- A pass requires
Constants test.
tests/unit/design-constants.test.tsasserts that DESIGN.md's §Cadence rules constants table equals the presenter's exports.Tests.
tests/unit/m1b-acceptance-check.test.tsfeeds each failing condition, including condition 0, a mismatched log header, a truncated log and a missing seed, to the checker's pure core, and expects a named failure.npx vitest run tests/unit/m1b-acceptance-check.test.ts tests/unit/presenter-sim.test.ts tests/unit/design-constants.test.ts→ all passed;npm run typecheckexits 0.[INTEGRATION-CRITICAL] Live probe on the deployed Worker (
health.version== HEAD):node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 20 --session m1b-118b-<yyyymmdd>→errors == 0.node scripts/m1b-acceptance-check.mjs --call 3339895706 --session <session_id> --projected --expect-seed demoruns end to end and prints every condition with its measured value. Quote the output;--projectedexits 2 by design, and the passing run is 118's.
Dependencies: COPILOT-102, COPILOT-105a, COPILOT-107, COPILOT-115, COPILOT-118a
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration
ID: COPILOT-118
Title: M1b verification: the acceptance check on the deployed Worker, the live-probe register, the measured table, the acceptance runbook [INTEGRATION-CRITICAL]
Description: As Stevan, I want one pass that re-runs every gate, replays the four samples on the deployed Worker, passes the acceptance check, records the figures against the 2026-09-27 baseline, and gives me a short acceptance walk, so that M1b is demonstrably done.
Acceptance Criteria:
- [INTEGRATION-CRITICAL] On the deployed Worker (
health.version== HEAD), with the published policy:- For each sample,
node scripts/replay-ws.mjs --call <id> --scenario <s> --speed 20(no--eval-uid, so it warms Stevan's cache) →errors == 0, one model id,max_input_tokens ≤ 12000. Quotereconnects. node scripts/replay-ws.mjs --call 3339895706 --scenario onboarding --speed 1 --wall-clock --eval-uid m1b-118-1x-<yyyymmdd> --session m1b-118-1x --log docs/verification/m1b-1x.jsonl→errors == 0, e2e p95 ≤ 3000 ms. It uses a fresh uid, so Jev answers every decision and none is served from the cache. Quote the transport figures, and the presented figures frompresenter-sim --log.node scripts/m1b-acceptance-check.mjs --call 3339895706 --session <that session_id> --log docs/verification/m1b-1x.jsonl --expect-seed demoexits 0.node scripts/design/shoot-dashboard.mjs docs/design/evidence '[["COPILOT-118-light","?call=3339895706","light",1440,900],["COPILOT-118-dark","?call=3339895706","dark",1440,900]]' --deployed --at 06:10 --expect '[data-section="covered"]' --expect-none '.tally'exits 0: the head's.tallyelement (web/src/plan.ts:415) is gone. 102's test proves that--expect-nonefails on a match.node scripts/probe-player.mjs --call 3339895706 --check seekexits 0.node scripts/eval.mjs --scenario both --layers l1,l4,l6 --run m1b-final-<yyyymmdd>→ L1 ≥ 90 % for both,L4: PASS, L6 PASS (cost per call ≤L6_LIMITS.cost_per_call_usd).node scripts/cost-report.mjs --since <M1b start ISO> --out /tmp/m1b-cost.md; its body is appended under an "M1b" heading indocs/reports/cost-latency.md, keeping the earlier sections.
- For each sample,
- Derived-artefact commit (§3 convention 1):
docs/demo-script.mdrows brought tom1b-final's timeline, withtests/unit/demo-check.test.tsre-pinned.node scripts/demo-check.mjs eval/reports/m1b-final-<yyyymmdd>.l4l6.json docs/demo-script.mdexits 0.docs/verification/M1b.md:- the register, one row per INTEGRATION-CRITICAL story, 118 included (102, 102a, 104a, 104b, 104c, 104d, 104e, 105, 105a, 106, 107, 108a, 108b, 110b, 110c, 110d, 110e, 110f, 111a, 111b, 111c, 112b, 113c, 114, 114b, 115b, 116, 117, 118, 118a, 118b, 119, 120, 121); the list is checked against the tagged stories, each with its command, quoted evidence and commit, and 113b listed as inert until a served list exists;
- the measured table from
node scripts/m1b-measure.mjs --call 3339895706 --session <id> --log docs/verification/m1b-1x.jsonl, against the baseline; - the jurisdiction row count;
- the uploaded copies' stitch version;
- the policy versions promoted in M1b, with their gate runs;
- the open decisions (12, 13).
docs/runbook/m1b-acceptance.md(never published; linked fromdocs/runbook/README.md): PRD §5.6 as a checklist, with its times checked against the final run.
- Gates, on the final commit, after the derived-artefact commit.
npm run typecheck,npx vitest run,npx vitest run -c vitest.workers.config.ts,node scripts/with-cf-env.mjs npx wrangler deploy --dry-run,npm run siteandnode scripts/check-secrets.mjsall exit 0. Quote the test counts. - INVARIANT: INV-COPILOT-006; INV-COPILOT-007; INV-COPILOT-013.
- [INTEGRATION-CRITICAL] On the deployed Worker (
Dependencies: COPILOT-102, COPILOT-104d, COPILOT-104e, COPILOT-106, COPILOT-107, COPILOT-108b, COPILOT-110e, COPILOT-111c, COPILOT-113b, COPILOT-114, COPILOT-114b, COPILOT-115b, COPILOT-116, COPILOT-117, COPILOT-118b, COPILOT-119, COPILOT-120
Priority: HIGH
Executor: claude:opus
Estimate: 1 otto iteration (retry budget: 1, eval)
Sizing and order
| Phase | Stories | Otto iterations | Retry budget (gate or Jev-dependent probe) |
|---|---|---|---|
| A Live granularity | 102a, 102, 103, 104a, 104f, 104b, 104c, 104d, 104e, 105a, 105, 106 | 12 | 2 (104e, 106) |
| B Card lifetime | 121, 107, 108a, 108b, 120 | 5 | 2 (108a, 108b) |
| C Plan semantics | 110a, 110f, 110g, 110h, 110c, 110d, 110b, 110e, 111a, 111b, 111c, 112, 112b, 113, 113c, 113b, 114, 114b, 115, 115b, 116 | 21 | 6 (110c, 110d, 110b, 110e, 111b, 116) |
| D Duplicate tick and rate false positive | 117, 119 | 2 | 2 |
| E Verification | 118a, 118b, 118 | 3 | 1 |
| Total | 43 stories | 43 | up to 13 more (56 at worst) |
The 43 stories are the proposal's 22 without the conditional 109, split to one concern each over five check rounds (§7), plus the design board 121.
- Starting points (no dependencies): 102a, 103, 104a, 110a, 112, 118a and 121.
- Policy chain: 110c → 110d → 117 → 119 → 106 → 108a → 108b → 110b → 116.
- Critical path (15 stories, 15 iterations plus up to 4 retries): 103 → 104f → 104c → 104e → 105 → 107 → 117 → 111a → 111b → 111c → 113 → 113c → 114 → 114b → 118.
5.6 Manual acceptance (Stevan; referenced by COPILOT-118; not acceptance criteria)
Nothing here blocks a story. COPILOT-118 copies it into docs/runbook/m1b-acceptance.md, with the times checked against the final run.
M-1 Design board (121). Open
docs/design/shotgun/M1b/board.htmlor theCOPILOT-121-*screenshots, and accept or change each recommendation. A change is a follow-up to one constant or one stylesheet in 107, 114 or 120.M-2 Open decisions.
- Decision 12: the line-55 wording, an approved line or a caution.
- Decision 13: whether an employer or HR letter counts as proof of address.
- The served-country list, if available.
§7 gives each one's change.
M-3 Acceptance walk. Sign in through Access and replay 3339895706 at 1x in the live view (call times):
- 00:05: seven chips read "Seeded for the demo", and none reads "From sign-up". "What they're sending, and why" and "Where they live and bank" never show as rows.
- 00:45-01:30: the client's long turn appears word by word, and the "SAR to NZD" chip arrives within a sentence of her naming the currencies.
- Around 01:40: the amount question arrives grey (already asked), never tinted.
- 02:22-04:00: no "Client thinks our rate is too high" row.
- 02:41: the Canadian leg is heard. "Amounts under the minimum" does not appear, because the main amount clears it (decision 2e).
- 05:08-06:04: one concern row, raised at 05:08, still open at 05:27, done at 06:04.
- 08:50: "is it gonna be safe". "Who holds their money" opens without waiting out the dwell, and its point stays at least 12 s.
- Throughout:
- the head reads "Guide for this call";
- Coming up has at most 3 rows; there is no tally; Covered is one collapsed line;
- "Whose account pays" never appears (a salary into her own account, decision 2d);
- at 20×, the player does not move when "catching up" shows.
- Review view:
- "Must-says: 2 of 3 that applied", with "Who holds their money" marked "Missed" and linked to 08:50;
- your three marks sit on the moments you marked, re-inserted after the re-import.
M-4 Optional. Check the AI Gateway credit balance before and after a replay, as in
docs/runbook/m1-acceptance.md§4 (there is no API for it).
6. Risks and Mitigations
| # | Risk | Mitigation |
|---|---|---|
| 1 | A forced re-import deletes every derived row of a call and renumbers i; marks, labels, L4 labels, the demo script and the runbook are keyed by i |
104a exports first, with expected_curation. 104b's claim refuses atomically when the curation changed. 104c restores idempotently, with ownership checks. 104e compares counts. INV-COPILOT-019 |
| 2 | Every state change invalidates the answer cache | A paid re-run costs about $0.05 a call; each gate is --fresh |
| 3 | The strong calls land at $0.047-0.053 after the split and the pre-judge | The gate is $0.06 (decision 16), and prejudge is the first drop. 108b quotes the delta |
| 4 | Word times inside a line are interpolated | An honest estimate; v2 parts bound spans at about 12 s. Real times arrive with nova-3 (§7 Q4, M3) |
| 5 | Hysteresis keeps an uncertain situation one decision longer | Only in [0.3, 0.6); a clear miss leaves at once; misses are counted once per decision (108a) |
| 6 | Three stacked points (about 18 s of reading) overload a listening rep | Decision 3 sets 3. The 121 board shows 2 against 3; the oldest goes first. Stevan signs off at M-1 |
| 7 | A pre-call fact is wrong | A confident call value replaces it ("updated"), and source_conflict surfaces the fact card (113, 111a) |
| 8 | No served-country list | The gate is inert on an empty list and never keyed on "known" alone (§7 Q3) |
| 9 | A published bundle lacks the new labels or fields | Optional fields (§3 convention 5); a deploy before each upload; the views fall back to today's words |
| 10 | More decisions reach the subrequest guard sooner at 20x | The COPILOT-058b reconnect path; reconnects is quoted |
| 11 | Uploads would keep 45 s turns | 104d moves the Worker ingest to v2 |
| 12 | Run-to-run Jev variance | Seeding removes residency variance. Probes assert bounds; 117 allows three replays for its screenshot; 111b fails honestly if decision 2c is not met live |
| 13 | Nine gated candidates in a row; a failed gate blocks the chain | The L1 margin is 91/91 and 79/81. A failure lists its cases; no threshold is lowered and no label is edited |
| 14 | Source figures were wrong (the baseline file, line 30, 08:16, "3 of 4") | The criteria use the measured values (§7 Q7-Q10) |
| 15 | The M2 PRD also claims migration 0004, and plans UPDATE calls SET pd_ct_id=NULL |
§7 Q6 |
| 16 | Retired checklist ids break L4's partition | 110d reports them as retired; the labels are untouched |
| 17 | A human-only step reaches an executor | §3 convention 2; §5.6 |
| 18 | Timing figures measured on capped pacing or invented arrivals | 118a's absolute deadlines and log; presenter-sim labels projections |
| 19 | Old decisions lack must_says/known |
115's legacy fallback |
| 20 | Live evidence depends on the sequence of code, deploy and derived files | §3 convention 1 defines the commit sequence and re-runs the suites on the final commit |
| 21 | A seed conflict clears on engine time (113); when the presenter defers the card, the time the rep sees it may be under 20 s | Accepted for M1b: the conflict still surfaces its card once and needs a client run end and a rep turn to clear. Measuring on displayed time needs a browser-reported shown event, left for later |
7. Unresolved Questions
Still open (Stevan). Each is marked in the criteria.
- Q1 [DECISION 12] The rep's line 55, "we don't share it with anyone else": an approved line or a caution? Nothing is drafted and no story encodes it. Either answer is a red-pen addition plus a candidate.
- Q2 [DECISION 13] Does an employer or HR letter count as proof of address for the activating partner? The situation
address_ties_me_to_tax_residencystays inpolicy/proposals/2026-09-27.jsonin full: its approved headline and hint, and its[VERIFY]point. The builder requires every situation'score_text_idamong its points (src/policy/build.ts:587).- This departs from the literal instruction to apply everything but the
[VERIFY]line. The alternative was a new "dormant situation" grammar with no behaviour. - On the answer, the point is approved, and a re-run of
apply-label-approvals.mjs(idempotent) adds the whole situation underclose_the_document_gap. Or the situation is dropped.
- This departs from the literal instruction to apply everything but the
- Q3 The served-country list per partner. Until it exists, the jurisdiction gate (113b) and the "outside the served list" surfacing (110b) are inert. Supplying it is a bank change (
fact_lists.served_countries) plus a candidate.
Open technical questions.
- Q4 Keep nova-3 word times for uploads now (4-6 h), for exact splits and gaps re-measured on real times? Not in M1b.
- Q5 "Joint" in decision 2d has no value in the funding-account closed list (
own_account,company_account,third_party). A joint account reads as unclear, so the item applies unless the purpose is salary. Addingjointis a bank change and is not in M1b. - Q6
prd-m2-corpus.mdclaims0004_scoring.sqlto0009_library.sql, and its0002/0003already exist onmain. M1b takes0004_precall.sql, so M2 must renumber from0005. Its0003_corpus.sqlalso runsUPDATE calls SET pd_ct_id=NULL, which would erase 3485591407's 109829.
Corrections to the sources, adopted in the criteria.
- Q7 At v1 line 30, the stored answers give
client_objecting0.82-0.84, not the 0.0 in FEEDBACK §2e. The proposal's engine gate for 119 would not stop this concern, so 119 fixes the criteria instead. - Q8 The L1 baseline is
eval/reports/final-20260927-{onboarding,cs}.l1.json(v4: 91/91, 79/81), notrecc-20260927(v2: 65.9 %/61.7 %). - Q9 "is it gonna be safe" is v1 line 107, at 08:50, not 08:16.
- Q10 Under decisions 2c-2e, the review tally for this call is "2 of 3 that applied". Who holds the money, booking and documents apply, and only booking and documents are said (
eval/labelled/must-say.json). FEEDBACK's "3 of 4" does not follow from the labels.
Changes against the proposal's §7 sketch (judgement recorded):
- 109 is dropped (decision 3 = A).
- 121 is added: the shotgun round before 107 and 114, extended to 120 per FEEDBACK §4.
- Decisions 2c, 2d, 2e and 2f were taken on 2026-09-27, so their markers are gone:
- 110c applies Stevan's approved file;
- the minimum stopgap is replaced by the correction rule (111c);
- the first-transfer walkthrough gains "Who holds their money" through
texts_from, which needs no new text id.
LIVE_LIST_MODEB/C are dropped.bank_countryhas noknown_fact. Theeval/cases/…paths becomeeval/labelled/….- 104a's exports go to
tests/fixtures/{curation,decisions}/, sincetests/fixtures/exports/holds the call-coach exports. The v2 fixtures come from the committed samples, because the worktree has nodata/.INGESTmoves with 106's candidate. - The proposal's
ws-probe --scenario --speedflags do not exist; the runs usereplay-ws, thenws-probe --resumeflags.
Adversarial-check notes (otto check, codex gpt-6-astra, xhigh):
Round 1 (2 BLOCKER, 38 MAJOR, 1 MINOR; FIX-FIRST). Adopted:
- Blockers: the archived M1 dependency ids (046, 058e) are removed and noted as completed prerequisites.
- Dependency fields hold ids only; the reasons moved to Notes.
- Splits of 104b, 108, 110a, 110b, 111b, 112, 113, 114, 115 and 118.
- Remap and restore: a bounded forced cut; speaker and fragment matching in the remap;
gold_json.i; an idempotent restore with route validation; an explicitly pinned baseline in 104e; 119 reading 104a's exported answers. - UI stories: a design anchor and deployed both-theme evidence on 105, 111b, 113 and 117.
- 107: the lifetime precedence.
- Missing edges added.
- Engine:
recordFitscounts a miss once per utterance;- compatible readers before data;
- checklist derivation per template;
- the approvals test sample built inline;
- the walkthrough item;
- precall provenance on the import path, and preservation on replacement;
- seed propagation through
set_weightsand rehydrate; - the UI-chrome rule;
- the legacy fallback.
- 118: timing and acceptance tooling.
Declined: splitting 103, 106 and 119 (each is one concern with its own verification).
Round 2 (0 BLOCKER, 28 MAJOR, 2 MINOR; FIX-FIRST). Adopted:
- The remap: v1 split parts sharing
tare handled by effective spans (lines 89-90), with a regression test. - The curation precondition is enforced atomically in the replacement claim (
expected_curation), with a first-import regression test;precall_jsonpreservation moves to 112b, which owns the column. - 104c validation: deterministic-id and ownership checks.
- Splits: 104b into 104b/104c/104d/104e; the simulator into its own story, 105a; 110a into 110a/110f; 118a into 118a/118b.
- The commit sequence: code, deploy, live run, derived-artefact commit, suites re-run.
- 105: switches only at run ends after the dwell.
- 106: the SQL path
$.jev.budget.dropped, with a non-zero denominator. - Edges: 110f→107, 107→111b, 107→117.
- Policy candidates: deploy before upload.
- The bank-country grammar: a country span with a template, an optional slot, and the precall accepting a country id; its approved template is applied by 110b from Stevan's file.
- 110c: mixed-file behaviour (approved entries validated all-or-nothing; drafts and homeless entries reported and skipped); the tax-residency situation kept out whole (Q2).
- The
not_needed_reasonrule relaxed (110a); approved reasons are set by 110c before 110d changesapplies. - Decision 2d in every template.
- Amount history with currency, correction provenance and fact reconciliation (111c).
source_conflictattached, with replacement provenance (113).reviewModeltakes the utterances for the last-rep-turn rule.- 117: the real
data-idselector, shot after promotion. - 118a: absolute-deadline pacing.
- 118b: asserts the full exit contract; 118 checks the tally in the browser (
--expect-none). - IC probes on 104b, 104c, 104d, 110f, 118a and 118b.
- 111b shoots its UI change.
- The two MINORs.
Declined: none.
- The remap: v1 split parts sharing
Round 3 (0 BLOCKER, 24 MAJOR, 1 MINOR; FIX-FIRST). Adopted:
- 104a: the export now comes from one consistent snapshot (before and after reads, with retries); the remap is split out as 104f.
- 104b: an omitted
expected_curationasserts zero curation atomically (the zero-to-one race);demo-check's existing test pins--stitch 1until 104e. - 104c: a field-level schema (identifiers, integers, audit metadata and
gold_jsonapart from prose);label_idas the identity; collision, duplicate and bound checks; a non-empty live restore on a disposable synthetic call. - 105: a boundary rule for late arrivals (a decision on a run's last part applies on arrival after the dwell).
- 107: the sticky line counts toward the cap of 3; the switch expectations are computed, not hard-coded.
- 110c: the invalid, draft, deferred and applicable classes.
- 110d:
satisfied_byderived for the mapped facts. - 110b: the chip's country rendering,
data-slot, and design and both-theme evidence. - 111a: a surfaced fact leaves
knownwhile current. - 111c: a seed is superseded by the call's first amount in its currency.
- 113: split into 113 (pure) and 113c (session paths); the chip view moves to 114b; history (
replaced) is separated from an activeconflictthat clears once surfaced. - 115: a missed must-say without points uses its hint.
- 115b: the "Known from sign-up" note is shown only for a
signupsource. - 102: a
--untilflag, so screenshots pause on the real rendered state rather than a projected time. - 120:
--view live. - 118: a fresh uid for the 1x run; the gate's 1x sample on wall-clock pacing (118a);
.tallyas the real selector; a--projecteddiagnostic mode (118b);cost-report --out, then append (the MINOR).
Declined, with reasons:
- The objection to the code commit plus derived-artefact commit (118, §3 convention 1). Round 2 asked for exactly this executable sequence, and round 3 reverses it: an oscillation, per the otto-plan rule.
scripts/deploy.mjsrefuses a dirty tree, so a story's deployed revision cannot also contain the evidence of its own deployment. M1 ran 122 stories this way. The sequence is kept, and the suites re-run on the final commit. - A three-way split of 104a. The two exports share one mechanism and one snapshot rule, so it is taken as a two-way split (104a exports, 104f remap).
Round 4 (0 BLOCKER, 20 MAJOR, 1 MINOR; FIX-FIRST; the last round allowed). Fixed after the round, and not re-checked:
- 110a carries the
plan.ts:422done_whenreader, so its typecheck stands alone. - Edges: 104b→103 and 104d→104b.
- 111a:
- the unsatisfied-fact assertion moved here from 110d, so there is no cycle;
- satisfaction is re-evaluated after surfacing, with a composite-fact test.
- Policy candidates: bundles are built before the code commit, because
policy/build/is tracked and--next-versionreads it. - 118a precedes the first gate (110c) and 104e's eval; the eval hot-file order is recorded.
- 104a: the consistency check includes
transcript_rev, the content hash and a hash of the rows. - 104c uses the real columns (
topic,jev_json,note,admin) and a real-row round trip. - 104d: both forced re-stitch paths carry the atomic curation condition.
- 111c: correction boundaries are client runs, not v2 parts, and legacy
amountsmigrate into history. - 114b links to
updated_at ?? at. - 115: the last-opportunity rule for the final rep run.
- 119 and 118b check
rate:*episodes and the concern type, not onlyconcern:rate:*. - Latency: transport latency (118a) is separated from displayed latency (118b).
- The arrival log carries an identity header, and the checker rejects mismatched or truncated logs.
- The acceptance check asserts the stitch counts and gaps, and requires the demo seed.
- The register lists all 31 INTEGRATION-CRITICAL stories (the MINOR); 102's and 118a's titles gain the tag.
Deferred at the time to
/prd-task-sizer: the splits of 102 and 110d (applied in round 5).Declined as oscillation (raised in rounds 3 and 4 against round 2's own demand): one commit per story with the evidence inside it.
deploy.mjsneeds a clean tree, so a story cannot commit the evidence of its own deployed revision. The code, deploy, run, derived-artefact, re-verify sequence of §3 convention 1 stays.- 110a carries the
Round 5 (0 BLOCKER, 14 MAJOR, 0 MINOR; FIX-FIRST; run by the coordinator). Each finding was checked against the code, and every one was applied except the one-commit objection:
- 102: split into 102a (the screenshot flags and tests) and 102 (the reveal).
- 110d: split into 110g (generators and L4
retired), 110h (the opt-in derived checklist and the hero) and 110d (the data conversion and gate). - 110a:
web/src/review-plan.ts:423mustSayOfreads an optionaldone_when;- text identity:
assertTextsaccepts canonical inherited ids fortexts_fromand a.known_noteid; src/policy/text.tsregistersknown_noteandview_labels, with rejection tests for drafts and lookup tests for approved ones.
- 110g:
scripts/redpen.mjs(describeApplies,describeDone) handles the new forms before 110d regenerates. - 110c: a must-say's
known_noteis written under the required.not_neededid, with the mapping printed. - 104c:
gold_jsonis validated against the closedFalsePositiveLabelshape. - 111a:
unknown_after_stageneeds a known, non-nonestage;- surfaced items compete in the ranking, and exactly one is current;
- judgement items keep their
applies.
- 111b and 115: a must-say whose
duenever fires hasdue_at: nulland is not counted ("Not due on this call"). - 113: a conflict clears only after
conflict_hold_scurrent, a client run end and a rep-turn decision, and survives consecutive mid-run decisions. - 105: the screenshot pauses on a
data-midrunpoint, not on the word reveal.
Declined: 118's one-commit objection. It is the same oscillation as rounds 3 and 4: it reverses round 2's mandate, and it is the one-commit-per-story sequence.
Round 6 (0 BLOCKER, 11 MAJOR, 0 MINOR; FIX-FIRST). Each finding was verified in the code before it was applied:
- Edges: 105 now depends on 102 (its probe expects
[data-partial]), and 110c on 110g. - 102a: tagged, with four deployed Playwright probes, two positive and two negative.
- 110g: resolves
texts_fromandchecklist: "derived"through the builder's helpers (redpen.mjs:108readitem.title.text;:285calledchecklist.map), tested on the converted shapes. - 110a: now also carries the widened
appliesreader (planApplies,plan.ts:136, handed the remaining union tochecklistApplies) andsignals_seen. 110f keeps only the wire types and the probe. - 110h:
checklist_by_templatejoinsHASH_FIELDS.weights(canonical.ts:18lacked it), with loader validation and mutation tests. - 111c: a correction marker supersedes whatever the currency; no marker means an addition.
- 113:
templateFor(plan.ts:72, which defaults toother) reads a new optionalplan.initial_template, which 110d sets tofirst_call_after_signupfor onboarding;- the conflict hold counts from a browser-reported shown time (a
card_shownevent path). This was reverted in round 7; see below.
- 104c: a label on an id that 110d retires is restored unchanged, with the export's
policy_at_exportas evidence, and reported asretired.
Declined: 118's one-commit objection, for the fifth time. It reverses round 2's mandate, and
deploy.mjsneeds a clean tree, so a story's deployed revision cannot contain its own evidence. 118's final gates re-run on the final commit, and the register records each story's deployed SHA.- Edges: 105 now depends on 102 (its probe expects
Round 7 (0 BLOCKER, 11 MAJOR, 1 MINOR; FIX-FIRST). This was the last round; the edits after it were not re-checked.
- Resolved by dropping the
card_shownpath (coordinator decision): the five findings on it (queue versus bypass, seek and event ordering, episode identity, headless presenter parity, and 113c's size). 113 measures the conflict hold on engine time again, with the limitation stated in the story and in §6 risk 21, and 113c is back to seed propagation only. - Applied:
- 110h:
rate_pressure_handledis exempt in the derived-checklist validation; - 121: tagged, with its Playwright render as the live probe;
- 118: the register includes 118 and is checked against the tagged set (the MINOR);
- 118b: displayed latency is attributed to its own decision, with no-change decisions counted separately;
- 114b: no "From CRM" label in M1b (a
crmchip has no suffix; CRM provenance arrives with the M2 join).
- 110h:
- Deferred to
/prd-task-sizerat convert time: the split ofsignals_seenout of 110a. - Declined: 118's one-commit objection, for the sixth time. It reverses round 2's mandate, and
deploy.mjsneeds a clean tree, so a story's deployed revision cannot contain its own evidence.
- Resolved by dropping the
Trend: 41 → 30 → 25 → 21 → 14 → 11 → 12 findings, 2 → 0 blockers. Of round 7's findings, five concerned the
card_shownpath added in round 6.
8. Implementation Steps
- Operator: read §7. Nothing blocks the start; decisions 12 and 13 affect only the held tax-residency situation and an uncoded line.
- Convert and size.
otto plan convert prd-m1b.md prd-implementation-m1b.json, then/prd-task-sizer prd-implementation-m1b.json. Letters stay reserved for its splits. - Run. On a branch off
main(for exampleotto/m1b-live-plan, in its own worktree), runotto implement <N> prd-implementation-m1b.json --llm claude. Every story isExecutor: claude:opus, and the run is sequential.summary.executionOrderfrom the convert step is authoritative. - When 121 lands: Stevan's board sign-off (§5.6 M-1).
- On decision 12, 13 or the served list (§5.6 M-2): the change named in §7 Q1-Q3, then a candidate through the gate (§3 convention 4).
- When 118 lands: Stevan's acceptance walk (§5.6 M-3).
- After M1b:
- renumber the M2 PRD's migrations and protect
pd_ct_id(§7 Q6); - run the one-day M3 transport spike before planning M3.
- renumber the M2 PRD's migrations and protect