Jev call copilot CurrencyTransfer research and architecture built 2026-09-28

Copilot — Milestone 1b: live granularity, card lifetime and the adaptive call plan for CurrencyTransfer replays

1. Overview

2. Requirements

Functional Requirements

Core

Secondary

Out of scope

Non-Functional Requirements

Decisions (Stevan, 2026-09-27; policy/approvals/2026-09-27-m1b-decisions.json)

# Decision Lands in
1a Cut at a sentence end after ≥ 6 words and ≥ 5 s since the last cut; forced at 12 s 103
1b Feedback (ticks, chips, caution, new points on the open card) moves at sentence ends; the card switches only at the end of the client's run; one flag, defaulting to that 105
2a Pre-call facts from the CT sign-up wizard via pd_ct_id, closed-list values only, chips labelled with their source; the join is M2; the sample call is seeded by hand as "Seeded for the demo" 112, 112b, 113, 114b
2b Known chips, Now (one card), Coming up (≤ 3 applying must-says), Covered (collapsed); no counts live 114
2c Who holds the money is a must-say that is due when the client raises safety, and otherwise belongs to the first-transfer walkthrough. It is never listed on a first call where the client does not raise safety 110d, 111b
2d Fund from your own account is a must-say only while the payer might be someone else (third party, joint, company or unclear). A salary into the client's own account never triggers it 110d, 111b
2e The GBP 5,000 minimum applies per client relationship or booking total, not per currency leg. A small second leg passes once the main amount clears the minimum: the item is not needed, and no batching line is written. Multiple amounts are tracked only so that a mid-call correction is not mistaken for a second leg 111c
2f The reworded labels are approved as drafted (policy/proposals/2026-09-27.json, approved_by stevan@currencytransfer.com). Its one [VERIFY] line waits for decision 13 110c
2g The rep sees their own review 115, 115b
2h Ranked situations; no Saudi overlay 116
3 Keep the like-for-like gate and add the six lifetime rules (12 s floor, stack to 3, sticky last suggestion 12 s, rep-turn grace, hysteresis per 17, pre-judge the opening item) 107, 108a, 108b
14 Live audio is M3 out of scope
15 The hard token cap stays 12,000 106, 108b, 118
16 L6_LIMITS.cost_per_call_usd = 0.06 106
17 A shown situation survives one answer in [0.3, 0.6) or one dropped question; it leaves on the second, at once under 0.3, or when another candidate qualifies 108a
12, 13 Still open. Stories carry [DECISION 12] / [DECISION 13] §7 Q1, Q2

3. Technical Architecture

flowchart LR
  subgraph devbox["devbox / otto worktree (no data/)"]
    SMP[("tests/fixtures/samples/*.json<br/>pseudonymised fragments")]
    ST2["stitch() STITCH_V2<br/>src/ingest/stitch.ts"]
    EXP["export-call-curation.mjs<br/>remap-i.mjs"]
    IMP["import-fixture.mjs --stitch 2"]
    PRE["precall-payload.mjs"]
    POL["policy/src → build-policy --next-version<br/>→ publish-policy (draft) → eval gate → publish"]
    EVAL["eval.mjs l1,l4,l6 · replay-ws · ws-probe<br/>presenter-sim · m1b-acceptance-check"]
    SMP --> ST2 --> IMP
  end
  subgraph CF["Cloudflare (Access on every hostname)"]
    W["Worker jev-copilot<br/>+ PUT /api/calls/:id/precall<br/>+ POST /api/calls/:id/curation"]
    D1[("D1 copilot<br/>+ 0004_precall.sql")]
    DO[("DO CallSession<br/>initialState(pin, seed) · step()")]
    AI[["typesafe/jev via AI Gateway jev-copilot"]]
  end
  subgraph WEB["Browser (owns the clock)"]
    TR["transcript.ts: word reveal"]
    PR["presenter.ts: run end, R1-R4"]
    PV["plan.ts / chips.ts: guide<br/>review-plan.ts: coaching"]
  end
  IMP -->|"POST /api/calls/import (force)"| W
  EXP -->|"D1 reads · POST curation"| W
  PRE -->|"PUT precall (service token)"| W
  POL -->|"POST /api/policy · eval-runs · publish"| W
  EVAL --> W
  W --> D1
  DO --> D1
  DO -->|"≤ 2 subrequests per decision"| AI
  TR --- PR --- PV
  WEB <-->|"WS decision{plan: items, facts[source], known, must_says, dropped_hard}"| DO

Key components. Line numbers are at 51dbb7e (code unchanged since a83e53e). The symbol is the anchor: executors re-locate by symbol if a line moves.

Area Existing (symbol, file:line) M1b change Stories
Stitch STITCH_V1 src/ingest/stitch.ts:53; maxWords: 120 :59; DECISION_MIN_WORDS = 4 :70; merge step :160-171; parts loop :186-213 (every part keeps t: turn.t); output order :230-233 STITCH_V2; each part's t is the previous part's t_end 103
Import DERIVED_TABLES src/routes/calls-import.ts:24 (a forced import deletes decisions, marks, moments, labels, answers, move_answers, rewrites); callRowValues :133 writes pd_ct_id from the payload; contentHash src/ingest/import-payload.ts:168 covers stitch_version; buildFixturePayload scripts/lib/fixture-payload.mjs:24 (no pd_ct_id); momentId src/engine/moments.ts:114 embeds start_i; labelsOf src/session/checkpoint.ts:279 reads gold_json, which embeds i curation export, i remap, an atomic curation precondition, a validated restore route 104a, 104b, 104c, 104e
Ingest stitch(roled) src/routes/ingest.ts:381; stitch(applyRoles(…)) src/routes/speaker-map.ts:308; INGEST src/policy/build.ts:119, bound into policy_hash by hashPolicy (src/policy/canonical.ts) uploads stitched under v2 (104d); INGEST.stitch_version 2 in the 106 candidate 104d, 106
Transcript LiveTranscript.update web/src/transcript.ts:207; UtteranceDriver.tick web/src/clock.ts:427 word reveal from u.t 102
Presenter isInsert web/src/presenter.ts:125; nextClientEnd :140; dueFor :153; releaseAt :164; apply :176; LIVE_MIN_DWELL :23 run end, sentence-end points, R1-R4, client-due must-says treated as inserted 105a, 105, 107
Situation gate fitsToAsk src/engine/situation.ts:152; recordFits :172 (idempotent, called from step() and again from updatePlan); retargetFit :205; shownSituationId :228; planRequest src/engine/budget.ts:97 (fact_values drop :121) hysteresis, pre-judge, a prejudge drop step 108a, 108b
Plan COVERED_DWELL = 4 src/engine/plan.ts:69; coveredToAsk :302; updatePlan :354; def.done_when.length :422; episode insert :443-459; asks :490-498; CallPlan src/engine/types.ts:341; FactValue :160 (value: string); FactChip :176 (at: Evidence, non-null); factChips src/engine/fact-values.ts:381 kinds, known, must_says, chip source, seeded evidence at i=-1, reopen, dropped_hard 106, 108a, 110f, 111a, 111b, 113, 117
Concern updateConcern src/engine/concern.ts:91: opens on client_objecting ≥ concern_open (:102), clears on any client line (:120) ≥ 4-word clear, reopen window, served-list gate 113b, 117
Amounts clientAmountsGbp src/engine/amount.ts:385; recordAmounts :362; when_amount_below_gbp in src/engine/hero.ts:40 (applies unless some client amount ≥ threshold) a correction replaces the corrected amount; a second leg is kept (decision 2e) 111c
State initialState(p: Pick<Policy,'scenario'|'version'>) src/engine/state.ts:17. Callers: src/session/handlers/load.ts:131, seek.ts:55, src/session/checkpoint.ts:254 (rehydrate, with a PolicyPin), src/engine/recompute.ts:73 (pure; called by src/session/handlers/set_weights.ts:61). buildJevState :63 initialState(pin, seed?), with the seed carried in DO meta and sessions.precall_json 113
Policy policy/src/plan-onboarding.json: templates (first_call_after_signup ten items; the per-template text ids plan.template.<template>.<item>.*), fact_slots, purpose trigger :176, purpose_max_decisions 8. rules-*.json: checklist :20-104, thresholds :251-277, token_budget.window 12. Bank onboarding.json: onboarding_objection_type :221, fact_lists.countries (35 ids incl. SA). customer-success.json cs_objection_type :116. shared.json client_objecting. policy/proposals/2026-09-27.json: Stevan's approved texts kinds, bank_country, served_countries, thresholds, criteria, the approved wording 106, 108a, 110b, 110c, 110d, 116, 117, 119
Eval L6_LIMITS scripts/lib/eval-l6.ts:21 (cost_per_call_usd: 0.05); the L4 partition in scripts/lib/eval-l4.ts; engine checks in scripts/lib/eval-l1.ts (checklist_unchanged :174); eval/labelled/{onboarding,cs,exemplar-moments,must-say}.json; scripts/demo-check.mjs:53 reads stitched/<call>.v1.json gate 0.06; retired ids in L4; the fact_value check; demo check on v2 104b, 104e, 106, 110d, 110e
Replay tooling paceDelays scripts/lib/replay-client.ts:41 (caps each gap at MAX_GAP_S even at 1x); scripts/replay-ws.mjs (--policy-version, --session; its summary prints reconnects and max_input_tokens); scripts/ws-probe.mjs (--resume; flags live in scripts/lib/probe-flags/); scripts/probe-player.mjs; scripts/design/shoot-dashboard.mjs (--deployed, --play) --wall-clock and --log; presenter-sim.mjs; --count-points and --known; --at, --until, --expect and --expect-none; --check status-slot 102a, 105a, 108b, 110f, 118a, 118b, 120
Player catchingUpLabel web/src/player.ts:71 reserved status slot 120

Data flow for one replayed part. Under v2, a 44.5 s client turn arrives as 4-6 kind:'turn' parts, each with its own t and t_end, so the protocol, the DO and the decision table are unchanged.

  1. The browser clock passes a part's u.t: its row appears and grows word by word.
  2. At the part's t_end, the browser sends utterance{i}.
  3. The DO builds the state: seeded facts at i=-1, a 16-unit window, and the pre-judge questions on client parts.
  4. The DO answers from the cache or from Jev and runs step(): kinds, known facts, must-say applies/due, hysteresis, and the concern guard and reopen.
  5. It commits in ≤ 2 subrequests and sends decision.
  6. The presenter applies ticks, chips, the caution and any new point on the open card as the decision arrives. It applies a card switch at the client's run end, after the 20 s dwell, and never during a rep turn.
  7. R1-R4 hold what is already on screen.

Integration points.

Conventions. The M1 conventions in prd.md §3 still apply: test commands, with-cf-env for live commands, pseudonymised fixture envelopes, replacement semantics, and null for unknown. M1b adds:

  1. Deploy and probe. node scripts/deploy.mjs refuses a dirty tree or a branch behind main, and deploys with --var GIT_SHA:<HEAD>. Every live script asserts GET /api/health version == git rev-parse HEAD first: replay-ws.mjs, ws-probe.mjs, eval.mjs, probe-player.mjs, presenter-sim.mjs, m1b-acceptance-check.mjs, and shoot-dashboard.mjs --deployed.

    • Commit sequence, per story:

      1. Commit the code. That commit is the deployed revision.
      2. node scripts/deploy.mjs.
      3. Run the live steps.
      4. Make at most one derived-artefact commit. It holds only what the live run produced or pinned: screenshots, eval/reports/*, docs/reports/*, docs/verification/*, and test pins, remap fixtures and label remaps derived from the run. It touches no runtime file under src/, web/, scripts/, policy/ or migrations/.
      5. Run npm run typecheck and npx vitest run on the final commit; both pass.

      This is M1's practice, and it is what makes health.version meaningful.

  2. Headless criteria only. Every criterion must be satisfiable by the executor under the jev-scripts service token. None needs a signed-in human, a Cloudflare dashboard figure, or a reply or sign-off from Stevan; those are in §5.6. In M1, COPILOT-023, 054, 055, 060 and 072 had to be amended mid-run because they did.

  3. Token caps on session requests only. The 12,000 cap is asserted on replay and eval max_input_tokens, and on the worst case (tests/unit/worst-case.ts, via npx vitest run tests/unit/engine-groups.test.ts tests/unit/policy-build.test.ts). Whole-bank figures and probe-questions.mjs figures are information only.

  4. Policy candidates. A story that changes policy/src (the policy chain in HOT FILES) follows these steps:

    1. Regenerate what the change makes stale: the policy fixtures (the scripts/build-policy.mjs header commands), the stub answers (node scripts/gen-stub-answers.mjs), policy/REDPEN.md (node scripts/redpen.mjs) and policy/LABELS.md (node scripts/labels-sheet.mjs).
    2. Build both scenarios with node scripts/build-policy.mjs --scenario <s> --next-version --out policy/build/<s>.v<N>.json. Never use a literal version. --next-version reads the tracked policy/build/, so no server is needed.
    3. Commit the code, the sources and the bundles together, then deploy, so the Worker's loader accepts the new shapes before any upload.
    4. Upload each draft with node scripts/publish-policy.mjs --scenario <s> --version <N>. It answers 409 not_evaluated, which is expected.
    5. Run the story's live probes pinned to the candidate (--policy-version <N>).
    6. Promote only through the gate: run node scripts/eval.mjs --scenario both --layers l1,l4,l6 --policy-version <N> --fresh --run <story>-<yyyymmdd>, then publish-policy.mjs --scenario <s> --version <N> for both scenarios. The gate is L1 ≥ 90 % per scenario, plus L4 and L6.
    7. List every case that passed in eval/reports/final-20260927-{onboarding,cs}.l1.json and now fails, with its score.

    A failed candidate stays a draft, and the story stays passes:false with its failing cases. No threshold is lowered and no label is edited to make a gate pass.

  5. New policy fields are optional. Every field M1b adds to src/policy/types.ts is optional; absent means M1 behaviour. This covers kind, applies/due/satisfied_by/surface_when, texts_from, known_note, view_labels, situation_fit_hard, concern_reopen_window_s, served_countries, priors and the prejudge drop step. New code deployed before its candidate is published therefore never breaks the published bundle. Each story tests its field absent on a clone of the committed fixture.

  6. Times, not line numbers, in live probes. Line numbers in this PRD are transcript_rev 1 (the v1 stitch); after 104e the four samples are v2, with new i. Unit tests on tests/fixtures/stitched/<call>.v1.json keep the v1 numbers. Live probes select by call time (utterances.t), or through tests/fixtures/remap/<call>.json (104e), never by a literal i.

  7. UI evidence from the deployed dashboard. Use node scripts/design/shoot-dashboard.mjs docs/design/evidence '<shots>' --deployed --at <mm:ss> --expect '<selector>', where each shot is [name, query, scheme, 1440, 900] and --at/--until/--expect/--expect-none come from 102a. Take a light shot and a dark shot; the review view adds view=review to the query. Before shooting, run a cache-warming replay-ws without --eval-uid, so the dashboard's session hits the same answers. A state that depends on Jev is shot with --until '<selector>', which pauses the page's own session the first time that state renders. A candidate story shoots after it publishes, because the page's session uses the published policy.

  8. Fixtures. The new kinds curation, decisions and remap join FIXTURE_KINDS (scripts/lib/fixtures.ts) and the schema-driven tests/unit/fixtures-clean.test.ts. They hold pseudonymised text only. tests/fixtures/exports/ keeps its call-coach exports. Test-only inputs that are not fixtures (a sample approvals reply, synthetic transcripts) are built inside the test and written to a temporary directory.

  9. Status notes. Block reasons and status notes use plain words, with no apostrophes or shell metacharacters. In M1 an apostrophe crashed otto (fixed in otto ab20ebb); nothing relies on that fix.

  10. Rep-facing words.

    • A policy text (item titles, hints and notes, plan.* labels, chips, situation texts) renders only when the pinned bundle holds it approved (INV-COPILOT-003, INV-COPILOT-015).
    • UI chrome is the fixed strings in web/src named by docs/design/DESIGN.md: pane and section names ("Now", "Coming up", "Covered", "Plan for this call" today), review-panel labels, and the rule reasons listed in DESIGN.md §Review view. Chrome is never built from call data, so it needs no red pen, as today.

HOT FILES. Stories that edit the same file are ordered by dependencies. Every edit is additive: append fields, rules and registrations; never reorder or rename. The run is sequential: otto implement, one story per iteration, as in M1.

File or group Stories, in order Rule
Policy chain: policy/src/*.json, tests/fixtures/policy/*, tests/fixtures/answers/*, policy/REDPEN.md, policy/LABELS.md 110c → 110d → 117 → 119 → 106 → 108a → 108b → 110b → 116 one candidate per story; rebase, regenerate, then --next-version
src/policy/{types,build,loader,text,canonical}.ts, src/policy/fact-slots.ts 110a (grammar, text identity and registry) → 110h (derived checklist, hash ownership) → 110c (the apply tool) → 110d → 117 and 108a (thresholds) → 110b (slot) → 116 (priors) additive types; the builder validates the new shapes
scripts/redpen.mjs, scripts/labels-sheet.mjs 110g → 110c 110g resolves texts_from and the derived checklist through the builder's helpers
web/src/review-plan.ts 110a (mustSayOf reader) → 115b
web/src/presenter.ts 105 → 107 105 owns the change moment; 107 owns the lifetime state and the client-due insert clause
web/src/plan.ts, web/src/shell.ts 105 → 107 → 117 → 114 105 adds data-midrun; 107 the last-suggestion line; 117 the reopened state word; 114 restructures the pane
web/src/chips.ts, web/src/render.ts 110b (data-slot, the country template) → 114b (data-source, no link when at.i < 0, suffix)
src/engine/plan.ts, src/engine/types.ts 110a (readers: done_when, applies) → 110f (types) → 117 → 111a → 111b → 113 (templateFor, the seed and its conflict); 108a adds only dropped_hard, and 108b only the pre-judge wiring typed fields appended
src/engine/concern.ts 117 → 113b
src/engine/amount.ts, and src/engine/fact-values.ts for the amount reconcile 111c
src/session/protocol.ts 110f re-exports the additive CallPlan fields only
src/routes/calls-import.ts 104b (curation precondition, pd_ct_id) → 112b (precall provenance and preservation)
src/index.ts 104c (curation route), 112b (precall route) one dispatch line each
scripts/lib/probe-flags/index.ts 110f (--known), 108b (--count-points) one handler file and one registration line each
scripts/design/shoot-dashboard.mjs 102a adds --at, --until, --expect and --expect-none; later stories only call it
scripts/lib/replay-client.ts, scripts/replay-ws.mjs 118a summary fields and flags appended
scripts/presenter-sim.mjs 105a → 107 → 118b output fields appended
docs/design/DESIGN.md 102 §Transcript pane; 105 → 107 → 111b §Cadence rules; 113 → 114b §Client facts chips; 117 §Plan item rows → 114 §Plan pane; 115b §Review view; 120 §Player and §States; 121 §Variants disjoint sections
eval/labelled/*.json 104e (must-say.json i remap), 110e → 116 (onboarding.json), 119 (exemplar-moments.json) append cases; labels are ground truth
scripts/lib/eval-l4.ts, scripts/lib/eval-l6.ts, scripts/lib/eval-l1.ts 118a (wall-clock 1x sample) → 110g (L4 retired) → 106 (L6 constant) → 110e (fact_value) 118a precedes the first gate (110c)
docs/reports/cost-latency.md 104e, 106, 108b, 118 append a section each
migrations/ 112b applies 0004_precall.sql, which 112 adds (0001-0003 exist). prd-m2-corpus.md also plans a 0004 and must renumber (§7 Q6)

4. Product Invariants

INV-COPILOT-001 to INV-COPILOT-014 carry forward unchanged; the full text is in prd.md §4, and stories cite them by id.

New in M1b:

5. User Stories

Every story is one code commit of one concern, plus at most one derived-artefact commit (§3 convention 1). "Estimate" counts otto iterations. A story with an eval gate or a Jev-dependent probe also carries a one-iteration retry budget, reported separately in the sizing table. Line numbers are v1 (transcript_rev 1) unless a criterion says otherwise. Ordering reasons for dependencies are in each story's Notes. Completed M1 stories on main (COPILOT-046, 058e, 065, 070 and the rest) are prerequisites, never dependency ids.

Phase A — Live granularity in replay

Phase B — Card lifetime

Phase C — Plan semantics and pre-call facts

Phase C — Plan semantics and pre-call facts