Latest evaluation run
Run final-20260927, recorded 2026-09-27T07:19:28.104Z against Worker version 2f03062938a0, layers L4 and L6, replayed at 20× with fresh model answers. Rendered from eval/reports/final-20260927.l4l6.json by scripts/render-eval-report.mjs; the newest committed run report is the one published.
Summary
| layer | result | checks passed |
|---|---|---|
| L4 replay against labels | PASS | 33/33 |
| L6 cost and latency | PASS | 6/6 |
| L1 customer_success cases | 79/81 cases (97.5%) | 164/166 |
| L1 onboarding cases | 91/91 cases (100.0%) | 200/200 |
Policies
| scenario | policy version | expected model |
|---|---|---|
| customer_success | 4 | jev-1.13.0 |
| onboarding | 4 | jev-1.13.0 |
Replayed calls
| role | call | scenario | decision points | requests | errors | cards shown | completeness | next step agreed |
|---|---|---|---|---|---|---|---|---|
| onboarding_strong | 3339895706 | onboarding | 96 | 96 | 0 | 48 | 0.286 | yes |
| onboarding_weak | 3485591407 | onboarding | 50 | 50 | 0 | 32 | 0.143 | no |
| cs_strong | 3303259297 | customer_success | 131 | 131 | 0 | 111 | – | – |
| cs_weak | 3347356034 | customer_success | 22 | 22 | 0 | 22 | – | – |
Completeness is must-say items done over items applicable for onboarding; customer-success calls have no must-say list.
L4: replay against labels
| role | check | result | detail |
|---|---|---|---|
| onboarding_strong | finished | PASS | every utterance answered |
| onboarding_strong | zero_errors | PASS | errors 0 |
| onboarding_strong | one_model | PASS | models [jev-1.13.0] |
| onboarding_strong | requests_eq_decision_points | PASS | requests 96, decision points 96 |
| onboarding_strong | card_shown_share | PASS | card shown on 48/96 (50.0%) |
| onboarding_strong | recall | PASS | recall 2/2 (1.000, ≥ 1) |
| onboarding_strong | precision | PASS | precision 1.000, no id outside said ∪ not_applicable persisted |
| onboarding_strong | completeness | PASS | engine completeness 0.286 (done 2 / applicable 7 = 0.286), label completeness 0.286 |
| onboarding_strong | next_step_agreed | PASS | next_step_agreed persisted true, labelled true |
| onboarding_strong | risk_flags | PASS | risk flags [], labelled [] |
| onboarding_weak | finished | PASS | every utterance answered |
| onboarding_weak | zero_errors | PASS | errors 0 |
| onboarding_weak | one_model | PASS | models [jev-1.13.0] |
| onboarding_weak | requests_eq_decision_points | PASS | requests 50, decision points 50 |
| onboarding_weak | card_shown_share | PASS | card shown on 32/50 (64.0%) |
| onboarding_weak | recall | PASS | recall 1/1 (1.000, ≥ 1) |
| onboarding_weak | precision | PASS | precision 1.000, no id outside said ∪ not_applicable persisted |
| onboarding_weak | completeness | PASS | engine completeness 0.143 (done 1 / applicable 7 = 0.143), label completeness 0.143 |
| onboarding_weak | next_step_agreed | PASS | next_step_agreed persisted false, labelled false |
| onboarding_weak | risk_flags | PASS | risk flags [], labelled [] |
| onboarding_weak | below_strong | PASS | completeness 0.143 (< strong's 0.286; labels 0.143 < 0.286) |
| cs_strong | finished | PASS | every utterance answered |
| cs_strong | zero_errors | PASS | errors 0 |
| cs_strong | one_model | PASS | models [jev-1.13.0] |
| cs_strong | requests_eq_decision_points | PASS | requests 131, decision points 131 |
| cs_strong | card_shown_share | PASS | card shown on 111/131 (84.7%) |
| cs_strong | resolved | PASS | resolution resolved, reached [owned, reported, resolved] |
| cs_weak | finished | PASS | every utterance answered |
| cs_weak | zero_errors | PASS | errors 0 |
| cs_weak | one_model | PASS | models [jev-1.13.0] |
| cs_weak | requests_eq_decision_points | PASS | requests 22, decision points 22 |
| cs_weak | card_shown_share | PASS | card shown on 22/22 (100.0%) |
| cs_weak | never_owned | PASS | reached [reported] |
L6: cost and latency
| measure | value |
|---|---|
| requests | 299 (0 without usage) |
| max input tokens | 10,511 |
| p95 input tokens | 9,395 |
| p95 provider latency | 498 ms |
| p95 end-to-end latency | 457 ms over 22 decision points of 3347356034 at 1× |
| highest cost per call | $0.0396 |
| check | result | detail |
|---|---|---|
| requests_measured | PASS | 299 requests with usage |
| max_input_tokens | PASS | max 10511 (≤ 12000) |
| p95_latency | PASS | p95 provider latency 498 ms (≤ 1500) |
| cost_per_call | PASS | 3303259297 $0.0396, 3339895706 $0.0345, 3347356034 $0.0068, 3485591407 $0.0172 (each ≤ $0.05) |
| model_expected | PASS | models [jev-1.13.0] (== jev-1.13.0) |
| e2e_p95 | PASS | e2e p95 457 ms (gate ≤ 3000, target ≤ 1500) over 22 decision points of 3347356034 at 1× |
L1 customer_success: single-decision cases
Run final-20260927-cs, policy version 4, model jev-1.13.0: 79 of 81 cases pass (97.5%), 164 of 166 checks.
| failing case | failed checks |
|---|---|
| lfl_cs_payment_mentioned_no_amount | fits::payment_mentioned_no_amount >= 0.6 |
| lfl_cs_several_payments_at_once | fits::several_payments_at_once >= 0.6 |
L1 onboarding: single-decision cases
Run final-20260927-onboarding, policy version 4, model jev-1.13.0: 91 of 91 cases pass (100.0%), 200 of 200 checks.