Insights D3: ratified defaults in generation
Arm A averages 2.33 correct completions per 122 attempts, 30.00 silent wrong answers and 39.00 refusals. Arm B averages 0.33 correct completions, 6.33 silent wrong answers and 76.67 refusals. Mean precision is 7.15% for A and 4.17% for B. The paired evidence and each acceptance target are reported below.
Current zai generator, measured with and without the D2 verifier under the corrected oracle. The fresh-v1-r2 slice has 122 attempts, 116 supported first turns, 31 families and only three conversations. It was previously exposed in D2; it is not a newly unseen holdout.
Generation and verification
Five owner defaults now have a mandatory generation-context lane with physical predicates/mappings: CT staff/house exclusions, Pipedrive non-employee people, CT-first person identity, CT-first source routing, and London calendar boundaries. Both federated leg prompts receive the same rules. Selection of the older founder-pinned curriculum remains unchanged. Explicit staff inclusion suppresses the exclusion and removes conflicting retrieved exclusion records. The final implementation names applicable defaults as unproved semantic obligations even with the verifier off. They are not treated as proved by a verifier acceptance. The tables measure frozen revision `2513a718`: its approved/federated execution path did not yet emit the added default records. The later recording correction and its offline non-interference checks are disclosed in the measurement addendum; the retained attempts remain the original measured records.
Tests: `test_assembled_prompt_contains_every_default_and_physical_rule`, `test_federated_leg_defaults_survive_zero_curriculum_budget`, `test_missing_staff_exclusion_is_unproved_with_verifier_off`, `test_explicit_include_staff_suppresses_default`, `test_do_not_include_staff_keeps_default`, `test_explicit_staff_override_removes_retrieved_exclusion_only`, `test_resolver_finds_each_record_by_name_and_synonym`, `test_default_obligations_follow_physical_cohort_without_people_words`, `test_approved_federated_leg_records_unproved_defaults_without_regeneration`, `test_recording_defaults_preserves_verifier_packet_and_prior_turn_view`. Founder selections are covered by `test_generation_curriculum.py`.
Fresh slice
Each cell is the mean [minimum, maximum] of three complete repeats. Fractional counts are repeat means. Precision = correct completed / (correct completed + silent wrong), excluding comparison-unverified answers. Across all repeats, A has 97 scored answers and 64 comparison-unverified answers; B has 20 and 27. Correct-completion coverage is 7/366 (1.91%) for A and 1/366 (0.27%) for B.
| Metric | D2 baseline | Arm A: defaults | Arm B: defaults + verifier |
|---|---|---|---|
| First-turn execution errors | 31.33 [30.00, 32.00] | 28.00 [27.00, 30.00] | 28.00 [25.00, 31.00] |
| Correct completed | 1.33 [0.00, 3.00] | 2.33 [1.00, 3.00] | 0.33 [0.00, 1.00] |
| Silent wrong | 43.33 [39.00, 47.00] | 30.00 [27.00, 34.00] | 6.33 [6.00, 7.00] |
| First-turn error rate | 27.01% [25.86%, 27.59%] | 24.14% [23.28%, 25.86%] | 24.14% [21.55%, 26.72%] |
| Precision among scored answers | 3.12% [0.00%, 7.14%] | 7.15% [3.33%, 10.00%] | 4.17% [0.00%, 12.50%] |
| Refusals | 9.00 [7.00, 10.00] | 39.00 [38.00, 41.00] | 76.67 [71.00, 80.00] |
| Guard refusals | 9.00 [7.00, 10.00] | 39.00 [38.00, 41.00] | 39.67 [36.00, 45.00] |
| Verifier refusals | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 37.00 [34.00, 42.00] |
| Comparison unverified | 35.67 [34.00, 37.00] | 21.33 [19.00, 23.00] | 9.00 [7.00, 10.00] |
| Full conversations | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] |
| p95 question latency (s) | 21.08 [20.20, 22.77] | 30.42 [27.29, 34.54] | 49.56 [43.77, 56.24] |
| verifier accept | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 15.67 [13.00, 18.00] |
| verifier reject | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 44.67 [41.00, 51.00] |
| verifier clarify | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.33 [0.00, 1.00] |
| Operational verifier verdicts (subset) | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 4.33 [2.00, 6.00] |
| Repair attempts (observed) | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 17.67 [16.00, 19.00] |
| Repair attempts (persisted) | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 16.00 [15.00, 17.00] |
| Repairs applied (persisted) | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 6.33 [6.00, 7.00] |
| Repairs with unrecorded application state | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 1.67 [1.00, 2.00] |
| Jobs with unresolved repair-attempt state | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] |
Verifier accept/reject/clarify totals include every returned verdict, even when a later turn error prevented persistence. Operational verdicts are a subset of these totals, not semantic defect detections. Observed repair attempts include persisted attempts plus generator dispatches after rejection when a later error dropped the application trace. Applied counts are persisted only; unrecorded application state and unresolved attempt state are separate. Applied does not mean correct. All verdict records and provider calls remain separately countable. The 19 persisted applied repairs on fresh end in 14 comparison-unverified refusals and five silent-wrong answers; none is correct. All D3 refusals have the refusal_unverified outcome: the oracle has not independently justified their refusal.
Paired differences and observed noise
Positive differences mean more of the named metric. Intervals resample whole families (10,000 draws, seed 2026090611). Counts are normalized to this slice: 122 attempts, 116 first turns, three conversations. Precision differences are percentage points. Precision pools counts before division; latency pools per-question times before taking p95. Their paired points can differ from differences of mean repeat statistics. A spread is an observed range, not a confidence bound. B−A uses A’s spread; historical comparisons use D2’s.
| Metric | B−A [95% CI] | A−D2 [95% CI] | B−D2 [95% CI] | A spread | D2 spread |
|---|---|---|---|---|---|
| First-turn execution errors | 0.00 [-4.00, 4.33] | -3.33 [-14.33, 6.67] | -3.33 [-11.87, 4.79] | 3.00 | 2.00 |
| Correct completed | -2.00 [-3.47, -0.67] | 1.00 [0.00, 2.11] | -1.00 [-2.21, 0.00] | 2.00 | 3.00 |
| Silent wrong | -23.67 [-33.39, -14.30] | -13.33 [-21.37, -4.82] | -37.00 [-46.28, -27.84] | 7.00 | 8.00 |
| Refusals | 37.67 [25.38, 49.52] | 30.00 [20.33, 39.98] | 67.67 [55.19, 79.33] | 3.00 | 3.00 |
| Guard refusals | 0.67 [-6.44, 7.33] | 30.00 [20.33, 39.98] | 30.67 [21.35, 40.00] | 3.00 | 3.00 |
| Verifier refusals | 37.00 [27.22, 47.63] | 0.00 [0.00, 0.00] | 37.00 [27.22, 47.63] | 0.00 | 0.00 |
| Comparison unverified | -12.33 [-17.73, -7.12] | -14.33 [-21.00, -7.79] | -26.67 [-36.08, -17.67] | 4.00 | 3.00 |
| Full conversations | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.00 | 0.00 |
| Precision among scored answers | -2.22% [-9.68%, 11.67%] | 4.23% [0.98%, 9.73%] | 2.01% [-3.38%, 16.43%] | 6.67% | 7.14% |
| p95 question latency (s) | 18.47 [13.05, 25.32] | 9.26 [4.37, 13.67] | 27.73 [22.75, 33.11] | 7.25 | 2.57 |
The degenerate full-conversation intervals describe these three conversations only. Undefined bootstrap draws (including samples containing no conversations) are counted in results.json, never replaced by zero.
Pass line
The final verdict retains D2’s composite rule: absolute thresholds plus the specified observed-spread and paired-interval evidence. A threshold pass can therefore still have a final FAIL. Three conversations are too few to establish the 20/50 target. The 30-second latency target is expected to fail for the verifier and is not waived.
| Target | A threshold | A final | B threshold | B final |
|---|---|---|---|---|
| precision >=80% | FAIL | FAIL | FAIL | FAIL |
| correct completed >=D2 | PASS | PASS | FAIL | FAIL |
| silent wrong <=half D2 | FAIL | FAIL | PASS | PASS |
| first-turn errors <=25/100 | PASS | FAIL | PASS | FAIL |
| full conversations >=20/50 | FAIL | FAIL | FAIL | FAIL |
| Zero observed failures in executed, present curated checks | PASS | PASS | PASS | PASS |
| every repeat p95 <=30s | FAIL | FAIL | FAIL | FAIL |
Skipped checks and the 19 absent inherited files remain unverified.
Real-usage sample
101 attempts, 86 supported first turns, five conversations; one repeat per arm. No repeat noise estimate is available.
| Metric | D2 | A | B |
|---|---|---|---|
| First-turn execution errors | 10.00 [10.00, 10.00] | 10.00 [10.00, 10.00] | 14.00 [14.00, 14.00] |
| Correct completed | 3.00 [3.00, 3.00] | 1.00 [1.00, 1.00] | 1.00 [1.00, 1.00] |
| Silent wrong | 35.00 [35.00, 35.00] | 22.00 [22.00, 22.00] | 2.00 [2.00, 2.00] |
| First-turn error rate | 11.63% [11.63%, 11.63%] | 11.63% [11.63%, 11.63%] | 16.28% [16.28%, 16.28%] |
| Precision among scored answers | 7.89% [7.89%, 7.89%] | 4.35% [4.35%, 4.35%] | 33.33% [33.33%, 33.33%] |
| Refusals | 10.00 [10.00, 10.00] | 36.00 [36.00, 36.00] | 67.00 [67.00, 67.00] |
| Guard refusals | 10.00 [10.00, 10.00] | 36.00 [36.00, 36.00] | 28.00 [28.00, 28.00] |
| Verifier refusals | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 39.00 [39.00, 39.00] |
| Comparison unverified | 39.00 [39.00, 39.00] | 25.00 [25.00, 25.00] | 13.00 [13.00, 13.00] |
| Full conversations | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] |
| p95 question latency (s) | 22.34 [22.34, 22.34] | 31.20 [31.20, 31.20] | 53.54 [53.54, 53.54] |
| verifier accept | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 16.00 [16.00, 16.00] |
| verifier reject | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 57.00 [57.00, 57.00] |
| verifier clarify | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] |
| Operational verifier verdicts (subset) | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 3.00 [3.00, 3.00] |
| Repair attempts (observed) | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 28.00 [28.00, 28.00] |
| Repair attempts (persisted) | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 26.00 [26.00, 26.00] |
| Repairs applied (persisted) | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 16.00 [16.00, 16.00] |
| Repairs with unrecorded application state | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 2.00 [2.00, 2.00] |
| Jobs with unresolved repair-attempt state | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] |
The 16 persisted applied repairs on real usage end in 13 comparison-unverified refusals, two comparison-unverified answers and one silent-wrong answer; none is correct. These final outcomes do not exclude partial corrections within an answer.
Paired real-usage intervals resample whole families within these single runs. They describe variation across sampled families, with no estimate of run-to-run noise. Count differences are normalized to 101 attempts, 86 first turns or five conversations. Of 10,000 draws, 472 have undefined precision in each B contrast; 46 have no conversations in each contrast. Those draws are excluded from the corresponding intervals and retained in the undefined-draw counts.
| Metric | B−A [95% CI] | A−D2 [95% CI] | B−D2 [95% CI] |
|---|---|---|---|
| First-turn execution errors | 4.00 [-1.16, 9.77] | 0.00 [-6.62, 6.06] | 4.00 [-3.25, 11.54] |
| Correct completed | 0.00 [0.00, 0.00] | -2.00 [-4.90, 0.00] | -2.00 [-4.90, 0.00] |
| Silent wrong | -20.00 [-30.00, -11.22] | -13.00 [-24.73, -2.91] | -33.00 [-43.77, -22.67] |
| Refusals | 31.00 [19.71, 41.64] | 26.00 [14.99, 37.31] | 57.00 [47.21, 65.65] |
| Guard refusals | -8.00 [-15.81, -0.94] | 26.00 [14.99, 37.31] | 18.00 [8.74, 27.97] |
| Verifier refusals | 39.00 [28.98, 48.48] | 0.00 [0.00, 0.00] | 39.00 [28.98, 48.48] |
| Comparison unverified | -12.00 [-22.33, -2.10] | -14.00 [-21.87, -5.61] | -26.00 [-36.73, -14.92] |
| Full conversations | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] |
| Precision among scored answers | 28.99% [0.00%, 95.45%] | -3.55% [-10.26%, 3.12%] | 25.44% [-9.76%, 92.86%] |
| p95 question latency (s) | 22.34 [9.14, 31.89] | 8.86 [0.99, 18.60] | 31.20 [20.07, 42.03] |
Outcome transitions from D2
Mean counts per 122 attempts, equally weighting all nine case-matched repeat pairs. These are descriptive transitions; a change to correct is not proof that a specific repair caused it.
| D2 outcome | D3 outcome | Arm A | Arm B |
|---|---|---|---|
| comparison_unverified | comparison_unverified | 17.00 | 7.56 |
| comparison_unverified | correct_completed | 1.11 | 0.11 |
| comparison_unverified | refusal_unverified | 11.22 | 23.00 |
| comparison_unverified | silent_wrong | 3.11 | 2.00 |
| comparison_unverified | task_error | 3.22 | 3.00 |
| correct_completed | comparison_unverified | 0.11 | 0.11 |
| correct_completed | correct_completed | 0.56 | 0.00 |
| correct_completed | refusal_unverified | 0.22 | 0.89 |
| correct_completed | silent_wrong | 0.11 | 0.00 |
| correct_completed | task_error | 0.33 | 0.33 |
| refusal_unverified | comparison_unverified | 0.22 | 0.22 |
| refusal_unverified | correct_completed | 0.00 | 0.22 |
| refusal_unverified | refusal_unverified | 5.67 | 7.00 |
| refusal_unverified | silent_wrong | 1.00 | 0.22 |
| refusal_unverified | task_error | 2.11 | 1.33 |
| silent_wrong | comparison_unverified | 2.67 | 0.67 |
| silent_wrong | correct_completed | 0.33 | 0.00 |
| silent_wrong | refusal_unverified | 13.56 | 34.67 |
| silent_wrong | silent_wrong | 22.00 | 4.11 |
| silent_wrong | task_error | 4.78 | 3.89 |
| task_error | comparison_unverified | 1.33 | 0.44 |
| task_error | correct_completed | 0.33 | 0.00 |
| task_error | refusal_unverified | 8.33 | 11.11 |
| task_error | silent_wrong | 3.78 | 0.00 |
| task_error | task_error | 18.89 | 21.11 |
Models, calls and timings
Both D3 arms request glm-5.2 on zai; the generator table reports response.model where available. D2 retained only the requested generator ID. Codex reports a CLI banner for gpt-6-astra; that does not independently attest the serving model. Codex reports total tokens only, not input/output.
| Run | Generator calls / failures / unclosed | Generator IDs | Verifier calls / failures / unclosed | Verifier IDs |
|---|---|---|---|---|
| a-fresh-r1 | 248 / 0 / 0 | {'glm-5.3': 248} | 0 / 0 / 0 | {} |
| a-fresh-r2 | 257 / 0 / 0 | {'glm-5.3': 257} | 0 / 0 / 0 | {} |
| a-fresh-r3 | 251 / 0 / 0 | {'glm-5.3': 251} | 0 / 0 / 0 | {} |
| b-fresh-r1 | 267 / 1 / 0 | {'glm-5.3': 266} | 60 / 6 / 0 | {'gpt-6-astra': 54} |
| b-fresh-r2 | 272 / 0 / 0 | {'glm-5.3': 272} | 64 / 5 / 0 | {'gpt-6-astra': 59} |
| b-fresh-r3 | 270 / 0 / 0 | {'glm-5.3': 270} | 57 / 1 / 0 | {'gpt-6-astra': 56} |
| a-real-r1 | 180 / 1 / 0 | {'glm-5.3': 179} | 0 / 0 / 0 | {} |
| b-real-r1 | 208 / 1 / 0 | {'glm-5.3': 207} | 73 / 3 / 0 | {'gpt-6-astra': 70} |
Families and observed defect classes
Arm A averages 2.33 correct completions per 122. The operational limit to address is incomplete family definitions and grounding. The next work is to bind those definitions to executable source mappings and answer contracts; a generator change or the added verifier is not a demonstrated remedy here. This is an operational priority, not an experimental exclusion of model effects: the fixed-generator design does not establish that changing models could not help.
Counts below pool all three Arm A repeats (366 attempts). A family is listed as wrong/error only when at least one attempt is `silent_wrong`, `task_error` or `infrastructure_failure`. Refusals and comparison-unverified answers remain separate. Classes are conservative observations from the saved result fields and terminal evidence; they do not claim an isolated semantic root cause. All 21 PlannerError cases lack retained generated SQL, but that absence does not establish why planning failed. The existing comparator already permits harmless aliases; differing labels alone are not classified as a defect.
| Family | Wrong answers | Execution errors | Correct | Refused | Comparison unverified | Observed classes |
|---|---|---|---|---|---|---|
| F01: Signups by acquisition source | 1 | 0 | 0 | 7 | 4 | Result row-set size differs; cohort/window/series cause not isolated |
| F02: Signups by country or nationality | 1 | 2 | 0 | 6 | 3 | Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed |
| F03: Commission total for a period | 1 | 6 | 0 | 13 | 1 | PlannerError; specific cause unconfirmed; Values, membership or column alignment disagree under the existing comparator; Source schema reference failed; Source SQL execution failed; specific cause unconfirmed |
| F04: Commission by month | 8 | 0 | 0 | 4 | 0 | Month labels differ from the reference axis; window/label cause not isolated; Result row-set size differs; cohort/window/series cause not isolated; Required zero-filled periods are missing from the returned series |
| F10: Quoted but not booked | 1 | 6 | 1 | 6 | 7 | Execution boundary failed; specific cause unconfirmed; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; Source schema reference failed |
| F11: Quote volume and quote-to-trade conversion | 3 | 0 | 0 | 6 | 3 | Result row-set size differs; cohort/window/series cause not isolated; Required zero-filled periods are missing from the returned series |
| F12: Churn tier classification | 8 | 0 | 0 | 0 | 4 | Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated; Values, membership or column alignment disagree under the existing comparator |
| F13: Quiet high-value clients | 3 | 1 | 0 | 1 | 7 | Federated plan contract failed; Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated |
| F14: Activation rate | 3 | 0 | 0 | 5 | 4 | Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated |
| F16: Signup counts by period, client type or stream | 7 | 0 | 0 | 4 | 1 | Result row-set size differs; cohort/window/series cause not isolated; Values, membership or column alignment disagree under the existing comparator; Required zero-filled periods are missing from the returned series |
| F17: Signup quality by source or country | 4 | 0 | 0 | 4 | 1 | Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated; Values, membership or column alignment disagree under the existing comparator |
| F19: Top clients by revenue | 6 | 2 | 0 | 9 | 4 | Missing expected output fields with incompatible projection width; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed |
| F20: Trades today or recent trades | 8 | 1 | 0 | 3 | 0 | Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed |
| F23: Bookings by payment partner or platform | 1 | 1 | 0 | 7 | 0 | PlannerError; specific cause unconfirmed; Values, membership or column alignment disagree under the existing comparator |
| F24: Period-over-period comparison | 2 | 9 | 0 | 0 | 1 | PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed |
| F26: Refer-and-earn programme | 1 | 0 | 1 | 5 | 5 | Result row-set size differs; cohort/window/series cause not isolated; Required zero-filled periods are missing from the returned series |
| F27: Corporate accounts sent to brokers | 0 | 3 | 0 | 6 | 0 | PlannerError; specific cause unconfirmed; Source SQL execution failed; specific cause unconfirmed |
| F31: Client listing with contact details | 6 | 1 | 1 | 0 | 1 | Missing expected output fields with incompatible projection width; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated |
| F33: First-time traders in a month | 2 | 6 | 0 | 2 | 2 | Federated plan contract failed; Values, membership or column alignment disagree under the existing comparator; Source SQL execution failed; specific cause unconfirmed |
| F34: Revenue by acquisition source or campaign | 4 | 0 | 0 | 4 | 1 | Values, membership or column alignment disagree under the existing comparator |
| F38: Sales lead counts | 3 | 5 | 0 | 1 | 0 | Month labels differ from the reference axis; window/label cause not isolated; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed; Required zero-filled periods are missing from the returned series |
| P01: Deal loss reasons ranking | 1 | 4 | 2 | 5 | 0 | Pipedrive tables lack required schema qualification; Result row-set size differs; cohort/window/series cause not isolated; SQL validation failed; specific cause unconfirmed |
| P02: Lost deals listing for one reason | 0 | 8 | 0 | 0 | 1 | Pipedrive tables lack required schema qualification |
| P03: Pipeline status and open deals | 0 | 11 | 1 | 0 | 0 | Pipedrive tables lack required schema qualification |
| P04: Win and loss rate by owner or industry | 0 | 7 | 0 | 0 | 2 | Pipedrive tables lack required schema qualification |
| P05: Sales activity volume and mix | 0 | 11 | 0 | 0 | 1 | Pipedrive tables lack required schema qualification |
| X01: Sales activity per trading client | 5 | 2 | 0 | 0 | 2 | Missing expected output fields with incompatible projection width; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; SQL validation failed; specific cause unconfirmed |
| X02: Deals versus trading reconciliation | 5 | 2 | 0 | 0 | 2 | Missing expected output fields with incompatible projection width; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated |
| X03: Lost reasons crossed with client profile | 6 | 0 | 0 | 3 | 0 | Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated |
Families without observed wrong/error outcomes:
- F22 (Commission by currency corridor): comparison_unverified=4, refusal_unverified=8
- F25 (Professional affiliate performance): comparison_unverified=3, correct_completed=1, refusal_unverified=8
The compact per-attempt classification ledger is `family-defects.jsonl`; the structural classification is reproducible with `scripts/insights_d3_family_audit.py` against full records. A deeper SQL-to-reference audit would be needed to distinguish cohort, measure and calendar causes where only a result mismatch or a generic source execution error is observed.
Measurement addendum and recording correction
All eight live runs completed on 2026-09-09, from 12:52 to 16:05 UTC, with no crashed repeat, restart or quality-motivated rerun. Every D3 application snapshot and evaluator source copy used revision `2513a7186df76693dd4085a38a4aca865fcefa1a`. The six fresh repeats use the same corrected corpus hash as D2; the two real-usage repeats likewise match their D2 corpus. Private copies of both corrected corpora are retained under `/tmp/lore-goal3-eval/goal-d3/oracles/`; `oracle-retention.json` records their hashes. No oracle was edited.
The controller admitted each run after two consecutive five-second samples with at least 95% CPU idle. `run-order-evidence.json` verifies all eight admissions against `host-observations.jsonl` and `sequence.jsonl`. This is admission evidence, not continuous monitoring or proof that no other process used the host during a run. The historical baseline and sequential D3 arms remain confounded with time, host activity and provider conditions.
Frozen results versus the final implementation
While examining completed Arm A records, an audit found that `execute_approved_query`, used by federated legs, omitted the new named default obligations from successful results. The frozen generation prompts already contained the defaults. The pre-measurement implementation review's statement that every execution path recorded them was too broad.
All application source remained unchanged throughout the eight live runs. After the final run and server shutdown, regression tests demonstrated the missing approved-path records and applicability gaps for person-table shorthand, CTE populations and explicit years. Three listing fixtures initially failed the existing LIMIT requirement; their limits were corrected before the recording fix, and both failed runs were retained. The valid fixtures then reproduced seven failures and four passes. After the fix, all eleven focused tests passed.
The final correction, checkpoint `d0bf8f2b`, adds named, unproved default records to approved/federated successful results and improves applicability detection. Revenue, trade and quote totals keep their own statistics filters; merely reading a CT fact table does not add the people-exclusion obligation. Explicit staff inclusion still suppresses it. The correction changes no generation prompt, SQL rewrite, verifier rule, oracle, refusal gate or provider configuration. Default applicability remains a conservative interpretation of question vocabulary and parsed sources, not a complete semantic proof.
`metadata-replay.json` records an offline perturbation of saved answer-record metadata: 934 of 934 original outcomes reproduced, 501 payloads with changed default metadata, zero outcome changes, and 934 unchanged verifier prior-turn views. Every other payload field and every raw attempt file remained unchanged. A focused integration test also compares the complete verifier prompts before and after adding default records. This is an offline non-interference check, not another benchmark or a live run of the final head. The final correction's latency has not been remeasured. The post-correction Python check passed 958 tests across 40 files, with zero failures and 88 skips; its curated subset passed 636 tests. The unchanged frontend retains the earlier 39 passing checks. Nineteen inherited curated files are absent and unverified. Results and source hashes are retained in `checks/final-summary.json` and `checks/final-check-receipt.json`; frozen checks remain separate in `checks/summary.json`.
Complete verdict and repair accounting
The evaluator's wrapper retained every verifier packet and returned verdict independently of application trace persistence. Seven executed jobs lost their application verification trace after a later error: five across the fresh repeats and two in real usage. Their returned non-operational rejections followed by generator dispatches establish seven additional repair attempts. The compact provider and verifier traces retain job IDs and monotonic timings so this inference can be checked independently.
Fresh repair attempts are 16, 19 and 18 (persisted: 15, 17 and 16); real usage has 28 attempts (26 persisted). Applied-repair counts include only persisted application state: six, seven and six on fresh, and sixteen on real usage. Application state for the seven recovered attempts is unrecorded and is not counted as applied or successful. No repair-attempt state remains unresolved under the dispatch rule. All returned accepts, rejects and clarifications are counted, including operational verdicts and turns whose later errors dropped the application trace. Free-form verifier signal text is hashed in compact records; the complete private records preserve it.
None of the persisted applied repairs ended in a correct completion. On fresh, the 19 applied repairs ended in 14 `refusal_unverified` and five `silent_wrong` outcomes. On real usage, the 16 ended in 13 `refusal_unverified`, two `comparison_unverified` and one `silent_wrong` outcome. These are final outcomes, not evidence that no partial correction occurred. Every D3 refusal is `refusal_unverified`; the corrected oracle has not independently justified the refusal.
Sampling labels and prior exposure
The leakage review identified an evaluator metadata omission: the `LiveProvider` wrapper left `LLMTextResult.sampling_pin` at its default. Consequently, 929 per-question SQL-generation summaries and 1,950 nested generation-attempt records say `off`; five questions have no summary label. `sampling-pin-audit.json` separates these counts. Source and retained configuration support requested temperature 0, seed 42 and disabled thinking. The misleading labels neither establish unpinned requests nor independently verify provider compliance.
After review, the wrapper was corrected to propagate `sqlgen_pin_state(config.provider)` into future result metadata. No request argument, measured record or result was changed, and no live evaluation was rerun. This telemetry correction, like the obligation-recording correction, is outside the measured revision. The frozen method has a dated qualification: inherited admission labels saying “no tuning after the slice was opened” establish neither an unseen holdout nor an absence of development after D2 exposure. The supported claim is unchanged source throughout these eight measurements.
Interpretation and reporting additions
Arm A's mean correct-completion gain over D2 is one answer per 122, inside D2's three-answer observed repeat range. Its precision interval is positive, but the gain is inside D2's precision range. Arm B's correct-completion difference against A is minus two: its family interval excludes zero, but the mean difference equals, and does not exceed, A's two-answer repeat range. Neither result establishes recovery beyond both forms of evidence required by the D2 method.
Silent wrong answers fall by 13.33 per 122 in A versus D2 and by 37 in B versus D2; both reductions exceed D2's observed range and have wholly negative paired intervals. Only B meets the requirement to halve them. B also has substantially more refusals. The data supports selective refusal as the main observed verifier contribution here, not a recovery in correct completions. The real-usage precision of 33.33% in B is one correct answer among only three scored answers, with no repeat-noise estimate.
The original paired-count method was extended during reporting to pooled per-question p95 differences and supplementary real-usage family intervals. These additions did not select a model, alter an answer, change a threshold or cause a rerun. Per-arm means/minima/maxima remain repeat statistics; paired precision and p95 pool first and therefore have different estimands. Completed-call latency summaries exclude failed calls; failed-call timings remain in the traces, and the per-question p95 includes all timed attempts.
The D2 baseline's application revision was `c4f76cf069abfc98f10a9b1fe8846555f2117990`. Its `src/` tree matches the fetched D2 tip `4bfc6022`; the intervening D2 changes were reporting/reconciliation work. Its historical served generator identity cannot be reconstructed: only the requested ID was retained. D3 records each completed zai response's `model` field; a Codex CLI banner remains configuration evidence rather than independent serving attestation.
Independent review and correction ledger
Three separate Codex CLI processes reviewed checkpoint `634a75d9044c7e68f81111734642194a60d8d630`, using `gpt-6-astra`, xhigh reasoning and `danger-full-access`, with self-contained prompts and strict no-modification instructions. All three exited successfully; before/after inventories detected zero changed files across 6,467 protected source, evidence and private-record files. Each prompt, reply and receipt is retained here; complete process transcripts remain in the private full-record directory. The CLI banner establishes requested configuration, not independent serving identity.
The reviewed results SHA-256 was `61f72ea449a136a8d50f4314653ef6ce648283358f2acae6660f5a4004a74ef6`; the reviewed Markdown SHA-256 was `e2843fae2e234ac29b9819a8819a0c2f8fba5a07209a92decf66cbd3ccd64d07`. The wording below was corrected after those reviews; it was not submitted for a fourth review.
| Review | Verdict | Confirmed corrections |
|---|---|---|
| Arithmetic and intervals | PASS; no numerical discrepancy | None. Independently reproduced raw structural metrics, all six family-bootstrap comparisons, rendered cells, transitions and repair accounting. |
| Leakage, model identity and noise | No blocking leakage identified; two recording/wording findings | Disclosed the misleading sampling labels, corrected future wrapper telemetry, and qualified inherited exposure labels. |
| Unsupported wording | Revise wording before sign-off; numbers stand | Narrowed the operational diagnosis; left PlannerError causes unresolved; clarified the measurement freeze, executed-check scope, scored-answer precision and applied-repair outcomes. |
1. **Operational diagnosis:** family definitions and executable grounding are the next operational limit to address. This fixed-generator experiment does not causally exclude a model change helping. The report no longer treats that exclusion as established.
2. **Planner errors:** renamed `planner_no_usable_sql` to `planner_error_unresolved` for 21 Arm A attempts. Their retained payloads lack SQL, but a generic PlannerError does not isolate why. This changes only the diagnostic ledger, not oracle outcomes.
3. **Measurement freeze:** appended a dated method qualification. All eight runs measured `2513a718`; later inspection informed the obligation-recording correction at `d0bf8f2b`. The final code was not live-benchmarked. The retained 934-attempt offline metadata replay supports non-interference, not a new latency measurement.
4. **Curated coverage:** the visible PASS label now says “Zero observed failures in executed, present curated checks.” The 88 skipped checks and 19 absent inherited files remain unverified. No acceptance rule was changed.
5. **Precision denominator:** labels now say “Precision among scored answers,” with the formula and separate comparison-unverified counts. Fresh coverage is A 7/366 (1.91%) and B 1/366 (0.27%), distinct from precision.
6. **Repairs and refusals:** added observed final outcomes for all persisted applied repairs: fresh 19 = 14 unverified refusals + five silent wrong; real 16 = 13 unverified refusals + two unverified answers + one silent wrong. None ended correct. Seven additional dispatched repairs retain unknown application state. All D3 refusals remain oracle-unverified.
7. **Sampling evidence:** independently counted 929 misleading per-question summary labels and 1,950 nested attempt labels. Added `sampling-pin-audit.json`; future evaluator result metadata now propagates the configured pin state. Original traces are unchanged; provider compliance is not independently attested.
8. **Exposure:** the slice was previously exposed during D2. Inherited “no tuning after the slice was opened” labels are retained but cannot establish an unseen holdout or absence of all post-D2 development.
No score, interval, transition, threshold, oracle or measured attempt changed as a result of review. Reviewers did not rerun the application, tests or semantic answer scoring; arithmetic review recomputed retained structural outcome fields. All substantive review findings were either confirmed and corrected above or retained as explicit limits.
Method and limitations
D3 method frozen before measurement
- Base: fetched origin/insights/goal-d2 at 4bfc6022. Only /tmp/lore-goal-d3 is writable for this work.
- Arm A: zai, requested glm-5.2 (the retained current generator endpoint), temperature 0, seed 42, thinking disabled, 25 s provider timeout. The provider response model ID is now retained independently of the requested ID.
- Arm B: identical code/generator plus the in-app Codex verifier, requested gpt-6-astra, xhigh, 30 s per call, 60 s turn deadline, definition-scoped rules, one bounded repair and unchanged shared invocation limits. The Codex banner identifies the requested configuration, not an independently attested serving identity.
- Order: A fresh repeats 1–3, B fresh repeats 1–3, A real usage, B real usage. Two workers throughout. No concurrent tests or evaluation arms. A crashed run gets one clean restart with an explicit restart label; the aborted records are retained. No quality-motivated reruns.
- D2 runtime, scorer, fixture, admission and frozen clock: 2026-09-06 12:00 UTC, Europe/London business periods. Corrected fresh-v1-r2 corpus, 122 attempts, 116 supported first turns, three conversations, 31 families. Corrected real usage: 101 attempts, 86 supported first turns, five conversations. No oracle edits.
- No D3 smoke opening. This slice was already exposed during D2, including its aborted run. D3 implements the five owner defaults, without using new slice answers to choose implementation changes. All source/prompt changes stop before measurement.
- Retain every attempt and timing, provider traces, and every verifier packet/verdict (including turns that later error) privately under /tmp/lore-goal3-eval/goal-d3. Export compact records, hashes, counters, conversation membership and timing; exclude question text, SQL, rows and free-text verifier reasoning from committed evidence.
- Baseline for acceptance and historical comparisons: D2 base-fresh-r1/r2/r3, current zai generator without mandatory generation defaults or verifier, rescored under the corrected oracle. Its requested identity only was recorded; a historical served identity cannot be reconstructed.
- Counts: per-repeat mean, min, max. Precision = correct completed / (correct completed + silent wrong); comparison-unverified is excluded. Missing/undefined values stay null.
- Pair whole intent families, with all derivative turns/conversations retained. Pool mean per-family counts across arm repeats, resample 31 families with replacement 10,000 times, seed 2026090611, nearest-rank percentile two-sided 95% intervals. Compute exact correct-completed intervals (not D2's first-turn proxy), count-normalized differences per 122 attempts or 116 first turns; full conversations per three. Precision differences pool counts before division, so they can differ from differences of repeat-mean precision. Undefined draws are counted, not replaced with zero.
- Compare B−A, A−D2, B−D2. Noise floor = observed max−min control repeat spread, not a confidence bound: A controls B−A; D2 controls historical comparisons. Sequential timing/provider conditions remain confounded with arm configuration.
- Pass line retains D2's composite rules: precision >=80%, gain beyond baseline spread and paired lower bound >0; correct completed >=baseline and interval not wholly below zero; silent wrong <=half baseline, reduction beyond spread and upper bound <0; first-turn error rate <=25%, reduction beyond spread and upper bound <0; full-conversation rate >=40%, gain beyond spread and lower bound >0; zero failures in present curated invariant suites; p95 per-question <=30 s on every repeat. Report literal threshold and composite verdict separately. The latency target is expected to fail for the verifier; it is not waived. Three conversations are too few to establish 20/50.
- Transitions: same case IDs, equal-weight Cartesian pairing of all historical/control repeats and candidate repeats, scaled to one 122-attempt repeat. This avoids pretending repeat index supplies a causal pairing. These tables describe outcomes, not proven corrections.
Post-measurement qualification (2026-09-09): the source-change statement above describes the measurement freeze. All eight measured runs remained on `2513a718`; subsequent inspection of saved Arm A records informed the recording correction at `d0bf8f2b`, which was not live-benchmarked. Inherited admission labels saying “no tuning after the slice was opened” do not establish an unseen holdout or the absence of development after D2 exposure. The later sampling-label correction changes future telemetry only; original labels remain in the retained records.
Receipt checks confirm identical frozen D3 application and evaluation-source revisions across all eight runs, two workers with no case limit, and matching corrected-corpus and evaluator hashes with the corresponding D2 baseline. The exact hashes are retained in results.json and each run receipt.
Curated checks: {"corrected_run": {"errors": 0, "failed": 0, "passed": 164, "skipped": 13}, "curated": {"errors": 0, "failed": 0, "passed": 636, "skipped": 88}, "errors": 0, "failed": 0, "initial_run": {"errors": 0, "failed": 9, "passed": 922, "skipped": 88}, "js_failed": 0, "js_passed": 39, "method": "Initial complete selection, then entire affected files replaced by corrected run; all before live evaluation. One prompt-context assertion updated for mandatory defaults; old SQL-only prover checks unchanged.", "missing_inherited_files": 19, "passed": 931, "skipped": 88}.
No application deployment or merge is part of D3. Sequential arms are confounded with time and provider conditions. The oracle is corrected but fixture-based; comparison-unverified answers and the small conversation sample limit conclusions. The staff/test marker remains partial as documented in D2. Named semantic obligations remain unproved; the added context does not provide a complete semantics proof. Host load observations are evidence of conditions, not proof of exclusive host use.
Full records: `/tmp/lore-goal3-eval/goal-d3/`. Compact recomputable records: `docs/evidence/insights-goal-d3-2026-09-09/runs/`. `scripts/insights_d3_report.py --compact` recomputes quality metrics, intervals and transitions from committed records alone. Completed-call token and latency summaries are in results.json; failed-call timings are also retained in the compact traces. Per-question p95 includes all timed attempts.