Insights D3: ratified defaults in generation

Arm A averages 2.33 correct completions per 122 attempts, 30.00 silent wrong answers and 39.00 refusals. Arm B averages 0.33 correct completions, 6.33 silent wrong answers and 76.67 refusals. Mean precision is 7.15% for A and 4.17% for B. The paired evidence and each acceptance target are reported below.

Current zai generator, measured with and without the D2 verifier under the corrected oracle. The fresh-v1-r2 slice has 122 attempts, 116 supported first turns, 31 families and only three conversations. It was previously exposed in D2; it is not a newly unseen holdout.

Generation and verification

Five owner defaults now have a mandatory generation-context lane with physical predicates/mappings: CT staff/house exclusions, Pipedrive non-employee people, CT-first person identity, CT-first source routing, and London calendar boundaries. Both federated leg prompts receive the same rules. Selection of the older founder-pinned curriculum remains unchanged. Explicit staff inclusion suppresses the exclusion and removes conflicting retrieved exclusion records. The final implementation names applicable defaults as unproved semantic obligations even with the verifier off. They are not treated as proved by a verifier acceptance. The tables measure frozen revision `2513a718`: its approved/federated execution path did not yet emit the added default records. The later recording correction and its offline non-interference checks are disclosed in the measurement addendum; the retained attempts remain the original measured records.

Tests: `test_assembled_prompt_contains_every_default_and_physical_rule`, `test_federated_leg_defaults_survive_zero_curriculum_budget`, `test_missing_staff_exclusion_is_unproved_with_verifier_off`, `test_explicit_include_staff_suppresses_default`, `test_do_not_include_staff_keeps_default`, `test_explicit_staff_override_removes_retrieved_exclusion_only`, `test_resolver_finds_each_record_by_name_and_synonym`, `test_default_obligations_follow_physical_cohort_without_people_words`, `test_approved_federated_leg_records_unproved_defaults_without_regeneration`, `test_recording_defaults_preserves_verifier_packet_and_prior_turn_view`. Founder selections are covered by `test_generation_curriculum.py`.

Fresh slice

Each cell is the mean [minimum, maximum] of three complete repeats. Fractional counts are repeat means. Precision = correct completed / (correct completed + silent wrong), excluding comparison-unverified answers. Across all repeats, A has 97 scored answers and 64 comparison-unverified answers; B has 20 and 27. Correct-completion coverage is 7/366 (1.91%) for A and 1/366 (0.27%) for B.

MetricD2 baselineArm A: defaultsArm B: defaults + verifier
First-turn execution errors31.33 [30.00, 32.00]28.00 [27.00, 30.00]28.00 [25.00, 31.00]
Correct completed1.33 [0.00, 3.00]2.33 [1.00, 3.00]0.33 [0.00, 1.00]
Silent wrong43.33 [39.00, 47.00]30.00 [27.00, 34.00]6.33 [6.00, 7.00]
First-turn error rate27.01% [25.86%, 27.59%]24.14% [23.28%, 25.86%]24.14% [21.55%, 26.72%]
Precision among scored answers3.12% [0.00%, 7.14%]7.15% [3.33%, 10.00%]4.17% [0.00%, 12.50%]
Refusals9.00 [7.00, 10.00]39.00 [38.00, 41.00]76.67 [71.00, 80.00]
Guard refusals9.00 [7.00, 10.00]39.00 [38.00, 41.00]39.67 [36.00, 45.00]
Verifier refusals0.00 [0.00, 0.00]0.00 [0.00, 0.00]37.00 [34.00, 42.00]
Comparison unverified35.67 [34.00, 37.00]21.33 [19.00, 23.00]9.00 [7.00, 10.00]
Full conversations0.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
p95 question latency (s)21.08 [20.20, 22.77]30.42 [27.29, 34.54]49.56 [43.77, 56.24]
verifier accept0.00 [0.00, 0.00]0.00 [0.00, 0.00]15.67 [13.00, 18.00]
verifier reject0.00 [0.00, 0.00]0.00 [0.00, 0.00]44.67 [41.00, 51.00]
verifier clarify0.00 [0.00, 0.00]0.00 [0.00, 0.00]0.33 [0.00, 1.00]
Operational verifier verdicts (subset)0.00 [0.00, 0.00]0.00 [0.00, 0.00]4.33 [2.00, 6.00]
Repair attempts (observed)0.00 [0.00, 0.00]0.00 [0.00, 0.00]17.67 [16.00, 19.00]
Repair attempts (persisted)0.00 [0.00, 0.00]0.00 [0.00, 0.00]16.00 [15.00, 17.00]
Repairs applied (persisted)0.00 [0.00, 0.00]0.00 [0.00, 0.00]6.33 [6.00, 7.00]
Repairs with unrecorded application state0.00 [0.00, 0.00]0.00 [0.00, 0.00]1.67 [1.00, 2.00]
Jobs with unresolved repair-attempt state0.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]

Verifier accept/reject/clarify totals include every returned verdict, even when a later turn error prevented persistence. Operational verdicts are a subset of these totals, not semantic defect detections. Observed repair attempts include persisted attempts plus generator dispatches after rejection when a later error dropped the application trace. Applied counts are persisted only; unrecorded application state and unresolved attempt state are separate. Applied does not mean correct. All verdict records and provider calls remain separately countable. The 19 persisted applied repairs on fresh end in 14 comparison-unverified refusals and five silent-wrong answers; none is correct. All D3 refusals have the refusal_unverified outcome: the oracle has not independently justified their refusal.

Paired differences and observed noise

Positive differences mean more of the named metric. Intervals resample whole families (10,000 draws, seed 2026090611). Counts are normalized to this slice: 122 attempts, 116 first turns, three conversations. Precision differences are percentage points. Precision pools counts before division; latency pools per-question times before taking p95. Their paired points can differ from differences of mean repeat statistics. A spread is an observed range, not a confidence bound. B−A uses A’s spread; historical comparisons use D2’s.

MetricB−A [95% CI]A−D2 [95% CI]B−D2 [95% CI]A spreadD2 spread
First-turn execution errors0.00 [-4.00, 4.33]-3.33 [-14.33, 6.67]-3.33 [-11.87, 4.79]3.002.00
Correct completed-2.00 [-3.47, -0.67]1.00 [0.00, 2.11]-1.00 [-2.21, 0.00]2.003.00
Silent wrong-23.67 [-33.39, -14.30]-13.33 [-21.37, -4.82]-37.00 [-46.28, -27.84]7.008.00
Refusals37.67 [25.38, 49.52]30.00 [20.33, 39.98]67.67 [55.19, 79.33]3.003.00
Guard refusals0.67 [-6.44, 7.33]30.00 [20.33, 39.98]30.67 [21.35, 40.00]3.003.00
Verifier refusals37.00 [27.22, 47.63]0.00 [0.00, 0.00]37.00 [27.22, 47.63]0.000.00
Comparison unverified-12.33 [-17.73, -7.12]-14.33 [-21.00, -7.79]-26.67 [-36.08, -17.67]4.003.00
Full conversations0.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]0.000.00
Precision among scored answers-2.22% [-9.68%, 11.67%]4.23% [0.98%, 9.73%]2.01% [-3.38%, 16.43%]6.67%7.14%
p95 question latency (s)18.47 [13.05, 25.32]9.26 [4.37, 13.67]27.73 [22.75, 33.11]7.252.57

The degenerate full-conversation intervals describe these three conversations only. Undefined bootstrap draws (including samples containing no conversations) are counted in results.json, never replaced by zero.

Pass line

The final verdict retains D2’s composite rule: absolute thresholds plus the specified observed-spread and paired-interval evidence. A threshold pass can therefore still have a final FAIL. Three conversations are too few to establish the 20/50 target. The 30-second latency target is expected to fail for the verifier and is not waived.

TargetA thresholdA finalB thresholdB final
precision >=80%FAILFAILFAILFAIL
correct completed >=D2PASSPASSFAILFAIL
silent wrong <=half D2FAILFAILPASSPASS
first-turn errors <=25/100PASSFAILPASSFAIL
full conversations >=20/50FAILFAILFAILFAIL
Zero observed failures in executed, present curated checksPASSPASSPASSPASS
every repeat p95 <=30sFAILFAILFAILFAIL

Skipped checks and the 19 absent inherited files remain unverified.

Real-usage sample

101 attempts, 86 supported first turns, five conversations; one repeat per arm. No repeat noise estimate is available.

MetricD2AB
First-turn execution errors10.00 [10.00, 10.00]10.00 [10.00, 10.00]14.00 [14.00, 14.00]
Correct completed3.00 [3.00, 3.00]1.00 [1.00, 1.00]1.00 [1.00, 1.00]
Silent wrong35.00 [35.00, 35.00]22.00 [22.00, 22.00]2.00 [2.00, 2.00]
First-turn error rate11.63% [11.63%, 11.63%]11.63% [11.63%, 11.63%]16.28% [16.28%, 16.28%]
Precision among scored answers7.89% [7.89%, 7.89%]4.35% [4.35%, 4.35%]33.33% [33.33%, 33.33%]
Refusals10.00 [10.00, 10.00]36.00 [36.00, 36.00]67.00 [67.00, 67.00]
Guard refusals10.00 [10.00, 10.00]36.00 [36.00, 36.00]28.00 [28.00, 28.00]
Verifier refusals0.00 [0.00, 0.00]0.00 [0.00, 0.00]39.00 [39.00, 39.00]
Comparison unverified39.00 [39.00, 39.00]25.00 [25.00, 25.00]13.00 [13.00, 13.00]
Full conversations0.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
p95 question latency (s)22.34 [22.34, 22.34]31.20 [31.20, 31.20]53.54 [53.54, 53.54]
verifier accept0.00 [0.00, 0.00]0.00 [0.00, 0.00]16.00 [16.00, 16.00]
verifier reject0.00 [0.00, 0.00]0.00 [0.00, 0.00]57.00 [57.00, 57.00]
verifier clarify0.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
Operational verifier verdicts (subset)0.00 [0.00, 0.00]0.00 [0.00, 0.00]3.00 [3.00, 3.00]
Repair attempts (observed)0.00 [0.00, 0.00]0.00 [0.00, 0.00]28.00 [28.00, 28.00]
Repair attempts (persisted)0.00 [0.00, 0.00]0.00 [0.00, 0.00]26.00 [26.00, 26.00]
Repairs applied (persisted)0.00 [0.00, 0.00]0.00 [0.00, 0.00]16.00 [16.00, 16.00]
Repairs with unrecorded application state0.00 [0.00, 0.00]0.00 [0.00, 0.00]2.00 [2.00, 2.00]
Jobs with unresolved repair-attempt state0.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]

The 16 persisted applied repairs on real usage end in 13 comparison-unverified refusals, two comparison-unverified answers and one silent-wrong answer; none is correct. These final outcomes do not exclude partial corrections within an answer.

Paired real-usage intervals resample whole families within these single runs. They describe variation across sampled families, with no estimate of run-to-run noise. Count differences are normalized to 101 attempts, 86 first turns or five conversations. Of 10,000 draws, 472 have undefined precision in each B contrast; 46 have no conversations in each contrast. Those draws are excluded from the corresponding intervals and retained in the undefined-draw counts.

MetricB−A [95% CI]A−D2 [95% CI]B−D2 [95% CI]
First-turn execution errors4.00 [-1.16, 9.77]0.00 [-6.62, 6.06]4.00 [-3.25, 11.54]
Correct completed0.00 [0.00, 0.00]-2.00 [-4.90, 0.00]-2.00 [-4.90, 0.00]
Silent wrong-20.00 [-30.00, -11.22]-13.00 [-24.73, -2.91]-33.00 [-43.77, -22.67]
Refusals31.00 [19.71, 41.64]26.00 [14.99, 37.31]57.00 [47.21, 65.65]
Guard refusals-8.00 [-15.81, -0.94]26.00 [14.99, 37.31]18.00 [8.74, 27.97]
Verifier refusals39.00 [28.98, 48.48]0.00 [0.00, 0.00]39.00 [28.98, 48.48]
Comparison unverified-12.00 [-22.33, -2.10]-14.00 [-21.87, -5.61]-26.00 [-36.73, -14.92]
Full conversations0.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
Precision among scored answers28.99% [0.00%, 95.45%]-3.55% [-10.26%, 3.12%]25.44% [-9.76%, 92.86%]
p95 question latency (s)22.34 [9.14, 31.89]8.86 [0.99, 18.60]31.20 [20.07, 42.03]

Outcome transitions from D2

Mean counts per 122 attempts, equally weighting all nine case-matched repeat pairs. These are descriptive transitions; a change to correct is not proof that a specific repair caused it.

D2 outcomeD3 outcomeArm AArm B
comparison_unverifiedcomparison_unverified17.007.56
comparison_unverifiedcorrect_completed1.110.11
comparison_unverifiedrefusal_unverified11.2223.00
comparison_unverifiedsilent_wrong3.112.00
comparison_unverifiedtask_error3.223.00
correct_completedcomparison_unverified0.110.11
correct_completedcorrect_completed0.560.00
correct_completedrefusal_unverified0.220.89
correct_completedsilent_wrong0.110.00
correct_completedtask_error0.330.33
refusal_unverifiedcomparison_unverified0.220.22
refusal_unverifiedcorrect_completed0.000.22
refusal_unverifiedrefusal_unverified5.677.00
refusal_unverifiedsilent_wrong1.000.22
refusal_unverifiedtask_error2.111.33
silent_wrongcomparison_unverified2.670.67
silent_wrongcorrect_completed0.330.00
silent_wrongrefusal_unverified13.5634.67
silent_wrongsilent_wrong22.004.11
silent_wrongtask_error4.783.89
task_errorcomparison_unverified1.330.44
task_errorcorrect_completed0.330.00
task_errorrefusal_unverified8.3311.11
task_errorsilent_wrong3.780.00
task_errortask_error18.8921.11

Models, calls and timings

Both D3 arms request glm-5.2 on zai; the generator table reports response.model where available. D2 retained only the requested generator ID. Codex reports a CLI banner for gpt-6-astra; that does not independently attest the serving model. Codex reports total tokens only, not input/output.

RunGenerator calls / failures / unclosedGenerator IDsVerifier calls / failures / unclosedVerifier IDs
a-fresh-r1248 / 0 / 0{'glm-5.3': 248}0 / 0 / 0{}
a-fresh-r2257 / 0 / 0{'glm-5.3': 257}0 / 0 / 0{}
a-fresh-r3251 / 0 / 0{'glm-5.3': 251}0 / 0 / 0{}
b-fresh-r1267 / 1 / 0{'glm-5.3': 266}60 / 6 / 0{'gpt-6-astra': 54}
b-fresh-r2272 / 0 / 0{'glm-5.3': 272}64 / 5 / 0{'gpt-6-astra': 59}
b-fresh-r3270 / 0 / 0{'glm-5.3': 270}57 / 1 / 0{'gpt-6-astra': 56}
a-real-r1180 / 1 / 0{'glm-5.3': 179}0 / 0 / 0{}
b-real-r1208 / 1 / 0{'glm-5.3': 207}73 / 3 / 0{'gpt-6-astra': 70}

Families and observed defect classes

Arm A averages 2.33 correct completions per 122. The operational limit to address is incomplete family definitions and grounding. The next work is to bind those definitions to executable source mappings and answer contracts; a generator change or the added verifier is not a demonstrated remedy here. This is an operational priority, not an experimental exclusion of model effects: the fixed-generator design does not establish that changing models could not help.

Counts below pool all three Arm A repeats (366 attempts). A family is listed as wrong/error only when at least one attempt is `silent_wrong`, `task_error` or `infrastructure_failure`. Refusals and comparison-unverified answers remain separate. Classes are conservative observations from the saved result fields and terminal evidence; they do not claim an isolated semantic root cause. All 21 PlannerError cases lack retained generated SQL, but that absence does not establish why planning failed. The existing comparator already permits harmless aliases; differing labels alone are not classified as a defect.

FamilyWrong answersExecution errorsCorrectRefusedComparison unverifiedObserved classes
F01: Signups by acquisition source10074Result row-set size differs; cohort/window/series cause not isolated
F02: Signups by country or nationality12063Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed
F03: Commission total for a period160131PlannerError; specific cause unconfirmed; Values, membership or column alignment disagree under the existing comparator; Source schema reference failed; Source SQL execution failed; specific cause unconfirmed
F04: Commission by month80040Month labels differ from the reference axis; window/label cause not isolated; Result row-set size differs; cohort/window/series cause not isolated; Required zero-filled periods are missing from the returned series
F10: Quoted but not booked16167Execution boundary failed; specific cause unconfirmed; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; Source schema reference failed
F11: Quote volume and quote-to-trade conversion30063Result row-set size differs; cohort/window/series cause not isolated; Required zero-filled periods are missing from the returned series
F12: Churn tier classification80004Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated; Values, membership or column alignment disagree under the existing comparator
F13: Quiet high-value clients31017Federated plan contract failed; Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated
F14: Activation rate30054Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated
F16: Signup counts by period, client type or stream70041Result row-set size differs; cohort/window/series cause not isolated; Values, membership or column alignment disagree under the existing comparator; Required zero-filled periods are missing from the returned series
F17: Signup quality by source or country40041Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated; Values, membership or column alignment disagree under the existing comparator
F19: Top clients by revenue62094Missing expected output fields with incompatible projection width; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed
F20: Trades today or recent trades81030Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed
F23: Bookings by payment partner or platform11070PlannerError; specific cause unconfirmed; Values, membership or column alignment disagree under the existing comparator
F24: Period-over-period comparison29001PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed
F26: Refer-and-earn programme10155Result row-set size differs; cohort/window/series cause not isolated; Required zero-filled periods are missing from the returned series
F27: Corporate accounts sent to brokers03060PlannerError; specific cause unconfirmed; Source SQL execution failed; specific cause unconfirmed
F31: Client listing with contact details61101Missing expected output fields with incompatible projection width; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated
F33: First-time traders in a month26022Federated plan contract failed; Values, membership or column alignment disagree under the existing comparator; Source SQL execution failed; specific cause unconfirmed
F34: Revenue by acquisition source or campaign40041Values, membership or column alignment disagree under the existing comparator
F38: Sales lead counts35010Month labels differ from the reference axis; window/label cause not isolated; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; Source SQL execution failed; specific cause unconfirmed; Required zero-filled periods are missing from the returned series
P01: Deal loss reasons ranking14250Pipedrive tables lack required schema qualification; Result row-set size differs; cohort/window/series cause not isolated; SQL validation failed; specific cause unconfirmed
P02: Lost deals listing for one reason08001Pipedrive tables lack required schema qualification
P03: Pipeline status and open deals011100Pipedrive tables lack required schema qualification
P04: Win and loss rate by owner or industry07002Pipedrive tables lack required schema qualification
P05: Sales activity volume and mix011001Pipedrive tables lack required schema qualification
X01: Sales activity per trading client52002Missing expected output fields with incompatible projection width; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated; SQL validation failed; specific cause unconfirmed
X02: Deals versus trading reconciliation52002Missing expected output fields with incompatible projection width; PlannerError; specific cause unconfirmed; Result row-set size differs; cohort/window/series cause not isolated
X03: Lost reasons crossed with client profile60030Missing expected output fields with incompatible projection width; Result row-set size differs; cohort/window/series cause not isolated

Families without observed wrong/error outcomes:

The compact per-attempt classification ledger is `family-defects.jsonl`; the structural classification is reproducible with `scripts/insights_d3_family_audit.py` against full records. A deeper SQL-to-reference audit would be needed to distinguish cohort, measure and calendar causes where only a result mismatch or a generic source execution error is observed.

Measurement addendum and recording correction

All eight live runs completed on 2026-09-09, from 12:52 to 16:05 UTC, with no crashed repeat, restart or quality-motivated rerun. Every D3 application snapshot and evaluator source copy used revision `2513a7186df76693dd4085a38a4aca865fcefa1a`. The six fresh repeats use the same corrected corpus hash as D2; the two real-usage repeats likewise match their D2 corpus. Private copies of both corrected corpora are retained under `/tmp/lore-goal3-eval/goal-d3/oracles/`; `oracle-retention.json` records their hashes. No oracle was edited.

The controller admitted each run after two consecutive five-second samples with at least 95% CPU idle. `run-order-evidence.json` verifies all eight admissions against `host-observations.jsonl` and `sequence.jsonl`. This is admission evidence, not continuous monitoring or proof that no other process used the host during a run. The historical baseline and sequential D3 arms remain confounded with time, host activity and provider conditions.

Frozen results versus the final implementation

While examining completed Arm A records, an audit found that `execute_approved_query`, used by federated legs, omitted the new named default obligations from successful results. The frozen generation prompts already contained the defaults. The pre-measurement implementation review's statement that every execution path recorded them was too broad.

All application source remained unchanged throughout the eight live runs. After the final run and server shutdown, regression tests demonstrated the missing approved-path records and applicability gaps for person-table shorthand, CTE populations and explicit years. Three listing fixtures initially failed the existing LIMIT requirement; their limits were corrected before the recording fix, and both failed runs were retained. The valid fixtures then reproduced seven failures and four passes. After the fix, all eleven focused tests passed.

The final correction, checkpoint `d0bf8f2b`, adds named, unproved default records to approved/federated successful results and improves applicability detection. Revenue, trade and quote totals keep their own statistics filters; merely reading a CT fact table does not add the people-exclusion obligation. Explicit staff inclusion still suppresses it. The correction changes no generation prompt, SQL rewrite, verifier rule, oracle, refusal gate or provider configuration. Default applicability remains a conservative interpretation of question vocabulary and parsed sources, not a complete semantic proof.

`metadata-replay.json` records an offline perturbation of saved answer-record metadata: 934 of 934 original outcomes reproduced, 501 payloads with changed default metadata, zero outcome changes, and 934 unchanged verifier prior-turn views. Every other payload field and every raw attempt file remained unchanged. A focused integration test also compares the complete verifier prompts before and after adding default records. This is an offline non-interference check, not another benchmark or a live run of the final head. The final correction's latency has not been remeasured. The post-correction Python check passed 958 tests across 40 files, with zero failures and 88 skips; its curated subset passed 636 tests. The unchanged frontend retains the earlier 39 passing checks. Nineteen inherited curated files are absent and unverified. Results and source hashes are retained in `checks/final-summary.json` and `checks/final-check-receipt.json`; frozen checks remain separate in `checks/summary.json`.

Complete verdict and repair accounting

The evaluator's wrapper retained every verifier packet and returned verdict independently of application trace persistence. Seven executed jobs lost their application verification trace after a later error: five across the fresh repeats and two in real usage. Their returned non-operational rejections followed by generator dispatches establish seven additional repair attempts. The compact provider and verifier traces retain job IDs and monotonic timings so this inference can be checked independently.

Fresh repair attempts are 16, 19 and 18 (persisted: 15, 17 and 16); real usage has 28 attempts (26 persisted). Applied-repair counts include only persisted application state: six, seven and six on fresh, and sixteen on real usage. Application state for the seven recovered attempts is unrecorded and is not counted as applied or successful. No repair-attempt state remains unresolved under the dispatch rule. All returned accepts, rejects and clarifications are counted, including operational verdicts and turns whose later errors dropped the application trace. Free-form verifier signal text is hashed in compact records; the complete private records preserve it.

None of the persisted applied repairs ended in a correct completion. On fresh, the 19 applied repairs ended in 14 `refusal_unverified` and five `silent_wrong` outcomes. On real usage, the 16 ended in 13 `refusal_unverified`, two `comparison_unverified` and one `silent_wrong` outcome. These are final outcomes, not evidence that no partial correction occurred. Every D3 refusal is `refusal_unverified`; the corrected oracle has not independently justified the refusal.

Sampling labels and prior exposure

The leakage review identified an evaluator metadata omission: the `LiveProvider` wrapper left `LLMTextResult.sampling_pin` at its default. Consequently, 929 per-question SQL-generation summaries and 1,950 nested generation-attempt records say `off`; five questions have no summary label. `sampling-pin-audit.json` separates these counts. Source and retained configuration support requested temperature 0, seed 42 and disabled thinking. The misleading labels neither establish unpinned requests nor independently verify provider compliance.

After review, the wrapper was corrected to propagate `sqlgen_pin_state(config.provider)` into future result metadata. No request argument, measured record or result was changed, and no live evaluation was rerun. This telemetry correction, like the obligation-recording correction, is outside the measured revision. The frozen method has a dated qualification: inherited admission labels saying “no tuning after the slice was opened” establish neither an unseen holdout nor an absence of development after D2 exposure. The supported claim is unchanged source throughout these eight measurements.

Interpretation and reporting additions

Arm A's mean correct-completion gain over D2 is one answer per 122, inside D2's three-answer observed repeat range. Its precision interval is positive, but the gain is inside D2's precision range. Arm B's correct-completion difference against A is minus two: its family interval excludes zero, but the mean difference equals, and does not exceed, A's two-answer repeat range. Neither result establishes recovery beyond both forms of evidence required by the D2 method.

Silent wrong answers fall by 13.33 per 122 in A versus D2 and by 37 in B versus D2; both reductions exceed D2's observed range and have wholly negative paired intervals. Only B meets the requirement to halve them. B also has substantially more refusals. The data supports selective refusal as the main observed verifier contribution here, not a recovery in correct completions. The real-usage precision of 33.33% in B is one correct answer among only three scored answers, with no repeat-noise estimate.

The original paired-count method was extended during reporting to pooled per-question p95 differences and supplementary real-usage family intervals. These additions did not select a model, alter an answer, change a threshold or cause a rerun. Per-arm means/minima/maxima remain repeat statistics; paired precision and p95 pool first and therefore have different estimands. Completed-call latency summaries exclude failed calls; failed-call timings remain in the traces, and the per-question p95 includes all timed attempts.

The D2 baseline's application revision was `c4f76cf069abfc98f10a9b1fe8846555f2117990`. Its `src/` tree matches the fetched D2 tip `4bfc6022`; the intervening D2 changes were reporting/reconciliation work. Its historical served generator identity cannot be reconstructed: only the requested ID was retained. D3 records each completed zai response's `model` field; a Codex CLI banner remains configuration evidence rather than independent serving attestation.

Independent review and correction ledger

Three separate Codex CLI processes reviewed checkpoint `634a75d9044c7e68f81111734642194a60d8d630`, using `gpt-6-astra`, xhigh reasoning and `danger-full-access`, with self-contained prompts and strict no-modification instructions. All three exited successfully; before/after inventories detected zero changed files across 6,467 protected source, evidence and private-record files. Each prompt, reply and receipt is retained here; complete process transcripts remain in the private full-record directory. The CLI banner establishes requested configuration, not independent serving identity.

The reviewed results SHA-256 was `61f72ea449a136a8d50f4314653ef6ce648283358f2acae6660f5a4004a74ef6`; the reviewed Markdown SHA-256 was `e2843fae2e234ac29b9819a8819a0c2f8fba5a07209a92decf66cbd3ccd64d07`. The wording below was corrected after those reviews; it was not submitted for a fourth review.

ReviewVerdictConfirmed corrections
Arithmetic and intervalsPASS; no numerical discrepancyNone. Independently reproduced raw structural metrics, all six family-bootstrap comparisons, rendered cells, transitions and repair accounting.
Leakage, model identity and noiseNo blocking leakage identified; two recording/wording findingsDisclosed the misleading sampling labels, corrected future wrapper telemetry, and qualified inherited exposure labels.
Unsupported wordingRevise wording before sign-off; numbers standNarrowed the operational diagnosis; left PlannerError causes unresolved; clarified the measurement freeze, executed-check scope, scored-answer precision and applied-repair outcomes.

1. **Operational diagnosis:** family definitions and executable grounding are the next operational limit to address. This fixed-generator experiment does not causally exclude a model change helping. The report no longer treats that exclusion as established.

2. **Planner errors:** renamed `planner_no_usable_sql` to `planner_error_unresolved` for 21 Arm A attempts. Their retained payloads lack SQL, but a generic PlannerError does not isolate why. This changes only the diagnostic ledger, not oracle outcomes.

3. **Measurement freeze:** appended a dated method qualification. All eight runs measured `2513a718`; later inspection informed the obligation-recording correction at `d0bf8f2b`. The final code was not live-benchmarked. The retained 934-attempt offline metadata replay supports non-interference, not a new latency measurement.

4. **Curated coverage:** the visible PASS label now says “Zero observed failures in executed, present curated checks.” The 88 skipped checks and 19 absent inherited files remain unverified. No acceptance rule was changed.

5. **Precision denominator:** labels now say “Precision among scored answers,” with the formula and separate comparison-unverified counts. Fresh coverage is A 7/366 (1.91%) and B 1/366 (0.27%), distinct from precision.

6. **Repairs and refusals:** added observed final outcomes for all persisted applied repairs: fresh 19 = 14 unverified refusals + five silent wrong; real 16 = 13 unverified refusals + two unverified answers + one silent wrong. None ended correct. Seven additional dispatched repairs retain unknown application state. All D3 refusals remain oracle-unverified.

7. **Sampling evidence:** independently counted 929 misleading per-question summary labels and 1,950 nested attempt labels. Added `sampling-pin-audit.json`; future evaluator result metadata now propagates the configured pin state. Original traces are unchanged; provider compliance is not independently attested.

8. **Exposure:** the slice was previously exposed during D2. Inherited “no tuning after the slice was opened” labels are retained but cannot establish an unseen holdout or absence of all post-D2 development.

No score, interval, transition, threshold, oracle or measured attempt changed as a result of review. Reviewers did not rerun the application, tests or semantic answer scoring; arithmetic review recomputed retained structural outcome fields. All substantive review findings were either confirmed and corrected above or retained as explicit limits.

Method and limitations

D3 method frozen before measurement

Post-measurement qualification (2026-09-09): the source-change statement above describes the measurement freeze. All eight measured runs remained on `2513a718`; subsequent inspection of saved Arm A records informed the recording correction at `d0bf8f2b`, which was not live-benchmarked. Inherited admission labels saying “no tuning after the slice was opened” do not establish an unseen holdout or the absence of development after D2 exposure. The later sampling-label correction changes future telemetry only; original labels remain in the retained records.

Receipt checks confirm identical frozen D3 application and evaluation-source revisions across all eight runs, two workers with no case limit, and matching corrected-corpus and evaluator hashes with the corresponding D2 baseline. The exact hashes are retained in results.json and each run receipt.

Curated checks: {"corrected_run": {"errors": 0, "failed": 0, "passed": 164, "skipped": 13}, "curated": {"errors": 0, "failed": 0, "passed": 636, "skipped": 88}, "errors": 0, "failed": 0, "initial_run": {"errors": 0, "failed": 9, "passed": 922, "skipped": 88}, "js_failed": 0, "js_passed": 39, "method": "Initial complete selection, then entire affected files replaced by corrected run; all before live evaluation. One prompt-context assertion updated for mandatory defaults; old SQL-only prover checks unchanged.", "missing_inherited_files": 19, "passed": 931, "skipped": 88}.

No application deployment or merge is part of D3. Sequential arms are confounded with time and provider conditions. The oracle is corrected but fixture-based; comparison-unverified answers and the small conversation sample limit conclusions. The staff/test marker remains partial as documented in D2. Named semantic obligations remain unproved; the added context does not provide a complete semantics proof. Host load observations are evidence of conditions, not proof of exclusive host use.

Full records: `/tmp/lore-goal3-eval/goal-d3/`. Compact recomputable records: `docs/evidence/insights-goal-d3-2026-09-09/runs/`. `scripts/insights_d3_report.py --compact` recomputes quality metrics, intervals and transitions from committed records alone. Completed-call token and latency summaries are in results.json; failed-call timings are also retained in the compact traces. Per-question p95 includes all timed attempts.