SDD Review-Dispatch Evals — grouped by experiment

Same 63 runs as the chronological matrix, regrouped by what each experiment asked. Chronological within groups.

Baseline — dev branch as shipped · 5 runs

Reference numbers the whole campaign is measured against: fractals $16.07 / 64.9 min, svelte $20.98 / 79.7 min.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 00:26sdd-go-fractalsbaselineclaudepass6521.2M16.07O$7.68 S$8.3921.812883
06-10 00:26sdd-rejects-extra-featuresbaselineclaudepass61.9M1.88O$1.15 S$0.722.33607b
06-10 00:26sdd-svelte-todobaselineclaudeindeterminate62.0M2.14O$1.40 S$0.742.759297
06-10 00:33spec-reviewer-catches-planted-flawsbaselineclaudepass10.2M0.50O$0.500.634b11
06-10 00:34sdd-svelte-todobaselineclaudepass8027.3M20.98O$11.14 S$9.8329.35e69c
subtotal15753M41.5656.87

Design iterations 1–5 — task-scoped review dispatch (the PR itself) · 15 runs

One task reviewer per task (spec + quality verdicts), review-package files, scope and escape-hatch tuning across five iterations on the branch. Converged on the frozen e355795 config: fractals −27% $, svelte −25% $ vs baseline. 'branch state' = the branch as it stood at launch time.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 00:34sdd-go-fractalsbranch stateclaudeindeterminate307.6M7.90O$4.93 S$2.979.9984d1
06-10 00:40sdd-svelte-todobranch stateclaudeindeterminate299.1M7.40O$3.99 S$3.25 H$0.179.882b6f
06-10 01:05sdd-go-fractalsbranch stateclaudepass4314.5M15.61O$13.44 S$2.1718.50c7b3
06-10 01:10sdd-svelte-todobranch stateclaudeindeterminate51.8M2.73O$2.733.371629
06-10 01:57sdd-svelte-todobranch stateclaudeindeterminate4715.6M12.60O$6.49 S$6.1220.864f68
06-10 02:46sdd-svelte-todobranch stateclaudeindeterminate82.6M2.61O$1.62 S$0.993.951c6f
06-10 02:55sdd-svelte-todobranch stateclaudepass5517.2M22.18O$22.1823.979af1
06-10 04:19sdd-go-fractalsbranch stateclaudepass7032.2M17.14O$9.29 S$5.89 H$1.9618.95ad31
06-10 04:19sdd-svelte-todobranch stateclaudepass8329.2M22.00O$13.82 S$7.53 H$0.6623.524033
06-10 05:43sdd-go-fractalsbranch stateclaudepass6822.9M17.59O$9.58 S$8.0118.74e161
06-10 06:58sdd-go-fractalsbranch stateclaudepass4815.7M13.55O$9.24 S$4.3114.484d41
06-10 07:53sdd-svelte-todobranch stateclaudepass7325.4M19.42O$11.96 S$7.4620.79ecfc
06-10 09:35sdd-svelte-todobranch stateclaudepass6622.4M18.16O$11.84 S$6.3219.60be5b
06-10 10:44sdd-svelte-todobranch stateclaudepass6319.7M15.76O$8.61 S$7.09 H$0.0617.069fb3
06-10 16:17sdd-go-fractalsbranch stateclaudepass5720.0M14.63O$7.87 S$6.53 H$0.2315.81309b
subtotal744256M209.29239.44

Planted-defect scenario development (v1 → v3) · 11 runs

Building the quality gate itself: a scenario whose plan plants an assertion-free test with a lying name plus a DRY violation. Failed through five distinct suppression mechanisms during development (controller pre-judging, severity pre-rating, reviewer calibration, implementer framing, and a wrong eval bar) — each fixed in the prompts generally, not by teaching to the test.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 01:06sdd-quality-reviewer-catches-planted-defectbranch stateclaudeindeterminate92.5M2.51O$1.52 S$0.993.4547df
06-10 01:48sdd-quality-reviewer-catches-planted-defectbranch stateclaudepass83.6M2.62O$2.24 H$0.373.310c68
06-10 04:19sdd-quality-reviewer-catches-planted-defectbranch stateclaudefail123.1M3.31O$2.07 S$1.243.779a13
06-10 04:49sdd-quality-reviewer-catches-planted-defectbranch stateclaudeindeterminate123.8M2.85O$2.04 S$0.47 H$0.343.51bc1c
06-10 07:53sdd-quality-reviewer-catches-planted-defectbranch stateclaudefail82.1M2.32O$1.74 S$0.582.58445a
06-10 09:09sdd-quality-reviewer-catches-planted-defectbranch stateclaudefail82.3M2.34O$1.80 S$0.552.69c127
06-10 09:20sdd-quality-reviewer-catches-planted-defectbranch stateclaudefail82.5M2.59O$2.05 S$0.542.848817
06-10 09:35sdd-quality-reviewer-catches-planted-defectbranch stateclaudepass112.6M2.14O$1.67 S$0.26 H$0.212.58719d
06-10 10:44sdd-quality-reviewer-catches-planted-defectbranch stateclaudefail71.6M1.73O$1.17 S$0.552.16b6ab
06-10 11:50sdd-quality-reviewer-catches-planted-defectbranch stateclaudepass133.2M3.12O$2.03 S$1.093.50d8c1
06-10 19:10sdd-quality-reviewer-catches-planted-defectcomboclaudepass102.9M2.77O$1.77 S$0.993.142ad0
subtotal10630M28.2933.52

Guardrail scenarios — validation matrix · 6 runs

sdd-rejects-extra-features ($1.31–1.37 vs $1.88 baseline) and spec-reviewer-catches-planted-flaws: quality gates every config iteration had to keep passing.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 00:33sdd-rejects-extra-featuresbranch stateclaudepass63.4M2.04O$1.70 H$0.332.4652b8
06-10 01:04spec-reviewer-catches-planted-flawsbranch stateclaudepass10.2M0.44O$0.440.600aa2
06-10 07:53sdd-rejects-extra-featuresbranch stateclaudepass52.1M1.37O$1.12 H$0.251.643a9d
06-10 08:00spec-reviewer-catches-planted-flawsbranch stateclaudepass20.2M0.50O$0.500.6175b9
06-10 10:53sdd-rejects-extra-featuresbranch stateclaudepass52.0M1.31O$1.09 H$0.231.518bf5
06-10 11:00spec-reviewer-catches-planted-flawsbranch stateclaudepass10.2M0.50O$0.500.6229f1
subtotal218M6.167.44

E1 — final-review package · 2 runs

Hand the FINAL whole-branch reviewer a package file too. WIN: 33 turns/23 tools → 6 turns/3 tools at opus prices; opus bill −19%. Landed.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 10:44sdd-go-fractalsbranch stateclaudepass4413.4M11.67O$7.13 S$4.5412.57a8d1
06-10 16:39sdd-go-fractalslean-ctlclaudepass5013.7M11.73O$5.78 S$5.9612.949a73
subtotal9527M23.4025.51

E2 — model-line decay · 1 run

Guidance-only model selection decays: dispatches from Task 3 onward inherited opus (+$5). Led to the REQUIRED model: template line. Landed.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 16:39sdd-go-fractalsfile-handoffsclaudepass4714.9M17.43O$16.18 S$1.2518.47a65c
subtotal4715M17.4318.47

E3/E3b — pipelined reviews · 2 runs

Parallel-call pipelining DEAD (0 paired dispatches in 29 — the controller emits exactly one tool call per message). Background pipelining works mechanically (7/28 dispatches) but benefit below run-to-run noise. Declined.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 17:17sdd-go-fractalspipelinedclaudepass4914.1M13.43O$8.27 S$5.1614.689d88
06-10 18:09sdd-go-fractalspipelinedclaudepass5417.8M14.84O$8.75 S$6.0816.08310e
subtotal10332M28.2630.76

E4 — durable progress ledger · 1 run

Ledger in <git-dir>/sdd/progress.md, maintained well, run cost in-band; guards against post-compaction re-dispatch (worst real-session failure: 269 dispatches for ~22 tasks). Landed.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 16:52sdd-go-fractalsledgerclaudepass4815.8M13.37O$8.85 S$4.43 H$0.0914.45b260
subtotal4816M13.3714.45

Combo gates — winners merged, re-validated · 3 runs

All surviving changes together (briefs, packages, model lines, ledger, recipes, risk budget): gates green, fractals band $11.67–14.84 over repeated runs.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 17:44sdd-go-fractalscomboclaudepass5514.4M12.81O$7.41 S$5.4013.9698b7
06-10 19:10sdd-go-fractalscomboclaudepass5416.6M14.31O$8.67 S$5.6415.51941a
06-10 19:10sdd-svelte-todocomboclaudepass5519.3M14.99O$7.76 S$7.2316.502824
subtotal16450M42.1145.98

A — lean vs combo (brief ablation) · 4 runs

Does the brief/report mechanism cost money? NO: lean $11.61/$12.27/$13.25 vs the combo band + the $14.10 same-night control — fully overlapping. Briefs stay, justified by fidelity + compaction durability. Also resolves: the 44.4-min e355795 run was a tail draw, not a better config.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 22:06sdd-go-fractalsleanclaudepass5117.0M13.25O$7.30 S$5.9514.72e478
06-10 22:06sdd-go-fractalsleanclaudepass4313.8M11.61O$7.17 S$4.4412.57ee4d
06-10 22:06sdd-go-fractalsleanclaudepass4915.8M12.27O$6.46 S$5.8113.41751f
06-10 22:06sdd-go-fractalscomboclaudepass6018.4M14.10O$6.95 S$7.1515.51d170
subtotal20265M51.2356.21

B — crisp-plan ceiling (strict-cost L1) · 3 runs

Hand-rewritten 7-task plan + Global Constraints header + per-task Interfaces lines: $9.51/$12.65/$12.65, dispatches 28 → 20-24 (−21%), fix waves flat, gates 3/3. VALIDATED in effect — the follow-up is making writing-plans elicit this shape.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 22:07sdd-go-fractals-crispcomboclaudepass5116.8M12.65O$6.33 S$6.22 H$0.1113.67f21c
06-10 22:07sdd-go-fractals-crispcomboclaudepass3812.1M9.51O$4.93 S$4.50 H$0.0810.3045f0
06-10 22:59sdd-go-fractals-crispcomboclaudepass5017.1M12.65O$6.70 S$5.88 H$0.0613.6815a0
subtotal13946M34.8037.65

C — sonnet controller (strict-cost L2 recon) · 3 runs

Controller session on sonnet: $6.68/$8.05 (~−40% vs band), no turn inflation, clean judgment trail (caught go-mod-tidy silently dropping cobra), heavier-and-sane haiku tiering. Zero BLOCKED/⚠️ events arose, so escalation behavior is untested — full N=5 + judgment-audit gates still owed. The $0.13 fail was a quorum seeding bug (API-key picker defaults No) — found, fixed, now superseded by upstream's equivalent fix.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 22:47sdd-go-fractalscomboclaude-sonnetindeterminate00.000.1348a0
06-10 22:54sdd-go-fractalscomboclaude-sonnetpass3112.6M6.68S$6.41 H$0.277.745038
06-10 23:27sdd-go-fractalscomboclaude-sonnetpass4116.1M8.05S$7.48 H$0.579.462175
subtotal7229M14.7217.33

D — haiku task reviewers (strict-cost L3) · 6 runs

DEAD as pre-registered: planted-defect 2 pass / 1 indet / 2 fail (baseline 5/5); 0 of 10 defects cleanly flagged — haiku ADVOCATES for defects (DRY praised as YAGNI; assert-nothing test 'plan-compliant'; found-then-downgraded with the exact prohibited rationale). The ~$2-3/run saving in the fractals run and the quality failure are the same mechanism: lax reviews trigger fewer fix waves.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 23:28sdd-quality-reviewer-catches-planted-defecthaiku-revclaudeindeterminate123.4M3.12O$2.50 S$0.49 H$0.133.7011cf
06-10 23:28sdd-quality-reviewer-catches-planted-defecthaiku-revclaudefail133.8M3.11O$2.37 S$0.55 H$0.193.88ed05
06-10 23:28sdd-quality-reviewer-catches-planted-defecthaiku-revclaudepass112.9M2.62O$1.95 S$0.55 H$0.123.13be87
06-10 23:44sdd-quality-reviewer-catches-planted-defecthaiku-revclaudefail123.4M2.76O$2.25 S$0.34 H$0.183.346a71
06-10 23:44sdd-quality-reviewer-catches-planted-defecthaiku-revclaudepass92.5M2.31O$1.72 S$0.47 H$0.122.594241
06-10 23:58sdd-go-fractalshaiku-revclaudepass3914.1M9.61O$5.09 S$4.10 H$0.4310.550297
subtotal9630M23.5327.19

E — svelte n=2 · 1 run

Second combo svelte run: $20.30 (9 fix waves — review-strictness variance; all 34 dispatches model-disciplined). PR claim becomes the honest range $14.99–20.30 vs $20.98 baseline — time/tokens clearly better, cost overlaps at the top.

When (UTC)ScenarioConfigAgentVerdictMinTokensCoding $SplitTotal $Run
06-10 22:07sdd-svelte-todocomboclaudepass7024.1M20.30O$12.35 S$7.9422.36c47f
subtotal7024M20.3022.36
Coding $ = agent under test; Total adds the QA agent. Batch A-E wall-clock ran under up to 7-way parallel load — trust $/tokens there, not minutes.