SDD review-dispatch optimization — experiment dashboard (2026-06-10)

46 eval runs today $410 eval spend 487M tokens 24.7h agent wall-clock ~$23 micro-tests (est.)

Variance note: identical configs re-run differ by ±20% (44.4 vs 57.1 min); compare ranges, not single rows. Row colors: baseline / iteration / experiment / final config.

go-fractals (primary test scenario)

UTCconfigverdictminMtok$cost
002653Zorigin/dev baselinepass64.921.216.07
003436Ztask-scoped v1 (driver flake)indeterminate29.87.67.90
010511Ztask-scoped v1pass42.814.515.61
041922Zquality-hardenedpass69.932.217.14
054305Ziter 1: turn guidancepass68.222.917.59
065838Ziter 2: merged task reviewerpass47.515.713.55
104434Ziter 3: frozen e355795pass44.413.411.67
161736Zvalidation re-run (d4dbf44)pass57.120.014.63
163928ZE1 lean-controllerpass50.213.711.73
163930ZE2 file-handoffs (model-decay accident)pass47.014.917.43
165219ZE4 durable-progresspass48.415.813.37
171746ZE3 pipelined (parallel calls)pass49.014.113.43
174422ZCOMBOpass54.714.412.81
180941ZE3b pipelined (background)pass53.617.814.84
191021ZCOMBO repeat (gate)pass54.116.614.31

svelte-todo

UTCconfigverdictminMtok$cost
002653Zscenario-dev iterationindeterminate5.82.02.14
003432Zorigin/dev baselinepass79.727.320.98
004050Zscenario-dev iterationindeterminate28.69.17.40
011032Zscenario-dev iterationindeterminate5.31.82.73
015739Zscenario-dev iterationindeterminate47.515.612.60
024622Zscenario-dev iterationindeterminate7.62.62.61
025508Ztask-scoped v1pass54.517.222.18
041922Zquality-hardenedpass82.829.222.00
075303Ziter 1pass73.525.419.42
093503Ziter 2pass66.222.418.16
104434Ziter 3: frozen e355795pass62.819.715.76
191023ZCOMBO (gate)pass55.019.314.99

planted-defect (scenario was being developed during early runs — fails are scenario iterations, not regressions)

UTCconfigverdictminMtok$cost
010638Zscenario-dev iterationindeterminate8.92.52.51
014843Zscenario-dev iterationpass8.13.62.62
041922Zscenario-dev iterationfail12.43.13.31
044959Zscenario-dev iterationindeterminate11.63.82.85
075303Zscenario-dev iterationfail7.82.12.32
090920Zscenario-dev iterationfail7.92.32.34
092037Zscenario-dev iterationfail8.42.52.59
093503Zscenario-dev iterationpass10.52.62.14
104434Zscenario-dev iterationfail6.81.61.73
115041Zscenario-dev iterationpass13.03.23.12
191024ZCOMBO (gate)pass10.22.92.77

rejects-extra-features

UTCconfigverdictminMtok$cost
002653Zorigin/dev baselinepass6.31.91.88
003352Ztask-scoped v1pass6.13.42.04
075303Ziter 1pass5.52.11.37
105312Ziter 3pass5.02.01.31

spec-reviewer-catches-planted-flaws

UTCconfigverdictminMtok$cost
003312Zscenario-dev iterationpass0.90.20.50
010440Zscenario-dev iterationpass0.90.20.44
080012Zscenario-dev iterationpass2.40.20.50
110001Zscenario-dev iterationpass0.90.20.50

Confirmed wins (landed on branch through 43a6ee2)

Tested & declined (negative results, full data in evals docs/experiments/)

Micro-tests (prompt-wording experiments — API calls, est. cost)

testquestionsamplesresult~$
micro 1dispatch composition: prohibition vs positive recipe25 × opusrecipe wins 3.0 (zero variance) vs prohibition 4.4 vs control 3.6 — prohibition worse than nothing; nuance clause regressed winner to 3.84
micro 2reviewer test-rerun directive: prohibition vs positive15 × opusboth 0/5 violations, control 3/5 — discrete prohibitions hold; kept shorter wording3
micro 3awriting-plans placeholders, no pressure20 × opusinconclusive — 0 placeholders in all variants incl. control8
micro 3bwriting-plans placeholders, 10 tasks + word-budget pressure20 × opusrunning8

Sources: evals/results/*-20260610T*; experiment log evals/docs/experiments/2026-06-10-sdd-cost-experiments.md; spec docs/superpowers/specs/2026-06-09-…-design.md (iterations 1-5). Generated Wed Jun 10 14:14:54 PDT 2026.