SDD review-dispatch optimization — experiment dashboard (2026-06-10)
46 eval runs today
$410 eval spend
487M tokens
24.7h agent wall-clock
~$23 micro-tests (est.)
Variance note: identical configs re-run differ by ±20% (44.4 vs 57.1 min); compare ranges, not single rows. Row colors: baseline / iteration / experiment / final config.
go-fractals (primary test scenario)
| UTC | config | verdict | min | Mtok | $ | cost |
| 002653Z | origin/dev baseline | pass | 64.9 | 21.2 | 16.07 | |
| 003436Z | task-scoped v1 (driver flake) | indeterminate | 29.8 | 7.6 | 7.90 | |
| 010511Z | task-scoped v1 | pass | 42.8 | 14.5 | 15.61 | |
| 041922Z | quality-hardened | pass | 69.9 | 32.2 | 17.14 | |
| 054305Z | iter 1: turn guidance | pass | 68.2 | 22.9 | 17.59 | |
| 065838Z | iter 2: merged task reviewer | pass | 47.5 | 15.7 | 13.55 | |
| 104434Z | iter 3: frozen e355795 | pass | 44.4 | 13.4 | 11.67 | |
| 161736Z | validation re-run (d4dbf44) | pass | 57.1 | 20.0 | 14.63 | |
| 163928Z | E1 lean-controller | pass | 50.2 | 13.7 | 11.73 | |
| 163930Z | E2 file-handoffs (model-decay accident) | pass | 47.0 | 14.9 | 17.43 | |
| 165219Z | E4 durable-progress | pass | 48.4 | 15.8 | 13.37 | |
| 171746Z | E3 pipelined (parallel calls) | pass | 49.0 | 14.1 | 13.43 | |
| 174422Z | COMBO | pass | 54.7 | 14.4 | 12.81 | |
| 180941Z | E3b pipelined (background) | pass | 53.6 | 17.8 | 14.84 | |
| 191021Z | COMBO repeat (gate) | pass | 54.1 | 16.6 | 14.31 | |
svelte-todo
| UTC | config | verdict | min | Mtok | $ | cost |
| 002653Z | scenario-dev iteration | indeterminate | 5.8 | 2.0 | 2.14 | |
| 003432Z | origin/dev baseline | pass | 79.7 | 27.3 | 20.98 | |
| 004050Z | scenario-dev iteration | indeterminate | 28.6 | 9.1 | 7.40 | |
| 011032Z | scenario-dev iteration | indeterminate | 5.3 | 1.8 | 2.73 | |
| 015739Z | scenario-dev iteration | indeterminate | 47.5 | 15.6 | 12.60 | |
| 024622Z | scenario-dev iteration | indeterminate | 7.6 | 2.6 | 2.61 | |
| 025508Z | task-scoped v1 | pass | 54.5 | 17.2 | 22.18 | |
| 041922Z | quality-hardened | pass | 82.8 | 29.2 | 22.00 | |
| 075303Z | iter 1 | pass | 73.5 | 25.4 | 19.42 | |
| 093503Z | iter 2 | pass | 66.2 | 22.4 | 18.16 | |
| 104434Z | iter 3: frozen e355795 | pass | 62.8 | 19.7 | 15.76 | |
| 191023Z | COMBO (gate) | pass | 55.0 | 19.3 | 14.99 | |
planted-defect (scenario was being developed during early runs — fails are scenario iterations, not regressions)
| UTC | config | verdict | min | Mtok | $ | cost |
| 010638Z | scenario-dev iteration | indeterminate | 8.9 | 2.5 | 2.51 | |
| 014843Z | scenario-dev iteration | pass | 8.1 | 3.6 | 2.62 | |
| 041922Z | scenario-dev iteration | fail | 12.4 | 3.1 | 3.31 | |
| 044959Z | scenario-dev iteration | indeterminate | 11.6 | 3.8 | 2.85 | |
| 075303Z | scenario-dev iteration | fail | 7.8 | 2.1 | 2.32 | |
| 090920Z | scenario-dev iteration | fail | 7.9 | 2.3 | 2.34 | |
| 092037Z | scenario-dev iteration | fail | 8.4 | 2.5 | 2.59 | |
| 093503Z | scenario-dev iteration | pass | 10.5 | 2.6 | 2.14 | |
| 104434Z | scenario-dev iteration | fail | 6.8 | 1.6 | 1.73 | |
| 115041Z | scenario-dev iteration | pass | 13.0 | 3.2 | 3.12 | |
| 191024Z | COMBO (gate) | pass | 10.2 | 2.9 | 2.77 | |
rejects-extra-features
| UTC | config | verdict | min | Mtok | $ | cost |
| 002653Z | origin/dev baseline | pass | 6.3 | 1.9 | 1.88 | |
| 003352Z | task-scoped v1 | pass | 6.1 | 3.4 | 2.04 | |
| 075303Z | iter 1 | pass | 5.5 | 2.1 | 1.37 | |
| 105312Z | iter 3 | pass | 5.0 | 2.0 | 1.31 | |
spec-reviewer-catches-planted-flaws
| UTC | config | verdict | min | Mtok | $ | cost |
| 003312Z | scenario-dev iteration | pass | 0.9 | 0.2 | 0.50 | |
| 010440Z | scenario-dev iteration | pass | 0.9 | 0.2 | 0.44 | |
| 080012Z | scenario-dev iteration | pass | 2.4 | 0.2 | 0.50 | |
| 110001Z | scenario-dev iteration | pass | 0.9 | 0.2 | 0.50 | |
Confirmed wins (landed on branch through 43a6ee2)
- Final-review package — final reviewer 33 turns/23 tools → 6/3, at opus prices (opus bill −19%).
- REQUIRED
model: template line — dispatch-model discipline 57/57 across gates; prose version decayed once (+$5 run).
- Progress ledger (.git/sdd/progress.md) — compaction-proof state; controller adopted it beyond spec (Minor-findings roll-up, final verdict).
- Unique collateral names (.git/sdd/review-<sha>..<sha>.diff) — worktree/submodule-safe, verified.
- Dispatch composition recipe + reviewer risk budget — micro-tested; gates green.
- Omnibus final fixer / scoped fix tests / anti-snowball — from real-session mining; no eval cost.
Tested & declined (negative results, full data in evals docs/experiments/)
- Controller turn batching — 0 multi-tool messages in every run, with and without guidance; 46% of controller turns are thinking/narration. Prompt-immune.
- Pipelining via parallel tool calls (E3) — 0 paired dispatches in 29.
- Pipelining via background dispatch (E3b) — adopted 7/28, but benefit < ±6 min noise floor; added coordination tokens.
- "Do not restate" prohibition — backfired (see micro 1); replaced by composition recipe.
- File-handoff context savings — mechanism adopted 100%, but savings ~$1/run, not the modeled ~50%: the controller's restatement is mostly genuine curation.
Micro-tests (prompt-wording experiments — API calls, est. cost)
| test | question | samples | result | ~$ |
| micro 1 | dispatch composition: prohibition vs positive recipe | 25 × opus | recipe wins 3.0 (zero variance) vs prohibition 4.4 vs control 3.6 — prohibition worse than nothing; nuance clause regressed winner to 3.8 | 4 |
| micro 2 | reviewer test-rerun directive: prohibition vs positive | 15 × opus | both 0/5 violations, control 3/5 — discrete prohibitions hold; kept shorter wording | 3 |
| micro 3a | writing-plans placeholders, no pressure | 20 × opus | inconclusive — 0 placeholders in all variants incl. control | 8 |
| micro 3b | writing-plans placeholders, 10 tasks + word-budget pressure | 20 × opus | running | 8 |
Sources: evals/results/*-20260610T*; experiment log evals/docs/experiments/2026-06-10-sdd-cost-experiments.md; spec docs/superpowers/specs/2026-06-09-…-design.md (iterations 1-5). Generated Wed Jun 10 14:14:54 PDT 2026.