Reference numbers the whole campaign is measured against: fractals $16.07 / 64.9 min, svelte $20.98 / 79.7 min.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 00:26 | sdd-go-fractals | baseline | claude | pass | 65 | 21.2M | 16.07 | O$7.68 S$8.39 | 21.81 | 2883 |
| 06-10 00:26 | sdd-rejects-extra-features | baseline | claude | pass | 6 | 1.9M | 1.88 | O$1.15 S$0.72 | 2.33 | 607b |
| 06-10 00:26 | sdd-svelte-todo | baseline | claude | indeterminate | 6 | 2.0M | 2.14 | O$1.40 S$0.74 | 2.75 | 9297 |
| 06-10 00:33 | spec-reviewer-catches-planted-flaws | baseline | claude | pass | 1 | 0.2M | 0.50 | O$0.50 | 0.63 | 4b11 |
| 06-10 00:34 | sdd-svelte-todo | baseline | claude | pass | 80 | 27.3M | 20.98 | O$11.14 S$9.83 | 29.35 | e69c |
| subtotal | 157 | 53M | 41.56 | 56.87 | ||||||
One task reviewer per task (spec + quality verdicts), review-package files, scope and escape-hatch tuning across five iterations on the branch. Converged on the frozen e355795 config: fractals −27% $, svelte −25% $ vs baseline. 'branch state' = the branch as it stood at launch time.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 00:34 | sdd-go-fractals | branch state | claude | indeterminate | 30 | 7.6M | 7.90 | O$4.93 S$2.97 | 9.99 | 84d1 |
| 06-10 00:40 | sdd-svelte-todo | branch state | claude | indeterminate | 29 | 9.1M | 7.40 | O$3.99 S$3.25 H$0.17 | 9.88 | 2b6f |
| 06-10 01:05 | sdd-go-fractals | branch state | claude | pass | 43 | 14.5M | 15.61 | O$13.44 S$2.17 | 18.50 | c7b3 |
| 06-10 01:10 | sdd-svelte-todo | branch state | claude | indeterminate | 5 | 1.8M | 2.73 | O$2.73 | 3.37 | 1629 |
| 06-10 01:57 | sdd-svelte-todo | branch state | claude | indeterminate | 47 | 15.6M | 12.60 | O$6.49 S$6.12 | 20.86 | 4f68 |
| 06-10 02:46 | sdd-svelte-todo | branch state | claude | indeterminate | 8 | 2.6M | 2.61 | O$1.62 S$0.99 | 3.95 | 1c6f |
| 06-10 02:55 | sdd-svelte-todo | branch state | claude | pass | 55 | 17.2M | 22.18 | O$22.18 | 23.97 | 9af1 |
| 06-10 04:19 | sdd-go-fractals | branch state | claude | pass | 70 | 32.2M | 17.14 | O$9.29 S$5.89 H$1.96 | 18.95 | ad31 |
| 06-10 04:19 | sdd-svelte-todo | branch state | claude | pass | 83 | 29.2M | 22.00 | O$13.82 S$7.53 H$0.66 | 23.52 | 4033 |
| 06-10 05:43 | sdd-go-fractals | branch state | claude | pass | 68 | 22.9M | 17.59 | O$9.58 S$8.01 | 18.74 | e161 |
| 06-10 06:58 | sdd-go-fractals | branch state | claude | pass | 48 | 15.7M | 13.55 | O$9.24 S$4.31 | 14.48 | 4d41 |
| 06-10 07:53 | sdd-svelte-todo | branch state | claude | pass | 73 | 25.4M | 19.42 | O$11.96 S$7.46 | 20.79 | ecfc |
| 06-10 09:35 | sdd-svelte-todo | branch state | claude | pass | 66 | 22.4M | 18.16 | O$11.84 S$6.32 | 19.60 | be5b |
| 06-10 10:44 | sdd-svelte-todo | branch state | claude | pass | 63 | 19.7M | 15.76 | O$8.61 S$7.09 H$0.06 | 17.06 | 9fb3 |
| 06-10 16:17 | sdd-go-fractals | branch state | claude | pass | 57 | 20.0M | 14.63 | O$7.87 S$6.53 H$0.23 | 15.81 | 309b |
| subtotal | 744 | 256M | 209.29 | 239.44 | ||||||
Building the quality gate itself: a scenario whose plan plants an assertion-free test with a lying name plus a DRY violation. Failed through five distinct suppression mechanisms during development (controller pre-judging, severity pre-rating, reviewer calibration, implementer framing, and a wrong eval bar) — each fixed in the prompts generally, not by teaching to the test.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 01:06 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | indeterminate | 9 | 2.5M | 2.51 | O$1.52 S$0.99 | 3.45 | 47df |
| 06-10 01:48 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | pass | 8 | 3.6M | 2.62 | O$2.24 H$0.37 | 3.31 | 0c68 |
| 06-10 04:19 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | fail | 12 | 3.1M | 3.31 | O$2.07 S$1.24 | 3.77 | 9a13 |
| 06-10 04:49 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | indeterminate | 12 | 3.8M | 2.85 | O$2.04 S$0.47 H$0.34 | 3.51 | bc1c |
| 06-10 07:53 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | fail | 8 | 2.1M | 2.32 | O$1.74 S$0.58 | 2.58 | 445a |
| 06-10 09:09 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | fail | 8 | 2.3M | 2.34 | O$1.80 S$0.55 | 2.69 | c127 |
| 06-10 09:20 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | fail | 8 | 2.5M | 2.59 | O$2.05 S$0.54 | 2.84 | 8817 |
| 06-10 09:35 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | pass | 11 | 2.6M | 2.14 | O$1.67 S$0.26 H$0.21 | 2.58 | 719d |
| 06-10 10:44 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | fail | 7 | 1.6M | 1.73 | O$1.17 S$0.55 | 2.16 | b6ab |
| 06-10 11:50 | sdd-quality-reviewer-catches-planted-defect | branch state | claude | pass | 13 | 3.2M | 3.12 | O$2.03 S$1.09 | 3.50 | d8c1 |
| 06-10 19:10 | sdd-quality-reviewer-catches-planted-defect | combo | claude | pass | 10 | 2.9M | 2.77 | O$1.77 S$0.99 | 3.14 | 2ad0 |
| subtotal | 106 | 30M | 28.29 | 33.52 | ||||||
sdd-rejects-extra-features ($1.31–1.37 vs $1.88 baseline) and spec-reviewer-catches-planted-flaws: quality gates every config iteration had to keep passing.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 00:33 | sdd-rejects-extra-features | branch state | claude | pass | 6 | 3.4M | 2.04 | O$1.70 H$0.33 | 2.46 | 52b8 |
| 06-10 01:04 | spec-reviewer-catches-planted-flaws | branch state | claude | pass | 1 | 0.2M | 0.44 | O$0.44 | 0.60 | 0aa2 |
| 06-10 07:53 | sdd-rejects-extra-features | branch state | claude | pass | 5 | 2.1M | 1.37 | O$1.12 H$0.25 | 1.64 | 3a9d |
| 06-10 08:00 | spec-reviewer-catches-planted-flaws | branch state | claude | pass | 2 | 0.2M | 0.50 | O$0.50 | 0.61 | 75b9 |
| 06-10 10:53 | sdd-rejects-extra-features | branch state | claude | pass | 5 | 2.0M | 1.31 | O$1.09 H$0.23 | 1.51 | 8bf5 |
| 06-10 11:00 | spec-reviewer-catches-planted-flaws | branch state | claude | pass | 1 | 0.2M | 0.50 | O$0.50 | 0.62 | 29f1 |
| subtotal | 21 | 8M | 6.16 | 7.44 | ||||||
Hand the FINAL whole-branch reviewer a package file too. WIN: 33 turns/23 tools → 6 turns/3 tools at opus prices; opus bill −19%. Landed.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 10:44 | sdd-go-fractals | branch state | claude | pass | 44 | 13.4M | 11.67 | O$7.13 S$4.54 | 12.57 | a8d1 |
| 06-10 16:39 | sdd-go-fractals | lean-ctl | claude | pass | 50 | 13.7M | 11.73 | O$5.78 S$5.96 | 12.94 | 9a73 |
| subtotal | 95 | 27M | 23.40 | 25.51 | ||||||
Guidance-only model selection decays: dispatches from Task 3 onward inherited opus (+$5). Led to the REQUIRED model: template line. Landed.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 16:39 | sdd-go-fractals | file-handoffs | claude | pass | 47 | 14.9M | 17.43 | O$16.18 S$1.25 | 18.47 | a65c |
| subtotal | 47 | 15M | 17.43 | 18.47 | ||||||
Parallel-call pipelining DEAD (0 paired dispatches in 29 — the controller emits exactly one tool call per message). Background pipelining works mechanically (7/28 dispatches) but benefit below run-to-run noise. Declined.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 17:17 | sdd-go-fractals | pipelined | claude | pass | 49 | 14.1M | 13.43 | O$8.27 S$5.16 | 14.68 | 9d88 |
| 06-10 18:09 | sdd-go-fractals | pipelined | claude | pass | 54 | 17.8M | 14.84 | O$8.75 S$6.08 | 16.08 | 310e |
| subtotal | 103 | 32M | 28.26 | 30.76 | ||||||
Ledger in <git-dir>/sdd/progress.md, maintained well, run cost in-band; guards against post-compaction re-dispatch (worst real-session failure: 269 dispatches for ~22 tasks). Landed.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 16:52 | sdd-go-fractals | ledger | claude | pass | 48 | 15.8M | 13.37 | O$8.85 S$4.43 H$0.09 | 14.45 | b260 |
| subtotal | 48 | 16M | 13.37 | 14.45 | ||||||
All surviving changes together (briefs, packages, model lines, ledger, recipes, risk budget): gates green, fractals band $11.67–14.84 over repeated runs.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 17:44 | sdd-go-fractals | combo | claude | pass | 55 | 14.4M | 12.81 | O$7.41 S$5.40 | 13.96 | 98b7 |
| 06-10 19:10 | sdd-go-fractals | combo | claude | pass | 54 | 16.6M | 14.31 | O$8.67 S$5.64 | 15.51 | 941a |
| 06-10 19:10 | sdd-svelte-todo | combo | claude | pass | 55 | 19.3M | 14.99 | O$7.76 S$7.23 | 16.50 | 2824 |
| subtotal | 164 | 50M | 42.11 | 45.98 | ||||||
Does the brief/report mechanism cost money? NO: lean $11.61/$12.27/$13.25 vs the combo band + the $14.10 same-night control — fully overlapping. Briefs stay, justified by fidelity + compaction durability. Also resolves: the 44.4-min e355795 run was a tail draw, not a better config.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 22:06 | sdd-go-fractals | lean | claude | pass | 51 | 17.0M | 13.25 | O$7.30 S$5.95 | 14.72 | e478 |
| 06-10 22:06 | sdd-go-fractals | lean | claude | pass | 43 | 13.8M | 11.61 | O$7.17 S$4.44 | 12.57 | ee4d |
| 06-10 22:06 | sdd-go-fractals | lean | claude | pass | 49 | 15.8M | 12.27 | O$6.46 S$5.81 | 13.41 | 751f |
| 06-10 22:06 | sdd-go-fractals | combo | claude | pass | 60 | 18.4M | 14.10 | O$6.95 S$7.15 | 15.51 | d170 |
| subtotal | 202 | 65M | 51.23 | 56.21 | ||||||
Hand-rewritten 7-task plan + Global Constraints header + per-task Interfaces lines: $9.51/$12.65/$12.65, dispatches 28 → 20-24 (−21%), fix waves flat, gates 3/3. VALIDATED in effect — the follow-up is making writing-plans elicit this shape.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 22:07 | sdd-go-fractals-crisp | combo | claude | pass | 51 | 16.8M | 12.65 | O$6.33 S$6.22 H$0.11 | 13.67 | f21c |
| 06-10 22:07 | sdd-go-fractals-crisp | combo | claude | pass | 38 | 12.1M | 9.51 | O$4.93 S$4.50 H$0.08 | 10.30 | 45f0 |
| 06-10 22:59 | sdd-go-fractals-crisp | combo | claude | pass | 50 | 17.1M | 12.65 | O$6.70 S$5.88 H$0.06 | 13.68 | 15a0 |
| subtotal | 139 | 46M | 34.80 | 37.65 | ||||||
Controller session on sonnet: $6.68/$8.05 (~−40% vs band), no turn inflation, clean judgment trail (caught go-mod-tidy silently dropping cobra), heavier-and-sane haiku tiering. Zero BLOCKED/⚠️ events arose, so escalation behavior is untested — full N=5 + judgment-audit gates still owed. The $0.13 fail was a quorum seeding bug (API-key picker defaults No) — found, fixed, now superseded by upstream's equivalent fix.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 22:47 | sdd-go-fractals | combo | claude-sonnet | indeterminate | 0 | — | 0.00 | 0.13 | 48a0 | |
| 06-10 22:54 | sdd-go-fractals | combo | claude-sonnet | pass | 31 | 12.6M | 6.68 | S$6.41 H$0.27 | 7.74 | 5038 |
| 06-10 23:27 | sdd-go-fractals | combo | claude-sonnet | pass | 41 | 16.1M | 8.05 | S$7.48 H$0.57 | 9.46 | 2175 |
| subtotal | 72 | 29M | 14.72 | 17.33 | ||||||
DEAD as pre-registered: planted-defect 2 pass / 1 indet / 2 fail (baseline 5/5); 0 of 10 defects cleanly flagged — haiku ADVOCATES for defects (DRY praised as YAGNI; assert-nothing test 'plan-compliant'; found-then-downgraded with the exact prohibited rationale). The ~$2-3/run saving in the fractals run and the quality failure are the same mechanism: lax reviews trigger fewer fix waves.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 23:28 | sdd-quality-reviewer-catches-planted-defect | haiku-rev | claude | indeterminate | 12 | 3.4M | 3.12 | O$2.50 S$0.49 H$0.13 | 3.70 | 11cf |
| 06-10 23:28 | sdd-quality-reviewer-catches-planted-defect | haiku-rev | claude | fail | 13 | 3.8M | 3.11 | O$2.37 S$0.55 H$0.19 | 3.88 | ed05 |
| 06-10 23:28 | sdd-quality-reviewer-catches-planted-defect | haiku-rev | claude | pass | 11 | 2.9M | 2.62 | O$1.95 S$0.55 H$0.12 | 3.13 | be87 |
| 06-10 23:44 | sdd-quality-reviewer-catches-planted-defect | haiku-rev | claude | fail | 12 | 3.4M | 2.76 | O$2.25 S$0.34 H$0.18 | 3.34 | 6a71 |
| 06-10 23:44 | sdd-quality-reviewer-catches-planted-defect | haiku-rev | claude | pass | 9 | 2.5M | 2.31 | O$1.72 S$0.47 H$0.12 | 2.59 | 4241 |
| 06-10 23:58 | sdd-go-fractals | haiku-rev | claude | pass | 39 | 14.1M | 9.61 | O$5.09 S$4.10 H$0.43 | 10.55 | 0297 |
| subtotal | 96 | 30M | 23.53 | 27.19 | ||||||
Second combo svelte run: $20.30 (9 fix waves — review-strictness variance; all 34 dispatches model-disciplined). PR claim becomes the honest range $14.99–20.30 vs $20.98 baseline — time/tokens clearly better, cost overlaps at the top.
| When (UTC) | Scenario | Config | Agent | Verdict | Min | Tokens | Coding $ | Split | Total $ | Run |
|---|---|---|---|---|---|---|---|---|---|---|
| 06-10 22:07 | sdd-svelte-todo | combo | claude | pass | 70 | 24.1M | 20.30 | O$12.35 S$7.94 | 22.36 | c47f |
| subtotal | 70 | 24M | 20.30 | 22.36 | ||||||