A.1Every run
One row per run, in both benchmarks and all three waves. Pool scores are bull-patched. Pool scores from the two benchmarks use different pools, so compare them within a benchmark only.
All runs
A.2Download the data
- runs_all.csv: one row per run. It gives the benchmark, wave, model, effort, harness, host, iterations, discovery iteration, pool score, the champion at the 115-iteration budget, ladder Elo and output tokens.
- curves_all.csv: one row per strategy-writing iteration, all waves. It gives the candidate, its mean win rate, the keep decision, the bull flag and the patched score.
- tokens.csv: output, input and cache tokens per run, with coverage and the transcript source.
- ladder_ratings.csv and ladder_h2h.csv: the ladder ratings and the full head-to-head matrix.
- Wave-1 detail: heuristic_arms.csv, heuristic_curves.csv, rl_arms.csv and the July ladder, elo_ladder_2026-07-02.csv.
A.3How the numbers were checked
- Read-only sources. Every run was read from its own repo: the journal, the commit log and the result files. No summary was trusted without its source.
- Every champion was re-run. Each strategy-writing champion was benched twice at 25,000 games per matchup: once on the scaffold engine and once with the bull bug patched. Every journal score reproduced.
- The bull flag is tested. A strategy that never aims a triple at the bull gives bit-identical results under the patch. So a changed result proves the strategy used the bug.
- Every ladder player passed a gate. Each had to reproduce its own recorded result in the ladder's harness before it was rated (chapter B.2).
- Iterations are benched candidates. Keep, discard and noted candidates all count. Results use each run's champion as of its 115th candidate.
- Tokens come from transcripts. Output tokens are summed from Claude Code transcripts, deduplicated by message, and from Codex session logs. They include thinking and reasoning tokens. The two engines count them differently.
A.4Deviations from protocol
Every deviation below is also noted where it matters. Together they are the main reason to read cross-model comparisons as descriptive.
| Run | Deviation | Effect |
|---|---|---|
| GPT-6 Astra, wave 1, strategy writing | Run by a driver script. The model proposed each strategy, and the script built, benched and kept it. | Measures proposal quality, not agentic work. Token counts are not comparable with full agent sessions. Waves 2 and 3 run Astra as a full agent. |
| GPT-6 Astra, wave 1, self-training | A coordinator framed the deliverable as "a reproducible environment first". | A different objective. All three efforts stopped far inside the time budget. |
| Claude Fable 5, wave 1 | After 25 iterations the run expanded its own opponent pool, seven times, to 19 bots. | Only iterations 1 to 25 are comparable. |
| Claude Opus 5, wave 1 | Counted multi-configuration sweeps as single iterations. | Only 106 of the 115 claimed iterations can be identified. |
| Claude Opus 5.5 high, wave 1 | The runner counted only keep and discard commits. The run also gave itself within-game opponent memory. | 125 strategies for a reported 115. Its champion at 115 is the same, X189. |
| Claude Opus 5.5 medium, wave 1 | A session hit a usage limit, and the runner started another full session. | 129 iterations. Its champion at 115 is X212 (74.6%). Its final champion X227 scores 76.6%. |
| Claude Fable 5.1 max, wave 1 | Stopped at 56 iterations after the run declared its rule family exhausted. | The fewest iterations of the modern runs. Wave 3 adds three full-length max runs. |
| All wave-2 strategy runs | The loop's 10-second hang rule discarded good candidates on the shared machines. The prompt now sets 180 seconds, and all 16 runs restarted from scratch at 16:15 ET on 24 September. | The aborted attempts are kept outside the runs as git bundles, tagged hang10s. |
| Claude Opus 5.5 medium-b, wave 2 | 16 candidates were committed as "noted", which the runner did not count. | 131 benched candidates. Its champion at 115 is X209 (76.4% patched). |
| GPT-5.6 Luna medium-b, wave 2 | One session scripted a batch of 14 candidates. All failed to compile at the same line, and the script did not stop. | 14 of 115 iterations produced nothing. Kept as a real result: loop.md counts a compile error as a discard. |
| Codex strategy runs, wave 2 | The Codex account hit its weekly usage limit at 17:37 ET on 24 September. | The runners waited. After a reset, work resumed at 17:48 ET. No iteration was lost. |
| GPT-5.6 Luna medium, wave 2, self-training | The first session stopped to ask for design approval, because Codex loaded a global skill that requires sign-off. | The continue prompt says no human is available. Only runs that needed a second session saw that line. |
| Wave-2 self-training runs | Four runs trained on the Mac at once, alongside five strategy runs and the ladder. Threads were capped at 4 per run. | The training budget is wall-clock time, so shared compute may have reduced effective training. |
A.5Known gaps
- Tokens. No transcripts survive for eight early runs, from before a machine migration on 1 August: the older Claude Opus, Fable 5 and Opus 5 strategy runs, and the July self-training runs, including IT3. The wave-1 Astra self-training logs are partial.
- Effort. Effort was not pinned before 6 September. Those runs used the global default, most likely medium.
- Replicates. Strategy-writing cells have one to three runs each, and self-training cells one or two. Cells with one run can mislead (chapter 17.2).
- Seeds. Each strategy run optimises against one fixed seed and pool. A score after 115 selections is partly fitted to that seed.
A.6Glossary
- Run
- One complete attempt by one model, at one effort, on one benchmark, from the pinned starting code. Older notes call it an arm.
- Wave
- A batch of runs launched together. Wave 1 ran from July to September. Wave 2 started on 24 September and wave 3 on 27 September.
- Pool
- The fixed bots a run is scored against. Each benchmark has its own pool (chapter 19.3).
- Pool score
- Mean win rate against the pool: per-opponent win rate, averaged over skill profiles, then over opponents.
- Iteration
- One strategy written and benched. Kept, discarded and noted candidates all count.
- Discovery
- The first iteration at which a run benched a clean candidate at 60% or more. It marks when the run found the lane-shutdown idea.
- Clean champion
- The best strategy a run kept that never aims a triple at the bull.
- Exploit champion
- A champion that aims a triple at the bull. "Patched" means re-scored with the bull bug fixed.
- Elo
- A Bradley–Terry rating from the head-to-head ladder, anchored at S1 = 1000. A 10-point gap is about a 51.4% edge.
- IT3
- Fable 5's league-trained self-training agent from July, and the ladder leader.
A.7History
The benchmark grew out of a seven-month project that started as an attempt to teach a computer to play bar darts. The history page tells that story.