A benchmark is only as good as what it holds fixed. In Cricket Bench that is the game, the dice, the opponents and the scoring code. Each run starts from a pinned copy of the same starting code and plays the same bots. This chapter describes those shared parts. Chapters 18 and 16 describe how each benchmark uses them.
19.1The game
Cricket uses seven targets: 15, 16, 17, 18, 19, 20 and the bull. The bull counts 25 points.
- Closing. A player needs three marks on a target to close it. A clean triple gives three marks, a double two, a single one.
- Scoring. Once you close a target that your opponent has not closed, every extra mark you land on it scores its face value. When both players have closed a target, nobody scores on it.
- Winning. You win when you have closed all seven targets and your score is at least your opponent's. If you close everything while behind, the game goes on.
- Turns. Each turn is three darts. The players alternate, and successive games alternate who throws first.
The miss mechanic is the heart of the game. A strategy chooses where to aim, not what it hits. Each dart samples its outcome from a skill profile. Aiming at a triple usually lands a single or a miss.
| Profile | Triple | Double | Single | Miss | Marks per round |
|---|---|---|---|---|---|
| pro | 0.41 | 0.20 | 0.25 | 0.14 | 5.64 |
| good | 0.30 | 0.22 | 0.30 | 0.18 | 4.92 |
| amateur | 0.15 | 0.20 | 0.35 | 0.30 | 3.60 |
Scale check. The amateur profile averages 3.6 marks per round. Real bar-league players usually average between 1 and 2. This benchmark is about play far above pub level.
Sources: autoresearch_strategies/loop.md "Game rules" at scaffold 7b098ec; profile table from config.py.
19.2Two engines
The two benchmarks play the same rules on different engines, because they need different speeds.
fast_sim.c · strategy writing
A strategy is a new branch inside the C engine. One benchmark call plays 825,000 games: 11 opponents, 3 skill profiles, 25,000 games each. The speed lets a run test one idea per iteration for 115 iterations.
game.py · self-training
A learning agent plays through the Python game API. It is slower, but it is the reference implementation, and any agent can call it without writing C.
The two engines agree to within 1.8 points at 5,000 games per matchup. The head-to-head ladder (chapter B) plays each pair on one engine and records which.
19.3The fixed opponents
Every run plays the same bots, and no run may change them.
| Pool | Bots | Skill | Games |
|---|---|---|---|
| A · strategy writing | E12, E10, E3, S2, PS, S6, S10, S14, E1, E11, S1 | amateur, good and pro, matched | 25,000 per matchup per profile |
| B · self-training | S2, S6, S10, S14, S16, E1, E3, E5, E9, E10, E11 | pro vs pro only | ≥ 2,000 per matchup |
The S-bots are Frongello's parameterised family. Each combines a scoring threshold, an "extra darts" redirect and a "chase" rule that follows the opponent's closed targets. The E-bots are later experiments. E12 was the best classic bot, so it is the baseline for strategy writing, at 54.6%. Pool B is the pool that a hand-tuned actor-critic failed against in February. On pool B the best hand-coded bot scores about 55%, so 55% is the self-training stretch goal.
Do not compare pool means across benchmarks. The two pools differ in bots, skill profiles and game counts, so 75% on pool A and 75% on pool B mean different things. The ladder in chapter B puts champions from both benchmarks on one scale.
19.4The bull bug and the patch
The engine had one real defect. In fast_sim_wrapper.py, the bull outcome builder removed the triple entry before it applied the 0.75 bull difficulty penalty. So a triple aimed at the bull skipped the penalty. That is worth about 4 points of mean win rate to any strategy that aims triple at the bull.
The Fable 5.1 max run found it from the code at its second iteration. Some runs flagged it and avoided it. Others used it, and one wrote it up as a law of the game. We did not fix the engine mid-benchmark, because that would have changed the task under the runs. Instead we score every champion twice:
- Raw. The score on the scaffold engine, as the run saw it.
- Patched. The same champion on a copy of the engine where a triple aim at the bull uses the double-aim bull distribution.
Clean champions, which never aim triple at the bull, reproduce exactly under the patch. That confirms the patch changes nothing else. Every result on this site uses the patched score unless it says otherwise.
Source: Darts-Cricket/benchmark_review/README.md, "Engine artifact" and "Bull-artifact rebench (2026-09-22)".
19.5Isolation rules for every run
- A fresh repo. Each run gets its own git repository, built from a pinned scaffold commit. Since wave 2 the repo holds only the scaffold's own history, with no other branches, so a run cannot read another run's work.
- A pinned model and effort. The model and reasoning effort are fixed per run and recorded in the run's first commit or its launch record.
- No memory tools. Claude runs start with
--strict-mcp-config, and Codex runs with--ignore-user-config. Both drop the user's MCP servers, including the long-term memory service. Codex memories are off by default. - No other runs. The prompt forbids reading other branches, worktrees or project folders. The self-training task says: "your independent judgment is the thing being measured."
- The journal is the only memory. Long runs are split into sessions. Each new session starts cold and re-reads the run's own journal.
- What still loads. Both engines load the user's global instruction file (
~/.claude/CLAUDE.mdor~/.codex/AGENTS.md) and the project'sCLAUDE.md. Those files are the same for every run of an engine.
Source: runs/README.md (protocol section).
19.6An independent review of the harness
Before the model comparisons, a different model audited the harness. It copied the engine to an isolated directory, rebuilt it, and re-ran everything. On 6 September GPT-6 Astra reproduced 29,550,000 games in 48 seconds. Every cell matched the checked-in results exactly. It filed six findings, two of them P1.
The review passed every reproduction check and still missed the bull bug. Reproducing a harness exactly is not the same as validating it. The bug came to light because a model under test read the engine closely.
Sources: benchmark_review/REVIEW.md; README "Engine artifact"; astra_efforts/high/review/REVIEW.md.