Cricket Bench

The setup: one game, fixed bots, a two-second scoreboard

Both benchmarks share one game, one set of dice and a fixed group of opponents. The model under test controls its own code and its own research. It does not control the engine, the opponents or the scoring.

Scaffolds 7b098ec (strategy writing) · 881e746 (self-training)Engines fast_sim.c · game.pyOpponents two fixed 11-bot pools

A benchmark is only as good as what it holds fixed. In Cricket Bench that is the game, the dice, the opponents and the scoring code. Each run starts from a pinned copy of the same starting code and plays the same bots. This chapter describes those shared parts. Chapters 18 and 16 describe how each benchmark uses them.

19.1The game

Cricket uses seven targets: 15, 16, 17, 18, 19, 20 and the bull. The bull counts 25 points.

The miss mechanic is the heart of the game. A strategy chooses where to aim, not what it hits. Each dart samples its outcome from a skill profile. Aiming at a triple usually lands a single or a miss.

Outcome probabilities when aiming at a triple (config.py; bull hits are further scaled by 0.75)
ProfileTripleDoubleSingleMissMarks per round
pro0.410.200.250.145.64
good0.300.220.300.184.92
amateur0.150.200.350.303.60

Scale check. The amateur profile averages 3.6 marks per round. Real bar-league players usually average between 1 and 2. This benchmark is about play far above pub level.

Sources: autoresearch_strategies/loop.md "Game rules" at scaffold 7b098ec; profile table from config.py.

19.2Two engines

The two benchmarks play the same rules on different engines, because they need different speeds.

fast_sim.c · strategy writing

C simulator · one full benchmark call in about two seconds

A strategy is a new branch inside the C engine. One benchmark call plays 825,000 games: 11 opponents, 3 skill profiles, 25,000 games each. The speed lets a run test one idea per iteration for 115 iterations.

game.py · self-training

Python rules of record · byte-identical in every self-training repo

A learning agent plays through the Python game API. It is slower, but it is the reference implementation, and any agent can call it without writing C.

The two engines agree to within 1.8 points at 5,000 games per matchup. The head-to-head ladder (chapter B) plays each pair on one engine and records which.

19.3The fixed opponents

Every run plays the same bots, and no run may change them.

The two fixed 11-bot pools
PoolBotsSkillGames
A · strategy writingE12, E10, E3, S2, PS, S6, S10, S14, E1, E11, S1amateur, good and pro, matched25,000 per matchup per profile
B · self-trainingS2, S6, S10, S14, S16, E1, E3, E5, E9, E10, E11pro vs pro only≥ 2,000 per matchup

The S-bots are Frongello's parameterised family. Each combines a scoring threshold, an "extra darts" redirect and a "chase" rule that follows the opponent's closed targets. The E-bots are later experiments. E12 was the best classic bot, so it is the baseline for strategy writing, at 54.6%. Pool B is the pool that a hand-tuned actor-critic failed against in February. On pool B the best hand-coded bot scores about 55%, so 55% is the self-training stretch goal.

Do not compare pool means across benchmarks. The two pools differ in bots, skill profiles and game counts, so 75% on pool A and 75% on pool B mean different things. The ladder in chapter B puts champions from both benchmarks on one scale.

19.4The bull bug and the patch

The engine had one real defect. In fast_sim_wrapper.py, the bull outcome builder removed the triple entry before it applied the 0.75 bull difficulty penalty. So a triple aimed at the bull skipped the penalty. That is worth about 4 points of mean win rate to any strategy that aims triple at the bull.

The Fable 5.1 max run found it from the code at its second iteration. Some runs flagged it and avoided it. Others used it, and one wrote it up as a law of the game. We did not fix the engine mid-benchmark, because that would have changed the task under the runs. Instead we score every champion twice:

Clean champions, which never aim triple at the bull, reproduce exactly under the patch. That confirms the patch changes nothing else. Every result on this site uses the patched score unless it says otherwise.

Source: Darts-Cricket/benchmark_review/README.md, "Engine artifact" and "Bull-artifact rebench (2026-09-22)".

19.5Isolation rules for every run

Source: runs/README.md (protocol section).

19.6An independent review of the harness

Before the model comparisons, a different model audited the harness. It copied the engine to an isolated directory, rebuilt it, and re-ran everything. On 6 September GPT-6 Astra reproduced 29,550,000 games in 48 seconds. Every cell matched the checked-in results exactly. It filed six findings, two of them P1.

⊗

The review passed every reproduction check and still missed the bull bug. Reproducing a harness exactly is not the same as validating it. The bug came to light because a model under test read the engine closely.

Sources: benchmark_review/REVIEW.md; README "Engine artifact"; astra_efforts/high/review/REVIEW.md.