In February, nineteen hand-tuned versions of an actor-critic agent topped out at 49.3% against this pool. The self-training benchmark asks a model to do better on its own. It gets the engine, the bots, a Python interpreter with PyTorch and one task file. The deliverable is a policy that was learned, plus a research journal that says how.
16.1The task
Every run reads the same TASK.md. The benchmark part is fixed:
Mean win rate across all 11 matchups vs this exact pool, at pro skill profile for both players: S2, S6, S10, S14, S16, E1, E3, E5, E9, E10, E11. ≥ 2,000 games per matchup at evaluation time. Alternate who throws first each game. Pass: mean WR > 50%. Stretch: ≥ 55%.
The method is open. The task names value-based RL, policy gradient, search plus learning and imitation plus improvement as examples, and then says "whatever". One line limits it: "Hand-coded heuristics are not the deliverable. The agent must learn its policy. Heuristics are allowed only as opponents, training curricula, or baselines."
Method rules
- Keep
RESEARCH_LOG.md: hypothesis, config, numbers and decision for every run, written before moving on. - Smoke-test the learning signal against one weak bot before any long run.
- Background long runs with
nohup, checkpoint, and poll the log. - Budget: "~4 hours wall-clock of training compute". Prefer three informative medium runs over one blind long one.
- Evaluate the full pool at every checkpoint that matters, per opponent. Commit after every meaningful step.
Hard rules
- Work only inside the run's own directory. Never read other project folders, and never use memory tools. "This is a clean-room experiment: your independent judgment is the thing being measured."
- Do not change the game rules in
game.py, the bots instrategies.py, or the benchmark definition. Wrappers for speed are fine if they are proven bit-identical. - "Report honestly. A negative result with a clear diagnosis beats an inflated claim."
Pool B is harder than the pool in the strategy benchmark. Both players are pros, and it contains the three "chase" bots S10, S14 and S16, which follow the opponent's closed targets. The best hand-coded bot, S2, scores 54.7% against the rest of this pool. That is why 55% is the stretch goal.
16.2What the model gets
| Piece | What it is |
|---|---|
game.py | The rules of record: 7 targets (15 to 20 and the bull), 3 marks to close, scoring on targets the opponent has not closed. A game is won by closing all seven while level or ahead on points. 21 actions: 7 targets × single, double or triple. |
config.py | Skill profiles. An aimed dart samples its outcome from the thrower's profile. The task says it plainly: "Aim ≠ outcome." |
strategies.py | Every bot in the pool, plus the rest of the S and E families. |
agent.py | A legacy tabular Q-learning player. The task calls it scaffolding: "reuse, replace, or ignore it." |
tests/ | 142 passing tests that document engine behaviour. Keep them green, and add tests for new code. |
| Python | PyTorch 2.13 with MPS, numpy and pytest, in a shared read-only virtual environment. |
The repository holds no prior RL work. It does not contain the strategy benchmark's C engine, its champions or any journal from another run.
16.3How we run it
A self-training run is one long agent session, not a loop of short ones. The model plans, writes the trainer, launches background training, polls it, evaluates and writes the journal, all in the same session.
Wave 1 (July to September)
The Claude runs were single Claude Code sessions on this Mac with the task pinned in the first message, and the reasoning effort pinned in a committed .claude/settings.json where it was recorded. The Opus 5.5 runs used this launch line:
claude --model claude-opus-5-5 --effort high --strict-mcp-config 'Read TASK.md end to end and carry out the task in full. …'
The three GPT-6 Astra runs ran under a Codex coordinator. It framed the deliverable as a "reproducible environment first", which is a different objective. All three stopped far inside the budget: the ultra run trained for 8 seconds. Each run was capped at four compute threads.
Wave 2 (24 September)
Four new runs, on the Mac, started by one script, runs/run_rl.sh. It gives every engine the same first prompt:
Read TASK.md end to end and carry out the task in full. You run on model, pinned to reasoning effort effort; state both in the first RESEARCH_LOG.md entry. Finish with a final benchmark (>= 2,000 games per matchup), a journal summary naming the final model and its per-opponent win rates, and a commit. After that final commit, create an empty file named RUN_COMPLETE in this directory.
| Engine | Command | Isolation |
|---|---|---|
| Claude | claude -p "$PROMPT" --model M --effort E --strict-mcp-config --dangerously-skip-permissions | No MCP servers, so no memory tools. No sandbox, the same as the wave-1 Claude runs. |
| Codex | codex exec -m M -c model_reasoning_effort=E --ignore-user-config -c approval_policy="never" | User config and its MCP servers ignored. Memories off. No sandbox, for parity with the Claude runs on the same host. |
- Resuming. If a session exits before
RUN_COMPLETEexists, usually at a usage limit, the runner waits 10 minutes and continues the same session with a second prompt. It says to re-read the task, the journal and the logs, and that "No human is available during this run: make design decisions yourself and do not stop to wait for approval." We added that line after one run stopped to ask for design sign-off. - Compute. Four runs trained on the Mac at once, so each got
OMP_NUM_THREADS,MKL_NUM_THREADSandVECLIB_MAXIMUM_THREADSset to 4. That is the same cap as the Astra wave-1 runs. The Mac was also running other work, so wall-clock budgets bought less compute than in wave 1. - Scratch space. Each run has its own
TMPDIRinside its directory.
16.4How a result is scored
- The run's own benchmark. The run picks its final model, then benchmarks it at 2,000 or more games per opponent, pro against pro, alternating the first throw. The best runs use a fresh seed that played no part in model selection. This number is the run's pool score.
- The adapter gate. For the ladder, each final model gets a small adapter that plays through our own game loop. It must reproduce the run's recorded pool score within 1.0 point before it is rated. All 16 self-training agents passed. The largest gap was 0.43 points, and July's IT3 and IT2_det matched all 11 official cells exactly.
- Head to head. The agent then plays every other champion from both benchmarks in the ladder, with the bull bug patched out. That rating is in chapter B.
A pool score and a ladder rating answer different questions. The pool score says how well an agent exploits eleven known bots. The rating says how strong it is against everything else. Chapter 15 shows where they disagree.
16.5The runs
| Wave | Model | Effort | Run | Harness |
|---|---|---|---|---|
| 1 | Fable 5 | unrecorded | rl-cleanslate (PPO run 2, then the IT3 league) | Claude Code session |
| 1 | Fable 5 | unrecorded | alphazero (a constrained cell: AlphaZero-style self-play only) | Claude Code session |
| 1 | Fable 5 | unrecorded | rl-interp (rules distilled from PPO run 2; analysis, no training) | Claude Code session |
| 1 | Opus (version unrecorded) | unrecorded | rl-opus (planned an actor-critic, trained no policy) | Claude Code session |
| 1 | Fable 5.1 | medium | rl-fable51, rl-fable51-b | Claude Code session |
| 1 | Opus 5.5 | medium | rl-opus55 | Claude Code session |
| 1 | Opus 5.5 | high | rl-opus55-high | Claude Code session |
| 1 | GPT-6 Astra | low, high, ultra | astra_rl low, high, ultra | Codex coordinator |
| 2 | Opus 5.5 | medium | rl-opus55-medium-b | run_rl.sh, one session |
| 2 | GPT-6 Astra | high | rl-astra-high-b | run_rl.sh, one session |
| 2 | Sonnet 5 | medium | rl-sonnet5-medium-a | run_rl.sh, one session |
| 2 | GPT-5.6 Luna | medium | rl-luna-medium-a | run_rl.sh, two sessions |
That is 15 runs, which produced 16 rated agents: the rl-cleanslate run shipped PPO run 2 and then the IT3 and IT2_det league champions. Three cells have two runs: Fable 5.1 medium, Opus 5.5 medium and GPT-6 Astra high. The two Astra high runs used different harnesses. The Luna run needed a second session because its first one stopped to ask for approval.