Files
coorl-lost-cities/docs/plans
coolguyandClaude Opus 4.8 4392ec255a Record the old-vs-new head to head in the real three-round game
The question was what any of this actually improved over the training method
that existed. Duplicate matches, 8192 a piece, same three deals and coins from
both seats:

- At 39.3M learner actions the match stack beats the Phase 0a gate agent
  (0.5842) which had 411M -- 10.5x the data. Sample efficiency is the headline.
- At 39.3M it *loses* to the league policy (0.3142). That is a budget gap, not a
  strength gap: league had 122.6M plus a league/exploiter structure.
- Scaled to a matched budget (131M vs league's 122.6M) it wins: 0.6094
  (CI 0.599-0.620), +22.0 points.

So: same compute, stronger agent, measured on the actual game.

Caveat kept honest in the plan -- league was trained with exploiters, and we have
measured average strength, not exploitability. "Harder to exploit" is not shown.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBQKgvBbxbheiTF1AVy1Sh
2026-07-15 03:30:23 +09:00
..
2026-07-14 20:09:03 +09:00