The earlier exploiters ran a quarter of their targets' budget. Rerun at a matched
~131M learner actions (measured: 130.4M and 131.1M):
ours exploiter reaches 0.4657 [0.455, 0.477], mean lead -5.5
league exploiter reaches 0.6295 [0.619, 0.640], mean lead +28.8
The CIs are nowhere near each other and the gap widened under funding, 0.094 ->
0.164. A qualitative line gets crossed too: a properly funded attacker trained
from scratch *beats* league outright, and still loses to ours.
Two things a reader has to carry with that, both of which cut against us:
- Neither exploiter plateaued -- both climb through the final update, so these
remain lower bounds. And ours' attacker has the steeper tail (+0.031 vs +0.016
per 100 updates) while league's curve is already bending, so a much larger
attacker could narrow the gap. Cannot rule out that it closes.
- League is a single-round policy blind to carry and match score, and both frozen
policies are played greedily, which is maximally exploitable by construction.
So this is not a clean "training method A vs B".
The claim that survives is the matched-budget one, and only that one.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBQKgvBbxbheiTF1AVy1Sh