Refactor docs/performance.md per librarian routing rule

docs/performance.md grew to 914 lines because dated experiments and
design analyses kept getting appended instead of routed to
docs/archive/ and docs/research/ as AGENTS.md prescribes. The
oversize check from librarian Stage 1 surfaced the file; this commit
acts on that finding by extracting the parts that belong elsewhere
and trimming the source to a focused current-state reference.

Extracts (verbatim from the original prose, with cross-link headers
and a brief routing note added at top):

- docs/archive/deep-cfr-performance-experiments-2026-05-07.md
  bundles torch.compile (regression), AMP (regression), GPU forward
  profiling (decision support), and Option B interleaved traversal
  (pass) — same date, same theme.
- docs/research/batched-traversal-inference-decision.md captures the
  durable A vs B vs C rationale with a closing "Outcome" pointer to
  the post-bench archive doc.
- docs/archive/post-a-optimization-calculus-2026-05-07.md preserves
  the forward-looking sequencing recorded pre-bench.
- docs/archive/option-a-bench-result-2026-05-07.md preserves the
  regression diagnosis and re-enable criteria.

docs/performance.md is now 345 lines, holds sections 1–9 (current
runtime / bottleneck / device / AMP status / batching / eval /
TensorRT / priorities), and ends with a "See Also" linking the four
extracts.

Also reworded the AGENTS.md soft-cap rule from a bare "~500-line
soft cap" to clarify the intent: the cap is a *routing trigger* (is
content piling up that should live in archive/research?), not a
split mandate. Reduces the risk of future agents shredding a useful
doc just to satisfy a number.

scripts/librarian.sh now exits 0 against the working tree.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-05-07 23:57:39 +09:00
co-authored by Claude Opus 4.7
parent 6ecb233bdd
commit 1cd9950bd3
7 changed files with 703 additions and 595 deletions
@@ -0,0 +1,187 @@
# Option A Bench Result and Structural Ceiling (2026-05-07)
**Source:** Extracted from `docs/performance.md` § "Option A Bench
Result and Structural Ceiling" on 2026-05-08.
Records the post-implementation Option A benchmark (regression: 0.21×
traversal), the diagnosis of the sync-blocking policy boundary as the
real ceiling, and the criteria under which to revisit Option A.
**Related:**
- Design rationale: `docs/research/batched-traversal-inference-decision.md`
- Forward-looking sequencing (pre-bench): `docs/archive/post-a-optimization-calculus-2026-05-07.md`
- Trainer-side experiments from the same date: `docs/archive/deep-cfr-performance-experiments-2026-05-07.md`
---
Option A (`traversal.inference_backend: server`) was implemented and
benchmarked. **Result: regression. A is deferred. `default.yaml` stays on
`local`. The implementation is preserved behind the flag for future revisit.**
## Bench numbers
`scripts/bench_inference_backend.py --device cuda --iterations 5 --warmup 1`,
RTX 3090, after a per-call IPC fix (replaced `multiprocessing.Manager()`
queues/events with spawn-context primitives, slot reuse per worker batch,
shared memory confirmed in use for state/response payloads).
| Backend | iter | traversal | adv_train | strat_train | mem_add | batch_tensor |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| `local` | 16.75s | 10.81s | 3.91s | 2.02s | 0.95s | 3.56s |
| `server` | 57.61s | 51.61s | 3.96s | 2.02s | 0.50s | 3.57s |
| Speedup | 0.29× | 0.21× | 0.99× | 1.00× | 1.89× | 1.00× |
Raw: `runs/bench/2026-05-07_193335_inference_backend/results.json`.
Training and eval phases are unchanged (as expected — A only touches the
traversal forward path). The regression is contained in `traversal_seconds`,
which is ~5× worse.
## Diagnosis
The server emits per-flush batch stats. **Mean batch size: ~7.27.9, max 8.**
This is the structural ceiling, not a tunable misconfiguration:
- Traversal recursion is **sync-blocking** at the policy call site. Each
worker has at most one in-flight policy request at a time.
- In-flight requests at the server ≤ `num_workers` = 8.
- The server further splits each batch by `(network_kind, network_index)`,
so the actual GPU forward group size is roughly half of that — about 4
rows per group.
Per-state cost at this realized batch size, from the GPU profile table:
| Realized batch | μs/state |
| ---: | ---: |
| 1 | 80.07 |
| 4 | 20.30 |
| 8 (extrapolated) | ~12 |
| 64 | 1.46 |
| 256 | 0.34 |
So the GPU is doing ~12μs per state instead of the projected ~1.5μs at
bs=64. The IPC round-trip per call (queue post + server scheduler + event
wakeup, even with shared-memory payload) is on the order of hundreds of μs
per call, which exceeds both the local CPU forward (~80200μs at bs=1 on
this small MLP) and the marginal GPU gain. Net: per-call cost roughly
doubles or triples, compounded across ~205k calls/iter, gives the observed
5× traversal regression.
`batch_window_us` and `max_batch` tuning cannot escape this ceiling —
there are simply not 64 concurrent in-flight requests to coalesce when only
8 workers are blocking-sync.
## What this means for the headline GPU profile (`scripts/profile_gpu_forward.py`)
The earlier "230× speedup at bs=256" is a **per-state GPU forward**
microbenchmark, not an end-to-end traversal speedup. Realizing that gain
requires *actually feeding the GPU* with bs=64+ batches. Sync-blocking
multi-worker traversal cannot do that. Reaching the bs=64 regime needs
either:
- Per-worker traversal interleaving (worker advances `worker_chunk_size`
traversals concurrently, suspending at each policy call — Option B
shape), which requires turning Cython traversal recursion into a
resumable state machine. Same scope as a partial Option C, localized to
worker scope.
- Option C proper (single-process vectorized traversal).
Both require restructuring traversal. Option A's "additive, no traversal
changes" property turned out to also mean "cannot drive the batch size up."
## Clarifying the traversal bottleneck: sync policy boundary, not SIMD
The tempting shorthand is "Python/GIL prevents traversal from using SIMD or
threads." The more precise diagnosis is narrower:
- Lost Cities game mechanics are already mostly Cython C-level operations.
`legal_actions`, action push/pop, and cached scoring are not Python list
walks on the hot path.
- The traversal recursion is Cython, but it synchronously crosses back into
Python/PyTorch at every policy-needed state: encode a single info state,
run one-row PyTorch forward, copy logits back to CPU/Numpy, then continue
recursion.
- This boundary makes every traversal worker **sync-blocking**. With
`num_workers=8`, the inference server can see at most eight in-flight
requests before per-network splitting, no matter how large `max_batch` is.
- GIL-free threading would help only after the same path is made
`nogil`-clean or after traversal is restructured so policy calls can be
batched. Simply "using SIMD" does not address the one-row policy boundary.
So the actionable bottleneck is **policy-call scheduling shape**, not scalar
game-rule arithmetic. The highest-leverage experiment is Option B:
per-worker interleaved traversal, where one worker advances many traversals,
suspends each at a policy request, batches those requests, and resumes the
corresponding continuations.
Microbench evidence (2026-05-07, `configs/deep_cfr/default.yaml`,
`experiments/traversal_policy_boundary/bench_policy_boundary.py`):
| Device | Component | median μs/call | p95 μs/call |
| --- | --- | ---: | ---: |
| CPU | encode + legal | 3.10 | 3.81 |
| CPU | push + pop | 0.15 | 0.22 |
| CPU | policy boundary bs=1 | 111.50 | 125.46 |
| CPU | torch forward bs=64 | 12.84 | 13.14 |
| CUDA | policy boundary bs=1 | 181.30 | 194.77 |
| CUDA | torch forward bs=64 | 2.55 | 2.75 |
This confirms the bottleneck is not Cython game-rule scalar work. The
single-request policy boundary is ~36× larger than encode+legal on CPU, while
CUDA bs=64 forward is ~71× cheaper than the current CUDA bs=1 boundary.
Expected upside is bounded by the fraction of traversal currently spent at
policy calls. Moving realized GPU forward from the current ~4-8 row regime
(~12-20μs/state) to bs=64 (~1.46μs/state) is an ~8-14× improvement on the
forward component, but not on game recursion, sample creation, or replay
writes. For the observed `local` traversal around 10-13s/iter, a realistic
first target is roughly **1.5-3× traversal speedup** if Option B reaches the
bs=64 regime without adding comparable scheduler overhead. Larger claims need
a prototype because traversal has substantial non-forward work.
## Why deferring A (not deleting) is the right call
- The plumbing (server process, shared-memory client, weight sync, config
flag) is complete and tested. Re-enabling is a config flip.
- The fundamental issue at this model size is that **GPU forward time is
too small to amortize IPC overhead** at any realistic batch size we can
drive without restructuring traversal. Bigger model changes that
arithmetic; the same plumbing then becomes useful.
- The `mem_add_seconds` row showed a real 1.89× win, suggesting the
shared-memory replay-write path adopted along the way is worth keeping
even with `local` backend. (Confirm separately; this is a side effect.)
## Re-enable A when one of these holds
1. **Model grows** to ~1024 hidden / ~6 layers (compile/TRT discussion in
`docs/archive/post-a-optimization-calculus-2026-05-07.md`). Forward
time scales with FLOPs while IPC overhead is fixed; at some point IPC
becomes a small fraction.
2. **Per-worker interleaved traversal** ships (Option B-shape refactor).
Drives realized batch toward 64 and reclaims the profile table's gains.
3. **Eval becomes the dominant phase** (`eval_every: 5`,
`evaluation.games: 1000`). Eval is not sync-blocking traversal; it is
already batch_size=64 in eval code. The same inference server can
serve eval directly without the worker-side ceiling.
## Free-threaded Python (3.13t) note
Free-threaded Python + Cython `nogil` would let many threads (well above
core count) run game logic concurrently in one process, with shared memory
and no IPC. With ~64 threads sync-blocking on policy calls, the server
would naturally see bs=64. **In principle this is the cleanest endpoint.**
In practice as of early 2026:
- Free-threaded Python is an opt-in build (`python3.13t`), still
experimental, with measurable single-thread overhead.
- PyTorch's free-threaded compatibility is partial.
- Cython `nogil`-cleanliness audit on the existing game engine is still
required and was the original reason `nogil` threading was deferred in
the design decision.
- No project-level adoption pressure on `python3.13t` today.
So free-threaded Python does change the architectural answer, but it does
not unblock A *now*. Track the ecosystem; revisit when (a) `python3.13t`
becomes mainstream or (b) the Cython engine is `nogil`-cleaned for other
reasons.