docs/performance.md grew to 914 lines because dated experiments and design analyses kept getting appended instead of routed to docs/archive/ and docs/research/ as AGENTS.md prescribes. The oversize check from librarian Stage 1 surfaced the file; this commit acts on that finding by extracting the parts that belong elsewhere and trimming the source to a focused current-state reference. Extracts (verbatim from the original prose, with cross-link headers and a brief routing note added at top): - docs/archive/deep-cfr-performance-experiments-2026-05-07.md bundles torch.compile (regression), AMP (regression), GPU forward profiling (decision support), and Option B interleaved traversal (pass) — same date, same theme. - docs/research/batched-traversal-inference-decision.md captures the durable A vs B vs C rationale with a closing "Outcome" pointer to the post-bench archive doc. - docs/archive/post-a-optimization-calculus-2026-05-07.md preserves the forward-looking sequencing recorded pre-bench. - docs/archive/option-a-bench-result-2026-05-07.md preserves the regression diagnosis and re-enable criteria. docs/performance.md is now 345 lines, holds sections 1–9 (current runtime / bottleneck / device / AMP status / batching / eval / TensorRT / priorities), and ends with a "See Also" linking the four extracts. Also reworded the AGENTS.md soft-cap rule from a bare "~500-line soft cap" to clarify the intent: the cap is a *routing trigger* (is content piling up that should live in archive/research?), not a split mandate. Reduces the risk of future agents shredding a useful doc just to satisfy a number. scripts/librarian.sh now exits 0 against the working tree. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
4.8 KiB
Post-A Optimization Calculus (2026-05-07)
Source: Extracted from docs/performance.md § "Post-A Optimization
Calculus (forward-looking, 2026-05-07)" on 2026-05-08.
Forward-looking sequencing recorded before Option A was benched. Has
not been measured. See
docs/archive/option-a-bench-result-2026-05-07.md for the actual bench
which deferred Option A — some assumptions below ("once Option A
lands") have to be re-evaluated in light of that result.
Once Option A lands, the bottleneck shape changes. This section records the expected sequencing for follow-up work. It is forward-looking and has not been measured yet — verify against bench numbers after A is benchmarked.
Why compile / TensorRT are negligible today but become meaningful later
Today (small model: 3-layer, 512 hidden):
torch.compileon the trainer's networks already regressed (see the 2026-05-07 experiment indocs/archive/deep-cfr-performance-experiments-2026-05-07.md). The model is too small for kernel fusion to beat compile dispatch overhead.torch.compile/ TensorRT on the inference-server forward (post-A) would shave ~30–50% off ~90μs/call → ~50–70μs/call. With forward share of an iter reduced to <1% by A's batching, the iter-level multiplier is ~1.00–1.01×. Negligible.
Two compounding shifts can flip this:
- Larger model. Going from 512 hidden / 3 layers to ~1024 hidden / ~6 layers pushes the forward call out of dispatch-bound territory into kernel-bound territory. Compile fusion and TensorRT both deliver real 1.5–2× on the forward call itself once the kernel is large enough to amortize launch overhead. Forward share of iter time also rebalances upward because per-call time scales with FLOPs while batching gain is fixed.
- Denser, larger evaluation. Moving toward
eval_every: 5andevaluation.games: 1000makes evaluation about half of iteration wall-clock (see the amortized eval table indocs/performance.md). Eval is pure inference, so TensorRT on the inference-server's forward path applies directly.
When both shifts happen together, an illustrative future iter (rough order of magnitude only):
| Configuration | Iter time (rough) |
|---|---|
| Today (small model, eval_every=25) | 17.85s |
| + A (batched traversal inference) | ~14s |
| + larger model (≈4× FLOPs), no compile/TRT | ~50s |
| + dense eval (eval_every=5, games=1000) | ~70s |
| + compile (trainer) + TensorRT (inference) | ~45s |
That last row is where compile/TensorRT contributes ~1.5× iter — the same tooling that is iter-neutral today. The numbers above are illustrative; real ratios depend on model size, kernel autotune outcomes, and the eval-vs-train balance.
Tooling split
- TensorRT: applies only to inference (no backward). Targets:
- inference-server forward in traversal,
- inference-server forward in evaluation. Both are served by the same A-era server, so a single TensorRT integration covers both.
torch.compile: applies to trainer's advantage/strategy training (forward+backward+optimizer). The 2026-05-07 regression on a small model does not generalize — it must be re-measured on whatever larger model config we settle on. Do not conclude "compile is bad" from the small-model data point.
Recommended sequencing
Do this in order. Skipping ahead is the failure mode that creates misleading "compile/TRT didn't help" data.
- Now: benchmark A (
scripts/bench_inference_backend.py) and confirm thelocalvsservermultipliers onhomeandremote. Validate the iter 1.2–1.3× / traversal 1.5–2× working estimate. - Next: experiment with a larger network config. Measure compute vs
learning-curve trade-off with the existing toolchain (no compile/TRT yet).
This step decides the model size that future optimizations target.
It is also the prerequisite for revisiting AMP,
torch.compile, and TensorRT: all three are dispatch-overhead-bound on the current small model. - Then: re-measure
torch.compileon the trainer at the chosen model size. The earlier regression was size-bound; expect a different result. - Then: integrate TensorRT into the inference server (covers traversal
and eval forward simultaneously). Bound the gain by the post-step-2
policy_network_secondsshare, not the headline TensorRT speedup. - In parallel with 2–4: if denser eval is operationally useful, raise
evaluation.gamesand lowerevaluation.eval_every. This step does not require code changes but sharply increases the value of step 4.
Out of scope until A bench numbers are in: Option C, nogil threading, async
inference client, compiled encoding.