Refactor docs/performance.md per librarian routing rule

docs/performance.md grew to 914 lines because dated experiments and
design analyses kept getting appended instead of routed to
docs/archive/ and docs/research/ as AGENTS.md prescribes. The
oversize check from librarian Stage 1 surfaced the file; this commit
acts on that finding by extracting the parts that belong elsewhere
and trimming the source to a focused current-state reference.

Extracts (verbatim from the original prose, with cross-link headers
and a brief routing note added at top):

- docs/archive/deep-cfr-performance-experiments-2026-05-07.md
  bundles torch.compile (regression), AMP (regression), GPU forward
  profiling (decision support), and Option B interleaved traversal
  (pass) — same date, same theme.
- docs/research/batched-traversal-inference-decision.md captures the
  durable A vs B vs C rationale with a closing "Outcome" pointer to
  the post-bench archive doc.
- docs/archive/post-a-optimization-calculus-2026-05-07.md preserves
  the forward-looking sequencing recorded pre-bench.
- docs/archive/option-a-bench-result-2026-05-07.md preserves the
  regression diagnosis and re-enable criteria.

docs/performance.md is now 345 lines, holds sections 1–9 (current
runtime / bottleneck / device / AMP status / batching / eval /
TensorRT / priorities), and ends with a "See Also" linking the four
extracts.

Also reworded the AGENTS.md soft-cap rule from a bare "~500-line
soft cap" to clarify the intent: the cap is a *routing trigger* (is
content piling up that should live in archive/research?), not a
split mandate. Reduces the risk of future agents shredding a useful
doc just to satisfy a number.

scripts/librarian.sh now exits 0 against the working tree.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-05-07 23:57:39 +09:00
co-authored by Claude Opus 4.7
parent 6ecb233bdd
commit 1cd9950bd3
7 changed files with 703 additions and 595 deletions
@@ -0,0 +1,98 @@
# Post-A Optimization Calculus (2026-05-07)
**Source:** Extracted from `docs/performance.md` § "Post-A Optimization
Calculus (forward-looking, 2026-05-07)" on 2026-05-08.
Forward-looking sequencing recorded *before* Option A was benched. Has
not been measured. See
`docs/archive/option-a-bench-result-2026-05-07.md` for the actual bench
which deferred Option A — some assumptions below ("once Option A
lands") have to be re-evaluated in light of that result.
---
Once Option A lands, the bottleneck shape changes. This section records the
expected sequencing for follow-up work. It is forward-looking and has not been
measured yet — verify against bench numbers after A is benchmarked.
## Why compile / TensorRT are negligible *today* but become meaningful later
Today (small model: 3-layer, 512 hidden):
- `torch.compile` on the trainer's networks already regressed (see the
2026-05-07 experiment in
`docs/archive/deep-cfr-performance-experiments-2026-05-07.md`).
The model is too small for kernel fusion to beat compile dispatch
overhead.
- `torch.compile` / TensorRT on the inference-server forward (post-A) would
shave ~3050% off ~90μs/call → ~5070μs/call. With forward share of an iter
reduced to <1% by A's batching, the iter-level multiplier is ~1.001.01×.
Negligible.
Two compounding shifts can flip this:
1. **Larger model.** Going from 512 hidden / 3 layers to ~1024 hidden /
~6 layers pushes the forward call out of dispatch-bound territory into
kernel-bound territory. Compile fusion and TensorRT both deliver real
1.52× on the forward call itself once the kernel is large enough to
amortize launch overhead. Forward share of iter time also rebalances upward
because per-call time scales with FLOPs while batching gain is fixed.
2. **Denser, larger evaluation.** Moving toward `eval_every: 5` and
`evaluation.games: 1000` makes evaluation about half of iteration wall-clock
(see the amortized eval table in `docs/performance.md`). Eval is pure
inference, so TensorRT on the inference-server's forward path applies
directly.
When both shifts happen together, an illustrative future iter (rough order of
magnitude only):
| Configuration | Iter time (rough) |
| --- | ---: |
| Today (small model, eval_every=25) | 17.85s |
| + A (batched traversal inference) | ~14s |
| + larger model (≈4× FLOPs), no compile/TRT | ~50s |
| + dense eval (eval_every=5, games=1000) | ~70s |
| + compile (trainer) + TensorRT (inference) | ~45s |
That last row is where compile/TensorRT contributes ~1.5× iter — the same
tooling that is iter-neutral today. The numbers above are illustrative; real
ratios depend on model size, kernel autotune outcomes, and the eval-vs-train
balance.
## Tooling split
- **TensorRT**: applies only to inference (no backward). Targets:
- inference-server forward in traversal,
- inference-server forward in evaluation.
Both are served by the same A-era server, so a single TensorRT integration
covers both.
- **`torch.compile`**: applies to trainer's advantage/strategy training
(forward+backward+optimizer). The 2026-05-07 regression on a small model
does **not** generalize — it must be re-measured on whatever larger model
config we settle on. Do not conclude "compile is bad" from the small-model
data point.
## Recommended sequencing
Do this in order. Skipping ahead is the failure mode that creates misleading
"compile/TRT didn't help" data.
1. **Now**: benchmark A (`scripts/bench_inference_backend.py`) and confirm the
`local` vs `server` multipliers on `home` and `remote`. Validate the iter
1.21.3× / traversal 1.52× working estimate.
2. **Next**: experiment with a larger network config. Measure compute vs
learning-curve trade-off with the existing toolchain (no compile/TRT yet).
This step decides the model size that future optimizations target.
It is also the prerequisite for revisiting AMP, `torch.compile`, and
TensorRT: all three are dispatch-overhead-bound on the current small model.
3. **Then**: re-measure `torch.compile` on the trainer at the chosen model
size. The earlier regression was size-bound; expect a different result.
4. **Then**: integrate TensorRT into the inference server (covers traversal
and eval forward simultaneously). Bound the gain by the post-step-2
`policy_network_seconds` share, not the headline TensorRT speedup.
5. **In parallel with 24**: if denser eval is operationally useful, raise
`evaluation.games` and lower `evaluation.eval_every`. This step does not
require code changes but sharply increases the value of step 4.
Out of scope until A bench numbers are in: Option C, `nogil` threading, async
inference client, compiled encoding.