Files
coorl-lost-cities/docs/archive/post-a-optimization-calculus-2026-05-07.md
coolguyandClaude Opus 4.7 1cd9950bd3 Refactor docs/performance.md per librarian routing rule
docs/performance.md grew to 914 lines because dated experiments and
design analyses kept getting appended instead of routed to
docs/archive/ and docs/research/ as AGENTS.md prescribes. The
oversize check from librarian Stage 1 surfaced the file; this commit
acts on that finding by extracting the parts that belong elsewhere
and trimming the source to a focused current-state reference.

Extracts (verbatim from the original prose, with cross-link headers
and a brief routing note added at top):

- docs/archive/deep-cfr-performance-experiments-2026-05-07.md
  bundles torch.compile (regression), AMP (regression), GPU forward
  profiling (decision support), and Option B interleaved traversal
  (pass) — same date, same theme.
- docs/research/batched-traversal-inference-decision.md captures the
  durable A vs B vs C rationale with a closing "Outcome" pointer to
  the post-bench archive doc.
- docs/archive/post-a-optimization-calculus-2026-05-07.md preserves
  the forward-looking sequencing recorded pre-bench.
- docs/archive/option-a-bench-result-2026-05-07.md preserves the
  regression diagnosis and re-enable criteria.

docs/performance.md is now 345 lines, holds sections 1–9 (current
runtime / bottleneck / device / AMP status / batching / eval /
TensorRT / priorities), and ends with a "See Also" linking the four
extracts.

Also reworded the AGENTS.md soft-cap rule from a bare "~500-line
soft cap" to clarify the intent: the cap is a *routing trigger* (is
content piling up that should live in archive/research?), not a
split mandate. Reduces the risk of future agents shredding a useful
doc just to satisfy a number.

scripts/librarian.sh now exits 0 against the working tree.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 23:57:39 +09:00

99 lines
4.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Post-A Optimization Calculus (2026-05-07)
**Source:** Extracted from `docs/performance.md` § "Post-A Optimization
Calculus (forward-looking, 2026-05-07)" on 2026-05-08.
Forward-looking sequencing recorded *before* Option A was benched. Has
not been measured. See
`docs/archive/option-a-bench-result-2026-05-07.md` for the actual bench
which deferred Option A — some assumptions below ("once Option A
lands") have to be re-evaluated in light of that result.
---
Once Option A lands, the bottleneck shape changes. This section records the
expected sequencing for follow-up work. It is forward-looking and has not been
measured yet — verify against bench numbers after A is benchmarked.
## Why compile / TensorRT are negligible *today* but become meaningful later
Today (small model: 3-layer, 512 hidden):
- `torch.compile` on the trainer's networks already regressed (see the
2026-05-07 experiment in
`docs/archive/deep-cfr-performance-experiments-2026-05-07.md`).
The model is too small for kernel fusion to beat compile dispatch
overhead.
- `torch.compile` / TensorRT on the inference-server forward (post-A) would
shave ~3050% off ~90μs/call → ~5070μs/call. With forward share of an iter
reduced to <1% by A's batching, the iter-level multiplier is ~1.001.01×.
Negligible.
Two compounding shifts can flip this:
1. **Larger model.** Going from 512 hidden / 3 layers to ~1024 hidden /
~6 layers pushes the forward call out of dispatch-bound territory into
kernel-bound territory. Compile fusion and TensorRT both deliver real
1.52× on the forward call itself once the kernel is large enough to
amortize launch overhead. Forward share of iter time also rebalances upward
because per-call time scales with FLOPs while batching gain is fixed.
2. **Denser, larger evaluation.** Moving toward `eval_every: 5` and
`evaluation.games: 1000` makes evaluation about half of iteration wall-clock
(see the amortized eval table in `docs/performance.md`). Eval is pure
inference, so TensorRT on the inference-server's forward path applies
directly.
When both shifts happen together, an illustrative future iter (rough order of
magnitude only):
| Configuration | Iter time (rough) |
| --- | ---: |
| Today (small model, eval_every=25) | 17.85s |
| + A (batched traversal inference) | ~14s |
| + larger model (≈4× FLOPs), no compile/TRT | ~50s |
| + dense eval (eval_every=5, games=1000) | ~70s |
| + compile (trainer) + TensorRT (inference) | ~45s |
That last row is where compile/TensorRT contributes ~1.5× iter — the same
tooling that is iter-neutral today. The numbers above are illustrative; real
ratios depend on model size, kernel autotune outcomes, and the eval-vs-train
balance.
## Tooling split
- **TensorRT**: applies only to inference (no backward). Targets:
- inference-server forward in traversal,
- inference-server forward in evaluation.
Both are served by the same A-era server, so a single TensorRT integration
covers both.
- **`torch.compile`**: applies to trainer's advantage/strategy training
(forward+backward+optimizer). The 2026-05-07 regression on a small model
does **not** generalize — it must be re-measured on whatever larger model
config we settle on. Do not conclude "compile is bad" from the small-model
data point.
## Recommended sequencing
Do this in order. Skipping ahead is the failure mode that creates misleading
"compile/TRT didn't help" data.
1. **Now**: benchmark A (`scripts/bench_inference_backend.py`) and confirm the
`local` vs `server` multipliers on `home` and `remote`. Validate the iter
1.21.3× / traversal 1.52× working estimate.
2. **Next**: experiment with a larger network config. Measure compute vs
learning-curve trade-off with the existing toolchain (no compile/TRT yet).
This step decides the model size that future optimizations target.
It is also the prerequisite for revisiting AMP, `torch.compile`, and
TensorRT: all three are dispatch-overhead-bound on the current small model.
3. **Then**: re-measure `torch.compile` on the trainer at the chosen model
size. The earlier regression was size-bound; expect a different result.
4. **Then**: integrate TensorRT into the inference server (covers traversal
and eval forward simultaneously). Bound the gain by the post-step-2
`policy_network_seconds` share, not the headline TensorRT speedup.
5. **In parallel with 24**: if denser eval is operationally useful, raise
`evaluation.games` and lower `evaluation.eval_every`. This step does not
require code changes but sharply increases the value of step 4.
Out of scope until A bench numbers are in: Option C, `nogil` threading, async
inference client, compiled encoding.