Files
coorl-lost-cities/docs/archive/post-a-optimization-calculus-2026-05-07.md
T
coolguyandClaude Opus 4.7 1cd9950bd3 Refactor docs/performance.md per librarian routing rule
docs/performance.md grew to 914 lines because dated experiments and
design analyses kept getting appended instead of routed to
docs/archive/ and docs/research/ as AGENTS.md prescribes. The
oversize check from librarian Stage 1 surfaced the file; this commit
acts on that finding by extracting the parts that belong elsewhere
and trimming the source to a focused current-state reference.

Extracts (verbatim from the original prose, with cross-link headers
and a brief routing note added at top):

- docs/archive/deep-cfr-performance-experiments-2026-05-07.md
  bundles torch.compile (regression), AMP (regression), GPU forward
  profiling (decision support), and Option B interleaved traversal
  (pass) — same date, same theme.
- docs/research/batched-traversal-inference-decision.md captures the
  durable A vs B vs C rationale with a closing "Outcome" pointer to
  the post-bench archive doc.
- docs/archive/post-a-optimization-calculus-2026-05-07.md preserves
  the forward-looking sequencing recorded pre-bench.
- docs/archive/option-a-bench-result-2026-05-07.md preserves the
  regression diagnosis and re-enable criteria.

docs/performance.md is now 345 lines, holds sections 1–9 (current
runtime / bottleneck / device / AMP status / batching / eval /
TensorRT / priorities), and ends with a "See Also" linking the four
extracts.

Also reworded the AGENTS.md soft-cap rule from a bare "~500-line
soft cap" to clarify the intent: the cap is a *routing trigger* (is
content piling up that should live in archive/research?), not a
split mandate. Reduces the risk of future agents shredding a useful
doc just to satisfy a number.

scripts/librarian.sh now exits 0 against the working tree.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 23:57:39 +09:00

4.8 KiB
Raw Blame History

Post-A Optimization Calculus (2026-05-07)

Source: Extracted from docs/performance.md § "Post-A Optimization Calculus (forward-looking, 2026-05-07)" on 2026-05-08.

Forward-looking sequencing recorded before Option A was benched. Has not been measured. See docs/archive/option-a-bench-result-2026-05-07.md for the actual bench which deferred Option A — some assumptions below ("once Option A lands") have to be re-evaluated in light of that result.


Once Option A lands, the bottleneck shape changes. This section records the expected sequencing for follow-up work. It is forward-looking and has not been measured yet — verify against bench numbers after A is benchmarked.

Why compile / TensorRT are negligible today but become meaningful later

Today (small model: 3-layer, 512 hidden):

  • torch.compile on the trainer's networks already regressed (see the 2026-05-07 experiment in docs/archive/deep-cfr-performance-experiments-2026-05-07.md). The model is too small for kernel fusion to beat compile dispatch overhead.
  • torch.compile / TensorRT on the inference-server forward (post-A) would shave ~3050% off ~90μs/call → ~5070μs/call. With forward share of an iter reduced to <1% by A's batching, the iter-level multiplier is ~1.001.01×. Negligible.

Two compounding shifts can flip this:

  1. Larger model. Going from 512 hidden / 3 layers to ~1024 hidden / ~6 layers pushes the forward call out of dispatch-bound territory into kernel-bound territory. Compile fusion and TensorRT both deliver real 1.52× on the forward call itself once the kernel is large enough to amortize launch overhead. Forward share of iter time also rebalances upward because per-call time scales with FLOPs while batching gain is fixed.
  2. Denser, larger evaluation. Moving toward eval_every: 5 and evaluation.games: 1000 makes evaluation about half of iteration wall-clock (see the amortized eval table in docs/performance.md). Eval is pure inference, so TensorRT on the inference-server's forward path applies directly.

When both shifts happen together, an illustrative future iter (rough order of magnitude only):

Configuration Iter time (rough)
Today (small model, eval_every=25) 17.85s
+ A (batched traversal inference) ~14s
+ larger model (≈4× FLOPs), no compile/TRT ~50s
+ dense eval (eval_every=5, games=1000) ~70s
+ compile (trainer) + TensorRT (inference) ~45s

That last row is where compile/TensorRT contributes ~1.5× iter — the same tooling that is iter-neutral today. The numbers above are illustrative; real ratios depend on model size, kernel autotune outcomes, and the eval-vs-train balance.

Tooling split

  • TensorRT: applies only to inference (no backward). Targets:
    • inference-server forward in traversal,
    • inference-server forward in evaluation. Both are served by the same A-era server, so a single TensorRT integration covers both.
  • torch.compile: applies to trainer's advantage/strategy training (forward+backward+optimizer). The 2026-05-07 regression on a small model does not generalize — it must be re-measured on whatever larger model config we settle on. Do not conclude "compile is bad" from the small-model data point.

Do this in order. Skipping ahead is the failure mode that creates misleading "compile/TRT didn't help" data.

  1. Now: benchmark A (scripts/bench_inference_backend.py) and confirm the local vs server multipliers on home and remote. Validate the iter 1.21.3× / traversal 1.52× working estimate.
  2. Next: experiment with a larger network config. Measure compute vs learning-curve trade-off with the existing toolchain (no compile/TRT yet). This step decides the model size that future optimizations target. It is also the prerequisite for revisiting AMP, torch.compile, and TensorRT: all three are dispatch-overhead-bound on the current small model.
  3. Then: re-measure torch.compile on the trainer at the chosen model size. The earlier regression was size-bound; expect a different result.
  4. Then: integrate TensorRT into the inference server (covers traversal and eval forward simultaneously). Bound the gain by the post-step-2 policy_network_seconds share, not the headline TensorRT speedup.
  5. In parallel with 24: if denser eval is operationally useful, raise evaluation.games and lower evaluation.eval_every. This step does not require code changes but sharply increases the value of step 4.

Out of scope until A bench numbers are in: Option C, nogil threading, async inference client, compiled encoding.