The experiment finds a model size where (a) learning-curve gains
justify compute, and (b) forward time is large enough to amortize
AMP/compile/TensorRT overhead. Outcome gates re-enabling those four
deferred optimizations and Option A.
Tested grid: hidden={512,768,1024,1536} × layers={3,4,6,8} subset.
200 iterations per config on home (RTX 3090), single seed initially,
second seed for boundary configs. Eval cadence held at default.
Decision tree included for: success → recommend new default and
trigger downstream plans; null → document and stay; expensive-but-
better → defer until AMP/compile/TRT land.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
performance.md additions:
- Batched Traversal Inference design decision (A vs B vs C with
rationale).
- Option A bench result and structural ceiling (realized batch ~7.2,
IPC overhead exceeds GPU gain at small model size).
- Post-A optimization calculus: why compile/TensorRT remain
iter-neutral today and become meaningful only after model growth
and/or denser eval. Sequencing matters; do not retest these on the
current small model.
- Free-threaded Python (3.13t/3.14t) note: cleanest endpoint in
principle, but PyTorch maturity + Cython nogil audit cost block
near-term adoption.
docs/plans/ (4 plan documents for Codex execution):
- batched_traversal_inference_server.md (executed; deferred).
- amp_trainer.md.
- torch_compile.md.
- cython_safe_heuristic_bots.md (executed; first-pass landed).
docs/reports/ (3 cost reports):
- cost_pytorch_free_threaded_2026-05-07.md: WAIT 3-6 months;
PyTorch wheels exist but our Cython is the gating cost.
- cost_cython_nogil_audit_2026-05-07.md: medium effort, traversal.pyx
carries 90% of blockers; Steps 1-3 (cfr_math/encoding nogil
keywords, TraversalStats cdef class) are safe and cheap, Steps
4-6 wait for triggers.
- cost_pytorch_cuda_multithread_2026-05-07.md: risky;
optimizer.step / load_state_dict race silently with concurrent
forward; per-thread default streams unset means naive threading
serializes on default stream anyway.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>