Add a microbench that separates game/encoding overhead, single-request policy boundary overhead, and batched PyTorch forward lower bounds. Record CPU/CUDA results and link the finding from the Julia port evaluation.
Co-Authored-By: Codex <codex@openai.com>
진단 가설/개입/측정 섹션은 docs/research/lost_cities_selectivity.md
에 verbatim으로 보존되어 있어 ideas.md 쪽 사본은 stale 위험 + 중복
유지 비용만 남는다. ideas.md는 brainstorm 인덱스로 축소하고, research/
의 세 thread를 단일 진입점으로 정리.
기존의 "All-negative fallback 가설 — 검증됨, 부분 풀림" 같은 stale
표현도 함께 사라진다 (research doc에선 여전히 open hypothesis로
다뤄지고 있음).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Captures the dependency graph between pending levers (model size,
Option B, AMP, compile, TensorRT, Option A re-enable, Julia port) and
the rule that infrastructure optimization precedes the model-size
experiment because every future training run benefits from the
infrastructure speedup, not just the one keystone experiment.
ideas.md gets a third Active Research Threads pointer.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Decision-in-advance for criteria 3 (multi-thread scaling), 4 (Flux/CUDA
MLP forward), and 5 (real-game-state slice). Each criterion lists PASS
/ PARTIAL / FAIL bands with concrete numerical thresholds, plus a
decision rule that maps {3, 4, 5} outcomes to a single action: port,
hybrid (Julia traversal + PyTorch networks), or stay on Python/Cython
and pursue Option B instead.
Cost-of-being-wrong asymmetry stated explicitly: port is months,
staying is zero work, so the GO bar is deliberately above 50% and the
STAY bar is permissive.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
experiments/julia_cfr_toy/ ports a synthetic CFR external-sampling
traversal to both Julia and Cython for direct head-to-head
measurement of the actual project hot-path pattern (recursive
tree + mutable regret state + branch-heavy legal-action logic).
Headline result (2026-05-07, single run): Julia ~1.9× faster than
Cython on this pattern, 0 MB allocation, 0% GC time. Root regret
parity ε ≤ 1e-9. The GC-pause concern that was the main argument
against Julia adoption did not materialize. Cython's 21.3 MB
allocation suggests its implementation can be tightened, so the
honest gap window is roughly 1.3×–1.9×.
Multi-thread scaling (bench_cfr_threaded.jl): 2.44× wall-clock at 8T
but only 31% efficiency — inconclusive, likely a toy-size artifact
(per-thread workload too small to amortize dispatch). A heavier
per-thread workload sweep is the remaining decisive test.
docs/research/julia_port_evaluation.md captures this evidence
alongside the earlier safe-heuristic single-thread parity result
and lists the remaining decision criteria (multi-thread scaling with
heavier workload, Flux.jl+CUDA.jl coverage, real-game-state slice).
Do not commit to porting until multi-thread scaling is conclusively
settled.
ideas.md gets a second Active Research Threads pointer.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Mechanism-level analysis of the slot_aware_playability iter 240
plateau (opened_colors 4.95+, bad_open_rate 88-92%, calibration
gap 6-9 → 2-4). Captures four expert consultations with diagnostic
hypotheses, intervention catalog (architectural / training-dynamics
/ game-specific), measurement plan, and a comparison table across
the four sources.
Key new directions surfaced:
- Current vs average vs league policy separation (Deep CFR average
strategy is the convergence target, not advantage current).
- All-negative fallback as Deep CFR ablation lever.
- Empirical r̃ partitioning by action class.
- Tabular Lost Cities oracle as a clean test of "is 5-color the
game-theoretic answer or an approximation artifact".
- Entry-gate target defined from traversal counterfactual values
instead of heuristic labels.
ideas.md gets an "Active Research Threads" pointer.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
OpenSpiel's Deep CFR records strategy samples at opponent nodes during
the traverser's tree walk; storing on traverser nodes under external
sampling drops the ρ_p reach factor and biases the average-policy
estimate. Reject that config combination at load time and add a research
note deriving why outcome sampling is unaffected while external is not.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Five notes covering outcome-sampling target correctness, package
architecture, v0 feature-parity vs legacy, opponent-policy network
divergence, and regret-matching fallback audit. Four are derived from
archive sources (cited via Source: lines); outcome-sampling-target is
a fresh write-up and serves as the style template.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Codex and other coding subagents have been creating branches like
experiments/foo and feature/bar without being asked. This fragments
review, hides work from the user, and requires manual cleanup. The
project intentionally develops on main with frequent small commits.
AGENTS.md adds an explicit "Git Branching Policy" section: no
checkout -b, switch -c, branch <name>, or PR-from-new-branch unless
the user asks for it in the current task. Includes a pass-through
clause so this propagates to subagents the main agent spawns.
CLAUDE.md adds a one-line pointer with the same pass-through note,
since CLAUDE.md mandates AGENTS.md is read at session start.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds torch.autocast(fp16) + GradScaler around _train_advantage and
_train_strategy when run.use_amp=true and device=cuda. CPU/non-CUDA
falls back to fp32 no-op. Mitigations:
- scaler.unscale_(optimizer) before grad_clip.
- nonfinite-loss guard skips overflowing batches and counts them.
- diff.float().square() in advantage loss to avoid fp16 overflow.
- strategy mask/log_softmax kept in fp32.
New metrics: amp/grad_scale, amp/nonfinite_loss_count.
Tests: AMP CUDA smoke + CPU fallback in test_deep_cfr_trainer.py.
Bench: scripts/bench_amp_trainer.py micro-benches train phases under
synthetic replay memory. smoke.yaml result is fp32 3.22ms / AMP 3.92ms
(0.82×, regression). 100-iter A/B on default.yaml deliberately
skipped: smoke regression mirrors the 2026-05-07 torch.compile
regression dynamic (dispatch overhead > kernel benefit at this model
size) and re-confirming on the same size adds no information.
Default stays run.use_amp: false. Re-enable trigger documented in
docs/performance.md: hidden_size >= 1024 or num_layers >= 6, then run
the bench script + 100-iter A/B before flipping default.
The experiment finds a model size where (a) learning-curve gains
justify compute, and (b) forward time is large enough to amortize
AMP/compile/TensorRT overhead. Outcome gates re-enabling those four
deferred optimizations and Option A.
Tested grid: hidden={512,768,1024,1536} × layers={3,4,6,8} subset.
200 iterations per config on home (RTX 3090), single seed initially,
second seed for boundary configs. Eval cadence held at default.
Decision tree included for: success → recommend new default and
trigger downstream plans; null → document and stay; expensive-but-
better → defer until AMP/compile/TRT land.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
performance.md additions:
- Batched Traversal Inference design decision (A vs B vs C with
rationale).
- Option A bench result and structural ceiling (realized batch ~7.2,
IPC overhead exceeds GPU gain at small model size).
- Post-A optimization calculus: why compile/TensorRT remain
iter-neutral today and become meaningful only after model growth
and/or denser eval. Sequencing matters; do not retest these on the
current small model.
- Free-threaded Python (3.13t/3.14t) note: cleanest endpoint in
principle, but PyTorch maturity + Cython nogil audit cost block
near-term adoption.
docs/plans/ (4 plan documents for Codex execution):
- batched_traversal_inference_server.md (executed; deferred).
- amp_trainer.md.
- torch_compile.md.
- cython_safe_heuristic_bots.md (executed; first-pass landed).
docs/reports/ (3 cost reports):
- cost_pytorch_free_threaded_2026-05-07.md: WAIT 3-6 months;
PyTorch wheels exist but our Cython is the gating cost.
- cost_cython_nogil_audit_2026-05-07.md: medium effort, traversal.pyx
carries 90% of blockers; Steps 1-3 (cfr_math/encoding nogil
keywords, TraversalStats cdef class) are safe and cheap, Steps
4-6 wait for triggers.
- cost_pytorch_cuda_multithread_2026-05-07.md: risky;
optimizer.step / load_state_dict race silently with concurrent
forward; per-thread default streams unset means naive threading
serializes on default stream anyway.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Implements the central inference server pattern: a dedicated GPU
process owns advantage/strategy/league networks, batches policy
requests across traversal workers via shared-memory tensor pool, and
returns logits. Workers route forward calls through InferenceClient /
NetworkProxy when traversal.inference_backend == "server".
Default remains traversal.inference_backend: local. The server
backend regresses iter time ~3.8× on the inspected default config
(small-model dispatch + sync-blocking traversal capping realized
batch at ~num_workers=8 instead of the bs=64-256 needed to amortize
IPC overhead). Keeping the implementation behind the flag lets us
re-enable when (a) model size grows, (b) per-worker interleaved
traversal lands, or (c) eval becomes dominant — see
docs/performance.md "Option A Bench Result and Structural Ceiling"
for the full diagnosis.
Plumbing included:
- inference_buffers.py: shared-memory tensor pool with slot
management.
- inference_client.py: per-worker client + NetworkProxy adapter for
the existing traversal.pyx call sites.
- inference_server.py: spawn-context server process with
batch-window aggregation, weight sync, shutdown sentinel.
- bench_inference_backend.py: A/B between local and server backends
with eval/checkpoint disabled.
- test_inference_server.py: round-trip and integration tests.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Cython implementation in heuristic_cy.pyx achieves ~2.55× speedup on
opponent_act_seconds (200-game eval: 59.20s → 23.24s). Original Python
implementation preserved verbatim in heuristic_py.py as the equivalence
reference. Action-sequence equivalence is verified by
test_safe_heuristic_equivalence.py against seeded game corpora.
Key implementation notes:
- File-local wraparound=True override required for negative discard
indexing; Cython global wraparound=False would segfault.
- annotation_typing=False preserves verbatim Python semantics.
- _CachedState materializes hands/expeditions/discards/deck once per
act() call — this is the dominant performance win.
Further C-array optimization of _card_value_for_me /
_card_value_for_opponent / _color_commitment / _bonus_potential is
deferred. The current 2.55× delivers most of the dense-eval future
benefit; further work is gated on actually adopting denser eval
schedules (eval_every=5, games=1000).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Measured DeepCFRMLP forward at bs={1,4,16,64,256,1024} on RTX 3090.
Per-state cost drops 232× from bs=1 (80 µs) to bs=256 (0.34 µs) while
per-call latency stays near 90 µs through bs=256. Policy-call supply
from a real run is ~368 states per traversal and ~200k per iteration,
well above the bs=64–256 plateau, so batched inference is not
supply-limited. GPU forward is not the limiter once batching exists.
Verdict: Optimization Priorities #5 (batched traversal inference) is
worth pursuing. End-to-end gain will still be bounded by encoding and
worker-GPU coordination overhead.
- scripts/profile_gpu_forward.py: standalone profiling script
- docs/performance.md: new "GPU forward profiling for batched traversal"
experiment section with table, supply estimate, and verdict
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Make the link from the failed torch.compile experiment to Optimization
Priorities #5 explicit, so a future revisit happens at the right time
(once compile is on the dominant phase, not just trainer optimization
steps).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wrapping trainer networks with torch.compile produced a 4.8% regression
in iteration time on default.yaml (17.93s → 18.79s). Two causes: (1)
the dominant phase is CPU traversal which bypasses the compiled
wrapper, (2) DeepCFRMLP is too small for compile dispatch overhead to
pay back. Implementation kept on experiments/torch-compile for future
revisits when the trainer model or inference path changes.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adopt a 5-namespace scheme so wandb groups related metrics in the
sidebar and capture-group regex (eval/(random|safe_heuristic)/win_rate0
vs eval/(?:random|safe_heuristic)/win_rate0) controls panel splitting:
- loss/{advantage,strategy}
- samples/{advantage,strategy,advantage_player_N}
- memory/{advantage,strategy,advantage_player_N}
- time/{iteration_seconds,traversal_seconds,advantage_train_seconds,
strategy_train_seconds,evaluation_seconds,memory_add_seconds,
checkpoint_seconds,batch_tensor_seconds,nodes_per_second,
advantage_player_N_sample_seconds,strategy_sample_seconds}
- traversal/{nodes,terminals,depth_cutoffs,node_limit_cutoffs,
max_depth_reached,endpoints,avg_endpoint_depth,
endpoint_depth_bucket_*,regret_fallback_*,sampled_actions}
- eval/<opponent>/<metric> (3-level so opponent can be the capture group)
`iteration` keeps no namespace (it's the wandb step axis). Internal
TraversalStats.to_dict() and benchmark.py's standalone result dict
keep their flat names — only the trainer's emitted metrics are
remapped, with the traversal_*→traversal/* translation done at
insertion into runtime_metrics.
analyze.py updated to read the new keys (PlotSpec metrics, color map,
opponent_names parser, _first_existing_eval lookup). Tests updated for
the new eval_metrics dict keys.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Previously the timer wrapped every call to _evaluate(), but on iterations
that skip eval (iteration % eval_every != 0) the function returns
immediately and the recorded value was just function-call overhead
(~3 µs), which made W&B show a wildly bimodal "evaluation_seconds"
metric. Now only set the key when eval_metrics is non-empty so
non-eval iterations have no data point.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Append a rough iters_per_hour estimate (3600/iteration_seconds) to the
console summary so users running long jobs can eyeball ETA without doing
the math.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Make traversal progress and iteration-complete summary lines easier to
visually scan during long runs by leading with [i=N]. Drop the redundant
"iteration=N" kv from the body to keep lines short.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The avg-strategy 1000iter file (with the recent +traversals/+LR/+LCFR
changes) is the canonical "best-known" config. Renamed it to
default.yaml so users start from a single, obvious entry point and
override one field per ablation via --set. Other 12 configs moved to
configs/archive/ — kept for historical reproduction, not for active use.
- configs/deep_cfr/{default.yaml, smoke.yaml} are the only active configs
- experiment_name shortened to "deep-cfr-default" (was a long mouthful)
- AGENTS.md examples and Project Layout section rewritten around
default.yaml; ablation example shows the override-one-field pattern
- Tests pointed at the archived slot-playability config for the legacy
reproduction assertions
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
If eval_every is positive but max_iterations falls before the next scheduled
eval iteration (including resume cases where current_iteration is already
past the last eval boundary), log a one-time warning at run start so the
user notices the misconfiguration. We deliberately do not force an
end-of-run eval, which would distort time budgets and reproducibility.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Default is one baseline + one treatment, sequential, same seed, with a
shared --wandb-tag hypothesis label for W&B Compare Runs filtering.
Multi-seed only on explicit request; never run two trainings on the same
GPU.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Brief guidance on roles (notes = purpose, tags = filter categories)
plus three anti-patterns to avoid. Otherwise free-form.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Exposes wandb.init's notes field through the CLI so each run can carry a
short description of its purpose, visible on the W&B run page alongside
tags and config.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- AGENTS.md: add Weights & Biases section covering install (extra),
online/offline modes, per-run wandb/ layout, sync command, and the
source-of-truth note (metrics.jsonl, not W&B).
- CLAUDE.md: replace soft "before making changes" wording with a
mandatory session-start instruction to read AGENTS.md in full.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Append eval_seconds and its share of iteration_seconds to the per-iteration
console summary when evaluation actually ran, so users watching the log can
see how much wall time eval is consuming without parsing JSON metrics.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Drop checkpoint.directory from config — config defines what an experiment
is, not where its outputs go. The CLI now computes the run directory from
run.experiment_name plus a timestamp, defaulting to runs/tmp/ for
throwaway runs and runs/ when --keep is passed.
- Remove CheckpointConfig.directory and DeepCFRConfig.checkpoint_path
- DeepCFRTrainer takes run_dir: Path explicitly
- CLI: add --keep boolean; --resume requires an explicit path (no shortcut)
- Auto path: runs/[tmp/]<YYYY-MM-DD_HHMMSS>_<experiment_name-kebab>/
- Rename 13 configs to kebab-case; strip directory: lines; kebab their
experiment_name values
- Rewrite AGENTS.md training/run sections; document
archive/tmp/<flat> layout, --keep, kebab-case scope
- Update tests for new run_dir flow and dropped --resume shortcut
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Remove legacy aliases, rename max_hours to max_minutes, collapse the
four checkpoint save flags into save_every + save_latest, and change
defaults to safer values (opponent_policy=self_play_league,
device=auto, eval_every=50, max_depth=null). Migrate all archived
yaml configs and tests to the new schema.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Mirror Deep CFR training metrics to W&B via a new WandbRunTracker
wired through CompositeRunTracker; wandb is an optional extra so
default installs and runs stay unchanged. Train CLI gains
--wandb/--wandb-project/--wandb-mode/--wandb-name/--wandb-tag, and
train() now closes the tracker in a finally block so runs finalize
even on early exit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Strategy network (학습 중인 average policy)를 traversal opponent로 사용하는
새 옵션 추가. Deep CFR 이론적 수렴이 average strategy에 대한 보장이라는
점에 착안 — opponent_policy=network의 발산 문제를 완화할 수 있는지 실증.
구현:
- config: opponent_policy validator에 average_strategy 추가
- traversal.pyx: opponent_policy_id=3, strategy_network 인자, softmax 기반
policy 도출 (_policy_from_strategy_network)
- workers.py: TraversalWorkerBatch에 strategy_network state_dict 추가
- trainer.py: 직렬/병렬 traversal call에 strategy_network 전달
- 1000-iter 실험 config 추가 (opponent_policy=network와 동일 hyperparam)
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
같은 hyperparameter (512x3, 2x updates, argmax_tiebreak)에 opponent_policy
만 self_play_league로 다른 run과의 iteration별 비교 표 추가. iter 350
시점에서 Random WR 38%p 격차 (network 34% vs league 72%) 확인 — 발산
가설을 실증.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
opponent_policy=network 설정으로 1000-iter 실험 두 개 (512x3, 1024x4)를
돌린 결과 두 실험 모두 policy collapse가 발생함을 확인. 큰 capacity는
plateau를 늘리지만 발산 자체를 막지 못함. 향후 학습은 self_play_league
기본값을 유지할 것을 권고.
- 실험에 사용한 config 두 개 추가
- 발견 분석 문서 추가 (원인, 비교, 권고)
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>