Files
coorl-lost-cities/docs/archive/deep-cfr-legacy-experiment-reproduction.md
T
2026-05-07 15:39:24 +09:00

214 lines
6.9 KiB
Markdown

# Deep CFR Legacy Experiment Reproduction Plan
This document tracks what is needed to reproduce the legacy `../coolrl`
experiment below with the same intended semantics in this repository:
`experiments/lost_cities/deep_cfr_pure_self_play_zero_pit_poc_full_depth_slot_aware_playability`
The goal is not to run a similar Deep CFR configuration. The goal is to map the
legacy experiment's hyperparameters and feature semantics into this repository's
own config schema so that the training run means the same thing.
Exact legacy YAML compatibility is not required. It is acceptable to create a
new config file under this repository as long as every relevant legacy
hyperparameter is represented explicitly and the differences are documented.
Mapped config in this repository:
`configs/deep_cfr/deep_cfr_selfplay_full_depth_slot_playability.yaml`
## Source Experiment
Legacy config path:
`/home/coolguy/dev/coolrl/experiments/lost_cities/deep_cfr_pure_self_play_zero_pit_poc_full_depth_slot_aware_playability/config.yaml`
Important legacy settings:
1. `seed: 79`
2. `max_hours: 4`
3. `max_iterations: null`
4. `device: CUDA`
5. `use_amp: false`
6. `rules.tier: tier3`
7. `encoding.derived_playability: true`
8. `encoding.slot_aware_playability: true`
9. `network.hidden_size: 256`
10. `network.num_layers: 3`
11. `network.activation: relu`
12. `traversal.traversals_per_player: 70`
13. `traversal.max_depth: null`
14. `traversal.max_nodes_per_traversal: 1000`
15. `traversal.opponent_policy: self_play_league`
16. `traversal.outcome_sampling_epsilon: 0.2`
17. `traversal.outcome_sampling_value_clip: 500`
18. `traversal.outcome_unsampled_regret: zero`
19. `optimization.advantage_batch_size: 1024`
20. `optimization.strategy_batch_size: 1024`
21. `optimization.advantage_updates_per_iteration: 256`
22. `optimization.strategy_updates_per_iteration: 256`
23. `optimization.learning_rate: 3.0e-5`
24. `optimization.weight_decay: 1.0e-4`
25. `optimization.grad_clip: 1.0`
## Required Reproduction Work
### Config
Represent the legacy hyperparameters in this repository's config schema:
1. Top-level:
- `experiment_name`
- `seed`
- `max_iterations`
- `max_hours`
- `device`
- `use_amp`
2. `rules`:
- `tier`
- optional direct `LostCitiesConfig` field overrides
3. `encoding`:
- `derived_playability`
- `slot_aware_playability`
4. `network`:
- `hidden_size`
- `num_layers`
- `activation`
5. `traversal` semantics:
- per-player traversal count
- max nodes per traversal
- traversal worker chunk size
- self-play league settings
6. `optimization` semantics:
- separate advantage and strategy batch sizes
- separate advantage and strategy update counts
- `weight_decay`
- `grad_clip`
7. `evaluation.on_max_steps`
8. `checkpoint`:
- `save_iteration_interval`
- `save_latest_only`
- `progress_interval_seconds`
### Encoding
Port the legacy feature semantics:
1. Base information-state encoding must remain deterministic.
2. Add `derived_playability` color-level features.
3. Add `slot_aware_playability` hand-slot action-local features.
4. Preserve feature order and dimensions from legacy where possible.
5. Add tests for input dimension and known-state feature values.
The slot-aware block is the core of the source experiment. Without this block,
the reproduction is not meaningful.
### Network
Make the MLP configurable:
1. `hidden_size`
2. `num_layers`
3. `activation`
The legacy experiment uses a 3-layer ReLU MLP with hidden size 256.
### Optimization
Match the legacy training knobs:
1. Separate advantage and strategy batch sizes.
2. Separate advantage and strategy update counts per iteration.
3. Adam `weight_decay`.
4. Gradient clipping.
### Traversal
Match legacy traversal semantics and metrics:
1. Per-player traversal count behavior.
2. Full-depth traversal when `max_depth: null`.
3. Max nodes per traversal.
4. Self-play league behavior.
5. Endpoint depth accounting:
- endpoint depth sum
- endpoint depth buckets
- depth bucket width
- depth bucket max
6. Optional worker progress logging.
7. Optional hotspot profiling.
### Run Loop And Checkpointing
Match legacy runtime behavior:
1. Stop by `max_hours`.
2. Stop by `max_iterations` when set.
3. Allow `max_iterations: null`.
4. Save every N iterations through `save_iteration_interval`.
5. Support `save_latest_only`.
6. Preserve useful config artifacts in the run directory.
7. Keep `metrics.jsonl`, `runtime_progress.json`, and `train.log` useful for
comparing this run against the legacy report. Exact artifact layout
compatibility is optional.
### Evaluation And Metrics
Match the legacy evaluation configuration:
1. `evaluation.on_max_steps`.
2. Opponent list:
- `random`
- `passive_discard`
- `safe_heuristic`
- `safe_heuristic_loose`
- `safe_heuristic_strict`
- `noisy_safe`
3. Metric suffixes used by the legacy analysis scripts, including opening
quality, opened-color distribution, discard take rates, expedition quality,
policy entropy, and timeout counts.
4. Preserve comparable semantics for the legacy experiment's decision criteria.
5. Exact metric key/schema compatibility is optional unless legacy `analyze.py`
is reused directly.
6. Confirm whether the current classic evaluator emits every required core
metric.
7. Add a small adapter or mapping table if this repository keeps different
metric names.
### Analysis Compatibility
The legacy `analyze.py` expects specific run artifacts and metric keys. Direct
reuse is optional. The required goal is that the generated run data can be
compared against the legacy report through documented metric semantics:
1. Core metric meanings are documented.
2. Latest iteration and latest eval iteration can be inferred.
3. Endpoint-depth metrics are present or explicitly marked as omitted.
4. Evaluation metric prefixes and suffixes are documented when they differ from
legacy output.
5. If we choose to reuse legacy `analyze.py` directly, then add a compatibility
adapter for keys and artifact layout.
## Suggested Commit Breakdown
1. Reproduction config schema and mapped experiment preset.
2. Configurable network and optimizer knobs.
3. Derived and slot-aware encoding.
4. Run loop, checkpoint, and evaluation compatibility.
5. Traversal endpoint metrics, progress, and profile compatibility.
6. Metric semantics and optional analysis adapter pass.
## Definition Of Done
The reproduction work is complete when:
1. A mapped config exists in this repository for the legacy experiment.
2. Every relevant legacy hyperparameter is either represented or explicitly
marked as intentionally irrelevant.
3. A smoke-sized version of the mapped config runs end to end.
4. The full mapped config starts training with the same key semantics.
5. The information-state input dimension matches the legacy slot-aware setup for
tier3.
6. The core training metrics and evaluation metrics are semantically comparable
with the legacy experiment report.