Deep CFR legacy 실험 재현 계획 문서화
This commit is contained in:
@@ -0,0 +1,209 @@
|
|||||||
|
# Deep CFR Legacy Experiment Reproduction Plan
|
||||||
|
|
||||||
|
This document tracks what is needed to reproduce the legacy `../coolrl`
|
||||||
|
experiment below with the same intended semantics in this repository:
|
||||||
|
|
||||||
|
`experiments/lost_cities/deep_cfr_pure_self_play_zero_pit_poc_full_depth_slot_aware_playability`
|
||||||
|
|
||||||
|
The goal is not to run a similar Deep CFR configuration. The goal is to map the
|
||||||
|
legacy experiment's hyperparameters and feature semantics into this repository's
|
||||||
|
own config schema so that the training run means the same thing.
|
||||||
|
|
||||||
|
Exact legacy YAML compatibility is not required. It is acceptable to create a
|
||||||
|
new config file under this repository as long as every relevant legacy
|
||||||
|
hyperparameter is represented explicitly and the differences are documented.
|
||||||
|
|
||||||
|
## Source Experiment
|
||||||
|
|
||||||
|
Legacy config path:
|
||||||
|
|
||||||
|
`/home/coolguy/dev/coolrl/experiments/lost_cities/deep_cfr_pure_self_play_zero_pit_poc_full_depth_slot_aware_playability/config.yaml`
|
||||||
|
|
||||||
|
Important legacy settings:
|
||||||
|
|
||||||
|
1. `seed: 79`
|
||||||
|
2. `max_hours: 4`
|
||||||
|
3. `max_iterations: null`
|
||||||
|
4. `device: CUDA`
|
||||||
|
5. `use_amp: false`
|
||||||
|
6. `rules.tier: tier3`
|
||||||
|
7. `encoding.derived_playability: true`
|
||||||
|
8. `encoding.slot_aware_playability: true`
|
||||||
|
9. `network.hidden_size: 256`
|
||||||
|
10. `network.num_layers: 3`
|
||||||
|
11. `network.activation: relu`
|
||||||
|
12. `traversal.traversals_per_player: 70`
|
||||||
|
13. `traversal.max_depth: null`
|
||||||
|
14. `traversal.max_nodes_per_traversal: 1000`
|
||||||
|
15. `traversal.opponent_policy: self_play_league`
|
||||||
|
16. `traversal.outcome_sampling_epsilon: 0.2`
|
||||||
|
17. `traversal.outcome_sampling_value_clip: 500`
|
||||||
|
18. `traversal.outcome_unsampled_regret: zero`
|
||||||
|
19. `optimization.advantage_batch_size: 1024`
|
||||||
|
20. `optimization.strategy_batch_size: 1024`
|
||||||
|
21. `optimization.advantage_updates_per_iteration: 256`
|
||||||
|
22. `optimization.strategy_updates_per_iteration: 256`
|
||||||
|
23. `optimization.learning_rate: 3.0e-5`
|
||||||
|
24. `optimization.weight_decay: 1.0e-4`
|
||||||
|
25. `optimization.grad_clip: 1.0`
|
||||||
|
|
||||||
|
## Required Reproduction Work
|
||||||
|
|
||||||
|
### Config
|
||||||
|
|
||||||
|
Represent the legacy hyperparameters in this repository's config schema:
|
||||||
|
|
||||||
|
1. Top-level:
|
||||||
|
- `experiment_name`
|
||||||
|
- `seed`
|
||||||
|
- `max_iterations`
|
||||||
|
- `max_hours`
|
||||||
|
- `device`
|
||||||
|
- `use_amp`
|
||||||
|
2. `rules`:
|
||||||
|
- `tier`
|
||||||
|
- optional direct `LostCitiesConfig` field overrides
|
||||||
|
3. `encoding`:
|
||||||
|
- `derived_playability`
|
||||||
|
- `slot_aware_playability`
|
||||||
|
4. `network`:
|
||||||
|
- `hidden_size`
|
||||||
|
- `num_layers`
|
||||||
|
- `activation`
|
||||||
|
5. `traversal` semantics:
|
||||||
|
- per-player traversal count
|
||||||
|
- max nodes per traversal
|
||||||
|
- traversal worker chunk size
|
||||||
|
- self-play league settings
|
||||||
|
6. `optimization` semantics:
|
||||||
|
- separate advantage and strategy batch sizes
|
||||||
|
- separate advantage and strategy update counts
|
||||||
|
- `weight_decay`
|
||||||
|
- `grad_clip`
|
||||||
|
7. `evaluation.on_max_steps`
|
||||||
|
8. `checkpoint`:
|
||||||
|
- `save_iteration_interval`
|
||||||
|
- `save_latest_only`
|
||||||
|
- `progress_interval_seconds`
|
||||||
|
|
||||||
|
### Encoding
|
||||||
|
|
||||||
|
Port the legacy feature semantics:
|
||||||
|
|
||||||
|
1. Base information-state encoding must remain deterministic.
|
||||||
|
2. Add `derived_playability` color-level features.
|
||||||
|
3. Add `slot_aware_playability` hand-slot action-local features.
|
||||||
|
4. Preserve feature order and dimensions from legacy where possible.
|
||||||
|
5. Add tests for input dimension and known-state feature values.
|
||||||
|
|
||||||
|
The slot-aware block is the core of the source experiment. Without this block,
|
||||||
|
the reproduction is not meaningful.
|
||||||
|
|
||||||
|
### Network
|
||||||
|
|
||||||
|
Make the MLP configurable:
|
||||||
|
|
||||||
|
1. `hidden_size`
|
||||||
|
2. `num_layers`
|
||||||
|
3. `activation`
|
||||||
|
|
||||||
|
The legacy experiment uses a 3-layer ReLU MLP with hidden size 256.
|
||||||
|
|
||||||
|
### Optimization
|
||||||
|
|
||||||
|
Match the legacy training knobs:
|
||||||
|
|
||||||
|
1. Separate advantage and strategy batch sizes.
|
||||||
|
2. Separate advantage and strategy update counts per iteration.
|
||||||
|
3. Adam `weight_decay`.
|
||||||
|
4. Gradient clipping.
|
||||||
|
|
||||||
|
### Traversal
|
||||||
|
|
||||||
|
Match legacy traversal semantics and metrics:
|
||||||
|
|
||||||
|
1. Per-player traversal count behavior.
|
||||||
|
2. Full-depth traversal when `max_depth: null`.
|
||||||
|
3. Max nodes per traversal.
|
||||||
|
4. Self-play league behavior.
|
||||||
|
5. Endpoint depth accounting:
|
||||||
|
- endpoint depth sum
|
||||||
|
- endpoint depth buckets
|
||||||
|
- depth bucket width
|
||||||
|
- depth bucket max
|
||||||
|
6. Optional worker progress logging.
|
||||||
|
7. Optional hotspot profiling.
|
||||||
|
|
||||||
|
### Run Loop And Checkpointing
|
||||||
|
|
||||||
|
Match legacy runtime behavior:
|
||||||
|
|
||||||
|
1. Stop by `max_hours`.
|
||||||
|
2. Stop by `max_iterations` when set.
|
||||||
|
3. Allow `max_iterations: null`.
|
||||||
|
4. Save every N iterations through `save_iteration_interval`.
|
||||||
|
5. Support `save_latest_only`.
|
||||||
|
6. Preserve useful config artifacts in the run directory.
|
||||||
|
7. Keep `metrics.jsonl`, `runtime_progress.json`, and `train.log` useful for
|
||||||
|
comparing this run against the legacy report. Exact artifact layout
|
||||||
|
compatibility is optional.
|
||||||
|
|
||||||
|
### Evaluation And Metrics
|
||||||
|
|
||||||
|
Match the legacy evaluation configuration:
|
||||||
|
|
||||||
|
1. `evaluation.on_max_steps`.
|
||||||
|
2. Opponent list:
|
||||||
|
- `random`
|
||||||
|
- `passive_discard`
|
||||||
|
- `safe_heuristic`
|
||||||
|
- `safe_heuristic_loose`
|
||||||
|
- `safe_heuristic_strict`
|
||||||
|
- `noisy_safe`
|
||||||
|
3. Metric suffixes used by the legacy analysis scripts, including opening
|
||||||
|
quality, opened-color distribution, discard take rates, expedition quality,
|
||||||
|
policy entropy, and timeout counts.
|
||||||
|
4. Preserve comparable semantics for the legacy experiment's decision criteria.
|
||||||
|
5. Exact metric key/schema compatibility is optional unless legacy `analyze.py`
|
||||||
|
is reused directly.
|
||||||
|
6. Confirm whether the current classic evaluator emits every required core
|
||||||
|
metric.
|
||||||
|
7. Add a small adapter or mapping table if this repository keeps different
|
||||||
|
metric names.
|
||||||
|
|
||||||
|
### Analysis Compatibility
|
||||||
|
|
||||||
|
The legacy `analyze.py` expects specific run artifacts and metric keys. Direct
|
||||||
|
reuse is optional. The required goal is that the generated run data can be
|
||||||
|
compared against the legacy report through documented metric semantics:
|
||||||
|
|
||||||
|
1. Core metric meanings are documented.
|
||||||
|
2. Latest iteration and latest eval iteration can be inferred.
|
||||||
|
3. Endpoint-depth metrics are present or explicitly marked as omitted.
|
||||||
|
4. Evaluation metric prefixes and suffixes are documented when they differ from
|
||||||
|
legacy output.
|
||||||
|
5. If we choose to reuse legacy `analyze.py` directly, then add a compatibility
|
||||||
|
adapter for keys and artifact layout.
|
||||||
|
|
||||||
|
## Suggested Commit Breakdown
|
||||||
|
|
||||||
|
1. Reproduction config schema and mapped experiment preset.
|
||||||
|
2. Configurable network and optimizer knobs.
|
||||||
|
3. Derived and slot-aware encoding.
|
||||||
|
4. Run loop, checkpoint, and evaluation compatibility.
|
||||||
|
5. Traversal endpoint metrics, progress, and profile compatibility.
|
||||||
|
6. Metric semantics and optional analysis adapter pass.
|
||||||
|
|
||||||
|
## Definition Of Done
|
||||||
|
|
||||||
|
The reproduction work is complete when:
|
||||||
|
|
||||||
|
1. A mapped config exists in this repository for the legacy experiment.
|
||||||
|
2. Every relevant legacy hyperparameter is either represented or explicitly
|
||||||
|
marked as intentionally irrelevant.
|
||||||
|
3. A smoke-sized version of the mapped config runs end to end.
|
||||||
|
4. The full mapped config starts training with the same key semantics.
|
||||||
|
5. The information-state input dimension matches the legacy slot-aware setup for
|
||||||
|
tier3.
|
||||||
|
6. The core training metrics and evaluation metrics are semantically comparable
|
||||||
|
with the legacy experiment report.
|
||||||
Reference in New Issue
Block a user