f398c9fc4c1121856dc802beb57198c650ada8b5
Strategy network (학습 중인 average policy)를 traversal opponent로 사용하는 새 옵션 추가. Deep CFR 이론적 수렴이 average strategy에 대한 보장이라는 점에 착안 — opponent_policy=network의 발산 문제를 완화할 수 있는지 실증. 구현: - config: opponent_policy validator에 average_strategy 추가 - traversal.pyx: opponent_policy_id=3, strategy_network 인자, softmax 기반 policy 도출 (_policy_from_strategy_network) - workers.py: TraversalWorkerBatch에 strategy_network state_dict 추가 - trainer.py: 직렬/병렬 traversal call에 strategy_network 전달 - 1000-iter 실험 config 추가 (opponent_policy=network와 동일 hyperparam) Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
coolrl-lost-cities
Focused Lost Cities extraction from the legacy coolrl repository.
The current implementation starts with the classic two-player card game:
- classic 5-expedition rules by default
- Python/Cython game engine
- env wrapper
- random, passive-discard, and safe-heuristic bots
- core rule, scoring, mask, env, canonical-state, bot, and GUI smoke tests
Training code, Deep CFR, learned-policy evaluation, GUI, and web client are intentionally outside the first port.
Development
uv run pytest tests/games/classic
uv run lost-cities-classic
For future GUI work, install the optional GUI dependencies:
uv sync --extra gui
Run the classic pygame GUI:
uv run lost-cities-classic-gui --mode pvc --bot safe-heuristic
The GUI uses the in-process Cython game engine.
Basic Usage
from coolrl_lost_cities.games.classic import GameState, build_bot, classic_config
state = GameState.new_game(classic_config(seed=1))
bot = build_bot("random", seed=1)
while not state.terminal:
state.apply_action(bot.act(state))
print(state.total_score(0), state.total_score(1))
See classic port notes for the current direction.
Languages
Python
73.3%
Cython
21.7%
Julia
4.8%
Shell
0.2%