50-iter sweep on default.yaml with stronger MCTS exploration: - c_puct: 3.0 -> 5.0 (UCB weight, more exploration of low-prior actions) - root_dirichlet_epsilon: 0.25 -> 0.4 (more noise injected at root prior) Standalone eval at iter 50 (30 games/opponent, all natural-end, timeouts=0): - vs heuristic-balanced: W=0/30 S=-70.5 (PA 0.15) - vs heuristic-aggressive: W=2/30 S=-65.1 (PA 0.14) [+10, +4] - vs heuristic-cautious: W=1/30 S=-48.0 (PA 0.14) [+2] 3 natural-end wins vs prev trapfix baseline iter 44 (which had 0 natural wins + 1 timeout-tie). Stall trap fixed remains true (timeouts=0 in c1). Trade-off observed: more exploration -> higher variance. Score avg vs cautious worsened (-32 -> -48), but win events appeared. For the non-terminal-win objective, exploration win > score-avg loss. Next: commit to long run (300 iter) with these params before tuning more.
31 lines
731 B
Bash
Executable File
31 lines
731 B
Bash
Executable File
#!/bin/bash
|
|
# Autonomous cycle eval helper.
|
|
# Usage: ./autonomous_cycle_eval.sh <run-prefix> [extra eval args...]
|
|
# Finds latest run matching prefix, runs eval --ckpt latest.pt with 30 games,
|
|
# and reports: timeouts, natural-end wins per opponent.
|
|
|
|
set -euo pipefail
|
|
|
|
PREFIX="${1:-}"
|
|
shift || true
|
|
|
|
if [ -z "$PREFIX" ]; then
|
|
echo "usage: $0 <run-prefix> [extra eval args...]"
|
|
exit 1
|
|
fi
|
|
|
|
RUN=$(ls -td runs/*${PREFIX}* 2>/dev/null | head -1)
|
|
if [ -z "$RUN" ]; then
|
|
echo "no run matching ${PREFIX}" >&2
|
|
exit 1
|
|
fi
|
|
|
|
CKPT="$RUN/latest.pt"
|
|
if [ ! -f "$CKPT" ]; then
|
|
echo "no checkpoint at $CKPT" >&2
|
|
exit 1
|
|
fi
|
|
|
|
echo "=== eval $CKPT ==="
|
|
uv run lost-cities-ismcts eval --ckpt "$CKPT" --games 30 --verbose "$@" 2>&1
|