Spot-check of the three drafts gemini produced in the --max 3
survey smoke test: all cited file paths exist, line numbers and
function names land within 1-2 lines of actual symbols
(game.pyx:217 cdef class GameState, evaluate.py:220 batched-entropy
block, trainer.py:892 _evaluate_parallel, action_distribution at
evaluate.py:238 with cited code at line 250 inside it). Numbers
cross-checked against archives match. Conclusions preserved.
The one systemic weakness was the Last-verified commit field:
gemini left a `<short-hash>` placeholder, a literal `HEAD`, or
omitted the commit entirely depending on the call. Fixed in two
places:
1. Manually patched the three drafts before acceptance and copied
them into docs/research/.
2. Added _current_commit_sha and _post_process_draft helpers to
both librarian_survey.py and librarian_promote.py. The drafts
now go through `**Last verified:**` line normalization that
substitutes today's date and `git rev-parse --short HEAD`
before being written to disk. Future runs converge
deterministically.
Net: docs/research/ gains classic-port-notes.md,
deep-cfr-batched-evaluation.md, and deep-cfr-evaluation-profile.md.
12 archive entries remain unprocessed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
scripts/librarian_survey.py walks docs/archive/*.md and dispatches
every entry without a research counterpart through the same prompt
assembly as librarian_promote.py. Outputs land under
runs/tmp/librarian-survey-<timestamp>/, classified into:
<stem>.md draft, ready to copy into docs/research/
<stem>.SKIP.txt LLM's one-line "not promotable" reason
<stem>.ERROR.txt CLI stderr if the call itself failed
Counterpart detection uses exact stem match plus a tail-match
heuristic so research notes that intentionally drop a domain prefix
still suppress their archive. Verified against the current tree:
docs/archive/deep-cfr-opponent-policy-network-divergence-* is
correctly recognized as already covered by
docs/research/opponent-policy-network-divergence.md.
--dry-run lists candidates and suggested research targets without
calling the LLM. --max N caps processed archives per run, useful as
a cost guard. Sequential dispatch; one LLM call per archive.
Plan updated to mark survey mode complete.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>