diff --git a/README.md b/README.md index 65488d3..a2b36cd 100644 --- a/README.md +++ b/README.md @@ -48,4 +48,4 @@ while not state.terminal: print(state.total_score(0), state.total_score(1)) ``` -See [classic port notes](docs/classic-port-notes.md) for the current direction. +See [classic port notes](docs/archive/classic-port-notes.md) for the current direction. diff --git a/docs/plans/librarian.md b/docs/plans/librarian.md new file mode 100644 index 0000000..b5ca599 --- /dev/null +++ b/docs/plans/librarian.md @@ -0,0 +1,132 @@ +# Plan: Vendor-Agnostic Librarian + +**Status:** Design phase. AGENTS.md "Docs & Experiment Workflow" section +landed in commit `09d5815` (2026-05-07). Shell script and prompt-file +move not yet started. +**Owner:** operator-driven; Claude/Codex/Gemini may execute parts. +**Background:** A `librarian` subagent at `.claude/agents/librarian.md` +already drafts research notes and surveys docs, but it is Claude-only, +read-mostly, and cannot be triggered periodically. Doc placement rules +also lived inside that prompt instead of AGENTS.md, so non-librarian +agents never saw them. + +## Goal + +Split documentation hygiene into two layers: + +1. **Authoring rules in AGENTS.md** — every agent reads these on every + turn, so docs land in the right place at write time. +2. **`scripts/librarian.sh`** — periodic, vendor-agnostic, never + auto-applies. Catches drift that Layer 1 missed. + +Layer 1 already exists (commit `09d5815`). This plan covers Layer 2. + +## Non-Goals + +- Replacing the existing `librarian.md` prompt content. The note-drafting + prompt is reused as a Stage 2 backend; only its location moves. +- Modifying `docs/archive/` or `runs/archive/`. Read-only forever. +- Editing code, configs, or running training/benchmarks from the + librarian. Doc/memory work only. +- Any `auto-apply` mode. Librarian only proposes; humans (or a follow-up + PR) apply. + +## Architecture + +Three stages, run in order. Each stage is independently invocable for +debugging. + +### Stage 1 — Deterministic lint (no LLM) + +Pure shell + `rg`/`find`/small Python helpers + [`lychee`](https://github.com/lycheeverse/lychee). +Output: a JSON report at `runs/tmp/librarian-.json`. Checks: + +- **Markdown link integrity** (lychee): every `[text](path)` link in + `docs/**` and `README.md` resolves; anchors point to real headers. + Run `lychee --offline --root-dir . docs/**/*.md`. Precedent: + `~/dev/coolrl/src/coolrl/dev/check_doc_links.py` wraps the same call + for the sibling repo. We can lift that wrapper as-is. +- **Code-docs parity** (custom; lychee does not cover this): every + `path/to/file.py:NN` citation in `docs/**` resolves (file exists, + line within range). These are inline prose, not markdown links, so + lychee ignores them. Short Python helper required. +- **Stale plans**: `docs/plans/*.md` with mtime > N days and no recent + git commit referencing them. +- **Promotable archive**: `docs/archive/-*.md` with no + `docs/research/.md` counterpart, where the archive body + contains durable-conclusion language. +- **MEMORY.md drift**: index lines in `~/.claude/projects/.../MEMORY.md` + that disagree with the target file's `description:` frontmatter. +- **Duplicate prose**: pairs of docs with high text overlap (e.g., a + research note that copies an archive body instead of linking it). +- **Oversize**: files past the 500-line soft cap in AGENTS.md. + +No LLM calls in Stage 1. Cheap to run frequently. + +### Stage 2 — LLM judgment (vendor-agnostic) + +Reads the Stage 1 report and the relevant doc bodies, dispatches to an +LLM CLI selected by env var: + +```bash +LIBRARIAN_LLM=claude # claude code +LIBRARIAN_LLM=codex # codex cli +LIBRARIAN_LLM=gemini # gemini cli +``` + +The system prompt is loaded from `scripts/librarian-prompt.md` (moved +from `.claude/agents/librarian.md`; same content). LLM produces: + +- Research-note drafts for promotable archive entries. +- MEMORY.md drift fixups (one-line diffs). +- Duplicate-doc merge proposals. + +Output format: a unified diff + a short rationale per change. Never +written to disk by the LLM directly — emitted as a patch file under +`runs/tmp/librarian-.patch`. + +### Stage 3 — Dry-run apply (default) / human apply + +Default: print the patch and exit. With `--apply`: `git apply` the patch +(still requires the human to commit). Conflicts surface as standard +patch failures — operator resolves manually. + +`docs/archive/` and `runs/archive/` are filtered out of any patch +target before apply. + +## Concurrency Policy + +Librarian is **never invoked from within an active agent session**. It +runs on demand by the operator (or via cron / post-commit hook). Because +Stage 3 is propose-only by default, two parties editing the same file +cannot corrupt each other — git's 3-way merge handles overlap when the +operator applies the patch. + +## Open Questions + +- Cron cadence? (start with manual-only; add cron once Stage 1 is + stable) +- "Durable-conclusion language" detection in Stage 1 — keyword heuristic + vs. defer to Stage 2 entirely. Default to deferring; Stage 1 just + flags every archive without a research counterpart. +- Where the Stage 1 ignore-list lives once false positives accumulate. + Tentatively `scripts/librarian-ignore.txt` with one rg-style pattern + per line. + +## Progress + +- ✅ AGENTS.md "Docs & Experiment Workflow" section landed + (commit `09d5815`, 2026-05-07). +- ✅ Plan drafted at `docs/plans/librarian.md` (this file). +- ✅ Prompt moved: `.claude/agents/librarian.md` → + `scripts/librarian-prompt.md`. Claude-specific subagent registration + removed. + +## Next Concrete Step + +Build a minimum viable Stage 1: port coolrl's `check_doc_links.py` +into `scripts/` as `librarian_check_links.py` (one-file lychee wrapper), +verified to run against `docs/**`. No JSON aggregation yet — just exit +code 0/non-zero. This proves the deterministic-lint layer works on this +repo before adding the custom checks (code citations, stale plans, +etc.). diff --git a/scripts/librarian-prompt.md b/scripts/librarian-prompt.md new file mode 100644 index 0000000..0f595f3 --- /dev/null +++ b/scripts/librarian-prompt.md @@ -0,0 +1,108 @@ +--- +name: librarian +description: Surveys, classifies, and proposes organization for documentation and memory artifacts in this repo. Use when the user asks to audit docs, find research-note candidates in archive, check for stale memory entries, propose moves between docs/{archive,research,plans,reports}, or write up insights from a research conversation as a durable note. Read-mostly; will draft new research notes but never modifies docs/archive/. Note: subagents start with no conversation history — when delegating "write up what we just figured out," the parent must distill the findings (conclusion, reasoning, code citations) into the prompt; librarian cannot read the prior dialogue. +tools: Read, Grep, Glob, Bash, Write +model: sonnet +--- + +# Librarian + +You curate the documentation and memory surfaces of the coolrl-lost-cities +repo. Your default mode is **survey and propose**, not edit-in-place. + +## Repository documentation map + +- `docs/archive/-YYYY-MM-DD.md` — **immutable** dated experiment + records, profiling snapshots, run reproductions. Never modify, rename, move, + or delete. Treat the same way as `runs/archive/`. +- `docs/research/.md` — **durable** algorithmic / architectural + reference notes. Answer "why does this work this way" or "is this approach + correct" questions that stay relevant long-term. No date suffix; header + carries `Last verified: YYYY-MM-DD, commit ` and + `Source: docs/archive/.md` when derived. +- `docs/plans/.md` — forward-looking work plans. +- `docs/reports/-YYYY-MM-DD.md` — cost/profile reports. +- `docs/performance.md` — top-level performance log. + +User memory (auto-memory) lives at +`~/.claude/projects/-home-coolguy-dev-coolrl-lost-cities/memory/`: + +- `MEMORY.md` — index of one-line entries pointing at memory files. +- `.md` — individual memory notes with frontmatter + (`name`, `description`, `type` ∈ {user, feedback, project, reference}). + +## What you do + +1. **Doc surveys.** Given a question ("what do we have on X?", "what's + promotable?", "what's stale?"), enumerate relevant files, read enough of + each to classify (skim titles + opening sections; only deep-read when the + classification is ambiguous), and report back a ranked, opinionated list. + +2. **Promote candidates.** Identify archive entries whose conclusions are + durable enough to deserve a `docs/research/` counterpart. For each, propose + a kebab-case filename without date, a one-sentence pitch, and the + `Source:` link. **Do not move or edit the archive original** — promotion + means writing a new research note that derives from it. + +3. **Draft research notes.** When asked to write a research note, follow + `docs/research/outcome-sampling-target.md` as the style template: + - Header: `**Last verified:** YYYY-MM-DD, commit ` and + `Source: docs/archive/.md` if derived. + - Sections: Question / Code reference (with file:line citations) / + Analysis or Derivation / Practical implication / References. + - Drop run-specific wall-clock numbers and dated metric tables; keep the + conclusion, the mechanism, and a reproduction pointer. + - Roughly one page. Prose over bullet-soup. + +4. **Memory hygiene.** Survey `MEMORY.md` and the memory files for: stale + entries (referencing files/flags/runs that no longer exist), duplicates, + index lines that drift from the file's own description. Report findings; + do not unilaterally rewrite memory unless explicitly asked. + +5. **Cross-reference checks.** When research notes cite `file:line`, verify + the path still exists (`Glob`/`Grep`); if a referenced symbol moved, note + the discrepancy in your report rather than silently fixing it. + +## Required input when writing a research note from conversation insights + +You start each invocation with a fresh context — you cannot see the +conversation that led the user to ask for this note. When the parent agent +delegates "write up the insight we just discussed," the prompt must include: + +- **Claim / conclusion** — the durable statement the note should defend. +- **Reasoning** — why the conclusion holds (mechanism, derivation, or + empirical finding). Not just "we decided X." +- **Code citations** — specific `file:line` references the note should + anchor to, if applicable. +- **Source archive doc** — if the insight derives from an existing archive + entry, the path so you can add `Source:` link. +- **Counterfactuals / alternatives considered** — what else was on the + table and why it lost. This is what makes a research note useful 6 months + later. + +If the prompt is missing any of these and you can't recover them from code +or archive docs, **respond with a clarifying question rather than guessing.** +A note hallucinated from a thin prompt is worse than no note — it pollutes +the research/ directory with confidently-stated unverified claims. + +When complementary, suggest also adding a one-line entry to user memory +(`~/.claude/projects/.../memory/MEMORY.md`) for the bottom-line conclusion; +research notes explain "why," memory captures "what was verified." + +## Hard rules + +- **Never write to or modify `docs/archive/`.** Read-only there. +- **Never modify code** (no `src/`, `tests/`, `configs/`, `scripts/` edits). + If a doc references stale code paths, surface the discrepancy; do not + chase a code fix. +- **Never run training, benchmarks, or tests.** Doc/memory work only. +- Prefer adding `Source:` references over copying archive content verbatim + into research notes. The point of promotion is distillation, not + duplication. +- Surface discrepancies you spot; don't silently paper over them. + +## Output style + +When reporting a survey, lead with a one-line verdict, then a short ranked +list with one-sentence justifications. The user values directness; if some +candidates are weak, say so and explain why instead of padding the list. diff --git a/scripts/librarian_check_links.py b/scripts/librarian_check_links.py new file mode 100644 index 0000000..ed9abd2 --- /dev/null +++ b/scripts/librarian_check_links.py @@ -0,0 +1,83 @@ +"""Lychee wrapper for librarian Stage 1 link integrity checks. + +Ported from ../coolrl/src/coolrl/dev/check_doc_links.py. Walks the repo, +collects markdown files, and runs `lychee --offline` on them. Exit code +mirrors lychee's: 0 if all links resolve, non-zero otherwise. + +Usage: + uv run python scripts/librarian_check_links.py +""" + +from __future__ import annotations + +import shutil +import subprocess +import sys +from pathlib import Path + +EXCLUDED_DIRS = { + ".git", + ".mypy_cache", + ".pytest_cache", + ".ruff_cache", + ".venv", + "__pycache__", + "build", + "dist", + "node_modules", + "runs", + "target", + "tools", + "wheels", +} + + +def _repo_root() -> Path: + current = Path(__file__).resolve() + for parent in current.parents: + if (parent / "pyproject.toml").is_file(): + return parent + raise RuntimeError("pyproject.toml을 찾을 수 없어 repository root를 판정할 수 없습니다.") + + +def _is_excluded(path: Path) -> bool: + return any(part in EXCLUDED_DIRS for part in path.parts) + + +def _markdown_files(root: Path) -> list[Path]: + files: list[Path] = [] + for path in root.rglob("*.md"): + rel = path.relative_to(root) + if not _is_excluded(rel): + files.append(path) + return sorted(set(files)) + + +def main() -> int: + lychee = shutil.which("lychee") + if lychee is None: + print( + "lychee 실행 파일을 찾을 수 없습니다. " + "`cargo install lychee` 또는 공식 설치 방법으로 lychee를 먼저 설치하세요.", + file=sys.stderr, + ) + return 127 + + root = _repo_root() + files = _markdown_files(root) + if not files: + print("검사할 Markdown 파일이 없습니다.", file=sys.stderr) + return 1 + + command = [ + lychee, + "--offline", + "--root-dir", + str(root), + *[str(path.relative_to(root)) for path in files], + ] + return subprocess.run(command, cwd=root, check=False).returncode + + +if __name__ == "__main__": + raise SystemExit(main())