Skip to content

feat: add capabilities and problems registries with captain skills - #98

Closed
harsh9200 wants to merge 4 commits into
kunchenguid:mainfrom
harsh9200:fm/capabilities-registry-c8
Closed

harsh9200 wants to merge 4 commits into
kunchenguid:mainfrom
harsh9200:fm/capabilities-registry-c8

Conversation

@harsh9200

Copy link
Copy Markdown

Intent

The developer (acting as captain managing the firstmate orchestration system) directed a crewmate to build a capabilities and problems registry for firstmate-on-itself, working from a written spec at .lavish/firstmate-ios/capabilities-registry-spec.md. The goal was three artifacts: a tracked CAPABILITIES.md living manifest accurately surveying the real environment's tools, MCPs, skills, and quality gates with their limits; a tracked PROBLEMS.md registry using an explicit consistent schema (problem, symptom, impact, suspected root cause, candidate fix/tool, status); and two captain-invocable skills, propose-tool and report-problem, each appending a structured entry to the right registry and queuing an evaluation. They required a minimal AGENTS.md edit (kept in its own section to rebase cleanly against a concurrent watcher fix touching sections 3/5/8) pointing to the manifest and both skills. They explicitly mandated exactly one Codex CLI review pass over the committed diff (checking schema consistency and that skills append valid entries), with real fixes applied and committed but no second pass, before validating through the no-mistakes pipeline.

What Changed

  • Added two tracked registries at the repo root: CAPABILITIES.md, a living manifest of tools, MCP servers, skills, and quality gates with their known limits, and PROBLEMS.md, a recurring-problem registry with an explicit schema (problem, symptom, impact, suspected root cause, candidate fix/tool, status).
  • Added the captain-invocable propose-tool and report-problem skills plus the shared bin/fm-registry.sh helper they call, which appends a schema-exact entry to the matching registry and queues an evaluation pointer; covered by the new hermetic tests/fm-registry.test.sh suite.
  • Wired the new artifacts into the docs: a new AGENTS.md section 14 pointing to both registries and skills, and updates to CONTRIBUTING.md, README.md, and docs/scripts.md, including a Document-gate sync of the tracked-file lists to include the two new registries.

Risk Assessment

✅ Low: Additive change of documentation registries plus one well-guarded append helper whose substantive issues were already fixed in a prior review pass; schemas are consistent and only minor latent fragilities remain.

Testing

Baseline checks (bash -n, --help) passed, then I drove the helper that backs both captain skills end-to-end in a sandboxed FM_ROOT and confirmed propose-tool/report-problem append entries whose fields match the declared registry schemas byte-for-byte and queue a corresponding evaluation line, including all Codex-review hardening paths (collision suffixing, slug fallback, --fix default, input validation, and the queue-preflight guard that prevents an orphaned registry entry). I also reviewed the three docs/skills artifacts and the contained AGENTS.md section 14 for correctness and link resolution. No existing test covered this script, so I added a hermetic tests/fm-registry.test.sh (matching the repo's tests/lib.sh convention and CI discovery) that locks in the schema-consistency, queue, and safety behaviors; it passes 6/6. This is a CLI/markdown change with no rendered UI surface, so the reviewer-visible evidence is a CLI transcript plus the generated registry/queue markdown rather than a screenshot. All testing wrote to the evidence dir and temp sandboxes only; the change under test is untouched and the sole working-tree addition is the intentional new test file.

Evidence: End-to-end CLI transcript: captain runs /propose-tool and /report-problem (commands, confirmations, generated CAPABILITIES.md/PROBLEMS.md entries, and queued evaluations)
# fm-registry end-to-end transcript (captain skill flow)

Exercised `bin/fm-registry.sh` (the helper backing /propose-tool and /report-problem)
against copies of the real tracked registries via FM_ROOT_OVERRIDE.

## 1. Captain runs /propose-tool
`` `
$ bin/fm-registry.sh propose-tool --tool "ripgrep (rg)" \
    --replaces "grep for fleet-wide code search" \
    --why "10-50x faster on large clones, respects .gitignore by default, fewer false hits" \
    --notes "already a common dependency; install via brew"
recorded T-20260627-030509-ripgrep-rg in CAPABILITIES.md and queued its evaluation in .../data/evaluation-queue.md
`` `

### -> appended to CAPABILITIES.md "Proposed tools" section
`` `
### T-20260627-030509-ripgrep-rg - ripgrep (rg)
- **Tool:** ripgrep (rg)
- **Replaces:** grep for fleet-wide code search
- **Why better:** 10-50x faster on large clones, respects .gitignore by default, fewer false hits
- **Notes:** already a common dependency; install via brew
- **Status:** proposed
- **Proposed:** 2026-06-27
`` `

## 2. Captain runs /report-problem
`` `
$ bin/fm-registry.sh report-problem --problem "Peeking panes too often burns context budget" \
    --symptom "repeated full-pane peeks during supervision inflate token use late in a session" \
    --impact "faster context growth, earlier decoder-flip risk on long sessions" \
    --cause "defaulting to pane peeks instead of reading the ~30-token status file first" \
    --fix "status-file-first discipline; cap peeks at 40 lines"
recorded P-20260627-030509-peeking-panes-too-often-burns-co in PROBLEMS.md and queued its evaluation in .../data/evaluation-queue.md
`` `

### -> appended to PROBLEMS.md "Registry" section
`` `
### P-20260627-030509-peeking-panes-too-often-burns-co - Peeking panes too often burns context budget
- **Problem:** Peeking panes too often burns context budget
- **Symptom:** repeated full-pane peeks during supervision inflate token use late in a session
- **Impact:** faster context growth, earlier decoder-flip risk on long sessions
- **Suspected root cause:** defaulting to pane peeks instead of reading the ~30-token status file first
- **Candidate fix / tool:** status-file-first discipline; cap peeks at 40 lines
- **Status:** reported
- **Reported:** 2026-06-27
`` `

## 3. Both queued an evaluation (fleet-local data/evaluation-queue.md)
`` `
- [ ] evaluate tool proposal T-20260627-030509-ripgrep-rg (ripgrep (rg), replaces grep for fleet-wide code search) -> CAPABILITIES.md (filed 2026-06-27)
- [ ] evaluate problem P-20260627-030509-peeking-panes-too-often-burns-co (Peeking panes too often burns context budget) -> PROBLEMS.md (filed 2026-06-27)
- [ ] evaluate tool proposal T-20260627-030536-dupe (dupe, replaces x) -> CAPABILITIES.md (filed 2026-06-27)
- [ ] evaluate tool proposal T-20260627-030536-dupe-2 (dupe, replaces x) -> CAPABILITIES.md (filed 2026-06-27)
- [ ] evaluate problem P-20260627-030536-entry (??? 上海 !!!) -> PROBLEMS.md (filed 2026-06-27)
- [ ] evaluate problem P-20260627-030536-nofix-case (nofix case) -> PROBLEMS.md (filed 2026-06-27)
`` `
Evidence: Generated proposed-tool entry appended to CAPABILITIES.md (schema-exact)
### T-20260627-030509-ripgrep-rg - ripgrep (rg)
- **Tool:** ripgrep (rg)
- **Replaces:** grep for fleet-wide code search
- **Why better:** 10-50x faster on large clones, respects .gitignore by default, fewer false hits
- **Notes:** already a common dependency; install via brew
- **Status:** proposed
- **Proposed:** 2026-06-27
Evidence: Generated problem entry appended to PROBLEMS.md (schema-exact) + queued evaluations
### P-20260627-030509-peeking-panes-too-often-burns-co - Peeking panes too often burns context budget
- **Problem:** Peeking panes too often burns context budget
- **Symptom:** repeated full-pane peeks during supervision inflate token use late in a session
- **Impact:** faster context growth, earlier decoder-flip risk on long sessions
- **Suspected root cause:** defaulting to pane peeks instead of reading the ~30-token status file first
- **Candidate fix / tool:** status-file-first discipline; cap peeks at 40 lines
- **Status:** reported
- **Reported:** 2026-06-27

# evaluation-queue.md:
- [ ] evaluate tool proposal T-20260627-030509-ripgrep-rg (ripgrep (rg), replaces grep for fleet-wide code search) -> CAPABILITIES.md (filed 2026-06-27)
- [ ] evaluate problem P-20260627-030509-peeking-panes-too-often-burns-co (Peeking panes too often burns context budget) -> PROBLEMS.md (filed 2026-06-27)
Evidence: New hermetic regression test (added, 6/6 pass)

Source: New hermetic regression test (added, 6/6 pass)

ok - propose-tool: appends the documented schema to CAPABILITIES.md and queues an evaluation
ok - report-problem: appends the documented schema (and TBD default) to PROBLEMS.md and queues an evaluation
ok - propose-tool: a colliding id is suffixed -N instead of overwriting
ok - fm-registry: a no-slug title gets the stable 'entry' id fallback
ok - fm-registry: missing flags, unknown subcommand, and unknown flags fail non-zero with a message
ok - fm-registry: a failed queue preflight leaves the tracked registry untouched

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 infos
  • ℹ️ bin/fm-registry.sh:101 - fm-registry.sh appends new entries to the end of the file (>> "$CAP" / >> "$PROB"), relying on "## Proposed tools" and "## Registry" each being the last section. The CAPABILITIES.md anchor marker ("propose-tool entries are appended below this line") is decorative - the script never seeks to it. If any section is ever added below these, generated entries will silently land in the wrong section. Consider inserting at the anchor marker (or asserting the target section is last) to make the contract robust.
  • ℹ️ bin/fm-registry.sh:142 - The help/usage output is produced by sed -n '13,15p' "$0", hard-coding the line range of the usage comment block. Editing the header comment shifts those lines and silently corrupts --help output. A heredoc usage() or matching the # Usage: anchor would be self-maintaining.
✅ **Test** - passed

✅ No issues found.

  • bash -n bin/fm-registry.sh and bin/fm-registry.sh --help (parse + usage)
  • End-to-end captain flow in a sandboxed FM_ROOT (copies of real registries): bin/fm-registry.sh propose-tool --tool ... --replaces ... --why ... --notes ... and bin/fm-registry.sh report-problem --problem ... --symptom ... --impact ... --cause ... --fix ... -> verified schema-exact entries appended to CAPABILITIES.md / PROBLEMS.md and pointer lines added to data/evaluation-queue.md
  • Robustness: id-collision suffixing (same-second dupe -> -2), slug fallback (??? 上海 !!! -> entry), omitted --fix -> TBD - to be evaluated, missing required flag / unknown subcommand / unknown flag all exit 1 with a message
  • Safety invariant: forced data/ creation failure (FM_DATA_OVERRIDE under a regular file) -> script exits 1 and the tracked registry is NOT mutated (no orphaned entry)
  • Programmatic label comparison: report-problem/propose-tool emitted **Field:** labels vs the schemas declared in PROBLEMS.md/CAPABILITIES.md headers (exact match; only optional Resolved omitted by design); skill doc relative links resolve to repo-root registries
  • Added and ran tests/fm-registry.test.sh (hermetic, matches existing tests/lib.sh convention, picked up by CI's tests/*.test.sh loop) - 6/6 cases pass
🔧 **Document** - 1 issue found → auto-fixed ✅
  • ℹ️ AGENTS.md:46 - AGENTS.md section 1's 'Shared, tracked material means ...' enumeration (line 46) and the section 2 layout tree both omit the two new tracked top-level files CAPABILITIES.md and PROBLEMS.md. Left unchanged because the change deliberately isolated its AGENTS.md edit to a single new section (14) for clean rebasing against a concurrent watcher fix; folding the two files into those enumerations is a judgment call about prioritizing accuracy over that isolation. The parallel list in CONTRIBUTING.md was updated since it carries no such constraint.

🔧 Fix: sync tracked-file lists with new registries
✅ Re-checked - no issues remain.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

- CAPABILITIES.md: living manifest of tools, MCPs, skills, gates, and known limits
- PROBLEMS.md: recurring-problem registry with explicit schema and status vocabulary
- propose-tool / report-problem captain skills, backed by bin/fm-registry.sh
- AGENTS.md section 14 points to the registries and both skills
- align PROBLEMS.md schema with helper/seed fields (Reported required, Resolved optional) instead of the Reported/Resolved mismatch
- make fm-registry.sh ids collision-free (-N suffix) with a non-empty slug fallback
- preflight the evaluation queue before mutating the tracked registry so a failed queue write never leaves an unqueued entry
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant