Skip to content

Skills: every model-invocable skill listed on and Codex-enabled, six re-pins, one retirement, the skill lifecycle guide and MCP_TIMEOUT parity (unit F3) - #553

Merged
seathatflowsinourveins merged 22 commits into
mainfrom
claude/sota-defaults-f3-20260930
Oct 1, 2026
Merged

seathatflowsinourveins merged 22 commits into
mainfrom
claude/sota-defaults-f3-20260930

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes, in one or two sentences: Every model-invocable skill is listed on for Claude and enabled for Codex; six pins are refreshed and resolving-merge-conflicts is retired after upstream removed it. adoption/skills/lifecycle.md becomes the single skill lifecycle document, and the Claude template sets MCP_TIMEOUT to 120000 to match the Codex lane's MCP startup allowance.
  • Base commit: 11227bfdf25b55b0481a9e78985c3fed26052472
  • Lane: lane:foundation
  • Owned paths touched: adoption/skills/{manifest.json,lifecycle.md}; adoption/templates/claude.settings.template.json (skillOverrides, skillListingBudgetFraction, env MCP_TIMEOUT); adoption/templates/codex.config.template.toml ([skills]); adoption/bootstrap.md (one step-4 line); blueprints/runtime-workers/skills/{manifest.json,README.md,research.md}; docs/decisions/{2026-09-25-skills-trial-and-usage.md (addendum), 2026-09-30-skills-llm-native-listing.md, 2026-09-30-native-skill-lifecycle.md}; evidence/artifacts/{native-skill-lifecycle-20260930,native-skill-finalization-20260930}/; tests/{test_skills_manifest,test_runtime_worker_skills,test_install_claude_profile,test_skill_usage}.py; manifests/evidence.json (final commit only).
  • Frozen surfaces touched (Gate A): adoption/skills/manifest.json and adoption/templates/claude.settings.template.json (the freeze is lifted for unit F3 only; this PR merges in the Gate A batch after U6 with the owner's go, and applies nothing to any host); adoption/templates/codex.config.template.toml [skills] (B1 input); manifests/evidence.json under the docs/lanes.md hot-file protocol (last commit only).
  • grants: none
  • Attribution: the source-review fold contains pre-existing uncommitted changes observed in the main checkout; original author not established (snapshot r2, tracked.diff sha256 314bd1b260da09396a0c9cbb708c81b1b280d3cab74d3206ff1a5a09ef3427fe), kept as docs/decisions/2026-09-30-native-skill-lifecycle.md.
  • Merge order: no later than unit F1 (its AGENTS.md line cites adoption/skills/lifecycle.md).

SOTA sources

  • Claude Code docs, read 2026-09-30: skills https://code.claude.com/docs/en/skills (fetched markdown sha256 adc20053…; L290-294 watcher, /reload-skills, /reload-plugins; L372 1,536-char cap; L1124-1130 1% budget, least-invoked dropped first, --debug warning, skillListingBudgetFraction); env vars https://code.claude.com/docs/en/env-vars (a908ea67…; L470 MCP_TIMEOUT "MCP server startup (default: 30000)"; L486 SLASH_COMMAND_TOOL_CHAR_BUDGET 1% with an 8,000-char fallback); MCP https://code.claude.com/docs/en/mcp (86b6d45a…; L403-404 the per-server timeout covers tool execution only); settings reference (skillOverrides, skillListingBudgetFraction). claude mcp add --help on 2.1.285 lists no timeout option.
  • Codex: https://developers.openai.com/codex/skills (d1579156…; L129 and L172-173 automatic detection with restart fallback, L180-185 [[skills.config]], L215 allow_implicit_invocation); openai/codex rust-v0.159.2 (ff6aec96948b70d94983af2641a6b67c94faeff5): codex-rs/ext/skills/src/render.rs L19-27, L126-152, L241-266; codex-rs/core/config.schema.json L4122-4127; codex-rs/models-manager/models.json (gpt-6-astra and gpt-6.1-sol context_window 272000); skills_config.rs, provider/host.rs and skills/src/lib.rs byte-identical to rust-v0.157.1 (36650394c5b38c2990ccf2a3457165ca3e9d9726).
  • Skills CLI: vercel-labs/skills v1.7.0 7407f3893ad4dceab546ac002c3ef806e4000c73 (newest tag; npm latest 1.7.0): src/cli.ts L394-401 (check/update/upgrade call runUpdate), src/list.ts L60-62, src/remove.ts L409-411, skills/find-skills/SKILL.md L35-103.
  • Skill sources, each verified from a blobless clone: affaan-m/ECC c70874fae9eb0e5ad0365beb7e2955899fd1d30f skills/search-first and 2b6e839771e53096d8451a213d40dc64ec8acac0 skills/iterative-retrieval; mattpocock/skills d81f3a183412e71a5b1e84ca21bc1a35eea03a60 (diagnosing-bugs, tdd, codebase-design, improve-codebase-architecture; .changeset/remove-resolving-merge-conflicts.md; removal commit daa01d8) and c55ee46073ed923f86ce59a5eb3b6d895095d1b7 (grill-me, writing-for-agents, the retired skill's historical pin); trailofbits/skills 82fe8226252622fa807643bdca1710901198553a plugins/static-analysis/skills/semgrep (commit 82fe822) and 0cc1c73a5e96749ab32d7ea5e14892fafa6972ae (eight others); anthropics/skills 8a1541c4a3ffa5a20a5a91de0dcf3f0bab1d1ef4 skills/skill-creator and 33375500bcea98d610eb30ce10ac4e59b89c390d (mcp-builder, frontend-design); openai/skills 49f948faa9258a0c61caceaf225e179651397431; typesafe-ai/skills 65a39f393687675ce170e6094757de20370365b9; vercel-labs/agent-browser d01253d9db28d75080e36da3c1c31ef89454731e; cloudflare/security-audit-skill c1c8a8c1471069fb0e188eeaff69b8e8db6564a8.
  • Evaluation method cited in the lifecycle guide: https://developers.openai.com/blog/eval-skills; https://developers.openai.com/blog/rethinking-skills-and-prompts-for-gpt-6-astra.

Evidence-class table

Claim Evidence class Command / receipt
All 28 pins match their upstream tree, SKILL.md sha256 and bytes, description length and invocation flag; each ref is on its default branch; each path is unchanged at HEAD source_review blobless clones plus a scratchpad pin checker (28 PASS, 0 failures)
resolving-merge-conflicts removed upstream (daa01d8), absent at d81f3a18 source_review git rev-parse and git log in the mattpocock/skills clone; changeset text
Listing states, budget sums, template mirror, retired "off", runtime-worker reuse and counts local_integration tests.test_skills_manifest, tests.test_runtime_worker_skills, tests.test_skills_status (failed first where new)
MCP_TIMEOUT "120000" in the template env local_integration McpStartupTimeoutTemplateTests (failed first: None != '120000')
Claude MCP_TIMEOUT semantics; no per-server startup knob source_review env-vars and mcp docs; claude mcp add --help (2.1.285)
Lifecycle guide commands and flags exist source_review install_skills.py and skills_status.py --help; Skills CLI v1.7.0 source
RenderTextAndCli no longer depends on the caller's XDG_STATE_HOME local_integration tests.test_skill_usage under a scratch XDG_STATE_HOME: 2 failures before, OK after
Codex lane's native activation runs and lifecycle receipt native_proven as recorded by the attributed receipts; not re-run in this PR evidence/artifacts/native-skill-finalization-20260930/results.json, native-skill-lifecycle-20260930/receipt.json
Listing overflow and invocation on any host not claimed (host run pending) measurement plan in the skills-trial addendum
Cross-family review (GPT-6 through the OmniRoute gateway, read-only) review verdict line posted as a PR comment when it completes

Description-character sums (declared equals computed): 28 skills, 26 on, 2 user-invocable-only (grill-me, improve-codebase-architecture: upstream disable-model-invocation, replacement pending the skills sweep and paired benchmark), 27 Codex-enabled; Claude on-listed sum 10,191 of the 10,500 cap; Codex-enabled sum 10,048, of which 9,872 reach Codex's automatic catalog. Listing fraction: at about 4 characters per token, 0.01 of a 200k window (8,000 chars) overflows the 10,191 listing, 0.05 (40,000 chars) holds it about four times over; Codex's default 2% of the 272,000-token window is 5,440 tokens and the template sets 6,000 (the 25 automatically shown skills render to about 2,976 tokens).

Local commands run

$ python3 -m unittest tests.test_skills_manifest tests.test_runtime_worker_skills tests.test_skills_status tests.test_landscape_sweep_harness tests.test_skill_usage tests.test_adoption_docs_consistency tests.test_install_claude_profile
exit 0; Ran 529 tests; OK (skipped=18, all environmental: PyYAML, tree-sitter parser, bash 3.2, shellcheck, context-mode security.js, one release-note profile)
$ python3 scripts/validate.py
exit 0; {"components": 69, "hashed_files": 8530, "profiles": 4, "receipts": 176, "status": "passed"}
$ python3 -m unittest tests.test_osv_lockfile_coverage.LockfileInventoryTests.test_every_tracked_lockfile_and_manifest_is_listed tests.test_blind_checkout.RepositoryClassificationTests.test_every_blueprint_value_under_a_label_key_is_classified tests.test_workflow_security_coverage.NewWorkflowSecurityCoverageTests.test_all_published_workflows_are_listed_and_covered
exit 0; Ran 3 tests; OK
$ python3 scripts/validate_convergence.py --all-recorded --root . --json
exit 0; valid, 25 records
$ python3 -m unittest tests.test_install_skills
exit 0; Ran 65 tests; OK
$ python3 scripts/build_ecosystem.py --check && python3 scripts/landscape.py --root . && python3 scripts/validate_catalogs.py
exit 0 each
$ python3 scripts/validate.py --scan-file <each of the 28 changed files>
exit 0; {"scanned_files": 28, "status": "passed"}; host user name 0 hits; personal paths 0 hits

Decision record

docs/decisions/2026-09-30-skills-llm-native-listing.md (alternatives, overturn conditions, sources); the 2026-09-30 addendum in docs/decisions/2026-09-25-skills-trial-and-usage.md; the attributed docs/decisions/2026-09-30-native-skill-lifecycle.md.

Host evidence

Not applicable: no files under evidence/hosts/ change. Nothing was installed or applied on a host; the host-batch steps (install the re-pins and skill-creator, remove resolving-merge-conflicts, apply the settings, render the Codex config, run the listing measure-then-lower plan) are in the addendum and belong to B1 after the Gate A owner's go.

Checklist

  • New/changed GitHub Actions are pinned to a full commit SHA with a version comment (no workflow changes).
  • New/changed workflows declare top-level permissions: contents: read (no workflow changes).
  • No secrets are printed, logged or committed; no new required secret was added.
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved (the main checkout was only read, via the r2 snapshot).

🤖 Generated with Claude Code

@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 30, 2026
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Cross-family review (GPT-6 through the OmniRoute gateway, read-only, max effort) of 589c227a: verdict: needs_changes. Coordinator adjudication: (1) accepted, medium: the name-keyed [[skills.config]] rule emitted for the disabled Anthropic skill-creator also disables Codex's bundled .system/skill-creator (Codex's matcher applies a name rule to every skill of that name; SkillConfig has path), so the emitter moves to path-keyed exclusions and the status check accepts them, with regression coverage keeping the bundled skill; (2) rejected: MCP_TIMEOUT, the bootstrap.md line and the installer test are in scope by the coordinator's delta brief and the Gate A owner's explicit acceptance (recorded in the unit brief's amendments); (3) accepted, low: the reporter still treats the 8,000-character fallback as the active Codex cap (codex_within_cap: false against a configured 6,000-token budget); the manifest gains the configured budget and the reporter compares the rendered catalog in tokens; (4) accepted in part, low: the provenance line keeps the neutral wording the Codex runtime lane asked for, with one sentence explaining why the fold briefs' earlier attribution (a process census of file writes) does not establish authorship; (5) accepted, low: the Codex template comment's skill count is synchronised with the listing record. A repair round follows on this PR before the Gate A batch.

🤖 Generated with Claude Code

@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-f3-20260930 branch from 589c227 to 657ff9d Compare September 30, 2026 18:44
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Rebased onto origin/main 8fc8611 (after #546) with the hot-file protocol: main's manifests/evidence.json taken and this unit's files re-registered in the last commit; python3 scripts/validate.py passed; new head 657ff9df, tree otherwise unchanged.

🤖 Generated with Claude Code

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Review repair, round 2 (GPT-6 cross-family review of 589c227). Branch rebuilt on origin/main 1f2cdce; head d6041854. manifests/evidence.json changes only in the last commit.

M1 fixed: install_skills.py --print-codex-config now disables by path, never by name. A name rule applies to every loaded skill of that name (openai/codex rust-v0.159.2 codex-rs/config/src/skills_config.rs L109-119), so name = "skill-creator" also hid Codex's bundled .system/skill-creator. Each table now names the installed <home>/.agents/skills/<name>/SKILL.md: Codex matches that against each skill's canonical path_to_skills_md (ext/skills/src/host_service.rs L366-371, host_outcome.rs L52-54) and loads user skills from $HOME/.agents/skills (host_roots.rs L103-108); skills@7407f389 installs a global Codex skill only there (src/installer.ts L392-402), and the Codex skills page shows the same form. Note: the path is the SKILL.md file under $HOME/.agents/skills, not a folder under the Codex home as the review wording suggested; a folder path matches nothing. scripts/skills_status.py accepts the path form (resolving ~, a path relative to the config folder, and symlinks as Codex does) and fails a name-keyed table for any name Codex's bundled skills use (imagegen, openai-docs, review-agent, skill-creator, skill-installer) as name_entry_hides_bundled_skill. The project-scoped runtime manifest now prints only with --project-dir.

L1 fixed: budget.codex_default_budget_chars is now codex_fallback_budget_chars (metadata only). codex_configured_budget_tokens = 6000 is tied by a test to the template's [skills] max_context_tokens. codex_catalog_description_chars = 9,872 counts the 25 skills Codex shows; upstream_allow_implicit_invocation: false marks the two it keeps to an explicit $name. The reporter estimates the shown catalog the way render.rs charges it (each line plus newline at ceil(bytes/4); L154-160, L25, L258-267) and reports that estimate beside the configured budget: about 2,990 tokens against 6,000.

L3 fixed: the template comment says 25 catalog-visible skills, about 11,900 bytes, about 2,990 tokens, and five .system skills at rust-v0.159.2; a test ties the count to the manifest.

L2 fixed: all three records use "pre-existing uncommitted changes observed in the main checkout; original author not established" with the snapshot hash, plus one sentence that the earlier attribution rested on a process census of file writes, which does not establish authorship, and that the lane concerned asked for the neutral wording.

M2: no change, per the coordinator's adjudication (MCP_TIMEOUT is in scope and accepted).

Also: a Codex path containing a NUL no longer crashes the status check; the lifecycle guide and the host steps say to restart Codex after changing config.toml (Codex skills page L188).

Evidence (our integration checks, not upstream tests or a Codex run): python3 -m unittest tests.test_skills_manifest tests.test_runtime_worker_skills tests.test_skills_status tests.test_install_skills tests.test_skill_usage tests.test_adoption_docs_consistency tests.test_install_claude_profile 487 OK (15 env skips; 468 before the repair); install_skills.py --print-codex-config with an example home prints the path-keyed skill-creator table; validate.py passed (8558 files, 177 receipts); the three registry tests OK; privacy scan of the 7,449 added lines 0. Every new test failed first against the pre-repair code and data (installer 9/14, reporter 13 failures + 4 errors, manifest 2 + 2, template count 26 != 27). Residuals: the reporter checks rule presence, not rule order or full schema validity; the bundled-name list is the rust-v0.159.2 set; the token estimate counts description characters as bytes and excludes bundled and plugin skills; no native Codex run.

🤖 Generated with Claude Code

seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…tables, never name-keyed (PR #553 review M1)

Codex applies a name rule to every loaded skill of that name (openai/codex rust-v0.159.2
codex-rs/config/src/skills_config.rs L109-119), so the printed `name = "skill-creator"` table also
disabled Codex's bundled .system/skill-creator. Each table now selects the installed SKILL.md by
`path`, which Codex matches against each loaded skill's canonical path_to_skills_md
(ext/skills/src/host_service.rs L366-371, host_outcome.rs L52-54; the Codex skills page L180-186
documents `path = "/path/to/skill/SKILL.md"`). A global install leaves a Codex skill only in the
canonical ~/.agents/skills/<name> (vercel-labs/skills@7407f389 src/installer.ts L392-402), which
Codex loads from its $HOME/.agents/skills root (host_roots.rs L103-108), so the path is built with
canonical_skill_dir under --home. A project-scoped manifest now prints only with --project-dir,
since its tables name <project>/.agents/skills/<name>/SKILL.md. Paths are TOML-escaped; a disabled
name that is not one path component is refused.

Tests (failing first against the unchanged installer): PrintCodexConfigTests (skill-creator by
path, ".." and quote handling, project copy, refused name), the scope test and ReuseRefGateTests
in tests/test_install_skills.py, and the runtime-worker print test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
… and a token estimate of the Codex catalog (PR #553 review M1, L1, L3)

M1: a path selector matches when it resolves, as Codex resolves it (~ against the home, a relative
path against the config folder, lexical normalization, then canonicalization; utils/absolute-path
lib.rs L28-58, absolutize.rs L22-45, config/src/loader/mod.rs L573-582 and L1424-1447,
skills_config.rs L188-192), to the installed ~/.agents/skills/<name>/SKILL.md; without a known
path it matches by the skill folder name. Entries Codex ignores (both selectors, neither, a blank
name) select nothing, and a name is trimmed (skills_config.rs L188-210). A name-keyed disable for a
name Codex's bundled skills carry (imagegen, openai-docs, review-agent, skill-creator,
skill-installer at rust-v0.159.2) is reported as name_entry_hides_bundled_skill.

L1: budget.codex_default_budget_chars becomes codex_fallback_budget_chars (metadata);
codex_configured_budget_tokens (6000) equals the Codex template's [skills] max_context_tokens by
test; codex_catalog_description_chars (9,872) counts the skills Codex shows the model, and
upstream_allow_implicit_invocation: false marks grill-me and improve-codebase-architecture (their
agents/openai.yaml; provider/host.rs L147-148, model.rs L22-28). The reporter estimates the shown
catalog as render.rs charges it, each line and its newline at ceil(bytes / 4) (L154-160, L25,
L258-267, L1158-1174), and reports it beside the configured budget instead of comparing 10,048
characters with the 8,000-character fallback.

L3: the Codex template comment now says 25 catalog-visible skills, about 11,900 bytes and about
2,990 tokens by that charge, and five .system skills at rust-v0.159.2; a test ties the count to the
manifest.

Tests (failing first against the unchanged reporter and data): the path-selector,
bundled-name, ignored-entry and catalog-estimate tests in tests/test_skills_status.py and the
budget and template-count tests in tests/test_skills_manifest.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…atalog budget fields and neutral fold provenance (PR #553 review M1, L1, L2, L3)

adoption/skills/lifecycle.md documents the path form (example table, upstream matching rules,
--project-dir for project manifests, the reporter's bundled-name failure) and the new budget
fields. The skills-trial addendum and the LLM-native listing record replace the name-collision
residual with the path decision, rewrite host step 3, record the budget rename and the recomputed
estimate (11,904 bytes, 2,987 tokens with a 28-character root), and cite the rust-v0.159.2 and
skills@7407f389 source lines read for this repair.

L2: the listing record and the addendum now use the exact provenance wording "pre-existing
uncommitted changes observed in the main checkout; original author not established" with the
snapshot hash, and all three records say that the fold briefs' earlier attribution rested on a
process census of file writes, which does not establish document authorship, and that the lane
concerned asked for the neutral wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…alog defaults (PR #553 review L3)

Main now renders the template's model from the Codex pin (gpt-6.1-sol from 0.159.1, gpt-6-astra
before it), so the unset-budget figure names both models' 272,000-token context_window
(codex-rs/models-manager/models.json at rust-v0.159.2) instead of gpt-6-astra alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-f3-20260930 branch from 657ff9d to d604185 Compare September 30, 2026 20:20
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…tables, never name-keyed (PR #553 review M1)

Codex applies a name rule to every loaded skill of that name (openai/codex rust-v0.159.2
codex-rs/config/src/skills_config.rs L109-119), so the printed `name = "skill-creator"` table also
disabled Codex's bundled .system/skill-creator. Each table now selects the installed SKILL.md by
`path`, which Codex matches against each loaded skill's canonical path_to_skills_md
(ext/skills/src/host_service.rs L366-371, host_outcome.rs L52-54; the Codex skills page L180-186
documents `path = "/path/to/skill/SKILL.md"`). A global install leaves a Codex skill only in the
canonical ~/.agents/skills/<name> (vercel-labs/skills@7407f389 src/installer.ts L392-402), which
Codex loads from its $HOME/.agents/skills root (host_roots.rs L103-108), so the path is built with
canonical_skill_dir under --home. A project-scoped manifest now prints only with --project-dir,
since its tables name <project>/.agents/skills/<name>/SKILL.md. Paths are TOML-escaped; a disabled
name that is not one path component is refused.

Tests (failing first against the unchanged installer): PrintCodexConfigTests (skill-creator by
path, ".." and quote handling, project copy, refused name), the scope test and ReuseRefGateTests
in tests/test_install_skills.py, and the runtime-worker print test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-f3-20260930 branch from d604185 to 1b95c04 Compare September 30, 2026 23:10
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
… and a token estimate of the Codex catalog (PR #553 review M1, L1, L3)

M1: a path selector matches when it resolves, as Codex resolves it (~ against the home, a relative
path against the config folder, lexical normalization, then canonicalization; utils/absolute-path
lib.rs L28-58, absolutize.rs L22-45, config/src/loader/mod.rs L573-582 and L1424-1447,
skills_config.rs L188-192), to the installed ~/.agents/skills/<name>/SKILL.md; without a known
path it matches by the skill folder name. Entries Codex ignores (both selectors, neither, a blank
name) select nothing, and a name is trimmed (skills_config.rs L188-210). A name-keyed disable for a
name Codex's bundled skills carry (imagegen, openai-docs, review-agent, skill-creator,
skill-installer at rust-v0.159.2) is reported as name_entry_hides_bundled_skill.

L1: budget.codex_default_budget_chars becomes codex_fallback_budget_chars (metadata);
codex_configured_budget_tokens (6000) equals the Codex template's [skills] max_context_tokens by
test; codex_catalog_description_chars (9,872) counts the skills Codex shows the model, and
upstream_allow_implicit_invocation: false marks grill-me and improve-codebase-architecture (their
agents/openai.yaml; provider/host.rs L147-148, model.rs L22-28). The reporter estimates the shown
catalog as render.rs charges it, each line and its newline at ceil(bytes / 4) (L154-160, L25,
L258-267, L1158-1174), and reports it beside the configured budget instead of comparing 10,048
characters with the 8,000-character fallback.

L3: the Codex template comment now says 25 catalog-visible skills, about 11,900 bytes and about
2,990 tokens by that charge, and five .system skills at rust-v0.159.2; a test ties the count to the
manifest.

Tests (failing first against the unchanged reporter and data): the path-selector,
bundled-name, ignored-entry and catalog-estimate tests in tests/test_skills_status.py and the
budget and template-count tests in tests/test_skills_manifest.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…atalog budget fields and neutral fold provenance (PR #553 review M1, L1, L2, L3)

adoption/skills/lifecycle.md documents the path form (example table, upstream matching rules,
--project-dir for project manifests, the reporter's bundled-name failure) and the new budget
fields. The skills-trial addendum and the LLM-native listing record replace the name-collision
residual with the path decision, rewrite host step 3, record the budget rename and the recomputed
estimate (11,904 bytes, 2,987 tokens with a 28-character root), and cite the rust-v0.159.2 and
skills@7407f389 source lines read for this repair.

L2: the listing record and the addendum now use the exact provenance wording "pre-existing
uncommitted changes observed in the main checkout; original author not established" with the
snapshot hash, and all three records say that the fold briefs' earlier attribution rested on a
process census of file writes, which does not establish document authorship, and that the lane
concerned asked for the neutral wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…alog defaults (PR #553 review L3)

Main now renders the template's model from the Codex pin (gpt-6.1-sol from 0.159.1, gpt-6-astra
before it), so the unset-budget figure names both models' 272,000-token context_window
(codex-rs/models-manager/models.json at rust-v0.159.2) instead of gpt-6-astra alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…llings in a session (PR #553 Gate A review, item 2)

The pinned find-skills body tells the model to run `npx skills add <owner/repo@skill> -g -y` and `npx skills
update` (vercel-labs/skills@7407f389 skills/find-skills/SKILL.md L28-29, L90, L100), and the template's
bypassPermissions default denied neither. Installation goes only through tools/adoption/install_skills.py,
whose own add and rollback remove run as subprocesses that Bash rules do not see.

41 Bash rules cover every writing spelling of skills@1.7.0 (src/cli.ts L336-402: add/a/i/install,
remove/rm/r, check/update/upgrade, experimental_*) in four invocation forms: bare `skills`, `npx [flags] skills`,
`npx [flags] skills@<version>` (without the one-letter aliases, which are common find-query words) and a
path ending in `bin/skills` (the pinned <tools-root>/skills-1.7.0/bin/skills, node_modules/.bin/skills).
check/update/upgrade end in `<word>*` where another `*` precedes, so the bare command matches too.
`Edit(~/.agents/**)` keeps the file tools off the canonical skill folders the installer hash-checks.
Discovery (`find`/`search`), `list`, `init`, `use` and `--version` stay allowed.

Rule form: https://code.claude.com/docs/en/permissions ("Wildcard patterns", "Compound commands",
"Wrappers", "What a Bash rule doesn't match", "Read and Edit", "Symlinks"; fetched 2026-09-30).
The new test failed first against the previous template (66 failures: 42 missing rules, 24 samples
not denied).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…ide one, and the listing wording (PR #553 Gate A review, items 2 and 5)

- Step 3 and Install and inspect: a session never installs a skill ad hoc; it records the candidate and the
  coordinator pins it. The Claude template denies the Skills CLI's writing commands in four forms and
  Edit(~/.agents/**); the rules match command text, not every route to the program (permissions page,
  "What a Bash rule doesn't match").
- Retire and remove: the host batch runs the CLI's remove from a shell outside a Claude session.
- Step 1 lists every model-invocable skill within the listing budgets, not every enabled skill: the two
  upstream user-only skills are not listed, and an overflowing listing loses descriptions.
- The absent docs/decisions/2026-09-30-sota-native-finalization.md citation now points at the committed
  evaluation artifact; the finalization record lands with unit F1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…it and names the skills deny rules (PR #553 Gate A review, item 3)

docs/decisions/2026-09-30-mcp-startup-timeout.md records MCP_TIMEOUT "120000": the alternatives (Claude
Code's 30 s default, env-vars L470; no per-server startup option, `claude mcp add --help` on 2.1.286 read
2026-09-30 with a scratch HOME, the F3 delta for 2.1.285; MCP_CONNECT_TIMEOUT_MS and
CLAUDE_CODE_MCP_STARTUP_WAIT_MS are different waits, L460 and L317), the cost (a `-p` run with
--mcp-config waits up to MCP_TIMEOUT for pending servers, headless L240) and the overturn conditions
(a measured start-up distribution under 30 s for every stack server on the reference host, or an
upstream per-server option). The step-4 note and McpStartupTimeoutTemplateTests' docstring link it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…553 Gate A review, item 6)

evidence/artifacts/skills-pin-verification-20260930/pin-verification.json keeps the checker's 28 rows
(codebase-design, improve-codebase-architecture, semgrep and skill-creator included) from the 18:02Z run
against blobless upstream clones, with the method, the checked manifest blob and a 22:53Z recheck after
a fresh fetch (every remote HEAD confirmed by git ls-remote) whose output was byte-identical. Our own
integration check of upstream git objects, not an upstream test or a client run. Registered in
manifests/evidence.json by the last commit (hot-file protocol).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…kills deny rules and the retained pin verification (PR #553 Gate A review, items 1, 2, 4 and 6)

- Listing record Scope: a "#381 interaction" paragraph. Codex arms B and A go from 11 to 25 visible skills
  (14 new, agent-browser, semgrep, codeql, security-audit and find-skills among them); arm N ignores the
  user configuration and, by the Gate A owner's count, already saw 26. M4 (>= 90% routed fetches; 5 of the
  fetch lane's 10 arm-B opportunities are seed-codex-web-table-1..5) and M8 (>= 80% per required lane;
  2 of symbol-references' 9, 1 of ai-memory's 7) could be hit by a diverted call that the B-versus-A gates
  cannot see. Counts from the #381 preregistration and the two manifests.
- Decision point 9 states the Gate A owner's decision as theirs: Amendment 4 names the catalog as part of
  the measured practice, no threshold is loosened, and diverted calls are reported, never reclassified.
- Decision point 4 records the Skills CLI deny rules; point 8 cites the retained pin verification output;
  a new overturn condition covers CLI command or rule-matching changes; Sources add the permissions page,
  skills@7407f389 src/cli.ts L336-402 and the find-skills body.
- Skills-trial addendum: the fraction plan's step 3 stays outside #381 window W, and a change after the
  Amendment 4 seal re-seals; host step 1 removes the retired skill from a shell outside a Claude session;
  two evidence-class rows for the deny-rule test and the pin verification.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Review repair, round 3 (the Gate A owner's review of d604185, ACCEPT WITH CHANGES). Branch rebased onto origin/main 29458b4 without conflicts; head 1b95c047. manifests/evidence.json changes only in the last commit.

1. #381 interaction recorded. The listing record's Scope has a "#381 interaction" paragraph, with counts read from the #381 preregistration (thresholds, arm_b_opportunity_counts, each task's family/arms/lane_tags) and the two manifests:

  • Codex arms B and A go from 11 to 25 visible skills; 14 are new, including agent-browser, semgrep, codeql, security-audit and find-skills. Arm N runs with --ignore-user-config and, by the Gate A owner's count, already saw 26.
  • A call diverted to a newly visible skill could miss M4 (at least 90% routed fetches; 5 of the fetch lane's 10 arm-B opportunities are seed-codex-web-table-1..5) or M8 (at least 80% correct lane use per required lane; 2 of symbol-references' 9 and 1 of ai-memory's 7 are Codex tasks). The B-versus-A gates cannot see it, because A changes the same way.
  • The applied settings are frozen values too (claude.settings.user.sha256; permissions_deny_count goes from 88 to 130).
  • Decision point 9 states the Gate A owner's decision as theirs: Amendment 4 names the catalog as part of the measured practice, no threshold is loosened, and diverted calls are reported and never reclassified.

2. No ad hoc installs from a session. The settings template now denies every Skills CLI command that writes installed skills (skills@7407f389 src/cli.ts L336-402: add, a, i, install, remove, rm, r, check, update, upgrade, experimental_*).

  • The rules cover four forms: a bare skills, npx [flags] skills, npx [flags] skills@<version> and a path ending in bin/skills (the pinned <tools-root>/skills-1.7.0/bin/skills). That is 41 Bash(...) rules, plus Edit(~/.agents/**).
  • The rules follow https://code.claude.com/docs/en/permissions: "Wildcard patterns", "Compound commands", "Wrappers" (npx is not stripped; a deny rule matches past a leading assignment) and "What a Bash rule doesn't match".
  • find, list, init, use and --version stay allowed. install_skills.py's own add and rollback remove run as subprocesses, which the rules do not see.
  • The rule set goes beyond add/remove to cover their aliases and check/update/upgrade: find-skills L29 advertises npx skills update, and at v1.7.0 check runs the same update code.
  • test_a_session_cannot_install_or_remove_skills_through_the_skills_cli checks the rules exist, that blocked samples match (find-skills' own npx skills add <owner/repo@skill> -g -y, versioned npx, the pinned path, DISABLE_TELEMETRY=1 skills remove …, a compound command), and that allowed samples (skills find, list, the installer, a commit message, a grep) do not. It failed first against the previous template: 71 failures (42 rules, 29 samples).
  • lifecycle.md records the rule: a session never installs a skill ad hoc; it records the candidate and the coordinator pins it. Removal now runs from a shell outside a session.

3. MCP_TIMEOUT decision record. docs/decisions/2026-09-30-mcp-startup-timeout.md records:

  • the alternatives: Claude Code's 30 s default (env-vars L470); no per-server startup option (claude mcp add --help on 2.1.286 lists none, as on 2.1.285); MCP_CONNECT_TIMEOUT_MS and CLAUDE_CODE_MCP_STARTUP_WAIT_MS bound different waits;
  • the cost: a -p run with --mcp-config waits up to MCP_TIMEOUT for pending servers (headless L240);
  • the overturn conditions: a measured start-up distribution under 30 s for every stack server on the reference host, or an upstream per-server option.

4. Fraction plan. Step 3 now reads: not inside #381 window W; a change after the Amendment 4 seal re-seals.

5. lifecycle.md. The absent finalization record's citation now points at evidence/artifacts/native-skill-finalization-20260930/README.md, and says the record lands with unit F1. Step 1 lists the model-invocable skills within the listing budgets: the two upstream user-only skills are not listed, and an overflowing listing loses descriptions.

6. Pin verification retained. evidence/artifacts/skills-pin-verification-20260930/pin-verification.json keeps the checker's 28 rows (codebase-design, improve-codebase-architecture, semgrep and skill-creator included) from the 18:02Z run, with the method and the manifest blob. A 22:53Z recheck after a fresh fetch, with every remote HEAD confirmed by git ls-remote, returned byte-identical output.

Also: a test fixture path /home/example/… became /srv/example-home/… so the privacy pattern stays clean.

Evidence (our integration checks, not upstream tests or a client run): the unittest set passed 490 (15 skipped); validate.py passed (8604 files, 177 receipts); the three registry tests passed; the F-tier invariant check passed 18 of 18; the privacy scan found 0 matches in 7,610 added lines.

Residuals:

  • The rules match command text: "$SKILLS_BIN" add, sh -c, other runners and file copies are not matched.
  • Context Mode (6f0cc684) applies the same rules server-side, but without bare-command matching or leading-assignment stripping.
  • npx skills@… find is blocked when the query contains a writer word.
  • No Claude Code run exercised the rules.
  • Applying the template changes frozen PR-H: E2E preregistration for the token-save-practice stack (frozen tasks, metrics, gate, procedure) #381 values (permissions_deny_count 88 → 130, claude.settings.user.sha256): it lands before the seal or the seal is re-taken (the Gate A owner's call).

🤖 Generated with Claude Code

seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…tables, never name-keyed (PR #553 review M1)

Codex applies a name rule to every loaded skill of that name (openai/codex rust-v0.159.2
codex-rs/config/src/skills_config.rs L109-119), so the printed `name = "skill-creator"` table also
disabled Codex's bundled .system/skill-creator. Each table now selects the installed SKILL.md by
`path`, which Codex matches against each loaded skill's canonical path_to_skills_md
(ext/skills/src/host_service.rs L366-371, host_outcome.rs L52-54; the Codex skills page L180-186
documents `path = "/path/to/skill/SKILL.md"`). A global install leaves a Codex skill only in the
canonical ~/.agents/skills/<name> (vercel-labs/skills@7407f389 src/installer.ts L392-402), which
Codex loads from its $HOME/.agents/skills root (host_roots.rs L103-108), so the path is built with
canonical_skill_dir under --home. A project-scoped manifest now prints only with --project-dir,
since its tables name <project>/.agents/skills/<name>/SKILL.md. Paths are TOML-escaped; a disabled
name that is not one path component is refused.

Tests (failing first against the unchanged installer): PrintCodexConfigTests (skill-creator by
path, ".." and quote handling, project copy, refused name), the scope test and ReuseRefGateTests
in tests/test_install_skills.py, and the runtime-worker print test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
… and a token estimate of the Codex catalog (PR #553 review M1, L1, L3)

M1: a path selector matches when it resolves, as Codex resolves it (~ against the home, a relative
path against the config folder, lexical normalization, then canonicalization; utils/absolute-path
lib.rs L28-58, absolutize.rs L22-45, config/src/loader/mod.rs L573-582 and L1424-1447,
skills_config.rs L188-192), to the installed ~/.agents/skills/<name>/SKILL.md; without a known
path it matches by the skill folder name. Entries Codex ignores (both selectors, neither, a blank
name) select nothing, and a name is trimmed (skills_config.rs L188-210). A name-keyed disable for a
name Codex's bundled skills carry (imagegen, openai-docs, review-agent, skill-creator,
skill-installer at rust-v0.159.2) is reported as name_entry_hides_bundled_skill.

L1: budget.codex_default_budget_chars becomes codex_fallback_budget_chars (metadata);
codex_configured_budget_tokens (6000) equals the Codex template's [skills] max_context_tokens by
test; codex_catalog_description_chars (9,872) counts the skills Codex shows the model, and
upstream_allow_implicit_invocation: false marks grill-me and improve-codebase-architecture (their
agents/openai.yaml; provider/host.rs L147-148, model.rs L22-28). The reporter estimates the shown
catalog as render.rs charges it, each line and its newline at ceil(bytes / 4) (L154-160, L25,
L258-267, L1158-1174), and reports it beside the configured budget instead of comparing 10,048
characters with the 8,000-character fallback.

L3: the Codex template comment now says 25 catalog-visible skills, about 11,900 bytes and about
2,990 tokens by that charge, and five .system skills at rust-v0.159.2; a test ties the count to the
manifest.

Tests (failing first against the unchanged reporter and data): the path-selector,
bundled-name, ignored-entry and catalog-estimate tests in tests/test_skills_status.py and the
budget and template-count tests in tests/test_skills_manifest.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…atalog budget fields and neutral fold provenance (PR #553 review M1, L1, L2, L3)

adoption/skills/lifecycle.md documents the path form (example table, upstream matching rules,
--project-dir for project manifests, the reporter's bundled-name failure) and the new budget
fields. The skills-trial addendum and the LLM-native listing record replace the name-collision
residual with the path decision, rewrite host step 3, record the budget rename and the recomputed
estimate (11,904 bytes, 2,987 tokens with a 28-character root), and cite the rust-v0.159.2 and
skills@7407f389 source lines read for this repair.

L2: the listing record and the addendum now use the exact provenance wording "pre-existing
uncommitted changes observed in the main checkout; original author not established" with the
snapshot hash, and all three records say that the fold briefs' earlier attribution rested on a
process census of file writes, which does not establish document authorship, and that the lane
concerned asked for the neutral wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scout and others added 16 commits September 30, 2026 21:41
… today's state

- 2026-09-30 LLM-native listing record: rebuilt on 11227bf; the user's
  restated directive quoted; Decision point 8 records the folded source
  review (six re-pins verified from blobless clones, the retirement with its
  "retired" marker and explicit off, the lifecycle guide as the single
  lifecycle document and the discovery order: listed skills, search-first,
  find-skills, the landscape sweep). Figures: 26 on (25 plus skill-creator),
  27 Codex-enabled (16 newly), Claude on 10191, Codex 10048 (9872 visible),
  5% of a 200k window about 3.9 times the on-listed sum; the Codex template's
  6000 tokens stays above the 5440-token default of gpt-6-astra and
  gpt-6.1-sol (272000-token windows at rust-v0.159.2); 25 visible skills
  render to about 11,900 bytes (about 2,980 tokens). New overturn condition
  for upstream drift or removal; sources for the clones, rust-v0.159.2
  lines, Skills CLI v1.7.0 lines and the docs pages with their hashes.
- Skills-trial addendum 2026-09-30: the same figures; a same-day source
  review paragraph (including the reviewed semgrep script change: explicit
  --max-target-bytes and an oversized report, metrics still off, no new
  network access); window effects for 16 newly enabled skills and
  search-first's new description; host steps remove the retired skill and
  its Codex table; evidence classes for the retirement tests, the
  blobless-clone pin checks and the lifecycle guide's commands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dict unchanged

Upstream affaan-m/ECC c70874fa rewrote one line of skills/search-first/SKILL.md,
its description. No catalogs/landscape record pins the kept winner's commit
(grep for the old 2b6e8397 ref and for search-first across
catalogs/landscape/*.json: the records name the winner only), so the kept
status and the verdict stay as recorded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tables, never name-keyed (PR #553 review M1)

Codex applies a name rule to every loaded skill of that name (openai/codex rust-v0.159.2
codex-rs/config/src/skills_config.rs L109-119), so the printed `name = "skill-creator"` table also
disabled Codex's bundled .system/skill-creator. Each table now selects the installed SKILL.md by
`path`, which Codex matches against each loaded skill's canonical path_to_skills_md
(ext/skills/src/host_service.rs L366-371, host_outcome.rs L52-54; the Codex skills page L180-186
documents `path = "/path/to/skill/SKILL.md"`). A global install leaves a Codex skill only in the
canonical ~/.agents/skills/<name> (vercel-labs/skills@7407f389 src/installer.ts L392-402), which
Codex loads from its $HOME/.agents/skills root (host_roots.rs L103-108), so the path is built with
canonical_skill_dir under --home. A project-scoped manifest now prints only with --project-dir,
since its tables name <project>/.agents/skills/<name>/SKILL.md. Paths are TOML-escaped; a disabled
name that is not one path component is refused.

Tests (failing first against the unchanged installer): PrintCodexConfigTests (skill-creator by
path, ".." and quote handling, project copy, refused name), the scope test and ReuseRefGateTests
in tests/test_install_skills.py, and the runtime-worker print test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and a token estimate of the Codex catalog (PR #553 review M1, L1, L3)

M1: a path selector matches when it resolves, as Codex resolves it (~ against the home, a relative
path against the config folder, lexical normalization, then canonicalization; utils/absolute-path
lib.rs L28-58, absolutize.rs L22-45, config/src/loader/mod.rs L573-582 and L1424-1447,
skills_config.rs L188-192), to the installed ~/.agents/skills/<name>/SKILL.md; without a known
path it matches by the skill folder name. Entries Codex ignores (both selectors, neither, a blank
name) select nothing, and a name is trimmed (skills_config.rs L188-210). A name-keyed disable for a
name Codex's bundled skills carry (imagegen, openai-docs, review-agent, skill-creator,
skill-installer at rust-v0.159.2) is reported as name_entry_hides_bundled_skill.

L1: budget.codex_default_budget_chars becomes codex_fallback_budget_chars (metadata);
codex_configured_budget_tokens (6000) equals the Codex template's [skills] max_context_tokens by
test; codex_catalog_description_chars (9,872) counts the skills Codex shows the model, and
upstream_allow_implicit_invocation: false marks grill-me and improve-codebase-architecture (their
agents/openai.yaml; provider/host.rs L147-148, model.rs L22-28). The reporter estimates the shown
catalog as render.rs charges it, each line and its newline at ceil(bytes / 4) (L154-160, L25,
L258-267, L1158-1174), and reports it beside the configured budget instead of comparing 10,048
characters with the 8,000-character fallback.

L3: the Codex template comment now says 25 catalog-visible skills, about 11,900 bytes and about
2,990 tokens by that charge, and five .system skills at rust-v0.159.2; a test ties the count to the
manifest.

Tests (failing first against the unchanged reporter and data): the path-selector,
bundled-name, ignored-entry and catalog-estimate tests in tests/test_skills_status.py and the
budget and template-count tests in tests/test_skills_manifest.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ing instead of crashing the check

TOML allows an escaped NUL in a string, and os.path.realpath raises ValueError on it, which would
have ended the whole status run on one malformed [[skills.config]] entry. Such a path cannot name
the installed SKILL.md, so it is treated as no match. Test (failing first with the ValueError):
test_a_path_no_file_can_have_selects_nothing_instead_of_crashing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…atalog budget fields and neutral fold provenance (PR #553 review M1, L1, L2, L3)

adoption/skills/lifecycle.md documents the path form (example table, upstream matching rules,
--project-dir for project manifests, the reporter's bundled-name failure) and the new budget
fields. The skills-trial addendum and the LLM-native listing record replace the name-collision
residual with the path decision, rewrite host step 3, record the budget rename and the recomputed
estimate (11,904 bytes, 2,987 tokens with a 28-character root), and cite the rust-v0.159.2 and
skills@7407f389 source lines read for this repair.

L2: the listing record and the addendum now use the exact provenance wording "pre-existing
uncommitted changes observed in the main checkout; original author not established" with the
snapshot hash, and all three records say that the fold briefs' earlier attribution rested on a
process census of file writes, which does not establish document authorship, and that the lane
concerned asked for the neutral wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ame, not by shifting line numbers

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…alog defaults (PR #553 review L3)

Main now renders the template's model from the Codex pin (gpt-6.1-sol from 0.159.1, gpt-6-astra
before it), so the unset-budget figure names both models' 272,000-token context_window
(codex-rs/models-manager/models.json at rust-v0.159.2) instead of gpt-6-astra alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
….toml change; base reference 1f2cdce

The Codex skills page (developers.openai.com/codex/skills, "Enable or disable local Codex
skills", read 2026-09-30, page sha256 d1579156...) says "Restart Codex after changing
~/.codex/config.toml" right after the [[skills.config]] example, which the guide's Codex bullet
and the addendum's host step 3 now say too. tools/adoption/apply_codex_lane.py still writes no
[skills] key at origin/main 1f2cdce, the branch's new base.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…llings in a session (PR #553 Gate A review, item 2)

The pinned find-skills body tells the model to run `npx skills add <owner/repo@skill> -g -y` and `npx skills
update` (vercel-labs/skills@7407f389 skills/find-skills/SKILL.md L28-29, L90, L100), and the template's
bypassPermissions default denied neither. Installation goes only through tools/adoption/install_skills.py,
whose own add and rollback remove run as subprocesses that Bash rules do not see.

41 Bash rules cover every writing spelling of skills@1.7.0 (src/cli.ts L336-402: add/a/i/install,
remove/rm/r, check/update/upgrade, experimental_*) in four invocation forms: bare `skills`, `npx [flags] skills`,
`npx [flags] skills@<version>` (without the one-letter aliases, which are common find-query words) and a
path ending in `bin/skills` (the pinned <tools-root>/skills-1.7.0/bin/skills, node_modules/.bin/skills).
check/update/upgrade end in `<word>*` where another `*` precedes, so the bare command matches too.
`Edit(~/.agents/**)` keeps the file tools off the canonical skill folders the installer hash-checks.
Discovery (`find`/`search`), `list`, `init`, `use` and `--version` stay allowed.

Rule form: https://code.claude.com/docs/en/permissions ("Wildcard patterns", "Compound commands",
"Wrappers", "What a Bash rule doesn't match", "Read and Edit", "Symlinks"; fetched 2026-09-30).
The new test failed first against the previous template (66 failures: 42 missing rules, 24 samples
not denied).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ide one, and the listing wording (PR #553 Gate A review, items 2 and 5)

- Step 3 and Install and inspect: a session never installs a skill ad hoc; it records the candidate and the
  coordinator pins it. The Claude template denies the Skills CLI's writing commands in four forms and
  Edit(~/.agents/**); the rules match command text, not every route to the program (permissions page,
  "What a Bash rule doesn't match").
- Retire and remove: the host batch runs the CLI's remove from a shell outside a Claude session.
- Step 1 lists every model-invocable skill within the listing budgets, not every enabled skill: the two
  upstream user-only skills are not listed, and an overflowing listing loses descriptions.
- The absent docs/decisions/2026-09-30-sota-native-finalization.md citation now points at the committed
  evaluation artifact; the finalization record lands with unit F1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…it and names the skills deny rules (PR #553 Gate A review, item 3)

docs/decisions/2026-09-30-mcp-startup-timeout.md records MCP_TIMEOUT "120000": the alternatives (Claude
Code's 30 s default, env-vars L470; no per-server startup option, `claude mcp add --help` on 2.1.286 read
2026-09-30 with a scratch HOME, the F3 delta for 2.1.285; MCP_CONNECT_TIMEOUT_MS and
CLAUDE_CODE_MCP_STARTUP_WAIT_MS are different waits, L460 and L317), the cost (a `-p` run with
--mcp-config waits up to MCP_TIMEOUT for pending servers, headless L240) and the overturn conditions
(a measured start-up distribution under 30 s for every stack server on the reference host, or an
upstream per-server option). The step-4 note and McpStartupTimeoutTemplateTests' docstring link it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…553 Gate A review, item 6)

evidence/artifacts/skills-pin-verification-20260930/pin-verification.json keeps the checker's 28 rows
(codebase-design, improve-codebase-architecture, semgrep and skill-creator included) from the 18:02Z run
against blobless upstream clones, with the method, the checked manifest blob and a 22:53Z recheck after
a fresh fetch (every remote HEAD confirmed by git ls-remote) whose output was byte-identical. Our own
integration check of upstream git objects, not an upstream test or a client run. Registered in
manifests/evidence.json by the last commit (hot-file protocol).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…kills deny rules and the retained pin verification (PR #553 Gate A review, items 1, 2, 4 and 6)

- Listing record Scope: a "#381 interaction" paragraph. Codex arms B and A go from 11 to 25 visible skills
  (14 new, agent-browser, semgrep, codeql, security-audit and find-skills among them); arm N ignores the
  user configuration and, by the Gate A owner's count, already saw 26. M4 (>= 90% routed fetches; 5 of the
  fetch lane's 10 arm-B opportunities are seed-codex-web-table-1..5) and M8 (>= 80% per required lane;
  2 of symbol-references' 9, 1 of ai-memory's 7) could be hit by a diverted call that the B-versus-A gates
  cannot see. Counts from the #381 preregistration and the two manifests.
- Decision point 9 states the Gate A owner's decision as theirs: Amendment 4 names the catalog as part of
  the measured practice, no threshold is loosened, and diverted calls are reported, never reclassified.
- Decision point 4 records the Skills CLI deny rules; point 8 cites the retained pin verification output;
  a new overturn condition covers CLI command or rule-matching changes; Sources add the permissions page,
  skills@7407f389 src/cli.ts L336-402 and the find-skills body.
- Skills-trial addendum: the fraction plan's step 3 stays outside #381 window W, and a change after the
  Amendment 4 seal re-seals; host step 1 removes the retired skill from a shell outside a Claude session;
  two evidence-class rows for the deny-rule test and the pin verification.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…atch

The fixture paths /home/example/... match the coordinator's committed-diff privacy pattern
(/home/[a-z]+/), although they name no host. /srv/example-home/... keeps the test's meaning: with no
known home, a path entry matches by its skill folder name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ntext Mode's own command path denies the bare command too

The Gate A owner's decision on #553 round 3: Context Mode (mksglu/context-mode 1.0.169, src/security.ts
evaluateCommandDenyOnly / matchesAnyPattern / globToRegex) applies the same user deny rules to ctx_execute and
ctx_batch_execute commands with a plain ^glob$ regex, so `Bash(skills update *)` let a bare `skills update` through
there. `Bash(skills check*)`, `Bash(skills update*)` and `Bash(skills upgrade*)` close that path; the short aliases
keep " *" because `skills init` is legitimate. A new test runs the installed security module (CONTEXT_MODE_SECURITY_JS,
skipped with a named reason when unset) against the template's rules: failing-first on the previous template, three
bare commands allowed; the existing rule test also failed first on the three renamed rules.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-f3-20260930 branch from 9dae8c4 to c422994 Compare October 1, 2026 01:42
@seathatflowsinourveins
seathatflowsinourveins merged commit 3361b34 into main Oct 1, 2026
25 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/sota-defaults-f3-20260930 branch October 1, 2026 03:09
seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…or named, the retired skill dropped

Unit F3 (#553), below this PR in the stack, pins skill-creator (anthropics/skills@8a1541c4, trial,
codex_enabled false), retires resolving-merge-conflicts and widens the Claude listings and Codex
enablement. Built against main's manifest, A1 failed 35 tests on the stack.

- build_args.py: TEMPLATE_SKILLS gains skill-creator, which the skills templates already name for its
  paired with-skill/without-skill benchmark; templates.json common's Skills paragraph names it for
  Claude Code only (Codex ships a different skill of that name), read and never run. Both prompts_sha256
  pins move (repository 9c34fa72...f14b -> b61956f3...726d, skills 2c2efbae...63d3 -> a76ee858...b460).
- catalogs/landscape/skills-lifecycle.json: skill-creator serves skills-skill-lifecycle;
  resolving-merge-conflicts leaves skills-ci-pr, retired by unit F3 (#553), whose open gap now says no
  installed skill resolves rebase conflicts; the open gaps that stated pre-F3 codex_enabled and listing
  values are removed or restated, and skills-agent-docs' requirement points skill authoring at
  skill-creator.
- tests: a catalog test holds each flag an open gap states to the manifest (12 stale statements fail it
  on the previous catalog); the whole-gap input test bounds security-audit's gap by the 600-character cut,
  since F3 shortened it to 930 characters.
- README and decision record: the repository run's prompts_sha256 moved once; the 28-skill path check
  is scoped to the manifest before F3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…or named, the retired skill dropped

Unit F3 (#553), below this PR in the stack, pins skill-creator (anthropics/skills@8a1541c4, trial,
codex_enabled false), retires resolving-merge-conflicts and widens the Claude listings and Codex
enablement. Built against main's manifest, A1 failed 35 tests on the stack.

- build_args.py: TEMPLATE_SKILLS gains skill-creator, which the skills templates already name for its
  paired with-skill/without-skill benchmark; templates.json common's Skills paragraph names it for
  Claude Code only (Codex ships a different skill of that name), read and never run. Both prompts_sha256
  pins move (repository 9c34fa72...f14b -> b61956f3...726d, skills 2c2efbae...63d3 -> a76ee858...b460).
- catalogs/landscape/skills-lifecycle.json: skill-creator serves skills-skill-lifecycle;
  resolving-merge-conflicts leaves skills-ci-pr, retired by unit F3 (#553), whose open gap now says no
  installed skill resolves rebase conflicts; the open gaps that stated pre-F3 codex_enabled and listing
  values are removed or restated, and skills-agent-docs' requirement points skill authoring at
  skill-creator.
- tests: a catalog test holds each flag an open gap states to the manifest (12 stale statements fail it
  on the previous catalog); the whole-gap input test bounds security-audit's gap by the 600-character cut,
  since F3 shortened it to 930 characters.
- README and decision record: the repository run's prompts_sha256 moved once; the 28-skill path check
  is scoped to the manifest before F3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Oct 1, 2026
…edger support and pinned-CLI source reviews (unit A1) (#541)

* Skills lifecycle catalog: 13 lifecycle tasks and 23 pinned skill sources

catalogs/landscape/skills-lifecycle.json keys each lifecycle task as a
skills-<task> layer for the landscape sweep's skills modality: its
requirement (from the adoption/skills/manifest.json gap fields), the
installed skills serving it (all 28 pinned skills, each once), its
source_ids, open gaps and overturn condition. Sources are every
source and excluded[].source of the skills manifest plus the named
additions, each pinned at the default-branch commit read with gh api on
2026-09-30; the skills.sh registry is pinned at the v1.7.0 tag commit
of the manifest's skills CLI. catalogs/README.md links the file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Landscape sweep: skills discovery modality keyed by lifecycle task

The packaged sweep gains a second discovery modality. build_inputs.py
--modality skills builds one skills-<task> layer per task of
catalogs/landscape/skills-lifecycle.json (installed skills with their
manifest pins and invocation flags, the task's pinned sources, known
skill refs), and --skills-scope prints the frozen scope for those layers
in saturation_ledger.py --scope's format with the ledger's own hash
functions. build_args.py resolves the modality at build time: a skills
run's discover and critic are the new discover_skills and critic_skills
templates, and facts and fit end in <<MODALITY>>, filled with
modality_skills (or with an empty string, so a repository run's frozen
templates and prompts_sha256 are unchanged). A run covers one modality.

schemas/discover-skills.json is the strict skills return (skill_ref,
source_id, lifecycle_task, pin, skill_md_sha256, license, source,
description_chars, model_invocable, codex_implicit, replaces, and the
verdict fields; the overturn comparison must name skill-creator or
promptfoo). sweep.js merges skill proposals by skill_ref and carries it
as the pipeline identity, so the refuters, convert.py, make_result.py
and the ledger compare it unchanged. source_reviews.py reviews a skill
survivor at its SKILL.md (sha256, disable-model-invocation, openai.yaml
allow_implicit_invocation), and make_result.py --decision-record writes
docs/decisions/<date>-skills-landscape-sweep.md from a skills RESULT.json.
skills_problems now scans every template for pinned skill names.

Tests: tests/test_landscape_sweep_skills.py (catalog, schema pass/fail/
malformed, inputs and scope, template resolution, staging, sweep.js under
node, skill source reviews, decision record); the harness suite adds the
skills prompts_sha256 pin and the all-template skill drift check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality: README section and decision record

tools/sota-convergence/landscape-sweep/README.md gains a Skills modality
section (inputs, scope, templates and schema, identity, one modality per
run, the run commands, outputs, and what the ledger does not bind yet).
docs/decisions/2026-09-30-skills-sweep-modality.md records the choice of a
modality inside the packaged sweep over a separate skills sweep script,
the runtime template selection rejected on measurement, the residuals
(saturation_ledger.py refuses the skills catalog; no live run) and the
overturn condition.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills inputs: known_skills as the manifest states them

The dry build on the real catalog showed known_skill_refs at 292 refs
(7.9 KB of every skills layer input, and of every GPT-6 prompt that embeds
it), many invented: an excluded manifest entry that names several
repositories was cross-multiplied into owner/repo@name pairs the manifest
never claims (microsoft/playwright-cli@webapp-testing). known_skills keeps
the installed skills by repository and the excluded skills by the
manifest's own source text (4.4 KB). The two skills templates that named
the field follow, so the skills-run prompts_sha256 pin moves to
5d9ae85a...48db; the repository pin is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Saturation ledger: report, append and check skills-* layers

The skills modality's layers had no ledger path: --report listed only
research-state.json rows, so build_args.py --due-report selected no
skills layer ("no layer selected"), and --append refused
skills/skills-debug as "not a layer in research-state.json".

saturation_ledger.py now reads catalogs/landscape/skills-lifecycle.json
as catalog "skills" layers (skills-<task>) beside the research-state
rows. skills_requirement_sha256 hashes a task's lifecycle_task,
requirement and overturn_when; build_inputs.py --skills-scope will call
it, so the frozen scope and the ledger share one definition. --report
lists every task after the research-state layers (research status "-"),
--append completes a skills RESULT.json, and --check accepts catalog
skills. A SOTA-convergence manifest has no skills section, so a skills
survivor binds to its retained votes and its source review instead of a
manifest row; repository layers keep their manifest binding. A changed
task requirement is a current requirement_changed trigger citing the
skills catalog. --scope is unchanged.

Tests: a skills sweep appends and checks, the report lists skills layers
as due (and the --report --json rows build_args.py reads), unknown
skills layers and mis-bound survivors are refused, and the committed
checkout's report lists all 13 tasks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills inputs: whole excluded names and texts; mark openai/skills stale

Review findings 1, 3, 5 and 7.

- known_skills: the committed adoption/skills/manifest.json states each
  excluded entry's skills as one comma-separated string, and the loop
  iterated its characters (363 of 363 excluded values were single
  characters in the dry build). excluded_names splits on commas and
  keeps each name or phrase whole; a list is taken item by item. The
  input fixtures now use the committed string shape, with one
  list-shaped entry, and a new test builds inputs from the committed
  catalog and manifest: no known_skills value has length 1, and
  obra/superpowers lists systematic-debugging.
- Whole texts: an installed skill's gap was cut at 300 characters with no
  marker (variant-analysis, writing-for-agents, security-audit lost text
  mid-word). The skills modality now passes the gap, the task's open gaps
  and its overturn_when whole; repository layers keep their caps.
- --skills-scope computes each requirement hash with
  saturation_ledger.skills_requirement_sha256, the ledger's own
  definition.
- Stale source: gh api 'repos/openai/skills/commits?sha=main&since=
  2026-07-02T00:00:00Z&per_page=1' returns no commit; the last main
  commit, 49f948fa, is dated 2026-06-24, 98 days before the check, so
  the common block's 90-day maintenance rule marks it stale. The catalog
  records that as a maintenance record on openai-skills (status, check
  date, API fact), keeps the source for its four installed skills, and
  adds "Stale source: replacement via this sweep" to skills-ci-pr and
  skills-security. The other 21 GitHub sources were checked the same way
  and are maintained (catalog notes). check_skills_catalog validates the
  optional record.
- The two model-invocable requirements say "no true-valued
  disable-model-invocation (true/yes/on/1)".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills templates: client-exact invocation flags, stale sources, asked-for doc edits

Review finding 4 and 7, GPT-6 review finding (c).

- Claude Code accepts true, yes, on and 1 (and false, no, off, 0) in any
  letter case for boolean frontmatter since v2.1.218
  (code.claude.com/docs/en/skills, Frontmatter reference). discover_skills
  and modality_skills now say "a true-valued disable-model-invocation
  (true/yes/on/1) in any letter case" instead of the literal
  "disable-model-invocation: true".
- Codex rust-v0.157.1 reads agents/openai.yaml with serde_yaml 0.9.34,
  whose bool is only a plain true/True/TRUE or false/False/FALSE, and
  ignores the whole file when it cannot deserialize it
  (codex-rs/ext/skills/src/loader/metadata.rs, "Fail open"). Both
  templates ask for a plain false and say that a quoted or other value
  leaves implicit invocation on.
- discover_skills labels the skills of a source whose catalog entry
  carries a stale maintenance record not_adopted, citing the record,
  unless a maintained fork or replacement is found (the common block's
  rule: "no commit on its default branch in the last 90 days");
  modality_skills tells the refuters such a source is stale under it.
- Writing CLAUDE.md or AGENTS.md conflicts with the canonical
  instructions only on the skill's own initiative: editing them when
  asked is skills-agent-docs' requirement, which the unconditional rule
  contradicted.

Failing first: the skills-run prompts_sha256 pin moved from
5d9ae85a...48db to 2c2efbae...63d3 and the template-phrase test failed on
the old literal; the repository-run pin (9c34fa72...f14b) passed
unchanged, since only skills templates changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skill source reviews: the pinned skills CLI's discovery order, at the adjudicated pin

Review finding 2, review finding 4 (Claude booleans), GPT-6 review
findings (a) and (b).

- Discovery order: skill_review refused any skill whose folder name
  occurs twice in the tree, which at the catalog pins refused 254 of
  affaan-m/ECC's 296 names (translations under docs/<lang>/skills, agent
  copies under .kiro/skills and .agents/skills), 2 of microsoft/skills'
  and 1 of openai/skills'. cli_skill_dir now follows vercel-labs/skills
  v1.7.0 discoverSkills (src/skills.ts at 7407f389, lines 180-329; README
  "Skill Discovery", lines 412-483): a valid root SKILL.md is the only
  skill; then the root's child folders, skills/ with .curated,
  .experimental and .system, and the 30 AGENT_PROJECT_SKILL_DIRS, each
  three levels deep with a shallower SKILL.md shadowing the folders below
  it; then the folders .claude-plugin manifests declare
  (plugin-manifest.ts getPluginSkillPaths); and only when none holds a
  skill, every folder five levels deep. The first location holding the
  name decides; only two same-named folders there are refused, as are a
  truncated tree and folders outside these locations (docs/ needs
  --full-depth). Where the README's folder list differs from the code
  (31 extra, 7 missing), the code decides. Checked on 2026-09-30: the
  resolver matches `npx skills@1.7.0 add <repo> --list` exactly on six
  sources (ECC 294, microsoft/skills 13, trailofbits/skills 83,
  openai/skills 43, mattpocock/skills 37, vercel-labs/agent-browser 1),
  and all 28 installed skills resolve to their manifest path at their
  manifest ref.
- Adjudicated pin: convert.py writes a skills survivor's pin and
  skill_md_sha256 into survivors.json, and the review reads the SKILL.md
  at that pin, falling back to the default branch only when gh cannot
  read the pin, and saying so in claim and observed.pin_fallback. A
  SKILL.md whose sha256 differs from the adjudicated one is refused; one
  skill judged at two pins gets one review per pin.
- agents/openai.yaml: openai_yaml_policy reads the documented policy
  mapping in block and flow form as Codex rust-v0.157.1 does (serde_yaml
  0.9.34: only a plain true/false is a bool; a quoted or other value, a
  repeated key or a tab-indented line makes Codex ignore the file, which
  leaves implicit invocation on), and records the value as written.
  PyYAML is not installed on the review host, so no loader is added.
- Frontmatter: a trailing comment no longer breaks a name, and
  disable-model-invocation reads Claude Code's boolean set (true, yes,
  on, 1 in any letter case).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills decision record: lost workers read as incomplete layers

Review finding 6. make_result.py --decision-record printed "No proposal
survived." for a layer whose proposals were refuted and for a layer that
lost workers alike. It now lists RESULT.json lost_workers in the context
(mapping <role>:<layer>[:followup] and the critic to layers with
sweep_common.deviation_rounds, and naming labels that map to no layer),
and per layer says whether it has survivors, lost workers (incomplete,
not a refutation; its reopen entries keep it from counting as clean),
only refuted proposals, or no proposal at all.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality docs: CLI-order reviews, stale sources, ledger binding

README (Skills modality): whole excluded names and texts; the stale-source
record; the invocation flags as Claude Code and Codex read them; the
unasked-write rule for CLAUDE.md and AGENTS.md; source reviews at the
adjudicated pin in the pinned skills CLI's discovery order; and the
"reported and skipped" cases stated with make_result.py's gate (a
skipped survivor blocks RESULT.json until resolved or the run is
recorded as stopped). The ledger bullet replaces "Not yet in the
ledger" and names the one remaining gap: ledger.schema.json's catalog
enum, which is outside this unit.

Decision record: live-page citations use section anchors and the fetched
snapshot's sha256 instead of line numbers (two same-day fetches of
code.claude.com's skills.md differed); the skills CLI and Codex/serde_yaml
facts with their tag commits; the rejected self-written review rule; the
stale source; the ledger binding; and named residuals (schema enum,
resolver limits, not run).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills staging test: the ledger's due report selects the skills layers

GPT-6 review finding (d) in the suite: build_args.select_layers over a
skills work directory's rows with the committed checkout's
saturation_ledger.py report selects exactly its due skills layers, and a
report without skills rows (the ledger before a5e830ab) still raises
"no layer selected". Failing first against the previous head's code: no
skills row was due ("[] is not true").

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Saturation ledger schema: allow the skills catalog for skills-* layers

The skills modality records its layers as catalog "skills"; the ledger's row
schema listed only foundation and us-equities. The README names the third value.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Sweep inputs: history and baseline from the run's own modality

build_inputs.py took previous_sweep and the default baseline manifest from the
ledger's last completed sweep of either modality, so a skills sweep after a
repository sweep emptied the next repository run's history and became its
baseline (GPT-6 review of #541, build_inputs.py:124). The ledger names a
sweep's modality by its layers' catalogs (sweep_modality: repository for
foundation and us-equities layers, skills for skills layers, none for a mixed
or empty sweep), and each run now reads only the last completed sweep of its
own modality. A modality with no completed sweep gets an empty history and no
baseline, and the summary line names the sweep it read.

Test: ModalityHistoryTests (repositories, skills, repositories: the third run
reads the first; the skills run reads none).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skill source reviews: the adjudicated pin only, validated copies, a folder hash and Codex's YAML reading

Three findings of the GPT-6 review of #541 (source_reviews.py:331, :451, :498):

- agents/openai.yaml: with PyYAML installed, yaml.compose_all (the libyaml
  CSafeLoader when present) builds the node tree and serde_yaml 0.9.34's rules
  from openai/codex rust-v0.157.1 decide: a plain true/false spelling only,
  anchors and aliases followed, a second document or a repeated field refused
  (Codex then ignores the file). Not safe_load: its YAML 1.1 constructor reads
  yes/no as booleans and drops the scalar style. Without PyYAML only the plain
  block-mapping subset is read; a flow mapping, an anchor or alias, a tag, a
  second document, a block scalar or any unrecognized line gives
  codex_implicit null with the reason, never a guessed true. The review records
  codex_implicit, unverified_reason and the reader.
- Discovery validates each candidate SKILL.md before the CLI's location order
  applies, as vercel-labs/skills v1.7.0 parseSkillMd does (a name and a
  description, both non-empty strings, typed as the CLI's yaml package reads
  YAML 1.2 core scalars), and a copy whose name field names another skill is not
  a copy (filterSkills). An invalid earlier copy is skipped and recorded in
  observed.skipped_skill_md, a valid later copy wins, the recursive fallback runs
  only when no location holds a valid skill, and a skill whose every copy is
  invalid is refused.
- A skill survivor is reviewed at its adjudicated pin and at no other commit: no
  pin, an unreadable pin or one that reads as another commit (pin_lookup
  failed), a null skill_md_sha256, or a later refusal (pin_lookup ok) gives no
  review and a stopped entry in the printed list. The review records the skill
  folder's git tree id (skill_folder_tree_sha, the CLI lock's skillFolderHash,
  src/blob.ts getSkillFolderHashFromTree), verified to cover the SKILL.md and
  agents/openai.yaml bytes read. make_result.py stops a layer with a stopped
  entry, and completes a skills layer only when each review's reviewed_commit is
  the adjudicated pin in the returns, over the judged skill_md_sha256, with the
  folder hash recorded.

Tests: the YAML table under both readers (the PyYAML class runs when PyYAML is
installed), git-oracle blob and tree ids, the three copy-validation cases, the
unreadable pin, a moved pin, no pin and a null hash, the folder hash changing
with agents/openai.yaml alone, a listing that does not cover the bytes read,
and make_result.py's pin, hash, folder-hash and stopped-layer refusals. The
existing review tests now give every survivor its pin and hash.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality docs: the delivered ledger schema and the pinned reviews

The README and the decision record said the ledger schema still left out the
skills catalog and needed a future edit, but this branch adds it at
catalogs/saturation/ledger.schema.json (GPT-6 review of #541, README.md:498).
Both now describe the delivered schema and drop the obsolete prerequisite, and
they describe what the repairs deliver: reviews at the adjudicated pin only,
stopped layers, candidate validation, the skill folder's tree hash, Codex's
reading of agents/openai.yaml with or without PyYAML, and each modality's own
history. The decision record's sources add the pinned files the repairs rely on
(vercel-labs/skills v1.7.0 parseSkillMd, filterSkills,
getSkillFolderHashFromTree and package.json; eemeli/yaml v2.9.0's default
schema; serde_yaml 0.9.34's bool, map and document rules), notes that the two
resolver measurements predate candidate validation, and records the stricter
review of a survivor with a null hash as a residual.

Test: SkillsDocsTests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* openai.yaml subset reader: an escape YAML does not define leaves codex_implicit unverified

GPT-6 round-3 review of #541 (medium 2): without PyYAML, `note: "C:\skills"` beside
`policy.allow_implicit_invocation: false` gave a definite codex_implicit false, while serde_yaml
0.9.34 scans with unsafe-libyaml 0.2.11 (src/scanner.rs lines 2195-2361), refuses the whole file
("found unknown escape character") and Codex then leaves implicit invocation on.

The subset reader now checks, anywhere in the file, that each quoted scalar closes on its line and
that each double-quoted one uses only YAML 1.2's escapes (\x, \u and \U with their 2, 4 and 8 hex
digits; no surrogate and nothing beyond U+10FFFF, as libyaml), and otherwise reports codex_implicit
null with the reason. With PyYAML the libyaml composer already refuses these files; the
pure-Python SafeLoader took a surrogate escape and raised ValueError beyond U+10FFFF, and both now
stay unverified as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skill source reviews: type SKILL.md frontmatter as the skills CLI's yaml package does

GPT-6 round-3 review of #541 (medium 1): a block-sequence description (`description:` then
`  - Find bugs.`) was classified as a string, so an invalid find-bugs/SKILL.md still won over a
valid skills/find-bugs/SKILL.md. The pinned CLI (vercel-labs/skills v1.7.0 src/skills.ts
parseSkillMd, lines 80-133) parses the frontmatter with the yaml package 2.9.0 and skips a copy
whose name or description is not a non-empty string.

skill_md_check now types the frontmatter as that package does (YAML 1.2, core schema): with PyYAML,
yaml.compose_all builds the tree and each scalar is typed by the core schema by its style, not by
PyYAML's YAML 1.1 resolver; without it a subset reader covers one-line plain, quoted and flow
values, multi-line plain scalars, literal and folded block scalars (strings, as the package reads
them) and nested lists and mappings. A list, mapping or flow collection in name or description is
an object, so that copy is skipped and a valid later copy wins. Neither reader guesses: a tag, an
anchor or alias (subset reader), a key given twice, an escape YAML does not define, a file PyYAML
cannot parse or syntax outside the subset leaves the copy unverified (SkillMdUnverified), and when
its verdict decides which copy the CLI takes, the survivor is stopped (pin_lookup ok) instead of
reviewed at a guessed copy; the fallback check waits for a verified SKILL.md before it trusts the
recursive search. In the same class, a SKILL.md gh cannot read is unverified rather than skipped
(the CLI reads it from its clone), and a copy's name is compared as the CLI records it
(sanitizeMetadata, src/sanitize.ts lines 61-65; an empty one matches its folder, lines 331-333).

Checked against yaml 2.9.0 run under node with parseSkillMd's checks: both readers agree on all
2,455 SKILL.md files of the catalog's GitHub sources at their pins, and give no contrary verdict on
102 edge cases; `npx skills@1.7.0 add <fixture> --list` skips the list and mapping copies with the
same message.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality docs: typed SKILL.md frontmatter, unverified copies and openai.yaml escapes

The sweep README and the decision record describe the frontmatter readers (PyYAML's composer typed
by the YAML 1.2 core schema, else the subset reader), the unverified outcome that stops a survivor
instead of deciding the CLI's pick by a guess, the name as sanitizeMetadata records it, and the
openai.yaml escape check. The decision record replaces the obsolete residual (the frontmatter read
line by line) with the readers' limits, records the 2026-09-30 check against yaml 2.9.0, and cites
eemeli/yaml v2.9.0, vercel-labs/skills v1.7.0 src/sanitize.ts and unsafe-libyaml 0.2.11.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skill source reviews: the pinned skills CLI's own parseSkillMd decides each SKILL.md copy

The hand-written SKILL.md frontmatter readers (the PyYAML composer reader and the subset reader) are gone:
an independent review of round 3 found both still giving a verdict on frontmatter the CLI's yaml package
reads otherwise. skill_md.mjs runs vercel-labs/skills v1.7.0's parseSkillMd, parseFrontmatter,
sanitizeMetadata and getSkillDisplayName (src/skills.ts L80-133 and L331-333, src/frontmatter.ts,
src/sanitize.ts L18-65 at 7407f389, ported line by line) with the yaml package skills-yaml.pin.json pins
(2.9.0 and its npm integrity, the CLI lockfile's), after checking every pinned file's sha256 and the
install's lockfile integrity, following the tree-sitter-bash pin pattern (shell-parser.pin.json). It
installs nothing; without node or a verified install every copy is unverified and the survivor stops.
A line break other than LF or CRLF in a frontmatter gives no verdict (libyaml, which Codex's serde_yaml
uses, breaks lines there and the yaml package does not).

Every SKILL.md copy read is tied to the regular-file blob the git tree lists (a symlink, missing content
or other bytes are unverified), plugin manifests likewise, and the license and disable-model-invocation
come from the same parse. agents/openai.yaml keeps its readers, behind a pre-check: bytes that are not
UTF-8, characters outside libyaml's printable set (unsafe-libyaml 0.2.11 src/reader.rs L381-395) and line
breaks other than LF and CRLF give codex_implicit null with the reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality docs: the reader is the CLI's own parseSkillMd, and what it leaves unverified

The README's Skills modality section and the decision record said neither frontmatter reader guesses,
which an independent review of round 3 refuted (finding 1). They now describe skill_md.mjs and
skills-yaml.pin.json (the CLI's parseSkillMd ported line by line, run with the pinned yaml 2.9.0 after its
files and lockfile integrity are checked, nothing installed at run time) and name what stays unverified:
no node or install, a line break other than LF or CRLF in a frontmatter, a copy that is not the listed
regular-file blob, symlinked folders, Claude Code's own parser, the openai.yaml checks, and CI, which does
not provision the pin yet. They record that npm resolves the CLI's yaml ^2.8.3 at install time (2.9.1 on
2026-09-30, also this host's install) and that 2.9.0 and 2.9.1 gave the same results, with the receipt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality receipt: the SKILL.md reader against the pinned CLI's own parseSkillMd and skills add --list

evidence/artifacts/skills-md-reader-20260930: skill_md.mjs against two oracles on all 2,455 SKILL.md files of
the skills catalog's 22 GitHub sources at their pins and on 139 edge cases (the round-3 review's inputs, line
breaks, the name rules, the yaml 2.9.1 changes): parseSkillMd and its helpers sliced byte for byte from the
published skills@1.7.0 dist/cli.mjs, run with the yaml its install resolved (2.9.1) and with the pinned 2.9.0
(local integration check), and the installed CLI's skills add <fixture> --list (independent observation).
Corpus: take/take 2,452, skip/skip 3, 0 contradictions, 0 field differences; the CLI lists the same 1,693
names and warns on the same 3 copies. Edge: 0 contradictions; the 9 cases with a line break other than LF or
CRLF get no verdict by design. yaml 2.9.0 and 2.9.1 gave identical results there and on 200,000 random
multi-line scalars. Four negative controls (two reader mutants, a flipped verdict, a tampered install) fail as
they must. run.sh reproduces it; log.txt and counts.json keep every step's exit status and the digests of the
raw outputs kept outside the repository.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skill source reviews: a search location behind a symlink is unknown; CI installs the yaml pin

Opus round-4 review of #541, D1: discoverSkills walks each search location with readdir, which follows a
symlink at or above it (published dist/cli.mjs L1339-1370), and skips a symlinked skill folder inside one
(entry.isDirectory() is false, src/skills.ts L284-298 and L141-149), while the git tree lists a symlink as one
blob and nothing below it. The review named a later copy the CLI throws away (seenNames, L1352).
location_doubt now marks a location at or below a symlink, and a folder whose SKILL.md is a symlink with skill
folders below it; plugin_manifests marks the plugin folders unknown when .claude-plugin is a symlink or a
manifest has no verified bytes. Reaching such a location before the pick is decided stops the survivor
(LocationUnverified); one after it cannot change the pick. Failing-first on 9e1bd44: all five symlink
layouts were reviewed at the wrong copy (exit 0).

.github/workflows/validate.yml installs skills-yaml.pin.json before the suite with the tree-sitter step's
pattern (a declared widening of A1's paths to this one workflow file), and tests/test_landscape_sweep_skills.py
holds it there as tests/test_shell_parser_ci.py holds the tree-sitter pin: a runtime tripwire, structure checks,
a ratchet over whole-suite jobs (validate-macos and freshness are recorded gaps) and mutation controls. Also
the round-4 test gaps: --skills-yaml, the environment over the HOME default, and the load_error refusal.
A comment at the gh read records the .gitattributes checkout transforms as a residual.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality docs: the CLI walks a symlinked search location; name de-duplication and the clone scoped

Opus round-4 review of #541. README and decision record: the CLI walks a search location through a symlink at
or above it and skips a symlinked skill folder inside one (the round-4 sentence had it backwards), and such a
location reached before the pick is decided stops the survivor. The record's name rule now states the actual
consequence of de-duplication (add keeps the first skill of a name, dist/cli.mjs L1352; only the update paths
at L7382 and L7588 keep duplicates, so the review can name a copy the CLI throws away), scopes "a GitHub source
is cloned before discovery" to a source named with a ref (L4101-4102, L5190-5212), lists .gitattributes
checkout transforms (LFS, eol, working-tree-encoding) as a residual, and records the CI provisioning of the
yaml pin with its two recorded gaps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality receipt: the symlink rule on the catalog's sources and the CI provisioning of the yaml pin

evidence/artifacts/skills-md-reader-20260930, rerun at the round-5 code (33 steps, 2026-09-30 23:45:46Z to
23:46:42Z; code under test by sha256: skill_md.mjs and skills-yaml.pin.json unchanged, source_reviews.py new).
The reader results are unchanged (corpus take/take 2,452, skip/skip 3; edge 0 contradictions, 9 line-break
refusals; skills add --list agreement; yaml 2.9.0 = 2.9.1). New: symlink_impact.py (8 of 22 sources hold
symlinks, 3 a symlinked search location, none a symlinked SKILL.md; all 678 picks unchanged by the rule) and
ci_step.py (the pin's 75 hashes are the bytes of the registry tarball whose sha512 is the pinned integrity;
validate.yml's step, run with npm 10.9.8, exits 0 and exports the directory; the reader under Node.js
22.23.2 verifies that install).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skill source reviews: a symlinked plugin manifest file stops only a pick that could depend on it

Two cases the round-5 symlink tests missed: .claude-plugin/marketplace.json as a symlink (readFile follows it;
the review reads no verified bytes) stops a survivor whose pick would come from the plugin folders or the
recursive search, and leaves the review of a pick decided at skills/. Round 4 stopped every review of a
repository with a symlinked manifest, with the reason "the git tree lists a symlink at .claude-plugin/...".
Decision record: the two includeDuplicateNames: true call sites (dist/cli.mjs L7382, L7588) are named as the
update paths, and the symlink residual names the layout it over-stops (.agents/skills -> ../skills searched
before a later pick).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality receipt: rerun with the plugin-manifest doubt in symlink_impact.py

symlink_impact.py now marks the plugin folders unknown for a manifest that is not a regular file, as
source_reviews.plugin_manifests does. Rerun at 6643ba60 (33 steps, 2026-09-30 23:56:37Z to 23:57:35Z): no
source has a plugin manifest behind a symlink, and every other result is unchanged (679 names, 678 picks equal
with and without the rule; the reader and CI provisioning results as before).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality on the stacked skills manifest (F3 #553): skill-creator named, the retired skill dropped

Unit F3 (#553), below this PR in the stack, pins skill-creator (anthropics/skills@8a1541c4, trial,
codex_enabled false), retires resolving-merge-conflicts and widens the Claude listings and Codex
enablement. Built against main's manifest, A1 failed 35 tests on the stack.

- build_args.py: TEMPLATE_SKILLS gains skill-creator, which the skills templates already name for its
  paired with-skill/without-skill benchmark; templates.json common's Skills paragraph names it for
  Claude Code only (Codex ships a different skill of that name), read and never run. Both prompts_sha256
  pins move (repository 9c34fa72...f14b -> b61956f3...726d, skills 2c2efbae...63d3 -> a76ee858...b460).
- catalogs/landscape/skills-lifecycle.json: skill-creator serves skills-skill-lifecycle;
  resolving-merge-conflicts leaves skills-ci-pr, retired by unit F3 (#553), whose open gap now says no
  installed skill resolves rebase conflicts; the open gaps that stated pre-F3 codex_enabled and listing
  values are removed or restated, and skills-agent-docs' requirement points skill authoring at
  skill-creator.
- tests: a catalog test holds each flag an open gap states to the manifest (12 stale statements fail it
  on the previous catalog); the whole-gap input test bounds security-audit's gap by the 600-character cut,
  since F3 shortened it to 930 characters.
- README and decision record: the repository run's prompts_sha256 moved once; the 28-skill path check
  is scoped to the manifest before F3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Skills modality receipt: log each command without trailing blanks, and rerun

run.sh's step() wrote the oracle-versions command's final newline into log.txt as a trailing space,
which git diff --check reported on line 26; the line is this receipt's own log format, not retained
tool output, so the logger now strips trailing blanks and the receipt was run again (2026-10-01
00:56:03Z to 00:57:01Z, 33 steps, the same exit statuses). Against the 2026-09-30 run, log.txt differs
only in its times and that trailing space, and counts.json only in the run times, the checkout commit
and the hashes of raw outputs kept outside the repository (the scratch path, node process ids and npm
timings inside them).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Re-register the changed hash-listed files in manifests/evidence.json (hot-file protocol)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…uide is on main

The skill-lifecycle bullet pointed to adoption/skills/lifecycle.md "(lands with unit F3)". F3 landed in
3361b34 (#553) and the guide is on main, so the parenthesis is stale. Shared hot file (docs/lanes.md):
AGENTS.md and its manifests/evidence.json re-registration are this branch's only commit, on main
dcae68b (#622). Handed off by the GitHub/CI lane (native-agent-stack-2f); taken at the user's request.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…uide is on main

The skill-lifecycle bullet pointed to adoption/skills/lifecycle.md "(lands with unit F3)". F3 landed in
3361b34 (#553) and the guide is on main, so the parenthesis is stale. Shared hot file (docs/lanes.md):
AGENTS.md and its manifests/evidence.json re-registration are this branch's only commit, rebuilt on main
e88d59e (#639) with AGENTS.md byte-identical to the acknowledged head 1283e4e.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…uide is on main

The skill-lifecycle bullet pointed to adoption/skills/lifecycle.md "(lands with unit F3)". F3 landed in
3361b34 (#553) and the guide is on main, so the parenthesis is stale. Shared hot file (docs/lanes.md):
AGENTS.md and its manifests/evidence.json re-registration are this branch's only commit, rebuilt on main
d2777ee (#619, #620) with AGENTS.md byte-identical to the acknowledged head 1283e4e.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…uide is on main

The skill-lifecycle bullet pointed to adoption/skills/lifecycle.md "(lands with unit F3)". F3 landed in
3361b34 (#553) and the guide is on main, so the parenthesis is stale. Shared hot file (docs/lanes.md):
AGENTS.md and its manifests/evidence.json re-registration are this branch's only commit, rebuilt on main
9b0b8d6 (#634) with AGENTS.md byte-identical to the acknowledged head 1283e4e.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Oct 3, 2026
…uide is on main (#636)

The skill-lifecycle bullet pointed to adoption/skills/lifecycle.md "(lands with unit F3)". F3 landed in
3361b34 (#553) and the guide is on main, so the parenthesis is stale. Shared hot file (docs/lanes.md):
AGENTS.md and its manifests/evidence.json re-registration are this branch's only commit, rebuilt on main
9b0b8d6 (#634) with AGENTS.md byte-identical to the acknowledged head 1283e4e.

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant