Repository navigation
Keep the root directory as one PATH component; refuse a host HOME of /; six anti-pattern log rows - #619
Conversation
…/; six anti-pattern log rows managed_block.shell_dir stripped trailing slashes so that an explicit --extra-dir / became the empty string, and the profile block wrote an empty PATH component, which names the current directory (POSIX XBD 8.3). The root now stays "/" (POSIX basename steps 3 and 4); an ecosystem root of "/" still gives /bin. The same strip in new_wsl_client_config.host_values made a declared HOME of "/" the empty prefix of every absolute path, so --apply moved the first HOST_PATH entry under the new home: such a host value file is now refused and nothing is written. Tests pin both cases and fail on the base code. Found by the Codex lane's read of #608 (comment 5961552065). Log rows: the defect; the gh suite's inherited UI variables (pager, NO_COLOR, TERM=dumb, sources verified at v2.102.0); a measurement window opened with a model still resident; a reused run label; and, handed over by the SDK lane, a full-registry print and per-block usage counting. Three changed files re-registered in manifests/evidence.json. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Cross-family read requested from the Codex lane at exact head |
|
Head is now |
|
Codex scoped source review at exact head The changed Native GitHub PR/head and fully paginated six-file delta reads returned0; original delta capture is26882bytes/SHA256 The advisory/required-CI hold and builder's destination readback remain separate. This comment supplies the requested exact-head cross-family source read, not a GitHub approval, merge or host qualification. |
…ock-root-dir-20261002
|
ACCEPT of the source carryover, exact HEAD Both comparisons have the same six PR-owned files and143 additions/13 deletions. Their added/removed lines match exactly; complete code/test/document blobs are unchanged at managed_block.py, new_wsl_client_config.py, their two test modules and harness-defaults. The registry retains the same three PR-owned SHA256/size replacements while incorporating main. The merge's two parents are the prior reviewed head anddcae68bd. All105 inherited changed paths match main's retirement source, with registry-only owned replacements. The independent bounded API comparison returned0 for all eight native source reads; its initial processor syntax failure1 occurred before any API invocation and is retained. The prior source judgment carries to these unchanged PR inputs. The reported188 local tests/validator checks remain owner-reported at this head; this review runs no tests or providers. Hosted required checks and exact-head/main/shared-path guards still govern merge. Native sign-ins/credentials and active client configurations were untouched, and root edited no Claude-owned source path. Publication minor: the current PR description still names old base56473e4b at its scope line. Update it to this actual base and keep historical/local command evidence attached to the head on which it ran. No material PR-owned source delta was found. |
…orrection - docs/decisions/2026-10-02-gpt-runtime-tracking.md: the gap by row kind, the tracked entries and their pin records, ownership (the Codex maintenance task owns upstream detection; this table is a report) and the relation to open #633's runtime-job catalog, the search-first sweep (wf_21fc37c5-123: extend the existing job; updatecli, Renovate and nvchecker trialed on copies of the records), the first local run against live GitHub metadata, verification, limits and follow-ups (tag-only rows, registry versions, advisories through OSV-Scanner). - docs/harness-defaults.md: anti-pattern row for stating what a catalog job tracks from where a repository's name appears, inserted at the top of the log table (open #619 appends at its end). - tests/test_catalog_freshness_runtime.py: the leak-gate test also captures stderr and checks that neither stream carries the matched text or the exception message; the reserved-id test checks that each id belongs to exactly one report kind. Both review nits; the stderr check fails when the caught exception is printed (mutation run). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…orrection - docs/decisions/2026-10-02-gpt-runtime-tracking.md: the gap by row kind, the tracked entries and their pin records, ownership (the Codex maintenance task owns upstream detection; this table is a report) and the relation to open #633's runtime-job catalog, the search-first sweep (wf_21fc37c5-123: extend the existing job; updatecli, Renovate and nvchecker trialed on copies of the records), the first local run against live GitHub metadata, verification, limits and follow-ups (tag-only rows, registry versions, advisories through OSV-Scanner). - docs/harness-defaults.md: anti-pattern row for stating what a catalog job tracks from where a repository's name appears, inserted at the top of the log table (open #619 appends at its end). - tests/test_catalog_freshness_runtime.py: the leak-gate test also captures stderr and checks that neither stream carries the matched text or the exception message; the reserved-id test checks that each id belongs to exactly one report kind. Both review nits; the stderr check fails when the caught exception is printed (mutation run). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…uide is on main The skill-lifecycle bullet pointed to adoption/skills/lifecycle.md "(lands with unit F3)". F3 landed in 3361b34 (#553) and the guide is on main, so the parenthesis is stale. Shared hot file (docs/lanes.md): AGENTS.md and its manifests/evidence.json re-registration are this branch's only commit, rebuilt on main d2777ee (#619, #620) with AGENTS.md byte-identical to the acknowledged head 1283e4e. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…orrection - docs/decisions/2026-10-02-gpt-runtime-tracking.md: the gap by row kind, the tracked entries and their pin records, ownership (the Codex maintenance task owns upstream detection; this table is a report) and the relation to open #633's runtime-job catalog, the search-first sweep (wf_21fc37c5-123: extend the existing job; updatecli, Renovate and nvchecker trialed on copies of the records), the first local run against live GitHub metadata, verification, limits and follow-ups (tag-only rows, registry versions, advisories through OSV-Scanner). - docs/harness-defaults.md: anti-pattern row for stating what a catalog job tracks from where a repository's name appears, inserted at the top of the log table (open #619 appends at its end). - tests/test_catalog_freshness_runtime.py: the leak-gate test also captures stderr and checks that neither stream carries the matched text or the exception message; the reserved-id test checks that each id belongs to exactly one report kind. Both review nits; the stderr check fails when the caught exception is printed (mutation run). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): the branch's last commit, on main d2777ee (#619, #620). The PR's own diff, without this registry, is byte-identical to the cross-family-accepted head f6a2f68's diff against its base. scripts/validate.py passed; validate_convergence --all-recorded valid (26 records); the verdict review gate passed against origin/main. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…recheck to the OSV exception's date and quote the authority Round 4 of PR #635, answering the security review of round 3: - tests/test_frozen_macos_variant_no_use.py: blueprints/**/*.json is no longer an excluded class, so a launch configuration there (an mcpServers entry such as socraticode-mcp.json) is scanned. The excluded classes are *.md, evidence/**, manifests/evidence.json and catalogs/**, and inside them only the configuration and scripts the module recognises are read. The 25 referencing lines this exposes are pinned (23 in experiment-macos-20260924.json, one in each retained OSV-Scanner JSON output of the 2026-09-30 relocks): 56 lines in 18 files. The name is matched in any ASCII letter case; pins stay exact. A new test reads .github/osv-scanner-frozen-macos.toml with tomllib and fails with the recheck message when the exception for GHSA-vcvr-r3jv-pc5j is absent or duplicated or its ignoreUntil is not 2026-12-24. The docstring describes the scope as the code applies it and names the limits: a launch configuration under evidence/** or catalogs/** with a name the module does not recognise is not scanned, and PINNED_LINES can be extended in the change that adds a use (the main ruleset requires no code-owner review). - Mutation checks in a scratch clone: round 3's 25 mutants keep their outcomes; N1 (blueprints/x/launch.json running pnpm in the variant), N2 (an upper-cased path) and N3 (ignoreUntil changed) fail now and pass against the round-3 module; X9-X12 (scope and case regressions, exception removed, reason edited) fail; L5 (the N1 configuration under evidence/**) passes, as the stated limit says. - evidence/receipts/dependabot-alert-16-dismissal-20261003.json: authorization quotes the user's message of 2026-10-03 verbatim with its limits; the alerts 7-15 precedent is analogous, not identical; the dismissal lasts until reopened, with its recheck tied to the exception's date test and to the tripwire. guard, review_date, overturn, limitations and guard_mutation_checks match the module; checked_commit moves to main d2777ee (#619 changed none of the cited files). - The closure record's alert-16 note gives the authority in one sentence that points to the receipt, quotes the dismissal comment's "Live recipe lock pins next 16.3.6+" and states the 16.3.8 fact separately, says exactly when the module fails on the lock, and describes the scan scope as applied. The branch is rebased onto main d2777ee (#619); round 3's registry commit was dropped first. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…log-freshness report (#634) * Track GPT runtime workers, SDKs and agents in the daily catalog-freshness report The daily catalog-freshness report covered the foundation stack and the trading pins only. The GPT-route runtime workers, agent SDKs, the GPT gateway, the evaluation harnesses and coding agents such as pi were not compared with upstream at the pins their runtime records install, so the foundation the trading north star runs on could fall behind without notice. extract_layers.py gains RUNTIME_PIN_SOURCES, 12 pins read at extraction time from the new-WSL install plan rows by slot (a composite row by part), the OpenHands runtime-worker recipe pin record and the native SDK constraints (PEP 503 name match), and RUNTIME_WATCH_SOURCES, 7 watch-only upstreams that a file on main names but no record on main pins (pi, oh-my-pi, the OpenAI Agents SDKs, Crawl4AI, Deep Agents and codex-action). Both resolve into runtime-pins.json. A malformed declaration raises. A record that moved or changed shape yields an entry with "pin": null and an error built only from the declared path and locator, so the daily job keeps its foundation and trading report. github_freshness.py fetches these repositories. build_manifest.py --runtime-freshness-out writes the report-only runtime-freshness.json; the manifest and the trading sidecar stay byte-identical. freshness_propose.py renders it as a separate drift.md table that never sets drift-status.txt and that the propose job never reads. The workflow writes and uploads the sidecar and prints its counts. These rows stay out of TRADING_PIN_SOURCES, whose pins must not repeat a selected card's repository (openai/codex, OmniRoute, inspect_ai and deer-flow are cards) and must sit in a trading taxonomy layer. tests/test_catalog_freshness_runtime.py covers extraction against the checked-in records, the resolver on synthetic records, the fetch, build and report stages and the workflow wiring, with no network. scripts/validate.py reports SHA-256 and byte-count mismatches for the six registered files edited here; they are re-registered in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Withhold only the runtime sidecar on a leak-gate trip and tighten the runtime tests Resolves the round-2 review findings on 6bc1199 (none blocking or major). F1: build_manifest.py catches LeakDetected around the runtime sidecar's gate only. A runtime upstream whose own data trips the gate (a third-party tag containing APCA, say) no longer fails the build step and so loses the workflow's diff step and drift report: the sidecar keeps its schema and keys with no entries, every count 0 and the fixed "gate_error": "leak_gate_tripped" (never the matched text or the exception message), the step prints {"runtime_freshness": {"gate_error": "leak_gate_tripped"}} and exits 0. The manifest's and the trading sidecar's leak checks stay fatal. freshness_propose.py renders that document as the runtime heading and one withheld line with no table, and returns runtime_unresolved ["leak_gate_tripped"], so the workflow's counts line shows it. F2: the install-plan test asserts each entry's declared slot and, for the two research-harnesses parts, the part_repository slug at its position. The codex and codex-sdk-and-codex-exec-app-server rows carry identical values, so a swapped slot passed before. F3: watch-only now means no install or runtime record on main pins the upstream (a catalog card may record an evaluated version). The alternative openai-agents-sdk card records v0.22.3 for openai/openai-agents-python, which no table compares. The comments, the build docstring, the README and the drift.md intro say so, and that the stack.json and selected (default/conditional) card pins stay in the other tables. F4: the synthetic constraints give openai-codex-cli-bin 9.9.9, so the PEP 503 test proves the openai_codex line matched. F5: the README says an unresolved row carries its upstream and dormancy only when it has a repository; one without has an empty upstream and the not_fetched dormancy. F6: a test runs extract_layers.main with a call-through spy on resolve_runtime_pins and checks that reserved_ids holds hftbacktest (a trading pin), nautilustrader (a trading card) and codex (a foundation component). F7: an unresolved row without a repository is no longer also listed as unfetched; the unresolved line reports it. A row that kept a declared repository is still listed when it was not fetched. Mutation runs in a scratch copy fail each new assertion on its target defect; the base test file passed the F2 slot-swap and F4 prefix-match mutants. scripts/validate.py reports SHA-256 and byte-count mismatches for the same six registered files as 6bc1199; they are re-registered in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Record the GPT runtime tracking decision and log the coverage-claim correction - docs/decisions/2026-10-02-gpt-runtime-tracking.md: the gap by row kind, the tracked entries and their pin records, ownership (the Codex maintenance task owns upstream detection; this table is a report) and the relation to open #633's runtime-job catalog, the search-first sweep (wf_21fc37c5-123: extend the existing job; updatecli, Renovate and nvchecker trialed on copies of the records), the first local run against live GitHub metadata, verification, limits and follow-ups (tag-only rows, registry versions, advisories through OSV-Scanner). - docs/harness-defaults.md: anti-pattern row for stating what a catalog job tracks from where a repository's name appears, inserted at the top of the log table (open #619 appends at its end). - tests/test_catalog_freshness_runtime.py: the leak-gate test also captures stderr and checks that neither stream carries the matched text or the exception message; the reserved-id test checks that each id belongs to exactly one report kind. Both review nits; the stderr check fails when the caught exception is printed (mutation run). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Compare tag-only runtime upstreams by declared tag patterns Inspect AI publishes no GitHub release and its first tag in name order is release/2025-11-28, so new-wsl:inspect-ai was not compared although its newest version tag is 0.3.276 against the pin 0.3.273. watch:codex-action (tags only) and watch:deepagents (a monorepo whose latest release can be another package's) relied on name order or on whichever package released last. - extract_layers.py: an optional "tags" declaration {prefix, pattern} on a runtime pin or watch source, checked at declaration time (a literal prefix, a pattern anchored with ^ and $ that has exactly one capture group), declared on new-wsl:inspect-ai (every tag), watch:codex-action (v) and watch:deepagents (deepagents==); each entry carries it or null. - github_freshness.py: one "gh api repos/{slug}/git/matching-refs/tags/{prefix} --paginate" call per declared prefix, tag names stored under matching_tags; a failure is a partial error that never aborts the batch, and a record without a list for a declared prefix stays pending. - build_manifest.py: build_runtime_freshness takes the highest version among the names the pattern fully matches (the capture as an integer tuple, never name order) as upstream.latest before the pin is compared, with latest_source and matching_tag_count; otherwise tag_pattern_unmatched or tag_pattern_unfetched. - freshness_propose.py: marks a matching-tag latest "(tag)" and lists the unmatched or unfetched rows in one line after the runtime table. References: nvchecker v2.22 nvchecker_source/github.py (use_max_tag with include_regex), Renovate's github-tags datasource, GitHub REST "List matching references"; gh v2.102.0 pkg/cmd/api/pagination.go (paginatedArrayReader). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep failed runtime tag lists out of partial errors and unpair matching tags from releases Fixes for the round-3 review findings on 0dadeab7 (no blocking or major defect): - G1: github_freshness.py records a failed matching-refs call under the record's matching_tags_errors ({prefix: reason, cut to 160 characters}) and counts it at the document level as matching_tags_errors, never in partial_errors. A partial error makes freshness_propose.py blank the repository's drift and trading rows (Inspect AI is a selected trading card) and holds the propose job. A record carrying it stays pending; the resume is per repository, so the next run fetches the whole repository again (the contract's permitted fallback). build_runtime_freshness gives that row tag_pattern_unfetched with compute_upstream's other fields, and the report blanks nothing. The fetch's last line also prints the new count. - G2: a matching tag sets released_at and prerelease to None and drops latest_flag; they describe the repository's latest GitHub release, which in a monorepo can be another package's. Dormancy still reads the repository's activity. - G3: the decision record's first-run table shows Inspect AI compared against the matching tag 0.3.276 and behind (round-3 local live run on 2026-10-03, 05:29Z to 05:30Z: 488 repositories, 0 errors; 19 entries, 4 behind, 7 not compared) and keeps the 6bc1199 run as the earlier observation. Verification gives 86 runtime tests, the four registered files of the round-3 hash failure and the round-3 and round-4 mutation runs; Limits says what a failed call does and that the list is read with --paginate within the per-call timeout. - G4: DOTTED_VERSION_RE and the declared pattern compile with re.ASCII, so a tag in fullwidth digits cannot rank; a missing or uncompilable pattern gives tag_pattern_unmatched instead of raising; at most 5,000 names are stored per (repository, prefix), with matching_tags_truncated when a list is cut; the tag-miss sentence and the stale releases/tags/commit text are corrected. Tests: 86 runtime tests (6 new, 8 changed). Mutation check: 29 mutants against the 14 new or changed tests; all 41 mutant/test pairings killed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Correct the round-4 record sentence and the review's wording nits - docs/decisions/2026-10-02-gpt-runtime-tracking.md: in the observed run, Deep Agents' released_at and prerelease came from its own release of the selected tag, so G2's nulling removed correct values there; it guards the monorepo case. Adds the independent live run at the round-4 code (488 repositories, 0 errors, 4 behind) and the forced tag-list failure check. - README and github_freshness docstring: the resume rule names matching_tags_errors and the per-repository refetch. - freshness_propose: the tag-miss sentence says both reasons keep the release or tag listing as latest; the runtime intro says the last-release and last-commit columns describe the repository's activity, not the tag in the latest column. The report test follows the sentence. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep runtime-only fetch failures out of the propose gate, withhold cut tag lists and check pointer syntax Fixes for the cross-family review of 2d28e04 (gpt-6-astra at max through native codex exec, read-only sandbox), which returned BLOCK: - X1 (major): github_freshness.py knows, per repository, whether a working file other than runtime-pins.json names it (collect_runtime_only_slugs, compared by normalized slug). The top-level errors and partial_errors, which hold the propose job through freshness_propose.py's upstream-errors.txt and upstream-partial-errors.txt and which saturation_ledger.py reads as the freshness input's completeness, count only those repositories (and any retained record that no working file names), as before runtime-pins.json existed. The failures of the runtime-only repositories (five today) count in the new runtime_only_errors and runtime_only_partial_errors, and runtime_only_repositories lists their slugs. Per-record fields and the resume are unchanged. drift.md's runtime section names, in one line, the rows on a runtime-only repository whose fetch failed; nothing in the drift or propose path reads the new counts. The fetch's last line prints them. - X2 (minor): build_manifest.apply_tag_declaration reads matching_tags_truncated. latest_source becomes tag_pattern_truncated, the selected tag stays the latest (compute_upstream's latest when none matched), and a pinned row is not_compared with reason tag_list_truncated; watch-only and unresolved rows keep their reasons. The tag-miss line names the row with its latest_source, and the "(tag)" marker is not shown for it. The flag is per record, so it applies to every declared prefix of that repository. - X3 (minor): extract_layers.py checks pin_pointer, repository_pointer and a row's array against RFC 6901 (JSON_POINTER_RE) when the declaration is checked, so "tag" for "/tag" raises instead of reading as a moved record. - X4 (nit): the decision record's limits say the runtime rows never trigger a proposal, set drift-status.txt or name receipt component ids, but appear in the committed drift.md; its Decision, Verification and Limits sections cover round 5. The README describes the counters and the cut-list rule. docs/github-automation.md's runtime-table paragraph says that a runtime-only failure does not hold the propose job. Tests: 99 runtime tests, 13 of them new; _fake_gh_api can now fail a repository's primary or releases call. Mutation check: 32 mutants against the 13 new tests; all 57 mutant/test pairings killed. Replay at the reviewed head: a runtime-only 503 wrote 1 to upstream-errors.txt (primary call) or upstream-partial-errors.txt (releases call); with round 5 both say 0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Say which repositories the runtime-only counters cover The runtime-only counters cover a repository that only runtime-pins.json names (no foundation, trading, trading-pin or star-candidate working file does), not every repository the runtime table is the only table for. The drift sentence and docs/github-automation.md now say that (round-5 review, minor). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Re-register the seven edited registered files in manifests/evidence.json Hot-file protocol (docs/lanes.md): the branch's last commit, on main d2777ee (#619, #620). The PR's own diff, without this registry, is byte-identical to the cross-family-accepted head f6a2f68's diff against its base. scripts/validate.py passed; validate_convergence --all-recorded valid (26 records); the verdict review gate passed against origin/main. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Scope
/as an empty PATH entry (an empty entry names the current directory), and the new-distribution client-configuration tool refuses a host value file whose HOME is/, where it moved the first PATH entry under the new home. Six dated rows join the anti-pattern log: the defect itself, a related gh qualification finding, two measurement-lane mistakes, and two rows handed over by the SDK lane.56473e4b840f(origin/main)lane:foundationtools/adoption/managed_block.py,tools/adoption/new_wsl_client_config.py,tests/test_managed_block.py,tests/test_new_wsl_client_config.py,docs/harness-defaults.md,manifests/evidence.json(re-registration of three files only, by the hot-file protocol)The defect was reported on merged #608 by the lane finalizing #610 (comment 5961552065) and confirmed independently by the Codex lane (comment 5961819287). The fix follows the reported control: no empty PATH component is written, and the root stays one component.
SOTA sources
basename, steps 3 and 4: a string made only of slashes becomes a single/(step 3), and only other strings lose their trailing slashes (step 4). https://pubs.opengroup.org/onlinepubs/9799919799/utilities/basename.htmlenviron(7)(man-pages), PATH: an empty entry means the current directory. Observed locally with dash and bash sourcing the block (fixture output kept with the builder's run).fc4b137= tag v2.102.0,internal/ghcmd/cmd.go#L387(pager precedence); AlecAivazis/survey v2.3.7core/template.go:27-29(NO_COLOR); charmbracelet/huh3c0116cform.go#L129-L132(TERM=dumbturns on accessible mode).1ef6d8c7message_parser.py#L210-L220(usageandmessage_idon each assistant message); Agent SDK cost tracking, "Track per-step usage" (messages from parallel tool use share an id and identical usage; count each id once).Evidence-class table
--extra-dir /(and//,/./) writes/as one PATH component and no empty component; a second sourcing changes nothingtests/test_managed_block.pytest_an_explicit_root_directory_is_one_component_and_no_component_is_empty; fails 4 times on the base code (negative control)test_ordinary_extra_directories_give_the_bytes_written_before,test_without_the_option_the_default_profile_gets_the_bytes_written_before; a differential over 1,014 cases changed only the 504 with a root extra directory/is refused, and nothing is moved or writtentests/test_new_wsl_client_config.pytest_a_host_value_file_whose_home_is_the_root_is_refused_and_nothing_is_moved_or_written; fails on the base codeLocal commands run
Build, review, verification, one repair round and a re-verification were done by separate agents. The review asked for the same pattern in
host_values, and the repair covered it. Residuals: the committed test sources the block with/bin/sh(dash) only, and bash ran in the fixture probe; thehost_valuesrefusal ran with stub clients and a temporary home; a declared HOME such as/.falls outside the pattern and is untested.Decision record
No new decision. The anti-pattern log in
docs/harness-defaults.mdrecords the defect and its prevention.Host evidence
None. After this merges, the destination's managed PATH block is read back once.
The required
osv-scannercheck fails on main's frozen historical lock (GHSA-vfj7-8cjw-p6xm,braces3.0.3, issue 384). This PR does not touch it.🤖 Generated with Claude Code