Repository navigation
Cover the managed-config path rule's normalisation and Windows clauses under ntpath - #973
Merged
seathatflowsinourveins merged 2 commits intoOct 11, 2026
Conversation
…ot a managed config The co-op's GPT read of #967 at 3de3dda (ACK, one non-blocking P3): the normalisation equality in _readable_system_file (os.path.abspath(name) == name) could be removed with every test green, because the relative-name test is already rejected by isabs and the default-location matrix only passes normalised strings. ManagedMcpDefaultLocations.test_an_absolute_path_that_is_not_normalised_is_not_a_managed_config uses real files (no fake host, so it runs on every platform): an existing child directory and the same managed-mcp.json reached through a dot-dot, a dot and a doubled-separator spelling. Each exists through the filesystem and each keeps the complete fenced argv, by parameter and by the default route; pathlib drops the dot and doubled-separator segments itself, so only the dot-dot spelling is also asserted as a Path. The normalised spelling of the same file drops only --strict-mcp-config, alone and next to a rejected candidate. Mutations (research record, pr925/p3-mutations-20261010): removing only the normalisation equality fails the new test (4 cases, none elsewhere), reverting the P3 fix still fails the relative-name test and now the new one; removing only isabs stays green because abspath(name) == name already implies an absolute name (an equivalent mutant, so isabs remains for readability). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The command center's note on landing #967: the most useful P3 is a CI-run test that patches the module's os with ntpath, so the Windows side of the check is covered under Python 3.12, the CI interpreter. os.path is posixpath on a POSIX host, so until now the Windows clauses of _readable_system_file were argued from ntpath's functions and never executed. ManagedMcpDefaultLocations.test_the_windows_path_rules_count_the_windows_string_and_not_the_posix_ones replaces the module's os, for the test only, with a namespace whose isabs and abspath are ntpath's and whose file checks answer yes, so the path rule alone decides: the Windows documented string drops --strict-mcp-config, the macOS and Linux strings do not (under 3.12 ntpath.isabs accepts a leading separator and only the equality rejects them; from 3.13 isabs rejects them first), and a dot-dot, dot, forward-slash and doubled-separator spelling of the Windows string, a relative name and a drive-relative name each keep every fence. One accepted candidate among rejected ones decides the form, by parameter and by the default route. Source-level mutants of _readable_system_file, run in the module's own globals so the mutated body sees the swapped os (research record, pr925/p3-source-mutations-*): removing only the normalisation equality fails the new test on 3.12 and 3.13 (the four non-normalised Windows spellings, plus the macOS and Linux strings on 3.12) and the POSIX-side test added in the previous commit; removing only isabs fails it through the drive-relative name, because ntpath's pure-Python abspath on a POSIX host leaves "C:managed-mcp.json" equal to itself (the Windows abspath resolves it, so there isabs is belt and braces); reverting to the presence-only check of #925 fails the POSIX and Windows tests. The harness of the earlier mutation tables replaced the whole function and would not see a swapped os, so these supersede it for this test. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
seathatflowsinourveins
force-pushed
the
claude/cc-native-practice-managed-path-normalised-20261010
branch
from
October 11, 2026 02:57
01fe0ae to
f2bebb4
Compare
seathatflowsinourveins
deleted the
claude/cc-native-practice-managed-path-normalised-20261010
branch
October 11, 2026 03:25
Merged
4 tasks done
seathatflowsinourveins
added a commit
that referenced
this pull request
Oct 11, 2026
…he old text The comment in tests/test_skill_usage.py said ntpath.abspath adds a drive to a POSIX string; the previous commit corrected it (CPython v3.12.4 Lib/ntpath.py:564-578, _abspath_fallback joins the working directory only when isabs is false, then calls normpath). The repository's correction rule asks for a regression test that fails before the fix, and a measurement of the standard library cannot: it passes on the old comment too. test_the_comment_of_the_windows_path_test_states_what_abspath_does reads the source of the test the comment belongs to and requires that the wrong claim is gone and the corrected one and its citation are present. On the module with the old comment restored it fails (the regex matches "abspath adds a drive"); on this head all eight tests of the class pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Oct 11, 2026
The RTK 0.51.0 qualification receipt cited three places that require rtk 0.50.0 as path:line without a commit. #973 added two imports to tests/test_skill_usage.py, so the unqualified `tests/test_skill_usage.py:33` named PARSER_DIR instead of the version regex it named when the receipt was written. The three locators are now pinned to the commit that added the receipt (14048b8), where each resolves; tests/test_rtk_receipt_locators.py reads the line at the pin through Git and requires the prerequisite's regex on it (it fails on the old receipt: line 33 of the working tree is PARSER_DIR). The comment in tests/test_skill_usage.py said ntpath.abspath adds a drive to a POSIX string; it turns the separators into backslashes and adds none (CPython v3.12.4 Lib/ntpath.py:564-578, _abspath_fallback). The comment says so and a new test measures it on a POSIX host. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Oct 11, 2026
…he old text The comment in tests/test_skill_usage.py said ntpath.abspath adds a drive to a POSIX string; the previous commit corrected it (CPython v3.12.4 Lib/ntpath.py:564-578, _abspath_fallback joins the working directory only when isabs is false, then calls normpath). The repository's correction rule asks for a regression test that fails before the fix, and a measurement of the standard library cannot: it passes on the old comment too. test_the_comment_of_the_windows_path_test_states_what_abspath_does reads the source of the test the comment belongs to and requires that the wrong claim is gone and the corrected one and its citation are present. On the module with the old comment restored it fails (the regex matches "abspath adds a drive"); on this head all eight tests of the class pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Oct 11, 2026
…TK receipt's locators (#1005) * Record the paired-run results of planning persistence and orchestration Planning persistence (paired run 2b, 12 runs): repository state files stay the default; the native task list is ranked second and the Ralph-style loop third, for fallback only, with the command center's three statements (an external done check, a post hoc and provisional criterion, a preregistered criterion for any rerun). Orchestration (paired run 1, 18 runs): the single session stays the routing for bounded fixes on north-star code; neither the verify Workflow nor the agent team clears the preregistered rule. F3's two pairs with the team are void because its diff was not captured. The compact artifact holds the per-run rows and the sha256 of every private paper; the rows, their layer pages and the record's addendum state the decisions; tests/test_paired_runs_results.py recomputes each decision from the artifact (19 single-edit mutations are caught); the baseline manifest's label lists the fields that now differ from its rows. The advisor run and the routing run are recorded as owed work; their results are later records. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Pin the RTK receipt's locators and correct the abspath comment of #973 The RTK 0.51.0 qualification receipt cited three places that require rtk 0.50.0 as path:line without a commit. #973 added two imports to tests/test_skill_usage.py, so the unqualified `tests/test_skill_usage.py:33` named PARSER_DIR instead of the version regex it named when the receipt was written. The three locators are now pinned to the commit that added the receipt (14048b8), where each resolves; tests/test_rtk_receipt_locators.py reads the line at the pin through Git and requires the prerequisite's regex on it (it fails on the old receipt: line 33 of the working tree is PARSER_DIR). The comment in tests/test_skill_usage.py said ntpath.abspath adds a drive to a POSIX string; it turns the separators into backslashes and adds none (CPython v3.12.4 Lib/ntpath.py:564-578, _abspath_fallback). The comment says so and a new test measures it on a POSIX host. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Pin the corrected abspath comment of #973 with a test that fails on the old text The comment in tests/test_skill_usage.py said ntpath.abspath adds a drive to a POSIX string; the previous commit corrected it (CPython v3.12.4 Lib/ntpath.py:564-578, _abspath_fallback joins the working directory only when isabs is false, then calls normpath). The repository's correction rule asks for a regression test that fails before the fix, and a measurement of the standard library cannot: it passes on the old comment too. test_the_comment_of_the_windows_path_test_states_what_abspath_does reads the source of the test the comment belongs to and requires that the wrong claim is gone and the corrected one and its citation are present. On the module with the old comment restored it fails (the regex matches "abspath adds a drive"); on this head all eight tests of the class pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Answer the Codex review of #1005 and record paired run 5 (model routing for review micros) Codex review in GitHub at bf864c3, four P2s, each reproduced before the change: - background-workflows: the route now says a saved Workflow in a headless run fenced with --setting-sources local is launched by scriptPath (its name was not found in 6 of 6 runs of paired run 1, arm B) and that a long verifier is cut by the wait ceiling; Overturn when says which measured failures the route carries, and the slots.md index cell follows (a new check compares every row's Overturn when across the three surfaces). - agent-teams note: "won no fixture" is replaced by the counted result against the single session: 1 of the 5 fixtures where its diff was captured (F2), none of F3 to F6. - tests/test_rtk_receipt_locators.py: a pin whose commit is present but whose tree is missing (a partial clone) is skipped; a bad path with the tree present still fails. - usage-cost-monitoring: the metering sentence is rewritten from a committed artifact (meter-windows-20261011.json) with a test that recomputes it. The "about 1.9 times" figure was one instant's (2.56 and 1.64 over two fixed 24-hour windows) and the clause that the events leave the advisor out contradicted the advisor run's Stage 0; both are gone. Paired run 5, as the command center directed: routing_5 in paired-runs.json (20 runs, the 12 decisive findings with the runs that found each, costs, hashes), RoutingRunTests recomputing the strict verdict from the rows (Opus stays), the record's addendum paragraph, run 4 marked stopped and run 5 removed from the owed table. No slot row, default or skill text changes. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Answer the Codex re-review and the Claude read of #1005: re-pin the ruling, count the totals in the tolerant floor, state the order of the second trial Answers the command center's Claude first read of #1005 at 5c5eeca (three P2s, five P3s) and Codex's re-review in GitHub of the same head (one P1, four P2s); each was reproduced before the change. - The ruling file the record cites was edited by the command center after the record cited it (CC P2-1): routing_5.ruling and papers.ruling carry the current sha256 and size, the record's paragraph is regenerated from them, and a test that is enabled by PAIRED_PAPER_ROOTS resolves every private paper of the artifact against its recorded hash, so a paper that changes after it is cited is noticed on a host that has it. - The ruling's condition for Sonnet at medium (the change is mechanical, or a second Claude-family read exists) was dropped from the artifact and the paragraph (CC P2-2); it is in both now, with the clause that the allowance is the command center's routing discretion and not something run 5 measured (CC P3-2). - The slot-manifest label named an older slots.json and its list lacked three pairs, so the exact-difference check skipped (CC P2-3): the label is the sha256 of the shipped file and the list is recomputed (background-workflows route and overturn_when, usage-cost-monitoring evidence, and the agent-teams overturn condition edited below). - Codex P1: the second trial was added by the command center after trial 1 had met the registered rule (addendum 3 of the run's papers says so). The paragraph and the decision now state that the two-trial verdict is a ruled extension and not a preregistered outcome, and that an extension after a pass can only move the verdict toward Opus. - Codex P2 and CC P3-1: the tolerant reading also bounds the totals, and the floor printed (0.62) modelled the per-finding clause only; the full reading's floor is 0.56 (exact convolution, checked against enumeration). Both figures are kept and labelled, the correction is recorded in a new private note (hash in papers.correction), and the strict rule's chance with each finding's own pooled rate (0.73, a plug-in calculation) is added. - Codex P2 (workers.md:29): the agent-teams overturn condition names the class the default is scoped to (work where workers must message each other); paired run 1 measured bounded fixes, which fall outside it. - Codex P2 (workers.md:76): the task-list alternative states that it has no completion condition of its own and needs the external done check of loops-completion, as the measured F5 runs used all 8 sessions without declaring done. - Codex P2 and CC P3-3: a blob that is absent in a blobless partial clone skips the locator test and a path that is not in the tree at the pin still fails. - CC P3-4: the meter test's tolerance is 0.01. CC P3-5: the addendum's heading names both dates. Each correction fails its own test when undone alone: 31 single-edit mutants of the run-5 data and text and 7 of the locator check are caught. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Answer the Claude micro and the new Codex threads of #1005 at 425240f: the registered trigger, the time bounds, the basis of the tolerant floor Answers the command center's Claude micro of 425240f (PASS_WITH_P3: P3-1 to P3-7) and Codex's three new threads of the same head (workers.md:61, test_paired_runs_results.py:91, scheduling-supervision.md:9). Each was reproduced before the change: five tests of this commit fail on the data of 425240f (retained output fail-before-micro3.txt). - P3-1: decision.two_trial_verdict_basis no longer says that the decision to run two trials was not fixed before any run. Addendum 1 (paired-routing/addendum-01-ruling-and-extension-20261011.md, sha256 1917091f, line 13, written before any run) fixed the trigger too: after a passing trial 1 the run ends (Sonnet takes review micros, no second trial). The command center ran trial 2 after seeing that result (addendum 3, pinned beside the basis with its path, size and hash), so the two-trial verdict stays a ruled extension and not a preregistered outcome. The paragraph of the record is regenerated from the artifact; the ruling's sidecar correction (sha256 338600a9) is pinned as papers.ruling_correction and the ruling's own bytes are unchanged; ruling.routing says "read or micro" like the ruling. - Codex P1 (test_paired_runs_results.py:91): the tolerant floor's reading and null model are the run's own, registered before any run (addendum 1, lines 17-18: the reading with the clause on the totals, and the noise floors of two identical reviewers who each find a finding with probability p in each run); no third-party source defines them, so the reference is the brute-force enumeration that already fails before the correction. The convolution of the per-finding differences is the one numpy 2.5.3 documents for numpy.convolve ("the sum of two independent random variables is distributed according to the convolution of their individual distributions"); a test cross-checks tolerant_floor against numpy.convolve where numpy is installed, the artifact records both bases (decision.tolerant_floor_basis), and the gated papers test reads lines 17-18 of the retained addendum. - Codex P2 (workers.md:61) and P3-3: the orchestration-frameworks row and the record said "a 45-minute ceiling per run". The registered design has two bounds, a 50-minute process cap for a headless session and a 45-minute cap for the team (paired-runs.json design.caps; preregistration.md line 156, sha256 531758c4), and the Workflow's verifier was cut by the client's own 45-minute wait for background tasks: the stderr of the runs F1-B, F3-B, F4-B and F5-B reads "Background tasks still running 45m after the last turn", after 2779 to 2794 seconds, inside the 50-minute cap. The row (slots.json and workers.md) and the record say so, and test_the_time_bounds_the_row_and_the_record_state_are_the_registered_ones computes the counts from the run rows (4 of 6 Workflow runs, 4 of 6 team runs of which 3 at the team cap). - Codex P2 (scheduling-supervision.md:9): that text is #1001's (commit ba6ab3a). This commit carries the same wording as #1001's third commit (efaa077): GitHub Actions schedules repository jobs, Claude runs inside GitHub are retired (2026-10-10, docs/decisions/2026-10-10-retire-in-github-claude-review.md:5-6, :43), so a scheduled Claude task runs locally. The test of the wording is that commit's. - P3-7: the RTK locator test cites git 2.53.0 git-ls-tree(1), OUTPUT FORMAT, for the entry it splits. The label of claude-native-slot-manifest.json equals the sha256 of slots.json again (relabel last). Its strict test arrives with #1001. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Correct the origin of the 45-minute background wait of paired run 1: the runner's setting, not the client's The command center's Claude micro of #1005 at 49bb649 (P2-1, P3-2, P3-3); each item was reproduced on that head before the change. - P2-1: the previous commit said the Workflow's verifier was cut by "claude -p's own 45-minute wait for background tasks". The runner set CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=2700000 for every headless run (paired/orch_run.py line 137); the client's default is 10 minutes (the retained env-vars.md line 357: "Default: `600000`, or 10 minutes"; headless.md line 77: "after 10 minutes of continuous idle waiting by default"). The decision is unaffected: the runs show exactly what a 45-minute ceiling produces (the four Workflow runs that ended there ran background work for 2,702 to 2,703 seconds, the ceiling being 2,700). One rewording pass over the orchestration-frameworks and background-workflows rows (slots.json), their pages (workers.md) and the record, with the setting in design.caps and the three sources as papers (the runner file, retained and hashed; the two documentation lines, hashed and checked on their lines by the gated test). - The test pins of the old wording are inverted and fail on the data of 49bb649 (fail-before-wave.txt); test_the_background_wait_ceiling_is_the_runners_setting_and_the_clients_default_is_ten_minutes computes the 2,700 seconds from the run rows; 13 single-edit mutants are caught and the control passes (mutations-wave.md). - P3-3: the git-ls-tree quote of the RTK locator test is exact (git 2.53.0, OUTPUT FORMAT and -z). The label of claude-native-slot-manifest.json equals the sha256 of slots.json again (relabel last). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Re-register the files of this pull request and recompute the manifest label after the rebase onto main Mechanical, after the rebase of the seven commits of this pull request onto main at b490361 (the landing of #1001, whose forward commits replace the first head ba6ab3a that this branch was stacked on). - The rebase conflicted in manifests/evidence.json and in claude-native-slot-manifest.json only, at seven of the seven replayed commits. Both were resolved by taking main's side: main already carries #1001's corrected manifest note and its strict label test, and this branch's own rows are re-registered below. - evidence_manifest.py --write and host_receipts.register_file re-register every file this branch adds or changes (12 files and the manifest) at its current bytes; relabel_manifest.py then sets the label of the slot manifest to the sha256 of slots.json (76d5eff8...) and rebuilds the list of differences, last. The note of the manifest keeps #1001's wording (GitHub Actions schedules repository jobs; a scheduled Claude task runs locally). - The diff against main is the diff of this pull request: 14 files, +3708/-57, no file of #1001 or #986. - validate.py and evidence_manifest.py --check pass (11654 files). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scope
tests.test_skill_usage.ManagedMcpDefaultLocationsfor the managed-config path rule of Count only an absolute managed-mcp.json path as a managed MCP config #967 (no production change): the normalisation equalityos.path.abspath(name) == name, which could be removed with every test green (the co-op's GPT read of Count only an absolute managed-mcp.json path as a managed MCP config #967 at3de3dda8, P3), and the Windows side of the rule underntpathon every host, so CI's Python 3.12 executes it (the command center's note on landing Count only an absolute managed-mcp.json path as a managed MCP config #967).9d4c00cd05d79b3fd5074df2e46bacab87ee67ff(main, which has Ground the Claude Code native practice in primary sources and one cross-client mapping #958 landed); two commits,e466b490e5502166eadec79e6ed4c2b199eee99band01fe0aecd939555ded316ceaf375ed90a1fb480f, replayed onto it from the follow-up prepared on Count only an absolute managed-mcp.json path as a managed MCP config #967's landed commit60004bb1d(no conflict).lane:foundationtests/test_skill_usage.py(two tests, two imports) andmanifests/evidence.json(re-registration of that file).Follows #967. Nothing here changes a setting, a hook, a host or
tools/skill-usage/skill_usage.py.SOTA sources
3de3dda8(ACK, one non-blocking P3; verdict sha256015f8308537fb21367254453ec47afe35ee8d6674ec27db728af25eba57ad13d): removing onlyand os.path.abspath(name) == namekept all 20 relevant tests green; for a readable file reached by an absolute spelling with an existingchild/../segment the head keeps--strict-mcp-configand that mutant drops it.ntpath.isabsatLib/ntpath.pyline 87 in 3.12.3 (accepts a leading single separator, so/etc/claude-code/managed-mcp.jsonis absolute and only the equality rejects it, becauseabspathadds a drive and turns the separators) and line 80 in 3.13.16 (rejects it first). On a POSIX hostntpath.abspathis its pure-Python fallback (no_getfullpathname), which is all the cases need and is why the drive-relative case is decided byisabshere.24972e3bc859fab2b46ed4c1e51f7d6130f06d3bd550811a114640de3370d0de,My()at byte 209080002, read 2026-10-10):/Library/Application Support/ClaudeCodeon macOS,C:\Program Files\ClaudeCodeon Windows,/etc/claude-codeotherwise; cited in Fence the headless /skill-doctor run #925 and Count only an absolute managed-mcp.json path as a managed MCP config #967.Evidence-class table
Path)local_integration(real files, every platform)tests.test_skill_usage.ManagedMcpDefaultLocations.test_an_absolute_path_that_is_not_normalised_is_not_a_managed_configosreplaced, for the test, by one whoseisabsandabspatharentpath's and whose file checks answer yes: the Windows documented string counts and the macOS and Linux strings do not; a dot-dot, dot, forward-slash and doubled-separator spelling of the Windows string, a relative name and a drive-relative name each keep every fence; one counted candidate decides the form, by parameter and by the default routelocal_integration(ntpath on a POSIX host; not run on Windows)tests.test_skill_usage.ManagedMcpDefaultLocations.test_the_windows_path_rules_count_the_windows_string_and_not_the_posix_ones, run under CPython 3.12.3 and 3.13.16_readable_system_file, executed in the module's own globals so the mutated body sees the swappedos(the earlier whole-function harness could not): removing only the normalisation equality fails both new tests on 3.12 and 3.13 (the four non-normalised Windows spellings, plus the macOS and Linux strings on 3.12); removing onlyisabsfails the ntpath test through the drive-relative name (equivalent on a real Windows host, whereabspathresolves it); reverting to the presence-only check fails the relative-name, POSIX and Windows tests; removing only the regular-file or the readability check fails the matrixlocal_integration(mutation harness; nothing on disk changes)mutations_p3_source.py, tablespr925/p3-source-mutations-py3.12.3.txtand-py3.13.16.txtin the lane's research recordDeclared test contract change
None: two tests are added and no existing test changes.
Local commands run
A first draft of the mutation table replaced the whole function, which cannot see a swapped
os; it reported theisabsmutant as equivalent and a missingnormpathin the shim looked like a catch. Both were corrected before the commit by executing the mutated source in the module's globals.Decision record
No new record: tests for the rule of #967 and its decision locator, which this PR does not move.
Host evidence
Not applicable: no file under
evidence/hosts/changes.Checklist
permissions: {}and grant each job only what it needs. (none changed)🤖 Generated with Claude Code