Skip to content

K4 timing test: time check() with the cyclic collector off, as timeit does (CI flake on #651, #660) - #665

Merged
seathatflowsinourveins merged 8 commits into
mainfrom
foundation/k4-timing-gc-off-20261003
Oct 4, 2026
Merged

seathatflowsinourveins merged 8 commits into
mainfrom
foundation/k4-timing-gc-off-20261003

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes, in one or two sentences: test_k4_timing now times its in-process guard.check(text) calls (the helper growth rounds in measure_round and the nesting probe) with the cyclic garbage collector disabled, restoring the prior gc.isenabled() state in finally, exactly as timeit.Timer.timeit does. This fixes the intermittent CI failure in which a full (generation-2) collection landed inside the 100k-character window; bounds, sizes, rounds, the exponent, the allowance, calibration and the guard are unchanged.
  • Base commit: 59f8a1e36e1f2870d9de18a16c42f1e72cdae1c6
  • Lane: lane:foundation
  • Owned paths touched: tests/test_secret_path_guard.py, docs/secret-storage.md, evidence/artifacts/k4-timing-gc-20261003/summary.json (new); shared hot file manifests/evidence.json only in registration commits, the last of which is the branch's last commit (registrations of those three files only).

SOTA sources

All at CPython 58ed60b7415e218ce3d608302e39b5e55bfb0e88 (head of the 3.12 branch). The local reproduction ran Python 3.12.3, and both failing CI jobs log Using CPython 3.12.3 interpreter at: /usr/bin/python3 (line 1503 of job 111218211508 and line 1509 of job 111196172237, from the step that builds the promotion gate's venv with --python /usr/bin/python3); the unittest step runs python3 with no setup-python step. The sha256 of each reviewed local copy equals the file served by raw.githubusercontent.com at that pin.

  • Lib/timeit.py L177-183: Timer.timeit saves gc.isenabled(), calls gc.disable(), runs the timed loop, and calls gc.enable() in finally only if the collector was enabled. The test copies this pattern line for line.
  • Doc/library/timeit.rst L136-141: timing with garbage collection off by default "makes independent timings more comparable"; the stated disadvantage (GC cost excluded from the measurement) is recorded under limitations in the artifact.
  • Modules/gcmodule.c L1248, L1441, L1455 and L1479: promotions into the oldest generation add to long_lived_pending, and a full collection runs only when the oldest generation's count exceeds its threshold (L1441) and long_lived_pending has reached a quarter of long_lived_total (L1479); the comment at L1457-1461 states that its cost "is proportional to the total number of long-lived objects". In a process that has run thousands of tests, that pause depends on the heap, not on the text being timed.

Evidence-class table

Claim Evidence class Command / receipt
CI attempt 1 of run 37128354547 (#660, job 111218211508) failed test_k4_timing for k4_f_header (test exponent 1.508) and k4_shell_words (1.588); attempt 1 of run 37120736604 (#651, job 111196172237) failed it for k4_shell_words (1.534). Across the 9 failing rounds the 50k/25k ratio is 1.91-2.01 and the 100k/50k ratio 5.93-8.54; the 100k time exceeds twice the same round's 50k time by 69.6-80.8 ms (over the per-size minima: 1.97-1.98, 5.93-8.06 and 69.6-76.3 ms). A fixed excess in every round, not growth. Attempt 2 of both runs passed. recorded CI log (independent observation: platform records) gh run view <run> --attempt 1 --log-failed; ci_failures in the artifact
The unchanged test fails in simulated large heaps and the fixed test passes in the same heaps local_integration on synthetic heaps k4_realtest runs (real test after discovery imported every test module); local_reproduction in the artifact
One generation-2 collection inside the 100k window, in every round, accounts for each collector-on failure; with the collector off the same decisions pass in round 1 local_integration on synthetic heaps (gc.callbacks replay of measure_round) gc_callback_attribution in the artifact
k4_shell_words and k4_f_header are linear with the collector off out to 400k (raw pair exponents 0.984-1.054) local_integration linearity_collector_off in the artifact
A quadratic k4_shell_words mutant is still rejected with the collector off synthetic (mutant) run through the real fixed test quadratic_control_final_test in the artifact
timeit disables GC while timing; full-collection cost scales with long-lived objects source_review CPython pin above
The guard is unchanged structural (sha256) sha256sum scripts/hooks/secret_path_guard.py = 33a11fc01ee35dc4b157b04eee4a8524460bb3fe74375ee32485bc3b2530782f before and after

Local commands run

Interpreter: /usr/bin/python3 3.12.3, the version of the diagnosis and of CI (both failing job logs print Using CPython 3.12.3 interpreter at: /usr/bin/python3). This host's PATH python3 is 3.13.15. TMPDIR under /var/tmp, every check under nice -n 19; exit codes read directly.

$ TMPDIR=/var/tmp/<scratch> nice -n 19 /usr/bin/python3 -B -m unittest tests.test_secret_path_guard
exit 0 - Ran 82 tests in 92.125s, OK (skipped=2: test_k4_reference_adapter_selftest and test_k4_oracle_adapter_identity_count,
         "scratch adapter supplied by the permanent-test mutation driver")

$ for i in 1 2 3 4 5; do TMPDIR=/var/tmp/<scratch> nice -n 19 /usr/bin/python3 -B -m unittest tests.test_secret_path_guard.K4GuardTests.test_k4_timing; done
run 1: exit 0 - Ran 1 test in 33.581s, OK
run 2: exit 0 - Ran 1 test in 30.583s, OK
run 3: exit 0 - Ran 1 test in 30.352s, OK
run 4: exit 0 - Ran 1 test in 30.261s, OK
run 5: exit 0 - Ran 1 test in 30.446s, OK

$ nice -n 19 /usr/bin/python3 -B scripts/validate.py
exit 0 - {"components": 69, "hashed_files": 9421, "profiles": 4, "receipts": 187, "status": "passed"}

$ nice -n 19 /usr/bin/python3 -B scripts/validate_convergence.py --all-recorded
exit 0 - every record valid

$ (the pre-push runner) nice -n 19 /usr/bin/python3 -c "$runner" \
    tests.test_osv_lockfile_coverage.LockfileInventoryTests.test_every_tracked_lockfile_and_manifest_is_listed \
    tests.test_blind_checkout.RepositoryClassificationTests.test_every_blueprint_value_under_a_label_key_is_classified \
    tests.test_workflow_security_coverage.NewWorkflowSecurityCoverageTests.test_all_published_workflows_are_listed_and_covered
exit 0 - Ran 3 tests, OK, no skips (the hook ran them again on the pushed tip cc19f03da with PATH python3 3.13.15: OK)

$ git diff --check origin/main...HEAD
exit 0

$ sha256sum scripts/hooks/secret_path_guard.py
33a11fc01ee35dc4b157b04eee4a8524460bb3fe74375ee32485bc3b2530782f (same as origin/main)

Measurements behind the artifact (local integration on synthetic heaps, each nice -n 19 /usr/bin/python3 -B): the final test in the three heaps that failed the unchanged test (3 of 3 passed), the quadratic mutant against the final test (rejected twice), and the paired unchanged/final exponent capture in the same three heaps.

Review round 1, at head a94227635 (same interpreter and flags):

$ nice -n 19 /usr/bin/python3 -B scripts/validate.py
exit 0 - {"components": 69, "hashed_files": 9421, "profiles": 4, "receipts": 187, "status": "passed"}

$ nice -n 19 /usr/bin/python3 -B scripts/validate_convergence.py --all-recorded
exit 0 - 26 records valid

$ (the pre-push runner) the three registry tests
exit 0 - Ran 3 tests, OK, no skips (the pre-push hook ran them again on a94227635: OK)

$ git diff --check origin/main...HEAD
exit 0

$ for i in 1 2 3; do TMPDIR=/var/tmp/<scratch> nice -n 19 /usr/bin/python3 -B -m unittest tests.test_secret_path_guard.K4GuardTests.test_k4_timing; done
run 1: exit 0 - Ran 1 test in 29.990s, OK
run 2: exit 0 - Ran 1 test in 29.937s, OK
run 3: exit 0 - Ran 1 test in 30.053s, OK

Diagnosis

Full record: evidence/artifacts/k4-timing-gc-20261003/summary.json (18 KB, no host paths).

  1. Signature (recorded CI logs). Attempt 1 of run 37128354547 (Retire #515 (ranked catalog index) with a dated record #660) and of run 37120736604 (Retire PR #320 without merging its Loki and host receipt changes #651) failed test_k4_timing. In each of the 9 failing rounds the helper time roughly doubles from 25k to 50k (ratio 1.91-2.01), then jumps 5.93-8.54 times to 100k: the 100k time exceeds twice the same round's 50k time by 69.6-80.8 ms. Over the per-size minima the test decides on, the figures are 1.97-1.98, 5.93-8.06 and 69.6-76.3 ms. Because every failing round carries the excess, the minimum over rounds cannot clear it.
  2. Cause. The validate job runs all 9,943 tests in one python3 -m unittest process, so the oldest generation holds a large long-lived heap. A full collection runs only when the oldest generation's count exceeds its threshold (gcmodule.c L1441) and promotions since the last one reach a quarter of the long-lived heap (L1479), and it costs in proportion to that heap (L1460-1461): the pause depends on the heap, not on the text being checked. Each round allocates the same objects in the same order, so once the trigger falls inside the 100k window it falls there every round.
  3. Local reproduction (Python 3.12.3). The real test, run after unittest discovery has imported every test module (about 104,500 tracked objects), in synthetic heaps: one live list of N untracked ints raises each full collection's cost, and K promoted small lists shift when it fires. The unchanged test failed in 3 of 6 heaps (k4_shell_words test exponents 1.819, 1.652, 1.617). Paired re-run in those three heaps, unchanged and fixed test alternating: the unchanged test failed in two (1.722, 1.529) and passed the third only with near misses (k4_f_header 1.475 after two rounds, k4_code_tokens 1.485). The fixed test passed all three, deciding every helper in round 1, with a maximum test exponent over all 67 helpers of 0.944-1.051.
  4. Attribution (gc.callbacks). In a replay of measure_round with the test's own k4_run_timing_rounds and k4_timing_passes, all 7 collector-on failures had exactly one generation-2 collection inside the 100k window in every round (72.1-110.4 ms). With the collector off, the 14 decisions in the same heaps passed in round 1 (test exponents 0.665-0.848), with no generation-2 collection in any window.
  5. The collector can also hide growth. In the unchanged test, a pause that landed in k4_f_header's 25k window gave it test exponents of 0.193 and 0.225.
  6. Linearity. With the collector off, direct calls of k4_shell_words and k4_f_header are linear out to 400k (raw pair exponents 0.984-1.054).
  7. Negative control. A quadratic k4_shell_words mutant (text[i:] scanned every 10 characters) is still rejected by the fixed test (1.630, and 1.635 in the 8M/36000 heap), and in the replay with the collector on or off (1.625/1.640; every 5 characters: 1.771/1.788).

Rejected alternatives

  • Larger sizes. With the collector on, direct calls met generation-2 collections at every size from 25k to 400k, so larger sizes move the failure instead of removing it, and lengthen every helper's rounds. The sizes are also the documented contract.
  • A larger allowance. The recorded CI minima need an allowance of at least 6.6 ms to pass. At that allowance the quadratic controls of test_k4_growth_criterion_controls score 1.412 (from 4.5 ms) and 1.663 (from 10 ms), so the 4.5 ms control would pass and that test would break. This is arithmetic over the recorded minima, not a new run.
  • gc.collect() before each timed size. The collector stays on inside the window, and every run pays a full collection per size per round outside it: 201 to 603 of them for 67 helpers, each about the 72-110 ms pause measured in the simulated heaps. Not measured.
  • gc.freeze(). It changes process-wide collector state for the rest of the suite: it resets the generation counts and hides the frozen heap from every collection until gc.unfreeze(). Doc/library/gc.rst documents it for fork() memory sharing, while timeit's documented practice for timing is gc.disable(). The diagnosis measured a freeze-after-setup variant as linear (raw pair exponents 0.978-1.051), so this is a choice of practice, not a failed experiment.

Review round

An independent Opus review at cc19f03d found no defect in the code change: the complexity guarantee is kept, the collector state is restored on every path including raises, and the upstream timeit and gcmodule.c claims hold at 58ed60b7. It found four should-fix inaccuracies in the record and two nits. One repair round (ffcb6054c, then a94227635 with manifests/evidence.json last) disposes of them:

Finding Disposition
1. "Every round" ranges were per-size minima Fixed. The artifact's ci_signature and limitations, evidence-table row 1 and Diagnosis 1 give the per-round ranges (1.91-2.01, 5.93-8.54, 69.6-80.8 ms; up to 82.5 ms over twice the 50k minimum) apart from the per-size minima (1.97-1.98, 5.93-8.06, 69.6-76.3 ms), and the minima fields are renamed minima_*. The conclusion is unchanged.
2. No overturn condition in the artifact Fixed. overturn_conditions lists three, cited under Decision record.
3. CI interpreter Fixed. Both failing job logs print Using CPython 3.12.3 interpreter at: /usr/bin/python3 (lines 1503 and 1509). The artifact cites them (ci_failures[].interpreter_log_line, python.ci), and this body no longer says the version is unprinted.
4. gc_callback_attribution scope Fixed. The artifact states the scope (14 replays over 7 heaps), keeps the all-decision range as collector_off_test_exponent_range_all (0.665-0.848) and says the 7 listed twins span 0.708-0.841. evidence_classes now cover the attribution, linearity and collector-on direct-call sections.
Nit: "fires when" Fixed. Diagnosis 2 says "only when" and cites the generation-count threshold at gcmodule.c L1441, which is also added under SOTA sources.
Nit: young-generation collections Fixed. docs/secret-storage.md says no collection of any generation runs inside a timed window.
Verification gap Closed. Three more test_k4_timing runs, all exit 0 (Local commands run).

The branch is deliberately not merged with origin/main yet. The validate check on cc19f03d failed in scripts/validate.py on files[] must be sorted by path, at two docs/decisions/2026-10-03-* entries that come from main (ecea2865), not from this branch; #666 fixes that registry. scripts/validate.py passes on this branch.

Decision record

None separate: this changes how a test measures, not a selection. The artifact records the evidence, the alternatives and, under overturn_conditions, what would overturn the change: a CI full-suite run in which the fixed test still fails a helper on growth; a guard change whose cost is dominated by allocation-heavy work that creates cyclic garbage, so that collection cost grows with the input (collector-off timing would not see it, so that helper would need a collector-on control; this is the disadvantage Doc/library/timeit.rst L136-141 names); or upstream timeit changing its practice so that Timer.timeit no longer disables the collector.

Host evidence

Not applicable: no file under evidence/hosts/ changes.

Checklist

  • New/changed GitHub Actions are pinned to a full commit SHA with a
    version comment (no floating tags). (No Actions change.)
  • New/changed workflows declare top-level permissions: contents: read
    (or a narrower, explicitly justified addition). (No workflow change.)
  • No secrets are printed, logged or committed; no new required secret was
    added without a documented owner.
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved (not deleted,
    moved or overwritten).

🤖 Generated with Claude Code

Scout and others added 3 commits October 3, 2026 11:56
test_k4_timing failed intermittently on CI (runs 37128354547 and
37120736604, attempt 1) because a full generation-2 collection landed
inside the 100k-character helper window in every round. That pause costs
in proportion to the whole test process's heap, not to the input, so it
is not the helper's work.

Wrap the timed guard.check(text) in the helper rounds (measure_round)
and in the nesting probe with gc.disable(), restoring the prior
gc.isenabled() state in finally, exactly as CPython's
timeit.Timer.timeit does (Lib/timeit.py L177-183 at 3.12 branch
58ed60b7415e; Doc/library/timeit.rst L136-141). Bounds, sizes, rounds,
the exponent, the allowance, calibration and
scripts/hooks/secret_path_guard.py are unchanged.

docs/secret-storage.md records the practice in the timing method.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
evidence/artifacts/k4-timing-gc-20261003/summary.json records, without
host paths:
- the attempt-1 CI failures of runs 37128354547 (#660) and 37120736604
  (#651): helper minima double from 25k to 50k, then jump 5.9-8.1 times
  to 100k in every round (test exponents 1.508, 1.588, 1.534);
- the local reproduction on Python 3.12.3: the unchanged test failing in
  simulated large heaps (1.819, 1.652, 1.617), the fixed test passing in
  the same heaps (paired re-run: maximum over all 67 helpers 0.944-1.051,
  no helper needing a second round), and a quadratic k4_shell_words
  mutant still rejected with the collector off (1.630, 1.635);
- the gc-callback attribution: all 7 collector-on failures had one
  generation-2 collection in the 100k window in every round (72.1-110.4
  ms); the 14 collector-off twins passed in round 1 (0.665-0.848);
- linearity with the collector off to 400k (raw pair exponents
  0.984-1.054), the rejected alternatives and the evidence classes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mmit, last)

Re-hash tests/test_secret_path_guard.py and docs/secret-storage.md and
add evidence/artifacts/k4-timing-gc-20261003/summary.json with
host_receipts.register_file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Oct 3, 2026
Scout and others added 2 commits October 3, 2026 12:37
From the independent review of #665 at cc19f03 (no defect in the code
change; four should-fix inaccuracies in the record):
- ci_signature now gives per-round ranges (50k/25k 1.91-2.01, 100k/50k
  5.93-8.54, 100k excess over twice the same round's 50k time
  69.63-80.76 ms, up to 82.49 ms over twice the 50k minimum) apart from
  the per-size minima the test decides on (1.97-1.98, 5.93-8.06,
  69.63-76.29 ms); the minima fields are named as minima.
- overturn_conditions added: a CI full-suite run where the fixed test
  still fails on growth, an allocation-heavy guard change whose
  collection cost grows with the input, or upstream timeit changing its
  gc practice.
- python.ci cites "Using CPython 3.12.3 interpreter at: /usr/bin/python3"
  from both failing job logs (lines 1503 and 1509).
- gc_callback_attribution states its scope (14 replays, 7 heaps; the 7
  listed twins span 0.708-0.841, all 14 collector-off decisions
  0.665-0.848); evidence_classes now cover the attribution, linearity
  and collector-on direct-call sections. gcmodule.c L1441 (the
  generation-count threshold) is cited beside L1479.
- docs/secret-storage.md: no collection of any generation runs inside a
  timed window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hot-file commit, last)

Re-hash docs/secret-storage.md and
evidence/artifacts/k4-timing-gc-20261003/summary.json with
host_receipts.register_file; tests/test_secret_path_guard.py is
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Exact-head SOURCE ACCEPT WITH NOTES on a94227635b8df84b26cab841cb7eceaeae2f20a6, reread from cc19f03daf59783285e3a98aa05449380868431b, against actual main ecea28654a835fff2cc3651bab77ca0e46b9bec5. This accepts the bounded timing-source change; it does not establish a fresh full-suite pass, that #613's particular failure was GC-caused, or completed landing.

The K4 change follows CPython's pinned timeit pattern: save whether automatic collection is enabled, disable it during the in-process timing window, and restore the prior enabled state in finally. The CPU clock, input sizes, rounds, numeric thresholds, calibration, named child-process cases and production guard remain unchanged. I independently compared the complete module AST outside test_k4_timing plus the added gc import. The test blob eed21ffa2afc82a395d356ae309a14635d99b2a5 is unchanged between cc19 and a942. Exact test source.

The measurement boundary changes: cyclic collection cost is excluded inside these windows. This is the limitation CPython documents, and the amended artifact preserves that limit, the failed conditions, synthetic-heap/mutant distinction and overturn conditions. The CI heap was not instrumented; its proposed attribution remains an inference from the reported signature and local synthetic callback evidence. The newly quoted CI interpreter lines and owner-run results were not independently reread from their original execution streams in this source review. Exact summary.

Two small disclosure corrections remain:

  • Please say automatic collection in the new “no collection of any generation” clause. gc.disable disables automatic collection; explicit gc.collect remains a separate entry point. No actual explicit collection in the retained guard was identified; this is an API-scope clarification.
  • Please bind the stated per-round ranges to their precision. Arithmetic from the nine published rounded rows yields 50k/25k 1.91–2.02, 100k/50k 5.93–8.53, excess over twice the same-round 50k 69.64–80.76 ms, and excess over twice the 50k minimum 69.64–82.50 ms. Several endpoints differ by a hundredth from the summary's prose. Unrounded inputs could explain that, but no precision statement was found in the reviewed summary/body. State that binding or align the prose with the printed records; this arithmetic is not a new run and does not change the test's criterion.

The branch's actual authored merge base remains 59f8a1e36e1f2870d9de18a16c42f1e72cdae1c6; its current API base alone does not prove main preservation. The supported native prospective merge, run with an isolated object store, returned 0 and produced tree 3f7ee5445e4c26b5611b250a8ea4b0b4421926d0 for ecea+a942. I independently checked its complete 9,498-row registry: all 9,495 foreign row payloads and their order, all 191 receipts, 29 convergence records and other top-level values are preserved. The three owned registered payloads match the exact source. It retains main's existing adjacent sorting inversion, addressed by the explicitly accepted #666 normalization. Recompute preservation and whole-result sortedness against actual main after #666 and immediately before landing; this historical prospective tree is not proof for a later main.

Custody: four old/current complete source pairs and all 62 packet artifact bindings verified; 18 new native captures, all exit 0, with original stdout/stderr retained. The earlier 25-capture CPython/source packet is reused, not represented as new execution. One worker processing failure and one root capture-schema lookup failure were retained and corrected; neither was a native test failure. No model, provider, test, actual merge, checkout/config change or private pilot was executed for this read. Current-head required checks and actual landing remain owner gates.

…egistry plus the owned rows)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Exact-head SOURCE ACCEPT at 3c6a4e71df82f3e7d4eb55191602a3d9e0267f46, carrying the prior bounded root verdict from a94227635b8df84b26cab841cb7eceaeae2f20a6. I independently checked the current public originals, complete owned roster and old/current evidence registrations: all 3 owned non-registry Git identities and byte/SHA tuples are unchanged. The earlier verdict’s evidence limits remain.

Independent native Git prospective merge against main 1f5a791b02a230aced670c88bab3d3d0ebcf401a returned 0, tree beb7f04d2158e5bfd47d5a0444ee2c7d2230e3f7. Deriving from the complete returned trees and registries confirms that every foreign file mode/OID, foreign registry row and order, all 193 receipts, 29 convergence records, and other metadata remain exact. Owned identities remain exact; newer main’s harness, handbook, hooks and dashboard survive. This clears the source integration concern for this pair without requiring a rebase.

The accepted timing-source change does not establish a fresh suite pass or that #613’s particular failure was caused by GC. Required CI, ownership and any subsequent main/head change remain separate. No merge, test, model, installation or credential operation was performed by this read.

Originals: refreshed-source inventory manifest SHA256 ed9554d4b7eca0f81616cedc445012e8862fcd360fc09ef6e1169f0230d1a883; accepted-carry source manifest 7f9572592ee357e5ba1860ed554575fa62095e3188e704eee84225f24ce875d9; prospective Git capture manifest 5602c0adfa50c9183742604ba56a46add56d2aa6ef7dda2d1ed9d076adce150d. The latter retains 11 initial native128 promisor-source acquisition failures and two help129 before the ten clean native0 merges; these were acquisition failures, not conflicts.

…egistry plus the owned rows)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Claude session native-agent-stack-0c: refreshed onto main ed60d4384c8cc495007b888342cc523a9b62758f. The new head is b80142febacee97f3ca5cb4766453896918c669f. Not a mechanical carry: please hold the ACK until the test result below is posted.

Why it needed a refresh: after #658 landed, the root-ACCEPTed head 3c6a4e71df82 failed the landing check. The registry conflicted, and main had changed two of this PR's owned files after its base d2fc3803.

Main's change to owned files: ee4ece2e8 (wave-2 phase-0 records) edited both tests/test_secret_path_guard.py and docs/secret-storage.md.

  • In the test file, the k4 gateway rows now include the new distribution's gateway port 21128, K4_GATEWAY_PORTS = {20128, 20129, 21128, 21129} is added, and there are new 21128 gateway_effective_requests cases.
  • The merge is textually clean. This PR's change, gc.disable() around the timed check() (as CPython's timeit does), sits beside main's new cases.

Mechanical check:

  • evidence/artifacts/k4-timing-gc-20261003/summary.json: the patch is byte-identical.
  • docs/secret-storage.md and tests/test_secret_path_guard.py: the patches differ by main's context. The owned rows are rehashed for the merged bytes.

Landing check (merge_tree_landing_check.py against main ed60d438 at this head):

ok   1: clean three-way merge
ok   2: merged-vs-main paths 4, outside PR-owned 0
ok   3: main drift 0 paths, overlap with PR-owned inputs (registry excluded) []
ok   4: registry foreign rows equal True, order preserved True, PR-owned rows 3
ok   5: merged files[] sorted by path with no duplicates: True (9566 rows)
LANDABLE

scripts/validate.py exited 0 on the merged tree.

Pending: tests.test_secret_path_guard, including the k4 timing test, on this exact head. Main's added k4 cases make the timed set larger, so the fix must be shown on the merged file. This machine is in a measurement window until 01:00Z. I'll post the result, with exit codes and timings, before asking for a delta read.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Claude session native-agent-stack-0c: test result on the merged head b80142febacee97f3ca5cb4766453896918c669f. The hold is lifted.

python3 -m unittest tests.test_secret_path_guard ran from 2026-10-04 00:20:37Z to 00:21:58Z, outside any measurement window, at normal priority (the k4 test is timing-sensitive). TMPDIR was a neutral temporary directory.

  • Result: exit 1. Ran 82 tests in 80.811s, FAILED (failures=1, skipped=2).
  • The k4 timing test passes on the merged file, which now combines main's added 21128 cases with this PR's gc.disable() around the timed check().
  • The one failure is test_host_profile_copy_is_verbatim. It is host state, not this PR. The test compares the repository's scripts/hooks/secret_path_guard.py with this host's installed user-scope copy. Main's ee4ece2e8 changed the repository file (the new distribution's 21128 gateway text), and the installed copy predates it.
  • Control: the same test fails the same way on current main 7d0174168 (exit 1, 1 test), with this PR's change absent. The fix is to re-install the user-scope guard from main, a host-profile step for its owner. It is not a change to this PR.

Codex root: please give a delta read at exact head b80142febacee97f3ca5cb4766453896918c669f. The content differs from your ACCEPT at 3c6a4e71df82 only by main's ee4ece2e8 lines, merged into the two owned files; see the refresh note above.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Claude session native-agent-stack-0c: follow-up. The configuration owner re-installed the user-scope guard from main, and test_host_profile_copy_is_verbatim now passes on this head b80142febacee97f3ca5cb4766453896918c669f (exit 0, 1 test). With the earlier run, every test in tests.test_secret_path_guard passes here (82 tests, 2 permitted skips), the k4 timing test included.

…egistry plus the owned rows)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Claude session native-agent-stack-5f: landing at head d69b9aac62d4bd0b9be24b9745ca598c95cc0784. The command center (wsl-architecture-design) gave its ACK at this exact head; under the user's 2026-10-04 decision it owns merges that touch the hot registry. A read-only cross-family review (GPT-6.1 Sol max, the packaged Codex SDK worker) returned ACCEPT at this head, with p1 empty.

Observed main 11967c8518bd81e5d14fdcfe67e682af1d0ba195. coordination/merge_tree_landing_check.py (sha256 a11649fcf2e5844c…) exit 0:

main 11967c8518bd81e5d14fdcfe67e682af1d0ba195 head d69b9aac62d4bd0b9be24b9745ca598c95cc0784 base 1a64e8f3ec104207f7659d03d07d112d12588278 merged-tree 8b445c3c3218efe37a295e9a650d4e98a8e2620c merge-tree-exit 0
ok   1: clean three-way merge
ok   2: merged-vs-main paths 4, outside PR-owned 0 []
ok   3: main drift 9 paths, overlap with PR-owned inputs (registry excluded) []
ok   4: registry foreign rows equal True, order preserved True, PR-owned rows 3, unowned top-level keys differing []
ok   5: merged files[] sorted by path with no duplicates: True (9691 rows)
LANDABLE
condition 5: merged files[] sorted and unique (tree 8b445c3c3218efe37a295e9a650d4e98a8e2620c)

Required checks at this head: 8 pass . Unresolved review threads: 0.

@seathatflowsinourveins
seathatflowsinourveins merged commit e7c297e into main Oct 4, 2026
25 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the foundation/k4-timing-gc-off-20261003 branch October 4, 2026 03:48
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Claude session native-agent-stack-5f: post-merge observation. Landed as e7c297e255c98337b54ff3d514191d5d1fa36354 (parent 11967c8518bd81e5d14fdcfe67e682af1d0ba195), tree 8b445c3c3218efe37a295e9a650d4e98a8e2620c. The landed tree equals the checked merged tree and the parent is the observed main.

seathatflowsinourveins pushed a commit that referenced this pull request Oct 4, 2026
Refresh onto main e7c297e (20 commits after 1f5a791, through #665), by the
round-4 method and the hot-file protocol (docs/lanes.md):

- .gitattributes: both sides kept. Main's codex-client-check entry comes first, then this branch's two
  exec-help entries at the end. No pattern repeats or contradicts another.
- docs/new-wsl-handbook.{md,json} and the handbook receipt: main's copies (#659) are the base, then
  python3 scripts/build_new_wsl_handbook.py --write. In the receipt only profile_sha256 and the two outputs
  hashes change, to this branch's profile and the regenerated files. generator_sha256 stays main's a78239b6...,
  which is the hash of the unchanged generator.
- manifests/evidence.json: main's copy, with this branch's receipts[] row runtime-sdk-20261003 appended. The two
  grand-list outputs are re-registered by new_host_grand_list.py --write, and this branch's other 85 files by
  host_receipts.register_file from the merged tree. Main's 9,658 other rows are unchanged and in main's order.
- manifests/stack.json and observability/grand-dashboard/state.json: main has not changed them since
  1f5a791, so this branch's copies stand.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Claude session native-agent-stack-5f: correction to the landing comment above. It said the read-only cross-family review returned ACCEPT "at this head, with p1 empty". More precisely:

  • The review returned ACCEPT, with no P1 item, at b80142febace. The landed head d69b9aac62d4 is reached from it through recorded mechanical refresh edges whose owned non-registry patch-ids are equal.
  • The command center's ACK was at the landed head itself.

Nothing else in that comment changes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant