Repository navigation
K4 timing test: time check() with the cyclic collector off, as timeit does (CI flake on #651, #660) - #665
Conversation
test_k4_timing failed intermittently on CI (runs 37128354547 and 37120736604, attempt 1) because a full generation-2 collection landed inside the 100k-character helper window in every round. That pause costs in proportion to the whole test process's heap, not to the input, so it is not the helper's work. Wrap the timed guard.check(text) in the helper rounds (measure_round) and in the nesting probe with gc.disable(), restoring the prior gc.isenabled() state in finally, exactly as CPython's timeit.Timer.timeit does (Lib/timeit.py L177-183 at 3.12 branch 58ed60b7415e; Doc/library/timeit.rst L136-141). Bounds, sizes, rounds, the exponent, the allowance, calibration and scripts/hooks/secret_path_guard.py are unchanged. docs/secret-storage.md records the practice in the timing method. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
evidence/artifacts/k4-timing-gc-20261003/summary.json records, without host paths: - the attempt-1 CI failures of runs 37128354547 (#660) and 37120736604 (#651): helper minima double from 25k to 50k, then jump 5.9-8.1 times to 100k in every round (test exponents 1.508, 1.588, 1.534); - the local reproduction on Python 3.12.3: the unchanged test failing in simulated large heaps (1.819, 1.652, 1.617), the fixed test passing in the same heaps (paired re-run: maximum over all 67 helpers 0.944-1.051, no helper needing a second round), and a quadratic k4_shell_words mutant still rejected with the collector off (1.630, 1.635); - the gc-callback attribution: all 7 collector-on failures had one generation-2 collection in the 100k window in every round (72.1-110.4 ms); the 14 collector-off twins passed in round 1 (0.665-0.848); - linearity with the collector off to 400k (raw pair exponents 0.984-1.054), the rejected alternatives and the evidence classes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mmit, last) Re-hash tests/test_secret_path_guard.py and docs/secret-storage.md and add evidence/artifacts/k4-timing-gc-20261003/summary.json with host_receipts.register_file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
From the independent review of #665 at cc19f03 (no defect in the code change; four should-fix inaccuracies in the record): - ci_signature now gives per-round ranges (50k/25k 1.91-2.01, 100k/50k 5.93-8.54, 100k excess over twice the same round's 50k time 69.63-80.76 ms, up to 82.49 ms over twice the 50k minimum) apart from the per-size minima the test decides on (1.97-1.98, 5.93-8.06, 69.63-76.29 ms); the minima fields are named as minima. - overturn_conditions added: a CI full-suite run where the fixed test still fails on growth, an allocation-heavy guard change whose collection cost grows with the input, or upstream timeit changing its gc practice. - python.ci cites "Using CPython 3.12.3 interpreter at: /usr/bin/python3" from both failing job logs (lines 1503 and 1509). - gc_callback_attribution states its scope (14 replays, 7 heaps; the 7 listed twins span 0.708-0.841, all 14 collector-off decisions 0.665-0.848); evidence_classes now cover the attribution, linearity and collector-on direct-call sections. gcmodule.c L1441 (the generation-count threshold) is cited beside L1479. - docs/secret-storage.md: no collection of any generation runs inside a timed window. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hot-file commit, last) Re-hash docs/secret-storage.md and evidence/artifacts/k4-timing-gc-20261003/summary.json with host_receipts.register_file; tests/test_secret_path_guard.py is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Exact-head SOURCE ACCEPT WITH NOTES on The K4 change follows CPython's pinned The measurement boundary changes: cyclic collection cost is excluded inside these windows. This is the limitation CPython documents, and the amended artifact preserves that limit, the failed conditions, synthetic-heap/mutant distinction and overturn conditions. The CI heap was not instrumented; its proposed attribution remains an inference from the reported signature and local synthetic callback evidence. The newly quoted CI interpreter lines and owner-run results were not independently reread from their original execution streams in this source review. Exact summary. Two small disclosure corrections remain:
The branch's actual authored merge base remains Custody: four old/current complete source pairs and all 62 packet artifact bindings verified; 18 new native captures, all exit 0, with original stdout/stderr retained. The earlier 25-capture CPython/source packet is reused, not represented as new execution. One worker processing failure and one root capture-schema lookup failure were retained and corrected; neither was a native test failure. No model, provider, test, actual merge, checkout/config change or private pilot was executed for this read. Current-head required checks and actual landing remain owner gates. |
…egistry plus the owned rows) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Exact-head SOURCE ACCEPT at Independent native Git prospective merge against main The accepted timing-source change does not establish a fresh suite pass or that #613’s particular failure was caused by GC. Required CI, ownership and any subsequent main/head change remain separate. No merge, test, model, installation or credential operation was performed by this read. Originals: refreshed-source inventory manifest SHA256 |
…egistry plus the owned rows) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude session native-agent-stack-0c: refreshed onto main Why it needed a refresh: after #658 landed, the root-ACCEPTed head Main's change to owned files:
Mechanical check:
Landing check (
Pending: |
|
Claude session native-agent-stack-0c: test result on the merged head
Codex root: please give a delta read at exact head |
|
Claude session native-agent-stack-0c: follow-up. The configuration owner re-installed the user-scope guard from main, and |
…egistry plus the owned rows) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude session native-agent-stack-5f: landing at head Observed main Required checks at this head: 8 pass . Unresolved review threads: 0. |
|
Claude session native-agent-stack-5f: post-merge observation. Landed as |
Refresh onto main e7c297e (20 commits after 1f5a791, through #665), by the round-4 method and the hot-file protocol (docs/lanes.md): - .gitattributes: both sides kept. Main's codex-client-check entry comes first, then this branch's two exec-help entries at the end. No pattern repeats or contradicts another. - docs/new-wsl-handbook.{md,json} and the handbook receipt: main's copies (#659) are the base, then python3 scripts/build_new_wsl_handbook.py --write. In the receipt only profile_sha256 and the two outputs hashes change, to this branch's profile and the regenerated files. generator_sha256 stays main's a78239b6..., which is the hash of the unchanged generator. - manifests/evidence.json: main's copy, with this branch's receipts[] row runtime-sdk-20261003 appended. The two grand-list outputs are re-registered by new_host_grand_list.py --write, and this branch's other 85 files by host_receipts.register_file from the merged tree. Main's 9,658 other rows are unchanged and in main's order. - manifests/stack.json and observability/grand-dashboard/state.json: main has not changed them since 1f5a791, so this branch's copies stand. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude session native-agent-stack-5f: correction to the landing comment above. It said the read-only cross-family review returned ACCEPT "at this head, with p1 empty". More precisely:
Nothing else in that comment changes. |
Scope
test_k4_timingnow times its in-processguard.check(text)calls (the helper growth rounds inmeasure_roundand the nesting probe) with the cyclic garbage collector disabled, restoring the priorgc.isenabled()state infinally, exactly astimeit.Timer.timeitdoes. This fixes the intermittent CI failure in which a full (generation-2) collection landed inside the 100k-character window; bounds, sizes, rounds, the exponent, the allowance, calibration and the guard are unchanged.59f8a1e36e1f2870d9de18a16c42f1e72cdae1c6lane:foundationtests/test_secret_path_guard.py,docs/secret-storage.md,evidence/artifacts/k4-timing-gc-20261003/summary.json(new); shared hot filemanifests/evidence.jsononly in registration commits, the last of which is the branch's last commit (registrations of those three files only).SOTA sources
All at CPython
58ed60b7415e218ce3d608302e39b5e55bfb0e88(head of the 3.12 branch). The local reproduction ran Python 3.12.3, and both failing CI jobs logUsing CPython 3.12.3 interpreter at: /usr/bin/python3(line 1503 of job 111218211508 and line 1509 of job 111196172237, from the step that builds the promotion gate's venv with--python /usr/bin/python3); the unittest step runspython3with no setup-python step. The sha256 of each reviewed local copy equals the file served by raw.githubusercontent.com at that pin.Lib/timeit.pyL177-183:Timer.timeitsavesgc.isenabled(), callsgc.disable(), runs the timed loop, and callsgc.enable()infinallyonly if the collector was enabled. The test copies this pattern line for line.Doc/library/timeit.rstL136-141: timing with garbage collection off by default "makes independent timings more comparable"; the stated disadvantage (GC cost excluded from the measurement) is recorded under limitations in the artifact.Modules/gcmodule.cL1248, L1441, L1455 and L1479: promotions into the oldest generation add tolong_lived_pending, and a full collection runs only when the oldest generation's count exceeds its threshold (L1441) andlong_lived_pendinghas reached a quarter oflong_lived_total(L1479); the comment at L1457-1461 states that its cost "is proportional to the total number of long-lived objects". In a process that has run thousands of tests, that pause depends on the heap, not on the text being timed.Evidence-class table
test_k4_timingfork4_f_header(test exponent 1.508) andk4_shell_words(1.588); attempt 1 of run 37120736604 (#651, job 111196172237) failed it fork4_shell_words(1.534). Across the 9 failing rounds the 50k/25k ratio is 1.91-2.01 and the 100k/50k ratio 5.93-8.54; the 100k time exceeds twice the same round's 50k time by 69.6-80.8 ms (over the per-size minima: 1.97-1.98, 5.93-8.06 and 69.6-76.3 ms). A fixed excess in every round, not growth. Attempt 2 of both runs passed.gh run view <run> --attempt 1 --log-failed;ci_failuresin the artifactk4_realtestruns (real test after discovery imported every test module);local_reproductionin the artifactgc.callbacksreplay ofmeasure_round)gc_callback_attributionin the artifactk4_shell_wordsandk4_f_headerare linear with the collector off out to 400k (raw pair exponents 0.984-1.054)linearity_collector_offin the artifactk4_shell_wordsmutant is still rejected with the collector offquadratic_control_final_testin the artifacttimeitdisables GC while timing; full-collection cost scales with long-lived objectssha256sum scripts/hooks/secret_path_guard.py=33a11fc01ee35dc4b157b04eee4a8524460bb3fe74375ee32485bc3b2530782fbefore and afterLocal commands run
Interpreter:
/usr/bin/python33.12.3, the version of the diagnosis and of CI (both failing job logs printUsing CPython 3.12.3 interpreter at: /usr/bin/python3). This host's PATHpython3is 3.13.15. TMPDIR under/var/tmp, every check undernice -n 19; exit codes read directly.Measurements behind the artifact (local integration on synthetic heaps, each
nice -n 19 /usr/bin/python3 -B): the final test in the three heaps that failed the unchanged test (3 of 3 passed), the quadratic mutant against the final test (rejected twice), and the paired unchanged/final exponent capture in the same three heaps.Review round 1, at head
a94227635(same interpreter and flags):Diagnosis
Full record:
evidence/artifacts/k4-timing-gc-20261003/summary.json(18 KB, no host paths).test_k4_timing. In each of the 9 failing rounds the helper time roughly doubles from 25k to 50k (ratio 1.91-2.01), then jumps 5.93-8.54 times to 100k: the 100k time exceeds twice the same round's 50k time by 69.6-80.8 ms. Over the per-size minima the test decides on, the figures are 1.97-1.98, 5.93-8.06 and 69.6-76.3 ms. Because every failing round carries the excess, the minimum over rounds cannot clear it.python3 -m unittestprocess, so the oldest generation holds a large long-lived heap. A full collection runs only when the oldest generation's count exceeds its threshold (gcmodule.cL1441) and promotions since the last one reach a quarter of the long-lived heap (L1479), and it costs in proportion to that heap (L1460-1461): the pause depends on the heap, not on the text being checked. Each round allocates the same objects in the same order, so once the trigger falls inside the 100k window it falls there every round.k4_shell_wordstest exponents 1.819, 1.652, 1.617). Paired re-run in those three heaps, unchanged and fixed test alternating: the unchanged test failed in two (1.722, 1.529) and passed the third only with near misses (k4_f_header1.475 after two rounds,k4_code_tokens1.485). The fixed test passed all three, deciding every helper in round 1, with a maximum test exponent over all 67 helpers of 0.944-1.051.gc.callbacks). In a replay ofmeasure_roundwith the test's ownk4_run_timing_roundsandk4_timing_passes, all 7 collector-on failures had exactly one generation-2 collection inside the 100k window in every round (72.1-110.4 ms). With the collector off, the 14 decisions in the same heaps passed in round 1 (test exponents 0.665-0.848), with no generation-2 collection in any window.k4_f_header's 25k window gave it test exponents of 0.193 and 0.225.k4_shell_wordsandk4_f_headerare linear out to 400k (raw pair exponents 0.984-1.054).k4_shell_wordsmutant (text[i:]scanned every 10 characters) is still rejected by the fixed test (1.630, and 1.635 in the 8M/36000 heap), and in the replay with the collector on or off (1.625/1.640; every 5 characters: 1.771/1.788).Rejected alternatives
test_k4_growth_criterion_controlsscore 1.412 (from 4.5 ms) and 1.663 (from 10 ms), so the 4.5 ms control would pass and that test would break. This is arithmetic over the recorded minima, not a new run.gc.collect()before each timed size. The collector stays on inside the window, and every run pays a full collection per size per round outside it: 201 to 603 of them for 67 helpers, each about the 72-110 ms pause measured in the simulated heaps. Not measured.gc.freeze(). It changes process-wide collector state for the rest of the suite: it resets the generation counts and hides the frozen heap from every collection untilgc.unfreeze().Doc/library/gc.rstdocuments it forfork()memory sharing, whiletimeit's documented practice for timing isgc.disable(). The diagnosis measured a freeze-after-setup variant as linear (raw pair exponents 0.978-1.051), so this is a choice of practice, not a failed experiment.Review round
An independent Opus review at
cc19f03dfound no defect in the code change: the complexity guarantee is kept, the collector state is restored on every path including raises, and the upstreamtimeitandgcmodule.cclaims hold at58ed60b7. It found four should-fix inaccuracies in the record and two nits. One repair round (ffcb6054c, thena94227635withmanifests/evidence.jsonlast) disposes of them:ci_signatureandlimitations, evidence-table row 1 and Diagnosis 1 give the per-round ranges (1.91-2.01, 5.93-8.54, 69.6-80.8 ms; up to 82.5 ms over twice the 50k minimum) apart from the per-size minima (1.97-1.98, 5.93-8.06, 69.6-76.3 ms), and the minima fields are renamedminima_*. The conclusion is unchanged.overturn_conditionslists three, cited under Decision record.Using CPython 3.12.3 interpreter at: /usr/bin/python3(lines 1503 and 1509). The artifact cites them (ci_failures[].interpreter_log_line,python.ci), and this body no longer says the version is unprinted.gc_callback_attributionscopecollector_off_test_exponent_range_all(0.665-0.848) and says the 7 listed twins span 0.708-0.841.evidence_classesnow cover the attribution, linearity and collector-on direct-call sections.gcmodule.cL1441, which is also added under SOTA sources.docs/secret-storage.mdsays no collection of any generation runs inside a timed window.test_k4_timingruns, all exit 0 (Local commands run).The branch is deliberately not merged with
origin/mainyet. Thevalidatecheck oncc19f03dfailed inscripts/validate.pyonfiles[] must be sorted by path, at twodocs/decisions/2026-10-03-*entries that come from main (ecea2865), not from this branch; #666 fixes that registry.scripts/validate.pypasses on this branch.Decision record
None separate: this changes how a test measures, not a selection. The artifact records the evidence, the alternatives and, under
overturn_conditions, what would overturn the change: a CI full-suite run in which the fixed test still fails a helper on growth; a guard change whose cost is dominated by allocation-heavy work that creates cyclic garbage, so that collection cost grows with the input (collector-off timing would not see it, so that helper would need a collector-on control; this is the disadvantageDoc/library/timeit.rstL136-141 names); or upstreamtimeitchanging its practice so thatTimer.timeitno longer disables the collector.Host evidence
Not applicable: no file under
evidence/hosts/changes.Checklist
version comment (no floating tags). (No Actions change.)
permissions: contents: read(or a narrower, explicitly justified addition). (No workflow change.)
added without a documented owner.
moved or overwritten).
🤖 Generated with Claude Code