Skip to content

Guard K4 timing test: repair of the cross-family findings on #594 (growth on raw timings, capped host factor, permanent controls) - #600

Merged
seathatflowsinourveins merged 4 commits into
mainfrom
claude/guard-k4-timing-repair-20261002
Oct 2, 2026
Merged

seathatflowsinourveins merged 4 commits into
mainfrom
claude/guard-k4-timing-repair-20261002

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Scope

What changed, by finding:

  1. Normalization weakened the growth criterion. The exponent is now always computed on raw per-size minima; the host factor applies to the absolute processor-time bound only, so the fixed 5 ms allowance is never scaled. Quadratic helper timings of 10, 40 and 160 ms are rejected at host factors 1.0, 2.6 and 4.0.
  2. One noisy reference could relax the bound without limit. The host factor is min(4.0, max(1.0, min(five reference samples) / 0.188)). The cap leaves margin over the recorded hosted macOS slowdown (about 2.8 times) and bounds the absolute limit at 2.0 s.
  3. No permanent control for the retry and minimum paths. Seven control tests call the same pure functions the live test calls (k4_timing_passes, k4_run_timing_rounds, k4_host_scale, k4_min_timings, k4_nesting_passes).
  4. Comments and documentation disagreed with the code. Both now state when calibration runs, five reference runs, the minimum, the cap, raw growth, and repetition only while a criterion fails.
  5. The cost claim was too broad. The comment gives exact measurement counts for a first-round pass and for the maximum.

The failure on hosted runners that this removes (for example run 36946042633 on #593): helper k4_shell_words, {25000: 0.009996, 50000: 0.019747, 100000: 0.117139}, exponent 1.5129 against 1.5. That is one noisy 100k sample: accepted once a second round supplies a clean one, still rejected when three rounds repeat it.

One judgment left for the keys lane: whole-check timings of 0.6, 0.8 and 1.0 s are accepted on a host whose measured factor is 2.6 (bound 1.3 s). That is the intended meaning of host scaling now that the factor is a capped minimum of five; it is stated as a control so the choice is visible.

SOTA sources

Evidence-class table

Claim Evidence class Command / receipt
Quadratic growth is rejected at host factors 1.0, 2.6 and 4.0; one noisy sample is tolerated only when a later round is clean synthetic test_k4_growth_criterion_controls, test_k4_timing_retry_controls, test_k4_timing_minimum_controls
A noisy reference sample cannot relax the bound; the factor is capped synthetic test_k4_host_calibration_controls, test_k4_reference_measurement_controls
The controls detect weakened decision logic synthetic thirteen in-memory reviewer mutants of the decision functions, all killed (listed below)
The live timing test passes on the workstation local_integration tests.test_secret_path_guard: 82 tests OK, 2 skipped; test_k4_timing three consecutive passes
Hosted runners no longer fail on the noisy sample none recorded yet this pull request's validate and validate-macos checks are the observation

Local commands run

$ python3 -B -m unittest tests.test_secret_path_guard
Ran 82 tests in 66.477s; OK (skipped=2); exit 0
$ python3 -B -m unittest tests.test_secret_path_guard.K4GuardTests.test_k4_timing   (three times)
OK in 27.2 s, 26.9 s, 26.9 s
$ reviewer mutants, in memory (decision functions replaced one at a time, seven controls run)
unmutated: 7 controls pass. Killed: normalize timings before the exponent; scale the 5 ms allowance by the host
factor; drop the absolute bound; drop the growth criterion; maximum of the reference samples; no cap; no floor of 1;
accept fewer than five reference samples; last round only; maximum over rounds; always three rounds; never repeat;
calibrate before the first round. 13 of 13 killed.
$ python3 -B scripts/validate.py
{"components": 69, "hashed_files": 8975, "profiles": 4, "receipts": 186, "status": "passed"}; exit 0
$ git diff --check dfdce8fc..HEAD
exit 0

How it was made: a GPT-6.1 Sol max worker (the Codex SDK worker through the loopback gateway) implemented a written contract in an owned worktree without committing; Claude Opus reviewed the diff, re-ran the commands above and wrote the reviewer mutants. The worker's own run reported the two blocking defects reproduced before editing and five consecutive passes of the K4 class.

Decision record

docs/secret-storage.md, the K4 paragraphs on timing bounds (rewritten in this change).

Checklist

  • No new or changed GitHub Actions or workflows.
  • No secrets are printed, logged or committed.
  • No new paid hosting, subscription or billing surface.
  • Peer-owned untracked files and worktrees were preserved.

🤖 Generated with Claude Code

Scout and others added 4 commits October 1, 2026 20:32
…ctor beside the exceeding measurement

After the guard reinstall on 2026-10-02 the workstation, at a load average of 25 to 50 from other lanes' measurement
workers, failed test_k4_timing on linear work: whole-check processor times of 0.58 to 0.65 s against a host-scaled
bound of 0.56 s, with a different set of helpers each run. A host factor taken once at the start missed contention
that moves during the run, and a helper repeated its rounds only when the growth exponent failed. Now:
- a helper repeats its round (up to three, elementwise minimum per size) when an absolute bound fails too, and a
  repeated round divides its times by the host factor measured just before and after it;
- a row's child keeps the minimum of its three spans (the least-disturbed run) and, only when that exceeds the
  workstation bound, asks for the host factor by two reference runs right after.
An unloaded run costs the same as before. Three runs at a load average of 25 to 30 passed (51 to 59 s). A negative
control that injects work quadratic in the command length into k4_join still fails the test: through the absolute
bound at the stronger setting, and through the growth exponent alone (1.62 against 1.5) at the weaker one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
manifests/evidence.json re-registers tests/test_secret_path_guard.py and docs/secret-storage.md;
scripts/validate.py passes with 186 receipts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ost factor, shared decision functions and permanent controls

Repairs the five findings of the cross-family read of dfdce8f (posted on pull request 594, 2026-10-02):

1. The growth exponent is always computed on raw per-size minima; the host factor applies to the absolute
   processor-time bound only, so the fixed 5 ms allowance is never scaled. Quadratic helper timings of 10, 40 and
   160 ms are rejected at host factors 1.0, 2.6 and 4.0.
2. The host factor is min(4.0, max(1.0, min(five reference samples) / 0.188)): one noisy reference sample can no
   longer relax the bound, and the bound cannot exceed 2.0 s. Reference runs use timeit.repeat.
3. Permanent controls exercise the retry path, the per-size minimum over rounds, calibration, the measurement
   budget and the nesting probe through the same pure functions the live test calls (k4_timing_passes,
   k4_run_timing_rounds, k4_host_scale, k4_min_timings, k4_nesting_passes).
4. Comments and docs/secret-storage.md describe what the code does: when calibration runs, five reference runs,
   minimum, cap, raw growth, repetition only while a criterion fails.
5. The cost statement gives exact measurement counts for a first-round pass and for the maximum.

The hosted-runner failure this removes: helper k4_shell_words, {25000: 0.009996, 50000: 0.019747,
100000: 0.117139}, exponent 1.5129 against 1.5 (run 36946042633): one noisy 100k sample, accepted once a
second round supplies a clean one, still rejected when three rounds repeat it.

Implemented by a GPT-6.1 Sol max worker (Codex SDK through the gateway) under a written contract in an owned
worktree; reviewed by Claude Opus: guard sha256 unchanged (33a11fc0...), thirteen reviewer mutants of the decision
functions all killed by the seven controls, tests.test_secret_path_guard 82 tests OK (2 skipped), test_k4_timing
three consecutive passes.

SOTA sources: CPython timeit, Timer.repeat (minimum of repeats):
https://docs.python.org/3/library/timeit.html#timeit.Timer.repeat

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
manifests/evidence.json: the sha256 and bytes of tests/test_secret_path_guard.py and docs/secret-storage.md only;
scripts/validate.py passes (69 components, 8975 hashed files, 4 profiles, 186 receipts).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Oct 2, 2026
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@seathatflowsinourveins
seathatflowsinourveins merged commit 6080214 into main Oct 2, 2026
25 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/guard-k4-timing-repair-20261002 branch October 2, 2026 04:33
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…ect the status section

- publication-checks.json: the failed lookup's recorded command now names the
  repository owner as <user>, the whole-token form scripts/host_receipts.py
  sanitize() writes for the user name. Exit 1, the 404 output and the
  correction fields are unchanged. No convergence record pins this file.
- Status at landing: drop the #600 attribution from the #556 timing bullet
  (#600 changed the Bash secret-path guard's K4 timing test, not the
  child-usage linearity check); state #626 as a dated read without a landing
  instruction; link the program's independent-review section (lines 480-484)
  beside the ownership split (line 178); reword the repository-variable
  observation without a "correction"; add a dated note for the redaction.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant