Skip to content

D1 slow-test residual sweep: retire stale warn row - #2817

Merged
briansrls merged 17 commits into
mainfrom
session/loyal-ferret-30
May 13, 2026
Merged

briansrls merged 17 commits into
mainfrom
session/loyal-ferret-30

Conversation

@briansrls

@briansrls briansrls commented May 13, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Wave-1 D1 slow-test residual sweep after the TestNodeCostDimension + JSONL manifest bridge landed in the #2761/#2764 stack.

  • removes one now-under-budget warn-policy row from scripts/test-node-wall-clock-ratchet.jsonl
  • adds a dated residual-sweep receipt documenting the name audit, partial timing evidence, and rows intentionally left in place
  • updates ROADMAP.md so the CI ratchet architecture row no longer describes the retired slow-test-exemptions.txt / TEST_TIMEOUT_MAX_EXEMPTIONS protocol as current state

Evidence

  • RUSTC_BOOTSTRAP=1 cargo test -p v3-compiler -- --list confirmed all 83 pre-cut manifest names existed in the current v3 test list.
  • ctrl-build -- env RUSTC_BOOTSTRAP=1 cargo test -p v3-compiler -- -Z unstable-options --report-time captured 391 lib-test timing rows before stopping on a remote-environment helper lookup failure for gunbc_execute_command_bootstrap.
  • The captured timing log measured bootstrap::tests::kernel_bool_path_a_attaches_diagnostic_when_boolean_algebra_unresolvable at 16 ms, below the 2000 ms Phase-0 budget, so its warn row was removed.
  • scripts/check-test-timeout.sh /tmp/v3-report-time.log parsed the captured log and exited cleanly: remaining over-budget rows were warn-listed, with no unexpected over-budget tests.
  • jq -e . scripts/test-node-wall-clock-ratchet.jsonl passed; manifest count is now 82 rows.

Notes

This PR is D1 residual paydown, not a new #102 closure claim. The timeout script still reads the JSONL bridge directly until slow-test policy is projected from modeled timing facts.

@briansrls briansrls changed the title T-WAD Slice 6: slow_test_exemptions_dissolved (test cost-dim cuts) D1 slow-test residual sweep: retire stale warn row May 13, 2026
@briansrls
briansrls marked this pull request as ready for review May 13, 2026 01:26

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review metadata

  • Provider / model: codex / unknown
  • Commit: d2c04387 · Trigger: schedule
  • Thinking: 149s wall

Non-blocking — Strengths

  • docs/debt/slow-test-residual-sweep-2026-05-13.md The sweep note documents the partial-run boundary and gives a concrete under-budget receipt for the removed warn row.

ROADMAP — Verified

  • CI ratchet architecture: The updated row preserves the #102 caveat, names the JSONL manifest bridge, and keeps fresh timing audits plus per-node budgets open.

✅ No blocking concerns.

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: openai-pro / gpt-5-5-pro
  • Commit: 207e676b · Trigger: manual
  • Comparison: main @ 4b0cca9f ... session/loyal-ferret-30 @ 207e676b
  • Conversation: View conversation

1. Story of the diff

This PR is a narrow residual cleanup for the slow-test timing ratchet. It updates the ROADMAP CI-ratchet debt row so the old free-form slow-test-exemptions.txt framing is no longer presented as live, and it records the current bridge shape: scripts/check-test-timeout.sh reads structured warn policy from scripts/test-node-wall-clock-ratchet.jsonl until #102 projects policy from modeled TestNodeCostDimension timing facts (ROADMAP.md:438). The new debt note then gives the receipt for this specific sweep: it says the sweep is not a #102 closure claim (docs/debt/slow-test-residual-sweep-2026-05-13.md:3-6), records a name audit and partial timing run (docs/debt/slow-test-residual-sweep-2026-05-13.md:8-16), and justifies retiring one warn row because that test measured 16 ms under the 2000 ms budget (docs/debt/slow-test-residual-sweep-2026-05-13.md:18-20). The only manifest data change is the deletion of that warn-policy row from scripts/test-node-wall-clock-ratchet.jsonl:49 (deleted), while the doc leaves the remaining over-budget rows and un-retimed integration rows explicitly tracked (docs/debt/slow-test-residual-sweep-2026-05-13.md:22-24).

2. Invariant categories

  1. LAYER MODEL (substrate vs implementation).

N/A — this diff touches ROADMAP/debt documentation and a test-time JSONL policy manifest; it does not alter Dag, substrate types, cross-pass carriers, or modeled .dag declarations.

  1. INVARIANTS.md + modeling-discipline.md.

Compliant — P5 Progress Is Dissolution is handled correctly: the PR removes one ad-hoc warn-policy row (scripts/test-node-wall-clock-ratchet.jsonl:49 deleted) and avoids overstating the bridge as closed by saying the ratchet still reads JSONL until projection from modeled timing facts (docs/debt/slow-test-residual-sweep-2026-05-13.md:3-6). That matches the invariant that scaffolds need dissolution paths and should reduce ad-hoc state rather than become steady state. chatgpt-review-5fc1e4ba-9f1b-44…

  1. CODING.md.

N/A — no Rust implementation, helper API, method/free-function shape, error carrier, or naming surface is introduced. The only machine-readable implementation artifact touched is deletion of one JSONL policy row.

  1. TESTING.md.

Compliant — no new test is needed because the subject is not new behavior; the PR supplies the right verification level for a manifest-row retirement: a libtest name audit (docs/debt/slow-test-residual-sweep-2026-05-13.md:8-10), a timing run accepted by the timeout parser with no unexpected over-budget tests (docs/debt/slow-test-residual-sweep-2026-05-13.md:12-16), and the specific under-budget measurement for the removed row (docs/debt/slow-test-residual-sweep-2026-05-13.md:18-20). This is consistent with TESTING’s Phase-0 timing-ratchet posture, where JSONL warn policy is an interim bridge and unknown over-budget tests fail closed. chatgpt-review-39a2934f-5ab9-42…

  1. LOCKED DESIGN DECISIONS.

Compliant — the diff does not weaken the zero-floor / tests-as-data direction. It explicitly says the JSONL manifest bridge remains pending #102 modeled timing projection (ROADMAP.md:438; docs/debt/slow-test-residual-sweep-2026-05-13.md:3-6) rather than presenting the side manifest as the final design.

  1. TRACKED vs UNTRACKED DEBT.

Compliant — the remaining debt is tracked with documentation, bounds, and a dissolution trigger. Documentation: the new debt receipt (docs/debt/slow-test-residual-sweep-2026-05-13.md:1-24). Bounds: 83 manifest names audited, 391 timing rows captured, seven lib-test warn rows still over budget, and integration-test rows explicitly not re-timed (docs/debt/slow-test-residual-sweep-2026-05-13.md:8-24). Dissolution trigger: #102 projection from modeled TestNodeCostDimension timing facts (ROADMAP.md:438; docs/debt/slow-test-residual-sweep-2026-05-13.md:3-6).

2.5. Top-down PM intent review

Compliant — the PR preserves the PM-level intent. The highest-level direction is that tests and timing policy should ultimately be structural .dag/TestClaim data with zero hand-authored residual, not permanent Rust or side-manifest authority. chatgpt-review-d26e13de-8fee-4b…

chatgpt-review-9197ea0c-307c-4a…

This diff moves in that direction by retiring a stale warn row while refusing to call the JSONL bridge final: docs/debt/slow-test-residual-sweep-2026-05-13.md:3-6 and ROADMAP.md:438 keep the modeled TestNodeCostDimension projection as the real closure condition. No PM-level must-have is made optional, and no dissolution goal is turned into permanent scaffolding.

3. Verdict

APPROVE. This is a clean residual paydown: it deletes one measured-under-budget warn row, documents the evidence, and keeps the remaining manifest bridge explicitly temporary rather than claiming #102 closure. I found no blocking or non-blocking findings tied to changed diff lines.

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: openai-pro / gpt-5-5-pro
  • Commit: 3c623738 · Trigger: manual
  • Comparison: main @ 9cd345a1 ... session/loyal-ferret-30 @ 3c623738
  • Conversation: View conversation

1. Story of the diff

This PR retires one stale slow-test warn-policy row from the Phase-0 wall-clock ratchet. The ROADMAP debt row is rewritten from the old slow-test-exemptions.txt/TEST_TIMEOUT_MAX_EXEMPTIONS framing to the newer JSONL manifest bridge, while explicitly keeping #102 open until timing policy is projected from modeled TestNodeCostDimension facts (ROADMAP.md:438). The new sweep receipt records the verification path: manifest-name audit, partial --report-time timing evidence, and the specific now-under-budget test that measured 16 ms (docs/debt/slow-test-residual-sweep-2026-05-13.md:8-20). The mechanical payload is the deletion of that test’s policy:"warn" row from scripts/test-node-wall-clock-ratchet.jsonl:49, so future over-budget behavior for that test falls back to the normal fail-closed timeout path.

2. Invariant categories

  1. LAYER MODEL (substrate vs implementation).

N/A — this diff does not alter substrate types, Dag, dag.rs, cross-pass data shapes, or modeled .dag carriers; it updates debt docs and one test-time JSONL policy row.

  1. INVARIANTS.md + modeling-discipline.md.

Compliant — P5/Strict Forward Progress is handled correctly: the diff removes one bridge row rather than adding a scaffold, and it explicitly says the JSONL bridge is not #102 closure because policy still reads from scripts/test-node-wall-clock-ratchet.jsonl until modeled timing facts project it (docs/debt/slow-test-residual-sweep-2026-05-13.md:3-6, ROADMAP.md:438). That matches the project rule that scaffolds need dissolution paths and must not become steady state. chatgpt-review-5a8e005c-fa32-4b…

  1. CODING.md.

N/A — no Rust implementation code, helper API, methods, error/result shapes, or new impurity surface are introduced; the diff is documentation plus JSONL data.

  1. TESTING.md.

Compliant — the changed test policy follows the Phase-0 wall-clock contract: the sweep documents the list audit and timing evidence (docs/debt/slow-test-residual-sweep-2026-05-13.md:8-16), removes only the under-budget test’s warn row (docs/debt/slow-test-residual-sweep-2026-05-13.md:18-20, scripts/test-node-wall-clock-ratchet.jsonl:49), and leaves remaining over-budget/incompletely-timed rows in warn backlog (docs/debt/slow-test-residual-sweep-2026-05-13.md:22-24). This aligns with TESTING’s current rule that the JSONL manifest is an interim bridge, unknown over-budget tests fail closed, and warn behavior needs an explicit manifest row. chatgpt-review-dae7b3cb-a43d-43…

  1. LOCKED DESIGN DECISIONS.

Compliant — the PR references the zero-residual/test-cost trajectory but does not dilute it: docs/debt/slow-test-residual-sweep-2026-05-13.md:3-6 explicitly avoids claiming #102 closure, and ROADMAP.md:438 keeps “projection from modeled TestNodeCostDimension timing facts” as the remaining target. That preserves the live Pure Bootstrap / testing direction rather than redefining the bridge as done. chatgpt-review-17cf125f-8d6c-4d…

  1. TRACKED vs UNTRACKED DEBT.

Compliant — no new untracked debt is added. The remaining side-manifest bridge is bounded and named: it is pending #102 projection from modeled timing facts (ROADMAP.md:438, docs/debt/slow-test-residual-sweep-2026-05-13.md:3-6), while this PR dissolves one concrete row by deleting the JSONL entry at scripts/test-node-wall-clock-ratchet.jsonl:49.

2.5. Top-down PM intent review

Compliant. The highest-level intent is that tests and correctness dimensions move toward structural data rather than hand-maintained behavioral exceptions: the thesis says the proof/test surface is structurally derived and that the previous Rust-side residual carve-out is retracted under the 0-floor target, with tests migrating to .dag TestClaim data. chatgpt-review-e926bd55-e046-49…

This PR moves in that direction at the PM level: it removes one manually maintained warn exception, documents why the remaining JSONL is still only a bridge, and leaves the modeled TestNodeCostDimension projection as the named closure target (ROADMAP.md:438, docs/debt/slow-test-residual-sweep-2026-05-13.md:3-6). It does not turn the bridge into a permanent scaffold or postpone the structural goal. chatgpt-review-e926bd55-e046-49…

3. Verdict

APPROVE

The PR is a narrow, well-scoped debt paydown: it removes a stale warn row, records the evidence for doing so, and keeps the larger #102/modeling target open rather than overstating closure. I did not find a diff-cited invariant or PM-intent mismatch.

@briansrls

Copy link
Copy Markdown
Contributor Author

Verified the relayed claude review against the current HEAD. It is an approval with no findings or requested changes, and the reviewed scope matches the current diff: one stale warn row retired plus ROADMAP/HANDOFF/debt receipt updates. No code or doc change is required for this feedback. — sent from loyal-ferret-30

@briansrls

Copy link
Copy Markdown
Contributor Author

Verified the cursor review against current HEAD 93a2169. The optional HANDOFF wording note is accurate: HANDOFF_PHASE0.md still says "Each exempt test" in the paydown bullet while nearby text now uses JSONL warn-policy language. I am not pushing a doc-only churn commit for this because the review marked it exploratory/optional, it does not affect policy, and moving HEAD would reset current CI/review state. I will leave it for the next substantive touch. — sent from loyal-ferret-30

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review metadata

  • Provider / model: codex / unknown
  • Commit: 1d973c0f · Trigger: schedule
  • Thinking: 154s wall

✅ No blocking concerns; the docs/manifest sweep keeps the JSONL bridge explicitly interim pending #102 and removes only the under-budget warn row with bounded timing evidence.

@briansrls

Copy link
Copy Markdown
Contributor Author

Verified the codex no-blocking-concerns review against the current PR state. It has no actionable findings, and it was for older commit 1d973c0; current HEAD is 93a2169 with the same docs/manifest hygiene scope plus green GitHub checks. No code or doc change is required for this feedback. Dashboard readiness still reports 1/2 recognized approval providers, so I am holding merge per manager instruction. — sent from loyal-ferret-30

@briansrls

Copy link
Copy Markdown
Contributor Author

Verified the claude approval against current HEAD 523f7f3. The branch head moved by merging origin/main, but the PR diff against origin/main remains the same four-file docs/manifest sweep: ROADMAP, the dated slow-test residual receipt, one JSONL warn-row deletion, and HANDOFF_PHASE0 alignment. The review has no actionable findings, so no code or doc change is required. CI is re-running on the new head and dashboard readiness is still false, so I am holding merge per manager instruction. — sent from loyal-ferret-30

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: openai-pro / gpt-5-5-pro
  • Commit: 523f7f3a · Trigger: manual
  • Comparison: main @ 938f334b ... session/loyal-ferret-30 @ 523f7f3a
  • Conversation: View conversation

1. Story of the diff

This PR is a small residual-paydown sweep for the slow-test timing ratchet. It removes one stale warn-policy entry from scripts/test-node-wall-clock-ratchet.jsonl after the test measured well below the 2s budget, and it adds docs/debt/slow-test-residual-sweep-2026-05-13.md to record the audit commands, partial timing evidence, and the specific retired row. The surrounding docs are updated to stop describing the older scripts/slow-test-exemptions.txt / row-count ratchet as live state: ROADMAP.md:438 now says the JSONL bridge is the current interim shape pending #102 projection from modeled TestNodeCostDimension facts, and src/v3/compiler/tests/HANDOFF_PHASE0.md:44-53 describes the current fail-closed behavior for unknown over-budget tests. The important semantic guard is that the new debt note explicitly says this is not a #102 closure claim while the script still reads the side manifest directly (docs/debt/slow-test-residual-sweep-2026-05-13.md:3-5).

2. Invariant categories

  1. LAYER MODEL (substrate vs implementation).

N/A — this diff does not touch Dag, substrate declarations, type/behavior variants, or compiler mutation APIs. The closest substrate-adjacent text is documentation that the ratchet still reads scripts/test-node-wall-clock-ratchet.jsonl “until policy is projected from modeled timing facts” (docs/debt/slow-test-residual-sweep-2026-05-13.md:3-5), so it does not introduce a new substrate authority.

  1. INVARIANTS.md + modeling-discipline.md.

Compliant — P5 / Progress Is Dissolution is handled correctly: the PR removes an ad-hoc warn row rather than adding one (scripts/test-node-wall-clock-ratchet.jsonl:49 deleted), records evidence that the row is now under budget (docs/debt/slow-test-residual-sweep-2026-05-13.md:17-19), and keeps the bridge’s dissolution trigger explicit via #102 / modeled timing facts (src/v3/compiler/tests/HANDOFF_PHASE0.md:170-172). This matches the invariant rule that scaffolds need named dissolution triggers and cannot become steady state. chatgpt-review-832a52d0-08dd-45…

  1. CODING.md.

N/A — no Rust implementation code, functions, methods, error carriers, or helper placement changed. The changed surface is documentation plus one JSONL policy-row deletion, so the src/v3/compiler/src/ coding-style rules are not exercised.

  1. TESTING.md.

Compliant — the diff does not add a new test, but for this kind of manifest cleanup the existing timing-ratchet is the enforcement surface: the debt note records the list audit command (docs/debt/slow-test-residual-sweep-2026-05-13.md:7-9), the partial --report-time run accepted by the ratchet parser (docs/debt/slow-test-residual-sweep-2026-05-13.md:11-15), and the measured 16 ms result that justified deleting the row (docs/debt/slow-test-residual-sweep-2026-05-13.md:17-19). This also preserves TESTING.md’s current rule that JSONL warn rows are interim, unknown over-budget tests fail closed, and rows are required only for tests intentionally allowed to warn. chatgpt-review-ab9429bc-3201-4a…

  1. LOCKED DESIGN DECISIONS.

Compliant — the diff does not alter a locked design doc and does not dilute the 0-floor / tests-as-data direction. It explicitly avoids claiming closure while the side manifest remains (docs/debt/slow-test-residual-sweep-2026-05-13.md:3-5) and keeps the next step as moving warn policy into modeled TestNodeCostDimension facts (src/v3/compiler/tests/HANDOFF_PHASE0.md:170-172), which is consistent with the zero-residual testing direction. chatgpt-review-1efb0921-86e6-4c…

  1. TRACKED vs UNTRACKED DEBT.

Compliant — no new scaffold is introduced. The existing JSONL bridge remains tracked with documentation (src/v3/compiler/tests/HANDOFF_PHASE0.md:44-49), bounds/fail behavior (src/v3/compiler/tests/HANDOFF_PHASE0.md:50-53), and a named dissolution trigger: “Gate #102 remains open until the ratchet consumes modeled TestNodeCostDimension timing facts instead of the JSONL bridge” (src/v3/compiler/tests/HANDOFF_PHASE0.md:170-172). The PR moves debt in the right direction by shrinking the manifest.

2.5. Top-down PM intent review

Compliant — the PM-level intent is that tests and CI/process claims keep migrating toward .dag / modeled structural facts, not permanent side manifests; TESTING.md and the thesis both frame the JSONL/warn-policy shape as interim on the way to 0 residual. This diff preserves that meaning: it says the current manifest is still a bridge, does not claim #102 closure (docs/debt/slow-test-residual-sweep-2026-05-13.md:3-5), removes one stale warn row (scripts/test-node-wall-clock-ratchet.jsonl:49 deleted), and leaves the future work as moving warn policy into modeled TestNodeCostDimension facts (src/v3/compiler/tests/HANDOFF_PHASE0.md:170-172). No semantic dilution or permanent-scaffold turn is visible in the changed lines.

3. Verdict

APPROVE — this is a narrow, well-scoped cleanup that tightens the slow-test ratchet by retiring one obsolete warn row and keeps the remaining bridge explicitly tracked. I did not find a diff-cited invariant violation or PM-intent mismatch.

@briansrls

Copy link
Copy Markdown
Contributor Author

Verified the current-head openai-pro review against HEAD 523f7f3. It reports APPROVE with no actionable findings; the PR diff remains the same four-file docs/manifest sweep, and GitHub checks are green. No code or doc change is required for this feedback. Dashboard readiness still reports ready=false / 0 recognized current-head approvals despite the review text, so I am holding merge per the dashboard source-of-truth policy. — sent from loyal-ferret-30

@briansrls

Copy link
Copy Markdown
Contributor Author

Verified the current-head codex review against HEAD 251bb1a. It reports APPROVE with no actionable findings, and the PR diff remains the same four-file docs/manifest sweep. No code or doc change is required for this feedback. Dashboard readiness still reports ready=false with only 1/2 recognized approval providers and v3 is re-running on the current head, so I am holding merge per manager instruction. — sent from loyal-ferret-30

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review metadata

  • Provider / model: codex / unknown
  • Commit: 523f7f3a · Trigger: schedule
  • Thinking: 106s wall

Non-blocking — Strengths

  • docs/debt/slow-test-residual-sweep-2026-05-13.md The sweep note is bounded to residual paydown, preserves #102 as open, and records concrete 16 ms evidence for the retired warn row.
  • scripts/test-node-wall-clock-ratchet.jsonl The manifest change shrinks the interim warn-policy bridge instead of expanding hand-maintained timing policy.

ROADMAP — Verified

  • CI ratchet architecture: The updated row keeps fresh CI-shaped timing audits, stale warn-row deletion, and per-node duration budgets open after this one-row retirement.

✅ No blocking concerns; this is a clean docs/manifest alignment with the thesis direction toward modeled timing facts.

@briansrls

Copy link
Copy Markdown
Contributor Author

Verified the latest no-finding reviews against the current PR state. Current HEAD is 348009c and the PR diff against origin/main remains the same four-file docs/manifest sweep: ROADMAP, the dated slow-test residual receipt, one JSONL warn-row deletion, and HANDOFF_PHASE0 alignment. The new claude review reports APPROVE with no actionable findings on current HEAD; the relayed codex no-blocking review was for older HEAD 523f7f3 but remains consistent with the unchanged PR diff. No code or doc change is required for either item. Dashboard readiness still reports ready=false with only 1/2 recognized approval providers and v3 is running on current HEAD, so I am holding merge per manager instruction. — sent from loyal-ferret-30

@briansrls

Copy link
Copy Markdown
Contributor Author

Verified the latest current-head codex approval against HEAD 1d5c2f5. It reports APPROVE with no actionable findings, and the PR diff remains the same four-file docs/manifest sweep. No code or doc change is required for this feedback. Dashboard readiness still reports ready=false with only 1/2 recognized approval providers; GitHub checks are green but mergeability is currently UNKNOWN, so I am holding merge per manager instruction and dashboard source-of-truth policy. — sent from loyal-ferret-30

@briansrls

Copy link
Copy Markdown
Contributor Author

Verified the current-head cursor approval against HEAD 2ccf663. It reports APPROVE with no actionable findings, and the PR diff remains the same four-file docs/manifest sweep. No code or doc change is required for this feedback. GitHub shows checks green and mergeStateStatus CLEAN; dashboard currently recognizes 2/2 providers but still reports checks_state=pending, so I am waiting for the dashboard source of truth to catch up before merging. — sent from loyal-ferret-30

@briansrls
briansrls merged commit d574fd5 into main May 13, 2026
5 checks passed
@briansrls
briansrls deleted the session/loyal-ferret-30 branch May 13, 2026 06:14

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review metadata

  • Provider / model: codex / unknown
  • Commit: 348009c8 · Trigger: schedule
  • Thinking: 139s wall

✅ No blocking concerns; the PR keeps #102 open, records bounded timing evidence, and shrinks the interim warn-policy manifest by one row.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant