Repository navigation
proud-owl-696 - #541
proud-owl-696#541
Conversation
|
What landedLayer 2 (per-test budget helper):
The infrastructure itself is well-designed. Matches Three gaps (blocking)1. Doesn't implement Layer 1 (CI wall-clock gate)The user's exact ask was "put a coarse ratchet on it (like 1 minute) to make sure we don't run into this again." That's the CI-level wall-clock gate — a single check in the Fix: add the - name: v3 tests with 120s budget
run: |
start=$(date +%s)
cargo test -p v3-compiler
elapsed=$(( $(date +%s) - start ))
echo "v3 test wall time: ${elapsed}s"
if [ $elapsed -gt 120 ]; then
echo "::error::v3 tests took ${elapsed}s (budget: 120s) — share bootstrap setup via OnceLock or collapse fine-grained tests"
exit 1
fi2. No first consumer — E-6 violationThe helper lands with zero tests opting in. Per E-6 (from this session's PR #533 Stage 1d §6 candidate invariants): "No target-spec field lands without a same-PR consumer." That discipline applies equally to test-infra helpers. Without a first consumer, nothing in the repo today actually benefits from this ratchet — a fresh regression of the ζ shape could land tomorrow and Fix: apply the macro to at least one test in this PR. The natural choice is the ζ regression itself — see #3. 3. Doesn't address the ζ regression — the motivating problem
Fix: either
Either way, the Layer 1 gate can't land until the baseline is restored. Recommended path
One more concernThe Layer 2 macro is SummaryGood foundation work on Layer 2. Blocked on shipping without Layer 1 + a first consumer + addressing the baseline regression that motivated the ratchet in the first place. The order to land this properly is (refactor ζ or apply budget to it) → Layer 1 CI gate → Layer 2 helper with first consumer — in one PR or tight sequence. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4520405399
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| ($ms:literal, $name:ident, $body:block) => { | ||
| #[test] | ||
| fn $name() { | ||
| $crate::common::budgeted::with_budget_ms($ms, || $body); |
There was a problem hiding this comment.
Call budget helpers through public path in macro
budgeted_test! is exported to the integration-test crate root, but it expands to $crate::common::budgeted::... even though budgeted is declared as a private module (mod budgeted;). As a result, any test that starts using this macro will fail to compile with E0603: module budgeted is private. Route macro calls through the existing public re-exports ($crate::common::with_budget_ms / DEFAULT_BUDGET_MS) or make the module public.
Useful? React with 👍 / 👎.
|
ChatGPT review in progress... (view conversation) Check back in ~30 minutes for the full review. |
|
Thanks for the comprehensive second pass — Layer 1 CI gate landed, ζ test file refactored, Layer 2 helper kept. All three gaps from my first review addressed. One trivial compile error blocking CI: The Fix: drop line 28 of pub use budgeted::{with_budget_ms, DEFAULT_BUDGET_MS};The Once CI is green we'll see whether the ζ test refactor dropped the v3 job back under the 120s budget — that's the real proof the Layer 1 gate + refactor combination works. |
briansrls
left a comment
There was a problem hiding this comment.
codex · gpt-5.4 · 45204053
BLOCKING (1)
Root Cause
src/v3/compiler/tests/lane2_stage_2d_symbolic_cost_test.rsThe shared-setup fix is modeled as one globalMutex<HashMap<...>>, which serializes all cache misses; switch to a per-keyOnceLock/LazyLockcache or compile outside the lock and insert afterward so the budget measures each test’s own work.
Non-blocking — Strengths
.github/workflows/ci.ymlThe suite-level 120s check is a reasonable fail-closed performance ratchet and fits the repo’s existing CI-budget pattern.
|
Violations (could not place on specific lines):
|
This comment has been minimized.
This comment has been minimized.
|
ChatGPT review in progress... (view conversation) Check back in ~30 minutes for the full review. |
briansrls
left a comment
There was a problem hiding this comment.
codex · gpt-5.4 · 6995abb8
✅ Review (blocking: 0, non-blocking: 2+/0-)
Non-blocking — Strengths
src/v3/compiler/tests/lane2_stage_2d_symbolic_cost_test.rsThe test helper change matches the root cause from the prior review: compile work is shared per fixture key, but lock scope is limited to map lookup/insert only..github/workflows/ci.ymlThe new 120s and 600s timing gates are scoped to the real regression surface, and the PR head’sv3GitHub Actions job passed with the full-suite step logging 513s.
✅ This looks clean: the previous contention bug is fixed, the budget helper stays in implementation-only test code, and the new CI ratchet is behaving as intended on the current head commit.
|
✅ Review (blocking: 0, non-blocking: 2+/0-) Non-blocking — Strengths
✅ This looks clean: the previous contention bug is fixed, the budget helper stays in implementation-only test code, and the new CI ratchet is behaving as intended on the current head commit. |
ChatGPT ReviewGenerated by gpt-5-4-pro Principle audit. Fail-closed. Grounded in the thesis’s “broken causal link → diagnostic / explicit failure” framing and the active review checklist, this round looks good. The new CI gates in chatgpt-review-3898741a-e40e-48… Illegal states unrepresentable. Satisfied. The new cache shape in Facts flow forward. Satisfied. The expensive fact this suite keeps re-deriving— Coproduct dissolution. Satisfied / not really in play. This diff does not introduce a new substrate enum or a new Rust enum that would need 🟢/🟡/🔴 classification. The additions are a helper module, a macro, a cache, and CI wiring. Single-authority metadata. Mostly satisfied. The budget mechanism now has one reusable authority in API-level enforcement. Partly satisfied. Within this suite, Design question. Should compile-sharing be a common test-fixture authority, or stay as an ad hoc cache inside each hot integration binary? What is at stake is not this one cache—it is whether the repo learns one pattern or several. This PR correctly centralizes the budget mechanism in Path to convergence. I do not see a must-fix-before-merge structural issue in the diff as shown. The current per-key Verdict. APPROVE. This looks like a good round. It turns an untracked performance pain into explicit ratchets with a real consumer, and I do not see it baking new substrate/modeling debt into the compiler itself. LOOP HEALTH: converging — this round makes forward progress by replacing repeated per-test full compiles with a bounded shared-setup path and explicit CI/test ratchets, without adding new substrate scaffolds or shifting debt across layers. |
|
Meta-review in progress... (view conversation) Loop-health check: is this review cycle making forward progress, or shifting debt? Posts in ~5-15 minutes. |
|
The conversation did not finish within the timeout window. The bot will start a fresh conversation on the next push. |
Opened from session-dashboard for session
proud-owl-696.