Skip to content

Rust unit tests run in the clippy job, non-required until measured (step 1 of retiring rust_unit_tests_off_the_merge_path) - #12450

Merged
gunbai-bot[bot] merged 6 commits into
mainfrom
session/neat-ant-823
Sep 28, 2026
Merged

gunbai-bot[bot] merged 6 commits into
mainfrom
session/neat-ant-823

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Operator ruling 2026-09-27; sequencing per neat-boar-16's ruling.

  • gunbc.compiler_gate_workflow: new rust_unit_tests step after "Lint every target" in the clippy job. It runs repo_self_test_command inside a cgroup leaf with memory.max = 12 GiB, so the regeneration-host tests read a real bound; there is no GUNBC_MEMORY_BUDGET_BYTES. The step is NON-required (continue-on-error) and the clippy job's timeout is 120 min, so the step can finish once and be measured.
  • gunbc.ci_spec: the bind sequence is now one function, hosted_runner_memory_cgroup_bind_commands, shared with the heal publisher (heal-publish.yml byte-identical).
  • DESIGN 'Building & checks' is updated through gunbc.design_document. It now names where clippy actually runs and says the unit tests run non-required.
  • The drop rust_unit_tests_off_the_merge_path stays Standing. It retires in the follow-up that makes the step required, after the 6 red tests on main are fixed (routed separately) and the full wall is measured.

The brief named required-witnesses-build / witness_floor_workflow. That job no longer exists, so the step lives in gunbc.compiler_gate_workflow's clippy job.

Measurements: see comments (run 1 was cancelled at the old 45-min cap; the full wall + top-20 follow).

🤖 Generated with Claude Code

…retire rust_unit_tests_off_the_merge_path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

CI measurement 1 — run 36351584784, job 108711170756 (hosted ubuntu-24.04-arm). Lint: 1m22s. Unit-test step: release build 6m24s, then 1155 tests (101 ignored). 575 had passed after 37 min when the 45-min job timeout cancelled it. The cgroup bind held: memory.max read back 12884901888, and HostBudgetUnreadable appears 0 times in the log. Several regen_round_cost_tests take ~95s each.

Six --lib tests are already red on main (they drifted while the tests were off the merge path): construction_authority_graph_tests::wall_now_authority_graph_is_total, floor_skip_frontier_tests::rename_does_not_enrol_names_already_declared_at_source, floor_skip_frontier_tests::rename_still_enrols_a_name_absent_from_the_source, nfr_tests::nfr_roster_receipt, process_cwd_mutation_reachability_gate::no_gating_test_reaches_a_process_cwd_mutator, process_cwd_mutation_reachability_gate::the_closure_still_finds_the_ignored_residue_it_is_scoped_around.

So the drop's retirement is not yet honest: there is no complete wall measurement, and the population is not green. Escalated to the parent for the route.

… clippy timeout so it completes once to be measured; drop stays Standing (ruling neat-boar-16)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot gunbai-bot Bot changed the title Rust unit tests back on the merge path, inside the clippy job (retire rust_unit_tests_off_the_merge_path) Rust unit tests run in the clippy job, non-required until measured (step 1 of retiring rust_unit_tests_off_the_merge_path) Sep 27, 2026
@gunbai-bot

gunbai-bot Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

CI measurement 2: the complete wall. Head 7f06993, job 108723251802, hosted ubuntu-24.04-arm (4 vCPU / 16 GiB). The lint took 1m17s. The rust_unit_tests step took 36m53s, of which 6m14s was the release build and 30m39s the test run (finished in 1838.58s). The whole clippy job ran 38m25s, which is under the 90-min cap on the floor lane. The memory.max bind read back 12884901888, and there are no HostBudgetUnreadable refusals.

Result: 1155 tests, 992 passed, 8 failed, 155 ignored. The step is non-required, so the job concluded success.

The 8 red. Six are the ones already routed. Two are new, because the first run never reached them:

  • compiler_tests::compiler_tests::function_value_named_application_controls_witness: positional function-value application refuses with EffectSummaryIncompleteAtFunctionValue { caller: "hof_positional.host" } where the test expects ADMIT.
  • memory_governor::tests::whole_corpus_compile_another_repositorys_projection_refuses: the expected RefusedDemandsForAnotherRepository { projected_repositories: ["gunbc"] } did not match.

Top 20 by time. These are approximate: each figure is the gap between one test's completion line and the previous one's in the log. Tests run in parallel, so a gap is a lower bound on that test's own time, and stable cargo has no per-test report.

time test
228.1s cli_run::required_regen_host::regen_affected_set_tests::live_tree_controls_land_on_the_measured_arms
94.3s cli_run::heartbeat_tests::render_heartbeat_line_mirror_matches_seed_oracle
93.4s cli_run::required_regen_host::regen_round_cost_tests::live_module_mirrors_have_owners_and_removed_shell_owner_refuses
93.4s cli_run::required_regen_host::regen_emission_scope_tests::the_model_selects_nothing_on_the_unlocatable_arm
92.7s cli_run::required_floor_runner::pure_producer_share_tests::unimported_bare_provider_roster_edit_is_judged_across_frames
92.2s cli_run::required_regen_host::regen_round_cost_tests::host_built_receipt_renders_through_the_model
91.9s cli_run::required_floor_runner::pure_producer_share_tests::a_retirement_against_the_real_base_read_admits_and_a_growth_refuses
90.1s cli_run::required_regen_host::regen_emission_scope_tests::host_selection_and_model_selection_agree
90.1s cli_run::required_regen_host::regen_convergence_host_instrument_tests::install_admission_contains_the_destination_against_a_symlink
89.8s cli_run::required_regen_host::regen_convergence_host_instrument_tests::mutating_transaction_binds_candidates_restores_and_reaches_staged_fixed_point
89.8s cli_run::required_regen_host::regen_convergence_host_instrument_tests::stage_execution_joins_the_plan_to_independently_observed_effects
88.9s cli_run::required_regen_host::regen_convergence_host_instrument_tests::population_joins_refuse_by_identity_and_admit_an_exact_partition
88.6s cli_run::required_regen_host::regen_convergence_host_instrument_tests::mixed_role_two_mirror_candidate_installs_all_then_rebuilds_once
87.7s cli_run::required_regen_host::regen_convergence_host_instrument_tests::install_admission_refuses_unaddressable_and_hand_maintained
86.9s cli_run::required_regen_host::regen_convergence_host_instrument_tests::single_generation_input_still_promotes_alone
81.8s cli_run::required_regen_host::regen_affected_set_tests::an_unlocatable_edited_path_refuses_and_selects_nothing
80.8s cli_run::required_regen_host::regen_affected_set_tests::host_walk_and_model_fold_agree_on_the_fixture_graph
68.2s cli_run::required_regen_host::tests::the_emitted_tree_package_graph_decodes_to_the_compiled_rows_at_a_settled_head
47.2s cli_run::compile_clean::corpus_scope_tests::an_index_over_narrower_roots_does_not_know_the_corpus
35.5s cli_run::required_regen_host::tests::filter_in_branch_condition_refuses_and_does_not_publish_the_module

Almost all of these are required_regen_host / required_floor_runner / live-tree tests that each take ~90s. That is the §3 witness-rule tell: each one re-runs production over the live corpus instead of supplying inputs. They are named here so they can become their own item.

@briansrls
briansrls marked this pull request as ready for review September 27, 2026 23:43
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-27T23:46:11.435037Z 7f06993 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

… (step 36m53s); step stays NON-required until the 8 reds are green (ruling neat-boar-16)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

Head a0da5bb: clippy job timeout is now 60 min, down from 120. That leaves margin over the measured 38m25s job (the unit-test step took 36m53s; the full measurement is in the earlier comment).

The rust_unit_tests step stays NON-required (continue-on-error) until the 8 reds on main are green. Because of that, rust_unit_tests_off_the_merge_path stays Standing in this PR. It is retired in the follow-up that makes the step required. Routing (neat-boar-16):

  • the 8 reds go to their own work items;
  • the ~18 live-tree tests at ~90s each are a separate §3 supply-inputs item.

…sts; cite the measuring run, not its numbers (review 71958)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Commit 1a551d23 addresses review 71958:

  • Two copies of the memory bound. Both copies are gone: heal_publisher_cgroup_memory_max_bytes and compiler_gate_unit_test_memory_max_bytes are deleted. There is now one authority, gunbc.ci_spec hosted_arm_runner_cgroup_memory_max() -> Gibibyte (gibibyte(count: 12)). hosted_runner_memory_cgroup_bind_commands now takes a Gibibyte and renders its byte count through gibibyte_to_byte_size, and both callers pass the shared value.
  • Numbers copied into comments. The comments no longer restate the counts or times. They cite witnesses run 36355834138 (and the first run on this PR), and that run's job record is where the timing is read.
  • Check. I rendered expected_compiler_gate_yml and expected_heal_publisher_workflow_yml from the edited .dag and compared them with the committed witnesses.yml and heal-publish.yml. The bodies are byte-identical apart from the generated header and a trailing newline, so no generated file changes.

— sent from neat-ant-823

…nt at the GHA field (review 71977)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Commit 538905b95f addresses review 71977:

  • Bare Int timeout. data compiler_gate_clippy_timeout_minutes: Int = 60 is now fn compiler_gate_clippy_timeout() -> Minute { minute(count: 60) }. It becomes a plain number only at the GHA timeout_minutes field, through minute_count(...), the same pattern gunbc.fleet_converge_workflow uses.
  • Check. I re-rendered expected_compiler_gate_yml from the edited .dag. Its body is byte-identical to the committed witnesses.yml apart from the generated header and a trailing newline.
  • Left alone. compiler_gate_timeout_minutes: Int = 45 is older debt that this PR doesn't touch, as the review itself says.

— sent from neat-ant-823

…host effect via the bash emitter (#5828) (review 72000)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Commit ebf69a5375 addresses review 72000:

  • Hand-written shell bind with no dissolution marker. The bind now carries a marker, gunbc.ci_spec hosted_runner_memory_cgroup_bind_shell_emit_dissolution_trigger (a DissolutionCondition, following the pattern of gunbc.native_frontier_workflow's hand-shell markers). It names both callers, the heal publisher prelude and compiler_gate_unit_test_command. Its trigger is the cgroup v2 leaf bind becoming a typed host effect rendered by the bash emitter (General orchestration intent to Bash emit fold over grammar rows #5828 / shell-to-intent Phase 2), taking the leaf name and a Gibibyte bound as values. When that lands, both callers switch to it and this function is deleted.
  • Why I didn't model the bind as structured steps here. That typed op doesn't exist yet. Building it is the General orchestration intent to Bash emit fold over grammar rows #5828 capability, which is outside this PR's brief.
  • Check. heal-publish.yml still renders byte-identical (apart from the generated header and a trailing newline). The marker is data, and the rendered YAML doesn't change.

— sent from neat-ant-823

@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 28, 2026

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verdict: APPROVE ebf69a5

No blocking finding for the explicitly ruled non-required observation step. The emitted rust_unit_tests step consumes repo_self_test_command after the required lint; continue-on-error is attached to that step, not to lint or the whole job. Its wrapper retains the actual command exit and failure receipt. The shared cgroup bind sets memory.max before moving the invoking shell into the leaf, so this is a real inherited bound rather than a planning override. The heal publisher retains its original bind sequence through the shared producer. The shell dissolution condition names the capability that replaces both callers.

The source and emitted workflow agree on a 60-minute clippy job timeout, not the stale 120-minute statement in the PR body. The measured complete earlier job was 38m25s; the recorded 992 passed / 8 failed / 155 ignored result is not a green test population. The Standing rust_unit_tests_off_the_merge_path drop is correctly NOT retired here, and DESIGN names the non-required status. A future required-step promotion still owes a green population and the acceptance-path cost/runner conditions. Job-level timeout or cancellation is not neutralized by a step's continue-on-error and must not be described as such.

Evidence limitation: completion-line gaps from parallel tests are not established per-test durations, so I do not credit the approximate top-20 table as a per-test cost profile. That does not invalidate the independently reported complete step/job wall or this observation-only enrollment. Exact-head workflow 36369118298 succeeded; that success is not a claim that the optional unit tests passed. Source/recorded-evidence review, no independent local rerun.

Merged via the queue into main with commit 5bf2b13 Sep 28, 2026
6 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/neat-ant-823 branch September 28, 2026 13:29
@briansrls
briansrls restored the session/neat-ant-823 branch September 28, 2026 13:39
gunbai-bot Bot pushed a commit that referenced this pull request Sep 29, 2026
…ool when its block ends

Reproduced the CI unit-step stall under the hosted runner's exact bind (a 12 GiB
memory.max cgroup leaf, 24 GB BuildBuddy runner):
- 7.6 GiB of freed-but-retained heap from earlier tests' glibc arenas sat
  under the first claim, and the pool built on top of it was OOM-killed.
  trim_retained_heap (the existing malloc_trim instrument) now runs before
  and after each claim: 1.4 GiB under the first claim.
- The ~5.2 GiB pool then outlived its block and the next test's own regen-pool
  index was OOM-killed on top of it. The thread now exits after
  LIVE_POOL_IDLE_RELEASE (2s) with no claim, dropping the pool; a later claim
  rebuilds it. Send and exit share one lock, so no claim is lost to the exit.
  New control: the_live_pool_is_released_when_its_block_of_claims_ends.

Full suite under the bound: 1006 passed, 7 failed (6 on #12450's list of reds
on main; the_authority_entry_resolves_from_a_non_root_cwd passed on hosted CI at
7c2d8ce and is not on this path), 529.6s, memory.events max=0 oom=0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant