Skip to content

test(kanban): make progress-stall policy tests probe-deterministic (t_5457397a) - #1040

Merged
Kyzcreig merged 1 commit into
mainfrom
fix/t_5457397a-stall-probe-deterministic
Sep 25, 2026
Merged

Kyzcreig merged 1 commit into
mainfrom
fix/t_5457397a-stall-probe-deterministic

Conversation

@Kyzcreig

@Kyzcreig Kyzcreig commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Kanban: t_5457397a

Problem. tests/hermes_cli/test_kanban_progress_stall.py waited for the REAL ps probe to read a freshly-spawned child as 0.0% CPU (6 consecutive samples, 30 s deadline) before every stall/reclaim policy test. On Linux, procps pcpu is lifetime cputime/elapsed, so a child whose startup burned ~30 ms needs ~30 s to round to 0.0 — much longer on a loaded runner. Result: 6/10 failed together (never read idle to the real probe) in merge-group batches 36060149219 and 36064421258, ejecting unrelated PRs #1002 and #994.

Fix.

  • kanban_db._worker_cpu_active split into _process_cpu_table() (the only I/O) + _cpu_active_in_table() (pure parse). Behaviour unchanged: unknown state → True (never authorizes a kill); children never veto.
  • Policy tests keep a real worker process (reclaim still really SIGTERMs it) but feed the probe a deterministic table through the seam. The table is parsed by production code, so children-never-veto stays covered.
  • Real-probe integration kept: busy process → active (load-immune direction) and new test_real_probe_reads_in_flight_worker_as_idle, which SKIPs with the measured reason (probe errors / loadavg1 > cores) and FAILs only on an unloaded host with a working ps.
  • Probe-failure tests still drive the real probe into a broken subprocess.run.
  • Reclaim/stall/escalation assertions unchanged.

Verified (Mac Studio, repo venv).

  • 11/11 pass (19 s).
  • 5/5 file runs green under 64 busy loops on 32 cores (loadavg1 up to 161), 54–73 s each, 0 skips.
  • 6/6 kanban_db mutants killed: any-child veto, probe-failure-authorizes, no escalation, CPU veto ignored, probe always-active, no reclaim.

Follow-up (separate card): on Linux the production probe's lifetime-average pcpu means a long-lived worker with meaningful accumulated CPU may read "active" long after it goes idle (veto, fail-safe direction) — worth measuring on ACE-AI.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

…_5457397a)

test_kanban_progress_stall.py waited for the REAL ps probe to read a fresh
child as 0.0% CPU inside 30 s before every policy test. On Linux procps pcpu
is lifetime cputime/elapsed, so a child whose startup burned ~30 ms needs
~30 s to round to 0.0 — far longer on a loaded runner. 6/10 tests failed in
merge-group batches 36060149219 and 36064421258, ejecting #1002 and #994.

- kanban_db: split _worker_cpu_active into _process_cpu_table() (the only
  I/O) + _cpu_active_in_table() (pure parse). Behaviour unchanged: unknown
  state still returns True; children still never veto.
- Policy tests keep a real worker process (reclaim really terminates it)
  but feed the probe a deterministic table via the seam; the table is still
  parsed by production code, so the children-never-veto rule stays covered.
- Real-probe integration kept: busy -> active (load-immune direction) and a
  new in-flight -> idle test that SKIPs with measured reason (probe errors /
  loadavg > cores) and FAILs only on an unloaded host with a working ps.
- Probe-failure tests still drive the real probe into a broken ps.

Verified: 11/11 locally; 5/5 runs green under 64 busy loops on 32 cores
(loadavg1 up to 161); 6/6 mutants of kanban_db killed (any-child veto,
failure-authorizes, no escalation, veto ignored, always-active, no reclaim).
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🤖 merged-by: apollo · lane: claude-bridge-session · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 · why: t_5457397a: probe seam extraction, production parse unchanged (_cpu_active_in_table), policy tests fed a deterministic ps table, real-probe integration test kept (SKIP-with-reason on host load); 38 checks green, Apollo self-review; ejected #1002/#994/#988 batches today

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🤖 merged-by: apollo · lane: claude-bridge-session · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 · why: re-arm after merge-queue ejection by the test_kanban_progress_stall load flake on the local pool (fix #1040); PR untouched

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🤖 merged-by: apollo · lane: claude-bridge-session · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 · why: re-arm (auto-merge disable+enable) after merge-queue ejection by the test_kanban_progress_stall load flake on the local pool; fix is #1040; PR untouched

@Kyzcreig
Kyzcreig enabled auto-merge September 24, 2026 23:44
@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 24, 2026
@Kyzcreig
Kyzcreig removed this pull request from the merge queue due to a manual request Sep 25, 2026
@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 25, 2026
Merged via the queue into main with commit 7fd339d Sep 25, 2026
100 of 103 checks passed
@Kyzcreig
Kyzcreig deleted the fix/t_5457397a-stall-probe-deterministic branch September 25, 2026 02:49
@Kyzcreig Kyzcreig added the fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent) label Sep 25, 2026
@ang-prism

ang-prism Bot commented Sep 27, 2026

Copy link
Copy Markdown

FleetReview

Review: post-merge · head 7fd339db7262 · duration 5m 05s
Profile: light (merit: default light: lines 164<800, files 2<1000000, hunks 12<1000000, no hot path) · policy: below-size-and-path-gates
Roster: B-state → gpt-6-sol (openai), C-assert-xhigh → claude-code-opus-5-5 (anthropic), F → gpt-6-sol (openai), G → grok-4.6 (xai)

Post-merge review (fleetreview:post-merge override): this reviewed the merge commit against its first parent — the bytes that already shipped. It is not a pre-merge gate pass.

profile: light (rule: default light: lines 164<800, files 2<1000000, hunks 12<1000000, no hot path) · round 0 · members: B-state, C-assert-xhigh, F, G · families: anthropic,openai,xai

Confidence: 4/5

No issues found.


FleetReview provenance · models: B=gpt-6-sol, C=claude-code-opus-5-5, D=grok-4.6, F=gpt-6-sol · cost: $0.66 · duration: 5m 01s · rounds: 1 · files examined: 2

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant