test(kanban): make progress-stall policy tests probe-deterministic (t_5457397a) - #1040
Conversation
…_5457397a) test_kanban_progress_stall.py waited for the REAL ps probe to read a fresh child as 0.0% CPU inside 30 s before every policy test. On Linux procps pcpu is lifetime cputime/elapsed, so a child whose startup burned ~30 ms needs ~30 s to round to 0.0 — far longer on a loaded runner. 6/10 tests failed in merge-group batches 36060149219 and 36064421258, ejecting #1002 and #994. - kanban_db: split _worker_cpu_active into _process_cpu_table() (the only I/O) + _cpu_active_in_table() (pure parse). Behaviour unchanged: unknown state still returns True; children still never veto. - Policy tests keep a real worker process (reclaim really terminates it) but feed the probe a deterministic table via the seam; the table is still parsed by production code, so the children-never-veto rule stays covered. - Real-probe integration kept: busy -> active (load-immune direction) and a new in-flight -> idle test that SKIPs with measured reason (probe errors / loadavg > cores) and FAILs only on an unloaded host with a working ps. - Probe-failure tests still drive the real probe into a broken ps. Verified: 11/11 locally; 5/5 runs green under 64 busy loops on 32 cores (loadavg1 up to 161); 6/6 mutants of kanban_db killed (any-child veto, failure-authorizes, no escalation, veto ignored, always-active, no reclaim).
|
🤖 merged-by: apollo · lane: claude-bridge-session · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 · why: t_5457397a: probe seam extraction, production parse unchanged (_cpu_active_in_table), policy tests fed a deterministic ps table, real-probe integration test kept (SKIP-with-reason on host load); 38 checks green, Apollo self-review; ejected #1002/#994/#988 batches today |
|
🤖 merged-by: apollo · lane: claude-bridge-session · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 · why: re-arm after merge-queue ejection by the test_kanban_progress_stall load flake on the local pool (fix #1040); PR untouched |
|
🤖 merged-by: apollo · lane: claude-bridge-session · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 · why: re-arm (auto-merge disable+enable) after merge-queue ejection by the test_kanban_progress_stall load flake on the local pool; fix is #1040; PR untouched |
FleetReviewReview: post-merge · head Post-merge review ( profile: light (rule: default light: lines 164<800, files 2<1000000, hunks 12<1000000, no hot path) · round 0 · members: B-state, C-assert-xhigh, F, G · families: anthropic,openai,xai Confidence: 4/5 No issues found. FleetReview provenance · models: B=gpt-6-sol, C=claude-code-opus-5-5, D=grok-4.6, F=gpt-6-sol · cost: $0.66 · duration: 5m 01s · rounds: 1 · files examined: 2 |
Kanban: t_5457397a
Problem.
tests/hermes_cli/test_kanban_progress_stall.pywaited for the REALpsprobe to read a freshly-spawned child as 0.0% CPU (6 consecutive samples, 30 s deadline) before every stall/reclaim policy test. On Linux, procpspcpuis lifetime cputime/elapsed, so a child whose startup burned ~30 ms needs ~30 s to round to 0.0 — much longer on a loaded runner. Result: 6/10 failed together (never read idle to the real probe) in merge-group batches 36060149219 and 36064421258, ejecting unrelated PRs #1002 and #994.Fix.
kanban_db._worker_cpu_activesplit into_process_cpu_table()(the only I/O) +_cpu_active_in_table()(pure parse). Behaviour unchanged: unknown state → True (never authorizes a kill); children never veto.test_real_probe_reads_in_flight_worker_as_idle, which SKIPs with the measured reason (probe errors / loadavg1 > cores) and FAILs only on an unloaded host with a workingps.subprocess.run.Verified (Mac Studio, repo venv).
Follow-up (separate card): on Linux the production probe's lifetime-average
pcpumeans a long-lived worker with meaningful accumulated CPU may read "active" long after it goes idle (veto, fail-safe direction) — worth measuring on ACE-AI.Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.