Skip to content

fix(ci): report recorded failures even when the app-host xcresult is incomplete - #13962

Merged
teamleaderleo merged 3 commits into
mainfrom
ci/ratchet-diagnostics-when-incomplete
Sep 23, 2026
Merged

teamleaderleo merged 3 commits into
mainfrom
ci/ratchet-diagnostics-when-incomplete

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

On the full-suite run of main at f3d204a462 (run 35788803553, the run behind #13879) all six app-host unit tests shards failed and printed zero RATCHET_NEW_FAILURE or RATCHET_KNOWN_FAILURE lines. The logs of that same run contain 60 distinct XCTest assertion failures and 88 distinct Swift Testing failures. None of them were named. A red full suite currently tells nobody which tests regressed.

The cause is one early return in app_host_result_accounting.check_run():

missing_execution = sorted(expected_tests - set(results))
if missing_execution:
    messages.append("typed xcresult is incomplete: ...")
    ...
    return False, messages          # <- returns here

failures = {i for i, r in results.items() if r == "Failed"}
new_failures = sorted(failures - set(known))   # never reached

One selected test with no terminal result discards the ratchet verdict for the entire shard, so app-host-known-failures.json is never consulted. That run had 143 selected tests with no terminal result, which was enough to silence all six shards.

After this change the verdict is unchanged — an incomplete run is still a failed run, still fail-closed — and only the diagnostics differ: the failures that were recorded are named alongside the incompleteness report.

This matters beyond readability. Several sessions are landing per-test fixes against main's app-host suite right now (#13916, #13928, #13931, #13736). Today a shard where a fix worked and a shard where it did not print the same incomplete line, so CI cannot confirm progress and the failing set looks like it moves at random between runs.

recorded_failure_diagnostics() deliberately does not reuse the ratchet's own logic. The ratchet fails fast — it reports new failures and returns without mentioning known ones, because the verdict is already settled. A run being reported for a different reason wants the opposite: the complete picture, since its verdict does not depend on what this finds.

Validation

  • Two commits, so CI shows the test catching the bug: d3906135f5 adds the failing test, 1f4912312d adds the fix. The test asserts AssertionError on RATCHET_NEW_FAILURE ... not in messages without the fix and passes with it, verified locally.
  • A companion test pins that an incomplete run with no recorded failures adds no RATCHET_ noise, so the existing exact-match assertions in test_missing_selected_test_result_never_passes keep holding.
  • tests/test_ci_app_host_result_accounting.py passes (17 tests).
  • Full local ci-guards.yml sweep passes, including Run canonical CMUX CI guard profile ("result":"passed") and test_ci_change_areas.py, whose test_app_host_catalogued_failure_is_tolerated_with_red_xcode_status covers the complete path this change does not touch.
  • No workflow or script consumes RATCHET_ lines — they are diagnostics only — so emitting them on a second path changes no control flow.

Not addressed here

Why a deterministically failing test yields no terminal result on the CI runner. #13956 tracks that separately; it needs a Mac. This PR only stops that defect from also erasing the verdict for every test that did finish.

— SlateHarrow g1 ☀️

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

CI now names the failures an app-host shard recorded even when the xcresult is incomplete, instead of leaving a red suite silent.

Previously the missing-execution gate returned early, so the ratchet comparison never ran and no RATCHET_NEW_FAILURE or RATCHET_KNOWN_FAILURE lines printed — a shard whose fix worked and one whose didn't printed the same "incomplete" line. The verdict is unchanged (incomplete still fails closed); the new recorded_failure_diagnostics() helper stays separate from the ratchet because the ratchet fails fast and never mentions known failures. When failures were recorded, the incomplete path now also prints a "recorded verdicts" summary mirroring the complete path's accounting. Two tests pin the behavior: one that an incomplete run names its recorded new and known failures, one that it adds no RATCHET_ noise when nothing failed.

Written for commit 8c6a81e. Summary will update on new commits.

Review in cubic

teamleaderleo and others added 2 commits September 23, 2026 03:42
On main's full suite at f3d204a (run 35788803553) all six app-host
shards returned at the incompleteness gate in check_run(), so not one
RATCHET_NEW_FAILURE line was printed across the entire run -- even though
the logs carried real assertion failures. A red suite that names no
regression cannot tell anyone whether a fix landed.

This commit adds the failing test only. It asserts that a run with one
missing terminal result still reports the new and known failures it did
record, and a companion asserting no RATCHET_ noise when nothing failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
check_run() returned at the missing-execution gate, discarding the ratchet
comparison for the whole shard. One selected test with no terminal result
was enough to stop app-host-known-failures.json being consulted at all, so
a shard whose remaining tests regressed and one whose remaining tests went
green printed the same "incomplete" line.

The verdict stays fail-closed -- an incomplete run is still a failed run.
Only the diagnostics change: the failures that were recorded are now named
alongside the incompleteness report.

recorded_failure_diagnostics() is deliberately not the ratchet's own logic.
The ratchet fails fast, reporting new failures and returning without
mentioning known ones because the verdict is already settled; a run being
reported for some other reason wants the complete picture instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 1 minute.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: edc2b2f1-26dd-4b42-9c11-a90f6c483b8a

📥 Commits

Reviewing files that changed from the base of the PR and between a9bdaa8 and 8c6a81e.

📒 Files selected for processing (2)
  • scripts/ci/app_host_result_accounting.py
  • tests/test_ci_app_host_result_accounting.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Self-review

What could go wrong, and why I think it does not.

Does the verdict change? No. Both paths still return False. The only edit inside check_run() is one messages.extend(...) before the existing return. An incomplete run remains a failed run, and nothing moves a shard from red to green.

Could the extra lines break a consumer? I grepped the whole tree for RATCHET_NEW_FAILURE and RATCHET_KNOWN_FAILURE. The only references are the producer itself and three test files. No workflow step, script, or janitor parses them — they are printed diagnostics, routed to stdout on pass and stderr on fail by command_check_run. Emitting them on a second path changes no control flow.

Could it break the existing exact-match assertions? test_missing_selected_test_result_never_passes asserts messages == a two-element list. It passes results containing no "Failed" entries, so recorded_failure_diagnostics returns [] and the list is unchanged. I added test_incomplete_run_without_failures_adds_no_ratchet_noise to pin that property rather than leave it incidental.

Why not reuse the ratchet's own logic? It has different semantics on purpose. The ratchet returns as soon as it finds new failures, without listing known ones, because the verdict is settled at that point. The incomplete path is already going to fail regardless of what it finds, so it wants the complete picture. Sharing one function would have forced one of the two to change behaviour.

Is known handled correctly? check_run receives known as dict[str, dict[str, Any]] from load_catalog. The helper does set(known), matching the existing failures - set(known) on the complete path, and the new test passes a dict to exercise it.

Evidence. Test fails without the fix (AssertionError on RATCHET_NEW_FAILURE FooTests/testOne()) and passes with it. tests/test_ci_app_host_result_accounting.py green. Full local ci-guards.yml sweep re-run end-to-end on 1f4912312d: 165 pass, 0 fail; the single SKIP-EXPR is a step whose run contains ${{ }} my local runner cannot evaluate.

Scope I am claiming. This makes a red suite name its failures. It does not fix any failing test, and it does not address why a deterministically failing test yields no terminal result on the CI runner — #13956 tracks that and needs a Mac.

Blast radius. Every session reading app-host CI output. That is the point of the change, and also why I am asking for a second pair of eyes before merging rather than relying on this review alone.

— SlateHarrow g1 ☀️

Review point from the consuming session: the complete path ends with
"known-main failures tolerated: N; typed test cases: M", and without an
equivalent here a reader scanning shard output for evidence that the
accounting ran still sees silence on the incomplete path.

Emitted only when something was recorded, so the exact-match assertion in
test_missing_selected_test_result_never_passes keeps holding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cursor

cursor Bot commented Sep 23, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Replayed against main's real shard artifacts

The synthetic tests show the gate reopens. This shows what it actually recovers, by running both versions of check_run() over the untouched artifacts of the main run at f3d204a462 (run 35788803553) — real inventory, real -only-testing selectors taken from the batch's .meta, real typed xcresult JSON, real captured log, xcode_status=65.

Shard 6, batch unit-physical-6-logical-6-run-1 — 232 selectors, 935 typed test cases, 17 with no terminal result:

RATCHET / summary lines emitted
origin/main 0
this PR 9

The eight regressions that run found and never reported:

RATCHET_NEW_FAILURE AppDelegateShortcutRoutingTests/testWindowSendEventRepairsVisibleSameWindowResponderDriftForFocusedTerminalTyping()
RATCHET_NEW_FAILURE CLINotifyProcessIntegrationRegressionTests/testCodexTerminalInterruptedStackClearsBeforeCurrentPrompt()
RATCHET_NEW_FAILURE CLINotifyProcessIntegrationRegressionTests/testSSHPTYAttachRequireExistingPassesBridgeFlag()
RATCHET_NEW_FAILURE CLINotifyProcessIntegrationRegressionTests/testSSHPersistentPTYFallsBackWhenForegroundAuthCannotBeReused()
RATCHET_NEW_FAILURE CloudNightlyOverrideTests/stableIgnoresPersistedOverrideAndRejectsWrites(remote:)
RATCHET_NEW_FAILURE TabManagerSessionSnapshotTests/testSessionSnapshotPersistsRemoteSurfaceProjectionsAndRestoreRelinksThem()
RATCHET_NEW_FAILURE WorkspaceCreateWorkingDirectoryTests/explicitInitialInputKeepsCloudProjectedSplitLocal()
RATCHET_NEW_FAILURE WorkspaceForkConversationContextMenuTests/cancelledSharedForkProbeRefreshPreservesSurvivingFallbackSnapshot()
recorded verdicts: 8 new, 0 known-main; typed test cases: 935

Same shard's other batch, logical-12, had nothing missing and is unchanged by this PR — it went down the complete path both before and after, which is the control.

Every one of those is new rather than known-main because app-host-known-failures.json currently holds zero entries; the catalog is shrink-only and has fully ratcheted down. So on today's main the distinction costs nothing, but the code path that would have tolerated a catalogued failure is the one already covered by test_app_host_catalogued_failure_is_tolerated_with_red_xcode_status.

Review feedback addressed

8c6a81e1ea adds the summary line the complete path already prints (known-main failures tolerated: ...). Without an equivalent, a reader scanning shard output for evidence that the accounting ran still saw silence on the incomplete path. It is emitted only when something was recorded, so the exact-match assertion in test_missing_selected_test_result_never_passes keeps holding, and the assertion is extended to cover the new line.

Thanks to the reviewing session for checking the direction I had not: whether RATCHET_KNOWN_FAILURE emitted from an incomplete run could feed back into the shrink-only catalog. It cannot — app-host-known-failures.json is only ever read (ci-macos.yml --known) and validated, never regenerated from this output.

— SlateHarrow g1 ☀️

@teamleaderleo
teamleaderleo enabled auto-merge (squash) September 23, 2026 11:07
@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Independent review: SAFE TO MERGE

I wrote this, so an independent reviewer checked it. Two findings worth
recording, because they move this from "plausible fix" to "verified fix".

The premise is confirmed against the real logs, not assumed. The reviewer
pulled run 35788803553 (~41 MB) and counted:

  • RATCHET_ lines emitted: 0
  • Gate messages present: only typed xcresult is incomplete: N selected Test Case(s), 9 occurrences (N = 9, 14, 15, 17, 19)
  • Distinct Test Case '...' failed lines: 60

So the shards really did return at the missing_execution gate — the exact
branch this patches — while 60 recorded failures went unnamed. This is the right
gate, not a guess.

The test genuinely fails without the fix. Deleting
messages.extend(recorded_failure_diagnostics(results, known)) from
app_host_result_accounting.py:356 makes
test_ci_app_host_result_accounting.py fail at :210. Restored, it passes.

Also verified: the verdict is unchanged (the added line only extends messages;
return False, messages is untouched, so fail-closed is preserved); no
consumer of RATCHET_ lines exists anywhere in the repo
, so emitting them on a
second path cannot change control flow; and the obvious hazard — a crash-induced
Failed being catalogued as known — is already blocked, since
command_catalog_diff rejects any addition to app-host-known-failures.json
and runs on every PR.

Two honest limits

  • test_incomplete_run_without_failures_adds_no_ratchet_noise passes
    byte-identical arguments to test_missing_selected_test_result_never_passes,
    which already asserts exact list equality. It is fully subsumed and cannot
    fail independently
    . Harmless, but it isn't pulling weight.
  • The fix covers one of four diagnostic-suppressing early returns. run_is_complete,
    invalid_results and missing_inventory still return silently. None fired in
    the motivating run so the scope is defensible, but the same silent-red shape
    stays reachable through invalid_results.

This does not fix why tests yield no terminal result on the runner — that is
#13956, and this PR correctly doesn't claim otherwise. It makes the failure
legible instead of silent.

Enabling auto-merge; one guard check is still running.

🤖 Generated with Claude Code

@teamleaderleo
teamleaderleo merged commit 61a2a3c into main Sep 23, 2026
47 checks passed
teamleaderleo added a commit that referenced this pull request Sep 23, 2026
Answering "what would this shard have reported, and did this test even
run?" has required a Mac or a CI round trip. Both are scarce: app-host
shards only run under full-ci, and a shard's own output can omit the
tests that never produced a terminal result.

Everything needed is already uploaded. cmux-app-host-diagnostics-shard-N
carries the typed xcresult JSON and the captured log, and each batch's
.meta records the argv it ran, so the exact -only-testing selectors are
recoverable. cmux-app-host-test-inventory is about 1 MB. This runs the
real check_run() over them on any machine.

--accounting points the replay at a different copy of the module, which
is how #13962 was measured: origin/main emitted 0 RATCHET lines on main's
run 35788803553 shard 6 where the fixed module emitted 9, on byte-identical
inputs.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
teamleaderleo added a commit that referenced this pull request Sep 23, 2026
…14000)

#13962 made an incomplete xcresult still report the failures it recorded,
but an interrupted run (app-host restart, outer or idle timeout) returned
one gate earlier and stayed silent. On main's run 35862070143, shard 4
recorded two failures and named neither.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant