Skip to content

Fix four unsatisfiable tests in the sidebar git suites - #8723

Merged
austinywang merged 2 commits into
manaflow-ai:mainfrom
ejc3:fix/git-index-test-fixtures
Aug 4, 2026
Merged

austinywang merged 2 commits into
manaflow-ai:mainfrom
ejc3:fix/git-index-test-fixtures

Conversation

@ejc3

@ejc3 ejc3 commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

Four tests in the sidebar git suites could never pass. On current main they fail for three
unrelated reasons, and this change fixes all four.

Two tests describe a file git cannot produce

testGitIndexVersionFourRefreshTracksIndexSignatureChanges and
testEmptyGitIndexRefreshTracksIndexSignatureChanges each wrote an index twice and changed only the
trailing 20-byte checksum, leaving the entry table byte-identical, then asserted the sidebar went
dirty.

The index trailer is the SHA-1 of the index content, so identical content cannot carry two different
trailers. No repository reaches that state. The product reads it deliberately:
gitIndexContentSignature hashes the entry count and each entry's path, mode and object id, and never
the trailer, so both writes produce the same content signature. The apply then takes the rebaseline
branch in SidebarGitMetadataService+Probe.swift — index signature changed, content signature
unchanged, so the stored clean signature moves forward and the panel stays clean. That is what
testCleanIndexSignatureRebaselinesWhenIndexRewriteKeepsTrackedContentClean pins, immediately after
the v4 test in the same file, and it passes. Two tests asked for opposite outcomes from the same input, so one had to fail, and the
rebaseline test is the correct one: it describes a stash-like rewrite.

Both now stage a real change — a new object id for the v4 index, an added entry for the empty one —
with a stat that still matches the worktree, so the stat scan stays clean and the dirty verdict has to
come from the content signature, which is the path a real staged change takes. The predicates are
unchanged. The empty-index test's scenario and message do change, from a staged delete to a staged
add, because staging the first entry out of an empty index is what actually moves the content
signature; its name still says "TracksIndexSignatureChanges", which remains true.

writeGitIndexVersion4 gains the objectIDBytes parameter that
writeGitIndexVersion2EntryFromStat already had. It defaults to the zero id, so the callers that do
not stage a change need no edit. writeGitIndexVersion3SkipWorktreeEntry still hard-codes a zero object id —
it has no need to express a staged change.

A scoped branch report went to a manager nothing could resolve

testDisablingGitWatchClearsCachedPullRequestBadgesWhenPullRequestsAreShownByDefault seeded a badge,
then sent report_git_branch with both --tab and --panel. That scoped path resolves its workspace
through AppDelegate's main-window contexts, not through TerminalController's active manager, so a
manager only the controller knew about was invisible to it: the handler returned before reaching the
clear and the seeded entry survived. A nearby report_pr test passes because the PR path resolves through controlSidebarTabForMutation, which checks the controller's own manager first and
then falls back to AppDelegate.

The test now registers a windowless context the way the other socket-routing tests do, and asserts the
workspace is resolvable before sending, so a future wiring break reports at that line instead of at
the far assertion. This is test wiring, not a product defect: tabManagerFor(tabId:) resolves every
real window, so no user-visible path depends on the difference.

An async test waited in a way that could only expire

testSameDirectoryInitialGitMetadataProbesShareOneSnapshotRead used a helper that blocks the thread on
XCTWaiter while polling through DispatchQueue.main. That works in a synchronous test, whose body
runs from an ordinary run-loop callback. This body is async on a @MainActor class, so it runs as a
main-actor job — already inside a main-queue drain, which libdispatch will not re-enter. The nested
run loop therefore ran neither the helper's own poll hops nor the snapshot's MainActor.run apply, and
the wait expired every time while the work it waited for landed immediately after the body suspended
again.

I confirmed the mechanism with a standalone binary rather than reasoning about it: a nested
CFRunLoopRunInMode entered from a main-actor async job does not run DispatchQueue.main blocks,
while the same nested loop entered from a run-loop callback does.

The test now awaits a suspending sibling helper with the same 3s budget and 0.05s interval. It is a
separate function rather than an inline loop so the next async test does not reintroduce the blocking
form, and its doc comment says why. It was the only async test in this target using the blocking
helper.

When reading a failure log in these suites, note that a timeout from the file-private helper is
reported twice, once for the XCTAssertTrue and once for the helper's own XCTFail, because its
line: UInt = #line resolves at the call site. Consecutive line numbers are one failure, not two.

Test plan

Measured on 4253cc2884, one GUI test host at a time, same checkout and DerivedData for both arms:

xcodebuild test -scheme cmux-unit -configuration Debug -destination 'platform=macOS' \
  -only-testing:cmuxTests/WorkspacePullRequestSidebarTests \
  -only-testing:cmuxTests/TabManagerPullRequestProbeTests \
  -only-testing:cmuxTests/TabManagerWorkspaceOwnershipTests \
  -only-testing:cmuxTests/CLICodexHookTimeoutRegressionTests
arm XCTest failures across a 36-test run distinct tests red
main unchanged 15 of 36 9
this branch 8 of 36 5

No host restarts in either arm, and CLICodexHookTimeoutRegressionTests reports
Test run with 8 tests in 1 suite passed in both.

The four tests this fixes are testGitIndexVersionFourRefreshTracksIndexSignatureChanges,
testEmptyGitIndexRefreshTracksIndexSignatureChanges,
testDisablingGitWatchClearsCachedPullRequestBadgesWhenPullRequestsAreShownByDefault and
testSameDirectoryInitialGitMetadataProbesShareOneSnapshotRead.

Pull-request CI on this repo runs review bots and security scanners, not the
test suite, so the arms above are the only test evidence this carries.

Still red on this branch, and not addressed here

Summary by CodeRabbit

  • Tests
    • Improved reliability of async UI assertions by adding a suspending wait helper that polls a main-thread condition until it succeeds or times out.
    • Expanded Git metadata refresh and pull request sidebar coverage, including scenarios where dirty/clean state changes based on index content signatures (now controllable per v4 index entry).
    • Updated workspace/scoped git-branch reporting tests to ensure consistent workspace resolution across different window contexts.

@coderabbitai

coderabbitai Bot commented Jul 23, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Tests now use suspending condition polling for asynchronous panel updates. Git index fixtures accept controlled object IDs, and sidebar tests cover app-delegate workspace resolution and staged index signature changes.

Changes

Tab manager async test

Layer / File(s) Summary
Suspending condition polling
cmuxTests/TabManagerUnitTests.swift
Adds an async main-actor polling helper and uses it to await Git branch updates across panels.

Sidebar Git metadata tests

Layer / File(s) Summary
Git index refresh coverage
cmuxTests/WorkspacePullRequestSidebarTests.swift
Adds configurable v4 object IDs, registers a windowless app-delegate context, and updates index signature refresh scenarios for staged content changes.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

  • manaflow-ai/cmux#8725: Both PRs update condition-waiting helpers to avoid blocking during asynchronous test polling.

Suggested reviewers: austinywang

🚥 Pre-merge checks | ✅ 24 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 11.11% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (24 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: fixing four failing sidebar Git tests.
Description check ✅ Passed The description has a strong Summary and Testing section and covers scope and rationale, though some template sections like Demo Video and Checklist are missing.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Cmux Swift Actor Isolation ✅ Passed PASS: the diff only changes cmuxTests/*.swift test code, and the actor-isolation rules explicitly allow tests; no production Swift isolation debt was introduced.
Cmux Swift Blocking Runtime ✅ Passed PASS: The only new sleep/polling logic is a @MainActor helper in cmuxTests, and the other edits are test-only Git fixtures; no production Swift sync was added.
Cmux Browser Automation Off-Main ✅ Passed PASS: the PR only changes test helpers/sidebar tests; no browser.* routing, WebKit/AppKit worker-lane commands, or policy files were touched.
Cmux Expensive Synchronous Load ✅ Passed Diff only changes two test files; no production Swift path adds RestorableAgentSessionIndex.load or similar expensive sync agent-history work on MainActor.
Cmux Cache Substitution Correctness ✅ Passed PR only changes test files; no production Swift/TS/JS cache-to-authoritative-read substitution in persistence/history/undo/snapshot paths.
Cmux No Hacky Sleeps ✅ Passed PASS: The rule excludes Swift/test-only code, and the touched files are cmuxTests Swift tests, not TS/JS/shell runtime scripts.
Cmux Algorithmic Complexity ✅ Passed Only test files changed; the complexity rule exempts test-only scaffolding, and no production hot-path collection scans were introduced.
Cmux Swift Concurrency ✅ Passed Patch is test-only and replaces blocking wait with an awaited helper; it adds no new background Dispatch, Combine, completion-handler, or fire-and-forget Task patterns.
Cmux Swift @Concurrent ✅ Passed The diff adds only UI-bound @MainActor async waiting and synchronous test helpers; no new nonisolated async work or misplaced @concurrent appears.
Cmux Swift Package Boundaries ✅ Passed Only test targets changed; no production Swift logic was added or left in app-target Sources/ code.
Cmux Swiftpm Lockfiles ✅ Passed Diff only changes two test files; no Package.swift, .gitignore, workflow, or Xcode package-reference changes, so no Package.resolved rule is triggered.
Cmux Swift Logging ✅ Passed The diff only changes tests/helpers and adds no print/debugPrint/dump/NSLog/Logger or ad hoc stdout/file diagnostics.
Cmux User-Facing Error Privacy ✅ Passed Only test files changed, and the rule explicitly allows tests and developer-only comments; no production user-facing error copy was modified.
Cmux Full Internationalization ✅ Passed PASS: HEAD only changes test files; no production Swift/UI strings, catalogs, Info.plist, or web locale files were touched.
Cmux Swiftui State Layout ✅ Passed Diff only changes test helpers and Git index test data; no SwiftUI view/state/layout code was introduced or modified.
Cmux Architecture Rethink ✅ Passed Diff is test-only; the new polling helper is explicitly test synchronization, which the rule allows, and no production ownership/lifecycle wiring changed.
Cmux Swift Auxiliary Window Close Shortcuts ✅ Passed Only cmuxTests files changed; no NSWindow/NSPanel/WindowGroup or cmuxAuxiliaryWindowIdentifiers code was introduced, so the rule doesn't apply.
Cmux Source Artifacts ✅ Passed Both changed paths are hand-written test sources; no logs, build output, caches, temp folders, or other artifacts appear in the diff.
Cmux No Test Or Debug Seam In Production Source ✅ Passed Only cmuxTests files changed; no production Sources/ code was modified to add a seam or debug hook.
Cmux No Ambient Global State ✅ Passed All new file-scope additions are private test helpers or test wiring; no new ambient singleton, mutable global, or public free-function API was introduced.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ejc3
ejc3 force-pushed the fix/git-index-test-fixtures branch from 172043a to 00ec71c Compare July 23, 2026 20:03
@ejc3
ejc3 force-pushed the fix/git-index-test-fixtures branch from 00ec71c to 198d18e Compare July 23, 2026 20:58
@ejc3 ejc3 changed the title Fix three git-index test fixtures and one async wait in the sidebar git suites Fix four unsatisfiable tests in the sidebar git suites Jul 23, 2026
@ejc3
ejc3 marked this pull request as ready for review July 23, 2026 21:51
@greptile-apps

greptile-apps Bot commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR fixes four previously unsatisfiable tests in the sidebar git suites. The changes are entirely test-scoped — no production Sources/ files are touched.

  • Async deadlock fix (TabManagerUnitTests.swift): waitForConditionSuspending replaces the blocking waitForCondition for the one async/@MainActor test, correctly using ContinuousClock.sleep to yield the main actor between polls instead of spinning a nested run loop that libdispatch won't re-enter.
  • Git index content-signature tests (WorkspacePullRequestSidebarTests.swift): Both v4 and empty-index tests now stage a real entry change (different object-ID bytes) so the content signature actually moves; the previous approach changed only the unreachable trailing checksum, which the product deliberately ignores and rebaselines.
  • Socket-routing test (WorkspacePullRequestSidebarTests.swift): The report_git_branch test now registers a windowless AppDelegate context matching how other socket-routing tests work, making the workspace resolvable by the scoped handler.

Confidence Score: 5/5

All changes are confined to test files with no production source modifications; the fixes correctly address the root causes described in the PR.

Every change is in cmuxTests/ with zero production-code modifications. The async helper is well-reasoned and documented. The git-index test fixes are mechanically correct — staging a real object-ID change is exactly what moves the content signature. The socket-routing wiring matches the existing pattern for other scoped-command tests. No new flakiness vectors are introduced.

Files Needing Attention: No files require special attention.

Important Files Changed

Filename Overview
cmuxTests/TabManagerUnitTests.swift Adds waitForConditionSuspending async helper and updates one test to use it, correctly replacing the blocking waitForCondition that deadlocked inside a main-actor async job.
cmuxTests/WorkspacePullRequestSidebarTests.swift Fixes three test issues: (1) adds objectIDBytes param to writeGitIndexVersion4 so tests can stage a real content change; (2) rewrites the empty-index test to add an actual entry rather than just changing the trailing checksum; (3) registers a windowless AppDelegate context so scoped report_git_branch can resolve the workspace.

Sequence Diagram

sequenceDiagram
    participant Test as async @MainActor test
    participant Helper as waitForConditionSuspending
    participant Clock as ContinuousClock
    participant Product as MainActor product code

    Note over Test,Product: New suspending path (this PR)
    Test->>Helper: "await waitForConditionSuspending { condition }"
    Helper->>Helper: "condition() == false"
    Helper->>Clock: try await clock.sleep(pollInterval)
    Clock-->>Helper: suspends yields main actor
    Note over Product: MainActor.run apply lands here
    Product-->>Helper: state updated
    Helper->>Helper: "condition() == true"
    Helper-->>Test: return true

    Note over Test,Product: Old blocking path (before this PR)
    Test->>Helper: "waitForCondition { condition } blocking"
    Helper->>Helper: spin nested run loop via XCTWaiter
    Note over Helper: libdispatch won't re-enter main-queue drain
    Note over Product: MainActor.run apply blocked never runs
    Helper->>Helper: deadline exceeded
    Helper-->>Test: XCTFail always times out
Loading

Reviews (5): Last reviewed commit: "test: use a monotonic clock in the suspe..." | Re-trigger Greptile

Comment thread cmuxTests/TabManagerUnitTests.swift Outdated
@ejc3
ejc3 force-pushed the fix/git-index-test-fixtures branch from 198d18e to 6e4f808 Compare July 23, 2026 22:19

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@cmuxTests/WorkspacePullRequestSidebarTests.swift`:
- Around line 427-430: Update the staged-index fixture generation around the
object ID and index-writing helpers to produce Git-valid indexes: derive blob
object IDs from the actual tracked fixture contents and append the correct
index-body checksum instead of zero-filled trailer bytes. Apply the same
correction to the related fixture paths, using local git add with forced index
v4 where appropriate or equivalent checksum generation.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: fee36609-7a43-419e-804b-6d3cbd975eae

📥 Commits

Reviewing files that changed from the base of the PR and between 198d18e and 6e4f808.

📒 Files selected for processing (2)
  • cmuxTests/TabManagerUnitTests.swift
  • cmuxTests/WorkspacePullRequestSidebarTests.swift

Comment thread cmuxTests/WorkspacePullRequestSidebarTests.swift
@ejc3
ejc3 force-pushed the fix/git-index-test-fixtures branch from 6e4f808 to 7710778 Compare July 23, 2026 22:32

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
cmuxTests/WorkspacePullRequestSidebarTests.swift (1)

427-430: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

The Git index fixtures remain invalid.

The new object-ID changes do not make these fixtures represent real staging: the IDs are arbitrary, and the trailer is still not the SHA-1 of the preceding index body.

  • cmuxTests/WorkspacePullRequestSidebarTests.swift#L427-L430: generate content-derived blob IDs and the correct index checksum.
  • cmuxTests/WorkspacePullRequestSidebarTests.swift#L1167-L1177: use a real staged-content fixture rather than an unchanged file with an arbitrary object ID.
  • cmuxTests/WorkspacePullRequestSidebarTests.swift#L1497-L1521: apply the same valid-index generation to the empty-index transition.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@cmuxTests/WorkspacePullRequestSidebarTests.swift` around lines 427 - 430, The
Git index fixtures in cmuxTests/WorkspacePullRequestSidebarTests.swift must use
valid content-derived blob IDs and SHA-1 trailers: update lines 427-430 to hash
each staged blob’s actual content and compute the checksum over the complete
preceding index body; update lines 1167-1177 to stage genuinely changed content
with its matching blob ID; and apply the same valid index generation and
checksum logic to the empty-index transition at lines 1497-1521.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@cmuxTests/TabManagerUnitTests.swift`:
- Around line 87-97: Update the test polling helper containing the deadline loop
to use a monotonic ContinuousClock (or injected test clock) for timeout
measurement instead of Date(). Replace fixed nanosecond Task.sleep polling with
the clock’s monotonic sleep or cooperative yielding, while preserving the
existing condition check and XCTFail timeout behavior.

---

Duplicate comments:
In `@cmuxTests/WorkspacePullRequestSidebarTests.swift`:
- Around line 427-430: The Git index fixtures in
cmuxTests/WorkspacePullRequestSidebarTests.swift must use valid content-derived
blob IDs and SHA-1 trailers: update lines 427-430 to hash each staged blob’s
actual content and compute the checksum over the complete preceding index body;
update lines 1167-1177 to stage genuinely changed content with its matching blob
ID; and apply the same valid index generation and checksum logic to the
empty-index transition at lines 1497-1521.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c10cfdd8-91d8-4cd9-9875-51925ea53ea9

📥 Commits

Reviewing files that changed from the base of the PR and between 6e4f808 and 7710778.

📒 Files selected for processing (2)
  • cmuxTests/TabManagerUnitTests.swift
  • cmuxTests/WorkspacePullRequestSidebarTests.swift

Comment thread cmuxTests/TabManagerUnitTests.swift Outdated
Two of them describe a git index that git cannot produce. The index trailer is the
SHA-1 of the index content, so rewriting only the trailing checksum while leaving
the entry table byte-identical is not a state a real repository reaches. The
product reads that shape deliberately: the index content signature covers the
entry count, path, mode and object id but not the trailer, so an index whose
content signature is unchanged is rebaselined as clean, which is exactly what
testCleanIndexSignatureRebaselinesWhenIndexRewriteKeepsTrackedContentClean pins.
The v4 and empty-index tests asserted the opposite for the same input, so one of
the two had to fail. They now stage a real change -- a new object id, and an added
entry -- whose stat still matches the worktree, so the dirty verdict comes from the
content signature the way it does for a real staged change.

The predicates are unchanged; the empty-index test's scenario and message move
from a staged delete to a staged add, because staging the first entry out of an
empty index is what actually moves the content signature.

writeGitIndexVersion4 gains the objectIDBytes parameter that
writeGitIndexVersion2EntryFromStat already had. It defaults to the zero id, so the
test call sites that do not stage a change need no edit; the convenience overload
threads it through. writeGitIndexVersion3SkipWorktreeEntry still
hard-codes a zero object id; it has no need to express a staged change.

testDisablingGitWatchClearsCachedPullRequestBadgesWhenPullRequestsAreShownByDefault
sent a scoped report_git_branch to a TabManager that only TerminalController knew
about. That path resolves its workspace through AppDelegate's main-window
contexts, so the report was dropped and the seeded badge survived. The test now
registers a windowless context like the other socket-routing tests, and asserts
the workspace is resolvable before sending, so a future wiring break reports there
instead of at the far assertion.

testSameDirectoryInitialGitMetadataProbesShareOneSnapshotRead waited with a helper
that blocks the main thread while pumping the main queue. That works in a
synchronous test, but this body is async: it runs as a main-actor job, inside a
main-queue drain that libdispatch will not re-enter, so the nested run loop ran
neither the helper's own poll hops nor the snapshot's MainActor.run apply. The
wait could only expire. It now awaits a suspending sibling helper with the same
timeout and interval.

The suspending helper propagates cancellation rather than swallowing it, so a
cancelled test unwinds instead of spinning the condition until its deadline.
@ejc3
ejc3 force-pushed the fix/git-index-test-fixtures branch from 7710778 to e1de81a Compare July 25, 2026 04:56
@ejc3

ejc3 commented Jul 25, 2026

Copy link
Copy Markdown
Contributor Author

Two of the three findings are already addressed on the current head; one I am deliberately not taking now.

  • The try? await Task.sleep cancellation swallow: the current helper uses do/catch and returns condition() on cancellation instead of spinning to the deadline, with a comment explaining exactly the failure mode this finding describes. The review saw the pre-rebase head.
  • Generate valid staged Git indexes: declined for this PR, with the reason on record. The synthetic index bytes are the entire fixture strategy of these tests, and replacing them with real git add-produced indexes is a redesign of that strategy, not a rider on a four-test repair — it changes what the tests exercise (hand-built corruption cases become impossible to express). If the suite should move to real-git fixtures, that wants its own change with its own red-first proof. It is a fair critique of the fixtures' realism and I have noted it as follow-up work rather than pretending it fits here.

@austinywang
austinywang merged commit ee23ff9 into manaflow-ai:main Aug 4, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants