Skip to content

fix(chat): align failed task progress test - #9172

Merged
maria-rcks merged 1 commit into
pingdotgg:mainfrom
maria-rcks:t3code/fix-failing-ci-checks
Sep 2, 2026
Merged

maria-rcks merged 1 commit into
pingdotgg:mainfrom
maria-rcks:t3code/fix-failing-ci-checks

fix(chat): align failed task progress test

e98056d
Select commit
Loading
Failed to load commit list.
MacroscopeApp / Macroscope - UI Consistency succeeded Sep 2, 2026 in 1m 31s

UI Consistency: All clear

All clear

Details

Note

Your check run agent prompt is: .macroscope/check-run-agents/ui-consistency.md
More information about how Check Run Agents work can be found in our Help Center.

Test-only change in apps/web/src/components/chat/MessagesTimeline.logic.test.ts. Verified against MessagesTimeline.logic.ts at the merge base (fix(chat): replace failed tools with thinking, #9165): a trailing failed task.progress entry sets latestToolFailed, which suppresses the work-live row and appends the thinking row with id live-activity-row, matching the new assertion. The completed, declined, and stopped cases still produce an inactive work-live row, so the remaining parameters are unchanged and correct. No component, Tailwind, theme, scroll, or environment-routing behavior is touched.


🔍 Agent activity — 11 tool calls

💭 This diff only touches a test file, so I want to verify whether that test change genuinely reflects real behavior—that a "failed" status correctly produces no work-live row—rather than just being adjusted to mask a regression. Since the PR claims to fix CI, I should check the actual code behind this.

1. browse_code — apps/web/src/components/chat/MessagesTimeline.logic.test.ts

2. browse_code — apps/web/src/components/chat/MessagesTimeline*

💭 I want to check MessagesTimeline.logic.ts to see how it handles failed task progress and whether that genuinely produces no work-live row. I also want to check git history to see if this behavior changed recently, since the PR title suggests it's fixing CI checks—so I need to figure out if the test was adjusted to match a regression or the correct current behavior.

3. git_tools — task.progress|toolLifecycleStatus|work-live|live-activity-row

💭 I'm wondering whether suppressing the work-live row in favor of a thinking row when the latest tool fails is intentional behavior from main, or if this test change is masking a real regression rather than just adapting to it. I want to trace the logic around lines 640-780 to see how the failure path is actually handled.

4. browse_code — apps/web/src/components/chat/MessagesTimeline.logic.ts

5. git_tools — apps/web/src/components/chat/MessagesTimeline.logic.ts

💭 I'm noticing there's a recent commit by the same author that reverted "thinking between actions" behavior, so I want to check whether that revert is an ancestor of this branch and examine its diff to see if it explains the failing test — whether the test update reflects a real regression or just matches the reverted prior behavior.

6. git_tools — f036b18c5d578c16d69ae2be0a4becccd843a355

7. git_tools

💭 The git log discrepancy is a bit puzzling, but I'll move past it. I want to check the diff of that specific commit to see exactly what the revert changed in the logic and test files, and whether it left the tests inconsistent.

8. git_tools — f036b18c5d578c16d69ae2be0a4becccd843a355

💭 That's odd — the revert commit doesn't match what's actually in the head file, since the head still has the active field and thinking rows the revert supposedly removed. This suggests the revert itself must have been reverted later, or I'm misreading which commit is actually most recent in the history. Since the PR diff only touches the test file, logic.ts must be identical at merge base and head — so I should trace where the thinking row gets appended at the end of the progress list to understand the failed-task behavior.

9. git_tools — kind: "thinking"

10. git_tools — apps/web/src/components/chat/MessagesTimeline.logic.test.ts

💭 I trace this back to the merge base PR that added suppressing the work-live row in favor of a thinking row when the latest tool fails — this test just wasn't updated to match, so I'm checking whether the new expectations for the failed trailing task.progress case are actually consistent with that intended behavior.

11. git_tools — function workEntryDisplayIndicatesToolFailure|function workEntryIndicatesToolSuccess|function workEntryIndicatesToolNeutralStatus

💭 This matches issue #9165's intent of replacing failed tools with thinking rows. The test uses null in the tuple to represent "no work-live row" which is a bit awkward but works fine with the type as boolean | null. I'm checking whether the failed case's assertion properly verifies the last row is a thinking row. This is a test-only change, so it shouldn't trigger concerns from rules about UI primitives, CSS, themes, or scroll behavior even though the file is in scope for the check. I'm also verifying the "declined" and "stopped" test cases still hold correctly under the current logic - declined excludes from failure detection so the work-live row stays inactive, and for stopped I'm tracing through whether toolLifecycleStatus and tone checks correctly keep latestToolFailed false, meaning the work-live row's active state depends on whether the tool succeeded. I'm working through the "completed" and "failed" cases too, checking whether tone and lifecycle status combinations produce the expected active/inactive states, and confirming that for "failed" the entry gets marked as visible but excluded from the active-turn check, so no work-live row appears and instead a thinking row gets pushed with the live activity indicator. Since there are no actual findings here, I'll report success with a summary indicating everything checks out clean, satisfying the schema's requirement to default to failure only when issues exist.


Agent Credits: 58 credits

Token usage: 16 uncached input · 229,924 cache-read input · 49,976 cache-write input · 4,688 output
Agent Credits may also include non-token charges from external tools such as web research.