Skip to content

fix(e2e): stabilize coverage suite failures - #3291

Merged
serrrfirat merged 1 commit into
nearai:mainfrom
serrrfirat:fix/e2e-plan-mode-ci
May 6, 2026
Merged

serrrfirat merged 1 commit into
nearai:mainfrom
serrrfirat:fix/e2e-plan-mode-ci

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

  • Thread-scope chat JobContext so plan_update emits current-thread SSE events and plan-mode checklists render in the web UI.
  • Stabilize the users settings-search E2E by using unique seeded users and waiting for API visibility before asserting UI rows.
  • Make the OAuth SSE auth-event scenario use the isolated auth fixture with deterministic credential host mapping and poll for SSE auth events.
  • Relax the read-only active-thread refresh E2E to assert the user-visible invariant: reload must not land on the patched read-only external thread and chat input remains enabled.

Verification

  • CARGO_TARGET_DIR=/tmp/ironclaw-fix-target CARGO_INCREMENTAL=0 cargo test --lib --no-default-features --features libsql agent::dispatcher::tests::test_chat_job_context_includes_thread_id_for_sse_scoped_tools -- --exact
  • CARGO_TARGET_DIR=/tmp/ironclaw-fix-target CARGO_INCREMENTAL=0 pytest scenarios/test_plan_mode.py scenarios/test_settings_search.py::test_search_filters_user_rows scenarios/test_skill_oauth_flow.py::TestSSEAuthEvents::test_auth_required_sse_event scenarios/test_sse_reconnect.py::test_refresh_skips_readonly_external_active_thread -v --timeout=120

@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) size: S 10-49 changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels May 6, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors JobContext creation in the agent dispatcher and enhances the reliability of end-to-end tests. It introduces a chat_job_context helper and updates several E2E tests with polling mechanisms and unique data generation to mitigate flakiness in CI environments. The review feedback suggests increasing the event buffer size in the SSE collection loop and expanding the set of monitored event types to ensure the tests remain stable and performant under various execution paths.

Comment on lines +395 to +396
if len(events_received) > 20:
break

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The hard limit of 20 events in collect_sse_events is quite low and may lead to flakiness in coverage-instrumented CI environments. Under heavy tool use or verbose logging, the engine can emit many thinking status updates. If the limit is reached before the authentication or approval event is received, the polling loop in the main task will time out and fail the test. Consider increasing this limit to provide more headroom for stabilization.

Suggested change
if len(events_received) > 20:
break
if len(events_received) > 100:
break

Comment on lines +420 to +426
while asyncio.get_running_loop().time() < deadline:
if any(
e.get("type") == "onboarding_state" and e.get("state") == "auth_required"
for e in events_received
) or "approval_needed" in [e.get("type", "") for e in events_received]:
break
await asyncio.sleep(0.5)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The polling loop is missing a check for the gate_required event type (specifically for Authentication resume kind), which is the standard auth signal for the engine v2 path. If the engine takes the v2 path, this loop will wait for the full 45-second deadline before proceeding to assertions, unnecessarily slowing down the test suite. Additionally, using a list comprehension inside the loop is less efficient than a generator expression within any().

        deadline = asyncio.get_running_loop().time() + 45
        while asyncio.get_running_loop().time() < deadline:
            if any(
                (e.get("type") == "onboarding_state" and e.get("state") == "auth_required") or
                (e.get("type") == "gate_required" and isinstance(e.get("resume_kind"), dict) and "Authentication" in e["resume_kind"]) or
                (e.get("type") == "approval_needed")
                for e in events_received
            ):
                break
            await asyncio.sleep(0.5)
References
  1. To improve performance, avoid redundant computations inside loops. For example, pre-calculate values like text.splitlines() before iterating.

@serrrfirat
serrrfirat enabled auto-merge May 6, 2026 11:06
@serrrfirat
serrrfirat merged commit bd165d8 into nearai:main May 6, 2026
33 checks passed
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
fix(e2e): stabilize coverage suite failures
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: agent Agent core (agent loop, router, scheduler) size: S 10-49 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant