Skip to content

[codex] Add Reborn WebUI legacy E2E harness - #5345

Merged
ilblackdragon merged 2 commits into
mainfrom
codex/reborn-e2e-harness-base
Jun 26, 2026
Merged

ilblackdragon merged 2 commits into
mainfrom
codex/reborn-e2e-harness-base

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

Summary

  • Adds the shared Reborn WebUI E2E harness used by the split legacy browser coverage ports.
  • Extends the mock LLM/test helpers needed by the Reborn WebUI scenarios.

Validation

  • tests/e2e/.venv/bin/python -m py_compile tests/e2e/helpers.py tests/e2e/mock_llm.py tests/e2e/reborn_webui_harness.py

Stack base for the follow-up Reborn runtime, OpenAI-compatible Responses API, and WebUI v2 browser coverage PRs.

@coderabbitai

coderabbitai Bot commented Jun 26, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 6f631169-76eb-494e-b259-2aa46f5147f9

📥 Commits

Reviewing files that changed from the base of the PR and between 667e3cc and 20babd7.

📒 Files selected for processing (3)
  • tests/e2e/helpers.py
  • tests/e2e/mock_llm.py
  • tests/e2e/reborn_webui_harness.py

📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes
    • Improved end-to-end coverage for the Reborn WebUI v2 experience, including message loading, sign-in, channel connection, Slack pairing, tool activity, and projects/workspace browsing.
    • Updated test flows to better handle multi-tool responses and recovery paths, reducing flaky behavior in complex chat interactions.
    • Added broader support for v2 server startup and page setup across different app modes, helping ensure more reliable UI testing.

Walkthrough

Adds Reborn WebUI v2 test support across DOM selectors, deterministic mock LLM tool behavior, and a shared Playwright/httpx harness for starting the server, opening the SPA, and driving chat thread polling.

Changes

Reborn WebUI v2 E2E support

Layer / File(s) Summary
v2 selectors
tests/e2e/helpers.py
Expands SEL_V2 with message list, auth gate, channel connect, Slack pairing, activity, and projects/workspace/filesystem selectors.
mock tool patterns
tests/e2e/mock_llm.py
Adds canned responses and Reborn tool-call patterns for approval files, parallel tools, and file-writing flows.
tool-name-aware recovery
tests/e2e/mock_llm.py
Strips attachment blocks before skill detection, reads advertised tool names, and selects Reborn or legacy tool names in special-response recovery.
server startup and fixtures
tests/e2e/reborn_webui_harness.py
Adds Reborn v2 server startup/shutdown helpers, auto-approve setup, and pytest fixtures for default, YOLO, restartable, loop-limited, and vision configurations.
browser and HTTP helpers
tests/e2e/reborn_webui_harness.py
Adds Playwright browser/page fixtures and HTTPX helpers for Reborn v2 threads, timelines, and finalized assistant polling.

Sequence Diagram(s)

sequenceDiagram
  participant reborn_v2_server_fixture
  participant start_reborn_webui_v2_server
  participant ironclaw_reborn_serve
  participant api_health

  reborn_v2_server_fixture->>start_reborn_webui_v2_server: create home dir and config.toml
  start_reborn_webui_v2_server->>ironclaw_reborn_serve: spawn serve subprocess on a free port
  start_reborn_webui_v2_server->>api_health: poll GET /api/health
  api_health-->>start_reborn_webui_v2_server: ready
  start_reborn_webui_v2_server-->>reborn_v2_server_fixture: base URL
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • nearai/ironclaw#5235 — touches the same tests/e2e/helpers.py selector map used by the Reborn v2 UI tests.

Poem

A v2 selector map unfurls,
Mock tools tap dance through little worlds.
Servers rise, then settle into light,
Threads reply and timelines hold tight.
The harness hums, and tests take flight.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5345 June 26, 2026 16:17 Destroyed
@github-actions github-actions Bot added size: XS < 10 changed lines (excluding docs) risk: low Changes to docs, tests, or low-risk modules labels Jun 26, 2026
@github-actions github-actions Bot added the contributor: core 20+ merged PRs label Jun 26, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for the Reborn WebUI v2 in the end-to-end test suite. It adds new selectors to helpers.py, updates the mock LLM to handle Reborn-specific namespaced builtin tools and strip attachment blocks for skill detection, and introduces a new Playwright test harness in reborn_webui_harness.py to manage the Reborn server process and helper functions. The review feedback suggests awaiting proc.wait() on process lookup errors to prevent zombie processes, and wrapping timeline polling calls in try-except blocks to make the test harness resilient against transient HTTP errors.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +57 to +60
try:
proc.send_signal(sig)
except ProcessLookupError:
return

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If proc.send_signal(sig) raises ProcessLookupError, the process has already exited but has not been reaped yet. Returning immediately without awaiting proc.wait() leaves the process as a zombie and prevents proc.returncode from being populated. Awaiting proc.wait() ensures proper cleanup.

Suggested change
try:
proc.send_signal(sig)
except ProcessLookupError:
return
try:
proc.send_signal(sig)
except ProcessLookupError:
await proc.wait()
return

Comment on lines +414 to +425
for _ in range(int(timeout * 2)):
last_timeline = await fetch_timeline(client, base_url, thread_id)
finalized = [
message
for message in last_timeline.get("messages", [])
if message.get("kind") == "assistant"
and message.get("status") == "finalized"
and (message.get("content") or "").strip()
]
if finalized:
return finalized[-1]
await asyncio.sleep(0.5)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If fetch_timeline raises a transient httpx.HTTPError (e.g., due to temporary server unreadiness during a restart or network blip), the polling loop will crash immediately. Wrapping the call in a try...except block and retrying makes the helper much more resilient against flakiness.

    for _ in range(int(timeout * 2)):
        try:
            last_timeline = await fetch_timeline(client, base_url, thread_id)
        except httpx.HTTPError:
            await asyncio.sleep(0.5)
            continue
        finalized = [
            message
            for message in last_timeline.get("messages", [])
            if message.get("kind") == "assistant"
            and message.get("status") == "finalized"
            and (message.get("content") or "").strip()
        ]
        if finalized:
            return finalized[-1]
        await asyncio.sleep(0.5)

Comment on lines +452 to +456
for _ in range(90):
timeline = await fetch_timeline(client, base_url, thread_id)
if finalized_assistant_count(timeline) >= expected:
return
await asyncio.sleep(0.5)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Transient network or HTTP errors during fetch_timeline can cause the polling loop to fail prematurely. Wrapping the call in a try...except block ensures the loop continues retrying until the timeout is reached.

Suggested change
for _ in range(90):
timeline = await fetch_timeline(client, base_url, thread_id)
if finalized_assistant_count(timeline) >= expected:
return
await asyncio.sleep(0.5)
for _ in range(90):
try:
timeline = await fetch_timeline(client, base_url, thread_id)
except httpx.HTTPError:
await asyncio.sleep(0.5)
continue
if finalized_assistant_count(timeline) >= expected:
return
await asyncio.sleep(0.5)

@railway-app

railway-app Bot commented Jun 26, 2026

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5345 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jun 26, 2026 at 4:23 pm

@ilblackdragon
ilblackdragon marked this pull request as ready for review June 26, 2026 17:36
Copilot AI review requested due to automatic review settings June 26, 2026 17:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

…ess-base

# Conflicts:
#	tests/e2e/mock_llm.py

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5345 — 20babd72 Deployed Jun 26, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XS < 10 changed lines (excluding docs)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants