Skip to content

Fix WebUI v2 live canary final-response waits - #5894

Merged
serrrfirat merged 3 commits into
mainfrom
codex/reborn-webui-v2-harness-drift
Jul 9, 2026
Merged

serrrfirat merged 3 commits into
mainfrom
codex/reborn-webui-v2-harness-drift

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 9, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • expose WebUI v2 assistant final-reply state as a DOM test hook
  • make the live QA harness wait for the final assistant reply before matching required text
  • keep the trigger-record side-effect wait for routine creation cases
  • record assistant reply wait reason/timing in canary artifacts

Validation

@ironloopai

ironloopai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: 7177c3d111d890cee121633ddd0c771f794dc894
Result: One or more review results were superseded by a newer PR head.
Next: Run @ironloopai review on the latest PR head.
Updated: 2026-07-09T17:49:59.298Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Superseded N/A N/A 2026-07-09T17:49:59.024Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Superseded by a newer PR head. New head: 7177c3d. Previous verdict: Approved.
Recent activity
Time Reviewer State Detail
2026-07-09T17:12:40.738Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-09T17:12:40.880Z ironloop/common-reviewer (reviewer) Queued Added to the local review work handoff.
2026-07-09T17:12:41.958Z ironloop/common-reviewer (reviewer) Started Reviewer worker started attempt 1.
2026-07-09T17:12:44.731Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at d5724fa.
2026-07-09T17:15:56.607Z ironloop/common-reviewer (reviewer) Superseded Old-head reviewer is still running after newer head 7177c3d replaced it. Codex is reviewing; process live; elapsed 3m 13s; timeout in 16m 47s; last heartbeat 2026-07-09T17:15:56.607Z. Codex emitted stderr output at 2026-07-09T17:15:54.152Z.
2026-07-09T17:16:13.296Z ironloop/common-reviewer (reviewer) Result captured Approved; 0 blocking findings.
2026-07-09T17:16:13.296Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
2026-07-09T17:49:59.024Z ironloop/common-reviewer (reviewer) Superseded A newer PR head replaced this review (7177c3d).
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
  • @ironloopai status
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: e9ee1e52-746b-42db-9bf8-341704a29ce3

📥 Commits

Reviewing files that changed from the base of the PR and between c61a2fb and 7177c3d.

📒 Files selected for processing (2)
  • scripts/reborn_webui_v2_live_qa/run_live_qa.py
  • scripts/reborn_webui_v2_live_qa/test_run_live_qa.py

📝 Walkthrough

Summary by CodeRabbit

  • New Features
    • Assistant messages now expose a final-reply indicator (data-final-reply) in the chat UI to distinguish streaming vs completed answers.
    • Live QA now captures additional completion timing and reason details, plus trigger-record wait measurements.
  • Bug Fixes
    • Improved assistant completion detection by waiting for the final-reply marker with quiet-period fallback and clearer timeouts.
    • Routine verification now waits for trigger-record count increases and reports clearer failures when they don’t appear after success.
  • Tests
    • Added/updated coverage for final-reply marker behavior, attribute-read error handling, and trigger-record timing assertions.

Walkthrough

Adds data-final-reply on assistant bubbles and teaches live QA to poll that marker plus trigger-record growth before finalizing routine results.

Changes

Frontend final reply marker

Layer / File(s) Summary
data-final-reply attribute and test
crates/ironclaw_webui_v2/frontend/src/pages/chat/components/message-bubble.tsx, crates/ironclaw_webui_v2/frontend/src/pages/chat/components/message-bubble.test.ts
Computes finalReplyState from message.isFinalReply for assistant messages and binds it to data-final-reply; a source-pattern test checks the wiring.

Live QA assistant reply and trigger record waiting

Layer / File(s) Summary
Assistant wait telemetry
scripts/reborn_webui_v2_live_qa/run_live_qa.py
AssistantReplyWaitResult gains final_reply_wait_ms and final_reply_reason; wait and record helpers add the new telemetry fields and polling constants.
Final-reply aware assistant polling
scripts/reborn_webui_v2_live_qa/run_live_qa.py
_wait_for_assistant_reply tracks data-final-reply during polling and uses it in required-text and semantic-judge completion paths, including timeout messaging.
Assistant wait tests
scripts/reborn_webui_v2_live_qa/test_run_live_qa.py
Test helpers add final_reply_state/get_attribute; new cases cover marker-driven completion and attribute-read failures.
Routine trigger-record polling
scripts/reborn_webui_v2_live_qa/run_live_qa.py
Adds _wait_for_trigger_record_after_count and wires _routine_creation_case to wait for a new trigger record after success, recording wait metrics and timeout details.
Trigger-record polling tests
scripts/reborn_webui_v2_live_qa/test_run_live_qa.py
Updates routine-creation mocks and assertions for explicit trigger-record polling and recorded wait duration.

Sequence Diagram(s)

sequenceDiagram
  participant ChatPage
  participant MessageBubble
  participant LiveQARunner
  participant PlaywrightPage

  ChatPage->>MessageBubble: render assistant message
  MessageBubble->>PlaywrightPage: expose data-final-reply
  LiveQARunner->>PlaywrightPage: poll assistant text and data-final-reply
  PlaywrightPage-->>LiveQARunner: final-reply state and text
Loading
sequenceDiagram
  participant LiveQARunner
  participant PlaywrightPage
  participant TriggerStore

  LiveQARunner->>PlaywrightPage: run routine creation case
  LiveQARunner->>TriggerStore: read trigger record count
  TriggerStore-->>LiveQARunner: count before success
  LiveQARunner->>TriggerStore: poll until count increases
  TriggerStore-->>LiveQARunner: updated count and wait time
Loading

Estimated code review effort: 3 (Moderate) | ~25 minutes

Possibly related PRs

Suggested reviewers: think-in-universe

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description covers summary and validation, but omits many required template sections like change type, linked issue, security, and rollback. Add the missing template sections: change type, linked issue, security impact, trust-boundary checklist, database impact, blast radius, rollback plan, and review follow-through.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly matches the PR's main change: fixing WebUI v2 live canary final-response waits.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5894 July 9, 2026 16:38 Destroyed
@github-actions github-actions Bot added size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 9, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces asynchronous polling to wait for a trigger record to be added after a routine is created, replacing a single immediate check. It adds the _wait_for_trigger_record_after_count helper function, integrates it into _routine_creation_case, and updates the corresponding unit tests. Feedback suggests wrapping the database count check in a try-except block to handle potential sqlite3.Error exceptions due to transient database locks during concurrent execution.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +1677 to +1683
last_count = _trigger_record_count(reborn_home, routine_name)
while last_count <= before_count:
now = time.monotonic()
if now >= deadline:
break
await asyncio.sleep(min(poll_interval, deadline - now))
last_count = _trigger_record_count(reborn_home, routine_name)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Since the ironclaw-reborn serve process runs concurrently and writes to the SQLite database, transient database locks (sqlite3.OperationalError: database is locked) can occur. Wrapping the _trigger_record_count call in a helper that catches sqlite3.Error and defaults to before_count (to keep polling) would make the helper much more resilient to these transient errors and prevent flaky test failures.

Suggested change
last_count = _trigger_record_count(reborn_home, routine_name)
while last_count <= before_count:
now = time.monotonic()
if now >= deadline:
break
await asyncio.sleep(min(poll_interval, deadline - now))
last_count = _trigger_record_count(reborn_home, routine_name)
def get_count() -> int:
try:
return _trigger_record_count(reborn_home, routine_name)
except sqlite3.Error:
return before_count
last_count = get_count()
while last_count <= before_count:
now = time.monotonic()
if now >= deadline:
break
await asyncio.sleep(min(poll_interval, deadline - now))
last_count = get_count()
References
  1. To prevent flaky tests, avoid assertions based on wall-clock time. Instead, verify state changes by comparing values (e.g., counts) before and after an action.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
✅ Approved 0 0 0 e4938afd9908

Head: e4938afd9908dbb3ba1178da6a2e03241adaa330
Next: No reviewer action needed.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

No concrete actionable issues found in the routine trigger-record wait change or its focused unit coverage.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.
  4. Use @ironloopai status to check queued/running/completed/failed/superseded state while reviewers run.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5894 July 9, 2026 16:56 Destroyed
@github-actions github-actions Bot added size: L 200-499 changed lines and removed size: M 50-199 changed lines labels Jul 9, 2026
@serrrfirat serrrfirat changed the title Fix WebUI v2 routine canary side-effect wait Fix WebUI v2 live canary final-response waits Jul 9, 2026
@github-actions

github-actions Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.13% (283982 / 333600 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 333600 lines now vs 320188 at floor capture (+13412 lines, +4.19%) — not a material change

⚠️ 3 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts, ironclaw_skill_learning

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.13% — 283982 / 333600 lines

Per-crate breakdown (65 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 347
ironclaw_skill_learning 0% 0 / 61
ironclaw_wasm_sandbox_core 7.37% 7 / 95
ironclaw_runtime_policy 33.2% 80 / 241
ironclaw_event_projections 43.34% 673 / 1553
ironclaw_run_state 52.36% 222 / 424
ironclaw_authorization 53.54% 461 / 861
ironclaw_triggers 59.79% 1740 / 2910
ironclaw_observability 61.54% 16 / 26
ironclaw_webui_v2 62.62% 2632 / 4203
ironclaw_reborn_cli 62.84% 3816 / 6073
ironclaw_mcp 63.15% 581 / 920
ironclaw_reborn_migration 67.01% 1172 / 1749
ironclaw_memory 67.12% 747 / 1113
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_filesystem 67.44% 3815 / 5657
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.08% 1658 / 2238
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.61% 958 / 1284
ironclaw_extractors 74.72% 538 / 720
ironclaw_first_party_extensions 77.62% 5411 / 6971
ironclaw_llm 78.16% 20071 / 25680
ironclaw_product_context 78.57% 11 / 14
ironclaw_wasm_product_adapters 80.58% 1510 / 1874
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_reborn_openai_compat 80.95% 956 / 1181
ironclaw_memory_native 81.86% 3226 / 3941
ironclaw_wasm 82.54% 950 / 1151
ironclaw_secrets 82.7% 2791 / 3375
ironclaw_events 83.47% 1762 / 2111
ironclaw_processes 84.06% 965 / 1148
ironclaw_turns 84.26% 13127 / 15579
ironclaw_host_api 85.17% 2549 / 2993
ironclaw_product_workflow 85.34% 10748 / 12594
ironclaw_projects 85.92% 659 / 767
ironclaw_network 86.12% 670 / 778
ironclaw_common 86.46% 1514 / 1751
ironclaw_threads 86.62% 4132 / 4770
ironclaw_slack_v2_adapter 86.79% 1806 / 2081
ironclaw_auth 86.94% 2995 / 3445
ironclaw_reborn_config 86.98% 1730 / 1989
ironclaw_product_adapters 86.98% 3207 / 3687
ironclaw_reborn_identity 87.03% 557 / 640
ironclaw_skills 87.35% 4336 / 4964
ironclaw_hooks 87.84% 9916 / 11289
ironclaw_product_adapter_registry 87.96% 526 / 598
ironclaw_reborn_traces 88.21% 11931 / 13526
ironclaw_extensions 88.26% 2631 / 2981
ironclaw_reborn_composition 88.76% 70139 / 79019
ironclaw_host_runtime 88.94% 17522 / 19700
ironclaw_reborn 89.15% 15846 / 17774
ironclaw_conversations 90% 2924 / 3249
ironclaw_approvals 90.51% 1507 / 1665
ironclaw_event_streams 91.48% 1009 / 1103
ironclaw_loop_support 92.44% 14731 / 15936
ironclaw_attachments 93.06% 630 / 677
ironclaw_resources 93.09% 4637 / 4981
ironclaw_reborn_webui_ingress 93.19% 2217 / 2379
ironclaw_telegram_v2_adapter 93.87% 2452 / 2612
ironclaw_agent_loop 94.58% 8776 / 9279
ironclaw_safety 94.8% 3668 / 3869
ironclaw_first_party_extension_ports 95% 3094 / 3257
ironclaw_outbound 95.59% 3556 / 3720

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (4 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_oauth v1-only: consumed only by root ironclaw (src/auth/oauth.rs); no crates/* dependents. Crate's own doc comment confirms v1-only. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@serrrfirat
serrrfirat marked this pull request as ready for review July 9, 2026 17:12

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
✅ Approved 0 0 0 c61a2fb70a10

Head: c61a2fb70a10a2668d3265799ed8d2bc3692cda3
Next: No reviewer action needed.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

No concrete blocking issues found in the PR changes. The diff adds a DOM-visible assistant final-reply state and updates the live QA waits/tests consistently.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.
  4. Use @ironloopai status to check queued/running/completed/failed/superseded state while reviewers run.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@scripts/reborn_webui_v2_live_qa/run_live_qa.py`:
- Around line 1451-1457: The attribute read fallback in run_live_qa.py is
clearing a previously known non-final state by setting last_final_reply_state to
None on exception, which later allows the quiet-period and semantic fallback
paths to run incorrectly. Update the try/except around
assistant.get_attribute("data-final-reply", ...) so a transient read failure
preserves the prior explicit value of last_final_reply_state rather than
overwriting it, and ensure the downstream fallback checks continue to respect
that state.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: e096609e-bc09-4684-b80b-03a0af191dff

📥 Commits

Reviewing files that changed from the base of the PR and between 8e05a37 and c61a2fb.

📒 Files selected for processing (4)
  • crates/ironclaw_webui_v2/frontend/src/pages/chat/components/message-bubble.test.ts
  • crates/ironclaw_webui_v2/frontend/src/pages/chat/components/message-bubble.tsx
  • scripts/reborn_webui_v2_live_qa/run_live_qa.py
  • scripts/reborn_webui_v2_live_qa/test_run_live_qa.py

Comment thread scripts/reborn_webui_v2_live_qa/run_live_qa.py Outdated
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5894 July 9, 2026 17:49 Destroyed
@railway-app

railway-app Bot commented Jul 9, 2026

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5894 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 9, 2026 at 5:50 pm

@serrrfirat
serrrfirat merged commit 0b9885e into main Jul 9, 2026
61 checks passed
@serrrfirat
serrrfirat deleted the codex/reborn-webui-v2-harness-drift branch July 9, 2026 18:50
@coderabbitai coderabbitai Bot mentioned this pull request Jul 10, 2026
6 of 9 tasks

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5894 — 7177c3d1 Deployed Jul 9, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant