Skip to content

feat(W18-A15): full regression runner — consolidated proof - #245

Merged
Ghenghis merged 1 commit into
developfrom
claude/w18-a15-regression-runner
May 11, 2026
Merged

Ghenghis merged 1 commit into
developfrom
claude/w18-a15-regression-runner

Conversation

@Ghenghis

Copy link
Copy Markdown
Owner

Summary

  • Adds the W18-A15 regression-runner subagent harness: watchdog (scripts/w18-a15-watchdog.mjs) that polls W18 PRs every N min and triggers at 8+ merged target lanes or 4 h, plus a consolidated report builder (scripts/w18-a15-build-regression-report.mjs) that ingests every Playwright reporter and the W15-A9 visual oracle summary.
  • Includes the artifacts from this sweep: e2e + visual + breadth + w18-a9 + w18-a11 stdout logs, the watchdog state JSON, the archived visual oracle stats, and a regression-summary.json machine snapshot that survives the gitignored test-results/ tree.
  • Drives the per-config + per-lane Playwright suite end-to-end against the live :8765 FastAPI / :5173 Vite stack the operator runs and reports honest pass/fail/skip per the W18-A14 no-skip-harness contract.

Test plan

  • Watchdog observed W18 PRs every 2 min until trigger (3/10 target lanes + 3 auxiliary merged at iter=6).
  • playwright.e2e.config.ts — 88 passed / 3 failed / 0 skipped (3.6 min). Failures: printer-panel Moonraker badge missing live peer; Source-OS Setup-Queue 30 s timeout; W18-A9 lane spec 30 s timeout under e2e config (needs test.setTimeout(10*60000) or lane-dedicated config).
  • playwright.visual.config.ts — 25 passed / 16 failed / 20 skipped (4.9 min). Visual oracle: 31 targets, 11 match, 0 diff, 0 error, 20 future placeholders. The 20 skipped are W15-A9 status: \"future\" test.skip() sites; the W18-A14 no-skip harness PR has not yet merged so the no-skip contract correctly flags them.
  • playwright.breadth.config.ts — 3 passed / 0 failed / 0 skipped (11 s) — PASS_REAL.
  • playwright.w18-a9.config.ts — 1 passed (33 s) — PASS_REAL when given the lane-dedicated 10-min timeout.
  • playwright.w18-a11.config.ts — 1 passed (22 s) — PASS_REAL.
  • Watchdog log + state JSON archived to 03_implementation/ui/tests/w18-a15-evidence/.
  • Hermes evidence chain extended with summary entry (ev_4fbb039c8717ba17).
  • hermes_run_gate invoked (git-status PASS, exit 0) on this worktree post-run.

Per-config table

Config Verdict Passed Failed Skipped Duration
playwright.e2e.config.ts FAIL 88 3 0 218s
playwright.visual.config.ts FAIL 25 16 20 294s
playwright.breadth.config.ts PASS_REAL 3 0 0 11s
playwright.w18-a9.config.ts PASS_REAL 1 0 0 33s
playwright.w18-a11.config.ts PASS_REAL 1 0 0 22s

Skip-aware verdicts: any skipped test counts as FAIL per the W18-A14 no-skip-harness contract.

Evidence chain

  • Hermes evidence chain: PASS
  • Task ID: W18-A15-FULL-REGRESSION-RUNNER-2026-05-11
  • hermes_run_gate: git-status PASS (exit 0); per-config Playwright stats stored in evidence ledger entry ev_4fbb039c8717ba17.
  • Locks: w18-a15 on scripts/w18-a15-watchdog.mjs, scripts/w18-a15-build-regression-report.mjs, docs/handoffs/W18-A15_FULL_REGRESSION_2026-05-11.md. Released on PR open (see post-merge release step below).

STRICT operator freeze

No printer hardware touched. Pinned verdicts unchanged.

  • GUI_PHYSICAL_PRINT_GREEN = OUT_OF_SCOPE_BY_OPERATOR
  • GUI_PRINTER_DRY_RUN_GREEN = OUT_OF_SCOPE_BY_OPERATOR

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented May 11, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@Ghenghis has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 9 minutes and 26 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: e3de733e-a7cf-447f-97e3-cd2103f70d33

📥 Commits

Reviewing files that changed from the base of the PR and between fe9bf03 and 601548d.

⛔ Files ignored due to path filters (6)
  • 03_implementation/ui/tests/w18-a15-evidence/breadth-run.log is excluded by !**/*.log
  • 03_implementation/ui/tests/w18-a15-evidence/e2e-run.log is excluded by !**/*.log
  • 03_implementation/ui/tests/w18-a15-evidence/visual-run.log is excluded by !**/*.log
  • 03_implementation/ui/tests/w18-a15-evidence/w18-a11-run.log is excluded by !**/*.log
  • 03_implementation/ui/tests/w18-a15-evidence/w18-a9-run.log is excluded by !**/*.log
  • 03_implementation/ui/tests/w18-a15-evidence/watchdog.log is excluded by !**/*.log
📒 Files selected for processing (6)
  • 03_implementation/docs/handoffs/W18-A15_FULL_REGRESSION_2026-05-11.md
  • 03_implementation/ui/scripts/w18-a15-build-regression-report.mjs
  • 03_implementation/ui/scripts/w18-a15-watchdog.mjs
  • 03_implementation/ui/tests/w18-a15-evidence/regression-summary.json
  • 03_implementation/ui/tests/w18-a15-evidence/visual-oracle-summary.json
  • 03_implementation/ui/tests/w18-a15-evidence/watchdog-state.json
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch claude/w18-a15-regression-runner

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a full regression runner system, including a watchdog script to monitor PR merges and a report builder to consolidate Playwright test results. It also adds a handoff document and evidence logs from a recent regression run that resulted in a failure. Feedback focused on improving the portability and efficiency of the new scripts, specifically by avoiding hardcoded Unix-style paths and replacing shell-based logging with native Node.js file system operations.

const E2E_RESULTS = path.join(UI_ROOT, 'test-results', 'e2e', 'results.json');
const BREADTH_RESULTS = path.join(UI_ROOT, 'test-results', 'e2e-breadth', 'results.json');
const VISUAL_SUMMARY = path.join(REPO_ROOT, '03_implementation', 'docs', 'evidence', 'visual_proof_2026-05-09', 'summary.json');
const VISUAL_PLAYWRIGHT_LOG = process.env.VISUAL_PLAYWRIGHT_LOG || '/tmp/visual-run.log';

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The default path for VISUAL_PLAYWRIGHT_LOG uses a hardcoded /tmp directory, which is not portable to Windows environments. Since SUMMARY_DIR is already defined and ensured to exist, it is a better location for the default log file.

Suggested change
const VISUAL_PLAYWRIGHT_LOG = process.env.VISUAL_PLAYWRIGHT_LOG || '/tmp/visual-run.log';
const VISUAL_PLAYWRIGHT_LOG = process.env.VISUAL_PLAYWRIGHT_LOG || path.join(SUMMARY_DIR, 'visual-run.log');

function logLine(msg) {
const line = `[${new Date().toISOString()}] ${msg}`;
console.log(line);
try { execSync(`echo "${line.replace(/"/g, '\\"')}" >> "${logPath}"`); } catch {}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using execSync with echo to append to a log file is inefficient and prone to shell-escaping issues across different platforms (e.g., how double quotes or special characters are handled in bash vs. cmd.exe). Using fs.writeFileSync with the append flag is more efficient, portable, and safer.

Suggested change
try { execSync(`echo "${line.replace(/"/g, '\\"')}" >> "${logPath}"`); } catch {}
try { writeFileSync(logPath, line + '\n', { flag: 'a' }); } catch {}

Ghenghis added a commit that referenced this pull request May 11, 2026
…evelop W18-A9 pollution

Round 3 merged zero PRs. Root cause: PR #239 (W18-A9 slicer-real-artifact,
merged in round 2) introduced a Layer D2 spec failure at
03_implementation/ui/tests/e2e/w18-a9-slicer-real-artifact.spec.ts:186:3
(line 264 `CadQuery` text visibility 5000ms timeout). All open W18 PRs
rebased on develop inherit this failure.

- #232 W18-A4 — own spec PASSES; only inherited W18-A9 fail
- #238 W18-A8 — own spec PASSES; only inherited W18-A9 fail
- #241 W18-A10p — own spec PASSES; only inherited W18-A9 fail
- #242 W18-A1p — own spec FAILS + inherits W18-A9 fail
- #243 W18-A12 — ruff format applied; ruff check still fails on test_slicer_route.py
- #244 W18-A13 — ruff-format fix-subagent has not pushed yet
- #245 — SKIP per brief (known-fail regression runner)

All 4 spec-only PRs (#232, #238, #241, #242) verified scope-safe:
- Zero new /api/printers/{id}/heat-*, /start-print, /upload-gcode endpoints.
- Zero new Moonraker / Octoprint dispatch.
- Zero flips of pinned GUI_PHYSICAL_PRINT_GREEN / GUI_PRINTER_DRY_RUN_GREEN.

Stop-criterion (>=3 PRs stuck in unresolvable conflict) met with 6 stuck.
Re-dispatch needed: fix-PR against develop repairing W18-A9 spec, then
the 4 scope-safe PRs auto-pass.

Hermes evidence chain: PASS (ev_4ea83b1da14191a8)
Task ID: W18-CASCADE-MERGER-2026-05-11 (round 3)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Watchdog + report-builder + handoff doc for the W18-A15 cross-lane
regression sweep. Watchdog polls W18 PRs on a configurable interval,
triggers when 8+ target lanes merge or 4 h elapses (whichever first,
hard cap 6 h), and writes a per-iter state file. Report builder
ingests Playwright JSON reporters (e2e, breadth, lane-specific
configs) plus the W15-A9 visual oracle summary + Playwright stdout,
emits Markdown + JSON.

Hermes evidence chain: PASS
Task ID: W18-A15-FULL-REGRESSION-RUNNER-2026-05-11
hermes_run_gate: git-status (PASS, exit 0)

Per-config verdicts when this PR was opened:
- playwright.e2e.config.ts:      88 P / 3 F / 0 S  FAIL
- playwright.visual.config.ts:   25 P / 16 F / 20 S FAIL
- playwright.breadth.config.ts:   3 P / 0 F / 0 S  PASS_REAL
- playwright.w18-a9.config.ts:    1 P / 0 F / 0 S  PASS_REAL
- playwright.w18-a11.config.ts:   1 P / 0 F / 0 S  PASS_REAL

No printer hardware writes. Pinned verdicts unchanged.
- GUI_PHYSICAL_PRINT_GREEN  = OUT_OF_SCOPE_BY_OPERATOR
- GUI_PRINTER_DRY_RUN_GREEN = OUT_OF_SCOPE_BY_OPERATOR

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Ghenghis
Ghenghis force-pushed the claude/w18-a15-regression-runner branch from 7ff7b7b to 601548d Compare May 11, 2026 16:11
@Ghenghis
Ghenghis merged commit 100398e into develop May 11, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant