feat(W18-A15): full regression runner — consolidated proof - #245
Conversation
|
Warning Rate limit exceeded
You’ve run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: ⛔ Files ignored due to path filters (6)
📒 Files selected for processing (6)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces a full regression runner system, including a watchdog script to monitor PR merges and a report builder to consolidate Playwright test results. It also adds a handoff document and evidence logs from a recent regression run that resulted in a failure. Feedback focused on improving the portability and efficiency of the new scripts, specifically by avoiding hardcoded Unix-style paths and replacing shell-based logging with native Node.js file system operations.
| const E2E_RESULTS = path.join(UI_ROOT, 'test-results', 'e2e', 'results.json'); | ||
| const BREADTH_RESULTS = path.join(UI_ROOT, 'test-results', 'e2e-breadth', 'results.json'); | ||
| const VISUAL_SUMMARY = path.join(REPO_ROOT, '03_implementation', 'docs', 'evidence', 'visual_proof_2026-05-09', 'summary.json'); | ||
| const VISUAL_PLAYWRIGHT_LOG = process.env.VISUAL_PLAYWRIGHT_LOG || '/tmp/visual-run.log'; |
There was a problem hiding this comment.
The default path for VISUAL_PLAYWRIGHT_LOG uses a hardcoded /tmp directory, which is not portable to Windows environments. Since SUMMARY_DIR is already defined and ensured to exist, it is a better location for the default log file.
| const VISUAL_PLAYWRIGHT_LOG = process.env.VISUAL_PLAYWRIGHT_LOG || '/tmp/visual-run.log'; | |
| const VISUAL_PLAYWRIGHT_LOG = process.env.VISUAL_PLAYWRIGHT_LOG || path.join(SUMMARY_DIR, 'visual-run.log'); |
| function logLine(msg) { | ||
| const line = `[${new Date().toISOString()}] ${msg}`; | ||
| console.log(line); | ||
| try { execSync(`echo "${line.replace(/"/g, '\\"')}" >> "${logPath}"`); } catch {} |
There was a problem hiding this comment.
Using execSync with echo to append to a log file is inefficient and prone to shell-escaping issues across different platforms (e.g., how double quotes or special characters are handled in bash vs. cmd.exe). Using fs.writeFileSync with the append flag is more efficient, portable, and safer.
| try { execSync(`echo "${line.replace(/"/g, '\\"')}" >> "${logPath}"`); } catch {} | |
| try { writeFileSync(logPath, line + '\n', { flag: 'a' }); } catch {} |
…evelop W18-A9 pollution Round 3 merged zero PRs. Root cause: PR #239 (W18-A9 slicer-real-artifact, merged in round 2) introduced a Layer D2 spec failure at 03_implementation/ui/tests/e2e/w18-a9-slicer-real-artifact.spec.ts:186:3 (line 264 `CadQuery` text visibility 5000ms timeout). All open W18 PRs rebased on develop inherit this failure. - #232 W18-A4 — own spec PASSES; only inherited W18-A9 fail - #238 W18-A8 — own spec PASSES; only inherited W18-A9 fail - #241 W18-A10p — own spec PASSES; only inherited W18-A9 fail - #242 W18-A1p — own spec FAILS + inherits W18-A9 fail - #243 W18-A12 — ruff format applied; ruff check still fails on test_slicer_route.py - #244 W18-A13 — ruff-format fix-subagent has not pushed yet - #245 — SKIP per brief (known-fail regression runner) All 4 spec-only PRs (#232, #238, #241, #242) verified scope-safe: - Zero new /api/printers/{id}/heat-*, /start-print, /upload-gcode endpoints. - Zero new Moonraker / Octoprint dispatch. - Zero flips of pinned GUI_PHYSICAL_PRINT_GREEN / GUI_PRINTER_DRY_RUN_GREEN. Stop-criterion (>=3 PRs stuck in unresolvable conflict) met with 6 stuck. Re-dispatch needed: fix-PR against develop repairing W18-A9 spec, then the 4 scope-safe PRs auto-pass. Hermes evidence chain: PASS (ev_4ea83b1da14191a8) Task ID: W18-CASCADE-MERGER-2026-05-11 (round 3) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Watchdog + report-builder + handoff doc for the W18-A15 cross-lane regression sweep. Watchdog polls W18 PRs on a configurable interval, triggers when 8+ target lanes merge or 4 h elapses (whichever first, hard cap 6 h), and writes a per-iter state file. Report builder ingests Playwright JSON reporters (e2e, breadth, lane-specific configs) plus the W15-A9 visual oracle summary + Playwright stdout, emits Markdown + JSON. Hermes evidence chain: PASS Task ID: W18-A15-FULL-REGRESSION-RUNNER-2026-05-11 hermes_run_gate: git-status (PASS, exit 0) Per-config verdicts when this PR was opened: - playwright.e2e.config.ts: 88 P / 3 F / 0 S FAIL - playwright.visual.config.ts: 25 P / 16 F / 20 S FAIL - playwright.breadth.config.ts: 3 P / 0 F / 0 S PASS_REAL - playwright.w18-a9.config.ts: 1 P / 0 F / 0 S PASS_REAL - playwright.w18-a11.config.ts: 1 P / 0 F / 0 S PASS_REAL No printer hardware writes. Pinned verdicts unchanged. - GUI_PHYSICAL_PRINT_GREEN = OUT_OF_SCOPE_BY_OPERATOR - GUI_PRINTER_DRY_RUN_GREEN = OUT_OF_SCOPE_BY_OPERATOR Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
7ff7b7b to
601548d
Compare
Summary
scripts/w18-a15-watchdog.mjs) that polls W18 PRs every N min and triggers at 8+ merged target lanes or 4 h, plus a consolidated report builder (scripts/w18-a15-build-regression-report.mjs) that ingests every Playwright reporter and the W15-A9 visual oracle summary.regression-summary.jsonmachine snapshot that survives the gitignoredtest-results/tree.:8765FastAPI /:5173Vite stack the operator runs and reports honest pass/fail/skip per the W18-A14 no-skip-harness contract.Test plan
playwright.e2e.config.ts— 88 passed / 3 failed / 0 skipped (3.6 min). Failures: printer-panel Moonraker badge missing live peer; Source-OS Setup-Queue 30 s timeout; W18-A9 lane spec 30 s timeout under e2e config (needstest.setTimeout(10*60000)or lane-dedicated config).playwright.visual.config.ts— 25 passed / 16 failed / 20 skipped (4.9 min). Visual oracle: 31 targets, 11 match, 0 diff, 0 error, 20 future placeholders. The 20 skipped are W15-A9status: \"future\"test.skip()sites; the W18-A14 no-skip harness PR has not yet merged so the no-skip contract correctly flags them.playwright.breadth.config.ts— 3 passed / 0 failed / 0 skipped (11 s) — PASS_REAL.playwright.w18-a9.config.ts— 1 passed (33 s) — PASS_REAL when given the lane-dedicated 10-min timeout.playwright.w18-a11.config.ts— 1 passed (22 s) — PASS_REAL.03_implementation/ui/tests/w18-a15-evidence/.ev_4fbb039c8717ba17).hermes_run_gateinvoked (git-statusPASS, exit 0) on this worktree post-run.Per-config table
playwright.e2e.config.tsplaywright.visual.config.tsplaywright.breadth.config.tsplaywright.w18-a9.config.tsplaywright.w18-a11.config.tsSkip-aware verdicts: any skipped test counts as FAIL per the W18-A14 no-skip-harness contract.
Evidence chain
git-statusPASS (exit 0); per-config Playwright stats stored in evidence ledger entryev_4fbb039c8717ba17.w18-a15onscripts/w18-a15-watchdog.mjs,scripts/w18-a15-build-regression-report.mjs,docs/handoffs/W18-A15_FULL_REGRESSION_2026-05-11.md. Released on PR open (see post-merge release step below).STRICT operator freeze
No printer hardware touched. Pinned verdicts unchanged.
GUI_PHYSICAL_PRINT_GREEN= OUT_OF_SCOPE_BY_OPERATORGUI_PRINTER_DRY_RUN_GREEN= OUT_OF_SCOPE_BY_OPERATOR🤖 Generated with Claude Code