Skip to content

ci: fail AI triage loudly when the session bails before posting (#338) - #340

Merged
allxsmith merged 2 commits into
mainfrom
claude/ai-auto-triage-gh-issues-cmbm7y
Jul 21, 2026
Merged

allxsmith merged 2 commits into
mainfrom
claude/ai-auto-triage-gh-issues-cmbm7y

Conversation

@allxsmith

@allxsmith allxsmith commented Jul 21, 2026 •

Copy link
Copy Markdown
Owner

Pull Request

Description

Fixes the silent no-op in AI auto-triage: a Claude Code CLI update made Task sub-agents background-by-default, so the triage session fanned out its search agents, said it would "wait for them to finish", and ended its turn — which terminates the headless session with no comment posted. The result record is a clean success (is_error: false, stop_reason: end_turn), so the existing watchdog stayed green (a new variant of the #317 silent-no-op shape; evidence in #338 from runs 29686174473 / issue #330 and 29625394616 / issue #327).

Changes:

  • .github/workflows/ai-triage.yml — the session prompt now requires every Task call to pass run_in_background: false and forbids ending the turn while any sub-agent is pending; it also requires a machine-checkable final message of only TRIAGE-RESULT: <command> commented|refreshed|skipped (<reason>) lines. The "Fail on silent session error" watchdog jq additionally requires that sentinel shape in the final result text, so any mid-task bailout — this one or a future runtime behavior shift — fails the job loudly instead of showing green. Header comment documents the ci: AI triage silently no-ops — headless session ends turn while background Task sub-agents still running #338 failure shape next to the existing [Bug] ai-triage session completes without posting its comment (3 permission denials, 11 turns — fan-out never runs) #317 note.

  • .claude/commands/triage-dedupe.md, triage-find-issues.md, triage-find-duplicate-prs.md — same synchronous fan-out rule, plus search hardening from the same logs: one prefix-allowlisted gh command per Bash call (shell loops/echo prefixes were being permission-denied) and a 6-search cap per agent (bursts were 403ing on the shared API rate limit).

  • bulma-ui (@allxsmith/bestax-bulma)

  • create-bestax (create-bestax)

  • docs (@allxsmith/bestax-docs)

  • Other (please specify): CI workflow (.github/workflows/ai-triage.yml) + triage command files (.claude/commands/)

Related Issue(s)

Fixes #338

Related to #317, #330, #327

Type of Change

  • Bug fix
  • Build tooling

Checklist

Screenshots / Demos

n/a — see the quoted session output in #338.

Additional Context

  • No release: ci: type, no package scope.
  • Verification caveat: issues: events run the workflow from main, and the checkout step pulls main's .claude/commands/, so the full fix only takes effect after merge. Pre-merge, applying the ai-triage label to this PR exercises the PR's copy of the workflow (prompt + watchdog) but still with main's command files. Post-merge test: apply ai-triage to an open issue (label runs are budget-exempt) and confirm a <!-- ai-triage:dedupe --> comment appears — or a loud red run if not.
  • This PR intentionally is not part of the AI loop (it touches .github/**, which the loop refuses).

🤖 Generated with Claude Code

https://claude.ai/code/session_01K1LrTrVRbtdh8paN7f67CB


Generated by Claude Code

Summary by CodeRabbit

  • Improvements
    • Improved automated triage reliability by ensuring searches complete before results are processed.
    • Strengthened search accuracy by limiting results to relevant, open items and excluding ineligible matches.
    • Added stricter validation to detect incomplete or missing triage results.
    • Improved handling of service limits to reduce failed or interrupted triage runs.
    • Standardized triage output for clearer, more consistent results.

A Claude Code CLI update made Task sub-agents background-by-default, so
the triage session fanned out its search agents, said it would wait for
them, and ended its turn — which terminates the headless session with no
comment posted. The result record is a clean success, so the is_error
watchdog stayed green (a new variant of the #317 silent-no-op shape).

- Command files + workflow prompt: Task calls must pass
  run_in_background: false, and the session must never end its turn
  while a sub-agent is pending.
- Workflow prompt now requires a machine-checkable final message
  (`TRIAGE-RESULT: <command> commented|refreshed|skipped (<reason>)`);
  the watchdog jq additionally requires the final result text to be
  nothing but such sentinel lines, so any future mid-task bailout fails
  the job instead of showing green.
- Search agent rules: one prefix-allowlisted gh command per Bash call
  (loops/echo prefixes were being permission-denied) and a 6-search cap
  per agent (bursts were 403ing on the shared API rate limit).

Fixes #338

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1LrTrVRbtdh8paN7f67CB
@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: ae6632b3-a3ac-450d-83f6-70fcaea01b5c

📥 Commits

Reviewing files that changed from the base of the PR and between 091ac47 and d2ffb51.

📒 Files selected for processing (4)
  • .claude/commands/triage-dedupe.md
  • .claude/commands/triage-find-duplicate-prs.md
  • .claude/commands/triage-find-issues.md
  • .github/workflows/ai-triage.yml

Walkthrough

The triage command specifications now require synchronous agents and constrained GitHub CLI searches. The AI triage workflow requires exact TRIAGE-RESULT: output and validates execution records with stricter watchdog logic.

Changes

Triage reliability controls

Layer / File(s) Summary
Synchronous triage search instructions
.claude/commands/triage-*.md
Dedupe, duplicate-PR, and issue triage commands require synchronous agents, direct single-command gh calls, capped searches, scoped results, and structured output.
Workflow execution and output contract
.github/workflows/ai-triage.yml
The workflow documents silent no-op cases and requires Claude to wait for agents and emit the expected number of TRIAGE-RESULT: lines.
Silent session watchdog validation
.github/workflows/ai-triage.yml
Watchdog parsing now validates successful result records, sentinel counts, and output shape, while skipping missing or empty execution files.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Claude
  participant ExecutionFile
  participant Watchdog
  Claude->>ExecutionFile: Write final triage result
  ExecutionFile->>Watchdog: Provide result record
  Watchdog->>Watchdog: Validate error status and sentinel count
Loading

Possibly related issues

  • allxsmith/bestax issue 338: Covers the same synchronous-agent, constrained-search, and triage completion-validation changes.

Possibly related PRs

  • allxsmith/bestax#277: Introduced the workflow Claude prompt and execution flow that this PR further validates.
  • allxsmith/bestax#292: Established the triage command and workflow constraints that these changes tighten.
  • allxsmith/bestax#319: Added related execution-output checks that this PR extends with sentinel validation.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the main CI triage fix and matches the changeset.
Description check ✅ Passed The description covers the required sections, affected package area, issue links, change type, checklist, demos, and context.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch claude/ai-auto-triage-gh-issues-cmbm7y

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

Preview Deployment

Preview URL: https://ca5fa99a.bestax.pages.dev

@allxsmith allxsmith added the deep-review label Jul 21, 2026 — with Claude

Copy link
Copy Markdown
Owner Author

@claude please do a deep adversarial review of this PR's diff. Focus on:

Please report findings as review comments or a summary comment — do not push fixes.


Generated by Claude Code

@claude

claude Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Claude finished @allxsmith's task in 7m 10s —— View job


Deep adversarial review complete

I pulled the actual execution logs for the two cited failures (runs 29686174473 / issue #330, 29625394616 / issue #327) to check the fix against real, not hypothetical, bailout text.

Confirmed: the core mechanism does close the exact #338 shape. Both real bailouts have is_error: false and a plain-string .result:

Neither starts with TRIAGE-RESULT:, so the new all(startswith("TRIAGE-RESULT:")) check in ai-triage.yml:404-411 correctly turns both green runs red. I also confirmed the .result field is genuinely a plain string in the real execution file (not an array/object), so the split("\n") assumption holds, and I traced through the NDJSON-vs-array flattening ([.[] | if type == "array" then .[] else . end]) — it correctly normalizes both shapes, and last on an empty/missing-result-record match correctly evaluates to null → fails closed. The for q in ...; do gh search issues ...; done loop denial visible in run 29686174473's permission_denials also independently confirms the motivation for the new "one gh command per Bash call" rule.

That said, I found gaps that mean a session can still bail (semi-)silently:

1. The watchdog can't detect a dropped second command on PR runs (ai-triage.yml:404-411)
PR triage always runs two commands (triage-find-issues.md then triage-find-duplicate-prs.md, per ai-triage.yml:319-323), but the jq only requires length > 0 non-blank lines, all prefixed TRIAGE-RESULT: — it never checks the count against IS_PR. A session that completes command 1, emits TRIAGE-RESULT: triage-find-issues commented, then hits a #338-style bailout partway through command 2 would produce exactly one valid sentinel line and pass. IS_PR is already a step output (steps.gate.outputs.is_pr) available to wire into this step's env: and assert line-count parity (1 for issues, 2 for PRs). As written, this is a real residual variant of the failure class the PR sets out to close, specific to the two-command PR path.

2. The sentinel is self-attested, not cross-checked (ai-triage.yml:404-411)
The check only validates that each line has the right prefix — it doesn't verify the claimed outcome (commented/refreshed/skipped) actually happened, e.g. by confirming a corresponding gh … comment tool call occurred. A confused (not necessarily malicious) session could emit TRIAGE-RESULT: triage-dedupe skipped (no credible duplicates) without having actually run the search agents, and the watchdog would show green. Likely not fully fixable with jq alone, but worth flagging as a limitation of "format-only" validation.

3. "No shell loops, echo prefixes, or ;/&& chains" omits pipes and substitution (.claude/commands/triage-dedupe.md:47-49, mirrored in the other two command files)
The real denial evidence (run 29686174473) is a for...do...done loop, which the new rule correctly bans. But the same prefix-matching Bash permission engine is also very likely to deny gh search issues foo | head, `gh ...`, or $(gh ...) constructs — none of which are called out. Search agents piping to head/jq for readability would hit the identical wasted-turn denial this PR set out to eliminate.

4. "Never end your turn while pending" has no actionable mechanism (ai-triage.yml:337-345)
The real logs show the Task tool_result explicitly telling the model "You will be notified automatically when it completes" — i.e. the notification model matches what's assumed. But the allowlist (ai-triage.yml:328) grants only Task, not any status/poll tool (no TaskOutput/BashOutput/Monitor equivalent). If run_in_background: false were ever silently ignored by the runtime (the literal premise of #338 — a CLI update changed the default), the model has no tool-based way to satisfy "don't end your turn" other than stalling on repeated no-op calls, which isn't specified. This is fine as belt-and-suspenders (the sentinel+watchdog is the real backstop either way), but the instruction itself is closer to aspirational than enforceable in that fallback scenario.

5. Minor: zero tolerance for chatter risks false failures
"your last output message must consist of exactly one line per command... and nothing else" (ai-triage.yml:347-355) is enforced purely through instruction-following. The captured logs show these sessions readily narrate themselves ("All five duplicate-search agents... I'll wait..."). A session that did complete its work correctly but adds even a one-line preamble/postamble to its final message now fails the job — trading the old silent-false-negative risk for a new (hopefully rarer, but real) noisy-false-positive risk. Worth watching after merge, per the PR's own noted verification caveat.

No contradictions found between the new synchronous-agent/rate-limit rules and the existing pre-check/idempotency sections of the three command files — the "silent exit still gets a sentinel" clarification (ai-triage.yml:352-355) correctly resolves the apparent tension with triage-find-duplicate-prs.md's "SILENT by default... post NOTHING" pre-check language.
· branch claude/ai-auto-triage-gh-issues-cmbm7y

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deep review — 0 finding(s)

# Severity Area Finding Location
— — — No blocking defects found. —

Overall: This is a focused CI/prompt hardening change (three .claude/commands/*.md triage prompts + the ai-triage.yml watchdog) with no component or public-API surface, so no story/docs/skills updates are owed. I exercised the new jq gate against clean-sentinel, trailing-chatter, is_error:true, null/empty result, missing-result-record, and array-vs-NDJSON inputs — it is correctly fail-closed (only a last result record with is_error==false whose text is nothing but TRIAGE-RESULT: lines passes; everything else, including the #338 orphaned-agent bailout message, fails the job). The one intentional strictness is that a leading-whitespace-indented sentinel line would be rejected, which matches the "and nothing else" contract in the prompt. The human's main focus should just be confirming the sentinel/idempotency contract is one the Sonnet triage session can reliably satisfy on the silent-skip and multi-command PR paths, since a compliant-but-empty final message is the only thing standing between green and a false CI failure.

🏄 Chill little config swell, dude — no gnarly bugs lurking under the surface, the watchdog holds its line and the sentinel logic paddles out clean. Good to send it.

…agent rules

Deep-review follow-up on #340: the watchdog accepted any nonzero number
of TRIAGE-RESULT lines, so a PR run that completed triage-find-issues
but bailed during triage-find-duplicate-prs passed on one sentinel. The
gate step's is_pr output now feeds the check, which requires exactly one
sentinel for issues and two for PRs. The search-agent command rules also
now name pipes and command substitution alongside loops/chains — the
same prefix allowlist denies those too.

Refs #338

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1LrTrVRbtdh8paN7f67CB

Copy link
Copy Markdown
Owner Author

Response to the deep adversarial review's five findings (addressed in d2ffb51):

Finding 1 (dropped second command on PR runs) — fixed. The gate step's is_pr output now feeds the watchdog, which requires exactly one TRIAGE-RESULT: line for issues and exactly two for PRs. Verified against the fixture set: a PR run with one sentinel now fails, two passes, and an issue run with two sentinels also fails (over-reporting is as suspicious as under-reporting).

Finding 3 (pipes/substitution omitted from the Bash rule) — fixed. All three command files now name pipes and command substitution alongside loops, echo prefixes, and ;/&& chains.

Finding 2 (sentinel is self-attested) — accepted as a limitation, no change. Cross-checking the claimed outcome would mean probing the item for marker comments with carve-outs for every legitimate silent exit (closed item, too-vague, no credible match) — that probing design was considered and rejected for exactly that complexity. Format+count validation catches the observed failure class (mid-task bailout); a session that fabricates a plausible sentinel without doing the work is a lying-agent problem no jq can solve.

Finding 4 ("never end your turn" lacks an enforcement mechanism) — accepted as belt-and-braces, no change. Correct that the instruction is aspirational if the runtime ignores run_in_background: false; the sentinel+count watchdog is the enforcement layer — in that scenario the run fails loudly instead of silently, which is the invariant this PR is actually establishing.

Finding 5 (zero-chatter rule risks false positives) — accepted deliberately, no change. A noisy false failure is strictly better than the silent false success it replaces: it's visible, re-runnable via the label, and diagnosable from the full session output (show_full_output: true). Will watch post-merge per the verification caveat; if chatter-induced failures show up in practice, loosening to "last line(s) must be the sentinels" is a one-line jq change.


Generated by Claude Code

@github-actions

Copy link
Copy Markdown
Contributor

Preview Deployment

Preview URL: https://79c62ccb.bestax.pages.dev

@allxsmith
allxsmith merged commit 48f0a60 into main Jul 21, 2026
10 checks passed
@allxsmith
allxsmith deleted the claude/ai-auto-triage-gh-issues-cmbm7y branch July 21, 2026 01:47
allxsmith added a commit that referenced this pull request Jul 21, 2026
…ot alone (#342)

The first live run after #340 (29794090279) false-failed: the session
correctly skipped closed issue #330 and emitted its sentinel, but
prefixed one explanatory line, which the nothing-but-sentinels check
rejected. The watchdog now requires the TRIAGE-RESULT lines to be the
LAST non-empty lines of the final message, exactly one per expected
command and none elsewhere. Bailouts never emit sentinels and still
fail; a bailout after a sentinel fails; over- and under-counts fail.

Refs #338

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1LrTrVRbtdh8paN7f67CB
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 3.5.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

allxsmith added a commit that referenced this pull request Jul 22, 2026
The deep-review label prompt validated diffs in isolation (0 findings on
PR #340) while a prompted @claude session on the same diff found 5,
including a real residual bug. Close the gap in what the reviewer is told
to do, not the model:

- Evidence phase: chase PR claims into linked issues and run/job logs,
  then verify the fix empirically — reading the diff is not verification.
- Residual-risk hunt: enumerate how the addressed failure class could
  still occur; refute with evidence or post as a finding. A Residual risk
  section is now required in the summary review.
- Advisory tier: 🔵 Advisory for real limitations/trade-offs worth
  putting on the record. Summary-table only — never inline, so advisory
  notes cannot become fixer work items and spin the AI loop.
- PR-type-aware lens: keep the component checklist for bulma-ui/docs
  diffs; add an infra lens (guard bypasses, shell+jq edge cases,
  untrusted input, silent failures, idempotency) for .github/.claude/
  scripts diffs.
- Optional focus steer: a triage+ user may pre-post a PR comment starting
  with deep-review: — a new gate step verifies each candidate author's
  live role (newest steer per distinct author, newest-first, max 5 role
  checks, so an outsider can neither inject text nor displace a
  maintainer's steer), and injects it as a FOCUS block via a
  random-delimiter heredoc output. Fail-soft by design: API failures
  degrade to an unfocused review, never a red run.

Unchanged: opus model, 120 turns, dedupe marker, one-review invariant,
allowed_bots, show_full_output debug-only, the is_error gate.

Docs: mention the steer in the AI development guide's label table and
CLAUDE.md's deep-review sentence.

Fixes #341
allxsmith added a commit that referenced this pull request Jul 22, 2026
The deep-review label prompt validated diffs in isolation (0 findings on
PR #340) while a prompted @claude session on the same diff found 5,
including a real residual bug. Close the gap in what the reviewer is told
to do, not the model:

- Evidence phase: chase PR claims into linked issues and run/job logs,
  then verify the fix empirically — reading the diff is not verification.
- Residual-risk hunt: enumerate how the addressed failure class could
  still occur; refute with evidence or post as a finding. A Residual risk
  section is now required in the summary review.
- Advisory tier: 🔵 Advisory for real limitations/trade-offs worth
  putting on the record. Summary-table only — never inline, so advisory
  notes cannot become fixer work items and spin the AI loop.
- PR-type-aware lens: keep the component checklist for bulma-ui/docs
  diffs; add an infra lens (guard bypasses, shell+jq edge cases,
  untrusted input, silent failures, idempotency) for .github/.claude/
  scripts diffs.
- Optional focus steer: a triage+ user may pre-post a PR comment starting
  with deep-review: — a new gate step verifies each candidate author's
  live role (newest steer per distinct author, newest-first, max 5 role
  checks, so an outsider can neither inject text nor displace a
  maintainer's steer), and injects it as a FOCUS block via a
  random-delimiter heredoc output. Fail-soft by design: API failures
  degrade to an unfocused review, never a red run.

Unchanged: opus model, 120 turns, dedupe marker, one-review invariant,
allowed_bots, show_full_output debug-only, the is_error gate.

Docs: mention the steer in the AI development guide's label table and
CLAUDE.md's deep-review sentence.

Fixes #341
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 5.8.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

allxsmith added a commit that referenced this pull request Jul 22, 2026
The deep-review label prompt validated diffs in isolation (0 findings on
PR #340) while a prompted @claude session on the same diff found 5,
including a real residual bug. Close the gap in what the reviewer is told
to do, not the model:

- Evidence phase: chase PR claims into linked issues and run/job logs,
  then verify the fix empirically — reading the diff is not verification.
- Residual-risk hunt: enumerate how the addressed failure class could
  still occur; refute with evidence or post as a finding. A Residual risk
  section is now required in the summary review.
- Advisory tier: 🔵 Advisory for real limitations/trade-offs worth
  putting on the record. Summary-table only — never inline, so advisory
  notes cannot become fixer work items and spin the AI loop.
- PR-type-aware lens: keep the component checklist for bulma-ui/docs
  diffs; add an infra lens (guard bypasses, shell+jq edge cases,
  untrusted input, silent failures, idempotency) for .github/.claude/
  scripts diffs.
- Optional focus steer: a triage+ user may pre-post a PR comment starting
  with deep-review: — a new gate step verifies each candidate author's
  live role (newest steer per distinct author, newest-first, max 5 role
  checks, so an outsider can neither inject text nor displace a
  maintainer's steer), and injects it as a FOCUS block via a
  random-delimiter heredoc output. Fail-soft by design: API failures
  degrade to an unfocused review, never a red run.

Unchanged: opus model, 120 turns, dedupe marker, one-review invariant,
allowed_bots, show_full_output debug-only, the is_error gate.

Docs: mention the steer in the AI development guide's label table and
CLAUDE.md's deep-review sentence.

Fixes #341
@bestax-release-bot

Copy link
Copy Markdown

🎉 This PR is included in version 2.0.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

@bestax-release-bot

Copy link
Copy Markdown

🎉 This PR is included in version 1.0.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: AI triage silently no-ops — headless session ends turn while background Task sub-agents still running

2 participants