Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 13 additions & 7 deletions .claude/commands/triage-dedupe.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,15 +16,21 @@ run locally), and `AUTOCLOSE` (`active` or `off`; assume `off` locally). If

1. `gh issue view NUMBER --repo REPO --json state,title,body,comments` — if
the issue is not open, stop.
2. Marker check — does any existing bot-authored comment contain
`<!-- ai-triage:dedupe -->`? (Match the marker + a bot author, not a
specific login — the workflow posts as github-actions[bot].)
2. Marker check — does any existing comment authored by bestaxbot or a
bot account contain `<!-- ai-triage:dedupe -->`? (Match the marker +
that author class, never one specific login — the workflow posts as
bestaxbot today; older comments are from github-actions[bot] or
claude[bot].)
- `TRIGGER=opened` and marker present → stop (already triaged).
- `TRIGGER=labeled` and marker present → continue; at the end REFRESH
that comment instead of posting a new one: if it is your most recent
comment on the issue, use
`gh issue comment NUMBER --repo REPO --edit-last --body ...`;
otherwise post a fresh comment (the old one stays as history).
that comment instead of posting a new one. Use
`gh issue comment NUMBER --repo REPO --edit-last --body ...` ONLY when
your most recent comment on the issue is itself the
`<!-- ai-triage:dedupe -->` comment: `--edit-last` selects by author,
not by marker, and every automation here posts as bestaxbot now, so a
newer repro draft or other machine comment would be the one it
overwrites. If anything else is newer, post a fresh comment (the old
one stays as history). See `.github/CLAUDE.md` rule 6.
3. If the issue is too vague to search meaningfully (no error text, no
component or file name, no concrete behavior), stop.

Expand Down
17 changes: 12 additions & 5 deletions .claude/commands/triage-find-duplicate-prs.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,13 +14,20 @@ assume `labeled` locally). If `NUMBER` is missing, ask.

1. `gh pr view NUMBER --repo REPO --json state,title,body,files,comments` —
if the PR is not open, stop.
2. Marker check — a bot-authored comment containing
`<!-- ai-triage:find-duplicate-prs -->` (match marker + bot author, not
a specific login — the workflow posts as github-actions[bot]):
2. Marker check — a comment authored by bestaxbot or a bot account
containing `<!-- ai-triage:find-duplicate-prs -->` (match marker + that
author class, never one specific login — the workflow posts as
bestaxbot today; older comments are from github-actions[bot] or
claude[bot]):
- `TRIGGER=opened` and marker present → stop.
- `TRIGGER=labeled` and marker present → continue; at the end refresh
that comment (`gh pr comment NUMBER --repo REPO --edit-last --body ...`
if it is your most recent comment on the PR; otherwise post fresh).
that comment with
`gh pr comment NUMBER --repo REPO --edit-last --body ...` ONLY when
your most recent comment on the PR is itself the
`<!-- ai-triage:find-duplicate-prs -->` comment. `--edit-last` selects
by author, not by marker, and the other triage commands post as the
same account — so a newer `find-issues` marker is what it would
overwrite. Otherwise post fresh. See `.github/CLAUDE.md` rule 6.

## Search — 3 parallel agents

Expand Down
17 changes: 12 additions & 5 deletions .claude/commands/triage-find-issues.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,13 +14,20 @@ EXACTLY ONE comment (or nothing). Context from the caller: `REPO`, `NUMBER`

1. `gh pr view NUMBER --repo REPO --json state,title,body,comments` — if the
PR is not open, stop.
2. Marker check — a bot-authored comment containing
`<!-- ai-triage:find-issues -->` (match marker + bot author, not a
specific login — the workflow posts as github-actions[bot]):
2. Marker check — a comment authored by bestaxbot or a bot account
containing `<!-- ai-triage:find-issues -->` (match marker + that
author class, never one specific login — the workflow posts as
bestaxbot today; older comments are from github-actions[bot] or
claude[bot]):
- `TRIGGER=opened` and marker present → stop.
- `TRIGGER=labeled` and marker present → continue; at the end refresh
that comment (`gh pr comment NUMBER --repo REPO --edit-last --body ...`
if it is your most recent comment on the PR; otherwise post fresh).
that comment with
`gh pr comment NUMBER --repo REPO --edit-last --body ...` ONLY when
your most recent comment on the PR is itself the
`<!-- ai-triage:find-issues -->` comment. `--edit-last` selects by
author, not by marker, and the other triage commands post as the same
account — so a newer `find-duplicate-prs` marker is what it would
overwrite. Otherwise post fresh. See `.github/CLAUDE.md` rule 6.
3. Note every issue already referenced in the PR body (`#N`, `Fixes #N`,
`Closes #N`, full URLs) — those are EXCLUDED from the results.

Expand Down
177 changes: 177 additions & 0 deletions .github/CLAUDE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,177 @@
# .github — CI and AI automation

Everything in `.github/workflows/` is **human-authored by design**. AI agents working in this
repository are denied write access to this directory (`Edit(.github/**)` and friends sit in
every write agent's `--disallowedTools`), and the AI fix loop refuses any PR that touches
`.github/**`. If you are an agent reading this: propose a diff in a comment, do not apply one.

This file is the security contract for these workflows. The rules below are not style
preferences — each one is load-bearing, and most were written after a review round or a
red-team found the failure it prevents. Where a rule has a documented origin, it is cited.

## The two invariants

Every AI workflow here is built to preserve these. If a change breaks one, the change is wrong,
regardless of how convenient it is.

- **I1 — the model-auth token never shares a job with code execution.** `CLAUDE_CODE_OAUTH_TOKEN`
pays for a session; anything that can run code in the same job can read it out of the
environment. `claude-repro.yml` splits drafting from publishing for exactly this reason: the
drafting job has no `Bash`, no `Task`, no network tool, so it cannot read env, fetch, or
execute. Note the direction of the guarantee — it raises the cost of exfiltration, it is not a
proof. The credential-shape check in `Collect draft` (literal, base64, and hex forms) is a
backstop for the same reason, and is documented as one.
- **I2 — no untrusted or model-authored free text reaches a re-trigger-capable identity.**
Comments posted with `GITHUB_TOKEN` do not emit workflow events. Comments posted with a PAT
**do**. So a drafted reproduction is sanitized deterministically and posted via
`GITHUB_TOKEN`, and every comment-triggered workflow independently gates out bestaxbot.

## Hard requirements

### 1. Pin every third-party action to a full commit SHA — the same SHA everywhere

`SECURITY.md` advertises SHA pinning as an active control, so this is a promise to users, not
housekeeping. One SHA per action across all workflow files, and the trailing comment must name
the version that SHA **actually is**.

Current pin: `actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7` (this is v7.0.1).

When bumping, bump every occurrence in a single commit and verify there is exactly one:

```bash
grep -rho "actions/checkout@[a-f0-9]\{40\}" .github/workflows/ | sort | uniq -c # expect one line
```

Origin: #361 shipped two new workflows pinned to `9c091bb…` (v7.0.0, twelve commits behind)
while nineteen other usages were on v7.0.1 — and both were commented `# v7`, so the drift was
invisible to a reader. Copilot caught it; nothing in CI would have.

### 2. Never widen `--allowedTools` on a session that holds a credential

**Treat every allowlist in this directory as security-critical.** It is a confinement boundary,
not a convenience list, and widening it grants capability with **no permissions diff for a
reviewer to notice** — the `permissions:` block looks identical before and after.

Two sessions where the allowlist is the _only_ thing between untrusted text and repository
write:

| Workflow | Credential in the job | What the allowlist is holding back |
| --------------- | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `ai-triage.yml` | `AI_LOOP_PAT` (bestaxbot) | Full repo write. Confined to GET-only `gh` reads plus the two comment commands. |
| `ai-scan.yml` | job `GITHUB_TOKEN`, **write**-scoped (`issues`, `pull-requests`) | The gate charges a budget marker and the labeler applies `needs-security-review`, so the token must be write-scoped. The session cannot use it _only_ because the Bash allowlist admits nothing that writes. |

Concrete rules:

- Adding **any** entry to those two allowlists is a security change. Say so in the PR
description and explain why the entry cannot write.
- Never add `Bash(gh api:*)` to a session that ingests untrusted issue/PR text — it is a
general-purpose write primitive wearing a read-shaped name.
- Never add `Edit`, `Write`, `MultiEdit`, or `Task` to `ai-scan.yml`.
- `--disallowedTools` is defense in depth, and its deny rules do take precedence over the
allows — but do not lean on it as the primary control. Narrow the allowlist.
Comment on lines +55 to +71

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail
rg -n -C 12 -e 'AI_LOOP_PAT|allowedTools|Bash\(gh|gh (issue|pr) comment|gh api' .github/workflows/ai-triage.yml

Repository: allxsmith/bestax

Length of output: 13303


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf 'CLAUDE.md relevant lines:\n'
sed -n '50,75p' .github/CLAUDE.md | cat -n

printf '\nai-triage relevant lines:\n'
sed -n '340,372p' .github/workflows/ai-triage.yml | cat -n
printf '\nAllowlist strings in ai-triage.yml:\n'
rg -n -C 2 -- '--allowedTools|--disallowedTools' .github/workflows/ai-triage.yml

Repository: allxsmith/bestax

Length of output: 6043


Remove PAT-backed comment writes from the model session.

ai-triage.yml gives the Claude session AI_LOOP_PAT plus Bash(gh pr comment:*) and Bash(gh issue comment:*). A prompt-injected request can reach a re-trigger-capable issue/PR comment as bestaxbot. Remove those comment tools from the model allowance; use a deterministic publisher with GITHUB_TOKEN, or pass only a fixed, validated payload through a separate publisher.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/CLAUDE.md around lines 55 - 71, Remove the PAT-backed comment
permissions from the Claude session configuration in ai-triage.yml by deleting
Bash(gh pr comment:*) and Bash(gh issue comment:*) from its allowlist. Route
comment publication through a deterministic publisher using GITHUB_TOKEN or a
separate publisher that accepts only a fixed, validated payload, while
preserving the model’s read-only gh access.

- Prefer removing the need for the boundary over hardening it. Splitting `ai-scan`'s labeler
into its own job so the model session can drop to `contents: read` is tracked in #455.

### 3. Opt in explicitly for anything that spends model usage

Gate on `== 'on'`, never on `!= 'off'`. Unset, empty, a typo, and every spelling of "stop" must
all mean **off** — no misspelling should be able to _start_ a control that spends usage on every
incoming item.

```yaml
if: vars.AI_LOOP_ENABLED == 'true' &&
(vars.AI_SCAN_MODE == 'on' || vars.AI_SCAN_MODE == 'y')
```

- GitHub Actions `==` on strings is **case-insensitive**, so `ON` and `Y` work too. Do not add
case variants to the expression.
- The one deliberate exception is `AI_TRIAGE_MODE`, whose label path is `!= 'off'` so that
human-applied labels still work with the variable unset. It is an exception because a label
requires a trusted human first.
- Every repository variable that steers this automation is tabulated in the ai-development docs
guide, **including its unset default**. Add new ones there in the same PR. Origin: that table
documented `AI_LOOP_ENABLED` backwards until Copilot caught it — the gates are `== 'true'`, so
unset means disabled.

### 4. Fail closed on verdicts, fail open on budgets

Both directions are deliberate and the split is the point:

- A **verdict** (is this item malicious?) fails **closed** — a crashed, missing, unparsable, or
forged sentinel flags rather than passes.
- A **budget or counter** fails **open** — if the daily-limit read errors, the item is simply
not scanned. A counter problem must not be able to wedge every incoming issue.

The cost of the second is real and should stay written down: downstream, an unscanned item is
indistinguishable from a clean one.

### 5. Secret masking scrubs logs, not comment bodies

Anything a session writes into a comment is published verbatim. Never build a comment body out
of model output without a deterministic sanitizer between them, and never assume masking will
save you — it applies to the job log only.

### 6. Scope comment edits by marker, never `--edit-last`

`gh ... --edit-last` picks the most recent comment by that author, which is not necessarily
yours. It could overwrite a historical `<!-- ai-triage:dedupe -->` marker that
`auto-close-duplicates.mjs` still reads — silently destroying an auto-close candidate. Select by
the marker your workflow owns.

More generally: when probing for a machine comment, match on **marker + (bestaxbot OR a
Bot-type author)**. Never probe one specific login. bestaxbot is a machine _User_ account, not a
Bot-type app, so a `type == 'Bot'` test alone misses it and a login test alone breaks the next
time the identity changes.

### 7. Fork PRs never run with secrets

Use plain `pull_request`, never `pull_request_target`, plus an explicit head-repo guard (#312).

### 8. Sender exclusions belong on every comment-triggered workflow

`contains()` matches a raw substring, so an `@mention` reproduced anywhere in a body — including
inside a quoted block or a code fence — is enough to re-trigger a write-capable session. Require
`sender.type == 'User'` **and** `sender.login != 'bestaxbot'`, and forbid the session from
writing the trigger string at all.

### 9. Logic worth testing does not belong in YAML

Shell embedded in a workflow step cannot be unit-tested without extracting it first. Put
non-trivial parsing in `scripts/*.mjs` with a `node --test` sibling — root `pnpm test` runs
`node --test "scripts/*.test.mjs"` (the glob is quoted so Node expands it, not the shell).
Two parsers are still inline and tracked in #454: the publish sanitizer and the scan-verdict
parser.

### 10. `harden-runner` on new jobs starts at `block`

New jobs ship with `egress-policy: block`. An **existing** live job may be introduced at `audit`
first, because flipping straight to block risks breaking it if the action's runtime egress needs
an un-allow-listed host — but that is a temporary state that owes a follow-up issue, not a
resting place. `ai-triage.yml` is currently the one job at `audit`.

## Review checklist for a workflow change

- [ ] Does any allowlist grow? If yes, name the credential in that job and justify each entry.
- [ ] Does a model session gain the ability to execute, fetch, or reach the network?
- [ ] Does model or issue text reach a comment body? Through what sanitizer?
- [ ] Is the posting identity `GITHUB_TOKEN` (inert) or a PAT (re-triggering)?
- [ ] New repository variable? Gate is `== 'on'`-shaped, and the docs table lists its unset default.
- [ ] Action SHAs match the repo-wide pin, and the version comment is truthful.
- [ ] New verdict path fails closed; new counter path fails open.
- [ ] Security comments claim exactly what the mechanism delivers — no more.

That last item is not padding. Three separate review rounds on #361 flagged comments that
overstated their mechanism: a flag described as blocking "every AI entry point" when third-party
reviewers never see it, a session called "read-only" when its token was write-scoped, and an
encoding check described as a proof when it is a backstop. A comment that overstates its control
is worse than no comment, because the next reader stops checking.

## Where the rest is documented

- **Design and operation of the AI loop, and the repository-variable table** — the
ai-development guide in `docs/docs/guides/getting-started/`.
- **User-facing security posture** — `SECURITY.md`.
- **Release and publish pipeline** — `ci.yml`, plus `VERSIONING.md` for the semantic-release
contract. Note that `main` is protected by a repository ruleset whose only automation bypass
is the release GitHub App, and its token is minted _after_ install and build so repo-owned
build code can never reach it.
Loading
Loading