Repository navigation
ci: triage-as-bestaxbot, safe robobun-style repro drafting + read-only security scan #361
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
15 commits
Select commit
Hold shift + click to select a range
c6f7c9f
ci: post AI triage comments as bestaxbot instead of github-actions[bot]
claude 5af2cc3
ci: add author-only repro and security-scan workflows
claude 1e2d502
ci: guard AI entry points against re-trigger and flagged content
claude 08a070b
docs: document repro and security-scan automation
claude 27fbb2d
ci: close the repro token-exfil path and correct the flag's blast radius
allxsmith 2fed974
test: cover the auto-close decision logic, and run root scripts in CI
allxsmith 7442268
docs: tabulate the AI repository variables and their unset defaults
allxsmith 989e141
ci: correct the variable defaults and stop overstating the security flag
allxsmith 29ee909
ci: scope the repro comment edit, and stop claiming a read-only scan …
allxsmith a301acf
ci: make the security scan explicit opt-in (AI_SCAN_MODE must be `on`)
allxsmith bf813ff
ci: the scan runs on `on` or `y`, and on nothing else
allxsmith b956097
docs: add inbound security triage to the SECURITY.md and README posture
allxsmith 04103cd
ci: add .github/CLAUDE.md workflow security contract
allxsmith c613cbd
ci: scope triage comment refreshes by marker, not by author
allxsmith d8f4b86
ci: drop the unverified label auto-create claim in ai-scan
allxsmith File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,177 @@ | ||
| # .github — CI and AI automation | ||
|
|
||
| Everything in `.github/workflows/` is **human-authored by design**. AI agents working in this | ||
| repository are denied write access to this directory (`Edit(.github/**)` and friends sit in | ||
| every write agent's `--disallowedTools`), and the AI fix loop refuses any PR that touches | ||
| `.github/**`. If you are an agent reading this: propose a diff in a comment, do not apply one. | ||
|
|
||
| This file is the security contract for these workflows. The rules below are not style | ||
| preferences — each one is load-bearing, and most were written after a review round or a | ||
| red-team found the failure it prevents. Where a rule has a documented origin, it is cited. | ||
|
|
||
| ## The two invariants | ||
|
|
||
| Every AI workflow here is built to preserve these. If a change breaks one, the change is wrong, | ||
| regardless of how convenient it is. | ||
|
|
||
| - **I1 — the model-auth token never shares a job with code execution.** `CLAUDE_CODE_OAUTH_TOKEN` | ||
| pays for a session; anything that can run code in the same job can read it out of the | ||
| environment. `claude-repro.yml` splits drafting from publishing for exactly this reason: the | ||
| drafting job has no `Bash`, no `Task`, no network tool, so it cannot read env, fetch, or | ||
| execute. Note the direction of the guarantee — it raises the cost of exfiltration, it is not a | ||
| proof. The credential-shape check in `Collect draft` (literal, base64, and hex forms) is a | ||
| backstop for the same reason, and is documented as one. | ||
| - **I2 — no untrusted or model-authored free text reaches a re-trigger-capable identity.** | ||
| Comments posted with `GITHUB_TOKEN` do not emit workflow events. Comments posted with a PAT | ||
| **do**. So a drafted reproduction is sanitized deterministically and posted via | ||
| `GITHUB_TOKEN`, and every comment-triggered workflow independently gates out bestaxbot. | ||
|
|
||
| ## Hard requirements | ||
|
|
||
| ### 1. Pin every third-party action to a full commit SHA — the same SHA everywhere | ||
|
|
||
| `SECURITY.md` advertises SHA pinning as an active control, so this is a promise to users, not | ||
| housekeeping. One SHA per action across all workflow files, and the trailing comment must name | ||
| the version that SHA **actually is**. | ||
|
|
||
| Current pin: `actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7` (this is v7.0.1). | ||
|
|
||
| When bumping, bump every occurrence in a single commit and verify there is exactly one: | ||
|
|
||
| ```bash | ||
| grep -rho "actions/checkout@[a-f0-9]\{40\}" .github/workflows/ | sort | uniq -c # expect one line | ||
| ``` | ||
|
|
||
| Origin: #361 shipped two new workflows pinned to `9c091bb…` (v7.0.0, twelve commits behind) | ||
| while nineteen other usages were on v7.0.1 — and both were commented `# v7`, so the drift was | ||
| invisible to a reader. Copilot caught it; nothing in CI would have. | ||
|
|
||
| ### 2. Never widen `--allowedTools` on a session that holds a credential | ||
|
|
||
| **Treat every allowlist in this directory as security-critical.** It is a confinement boundary, | ||
| not a convenience list, and widening it grants capability with **no permissions diff for a | ||
| reviewer to notice** — the `permissions:` block looks identical before and after. | ||
|
|
||
| Two sessions where the allowlist is the _only_ thing between untrusted text and repository | ||
| write: | ||
|
|
||
| | Workflow | Credential in the job | What the allowlist is holding back | | ||
| | --------------- | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | ||
| | `ai-triage.yml` | `AI_LOOP_PAT` (bestaxbot) | Full repo write. Confined to GET-only `gh` reads plus the two comment commands. | | ||
| | `ai-scan.yml` | job `GITHUB_TOKEN`, **write**-scoped (`issues`, `pull-requests`) | The gate charges a budget marker and the labeler applies `needs-security-review`, so the token must be write-scoped. The session cannot use it _only_ because the Bash allowlist admits nothing that writes. | | ||
|
|
||
| Concrete rules: | ||
|
|
||
| - Adding **any** entry to those two allowlists is a security change. Say so in the PR | ||
| description and explain why the entry cannot write. | ||
| - Never add `Bash(gh api:*)` to a session that ingests untrusted issue/PR text — it is a | ||
| general-purpose write primitive wearing a read-shaped name. | ||
| - Never add `Edit`, `Write`, `MultiEdit`, or `Task` to `ai-scan.yml`. | ||
| - `--disallowedTools` is defense in depth, and its deny rules do take precedence over the | ||
| allows — but do not lean on it as the primary control. Narrow the allowlist. | ||
| - Prefer removing the need for the boundary over hardening it. Splitting `ai-scan`'s labeler | ||
| into its own job so the model session can drop to `contents: read` is tracked in #455. | ||
|
|
||
| ### 3. Opt in explicitly for anything that spends model usage | ||
|
|
||
| Gate on `== 'on'`, never on `!= 'off'`. Unset, empty, a typo, and every spelling of "stop" must | ||
| all mean **off** — no misspelling should be able to _start_ a control that spends usage on every | ||
| incoming item. | ||
|
|
||
| ```yaml | ||
| if: vars.AI_LOOP_ENABLED == 'true' && | ||
| (vars.AI_SCAN_MODE == 'on' || vars.AI_SCAN_MODE == 'y') | ||
| ``` | ||
|
|
||
| - GitHub Actions `==` on strings is **case-insensitive**, so `ON` and `Y` work too. Do not add | ||
| case variants to the expression. | ||
| - The one deliberate exception is `AI_TRIAGE_MODE`, whose label path is `!= 'off'` so that | ||
| human-applied labels still work with the variable unset. It is an exception because a label | ||
| requires a trusted human first. | ||
| - Every repository variable that steers this automation is tabulated in the ai-development docs | ||
| guide, **including its unset default**. Add new ones there in the same PR. Origin: that table | ||
| documented `AI_LOOP_ENABLED` backwards until Copilot caught it — the gates are `== 'true'`, so | ||
| unset means disabled. | ||
|
|
||
| ### 4. Fail closed on verdicts, fail open on budgets | ||
|
|
||
| Both directions are deliberate and the split is the point: | ||
|
|
||
| - A **verdict** (is this item malicious?) fails **closed** — a crashed, missing, unparsable, or | ||
| forged sentinel flags rather than passes. | ||
| - A **budget or counter** fails **open** — if the daily-limit read errors, the item is simply | ||
| not scanned. A counter problem must not be able to wedge every incoming issue. | ||
|
|
||
| The cost of the second is real and should stay written down: downstream, an unscanned item is | ||
| indistinguishable from a clean one. | ||
|
|
||
| ### 5. Secret masking scrubs logs, not comment bodies | ||
|
|
||
| Anything a session writes into a comment is published verbatim. Never build a comment body out | ||
| of model output without a deterministic sanitizer between them, and never assume masking will | ||
| save you — it applies to the job log only. | ||
|
|
||
| ### 6. Scope comment edits by marker, never `--edit-last` | ||
|
|
||
| `gh ... --edit-last` picks the most recent comment by that author, which is not necessarily | ||
| yours. It could overwrite a historical `<!-- ai-triage:dedupe -->` marker that | ||
| `auto-close-duplicates.mjs` still reads — silently destroying an auto-close candidate. Select by | ||
| the marker your workflow owns. | ||
|
|
||
| More generally: when probing for a machine comment, match on **marker + (bestaxbot OR a | ||
| Bot-type author)**. Never probe one specific login. bestaxbot is a machine _User_ account, not a | ||
| Bot-type app, so a `type == 'Bot'` test alone misses it and a login test alone breaks the next | ||
| time the identity changes. | ||
|
|
||
| ### 7. Fork PRs never run with secrets | ||
|
|
||
| Use plain `pull_request`, never `pull_request_target`, plus an explicit head-repo guard (#312). | ||
|
|
||
| ### 8. Sender exclusions belong on every comment-triggered workflow | ||
|
|
||
| `contains()` matches a raw substring, so an `@mention` reproduced anywhere in a body — including | ||
| inside a quoted block or a code fence — is enough to re-trigger a write-capable session. Require | ||
| `sender.type == 'User'` **and** `sender.login != 'bestaxbot'`, and forbid the session from | ||
| writing the trigger string at all. | ||
|
|
||
| ### 9. Logic worth testing does not belong in YAML | ||
|
|
||
| Shell embedded in a workflow step cannot be unit-tested without extracting it first. Put | ||
| non-trivial parsing in `scripts/*.mjs` with a `node --test` sibling — root `pnpm test` runs | ||
| `node --test "scripts/*.test.mjs"` (the glob is quoted so Node expands it, not the shell). | ||
| Two parsers are still inline and tracked in #454: the publish sanitizer and the scan-verdict | ||
| parser. | ||
|
|
||
| ### 10. `harden-runner` on new jobs starts at `block` | ||
|
|
||
| New jobs ship with `egress-policy: block`. An **existing** live job may be introduced at `audit` | ||
| first, because flipping straight to block risks breaking it if the action's runtime egress needs | ||
| an un-allow-listed host — but that is a temporary state that owes a follow-up issue, not a | ||
| resting place. `ai-triage.yml` is currently the one job at `audit`. | ||
|
|
||
| ## Review checklist for a workflow change | ||
|
|
||
| - [ ] Does any allowlist grow? If yes, name the credential in that job and justify each entry. | ||
| - [ ] Does a model session gain the ability to execute, fetch, or reach the network? | ||
| - [ ] Does model or issue text reach a comment body? Through what sanitizer? | ||
| - [ ] Is the posting identity `GITHUB_TOKEN` (inert) or a PAT (re-triggering)? | ||
| - [ ] New repository variable? Gate is `== 'on'`-shaped, and the docs table lists its unset default. | ||
| - [ ] Action SHAs match the repo-wide pin, and the version comment is truthful. | ||
| - [ ] New verdict path fails closed; new counter path fails open. | ||
| - [ ] Security comments claim exactly what the mechanism delivers — no more. | ||
|
|
||
| That last item is not padding. Three separate review rounds on #361 flagged comments that | ||
| overstated their mechanism: a flag described as blocking "every AI entry point" when third-party | ||
| reviewers never see it, a session called "read-only" when its token was write-scoped, and an | ||
| encoding check described as a proof when it is a backstop. A comment that overstates its control | ||
| is worse than no comment, because the next reader stops checking. | ||
|
|
||
| ## Where the rest is documented | ||
|
|
||
| - **Design and operation of the AI loop, and the repository-variable table** — the | ||
| ai-development guide in `docs/docs/guides/getting-started/`. | ||
| - **User-facing security posture** — `SECURITY.md`. | ||
| - **Release and publish pipeline** — `ci.yml`, plus `VERSIONING.md` for the semantic-release | ||
| contract. Note that `main` is protected by a repository ruleset whose only automation bypass | ||
| is the release GitHub App, and its token is minted _after_ install and build so repo-owned | ||
| build code can never reach it. | ||
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift
🧩 Analysis chain
🏁 Script executed:
Repository: allxsmith/bestax
Length of output: 13303
🏁 Script executed:
Repository: allxsmith/bestax
Length of output: 6043
Remove PAT-backed comment writes from the model session.
ai-triage.ymlgives the Claude sessionAI_LOOP_PATplusBash(gh pr comment:*)andBash(gh issue comment:*). A prompt-injected request can reach a re-trigger-capable issue/PR comment as bestaxbot. Remove those comment tools from the model allowance; use a deterministic publisher withGITHUB_TOKEN, or pass only a fixed, validated payload through a separate publisher.🤖 Prompt for AI Agents