Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
dfa1452
ci: add @aiter-bot review workflow (self-hosted GLM PR reviewer)
zufayu Sep 18, 2026
bc6ec70
ci: add .review-loop orchestration for @aiter-bot review
zufayu Sep 20, 2026
a1ceae3
ci: satisfy Black/Ruff/actionlint on review-loop tooling
zufayu Sep 20, 2026
995b077
ci: house review-loop tooling in the review-pr skill dir (aiter layout)
Sep 20, 2026
32a660c
review-pr: harden prompts against SKILL.md drift
Sep 20, 2026
e036ef6
review-pr: make check_prompts.py layout-independent (git-resolve SKIL…
Sep 20, 2026
4759b9e
aiter-review-bot: security + cleanup hardening
Sep 20, 2026
0d7dbba
aiter-review-bot: surface a held HIGH RISK review (phase-1 policy)
Sep 20, 2026
abd1e72
review-pr: make the Claude agent entrypoint configurable (AITER_REVIE…
Sep 20, 2026
f3a7052
review-pr: preflight gates gh version (fetch needs baseRefOid)
zufayu Sep 20, 2026
e02d72e
review-pr: run_one extracts WORK path-agnostically (not hardcoded /tmp)
zufayu Sep 20, 2026
f82b0e8
review-pr: fetch honors TMPDIR for the scratch WORK dir and its GC
zufayu Sep 20, 2026
537691a
review-pr: add RUNNER-SETUP.md (generic self-hosted runner contract)
zufayu Sep 20, 2026
eeaa35f
review-pr: single @aiter-bot review command (drop re-review)
zufayu Sep 20, 2026
201bba7
review-pr: retry the GLM worker/refuter on timeout
zufayu Sep 20, 2026
451e5f0
review-pr: GLM health-gate + outage notice to the PR (@-mention the G…
zufayu Sep 21, 2026
ce3965b
review-pr: classify failures and route each to its owner, drift-guarded
zufayu Sep 21, 2026
8666308
review-pr: guarantee every failed review names an owner (catch-all trap)
zufayu Sep 21, 2026
2dac99a
review-pr: fix findings from the bot's own review — post every verdic…
zufayu Sep 21, 2026
a1e670f
review-pr: fix the bot's second self-review findings (bugs in my own …
zufayu Sep 21, 2026
362054d
review-pr: fix third self-review findings (data-loss risk + stale refs)
zufayu Sep 21, 2026
4efc372
review-pr: two-job gate + real allowlist + real owner logins
zufayu Sep 22, 2026
1fabb0c
Merge branch 'main' into ci/aiter-review-bot (pick up CI pipeline fixes)
zufayu Sep 23, 2026
5d1900f
review-pr: fix EXIT trap forcing exit 1 on success (card never posted…
zufayu Sep 23, 2026
ff046c3
review-pr: _publish.py honors GITHUB_REPOSITORY (was hardcoding ROCm/…
zufayu Sep 23, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
48 changes: 48 additions & 0 deletions .claude/skills/review-pr/RUNNER-SETUP.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Runner setup — aiter-review-bot

The `@aiter-bot review` workflow (`.github/workflows/aiter-review-bot.yml`) runs on a
self-hosted runner labeled `self-hosted, box308`. This is what that runner must provide.
Verify the box with `bash .claude/skills/review-pr/preflight.sh` — it checks every item below
and prints a fix hint for each red. This doc explains how to make those checks pass; it holds
no machine-specific paths (each site fills in its own).

## 1. Register the runner (needs repo admin)

Creating a runner registration token requires **repo admin** on this repository (a `maintain`
or `push` role is not enough — the `registration-token` API returns 403). Either get an admin
to add the runner, or reuse a runner your org already has by giving it the labels above.

Install the runner on a **data volume with room to spare**, not `$HOME` or `/`: each review
checks out the repo and builds two worktrees (~180 MB apiece) under the runner's `_work`, and
a box whose root fs fills up will fail mid-review. Register with the labels `box308`.

## 2. Box config via the runner `.env` (loaded into every job)

The workflow's review step sets no environment on purpose, so the box-specific config lives in
`<runner-dir>/.env`, which the GitHub Actions runner injects into every job. Set at least:

TMPDIR=<data-volume>/tmp # review scratch (WORK) lands here, off the root fs
AITER_REVIEW_AGENT=<path>/claude # the headless Claude entrypoint for this box
ANTHROPIC_BASE_URL=http://localhost:30000 # the model endpoint (a local GLM here)
ANTHROPIC_MODEL=<model> # e.g. a local GLM served by the box
ANTHROPIC_SMALL_FAST_MODEL=<model> # same model, or the session-title call warns
PATH=<data-volume>/bin:/usr/local/bin:/usr/bin:/bin # a gh >= 2.24 must be first (fetch needs baseRefOid)
GH_CONFIG_DIR=<data-volume>/gh-config # gh auth stored here, NOT as GH_TOKEN in the env

## 3. gh auth off the job environment

`fetch.sh` calls `gh` to read the PR. Authenticate it under `GH_CONFIG_DIR` (from a token with
`public_repo` read) rather than exporting `GH_TOKEN`, so the headless review agent never
inherits a token from its environment:

GH_CONFIG_DIR=<data-volume>/gh-config gh auth login --with-token < token-file

## 4. Bot identity

Set the repo secret `AITER_BOT_TOKEN` to the bot account's PAT (`public_repo`). The workflow
uses it only in the claim and publish steps — never in the review step — so the review agent
cannot read it.

## 5. Verify

bash .claude/skills/review-pr/preflight.sh # all green = @aiter-bot review runs end to end
51 changes: 51 additions & 0 deletions .claude/skills/review-pr/_apply_refutation.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
"""Drop the findings the independent refuter KILLED from the card, so the seven gates see only
what survived. In the interactive review flow a human removes KILLED findings by hand before
finalizing; run_one.sh calls this to do it headlessly -- without it, the very first review whose
refuter kills a finding (its designed job) fails triage.py's independent gate and loses the whole
review.

independent.txt has one `SURVIVED|KILLED -- ...` line per card finding, in the card's order
(triage.py's independent gate relies on that same 1:1 order). We map by position and, only when
the counts line up 1:1, remove the KILLED findings; on any mismatch we change nothing and let the
gate report it, rather than guess.

Usage: _apply_refutation.py <card.md> <independent.txt>
"""

import pathlib
import re
import sys

FIND_PREFIX = ("\U0001F534", "⚠", "\U0001F4DD") # 🔴 ⚠️ 📝
VERDICT = re.compile(r"^\s*(SURVIVED|KILLED)\b")


def main():
if len(sys.argv) < 3:
print("usage: _apply_refutation.py <card.md> <independent.txt>"); return 2
card_p = pathlib.Path(sys.argv[1])
indep = pathlib.Path(sys.argv[2]).read_text(encoding="utf-8")
verdicts = [m.group(1) for m in (VERDICT.match(l) for l in indep.splitlines()) if m]

lines = card_p.read_text(encoding="utf-8").splitlines(keepends=True)
finding_lines = [i for i, l in enumerate(lines) if l.lstrip().startswith(FIND_PREFIX)]

if not finding_lines or not verdicts:
print("nothing to apply (no findings or no verdicts)"); return 0
if len(verdicts) != len(finding_lines):
# Ambiguous mapping -- do not guess which finding a verdict refers to. Leave the card as
# is; the independent gate will report the count mismatch.
print(f"skip: {len(finding_lines)} card findings vs {len(verdicts)} verdicts -- mismatch")
return 0

drop = {finding_lines[i] for i, v in enumerate(verdicts) if v == "KILLED"}
if not drop:
print("no KILLED findings; card unchanged"); return 0
kept = [l for i, l in enumerate(lines) if i not in drop]
card_p.write_text("".join(kept), encoding="utf-8")
print(f"applied refutation: dropped {len(drop)} KILLED finding(s), kept {len(finding_lines) - len(drop)}")
return 0


if __name__ == "__main__":
sys.exit(main())
100 changes: 100 additions & 0 deletions .claude/skills/review-pr/_collect.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,100 @@
import os
import pathlib
import shutil
import subprocess
import sys

sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import _lib

HERE = pathlib.Path(os.path.dirname(os.path.abspath(__file__)))
if len(sys.argv) < 2:
sys.exit("usage: collect.sh <WORK_DIR> [PR_number]")
W = sys.argv[1]
num, _ = _lib.pr_meta(W)
pr = sys.argv[2] if len(sys.argv) > 2 else (str(num) if num else None)
if not pr:
sys.exit(
f"cannot read PR number ({W}/pr_meta.json missing or corrupt); pass it explicitly: collect.sh {W} <PR_number>"
)

missing = [f for f in _lib.REQUIRED if not _lib.nonempty(os.path.join(W, f))]
if missing:
# Validate before making the dir: a refused collect must not leave an empty dir, or it
# gets picked up by the report index.
sys.exit(
f"refusing to collect: {W} is missing required artifacts {', '.join(missing)}.\n"
f"{W} is an incomplete WORK dir (fetch likely failed) -- re-run the review."
)

D = HERE / "reports" / f"PR-{pr}"
D.mkdir(parents=True, exist_ok=True)

# SKILL.md's `not independently refuted` is a conditional annotation (written only "when
# there is no independent reader"). We split Step 7.7 to a separate agent, so when the worker
# wrote the card it did not yet know whether a refuter would run and had to annotate "none".
# If independent.txt exists, an independent refutation did happen and the suffix must come off,
# or the card is lying. Mechanically restore the state SKILL.md specifies; no review judgement.
_card = pathlib.Path(W) / "card.md"
if _lib.nonempty(os.path.join(W, "independent.txt")):
_t = _card.read_text(encoding="utf-8")
if "not independently refuted" in _t:
_t2 = _t.replace(" — not independently refuted", "").replace(
" -- not independently refuted", ""
)
_card.write_text(_t2, encoding="utf-8")
print(
" removed 'not independently refuted' from the card (independent.txt exists)"
)

copied = 0
for f in _lib.REQUIRED + _lib.OPTIONAL:
src = os.path.join(W, f)
if _lib.nonempty(src):
shutil.copy2(src, D / f)
copied += 1
elif f in _lib.OPTIONAL:
(D / f).unlink(missing_ok=True)


# GATES.txt is this report's credibility receipt; keep it with the artifacts — write failures in too
_proj = subprocess.run(
["git", "-C", str(HERE), "rev-parse", "--show-toplevel"],
capture_output=True,
text=True,
check=False,
).stdout.strip() or os.path.dirname(os.path.dirname(os.path.dirname(str(HERE))))
res = _lib.run_gates(W, _proj)
npass = sum(1 for _, rc, _, _ in res if rc == 0)
lines = [f"======== SEVEN GATES ({W}) ========"]
for name, rc, first, full in res:
lines.append(
f" {'✅' if rc == 0 else '❌'} {name:<12} "
+ (
first
if rc == 0
else f"exit={rc}\n" + "\n".join(" " + l for l in full.split("\n"))
)
)
lines.append(f"======== {npass} green / {len(res)-npass} red ========")
(D / "GATES.txt").write_text("\n".join(lines) + "\n", encoding="utf-8")

print(f" collected into {D} ({copied} items, {npass}/7 green)")

# Refresh the aggregate report right after collect (a manual step gets forgotten, and a
# forgotten one leaves the report inconsistent with reality). _report.py is a dev-only tool
# and may not ship; skip it silently when absent.
_rep = HERE / "_report.py"
if _rep.is_file():
subprocess.run([sys.executable, str(_rep)], check=False)
if npass != 7:
print(
f" ⚠️ this report has {7-npass} gate(s) unpassed — the index marks it; do not treat it as a conclusion"
)

# collect is the last automatic step; the next step, git commit, is manual, and "a manual
# step gets forgotten" is exactly why #4860 sat unstaged in the worktree. Remind explicitly
# and point at the breakpoint detector.
print(
f" ↳ not landed yet: git add reports/PR-{pr}/ && git commit to finish."
)
40 changes: 40 additions & 0 deletions .claude/skills/review-pr/_gates.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
import os
import subprocess
import sys

sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import _lib

if len(sys.argv) < 2:
sys.exit("usage: gates.sh <WORK_DIR> [PROJECT_ROOT]")
W = sys.argv[1]
if len(sys.argv) > 2:
PROJ = sys.argv[2]
else:
PROJ = subprocess.run(
[
"git",
"-C",
os.path.dirname(os.path.abspath(__file__)),
"rev-parse",
"--show-toplevel",
],
capture_output=True,
text=True,
check=False,
).stdout.strip()

print(f"======== SEVEN GATES ({W}) ========")
res = _lib.run_gates(W, PROJ)
npass = 0
for name, rc, first, full in res:
if rc == 0:
npass += 1
print(f" ✅ {name:<12} {first}")
else:
print(f" ❌ {name:<12} exit={rc}")
for l in full.split("\n"):
print(f" {l}")
nfail = len(res) - npass
print(f"======== {npass} green / {nfail} red ========")
sys.exit(0 if nfail == 0 else 1)
Loading
Loading