Repository navigation
ci: add engine version watch and nightly failure triage automation - #1973
Conversation
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
📝 WalkthroughWalkthroughAdds scheduled engine-version monitoring and nightly workflow triage. The workflows collect upstream updates or failed runs, then use Claude-assisted jobs to create or update labeled GitHub issues. ChangesEngine version watch
Nightly workflow triage
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant Scheduler
participant DetectionJobs
participant ExternalSources
participant Claude
participant GitHubIssues
Scheduler->>DetectionJobs: start scheduled workflows
DetectionJobs->>ExternalSources: fetch engine versions or workflow runs
ExternalSources-->>DetectionJobs: return update and failure data
DetectionJobs->>Claude: process pending updates or failures
Claude->>GitHubIssues: create or update labeled issues
Possibly related PRs
Suggested labels: Suggested reviewers: Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a86bca5402
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
|
👋 The PR description doesn't fully follow
Please update the PR description so reviewers have the context they need. |
| emit sglang "$sglang_current" "$sglang_latest" | ||
|
|
||
| vllm_current=$(sed -n "s/.*default: 'vllm\/vllm-openai:v\([0-9.]*\)'.*/\1/p" .github/workflows/release-vllm-docker.yml | head -1) | ||
| vllm_latest=$(curl -fsS https://pypi.org/pypi/vllm/json | jq -r .info.version) | ||
| emit vllm "$vllm_current" "$vllm_latest" | ||
|
|
||
| trtllm_current=$(sed -n 's/^TRTLLM_VERSION="\(.*\)"$/\1/p' scripts/ci_install_trtllm.sh | head -1) | ||
| trtllm_latest=$(curl -fsS https://pypi.nvidia.com/tensorrt-llm/ \ | ||
| | grep -o 'tensorrt_llm-[0-9][^-]*' | sed 's/tensorrt_llm-//' | sort -uV | tail -1) | ||
| emit tensorrt-llm "$trtllm_current" "$trtllm_latest" | ||
|
|
||
| tokenspeed_current=$(sed -n 's/.*TOKENSPEED_REF:-\([0-9a-f]*\)}.*/\1/p' scripts/ci_install_tokenspeed.sh | head -1) | ||
| tokenspeed_latest=$(gh api repos/lightseekorg/tokenspeed/commits/main --jq .sha) | ||
| emit tokenspeed "$tokenspeed_current" "$tokenspeed_latest" |
There was a problem hiding this comment.
🟡 Nit: The sed -n extraction patterns for *_current values return empty strings silently if the source file format changes (sed exits 0 with no match, unlike grep). An empty current would make emit report update: true with current: "", causing Claude to file a confusing upgrade issue with no pin info.
Consider guarding each extraction, e.g.:
sglang_current=$(sed -n 's/.*"sglang\[all\]==\([^"]*\)".*/\1/p' scripts/ci_install_sglang.sh | head -1)
if [ -z "$sglang_current" ]; then
echo "WARNING: failed to extract sglang pin" >&2
else
sglang_latest=$(curl -fsS https://pypi.org/pypi/sglang/json | jq -r .info.version)
emit sglang "$sglang_current" "$sglang_latest"
fiThis way a changed file format fails visibly in the Actions log rather than producing a bogus issue.
Separately: the TRT-LLM extraction uses grep -o in a pipeline (line 35), which unlike sed -n exits 1 on no match. Under pipefail that crashes the entire script before tokenspeed is checked — a single engine's scrape failure blocks all engines downstream.
There was a problem hiding this comment.
Clean CI automation — well-scoped permissions, deterministic detection gates before Claude runs, idempotent label creation, and proper dedup against open issues. Two minor nits posted inline (empty-value guards in the version checker, and allowed-tools scope for triage). Nothing blocking.
Summary: 0 🔴 Important · 2 🟡 Nit · 0 🟣 Pre-existing
There was a problem hiding this comment.
Actionable comments posted: 13
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/engine-version-watch.yml:
- Around line 24-25: Update both actions/checkout steps in the workflow,
including the checkout in the research-and-file job, to set persist-credentials
to false while preserving their existing checkout behavior.
- Around line 78-81: Validate that ANTHROPIC_API_KEY is non-empty in the “Export
API key from pod env” step before masking or writing it to GITHUB_ENV. If it is
unset, emit a clear error message and exit nonzero; otherwise preserve the
existing mask and export behavior for claude-code-action.
- Around line 7-10: Add a workflow-level concurrency guard to the engine-version
watch workflow, using a stable group key shared by scheduled and manually
dispatched runs and cancel-in-progress behavior that prevents overlapping
executions. Keep the existing dedup and research-and-file steps unchanged.
- Around line 12-14: Move issues: write out of the workflow-level permissions
block and declare it only in the job-level permissions for the jobs that create
or update issues, including research-and-file. Keep contents: read scoped as
currently required and ensure jobs without issue access do not inherit issues:
write.
- Around line 119-131: Ensure the issue-creation setup makes the enhancement
label available before the gh issue create command runs. Update the existing
label initialization near gh label create --force to idempotently create or
restore both engine-watch and enhancement, while preserving the current labels
used by the issue command.
- Around line 46-68: Replace the quoted title search in the dedup loop with a
structured marker or metadata-based check that uniquely identifies the engine
and latest version, avoiding GitHub search parsing of punctuation in names and
versions. Update the `existing` calculation while preserving the current
behavior of excluding rows with an open matching `engine-watch` issue.
In @.github/workflows/nightly-triage.yml:
- Around line 35-36: Update both gh installation sites in
.github/workflows/nightly-triage.yml at lines 35-36 and 79-80: download the
release archive and its matching checksum, verify the archive with sha256sum
--check, and only then extract it. Apply the same verified-install flow in both
the collect and triage jobs, preserving the existing versioned archive and
destination paths.
- Around line 14-17: Update the permissions configuration in the nightly triage
workflow so the collect job receives only actions: read, while the triage job
retains actions: write and issues: write. Scope these permissions per job rather
than granting both write scopes globally.
- Around line 94-128: Replace the shell-specific commands and pipelines in the
prompt block with declarative natural-language instructions for fetching
failed-run evidence, applying the conservative rerun policy, and finding or
updating the rolling triage issue. Preserve the repository/workflow context,
classification criteria, attempt restrictions, recurring-signature reporting,
concise format, and prohibition on file edits or pull requests.
- Around line 9-12: Add a concurrency group to the nightly triage workflow
associated with the existing schedule and workflow_dispatch triggers, using a
stable workflow-specific group key and disabling cancellation so queued runs
execute serially. Leave the trigger definitions and triage job behavior
unchanged.
- Around line 48-54: Remove the `|| rows='[]'` fallback from the `gh run list`
collection command so authentication, rate-limit, and network errors propagate
and fail the collection job. Keep the existing successful-result filtering and
`failures` aggregation unchanged.
In `@scripts/check_engine_versions.sh`:
- Around line 27-36: Add explicit connection and overall request timeouts to
every curl invocation in the version checks for sglang_latest, vllm_latest, and
trtllm_latest. Preserve the existing failure and output flags while ensuring
stalled upstream requests fail promptly instead of waiting for the job-level
timeout.
- Around line 15-41: Validate every extracted current and latest version before
calling emit, including sglang_current, sglang_latest, vllm_current,
vllm_latest, trtllm_current, trtllm_latest, tokenspeed_current, and
tokenspeed_latest. Fail immediately with a clear error when any value is empty,
preventing emit from comparing invalid data or producing misleading update
results; preserve the existing extraction and emit behavior for valid values.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: a31ad37c-567b-40c0-a72c-b07fc2ee303b
📒 Files selected for processing (3)
.github/workflows/engine-version-watch.yml.github/workflows/nightly-triage.ymlscripts/check_engine_versions.sh
| - name: Drop engines that already have an open issue | ||
| id: dedup | ||
| env: | ||
| GH_TOKEN: ${{ github.token }} | ||
| run: | | ||
| gh label create engine-watch --repo "$GITHUB_REPOSITORY" \ | ||
| --description "Automated engine release tracking" --color 0e8a16 --force | ||
| pending='[]' | ||
| while IFS= read -r row; do | ||
| engine=$(jq -r .engine <<<"$row") | ||
| latest=$(jq -r .latest <<<"$row") | ||
| existing=$(gh issue list --repo "$GITHUB_REPOSITORY" --state open \ | ||
| --label engine-watch --search "\"$engine $latest\" in:title" \ | ||
| --json number --jq length) | ||
| if [ "$existing" = "0" ]; then | ||
| pending=$(jq -c --argjson r "$row" '. + [$r]' <<<"$pending") | ||
| else | ||
| echo "Skipping $engine $latest: open issue exists" | ||
| fi | ||
| done < <(jq -c '.[]' updates.json) | ||
| echo "pending=$pending" >> "$GITHUB_OUTPUT" | ||
| echo "has_pending=$(jq 'length > 0' <<<"$pending")" >> "$GITHUB_OUTPUT" | ||
|
|
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🌐 Web query:
Does GitHub's issue search "in:title" with a quoted phrase containing dots or hyphens reliably match only the exact literal substring?
💡 Result:
No, GitHub's issue search does not reliably match quoted phrases containing dots or hyphens as exact literal substrings [1]. While GitHub issue search supports double quotes for multi-word phrases [2][3], the underlying search mechanism typically processes queries by normalizing text, which often involves treating punctuation marks like dots and hyphens as noise or delimiters [1]. Consequently, a search for a phrase such as "item-name" or "version.1" frequently results in the search engine treating those characters as separators rather than literal parts of the string, causing the search to return matches for the individual words (e.g., "item" and "name") rather than the specific, combined string [4][1]. This limitation is a known point of frustration for users searching for identifiers, function names, or version strings that rely on these characters [1]. It is distinct from GitHub's separate Code Search functionality, which does support more advanced exact matching and special character handling [5][6]. For issue search, no native syntax currently guarantees that these special characters will be treated as part of an immutable, literal substring [1].
Citations:
- 1: Support for literal (quoted) search queries cli/cli#7573
- 2: https://docs.github.com/en/search-github/searching-on-github/searching-issues-and-pull-requests
- 3: https://docs.github.com/en/issues/tracking-your-work-with-issues/using-issues/filtering-and-searching-issues-and-pull-requests
- 4: https://stackoverflow.com/questions/64173152/how-search-github-issue-by-exact-match-like-double-quotes-for-google
- 5: https://github.com/github/docs/blob/main/content/search-github/github-code-search/understanding-github-code-search-syntax.md
- 6: https://stackoverflow.com/questions/26433561/how-to-search-on-github-to-get-exact-string-matches-including-special-character
🌐 Web query:
GitHub issue search quoted phrase punctuation dots hyphens exact match docs
💡 Result:
When searching for issues and pull requests on GitHub, the search engine does not support exact string matching for punctuation, including dots and hyphens [1][2]. While you can use quotation marks to group multi-word terms (e.g., "in progress") to ensure those specific words appear in the result [1][2], GitHub's standard search behavior treats punctuation marks as delimiters or ignores them entirely [3][4]. Consequently, a search for a phrase containing punctuation will typically return results that match the words themselves, regardless of the punctuation [3][4]. For searching source code on GitHub, the experience differs due to GitHub's specialized Code Search engine [5]. GitHub Code Search supports exact string matching by surrounding the string in quotes [5]. It also allows for more advanced pattern matching using regular expressions enclosed in forward slashes (e.g., /my-exact-string/) [5][6]. Unlike the standard issue/pull request search, this code-specific engine is designed to handle non-alphanumeric characters more effectively [5]. In summary: - For issues and pull requests: Use quotes to group words, but be aware that punctuation like hyphens and dots will not be treated as part of an exact string match [1][2][3]. - For code: Use quotes for exact string matching or regular expressions (between slashes) to include specific punctuation marks in your search criteria [5][6].
Citations:
- 1: https://docs.github.com/en/search-github/searching-on-github/searching-issues-and-pull-requests
- 2: https://docs.github.com/en/search-github/getting-started-with-searching-on-github/understanding-the-search-syntax
- 3: https://docs.github.com/en/search-github/searching-on-github/searching-code
- 4: https://docs.github.com/en/search-github/searching-on-github/searching-code.md
- 5: https://docs.github.com/en/search-github/github-code-search/understanding-github-code-search-syntax
- 6: https://github.com/github/docs/blob/main/content/search-github/github-code-search/understanding-github-code-search-syntax.md
Avoid relying on quoted title search for dedup
gh issue list --search "\"$engine $latest\" in:title" is brittle for versions and names containing punctuation like . or -. GitHub issue search treats those characters loosely, so this can miss an existing issue or match the wrong one. Use a structured marker or metadata check instead.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/engine-version-watch.yml around lines 46 - 68, Replace the
quoted title search in the dedup loop with a structured marker or metadata-based
check that uniquely identifies the engine and latest version, avoiding GitHub
search parsing of punctuation in names and versions. Update the `existing`
calculation while preserving the current behavior of excluding rows with an open
matching `engine-watch` issue.
| prompt: | | ||
| REPO: ${{ github.repository }} | ||
| FAILED NIGHTLY RUNS (last 24h, JSON): ${{ needs.collect.outputs.failures }} | ||
|
|
||
| Triage every run listed. For each: | ||
| 1. Fetch evidence: | ||
| `gh run view <run_id> --repo ${{ github.repository }} --log-failed | tail -300` | ||
| (fall back to `gh api .../jobs` for job names/conclusions if | ||
| logs are unavailable). Read the workflow file and the scripts | ||
| it calls when the failure points into them. | ||
| 2. Classify with the log lines as evidence: | ||
| - infra/flake: runner lost, OOM-killed runner, network timeout, | ||
| registry/HF download failure, upstream API 5xx | ||
| - regression: assertion/score threshold failures (bfcl/tau2 | ||
| score drops, benchmark deltas), build breaks in our code | ||
| - config/env: version conflicts, missing secrets, disk full | ||
| 3. Rerun policy (conservative): if AND ONLY IF the failure is a | ||
| clear infra/flake and the run's attempt == 1, rerun it once: | ||
| `gh run rerun <run_id> --repo ${{ github.repository }} --failed`. | ||
| Never rerun regressions or second attempts. | ||
| 4. Report per workflow in ONE rolling issue: | ||
| - Find it: gh issue list --repo ${{ github.repository }} | ||
| --state open --label nightly-triage | ||
| --search "\"<workflow name>\" in:title" | ||
| - If it exists, add a comment titled with today's date holding | ||
| the day's triage (classification, key log lines, action | ||
| taken, run links). If the same failure signature already | ||
| appears in recent comments, say so — recurring failures are | ||
| the signal these issues exist to surface. | ||
| - Else create it: | ||
| gh issue create --label nightly-triage | ||
| --title "[nightly-triage] <workflow name>" | ||
| --body <first day's triage> | ||
| Keep each day's entry short: verdict first, then evidence. | ||
| Do NOT edit repository files or open pull requests. |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win
Keep the Claude prompt declarative.
Replace embedded gh commands and shell pipelines with natural-language instructions describing the evidence to retrieve, rerun policy, and issue updates. Based on learnings, LLM prompt: fields should use natural-language guidance rather than shell-specific constructs.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/nightly-triage.yml around lines 94 - 128, Replace the
shell-specific commands and pipelines in the prompt block with declarative
natural-language instructions for fetching failed-run evidence, applying the
conservative rerun policy, and finding or updating the rolling triage issue.
Preserve the repository/workflow context, classification criteria, attempt
restrictions, recurring-signature reporting, concise format, and prohibition on
file edits or pull requests.
Source: Learnings
cdac414 to
0a82a2b
Compare
|
Addressed the review findings; all fixes are in the current head. Fixed
Skipped
Validation
|
0a82a2b to
2762327
Compare
Two scheduled Claude-driven workflows on the existing review-action rails: every 3 days a deterministic detector compares engine pins against upstream (PyPI, pypi.nvidia.com, tokenspeed main) and, deduped against open issues, Claude researches each update and files one upgrade issue per engine; daily after the bfcl/tau2 windows, failed nightly runs are collected and Claude classifies each failure, reruns clear first-attempt infra flakes once, and maintains one rolling triage issue per workflow. Claude only runs when there is something to do. Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
2762327 to
b6d4935
Compare
|
Pre-merge test-run results (via a temporary trigger, since new workflows can't be dispatched before they exist on the default branch):
Post-merge: |
There was a problem hiding this comment.
Actionable comments posted: 6
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/engine-version-watch.yml:
- Around line 143-154: Update the prompt text in the workflow so it
declaratively instructs the agent to research the tokenspeed commit comparison,
re-check the torch pin, verify the release artifact, and create an issue in the
current repository with the engine-watch and enhancement labels, specified
title, and researched body. Remove embedded gh command syntax, shell
continuations, and the <researched body> placeholder while preserving the
required lookup and issue fields as natural-language requirements.
- Around line 7-14: Remove the temporary push trigger and its branch filter from
the workflow’s on configuration, leaving only the schedule and workflow_dispatch
triggers in the engine-version watch workflow.
- Around line 72-78: Update the issue query in the engine-version watch workflow
to fetch all open engine-watch issues rather than limiting results to 200, so
deduplication checks the complete issue set. Preserve the existing title JSON
and jq filtering used by the pending loop.
In @.github/workflows/nightly-triage.yml:
- Around line 13-15: Remove the temporary push trigger under the workflow’s on
configuration, including the feat/engine-watch-nightly-triage branch filter.
Keep only the existing cron and manual workflow_dispatch triggers so Claude
triage cannot run on normal branch pushes.
- Line 161: Update the --allowedTools configuration in the nightly triage
workflow to remove WebFetch and WebSearch and replace unrestricted Bash with
allowlisted gh commands matching only those invoked by the triage prompt.
Preserve the other required tools and ensure no general shell commands remain
permitted.
In `@scripts/check_engine_versions.sh`:
- Around line 17-22: Update the require() validation to reject both empty values
and the literal string "null" emitted by jq -r for missing JSON fields. Preserve
the existing error message and exit behavior so invalid pin or upstream values
fail before creating update or issue records.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: a6d8737a-c93f-450f-b5fd-a58533f39024
📒 Files selected for processing (3)
.github/workflows/engine-version-watch.yml.github/workflows/nightly-triage.ymlscripts/check_engine_versions.sh
| titles=$(gh issue list --repo "$GITHUB_REPOSITORY" --state open \ | ||
| --label engine-watch --limit 200 --json title --jq '[.[].title]') | ||
| pending='[]' | ||
| while IFS= read -r row; do | ||
| engine=$(jq -r .engine <<<"$row") | ||
| latest=$(jq -r .latest <<<"$row") | ||
| existing=$(jq --arg t "$engine $latest" '[.[] | select(contains($t))] | length' <<<"$titles") |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Inspect the workflow around the cited lines
FILE=".github/workflows/engine-version-watch.yml"
wc -l "$FILE"
sed -n '1,160p' "$FILE"
# Locate all uses of `gh issue list` and related dedup logic
rg -n "gh issue list|engine-watch|pending='\\[\\]'" .github/workflows -SRepository: lightseekorg/smg
Length of output: 7832
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Read the gh CLI docs from the installed help text if available
gh issue list --help | sed -n '1,220p'Repository: lightseekorg/smg
Length of output: 2785
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Probe whether gh issue list defaults imply pagination limits or whether --limit is only a display cap
python3 - <<'PY'
print("noop")
PYRepository: lightseekorg/smg
Length of output: 159
Fetch all open engine-watch issues for deduping. --limit 200 only checks the first 200 matches, so an older open issue can be skipped and a duplicate created once the label grows past that cap.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/engine-version-watch.yml around lines 72 - 78, Update the
issue query in the engine-version watch workflow to fetch all open engine-watch
issues rather than limiting results to 200, so deduplication checks the complete
issue set. Preserve the existing title JSON and jq filtering used by the pending
loop.
| For tokenspeed, `current`/`latest` are commit SHAs — summarize | ||
| `gh api repos/lightseekorg/tokenspeed/compare/<current>...<latest>` | ||
| (commit subjects only) and note that the torch pin in | ||
| ci_install_tokenspeed.sh must be re-checked against upstream. | ||
| 3. Verify the release image/wheel actually exists (Docker Hub tag, | ||
| NGC wheel listing) before recommending it. | ||
|
|
||
| Then create the issue: | ||
| gh issue create --repo ${{ github.repository }} \ | ||
| --label engine-watch --label enhancement \ | ||
| --title "[engine-watch] <engine> <latest> available (pinned: <current>)" \ | ||
| --body <researched body> |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win
Keep the Claude prompt declarative.
The embedded gh api/gh issue create shell syntax and <researched body> placeholder can be interpreted as executable shell rather than instructions. Describe the required API lookup and issue fields in natural language.
Proposed fix
- `gh api repos/lightseekorg/tokenspeed/compare/<current>...<latest>`
- (commit subjects only) and note that the torch pin in
+ compare the current and latest commits through the GitHub API,
+ summarize commit subjects only, and note that the torch pin in
ci_install_tokenspeed.sh must be re-checked against upstream.
@@
- Then create the issue:
- gh issue create --repo ${{ github.repository }} \
- --label engine-watch --label enhancement \
- --title "[engine-watch] <engine> <latest> available (pinned: <current>)" \
- --body <researched body>
+ Then create an issue in the repository with the `engine-watch`
+ and `enhancement` labels. Its title must be:
+ `[engine-watch] <engine> <latest> available (pinned: <current>)`.Based on learnings, prompt: fields intended for an LLM should use natural-language instructions rather than shell-specific constructs.
📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| For tokenspeed, `current`/`latest` are commit SHAs — summarize | |
| `gh api repos/lightseekorg/tokenspeed/compare/<current>...<latest>` | |
| (commit subjects only) and note that the torch pin in | |
| ci_install_tokenspeed.sh must be re-checked against upstream. | |
| 3. Verify the release image/wheel actually exists (Docker Hub tag, | |
| NGC wheel listing) before recommending it. | |
| Then create the issue: | |
| gh issue create --repo ${{ github.repository }} \ | |
| --label engine-watch --label enhancement \ | |
| --title "[engine-watch] <engine> <latest> available (pinned: <current>)" \ | |
| --body <researched body> | |
| For tokenspeed, `current`/`latest` are commit SHAs — summarize | |
| compare the current and latest commits through the GitHub API, | |
| summarize commit subjects only, and note that the torch pin in | |
| ci_install_tokenspeed.sh must be re-checked against upstream. | |
| 3. Verify the release image/wheel actually exists (Docker Hub tag, | |
| NGC wheel listing) before recommending it. | |
| Then create an issue in the repository with the `engine-watch` | |
| and `enhancement` labels. Its title must be: | |
| `[engine-watch] <engine> <latest> available (pinned: <current>)`. |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/engine-version-watch.yml around lines 143 - 154, Update
the prompt text in the workflow so it declaratively instructs the agent to
research the tokenspeed commit comparison, re-check the torch pin, verify the
release artifact, and create an issue in the current repository with the
engine-watch and enhancement labels, specified title, and researched body.
Remove embedded gh command syntax, shell continuations, and the <researched
body> placeholder while preserving the required lookup and issue fields as
natural-language requirements.
Source: Learnings
| claude_args: | | ||
| --model claude-opus-4-6 | ||
| --max-turns 60 | ||
| --allowedTools "Read,Glob,Grep,Bash,WebFetch,WebSearch,TaskCreate,TaskUpdate,TaskGet" |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/usr/bin/env bash
set -euo pipefail
curl -fsSL \
https://raw.githubusercontent.com/anthropics/claude-code-action/v1/docs/configuration.md |
grep -nE 'Bash\(|allowedTools'Repository: lightseekorg/smg
Length of output: 871
🏁 Script executed:
#!/usr/bin/env bash
set -euo pipefail
# Show the relevant section of the workflow.
sed -n '120,190p' .github/workflows/nightly-triage.yml | cat -n
# Also show the job permissions block if present.
rg -n "permissions:|actions: write|issues: write|allowedTools|claude" .github/workflows/nightly-triage.ymlRepository: lightseekorg/smg
Length of output: 3483
Restrict Claude’s Bash access to the gh commands this job needs.
Bash is still unrestricted here while the job has actions: write and issues: write; drop WebFetch/WebSearch and allow only the specific gh commands used by the triage prompt.
Suggested restriction
- --allowedTools "Read,Glob,Grep,Bash,WebFetch,WebSearch,TaskCreate,TaskUpdate,TaskGet"
+ --allowedTools "Read,Glob,Grep,Bash(gh run view:*),Bash(gh run rerun:*),Bash(gh issue list:*),Bash(gh issue create:*),Bash(gh issue comment:*),Bash(gh api:*),TaskCreate,TaskUpdate,TaskGet"📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| --allowedTools "Read,Glob,Grep,Bash,WebFetch,WebSearch,TaskCreate,TaskUpdate,TaskGet" | |
| --allowedTools "Read,Glob,Grep,Bash(gh run view:*),Bash(gh run rerun:*),Bash(gh issue list:*),Bash(gh issue create:*),Bash(gh issue comment:*),Bash(gh api:*),TaskCreate,TaskUpdate,TaskGet" |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/nightly-triage.yml at line 161, Update the --allowedTools
configuration in the nightly triage workflow to remove WebFetch and WebSearch
and replace unrestricted Bash with allowlisted gh commands matching only those
invoked by the triage prompt. Preserve the other required tools and ensure no
general shell commands remain permitted.
| require() { # name value | ||
| if [ -z "$2" ]; then | ||
| echo "ERROR: could not determine $1 (pin or upstream format changed?)" >&2 | ||
| exit 1 | ||
| fi | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
Reject literal null values from upstream JSON.
jq -r emits null for a missing field, and this non-empty check accepts it. That can create a bogus TokenSpeed update/issue or corrupt a package update record instead of failing.
Proposed fix
require() { # name value
- if [ -z "$2" ]; then
+ if [ -z "$2" ] || [ "$2" = "null" ]; then
echo "ERROR: could not determine $1 (pin or upstream format changed?)" >&2
exit 1
fi
}📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| require() { # name value | |
| if [ -z "$2" ]; then | |
| echo "ERROR: could not determine $1 (pin or upstream format changed?)" >&2 | |
| exit 1 | |
| fi | |
| } | |
| require() { # name value | |
| if [ -z "$2" ] || [ "$2" = "null" ]; then | |
| echo "ERROR: could not determine $1 (pin or upstream format changed?)" >&2 | |
| exit 1 | |
| fi | |
| } |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@scripts/check_engine_versions.sh` around lines 17 - 22, Update the require()
validation to reject both empty values and the literal string "null" emitted by
jq -r for missing JSON fields. Preserve the existing error message and exit
behavior so invalid pin or upstream values fail before creating update or issue
records.
| while IFS= read -r row; do | ||
| engine=$(jq -r .engine <<<"$row") | ||
| latest=$(jq -r .latest <<<"$row") | ||
| existing=$(jq --arg t "$engine $latest" '[.[] | select(contains($t))] | length' <<<"$titles") |
There was a problem hiding this comment.
🔴 Important: jq contains does substring matching, which causes false dedup when a final release's version string is a prefix of an existing RC issue's version.
Concrete scenario: an open issue is titled [engine-watch] tensorrt-llm 1.3.0rc18 available (pinned: 1.2.0). When the stable 1.3.0 ships, the search term becomes "tensorrt-llm 1.3.0", and contains("tensorrt-llm 1.3.0") returns true against the RC title because "1.3.0rc18" contains "1.3.0" as a substring. The stable release notification is silently dropped.
Same class of bug for minor-version prefixes: "sglang 0.5.1" would match an open "sglang 0.5.16" issue.
Fix: construct and match the full expected title prefix instead of a bare substring:
| existing=$(jq --arg t "$engine $latest" '[.[] | select(contains($t))] | length' <<<"$titles") | |
| expected="[engine-watch] $engine $latest available" | |
| existing=$(jq --arg t "$expected" '[.[] | select(startswith($t))] | length' <<<"$titles") |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b6d4935788
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| env: | ||
| GH_TOKEN: ${{ github.token }} | ||
| run: | | ||
| bash scripts/check_engine_versions.sh | tee versions.jsonl |
There was a problem hiding this comment.
Fail the version check when the detector fails
When any upstream lookup in scripts/check_engine_versions.sh fails (for example a curl timeout, NVIDIA index format change, or gh token issue), the script exits non-zero but this pipeline returns tee's status. GitHub documents an unspecified Linux/macOS run shell as bash -e {0}, while only shell: bash adds -o pipefail (docs), so the scheduled watch can continue with updates=[] or a partial JSONL file instead of alerting maintainers. Add shell: bash/set -o pipefail or avoid piping the detector before relying on its result.
Useful? React with 👍 / 👎.
| If that issue already mentions this run_id in its body or | ||
| comments, the run was triaged on a previous day — skip it. |
There was a problem hiding this comment.
Track rerun attempts instead of only run IDs
When Claude reruns a clear flake and that second attempt also fails, GitHub keeps it as another attempt of the same workflow run (the collector already carries attempt, and GitHub shows prior attempts under the same run in its rerun docs). The next collection will therefore see the same run_id with attempt == 2, but these instructions tell Claude to skip solely because that run_id was mentioned in the first day's comment, so failed reruns are never triaged or reported; dedup by (run_id, attempt) or add an explicit follow-up after reruns complete.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/engine-version-watch.yml:
- Around line 109-120: Extract the duplicated “Install gh CLI” step from both
the detect and engine-version jobs into a local composite action. Move the
pinned GH_VERSION and GH_SHA256 values and installation logic into that action,
then replace both inline blocks with action invocations so future version
updates require a single edit.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: af5b1fd9-295e-41f2-a578-1d708ee5cdf6
📒 Files selected for processing (3)
.github/workflows/engine-version-watch.yml.github/workflows/nightly-triage.ymlscripts/check_engine_versions.sh
| - name: Install gh CLI | ||
| run: | | ||
| if ! command -v gh &>/dev/null; then | ||
| mkdir -p "$HOME/.local/bin" | ||
| GH_VERSION="2.74.0" | ||
| GH_SHA256="e55c9d49dc49c0b0fef0a9acd3510482fd9e27ff52ae80f8a6e838cd25b4cd89" | ||
| curl -fsSL --connect-timeout 10 --max-time 120 -o /tmp/gh.tgz \ | ||
| "https://github.com/cli/cli/releases/download/v${GH_VERSION}/gh_${GH_VERSION}_linux_amd64.tar.gz" | ||
| echo "${GH_SHA256} /tmp/gh.tgz" | sha256sum --check --quiet | ||
| tar xzf /tmp/gh.tgz --strip-components=2 -C "$HOME/.local/bin" "gh_${GH_VERSION}_linux_amd64/bin/gh" | ||
| echo "$HOME/.local/bin" >> "$GITHUB_PATH" | ||
| fi |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win
Duplicated gh CLI install block across both jobs.
Lines 109-120 are byte-identical to lines 35-46 in the detect job. Extracting this into a local composite action would centralize the pinned GH_VERSION/GH_SHA256 pair so future bumps only need one edit.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/engine-version-watch.yml around lines 109 - 120, Extract
the duplicated “Install gh CLI” step from both the detect and engine-version
jobs into a local composite action. Move the pinned GH_VERSION and GH_SHA256
values and installation logic into that action, then replace both inline blocks
with action invocations so future version updates require a single edit.
There was a problem hiding this comment.
Actionable comments posted: 5
♻️ Duplicate comments (4)
.github/workflows/nightly-triage.yml (1)
117-155: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valuePrompt still embeds
ghcommand sequences instead of natural-language instructions.This carries over the prior (unaddressed) review comment: the
prompt:block still specifies literalgh issue list,gh run view ... --log-failed,gh run rerun, andgh issue createinvocations rather than describing the evidence, rerun policy, and issue-update goal in natural language for Claude to reason about.Based on learnings,
prompt:fields for LLMs should express guidance in natural language rather than embedding shell/CLI-specific constructs.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/nightly-triage.yml around lines 117 - 155, The prompt block should provide natural-language guidance instead of prescribing literal gh CLI commands. Update the instructions around rolling-issue lookup, run-evidence collection, rerun decisions, and issue creation/comments to describe the required outcomes and constraints while retaining the repository context, run IDs, classification rules, conservative rerun policy, and no-file-edit restriction.scripts/check_engine_versions.sh (1)
17-22: 🗄️ Data Integrity & Integration | 🟠 MajorReject literal
nullfromjq -r.
require()only rejects empty strings. Missing.info.versionor.shavalues becomenull, which can produce invalid update records or false comparisons.Proposed fix
- if [ -z "$2" ]; then + if [ -z "$2" ] || [ "$2" = "null" ]; then🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@scripts/check_engine_versions.sh` around lines 17 - 22, Update require() to reject both empty values and the literal “null” emitted by jq -r when fields are missing, while preserving its existing error and exit behavior. Ensure missing version or SHA values cannot proceed to update-record creation or comparisons..github/workflows/engine-version-watch.yml (2)
141-152: 📐 Maintainability & Code Quality | 🔵 TrivialKeep the Claude prompt declarative.
The prompt still embeds
gh api,gh issue create, shell continuations, and a<researched body>placeholder. Express these as natural-language requirements so the model does not interpret shell syntax as part of the task.Based on learnings, LLM
prompt:fields should use natural-language instructions rather than shell-specific command syntax.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/engine-version-watch.yml around lines 141 - 152, Rewrite the Claude prompt text around the token comparison and issue creation steps as declarative natural-language requirements. Remove embedded gh api and gh issue create commands, shell continuation syntax, and the <researched body> placeholder; describe the required research, verification, and issue contents in prose while preserving the existing engine-watch workflow behavior.Source: Learnings
69-75: 🗄️ Data Integrity & Integration | 🟠 MajorMake issue deduplication complete and exact.
--limit 200can omit older open issues, whilecontains("$engine $latest")can match unrelated title fragments. Paginate the full openengine-watchset and compare a stable exact marker/title for the engine and release.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/engine-version-watch.yml around lines 69 - 75, Update the issue lookup in the workflow’s pending-issue loop to retrieve the complete open engine-watch issue set by paginating beyond the current 200-item limit, then replace the substring-based contains check with an exact comparison against the stable engine-and-release marker/title so unrelated issues cannot be treated as duplicates.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/engine-version-watch.yml:
- Around line 35-46: Update the “Install gh CLI” step to always install and
expose the verified gh release, removing the command -v gh bypass so both jobs
use the pinned GH_VERSION binary validated by GH_SHA256.
- Around line 161-166: Restrict shell access in the Claude invocation’s
claude_args by removing unrestricted Bash and allowing only the specific gh
command patterns required for this job, or move issue creation to a
deterministic workflow step. Preserve the existing GitHub and Anthropic
credentials while ensuring prompt-injected responses cannot execute arbitrary
repository commands.
In @.github/workflows/nightly-triage.yml:
- Around line 34-45: Extract the duplicated checksum-verified gh CLI
installation blocks from the workflow’s install step and the triage step into a
local composite action. Invoke that action from both locations, keeping the
pinned GH_VERSION and GH_SHA256 definitions centralized while preserving the
existing PATH setup and installation behavior.
- Around line 156-159: Restrict the claude_args allowedTools configuration in
the nightly triage job by removing unrestricted Bash, WebFetch, and WebSearch,
and grant only the narrowly required GitHub CLI operations for viewing and
rerunning runs plus listing, creating, and commenting on issues. Preserve the
existing model and turn-limit settings.
In `@scripts/check_engine_versions.sh`:
- Around line 24-28: The vmax function must compare versions using full PEP 440
ordering rather than the current rc-only GNU sort transformation. Replace its
comparison logic with a real PEP 440-compatible comparator that correctly orders
dev, alpha, rc, final, and epoch versions while returning the newer input.
---
Duplicate comments:
In @.github/workflows/engine-version-watch.yml:
- Around line 141-152: Rewrite the Claude prompt text around the token
comparison and issue creation steps as declarative natural-language
requirements. Remove embedded gh api and gh issue create commands, shell
continuation syntax, and the <researched body> placeholder; describe the
required research, verification, and issue contents in prose while preserving
the existing engine-watch workflow behavior.
- Around line 69-75: Update the issue lookup in the workflow’s pending-issue
loop to retrieve the complete open engine-watch issue set by paginating beyond
the current 200-item limit, then replace the substring-based contains check with
an exact comparison against the stable engine-and-release marker/title so
unrelated issues cannot be treated as duplicates.
In @.github/workflows/nightly-triage.yml:
- Around line 117-155: The prompt block should provide natural-language guidance
instead of prescribing literal gh CLI commands. Update the instructions around
rolling-issue lookup, run-evidence collection, rerun decisions, and issue
creation/comments to describe the required outcomes and constraints while
retaining the repository context, run IDs, classification rules, conservative
rerun policy, and no-file-edit restriction.
In `@scripts/check_engine_versions.sh`:
- Around line 17-22: Update require() to reject both empty values and the
literal “null” emitted by jq -r when fields are missing, while preserving its
existing error and exit behavior. Ensure missing version or SHA values cannot
proceed to update-record creation or comparisons.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: e46de35b-4d0e-41db-aeeb-6f578ceb8790
📒 Files selected for processing (3)
.github/workflows/engine-version-watch.yml.github/workflows/nightly-triage.ymlscripts/check_engine_versions.sh
| - name: Install gh CLI | ||
| run: | | ||
| if ! command -v gh &>/dev/null; then | ||
| mkdir -p "$HOME/.local/bin" | ||
| GH_VERSION="2.74.0" | ||
| GH_SHA256="e55c9d49dc49c0b0fef0a9acd3510482fd9e27ff52ae80f8a6e838cd25b4cd89" | ||
| curl -fsSL --connect-timeout 10 --max-time 120 -o /tmp/gh.tgz \ | ||
| "https://github.com/cli/cli/releases/download/v${GH_VERSION}/gh_${GH_VERSION}_linux_amd64.tar.gz" | ||
| echo "${GH_SHA256} /tmp/gh.tgz" | sha256sum --check --quiet | ||
| tar xzf /tmp/gh.tgz --strip-components=2 -C "$HOME/.local/bin" "gh_${GH_VERSION}_linux_amd64/bin/gh" | ||
| echo "$HOME/.local/bin" >> "$GITHUB_PATH" | ||
| fi |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
file=".github/workflows/engine-version-watch.yml"
echo "== outline =="
ast-grep outline "$file" --view expanded || true
echo
echo "== lines 1-170 =="
nl -ba "$file" | sed -n '1,170p'Repository: lightseekorg/smg
Length of output: 285
🏁 Script executed:
#!/bin/bash
set -euo pipefail
file=".github/workflows/engine-version-watch.yml"
echo "== lines 1-170 =="
sed -n '1,170p' "$file" | cat -n
echo
echo "== gh occurrences =="
rg -n '\bgh\b|GH_TOKEN|PATH|Install gh CLI|command -v gh' "$file"Repository: lightseekorg/smg
Length of output: 9753
Force the pinned gh binary
if ! command -v gh lets any preinstalled gh bypass the version/SHA pin, so the gh calls in both jobs can run against an untrusted binary. Install the verified tarball unconditionally, or check the existing binary’s version and provenance before skipping.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/engine-version-watch.yml around lines 35 - 46, Update the
“Install gh CLI” step to always install and expose the verified gh release,
removing the command -v gh bypass so both jobs use the pinned GH_VERSION binary
validated by GH_SHA256.
| claude_args: | | ||
| --model claude-opus-4-6 | ||
| --max-turns 40 | ||
| --allowedTools "Read,Glob,Grep,Bash,WebFetch,WebSearch,TaskCreate,TaskUpdate,TaskGet" | ||
| env: | ||
| GH_TOKEN: ${{ github.token }} |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Inspect the workflow around the referenced lines
sed -n '130,190p' .github/workflows/engine-version-watch.yml
# Find references to claude-code-action and any related docs/config in the repo
rg -n "claude-code-action|allowedTools|ANTHROPIC_API_KEY|GH_TOKEN|Bash,WebFetch|Do not edit|engine-version-watch" .github/workflows . -g '!**/node_modules/**' -g '!**/dist/**' -g '!**/build/**'Repository: lightseekorg/smg
Length of output: 8132
🌐 Web query:
claude-code-action@v1 allowedTools Bash WebFetch WebSearch command-scoped shell permissions GH_TOKEN ANTHROPIC_API_KEY documentation
💡 Result:
In the claude-code-action@v1 GitHub Action, WebSearch and WebFetch are disabled by default for security reasons [1][2][3]. To enable them, you must explicitly include them in the allowed_tools configuration [2][4]. The action logic is designed so that if these tools appear in your provided allowed_tools list, they are removed from the internal disallowed list, effectively overriding the default restriction [5][2][4]. Configuration Example: To allow these tools, update your workflow YAML: - uses: anthropics/claude-code-action@v1 with: allowed_tools: "WebSearch,WebFetch,Bash" anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }} Regarding permissions and security: 1. Environment Variables: Never hardcode your ANTHROPIC_API_KEY or tokens in workflow files [6]. Always use repository secrets and reference them via ${{ secrets.SECRET_NAME }} [6][7]. 2. Shell Permissions: Bash tool permissions can be restricted using command-pattern matching [8]. You can limit access to specific commands by specifying them in the allowed_tools string (e.g., allowed_tools: "Bash(git:*),Bash(npm run *)") [8][9]. 3. GitHub Tokens: If using allowed_non_write_users to bypass standard write permission requirements, it is strongly recommended to use a scoped github_token (e.g., ${{ secrets.GITHUB_TOKEN }}) rather than a personal access token, and to restrict allowed_tools to the minimum necessary for the workflow [6]. 4. Global Settings: Tools denied at any level (e.g., through managed settings or explicit deny rules) cannot be overridden by allowed_tools [10]. Documentation on specific tool behavior and security best practices can be found in the official Claude Code documentation [8][7] and the repository's security guide [6].
Citations:
- 1:
--allowedToolsinclaude_argsdoesn't prevent default disabling of WebSearch and WebFetch anthropics/claude-code-action#690 - 2: WebSearch and WebFetch Disallowed Tools Override anthropics/claude-code-action#49
- 3: https://deepwiki.com/anthropics/claude-code-action/5.3-tool-and-permission-configuration
- 4: fix: respect --allowedTools when building disallowed tools list anthropics/claude-code-action#1033
- 5: feat: allow user override of hardcoded disallowed tools anthropics/claude-code-action#71
- 6: https://github.com/anthropics/claude-code-action/blob/refs/heads/main/docs/security.md
- 7: https://code.claude.com/docs/en/github-actions?_rsc=4Vrbm42DZc3Y7r9j
- 8: https://code.claude.com/docs/en/tools
- 9: https://github.com/anthropics/claude-code-base-action?tab=readme-ov-file
- 10: https://code.claude.com/docs/en/permissions
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Show the workflow file with line numbers around the referenced section
nl -ba .github/workflows/engine-version-watch.yml | sed -n '145,180p'
# Locate any job-level env/permissions near this workflow
rg -n -C 3 "permissions:|ANTHROPIC_API_KEY|GH_TOKEN|allowedTools|claude_args" .github/workflows/engine-version-watch.ymlRepository: lightseekorg/smg
Length of output: 194
Restrict Claude’s shell access. Bash is still enabled alongside GH_TOKEN and ANTHROPIC_API_KEY, so a prompt-injected response can still drive arbitrary shell commands with repo write access. Use command-pattern Bash allowlists for only the GitHub CLI operations this job needs, or move issue creation into a deterministic workflow step.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/engine-version-watch.yml around lines 161 - 166, Restrict
shell access in the Claude invocation’s claude_args by removing unrestricted
Bash and allowing only the specific gh command patterns required for this job,
or move issue creation to a deterministic workflow step. Preserve the existing
GitHub and Anthropic credentials while ensuring prompt-injected responses cannot
execute arbitrary repository commands.
| - name: Install gh CLI | ||
| run: | | ||
| if ! command -v gh &>/dev/null; then | ||
| mkdir -p "$HOME/.local/bin" | ||
| GH_VERSION="2.74.0" | ||
| GH_SHA256="e55c9d49dc49c0b0fef0a9acd3510482fd9e27ff52ae80f8a6e838cd25b4cd89" | ||
| curl -fsSL --connect-timeout 10 --max-time 120 -o /tmp/gh.tgz \ | ||
| "https://github.com/cli/cli/releases/download/v${GH_VERSION}/gh_${GH_VERSION}_linux_amd64.tar.gz" | ||
| echo "${GH_SHA256} /tmp/gh.tgz" | sha256sum --check --quiet | ||
| tar xzf /tmp/gh.tgz --strip-components=2 -C "$HOME/.local/bin" "gh_${GH_VERSION}_linux_amd64/bin/gh" | ||
| echo "$HOME/.local/bin" >> "$GITHUB_PATH" | ||
| fi |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value
Duplicated gh CLI install block.
The checksum-verified gh install (lines 34-45) is repeated verbatim in triage (94-105). Consider extracting to a local composite action to keep the pinned version/SHA256 in one place.
Also applies to: 94-105
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/nightly-triage.yml around lines 34 - 45, Extract the
duplicated checksum-verified gh CLI installation blocks from the workflow’s
install step and the triage step into a local composite action. Invoke that
action from both locations, keeping the pinned GH_VERSION and GH_SHA256
definitions centralized while preserving the existing PATH setup and
installation behavior.
| claude_args: | | ||
| --model claude-opus-4-6 | ||
| --max-turns 60 | ||
| --allowedTools "Read,Glob,Grep,Bash,WebFetch,WebSearch,TaskCreate,TaskUpdate,TaskGet" |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | ⚡ Quick win
Claude retains unrestricted Bash plus WebFetch/WebSearch while holding actions:write + issues:write.
Carried over from the prior review: this job can rerun workflows and write issues, yet --allowedTools still grants full Bash and web access rather than an allowlist of the specific gh run view/rerun/issue list/create/comment commands the prompt actually needs.
Suggested restriction
- --allowedTools "Read,Glob,Grep,Bash,WebFetch,WebSearch,TaskCreate,TaskUpdate,TaskGet"
+ --allowedTools "Read,Glob,Grep,Bash(gh run view:*),Bash(gh run rerun:*),Bash(gh issue list:*),Bash(gh issue create:*),Bash(gh issue comment:*),Bash(gh api:*),TaskCreate,TaskUpdate,TaskGet"🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/nightly-triage.yml around lines 156 - 159, Restrict the
claude_args allowedTools configuration in the nightly triage job by removing
unrestricted Bash, WebFetch, and WebSearch, and grant only the narrowly required
GitHub CLI operations for viewing and rerunning runs plus listing, creating, and
commenting on issues. Preserve the existing model and turn-limit settings.
| # Newest of two versions under PEP 440 pre-release ordering: rcN sorts | ||
| # before its final release (GNU sort -V treats '~' as lowest). | ||
| vmax() { | ||
| printf '%s\n%s\n' "$1" "$2" | sed 's/rc/~rc/' | sort -V | tail -1 | sed 's/~rc/rc/' | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Inspect the referenced file and surrounding lines.
git ls-files scripts/check_engine_versions.sh
wc -l scripts/check_engine_versions.sh
cat -n scripts/check_engine_versions.sh | sed -n '1,120p'
# Probe whether the repo mentions the same version selection logic elsewhere.
rg -n "sort -V|rc/~rc|PEP 440|vmax\(\)" scripts . || true
# Run a small read-only Python probe comparing the current vmax logic on representative PEP 440-like inputs.
python3 - <<'PY'
from itertools import product
import subprocess, textwrap, os, tempfile, shlex, sys
def vmax(a, b):
script = f"printf '%s\\n%s\\n' {shlex.quote(a)} {shlex.quote(b)} | sed 's/rc/~rc/' | sort -V | tail -1 | sed 's/~rc/rc/'"
out = subprocess.check_output(["bash", "-lc", script], text=True).strip()
return out
cases = [
("1.0a1", "1.0"),
("1.0b1", "1.0a1"),
("1.0rc1", "1.0b1"),
("1.0.dev1", "1.0a1"),
("1.0.post1", "1.0"),
("1!1.0", "2.0"),
("1.0", "1.0+local"),
("1.0rc1", "1.0"),
("1.0.post1", "1.0rc1"),
]
try:
from packaging.version import Version
except Exception as e:
print("packaging unavailable:", e)
Version = None
for a,b in cases:
got = vmax(a,b)
py = None
if Version is not None:
py = a if Version(a) >= Version(b) else b
print(f"{a:12} vs {b:12} -> vmax={got:12} python={py}")
PYRepository: lightseekorg/smg
Length of output: 5063
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Focus on the exact file if it exists.
if [ -f scripts/check_engine_versions.sh ]; then
cat -n scripts/check_engine_versions.sh | sed -n '1,120p'
fi
python3 - <<'PY'
try:
from packaging.version import Version
except Exception as e:
print("packaging unavailable:", e)
raise SystemExit(0)
def vmax(a, b):
import subprocess, shlex
script = f"printf '%s\\n%s\\n' {shlex.quote(a)} {shlex.quote(b)} | sed 's/rc/~rc/' | sort -V | tail -1 | sed 's/~rc/rc/'"
return subprocess.check_output(["bash", "-lc", script], text=True).strip()
cases = [
("1.0a1", "1.0"),
("1.0b1", "1.0a1"),
("1.0rc1", "1.0b1"),
("1.0.dev1", "1.0a1"),
("1.0.post1", "1.0"),
("1!1.0", "2.0"),
("1.0", "1.0+local"),
("1.0rc1", "1.0"),
("1.0.post1", "1.0rc1"),
]
for a,b in cases:
got = vmax(a,b)
py = a if Version(a) >= Version(b) else b
print(f"{a:12} vs {b:12} -> vmax={got:12} python={py}")
PYRepository: lightseekorg/smg
Length of output: 3372
Use a real PEP 440 comparator. vmax() only rewrites rc before sort -V, so it still misorders alpha/dev and epoch versions (1.0a1 vs 1.0, 1.0.dev1 vs 1.0a1, 1!1.0 vs 2.0). TensorRT-LLM uses the same logic, so upstream releases can be classified incorrectly.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@scripts/check_engine_versions.sh` around lines 24 - 28, The vmax function
must compare versions using full PEP 440 ordering rather than the current
rc-only GNU sort transformation. Replace its comparison logic with a real PEP
440-compatible comparator that correctly orders dev, alpha, rc, final, and epoch
versions while returning the newer input.
Description
Problem
Engine upgrades (SGLang/vLLM/TensorRT-LLM/TokenSpeed) only happen when someone remembers to check upstream, and the nightly benchmark/eval workflows (bfcl, tau2, benchmark, docker builds) fail with nobody looking.
Solution
Two scheduled workflows built on the same rails as the existing Claude code-review action (
k8s-runner-cpupod-env API key,claude-opus-4-6, gh CLI), both structured so Claude only runs when there's something to act on — detection is deterministic and token-free.engine-version-watch.yml(every 3 days + manual):scripts/check_engine_versions.shreads the pins from the same files upgrade PRs edit (installer scripts, release-docker defaults, TokenSpeed REF) and compares against PyPI, pypi.nvidia.com, and tokenspeedmain. Verified live: it currently reports exactly the four known-behind engines.engine-watchissue for that version are dropped (dedup), so each release is reported once.nightly-triage.yml(daily 15:30 UTC, after the ~7h bfcl/tau2 windows + manual):--log-failed, classifies (regression vs infra/upstream flake vs config), reruns only clear first-attempt infra flakes exactly once, and appends the day's verdict to one rolling[nightly-triage]issue per workflow — recurring failure signatures surface instead of vanishing.Both are cron/dispatch-only (no fork-triggerable inputs reach the prompts), labels are created idempotently, and permissions are scoped (
issues: write, plusactions: writeonly for the triage rerun).Test Plan
scripts/check_engine_versions.shrun live: emits correct current/latest for all four engines against today's upstream state.gh run list/gh issue listJSON paths exercised manually.workflow_dispatcheach once to shake out the runner environment (I can babysit the first runs).Checklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspassesSummary by CodeRabbit