Skip to content

feat(skills): add nemoclaw-maintainer-ci-flake-triage skill - #3526

Closed
cjagwani wants to merge 2 commits into
mainfrom
ship-skill-ci-flake-triage
Closed

feat(skills): add nemoclaw-maintainer-ci-flake-triage skill#3526
cjagwani wants to merge 2 commits into
mainfrom
ship-skill-ci-flake-triage

Conversation

@cjagwani

@cjagwani cjagwani commented May 14, 2026

Copy link
Copy Markdown
Collaborator

Summary

Classifies a failing CI check as PR_CAUSED / PRE_EXISTING_FLAKE / INFRASTRUCTURE / UNCLEAR. Combines diff-overlap detection, recent main-history comparison, multi-rerun reproducibility analysis, and infrastructure-signature greps.

Behavior

  • Local-only by default — drafts only, never posts to GitHub.
  • Emits a JSON sidecar (/tmp/nemoclaw-skill-output-ci-flake-triage-<run_id>.json) for chaining with sibling maintainer skills in the suite.

Conformance audit — Claude Agent Skills best practices

This skill was audited against the official Skill authoring best practices before draft. Per-item evidence:

Core quality

Item Status Evidence
Description specific + key terms description is 540 chars (under 1024 cap), first word Classifies (third-person, per spec)
Description has WHAT + WHEN Explicit Use when… trigger phrase present in the description
SKILL.md body under 500 lines Currently 168 lines (34%)
Additional details in separate files Supporting files: MULTI-MODEL-TESTING.md
Progressive disclosure used appropriately Heavyweight content extracted to one-level-deep supporting files where applicable
No time-sensitive info No absolute month/year cutoffs ("before/after MONTH 20YY" patterns) — all references are anchored to events or commits
Consistent terminology Audited for variant spellings (open issue vs open-issue, skill vs Skill, etc.)
Examples are concrete 3 real PR/issue references in SKILL.md: #3409, #3498, #3501
File references one level deep All supporting files linked directly from SKILL.md, never nested-deeper
Workflows have clear steps Numbered steps with explicit halt/stop conditions

Code and scripts

Item Status Notes
Scripts solve problems vs punt This is a markdown-only skill — no executable scripts in scripts/
Error handling explicit "Halt conditions" section enumerates non-obvious failure modes
No voodoo constants Thresholds (e.g. --min-confidence 0.6, --top N) documented with rationale
No Windows-style paths All paths use forward slashes
Validation/verification steps Critical operations gated by per-rule preflights
Feedback loops Calibration log / audit log where applicable for iteration based on real outcomes

Testing

Item Status Evidence
≥3 evaluations evals/ contains 3 JSON scenarios following the docs' eval schema
Multi-model test plan MULTI-MODEL-TESTING.md — Haiku / Sonnet / Opus expectations, pass criteria per eval, known model-size risks
Tested across all 3 models Test plan documented, not yet executed. PR is draft for visibility; the team can run the eval suite during adoption review
Tested with real usage Skill exercised on the live NemoClaw queue during 2026-05 maintainer sessions; reference cases in SKILL.md

Frontmatter constraints (validated)

  • name: nemoclaw-maintainer-ci-flake-triage — under 64 chars, lowercase + hyphens, no reserved words ("anthropic" / "claude")
  • description: third-person verb-initial, under 1024 chars, no XML tags, includes explicit Use when… trigger

Notes for reviewers

Part of an 11-skill maintainer suite. Draft for visibility. The team's <10 open-PR policy means 6 are open and 5 are closed-but-branch-preserved; reopen via gh pr reopen <num> as slots free up.

🤖 Generated with Claude Code

Classifies a failing CI check as PR_CAUSED, PRE_EXISTING_FLAKE,
INFRASTRUCTURE, or UNCLEAR. Combines diff-overlap detection, recent
main-history comparison, multi-rerun reproducibility analysis, and
infrastructure-signature greps. Tracks flake history at
/tmp/flake-history.jsonl and auto-elevates chronic flakes (7+ hits
in 14 days) to INFRASTRUCTURE with a draft fix-up issue.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented May 14, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@github-actions

Copy link
Copy Markdown
Contributor

This repository limits contributors to 10 open pull requests. Please close or merge existing PRs before opening new ones.

@github-actions github-actions Bot closed this May 14, 2026
@github-actions

github-actions Bot commented May 14, 2026

Copy link
Copy Markdown
Contributor

E2E Advisor Recommendation

Required E2E: None
Optional E2E: None

Workflow run

Full advisor summary

E2E Recommendation Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required E2E

  • None. No NemoClaw E2E is recommended. The changes are limited to .agents skill documentation and eval JSON for maintainer CI flake triage, with no impact on runtime/user-facing NemoClaw behavior or security/deployment/sandbox paths.

Optional E2E

  • None.

New E2E recommendations

  • None.

Adds the following to satisfy the Claude Agent Skills best-practices
checklist (https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices):

- Three evaluation scenarios in evals/ following the docs' eval schema
- Multi-model test plan in MULTI-MODEL-TESTING.md (Haiku / Sonnet /
  Opus expectations, pass criteria, known risks)
- Terminology normalized to single canonical form
- Concrete reference cases (real-but-anonymized examples) where the
  prior SKILL.md was abstract
- Progressive-disclosure splits where SKILL.md was approaching the
  500-line soft limit (issue-autopilot, scope-issues)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>
@cjagwani cjagwani reopened this May 15, 2026
@coderabbitai

coderabbitai Bot commented May 15, 2026

Copy link
Copy Markdown
Contributor

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 4039bb06-1d36-4ec7-bb42-6fb0e28969e7

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ship-skill-ci-flake-triage

Comment @coderabbitai help to get the list of available commands and usage tips.

@jyaunches

Copy link
Copy Markdown
Contributor

Migrated this skill to the NemoClaw team-skills GitLab repo: https://gitlab-master.nvidia.com/jyaunches/nemoclaw-team-skills. Closing this NemoClaw PR because team skills now live there rather than in the NemoClaw repo.

@jyaunches jyaunches closed this May 15, 2026
@wscurran wscurran added the feature PR adds or expands user-visible functionality label Jun 8, 2026
@cv
cv deleted the ship-skill-ci-flake-triage branch June 28, 2026 00:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feature PR adds or expands user-visible functionality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants