Skip to content

feat: add grill-me skill — adversarial plan interview - #43694

Closed
rafaumeu wants to merge 2 commits into
NousResearch:mainfrom
rafaumeu:feat/grill-me-skill
Closed

rafaumeu wants to merge 2 commits into
NousResearch:mainfrom
rafaumeu:feat/grill-me-skill

Conversation

@rafaumeu

Copy link
Copy Markdown

What does this PR do?

Adds a new bundled skill: grill-me — an adversarial plan interview that stress-tests ideas through structured questioning BEFORE any code is written.

Inspired by Matt Pocock's grill-me concept for Claude Code and ChaseAI's Codex variant, adapted natively for Hermes with:

  • One question at a time with a recommendation for each
  • Codebase exploration — reads existing code instead of asking questions the code can answer
  • Four-phase structure: Understanding → Technical Decisions → Edge Cases → Synthesis
  • Language-adaptive — skill text is English, but the interview adapts to the user's language
  • Integrates with existing skills: suggests plan, subagent-driven-development, requesting-code-review as next steps

Why bundled (not optional-skills or hub)?

Plan validation is universally useful. Every developer benefits from stress-testing complex ideas before implementation. The skill has zero dependencies and fits naturally alongside plan, spike, and requesting-code-review in the software-development category.

Type of Change

  • 🎯 New skill (bundled or hub)

Changes Made

  • skills/software-development/grill-me/SKILL.md — new skill (136 lines, ~6.3k chars)

Verification

  • Frontmatter validated (name, description ≤1024 chars, version, author, license, metadata)
  • Description starts with "Use when ..." (peer-matched pattern)
  • related_skills references resolve in-repo (plan, requesting-code-review, subagent-driven-development, test-driven-development)
  • Structure follows peer convention: Overview → When to Use → body → Pitfalls → Verification Checklist
  • No external dependencies or tool requirements

Testing

python3 -c "
import yaml, re, pathlib
content = pathlib.Path(\"skills/software-development/grill-me/SKILL.md\").read_text()
assert content.startswith(\"---\")
m = re.search(r\"\n---\s*\n\", content[3:])
fm = yaml.safe_load(content[3:m.start()+3])
assert \"name\" in fm and \"description\" in fm
assert len(fm[\"description\"]) <= 1024
assert len(content) <= 100_000
print(\"ALL CHECKS PASSED\")
"
# → ALL CHECKS PASSED

@ricardocamiloconsir ricardocamiloconsir left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review — grill-me Skill

Revisão completa da skill grill-me conforme critérios de qualidade do projeto.


✅ Aprovações

1. Adequação da solução
A abordagem de entrevista adversarial pré-implementação é excelente e preenche uma lacuna real no workflow do Hermes. A distinção clara entre plan (documentar), grill-me (stress-testar) e requesting-code-review (revisar código) está bem articulada na seção "This skill vs related skills".

2. Estrutura e convenções

  • Frontmatter segue exatamente o padrão das skills bundled (spike, plan, requesting-code-review): name, description ≤1024 chars, version, author, license, platforms, metadata com tags e related_skills
  • Estrutura do corpo segue a convenção peer: Overview → When to Use → body → Pitfalls → Verification Checklist
  • Description começa com o padrão "Use when..." — correto

3. Dev-friendly

  • Nomes de seções claros e autoexplicativos
  • "Mandatory Rules" como lista numerada — fácil de seguir
  • "Common Pitfalls" com anti-patterns explícitos — muito útil pra evitar erros comuns
  • "Verification Checklist" com checkboxes — permite ao agent auto-validar cada execução

4. Reutilizabilidade
A skill é 100% genérica — não depende de stack, framework ou linguagem. Pode ser usada em qualquer projeto que o Hermes gerencie. As 4 fases (Understanding → Technical → Edge Cases → Synthesis) são aplicáveis a qualquer plano complexo.

5. Sem redundâncias

  • Não recria funcionalidade de plan (grill-me não escreve planos)
  • Não recria funcionalidade de requesting-code-review (grill-me não analisa código)
  • Não recria spike (grill-me não faz experimentos descartáveis)
  • Cada skill tem escopo bem definido

⚠️ Apontamentos e Sugestões de Melhoria

Finding 1 — related_skills referencia skill em optional-skills/ (Severidade: Baixa)
subagent-driven-development está em optional-skills/software-development/, não em skills/software-development/. Isso significa que nem todo usuário terá essa skill instalada. O related_skills aponta para ela como se fosse bundled, o que pode causar confusão se o Hermes tentar resolver a referência e não encontrar.

Sugestão: Considerar se a referência é aceitável como "soft reference" (skill existe no repositório, mas é opt-in), ou se deveria ser removida do related_skills do frontmatter, mantendo apenas na seção "Integration with Other Skills" do corpo.

Finding 2 — Description pode exceder o limite em runtime (Severidade: Baixa)
A description tem ~260 chars, bem dentro do limite de 1024. No entanto, o PR menciona que foi validado — ✅ confirmado, sem problema real aqui.

Finding 3 — Ausência de trigger pattern explícito para /grill-me (Severidade: Baixa)
O description menciona /grill-me como trigger, mas o Hermes não tem um sistema de slash commands automático baseado em description de skills. O trigger real depende do skill matching heurístico do agente. A descrição cobre bem os patterns de linguagem natural, mas seria útil documentar no PR se há necessidade de integração com o sistema de slash commands do Hermes CLI.

Finding 4 — Fase 2 "Technical Decisions" — alcance muito amplo (Severidade: Informacional)
O range de "4-8 questions" na Fase 2 é bastante elástico. Para planos muito complexos, 8 perguntas podem não ser suficientes; para planos simples, 4 pode ser excessivo.

Sugestão: Considerar reformular para "3-6 questions, expandir se o plano tiver múltiplas decisões arquiteturais independentes" — dá mais flexibilidade sem comprometer a estrutura.

Finding 5 — "Explore the codebase when possible" — sem fallback explícito (Severidade: Baixa)
Regra 3 diz para explorar o codebase com search_files, read_file, terminal. Mas não cobre o caso em que o grill acontece antes de qualquer codebase existir (ex: projeto greenfield). A regra poderia mencionar: "If no codebase exists yet, skip to asking the user directly."


📊 Veredicto

Critério Avaliação
Solução adequada ✅ Excelente
Padrões do projeto ✅ Conforme
Reutilizável ✅ 100% genérica
Dev-friendly ✅ Clara e bem estruturada
Sem redundâncias ✅ Sem overlap
Otimizada ✅ N/A (skill declarativa)
Interface/UX ✅ N/A (skill de workflow)
Responsiva ✅ N/A

Veredicto: APPROVE com sugestões menores.

Os 5 findings são todos de severidade baixa ou informacional. A skill está bem escrita, segue as convenções do projeto, preenche uma lacuna real no workflow, e não tem overlap com skills existentes. Os apontamentos são melhorias incrementais que podem ser endereçados em follow-ups sem bloquear o merge.

Parabéns ao @rafaumeu pela qualidade da implementação — a estrutura de 4 fases com Mandatory Rules explícitas e Common Pitfalls é um padrão que outras skills poderiam adotar.

@rafaumeu
rafaumeu force-pushed the feat/grill-me-skill branch from 990ca99 to 05f8aaf Compare June 10, 2026 20:00
@alt-glitch alt-glitch added type/feature New feature or request P3 Low — cosmetic, nice to have labels Jun 15, 2026

@teknium1 teknium1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for contributing a focused pre-implementation interview workflow. Current main does not contain grill-me, so the feature is still relevant, but this draft needs current skill-standard updates before salvage.

Problems

  • skills/software-development/grill-me/SKILL.md:3 has a 134-character description; AGENTS.md:888-900 requires one sentence of at most 60 characters.
  • skills/software-development/grill-me/SKILL.md:13-136 does not use the modern required skill outline in AGENTS.md:933-940, including the # <Skill> Skill title and Prerequisites / How to Run / Quick Reference / Procedure sections.
  • The diff adds only SKILL.md; AGENTS.md:948-950 requires tests/skills/test_<skill>_skill.py coverage for a new skill.
  • The linked #62874 discussion identifies this as part of an open grill-me cluster with #37639; consolidate the selected content rather than landing parallel variants.

Suggested changes

  • Tighten the description to the 60-character contract.
  • Recast the existing interview material into the current section order and add focused structural tests.
  • Coordinate the overlap with #37639/#62874 during salvage.

This is an automated hermes-sweeper review.

@@ -0,0 +1,136 @@
---
name: grill-me
description: "Adversarial plan interview: one question at a time, with recommendations, resolving the full decision tree before any code is written."

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Current hardline skill standards require a one-sentence description of at most 60 characters (AGENTS.md:888-900); this description is 134 characters. Please shorten it to a trigger-focused summary.

tags: [planning, adversarial, interview, decision-tree, pre-implementation, review, alignment]
related_skills: [plan, requesting-code-review, subagent-driven-development, test-driven-development]
---

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please reshape this draft to the required modern skill outline in AGENTS.md:933-940: # <Skill> Skill, then When to Use, Prerequisites, How to Run, Quick Reference, Procedure, Pitfalls, and Verification.

@teknium1 teknium1 added the sweeper:blast-contained Sweeper blast radius: contained — one narrow path / opt-in / few users label Jul 14, 2026
SquadOps AI and others added 2 commits July 14, 2026 23:12
New bundled skill that stress-tests plans through structured adversarial
questioning. One question at a time, each with a recommendation, resolving
the full decision tree before any code is written.

Four-phase structure: Understanding -> Technical Decisions -> Edge Cases -> Synthesis.
Integrates with plan, subagent-driven-development, and requesting-code-review.
@teknium1

Copy link
Copy Markdown
Collaborator

Merged via PR #80708 — your commits were cherry-picked onto current main with your authorship preserved in git log, and the skill now ships in the optional catalog (hermes skills install official/software-development/grill-me). Thanks @rafaumeu!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

P3 Low — cosmetic, nice to have sweeper:blast-contained Sweeper blast radius: contained — one narrow path / opt-in / few users type/feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants