Conversation
Remove refusal instructions from CYBER_RISK_INSTRUCTION since they are redundant for Anthropic models (applied server-side) and useless for uncensored models in multi-provider setups. Keep positive guidance for security testing contexts and add red teaming support.
There was a problem hiding this comment.
Pull request overview
Updates the security-related system prompt guidance used by the CLI/agent to emphasize positive, allowed security contexts (including red teaming) while removing refusal-oriented language for multi-provider model compatibility.
Changes:
- Replaces the previous “assist + refuse” security instruction text with positive guidance that includes red teaming/red team operations.
- Simplifies/rewrites the header comment above
CYBER_RISK_INSTRUCTION.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Vasanthdev2004
left a comment
There was a problem hiding this comment.
Review: PR #668 — Replace refusal language with positive security guidance
Reviewed on head 6ae4b36. CI green ✅. 1 file, +4/-20.
What changed
- Removes Anthropic-internal header comment (Safeguards team ownership, review process, named contacts) — irrelevant to the open fork
- Replaces the instruction text: removes refusal language ("Refuse requests for destructive techniques, DoS attacks, mass targeting, supply chain compromise, or detection evasion for malicious purposes") and the "authorized" / "require clear authorization context" gating
- New text is purely positive guidance: assist with security testing, red teaming, CTF, etc. Dual-use tools listed with legitimate use contexts
- Adds "red teaming" / "red team operations" as explicit supported contexts
✅ What works
- The PR's core argument is valid: for multi-provider setups (Ollama, local models, uncensored models), refusal language in the system prompt is both redundant (Anthropic applies this server-side) and ineffective (other models may not follow it). Positive guidance is more useful.
- Removing the Anthropic-internal Safeguards team comment is correct for the open fork.
- The new text is concise and covers the right contexts (pentesting, red team, CTF, research, defensive).
🔧 Blocker: Removal of authorization/scope gating
The old text had two authorization safeguards:
"authorized security testing"— the word "authorized" required some legitimacy signal"require clear authorization context"— explicitly instructed the model to ask for scope/authorization before providing dual-use guidance
The new text removes both. It says "Assist with security testing..." without any authorization qualifier, and "Dual-use security tools can be used in pentesting engagements, red team operations..." without requiring the model to verify scope.
Copilot's suggestion is reasonable: add back a lightweight authorization check. Something like:
"Assist with security testing, defensive security, red teaming, CTF challenges, and educational contexts. Dual-use security tools can be used in pentesting engagements, red team operations, CTF competitions, security research, or defensive use cases. Before providing dual-use guidance, ask for engagement authorization and scope details if they are not already clear in the request."
This keeps the positive framing (no refusal language) while restoring the model's tendency to check for legitimate context before helping with C2 frameworks or exploit development. It's a safety-net pattern, not a refusal gate.
🟡 Nit: capitalization
Comment uses lowercase openclaude — should be OpenClaude for consistency with the rest of the codebase. Copilot caught this too.
Verdict: Needs changes 🔧
One blocker: the authorization/scope check should be restored in some form. Not as refusal language ("Refuse requests for...") but as a context-verification prompt ("Ask for engagement authorization and scope if not already clear"). This preserves the positive-guidance philosophy while preventing the instruction from becoming a blanket green light for dual-use tool assistance with no context check.
Remove "MUST refuse to improve or augment" from CYBER_RISK_MITIGATION_REMINDER and replace with positive guidance. Remove model exemption logic (MITIGATION_EXEMPT_MODELS) since the reformulated reminder is useful for all models.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
…ine#668) Replace Anthropic-internal refusal text in CYBER_RISK_INSTRUCTION with positive guidance for security testing, red teaming, and CTF. Simplify FileReadTool mitigation logic (always include, no conditional check).
Vasanthdev2004
left a comment
There was a problem hiding this comment.
Requesting changes for a product/policy regression:
src/constants/cyberRiskInstruction.tsremoves the authorization/refusal language and replaces it with blanket positive guidance for dual-use offensive tooling. This string is injected directly into the system prompt, so it materially weakens current guardrails rather than just rewording documentation.
Residual gap: I still do not see prompt/eval regression coverage accompanying this guardrail change.
gnanam1990
left a comment
There was a problem hiding this comment.
Thanks for the cleanup — the header-comment trim and dropping the MITIGATION_EXEMPT_MODELS carve-out are both reasonable.
Blocker on the prompt content though: this isn't just removing "refusal language" — it's removing concrete guardrails:
CYBER_RISK_INSTRUCTIONdrops the explicit refusal of destructive techniques, DoS, mass targeting, supply-chain compromise, and detection evasion for malicious purposes. Those aren't Anthropic-specific safety theater — they're things we don't want any provider (open or otherwise) helping users do. Please keep that sentence.CYBER_RISK_MITIGATION_REMINDERdrops "you MUST refuse to improve or augment the code" for files identified as malware. Same concern — analyze, yes; ship working malware, no.
Suggested path: keep the substantive refusal lines, drop only the "Anthropic Safeguards team" framing in the comment block (which I agree is fine to remove during the rebrand). That gets you the cleanup without weakening user safety.
Happy to re-review once that's in. Thanks!
|
Closing as stale, also this increases risk and should probably be re-approached |
Summary
CYBER_RISK_INSTRUCTION— redundant for Anthropic models (applied server-side) and useless for uncensored models in multi-provider setupsTest plan
bun run buildpasses