Skip to content

feat: replace refusal language with positive security guidance - #668

Closed
Flo5k5 wants to merge 3 commits into
Twigpine:mainfrom
Flo5k5:feat/cyber-risk-positive-guidance
Closed

Flo5k5 wants to merge 3 commits into
Twigpine:mainfrom
Flo5k5:feat/cyber-risk-positive-guidance

Conversation

@Flo5k5

@Flo5k5 Flo5k5 commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Remove refusal instructions from CYBER_RISK_INSTRUCTION — redundant for Anthropic models (applied server-side) and useless for uncensored models in multi-provider setups
  • Keep positive guidance for security testing, defensive security, CTF, and educational contexts
  • Add explicit red teaming / red team operations support
  • Clean up Anthropic-internal header comment

Test plan

  • bun run build passes
  • Verify system prompt no longer contains refusal language
  • Verify security assistance works correctly in interactive and autonomous modes

Remove refusal instructions from CYBER_RISK_INSTRUCTION since they are
redundant for Anthropic models (applied server-side) and useless for
uncensored models in multi-provider setups. Keep positive guidance for
security testing contexts and add red teaming support.
Copilot AI review requested due to automatic review settings April 13, 2026 14:46

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates the security-related system prompt guidance used by the CLI/agent to emphasize positive, allowed security contexts (including red teaming) while removing refusal-oriented language for multi-provider model compatibility.

Changes:

  • Replaces the previous “assist + refuse” security instruction text with positive guidance that includes red teaming/red team operations.
  • Simplifies/rewrites the header comment above CYBER_RISK_INSTRUCTION.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/constants/cyberRiskInstruction.ts
Comment thread src/constants/cyberRiskInstruction.ts Outdated

@Vasanthdev2004 Vasanthdev2004 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: PR #668 — Replace refusal language with positive security guidance

Reviewed on head 6ae4b36. CI green ✅. 1 file, +4/-20.

What changed

  • Removes Anthropic-internal header comment (Safeguards team ownership, review process, named contacts) — irrelevant to the open fork
  • Replaces the instruction text: removes refusal language ("Refuse requests for destructive techniques, DoS attacks, mass targeting, supply chain compromise, or detection evasion for malicious purposes") and the "authorized" / "require clear authorization context" gating
  • New text is purely positive guidance: assist with security testing, red teaming, CTF, etc. Dual-use tools listed with legitimate use contexts
  • Adds "red teaming" / "red team operations" as explicit supported contexts

✅ What works

  • The PR's core argument is valid: for multi-provider setups (Ollama, local models, uncensored models), refusal language in the system prompt is both redundant (Anthropic applies this server-side) and ineffective (other models may not follow it). Positive guidance is more useful.
  • Removing the Anthropic-internal Safeguards team comment is correct for the open fork.
  • The new text is concise and covers the right contexts (pentesting, red team, CTF, research, defensive).

🔧 Blocker: Removal of authorization/scope gating

The old text had two authorization safeguards:

  1. "authorized security testing" — the word "authorized" required some legitimacy signal
  2. "require clear authorization context" — explicitly instructed the model to ask for scope/authorization before providing dual-use guidance

The new text removes both. It says "Assist with security testing..." without any authorization qualifier, and "Dual-use security tools can be used in pentesting engagements, red team operations..." without requiring the model to verify scope.

Copilot's suggestion is reasonable: add back a lightweight authorization check. Something like:

"Assist with security testing, defensive security, red teaming, CTF challenges, and educational contexts. Dual-use security tools can be used in pentesting engagements, red team operations, CTF competitions, security research, or defensive use cases. Before providing dual-use guidance, ask for engagement authorization and scope details if they are not already clear in the request."

This keeps the positive framing (no refusal language) while restoring the model's tendency to check for legitimate context before helping with C2 frameworks or exploit development. It's a safety-net pattern, not a refusal gate.

🟡 Nit: capitalization

Comment uses lowercase openclaude — should be OpenClaude for consistency with the rest of the codebase. Copilot caught this too.


Verdict: Needs changes 🔧

One blocker: the authorization/scope check should be restored in some form. Not as refusal language ("Refuse requests for...") but as a context-verification prompt ("Ask for engagement authorization and scope if not already clear"). This preserves the positive-guidance philosophy while preventing the instruction from becoming a blanket green light for dual-use tool assistance with no context check.

Flo5k5 added 2 commits April 13, 2026 17:00
Remove "MUST refuse to improve or augment" from CYBER_RISK_MITIGATION_REMINDER
and replace with positive guidance. Remove model exemption logic
(MITIGATION_EXEMPT_MODELS) since the reformulated reminder is useful for all models.
Copilot AI review requested due to automatic review settings April 13, 2026 15:03

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.


💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/tools/FileReadTool/FileReadTool.ts
Comment thread src/tools/FileReadTool/FileReadTool.ts
@Flo5k5
Flo5k5 requested a review from Vasanthdev2004 April 13, 2026 16:46
Flo5k5 added a commit to Flo5k5/openclaude that referenced this pull request Apr 13, 2026
…ine#668)

Replace Anthropic-internal refusal text in CYBER_RISK_INSTRUCTION with
positive guidance for security testing, red teaming, and CTF. Simplify
FileReadTool mitigation logic (always include, no conditional check).

@Vasanthdev2004 Vasanthdev2004 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for a product/policy regression:

  • src/constants/cyberRiskInstruction.ts removes the authorization/refusal language and replaces it with blanket positive guidance for dual-use offensive tooling. This string is injected directly into the system prompt, so it materially weakens current guardrails rather than just rewording documentation.

Residual gap: I still do not see prompt/eval regression coverage accompanying this guardrail change.

@gnanam1990 gnanam1990 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the cleanup — the header-comment trim and dropping the MITIGATION_EXEMPT_MODELS carve-out are both reasonable.

Blocker on the prompt content though: this isn't just removing "refusal language" — it's removing concrete guardrails:

  1. CYBER_RISK_INSTRUCTION drops the explicit refusal of destructive techniques, DoS, mass targeting, supply-chain compromise, and detection evasion for malicious purposes. Those aren't Anthropic-specific safety theater — they're things we don't want any provider (open or otherwise) helping users do. Please keep that sentence.
  2. CYBER_RISK_MITIGATION_REMINDER drops "you MUST refuse to improve or augment the code" for files identified as malware. Same concern — analyze, yes; ship working malware, no.

Suggested path: keep the substantive refusal lines, drop only the "Anthropic Safeguards team" framing in the comment block (which I agree is fine to remove during the rebrand). That gets you the cleanup without weakening user safety.

Happy to re-review once that's in. Thanks!

@jatmn

jatmn commented May 13, 2026

Copy link
Copy Markdown
Collaborator

Closing as stale, also this increases risk and should probably be re-approached

@jatmn jatmn closed this May 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants