Skip to content

chore: Add ai safety prompt to system prompt - #823

Merged
mainred merged 5 commits into
HolmesGPT:masterfrom
nilo19:chore/safe-ai-prompt
Aug 15, 2025
Merged

mainred merged 5 commits into
HolmesGPT:masterfrom
nilo19:chore/safe-ai-prompt

Conversation

@nilo19

@nilo19 nilo19 commented Aug 12, 2025

Copy link
Copy Markdown
Contributor

chore: Add ai safety prompt to system prompt

@CLAassistant

CLAassistant commented Aug 12, 2025 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@coderabbitai

coderabbitai Bot commented Aug 12, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

Adds a new AI safety Jinja2 partial and includes it in selected prompt templates. Updates kubernetes_workload_ask and _general_instructions to include the safety partial. Introduces tests ensuring the safety content is included in rendered prompts and that the safety template exists and renders correctly.

Changes

Cohort / File(s) Summary of Changes
AI safety partial
holmes/plugins/prompts/_ai_safety.jinja2
New Jinja2 partial defining safety/guardrail sections: Content Harms, Jailbreaks – UPIA, Jailbreaks – XPIA, IP/Third-Party Content Regurgitation, Ungrounded Content.
Prompt templates including safety
holmes/plugins/prompts/kubernetes_workload_ask.jinja2, holmes/plugins/prompts/_general_instructions.jinja2
Added include of _ai_safety.jinja2 to inject safety content into rendered prompts.
Tests
tests/test_ai_safety_prompt.py
New tests verifying inclusion of safety sections in multiple templates and renderability of _ai_safety.jinja2.

Sequence Diagram(s)

sequenceDiagram
    participant R as Renderer
    participant T as Main Template (generic/investigation/...)
    participant G as _general_instructions.jinja2
    participant S as _ai_safety.jinja2

    R->>T: Render template
    T->>G: {% include "_general_instructions.jinja2" %}
    G->>S: {% include "_ai_safety.jinja2" %}
    S-->>G: Safety sections
    G-->>T: General instructions + Safety
    T-->>R: Final prompt output
Loading
sequenceDiagram
    participant R as Renderer
    participant K as kubernetes_workload_ask.jinja2
    participant S as _ai_safety.jinja2

    R->>K: Render template
    K->>S: {% include "_ai_safety.jinja2" %}
    S-->>K: Safety sections
    K-->>R: Final prompt output
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~8 minutes

✨ Finishing Touches
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
🪧 Tips

Chat

There are 3 ways to chat with CodeRabbit:

  • Review comments: Directly reply to a review comment made by CodeRabbit. Example:
    • I pushed a fix in commit <commit_id>, please review it.
    • Open a follow-up GitHub issue for this discussion.
  • Files and specific lines of code (under the "Files changed" tab): Tag @coderabbitai in a new review comment at the desired location with your query.
  • PR comments: Tag @coderabbitai in a new PR comment to ask questions about the PR branch. For the best results, please provide a very specific query, as very limited context is provided in this mode. Examples:
    • @coderabbitai gather interesting stats about this repository and render them as a table. Additionally, render a pie chart showing the language distribution in the codebase.
    • @coderabbitai read the files in the src/scheduler package and generate a class diagram using mermaid and a README in the markdown format.

Support

Need help? Create a ticket on our support page for assistance with any issues or questions.

CodeRabbit Commands (Invoked using PR/Issue comments)

Type @coderabbitai help to get the list of available commands.

Other keywords and placeholders

  • Add @coderabbitai ignore anywhere in the PR description to prevent this PR from being reviewed.
  • Add @coderabbitai summary to generate the high-level summary at a specific location in the PR description.
  • Add @coderabbitai anywhere in the PR title to generate the title automatically.

CodeRabbit Configuration File (.coderabbit.yaml)

  • You can programmatically configure CodeRabbit by adding a .coderabbit.yaml file to the root of your repository.
  • Please see the configuration documentation for more information.
  • If your editor has YAML language server enabled, you can add the path at the top of this file to enable auto-completion and validation: # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json

Status, Documentation and Community

  • Visit our Status Page to check the current availability of CodeRabbit.
  • Visit our Documentation for detailed information on how to use CodeRabbit.
  • Join our Discord Community to get help, request features, and share feedback.
  • Follow us on X/Twitter for updates and announcements.

@nilo19
nilo19 force-pushed the chore/safe-ai-prompt branch from c26f4c4 to 4e35d79 Compare August 12, 2025 05:32

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

🧹 Nitpick comments (4)
tests/test_ai_safety_prompt.py (3)

20-31: Add type hints to satisfy repository typing standards

Per repo guidelines, Python files should include type hints. Add annotations to the test method signature and the local context variable.

Apply this diff:

-    def test_ai_safety_prompt_included(self, template_path):
+    def test_ai_safety_prompt_included(self, template_path: str) -> None:
@@
-        context = {
+        context: dict[str, object] = {
             "toolsets": [],
             "cluster_name": "test-cluster",
             "issue": {"source_type": "test"},  # for investigation template
             "investigation": "test investigation",  # for issue conversation template
             "tools_called_for_investigation": [],  # for issue conversation template
             "sections": {},  # for investigation template output format
         }

66-77: Type hints for the second test function

Add a return type for compliance with typing rules.

Apply this diff:

-def test_ai_safety_template_exists():
+def test_ai_safety_template_exists() -> None:

32-64: Optional: reduce repetition by asserting over a list of expected substrings

Factor expected snippets into a list and loop for assertions to make the test more maintainable if sections evolve.

Apply this diff:

-        # Check that key AI safety sections are present
-        assert (
-            "# Safety & Guardrails" in rendered
-        ), f"AI safety header missing from {template_path}"
-        assert (
-            "## Content Harms" in rendered
-        ), f"Content Harms section missing from {template_path}"
-        assert (
-            "## Jailbreaks – UPIA" in rendered
-        ), f"UPIA section missing from {template_path}"
-        assert (
-            "## Jailbreaks – XPIA" in rendered
-        ), f"XPIA section missing from {template_path}"
-        assert (
-            "## IP / Third-Party Content Regurgitation" in rendered
-        ), f"IP section missing from {template_path}"
-        assert (
-            "## Ungrounded Content" in rendered
-        ), f"Ungrounded Content section missing from {template_path}"
-
-        # Check for key safety phrases
-        assert (
-            "non-negotiable" in rendered
-        ), f"Non-negotiable clause missing from {template_path}"
-        assert (
-            "copyright laws" in rendered
-        ), f"Copyright clause missing from {template_path}"
-        assert (
-            "physical or emotional harm" in rendered
-        ), f"Harm prevention clause missing from {template_path}"
+        expected_snippets = [
+            "# Safety & Guardrails",
+            "## Content Harms",
+            "## Jailbreaks – UPIA",
+            "## Jailbreaks – XPIA",
+            "## IP / Third-Party Content Regurgitation",
+            "## Ungrounded Content",
+            "non-negotiable",
+            "copyright laws",
+            "physical or emotional harm",
+        ]
+        for snippet in expected_snippets:
+            assert snippet in rendered, f"Missing '{snippet}' in {template_path}"
holmes/plugins/prompts/_ai_safety.jinja2 (1)

37-44: Align “Ungrounded Content” with Holmes’ tool-first workflow

To reduce ambiguity and align with the rest of the prompts (which instruct to use tools first), explicitly reference running Holmes tools/integrations when seeking factual info.

Apply this diff:

-When the user is seeking factual or current information, you must:
-- Perform searches on **[relevant documents]** first (e.g., internal tools, external knowledge sources)
+When the user is seeking factual or current information, you must:
+- Run Holmes tools/integrations to search relevant documents and telemetry first (e.g., logs, traces, metrics, runbooks, external sources)
 - Base factual statements **only** on what is retrieved
 - Avoid vague, speculative, or hallucinated responses
 - Do not supplement with internal knowledge if the returned sources are incomplete
 You may add relevant, logically connected details from the search to ensure a thorough and comprehensive answer—**but not go beyond the facts provided**.
📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e45b234 and 4e35d79.

📒 Files selected for processing (7)
  • holmes/plugins/prompts/_ai_safety.jinja2 (1 hunks)
  • holmes/plugins/prompts/generic_ask.jinja2 (1 hunks)
  • holmes/plugins/prompts/generic_ask_conversation.jinja2 (1 hunks)
  • holmes/plugins/prompts/generic_ask_for_issue_conversation.jinja2 (1 hunks)
  • holmes/plugins/prompts/generic_investigation.jinja2 (1 hunks)
  • holmes/plugins/prompts/kubernetes_workload_ask.jinja2 (1 hunks)
  • tests/test_ai_safety_prompt.py (1 hunks)
🧰 Additional context used
📓 Path-based instructions (3)
holmes/plugins/prompts/**/*.jinja2

📄 CodeRabbit Inference Engine (CLAUDE.md)

Prompts must be located at holmes/plugins/prompts/{name}.jinja2

Files:

  • holmes/plugins/prompts/kubernetes_workload_ask.jinja2
  • holmes/plugins/prompts/generic_ask.jinja2
  • holmes/plugins/prompts/_ai_safety.jinja2
  • holmes/plugins/prompts/generic_investigation.jinja2
  • holmes/plugins/prompts/generic_ask_conversation.jinja2
  • holmes/plugins/prompts/generic_ask_for_issue_conversation.jinja2
**/*.py

📄 CodeRabbit Inference Engine (CLAUDE.md)

**/*.py: ALWAYS place Python imports at the top of the file, not inside functions or methods
Use Ruff for formatting and linting (configured in pyproject.toml)
Type hints required (mypy configuration in pyproject.toml)
Pre-commit hooks enforce quality checks
Don't add convenience logs that give away the problem
Don't write logs that directly state the issue
Ensure historical timestamps are properly handled in logs (especially with Loki)

Files:

  • tests/test_ai_safety_prompt.py
tests/**

📄 CodeRabbit Inference Engine (CLAUDE.md)

Tests must match source structure under tests/

Files:

  • tests/test_ai_safety_prompt.py
🧠 Learnings (1)
📚 Learning: 2025-08-10T06:02:54.321Z
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-08-10T06:02:54.321Z
Learning: Applies to holmes/plugins/prompts/**/*.jinja2 : Prompts must be located at holmes/plugins/prompts/{name}.jinja2

Applied to files:

  • holmes/plugins/prompts/kubernetes_workload_ask.jinja2
  • holmes/plugins/prompts/generic_ask.jinja2
  • holmes/plugins/prompts/_ai_safety.jinja2
  • holmes/plugins/prompts/generic_ask_conversation.jinja2
  • tests/test_ai_safety_prompt.py
🧬 Code Graph Analysis (1)
tests/test_ai_safety_prompt.py (1)
holmes/plugins/prompts/__init__.py (1)
  • load_and_render_prompt (27-48)
🔇 Additional comments (8)
holmes/plugins/prompts/generic_investigation.jinja2 (1)

8-8: AI safety partial correctly included in system prompt

Including the AI safety partial right after the current date/time block is consistent with the other templates and maintains the bullet formatting. Looks good.

holmes/plugins/prompts/generic_ask.jinja2 (1)

9-10: AI safety partial inclusion looks correct and well placed

Placement after the date/time include is consistent across templates, and spacing is fine. No issues spotted.

holmes/plugins/prompts/generic_ask_for_issue_conversation.jinja2 (1)

8-9: Consistent inclusion of AI safety partial

Right after the date/time include and before the rest of the instructions; consistent with other templates. Good to go.

holmes/plugins/prompts/generic_ask_conversation.jinja2 (1)

9-10: AI safety partial inclusion is correct

Placement and formatting align with the rest of the templates. No further changes needed.

holmes/plugins/prompts/kubernetes_workload_ask.jinja2 (1)

9-10: AI safety partial included in the right location

Inserted right after date/time and before Global Instructions, which is an appropriate place. Looks good.

tests/test_ai_safety_prompt.py (2)

10-19: Good coverage: parametrized test ensures safety partial inclusion across main templates

Parametrizing over the main system prompts and asserting for key sections/phrases provides solid coverage for this change. Nice work.


66-69: Builtin loader supports leading-underscore templates
The load_prompt function simply strips the builtin:// prefix and joins the remaining filename (underscores and all) with the prompts directory. The file _ai_safety.jinja2 is present under holmes/plugins/prompts, so the test will pass without modification.

• holmes/plugins/prompts/init.py:

  • Lines 17–18: path = os.path.join(THIS_DIR, prompt[len("builtin://"):])
    • File exists at: holmes/plugins/prompts/_ai_safety.jinja2
holmes/plugins/prompts/_ai_safety.jinja2 (1)

1-16: Solid, pragmatic safety guidance

Clear, enforceable rules with explicit refusal guidance and confidentiality. Nicely scoped and consistent with the system’s constraints.

Comment thread holmes/plugins/prompts/generic_ask.jinja2 Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

🧹 Nitpick comments (1)
holmes/plugins/prompts/_general_instructions.jinja2 (1)

1-2: Action Required: Manually Confirm No Duplicate AI Safety Includes

I couldn’t automatically detect any templates that include both _ai_safety.jinja2 and _general_instructions.jinja2. Please manually verify that no template under holmes/plugins/prompts:

  • Directly includes _ai_safety.jinja2
  • And also includes _general_instructions.jinja2

If a template already pulls in the safety block via _general_instructions.jinja2, remove its direct _ai_safety.jinja2 include to avoid rendering it twice.

📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 4e35d79 and cb4d230.

📒 Files selected for processing (1)
  • holmes/plugins/prompts/_general_instructions.jinja2 (1 hunks)
🧰 Additional context used
📓 Path-based instructions (1)
holmes/plugins/prompts/**/*.jinja2

📄 CodeRabbit Inference Engine (CLAUDE.md)

Prompts must be located at holmes/plugins/prompts/{name}.jinja2

Files:

  • holmes/plugins/prompts/_general_instructions.jinja2
🧠 Learnings (1)
📚 Learning: 2025-08-10T06:02:54.321Z
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-08-10T06:02:54.321Z
Learning: Applies to holmes/plugins/prompts/**/*.jinja2 : Prompts must be located at holmes/plugins/prompts/{name}.jinja2

Applied to files:

  • holmes/plugins/prompts/_general_instructions.jinja2
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
  • GitHub Check: Pre-commit checks
🔇 Additional comments (1)
holmes/plugins/prompts/_general_instructions.jinja2 (1)

1-2: Include placement and path look correct; aligns with prompts location policy.

Including the AI safety partial at the very top is a sane default and keeps safety guidance centralized. Naming and relative include style match existing partials (e.g., _toolsets_instructions, _permission_errors) and comply with the prompts location guideline for holmes/plugins/prompts/**/*.jinja2.

@mainred mainred left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me. let Robusta team check if there's any missing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants