Repository navigation
Add toolset to trigger intentional OOM kill #1225
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: master
Are you sure you want to change the base?
Changes from all commits
bdabdf0
199ee26
0058135
15f45c9
735d861
465745d
b96b092
f5ef36d
d63ee32
09dedea
50bfb77
7a5b967
08418f0
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,110 @@ | ||
| # Stress Testing Holmes with Intentional OOM Kills | ||
|
|
||
| > ⚠️ **Never enable this toolset in production.** It allocates ~30 GB of RAM on the Holmes host and is intended only for controlled, non-production stress tests. | ||
|
|
||
| This guide explains how to intentionally trigger OOM kills using the built-in OOM toolsets, how to enable them safely, and how to confirm they are available. | ||
|
|
||
| ## Available Toolsets | ||
|
|
||
| Holmes ships with two disabled-by-default toolsets for inducing an OOM kill: | ||
|
|
||
| - **`oom_kill` (Python)** – `trigger_oom_kill` allocates ~30 GB and sleeps for a configurable duration. | ||
| - **`oom_kill_bash` (YAML/bash)** – `trigger_oom_kill_bash` does the same via a bash-executed Python snippet and applies `ulimit -v 2097152` (2 GiB virtual memory cap) before allocation to reduce blast radius. | ||
|
|
||
| Both toolsets require the environment variable `ALLOW_HOLMES_OOMKILL_TOOLSET` to pass prerequisites and must be explicitly enabled in configuration. They are **disabled by default** and will not be loaded unless you opt in. | ||
|
|
||
| ## Enabling via Helm/ArgoCD (cluster install) | ||
|
|
||
| Use the correct values path for your deployment method. | ||
|
|
||
| === "Robusta Helm Chart (Holmes as subchart)" | ||
|
|
||
| 1. **Set the env guard** (required): | ||
| ```bash | ||
| argocd app set <APP_NAME> \ | ||
| --helm-set-string holmes.additionalEnvVars[0].name=ALLOW_HOLMES_OOMKILL_TOOLSET \ | ||
| --helm-set-string holmes.additionalEnvVars[0].value=true | ||
| ``` | ||
|
|
||
| 2. **Enable the toolsets**: | ||
| ```bash | ||
| argocd app set <APP_NAME> \ | ||
| --helm-set-string holmes.toolsets.oom_kill.enabled=true \ | ||
| --helm-set-string holmes.toolsets.oom_kill_bash.enabled=true | ||
| ``` | ||
|
|
||
| 3. **Sync to apply**: | ||
| ```bash | ||
| argocd app sync <APP_NAME> | ||
| ``` | ||
|
|
||
| 4. **Verify** (optional): | ||
| ```bash | ||
| kubectl -n <holmes-namespace> exec -it <holmes-pod> -- \ | ||
| cat /app/custom_toolset.yaml | ||
| # Expect oom_kill and oom_kill_bash present and enabled | ||
| ``` | ||
|
|
||
| === "Holmes Helm Chart (direct)" | ||
|
|
||
| 1. **Set the env guard** (required): | ||
| ```bash | ||
| argocd app set <APP_NAME> \ | ||
| --helm-set-string additionalEnvVars[0].name=ALLOW_HOLMES_OOMKILL_TOOLSET \ | ||
| --helm-set-string additionalEnvVars[0].value=true | ||
| ``` | ||
|
|
||
| 2. **Enable the toolsets**: | ||
| ```bash | ||
| argocd app set <APP_NAME> \ | ||
| --helm-set-string toolsets.oom_kill.enabled=true \ | ||
| --helm-set-string toolsets.oom_kill_bash.enabled=true | ||
| ``` | ||
|
|
||
| 3. **Sync to apply**: | ||
| ```bash | ||
| argocd app sync <APP_NAME> | ||
| ``` | ||
|
|
||
| 4. **Verify** (optional): | ||
| ```bash | ||
| kubectl -n <holmes-namespace> exec -it <holmes-pod> -- \ | ||
| cat /app/custom_toolset.yaml | ||
| # Expect oom_kill and oom_kill_bash present and enabled | ||
| ``` | ||
|
|
||
| ## Enabling in Local CLI Mode | ||
|
|
||
| Add to your local config (e.g., `config.yaml`) and set the env guard before running the CLI: | ||
|
|
||
| ```yaml | ||
| toolsets: | ||
| oom_kill: | ||
| enabled: true | ||
| oom_kill_bash: | ||
| enabled: true | ||
| ``` | ||
|
|
||
| Then run: | ||
| ```bash | ||
| export ALLOW_HOLMES_OOMKILL_TOOLSET=true | ||
| holmes --config ./config.yaml ... | ||
| ``` | ||
|
|
||
| Because both toolsets are disabled by default and gated by `ALLOW_HOLMES_OOMKILL_TOOLSET`, they will **not** be auto-enabled in local mode unless you explicitly enable them and set the env variable. | ||
|
|
||
| ## Using the Tools | ||
|
|
||
| - **Python toolset**: `trigger_oom_kill` (param: `hold_seconds`, default 300). | ||
| - **Bash toolset**: `trigger_oom_kill_bash` (param: `hold_seconds`, default 300). | ||
|
|
||
| Example invocation (conceptual): | ||
| ``` | ||
| trigger_oom_kill: allocate ~30GB and sleep for 120s | ||
| ``` | ||
|
|
||
| ## Safety Considerations | ||
|
|
||
| - Keep this toolset out of production environments. | ||
| - Ensure hosts have proper isolation; the process is expected to be OOM-killed. | ||
| - Consider running in a dedicated test cluster or namespace. | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,91 @@ | ||
| import time | ||
| from typing import Any, Dict | ||
|
|
||
| from holmes.core.tools import ( | ||
| CallablePrerequisite, | ||
| StructuredToolResult, | ||
| StructuredToolResultStatus, | ||
| Tool, | ||
| ToolInvokeContext, | ||
| ToolParameter, | ||
| Toolset, | ||
| ToolsetTag, | ||
| ToolsetEnvironmentPrerequisite, | ||
| ) | ||
|
|
||
|
|
||
| class TriggerOOMKill(Tool): | ||
| toolset: "OOMKillToolset" | ||
|
|
||
| def __init__(self, toolset: "OOMKillToolset"): | ||
| super().__init__( | ||
| name="trigger_oom_kill", | ||
| description=( | ||
| "Allocates approximately 30GB of memory on the Holmes host to provoke the " | ||
| "OOM killer. This is intended for stress testing only and will likely crash " | ||
| "the running process. No confirmation is required because this is meant for " | ||
| "automated stress scenarios." | ||
| ), | ||
| parameters={ | ||
| "hold_seconds": ToolParameter( | ||
| description=( | ||
| "How long to keep the memory allocated before exiting. Defaults to 300 seconds." | ||
| ), | ||
| type="integer", | ||
| required=False, | ||
| ), | ||
| }, | ||
| toolset=toolset, # type: ignore[call-arg] | ||
| ) | ||
|
|
||
| def _invoke(self, params: dict, context: ToolInvokeContext) -> StructuredToolResult: | ||
| hold_seconds = params.get("hold_seconds", 300) | ||
| if not isinstance(hold_seconds, int) or hold_seconds <= 0: | ||
| return StructuredToolResult( | ||
| status=StructuredToolResultStatus.ERROR, | ||
| error="hold_seconds must be a positive integer.", | ||
| params=params, | ||
| ) | ||
|
|
||
| size_bytes = 30 * 1024 * 1024 * 1024 | ||
| print( | ||
| f"Allocating {{size_bytes / 1024 / 1024 / 1024:.0f}} GB of memory to intentionally trigger OOM kill; sleeping for {hold_seconds}s" | ||
| ) | ||
| data = bytearray(size_bytes) # type: ignore | ||
| time.sleep({hold_seconds}) | ||
|
|
||
|
Comment on lines
+41
to
+56
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Critical syntax errors and missing return statement. The
Additionally, as noted in past reviews, 🔎 Proposed fix for syntax errors, missing return, and reliable OOM triggering def _invoke(self, params: dict, context: ToolInvokeContext) -> StructuredToolResult:
hold_seconds = params.get("hold_seconds", 300)
if not isinstance(hold_seconds, int) or hold_seconds <= 0:
return StructuredToolResult(
status=StructuredToolResultStatus.ERROR,
error="hold_seconds must be a positive integer.",
params=params,
)
size_bytes = 30 * 1024 * 1024 * 1024
print(
- f"Allocating {{size_bytes / 1024 / 1024 / 1024:.0f}} GB of memory to intentionally trigger OOM kill; sleeping for {hold_seconds}s"
+ f"Allocating {size_bytes / 1024 / 1024 / 1024:.0f} GB of memory to intentionally trigger OOM kill; sleeping for {hold_seconds}s"
)
data = bytearray(size_bytes) # type: ignore
- time.sleep({hold_seconds})
+ # Touch every page to force physical allocation
+ page_size = 4096
+ for i in range(0, size_bytes, page_size):
+ data[i] = 1
+ print("Memory allocation complete")
+ time.sleep(hold_seconds)
+
+ return StructuredToolResult(
+ status=StructuredToolResultStatus.SUCCESS,
+ data="OOM kill triggered successfully",
+ params=params,
+ )🧰 Tools🪛 Ruff (0.14.10)41-41: Unused method argument: (ARG002) 54-54: Local variable Remove assignment to unused variable (F841) 🤖 Prompt for AI Agents |
||
| def get_parameterized_one_liner(self, params: Dict[str, Any]) -> str: | ||
| hold_seconds = params.get("hold_seconds", 300) | ||
| return ( | ||
| "python - <<'PY' ... # allocates ~30GB and sleeps for " | ||
| f"{hold_seconds}s to trigger OOM" | ||
| ) | ||
|
|
||
|
|
||
| class OOMKillToolset(Toolset): | ||
| def __init__(self): | ||
| super().__init__( | ||
| name="oom_kill", | ||
| enabled=False, | ||
| description=( | ||
| "Dangerous toolset that intentionally exhausts memory on the Holmes host to trigger an OOM kill. " | ||
| "Use only in controlled stress tests." | ||
| ), | ||
| docs_url=None, | ||
| icon_url=None, | ||
| prerequisites=[ | ||
| ToolsetEnvironmentPrerequisite(env=["ALLOW_HOLMES_OOMKILL_TOOLSET"]), | ||
| CallablePrerequisite(callable=self.prerequisites_callable), | ||
| ], | ||
| tools=[TriggerOOMKill(self)], | ||
| experimental=True, | ||
| tags=[ToolsetTag.CORE], | ||
| is_default=False, | ||
| ) | ||
|
|
||
| def prerequisites_callable(self, config: dict[str, Any]) -> tuple[bool, str]: | ||
| # No special configuration is required for this toolset. | ||
| return True, "" | ||
|
|
||
| def get_example_config(self) -> Dict[str, Any]: | ||
| return {} | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,28 @@ | ||
| toolsets: | ||
| oom_kill_bash: | ||
| description: "DANGEROUS: intentionally allocate ~30GB on the Holmes host to trigger OOM kill for stress testing (no confirmation required)." | ||
| docs_url: "" | ||
| icon_url: "" | ||
| tags: | ||
| - core | ||
| prerequisites: | ||
| - env: | ||
| - ALLOW_HOLMES_OOMKILL_TOOLSET | ||
| additional_instructions: | | ||
| ⚠️ This toolset is intentionally destructive. Only use in controlled environments. | ||
| tools: | ||
| - name: "trigger_oom_kill_bash" | ||
| description: "Allocate ~30GB of memory and hold it for a period to provoke the OOM killer. No confirmation required; intended for automated stress tests." | ||
| command: | | ||
| python - <<'PY' | ||
| import time | ||
|
|
||
| hold_seconds = int("{{ hold_seconds|default(300) }}") | ||
| if hold_seconds <= 0: | ||
| raise SystemExit("hold_seconds must be positive.") | ||
|
|
||
| size_bytes = 30 * 1024 * 1024 * 1024 | ||
| print(f"Allocating {size_bytes / 1024 / 1024 / 1024:.0f} GB of memory to intentionally trigger OOM kill; sleeping for {hold_seconds}s") | ||
| data = bytearray(size_bytes) | ||
| time.sleep(hold_seconds) | ||
|
aantn marked this conversation as resolved.
|
||
| PY | ||
Uh oh!
There was an error while loading. Please reload this page.