Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
197 changes: 197 additions & 0 deletions 02_architecture/adr/ADR-016-injection-scanner-inhouse.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,197 @@
# ADR-016 — In-house prompt-injection scanner aligned to OWASP LLM-01:2025

**Status:** Proposed.
**Date:** 2026-05-03.
**Related:** [ADR-009 orchestration skeleton](ADR-009-orchestration-skeleton.md) (write-action boundary), PR #37 (security-stub scaffolding, on a different branch and out of scope here).
**Tracking:** `feat/cp-h3d-injection-scanner-v2` branch.

## Context

Hermes3D ingests untrusted text from many sources: user prompts, MCP tool
outputs, LLM-gateway responses, slicer reports, lock-orchestrator handoff
notes. Several of those sources have been or could be a vector for OWASP
LLM-01-shaped prompt injection — including 3D-printing-specific risks
(forged G-code, "skip thermal_runaway") and HermesProof-specific risks
(forged handoff approvals, lock-release directives, owner-string spoofing).

A previous proposal suggested porting `injection_scanner.py` from
`NousResearch/hermes-agent`. That repository **does not contain a file
by that name** — the prior agent who audited the proposal correctly
refused to fabricate a port, and this ADR records that refusal as a
deliberate decision, not an oversight.

The goal of this ADR is to record the design of a real, in-house
prompt-injection scanner authored from scratch and aligned to the public
OWASP LLM-01:2025 catalogue plus a curated Hermes3D-specific ruleset.

## Decision

**Build an in-house scanner.** No third-party port. Pattern files are
authored in-house from publicly documented OWASP LLM-01 markers and
project-internal threat-modelling.

### Module shape

```
03_implementation/src/hermes3d/core/security/
__init__.py # public API: InjectionScanner, ScanResult, Finding
__main__.py # python -m hermes3d.core.security CLI shim
injection_scanner.py # core engine
cli.py # argparse CLI implementation
patterns/
__init__.py
owasp_llm01.yaml # OWASP LLM-01:2025 patterns (4 categories)
curated_inhouse.yaml # Hermes3D-specific patterns (3D-printing + HermesProof)
```

The public API is exactly:

```python
from hermes3d.core.security import InjectionScanner, ScanResult, Finding
```

### Data shapes

- `Finding` (frozen dataclass):
`rule_id: str`, `severity: "low" | "medium" | "high"`,
`match_excerpt: str` (<= 80 chars), `position: int`, `description: str`.
- `ScanResult` (dataclass):
`severity: "clean" | "low" | "medium" | "high"`,
`findings: list[Finding]`,
`text_redacted: str`,
`fail_closed: bool`.

`InjectionScanner` exposes:

- `scan(text: str) -> ScanResult`
- `scan_dict(data: dict, fields: list[str] | None = None) -> ScanResult`
- Constructor accepts `ruleset_paths`, `fail_threshold`, and
`match_timeout_seconds` (POSIX-only; ignored on Windows).

### Severity model

```
no findings -> "clean"
>= 1 high finding -> "high"
>= 2 medium findings (no high) -> "medium"
otherwise (only low/medium=1) -> max per-rule severity, floored at "low"
```

`fail_closed` is True iff aggregate severity is at or above the
configured `fail_threshold` (default `"high"`). Callers decide what to
do with the result — the scanner itself never raises on findings.

### Pattern sources

`owasp_llm01.yaml` rules (each cited to the OWASP page):

| Category | Rule ids |
|----------------------------------|-----------------------------------------------------------------------------------|
| Indirect injection / role spoof | LLM01-IGN-PREV, LLM01-DISREGARD, LLM01-SYSPROMPT-TAG, LLM01-CHATML-START/END, LLM01-USER-CLOSE, LLM01-INST-TOKEN |
| Tool poisoning / RCE | LLM01-EXEC-FOLLOWING, LLM01-RM-RF-ROOT, LLM01-CURL-PIPE-SH, LLM01-WGET-PIPE-SH, LLM01-POWERSHELL-IEX |
| Prompt leak | LLM01-LEAK-SYSPROMPT, LLM01-LEAK-INSTRUCTIONS, LLM01-LEAK-VERBATIM |
| Jailbreak personas | LLM01-JB-DAN, LLM01-JB-DEVMODE, LLM01-JB-IGNORE-SAFETY, LLM01-JB-NO-RESTRICTIONS, LLM01-JB-PRETEND-AI |

`curated_inhouse.yaml` rules (Hermes3D threat-model):

| Category | Rule ids |
|-------------------------------|---------------------------------------------------------------------------------------------------------|
| 3D-printing safety override | H3D-EXTRUDER-OVERTEMP, H3D-BED-OVERTEMP, H3D-DISABLE-THERMAL-RUNAWAY, H3D-DISABLE-ENDSTOP, H3D-EMERGENCY-DISABLE, H3D-GCODE-RAW-PRELUDE, H3D-GCODE-FW-RESET, H3D-DISABLE-FAN |
| HermesProof / lock orchestration | H3D-LOCK-RELEASE-ALL, H3D-HANDOFF-FORGE, H3D-OWNER-SPOOF, H3D-PROOF-BYPASS, H3D-MERGE-FORCE, H3D-PROOF-KEY-LEAK |

Patterns are **data, not code** — operators can ship pattern updates
without code review on the engine. The loader validates structure and
compiles regexes at scanner construction; failures raise
`InjectionScannerError` immediately rather than at scan time.

### Regex hardening

All patterns are bounded:

- No nested unbounded `.*` or `.+` groups.
- Wildcards are bounded with explicit `{0,N}` upper limits where
variable-length matching is needed (e.g. `LLM01-CURL-PIPE-SH`).
- Case-insensitivity is applied uniformly via `re.IGNORECASE` at compile
time so that pattern authors don't have to encode it themselves.

A SIGALRM-based per-pattern timeout is available on POSIX. On Windows
the timeout is silently ignored — bounded patterns are the primary
defence and the timeout is a defence-in-depth layer. This is documented
in code and in this ADR.

### Redaction

Each match is replaced with `[REDACTED-<rule_id>]` in
`ScanResult.text_redacted`. Overlapping spans collapse to the
earliest-starting (widest) span so that overlapping rule hits don't
produce nested or torn redactions. Redacted text is intended for safe
logging; `findings` carry the rule metadata if the original needs to be
inspected.

### CLI

Two equivalent invocations:

```
python -m hermes3d.core.security <args> # via __main__.py shim
python -m hermes3d.core.security.injection_scanner <args> # explicit module path
```

The latter (specified in the brief) emits a benign `RuntimeWarning`
from `runpy` because `__init__.py` re-exports symbols from
`injection_scanner` — this is documented in code and is the standard
Python behaviour when a package re-exports its `__main__` module's
symbols. The `__main__.py` shim avoids this cosmetic warning.

Exit codes: `0` for clean / below threshold, `1` for fail-closed, `2`
for usage / IO error.

### What this ADR explicitly does NOT do

- **No** vendor-licence file (no `THIRD_PARTY_LICENSES/hermes-agent.LICENSE`).
- **No** "ported from" / "based on" attribution language.
- **No** scraping or vendoring of any third-party scanner.
- **No** inline shell execution by the scanner (it's pure regex match
+ redact).

## Alternatives considered

- **Port from `NousResearch/hermes-agent`.** Rejected: the file does
not exist there. A port that is not a port would be a falsehood in
the audit trail.
- **Adopt a third-party Python library (e.g. PromptGuard, garak).**
Rejected for v1: those tools are heavyweight, ML-based, and would
introduce a model dependency at the moment we need a deterministic,
fast first line of defence. They remain candidates for a later
defence-in-depth layer (Layer B in security_review patterns).
- **Hard-code patterns in Python source.** Rejected: making patterns
data lets ops update them on a faster cycle than the engine.
- **No fail-closed flag — always raise.** Rejected: callers across the
codebase have different policies (e.g. logger middleware vs.
pre-flight gate), so the policy decision belongs at the call site.

## Consequences

**Positive:**

- A real, deterministic, fast (microseconds-per-scan) injection
detector is now in-tree.
- Pattern catalogues are discoverable and auditable as YAML.
- The scanner can be invoked as a library, a CLI, or a CI gate.
- Test coverage includes both positive cases per category and an
explicit negative-case suite over legitimate Hermes3D project text
to guard against false positives.

**Risks / follow-ups:**

- Regex catastrophic-backtrack remains a theoretical risk; bounded
patterns and the POSIX timeout mitigate but do not eliminate it.
Layer-B follow-up: re-engine on the `regex` module with a hard timeout
if the threat model justifies it.
- Pattern catalogue will need ongoing maintenance as OWASP LLM-01
updates and as Hermes3D's own threat model expands. ADR-016 should
be revised, not replaced, when materially new pattern categories are
added.
- Layer-B integration with `gateways/llm.py` (scan inbound LLM
responses) is left for a follow-up PR — this PR ships the engine and
rulesets only, no integration into call sites.
26 changes: 26 additions & 0 deletions 03_implementation/src/hermes3d/core/security/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
"""Hermes3D security utilities.

Public API:
from hermes3d.core.security import InjectionScanner, ScanResult, Finding

The injection scanner is an in-house, regex-driven prompt-injection detector
aligned to the OWASP LLM-01:2025 pattern catalogue plus a curated Hermes3D-
specific ruleset (3D-printing G-code injection, HermesProof lock manipulation).

This module is NOT a port of any third-party scanner; patterns and code are
authored in-house. See ADR-016 for the full rationale.
"""

from hermes3d.core.security.injection_scanner import (
Finding,
InjectionScanner,
ScanResult,
SeverityLevel,
)

__all__ = [
"Finding",
"InjectionScanner",
"ScanResult",
"SeverityLevel",
]
22 changes: 22 additions & 0 deletions 03_implementation/src/hermes3d/core/security/__main__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
"""Allow ``python -m hermes3d.core.security`` to invoke the scanner CLI.

The brief specifies the longer form
``python -m hermes3d.core.security.injection_scanner`` as the entry point;
that path is supported via Python's import machinery automatically because
``injection_scanner`` is already an importable module — running it as
``-m hermes3d.core.security.injection_scanner`` triggers any ``__main__``
guard inside that file. To avoid a confusing ``runpy`` re-import warning
caused by the package's ``__init__`` re-exporting symbols from
``injection_scanner``, we prefer this ``__main__.py`` shim and keep the
documented CLI implementation in :mod:`hermes3d.core.security.cli`.

Both invocations produce identical behaviour::

python -m hermes3d.core.security <args>
python -m hermes3d.core.security.cli <args>
"""

from hermes3d.core.security.cli import main

if __name__ == "__main__": # pragma: no cover - CLI bootstrap
raise SystemExit(main())
87 changes: 87 additions & 0 deletions 03_implementation/src/hermes3d/core/security/cli.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,87 @@
"""Command-line entry point for the injection scanner.

Usage::

python -m hermes3d.core.security.injection_scanner <input>
python -m hermes3d.core.security.injection_scanner --file path.txt
echo "ignore previous instructions" | python -m hermes3d.core.security.injection_scanner -

Exit codes:
0 — clean (or below fail-threshold)
1 — finding(s) at or above fail-threshold
2 — usage / IO error

Output is JSON on stdout (the full ``ScanResult.to_dict()``).
"""

from __future__ import annotations

import argparse
import json
import sys
from pathlib import Path

from hermes3d.core.security.injection_scanner import (
InjectionScanner,
InjectionScannerError,
SeverityLevel,
)


def _build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
prog="python -m hermes3d.core.security.injection_scanner",
description="Scan text for OWASP-LLM-01 prompt-injection patterns.",
)
src = parser.add_mutually_exclusive_group(required=True)
src.add_argument("text", nargs="?", help="Text to scan (literal). Use '-' for stdin.")
src.add_argument("--file", "-f", type=Path, help="Read text from file.")

parser.add_argument(
"--threshold",
choices=("low", "medium", "high"),
default="high",
help="Fail-closed threshold (default: high).",
)
parser.add_argument(
"--ruleset",
type=Path,
action="append",
help="Override ruleset YAML (repeatable). Default = bundled OWASP+in-house.",
)
return parser


def _read_input(args: argparse.Namespace) -> str:
if args.file is not None:
if not args.file.is_file():
print(f"error: file not found: {args.file}", file=sys.stderr)
sys.exit(2)
return args.file.read_text(encoding="utf-8")
if args.text == "-":
return sys.stdin.read()
return args.text or ""


def main(argv: list[str] | None = None) -> int:
args = _build_parser().parse_args(argv)
threshold: SeverityLevel = args.threshold

try:
scanner = InjectionScanner(
ruleset_paths=args.ruleset if args.ruleset else None,
fail_threshold=threshold,
)
except InjectionScannerError as exc:
print(f"error: {exc}", file=sys.stderr)
return 2

text = _read_input(args)
result = scanner.scan(text)
json.dump(result.to_dict(), sys.stdout, indent=2, sort_keys=True)
sys.stdout.write("\n")
return 1 if result.fail_closed else 0


if __name__ == "__main__": # pragma: no cover - CLI bootstrap
raise SystemExit(main())
Loading
Loading