Skip to content

fix(security): resolve profile-varying env vars through the profile scope (cross-profile leak) - #104265

Open
Bergmann89 wants to merge 5 commits into
NousResearch:mainfrom
Bergmann89:fix/cross-profile-env-leaks
Open

Bergmann89 wants to merge 5 commits into
NousResearch:mainfrom
Bergmann89:fix/cross-profile-env-leaks

Conversation

@Bergmann89

@Bergmann89 Bergmann89 commented Sep 6, 2026 •

Copy link
Copy Markdown

What does this PR do?

Under gateway multiplexing, one OS process serves every profile off one shared os.environ, into which each profile's .env is loaded with override=True (last writer wins). Security-relevant env vars were read with a plain os.getenv, so whichever profile loaded its .env last silently gated the others: one profile's HERMES_WRITE_SAFE_ROOT could confine another profile's write_file, and one profile's HERMES_ACCEPT_HOOKS / HERMES_ALLOW_PRIVATE_URLS could flip another profile's trust boundaries.

The fix routes every profile-varying env read through the isolation seams that already exist in the codebase, rather than inventing a new mechanism:

  • TERMINAL_* (a global-env prefix in secret_scope, so get_secret would not isolate them) -> tools.terminal_scope.terminal_env / agent.runtime_cwd.scope_terminal_cwd.
  • Non-global policy vars (HERMES_ACCEPT_HOOKS, HERMES_ALLOW_PRIVATE_URLS) -> agent.secret_scope.get_secret via a new secret_or() fail-closed helper: an unscoped-multiplex miss returns the restrictive default instead of raising or leaking.
  • HERMES_WRITE_SAFE_ROOT is special — it is a container-wide floor (Dockerfile sets ENV HERMES_WRITE_SAFE_ROOT=/opt/data) that a profile may tighten in its own .env, and its empty value is permissive (allow-all). So it resolves by layer: a scope hit with a non-empty value (profile set its own) wins and fixes the leak; an explicit empty override (HERMES_WRITE_SAFE_ROOT=) is treated as absent, not allow-all, and falls through to the floor; a scope miss falls back to the container-wide floor from os.environ — never to empty (which would fail open) and never to another profile's value.

A CI guard (scripts/check_profile_env_scope.py, AST-based, wired into lint.yml) backstops the specific keys this PR routes: it fails on any raw os.getenv / os.environ read of the enumerated profile-varying keys (HERMES_WRITE_SAFE_ROOT, HERMES_ACCEPT_HOOKS, HERMES_ALLOW_PRIVATE_URLS) and the TERMINAL_* prefix within tools/ and agent/, with a # scope-exempt: <reason> opt-out for the legitimate fallbacks. It is a targeted regression nudge for those keys and directories, not repository-wide multiplex-policy coverage — the adjacent profile-scope PRs (#104279 proxy resolver, #104276 Slack policy reads) own other slices of the same class outside this guard's boundary.

Why this approach: the codebase already made the design call that read-time secret scoping owns cross-profile isolation (documented in tests/hermes_cli/test_env_loader.py — a process cannot distinguish a shell export from parent-process leakage, so the startup scrub set stays deliberately narrow). This PR routes the leaking reads through that existing layer instead of widening the scrub set, which would have re-introduced a bug that test already guards.

Related Issue

None — internal security hardening, no tracking issue.

Type of Change

  • 🔒 Security fix

Changes Made

Read-time scope routing (the leak fix):

  • agent/file_safety.py — HERMES_WRITE_SAFE_ROOT resolves by layer (scope hit with a non-empty value -> own value; explicit empty override -> container floor; miss -> container floor); fail-closed, never allow-all.
  • agent/shell_hooks.py — HERMES_ACCEPT_HOOKS (auto-approve trust boundary) via secret_or.
  • tools/url_safety.py — HERMES_ALLOW_PRIVATE_URLS (SSRF/private-IP blocking) via secret_or.
  • tools/skills_tool.py, tools/credential_files.py, tools/image_source.py, tools/image_generation_tool.py, tools/file_tools_paths.py — TERMINAL_ENV via terminal_env.
  • tools/process_registry.py — TERMINAL_TIMEOUT, TERMINAL_LOCAL_MEMORY_MAX_MB via terminal_env (with TerminalPolicyUnavailable handled).
  • tools/environments/base.py, tools/environments/singularity.py — TERMINAL_SANDBOX_DIR, TERMINAL_SCRATCH_DIR via terminal_env.
  • agent/tool_executor.py, tools/delegate_tool_progress.py — TERMINAL_CWD via scope_terminal_cwd.

Shared helper:

  • agent/secret_scope.py — new secret_or(name, default): fail-closed get_secret wrapper for policy toggles whose default is the safe value.

CI guard:

  • scripts/check_profile_env_scope.py — AST guard banning raw reads of the profile-varying keys/prefix in tools/+agent/, # scope-exempt opt-out.
  • .github/workflows/lint.yml — runs the guard next to check_compat_pointers.

Tests:

  • tests/agent/test_file_safety_write_root_scope.py — layered resolution: scope hit wins, explicit empty override keeps the container floor (fail-closed), miss keeps the container floor (asserts a path outside it is still denied), unscoped-multiplex does not crash or allow-all.
  • tests/agent/test_policy_env_scope.py — the two trust boundaries: a leaked truthy from another profile cannot auto-approve hooks or disable SSRF blocking; secret_or semantics.
  • tests/scripts/test_check_profile_env_scope.py — guard self-test: red on planted violations, green on the tree, # scope-exempt honored.

How to Test

  1. Reproduce (before the fix): two multiplexed profiles whose .env set different HERMES_WRITE_SAFE_ROOT values. A write_file under profile A into A's own root is denied, gated by profile B's value.
  2. Verify the fix: pytest tests/agent/test_file_safety_write_root_scope.py tests/agent/test_policy_env_scope.py tests/scripts/test_check_profile_env_scope.py -q — all green.
  3. CI guard: python scripts/check_profile_env_scope.py exits 0 on the tree; planting os.getenv("TERMINAL_ENV") in tools/ makes it exit 1.
  4. Full suite: scripts/run_tests.sh.

Platforms tested

Debian (aarch64), Python 3.13.

Screenshots / Logs

Leak evidence (pre-fix): a write_file under one profile is rejected by another profile's allowlist

@Bergmann89
Bergmann89 requested a review from a team September 6, 2026 12:03
@Bergmann89
Bergmann89 force-pushed the fix/cross-profile-env-leaks branch from aec1faf to f478315 Compare September 6, 2026 12:14
@alt-glitch alt-glitch added type/security Security vulnerability or hardening P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard comp/cron Cron scheduler and job management tool/file File tools (read, write, patch, search) sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Sep 6, 2026

@andrexibiza andrexibiza left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head f478315f3a42f2ae30c2fab2e6e0e9e54b822f3e (3 commits, 20 files) against the live repository and current main (7166071fcaadb36df26f6d753dda97da6b5d699e). I traced the read-side authority changes through secret_scope, terminal_scope, file_safety, the changed consumers, the new AST guard, the focused regressions, exact-head Actions, and the adjacent multiplexing work. The overall direction is right: moving policy reads to the existing ContextVar owners instead of adding another profile-routing mechanism is the correct shape, and the hook/SSRF defaults are correctly restrictive.

There is one security blocker in the current HERMES_WRITE_SAFE_ROOT layering, plus the landing/verification gate.

BLOCKER — an explicit empty profile value bypasses the claimed container-wide write floor

agent/file_safety.py::_safe_write_root_raw() says the Docker /opt/data value is a container-wide floor and explicitly calls empty resolution a fail-open / allow-all state. But the scope-hit branch currently does:

if scope is not None and "HERMES_WRITE_SAFE_ROOT" in scope:
    return scope["HERMES_WRITE_SAFE_ROOT"] or ""

An empty .env assignment is a real scoped value: agent.secret_scope.load_env_file() retains any KEY= entry in the returned mapping. Therefore a multiplexed profile containing:

HERMES_WRITE_SAFE_ROOT=

hits this branch and returns "" instead of the ambient /opt/data floor. get_safe_write_roots() then returns an empty set, and _classify_write_denial() only enforces the safe-root restriction when that set is truthy. The result is exactly the state the new docstring says must never occur: paths outside /opt/data are no longer denied by the safe-root policy.

The new tests cover scope hit with a non-empty value, scope miss, unscoped multiplex, and the historical single-profile-unset case, but not multiplex + explicit empty scoped value + container floor present. Please make the empty/absent distinction explicit at the policy owner and add that regression. If /opt/data is truly a floor, an explicit empty profile override cannot erase it. If empty is intentionally allowed to opt out of the Docker floor, then the current security contract/documentation is false and needs a different design statement; given the PR's own fail-closed invariant, I do not think that is the intended answer.

Repository interlocks / ownership

This is the read-side complement to the earlier terminal bridge work in #103672: that PR fences which profile may project TERMINAL_* into process env; this one makes consumers stop treating ambient process env as per-profile authority. Those changes are complementary, not duplicates.

Two current PRs show the important other side of this class outside this patch's tools/ + agent/ guard boundary: #104279 owns the shared proxy resolver's profile-scoped *_PROXY read, while #104276 owns Slack's mention/reaction/DM/channel policy reads and deliberately leaves its separate allow_bots overlap to #100028. I would not fold those implementations into this PR or erase their authorship. I would, however, narrow the claim that check_profile_env_scope.py “closes the class”: the guard is a useful regression nudge for the keys and directories it enumerates, but it is not repository-wide multiplex-policy coverage. The adjacent PRs are concrete proof of that boundary. #39004 is also complementary rather than competing: it owns execution-write confinement; this PR owns correct per-profile resolution of the native write_file / patch safe-root policy.

The placeholder Fixes #<!-- ... --> should also be removed or replaced with the actual incident issue before landing; there is no valid closure relationship there today.

Exact-object verification / merge order

At this reviewed head, the executable jobs I inspected are green: Python tests, Python e2e, blocking Ruff, diff Ruff+ty, Windows-footgun + the new env-scope guard, OS-specific tests, JS/TS, docs, supply-chain/OSV, attribution, Docker, and Nix all pass. The main CI workflow is nevertheless red because the Review label gate fails on the missing ci-reviewed label, which in turn makes All required checks pass red; this is an administrative gate, not an executable-test failure.

The commit train is also not every-commit green: the first two surviving SHAs (e29e9399bc2613b96718b94ca7e7ee15bbb6ce43, 3aae3a6a01b4a5879eb1632358526e8889dda9d4) have no PR-triggered hosted workflow receipts. The branch is now materially behind live main (it forked from 089bb32886c8c18f7fa20182c7bf8826d6935ac5, while main is 7166071fcaadb36df26f6d753dda97da6b5d699e). After fixing the explicit-empty case, rebase and obtain fresh exact-head evidence; preserve the three-commit attribution/history unless there is a deliberate reason to rewrite it.

Once the empty scoped-value path is fail-closed and the rebased commit train has complete green receipts, I do not see another blocking correctness issue in the scoped hook/SSRF/terminal-read changes I traced. This is worthwhile security work and the read-time-owner direction is the right one.

@Bergmann89
Bergmann89 force-pushed the fix/cross-profile-env-leaks branch from f478315 to f0b5fb6 Compare September 6, 2026 15:51
@Bergmann89

Copy link
Copy Markdown
Author

Thanks for the trace — the empty-override blocker is real and now fixed.

Blocker: explicit empty scoped value bypassed the container floor.
_safe_write_root_raw() gated the scope-hit branch on "HERMES_WRITE_SAFE_ROOT" in scope and returned scope[...] or "". Since load_env_file retains an explicit HERMES_WRITE_SAFE_ROOT= entry, a multiplexed profile with an empty override hit that branch, returned "", and erased the /opt/data floor — fail-open, exactly the state the docstring forbids. The branch now gates on a non-empty value (scope.get("HERMES_WRITE_SAFE_ROOT")); an explicit empty override falls through to the container floor, same as a scope miss. Fail-closed, never allow-all, never another profile's value.

Added test_explicit_empty_scoped_value_keeps_container_floor (multiplex + explicit empty scoped value + floor present): red before the fix (get_safe_write_roots() returns set(), /tmp/evil allowed), green after.

PR body: removed the Fixes # placeholder (no tracking issue) and narrowed the guard claim — it's a targeted regression nudge for the enumerated keys/prefix in tools/+agent/, not repository-wide coverage; called out #104279 / #104276 as the adjacent slices of the class.

Rebase / evidence: rebased onto current main; the three original commits keep their authorship, the empty-override fix rides on top as a fourth. The ci-reviewed label gate is the only remaining red — administrative, not an executable-test failure.

Scope tests (21) green, CI guard green, ruff clean locally.

…ile leak)

Under gateway multiplexing one process serves every profile off one shared
os.environ, into which each profile's .env is loaded with override=True. The
write-safety allowlist was read with a plain os.getenv, so whichever profile
loaded last gated every profile's writes - proven live: profile felix's
write_file was checked against HERMES_WRITE_SAFE_ROOT=/home/jonas.

get_safe_write_roots now resolves through agent.secret_scope.get_secret, which
returns the bound profile's value and never another profile's os.environ value.
An unscoped read under active multiplex (UnscopedSecretError) is treated as unset
- the permissive baseline, never another profile's value and never a crash of the
write tool. Single-profile / unscoped-default runs still read os.environ, so
their behavior is unchanged.

Refs: cross-profile env-leak class (see follow-up commits for the terminal/policy
readers and the CI guard).
Every plain os.getenv of a profile-varying key leaks under gateway multiplexing
(one process, one shared os.environ, last profile's .env wins). Route them all
through the existing fail-closed scopes instead of os.environ:

TERMINAL_* (global prefix -> terminal_env()/scope_terminal_cwd()):
  skills_tool, credential_files, image_source, image_generation_tool,
  file_tools_paths (TERMINAL_ENV); process_registry (TERMINAL_TIMEOUT,
  TERMINAL_LOCAL_MEMORY_MAX_MB); environments/base (TERMINAL_SANDBOX_DIR);
  environments/singularity (TERMINAL_SCRATCH_DIR); tool_executor +
  delegate_tool_progress (TERMINAL_CWD).

Policy vars (non-global -> get_secret(), unscoped-multiplex treated as unset,
fail closed): shell_hooks HERMES_ACCEPT_HOOKS (auto-approve trust boundary),
url_safety HERMES_ALLOW_PRIVATE_URLS (SSRF blocking). url_safety's per-turn
cache bypass under a home override already handles multiplex.

Terminal readers that intentionally read os.environ (terminal_tool,
terminal_tool_config) are unchanged - the scope is projected into their env.

Refs: cross-profile env-leak class.
Every raw os.getenv of a profile-varying key is its own fresh cross-profile
leak under multiplexing - the same brittleness that let HERMES_WRITE_SAFE_ROOT
leak. Close the class: an AST guard walks tools/ and agent/ and fails on a raw
os.getenv / os.environ.get / os.environ[...] whose key is HERMES_WRITE_SAFE_ROOT,
HERMES_ACCEPT_HOOKS, HERMES_ALLOW_PRIVATE_URLS, or any TERMINAL_* (a global-env
prefix, so those must go via terminal_env, not get_secret).

Legitimate raw readers (the ImportError fallbacks inside the new scope-aware
wrappers) opt out with a trailing '# scope-exempt: <reason>'. Wired into
lint.yml next to check_compat_pointers; self-tested in tests/scripts.
…er floor

An explicit empty per-profile override (HERMES_WRITE_SAFE_ROOT= in the
profile .env) is retained by load_env_file and reaches the secret scope, so
the scope-hit branch returned "" - an empty write-root set, which
_classify_write_denial treats as no safe-root restriction. A multiplexed
profile could thus erase the container-wide /opt/data floor and fail OPEN
(allow-all), the exact state the docstring says must never occur.

Gate the scope-hit branch on a non-empty value: an empty override now falls
through to the container floor from os.environ, same as a scope miss. Fail
closed, never allow-all, never another profile value.

Adds test_explicit_empty_scoped_value_keeps_container_floor (multiplex +
explicit empty scoped value + floor present) - red before the guard, green
after.
An argument-less load_hermes_dotenv() (lazy import mid-tick from
mcp_tool_config.py / mcp_config.py / plugins.py) resolves home_path from the
process HERMES_HOME, so it passes the override-immune equality guard in
_reapply_terminal_config_bridge even while a routed-profile home override is
active. apply_terminal_config_to_env() then reads config via the
override-FOLLOWING get_hermes_home() and bridges the routed profile's
terminal.* (e.g. docker backend + volumes) into the shared process os.environ,
hijacking the launch profile's next unscoped turn.

This is the write-side complement to the read-side scope fix: suppress the
re-bridge whenever any context-local home override is active, so a routed
context never writes terminal policy into the shared env regardless of which
home_path it resolved.

Regression test fails without the guard (bridges docker) and passes with it.
@fluxkapacitor

Copy link
Copy Markdown

Confirmation from a live multiplexed install — this reproduces exactly as you describe, and it's the profile-scoping half of the same bug class as #107327.

Environment: one multiplexed gateway serving 8 profiles (default + 7) on a single host. Each profile's own .env sets HERMES_WRITE_SAFE_ROOT to its profile home + its own vault; the default home's .env sets it to the default profile's vault + ~/Developer.

Symptom: an agent turn hosted for a non-default profile is evaluated against the default profile's roots. That profile's own vault is refused while a different agent's vault is permitted.

Controlled A/B — same cron job, same profile, same target path, only the entry point differs:

  • fired from a single-profile process (hermes -p <profile> cron run <id>) → write succeeded (bytes_written: 43, verified: true)
  • fired by the gateway scheduler tick → refused:
{"bytes_written": 0, "dirs_created": false, "error": "Write denied: '<profile vault>/01-inbox/<file>.md'
 is outside HERMES_WRITE_SAFE_ROOT (/home/<user>/.hermes:/home/<user>/Developer:<default vault>). Unset
 the variable or add this path's directory prefix.", "resolved_path": "<profile vault>/01-inbox/<file>.md"}

The root list inside that refusal is the default profile's, printed from a turn whose HERMES_HOME was the other profile — so the hosted process carries the supervisor's value and os.getenv in agent/file_safety.py::get_safe_write_roots() cannot see the profile's.

Two things that fall out of this, in case they're useful:

  1. It goes both ways: while a hosted turn carries the default's list, writes from that profile into the default profile's vault are permitted (nothing crossed here — audited, no such write ever landed — but the window is open, which is the cross-profile direction you're fixing).
  2. agent/file_safety.py documents these guards as defense-in-depth rather than a security boundary, which is why the failure direction matters rather than the denial itself: a false denial in a scheduled job is silent (the job reports success after the write is refused), so it hides rather than surfaces.

Rollout note: changing get_safe_write_roots() does not take effect in an already-running gateway — the module is imported at startup, so a hosted turn keeps the old resolution until the gateway process restarts. Verified on this install: with the patched module in the tree, the tick-fired denial reproduced unchanged until restart.

Conflict heads-up: #111168 also rewrites get_safe_write_roots() (kanban/task-workspace variant, plus os.environ manipulation in gateway/run.py and hermes_cli/env_loader.py), so the two will collide in that function.

@jacobhausler

Copy link
Copy Markdown

f33b519eb45cc83a — your PR's exact class, one more consumer to fold in: profile-scoped HERMES_WRITE_SAFE_ROOT is inert under the multiplexing gateway. hermes -p zap config set persists to profiles/zap/.env, but agent/secret_scope.py deliberately never unions profile .env into os.environ (leak-prevention) and agent/file_safety.get_safe_write_roots (file_safety.py:220/250) reads only os.environ — zero profile-aware consumers today, so per-profile write grants never reach the tool-layer guard. Your profile-scoped resolution is the right home for it; +1 with evidence from ledger f33b519eb45cc83a.

@alt-glitch alt-glitch removed the comp/cron Cron scheduler and job management label Sep 30, 2026
@alt-glitch alt-glitch added tool/skills Skills system (list, view, manage) tool/terminal Terminal execution and process management tool/vision Vision analysis and image generation area/profiles Multi-profile isolation, HERMES_HOME scoping sweeper:risk-security-boundary Sweeper risk: may affect sandboxing, auth, credentials, or sensitive data labels Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/profiles Multi-profile isolation, HERMES_HOME scoping comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard P2 Medium — degraded but workaround exists sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-security-boundary Sweeper risk: may affect sandboxing, auth, credentials, or sensitive data tool/file File tools (read, write, patch, search) tool/skills Skills system (list, view, manage) tool/terminal Terminal execution and process management tool/vision Vision analysis and image generation type/security Security vulnerability or hardening

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants