fix(cron): handle passthrough secrets in multiplexed dispatch - #107413
fangliquanflq wants to merge 4 commits into
Conversation
Related: #106050 addresses the same restart-safe multiplexed cron dispatch symptom (#107399) from the other side — installing the profile secret scope around the external worker handoff in |
andrexibiza
left a comment
There was a problem hiding this comment.
Reviewed the exact current object: head 53f459e40e2288788558c576463935510c0bec45, base/live main@67764dc0863349a384c16425e73ee8571f3a94b7 (1 ahead / 0 behind), both changed files, the resolver and subprocess-env call chain, agent/secret_scope.py, the cron handoff, #107399, the existing discussion, #106050, the broader profile-isolation work in #73051, and exact-head CI.
The production failure in #107399 is real, and I like that this patch tries to close the class at a shared choke point rather than whack-a-mole individual keys. I do think the shared resolver is the wrong place to relax the invariant, though: this turns missing profile authority into permission to consume ambient process authority. That is a security-boundary regression under multiplexing, not just a policy choice. I left the concrete P1 inline.
Required closure
Please keep resolve_passthrough_value() fail-closed when multiplexing is active and no profile scope exists. Then solve the process-injected passthrough use case with an authority-bearing boundary, for example one of these shapes:
- bind the actual firing profile at the spawn seam and project only values authorized for that profile; or
- if Hermes intentionally supports process-global secret passthrough, represent that as a distinct typed/operator capability rather than inferring global ownership from the existing
terminal.env_passthroughname allowlist.
The regression needs the opposite side of the shape, not only the positive unscoped case currently added: two profiles, an ambient SERVICE_TOKEN/BW_SESSION, multiplexing enabled, and an unscoped restart-safe child build for profile B. Prove that B cannot receive A/the launch process value merely because the name was allowlisted elsewhere. Then add the positive witness for the intended process-injected/global policy through the actual restart-safe cron spawn boundary. The existing TestTerminalIntegration case currently proves the behavior that needs an authority distinction; it does not prove cross-profile isolation.
Interlock / attribution
#107399 (zicochaos) is the direct defect provenance and its measured process-env use case should be preserved. #106050 (MichaelNewham) is a direct competing carrier for the same cron dispatch symptom from the other side: it installs the firing profile's secret scope around the external-worker handoff, preserving the current fail-closed resolver. It should not be erased as “duplicate”; its profile-binding contribution is exactly the missing authority edge here. Conversely, #106050 alone does not obviously satisfy #107399's intentionally process-injected BW_SESSION case when that key is absent from the profile scope, so the right consolidation has to preserve both requirements rather than blindly choosing either patch.
#73051 (franco314) is broader adjacent profile-isolation work and already identifies the other structural half relevant here: config-based passthrough caching must be profile-aware in an app-global multiplexer, and a missing profile credential must remain authoritative. It also touches this same env_passthrough surface, so any canonical carrier should absorb or reconcile that invariant explicitly rather than creating two divergent definitions of what the allowlist authorizes.
Merge-order implication: I would not land this shared resolver relaxation ahead of the authority decision. If it lands first, every existing/future unscoped multiplexed build_subprocess_env() caller gets the ambient fallback, including the non-cron child-spawn sites #107399 itself identified. A cron-local scope fix can make that one path safer afterward, but it cannot restore the global fail-closed invariant for those other callers.
Exact-head proof
The PR body names scripts/run_tests.sh tests/tools/test_env_passthrough.py, but the exact head has no hosted check-runs or commit statuses. The three substantive workflows are all action_required with zero jobs executed: CI 34488158529, Docker 34488157337, and Nix 34488157445. So this one-commit train is currently 0/1 commits proven hosted-green.
This is a worthwhile bug to close—the restart-safe failure data in #107399 is unusually useful, and the shared-call-chain diagnosis is good. The last step is making the fix carry the same profile authority that the rest of multiplexing is deliberately built around, rather than making absence of that authority silently acceptable.
| if current_secret_scope() is None: | ||
| if multiplex_active and fallback is not None and is_env_passthrough(name): | ||
| return fallback | ||
| return get_secret(name) if multiplex_active else fallback |
There was a problem hiding this comment.
P1 — Preserve the profile-authority boundary here. terminal.env_passthrough proves that a name is eligible for projection; it does not prove that the ambient process value belongs to the currently firing profile. In multiplex mode agent/secret_scope.py deliberately raises on exactly this state because os.environ can contain another/launch profile's credential, and this module's config allowlist is still cached once per process. With this branch, an unscoped profile-B child can therefore consume a process SERVICE_TOKEN merely because that name was registered/loaded elsewhere. Because this is the shared resolver, the relaxation also reaches other unscoped subprocess callers, not only cron. Please keep unscoped multiplex resolution fail-closed and bind an actual profile-scoped projection at the spawn boundary (or introduce a separately typed process-global-secret authority). Add an A/B negative regression through the real restart-safe spawn path so a launch/A process value cannot reach B on an allowlist-only proof.
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (NousResearch#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is NousResearch#107413 / NousResearch#106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (NousResearch#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test.
|
Addressed the review points on the current head:
This incorporates the profile-scoped handoff direction identified in PR #106050 while retaining issue #107399's process-injected passthrough behavior for the profile that owns the gateway process environment. PR #73051's current diff is limited to Desktop session-profile propagation and does not overlap this env-passthrough path. Verification:
The restart-safe regression now covers the positive profile-secret case, the negative A/B isolation case, the launch-profile process-injected case, and scope reconstruction inside the external worker. |
|
Landed on |
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (NousResearch#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is NousResearch#107413 / NousResearch#106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (NousResearch#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test.
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (NousResearch#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is NousResearch#107413 / NousResearch#106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (NousResearch#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test.
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (NousResearch#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is NousResearch#107413 / NousResearch#106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (NousResearch#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test.
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (NousResearch#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is NousResearch#107413 / NousResearch#106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (NousResearch#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test. (cherry picked from commit 2f87677)
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test. (cherry picked from commit 2f87677)
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (NousResearch#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is NousResearch#107413 / NousResearch#106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (NousResearch#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test. (cherry picked from commit 2f87677)
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test. (cherry picked from commit 2f87677) (cherry picked from commit dbede34) (cherry picked from commit a7ac197)
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test. (cherry picked from commit 2f87677) (cherry picked from commit dbede34) (cherry picked from commit a7ac197)
…nder multiplex semantics The desktop backend ticks EVERY local profile's cron store from one process — its own docstring says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared `os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed there after the tick, and a scope miss read the launch profile's tokens (#107692). Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py) is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam, left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation applies inside the routed fire with no per-site patching; the launch profile's own fires and the backend's turns keep single-profile semantics; marker and override both reach the pool worker via `copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through `is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970). Two consequences of suppressing the write are handled rather than left as regressions: - a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`; the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub / passthrough rules apply to those values and the parent process is never mutated; - plugin secret sources are discovered on the fire's first agent build, after the scope froze, and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern `_publish_env_value` already uses for `.env` writes under multiplex). And the profile's external secret sources are hydrated before the scope is frozen, the order gateway/run.py and the external cron worker already use. Tests pin each direction: the marker without the semantics before the scope, the semantics on and off exactly with it, the marker reaching a copy_context worker; the process's own profile staying single-profile; the restart-safe handoff's child env building without raising under a routed tick with a passthrough key registered; a real child process receiving the routed values while `os.environ` keeps the launch value; a source registered after the freeze reaching the fire through the real PluginManager refresh. Reverting any one direction fails a distinct test. (cherry picked from commit 2f87677) (cherry picked from commit dbede34) (cherry picked from commit a7ac197)
Summary
Testing
scripts/run_tests.sh tests/cron/test_restart_safe_worker.py -k test_launch_external_worker_uses_restart_safe_scope_and_acknowledgesscripts/run_tests.sh tests/tools/test_env_passthrough.py tests/cron/test_restart_safe_worker.pyscripts/run_tests.sh tests/tools/test_local_env_blocklist.py -k TestProfileScopedPassthroughRelated Issue
Closes #107399