Skip to content

feat(a2a): expose inbound message metadata to sessions - #2

Closed
RuniThomsen wants to merge 277 commits into
mainfrom
feat/a2a-inbound-metadata-session
Closed

RuniThomsen wants to merge 277 commits into
mainfrom
feat/a2a-inbound-metadata-session

Conversation

@RuniThomsen

Copy link
Copy Markdown
Collaborator

Summary

  • advertise the optional https://runi.services/a2a/ext/task/v1 extension in the A2A Agent Card
  • render inbound Message.metadata generically beside message Parts in the receiving Hermes session
  • preserve the existing untrusted-peer framing and text-only audit/persistence boundary

Verification

  • TDD RED: 3 targeted tests failed before implementation
  • scripts/run_tests.sh tests/plugins/test_a2a_plugin.py tests/plugins/test_a2a_phase23.py — 154 passed
  • python -m pytest tests/plugins/test_a2a_plugin.py::TestInboundRoundTrip::test_live_server_passes_message_metadata_into_agent_session -m integration -q — 1 passed
  • git diff --check — passed

Scope

This does not interpret extension keys or add a Runi-specific adapter path. Any JSON-object Message.metadata is serialized next to the Parts and passes through the same inbound safety frame. Outbound metadata remains covered separately by NousResearch#90396.

victor-kyriazakos and others added 30 commits August 20, 2026 19:56
… ON the read loop self-deadlocked the transport

Round 2 of the approval-turn stuck-stream hunt. Round 1 (interim-marked
acks) fixed the draft-hijack-by-matching path — live logs confirm the
absorption fallback no longer fires — but the freeze persisted because
of a second, deeper defect on the same codepath:

_consume_prompt_response executes ON the transport read loop (inbound
frame -> _handle_frame -> _inbound handler). The handler awaited
self.send() for its '✅ Approved once' ack — but send() blocks on an
outbound_result future that ONLY the read loop can resolve, and the
read loop is blocked inside this very handler. Guaranteed self-deadlock
for the full outbound timeout (30s) on EVERY button tap. While wedged,
everything on the transport starved: draft appends (the frozen stream
right after approving), sibling approval-card sends (timed out into
'possibly-delivered' — the observed double-approval ambiguity), and the
turn's seal (timed out ambiguous -> plain-send fallback -> duplicate
final). Log signature was the tell: card-send timeout at tap time, no
absorption INFO, no seal-failed WARNING, no suppression line.

Fix: _send_lifecycle_ack() — acks ride a background task with strong
ref retention; the handler returns immediately and the read loop keeps
consuming, so the ack's own result frame resolves normally. Applied to
all six lifecycle sends (approval ack, slash-confirm ack + result text,
clarify acks, expiry notice). Acks are cosmetic by contract; failure
logs at debug and never breaks the reader.

Tests: new deadlock-shape test (gated transport send; handler must
return within 1s and the ack must still egress afterwards — RED on the
awaited version via TimeoutError at the exact deadlock), prior 3 tests
green with a yield for the background task. Targeted sweep 203/203.
…ck task

test_expired_own_prompt_notifies_instead_of_unknown_command asserted the
expiry notice synchronously after _consume_prompt_response returned. The
notice now rides a background task (read-loop self-deadlock fix: awaiting
a send from the prompt_response handler blocks the very read loop that
resolves the send's result future), so the test yields one tick before
asserting egress. Behavior contract unchanged: exactly one notice, no
chat dispatch.
…dling as approvals

The clarify caller treated a send-scheduling timeout as a definitive
failure: clear_session() + '[clarify prompt could not be delivered]'.
Same physics as the approval card fixed earlier in this PR — the card
may well have posted with a late connector ack — so the teardown ran
out from under a rendered clarify card and the user's answer resolved
nothing.

New _clarify_send_disposition() routes the outcome through
_approval_send_outcome: only a DEFINITIVE failure (error result,
non-timeout exception, no future) clears the registration and aborts;
ambiguous logs a warning and falls through to wait_for_response, whose
existing bounded wait already handles the truly-lost-card case. This
makes the boundary rule stated in the ambiguity test docstring hold
for the clarify lane, not just approvals.

Tests: 5 disposition tests mirroring the approval suite, including
clear_session-not-called on timeout. Mutation-verified: folding
ambiguous into the failed branch sends
test_timeout_keeps_registration_armed_and_proceeds_to_wait red.
Relay + gateway sweep 283/283.
…ify caller path gets contract tests

Two review nits on the clarify ambiguity fix:

- _approval_send_outcome swallowed the failure detail the old inline
  callers logged (scheduling exception text / SendResult error). Both
  the approval and clarify lanes now share one warning with the detail,
  logged in the classifier itself.

- The clarify tests pinned the disposition helper but nothing proved
  the ambiguous branch actually reaches the bounded wait. The
  send-then-wait sequence is extracted to _clarify_send_then_wait (the
  callback closure now just binds context onto it) and the suite gains
  caller-path tests: ambiguous/sent -> wait_for_response with the
  generated clarify_id and configured timeout; definitive failure ->
  sentinel without waiting; no-response timeout sentinel preserved;
  plus caplog assertions that failed sends log their detail.

Relay + gateway sweep 289/289.
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.

- cron/jobs.py: new optional job field, validated at the storage choke
  point against the canonical grammar via the shared
  hermes_constants.parse_reasoning_effort (spelling-only; capability
  clamping stays owned by the provider transports at send time, same as
  config-set effort). Empty string clears on update; invalid values
  raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
  > agent.reasoning_overrides > agent.reasoning_effort at fire time,
  after the auth-fallback model swap (the pin is model-independent by
  design). A stored value that no longer parses warns and falls back to
  config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
  (create and update), conditional key in _format_job, schema documents
  grammar/precedence/transport clamping/clear semantics. Agent-settable,
  unlike model/provider pins: it cannot redirect spend to a different
  model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.

Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
…ol schema

Standing policy: models do not make model-configuration decisions (the
only exception is user-defined profile selection in Bot Mode/kanban).
The per-job reasoning pin stays fully functional via
`hermes cron create/edit --reasoning-effort` and the job store; the
cronjob tool still SURFACES the pin in listings but cannot set it.
A schema-absence test pins the policy.
…odel dispatch still drops it

The CLI (hermes cron create/edit) routes through cronjob(); removing the
parameter outright broke that lane (CI slices 6/9). The parameter is back
on the function, but CRONJOB_SCHEMA and the registry handler still omit
it — same pattern as the intentional model/provider/base_url omission.
New test proves a hallucinated reasoning_effort arg through the model
dispatch is dropped.
… socket stalls

A single urlopen(timeout=5) TimeoutError from the PS runspace listener
failed the test on a loaded runner (run 32440286339) even though the
listener recovered moments later — a transient stall is not the hang
this test guards. /progress sampling now retries until a deadline
(only a persistently unresponsive listener fails), the self-test hold
grows 10s -> 30s so retry time cannot push sampling past the held
stage, and the exit wait gets matching headroom.
Reuse the built-in store predicate during agent initialization and evaluate the config-backed memory tool check immediately after edits instead of applying the generic external-probe TTL.
Use Hermes's shared truthy-value parser so quoted false memory flags disable both built-in stores as expected.
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
Route the model-supplied target through _bound_error_text so a huge
bogus target can't bloat context, and restore the "Use 'memory' or
'user'" hint. Follow-up to HexLab98's review note on the salvage.
…iases for custom opencode-* providers

Builds on @Lesnak1's NousResearch#85619 (issue NousResearch#85589):

- New opencode_provider_family() single-owner predicate in
  hermes_cli/models.py — resolves built-in AND custom family providers
  (opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
  Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
  x4 from the salvaged commits) plus 4 sibling sites the PR missed:
  cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
  model_normalize.py flat-namespace strip, model_switch.py base_url
  normalization.
- Responses transport: alias OpenCode-reserved function names
  (web_search, search_files -> hermes_*) on the wire and map them back on
  dispatch — same pattern as the xAI web_search collision fix. Matches
  family providers and any base_url on opencode.ai. Fixes the HTTP 400
  'custom function name X is reserved' half of NousResearch#85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
…mily providers

The named-custom-provider runtime path returned a static api_mode, so a
providers: entry like opencode-go-bridge -> https://opencode.ai/zen/go/v1
sent responses-only models (grok-4.5, gpt-5.6-luna) to /chat/completions
and got HTTP 503 (NousResearch#85589 repro). Now: when the provider name is in the
OpenCode family or the base_url is hosted on opencode.ai, derive api_mode
from the effective model and run the symmetric /v1 normalization — unless
the user declared an explicit transport, which stays authoritative.

5 new regression tests against a real temp HERMES_HOME config.
…ent config

sol-reviewer round-2 IMPORTANT: relay env vars had no scope
classification, so the two readers disagreed under a multiplexed
profile scope — gateway/config.py (scope-aware getenv) dropped a
process-env GATEWAY_RELAY_URL during the scoped runner reload while
gateway/relay's relay_url()/register_relay_adapter()/self-provision
(direct os.environ reads) still saw it. Result: adapter registered but
Platform.RELAY absent from config, so the connect loop never dialed
and direct adapters stayed up. The inverse split (profile-only stamp:
config enables RELAY, registration finds no URL) was equally dead.

GATEWAY_RELAY_* ROUTING stamps (URL, ENDPOINT, ALLOW_DIRECT_PLATFORMS,
PLATFORMS, BOT_IDS, ROUTE_KEYS, INSTANCE_ID, WAKE_URL, DISPLAY_NAME)
are now in _GLOBAL_ENV_EXACT: deployment config read from os.environ
under any scope, exactly like the API_SERVER listener settings
(NousResearch#69379), so every reader resolves the same value. Relay AUTH material
(SECRET, ID, DELIVERY_KEY, IDP_*) is deliberately NOT global — it
stays profile-scoped with the fail-closed multiplex guard, mirroring
the non-secret/secret line the terminal env blocklist already draws
(tools/environments/local.py).

The round-1 multiplex regression test asserted the now-rejected
semantic (profile-scoped stamps win); it is inverted to pin the
global-stamp contract: a process-env stamp survives the profile scope
(sweep runs, matching registration), and a profile-only .env stamp
does NOT activate relay.
…ultiplex profile

sol-reviewer round-3 IMPORTANT: gateway enroll --connector-url /
--wake-url persist GATEWAY_RELAY_URL / GATEWAY_RELAY_WAKE_URL into the
active profile's .env and tell the user a restart activates them. For
a SECONDARY profile of a multiplexed gateway that is silently untrue:
the routing stamps are process-global, and a secondary profile's .env
is loaded into an isolated secret scope, never exported to os.environ,
so the gateway can never read them from there. Enrollment (the
credential exchange) still succeeds and the creds are still valid, so
warn rather than refuse, pointing at the process environment or the
default profile. Single-profile gateways and the launch profile are
unaffected (load_hermes_dotenv exports their .env at startup).
…iplex scope

sol-reviewer round-3 findings: the regression tests stopped at the
config boundary, so config/registration agreement was only manually
verified. New TestConfigRegistrationAgreementUnderMultiplexScope
exercises both sides under an active profile scope: a process-env
stamp yields Platform.RELAY enabled AND relay_url()/
register_relay_adapter() agreement AND a constructed RelayAdapter with
a live transport; a profile-only .env stamp is inert on BOTH sides (no
half-enabled state). Also narrows the test_config docstring that
overclaimed profile .env stamps are unsupported — the launch profile's
.env still activates relay via load_hermes_dotenv's os.environ export;
only an isolated multiplex scope is never consulted.
…ning

sol-reviewer round-4 IMPORTANT (reproduced by execution): the enroll
warning read multiplex_profiles via load_gateway_config() under the
SECONDARY profile's HERMES_HOME, but the flag normally lives in the
DEFAULT root's config.yaml — so the warning never fired in the real
topology, preserving the round-3 defect it claimed to fix.

The topology decision now mirrors the multiplexer-conflict guard in
hermes_cli/gateway.py: secondary detection is the resolved-path
relationship to <default_root>/profiles/ (not a directory-name
heuristic — also fixes the round-4 MINOR false positive on unrelated
dirs named 'profiles'), and the multiplex flag comes from the
GATEWAY_MULTIPLEX_PROFILES env override or a raw read of the default
root's config.yaml. The raw read also avoids running the full
enablement pass (round-4 MINOR: load_gateway_config() emitted the
relay-exclusive sweep's own warnings into enroll output).

The warning now replaces the generic 'restart to pick up the new env'
line instead of following it (round-4 NIT: the two messages were
contradictory), and the helper returns whether it fired.

New test file pins all six topology cases, including the exact
false-negative reproduction (flag in default root only) and env
override in both directions.
…-exclusive-messaging

fix(gateway): GATEWAY_RELAY_URL env stamp disables direct messaging platforms
…loud-mcp

docs: guide for managing Hermes Cloud via the Portal MCP server
Two follow-ups to the salvaged NousResearch#89369 base against current main:
- compact projection entries no longer overwrite the local rich copy
  (attachments survive; watermark accounting stays stable)
- synthetic legacy-N thread ids collapse to one bucket in the entry key,
  so id-less entries don't duplicate after a pull and manufacture
  phantom member turns into busy sessions
…p-room sync

Resolves the room-lifecycle class on top of the salvaged NousResearch#89369 projection:

- v3 projection keys rooms by immutable roomId (id:<roomId>) with
  name:<name> fallback for legacy rooms; v1/v2 envelopes are normalized
  on read so mixed-version fleets share one merge path
- rename is now a same-key field update — no distributed delete+create,
  no old-name resurrection from lagging gateways
- id tombstones are FINAL (ids are never reused), so a gateway that was
  offline during a disband can never resurrect the room, regardless of
  the revision its stale copy carries; same-name recreation is unaffected
  because it mints a fresh roomId
- the projection fans out to EVERY reachable default-profile gateway
  (per-gateway job queues, CAS revision streams, backoff and retry caps),
  so rooms survive any single gateway dying and surface on gateway-only
  clients without waiting for a Desktop to foreground that gateway
- cold hydrate follows a remote rename via roomId instead of duplicating
  the room under both names

New tests: id-keyed rename continuity, final id-tombstones vs lagging
high-revision copies, rename-job shape (changed+deleted same key),
cold-hydrate re-keying, multi-gateway fan-out. Sabotage-verified: each
new test fails against the pre-class behavior.
kshitijk4poor and others added 20 commits August 22, 2026 14:55
Review follow-ups for the salvaged picker partitioning:
- Drop the _model_chunks stash — it went stale when navigating to a
  provider with an empty model list (early return skipped the
  reassignment), producing a wrong 'N more available' count. Derive
  shown = min(len(models), 75) directly instead.
- Add _DISCORD_SELECT_MAX_OPTIONS / _DISCORD_SELECT_MAX_ROWS constants
  per the file's named-limits convention; replaces 4 bare literals.
- Remove dead total_rows variable.
…s-profile SSH leakage

_resolve_container_task_id always returned "default", so _active_environments
shared a single SSHEnvironment across all WebUI sessions. When a user switched
from profile A (ssh_host=10.0.0.1) to profile B (ssh_host=10.0.0.2), the new
session found _active_environments["default"] already set to A's SSHEnvironment
and reused it — silently running every command on the wrong remote host.

Fix: when HERMES_SESSION_KEY is present (set per-session by the WebUI streaming
layer and per-message by the gateway via contextvars), return "session:<key>"
as the cache key instead of "default". Each session now owns its own slot in
_active_environments and always creates an environment from its own profile's
TERMINAL_SSH_HOST / TERMINAL_ENV config.

Behaviour unchanged in CLI mode (no HERMES_SESSION_KEY → still "default").
RL/benchmark task overrides (register_task_env_overrides) are unaffected.
Subagent task_ids inside a WebUI session collapse to "session:<key>" so they
continue to share the parent session's container.

Five new regression tests added to test_shared_container_task_id.py.
The existing session-key regressions set HERMES_SESSION_KEY via os.environ,
which only exercises the os.getenv() fallback branch. Real gateway turns bind
the identity through gateway.session_context.set_session_vars() (a ContextVar)
and never write the process-global env var. Add two companion regressions that
bind via set_session_vars() with HERMES_SESSION_KEY absent from os.environ:

- test_session_key_from_contextvar_without_environ: container slot scopes to
  session:<key> purely through the ContextVar (subagent inheritance covered).
- test_contextvar_session_key_wins_over_environ: with a different value left in
  os.environ, the ContextVar-bound session wins, so two concurrent gateway
  sessions in one process cannot cross-contaminate via the process global.

Cleanup via clear_session_vars(tokens) in finally.
…-key lookups

Follow-up to the session-scoping fix: _get_sudo_password_cache_scope()
and _resolve_container_task_id() carried byte-identical copies of the
HERMES_SESSION_KEY lookup (contextvar + os.environ fallback). Collapse
both onto one helper adopting the bare-import convention approval.py
already uses — get_session_env() implements the fallback internally, so
the old try/except could only fire on import failure, where silently
degrading to process-global semantics would reintroduce exactly the
cross-session contamination the fix prevents.
Telegram RetryAfter on send() slept the server retry_after with no
ceiling, so a 97-minute penalty pinned the coroutine. Mirror the edit
path: waits over 5s return immediately; short waits still retry inline.
Restart notification and obligation redelivery ran before the
startup-restore gate opened, so one hung Telegram send queued inbound
on every platform. Bound those sends with the same timeout the resume
gate already uses, and clear resume_pending before send so a timed-out
redelivery cannot also replay the turn.
… any redelivery send

Follow-up to the salvaged NousResearch#91986: the per-row clear still left rows the
loop had not reached exposed — a slow send ahead of them could hold the
loop past the inbound-gate timeout and let
_schedule_resume_pending_sessions replay those turns. Clearing every
claimed row up front closes the duplicate window; claiming already
spent the redelivery attempt, so the ledger retry path is unchanged.
The eager session.resume path called _transfer_db_to_agent(agent, db)
unconditionally. With no non-launch profile selected, db resolves to the
SHARED launch handle (_get_db()), so the transfer succeeded on identity
alone — the agent IS holding that handle — and session.close() then
closed the process-wide database under every unrelated session:
subsequent writes failed with "'NoneType' object has no attribute
'execute'" and the Desktop could not open chats until restart (NousResearch#91610).
This directly violated _transfer_db_to_agent's own contract ("Never
called for the shared launch handle", introduced with the ownership
lifecycle in NousResearch#81071).

Gate the transfer on owns_db (dedicated handles only), and add defense
in depth: _transfer_db_to_agent now refuses db is _get_db() even when a
caller invokes it incorrectly.
…ady queued

_ensure_reconnect_watcher_running() exists for one situation: the reconnect
watcher has exhausted _MAX_SUPERVISED_RESTARTS, so _spawn_supervised has logged
"giving up restarts" and will never bring it back on its own (NousResearch#70344, and the
supervised-restart half of NousResearch#71758). It had exactly one call site, inside the
newly-queued branch of _queue_retryable_fatal_platform.

That branch is unreachable for a platform already in _failed_platforms, which
is the only kind of platform the watcher can have been retrying long enough to
burn five rapid restarts on. So the backstop could not fire in the one state it
was written for.

The failure is silent by construction. The early return logs nothing, so there
is no "queued for background reconnection" line. The stranded check in
_handle_adapter_fatal_error_detached deliberately treats a queued platform as
safe, so the gateway does not exit for the service manager either. With another
platform still connected, self.adapters is non-empty and the "gateway staying
alive, watcher will retry in background" branch is skipped too. A retryable
fatal error can therefore produce a single ERROR line and then nothing: the
platform sits in the queue that nobody is draining until someone restarts the
process by hand (NousResearch#90386 reports 4h17m of that, with cron unaffected throughout).

Call the ensure on the already-queued path as well. It is already idempotent
and already cheap: it returns immediately unless the tracked task is done, and
it routes through the same on_spawn handle tracking, so a live watcher is never
duplicated.

The queue entry itself is deliberately left untouched. Re-enqueueing would
reset attempts and next_retry, restarting the backoff ladder on every fatal
error and hammering a provider that is already refusing the connection.
Review of NousResearch#90448 by @andrexibiza: adding _ensure_reconnect_watcher_running()
to the already-queued branch of a fatal callback is still an event-coupled
check. It needs a later fatal error from some other platform to arrive, and
NousResearch#81036 makes that less likely rather than more -- it publishes the queue
before disconnect and drops the failed adapter from the live map, so after
the watcher's supervised restart budget is spent there may be no adapter
left to emit the event recovery is waiting on.

That is the state NousResearch#72366 (salvage of NousResearch#71867 by @ygd58) restored supervision
to close: queued work exists, the watcher is dead, and nobody owns the
invariant. Supervision being finite is correct; having no owner past the
budget is not.

_spawn_supervised now takes on_give_up, invoked when it abandons a task --
the supervisor is the only thing that knows it has. The reconnect watcher
uses it to hold:

  while _running and _failed_platforms is non-empty, either a reconnect
  watcher is live or a bounded respawn is scheduled.

Empty queue: leave it down and log; the enqueue path spawns a fresh watcher
the moment something depends on one. Non-empty: a bounded slow tier at
_RECONNECT_WATCHER_SLOW_RETRY_SECS (300s) for _MAX_SLOW_WATCHER_RESPAWNS (6)
attempts, standing down early if the queue drains or a watcher returns on
its own. Exhausted: one loud error naming the platforms left unattended.

The ceiling is (1 + _MAX_SUPERVISED_RESTARTS) x (1 + _MAX_SLOW_WATCHER_RESPAWNS)
spawns -- 42 across at least half an hour -- because each slow attempt hands
the watcher a fresh supervised budget. A test asserts that ceiling so it
cannot quietly become a restart loop.

Deliberately NOT included: requesting a process restart when the slow tier
is also exhausted. Taking down every healthy platform to heal a sick one is
a blast-radius policy decision for a maintainer.

Two things this turned up:

- _spawn_supervised did not thread on_give_up through its own backoff
  respawn, so the callback was lost after the first restart and the give-up
  branch had no owner at exactly the moment it needed one -- the same defect
  the on_spawn docstring warns about, one parameter over.
- Three call sites repeated the (factory, name, on_spawn) triple, whose
  on_spawn half is load-bearing. They now go through
  _spawn_reconnect_watcher().

_supervised_backoff() names the previously-inline exponential schedule so
the exhaustion tests can collapse it; production behaviour is unchanged.

Refs NousResearch#90386
Spawn-time Context isolation cannot rewrite an already-running watcher task. Run dispatcher SQLite offloads in an empty Context so write_txn no longer false-trips after delegate_task, while real child callers still hit the mutation guard.
@RuniThomsen

Copy link
Copy Markdown
Collaborator Author

Superseded by #3; #2 compared the upstream-based branch directly with the fork's intentionally patched main and therefore included unrelated divergence.

RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
Kanban cards have no length limit, but the session title store rejects
titles past SessionDB.MAX_TITLE_LENGTH with ValueError. _persist_session_title
reads that as a unique-title collision, retries with a "#N" suffix (longer
still), and the caller suppresses the second failure - so a worker spawned on
a >100-char card ended up with no title at all, where main at least gave it a
derived one. Trim the card title (with room for the "#N" retry suffix) before
persisting; a retried card now gets "<trimmed> #2" within the cap.

Review finding: >100-char card title left the kanban worker session untitled.
RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
…, with or without the multiplex flag

Two authority gaps in served_profile_child_env (NousResearch#111617 review, andrexibiza P1 #1/#2,
kvnloo finding 1):

- The base was hermes_subprocess_env(inherit_credentials=True) = the launch environ's
  provider credentials; strip_launch_profile_env only knows names with .env/source
  provenance, so a key systemd/Compose/the shell injected into the launch process
  survived into profile B's child whenever B did not define the same name. Now a ROUTED
  target scrubs every Tier-1/Tier-2 credential from the base regardless of provenance
  before B's own scope is overlaid (the child boundary gets get_secret's contract: a
  scoped miss is no credential, never ambient fallback). The launch profile's own child
  keeps its env. bot_relay's base=os.environ goes through the same scrub.
- strip_launch_profile_env / the scrub keyed on is_multiplex_active(); the Desktop and
  dashboard backends serve ?profile=B by installing the HERMES_HOME override without
  that flag, so B's slash worker / helper children kept A's .env and settings. The
  authority test is now "is the target a routed home" (target != process home).
- _build_browser_env resolved the passthrough keys via get_secret, which falls through
  to os.environ on a scoped miss while multiplexing is inactive: a routed B with no
  Firecrawl key got A's. Under serves_routed_profile() the bound scope is the only source.
- served_profile_child_env(inherit_credentials=True) with no target and no scope bound
  under multiplex minted with the launch credentials (key_cmd TTL refresh on a worker
  thread); it now raises UnscopedSecretError like get_secret.

tests/tui_gateway/test_served_profile_child_env_authority.py: ambient-only A key + B
missing it (mux on), flag-off routed B (helper child + browser), real child observation.
3/3 red on base.
RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
`_TERMINAL_KANBAN_TOOLS` listed only `kanban_complete` / `kanban_block`,
so the turn-end guard fired at workers that had already handed the card
off correctly:

- A build worker that calls `kanban_request_review` moves the card from
  `running` to `review` (tools/kanban_tools.py `_handle_request_review`),
  yet `session_called_kanban_terminal()` returned False and the nudge
  told it "Task is still `running`" — false by then — and to call
  `kanban_complete`, which would close a card that must go through
  review. Both goal-mode prompts name that tool explicitly
  (hermes_cli/goals.py `KANBAN_GOAL_CONTINUATION_TEMPLATE` and
  `KANBAN_GOAL_FINALIZE_TEMPLATE`), so the prompt and the guard
  contradicted each other.

- Review agents hit the same wall. The review lane spawns through the
  same `_default_spawn`, which sets `HERMES_KANBAN_TASK`, so the guard is
  active for them, and the force-loaded sdlc-review skill's decision
  table ends the request-changes path on `kanban_request_changes`.

Adds both handoff tools to the set. Workers that obey their prompt now
exit cleanly; the guard still fires for a non-terminal board tool, which
is the case it exists for.

Covers lifecycle mismatch #2 of NousResearch#94916 only — the `dispatch --dry-run`
half is a separate change and is discussed on the issue.
RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
… in the flag table

- Dialog 2 now shows the read-then-answer step for one benign prompt and
  warns against blind timed Enter, and names --permission-mode acceptEdits
  as the narrower opt-in (idea from PR NousResearch#113456).
- Quick Reference row and pitfall #2 carry the opt-in wording instead of
  teaching Down+Enter as the expected move.
- Regenerated the claude-code docs page.
RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
Adopts the DNS-rebinding-pinned transport hardening (repo issues #2/#3,
PR #3) and the starter-feed/settings failure-surfacing + SSRF-gated icon
proxy fix (issue #6, PR #8). Full range in
tony-simons-aiowa/hermes-newswire deccdc4..e6b438e (13 commits):

Security-relevant highlights:
- All outbound fetches (feeds, redirects, icons) now go through a pinned
  transport: the SSRF gate's validated address set is bound to the actual
  connection — no second DNS lookup, so DNS rebinding/TOCTOU has no
  window; the plugin fails closed if the pin seam changes.
- New GET /icon.json proxies favicons through the same gate and returns
  base64 data URLs — the renderer's <img> no longer performs unpinned
  DNS resolutions of feed-controlled hostnames. 64 KB cap enforced
  mid-transfer; image content-type allowlist; bounded, normalized TTL
  cache.
- Renderer surfaces backend failures (settings/sources banners,
  starter-feed inline errors) instead of silent no-ops.

Capabilities unchanged (all empty — dashboard plugin, no tools/hooks/
env). Verification at the new pin: 129 pytest, 33 renderer interaction
checks, 26 ESM render smoke, hermes plugins validate clean.
RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
…xes)

Picks up the plugin's merged fixes: empty-input tool calls no longer fail the
request (#2), Haiku 4.5 requests no longer send adaptive thinking, top-level
schema combinators are stripped before the native validator, the claude CLI is
resolved on the env PATH, and relay/native failures now name their real cause
instead of only the admission denial. Plugin version unchanged (0.3.0); the
card art URL follows the pin.
RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
gc_blobs carried two byte-identical `warning("malformed ledger line; blob
GC skipped"); return 0, 0` blocks — one under `except json.JSONDecodeError`,
one under `if not isinstance(row, dict)` (tools/skill_ledger.py:350-357;
simplify reuse #2, efficiency note, re-gate G2 S2). Funnel the decode
failure into the type guard (`row = None`) so the abort exists once.
Behaviour is unchanged for both the [malformed-json] and [non-dict-row]
parametrizations: any line that is not a JSON object still aborts the
sweep with the same warning.
RuniThomsen pushed a commit that referenced this pull request Sep 24, 2026
argparse's _get_value always calls a `type=` callable with the raw str
token, and int(<str>) can only raise ValueError, so the TypeError arm in
`except (TypeError, ValueError)` is unreachable. The try itself stays:
without it argparse would print "invalid _nonnegative_int value: 'thirty'",
leaking the private helper name into the usage error.

Invariant: `_nonnegative_int` is only referenced as an argparse `type=`
(hermes_cli/kanban_parser.py) and by the unit test, which passes str.

Finding: simplify/D.quality.md #2 (hermes_cli/kanban_parser.py:50).
Dead-branch deletion; existing test_gc_parser_rejects_negative_retention_days
still covers the "thirty" -> "must be an integer" path (green).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.