Skip to content

Plugin loader: a hung import/register() is skipped after a short deadline instead of hanging startup (#108139) - #118915

Merged
teknium1 merged 1 commit into
mainfrom
register-deadline
Sep 22, 2026
Merged

teknium1 merged 1 commit into
mainfrom
register-deadline

Conversation

@teknium1

Copy link
Copy Markdown
Collaborator

A plugin whose import or register() never returns no longer hangs Hermes startup: each plugin's load runs under a short deadline, the offender alone is skipped with a named reason, and the rest keep loading.

Fixes #108139 — thanks @Graemezee1 for the faulthandler-backed report. Supersedes #108144 by @KoNit-K: that PR bounded the join onto an in-flight background discovery (main already carries a 30 s bound there), while the hang lives in the load itself and also hits the first synchronous discovery, as @kvnloo's review noted; this PR bounds it at the source.

Changes

  • plugins.load_timeout_seconds (config.yaml, default 10, 0 disables, max 600) — deadline on one plugin's import + register() (hermes_cli/plugins_loader.py::run_with_load_deadline). The load runs on a daemon worker that inherits the caller's context (the Hermes-home ContextVar); on overrun the worker is abandoned and the plugin is recorded as failed with load timed out after 10s (import + register() never returned) — the same LoadedPlugin.error channel every other load failure uses (startup WARNING, /plugins, list_plugins()), and discovery continues with the next plugin. Pre-hang registrations are disposed through the existing failure path (Plugin loader: version gate reads running code, SystemExit isolated, uv quarantine from any cwd, impostor dirs refused, range pins (#72052 #104404 #101962 #112096 #108371 #71650 #86992 #98407) #118841's SystemExit/Exception isolation is unchanged; KeyboardInterrupt still propagates).
  • Late registrations ignored — the plugin's PluginContext is marked abandoned on timeout; every ctx.register_* / subscribe / on_unload from the still-running worker is refused with a WARNING (called register_hook() after its load timed out; ignored), so nothing lands in a registry the failure path already swept.
  • Bounded abandoned loaders (Concurrent observer-hook invocations are dropped as if a callback had timed out #98382 shape) — at most 8 live abandoned loader threads per process; past the cap a load is refused with not loaded: 8 abandoned plugin loader thread(s) are still running … restart Hermes to retry rather than run inline (inline would recreate the hang).
  • Lock discipline for the worker (it cannot own the caller's RLocks): the deferred-platform eager fallback now runs outside the replacement transaction; discover_and_load() re-entered from a loader worker (a plugin importing model_tools does that) returns on the already-set _discovered flag instead of blocking on the sweep's lock; _join_background_discovery is a no-op from a loader worker (its parent is the thread it would join).
  • Config default + dashboard schema entry (hermes_cli/config_defaults.py, hermes_cli/web_server_config.py), docs in website/docs/user-guide/features/plugins.md.

Root cause in one sentence: PluginManager._load_plugin_scoped called the plugin's register() synchronously on the discovering thread with no deadline, so one blocking plugin held discover_and_load() — and every synchronous caller (hermes chat, gateway boot, ACP session/new) — forever.

Validation

Live repro (fake home, hangplug with while True: pass in register() + healthy okplug, both enabled):

hermes chat -q hi
origin/main d855200 hangs; killed by timeout 40 (exit 124), no output
this branch WARNING … Failed to load plugin 'hangplug': load timed out after 10s (import + register() never returned) exactly 10.0 s after the load began, okplug loads, Plugin discovery complete: 60 found, 54 enabled, chat runs to completion (exit 0, Messages: 2)

Cap E2E: 10 hanging plugins + 1 healthy at load_timeout_seconds: 0.2 → discovery finishes in 1.7 s, 8 loaders abandoned, plugins 9–11 refused with the cap reason, list_plugins() reports every reason, 8 live plugin-load:* threads (never more).

Tests (tests/hermes_cli/test_plugin_manifest_v2.py::TestLoadIsolation): test_register_overrunning_load_timeout_skips_only_that_plugin — red on base (hangs past the runner timeout), green here; test_load_timeout_zero_runs_register_inline — 0 keeps the inline path. scripts/run_tests.sh tests/hermes_cli tests/plugins: 15323 passed; 4 test_dashboard_auth_gate reds are host-environmental (BACKEND_PORT_IN_USE port=9119, the live serve) and the one test_relay_shared_metrics / test_cmd_update red each pass in isolation (host load 27–48). ruff, check_no_tmp_literals, check-windows-footguns, check_compat_pointers, git diff --check clean.

Known limit (inherent to abandoning a thread): a plugin that hangs in a pure-Python busy loop keeps competing for the GIL after it is skipped, so the rest of startup runs slower than usual; a hang in blocking I/O (the reported MCP case) releases the GIL and costs nothing after the skip.

Infographic

Plugin load deadline

@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 22, 2026
@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 0cdbd29 — fix(plugins): per-plugin load deadline so a hung register()

debug info

CI timings

CI timings · View report · View job

Wall time 7m17s vs 5m59s (+21.7%). 6 job(s) slower, 6 faster, 1 unchanged.

  • Docs Site / docs-site-checks: +39.0s
  • Python tests / Run tests: -21.0s
  • Python lints / Windows footguns (blocking): -8.0s
  • Check contributors / check-attribution: -7.0s
  • Profile artifact check / Reject profile archives: -6.0s

… hangs startup

A plugin whose import or register() never returns (an infinite loop, a blocking
network call) held PluginManager.discover_and_load() forever, and with it every
synchronous caller: `hermes chat`, gateway startup, ACP session/new (#108139).

Each plugin's import + register() now runs under `plugins.load_timeout_seconds`
(default 10, 0 disables, max 600) on a daemon worker. On overrun the plugin is
recorded as failed with "load timed out after Ns" (same channel as every other
load failure: startup WARNING, `/plugins`, `list_plugins()`), its pre-hang
registrations are disposed, and discovery continues with the next plugin. The
abandoned worker's later `ctx.register_*`/`subscribe`/`on_unload` calls are
refused with a WARNING (the context is marked abandoned), so a late registration
can never land in a registry the failure path already swept. Abandoned loaders
are capped per process (8); past the cap further loads are refused with a named
reason rather than run inline, which would recreate the hang (#98382 shape).

Because the worker cannot own the caller's RLocks: the deferred-platform eager
fallback now runs outside the replacement transaction, discovery re-entered from
a loader worker returns on the already-set discovered flag instead of blocking on
the sweep's lock, and such a worker never joins the background discovery thread
that is waiting on it.

@arkheioncorp arkheioncorp left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: Plugin loader deadline (#118915)

Verdict: APPROVE (with one observation)

This addresses #108139 — a hung register() no longer hangs startup. The implementation is robust.

Strengths

  • run_with_load_deadline: runs import+register on a daemon thread via contextvars.copy_context().run (preserves the Hermes-home override ContextVar). On timeout, the worker is abandoned, ctx._abandon_load() is set, and PluginLoadTimeout propagates to the calling thread. Correct.
  • _ignore_after_abandoned_load: decorator that wraps every register_*, subscribe, and on_unload method on PluginContext. Late registrations from the abandoned worker are logged and dropped. This is the critical safety net.
  • _reserve_abandoned_loader_slot: caps abandoned loaders at 8. When the cap is reached, further loads are refused (not run inline) — this is the right call, since at the cap the process already has several hung loaders and an inline load would recreate the original hang.
  • in_plugin_load_worker(): used in discover_and_load to skip re-entrant discovery blocking, and in _join_background_discovery to avoid joining the worker from itself. Correct.
  • load_timeout_seconds: 0 disables the deadline and runs register() inline — tested and documented.
  • Deferred platform fallback: if _lease_deferred_platform fails, _load_plugin is called as an eager fallback. The eager load runs on a deadline worker, whose registrations need the coordinator lock — the comment explains this correctly.

Observation

  • _register_deferred_platform lost its @_serialized_replacement decorator in the refactor; _lease_deferred_platform now carries it. The fallback path calls _load_plugin outside the serialized section. If two threads concurrently try to load the same deferred platform and both leases fail, _load_plugin could run twice. Verify that _load_plugin / _load_plugin_scoped has its own synchronization (likely the _discovery_lock or an internal lock) that makes this safe. If not, the eager fallback could race.

Tests are thorough: timed-out plugin skips, late registration blocked, zero timeout runs inline, other plugins still load. No security concerns. Approved with the above observation noted.

@teknium1
teknium1 merged commit 9863e31 into main Sep 22, 2026
34 checks passed
@teknium1
teknium1 deleted the register-deadline branch September 22, 2026 08:11
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/cli CLI entry point, hermes_cli/, setup wizard comp/plugins Plugin system and bundled plugins area/config Config system, migrations, profiles sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/config Config system, migrations, profiles ci-reviewed applied to manually approve dangerous changes comp/cli CLI entry point, hermes_cli/, setup wizard comp/plugins Plugin system and bundled plugins P2 Medium — degraded but workaround exists sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: discover_plugins() has no bounded timeout on MCP-backed plugins, blocks ACP session/new (and 30+ other call sites) indefinitely

3 participants