Skip to content

Two plugins installed together both go live with their MCP tools and skills - #119960

Merged
alt-glitch merged 1 commit into
mainfrom
fix/plugin-go-live-race
Sep 23, 2026
Merged

alt-glitch merged 1 commit into
mainfrom
fix/plugin-go-live-race

Conversation

@alt-glitch

Copy link
Copy Markdown

Two plugins that finish installing at the same time now both go live with their MCP tools and skills. Before, one of them usually came up with no tools until the app restarted.

The problem

In Hermes (Nous Research's agent app), installing a plugin ends with a "go-live" step, load_and_go_live. It does three things in the running backend:

  1. It rescans the plugins on disk (a forced rediscovery).
  2. It reads the new plugin's MCP server configs and skills from the plugin manager.
  3. It connects those MCP servers and tells the open chats about them.

An MCP server is the local process that exposes a plugin's tools. NVIDIA App's server exposes 12 tools, for example.

A forced rediscovery first unloads every plugin, then loads them all again. While it runs, the plugin manager is empty. Two go-lives can run at the same time, because the Desktop onboarding install card installs its rows in parallel. Then the second go-live's rescan empties the manager while the first go-live reads it or connects. The first plugin reads nothing, connects nothing, and reports success with no tools.

sequenceDiagram
  participant A as go-live nvidia-app
  participant M as plugin manager
  participant B as go-live nvidia-broadcast
  A->>M: rescan (loads both plugins)
  B->>M: rescan: unload all...
  A->>M: read nvidia-app servers + skills
  M-->>A: {} (manager is empty)
  B->>M: ...load all again
  B->>M: read + connect nvidia-broadcast (10 tools)
  Note over A: nvidia-app: "Installed", 0 tools, 0 skills
Loading

Seen in a cold onboarding run on a Windows RTX 5090 test PC (main at 5f47c35d37):

  • The install card row said "Installed" for NVIDIA App with no tool count. The NVIDIA Broadcast row said "Installed · 10 tools".
  • The backend log has no connect attempt for NVIDIA App. The two rescans ran 200 ms apart ("59 found" at 12:10:10.003, "60 found" at 12:10:10.205).
  • In the build chat that followed, tool_search "nvidia app gpu driver status" found nothing. The model used nvidia-smi instead. The Plugins tab said "nvidia-app is not running", which was also wrong, because only the MCP connection was missing.
  • Nothing fixed it until the app restarted. A new chat did not help.

The per-install rescan came from #119266 (late-loaded plugins wire their handlers live). The connect after the rescan came from #119644 (installed plugins' MCP tools are live in open chats).

What changes

hermes_cli/plugins_activation.py and hermes_cli/plugins_activation_live.py (+27/−9):

  • load_and_go_live holds a process-wide lock, _GO_LIVE_LOCK, from the rescan to the chat refresh. One go-live runs at a time. The second install's go-live waits for the first, which takes about 1 s per server.
  • The connect is inside that lock too. The connect reads the server's liveness declaration (how Hermes finds the server's live URL and token, tools/mcp_tool_transport.py::_live_endpoint), and a rescan clears that declaration as well.
  • The rescan and the reads of servers and skills also hold the plugin manager's own _discovery_lock. Other forced rescans (the gateway's reload-plugins verb, the dashboard) do not take the go-live lock, but they do take that one.
  • The MCP connect stays outside _discovery_lock. Chats call discover_plugins(), and they must not wait for a network connect.
  • connect_plugin_mcp(activation, portable) receives the server configs read under the lock, instead of reading the manager again. It has one caller.
  • A background discovery is joined before any lock is taken, so the go-live cannot deadlock on it.

What the user experiences

Situation Before After
Onboarding card: Install clicked on two rows back to back One plugin usually shows "Installed" with no tools; its tools and skill are missing in every chat until restart Both rows show "Installed · N tools"; both plugins' tools and skills are live
Two installs from any surface that overlap (Plugins tab, CLI + Desktop) Same race Same fix: the go-lives run one after the other
One install at a time Works Works, same speed

What this does not do

  • It does not split "install" from "refresh". Today every install ends with its own rescan and connect. A cleaner design installs rows in parallel, then refreshes once for all new plugins at a single trigger, such as the card settling. That design touches every install surface and is a separate change. This lock stays useful after it, because two separate refreshes can still overlap.
  • It does not fix the related race in fix(tui): plugin activation refreshes sessions under the reload lock #119862: a plugin's chat refresh that reads the tool registry while reload.mcp has torn it down. That PR is independent of this one.
  • It adds no test. Sid (maintainer) asked for no new tests in this batch of fixes from the Windows onboarding runs. The live check below reproduces the race on demand.

How to test

scripts/run_tests.sh tests/hermes_cli/test_plugins_cmd_activation_keys.py tests/tui_gateway/test_plugins_manage_late_activation.py tests/gateway/test_late_plugin_rewire.py tests/hermes_cli/test_plugins.py
# 4 files, 106 passed, 0 failed

ruff check, scripts/check-windows-footguns.py --all and scripts/check_compat_pointers.py are clean.

Live verification

Harness on the Windows RTX 5090 test PC. It copies the onboarding run's installed nvidia-app and nvidia-broadcast plugins into a temporary HERMES_HOME, starts load_and_go_live("nvidia-app") and load_and_go_live("nvidia-broadcast") on two threads, and counts each plugin's live tools, skills, and tools in tools.registry. The MCP servers are the real NVIDIA App and NVIDIA Broadcast servers.

Tree Start Runs where one plugin got 0 tools, 0 skills
main 5f47c35d37 together 13 of 13
main 5f47c35d37 second 100 ms later (the onboarding timing) 6 of 6, always NVIDIA App
this PR bcb9d5ba4f together 0 of 8
this PR bcb9d5ba4f second 100 ms later 0 of 3

In every run on this PR, both plugins end with all their tools registered (NVIDIA App 12, NVIDIA Broadcast 10) and 1 skill each. A start 300 ms apart does not reproduce the race on main, which is why an earlier onboarding run passed.

Not yet observed: a full cold Desktop onboarding run on this commit.

The onboarding install card installs its rows concurrently, and each install
runs load_and_go_live: a forced plugin rediscovery, then a read of the plugin's
MCP server configs and skills, then an MCP connect. A forced rediscovery
unloads every plugin first (PluginManager.unload clears _portable_mcp_servers,
_plugin_skills and each server's liveness declaration) and only then loads them
again. When NVIDIA App and NVIDIA Broadcast finished installing together, the
Broadcast pass unloaded the manager while the NVIDIA App go-live was reading
it: NVIDIA App went live with no MCP servers and no skill (the card row said
"Installed" with no tool count, and the build chat had no NVIDIA App tools
until the app restarted).

load_and_go_live now runs one at a time in the process (_GO_LIVE_LOCK), so a
second install's rediscovery cannot run while the first reads or connects.
The connect is inside that lock because it reads the server's liveness
declaration (tools/mcp_tool_transport.py::_live_endpoint), which a forced pass
also clears. The reads share the manager's discovery lock with the pass that
produced them, for forced passes that do not go through load_and_go_live, and
connect_plugin_mcp takes the server configs read there instead of reading the
manager again. The MCP connect stays outside the discovery lock, so chats that
call discover_plugins() are not held for the length of a connect.
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on bcb9d5b — fix: two plugins installed together both go live with their

debug info

CI timings

CI timings · View report · View job

Wall time 5m54s vs 6m12s (-4.8%). 8 job(s) slower, 4 faster,

  • OS-specific tests / macOS-only tests: +60.0s
  • Profile artifact check / Reject profile archives: +38.0s
  • Detect affected areas: -33.0s
  • OS-specific tests / Windows-only tests: +16.0s
  • Check contributors / check-attribution: -14.0s

@alt-glitch
alt-glitch merged commit c07501e into main Sep 23, 2026
34 checks passed
@alt-glitch
alt-glitch deleted the fix/plugin-go-live-race branch September 23, 2026 07:15
@alt-glitch alt-glitch added type/bug Something isn't working P3 Low — cosmetic, nice to have comp/cli CLI entry point, hermes_cli/, setup wizard comp/plugins Plugin system and bundled plugins tool/mcp MCP client and OAuth labels Sep 23, 2026
@alt-glitch

Copy link
Copy Markdown
Author

Desktop after-evidence, on main at c07501ec41 (this fix merged), cold onboarding run on the Windows RTX 5090 test PC, 2026-09-23 07:20Z:

  • On the install card, Install was clicked on both rows 170 ms apart (07:23:04.64Z and 07:23:04.82Z).
  • Both rows then showed tools: "Installed · 12 tools · skill …nvidia-app" and "Installed · 10 tools · skill …nvidia-broadcast".
  • The backend log shows both servers registered, 350 ms apart: MCP server 'nvidia-app' (HTTP): registered 12 tool(s) at 12:53:08.229 and 'nvidia-broadcast' … registered 10 tool(s) at 12:53:08.578.
  • In the first build chat, the model loaded both NVIDIA skills, then called nvapp_client_get_driver_status, nvapp_client_get_applications, nvapp_overlay_get_status and get_broadcast_state. All four returned results.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cli CLI entry point, hermes_cli/, setup wizard comp/plugins Plugin system and bundled plugins P3 Low — cosmetic, nice to have tool/mcp MCP client and OAuth type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant