Two plugins installed together both go live with their MCP tools and skills - #119960
Merged
Merged
Conversation
The onboarding install card installs its rows concurrently, and each install runs load_and_go_live: a forced plugin rediscovery, then a read of the plugin's MCP server configs and skills, then an MCP connect. A forced rediscovery unloads every plugin first (PluginManager.unload clears _portable_mcp_servers, _plugin_skills and each server's liveness declaration) and only then loads them again. When NVIDIA App and NVIDIA Broadcast finished installing together, the Broadcast pass unloaded the manager while the NVIDIA App go-live was reading it: NVIDIA App went live with no MCP servers and no skill (the card row said "Installed" with no tool count, and the build chat had no NVIDIA App tools until the app restarted). load_and_go_live now runs one at a time in the process (_GO_LIVE_LOCK), so a second install's rediscovery cannot run while the first reads or connects. The connect is inside that lock because it reads the server's liveness declaration (tools/mcp_tool_transport.py::_live_endpoint), which a forced pass also clears. The reads share the manager's discovery lock with the pass that produced them, for forced passes that do not go through load_and_go_live, and connect_plugin_mcp takes the server configs read there instead of reading the manager again. The MCP connect stays outside the discovery lock, so chats that call discover_plugins() are not held for the length of a connect.
૮ >ﻌ< ა ci reviewran on bcb9d5b — fix: two plugins installed together both go live with their debug infoCI timingsCI timings · View report · View jobWall time 5m54s vs 6m12s (-4.8%). 8 job(s) slower, 4 faster,
|
Author
|
Desktop after-evidence, on
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two plugins that finish installing at the same time now both go live with their MCP tools and skills. Before, one of them usually came up with no tools until the app restarted.
The problem
In Hermes (Nous Research's agent app), installing a plugin ends with a "go-live" step,
load_and_go_live. It does three things in the running backend:An MCP server is the local process that exposes a plugin's tools. NVIDIA App's server exposes 12 tools, for example.
A forced rediscovery first unloads every plugin, then loads them all again. While it runs, the plugin manager is empty. Two go-lives can run at the same time, because the Desktop onboarding install card installs its rows in parallel. Then the second go-live's rescan empties the manager while the first go-live reads it or connects. The first plugin reads nothing, connects nothing, and reports success with no tools.
sequenceDiagram participant A as go-live nvidia-app participant M as plugin manager participant B as go-live nvidia-broadcast A->>M: rescan (loads both plugins) B->>M: rescan: unload all... A->>M: read nvidia-app servers + skills M-->>A: {} (manager is empty) B->>M: ...load all again B->>M: read + connect nvidia-broadcast (10 tools) Note over A: nvidia-app: "Installed", 0 tools, 0 skillsSeen in a cold onboarding run on a Windows RTX 5090 test PC (
mainat5f47c35d37):tool_search "nvidia app gpu driver status"found nothing. The model usednvidia-smiinstead. The Plugins tab said "nvidia-app is not running", which was also wrong, because only the MCP connection was missing.The per-install rescan came from #119266 (late-loaded plugins wire their handlers live). The connect after the rescan came from #119644 (installed plugins' MCP tools are live in open chats).
What changes
hermes_cli/plugins_activation.pyandhermes_cli/plugins_activation_live.py(+27/−9):load_and_go_liveholds a process-wide lock,_GO_LIVE_LOCK, from the rescan to the chat refresh. One go-live runs at a time. The second install's go-live waits for the first, which takes about 1 s per server.tools/mcp_tool_transport.py::_live_endpoint), and a rescan clears that declaration as well._discovery_lock. Other forced rescans (the gateway'sreload-pluginsverb, the dashboard) do not take the go-live lock, but they do take that one._discovery_lock. Chats calldiscover_plugins(), and they must not wait for a network connect.connect_plugin_mcp(activation, portable)receives the server configs read under the lock, instead of reading the manager again. It has one caller.What the user experiences
What this does not do
reload.mcphas torn it down. That PR is independent of this one.How to test
scripts/run_tests.sh tests/hermes_cli/test_plugins_cmd_activation_keys.py tests/tui_gateway/test_plugins_manage_late_activation.py tests/gateway/test_late_plugin_rewire.py tests/hermes_cli/test_plugins.py # 4 files, 106 passed, 0 failedruff check,scripts/check-windows-footguns.py --allandscripts/check_compat_pointers.pyare clean.Live verification
Harness on the Windows RTX 5090 test PC. It copies the onboarding run's installed
nvidia-appandnvidia-broadcastplugins into a temporaryHERMES_HOME, startsload_and_go_live("nvidia-app")andload_and_go_live("nvidia-broadcast")on two threads, and counts each plugin's live tools, skills, and tools intools.registry. The MCP servers are the real NVIDIA App and NVIDIA Broadcast servers.main5f47c35d37main5f47c35d37bcb9d5ba4fbcb9d5ba4fIn every run on this PR, both plugins end with all their tools registered (NVIDIA App 12, NVIDIA Broadcast 10) and 1 skill each. A start 300 ms apart does not reproduce the race on
main, which is why an earlier onboarding run passed.Not yet observed: a full cold Desktop onboarding run on this commit.