fix(gateway): multiplexed profiles use their own prefill_messages_file and provider_routing - #124630
Open
AhmetArif0 wants to merge 1 commit into
Open
AhmetArif0 wants to merge 1 commit into
AhmetArif0 wants to merge 1 commit into
Conversation
…e and provider_routing GatewayRunner read prefill_messages_file and provider_routing once, in _init_runtime_settings, from the launch profile's config, and every agent the gateway built took that copy. Under multiplexing a routed profile's turns therefore ran with the default profile's prefill messages and OpenRouter routing: its own only/ignore/order never applied, and a data_collection: deny set only on that profile was dropped. NousResearch#108453 made _load_prefill_messages resolve against the active home, but the loader only ever ran at boot, outside any profile scope. Read both where the turn runs, inside the routed profile's _profile_runtime_scope, the way reasoning and service tier are already resolved per turn and the ephemeral prompt was fixed in NousResearch#89161: _build_fresh_agent loads the prefill messages, and the main turn and /background load provider_routing. The boot copies are removed so nothing can hand the launch profile's values to another profile again. Cron already reads both per job, and the TUI/Desktop backend per agent build. A torn config.yaml write does not drop routing on a per-turn read: load_user_config_effective serves the last good parse.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
On a multiplexed gateway, a routed profile's turns run with the launch (default) profile's
prefill_messages_fileandprovider_routinginstead of their own.GatewayRunner._init_runtime_settingsreads both once at boot, outside any profile scope, intoself._prefill_messages/self._provider_routing, and every agent the gateway builds takes that copy:gateway/run_turn_runner.py::_build_fresh_agent→prefill_messages=runner._prefill_messagesgateway/run_turn_runner.py::run_sync→pr = runner._provider_routing(→providers_allowed/ignored/order,provider_sort,provider_require_parameters,provider_data_collection)gateway/run_turn.py::_run_background_task_inner(/background) →pr = self._provider_routingSo a secondary profile's own
only/ignore/ordernever apply, and aprovider_routing.data_collection: denyset on that profile is silently dropped: its requests go out with the default profile's data-collection preference.This contradicts
multi-profile-gateways.md("Per-turn runtime settings follow the routed profile … never from the profile the gateway was launched under"). It is also a gap in #108453: that PR listed_load_prefill_messagesamong the per-turn reads it moved to the active home (819517f), but the loader is only ever called at boot, so the change never took effect for a routed turn.The fix reads both where the turn runs, inside the routed profile's
_profile_runtime_scope. That matches how the same turn already resolves reasoning and service tier, and how the ephemeral system prompt was fixed for the same boot-snapshot problem in #89161. Cron already reads both per job, and the TUI/Desktop backend per agent build; the gateway was the only surface that snapshotted them.Related Issue
No issue filed. Found while checking the per-turn settings #108453 moved to the routed profile.
Type of Change
Changes Made
gateway/run_turn_runner.py:_build_fresh_agentcallsrunner._load_prefill_messages(); the turn computespr = runner._load_provider_routing()next to the per-turn reasoning / service-tier resolution.gateway/run_turn.py:/backgroundcomputespr = self._load_provider_routing()(it already runs inside_profile_scope_for_source).gateway/run.py:_init_runtime_settingsno longer snapshots_prefill_messages/_provider_routing. No production code reads them any more, and keeping a launch-profile copy around would invite the same bug.gateway/run_config_loaders.py:_load_prefill_messagesdocstring now says it resolves from the active gateway home and must be called per agent build (it said~/.hermes/).website/docs/user-guide/multi-profile-gateways.md: addsprovider_routing(includingdata_collection) andprefill_messages_fileto the list of per-turn settings that follow the routed profile.tests/gateway/test_routed_profile_prefill_routing.py(new).How to Test
Two profiles, both
prefill_messages_file: prefill.json, with differentprovider_routing(the routed one addsdata_collection: deny). The test drives the realGatewayRunner._run_agent→TurnRunner.run_sync→_build_fresh_agentpath, and the real_run_background_task, inside the routed profile's_profile_runtime_scope(the scope_profile_scope_for_sourceenters under multiplexing). The runner also carries the launch profile's boot copy, so the test proves a turn no longer uses it.mainprefill_messagesdefault-PREFILL❌beta-PREFILL✅providers_allowed["default-only"]❌["beta-only"]✅provider_data_collectionNone(dropped) ❌"deny"✅/backgroundproviders_allowed/data_collection["default-only"]/None❌["beta-only"]/"deny"✅Mutation check. Each change was reverted on its own with the test file kept:
prefill_messagespr→ the main-turn test fails onproviders_allowed/backgroundpr→ the background test failsWith all three restored, both tests pass.
Notes / scope
provider_routing/prefill_messages_filereaches newly built agents without a restart, the same asservice_tierandfallback_providerstoday. A cached agent keeps what it was built with. Rebuilding cached agents when routing changes is what fix(gateway): reload provider routing and bust cached agents on routing changes #32088 proposes; this PR does not touch the cache signature.config.yamlis being written.load_user_config_effectiveserves the last good parse for a broken file (in-process, then the.goodbackup).HERMES_PREFILL_MESSAGES_FILEstays a process-wide override, likeHERMES_EPHEMERAL_SYSTEM_PROMPT.getattr(self, "_provider_routing", {}). After this change that attribute no longer exists, so the display would show{}. Callingself._load_provider_routing()fixes that, and also shows the routed profile's routing under multiplexing.Checklist
Code
run.py; it does not cover prefill or multiplexing.tests/gateway+ related files at base9bb2ea1b5c: 16 failures on this branch, all of which also fail onmainat the same base.mainhas one extra flaky failure (test_control_socket::test_collect_fleet_versions_falls_back_to_state_file). No new failures.main: every test file that builds a runner with these attributes (38 files, plus the new one andtest_personality_routed_profile.py) passes, 258 tests.ruff checkis clean on the changed files, and so isscripts/check-windows-footguns.py.Documentation & Housekeeping
multi-profile-gateways.md, loader docstring)cli-config.yaml.example: N/A (no config keys added or changed)CONTRIBUTING.md/AGENTS.md: N/Aencoding="utf-8".