Skip to content

fix(gateway): admit exactly one /update per profile - #78977

Open
briandevans wants to merge 2 commits into
NousResearch:mainfrom
briandevans:fix/gateway-update-atomic-admission-15539
Open

briandevans wants to merge 2 commits into
NousResearch:mainfrom
briandevans:fix/gateway-update-atomic-admission-15539

Conversation

@briandevans

@briandevans briandevans commented Aug 5, 2026 •

Copy link
Copy Markdown

Supersedes #15539

@hharry11's #15539 identified this collision correctly and the underlying overwrite path is still live on main. This PR re-does that fix against the code as it exists today, and closes the three specific deficiencies called out in review.

  • What fix: prevent concurrent gateway updates from clobbering shared IPC state #15539 covered: the diagnosis — two /update invocations clobber the shared .update_pending.json IPC state, and a second updater is spawned.
  • What fix: prevent concurrent gateway updates from clobbering shared IPC state #15539 did not do: it patched gateway/run.py:7821, but the handler moved to gateway/slash_commands.py in 619bd7827, so its guard never executes on HEAD (that PR is also CONFLICTING/DIRTY today). Its guard was if pending_path.exists() or claimed_path.exists(): return, a check-then-write that both callers can pass. Its two tests pre-seeded the marker, so they never reached the check-to-create gap.
  • What this adds: the guard in the file the handler actually lives in, an atomic reservation instead of check-then-write, a TTL so a crashed updater cannot wedge /update forever, and a synchronized two-caller regression that pre-seeds nothing.

What does this PR do?

_handle_update_command in gateway/slash_commands.py writes .update_pending.json and spawns a detached hermes update --gateway with no admission control of any kind. A grep -niE "claimed|lock|already|in_progress|is_running|O_EXCL" over the whole handler body on main returns nothing.

/update is profile-global: it rewrites the checkout and the virtualenv every session on the host shares, and the .update_* markers are a single-slot mailbox holding one requester's routing metadata. Two concurrent invocations — a user double-tapping because the update runs for minutes with no acknowledgement, or two platforms on one multiplexed gateway both triggering it — each write the marker and each spawn an updater against the same checkout and venv. The second write replaces the first requester's routing metadata, so that user never learns their update finished.

There is a second, louder symptom: both handlers stage through the same .update_pending.tmp path, so the losing Path.replace() raises FileNotFoundError straight out of the handler. That is reproduced by the new concurrency test against unpatched main — see How to Test.

The fix reserves the profile-wide slot with os.open(..., O_CREAT | O_EXCL) before any routing metadata is written and before the updater is spawned, collapsing observe-and-claim into a single atomic syscall so exactly one caller can ever win. O_EXCL is the same primitive on Windows (CPython raises FileExistsError there too), so the win32 spawn branch needs no separate handling.

Related Issue

No separate issue — this supersedes PR #15539, whose review states the acceptance criteria. Deliberately not using a closing keyword, so #15539 stays open for its author to close.

Acceptance checklist from the #15539 review

"gateway/run.py:7829 checks for a marker before the later write at PR lines 7843-7845. Two simultaneous handlers can both pass that check and both spawn an update process, so the reported race remains."

Answered by _claim_update_slot, which uses os.open(O_CREAT | O_EXCL). There is no window between observing the slot is free and taking it — it is one syscall, and the kernel picks the winner. test_claim_update_slot_is_exclusive_under_parallel_callers releases 16 threads from a barrier onto one path and asserts exactly one True.

"The handler moved to gateway/slash_commands.py in 619bd7827; current main still overwrites .update_pending.json at gateway/slash_commands.py:4531-4549."

The guard is in gateway/slash_commands.py, in _handle_update_command. The cited lines have drifted since the review — the handler now starts at :5393 and the unguarded write is at :5450-5453 on 36cb5ae55. Nothing is changed in gateway/run.py.

"The new tests at PR tests/gateway/test_update_command.py:218 and :257 pre-seed a marker, so they do not exercise the concurrent check-to-create gap."

No test in this PR pre-seeds a marker to stand in for the race. Every marker the guard is handed is created by the handler or by the primitive under test. test_concurrent_update_commands_spawn_exactly_one_updater runs two callers in real threads against the real filesystem, released together from a barrier placed at the last step before the claim, and asserts exactly one spawn plus one "already running" reply. test_second_update_is_rejected_without_preseeding_a_marker drives the reject path purely from state the first call produced, and asserts the winner's routing metadata survived byte-for-byte.

The one test that does write a marker, test_rejects_when_notifier_claims_a_live_update_mid_admission, writes it from inside the admission window to stand in for the concurrent notifier — that is the event being simulated, not state handed to the guard up front. See Follow-up below.

Sibling-site sweep

git grep -n "update_pending" origin/main -- '*.py' (non-test), every site classified:

Site Role Changed?
gateway/slash_commands.py:5450-5453 the only writer + spawner, unguarded yes
gateway/run.py:11281-11282 startup re-arms the notification watch if either marker exists no — reader
gateway/run.py:20564-20565 _watch_update_progress reads claimed→pending for routing no — reader
gateway/run.py:20806-20824 _send_update_notification — already claims atomically via pending_path.replace(claimed_path) no — reader, and already correct

The three run.py sites only ever consume the marker; none of them creates one, so none of them can produce a second updater. The reservation deliberately mirrors the claim idiom the reader at gateway/run.py:20819 already uses, so the two halves of the protocol match. tests/gateway/test_update_streaming.py exercises the watcher (a reader) and needed no change; it still passes unmodified.

Changes Made

  • gateway/slash_commands.py
    • New module-level _claim_update_slot(pending_path, ttl_seconds) — atomic O_CREAT | O_EXCL reservation, with takeover of a reservation older than the TTL.
    • New _UPDATE_RESERVATION_TTL_S = 3600.0. An updater killed before it writes its exit code (host reboot, OOM kill) leaves a marker nothing will clean up; without a ceiling the new guard would wedge /update for the lifetime of the profile — a worse bug than the one being fixed. Sized above _watch_update_progress's own 1800s watch timeout so a live, still-watched update is never stolen.
    • _handle_update_command claims the slot before writing routing metadata and before spawning. The loser gets t("gateway.update.already_running") and _schedule_update_notification_watch(), so it still learns how the running update turned out.
    • The metadata write moved inside the existing try, and the failure path releases the reservation (and any orphaned .update_pending.tmp), so a spawn that never started does not block the next /update.
  • locales/*.yaml (all 17) — new gateway.update.already_running key. tests/agent/test_i18n.py asserts catalog parity, so the key has to land in every locale in the same commit.
  • tests/gateway/test_update_command.py — new TestUpdateAdmissionControl with 6 tests.

How to Test

  1. Focused suites, all green:
    pytest tests/gateway/test_update_command.py tests/gateway/test_update_streaming.py \
           tests/agent/test_i18n.py tests/gateway/test_startup_restart_race.py -q
    # 64 passed
    
  2. Fails before / passes after. Restore only gateway/slash_commands.py to origin/main and re-run the new class:
    git checkout origin/main -- gateway/slash_commands.py
    pytest tests/gateway/test_update_command.py::TestUpdateAdmissionControl -q
    # 5 failed, 1 passed
    
    The concurrency test additionally surfaces the unhandled crash on unpatched main:
    FileNotFoundError: [Errno 2] No such file or directory:
      '.../hermes/.update_pending.tmp' -> '.../hermes/.update_pending.json'
      at gateway/slash_commands.py:5452 in _handle_update_command
    
    With the fix restored, all 6 pass. (test_failed_spawn_releases_the_update_slot is the one that passes both ways by design — it guards the new release path, i.e. that this PR cannot itself wedge /update, so it has nothing to fail against on main.)
  3. Full tests/gateway/ suite, run serially on this branch and on clean origin/main for comparison: the identical 7 failures on both, in test_discord_send.py, test_send_multiple_images.py, test_session_store_prune.py, test_shutdown_forensics.py, test_systemd_notify.py. All are pre-existing and order-dependent (they pass in isolation), none is in touched code, and the set does not change with this PR applied.
  4. Manual: from a messaging platform, send /update twice in quick succession. Before, two hermes update processes appear (pgrep -fa "hermes update") and only the second chat is notified. After, the second call answers "A Hermes update is already running for this profile" and one updater runs.

Follow-up: second-updater window closed in d15ce760e

Copilot's review caught a real defect in the first commit, and it is worth recording because it is the same class of bug as the one this PR fixes. The claimed_path.exists() pre-check was not synchronized with the exclusive create. _send_update_notification renames .update_pending.json → .update_pending.claimed.json while an update is in flight; landing that rename between the check and the create leaves pending momentarily absent, so the claim succeeded and admitted a second updater against the live one — and the notifier's later claimed.replace(pending) clobbered the new reservation. Reproduced, not theoretical:

AssertionError: Expected 'Popen' to not have been called. Called 1 times.
Calls: [call(['/usr/bin/setsid', 'bash', '-c',
        'PYTHONUNBUFFERED=1 /usr/bin/hermes update --gateway > .../.update_output.txt ...'])]

d15ce760e re-checks the claimed marker immediately after a successful claim and hands the reservation back if it is present. The release itself is narrowed: _release_update_slot only removes the marker while it is still zero-length — the state the exclusive create leaves it in — because if the notifier restored a live update's metadata in between, deleting it would lose that user's completion notice.

Both new tests (test_rejects_when_notifier_claims_a_live_update_mid_admission, test_release_update_slot_keeps_a_marker_that_holds_metadata) fail against 58ac06ac6 and pass on d15ce760e. 8 tests in TestUpdateAdmissionControl, 20 in the file.

Related / Positioning

Both dedup nets were run; stating what was checked rather than claiming the field is empty.

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate — see Related / Positioning
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run the tests: the full tests/gateway/ package (which contains every test this PR touches) plus tests/agent/test_i18n.py for catalog parity, run serially against both this branch and clean origin/main — identical results, zero new failures. Details in How to Test item 3.
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: macOS (Darwin 25.4, arm64), Python 3.11.15

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — the new helper and the TTL constant carry docstrings/comments explaining the invariant; no user docs describe /update concurrency
  • N/A — no config keys added or changed
  • N/A — no architecture or workflow change
  • I've considered cross-platform impact (Windows, macOS) — O_CREAT | O_EXCL raises FileExistsError on Windows as on POSIX, so the win32 spawn branch needs no separate guard; only the metadata write moved inside the existing try, which is platform-agnostic. Tested on macOS.
  • N/A — no tool descriptions or schemas changed

Copilot AI lite review requested due to automatic review settings August 5, 2026 00:21

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR hardens the gateway’s /update slash command so that exactly one update process can run per profile at a time, preventing concurrent invocations from clobbering the shared .update_* IPC markers and spawning duplicate updaters. It implements atomic admission control in the live handler (gateway/slash_commands.py), adds an “already running” localized user message across all locales, and introduces concurrency-focused regression tests.

Changes:

  • Add _claim_update_slot() using os.open(..., O_CREAT | O_EXCL) + TTL takeover to atomically reserve the profile-wide update slot before writing routing metadata and spawning the updater.
  • Return a new localized “already running” response when a second /update is rejected, while still scheduling the notification watcher.
  • Add a new TestUpdateAdmissionControl suite covering exclusivity, TTL reclaim, and real concurrent callers without pre-seeding markers.

Reviewed changes

Copilot reviewed 19 out of 19 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
gateway/slash_commands.py Adds atomic /update slot claiming with TTL, integrates admission control into _handle_update_command, and returns the new gateway.update.already_running message on rejection.
tests/gateway/test_update_command.py Adds a concurrency-focused test class validating exclusive admission, TTL reclaim, and single-updater spawning under parallel callers.
locales/af.yaml Adds gateway.update.already_running translation.
locales/ar.yaml Adds gateway.update.already_running translation.
locales/de.yaml Adds gateway.update.already_running translation.
locales/en.yaml Adds gateway.update.already_running translation.
locales/es.yaml Adds gateway.update.already_running translation.
locales/fr.yaml Adds gateway.update.already_running translation.
locales/ga.yaml Adds gateway.update.already_running translation.
locales/hu.yaml Adds gateway.update.already_running translation.
locales/it.yaml Adds gateway.update.already_running translation.
locales/ja.yaml Adds gateway.update.already_running translation.
locales/ko.yaml Adds gateway.update.already_running translation.
locales/pt.yaml Adds gateway.update.already_running translation.
locales/ru.yaml Adds gateway.update.already_running translation.
locales/tr.yaml Adds gateway.update.already_running translation.
locales/uk.yaml Adds gateway.update.already_running translation.
locales/zh.yaml Adds gateway.update.already_running translation.
locales/zh-hant.yaml Adds gateway.update.already_running translation.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread gateway/slash_commands.py Outdated
Comment on lines +5516 to +5523
if claimed_path.exists() or not _claim_update_slot(
pending_path, _UPDATE_RESERVATION_TTL_S
):
# The loser still gets the outcome: the watcher is profile-wide and
# reports the result into the chat that started the update.
self._schedule_update_notification_watch()
return t("gateway.update.already_running")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in d15ce760e (current head). This was a real second-updater admission, not a theoretical window — reproduced before the fix, with the exact interleaving you described:

AssertionError: Expected 'Popen' to not have been called. Called 1 times.
Calls: [call(['/usr/bin/setsid', 'bash', '-c',
        'PYTHONUNBUFFERED=1 /usr/bin/hermes update --gateway > .../.update_output.txt ...'])]

The fix follows your suggestion: re-check claimed_path immediately after a successful claim and release the reservation if it is present.

admitted = not claimed_path.exists() and _claim_update_slot(
    pending_path, _UPDATE_RESERVATION_TTL_S
)
if admitted and claimed_path.exists():
    _release_update_slot(pending_path)
    admitted = False

One thing worth calling out on the release side. A plain pending_path.unlink() reintroduces a smaller version of the same problem in the opposite direction: if the notifier's claimed_path.replace(pending_path) lands between the re-check and the unlink, the marker is no longer this caller's empty reservation but the live update's restored routing metadata, and deleting it would silently lose that user's completion notice. So _release_update_slot only removes the marker while it is still zero-length — the state the exclusive create leaves it in, and one nothing else produces:

def _release_update_slot(pending_path: Path) -> None:
    try:
        if pending_path.stat().st_size == 0:
            pending_path.unlink()
    except OSError:
        pass

Two regression tests, both of which fail against the previous commit 58ac06ac6:

  • tests/gateway/test_update_command.py::TestUpdateAdmissionControl::test_rejects_when_notifier_claims_a_live_update_mid_admission — writes the claimed marker from inside the admission window to stand in for the concurrent notifier, then asserts no spawn, an "already running" reply, the reservation handed back, and the live update's claimed marker untouched.
  • tests/gateway/test_update_command.py::TestUpdateAdmissionControl::test_release_update_slot_keeps_a_marker_that_holds_metadata — pins the zero-length condition, so a marker carrying metadata survives a release.

tests/gateway/ is green apart from seven failures that are identical on clean origin/main (test_discord_send, test_send_multiple_images, test_session_store_prune, test_shutdown_forensics, test_systemd_notify) and are order-dependent — they pass in isolation.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correction to the SHA above: d15ce760e was orphaned when this branch was rebased onto current main. The fix is live on head as 0e2af35db73 ("fix(gateway): re-check the claimed marker after reserving the update slot"), sitting on 2ddeed6dcb6 ("fix(gateway): admit exactly one /update per profile").

The mechanism is unchanged. The claimed_path.exists() pre-check is no longer load-bearing on its own — after _claim_update_slot wins the exclusive create, the claimed marker is re-checked, and if the notifier's pending -> claimed rename landed in the window you identified, the reservation is handed back and the caller takes the loser path:

admitted = not claimed_path.exists() and _claim_update_slot(
    pending_path, _UPDATE_RESERVATION_TTL_S
)
if admitted and claimed_path.exists():
    _release_update_slot(pending_path)
    admitted = False

_release_update_slot only unlinks the marker while it is still the empty file the claim created, so a marker the notifier has already restored via claimed_path.replace(pending_path) — someone else's routing metadata — is left alone rather than deleted.

Rebase-proof anchor, since the SHA will move again: the regression is test_rejects_when_notifier_claims_a_live_update_mid_admission in tests/gateway/test_update_command.py, which drives the rename precisely between the check and the create. Grep that name rather than the commit id.

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/gateway Gateway runner, session dispatch, delivery area/install-update Installer, updater, packaging, wheels, doctor sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Aug 5, 2026
`/update` wrote `.update_pending.json` and spawned a detached
`hermes update --gateway` with no admission control of any kind. Two
concurrent invocations — a user double-tapping because the update runs
for minutes with no acknowledgement, or two platforms on one multiplexed
gateway both triggering it — each wrote the marker and each spawned an
updater against the same checkout and virtualenv. The second write also
replaced the first requester's routing metadata, so that user never
learned their update finished. Both handlers also stage through the same
`.update_pending.tmp` path, so the losing rename raises FileNotFoundError
out of the handler.

Reserve the profile-wide slot with os.open(O_CREAT | O_EXCL) before any
routing metadata is written and before the updater is spawned. That
collapses observe-and-claim into a single atomic syscall, so exactly one
caller can ever win. The losing caller gets an "update already running"
reply and still gets `_schedule_update_notification_watch()`, so it
learns the outcome of the update that is actually running.

The reservation carries a TTL: an updater killed before it can write its
exit code (host reboot, OOM kill) would otherwise wedge /update for the
lifetime of the profile. Any failure between the claim and the spawn
releases the slot immediately so the user can retry.

Supersedes NousResearch#15539, which identified this collision. That guard patched
`gateway/run.py`, where the handler no longer lives after 619bd78, and
used a check-then-write `if pending_path.exists(): return` that both
callers can pass; its tests pre-seeded the marker, so they never
exercised the check-to-create gap.
…slot

The `claimed_path.exists()` pre-check was not synchronized with the
exclusive create. `_send_update_notification` renames
`.update_pending.json` -> `.update_pending.claimed.json` while an update
is in flight; landing that rename between the check and the create leaves
`pending` momentarily absent, so `_claim_update_slot` succeeds and admits
a second updater against the live one. The notifier's later
`claimed.replace(pending)` then clobbers the new reservation.

Re-check the claimed marker immediately after a successful claim and hand
the reservation back if it is present. `_release_update_slot` only removes
the marker while it is still the empty file the claim created, so a marker
the notifier has already restored — someone else's routing metadata —
survives instead of being deleted.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/install-update Installer, updater, packaging, wheels, doctor comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants