Skip to content

fix(update): bound the update prompts so an unattended run cannot hang - #92410

Closed
jackulau wants to merge 2 commits into
NousResearch:mainfrom
jackulau:fix/92303-bounded-update-prompts
Closed

jackulau wants to merge 2 commits into
NousResearch:mainfrom
jackulau:fix/92303-bounded-update-prompts

Conversation

@jackulau

Copy link
Copy Markdown

What does this PR do?

hermes update can block forever on a prompt nobody is able to answer. The
report is a Windows Scheduled Task with -WindowStyle Hidden: 44 minutes parked
on Restore local changes now? [Y/n], ending in exit code 0x40010004, which is
an external kill rather than an exit. The update neither completed nor failed.

Why every existing guard misses it. A hidden console is still a real
console. sys.stdin.isatty() returns True, so _non_interactive_update and
the prompt_for_restore predicate both classify the run as interactive:

prompt_for_restore = (
    auto_stash_ref is not None
    and not assume_yes
    and (gateway_mode or (sys.stdin.isatty() and sys.stdout.isatty()))
)

And the handle is open, so input() never raises EOFError either. The
existing except (EOFError, UnicodeDecodeError) guard is not wrong, it is simply
never reached. There is no property of the process that distinguishes "console
a human is watching" from "console nobody will ever type into".

So the prompts are bounded rather than predicted. An unanswered prompt takes
its documented default, says so in the log, and latches the run as unattended.
The latch is the interesting part: a prompt that went unanswered is an
observation that nobody is at the keyboard, which is strictly stronger than
anything isatty can infer, so later prompts in the same run stop asking.

The route I did not take

My first instinct was GetConsoleWindow() + IsWindowVisible(): a console
nobody can see is a console nobody can type into. I dropped it because ConPTY
hosts (Windows Terminal, the VS Code terminal) are reported to leave a hidden
console window on the process, which would misclassify a large fraction of
Windows users as unattended, and the failure would be silent: their local
changes would just stop being offered for restore.

I could not verify that ConPTY behaviour from where I was working (with stdin
redirected, GetConsoleWindow() returns 0 and isatty() is already False,
which is the case the existing guards already handle). I would rather not ship a
silent-failure detector on a behaviour I am taking on trust. Flagging it in case
a maintainer knows it cold and prefers that route.

Two calls I am handing to a maintainer rather than making quietly

  1. The timeout value, 300s. Long enough that a human reading a four-line
    prompt is never affected, short enough that "forever" becomes "five minutes".
    It is a judgement call, not a derived number. Happy to change it, or to put it
    behind config/env if you want it tunable; I did not add a config key because
    that is a new public surface for what is currently one constant.
  2. What a timeout should default to. I chose skip-restore, matching the
    EOFError path that already exists. It is the only default that cannot lose
    work: the stash stays on disk and the existing git stash apply <ref>
    guidance still prints. The argument for the other side is real, though, and I
    want it on the record: [Y/n] means a human pressing enter gets restore, so
    an argument exists that a timeout should match the visible default. I think
    "unattended" is a different situation from "pressed enter" and should be
    allowed a more conservative answer, but this is your semantics to set.

Related Issue

Fixes #92303

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Changes Made

All in hermes_cli/update_cmd.py.

  • _read_line_with_timeout(default, timeout=None, read_fn=None) (new).
    Reads one line on a daemon thread and gives up after timeout, returning
    (response, timed_out). timeout <= 0 restores today's unbounded blocking
    read, so the old behaviour stays reachable from a test rather than only from a
    revert. daemon=True is load bearing: the reader stays parked on stdin for the
    life of the process, and a non-daemon thread there would block interpreter
    shutdown, which is the same unbounded wait wearing a different hat.
  • _prompt_proven_unattended (new) plus _reset_prompt_interactivity(),
    called at the top of _cmd_update_impl so the latch scopes to one update
    rather than to the interpreter. It also closes a hazard the timeout would
    otherwise open: without it, a later prompt could arm a second reader on the
    same stdin and have its answer swallowed by the first, still-parked one.
  • _report_unanswered_prompt (new). Names the prompt, the consequence, and
    --yes / --keep-stash, so an unattended transcript explains itself instead
    of ending at a bare prompt line. This is the reporter's option 3, folded in.
  • _restore_stashed_changes: the reported site, now bounded. Default "n",
    identical to the EOFError path it replaces.
  • _sync_with_upstream_if_needed: bounded too, and this one is worth calling
    out because it was not in the report. It runs at line 6046 while the stash
    restore runs at 6053, so on a fork with no upstream remote an unattended
    run hangs here and never reaches the reported hang at all. Fixing only what
    was reported would have left the same bug one prompt earlier.
  • The config-migration prompt now reads
    elif _prompt_proven_unattended or not (sys.stdin.isatty() and ...). That is
    deliberately routed into the branch that already exists rather than given its
    own: the "auto" default and everything downstream are unchanged, so the only
    new thing is a better signal. This is the generalisation asked for at the end
    of the report, without changing any prompt's semantics under cover of a stash
    fix.

How to Test

pytest tests/hermes_cli/test_update_prompt_timeout.py -q
# 16 passed

16 new tests in a new file, covering three things: the bound exists and takes the
safe default; the answered paths are unchanged, including the EOFError,
UnicodeDecodeError and KeyboardInterrupt paths that predate this fix; and the
latch.

Sabotage proof. Each mutation applied to hermes_cli/update_cmd.py alone,
with the new suite re-run:

Mutation Tests that fail
Remove the bound entirely (back to a blocking read) test_an_unanswerable_prompt_gives_up, test_the_reader_thread_cannot_keep_the_process_alive, test_a_timeout_proves_the_run_unattended, test_a_later_prompt_does_not_ask_again, test_the_reported_hang_now_ends, test_it_stops_waiting_too
Time out, but do not latch the run unattended test_a_timeout_proves_the_run_unattended, test_a_later_prompt_does_not_ask_again
daemon=False on the reader thread test_the_reader_thread_cannot_keep_the_process_alive
Return "" on timeout instead of the default test_an_unanswerable_prompt_gives_up, test_it_stops_waiting_too
Print the blank line on the answered path instead of the exception path test_eof_keeps_the_blank_line_it_always_printed

Two rows worth a second look.

Row 4 is the reason the helper returns an explicit default rather than
"". Both call sites test response in {"", "y", "yes"}, so an empty string
does not mean "no answer", it means yes. A timeout that returned "" would
have silently added an upstream remote and silently restored a stash on exactly
the unattended runs this PR exists to make safe, and every "did it time out?"
assertion would still have passed.

Row 5 is a spacing inversion, and it is in here because I made it. Routing
that prompt through the shared helper made it natural to print its blank line on
the answered path instead of the exception path, which is backwards and which
no assertion about a return value would ever notice.

Existing suites. test_update_yes_flag.py, test_update_autostash.py and
test_update_streaming.py: 15 passed, 7 failed, 1 skipped, and the same 7 fail
identically on the unpatched tree
. They are Windows-only console-encoding
crashes (UnicodeEncodeError: 'charmap' codec can't encode character '✓'
out of a print()), unrelated to this change. Verified by stashing only
hermes_cli/update_cmd.py and re-running the same command.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run pytest tests/ -q and all tests pass
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: Windows 11 (26200), Python 3.13

Leaving that box honest rather than ticked: pytest tests/ -q does not
collect on Windows at all. tests/hermes_cli/test_doctor_journal_modes.py raises
AttributeError: module 'os' has no attribute 'geteuid' at import and aborts the
run. What I did run is the new suite plus every existing suite that touches these
functions, with a stashed baseline for the pre-existing failures, as above. Linux
CI on this PR is the authority for the full-tree line.

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings): the new
    helpers carry the reasoning, including why daemon=True matters and why
    the latch exists
  • I've updated cli-config.yaml.example if I added/changed config keys: N/A.
    The timeout is a module constant, deliberately not a config key yet (see
    the maintainer calls above)
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows: N/A
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide:
    the whole point is that it is not platform-specific. threading plus a
    bounded join behaves the same everywhere, which is exactly why I rejected
    the Win32 console-handle route. The bug is reported on Windows but the same
    shape reaches CI and remote automation on any OS whenever stdin is a live
    handle nobody answers
  • I've updated tool descriptions/schemas if I changed tool behavior: N/A

Screenshots / Logs

Before, from the report:

⚠ Local changes were stashed before updating.
  Restoring them may reapply local customizations onto the updated codebase.
  Review the result afterward if Hermes behaves unexpectedly.
Restore local changes now? [Y/n]
<44 minutes of nothing, then exit code 1073807364>

After:

⚠ Local changes were stashed before updating.
  Restoring them may reapply local customizations onto the updated codebase.
  Review the result afterward if Hermes behaves unexpectedly.
Restore local changes now? [Y/n]

⚠ No answer to 'Restore local changes now?' after 300s, so this run is being treated as unattended.
  Not restoring. Your changes are preserved in git stash.
  Pass --yes (and --keep-stash where you want the stash left parked) to skip these prompts in scripted runs.
Skipped restoring local changes.
Your changes are still preserved in git stash.
Restore manually with: git stash apply stash@{0}

`hermes update` blocks forever on a prompt nobody can answer. The report is a
Windows Scheduled Task with `-WindowStyle Hidden`: 44 minutes parked on
"Restore local changes now? [Y/n]", ending in exit code 0x40010004, an external
kill rather than an exit.

Every existing guard misses it, and not by accident. A hidden console is a real
console, so `sys.stdin.isatty()` is True and both `_non_interactive_update` and
the `prompt_for_restore` predicate classify the run as interactive. The handle
is open, so `input()` never raises EOFError either; the existing
`except (EOFError, UnicodeDecodeError)` guard is not wrong, it is unreachable.
No property of the process separates "console a human is watching" from
"console nobody will ever type into".

So bound the prompts instead of predicting them. An unanswered prompt takes its
documented default, says so, and latches the run as unattended, which is an
observation rather than the inference every isatty check makes.

- `_read_line_with_timeout` reads on a daemon thread and gives up after the
  timeout, returning (response, timed_out). `timeout <= 0` restores today's
  blocking read, so the old behaviour is reachable from a test rather than only
  from a revert. daemon=True is load bearing: the reader stays parked on stdin
  for the life of the process, and a non-daemon thread there would block
  interpreter shutdown, which is the same unbounded wait wearing a different
  hat.
- `_prompt_proven_unattended`, reset at the top of `_cmd_update_impl` so it
  scopes to one update rather than to the interpreter. It also closes a hazard
  the timeout opens: without it a later prompt could arm a second reader on the
  same stdin and have its answer swallowed by the first, still-parked one.
- `_report_unanswered_prompt` names the prompt, the consequence and `--yes`, so
  an unattended transcript explains itself instead of ending at a bare prompt.
- `_restore_stashed_changes`: the reported site. Default "n", identical to the
  EOFError path it replaces, and safe because the stash stays on disk with its
  existing `git stash apply` guidance.
- `_sync_with_upstream_if_needed`: bounded too, though it was not reported. It
  runs before the stash restore, so on a fork with no upstream remote an
  unattended run hangs there and never reaches the reported hang at all.
- The config-migration prompt consults the latch through the non-interactive
  branch it already has. Its "auto" default and everything downstream are
  unchanged; only the signal feeding the branch is new.

Deliberately not using `GetConsoleWindow()` + `IsWindowVisible()`. ConPTY hosts
are reported to leave a hidden console window on the process, which would
misclassify Windows Terminal and VS Code users as unattended, and the failure
would be silent. I could not verify that behaviour, and a silent-failure
detector is not something to ship on trust.

Two choices are named for a maintainer rather than made quietly: the 300s value,
and whether a timeout should default to skip-restore (chosen, cannot lose work)
or to restore (matching the `[Y/n]` default a human gets by pressing enter).

16 regression tests. Sabotage proof, five mutations, each caught: removing the
bound fails 6; not latching fails 2; daemon=False fails 1; returning "" instead
of the default fails 2; inverting the blank-line spacing fails 1. The fourth is
why the helper returns an explicit default: both call sites test
`response in {"", "y", "yes"}`, so "" does not mean "no answer", it means yes.

Existing update suites: 15 passed, 7 failed, and the same 7 fail identically on
the unpatched tree (Windows-only console-encoding crashes in `print()`).

Fixes NousResearch#92303
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/cli CLI entry point, hermes_cli/, setup wizard platform/windows Native Windows-specific behavior or breakage sweeper:risk-platform-windows Sweeper risk: may break or behave differently on native Windows sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Aug 22, 2026
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference, author can ignore or act on any point.

Reviewed hermes_cli/update_cmd.py (+478/-16) and the new 310-line test file.

What's good

  • Correct diagnosis and honest framing: a hidden Windows console is a real console (isatty() True, handle open, no EOFError ever), so no up-front predicate can distinguish "human watching" from "nobody will ever type". Bounding prompts and treating silence as evidence is the portable answer.
  • The implementation details are unusually well-reasoned: daemon=True documented as load-bearing (a non-daemon reader parked on input() would recreate the hang at interpreter exit), the unattended latch prevents a second reader from having its answer swallowed by an abandoned first one, the latch resets per update run, and timeout<=0 keeps the legacy path reachable from tests. Logging the consequence plus the --yes hint turns the transcript into an explanation instead of a cliff.
  • Treating an unanswered prompt as stronger evidence than any isatty inference is a genuinely nice epistemic upgrade over heuristic guards.

Suggestions

  1. Coordinate with fix(update): never block on prompts when stdin is unattended #92448 (open, same file, same symptom): it adds _prompt_answerable() TTY-guards to these exact two functions (_restore_stashed_changes, _sync_with_upstream_if_needed). The approaches are complementary — avoid asking when detectably unattended; bound the ask otherwise — but whichever lands second needs a deliberate rebase, and running both could double-report ("no interactive terminal" + "no answer after 300s"). Worth a note on either PR.
  2. After a timeout, the abandoned daemon thread still holds the console's stdin buffer; a subsequent subprocess (git) inheriting stdin could theoretically consume the user's late-typed answer as git input. Harmless in unattended runs by definition, but worth one comment acknowledging the abandoned-reader lifetime.
  3. Nit: _report_unanswered_prompt hardcodes %ds against the module constant while taking what/consequence as parameters — pass the timeout through so a future per-prompt override stays honest.

Strong engineering; #1 is process, not product, but it matters here.

Review follow-up, two points.

_report_unanswered_prompt read the module constant while _read_line_with_timeout
already accepted a per-prompt timeout, so the first call site to override the
bound would have printed a number the run never waited. It now takes the same
optional timeout and resolves it the same way, with two tests pinning the
default and the override apart; the override test fails against the old
hardcoded form.

Also documented the abandoned reader's lifetime at the point where it is
abandoned. A thread parked on input() cannot be cancelled portably, so it keeps
its claim on stdin for the life of the process and a late-typed line races
between it and any git subprocess inheriting the same handle. That is accepted
rather than fixed, and the comment says why: the run has just produced evidence
that nobody is typing, and sending every later subprocess to DEVNULL on the
strength of one timeout would change how git behaves in a run that merely
paused.
@jackulau

Copy link
Copy Markdown
Author

All three acted on. 96836bed5a.

1. #92448. Noted on that PR rather than only here, since coordination that lives on one side of an overlap is not coordination. #92448 (comment)

Two things came out of reading their diff properly.

The double-report worry does not survive the merge, as far as I can tell. Their guard sits ahead of the read in both functions: _restore_stashed_changes gains an elif not _prompt_answerable(): arm before the input path, so a non-TTY short-circuits and no bounded reader ever arms, while a TTY skips their arm and leaves only the bound to speak. _sync_with_upstream_if_needed is the same shape as if _prompt_answerable(): <ask> else: <skip>, which merges with the ask being bounded instead of blocking. One message on every path.

The more useful finding is that the two approaches do not overlap as much as the file diff suggests. _prompt_answerable() requires both stdin and stdout to be TTYs, which is exactly the inference that fails for the reported configuration: -WindowStyle Hidden gives a real console that is simply not drawn, so both are True and the guard passes the run through to the same blocking input(). Their patch covers the console-less shapes (a task registered "run whether user is logged on or not", CI with redirected stdin) immediately and at zero risk, which is a real population and worth having. It just is not the hidden-window one, though their docstring currently names it, which I flagged.

So: complementary, and I offered to do the rebase from either side depending on which lands first. If theirs goes in ahead of mine, #92410 reduces to the bounding with their guard kept as the fast path.

2. Abandoned reader. Documented at the point where the reader is abandoned. The comment states the race you describe (the parked thread keeps its claim on stdin for the life of the process, so a late-typed line races between it and any git subprocess inheriting the same handle) and says it is accepted rather than fixed, with the reason: the run has just produced evidence that nobody is typing, and routing every later subprocess to DEVNULL on the strength of one timeout would change how git behaves in a run that merely paused. That is a bigger claim than one unanswered prompt supports.

3. Not a nit. _read_line_with_timeout already takes a per-prompt timeout, so the hardcode was not merely inelegant, it was a message that would start lying the first time a call site used the parameter that already exists. _report_unanswered_prompt now takes the same optional timeout and resolves it identically, and two tests pin the default and the override apart. The override test fails against the old form with after 300s where it expects after 12s, so it is a real guard rather than a restatement.

18 passed; ruff clean.

salch-cred added a commit to salch-cred/hermes-agent that referenced this pull request Aug 24, 2026
The hidden-window case (-WindowStyle Hidden) gives a real console so isatty() returns True and this guard does not fire. Document the limitation honestly and point at the bounded-wait approach (NousResearch#92410) as the complement.
ethernet8023 pushed a commit that referenced this pull request Aug 28, 2026
_sync_with_upstream_if_needed called bare input() with no assume_yes parameter and no tty check, so a fork checkout without an upstream remote wedged hermes update forever in any non-interactive context (CI, cron, the desktop updater hand-off): stdin stays open, EOFError never fires. Thread assume_yes and the gateway input_fn into the helper and skip the prompt as a decline under assume_yes or a non-tty stdio pair, without writing the decline marker or touching git remotes, so interactive runs still get asked later. Both call sites forward the interaction state; the config-migration and stash-restore prompts already carry this gate.

Closes #60240 (prompt half). Supersedes #78678, #92448, #92410.
Co-authored-by: BlackishGreen33 <BlackishGreen33@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: jackulau <jackulau@users.noreply.github.com>
@yoniebans

Copy link
Copy Markdown

The fork-upstream prompt-hang portion of this is superseded by #97052, now merged (00bbfc6): assume_yes plus a TTY gate covers the unattended populations there without reader threads. The residual case this PR also bounds (a console that reports a TTY pair, no --yes, nobody present) and the timeout bounding of the other prompts are intentionally left open by #97052, so closing this as superseded for the prompt that was hanging in practice. You're credited on the merged commit.

@yoniebans yoniebans closed this Aug 31, 2026
melon-xf added a commit to melon-xf/hermes-agent that referenced this pull request Sep 3, 2026
_sync_with_upstream_if_needed called bare input() with no assume_yes parameter and no tty check, so a fork checkout without an upstream remote wedged hermes update forever in any non-interactive context (CI, cron, the desktop updater hand-off): stdin stays open, EOFError never fires. Thread assume_yes and the gateway input_fn into the helper and skip the prompt as a decline under assume_yes or a non-tty stdio pair, without writing the decline marker or touching git remotes, so interactive runs still get asked later. Both call sites forward the interaction state; the config-migration and stash-restore prompts already carry this gate.

Closes NousResearch#60240 (prompt half). Supersedes NousResearch#78678, NousResearch#92448, NousResearch#92410.
Co-authored-by: BlackishGreen33 <BlackishGreen33@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: jackulau <jackulau@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cli CLI entry point, hermes_cli/, setup wizard P2 Medium — degraded but workaround exists platform/windows Native Windows-specific behavior or breakage sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-platform-windows Sweeper risk: may break or behave differently on native Windows type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

hermes update hangs forever on 'Restore local changes now? [Y/n]' when run unattended (Scheduled Task / hidden stdin)

4 participants