Skip to content

feat(api-server): bind clarify responses on /v1/runs - #89880

Open
meiqinsi wants to merge 4 commits into
NousResearch:mainfrom
meiqinsi:feat/api-server-runs-clarify
Open

meiqinsi wants to merge 4 commits into
NousResearch:mainfrom
meiqinsi:feat/api-server-runs-clarify

Conversation

@meiqinsi

@meiqinsi meiqinsi commented Aug 19, 2026 •

Copy link
Copy Markdown

What does this PR do?

Makes /v1/runs support the existing clarify tool as a real hard gate: the agent thread blocks on wait_for_response, clients get a bounded SSE clarify.request, and they answer via POST /v1/runs/{run_id}/clarification.
This salvages #68105 onto current main (keeps /steer, session-model lock, and stop semantics), then adds:

  1. Pass multi_select through the Runs callback → register → SSE prompt, and accept response.type: "choices".
  2. Stamp process-local awaiting_user on run status while clarification is pending; clear on answer, timeout, /stop, or cleanup.
    Clarify stays a Runs-only overlay (enable_clarify=True + callback) — it is not added to the default hermes-api-server toolset.
    Scope is ask/answer on /v1/runs only; defer/park/wake is out of scope.

Related Issue

Related to #2971 and #68105 (credit: @hazeion).

Fixes #

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Changes Made

  • gateway/platforms/api_server.py — Runs clarify callback, SSE clarify.request / clarify.responded, POST /v1/runs/{run_id}/clarification, waiting_for_clarification as in-flight, multi_select + awaiting_user
  • tools/clarify_gateway.py — session-bound get_pending_by_id / resolve_gateway_clarify
  • tests/gateway/test_api_server_runs.py — ask/answer, binding, multi-select, stop-while-waiting, awaiting_user enter/exit
  • tests/gateway/test_api_server.py / test_api_server_toolset.py / tests/tools/test_clarify_gateway.py — capabilities + overlay + gateway binding coverage
  • website/docs/user-guide/features/api-server.md (+ programmatic integration endpoint list) — document clarification API

How to Test

  1. scripts/run_tests.sh tests/gateway/test_api_server_runs.py
  2. scripts/run_tests.sh tests/gateway/test_api_server_toolset.py tests/tools/test_clarify_gateway.py tests/gateway/test_api_server.py
  3. Manual: start a run that calls clarify → SSE clarify.request → GET /v1/runs/{id} shows waiting_for_clarification + awaiting_user: true → POST .../clarification resumes the same turn
  4. Manual: multi_select=true → response.type=choices with choice_ids → callback receives a JSON array string
  5. Manual: /stop while waiting clears pending clarify and does not hang the agent forever

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run scripts/run_tests.sh on the touched suites and all tests pass
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: macOS (darwin)

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — or N/A
  • I've updated cli-config.yaml.example if I added/changed config keys — or N/A
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — or N/A
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide — or N/A
  • I've updated tool descriptions/schemas if I changed tool behavior — or N/A

Screenshots / Logs

N/A — covered by gateway unit/integration tests around Runs clarification SSE + POST.

@alt-glitch alt-glitch added type/feature New feature or request P3 Low — cosmetic, nice to have comp/gateway Gateway runner, session dispatch, delivery sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Aug 19, 2026

@andrexibiza andrexibiza left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head 31fe07a9c0daf770dabf97aabf60117b09883cc6 against base/current main 9162ea6db1fe0f57d6fc4de5120fac5c5a1938be. I inspected the Runs lifecycle changes, clarify_gateway binding primitive, stop/orphan cleanup, run/session status updates, exact-head tests/CI state, the source PR #68105, the now-merged batch-clarify contract in #89467, and the adjacent open Runs work in #89748 and #89754.

The core direction is sound. Using run_id as the clarify queue session key mirrors the approval queue's per-run isolation, the response endpoint validates stable server-issued choice IDs rather than trusting labels, free text is bounded and never echoed in clarify.responded, secret redaction happens before HTTP egress, and stop/cleanup release the blocking Event.wait rather than leaving an unanswerable agent thread parked. Preserving #68105 / @hazeion as the substantive source is also the right provenance model; this PR is a salvage/superseder, not a clean-room replacement.

I found two merge blockers at the ownership/state boundaries.

1. The new clarification mutation is authenticated to the URL-selected profile, but not authorized to the run's owning profile

The API listener is multiplexed: /p/<profile>/... sets _api_request_profile, enters that profile's runtime scope, and _expected_api_key() explicitly returns the API key authorized for that URL-selected profile. That gives the request a concrete profile principal.

But Runs are stored in adapter-global maps keyed only by run_id, and _handle_run_clarification() does not compare that principal with the profile that created the run. In fact the queued run status currently records created_at, session_id, and model, but not request_profile. The new handler does:

run_id = request.match_info["run_id"]
if run_id not in self._run_statuses: ...
...
clarify_session_key = self._run_clarify_sessions.get(run_id)
pending = clarify_gateway.get_pending_by_id(request_id, session_key=clarify_session_key)
...
clarify_gateway.resolve_gateway_clarify(...)

So a request authenticated as profile B can address a run created under profile A if it knows A's run_id + clarification request_id; the queue binding proves which run, but never which profile principal owns that run. The mutation then resumes A's blocked agent from B's authenticated URL scope.

This is an inherited gap from #68105's handler, not new authorship by @meiqinsi, but this salvage is the point where the endpoint would enter current main. Please bind every run to its creation profile and fail closed before reading or mutating clarification state when the current request profile differs. A regression witness should create a run through /p/alpha/v1/runs with alpha's API key, capture its pending clarification, then POST that exact run/request pair through /p/beta/... using beta's valid key and prove it cannot resolve or alter alpha's wait. The positive alpha→alpha path should still resume normally.

This also exposes the broader defect-class question for the existing run-control family (GET, events, approval, steer, stop): they use the same global run_id namespace. I am not asking this PR to silently claim class closure unless the whole family is audited, but the new mutating clarification endpoint should not deepen a known profile-ownership hole. #89754's idempotency work is adjacent evidence of the same ownership axis: run admission bookkeeping also needs profile scope.

2. _session_awaiting_user: Dict[str, bool] loses legitimate concurrent waiters

Current main already states the load-bearing invariant in _handle_runs: client-provided session_id is a conversation scope, not a run/authorization namespace, and multiple concurrent Runs can intentionally share it. That is exactly why approvals key their blocking state by run_id.

This PR gets the clarify request queue right (clarify_session_key = run_id) but projects the new session-level state into a single bit:

self._session_awaiting_user: Dict[str, bool] = {}
...
self._session_awaiting_user[session_id] = True
...
self._session_awaiting_user.pop(session_id, None)

If runs A and B share session s and both are waiting for clarification, A answering, timing out, stopping, or merely finishing cleanup removes s from the map while B is still waiting_for_clarification. The per-run status for B remains truthful, but the session-level contract this PR is explicitly adding becomes false. The same textual session_id can also exist under separate multiplexed profile homes/state DBs, so the raw-string key collapses profile ownership as well as run multiplicity.

Please make session waiting state ownership-preserving: e.g. (profile, session_id) -> set[run_id/request_id] / refcount, or derive it from the owned per-run state rather than maintaining a lossy boolean. Add a witness with two concurrent runs sharing one session: put both into clarify wait, resolve/stop/timeout one, and assert the session still reports awaiting-user until the second waiter exits. Add the cross-profile same-session-id twin if the session map remains adapter-global.

Composition / merge order

  • #68105 / @hazeion — substantive source. This PR correctly credits and supersedes it; the profile-authorization blocker above is inherited from that implementation and should be repaired without losing attribution.
  • #89467 / @ethernet8023 — merged current-main clarify contract. It added questions batches. This Runs callback does not accept questions, so the core deliberately treats it as a legacy callback and decomposes a batch into sequential prompts; that is compatible, not a blocker. Please add one composed batch witness, though, because this is the exact kind of current-main semantic that a salvage can miss.
  • #89748 / @RoySRose — complementary event-ordering fix on the same Runs/SSE seam. If it lands first, this branch should rebase and ensure clarify.request / clarify.responded obey the same ordering primitive rather than recreating a second cross-thread enqueue rule.
  • #89754 / @RoySRose — complementary run/approval idempotency work on the same file/tests. It needs composition with clarification admission/control ownership rather than a textual conflict-only rebase.
  • #2971 — remains the umbrella request. This PR closes ask/answer for Runs, but correctly does not claim defer/park/wake or the full resumable-interaction class.

Verification / CI

The branch is exactly on current main (no base drift at review time). The PR reports its touched Runs/API/toolset/clarify suites green on macOS and the new tests cover exact request binding, single-use resolution, multi-select shape, auth, redaction/bounds, and stop release.

Repository-hosted exact-head evidence is currently incomplete: the CI workflow for 31fe07a... ended in failure before creating any jobs, while Nix/Docker and label-rerun workflows are action_required; fetching the CI run returns an empty job list. So there is no executed exact-head matrix to independently corroborate the local suite yet.

Re-review gate: bind run control to the creating profile; replace the session-level boolean with multiplicity/profile-safe ownership and add the concurrent-waiter witness; add a current-main batch-clarify composition test; then obtain an executed exact-head CI matrix after the repository workflow issue clears.

@meiqinsi

Copy link
Copy Markdown
Author

@andrexibiza Thanks for the review. Addressed the two blockers plus the #89467 composition test.

  1. Profile ownership on clarification. Each run now records request_profile at create time. _handle_run_clarification fails closed with 404 run_not_found before reading or resolving pending state when the URL profile is not the owner. Witness: /p/alpha creates and waits; /p/beta with a valid beta key cannot resolve the same run_id/request_id; alpha→alpha still resumes. This PR does not claim the same gate for GET / events / approval / steer / stop.

  2. Session awaiting bool. Chose derive-from-per-run and dropped _session_awaiting_user; there is no session-level awaiting report. Waiting truth stays on the run (status / awaiting_user / pending clarify). Same session_id, two concurrent runs in clarify: resolving one leaves the sibling run still awaiting_user: true with its pending intact. No adapter-global session map remains, so the cross-profile same-session_id twin does not apply.

  3. feat(clarify): ask multiple independent questions in one call #89467 batch. Runs callback is still the legacy (no questions) form, so a questions batch is sequential prompts on the same run. Witness: two clarify.request events and two POSTs, distinct request_ids; a spent id is 409.

Will rebase if #89748 / #89754 land first. Exact-head CI still depends on the repo workflow issue you noted.

@andrexibiza andrexibiza left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed exact head dd38b401ced229ea5a5c024446aaa20b6cca32f9 and GitHub's current-main test merge 2feff65e2ede5f75c65ec5d5503ef873cbf1468c (f43eabee5f36e11448086ee8ee17c499958e81bf + this head).

The two code blockers and the requested #89467 composition witness are closed. I found no remaining code blocker from my prior review.

  • Profile ownership: run admission records request_profile; clarification checks the URL-selected principal before any pending-state lookup or mutation and returns the indistinguishable 404 run_not_found on mismatch. The alpha/beta regression proves beta's valid key cannot resolve alpha's exact run/request pair, leaves the pending entry intact, and preserves the alpha→alpha positive path.
  • Concurrent wait ownership: the lossy adapter-global _session_awaiting_user boolean is gone. Waiting truth is per run, and the same-session_id concurrency witness proves resolving one run does not clear or consume its sibling's wait.
  • Merged batch contract: the real clarify_tool(..., questions=[...]) path is exercised as sequential prompts on one run with distinct request IDs; replay of the spent first ID fails with 409 before the second prompt is answered.

#89748 and #89754 are still open, so their prospective composition condition has not triggered. The synthesized test merge is clean against current main, and I found no new semantic conflict in the composed Runs lifecycle.

The only remaining gate is repository-hosted verification. The exact-head CI/Docker/Nix workflows are action_required; the CI run references test-merge SHA 2feff65e... but created no jobs. Do not represent the matrix as green or merge until a maintainer approves/runs the fork workflows and the exact test-merge checks execute successfully.

Non-blocking housekeeping: the PR description still says awaiting_user is stored in a session map. That sentence is stale; the implementation correctly removed that map in favor of per-run truth.

Copy link
Copy Markdown

Re-reviewed current head dd38b401ced229ea5a5c024446aaa20b6cca32f9 against the two blockers and the #89467 composition requirement from my earlier review.

The requested source repairs are closed:

  1. Clarification ownership now fails closed by creating profile. Each run records request_profile; _handle_run_clarification() calls _run_profile_forbidden() immediately after authentication and before parsing, reading, or resolving pending clarification state. The multiplex witness creates under /p/alpha, proves a valid /p/beta principal receives indistinguishable 404 run_not_found without altering the pending wait, and proves alpha→alpha still resumes.
  2. The lossy session-level awaiting bit is gone. Waiting truth is per run only. The same-session_id, two-run witness proves resolving one run leaves its sibling in waiting_for_clarification with awaiting_user: true and its pending request intact.
  3. Current-main batch composition is pinned. The questions form is decomposed through the legacy callback into sequential prompts on the same run, with distinct request IDs; replay of the spent first ID returns 409 and both answers compose into the final batch result.

I also checked the request-id/choice validation, exact run binding in clarify_gateway, single-use resolution, stop/orphan cleanup, bounded/redacted prompt projection, and the current adjacent topology. #89748 and #89754 are still open, so no already-landed composition has been missed on this head; whichever lands first will still require the declared rebase rather than conflict-only resolution.

Disposition: my two source blockers are cleared on dd38b401. I do not see a remaining clarification-lifecycle blocker in the current diff. The only unresolved merge gate is external execution: exact-head CI 32377400853, Docker 32377399894, and Nix 32377399929 are all completed/action_required, so GitHub has not run a hosted matrix for this commit.

simeiqin and others added 4 commits September 8, 2026 06:49
(cherry picked from commit 6c7fdfc)
Co-authored-by: Cursor <cursoragent@cursor.com>
Port multi_select, awaiting_user isolation, profile binding, and sequential
batch coverage onto api_server_runs (implementation already in prior commit)
plus matching tests/docs. Original PR commits consolidated during rebase
onto main's extracted runs module.

Co-authored-by: Cursor <cursoragent@cursor.com>
Split the RequestKey import so older aiohttp releases do not null out
`web` and break /v1/runs handlers.

Co-authored-by: Cursor <cursoragent@cursor.com>
The clarification route delegated to a missing _handle_run_clarification
during the main rebase, which 500'd POSTs and left clarify waiters hung
until the suite timeout. Wire the handler back and adapt unit tests to
main's run-ownership stamps.

Co-authored-by: Cursor <cursoragent@cursor.com>
@meiqinsi
meiqinsi force-pushed the feat/api-server-runs-clarify branch from ed17637 to 472f440 Compare September 7, 2026 22:50

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/gateway Gateway runner, session dispatch, delivery P3 Low — cosmetic, nice to have sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants