Skip to content

fix(mcp): heartbeat activity during long synchronous MCP/hive tool calls (t_cc4a1b4f) - #44

Merged
sahilm-ai merged 1 commit into
mainfrom
kanban/t_cc4a1b4f
Jun 3, 2026
Merged

sahilm-ai merged 1 commit into
mainfrom
kanban/t_cc4a1b4f

Conversation

@sahilm-ti

@sahilm-ti sahilm-ti commented Jun 2, 2026 •

Copy link
Copy Markdown
Owner

Problem

A single synchronous MCP/hive tool call — e.g. `ns_runreport` against the NetSuite hive — blocks the agent's tool-execution thread inside `_run_on_mcp_loop` for many minutes. The poll loop already honors user interrupts, but it never fired the thread-local activity callback the agent registers before dispatching a tool.

For dispatcher-spawned kanban workers, the agent's last-activity timestamp is bridged to the board heartbeat via `AIAgent._touch_activity` → `heartbeat_current_worker_from_env`. So a long hive call left the heartbeat stale, the dispatcher's stuck-worker watchdog (`stuck_after_seconds`, default 900s) classified the live worker as stale, and killed + re-queued it mid-call.

Observed impact (task t_c90d83c9): 17 consecutive stuck-kills (runs 803–821, heartbeat_age 1174–3077s). The task only completed by abandoning live hive fetches.

Diagnosis — which layer

The hang is in the Hermes worker transport, not the proxy or the NetSuite report engine. Evidence:

  • `tools/environments/base.py` — the terminal tool's `_wait_for_process` fires `touch_activity_if_due()` every ~10s during a long command, keeping the heartbeat alive.
  • `tools/mcp_tool.py` `_run_on_mcp_loop` (the path every MCP/hive tool call takes) polled `future.result(timeout=0.1)` in a loop checking only `is_interrupted()` — zero activity touches for the entire blocking wait.

So the terminal path was survivable and the MCP path was not, for the same wall-clock duration. That asymmetry is the bug.

Fix

Mirror the terminal path: fire `touch_activity_if_due()` on each poll iteration at a 30s cadence. 30s is well inside both the gateway inactivity window and the 15-min kanban stuck default; the underlying heartbeat bridge is itself rate-limited to one DB write per 60s, so an over-eager cadence is harmless. The import is guarded so the MCP loop still works if `tools.environments.base` is unavailable on a niche surface.

This makes one long `ns_runreport` (or any slow hive call) survivable without a per-task `stuck_after_seconds` bump. The proxy/per-tool `timeout` (default 120s, configurable per server) still fires and returns a clean `TimeoutError` — that path is unchanged.

Test

`tests/tools/test_mcp_tool.py::TestRunOnMcpLoop::test_activity_callback_fires_during_long_call` — real background event loop + a coroutine that blocks across several poll iterations; asserts the activity callback fires during the wait with the expected label and 30s cadence.

```
209 tests passed (test_mcp_tool.py + test_mcp_tool_401_handling.py + test_mcp_structured_content.py)
```

Notes / boundary

  • Infra fix in hermes-agent only. Not a BSB-agent or tool-correctness change.
  • The proxy-side timeout-enforcement gap raised in the card (cross-ref t_0b40f446 — a slow call that never returns) is a separate mcp-proxy-server concern; this PR guarantees the Hermes side stays alive and honors its own per-call timeout, so even an un-bounded upstream now fails cleanly at `tool_timeout` instead of hanging the worker to death.

AC: a kanban worker can make one long `ns_runreport` hive call without being stuck-killed; the MCP poll loop heartbeats throughout the wait.

Summary by CodeRabbit

  • Bug Fixes

    • Improved activity tracking during long-running MCP tool calls to prevent false inactivity detection.
  • Tests

    • Added regression test validating activity callback invocation while processing extended MCP operations.

@coderabbitai

coderabbitai Bot commented Jun 2, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

The PR adds activity heartbeat tracking to prevent timeout during long-running MCP tool calls. _run_on_mcp_loop() now imports and periodically invokes touch_activity_if_due() with a 30-second cadence while polling the scheduled MCP future. A regression test validates the callback fires with correct label and state bookkeeping.

Changes

Activity heartbeat during long MCP calls

Layer / File(s) Summary
Activity heartbeat implementation
tools/mcp_tool.py
_run_on_mcp_loop() imports touch_activity_if_due with defensive no-op fallback, initializes a heartbeat state with 30-second interval from call timing, and invokes the activity callback on each polling iteration inside the wait/cancel/timeout loop.
Activity callback regression test
tests/tools/test_mcp_tool.py
New test test_activity_callback_fires_during_long_call runs a real event loop thread, patches the MCP loop and activity callback to record invocations, asserts the long-running coroutine completes, and verifies the callback fires at least once with expected "waiting for MCP tool call" label and heartbeat state fields.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~12 minutes

Poem

🐰 A heartbeat flutters while the MCP runs long,
No timeout fears when the pulse stays strong,
Thirty seconds steady, callback sings its song,
Agent stays alive—activity can't go wrong! 💓

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically describes the main change: adding periodic activity heartbeating during long synchronous MCP/hive tool calls to prevent worker timeout/disconnection.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch kanban/t_cc4a1b4f

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jun 2, 2026 •

Copy link
Copy Markdown

🔎 Lint report: kanban/t_cc4a1b4f vs origin/main

ruff

Total: 0 on HEAD, 0 on base (➖ 0)

🆕 New issues: none

✅ Fixed issues: none

Unchanged: 0 pre-existing issues carried over.

ty (type checker)

Total: 9760 on HEAD, 9760 on base (➖ 0)

🆕 New issues: none

✅ Fixed issues: none

Unchanged: 5053 pre-existing issues carried over.

Diagnostics are surfaced as warnings — this check never fails the build.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
tests/tools/test_mcp_tool.py (2)

469-517: ⚡ Quick win

Consider verifying multiple callback invocations to confirm polling behavior.

The test currently asserts assert recorded, verifying the callback fired at least once. Since the coroutine sleeps for 0.35s and the poll loop iterates every 0.1s, we should expect approximately 3-4 invocations of touch_activity_if_due. Strengthening the assertion to verify multiple invocations would better confirm that the poll loop calls the activity callback on each iteration, not just once.

🔍 Suggested assertion enhancement
     assert result == "done"
     # The callback must have been invoked at least once while waiting.
     assert recorded, "activity callback never fired during the MCP wait"
+    # With 0.35s sleep and 0.1s poll interval, expect 3-4 invocations
+    assert len(recorded) >= 2, f"Expected multiple callback invocations, got {len(recorded)}"
     # State carries the heartbeat cadence and start bookkeeping.
     first_state, first_label = recorded[0]
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/tools/test_mcp_tool.py` around lines 469 - 517, Update the
test_activity_callback_fires_during_long_call to assert multiple invocations of
the activity callback: after calling mcp._run_on_mcp_loop and confirming result
== "done", add a stronger assertion like assert len(recorded) >= 3 (or another
small integer reflecting the expected 3–4 poll iterations) to ensure
touch_activity_if_due was called repeatedly; keep the existing checks for
first_state/first_label and their contents and reference the recorded list, the
_run_on_mcp_loop call, and the patched touch_activity_if_due to locate where to
insert the new assertion.

469-517: ⚡ Quick win

Consider adding test coverage for the guarded import fallback.

The PR description mentions "The import is guarded so the MCP loop still works if tools.environments.base is unavailable." However, there's no test verifying this graceful degradation behavior. A complementary test could verify that when tools.environments.base is unavailable, the MCP loop still completes calls successfully without crashing, even though activity callbacks won't fire.

Would you like me to generate a test that verifies graceful degradation when the import fails?

💡 Suggested test structure
def test_activity_callback_degrades_gracefully_when_import_unavailable(self):
    """MCP loop completes successfully even when touch_activity_if_due unavailable."""
    import tools.mcp_tool as mcp
    
    loop = asyncio.new_event_loop()
    t = threading.Thread(target=loop.run_forever, daemon=True)
    t.start()
    
    async def _fast():
        return "done"
    
    try:
        with patch.object(mcp, "_mcp_loop", loop):
            # Simulate the guarded import failing by making the module unavailable
            with patch.dict('sys.modules', {'tools.environments.base': None}):
                result = mcp._run_on_mcp_loop(_fast, timeout=10)
    finally:
        loop.call_soon_threadsafe(loop.stop)
        t.join(timeout=2)
        loop.close()
    
    # Should complete successfully without crashing
    assert result == "done"
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/tools/test_mcp_tool.py` around lines 469 - 517, Add a test ensuring the
guarded import fallback for tools.environments.base doesn't break the MCP loop:
create a new test (e.g.,
test_activity_callback_degrades_gracefully_when_import_unavailable) that imports
tools.mcp_tool, starts a real asyncio loop and thread like
test_activity_callback_fires_during_long_call, then patch.object(mcp,
"_mcp_loop", loop) and use patch.dict('sys.modules', {'tools.environments.base':
None}) to simulate the missing module, call mcp._run_on_mcp_loop with a simple
coroutine and assert it returns successfully; this verifies the guarded import
around touch_activity_if_due (and related logic) lets _run_on_mcp_loop complete
even when the activity callback module is absent.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@tests/tools/test_mcp_tool.py`:
- Around line 469-517: Update the test_activity_callback_fires_during_long_call
to assert multiple invocations of the activity callback: after calling
mcp._run_on_mcp_loop and confirming result == "done", add a stronger assertion
like assert len(recorded) >= 3 (or another small integer reflecting the expected
3–4 poll iterations) to ensure touch_activity_if_due was called repeatedly; keep
the existing checks for first_state/first_label and their contents and reference
the recorded list, the _run_on_mcp_loop call, and the patched
touch_activity_if_due to locate where to insert the new assertion.
- Around line 469-517: Add a test ensuring the guarded import fallback for
tools.environments.base doesn't break the MCP loop: create a new test (e.g.,
test_activity_callback_degrades_gracefully_when_import_unavailable) that imports
tools.mcp_tool, starts a real asyncio loop and thread like
test_activity_callback_fires_during_long_call, then patch.object(mcp,
"_mcp_loop", loop) and use patch.dict('sys.modules', {'tools.environments.base':
None}) to simulate the missing module, call mcp._run_on_mcp_loop with a simple
coroutine and assert it returns successfully; this verifies the guarded import
around touch_activity_if_due (and related logic) lets _run_on_mcp_loop complete
even when the activity callback module is absent.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 4477dfc7-6c41-47ca-8f01-da77754f3e19

📥 Commits

Reviewing files that changed from the base of the PR and between 62a2d5886d1b0dd0b76d3589066dd560dd03df3f and 0cdcd2973eab73740fa13a6efc3487606e380a2c.

📒 Files selected for processing (2)
  • tests/tools/test_mcp_tool.py
  • tools/mcp_tool.py

@sahilm-ti

Copy link
Copy Markdown
Owner Author

auto-review: approved, awaiting human merge + kanban_approve.

Matrix checks (U1–U5, C1–C5): pass.

  • In-scope: diff touches only tools/mcp_tool.py (+29) and tests/tools/test_mcp_tool.py (+49/-3) — matches AC exactly.
  • CI: lint, ruff, ty-diff, nix (ubuntu+macos), e2e, supply-chain, attribution all green. test (4) and test (5) fail on tests/hermes_cli/test_kanban_default_assignee.py + test_kanban_core_functionality.py — pre-existing failures, confirmed reproducible on HEAD~1 (PR parent), disjoint from this diff (PR touches zero dispatcher code). Not blocking.

Code-quality judgment (role-reviewer): pass.

  • Fix mirrors the proven terminal-path heartbeat (tools/environments/base.py:_wait_for_process → touch_activity_if_due). Signature verified: _activity_state carries required start/last_touch + interval=30.0. 30s cadence is well inside the 900s stuck window; underlying bridge is rate-limited to 1 write/60s so over-eager cadence is harmless.
  • Guarded import for niche surfaces. Purely additive in the poll loop — no change to interrupt/timeout semantics. Per-call tool_timeout path unchanged and now genuinely reachable.
  • Regression test uses a real background event loop + blocking coroutine, asserts callback label + 30s cadence. Good shape.

@sahilm-ti
sahilm-ti force-pushed the kanban/t_cc4a1b4f branch from 0cdcd29 to 5ccfd69 Compare June 3, 2026 15:34
sahilm-ai added a commit that referenced this pull request Jun 3, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
…e tool calls

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.
@sahilm-ti
sahilm-ti force-pushed the kanban/t_cc4a1b4f branch from 5ccfd69 to 749accb Compare June 3, 2026 18:11
@sahilm-ti

Copy link
Copy Markdown
Owner Author

auto-review: approved, awaiting human merge + kanban_approve.

Matrix checks (U1–U6, C1–C6): all pass.

  • U1 in-scope ✓ U2 no out-of-scope deletions ✓ U3 no secrets ✓ U4 AC covered ✓ U5 mergeStateStatus CLEAN ✓ U6 no UI-emitter path ✓
  • C1 CI green (all 6 test shards incl test(3)+test(4)) ✓ C2 no new type:ignore/cast/Any ✓ C3 ruff+ty green ✓ C4 regression tests added ✓ C5 worker identity sahilm-ai ✓ C6 no UI files ✓
  • CodeRabbit: no unresolved actionable threads.

Code-quality judgment (role-reviewer): APPROVED. Single commit, clean rebase onto current main, black-box tests at public boundaries, dependencies injected (activity-state dict / notifier+clock), defensive guards sit on optional liveness/result paths (not required-value fallbacks).

@sahilm-ai
sahilm-ai merged this pull request into main Jun 3, 2026
24 checks passed
@sahilm-ai
sahilm-ai deleted the kanban/t_cc4a1b4f branch June 3, 2026 18:40
sahilm-ti pushed a commit that referenced this pull request Jun 5, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jun 5, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jun 15, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jun 15, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jun 17, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jun 17, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jun 22, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jun 22, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jul 3, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jul 3, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jul 9, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jul 13, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jul 15, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jul 15, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jul 17, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jul 17, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jul 21, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jul 21, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jul 23, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jul 23, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti pushed a commit that referenced this pull request Jul 28, 2026
…artbeat enforcement (#46)

test_dispatch_once_stale_disabled_when_timeout_zero stored os.getpid() as
worker_pid with a 5h-old started_at and no heartbeat, then called
dispatch_once(stale_timeout_seconds=0). dispatch_once runs
enforce_missing_heartbeat independently of stale_timeout_seconds, which
os.kill(SIGTERM)'d the stored PID = the pytest process itself, killing
pytest before it printed its summary (raw RC=143). The parallel harness
then scraped 0 passed/0 failed and bucketed the file as 'no tests ran',
turning test(4) red on #43/#44/#45 — broken-main from the upstream rebase.

Fix: set a recent last_heartbeat_at on the run so enforce_missing_heartbeat
skips the task. This test isolates STALE detection, not heartbeat
enforcement. Also hardened test_enforce_max_runtime_integrates_with_dispatch
(same os.getpid() footgun on its real-os.kill dispatch_once call).

Test-only change. Verified: raw pytest RC=0, 166 passed 1 skipped.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Jul 28, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Aug 24, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
sahilm-ti added a commit that referenced this pull request Sep 2, 2026
…e tool calls (#44)

A single synchronous MCP tool call (e.g. ns_runreport against the NetSuite
hive) blocks _run_on_mcp_loop for many minutes. The poll loop honored user
interrupts but never fired the thread-local activity callback the agent sets
before dispatch, so the agent's last-activity timestamp went stale. For
dispatcher-spawned kanban workers that timestamp is bridged to the board
heartbeat via _touch_activity, so a long hive call starved the heartbeat and
the stuck-worker watchdog killed the worker mid-call (17 consecutive
stuck-kills observed on one report-heavy task).

Mirror the terminal tool's _wait_for_process heartbeat: fire
touch_activity_if_due() on each poll iteration at a 30s cadence (well inside
both the gateway inactivity window and the 15-min kanban stuck default; the
underlying heartbeat bridge is itself rate-limited to one write per 60s).

Adds a regression test exercising a real background loop + blocking coroutine
that asserts the callback fires during the wait.

Co-authored-by: Sahil (AI) <266772320+sahilm-ai@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants