Skip to content

fix(kanban): coerce foreign timestamp values when reading task/run/event rows - #97462

Open
liuhao1024 wants to merge 1 commit into
NousResearch:mainfrom
liuhao1024:liuhao/cron-bugfix-97455
Open

liuhao1024 wants to merge 1 commit into
NousResearch:mainfrom
liuhao1024:liuhao/cron-bugfix-97455

Conversation

@liuhao1024

Copy link
Copy Markdown

What does this PR do?

kanban show (text and JSON), the kanban_show worker context, and the dashboard's run history all crash with invalid literal for int() with base 10: '2026-08-28T14:35:22+00:00' once a timestamp column in tasks / task_runs / task_events holds an ISO-8601 string, making the affected card unreadable (#97455).

SQLite's type affinity accepts those writes silently, and every mutator in kanban_db already writes int(time.time()) — the stray strings come from rows written outside the mutators (e.g. a worker shell running raw SQL). Since the writer can't be pinned down, the fix is on the reader side: a new _coerce_epoch helper passes numeric values through, converts timezone-aware ISO-8601 strings (both +00:00 and Z spellings observed in the wild) to epoch seconds, and degrades anything untrustworthy (naive ISO, garbage) to None so the card stays readable. Naive ISO strings are deliberately dropped rather than guessed at a timezone, so a malformed value is never rendered as a wrong time.

Applied at the row mappers (Task.from_row, Run.from_row, both Event constructors), which covers every consumer of those objects, plus the two raw-SQL reader spots in the dispatcher (respawn guard, rate-limit cooldown) and the worker-context "recent work" renderer that called int() directly.

Related Issue

Fixes #97455

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)

Changes Made

  • hermes_cli/kanban_db.py: added _coerce_epoch() — epoch/float/epoch-as-str pass through, timezone-aware ISO-8601 converts to epoch seconds, naive ISO and garbage degrade to None
  • hermes_cli/kanban_db.py: Task.from_row / Run.from_row / both Event constructions now coerce created_at / started_at / completed_at / ended_at / claim_expires / last_heartbeat_at instead of passing strays through or calling bare int()
  • hermes_cli/kanban_db.py: check_respawn_guard (rate-limit cooldown + recent-completion window) and the build_worker_context "recent work" renderer coerce raw ended_at values instead of int()
  • tests/hermes_cli/test_kanban_db.py: _coerce_epoch matrix test + row-mapper regression that seeds the exact corruption from the issue
  • tests/hermes_cli/test_kanban_cli.py: end-to-end kanban show regression (text + JSON) reproducing the issue's UPDATE scenario

How to Test

  1. uv run --extra dev python -m pytest tests/hermes_cli/test_kanban_db.py tests/hermes_cli/test_kanban_cli.py tests/tools/test_kanban_tools.py -q — Observed result: 70 passed, 1 skipped (the skip is a pre-existing environment-conditional test, unrelated)
  2. uv run --extra dev python -m pytest tests/hermes_cli/test_kanban_db.py::test_readers_coerce_iso_strings_in_integer_timestamp_columns tests/hermes_cli/test_kanban_cli.py::test_show_survives_iso_strings_in_integer_timestamp_columns -q — Observed result: 2 passed; the tests store '2026-08-28T14:35:22+00:00' into task_runs.ended_at and '2026-08-28T16:58:56.613Z' into tasks.completed_at / task_events.created_at exactly as in the issue, then assert show renders and the JSON payload carries epoch ints
  3. On the issue's repro (UPDATE task_runs SET ended_at = '2026-08-28T14:35:22+00:00' ... then hermes kanban show <task_id>): should pass — the card renders and runs[].ended_at is the coerced epoch int 1756384522 instead of raising invalid literal for int()

Checklist

Code

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — or N/A (behavior rationale documented in the _coerce_epoch docstring)
  • I've updated cli-config.yaml.example if I added/changed config keys — or N/A
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — or N/A
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide — or N/A (pure-Python datetime parsing, no platform-specific code)
  • I've updated tool descriptions/schemas if I changed tool behavior — or N/A

…ent rows

SQLite's type affinity lets raw SQL store ISO-8601 strings into the
INTEGER timestamp columns silently. Every mutator in kanban_db writes
int(time.time()), but rows written outside the mutators (e.g. a worker
shell running raw SQL) leave ISO strings behind, and Run.from_row's
int(row["ended_at"]) then crashes every later read of the affected card
with `invalid literal for int() with base 10` — kanban show (text and
JSON), the kanban_show worker context, and the dashboard's list_runs
all fail (NousResearch#97455).

Readers now tolerate those strays via _coerce_epoch: numeric values
pass through, timezone-aware ISO-8601 strings convert to epoch seconds,
and anything untrustworthy (naive ISO, garbage) degrades to None so the
card stays readable instead of crashing. Also covers the worker-context
"recent work" renderer and the dispatcher respawn/cooldown reads.
@alt-glitch alt-glitch added type/bug Something isn't working comp/cron Cron scheduler and job management P3 Low — cosmetic, nice to have labels Aug 28, 2026
@liuhao1024

Copy link
Copy Markdown
Author

The Windows-only failure in this run is a known flaky timing test, unrelated to this diff.

Failing test: tests/test_desktop_update_windows_progress.py::test_progress_advances_while_the_orchestrator_blocks — AssertionError: /progress unresponsive until deadline (last error: TimeoutError('timed out')).

Why it's unrelated to this change:

  • This PR only touches hermes_cli/kanban_db.py and its two kanban test files — no shared code path with the Desktop update /progress socket/HTTP chain.
  • This test has a documented flake history on CI: 6d1284a (Aug 19, "deflake the Windows progress self-test (both race directions seen in CI)") and 18a15a4 (Aug 20, "Windows progress self-test survives transient /progress socket stalls"). The failure here (/progress unresponsive until deadline) is the same transient-stall shape the second deflake commit targeted, so a stall can still slip through under runner load.
  • Everything else in the same run is green: the full Python tests / Run tests job passed, and 153/154 selected Windows tests passed — only this deadline-sensitive test timed out.

No rerun attempted since external contributors can't trigger one; happy to rebase if that's preferred.

@KeyArgo

KeyArgo commented Aug 28, 2026

Copy link
Copy Markdown

Verified this covers the issue completely. Reproduced the crash independently: UPDATE task_runs SET ended_at='2026-08-28T14:35:22+00:00' then hermes kanban show raises ValueError: invalid literal for int() from Run.from_row (kanban_db.py:1287), exactly as reported. The reader-side _coerce_epoch approach is the right call — it matches the reporter's own proposed fix (route all timestamp reads through _to_epoch()), and this PR covers every affected column the issue enumerated (task_runs.ended_at, task_events.created_at, tasks.completed_at) plus the dispatcher/worker-context raw int() call sites. Nice to see both text and JSON show covered with the exact seeded corruption. LGTM.

@liuhao1024

Copy link
Copy Markdown
Author

Thanks for reproducing the crash independently and confirming the reader-side coercion covers the issue — matching the reporter's intent with a defensive read path was exactly the goal.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cron Cron scheduler and job management P3 Low — cosmetic, nice to have type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Non-epoch timestamps written into INTEGER columns make kanban show crash with `invalid literal for int()

3 participants