Conversation
This was referenced Sep 23, 2026
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.
test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.
The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.
test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.
The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.
test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.
The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.
test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.
The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.
test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.
The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…xfail scoped to its deviation ev_long_run used to flag a missing claim refresh as stalled_heartbeat and carry on, so neutralising heartbeat_fire_claim in _heartbeat_loop left all five scenarios green. Every virtual 30 s of the 15-minute hold now (a) waits for the run's heartbeat to restamp the claim and fails if it never does, and (b) has a contender call the real claim_job_for_fire and asserts it loses, also once the first stamp is older than the TTL. With the heartbeat replaced by True all 5 scenarios are red on the refresh assertion; on main they are green. The #119970 cell no longer carries a whole-scenario strict xfail. While its behavioural probe reproduces, only the oracle's fired-set and next_run_at checks may deviate: each deviation is recorded, the model follows the stored slot, and the soak runs every virtual day with all other checks live, XFAILing at the end only if a deviation was seen (and failing if the probe says open but none was). The #120314 entry is gone (merged). A scenario that fails mid-hold now releases its held run before teardown so it cannot bleed into the next scenario.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.
test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.
The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…xfail scoped to its deviation ev_long_run used to flag a missing claim refresh as stalled_heartbeat and carry on, so neutralising heartbeat_fire_claim in _heartbeat_loop left all five scenarios green. Every virtual 30 s of the 15-minute hold now (a) waits for the run's heartbeat to restamp the claim and fails if it never does, and (b) has a contender call the real claim_job_for_fire and asserts it loses, also once the first stamp is older than the TTL. With the heartbeat replaced by True all 5 scenarios are red on the refresh assertion; on main they are green. The #119970 cell no longer carries a whole-scenario strict xfail. While its behavioural probe reproduces, only the oracle's fired-set and next_run_at checks may deviate: each deviation is recorded, the model follows the stored slot, and the soak runs every virtual day with all other checks live, XFAILing at the end only if a deviation was seen (and failing if the probe says open but none was). The #120314 entry is gone (merged). A scenario that fails mid-hold now releases its held run before teardown so it cannot bleed into the next scenario.
|
Checked this against current main ( Two notes in case they help review:
|
With no `timezone` configured (the shipped default), hermes_time.now() falls back to datetime.now().astimezone(), a fixed UTC offset, and cron treated that offset as the zone. compute_next_run attached the last run's offset to croniter's next wall clock, parse_schedule attached the creation-time offset to a naive one-shot, and _ensure_aware read legacy naive timestamps with today's offset. Every wall clock on the far side of a DST change landed an hour off. On a New York host the first `0 9 * * *` run after spring-forward was stored at 09:00-05:00 and fired at 10:00 EDT, logging a false timezone_migration.catch_up warning, and a naive one-shot typed in February for 1 April fired at 10:00 EDT. All three sites now resolve naive wall clocks through one helper. A configured zone keeps the existing fold=0/fold=1 readings. With none, the helper takes the host offset a day either side of the wall clock and keeps the readings the host agrees with at that instant. A spring-forward gap matches neither and keeps both, earlier offset first, the way zoneinfo orders them. Server-local and configured modes now resolve every wall clock to the same instant: checked every 15 minutes across 2026 in seven zones, including Lord Howe's 30-minute shift and southern-hemisphere rules. parse_schedule reports a timestamp at the datetime bounds as an invalid timestamp, so the helper's day-either-side probe cannot surface an OverflowError. Six existing tests stood in for a configured zone by patching the clock while leaving the zone unset. They now configure the zone they describe, and their assertions are unchanged. One oracle built its expected value from today's host offset; it now uses the host's reading of that date.
…x is in the tree
jonpol01
force-pushed
the
sweep/cron-server-local-tz-dst-fixed-offset
branch
from
October 4, 2026 03:44
bc18813 to
848a4a9
Compare
Author
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
With
timezoneunset (the default),hermes_time.now()falls back todatetime.now().astimezone(), which has a fixed UTC offset. Cron treated that offset as the zone in three places:compute_next_runattached the last run's offset to croniter's next wall clock.parse_scheduleattached the creation-time offset to a naive one-shot._ensure_awareread legacy naive timestamps with today's offset.As a result, every wall clock on the far side of a DST change landed an hour off. On a New York host the first
0 9 * * *run after spring-forward was stored at 09:00-05:00 and fired at 10:00 EDT, with a falsecron.timezone_migration.catch_upWARNING. A naive one-shot typed in February for 1 April also fired at 10:00 EDT.All three sites now go through one helper,
_wall_clock_readings(wall, zone):Server-local mode and the same zone configured now resolve every wall clock to the same instant. I checked this every 15 minutes across 2026 in seven zones, including Lord Howe's 30-minute shift and southern-hemisphere rules, with 0 mismatches.
Why not plain
naive.astimezone()? On CPython 3.11 its fold=0 reading of a skipped wall clock is the earlier instant, the reverse of zoneinfo. A30 2 * * *job would have moved from 03:30 EDT to 01:30 EST on spring-forward day.The fix stays in
cron/jobs.pyand does not changehermes_time.now(), so no other clock consumer sees a different tzinfo.Related: #66436 / #66456 are the configured-zone half of this bug (croniter reusing the base offset), which main fixed in d6d29b0 by anchoring croniter to the configured IANA zone. That fix cannot reach the default: with no timezone configured, the zone is a fixed UTC offset, so re-attaching it keeps the wrong offset. This PR covers that default case.
Related Issue
Fixes #119969
Type of Change
Changes Made
cron/jobs.py:_wall_clock_readings()helper.compute_next_runuseszone = get_timezone(), so an unset zone means the host's zone rather thanbase_time's fixed offset, and builds its strictly-after candidates from the helper.parse_schedulereads naive one-shots through the helper._ensure_awarereads legacy naive values through the helper.parse_schedule's timestamp guard also catchesOverflowError, so a naive timestamp atdatetime.min/maxstays an "Invalid timestamp"ValueError. Base rejects those inputs later increate_jobwith aValueError, and this keeps that contract.tests/cron/test_server_local_dst.py(new): sets up a host in America/New_York the way a real host does (TZ+time.tzset()), with no Hermes zone configured, and pins the scheduler clock to that host clock. It covers:30 2 * * *,30 1 * * *and*/20 * * * *It is skipped where
time.tzsetdoes not exist (native Windows).tests/cron/test_cron_timezone_migration_catchup.pyandtests/cron/test_jobs.py: six tests stood in for a configured zone by patching the clock with a fixed offset while leaving the zone unset. Unset now means the host's zone (UTC in the suite), so they configure the zone they describe: Europe/Brussels and Asia/Kolkata. Their assertions are unchanged, and they pass on main and with this change.tests/cron/test_timezone.py: the legacy-naive oracle built its expected value from today's host offset, the reading under fix. It now uses the host's reading of that date.How to Test
TZ=UTC pytest tests/cron/test_server_local_dst.py -q: 11 pass. Withcron/jobs.pyfrom main, 9 fail. The 2 bounds-guard tests pass on both by design.TZ=America/New_Yorkwith a scratchHERMES_HOME. Main prints2026-03-08T09:00:00-05:00and2026-04-01T09:00:00-05:00(10:00 EDT). This branch prints-04:00for both (09:00 EDT), the same as main withHERMES_TIMEZONE=America/New_York.compute_next_runhunk → the 6 recurring testsparse_schedulehunk → the 2 one-shot tests_ensure_awarehunk → the legacy test[30 2 * * *]parityOverflowErrorfrom the guard → the 2 bounds testsChecklist
Code
fix(scope):,feat(scope):, etc.)pytest tests/ -qand all tests pass. Not the full suite. I ran every test file that exercises the cron scheduling functions, one file at a time under the hermetic runner env: all 85 files in tests/cron, the cron-related files in tests/hermes_cli, tests/gateway, tests/tools, tests/plugins and tests/agent, and the new file. 1562 passed, 0 failed.Documentation & Housekeeping
docs/, docstrings) — or N/A. The configuration docs already say empty means server-local time; this makes that hold across DST.cli-config.yaml.exampleif I added/changed config keys — or N/ACONTRIBUTING.mdorAGENTS.mdif I changed architecture or workflows — or N/Adatetime.astimezone(), which works on every platform. The new tests skip wheretime.tzsetis missing.scripts/check-windows-footguns.py --diff origin/mainreports only a pre-existingread_text()attests/cron/test_jobs.py:1327, which this PR does not touch.Screenshots / Logs
The real store loop on a New York host with no timezone configured: the Saturday 09:00 run is recorded, then the clock is stepped to Sunday.