fix(bin): stop transient forge read timeouts from waking the contributions watcher - #5860
Closed
belyaev-den wants to merge 1 commit into
Closed
belyaev-den wants to merge 1 commit into
belyaev-den wants to merge 1 commit into
Conversation
Author
|
Superseded by #5900, which already treats read-bound forge timeouts as non-evidence; closing. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Stop the repeating "contributions: observation unavailable" wake for a healthy watched issue.
Observed on 2026-09-26 in one home: https://github.com/mvdlabtech/manvaig/issues/56 (open, unchanged, 3 comments) produced a
check: contributions: observation unavailablewake roughly every hour for most of a day. Every manualgh apiread of the issue, its comments and its events succeeded, typically in 0.7-1 s, but one events read in eighteen took 5.2 s.bin/fm-contributions.shcaps every forge read at 5 s, a slow read counts as a genuine failure, and each later successful poll ends the failure "episode", so every occasional slow read starts a new episode and wakes the supervisor again. The wake carries no contribution signal; it is noise that costs a supervisor turn each time.What Changed
bin/fm-contributions.sh: a forge read that hits the 5 s timeout outside the budget deadline is now tracked separately (forge-timeoutmarker) from a genuine forge failure. If the URL still has a fresh prior good observation (error-free, withinFM_CONTRIBUTIONS_MAX_AGE), the poll keeps that record untouched and retries next poll, with noobservation unavailableerror and no wake.tests/fm-contributions.test.sh: adds coverage for both cases. A timeout with a fresh observation is suppressed. A timeout once the observation has gone stale still records the error and wakes.Risk Assessment
✅ Low: This is a small change. It only suppresses a wake when a single non-budget forge read times out (rc 124), no genuine failure happened in the same observation, and some owner still has an error-free observation within FM_CONTRIBUTIONS_MAX_AGE. Stale, missing or errored evidence still records an error and wakes once per episode, and behavioral tests cover both paths.
Testing
I ran the six targeted contribution tests; all pass on this change, and the two new ones fail on the base commit. I then ran
fm-contributions.sh polllive, seven times, in a disposable lab home. It used the real GitHub issue from the report with the realghlogin, with the events read slowed past the 5 s cap. On the base script, one-off slow reads between healthy polls woke the supervisor each time; on the fixed script they stay silent and the saved observation is left unchanged. Once the observation goes stale, a real outage still wakes exactly once, and the next healthy poll ends the failure episode. Both transcripts are in the evidence directory. This is a CLI/check change, so there is no visual UI to capture. The unit-test-only scenarios are reported as untested for live purposes.Evidence: Live poll transcript with the fix (real issue #56, events read delayed 6 s)
Source: Live poll transcript with the fix (real issue #56, events read delayed 6 s)
Evidence: Live poll transcript on base commit (reproduces the repeated wake)
Source: Live poll transcript on base commit (reproduces the repeated wake)
Evidence: Targeted regression tests on this change
Source: Targeted regression tests on this change
Evidence: New regression tests failing on base commit
Source: New regression tests failing on base commit
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
sleepor from local lock/snapshot timing; it is outside the scope of this change and worth a separate look.tests/fm-contributions.test.shsubset: test_issue_read_timeout_between_successes_stays_silent, test_persistent_issue_read_timeout_wakes_when_stale, test_unavailable_forge_records_error_and_wakes_once_per_episode, test_budget_bounded_call_timeout, test_genuine_failure_near_deadline_is_unavailable, test_late_owner_keeps_failure_episode_suppressed (all pass at cd9d463)Same two new regression tests against a base-commit (30ef650)bin/export: both fail with the spurious wake line, which confirms they reproduce the defectLive: disposable lab home frombin/fm-lab-home.sh create, backlog row watching https://github.com/mvdlabtech/manvaig/issues/56, realghbehind a PATH shim that sleeps 6 s on the events read when enabled,bin/fm-contributions.sh pollrun seven times (healthy / slow within freshness / recovery / slow / slow after 16 min / slow again / recovery)Same seven-poll live sequence against the base-commit script for a before/after comparison✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.