Skip to content

fix: reconcile keepalive authority receipts using exact worker evidence - #3601

Merged
stranske merged 19 commits into
mainfrom
codex/keepalive-authority-recovery-20260927
Sep 28, 2026
Merged

stranske merged 19 commits into
mainfrom
codex/keepalive-authority-recovery-20260927

Conversation

@stranske

@stranske stranske commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Source: Issue #3595

Closes #3595

Automated Status Summary

Scope

The current Workflows-source keepalive authority protocol can strand a prepared or consumed receipt for up to 24 hours after a failed/cancelled preflight or mark-running attempt. Conversely, the failure reporter currently assumes no agent executed, which could replay a single-use authority generation if liveness recovery is broadened without an execution-evidence guard. The active generated canary stranske/Portable-Alpha-Extension-Model#2318 has two unresolved P1 threads on exact head 09b3ae8f4b98dc1530636b091d51e7c2afaa5fc0. Related findings on stranske/trip-planner#1869 and stranske/Travel-Plan-Permission#1638 concern the same Workflows source.

Context for Agent

Related Issues/PRs

Tasks

  • In scripts/runner_lib/core.py, make release_authority_challenge retryable for the same exact receipt, reservation and owner attempt when the first ledger release failed after completion was recorded; reject stale or replacement receipts.
  • In templates/consumer-repo/.github/workflows/agents-81-gate-followups.yml and .github/workflows/agents-keepalive-loop.yml, cover cancellation and failure between preparation and finalization, and move finalization after failure-prone mark-running setup while keeping worker execution gated on successful finalization.
  • In templates/consumer-repo/.github/workflows/agents-keepalive-loop-reporter.yml and .github/workflows/agents-keepalive-loop-reporter.yml, derive agent_execution_started from exact originating worker-job evidence with started, definitely-not-started, and unknown states; make absent PR association and API errors retryable rather than asserting false.
  • In .github/scripts/keepalive_loop.js and .github/scripts/keepalive_authority_state.js, reconcile the exact receipt and owner attempt independently of fresh authority-failure classification and presentation state.running, and never reopen a consumed or confirmed receipt when worker execution is started or unknown.
  • Keep templates/consumer-repo/.github/scripts/ copies byte-aligned, update .github/sync-manifest.yml descriptions or entries as appropriate, document recovery ownership in docs/keepalive/Agents.md and docs/keepalive/GoalsAndPlumbing.md, and add focused regression tests.

Acceptance criteria

  • Targeted authority helper, runner-lib, keepalive-loop, and workflow delivery tests pass, including pytest -q tests/workflows/test_keepalive_authority_delivery.py and the applicable authority and runner suites.
  • A deliberate regression test fails when execution-start protection is removed and passes restored; tests cover preflight cancellation, mark-running setup failure, release-write failure followed by same-attempt retry, response-lost idempotence, generic reporter failure, worker-started summary failure, and unknown execution evidence.
  • Exact-head source PR review threads and required checks are clear before source merge; Maint 68 canary and Maint 71 reconciliation create current, review-clear, green delivery evidence before promotion.

Head SHA: 6c314e2
Latest Runs: ✅ success — Gate
Required: gate: ✅ success

Workflow / Job Result Logs
Gate ✅ success View run
Health 40 Sweep ✅ success View run
Health 44 Gate Branch Protection ✅ success View run
Health 45 Agents Guard ✅ success View run
Health 50 Security Scan ✅ success View run
Health 51 Actions SAST (zizmor) ✅ success View run
Health 52 Semgrep Scan ✅ success View run
Health 69 Consumer Sync Shadow Evidence ✅ success View run
Health 72 Template Sync ✅ success View run
Health 73 Template Completeness ✅ success View run
Health 74 Template Drift ✅ success View run
Keepalive E2E ✅ success View run
Maint 52 Validate Workflows ✅ success View run
PR 11 - Minimal invariant CI ✅ success View run
PR 46 Dependency Repair Contract ⏭️ skipped Last completed result; current run in PR checks
Selftest CI ✅ success View run
Validate Sync Manifest ✅ success View run

Summary by CodeRabbit

  • Reliability
    • Failed or cancelled setup can now release a prepared challenge, with safe retries for the same workflow attempt.
    • Recovery uses evidence from the exact worker attempt: challenges are reopened only when evidence confirms the worker did not start. Uncertain evidence preserves the existing state.
    • Failed runs can recover their associated pull request even when the workflow event does not include one, and settled recovery is reflected in the keepalive summary.
    • Confirmed challenges remain confirmed, and uncertain confirmation does not refund a consumed challenge.

@coderabbitai

coderabbitai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Warning

Review limit reached

Next included review available in 12 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available. Your 122 included PR review attempts over the past 7 days set your current allowance at 1 review per hour.

Your organization has reached its usage spending cap. Adjust your spending cap in the billing tab.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: stranske/Workflows/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 20f40ca0-a04c-44bd-b9e8-7e450399349a

📥 Commits

Reviewing files that changed from the base of the PR and between d40168f and 6c314e2.

📒 Files selected for processing (26)
  • .github/scripts/__tests__/keepalive-authority-state.test.js
  • .github/scripts/__tests__/keepalive-loop.test.js
  • .github/scripts/__tests__/keepalive-reporter-applicability.test.js
  • .github/scripts/__tests__/keepalive-state.test.js
  • .github/scripts/__tests__/keepalive-worker-evidence.test.js
  • .github/scripts/keepalive_authority_state.js
  • .github/scripts/keepalive_loop.js
  • .github/scripts/keepalive_reporter_applicability.js
  • .github/scripts/keepalive_state.js
  • .github/scripts/keepalive_worker_evidence.js
  • .github/sync-manifest.yml
  • .github/workflows/agents-keepalive-loop-reporter.yml
  • .github/workflows/agents-keepalive-loop.yml
  • config/template-drift-allowlist.txt
  • docs/keepalive/Agents.md
  • docs/keepalive/GoalsAndPlumbing.md
  • renovate-presets/consumer-managed-paths.json
  • templates/consumer-repo/.github/scripts/keepalive_authority_state.js
  • templates/consumer-repo/.github/scripts/keepalive_loop.js
  • templates/consumer-repo/.github/scripts/keepalive_reporter_applicability.js
  • templates/consumer-repo/.github/scripts/keepalive_state.js
  • templates/consumer-repo/.github/scripts/keepalive_worker_evidence.js
  • templates/consumer-repo/.github/workflows/agents-81-gate-followups.yml
  • templates/consumer-repo/.github/workflows/agents-keepalive-loop-reporter.yml
  • tests/scripts/test_sync_manifest_compiler.py
  • tests/workflows/test_keepalive_authority_delivery.py
📝 Walkthrough

Walkthrough

The keepalive authority flow now indexes receipts by owner attempt, classifies worker execution for exact workflow attempts, and reconciles eligible failures. Workflow finalization and reporting handle failure or cancellation, and settled recoveries are projected into keepalive summaries.

Changes

Authority receipt recovery

Layer / File(s) Summary
Attempt-bound receipt lifecycle
.github/scripts/keepalive_authority_state.js, templates/consumer-repo/.github/scripts/keepalive_authority_state.js, scripts/runner_lib/core.py, tests
Preparation and finalization validate the immutable attempt index. Release retains receipt evidence and recognizes exact retries. Runner cleanup can retry a matching runner-preflight-failed reservation.
Evidence-gated reconciliation and state projection
.github/scripts/keepalive_worker_evidence.js, .github/scripts/keepalive_authority_state.js, .github/scripts/keepalive_state.js, .github/scripts/keepalive_loop.js, templates/consumer-repo/.github/scripts/*, tests
Worker evidence is classified as started, not-started, or unknown for an exact attempt. Only not-started evidence allows reconciliation; confirmed receipts are not reopened. Settled recoveries update the trusted summary.
Workflow finalization and failed-run reporting
.github/workflows/agents-keepalive-loop.yml, .github/workflows/agents-keepalive-loop-reporter.yml, templates/consumer-repo/.github/workflows/*, tests, docs, .github/sync-manifest.yml, config/template-drift-allowlist.txt
Finalization follows mark-running setup, and failure or cancellation triggers release cleanup. Reporters obtain worker evidence, resolve missing PR targets through the attempt index, reconcile failed attempts, and project settled recoveries.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Bug fix · Severity of issue fixed: Medium

Sequence Diagram(s)

sequenceDiagram
  participant Reporter as Reporter workflow
  participant Evidence as Worker evidence helper
  participant Jobs as Workflow jobs API
  participant Authority as Authority state
  participant Summary as Keepalive summary
  Reporter->>Evidence: Request evidence for exact run attempt
  Evidence->>Jobs: List jobs for that attempt
  Jobs-->>Evidence: Return job and step outcomes
  Evidence-->>Reporter: Return started, not-started, or unknown
  Reporter->>Authority: Resolve target and reconcile failed attempt
  Authority-->>Reporter: Return recovery status
  Reporter->>Summary: Project released or reopened recovery
Loading

Merge Risk: 🟡 Moderate · up to d4016

The authority recovery flow can refund a consumed single-use receipt based on an unverified summary value. It can also leave receipts stuck when a worker is cancelled before it starts, and it reports failures for skipped runs. The refund path should be fixed before merging.

🚥 Pre-merge checks | ✅ 3 | ❌ 1 | ❓ 1

❌ Failed checks (1 warning, 1 inconclusive)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 5.77% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 52 functions across 15 files. (9 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
Linked Issues check ❓ Inconclusive The implementation addresses the coding objectives in #3595. It adds exact-attempt indexing and retryable release, cancellation and failure cleanup, post-setup finalization, worker-job evidence states… Provide exact-head results for the required targeted test suites, including pytest -q tests/workflows/test_keepalive_authority_delivery.py and the applicable authority and runner suites. Include results for the added JavaScript regression…
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: reconciling keepalive authority receipts using exact worker execution evidence.
Out of Scope Changes check ✅ Passed The changed source, workflows, consumer-template copies, tests, documentation, sync manifest, and template-drift allowlist all support the authority recovery and delivery objectives in #3595. No unrel…
Full details: Linked Issues check

Explanation

The implementation addresses the coding objectives in #3595. It adds exact-attempt indexing and retryable release, cancellation and failure cleanup, post-setup finalization, worker-job evidence states, attempt-bound reconciliation, recovery projection, aligned consumer copies, documentation, manifest updates, and regression tests. The delivery tests also assert workflow ordering and root/template coverage. The provided evidence reports checks as pending or in progress. It does not establish that the required targeted authority, runner, keepalive-loop, and delivery test suites pass.

Resolution

Provide exact-head results for the required targeted test suites, including pytest -q tests/workflows/test_keepalive_authority_delivery.py and the applicable authority and runner suites. Include results for the added JavaScript regression tests.

Full details: Docstring Coverage

Explanation

Docstring coverage is 5.77% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 52 functions across 15 files. (9 skipped: 9 unsupported.)

✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-27T18:29:21.959843Z 6c314e2 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@stranske
stranske deployed to agent-high-privilege September 27, 2026 16:55 — with GitHub Actions Active
@stranske-keepalive

stranske-keepalive Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Workflow source detected

PR #3601 now has valid workflow source context (origin=github_issue ref=#3595).

A linked GitHub issue is present for this PR.

@agents-workflows-bot

agents-workflows-bot Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Automated Status Summary

Head SHA: 43153ee
Latest Runs: ⏳ pending — Gate
Required contexts: summary
Required: core tests (3.12): ⏳ pending, core tests (3.13): ⏳ pending, docker smoke: ⏳ pending, gate: ⏳ pending

Workflow / Job Result Logs
(no jobs reported) ⏳ pending —

Coverage Overview

  • Coverage history entries: 1

Coverage Trend

Metric Value
Current 80.56%
Baseline 85.00%
Delta -4.44%
Minimum 70.00%
Status ✅ Pass

Top Coverage Hotspots (lowest coverage)

File Coverage Missing
scripts/issue_dedup_smoke.py 0.0% 4
scripts/runner_lib/__main__.py 0.0% 3
scripts/prune_agent_stubs.py 39.7% 26
tools/ensure_workflow_timeout_variables.py 42.1% 74
scripts/sync_label_docs.py 42.9% 64
scripts/repo_review_round2_runner.py 44.3% 348
scripts/validate_template_sync.py 45.1% 51
scripts/repo_review_backlog_scan.py 45.3% 116
tools/codex_session_analyzer.py 47.9% 59
scripts/create_verifier_labels.py 48.3% 58
tools/ci_failure_triage.py 49.7% 113
scripts/langchain/verdict_extract.py 54.1% 21
scripts/langsmith_observability_health.py 55.3% 83
scripts/select_consumer_sync_phase.py 55.6% 62
scripts/audit_belt_ledger_completion.py 57.1% 14

Low Coverage Files (<50.0%)

File Coverage Missing
scripts/issue_dedup_smoke.py 0.0% 4
scripts/runner_lib/__main__.py 0.0% 3
scripts/prune_agent_stubs.py 39.7% 26
tools/ensure_workflow_timeout_variables.py 42.1% 74
scripts/sync_label_docs.py 42.9% 64
scripts/repo_review_round2_runner.py 44.3% 348
scripts/validate_template_sync.py 45.1% 51
scripts/repo_review_backlog_scan.py 45.3% 116
tools/codex_session_analyzer.py 47.9% 59
scripts/create_verifier_labels.py 48.3% 58
tools/ci_failure_triage.py 49.7% 113

Updated automatically; will refresh on subsequent CI/Docker completions.


Keepalive checklist

Scope

The current Workflows-source keepalive authority protocol can strand a prepared or consumed receipt for up to 24 hours after a failed/cancelled preflight or mark-running attempt. Conversely, the failure reporter currently assumes no agent executed, which could replay a single-use authority generation if liveness recovery is broadened without an execution-evidence guard. The active generated canary stranske/Portable-Alpha-Extension-Model#2318 has two unresolved P1 threads on exact head 09b3ae8f4b98dc1530636b091d51e7c2afaa5fc0. Related findings on stranske/trip-planner#1869 and stranske/Travel-Plan-Permission#1638 concern the same Workflows source.

Context for Agent

Related Issues/PRs

Tasks

  • In scripts/runner_lib/core.py, make release_authority_challenge retryable for the same exact receipt, reservation and owner attempt when the first ledger release failed after completion was recorded; reject stale or replacement receipts.
  • In templates/consumer-repo/.github/workflows/agents-81-gate-followups.yml and .github/workflows/agents-keepalive-loop.yml, cover cancellation and failure between preparation and finalization, and move finalization after failure-prone mark-running setup while keeping worker execution gated on successful finalization.
  • In templates/consumer-repo/.github/workflows/agents-keepalive-loop-reporter.yml and .github/workflows/agents-keepalive-loop-reporter.yml, derive agent_execution_started from exact originating worker-job evidence with started, definitely-not-started, and unknown states; make absent PR association and API errors retryable rather than asserting false.
  • In .github/scripts/keepalive_loop.js and .github/scripts/keepalive_authority_state.js, reconcile the exact receipt and owner attempt independently of fresh authority-failure classification and presentation state.running, and never reopen a consumed or confirmed receipt when worker execution is started or unknown.
  • Keep templates/consumer-repo/.github/scripts/ copies byte-aligned, update .github/sync-manifest.yml descriptions or entries as appropriate, document recovery ownership in docs/keepalive/Agents.md and docs/keepalive/GoalsAndPlumbing.md, and add focused regression tests.

Acceptance criteria

  • Targeted authority helper, runner-lib, keepalive-loop, and workflow delivery tests pass, including pytest -q tests/workflows/test_keepalive_authority_delivery.py and the applicable authority and runner suites.
  • A deliberate regression test fails when execution-start protection is removed and passes restored; tests cover preflight cancellation, mark-running setup failure, release-write failure followed by same-attempt retry, response-lost idempotence, generic reporter failure, worker-started summary failure, and unknown execution evidence.
  • Exact-head source PR review threads and required checks are clear before source merge; Maint 68 canary and Maint 71 reconciliation create current, review-clear, green delivery evidence before promotion.

@stranske
stranske deployed to agent-high-privilege September 27, 2026 16:58 — with GitHub Actions Active

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0d7ac8ca29

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +34 to +38
- name: Require PR association for failed originating run
if: >-
github.event.workflow_run.conclusion != 'success' &&
!github.event.workflow_run.pull_requests[0].number
run: |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Recover the target PR for dispatched authority runs

Hourly authority challenges are started by agents-keepalive-sweep.yml as workflow_dispatch runs on the default branch, with the target carried only in the pr_number input. Their subsequent workflow_run payload can therefore have an empty pull_requests array, so this guard exits before collecting worker evidence or reopening the consumed receipt. A failure or cancellation after finalization but before the worker starts—the exact case this reporter is meant to repair—then leaves the receipt spent and the PR stuck. Persist or otherwise recover the dispatch input instead of requiring a workflow-run PR association; the synced consumer reporter has the same guard.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed on b8ee068: both reporters now recover an unassociated workflow_dispatch run from the unique authority-ledger receipt matching its exact owner attempt, fail closed on ambiguity/no target, and use the ledger PR head for summary projection. Added source/template coverage and validated focused tests, template drift, and actionlint. Please reassess this finding on the current head.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fresh review completed on b8ee068 without an in-thread disposition here. The missing-PR recovery remains in current head 4176c45; the later change only skips the ledger scan when the webhook already identifies a PR. Please explicitly confirm whether this finding is satisfied on the current head, or identify a remaining defect.

@stranske
stranske deployed to agent-high-privilege September 27, 2026 17:00 — with GitHub Actions Active
@stranske
stranske deployed to agent-high-privilege September 27, 2026 17:03 — with GitHub Actions Active
@stranske

Copy link
Copy Markdown
Owner Author

@codex review

@stranske
stranske deployed to agent-high-privilege September 27, 2026 17:08 — with GitHub Actions Active
@stranske
stranske deployed to agent-high-privilege September 27, 2026 17:10 — with GitHub Actions Active

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b8ee06851c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +164 to +168
authorityTarget = await findAuthorityPrForAttempt({
request: authorityRequest,
repository: `${context.repo.owner}/${context.repo.repo}`,
ownerAttempt,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Skip the full ledger scan when the PR is known

For normal failed runs whose workflow_run.pull_requests[0] already identifies the PR, this unconditional lookup still lists the entire authority directory and findAuthorityPrForAttempt then reads every historical *.json ledger file. Reporting therefore costs O(all previously challenged PRs) API calls, aborts if any unrelated ledger file is unreadable or malformed, and is guaranteed to fail once the listing reaches the explicit 1,000-entry guard. Only perform this repository-wide lookup when prNumber is absent; an associated run can reconcile directly against its known PR's ledger.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed on 4176c45: the ledger-wide lookup now runs only when workflow_run has no PR association; associated runs use their known PR directly. Both source and consumer reporters are updated, with the focused delivery assertion, template drift, and actionlint passing. Please disposition this thread on the current head.

@stranske

Copy link
Copy Markdown
Owner Author

@codex review

@stranske
stranske deployed to agent-high-privilege September 27, 2026 17:14 — with GitHub Actions Active

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4176c4578f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +416 to +417
if (!Array.isArray(entries) || entries.length >= 1000) {
throw new Error('Authority directory listing unavailable or incomplete');

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid the capped directory scan for dispatch recovery

Fresh evidence after the associated-run optimization is that authority challenges are still workflow_dispatch runs whose webhook commonly has no PR association, so their failed-run recovery necessarily reaches this branch. Because authority files are retained per PR with no cleanup, once the directory reaches 1,000 entries every such recovery is rejected here, while even earlier failures require one API read per historical PR and can exhaust the App's rate limit. Persist an attempt-to-PR lookup or another direct target reference instead of depending on a full Contents API directory scan.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed on 92eeb50: preparation now creates an immutable SHA-keyed exact-attempt locator before the PR ledger reservation, and finalization verifies it. Missing-PR reporter recovery uses one direct index GET plus one PR ledger GET with receipt-ID/binding validation; there is no historical directory scan. Lost create responses deny that invocation, and released attempts need a new run attempt. Source/template mirror, manifest, docs, and 228 JS + 273 Python focused tests pass. Please disposition this finding on the current head.

@stranske
stranske deployed to agent-high-privilege September 27, 2026 17:20 — with GitHub Actions Active
@stranske

Copy link
Copy Markdown
Owner Author

@codex review

@stranske
stranske deployed to agent-high-privilege September 27, 2026 17:25 — with GitHub Actions Active
@stranske
stranske deployed to agent-high-privilege September 27, 2026 18:11 — with GitHub Actions Active

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 50ff464445

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread .github/workflows/agents-keepalive-loop-reporter.yml
@stranske

Copy link
Copy Markdown
Owner Author

@codex review

@stranske
stranske deployed to agent-high-privilege September 27, 2026 18:13 — with GitHub Actions Active

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4481165dc9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

if (settled.state.status !== 'available' ||
!receiptMatches(settled.state.released_receipt, claim, ownerAttempt,
receipt.provider, state.head_sha)) return { status: 'uncertain' };
return { status: 'released', reason: result.reason, state: settled.state };

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Refresh the expiry when releasing a delayed preparation

If this reporter runs or is manually retried after the prepared generation's expires_at, this branch returns the released state with its original, already-expired due/expiry window. projectRecoveredAuthorityState then publishes those timestamps, while selectDueAuthorityChallenge rejects the projection when now >= expires_at, so the promised retry is never dispatched and the PR remains stuck. Rotate to a fresh generation/window here when the released state has expired, as the consumed-receipt recovery path already does.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed on head 6c314e2: exact non-start release now rotates an expired prepared window under one ledger CAS, retaining the immutable released receipt, original generation, and lineage. Replays of an expired release refresh the window only for the same indexed attempt and open same-head automation-labeled PR without a hard human blocker; projection accepts exact release lineage, and expiry alone no longer reinitializes a same-head prepared/consumed/confirmed receipt. Source/template code and docs updated. Focused validation: 256 JS and 73 Python tests passed, including expiry, replay, response-loss, confirmed-state and multi-window projection regressions. Please reassess and disposition this finding on the exact current head.

@stranske
stranske deployed to agent-high-privilege September 27, 2026 18:26 — with GitHub Actions Active
@stranske

Copy link
Copy Markdown
Owner Author

@codex review

@stranske

Copy link
Copy Markdown
Owner Author

Reviewer disposition handoff for current head 6c314e2: completed Codex reviews have not resolved five active non-outdated threads despite owner fix replies. Originating reviewer, please give explicit thread-level acceptance or remaining changes for #3601 (comment), #3601 (comment), #3601 (comment), #3601 (comment), and #3601 (comment). The expiry finding has a fresh exact-head reply at #3601 (comment). If the bot cannot disposition its threads, an authorized human reviewer must do so; the source owner will not self-resolve or merge through active threads. Recheck after the current review/CI, no later than the 20:25 UTC campaign run.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 🚀

Reviewed commit: 6c314e2e5b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@stranske

Copy link
Copy Markdown
Owner Author

Repository-owner reviewer handoff: the exact-head Codex review of 6c314e2 reported no major issues, but five original non-outdated finding threads remain active without explicit same-head acceptance. The existing thread list and fixes are at #3601 (comment). No independent human reviewer is currently configured among repository collaborators (only stranske and stranske-automation-bot). Please designate or authorize an independent human reviewer to state accepted or remaining defect in each original thread, or have the originating Codex reviewer do so. I will not self-resolve or merge through active threads. Next automatic check: 2026-09-28 02:25 UTC.

@stranske

Copy link
Copy Markdown
Owner Author

Closer exact-head review disposition for 6c314e2: all five active findings are satisfied by the current source/template changes. Attempt-aware reporter fingerprinting binds run ID plus run attempt; both mark-running cleanup paths pass the signed fingerprint; worker evidence loads the registry at the originating run head and returns unknown on unavailable evidence; ordinary unassociated dispatches skip only after exact producer/index classification while authority-candidate or uncertain evidence fails closed; root reporter Node/API setup is guarded by the applicability skip. Independent focused validation passed: node --test over worker evidence, reporter applicability, and authority state: 40/40; PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 pytest -q tests/workflows/test_keepalive_authority_delivery.py -p no:cacheprovider: 3/3. A bounded Astra Medium assessment independently reached the same five satisfied verdicts and identified only two future test-strength opportunities (behavioral fingerprint hashing and expression-equality coverage), not remaining production defects. Resolving the five threads; no source change or push was required.

@stranske

Copy link
Copy Markdown
Owner Author

Final merge gate for exact head 6c314e2: CLEAN and MERGEABLE; required summary check passed; no failed or pending checks; complete thread pagination found zero active non-outdated threads. The named expected-check helper is absent and was not treated as a pass. Direct same-repository topology comparison against merged workflow-heavy PRs #3600 and #3603 shows the identical 18 non-skipped Gate jobs, all successful here, including privilege gate, Python 3.12/3.13, scripts, ledger validation, test-quality, and summary. The head predates this round and no push occurred, so the seven-minute exact-head review floor is satisfied. Squash merging without deleting the branch, then applying verify:compare.

@stranske
stranske merged commit 43153ee into main Sep 28, 2026
66 checks passed
@stranske
stranske deleted the codex/keepalive-authority-recovery-20260927 branch September 28, 2026 09:31
@stranske stranske added the verify:compare Compare multiple LLM evaluations label Sep 28, 2026
@stranske
stranske deployed to agent-high-privilege September 28, 2026 09:31 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

Provider Comparison Report

Provider Summary

Provider Model Verdict Confidence Summary
openai gpt-5.6-terra PASS 91% The merged changes comprehensively implement the requested keepalive authority-recovery behavior. Authority release now preserves exact receipt, reservation, and owner-attempt identity for safe sam...
anthropic claude-sonnet-5 CONCERNS N/A Review the PR manually or re-run once LLM credentials are available.
📋 Full Provider Details (click to expand)

openai

  • Model: gpt-5.6-terra
  • Verdict: PASS
  • Confidence: 91%
  • Scores:
    • Correctness: 9.0/10
    • Completeness: 9.0/10
    • Quality: 9.0/10
    • Testing: 9.0/10
    • Risks: 9.0/10
  • Summary: The merged changes comprehensively implement the requested keepalive authority-recovery behavior. Authority release now preserves exact receipt, reservation, and owner-attempt identity for safe same-attempt retry while rejecting stale or replacement data. Reconciliation is expanded to operate from exact receipt/owner evidence rather than presentation state or fresh authority-failure classification, and protects consumed/confirmed receipts whenever execution is started or indeterminate. Reporter handling adds explicit worker-execution evidence states (started, definitely not started, unknown), treating missing PR association and API failures as retryable rather than assuming no worker executed. The keepalive workflows move finalization behind failure-prone mark-running setup and retain worker gating on successful finalization, including cancellation/failure recovery paths. Root and consumer-template script copies are updated in parallel, with manifest, documentation, and drift-related metadata adjusted. Focused runner, authority-state, loop/state, reporter applicability, worker-evidence, workflow-delivery, and synchronization tests were added or extended to cover the stated failure and idempotence scenarios. The implementation is readable, separates evidence classification from reconciliation decisions, and does not introduce an apparent unsafe reopening path for potentially executed authority generations.

anthropic

  • Model: claude-sonnet-5
  • Verdict: CONCERNS
  • Confidence: N/A
  • Summary: Review the PR manually or re-run once LLM credentials are available.
  • Concerns:
    • LLM evaluation could not run.
  • Error: LLM invocation failed: Error code: 400 - {'type': 'error', 'error': {'type': 'invalid_request_error', 'message': 'You have reached your specified API usage limits. You will regain access on 2026-10-01 at 00:00 UTC.'}, 'request_id': 'req_011CfVcUeZ24CSByE9dYH64K'}

Agreement

  • No clear areas of agreement.

Disagreement

Dimension openai anthropic
Verdict PASS CONCERNS

Unique Insights

  • openai: The merged changes comprehensively implement the requested keepalive authority-recovery behavior. Authority release now preserves exact receipt, reservation, and owner-attempt identity for safe same-attempt retry while rejecting stale or replacement data. Reconciliation is expanded to operate fro...
  • anthropic: LLM evaluation could not run.

🔍 LangSmith Traces

This branch was successfully deployed

1 active deployment
agent-high-privilege — 6c314e2e Deployed Sep 28, 2026 by stranske via privilege environment gate #14168
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

autofix:escalated verify:compare Compare multiple LLM evaluations

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[sync-review] Fix upstream manifest-synced paths blocking stranske/Portable-Alpha-Extension-Model#2318

2 participants