Skip to content

MAINT-78: propose provisional Luna verifier selection from paired evidence - #3689

Draft
stranske wants to merge 3 commits into
mainfrom
codex/maint78-provisional-luna-20261002
Draft

stranske wants to merge 3 commits into
mainfrom
codex/maint78-provisional-luna-20261002

Conversation

@stranske

@stranske stranske commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

HOLD: complexity evidence does not justify a global switch

Keep this PR draft and unmerged. The original eight-case comparison covered six distinct PRs: four one-file evidence documents, one two-file 17-line fix, and one three-file 158-line integration change. The production verifier caps context at 8,000 characters; that integration case was truncated.

A zero-API-call complexity capture reconstructed production-style inputs for three larger previously verified PRs: Workflows #3601 (30 files, 172,022 context characters), Manager-Database #1703 (11 files, 80,174), and Pension-Data #912 (9 files, 30,861). All exceed the prompt cap, so much of their code is absent from the actual model input. A fourth candidate had no usable verifier acceptance context and was excluded. On a subscription-only Terra/Luna pair for Manager-Database #1703, both models returned CONCERNS citing missing acceptance evidence and truncated code, while the source issue records a durable earlier OpenAI PASS and no substantive follow-up debt. This shows a prompt-coverage limitation, not a quality win for either model. No further API confirmation was run.

Resolve prompt coverage issue #3701 and source-evidence issue #3700, then evaluate representative complex PASS and NON_PASS cases before reviving the global selection proposal. The earlier 8/8 result remains valid for its narrow cases.

Proposed provisional change, if the hold clears

Select gpt-6-luna for the OpenAI verifier-balanced registry profile, retain Terra in selection history for rollback, and review the first ten live verifier outcomes or by October 9. Merging the selection PR would be the human approval required by config/model_selection_policy.json; do not merge it while this hold remains.

The original subscription screen and 16-call capped API confirmation found both models 8/8 correct on four PASS and four NON_PASS inputs, with zero false PASS or schema errors. The six natural captures were retrospective and two NON_PASS inputs were controlled defects. Measured-token cost estimates were $0.000561 per accepted review for Luna versus $0.011036 for Terra; total modeled cost was $0.092777, not an invoice charge. These facts support a narrow provisional comparison, not population quality or complex-PR coverage.

Verification

The selection branch's exact-head Gate passed all 56 checks before this complexity finding. The draft hold is an evidence limitation, not a CI failure. The Terra API compatibility fix is merged in #3688. Consumer runtime changes would require Maint 68/71 sync delivery after any future approved selection.

@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Your organization has reached its usage spending cap. Adjust your spending cap in the billing tab.

Next included review available in 29 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available. Your 101 included PR review attempts over the past 7 days set your current allowance at 1 review per hour.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: stranske/Workflows/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 0d3cf75b-c09d-4f74-a1ce-0c5f1ea27242

📥 Commits

Reviewing files that changed from the base of the PR and between c90abeb and 923d73c.

📒 Files selected for processing (8)
  • README.md
  • config/model_eval_candidates.json
  • config/model_registry.json
  • docs/MODEL_SELECTION_POLICY.md
  • docs/ops/ASTRA_ROLLOUT.md
  • templates/consumer-repo/config/model_registry.json
  • tests/tools/test_maint78_cost_quality.py
  • tools/plan_model_eval.py
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@stranske
stranske deployed to agent-standard October 2, 2026 07:19 — with GitHub Actions Active
@stranske-keepalive

stranske-keepalive Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Automated Status Summary

Head SHA: dcd4647
Latest Runs: ⏳ pending — Gate
Required contexts: summary
Required: core tests (3.12): ⏳ pending, core tests (3.13): ⏳ pending, docker smoke: ⏳ pending, gate: ⏳ pending

Workflow / Job Result Logs
(no jobs reported) ⏳ pending —

Coverage Overview

  • Coverage history entries: 1

Coverage Trend

Metric Value
Current 80.51%
Baseline 85.00%
Delta -4.49%
Minimum 70.00%
Status ✅ Pass

Top Coverage Hotspots (lowest coverage)

File Coverage Missing
scripts/issue_dedup_smoke.py 0.0% 4
scripts/runner_lib/__main__.py 0.0% 3
scripts/prune_agent_stubs.py 39.7% 26
tools/ensure_workflow_timeout_variables.py 42.1% 74
scripts/sync_label_docs.py 42.9% 64
scripts/repo_review_round2_runner.py 44.3% 348
scripts/validate_template_sync.py 45.1% 51
scripts/repo_review_backlog_scan.py 45.3% 116
tools/codex_session_analyzer.py 47.9% 59
scripts/create_verifier_labels.py 48.3% 58
tools/ci_failure_triage.py 49.7% 113
scripts/langchain/verdict_extract.py 54.1% 21
scripts/langsmith_observability_health.py 55.3% 83
scripts/select_consumer_sync_phase.py 55.6% 62
scripts/audit_belt_ledger_completion.py 57.1% 14

Low Coverage Files (<50.0%)

File Coverage Missing
scripts/issue_dedup_smoke.py 0.0% 4
scripts/runner_lib/__main__.py 0.0% 3
scripts/prune_agent_stubs.py 39.7% 26
tools/ensure_workflow_timeout_variables.py 42.1% 74
scripts/sync_label_docs.py 42.9% 64
scripts/repo_review_round2_runner.py 44.3% 348
scripts/validate_template_sync.py 45.1% 51
scripts/repo_review_backlog_scan.py 45.3% 116
tools/codex_session_analyzer.py 47.9% 59
scripts/create_verifier_labels.py 48.3% 58
tools/ci_failure_triage.py 49.7% 113

Updated automatically; will refresh on subsequent CI/Docker completions.


Keepalive checklist

Scope

Scope section missing from source issue.

Context for Agent

Related Issues/PRs

Tasks

  • Update .github/workflows/maint-77-model-registry-freshness.yml to dispatch .github/workflows/maint-78-model-evaluation-pilot.yml when a new catalog candidate passes freshness screening.
  • Extend tools/harvest_verifier_corpus.py and .github/workflows/maint-79-verifier-corpus-harvest.yml to join verifier decisions to realized PR outcomes with stable case identities.
  • Extend tools/prepare_model_promotion.py and .github/workflows/maint-86-model-promotion-prepare.yml so same-family, non-increasing-cost candidates can prepare bounded promotion PRs while riskier swaps remain approval-gated.
  • Add rollback metadata and quality_gate_breach handling to .github/workflows/maint-86-model-promotion-prepare.yml without weakening config/model_selection_policy.json.
  • Add the end-to-end regression cases to tests/tools/test_harvest_verifier_corpus.py, tests/tools/test_prepare_model_promotion.py, and tests/workflows/test_model_eval_pilot_workflow.py.

Acceptance criteria

  • python -m pytest tests/tools/test_harvest_verifier_corpus.py tests/tools/test_prepare_model_promotion.py tests/workflows/test_model_eval_pilot_workflow.py -q passes with non-zero collection.
  • A new catalogued model results in an auto-dispatched pilot with itself added as a candidate — no human trigger, no manual candidate edit.
  • Live pr_verifier decisions + realized PR outcomes accumulate labeled paired cases into the approval corpus automatically (test: N simulated decisions+outcomes produce N corpus cases with correct labels).
  • A candidate meeting all gates on ≥75 auto-harvested cases, same-family + cost≤, produces an auto-promotion PR; a cross-family/pricier candidate produces an approval-required PR.
  • A post-promotion quality_gate_breach auto-reverts to the prior selection.
  • Policy gates + human-approval-for-risky-swaps remain unchanged.

@stranske
stranske deployed to agent-standard October 2, 2026 07:22 — with GitHub Actions Active
Base automatically changed from codex/maint78-api-compat-20261002 to main October 2, 2026 07:24
@stranske
stranske force-pushed the codex/maint78-provisional-luna-20261002 branch from b97d424 to a4140d0 Compare October 2, 2026 07:24
@stranske
stranske marked this pull request as ready for review October 2, 2026 07:24
Copilot AI balanced review requested due to automatic review settings October 2, 2026 07:24
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-02T07:27:27.116550Z a4140d0 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a4140d082f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tools/plan_model_eval.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The generated candidate registry remains stale and will fail its enforced drift test.

Review effort: Balanced
Findings: 1 High severity

Open (1)
What changed in this PR

Proposes GPT-6 Luna as the provisional OpenAI verifier based on paired MAINT-78 evidence.

Changes:

  • Switches the verifier selection from Terra to Luna.
  • Adds monitoring-oriented planning output and tests.
  • Synchronizes registry and policy documentation.
File Description
tools/​plan_model_eval.py Reports provisional evidence and monitoring actions.
tests/​tools/​test_maint78_cost_quality.py Updates plan expectations for Luna.
config/​model_registry.json Selects Luna and records evidence/history.
templates/​consumer-repo/​config/​model_registry.json Mirrors the registry for consumers.
docs/​MODEL_SELECTION_POLICY.md Documents evidence and rollback criteria.
docs/​ops/​ASTRA_ROLLOUT.md Updates the auxiliary verifier selection.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread config/model_registry.json
@stranske
stranske deployed to agent-standard October 2, 2026 07:33 — with GitHub Actions Active
@stranske stranske added agents:keepalive Use to initiate keepalive functionality with agents autofix Opt-in automated formatting & lint remediation agent:codex Agent-created issues from Codex agent:auto Delegates agent routing to the auto-delegation policy agent:retry Add to trigger agent retry after rate limit or pause labels Oct 2, 2026
@stranske
stranske deployed to agent-standard October 2, 2026 08:42 — with GitHub Actions Active
@stranske
stranske deployed to agent-standard October 2, 2026 08:42 — with GitHub Actions Active
@agents-workflows-bot

Copy link
Copy Markdown
Contributor

🤖 Keepalive Loop Status

PR #3689 | Agent: Codex | Iteration 0/12

Current State

Metric Value
Iteration progress [----------] 0/12
Action run (agent-run-skipped)
Gate unknown
Tasks 0/11 complete
Timeout 45 min (default)
Timeout usage 0m elapsed (2%, 45m remaining)
Keepalive ✅ enabled
Autofix ❌ disabled

Agent Delegation (auto mode)

Field Value
Selected agent Codex
Reason initial-selection-label
Delegation source static

Last Codex Run

Result Value
Status ⏭️ Skipped
Reason agent-run-skipped

To retry:

  • Add the agent:retry label, OR
  • Wait for conditions to resolve (e.g., Gate success, labels present)

🔍 Failure Classification

| Error type | infrastructure |
| Error category | transient |
| Suggested recovery | Capture logs and context; retry once and escalate if the issue persists. |

@agents-workflows-bot

agents-workflows-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor
Keepalive Work Log (click to expand)
# Time (UTC) Agent Action Result Files Tasks Progress Commit Gate
0 2026-10-02 08:43:03 Codex run (agent-run-skipped) retry skipped — 0 0/11 — —
0 2026-10-02 08:43:58 Codex run (agent-run-skipped) retry skipped — 0 0/11 — cancelled
0 2026-10-02 08:48:37 Codex run (agent-run-skipped) retry skipped — 0 0/11 — success
0 2026-10-02 09:36:39 Codex run (agent-run-skipped) retry skipped — 0 0/11 — success
0 2026-10-02 10:33:21 Codex run (agent-run-skipped) retry skipped — 0 0/11 — success
0 2026-10-02 11:32:31 Codex run (agent-run-skipped) retry skipped — 0 0/11 — success
0 2026-10-02 12:46:19 Codex run (agent-run-skipped) retry skipped — 0 0/11 — success
0 2026-10-02 13:34:22 Codex run (agent-run-skipped) retry skipped — 0 0/11 — success

@stranske
stranske deployed to agent-standard October 2, 2026 08:43 — with GitHub Actions Active
@stranske-keepalive

stranske-keepalive Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

🤖 Keepalive Loop Status

PR #3689 | Agent: Codex | Iteration 0/12

Current State

Metric Value
Iteration progress [----------] 0/12
Action run (agent-run-skipped)
Gate success
Tasks 0/11 complete
Timeout 45 min (default)
Timeout usage 0m elapsed (1%, 45m remaining)
Keepalive ✅ enabled
Autofix ❌ disabled

Agent Delegation (auto mode)

Field Value
Selected agent Codex
Reason cooldown (5 rounds remaining)
Delegation source static

Last Codex Run

Result Value
Status ⏭️ Skipped
Reason agent-run-skipped

To retry:

  • Add the agent:retry label, OR
  • Wait for conditions to resolve (e.g., Gate success, labels present)

🔍 Failure Classification

| Error type | infrastructure |
| Error category | transient |
| Suggested recovery | Capture logs and context; retry once and escalate if the issue persists. |

@stranske

stranske commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Source-sync canary reviews found a policy conflict on this same file: the synced header says next decision review 2026-10-31 while config/model_registry.json still has overdue 2026-08-30 global, OpenAI, and GitHub decision deadlines. Issue #3695 now owns a bounded compatibility/documentation repair plus the publisher-prefixed reasoning-model guard; please coordinate its policy edit with this PR and do not silently advance review_by without decision evidence. This PR currently changes docs/MODEL_SELECTION_POLICY.md and both registry copies, so rebase/review against any #3695 source fix before merging. Reviewer evidence: stranske/Portable-Alpha-Extension-Model#2328 (comment).

This branch was successfully deployed

1 active deployment
agent-standard — 923d73cb Deployed Oct 2, 2026 by stranske via privilege environment gate #14577
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agent:auto Delegates agent routing to the auto-delegation policy agent:codex Agent-created issues from Codex agent:retry Add to trigger agent retry after rate limit or pause agents:keepalive Use to initiate keepalive functionality with agents autofix:escalated autofix Opt-in automated formatting & lint remediation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MAINT-78: verifier prompt silently truncates complex PR code

2 participants