Skip to content

feat(OMN-14889): git-tag-driven release-train trigger for dev+stability lab lanes (re-cut onto dev) - #2388

Merged
jonahgabriel merged 3 commits into
devfrom
jonah/omn-14889-release-train-recut
Jul 21, 2026
Merged

jonahgabriel merged 3 commits into
devfrom
jonah/omn-14889-release-train-recut

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Jul 21, 2026 •

Copy link
Copy Markdown
Collaborator

OMN-14889 — release-train lab-lane trigger (re-cut onto dev)

Re-cut of closed PR #2377 (single commit c2ec0d2e, branch deleted) onto current dev.
#2377 was a stacked branch on top of #2370 (OMN-14873); #2370 merged to dev by squash as
54938fd3 at 2026-07-21T19:35:17Z and #2377 was closed unmerged 2 seconds later. Because the
parent landed squashed, a cherry-pick/rebase of c2ec0d2e produces a spurious add/add conflict on
scripts/runtime_build/refresh_stability_lane.sh. The first commit here therefore used a
file-scoped restore of only the 6 paths that are actually #2377's own diff — conflict-free by
construction. Two later commits deliberately diverge from #2377 (see "Commit provenance").

Closes OMN-14889

⚠️ Safety correction (2026-07-21)

An earlier revision of this PR body and of three files in the diff asserted the
omnibase-deploy runner was unprovisioned and that the deploy job "queues and never picks
up." That is false. The runner is registered and ONLINE, so pushing a lab/dev/** or
lab/stability/** tag triggers a real lane refresh in execute mode.
The stale text has been
corrected in this push — see "Stale-claim repair" below. Do not read a lab tag push as inert.

What this ships

Git-tag-driven release-train trigger for the dev and stability lab lanes. cut-tag
(workflow_dispatch) pushes a lab/<lane>/<utc>-<shortsha> tag; that tag push fires deploy,
which calls refresh_stability_lane.sh (OMN-14873, reused verbatim) or the new
refresh_dev_lane.sh, then publishes the lane's JSON receipt to the job summary.
Prod is untouched — lab/** is a disjoint tag namespace from v*, and prod promotion still
routes through node_redeploy_orchestrator's grant gate (CLAUDE.md 2a/12). Nothing here can
satisfy or bypass that gate.

Files (7)

Counts are cumulative additions/deletions vs the merge base (git diff --stat <merge-base>..c9b1cc6f),
not per-commit.

File Change +/-
.github/workflows/release-train-lab.yml new — trigger workflow + stale-claim repair + unrecognized-lane fallback +222 / -0
docker/docker-compose.runners.yml modified, additive only (new omninode-deploy-runner service + its named volume) + stale-claim repair +95 / -0
docs/runbooks/release-train-lab.md new — runbook + stale-claim repair +187 / -0
scripts/runtime_build/cut_release_train_tag.sh new +141 / -0
scripts/runtime_build/refresh_dev_lane.sh new — dev-lane refresh + OMN-14900 hardening +576 / -0
scripts/runtime_build/verify_dev_refresh.py new — dev-lane health gate + broker-address parameterization +460 / -0
scripts/runtime_build/tests/test_verify_dev_refresh.py new — unit test for the broker-address seam +42 / -0

Total: 7 files, 1723 insertions, 0 deletions.

scripts/runtime_build/refresh_stability_lane.sh is NOT in this diff. dev's copy (from #2370)
is newer than the one #2377's stacked branch carried (+13/-2 vs refs/pull/2377/head:
git checkout dev → git checkout --force --detach "${RESOLVED_REF_SHA}", fixing the OMN-12618
sibling-worktree collision). dev's version is kept; this branch does not touch that file.

Commit provenance — the branch is NOT byte-exact to #2377

Three commits, only the first of which is a restore:

Commit What it is Relationship to #2377
792c3b7a file-scoped restore of #2377's 6 paths byte-exact. git diff refs/pull/2377/head 792c3b7a -- <the 6 paths> is empty. git status --short at that commit listed exactly those 6 paths, nothing else.
3d4dee67 stale-claim repair (3 files) diverges — corrects the false "runner unprovisioned" text. Docs/comments only.
c9b1cc6f functional hardening + 3 unrelated changes diverges — functional, not editorial. Enumerated below.

c9b1cc6f contains four distinct changes. Only (a) is the scoped work:

(a) refresh_dev_lane.sh hardening — the scoped change. Replaces the git checkout dev +
git reset --hard "${REF}" pair with rev-parse "${REF}^{commit}" followed by
checkout --force --detach "${RESOLVED_REF_SHA}", and routes every ambient-clone git call through a
new git_clone() wrapper that applies per-invocation git -c safe.directory=<clone>. It also adds
a new loop that refreshes the four sibling ambient clones (see "Reviewer attention" below).
This addresses OMN-14900 failure modes 3 and 1 for the dev lane only.

(b) .github/workflows/release-train-lab.yml — unrecognized-lane fallback. The receipt-summary
step's if/else on LANE became if/elif/else with state_dir="" in the terminal branch, plus a
"No receipt is associated with an unrecognized release-train lane." message. Previously an
unrecognized lane silently reused the dev receipt path. Unrelated to (a).

(c) verify_dev_refresh.py broker-address parameterization. Adds
DEFAULT_BROKER_ADDRESS = os.environ.get("OMNIBASE_DEV_BROKER_ADDRESS", "redpanda:9092") and a new
broker_address parameter on check_cluster_health(), replacing the hardcoded
brokers=redpanda:9092. Ships with the new tests/test_verify_dev_refresh.py
(test_cluster_health_uses_configured_broker_address). Unrelated to (a).

(d) ~/.omnibase/.env sourcing restructure. The .env source block in refresh_dev_lane.sh
was de-nested from inside the docker/runtime-policy.env guard into its own top-level if, with
set -a / set +a rebalanced so each block brackets its own source. Behavior change: .env is now
sourced even when runtime-policy.env is absent. Unrelated to (a).

Reviewer attention — sibling-clone refresh loop (refresh_dev_lane.sh)

c9b1cc6f adds a loop that runs, for every repo in
ALL_TRACKED_REPOS=(omnibase_infra omnibase_core omnibase_compat onex_change_control omnimarket):

git_clone "${clone}" fetch origin --prune
sibling_ref="${REF}"
if ! sibling_sha="$(git_clone "${clone}" rev-parse "${sibling_ref}^{commit}" 2>/dev/null)"; then
    sibling_ref="origin/dev"
    sibling_sha="$(git_clone "${clone}" rev-parse "${sibling_ref}^{commit}")"
fi
git_clone "${clone}" checkout --force --detach "${sibling_sha}"

Two risks a reviewer must weigh before approving:

  1. --force against shared ambient clones. This force-checks-out four additional shared
    clones under $OMNI_HOME, discarding any local modifications in them and leaving each on a
    detached HEAD rather than the branch its other consumers expect. These clones are contended
    multi-consumer resources — the 04:14/04:19 run failures below name a live sibling worktree on the
    deploy host's omnibase_infra clone (runtime-sync-worktrees/OMN-12618/...), and on the
    workstation clone this branch was cut from git worktree list reports 23 entries (the clone
    itself plus 22 linked worktrees).
    The sibling clones' worktree state on the deploy host has not been probed. This is in
    tension with OMN-14900's ranked-feat: PostgreSQL Adapter with Comprehensive Tests and Structured Logging #1 proposed fix ("stop writing to the shared ambient clone") —
    it expands CI's write surface on shared clones from one repo to five.
  2. Silent fallback permits a MIXED-ref lane behind a single-REF receipt. A lab/dev/** tag
    exists only in omnibase_infra, so in every sibling clone the rev-parse fails and the loop
    falls back to origin/dev with no warning, no receipt field, and no refusal. The lane can be
    assembled from omnibase_infra@<tag-sha> plus four siblings at whatever origin/dev was at
    fetch time, while the receipt records a single REF.

Both are filed as OMN-14901 (child of OMN-14889) rather than fixed here.

Stale-claim repair (commit 3d4dee67)

Three files asserted the runner was unprovisioned. Corrected to live-verified fact:

  • .github/workflows/release-train-lab.yml — PROVISIONING STATUS header block
  • docs/runbooks/release-train-lab.md — provisioning section + "manual canary" preamble
  • docker/docker-compose.runners.yml — omninode-deploy-runner comment block

Live facts, re-verified 2026-07-21:

gh api repos/OmniNode-ai/omnibase_infra/actions/runners
-> {"name":"omninode-deploy-runner","status":"online","busy":false,
    "labels":["self-hosted","Linux","X64","omnibase-deploy"]}
  • The runner carries the exact label set both jobs target (runs-on: [self-hosted, omnibase-deploy]).
    It is registered at the repository level, not the org level the original runbook recipe
    assumed — the docs now point at the repos/... endpoint for verification.
  • The trigger only requires release-train-lab.yml to exist at the tagged commit, so a tag cut
    from a feature branch fires it exactly as a tag on dev would. That is how the runs below
    happened while this workflow is still unmerged.

dod_evidence

Verifiable from remote surfaces (re-checked at head c9b1cc6f):

  • Restore commit 792c3b7a is byte-exact to refs/pull/2377/head over the 6 restored paths
    (empty git diff), and scoped — git status --short listed exactly those 6 paths.
  • Live CI on c9b1cc6f at the time of this body edit: 130 pass, 5 skipping, 0 fail;
    CI Summary, verify / verify,
    deploy-gate / deploy-gate, occ-preflight / eligibility and receipt-honesty all pass.
  • Evidence-Source: OCC#4555 — onex_change_control#4553, branch
    codex-letter/omn-14889-infra-2388-occ, MERGED 2026-07-21T20:36:27Z.

Author-reported local runs (not independently re-verified in this body):

  • uv run ruff format scripts/ → 196 files unchanged; uv run ruff check --fix scripts/ → all
    checks passed.
  • pre-commit run --files <changed paths> → zero failures (all hooks Passed/Skipped).
  • Pre-push governed selector escalated and ran 845 passed, 1 skipped (no -k, no bypass flag,
    no --no-verify) at the 3d4dee67 push. The authoritative post-c9b1cc6f signal is the live CI
    line above.
  • verify_dev_refresh.py --help exits 0 — the module imports and its argparse is well-formed.
  • Both edited YAML files still parse (yaml.safe_load).

Trigger path IS live (corrected — was previously claimed dormant):

  • The omnibase-deploy runner is registered and online; the deploy job does pick up.
  • 7 lab/stability/* tags exist on origin, 5 of them pointing at c2ec0d2e. 5 real
    release-train-lab runs executed
    2026-07-21T03:59–04:19Z — run ids 29799976126,
    29800099028, 29800399010, 29800684577, 29800954943. All 5 conclusion=failure.
  • The two earliest tags (43935c84…, e5404e36…) produced no runs — the workflow file does not
    exist at those tagged commits.
  • Zero lab/dev/** tags exist. The dev lane has never been triggered.

All 5 runs failed BEFORE any mutation — benign, but not "nothing happened":

Runs Step Error Exit
03:59 Refresh stability lane fatal: detected dubious ownership in repository at /data/omninode/omni_home/omnibase_infra 128
04:01, 04:08 Refresh stability lane error: cannot open .git/FETCH_HEAD: Permission denied 255
04:14, 04:19 Refresh stability lane fatal: 'dev' is already checked out at /data/omninode/runtime-sync-worktrees/OMN-12618/... 128

Zero docker compose / up -d / --force-recreate lines appear in any of the 5 job logs, so no
lane was mutated
. The remaining blocker is the runner container's ownership/permissions and
worktree contention against the bind-mounted $OMNI_HOME clone — not registration.

OMN-14900 status after c9b1cc6f — the deploy hop is still blocked

Mode Symptom Status
1 dubious ownership fixed for the dev lane only — the git_clone() wrapper applies -c safe.directory=<clone> per invocation. refresh_stability_lane.sh on dev still contains zero safe.directory handling, so mode 1 remains open for the stability lane.
2 .git/FETCH_HEAD: Permission denied OPEN — not addressed by this PR at all. fetch origin --prune still writes the shared clone as the runner uid, and a detached checkout --force needs strictly more write access than the fetch that already failed.
3 'dev' is already checked out at … fixed for the dev lane only — detached checkout carries no branch identity.

Mode 2 blocks the deploy hop on both lanes. Do not read this PR as making the deploy hop
work.
The three-mode analysis, ranked fixes, and acceptance criteria live in OMN-14900.

NOT proven — do not read this PR as a successful deploy:

  • The dev lane refresh has never been exercised, warm or cold. No lab/dev/** tag has been
    cut, no dev lane refreshed, no dev receipt produced, zero dev-lane runtime readback. (The
    stability-lane runs above all died pre-mutation, so there is no successful lane readback of any
    kind.)
  • The c9b1cc6f hardening in (a) is untested against a real runner. It is a fix for observed
    failure modes, not a fix proven by a green run.
  • verify_dev_refresh.py deliberately omits consumer-group Stable checking (present in the
    stability-test gate). It was built against a lane whose core services were not live, so real
    group names could not be captured; declaring unverified names would be a
    plausible-but-unverified check. Flagged as follow-up, not faked.
  • verify_dev_refresh.py test coverage is one unit test on one seam
    (test_cluster_health_uses_configured_broker_address, added in c9b1cc6f). It covers the
    broker-address parameterization only — not gate behavior, not the manifest/health/contract-count
    checks, not refresh_dev_lane.sh at all.
  • proof_class: code-only.

Pre-existing dev breakage (not introduced here)

pre-commit run --all-files fails one hook on a file this PR does not touch:
tests/ci/test_runner_routing_audit.py has SPDX-FileCopyrightText: 2026 where the validator
requires 2025. It is unchanged since 2f6fc271 on dev and is absent from this diff. Left alone
rather than smuggled into an unrelated PR.

Notes

  • No OCC evidence companion is authored by this PR — the implementer must not author their own
    evidence. The pre-push deploy-scope DoD hook returned NOTICE_COMPANION_UNMERGED for
    docker/docker-compose.runners.yml; the hosted deploy-gate resolves it via Evidence-Source.
  • Follow-ups filed, not fixed here: OMN-14900 (deploy hop blocked, mode 2 open),
    OMN-14901 (sibling-clone --force + mixed-ref receipt).
  • Do not auto-merge. Merges are Codex's.

Summary by CodeRabbit

  • New Features

    • Added release-train workflows for cutting and deploying development and stability lab releases.
    • Added dedicated deployment runner support.
    • Added health-gated development refreshes with warm and cold recovery paths.
    • Added automatic rollback when warm refresh health checks fail.
    • Added deployment receipts and workflow summaries with refresh results.
    • Added dry-run and manual canary procedures for safer operations.
  • Documentation

    • Added a runbook covering lab-lane releases, runner setup, validation, rollback, and operational troubleshooting.

Evidence-Ticket: OMN-14889
Evidence-Source: OCC#4555

…ty lab lanes

Re-cut of closed PR #2377 (commit c2ec0d2) onto current dev after #2370
(OMN-14873) merged as 54938fd. File-scoped restore of the 6 paths only; dev's
newer refresh_stability_lane.sh (detached-checkout fix) is left untouched.

- .github/workflows/release-train-lab.yml (new)
- docker/docker-compose.runners.yml (+74, additive)
- docs/runbooks/release-train-lab.md (new)
- scripts/runtime_build/cut_release_train_tag.sh (new)
- scripts/runtime_build/refresh_dev_lane.sh (new)
- scripts/runtime_build/verify_dev_refresh.py (new)
@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds a tag-driven release-train workflow for dev and stability lab lanes, a dedicated deployment runner, tag-cutting automation, a cold-aware dev refresh with rollback, and a fail-closed health gate producing durable JSON receipts.

Changes

Release-train lab deployment

Layer / File(s) Summary
Tag cutting and workflow dispatch
.github/workflows/release-train-lab.yml, scripts/runtime_build/cut_release_train_tag.sh, docs/runbooks/release-train-lab.md
Manual inputs and lab tag pushes dispatch lane-specific jobs; the tag script creates tags across sibling clones and pushes from the anchor repository.
Dedicated runner provisioning
docker/docker-compose.runners.yml, docs/runbooks/release-train-lab.md
Adds the omninode-deploy-runner service, dedicated credentials, host-path and Docker access, provisioning steps, canary execution, and documented scope boundaries.
Dev lane refresh and rollback
scripts/runtime_build/refresh_dev_lane.sh
Adds warm, cold-aware restart, and cold-aware full refresh paths with ancestry checks, health-gate execution, rollback/reverification, receipts, and exit codes.
Health-gate checks and result contract
scripts/runtime_build/verify_dev_refresh.py
Adds structured reports and fail-closed checks for image changes, revisions, manifest contracts, HTTP health, and broker cluster health.

Estimated code review effort: 5 (Critical) | ~120 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Operator
  participant TagScript
  participant GitHubWorkflow
  participant DeployRunner
  participant RefreshScript
  participant HealthGate

  Operator->>TagScript: cut and push lab tag
  TagScript->>GitHubWorkflow: trigger tag-push workflow
  GitHubWorkflow->>DeployRunner: run lane deployment
  DeployRunner->>RefreshScript: invoke refresh script
  RefreshScript->>HealthGate: verify refreshed runtime
  HealthGate-->>RefreshScript: return gate result
  RefreshScript-->>GitHubWorkflow: write receipt and exit status
Loading
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: a git-tag-driven release-train trigger for dev and stability lab lanes.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch jonah/omn-14889-release-train-recut

Comment @coderabbitai help to get the list of available commands.

The workflow header, runbook, and runners compose block all asserted the
omnibase-deploy runner was unprovisioned and that the deploy job "will queue
but never pick up." That is FALSE and safety-relevant: a reviewer would
conclude a lab/** tag push is inert when it in fact triggers a real lane
refresh in execute mode.

Verified live 2026-07-21:
  gh api repos/OmniNode-ai/omnibase_infra/actions/runners
  -> omninode-deploy-runner status=online,
     labels [self-hosted, Linux, X64, omnibase-deploy]

That is the exact label set both jobs target, and registration is at the
REPOSITORY level (not the org level the runbook recipe assumed).

Corrected all three sites to state: the runner is online; a lab/dev/** or
lab/stability/** tag push DOES refresh the lane in execute mode (intended,
since dev+stability are pre-authorized, but not a no-op); and the remaining
blocker is host-side git access to the ambient $OMNI_HOME clone, not
registration.

Five runs on 2026-07-21T03:59-04:19Z all failed in the "Refresh stability
lane" step, in three distinct forms as intermediate fixes were attempted:
dubious ownership (128), .git/FETCH_HEAD Permission denied (255), and 'dev'
already checked out in a sibling worktree (128). All died before any
container action -- zero docker compose / up -d / --force-recreate lines in
any of the five logs -- so no lane was mutated.

Prod boundary statements preserved and reinforced: lab/** and v* remain
disjoint namespaces and prod promotion still routes through
node_redeploy_orchestrator's grant gate.

No functional change; comments and docs only.

Closes OMN-14889

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🧹 Nitpick comments (4)
scripts/runtime_build/verify_dev_refresh.py (2)

318-367: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Fail-closed health gate has no test coverage yet.

run_health_gate/HealthGateReport.overall gate a production-adjacent deploy path and already accept injectable runner/opener for testability, but per PR objectives this file has no test coverage. Given the criticality (fail-closed AND-of-all-checks) and that it hasn't been exercised end-to-end yet, unit tests for overall's PASS/FAIL/INFRA_ERROR branching would be cheap to add given the existing DI seams.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/runtime_build/verify_dev_refresh.py` around lines 318 - 367, **
Add unit tests in the existing test area for run_health_gate using injected
runner and opener doubles, covering HealthGateReport.overall’s PASS, FAIL, and
INFRA_ERROR branches. Exercise the fail-closed AND behavior by varying service,
manifest, health, and cluster checks, and assert the resulting overall status
and relevant report fields.

264-282: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

status_code != 200 branch is effectively unreachable — urlopen raises HTTPError (a subclass of URLError) for non-2xx responses before this check runs.

Since except (urllib.error.URLError, OSError) already catches HTTPError, a real non-200 health response surfaces as the generic "health fetch failed: ..." message instead of the dedicated "health endpoint returned HTTP {status_code}" message. PASS/FAIL correctness is unaffected, but this reduces diagnostic clarity right as this gate is about to be exercised for the first time in production.

♻️ Proposed fix: distinguish HTTPError explicitly
     try:
         with open_fn(health_url, timeout=10) as resp:
             status_code = getattr(resp, "status", 200)
             raw = resp.read()
+    except urllib.error.HTTPError as exc:
+        return False, f"health endpoint returned HTTP {exc.code}"
     except (urllib.error.URLError, OSError) as exc:
         return False, f"health fetch failed: {exc}"
-    if status_code and status_code != 200:
-        return False, f"health endpoint returned HTTP {status_code}"
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/runtime_build/verify_dev_refresh.py` around lines 264 - 282, Update
check_health to catch urllib.error.HTTPError separately from the broader
URLError/OSError handler, return the dedicated “health endpoint returned HTTP”
diagnostic using the exception’s HTTP status, and retain the existing generic
fetch-failure handling for other URL or OS errors.
docker/docker-compose.runners.yml (2)

948-949: 🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick win

Hardcoded Docker socket GID is host-specific and fragile.

group_add: - "984" bakes in one host's Docker group GID. If this compose file is ever run on a different host/distro (or the host's docker group GID changes), DinD access silently breaks with socket permission errors, which would fail exactly the builds/deploys this runner exists to perform.

🔧 Suggested fix
+    environment:
+      DOCKER_GID: ${DOCKER_GID:-984}
     group_add:
-      - "984" # Match host Docker socket GID for DinD access
+      - "${DOCKER_GID:-984}" # Match host Docker socket GID for DinD access; override via DOCKER_GID if host differs
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docker/docker-compose.runners.yml` around lines 948 - 949, Replace the
hardcoded Docker group ID in the runner service’s group_add configuration with a
runtime-configurable value sourced from the host environment, while preserving
the existing Docker socket access behavior and providing an appropriate fallback
or required-variable validation.

925-947: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Reuse the shared runner anchors here. omninode-deploy-runner duplicates the shared runner config instead of merging *runner-base / *runner-env, so future changes to common settings can drift out of sync. Merge the anchors and keep only the deploy-specific overrides.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docker/docker-compose.runners.yml` around lines 925 - 947, Update the
omninode-deploy-runner service to merge the shared runner-base and runner-env
anchors, removing duplicated common configuration while retaining only
deploy-specific overrides such as the runner name, labels, group, token, and
deploy-specific mounts or settings.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/release-train-lab.yml:
- Around line 151-184: Update the “Publish receipt to job summary” step to
handle an unset or unrecognized LANE explicitly before selecting a state
directory. For an unknown lane, report that no receipt can be associated with
the run and skip reading any receipt file; retain the existing stability and dev
directory mappings only for recognized lane values.

In `@scripts/runtime_build/refresh_dev_lane.sh`:
- Around line 60-70: Move the ~/.omnibase/.env existence check and source
operation outside the runtime-policy.env conditional in the refresh script. Keep
runtime-policy.env loading guarded by its own check, while ensuring
~/.omnibase/.env is sourced independently whenever present.
- Around line 282-341: The warm-path ancestry check in the NEW_REFS loop reads
stale sibling clone heads because only INFRA_CLONE is refreshed. Before
collecting NEW_REFS when BRANCH is warm, refresh every sibling in
ALL_TRACKED_REPOS to the requested REF using the same fetch/checkout/reset flow,
then perform the existing ancestry comparisons; preserve the cold-aware path
unchanged.

In `@scripts/runtime_build/verify_dev_refresh.py`:
- Around line 285-315: Update check_cluster_health to obtain the broker address
from the appropriate environment variable instead of hardcoding "redpanda:9092"
in the rpk arguments. Preserve the existing health-check command and behavior,
including the current broker-container parameter.

---

Nitpick comments:
In `@docker/docker-compose.runners.yml`:
- Around line 948-949: Replace the hardcoded Docker group ID in the runner
service’s group_add configuration with a runtime-configurable value sourced from
the host environment, while preserving the existing Docker socket access
behavior and providing an appropriate fallback or required-variable validation.
- Around line 925-947: Update the omninode-deploy-runner service to merge the
shared runner-base and runner-env anchors, removing duplicated common
configuration while retaining only deploy-specific overrides such as the runner
name, labels, group, token, and deploy-specific mounts or settings.

In `@scripts/runtime_build/verify_dev_refresh.py`:
- Around line 318-367: **
Add unit tests in the existing test area for run_health_gate using injected
runner and opener doubles, covering HealthGateReport.overall’s PASS, FAIL, and
INFRA_ERROR branches. Exercise the fail-closed AND behavior by varying service,
manifest, health, and cluster checks, and assert the resulting overall status
and relevant report fields.
- Around line 264-282: Update check_health to catch urllib.error.HTTPError
separately from the broader URLError/OSError handler, return the dedicated
“health endpoint returned HTTP” diagnostic using the exception’s HTTP status,
and retain the existing generic fetch-failure handling for other URL or OS
errors.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 6a2915a9-1db9-4af8-bfc7-7361c0918875

📥 Commits

Reviewing files that changed from the base of the PR and between 54938fd and 792c3b7.

📒 Files selected for processing (6)
  • .github/workflows/release-train-lab.yml
  • docker/docker-compose.runners.yml
  • docs/runbooks/release-train-lab.md
  • scripts/runtime_build/cut_release_train_tag.sh
  • scripts/runtime_build/refresh_dev_lane.sh
  • scripts/runtime_build/verify_dev_refresh.py

Comment thread .github/workflows/release-train-lab.yml
Comment thread scripts/runtime_build/refresh_dev_lane.sh
Comment thread scripts/runtime_build/refresh_dev_lane.sh
Comment thread scripts/runtime_build/verify_dev_refresh.py
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant