Skip to content

fix(OMN-15617): resolve bash>=5 explicitly instead of trusting PATH order - #2651

Merged
jonahgabriel merged 1 commit into
devfrom
jonah/omn-15617-fix-200-gate-host-ssh-path-resolves-bash-3257-failing-15
Aug 4, 2026
Merged

jonahgabriel merged 1 commit into
devfrom
jonah/omn-15617-fix-200-gate-host-ssh-path-resolves-bash-3257-failing-15

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Aug 4, 2026 •

Copy link
Copy Markdown
Collaborator

OMN-15617 — .200 gate host bash-version resolution

Fixes the failure OMN-15617 documents: on stickybeatz-studio (.200, the
rule-11a default gate host), non-interactive ssh sessions resolve bash to
the system 3.2.57 shell even though a modern bash 5.x sits at
/opt/homebrew/bin/bash — it just is not first on PATH for that session
class. docker/runners/runner-monitor.sh uses declare -A (bash>=4), so
all 15 tests in test_runner_monitor_wedge_detection.py (the exact count
the ticket names) fail there on every commit — a bash syntax error deep
inside a subprocess, not a resolvable "wrong interpreter" diagnostic.

Fix shape (per ticket AC — explicit resolution + fail-closed canary, no silent fallback/skip)

  • scripts/ci/resolve_modern_bash.sh — new bash-3.2-safe resolver, the
    single source of truth for finding a bash interpreter >=5, independent
    of PATH order (checks OMNIBASE_INFRA_BASH_BIN override, the two brew
    prefixes, then every bash on $PATH). Fails loud with a pointed
    remediation message when none is resolvable — never a silent fallback,
    never a quiet skip. Must itself run under bash 3.2 (no declare -A) since
    it cannot presuppose the thing it's resolving.
  • tests/unit/observability/runner_health/_resolve_modern_bash.py — thin
    pytest wrapper shelling out to the resolver; pytest.fails (not skip) if
    unresolvable.
  • tests/unit/observability/runner_health/test_runner_monitor_wedge_detection.py
    (the flagged 15 tests) and test_runner_monitor_auto_bounce.py (same bug
    pattern, same fix, applied for consistency though only the former is in
    the ticket's named count) — both now invoke the resolved interpreter
    explicitly instead of bare "bash" for the wrapper script and the inner
    monitor.sh invocation.
  • scripts/hooks/prepush_smart_tests.sh — bash>=5 canary added before the
    governed selector runs, using the same resolver script (DRY — single
    source of truth with the pytest harness, so they cannot drift). Exports
    OMNIBASE_INFRA_BASH_BIN so pytest does not have to re-discover it. Fails
    loud with remediation when unresolvable.

Seams

Seam Detail
New env var OMNIBASE_INFRA_BASH_BIN — optional override consumed by scripts/ci/resolve_modern_bash.sh; exported by prepush_smart_tests.sh after resolution so pytest inherits it without re-discovery
Shared script scripts/ci/resolve_modern_bash.sh — stdout contract: absolute interpreter path on success (exit 0), nothing on stdout + pointed message on stderr on failure (exit 1)
Test import tests/unit/observability/runner_health/_resolve_modern_bash.py::resolve_modern_bash() — imported by both test_runner_monitor_wedge_detection.py and test_runner_monitor_auto_bounce.py
Hook call site scripts/hooks/prepush_smart_tests.sh — new block right after REPO_ROOT resolution, before BASE_REF; uses die() (existing helper) on failure
No change to docker/runners/runner-monitor.sh itself, .pre-commit-config.yaml hook wiring, CI workflow files

Verification

Local proxy (macOS ships stock bash 3.2 by default — same class of bug):
under a stock-only PATH (env -i PATH=/usr/bin:/bin), pre-fix code fails
14/15 in test_runner_monitor_wedge_detection.py (1 static/non-bash test
passes); post-fix, 15/15 pass. Full related scope (119 tests,
tests/unit/observability/runner_health/ +
tests/ci/test_prepush_hook_host_identity_guard.py) passes with only the
pre-existing missing-flock skip (unrelated).

.200 (stickybeatz-studio) live validation could not be completed this
session
— ssh stickybeatz-studio returned Permission denied (publickey,password,keyboard-interactive) for this session's identity
(no ssh-agent identities, local id_ed25519 rejected by the remote host).
Gates instead ran on this Mac as a documented rule-11a exception; the local
Mac's stock /bin/bash is also 3.2.57 (identical to .200's), so the
before/after repro above is a faithful proxy for the exact failure mode,
though it is not the designated host itself. Flagging for a follow-up
verification pass with .200 access, or operator confirmation of PATH-order
behavior there post-merge.

Push itself DID exercise the real pre-push hook end-to-end on this Mac and
confirmed the canary fires and logs correctly:

[prepush-smart-tests] bash>=5 canary: resolved /opt/homebrew/bin/bash

followed by the full governed-selector impacted-subset run (3198 passed, 9
skipped, unrelated).

Ticket: OMN-15617

Evidence-Ticket: OMN-15617
Evidence-Source: OCC#6047

…rder

On stickybeatz-studio (.200, the rule-11a default gate host),
non-interactive ssh resolves `bash` to the system 3.2.57 shell even
though a modern bash 5.x sits at /opt/homebrew/bin/bash — it just is
not first on PATH for that session class. runner-monitor.sh uses
`declare -A` (bash>=4), so all 15 tests in
test_runner_monitor_wedge_detection.py fail silently there — a bash
syntax error inside a subprocess, not a resolvable "wrong interpreter"
diagnostic.

- scripts/ci/resolve_modern_bash.sh: new bash-3.2-safe resolver, the
  single source of truth for finding a bash>=5 interpreter
  independent of PATH order (checks OMNIBASE_INFRA_BASH_BIN override,
  the two brew prefixes, then every "bash" on PATH). Fails loud with a
  pointed remediation message when none is resolvable — never a
  silent fallback, never a quiet skip.
- tests/unit/observability/runner_health/_resolve_modern_bash.py: thin
  pytest wrapper shelling out to the resolver script; used by both
  test_runner_monitor_wedge_detection.py (the flagged 15) and
  test_runner_monitor_auto_bounce.py (same bug pattern, same fix).
  Both now invoke the resolved interpreter explicitly instead of bare
  "bash".
- scripts/hooks/prepush_smart_tests.sh: bash>=5 canary added before the
  governed selector runs, using the same resolver script (DRY — single
  source of truth with the pytest harness). Exports
  OMNIBASE_INFRA_BASH_BIN so pytest does not have to re-discover it.
  Fails loud with remediation when unresolvable.

Local proxy verification (macOS ships stock bash 3.2, reproducing the
identical PATH-order bug): under a stock-only PATH
(env -i PATH=/usr/bin:/bin), pre-fix code fails 14/15 in
test_runner_monitor_wedge_detection.py; post-fix, 15/15 pass. Full
related scope (119 tests) passes with only the pre-existing
missing-`flock` skip.
@coderabbitai

coderabbitai Bot commented Aug 4, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 34 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: e66ba7d5-2b76-4648-8773-c90d9e15cc42

📥 Commits

Reviewing files that changed from the base of the PR and between 2e5ef7d and fa5264c.

📒 Files selected for processing (5)
  • scripts/ci/resolve_modern_bash.sh
  • scripts/hooks/prepush_smart_tests.sh
  • tests/unit/observability/runner_health/_resolve_modern_bash.py
  • tests/unit/observability/runner_health/test_runner_monitor_auto_bounce.py
  • tests/unit/observability/runner_health/test_runner_monitor_wedge_detection.py

Comment @coderabbitai help to get the list of available commands.

jonahgabriel pushed a commit to OmniNode-ai/onex_change_control that referenced this pull request Aug 4, 2026
#6047)

* evidence: OCC companion pass 1 for OmniNode-ai/omnibase_infra#2651

* evidence: OCC companion self-bind for #6047

---------

Co-authored-by: node-occ-companion-effect <occ-companion-effect@omninode.ai>
@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

⚠️ Hostile Reviewer — DEGRADED (informational)

Blocking findings (critical): 0
Total findings: 0
Models succeeded: none

Note: All reviewer models failed or were unavailable. Degraded results are informational during the pilot phase (OMN-8468/OMN-8524) and do not block merge. Error: all review endpoints [192.168.86.201:8000 192.168.86.201:8001 ] unreachable — preflight short-circuit (no models available)


Gate semantics (pilot phase)

Verdict Meaning Blocks merge?
passed No critical findings No
blocked CRITICAL findings found Yes
degraded All models unavailable (infra) No (pilot)

Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524)

@jonahgabriel
jonahgabriel merged commit 16b1d7a into dev Aug 4, 2026
148 of 155 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-15617-fix-200-gate-host-ssh-path-resolves-bash-3257-failing-15 branch August 4, 2026 16:40
jonahgabriel added a commit that referenced this pull request Aug 4, 2026
…d-split (#2652)

Addresses two adversarial-verify defects on PR #2651:

- Reorder resolve_modern_bash() before the jq/flock tool-availability skip
  in test_runner_monitor_wedge_detection.py and
  test_runner_monitor_auto_bounce.py. Previously a missing jq (or flock)
  triggered pytest.skip() before the bash>=5 canary ran, so a host with the
  wrong interpreter AND a missing secondary tool would report green-by-
  absence instead of the RED the ticket's AC2 requires.
- Fix scripts/ci/resolve_modern_bash.sh to build CANDIDATES as a bash-3.2-
  safe array instead of a space-joined string. The prior unquoted
  word-splitting silently dropped any interpreter path or PATH entry
  containing a space, risking a silent-wrong-answer resolution in the exact
  script whose purpose is to eliminate that failure mode.

Verified: 102 passed / 4 skipped (flock-only, unrelated) under
env -i PATH=/usr/bin:/bin; resolver returns exit 1 with pointed stderr and
empty stdout when OMNIBASE_INFRA_MIN_BASH_MAJOR is set unreachably high
(proves genuine fail-closed RED); resolver correctly resolves a
space-containing OMNIBASE_INFRA_BASH_BIN path post-fix (previously would
silently drop the fragment).

Ticket: OMN-15617
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant