Skip to content

fix(OMN-15131): stage_workspace.sh invokes check_sibling_lock_pins.py with an interpreter that has pydantic - #2446

Merged
jonahgabriel merged 1 commit into
devfrom
jonah/omn-15131-fix-runners-bare-python3-lacks-pydantic
Jul 25, 2026
Merged

jonahgabriel merged 1 commit into
devfrom
jonah/omn-15131-fix-runners-bare-python3-lacks-pydantic

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Jul 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Fixes OMN-15131: the OMN-14900 stability deploy re-fire (run 30170235015, tag lab/stability/20260725T184501Z-f81d82686973) got past the OMN-15122 step-3b fix (PR #2444 -- contract copy succeeded) but crashed one step later in check_sibling_lock_pins.py:

Traceback (most recent call last):
  File ".../scripts/runtime_build/check_sibling_lock_pins.py", line 83, in <module>
    from pydantic import BaseModel, ConfigDict, Field
ModuleNotFoundError: No module named 'pydantic'
ERROR: sibling-pin preflight failed against .../omnimarket/uv.lock

This was misreported downstream as an OMN-12977 lock-drift condition. It is not -- the check never ran far enough to compare a single pin.

Root cause (verified live, not inferred)

stage_workspace.sh invoked the preflight with the bare python3 resolved off PATH:

python3 "${SCRIPT_DIR}/check_sibling_lock_pins.py" ...

Probed directly on the omninode-deploy-runner container (ssh omni-201-ts, read-only):

  • bare python3 on that container: ModuleNotFoundError: No module named 'pydantic'
  • ${repo_root}/.venv/bin/python (the venv deploy-runtime.sh's own uv sync builds earlier in the same job): pydantic 2.13.4 importable
  • uv run python from the repo root: pydantic 2.13.4 importable

deploy-runtime.sh's own check_sibling_lock_pins() bash function already resolves the interpreter correctly for this exact script (repo-venv python -> uv run -> bare python3 last resort) -- but stage_workspace.sh calls the same script directly, earlier in the pipeline, with none of that resolution.

Fix (class a: wrong interpreter at the call site)

Added resolve_sibling_lock_pins_python() to stage_workspace.sh, mirroring the precedence deploy-runtime.sh's check_sibling_lock_pins() already uses for the identical script, and routed the call site through it instead of the unconditional bare python3.

Rejected alternatives

  • (b) install pydantic into the runner image's system python3 -- rejected: would require a runner-image rebuild for a dependency the repo already vendors in its own venv one directory over; adds a second place pydantic's version has to stay in sync.
  • (c) drop the pydantic import from check_sibling_lock_pins.py -- rejected: the models it defines are load-bearing for the JSON provenance --output other steps (compute_workspace_provenance.py) consume; downgrading to hand-rolled dict validation trades a real typed contract for an untyped one to work around an invocation bug, not a real constraint.

No runner-image rebuild is required -- the fix routes to interpreters that already exist on the runner (the repo's own .venv or uv run), not a new dependency.

Test coverage

Extended tests/scripts/ with test_stage_workspace_sibling_lock_pins_interpreter.py (6 tests, same extract-and-execute-the-bash-function-in-isolation harness OMN-15122/PR #2444 used for resolve_core_contracts_dir()):

  • resolver prefers repo .venv/bin/python when present
  • falls back to uv run when no repo venv
  • falls back to bare python3 only as a last resort
  • regression guard: the call site must not contain an unconditional bare python3 "${SCRIPT_DIR}/check_sibling_lock_pins.py" literal
  • reproduces the exact live-failure condition (no venv staged yet, uv present) and asserts the resolver does NOT select bare python3

Gates run on .200 (this Mac's canonical omnibase_core clone is mid an unrelated unresolved merge, which poisons local uv run/pre-commit -- confirmed unrelated to this change): uv run pytest tests/scripts/ (373 passed), ruff format --check, ruff check, mypy src/omnibase_infra/ (clean), pre-commit run on changed files (clean after an SPDX-header-year fix).

Refs: OMN-15131, OMN-15122 (immediately preceding fix, PR #2444), OMN-14900 (parent stability deploy hop), OMN-12977 (the preflight this crash was misattributed to).

Evidence-Ticket: OMN-15131
Evidence-Source: OCC#4856

Evidence-Commit: c71cbcd0ff7145a5c507dcaa344e6979c36d18da

… with an interpreter that has pydantic

The deploy runner's bare python3 has zero packages installed (not even
pydantic), so check_sibling_lock_pins.py crashed with ModuleNotFoundError
one step after the OMN-15122 contract-resolution fix, misreported
downstream as a sibling-pin drift condition (it never ran far enough to
compare a single pin). Route through the same interpreter precedence
(repo .venv python -> uv run -> bare python3) deploy-runtime.sh's own
check_sibling_lock_pins() bash function already uses for this exact
script, instead of the unconditional bare python3 stage_workspace.sh was
using.

Verified live on the omninode-deploy-runner container: repo .venv/bin/python
and uv run python both have pydantic 2.13.4 importable; bare python3 does
not.

Rejected alternatives:
- install pydantic into the runner image system python3: requires a
  runner-image rebuild for a dependency already vendored in the repo's
  own venv one directory over.
- drop the pydantic import from check_sibling_lock_pins.py: its models are
  load-bearing for the JSON provenance --output other steps consume;
  trades a typed contract for an untyped one to work around an invocation
  bug, not a real constraint.

Refs: OMN-15131

Co-Authored-By: Claude <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jul 25, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 41 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 37d5730f-e986-4f8a-bcfb-97fb0a3bf572

📥 Commits

Reviewing files that changed from the base of the PR and between f81d826 and 4a598dd.

📒 Files selected for processing (2)
  • scripts/runtime_build/stage_workspace.sh
  • tests/scripts/test_stage_workspace_sibling_lock_pins_interpreter.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch jonah/omn-15131-fix-runners-bare-python3-lacks-pydantic

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

⚠️ Hostile Reviewer — DEGRADED (informational)

Blocking findings (critical): 0
Total findings: 0
Models succeeded: none

Note: All reviewer models failed or were unavailable. Degraded results are informational during the pilot phase (OMN-8468/OMN-8524) and do not block merge. Error: all review endpoints [192.168.86.201:8000 192.168.86.201:8001 ] unreachable — preflight short-circuit (no models available)


Gate semantics (pilot phase)

Verdict Meaning Blocks merge?
passed No critical findings No
blocked CRITICAL findings found Yes
degraded All models unavailable (infra) No (pilot)

Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524)

jonahgabriel added a commit to OmniNode-ai/onex_change_control that referenced this pull request Jul 25, 2026
#4856)

* evidence: OCC companion pass 1 for OmniNode-ai/omnibase_infra#2446

* evidence: OCC companion self-bind for #4856

---------

Co-authored-by: node-occ-companion-effect <occ-companion-effect@omninode.ai>
@jonahgabriel
jonahgabriel merged commit ead076a into dev Jul 25, 2026
110 of 123 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-15131-fix-runners-bare-python3-lacks-pydantic branch July 25, 2026 19:30
jonahgabriel added a commit that referenced this pull request Jul 25, 2026
…ris (#2448)

* fix(OMN-15134): self-heal workspace-reset hook against root-owned debris

Root-owned .venv artifacts under the omninode-deploy-runner Actions job
workspace jammed the built-in job-started cleanup hook (run 30171892373),
blocking the OMN-14900 stability deploy hop before any workflow step ran.

Root cause (verified live, not inferred): this image's ENTRYPOINT
legitimately starts as root (docker-socket GID fix) before gosu-dropping to
`runner` for every job step -- so a bare `docker exec omninode-deploy-runner
...` without `-u runner` (a same-session `uv run` verification probe cited
in PR #2446's body, mistakenly described as read-only) silently ran as
root and left a root-owned `.venv` (confirmed: `/root/.cache/uv` populated
with a matching timestamp) that the unprivileged workspace-reset hook could
not `rm -rf`.

Fix: harden runner-job-started.sh to fail loud (name the offending paths,
never silently succeed or hang) then self-heal via a narrowly scoped
NOPASSWD sudo rule (Dockerfile, one command + one argument pattern, confined
to this runner's own `_work` tree -- never a general root shell). Rejected
alternatives: (a) flip the image's default USER to `runner` -- rejected,
the ENTRYPOINT's root-phase init (docker-socket GID fix, OMNI_HOME chown,
operator-env copy) genuinely needs root before it gosu-drops, so this would
either break that init or not change docker exec's default identity (a
compose-level `user:` override still shows up as the container's effective
Config.User); (c) forbid `.venv` in the bind-mounted job workspace -- N/A
here, the workspace is NOT bind-mounted from host (verified via a
write-visibility probe), so this was never a DooD path-mismatch bug.

Verified live on omninode-deploy-runner (ssh omni-201-ts): cleaned the 18
confirmed root-owned paths + /root/.cache/uv (count_after=0); reproduced the
exact RED condition (simulated root-owned debris) against the real Ubuntu
22.04 GNU-realpath container and confirmed the hardened hook fails loud with
the offending path + a clear "sudo not installed" diagnostic (pre-rebuild,
no sudoers rule yet) and exits non-zero with a manual-remediation
instruction -- never silently wedging or succeeding. Added
test_runner_job_started_root_owned_debris.py (happy path + undeletable-debris
fail-loud + pre-existing path-confinement regression guard); skips on
non-GNU-realpath hosts (this repo's macOS gate host), runs for real on
Linux CI.

Gates run on .200 (patch-transfer + sha256 verify): ruff format/check,
mypy, full pre-commit (shellcheck included) -- all clean.

Refs: OMN-15134, OMN-14900 (parent stability deploy hop), OMN-15131/PR #2446
(the same-session probe that produced the debris).

Evidence-Ticket: OMN-15134

* chore(OMN-15134): retrigger occ-preflight (Evidence-Source now present on PR body)

---------

Co-authored-by: Jonah Gray <jonah.g.gray@gmail.com>
jonahgabriel added a commit that referenced this pull request Jul 26, 2026
…t catalog, T0/T1 readback, dev gotchas, prod pointer (#2454)

Extends docs/runbooks/release-train-lab.md (existing mechanism doc, not a
new file) with the operator-facing procedure proven live on run
30180376657 (2026-07-26, tag lab/stability/20260725T235956Z-87ec5b3165ce,
overall: PASS) — the terminal success of the six-iteration OMN-14900
hardening chain.

- Copy-paste tag-cut + watch procedure with expected output and a
  tag-content WARNING (never reuse a parked tag name).
- Preflight catalog: what each of the 6 landed fixes (#2450/#2446/#2444/
  #2452/#2434/#2448) catches, with pre-fix failure signatures.
- T0/T1 readback discipline: health 18085/18086, contract-count floor vs.
  regression, discovery_errors baseline, consumer groups, vcs_ref
  ancestry; FAILED_ROLLED_BACK equality-proof discipline.
- Dev-lane gotchas: stale workspace-build config YAML, OMN-14968 false
  FAILED on runtime-worker; pointer (not duplicate) to
  cold-lane-full-bringup.md for cold bring-up.
- Prod: pointer-only section citing CLAUDE.md rules 2a/12 and the
  onex_change_control#4892 prep-only grant-PR pattern; explicit
  raw-docker-mutation prohibition citing the no-raw-prod-bypass gate.
- Verification checklist + rollback section.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant