Skip to content

ci(OMN-18291): route one container-backed job class to the cloud CI fleet - #3478

Merged
jonahgabriel merged 4 commits into
devfrom
jonah/omn-18291-cloud-canary-route
Sep 13, 2026
Merged

jonahgabriel merged 4 commits into
devfrom
jonah/omn-18291-cloud-canary-route

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

Closes OMN-18291 (part 2 of 2). Parent epic OMN-18205. Part 1 is omninode_infra#1414, which builds the capacity trigger this routing depends on.

What changes

One job class moves to the AWS cloud runner fleet: rebuilt-postgres16-proof, the single job of .github/workflows/application-acl-postgres16-proof.yml.

It is routed through a new dedicated variable read ahead of the docker seam, not by flipping the docker seam itself.

fork PR -> OMNI_PUBLIC_PR_RUNS_ON_JSON            (unchanged, still first)
        -> OMNI_CLOUD_CANARY_RUNS_ON_JSON         (new, this job class only)
        -> OMNI_DOCKER_CI_RUNS_ON_JSON            (unchanged, 14 other job definitions)
        -> OMNI_TRUSTED_CI_RUNS_ON_JSON           (unchanged)
        -> literal fallback                        (unchanged)

Why not the docker seam

Measured from each job's parsed runner selection on dev, not from a file grep: the docker seam governs fifteen job definitions here. Several run on every pull request and one feeds the required CI Summary umbrella. The cloud fleet's maximum is four ephemeral instances, each of which runs exactly one job and then terminates, and a boot costs minutes.

Pointing fifteen job definitions at four single-job instances is not a canary. It is the saturation failure already recorded in this file's trusted-seam entry, repeated on a twentieth of the capacity. So the docker seam is left byte-identical and the canary gets a variable of its own.

Why this job class

  • It is a real job on a normal trigger, not a synthetic probe.
  • It is genuinely container-backed: docker compose up --build against PostgreSQL 16, proving an ACL matrix and its rollback. That is the one capability this fleet has and the lab fleet does not, because a lab runner is itself a container reaching a shared daemon through a mounted socket.
  • Its trigger is path-filtered rather than every-pull-request.
  • It appears in no required-context list and in no external-context tuple, checked live rather than assumed.
  • Its fork guard already sends fork pull requests to hosted runners, and it is untouched and still evaluated first.

The honest cost

This repository is public, so its hosted minutes are free and a cloud instance is not. This flip costs money where the omnibase_core move recorded in the same file saved it. That is accepted deliberately and written into the policy entry: the purpose is to prove the trigger and the routing path end to end on a real job with a small blast radius. The repositories where the fleet pays for itself are the private ones the runner group grants, and moving one of those is a separate claim with its own capacity measurement.

Routing rule, every step

Scopes enumerated live immediately before the first write, at 09:31:31Z: the organisation and all eleven audited repositories, all absent, with a positive control against a variable known to exist so the absence is a finding rather than a broken query.

One write, at 09:35Z: the repository-scoped variable on this repository. Every other scope read back individually afterwards and still absent. The out-of-scope variables are unchanged, proven by their own last-updated timestamps, all of which predate this work by days or weeks.

Variable Value Last updated
OMNI_TRUSTED_CI_RUNS_ON_JSON (org) ["ubuntu-latest"] 2026-08-27
OMNI_PUBLIC_PR_RUNS_ON_JSON (org) ["ubuntu-latest"] 2026-05-02
OMNI_TRUSTED_CI_RUNS_ON_JSON (repo) ["ubuntu-latest"] 2026-09-07
OMNI_DOCKER_CI_RUNS_ON_JSON (repo) ["ubuntu-latest"] 2026-08-28
OMNI_SECURITY_SCAN_RUNS_ON_JSON (repo) ["ubuntu-latest"] 2026-08-25

Placement counts, from parsed runner selection on the default branch:

Class Before After
cloud canary variable 0 1
docker seam 15 14
hosted literal 18 19
trusted seam 136 136
lab fleet literal 13 13
security scan seam 1 1
delegated to callee 25 25
other expression 6 6
total 214 215

The hosted literal gains one because of the capacity requester described below; the total gains one job definition, not one placement change.

Asserted intent updated in the same change. config/runner_routing_policy.yaml declares the new variable and scripts/audit-runner-routing.py reads it, so an undeclared shadow appearing on any other repository is a reported finding rather than an invisible edit. The absence half matters more here than for the two existing scoped variables: this one points at a fleet that bills per instance, so a stray shadow spends money rather than merely moving a job.

The audit was run live in both directions. Before the write it correctly reported the declared-but-absent shadow, which is the positive control proving the pass is not vacuous. After the write both passes report Runner routing audit passed.

The capacity requester

The fleet is scale-to-zero, so a job routed there with nothing to start it simply queues. A request-cloud-capacity job asks for one runner over OIDC before the routed job is scheduled, and the routed job declares needs: on it.

It is hosted by literal, which is load-bearing rather than lazy: it is the thing that causes the fleet to exist for this run, so it cannot resolve through the variable it is serving, and a literal pin means no future flip of any seam can relocate the requester onto the capacity it is trying to request. That is why this workflow joins the hosted allowlist, and the reason is recorded there.

It holds no autoscaling permission. The decision lives in the function described in part 1, which clamps to the group's live maximum and refuses to lower the dial.

For a fork pull request the requester is skipped and the routed job still runs, on hosted runners, exactly as it did before this change.

Proof this job lands where the routing says

A first step reads the runner name and, on a cloud runner, the instance metadata service, and prints the placement. A green check does not say where a job ran, and where is the entire claim. The step is deliberately non-fatal: landing elsewhere is a finding to read in the log, not a reason to fail a proof that is otherwise valid.

Tests

tests/ci/test_runner_routing_audit.py gains five: that the script's audited keys and the shipped declarations cannot drift apart in either direction, that the variable names the cloud fleet label (the same variable holding a hosted value would read as a live canary everywhere while routing nothing), that the precedence order is real when read from the parsed runner selection rather than from a mention in a comment, with the fork branch asserted to still come first, that the requester cannot resolve onto the fleet it requests, and that its file is allowlisted.

Evidence-Ticket: OMN-18291
Evidence-Source: OCC#9324

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hostile Reviewer — adversarial findings (OMN-17492)

Models succeeded: glm-review
Models failed: codex
New finding threads: 0
Deduped (already posted on this PR): 0
Nit-level findings suppressed: 1

The model is the FINDER, never the gate: merge is gated only by the
deterministic Hostile Review Thread Gate, which blocks while
hostile-reviewer threads are unresolved. Resolve each thread after
addressing (or rejecting, with a reply) its finding.

Findings not anchored to a changed file

  • [CRITICAL] hostile-reviewer (glm-review)

    Error detection greps for a field the query strips out | The aws lambda invoke command uses --query '{status:StatusCode,error:FunctionError}' to project only two keys into scaler-response.json. The subsequent failure check greps that file for '"errorMessage"', which can never match because errorMessage is not part of the projected output. A lambda that raises returns exit code 0 with FunctionError set to 'Unhandled', and the projected JSON will contain '"error": "Unhandled"' but no errorMessage. The job the

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MAJOR] hostile-reviewer (glm-review)

    Scale-up request is fire-and-forget against a minutes-long boot | The Lambda is invoked once and its 200 response is treated as proof of capacity, but the comments elsewhere in this diff state that a fleet boot costs minutes and the fleet is scale-to-zero. Completing the requester job does not mean a runner with the omni-cloud-ci label has registered; the routed job can still sit queued and burn its 30-minute timeout against fleet boot time plus queue time. The revert condition in the policy file even names

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Runner-name prefix is a weak discriminator for IMDS access | The placement check trusts RUNNER_NAME matching 'omni-cloud-i-*' before querying the instance metadata service. The comment claims IMDS is an independent discriminator, but the gate to reach it is runner-controlled naming; any runner image or bootstrap that sets RUNNER_NAME to the cloud prefix, or any misrouted job on a machine that does expose IMDS, will report PLACEMENT: cloud fleet. Because the check is deliberately non-fatal, the risk is limit

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    No guard on unset routing OIDC variables | role-to-assume interpolates vars.AWS_ACCOUNT_ID and vars.AWS_CI_RUNNER_SCALER_ROLE_NAME with no fallback or validation. If either variable is unset the ARN is malformed, the credential step fails, and under needs the downstream job is silently skipped; the error message will not say why. Combined with the requester's own broad skip/failure semantics, an environment misconfiguration presents as a vanished proof job. | Evidence: role-to-assume: arn:aws:iam::${{ vars.

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Cancellation between scale-up and job start leaks a billed instance | The workflow sets cancel-in-progress concurrency. If the run is cancelled after the Lambda has raised the fleet maximum but before or during the routed job's queue wait, the instance boots and terminates unused. The diff mentions scale-down as automatic and bounded by guard alarms, but the requester performs no compensating scale-down on cancellation and no wait mechanism exists to shorten the exposure window. | Evidence: cancel-in-progre

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Workflow-level tests parse YAML that GitHub will never parse as YAML | test_the_canary_variable_is_read_ahead_of_the_docker_seam asserts substring ordering inside the raw runs-on template string. GitHub evaluates runs-on as an expression; YAML parsing proves only source-order of the variable names, not evaluation order. The two coincide today because the expression uses || short-circuiting in the stated precedence, but the test would pass equally against an expression where evaluation order differs from tex

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Downstream skip semantics conflate skipped requester with viable run | The routed job's if excludes only failure and cancelled requester results, so 'skipped' proceeds. This is intentional for fork PRs, but it means any future change that causes the requester to skip for an unintended reason (a new event type, an org-level condition) routes the proof to the cloud label with no capacity request behind it. The condition encodes one expected skip reason implicitly rather than explicitly. | Evidence: if: >- alw

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

@github-actions

github-actions Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ Hostile Reviewer — DEGRADED (informational)

Critical findings: 1
Major findings: 2
Total findings: 7
Models succeeded: glm-review

Note: Fewer than 2 reviewer models succeeded. Degraded results are informational (OMN-8468/OMN-8524) and do not block merge. Error: cli_review exit 2 (fewer than 2 models succeeded — partial/total outage)


Semantics (OMN-17492 — the model finds, thread resolution gates)

Surface Meaning Blocks merge?
Review threads Per-finding, posted by the reviewer No (informational)
Hostile Review Thread Gate Deterministic: unresolved hostile-reviewer threads exist Fails until resolved (not yet a required context)
degraded verdict Fewer than 2 models succeeded (infra) No

Powered by omniintelligence.review_pairing.cli_review — multi-model adversarial review: qwen3-review, qwen3-review-b, glm-review (OMN-8468/OMN-8524/OMN-17492)

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hostile Reviewer — adversarial findings (OMN-17492)

Models succeeded: glm-review
Models failed: codex
New finding threads: 0
Deduped (already posted on this PR): 0
Nit-level findings suppressed: 0

The model is the FINDER, never the gate: merge is gated only by the
deterministic Hostile Review Thread Gate, which blocks while
hostile-reviewer threads are unresolved. Resolve each thread after
addressing (or rejecting, with a reply) its finding.

Findings not anchored to a changed file

  • [MAJOR] hostile-reviewer (glm-review)

    Error detection greps for the wrong field and inspects only query metadata | The aws lambda invoke command uses --query to emit only {status: StatusCode, error: FunctionError} into scaler-response.json. The payload, where a handled error's errorMessage would live, is discarded. FunctionError carries the value 'Handled' or 'Unhandled', never the message text, so grep -q '"errorMessage"' can never match anything the command writes. A handled Lambda failure therefore reports success, the guard exits 0, and t

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MAJOR] hostile-reviewer (glm-review)

    Job ordering does not guarantee runner registration; race with fleet boot | needs: sequences GitHub's scheduling of the dependent job, not the fleet's readiness. The Lambda returns as soon as it has raised the ASG desired capacity; a cold instance takes minutes to boot, register, and accept jobs. During that window the routed job is offered to a group with zero registered runners and queues. The workflow's own policy file treats 'observed QUEUED rather than running' as a revert condition, so the design sh

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Cancelled or failed runs leave an already-requested instance to bill unused | The Lambda invocation is fire-and-forget. If the requester job is cancelled after the invoke succeeds, or the run is cancelled by the concurrency group while the requester is in flight, the raised capacity has no dependent job to consume it. The stated scale-down mechanism (OMN-18224) is an exit trap on the runner job; nothing here scales back down when the consumer never starts. On a public repository this is a cost path, and the

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Guard step loses set -e and tolerates silent metadata-probe failure | The placement step uses set -uo pipefail without -e, which is deliberate for the non-fatal design, but the consequence is unexamined: a failed IMDSv2 token request leaves TOKEN empty and the instance-id curl emits an AWS error page into the log with no annotation, indistinguishable from a placement proof. A future fleet misconfiguration (IMDS hop limit, metadata options disabled) yields a log that reads as 'not the cloud fleet' rather t

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Fork-guard interplay leaves failure semantics of always() under-specified | The dependent job's condition admits the requester's 'skipped' result only, excluding failure and cancelled. If the scaler role assumption fails (expired trust, region outage), the requester fails, the proof job is skipped, and a path-filtered required-optional job silently does not run. That is defensible, but no test asserts it and the revert clause conflates 'capacity trigger did not fire' with the queued-job symptom; an auth fai

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    New tests validate workflow shape but nothing exercises the scaler invocation contract | The added tests thoroughly pin routing precedence, allowlist entry, and the literal hosted pin, all static structure. Nothing covers the actual runtime contract with the Lambda: payload schema, the error-detection guard (which, per the separate correctness finding, does not work), or the readiness window. The comment block in the policy file records a prior incident where placement proof and job capability diverged; the

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hostile Reviewer — adversarial findings (OMN-17492)

Models succeeded: glm-review
Models failed: codex
New finding threads: 0
Deduped (already posted on this PR): 0
Nit-level findings suppressed: 1

The model is the FINDER, never the gate: merge is gated only by the
deterministic Hostile Review Thread Gate, which blocks while
hostile-reviewer threads are unresolved. Resolve each thread after
addressing (or rejecting, with a reply) its finding.

Findings not anchored to a changed file

  • [CRITICAL] hostile-reviewer (glm-review)

    Lambda failure detection greps a key the query never emits | The step renames FunctionError to a key named 'error' via --query '{status:StatusCode,error:FunctionError}'. The subsequent failure check greps for '"errorMessage"', which cannot appear in the query-projected output. A lambda that raises or throws yields {"status":200,"error":"Unhandled"}; the grep finds nothing, the job reports success, and the routed job queues until its runner never arrives. This is precisely the silent-success mode the inline

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MAJOR] hostile-reviewer (glm-review)

    No verification that a runner actually registered after scale_up | The lambda accepting the invocation does not prove the fleet produced a runnable runner. If the group is already at its maximum of four, or the ASG launch fails, the function can return 200 with no scaling performed (its contract is to clamp and refuse to lower, not to fail when it cannot raise). The workflow then marks the requester green and the routed job queues for the full window with no diagnostic. The revert condition in the policy fi

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MAJOR] hostile-reviewer (glm-review)

    Scale-up requested but fleet readiness not awaited; serial dependency adds queue latency | aws lambda invoke returns when the function returns, not when the ASG has launched, bootstrapped, and registered the runner. Instance boot 'costs minutes' per the policy comment, so the downstream job will routinely sit in queue after its dependency completes. On a four-instance single-job fleet with cancel-in-progress concurrency, rapid successive pushes serialise: each cancelled run may leave a just-requested instan

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Capacity request runs on every push to every branch | The requester's only gate is the fork guard. Any push to any branch, plus workflow_dispatch and other events, invokes the scaler and potentially bills an instance, since the downstream proof job appears path-filtered but the requester is not. The policy file's cost discussion addresses the canary job but not this unconditional trigger surface. | Evidence: if: >- github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == githu

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Unset AWS_ACCOUNT_ID or role-name variables produce a malformed ARN late and opaquely | vars.AWS_ACCOUNT_ID and vars.AWS_CI_RUNNER_SCALER_ROLE_NAME are interpolated directly into the role ARN. If either is unset or blank in a forked or rehosted context, configure-aws-credentials fails with an AssumeRole error against an ARN like 'arn:aws:iam:::role/', which does not point the operator at the missing variable. The canary variable, by contrast, has explicit fallbacks in the runs-on expression. | Evidence: rol

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Placement proof suppresses exit-on-error and ignores token fetch failure | The step uses 'set -uo pipefail' without -e deliberately, but within the cloud branch a failed IMDS token request yields an empty TOKEN that is then sent as a header, producing a 401 response echoed as the instance id. The log then asserts 'PLACEMENT: cloud fleet' on the strength of a RUNNER_NAME pattern match while the independent discriminator cited in the comment never actually corroborated anything. | Evidence: TOKEN="$(curl -sS

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    No test covers the workflow's failure path for the scaler response | The new tests assert ordering, allowlisting, and the needs edge, but the only runtime branching logic introduced in the diff (the errorMessage grep) has no test at any level. Given the grep checks the wrong key, a single fixture asserting the projected JSON shape would have caught it. | Evidence: if grep -q '"errorMessage"' scaler-response.json; then ... exit 1 | Fix: Extract the scaler response handling into a small script or composite ac

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Cancelled requester leaves fleet capacity raised with no in-workflow compensation | The diff states scale-down is handled elsewhere (OMN-18224 exit trap), but a run cancelled between lambda invocation and runner consumption, or a runner that fails to register and consume the slot, leaves the raised dial with no job attached. The policy revert conditions cover a queued job but not an orphaned capacity request from a cancelled run. | Evidence: Scale-DOWN has been automatic since OMN-18224 (an exit trap, bound

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hostile Reviewer — adversarial findings (OMN-17492)

Models succeeded: glm-review
Models failed: codex
New finding threads: 1
Deduped (already posted on this PR): 0
Nit-level findings suppressed: 1

The model is the FINDER, never the gate: merge is gated only by the
deterministic Hostile Review Thread Gate, which blocks while
hostile-reviewer threads are unresolved. Resolve each thread after
addressing (or rejecting, with a reply) its finding.

Findings not anchored to a changed file

  • [CRITICAL] hostile-reviewer (glm-review)

    Failure detection greps for a key the queried output cannot contain | The step queries the invocation result into {status: StatusCode, error: FunctionError}, so the scaler-response.json written to disk contains at most a key named "error" with a value such as "Unhandled". The subsequent guard greps for '"errorMessage"', which is a key of the raw Lambda JSON, not of the projected query output. A Lambda that raises will therefore pass the guard, exit 0, and the downstream job will queue silently for up to its

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MAJOR] hostile-reviewer (glm-review)

    Scaler response payload is never validated | The comment states the function 'owns the decision: it reads the group's live maximum, clamps to it, and refuses to lower the dial'. A Lambda can return HTTP 200 with a structured body indicating it declined or no-op'ed (clamped to zero, throttled internally, rejected the group) without raising FunctionError. The workflow inspects only the error marker and never parses the body for a confirmation that desired capacity was actually raised. The claim 'this job hold

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    IMDSv2 token fetched without -f, response never checked | In the placement proof, the token PUT uses curl -sS --max-time 5 without --fail; on failure curl exits nonzero, but the step uses set -uo pipefail without -e, so the script continues and sends an empty or error-body token header to the metadata endpoint. The step is deliberately non-fatal, but as written a transient IMDS failure produces a misleading 'PLACEMENT: cloud fleet' branch preceded by unvalidated output, weakening the independent-discrimin

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Race between scale-up acceptance and runner registration is unbounded | The Lambda sets desired capacity and returns; EC2 boot, bootstrap (plugin installs, digest verification per the policy comments), and runner registration take minutes with no health signal returned to the workflow. The routed job then queues on GitHub's side with no feedback loop: if the boot fails, the run waits out its 30-minute timeout rather than failing fast. The policy file lists 'observed QUEUED rather than running' as a revert c

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

  • [MINOR] hostile-reviewer (glm-review)

    Downstream runs on scaler failure only via explicit result checks; 'action_required' and other states unhandled | The downstream if checks result != 'failure' && result != 'cancelled' under always(). A 'skipped' result correctly proceeds (fork case), and 'success' proceeds, but the expression implicitly permits 'action_required' (environments awaiting approval) and any future state. Currently no environment is used, so impact is nil, but the guard encodes an allowlist-by-blacklist rather than whitelist. | E

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

Findings demoted from threads (anchor rejected)

  • [MAJOR] hostile-reviewer (glm-review)

    No test covers the capacity-request failure path | The new tests assert structural properties (pinning, ordering, allowlisting, policy drift) but nothing exercises the shell logic of the requester step: the grep guard, the query projection, or behavior when the Lambda returns FunctionError. Given finding #1, a test invoking the guard against a realistic Lambda error envelope would have caught the bug before merge. The test file's own stated philosophy ('a routing change is finished when a real job of that c

    Resolve this thread when addressed — the Hostile Review Thread Gate blocks while hostile-reviewer threads are unresolved (OMN-17492).

jonahgabriel pushed a commit to OmniNode-ai/onex_change_control that referenced this pull request Sep 13, 2026
#9324)

* evidence: OCC companion pass 1 for OmniNode-ai/omnibase_infra#3478

* evidence: OCC companion self-bind for #9324

---------

Co-authored-by: node-occ-companion-effect <occ-companion-effect@omninode.ai>
@jonahgabriel
jonahgabriel merged commit a77df3a into dev Sep 13, 2026
123 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-18291-cloud-canary-route branch September 13, 2026 10:30
Patel230 pushed a commit to OmniNode-ai/onex_change_control that referenced this pull request Sep 16, 2026
#9324)

* evidence: OCC companion pass 1 for OmniNode-ai/omnibase_infra#3478

* evidence: OCC companion self-bind for #9324

---------

Co-authored-by: node-occ-companion-effect <occ-companion-effect@omninode.ai>
Patel230 pushed a commit that referenced this pull request Sep 16, 2026
…leet (#3478)

* ci(OMN-18291): route one container-backed job class to the cloud CI fleet

* docs(OMN-18291): record what the first real routed run found

* docs(OMN-18291): the second capability gap the routed run found

* docs(OMN-18291): the third capability gap, and the general lesson
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant