Repository navigation
Check active model selectors against a versioned native catalog - #943
Conversation
… rows Leave CLAUDE_CODE_WORKFLOW_SUBAGENT_MODEL unset (declined); decline ultrafast_mode (plan-gated, dropped client-side by the OmniRoute profile); keep the Codex cyber access programs as owner information only; decline the model_providers capabilities rows until the currency lane's trial after #943. Rows, decision record and the re-registered catalog hash change together. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
b50f740 to
46cbd7d
Compare
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
…ecord The page-URL diffs go to the command center (through the daily watch service it installs after #943), the changelog-name report to the currency lane, and the settings-drift inventory to cc-native-practice. The weekly practice pass, the TeammateIdle gate, the cost-data check and the WebSearch interaction now say who acts or that nothing is scheduled. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
trading-cc ack, scope limited to the hunks in What the hunks do: they change the Claude model id What I checked at that head: This is an ownership ack for those hunks only. It is not a read of the rest of the PR, and it names this head; a later head needs a fresh ack if those two files change. |
|
trading-cc re-affirms its earlier scoped ack (comment 6094414078) at head 042eb8c. The two files it covers, |
042eb8c to
361f417
Compare
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 94fbc29ab3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
94fbc29 to
a2a873b
Compare
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7e4a0c94fd
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ced5ccf43d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The committed-path proposal at 46cbd7d requires the new model checker, so the reviewed repository inventory no longer matches its native selector. Add exactly the proposed scripts/active_model_currency.py repo-root row and advance this base's hard count from 90 to 91. Preserve every other policy field, existing row order and independently reviewed receipt/user grant. Register the corrected policy and test through the native evidence helper. The full policy module reproduces one failure before the repair and passes all 21 cases afterward. FULL validation passed before this forward commit. At a later authorized landing rebase, derive the next count from every already-landed row rather than reusing this branch-local count. Primary model and supported derivation: 8c57010 https://github.com/seathatflowsinourveins/native-agent-stack/blob/46cbd7de292eb9b82062bd49e65b832f1e0c6902/scripts/local_pages_policy_grants.py
ced5ccf to
2cd2e7b
Compare
… rows Leave CLAUDE_CODE_WORKFLOW_SUBAGENT_MODEL unset (declined); decline ultrafast_mode (plan-gated, dropped client-side by the OmniRoute profile); keep the Codex cyber access programs as owner information only; decline the model_providers capabilities rows until the currency lane's trial after #943. Rows, decision record and the re-registered catalog hash change together. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ecord The page-URL diffs go to the command center (through the daily watch service it installs after #943), the changelog-name report to the currency lane, and the settings-drift inventory to cc-native-practice. The weekly practice pass, the TeammateIdle gate, the cost-data check and the WebSearch interaction now say who acts or that nothing is scheduled. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Scope
Compare active model selectors with retained native Codex, OmniRoute and Anthropic catalog captures, and include stale or incomplete coverage in the existing currency due-file. The checker extracts text model fields, Python defaults/argv, JSON command/argv/allowlist values and named launch/template surfaces. Startup reads the precomputed notice; the audit timer remains a template.
Codex's short
-mcounts at its executable position after supported wrappers. Nice/ionice and hcom launches preserve native arguments while wrapper operands, Git directories, Python modules and ordinary message operands cannot establish client context. Claude uses--model. Quoted or escaped punctuation remains data; genuine shell operators establish command boundaries in strings, while argv operators remain literal.File, post-read and line admission now let model-bearing commands reach structured/POSIX decoding even when their raw model spelling is escaped. Model identity comes from the decoded selector argument. Candidate normalization is only an admission hint; it cannot establish executable context or select a model. Command strings decode escapes, while supplied argv remains literal. When POSIX parsing of a text field or a non-prose command line fails, the checker skips that field or line only when no raw or potentially escaped model-ID candidate is present. Parse failures with such a candidate keep coverage unknown, except on prose command lines, where a failed command parse falls back to sentence matching of the raw line. A readable first field value keeps its existing scope.
Markdown extraction preserves a formatted executable and its operands in one surrounding command. A single-token span decodes POSIX quotes and escapes with upstream
shlex.split, then encodes withshlex.quote. Quoted wrapper names retain their identity. Full command spans keep their independent parse; other span text stays one opaque surrounding token. Multiword comments cannot suppress an earlier formatted model. Literal punctuation stays data, and fences stay literal. Text fields contribute only their first lexical value. Unknown launchers, custom shell functions and shell keywords stop the executable walk; timeout requires a literal duration.The landing inventory is 94 (main's 93 plus the checker); the checker is the only added repo inventory grant. The exact dated G5 comparison record has zero active selectors under the full grammar and retains its unchanged bytes and ordinary active-copy/size guards.
f1f6c0c20e214f1cee9aba3c5f0c1430d66e75ed.ced5ccf43d7206451d3eb5c8693a9bc0389f8566, the ONE final docs/test-scope correction of7e4a0c94fd42879d3b41adae1042a048fff5c6a4under the CC's option (a). Runtime checker/collector and the grant policy remain byte-identical; only the operating note, tests and two evidence bindings change. The conflict-only rebase remains for the CC's landing cue.lane:shared. Trading acknowledgement 6094414078, reaffirmed in 6094723019, covers the byte-identical run_worker.py and research-runtime README hunks.SOTA sources
-m/--model; installed Codex 0.162.1 help agrees. Installed Claude Code 2.1.296 help declares--modelwithout-m.+n:c:p:P:u:tVh, and its command-versus-targeting/execvp branches; flock.c defines lock/option operands. Installed wrapper/Python controls preserve literal argv without launching a model client.c1382380de69521303b416720a52f42d51af6248, OmniRoute v3.8.51's models route and Anthropic models API2023-06-01. Observation identities are preserved; comparison does not certify live availability or delivered effort. RFC 6901 supplies JSON locators.Evidence-class table
The historical repair measurements below refer to reviewed
94fbc29aand its pure landing rebase. The final B1 repair has separate validation below.synthetic/local_integrationsyntheticsynthetic/local_integrationsyntheticsyntheticsyntheticsynthetic/local_integrationlocal_integrationlocal_integrationlocal_integrationlocal_integrationThe repository-only census excludes host/settings files and protected peer lanes. The previous safe host corpus restored all 26 readiness selectors and hcom; its remaining custom
run native codexfunction row is explicitly outside the static grammar. Claude's unadvertised-mremains excluded. No claim of complete custom-wrapper or settings-inclusive host coverage.Local commands run at reviewed 94fbc29
Decision record
Active-model currency decision and operational scope. The latest correction admits potentially escaped selections before upstream decoding, preserving full-span separation, literal argv and punctuation protections. Earlier Markdown controls remain. Native registration regenerates exactly three source bindings; all other evidence metadata is exact. An attempted inherited receipt display canonicalization failed native receipt fidelity and was discarded before commit. Publication scans check the complete PR body, new committed text and new messages with count-only output; no absolute home path is added or quoted.
Host evidence
No host receipt or platform-status change. Previously measured wrapper/Python and Bash echo controls establish local command behavior. Model/provider execution and host application are unrun. CC owns landing and post-landing units.
Checklist
Prior landing rebase a2a873b validation
The exact prepared head passes 229 tests without skips (active-model, currency-due, research-workflow, grants and evidence modules), FULL validation of 70 components / 11,482 hashed files / four profiles / 239 receipts, evidence normalization and diff checks. PR-range-only gitleaks scans the ten rebased commits with 100% redaction and no leaks. The CC's recompute tool derives the shared inventory count; the only grant addition remains the checker.
final_stackis absent at this pinned main and was not run. The exact-lease landing push replaces reviewed94fbc29awith this pure rebase; no functional repair was added.Final adjudicated B1 repair
Wider admission previously passed model-option/default-expansion lines into a parser that raised on incomplete quoting. Three permanent real-check cases fail at exact a2a873b with
unknown/ 2 and pass after repair withcurrent/ 0 and no errors: the native shell launch shape, the same shape in a sh fence, and model-option plus alias/apostrophe prose. A preservation method verifies that raw and potentially escaped model-ID parse failures in.shcommands still return unknown. The guard reuses existing model_value_candidate; CPython v3.13.16 at cbc944f4bc59639a444dd971c737788ba2283a91:Lib/shlex.py:L185-L191 supplies the incomplete-quotation failure being classified.The touched currency module passes 97 tests without skips; FULL validation passes 70 components / 11,482 hashed files / four profiles / 239 receipts. Evidence/diff checks pass and the actual eleven-commit PR range scans with 100% redaction and no leaks. The native
check --root <owned repository> --host --jsonshape changes unknown/2 (one error) to current/0 (no errors, no stale selectors) on the same 3,690-file admitted inventory. Both host settings files and peer repository contents are excluded before file reads under the standing privacy/protected-scope rules; a settings-inclusive host run is unrun. Only three existing evidence bindings update; inventory remains 91.Final documentation-scope correction
The CC's option (a) treats the changed operating-note guarantee as a blocker under rule 1. The docs now state the actual field/non-prose unknown-coverage scope and the prose-command raw-sentence fallback, including docs sh fences, with pinned source links to 7e4a0c9: scripts/active_model_currency.py:L614-L617 and L633-L639. The former broad preservation-method name is scoped to non-prose commands.
Two documented-behavior methods fail at exact 7e4a0c9 on the contradictory/missing documentation contract (two failures, zero errors). Their runtime controls already pass at that head: Astra's docs sh fence containing
codex exec -m gpt-6.1-sol 'unclosedreturns current/0/no errors; field and non-prose malformed-ID controls return unknown/2. After the doc correction both methods pass. This is documentation-alignment evidence, not a runtime before/after behavior change.The final currency module passes 99 tests without skips, FULL validation passes 70 components / 11,482 hashed files / four profiles / 239 receipts, and evidence/diff checks pass. Only the two existing doc/test bindings update. PR-range-only gitleaks scans twelve commits with 100% redaction and no leaks. The prior body correction and all ten individually replied/resolved Codex Follow-ups remain delivered; no additional implementation changes or landing rebase are included.
Follow-ups
F1 / P3: validate executable position for long model options, retaining echo/printf ordinary-operand negatives.
F2 / P3: refine option-terminator handling for
codex exec -- -m.F3 / P3: decide and test extraction of
${VAR:-<model>}defaults; they are not extracted currently.F5: hcom mirror coverage is owned by the CC's MIRRORS update.
Codex follow-up 4240876145: Align the stale-model collector exit status with the systemd timer-unit success contract. (
scripts/currency_due.py:913).Codex follow-up 4240876148: Review effective project-local Claude settings excluded from Git inventory. (
scripts/active_model_currency.py:646).Codex follow-up 4240876151: Model Claude fallback-model flags and their ordered values. (
scripts/active_model_currency.py:44).Codex follow-up 4240876152: Scope Codex-native routing aliases to the client/configuration where they are valid. (
scripts/active_model_currency.py:721).Codex follow-up 4240876155: Preserve and validate observed provider/routing prefixes during model-ID extraction. (
scripts/active_model_currency.py:29).Codex follow-up 4240876157: Distinguish Claude availableModels prefix semantics from exact runtime-model selections. (
scripts/active_model_currency.py:722).Codex follow-up 4242188824: Include applicable Codex profile configuration files in an agreed host-coverage policy. (
scripts/active_model_currency.py:653).Codex follow-up 4242188828: Review recursive discovery of nested agent files. (
scripts/active_model_currency.py:658).Codex follow-up 4242188832: Handle YAML block-list/plural selector fields under the recorded grammar policy. (
scripts/active_model_currency.py:601).Codex follow-up 4242188833: Handle PowerShell environment-assignment prefixes for model selectors. (
scripts/active_model_currency.py:604).F5b / P3: model hcom's
--systemvalue flag, verified at the new hcom v0.7.28 mirror (src/commands/launch.rs:639-642).Wider: decide the behavior of escaped model IDs in prose when command parsing falls back to raw sentence matching.
Wider: measure the 30-second ACTIVE_MODELS budget under load before changing its policy; the confirmation micro's timeout explanation remains a hypothesis.
Wider: retain the hostname-hygiene item for its dedicated follow-up.
CC terminal-rule follow-up 4242590024: Measure the complete host-scan budget and align the collector timeout with inventory subprocess limits and subsequent parsing (scripts/currency_due.py:187). The CC classified this as not P1; deferred after the correction allowance was used.
CC terminal-rule follow-up 4242590030: Require Anthropic catalog pagination to declare has_more as explicit boolean false before publishing a complete catalog (scripts/active_model_currency.py:80). The CC classified this as not P1; deferred after the correction allowance was used.
CC terminal-rule follow-up 4242590036: Handle language block comments before extracting selectors, including multiline JavaScript/TypeScript comments (scripts/active_model_currency.py:594). The CC classified this as not P1; deferred after the correction allowance was used.
CC terminal-rule follow-up 4242590042: Classify Anthropic model rows using the API's declared line field, including future or differently spelled model IDs (scripts/active_model_currency.py:92). The CC classified this as not P1; deferred after the correction allowance was used.
CC terminal-rule follow-up 4242590046: Carry a structured command's sibling executable context into args traversal so short model flags are recognized (scripts/active_model_currency.py:503). The CC classified this as not P1; deferred after the correction allowance was used.