feat(reborn): route caller-requested model on OpenAI-compatible API (Phase 2) - #5985
Conversation
…ponses API The OpenAI-compatible Responses (and Chat) surface hard-coded `usage: None` and Reborn captured no per-run token totals anywhere — token counts existed only per-LLM-call, were spent transiently on budget/stop heuristics, and were never aggregated, persisted, or projected. Callers had no view of tokens or cost. This lands the shared per-run usage backbone and surfaces it as `usage` + an IronClaw `cost` extension. - ironclaw_turns: LoopModelUsage gains cache token fields (previously dropped at the gateway) + add_assign/total_tokens. New additive `model_usage: Option<LoopModelUsage>` on TurnRunState/TurnRunRecord/RunRecord rides the JSON-blob snapshot like resolved_model_route — no table, column, migration, or per-backend change. The dead usage_summary_ref/LoopUsageSummaryRef (never set/read; pointed at a store that never existed) is replaced by an inline model_usage value on LoopCompleted/LoopFailed, extracted in LoopExitApplier::apply and accumulated onto the run record at the terminal transition (block/resume legs sum). - ironclaw_runner: reply paths preserve provider cache token counts. - ironclaw_agent_loop: LoopExecutionState accumulates cumulative usage at both assistant-reply finalize paths; completed/failed exits carry it. - ironclaw_reborn_openai_compat: OpenAiResponseUsage/OpenAiUsage gain input_tokens_details.cached_tokens (OpenAI-standard) + a namespaced cost object (input/cached-input/output/total USD, decimal strings). No new rust_decimal/ironclaw_llm dependency on the route crate. - ironclaw_reborn_composition: the projection reader reads persisted model_usage via get_run_state, prices it through ironclaw_llm::costs (cache-read at the provider discount, unknown models fall back to default rate not zero), and fills usage. Cost gated behind root-llm-provider; tokens always reported. Follow-up (Phase 2): route the requested model through the turn so model selection actually takes effect (and cost prices the model that ran). Tests: token-breakout, cost pricing incl. cache discount, and unknown-model default-rate fallback; existing responses/chat/DTO/streaming contract suites updated for the new optional usage fields. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…Phase 2) Stacked on the usage/cost PR. Makes the OpenAI-compatible Responses/Chat `model` field actually route the turn's LLM call instead of being validated-and-echoed but ignored. Semantics: route if the provider can serve it, else silently fall back to the deployment's active model. The requested model is threaded from the request to the model call as an advisory route, honored at the provider boundary: - product_adapters: UserMessagePayload gains optional requested_model (wire-defaulted; with_requested_model filters empty). - openai_compat: chat + responses workflows set requested_model from request.model on the submitted payload (previously dropped). - product_workflow: AcceptedProductInboundTurn::submit carries it onto the new SubmitTurnRequest.requested_model. - ironclaw_turns: submit_turn records it as an advisory LoopModelRouteSnapshot (LoopModelRouteSnapshot::advisory — only model_id meaningful, is_advisory() distinguishes it, None for empty/invalid). Reuses the existing resolved_model_route field that already flows run -> loop context -> gateway, so no new run-state field. - ironclaw_runner: LlmProviderModelGateway::request_model_override prefers the request's route model_id over the profile default, falling back to the active model when absent (providers that honor per-request overrides serve it, others fall back). attach_model_route_snapshot passes an advisory snapshot through unvalidated when no route resolver is wired (default runtime); routed fail-closed hosts are unchanged. Child/subagent runs and idempotent replays carry no requested model. Known limitation: per-request routing only takes effect for providers that honor CompletionRequest.model (NEAR AI); RigAdapter-backed providers ignore it and fall back. Strict operator-configured-only validation would need the model catalog wired into the default runtime (follow-up). Tests: gateway honors requested route over profile default + falls back when absent (recording-provider seam); submit records advisory route from requested_model (and none when absent); advisory() construction/validation; UserMessagePayload requested_model serde round-trip + empty filtering. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (3)
📝 WalkthroughSummary by CodeRabbit
WalkthroughAdds an optional requested-model hint to inbound payloads, propagates it into turn state, represents it as an advisory route, and applies it during gateway model selection. Existing submission call sites explicitly set the field to ChangesRequested model routing
Estimated code review effort: 3 (Moderate) | ~25 minutes Sequence Diagram(s)sequenceDiagram
participant OpenAIClient
participant UserMessagePayload
participant TurnCoordinator
participant RunState
participant LlmProviderModelGateway
OpenAIClient->>UserMessagePayload: provide model hint
UserMessagePayload->>TurnCoordinator: submit requested_model
TurnCoordinator->>RunState: record advisory model route
RunState->>LlmProviderModelGateway: provide resolved route
LlmProviderModelGateway->>LlmProviderModelGateway: select requested model or fallback
Possibly related issues
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 3 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (3 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
🚅 Deployed to the ironclaw-pr-5985 environment in ironclaw-ci-preview
|
Addresses PR #5976 review findings (Copilot, ironloopai, CodeRabbit): - Stop double-counting cache tokens in the OpenAI-compatible usage/cost projection. `cache_read_input_tokens` is a subset of `input_tokens` (per the type contract + how nearai/OpenAI populate it), so it is no longer added on top of the wire `input_tokens`, and the full-rate billable input is now `input_tokens - cache_read` (cache_read stays discounted). `cache_creation` remains additive. (Copilot, ironloopai) - Accumulate model-response usage for EVERY model turn in the canonical executor, before branching on the output, so tool-using (`CapabilityCalls`) turns no longer drop their usage/cost. Removed the now-duplicate accumulation in `AssistantReplyStage`. (ironloopai) - Clear `cumulative_model_usage` when `rebase_for_run` rebases onto a different run, so a retry does not re-report the failed run's tokens. Preserved for same-run gate resume. (CodeRabbit) - Replace (not accumulate) run usage in the validated loop-exit transition: the loop reports its per-run cumulative at every exit, so a block→resume→complete sequence no longer double-counts pre-block legs. (CodeRabbit) - Annotate the best-effort `.ok()?` DB/IO reads in `read_run_usage` with `// silent-ok:` comments. (CodeRabbit) Tests: updated the usage/cost projection tests to the corrected cache semantics and added non-zero `cache_creation` + Claude 10x cache-read discount coverage; caller-level regression tests for the rebase reset, the block→resume→complete cumulative usage, and capability-turn usage accumulation. Also fixed two root integration-test `OpenAiResponseUsage` literals missing the new `cost`/`input_tokens_details` fields (the CI compile break). Deferred: pricing by the resolved model route (ironloopai) lands in the stacked model-routing PR #5985, where `resolved_model_route` is actually populated and testable. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ponses API (#5976) * feat(reborn): per-run token usage + USD cost on OpenAI-compatible Responses API The OpenAI-compatible Responses (and Chat) surface hard-coded `usage: None` and Reborn captured no per-run token totals anywhere — token counts existed only per-LLM-call, were spent transiently on budget/stop heuristics, and were never aggregated, persisted, or projected. Callers had no view of tokens or cost. This lands the shared per-run usage backbone and surfaces it as `usage` + an IronClaw `cost` extension. - ironclaw_turns: LoopModelUsage gains cache token fields (previously dropped at the gateway) + add_assign/total_tokens. New additive `model_usage: Option<LoopModelUsage>` on TurnRunState/TurnRunRecord/RunRecord rides the JSON-blob snapshot like resolved_model_route — no table, column, migration, or per-backend change. The dead usage_summary_ref/LoopUsageSummaryRef (never set/read; pointed at a store that never existed) is replaced by an inline model_usage value on LoopCompleted/LoopFailed, extracted in LoopExitApplier::apply and accumulated onto the run record at the terminal transition (block/resume legs sum). - ironclaw_runner: reply paths preserve provider cache token counts. - ironclaw_agent_loop: LoopExecutionState accumulates cumulative usage at both assistant-reply finalize paths; completed/failed exits carry it. - ironclaw_reborn_openai_compat: OpenAiResponseUsage/OpenAiUsage gain input_tokens_details.cached_tokens (OpenAI-standard) + a namespaced cost object (input/cached-input/output/total USD, decimal strings). No new rust_decimal/ironclaw_llm dependency on the route crate. - ironclaw_reborn_composition: the projection reader reads persisted model_usage via get_run_state, prices it through ironclaw_llm::costs (cache-read at the provider discount, unknown models fall back to default rate not zero), and fills usage. Cost gated behind root-llm-provider; tokens always reported. Follow-up (Phase 2): route the requested model through the turn so model selection actually takes effect (and cost prices the model that ran). Tests: token-breakout, cost pricing incl. cache discount, and unknown-model default-rate fallback; existing responses/chat/DTO/streaming contract suites updated for the new optional usage fields. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reborn): correct per-run usage/cost accounting from PR #5976 review Addresses PR #5976 review findings (Copilot, ironloopai, CodeRabbit): - Stop double-counting cache tokens in the OpenAI-compatible usage/cost projection. `cache_read_input_tokens` is a subset of `input_tokens` (per the type contract + how nearai/OpenAI populate it), so it is no longer added on top of the wire `input_tokens`, and the full-rate billable input is now `input_tokens - cache_read` (cache_read stays discounted). `cache_creation` remains additive. (Copilot, ironloopai) - Accumulate model-response usage for EVERY model turn in the canonical executor, before branching on the output, so tool-using (`CapabilityCalls`) turns no longer drop their usage/cost. Removed the now-duplicate accumulation in `AssistantReplyStage`. (ironloopai) - Clear `cumulative_model_usage` when `rebase_for_run` rebases onto a different run, so a retry does not re-report the failed run's tokens. Preserved for same-run gate resume. (CodeRabbit) - Replace (not accumulate) run usage in the validated loop-exit transition: the loop reports its per-run cumulative at every exit, so a block→resume→complete sequence no longer double-counts pre-block legs. (CodeRabbit) - Annotate the best-effort `.ok()?` DB/IO reads in `read_run_usage` with `// silent-ok:` comments. (CodeRabbit) Tests: updated the usage/cost projection tests to the corrected cache semantics and added non-zero `cache_creation` + Claude 10x cache-read discount coverage; caller-level regression tests for the rebase reset, the block→resume→complete cumulative usage, and capability-turn usage accumulation. Also fixed two root integration-test `OpenAiResponseUsage` literals missing the new `cost`/`input_tokens_details` fields (the CI compile break). Deferred: pricing by the resolved model route (ironloopai) lands in the stacked model-routing PR #5985, where `resolved_model_route` is actually populated and testable. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…cate The is_advisory() method had zero production consumers: the model gateway reads snapshot.model_id uniformly regardless of advisory-ness, and loop_driver_host gates on route-resolver presence rather than the advisory flag. The predicate (and the operator_resolved_route_is_not_advisory test that existed only to exercise it) dressed the three "requested" sentinel placeholder components up as a first-class concept nothing acts on. Keep advisory() — the actual reuse seam that stores a caller-requested model hint in resolved_model_route — and have the remaining tests assert on model_id, the only component that carries meaning. Addresses thermo-nuclear code-quality review of PR #5985. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…-model-select # Conflicts: # crates/ironclaw_agent_loop/src/executor/assistant_reply.rs # crates/ironclaw_reborn_composition/src/llm_admin/openai_compat_serve.rs # crates/ironclaw_reborn_composition/src/llm_admin/openai_compat_serve/tests.rs # crates/ironclaw_turns/src/memory/mod.rs
…host Merging main surfaced a semantic conflict between two independently-correct changes: - #5985 made a resolver-less host (the default product runtime) pass a persisted model-route snapshot through unvalidated, so an OpenAI-compatible caller's requested model is honored by the non-routed gateway. - main independently hardened the same host to FAIL CLOSED when a model-route snapshot is present but no resolver is wired (an operator route we cannot validate is a misconfiguration), with tests locking that behavior. The auto-merge collapsed these into an unconditional pass-through, which broke `text_only_host_factory_rejects_persisted_model_route_snapshot_without_resolver`. Both intentions are correct and coexist by discriminating the snapshot kind — exactly what LoopModelRouteSnapshot::is_advisory() expresses: - advisory snapshot (caller-requested hint) + no resolver -> pass through - operator route + no resolver -> fail closed ("resolver required") This reverses the earlier "drop dead is_advisory()" commit on this branch: the predicate looked unused in #5985 in isolation, but main's stricter guard makes it load-bearing. is_advisory() and its unit tests are restored, the guard now branches on it, and a caller-path regression test (`text_only_host_factory_passes_advisory_model_route_snapshot_without_resolver`) pins the pass-through side that main's reject test does not cover. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
crates/ironclaw_product_adapters/src/inbound.rs (1)
153-162: 🔒 Security & Privacy | 🔴 Critical | ⚡ Quick winIngress length bounds are bypassed during deserialization.
UserMessagePayload::newvalidates the payload whilerequested_modelis stillNone. Thedeserializeimplementation then attaches the untrustedwire.requested_modelusingwith_requested_model, but never callsvalidate()again. As a result, theREQUESTED_MODEL_MAX_BYTESlimit is completely bypassed for network ingress payloads.As per repository invariants, you must validate and bound original ingress payloads before storage or dispatch.
Proposed fix
fn deserialize<D>(deserializer: D) -> Result<Self, D::Error> where D: Deserializer<'de>, { let wire = UserMessagePayloadWire::deserialize(deserializer)?; - Self::new(wire.text, wire.attachments, wire.trigger) - .map(|payload| payload.with_requested_model(wire.requested_model)) - .map_err(serde::de::Error::custom) + let payload = Self::new(wire.text, wire.attachments, wire.trigger) + .map(|payload| payload.with_requested_model(wire.requested_model)) + .map_err(serde::de::Error::custom)?; + payload.validate().map_err(serde::de::Error::custom)?; + Ok(payload) }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/ironclaw_product_adapters/src/inbound.rs` around lines 153 - 162, Update UserMessagePayload::deserialize so the deserialized wire fields, including wire.requested_model, are validated against ingress bounds before the payload is stored or dispatched. After applying with_requested_model, invoke the payload validation path again and convert any validation failure through serde::de::Error::custom, preserving UserMessagePayload::new validation for the other fields.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@crates/ironclaw_reborn_openai_compat/src/chat_workflow.rs`:
- Around line 798-799: The payload builders in
crates/ironclaw_reborn_openai_compat/src/chat_workflow.rs lines 798-799 and
crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs lines 1389-1390
must revalidate after applying with_requested_model: bind the completed
UserMessagePayload, call payload.validate()?, then return it. Preserve the
existing builder inputs and error propagation.
---
Outside diff comments:
In `@crates/ironclaw_product_adapters/src/inbound.rs`:
- Around line 153-162: Update UserMessagePayload::deserialize so the
deserialized wire fields, including wire.requested_model, are validated against
ingress bounds before the payload is stored or dispatched. After applying
with_requested_model, invoke the payload validation path again and convert any
validation failure through serde::de::Error::custom, preserving
UserMessagePayload::new validation for the other fields.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: c5c3a9e7-cc22-4ea9-a9aa-852a5f4be769
📒 Files selected for processing (34)
crates/ironclaw_conversations/src/inbound.rscrates/ironclaw_host_runtime/tests/support/host_runtime_harness.rscrates/ironclaw_loop_host/tests/turn_event_publisher_contract.rscrates/ironclaw_product_adapters/src/inbound.rscrates/ironclaw_product_workflow/src/auth_continuation.rscrates/ironclaw_product_workflow/src/inbound_turn.rscrates/ironclaw_product_workflow/src/reborn_services.rscrates/ironclaw_product_workflow/tests/product_workflow_contract.rscrates/ironclaw_reborn_composition/src/factory.rscrates/ironclaw_reborn_composition/src/factory/auth_tests.rscrates/ironclaw_reborn_composition/src/runtime.rscrates/ironclaw_reborn_composition/src/runtime/tests/auth_interaction.rscrates/ironclaw_reborn_openai_compat/src/chat_workflow.rscrates/ironclaw_reborn_openai_compat/src/responses_workflow.rscrates/ironclaw_runner/src/loop_driver_host.rscrates/ironclaw_runner/src/model_gateway.rscrates/ironclaw_runner/src/subagent/await_edge/boot_recovery.rscrates/ironclaw_runner/src/subagent/await_edge/resolver.rscrates/ironclaw_runner/tests/concurrent_workers.rscrates/ironclaw_runner/tests/llm_gateway.rscrates/ironclaw_runner/tests/loop_driver_host.rscrates/ironclaw_runner/tests/turn_scheduler_contract.rscrates/ironclaw_turns/src/memory/mod.rscrates/ironclaw_turns/src/request.rscrates/ironclaw_turns/src/run_profile/host.rscrates/ironclaw_turns/tests/active_run_ref_state_contract.rscrates/ironclaw_turns/tests/agent_loop_host_contract.rscrates/ironclaw_turns/tests/filesystem_turn_state_contract.rscrates/ironclaw_turns/tests/per_inbound_type_concurrency_cap.rscrates/ironclaw_turns/tests/per_user_concurrency_cap.rscrates/ironclaw_turns/tests/retry_failed_turn_store_contract.rscrates/ironclaw_turns/tests/turn_coordinator_contract.rstests/integration/subagent_await_edge.rstools/ironclaw_stress/src/user_turn.rs
The Responses/Chat OpenAI-compat surfaces forwarded request.model verbatim as a per-run requested-model hint. But "default" is the server's alias for "use the active model" (the models listing advertises it), not a concrete model id. Forwarding it created an advisory route with model_id "default", which request_model_override rejects as non-concrete (PolicyDenied) — failing every run whose client sent model="default", including 6 legacy Responses API E2E scenarios (status "failed" instead of "completed"). Map the wire model through model_validation::requested_model_hint before forwarding: the "default" sentinel (and, defensively, empty) yields None so the run falls back to normal resolution — the active model on the non-routed gateway, the resolver's default route on routed hosts — while a concrete model name is still forwarded. Keeps the model_gateway "default"-is-not-concrete guard intact for genuine route/active-model misconfiguration. Regression: model_validation unit tests for the sentinel/whitespace/concrete cases; the legacy Responses API E2E suite exercises the full path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@crates/ironclaw_reborn_openai_compat/src/chat_workflow.rs`:
- Around line 798-801: In
crates/ironclaw_reborn_openai_compat/src/chat_workflow.rs lines 798-801, bind
the UserMessagePayload builder result, call payload.validate()? after
with_requested_model, and return the validated payload. Apply the same change in
crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs lines 1389-1392,
ensuring both ingress paths validate the model hint before dispatch.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: b2702106-aa47-4d0e-92e4-c5e625380a37
📒 Files selected for processing (3)
crates/ironclaw_reborn_openai_compat/src/chat_workflow.rscrates/ironclaw_reborn_openai_compat/src/model_validation.rscrates/ironclaw_reborn_openai_compat/src/responses_workflow.rs
Coverage ratchetReborn integration-tier coverageLine coverage (Reborn crates): 85.54% — 302060 / 353108 lines Per-crate breakdown (63 crates, lowest-covered first)
This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors. Exemptions (3 entry/entries excluded from the accounting above)
|
…uction CodeRabbit review flagged that UserMessagePayload::with_requested_model attaches the model hint AFTER new() ran validate() (with requested_model still None), so the 256-byte REQUESTED_MODEL_MAX_BYTES bound was bypassed. The crate contract is explicit: "Validated DTOs must validate both constructors and serde deserialization." - The custom Deserialize (untrusted wire path) now re-validates the assembled payload, so a wire-supplied requested_model is bounded like every other ingress field. This is the real bypass — an unbounded model string could otherwise deserialize and flow to persistence/dispatch. - The Responses and Chat OpenAI-compat builder call sites now validate the assembled payload before submitting (defense in depth; request.model is already capped at parse by validate_model_name, but the type invariant must hold locally). Regression: user_message_payload_bounds_requested_model_on_every_path asserts an over-limit hint is rejected via both the builder+validate path and deserialization, and that a hint at the cap is accepted. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
@coderabbitai Addressed the ingress-validation finding in f5aff5d. Critical (outside-diff) — Major (inline) — builder call sites: the Chat and Responses OpenAI-compat workflows now Regression: |
|
Tip For best results, initiate chat on the files or code changes. 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
gh pr diff 5985 --repo nearai/ironclaw | sed -n '1,400p' | grep -n "requested_model\|with_requested_model\|fn deserialize\|validate(" -A3 -B3 crates/ironclaw_product_adapters/src/inbound.rs || true
echo "---direct file fetch---"
gh api repos/nearai/ironclaw/commits/f5aff5dc7 --jq '.files[].filename'Length of output: 7654 Confirmed against f5aff5d. The Deserialize impl now calls Nothing further to add — this matches the crate's validated-DTO contract. (=^・ω・^=) If you found this review helpful, would you consider giving us a shout-out on X? Thank you for using CodeRabbit! |
What
Makes the OpenAI-compatible Responses/Chat API's
modelfield actually route the turn's LLM call, instead of being validated-and-echoed but ignored. Stacked on top of #5976 (per-run usage + cost); merge that first.Semantics (as agreed): route if the provider can serve it, else silently fall back to the deployment's active model — no hard 400.
Why
Before this, a client's
modelwas dropped at submit; every turn ran on the deployment's globally-active provider model. The routing seam (resolved_model_route→LlmProviderModelGateway) already threaded a per-run model to the gateway, but the gateway ignored it and product turns always resolved toprovider.active_model_name().How
requested_modelis threaded from the request to the model call as an advisory route, honored by the gateway at the provider boundary:ironclaw_product_adapters—UserMessagePayloadgains an optionalrequested_model(wire-defaulted;with_requested_modelfilters empty). Surfaces that don't pick a model leave itNone.ironclaw_reborn_openai_compat— the chat + responses workflows setrequested_modelfromrequest.modelon the submitted payload (previously dropped).ironclaw_product_workflow—AcceptedProductInboundTurn::submitcarries it ontoSubmitTurnRequest.requested_model(new field).ironclaw_turns— the store'ssubmit_turnrecords it as an advisoryLoopModelRouteSnapshoton the new run (LoopModelRouteSnapshot::advisory(model): onlymodel_idis meaningful, provider/config/auth are"requested"placeholders;is_advisory()distinguishes it; returnsNonefor empty/invalid model strings → fallback). This reuses the existingresolved_model_routefield that already flows run → loop context → gateway, so no new run-state field and no fixture churn beyond setting it.ironclaw_runner—LlmProviderModelGateway::request_model_overridenow prefersrequest.resolved_model_route.model_idover the profile default, falling back to the active model when absent. Providers that honor per-request overrides (e.g. NEAR AI) serve the requested model; providers that bake the model at construction ignore it and fall back — the "route-if-serveable-else-fallback" decision happens at the provider boundary.attach_model_route_snapshotpasses an advisory snapshot through unvalidated when no route resolver is wired (the default product runtime that serves the OpenAI-compat surface). Routed hosts (resolver present, fail-closedRoutedLlmProviderModelGateway) are unchanged and still validate — and never receive an advisory snapshot, since only the OpenAI-compat surface (on the default runtime) sets one.Child/subagent runs and idempotent replays carry no requested model (fall back to the active model); the model is not persisted in the message store, so replays intentionally don't recover it.
Tests
ironclaw_runner(llm_gateway, recording-provider integration seam): the per-run requested model overrides the profile default on the capturedCompletionRequest.model; absent a route it falls back to the profile default.ironclaw_turns:submit_turnrecords an advisory route fromrequested_model(and none when absent);advisory()carries the model / marks itself advisory / trims and rejects empty or invalid model strings; an operator-resolved route is not advisory.ironclaw_product_adapters:UserMessagePayloadround-tripsrequested_modelover the wire, omits it whenNone, and filters empty.Known limitation
Per-request model routing only takes effect for providers that honor
CompletionRequest.model(NEAR AI today);RigAdapter-backed providers (OpenAI/Anthropic/etc.) bake the model at construction and ignore the override, falling back to their configured model. The response still echoes the requested model. Enforcing "operator-configured only" with a strict catalog check (vs. the provider-boundary fallback here) would require wiring the model catalog/resolver into the default runtime — a follow-up.🤖 Generated with Claude Code