Skip to content

feat(loop): derive the prompt context budget from the model's advertised window - #8053

Open
henrypark133 wants to merge 29 commits into
mainfrom
context-length
Open

henrypark133 wants to merge 29 commits into
mainfrom
context-length

Conversation

@henrypark133

Copy link
Copy Markdown
Collaborator

Summary

  • The agent loop's prompt context budget was a compiled-in 128k/20k PromptContextTokenBudget regardless of model. Runs now derive it from the provider-advertised context window: PromptContextTokenBudget::from_advertised_window takes 90% of the window as the limit and keeps the flat 20k response reserve, clamped to a quarter of the limit so small-window models keep a usable transcript. None / 0 reproduce today's default exactly.
  • The turn-runner host build resolves the window once per run through a new defaulted HostManagedModelGateway::advertised_context_window_tokens and carries the result on LoopRunContext.resolved_context_budget (not persisted). The production LlmProviderModelGateway reads ModelMetadata.context_length only when the metadata's id matches the model the request will actually be served by, so a route override never borrows another model's window.
  • All five default sites consume that one value: both compaction strategies (ironclaw_agent_loop, which stays contracts-only), the context port, the model port via ThreadResolvingLoopModelGateway, and structured finalization. The four family replay fingerprints drop the hard-coded context_limit=128000,reserve=20000 literals (digests recomputed) so the replay identity stays stable across models.
  • Found while proving the critical path: TokenRefreshingProvider::model_metadata() ran a token-refresh HTTP call (under the lock shared with in-flight complete()) before delegating to a static answer. Removed, pinned by test, and made a CONTRACT rule: model_metadata() is a static, I/O-free description.
  • ironclaw_loop_contracts grew by ~41 genuine lines against a ceiling main already sat 3 lines under; tests moved to the crate's out-of-line tests.rs layout and the ceiling re-pinned at the measured 13,773 with the reason in the ladder comment.

Design note: docs/internal/reborn/design/model-derived-context-budget.md. Plan: docs/internal/superpowers/plans/2026-09-01-model-derived-context-budget.md.

Change Type

  • Bug fix (TokenRefreshingProvider::model_metadata token-refresh I/O)
  • New feature
  • Refactor (context_budget tests out of line; compaction budget passed as an argument)
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

None — no tracking issue exists for this; opened from a direct request. Flagging for the maintainer whether one should be filed and linked.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings (zero warnings, run twice: pre- and post-fixup tree)
  • cargo build (as part of the test runs)
  • Relevant tests pass: cargo test -p ironclaw_llm -p ironclaw_loop_contracts -p ironclaw_agent_loop -p ironclaw_loop_host -p ironclaw_turn_runner -p ironclaw_architecture_tests; cargo test -p ironclaw_integration_tests --test reborn_integration_context_budget --test reborn_integration_model_recovery --test reborn_integration_greeting; full cargo test -p ironclaw_integration_tests and -p ironclaw_composition with RUST_MIN_STACK=67108864.
  • cargo test -p <owning-crate> --features integration — Not applicable: no database-backed behavior changed.
  • Manual testing — Not applicable: behavior is pinned at crate and integration tiers below.
  • pr-shepherd — not yet run; will run on CI feedback.

Pre-existing failures observed and confirmed identical on main at 99457e152, none in files this branch touches: ironclaw_turn_runner trace_capture::tests::capture_skips_when_policy_missing_or_disabled; five trace-commons tests that read a stray real ~/.ironclaw/trace_contributions/policy.json on the dev machine; ironclaw_composition capability_port_omits_host_disclosure_without_confirmed_host_mount (FilesystemDenied vs InputEncode); reborn_user_submit_completes_while_another_turn_state_write_is_blocked flakes 2–4 of 8 isolated runs on both branches on its 5 s window (its harness never reaches this branch's code path).

Test Strategy

User behavior: a run served by a model that advertises a context window gets a prompt budget sized to that model — a 40k-window model is sent a smaller transcript and compacts earlier than the 128k default; a model that advertises nothing behaves exactly as before.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: ironclaw_loop_contracts context_budget/tests.rs (derivation: none/zero → default, large window keeps flat reserve, small window clamps reserve, today's constant is reduced by the margin) and run_context/tests.rs (field default, builder, wire round trip). ironclaw_agent_loop compaction.rs / active_task_compaction.rs (run-context budget overrides the strategy default; absent budget falls back) plus the four digest self-consistency tests. ironclaw_loop_host model_gateway.rs (gateway reports the provider window / none / none on route override) and tests/thread_loop_host_contract.rs::thread_resolving_gateway_applies_its_prompt_context_budget_to_message_selection. ironclaw_turn_runner loop_driver_host/context_budget_tests.rs (resolved budget reaches the run context; nothing advertised leaves it None; derived_budget_sizes_the_request_the_host_sends_to_the_gateway). ironclaw_llm token_refreshing.rs::model_metadata_does_not_refresh_the_token.
  • Reborn integration: tests/integration/context_budget.rs (small_advertised_window_shrinks_what_the_model_receives, unadvertised_window_keeps_the_compiled_in_ceiling) through the real LlmProviderModelGateway and ironclaw_llm decorator chain; harness gains advertised_context_window(tokens) and captured_request_message_count(index). Coverage row added to tests/AGENTS.md.
  • Recorded fixture: Not applicable: no provider payload shape changed.
  • Browser E2E: Not applicable: no UI.
  • Backend or runtime: Not applicable: no database or runtime-lane change.
  • Live canary: Not applicable: scripted provider exercises the same gateway path.

What the tests prove: every link of the wiring chain is mutation-verified — each new test was shown failing with its link removed (gateway override returning None → integration 13 vs 13; wrapper builder call removed → all 5 seeded messages forwarded; driver-host derivation replaced by the config default → all 5 forwarded; token refresh restored → 1 connection attempt on the counting endpoint). The integration test deliberately does not separate compaction from request sizing (they share visible_transcript_tokens, so compaction always fires first); that separation is what the two crate-tier tests above pin.

Commands run: see Validation. All with RUST_MIN_STACK=67108864 (what CI and scripts/ci/quality_gate.sh set; without it ~46 full-runtime test binaries overflow libtest's default stack on this machine).

Security Impact

None. The advertised window is advisory: a provider that cannot report one, or reports one for a different model than the route serves, leaves the run on the compiled-in budget (// silent-ok marked). No new network calls — one existing unnecessary network call (token refresh inside model_metadata) is removed.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: none added. PromptContextTokenBudget and the new Option field are plain DTOs constructed by the turn runner.
  • Untrusted content enters prompts only through an envelope/escaping primitive — N/A, no prompt content path changed.
  • Hashes declare purpose — the four ComponentDigest BLAKE3 replay fingerprints changed content, not purpose; self-consistency tests pin them.
  • New/changed variants: none. New defaulted trait method on HostManagedModelGateway; all implementors inherit None. Command: rg -n "impl.*HostManagedModelGateway for" crates tests — only LlmProviderModelGateway overrides.
  • serde(default) field: resolved_context_budget: Option<_> with #[serde(default, skip_serializing_if = "Option::is_none")] — missing means None means today's default (fails to the conservative pre-existing behavior). Round-trip test in run_context/tests.rs.
  • Bounds/overflow: saturating_mul / integer division in the derivation; saturating_sub in visible_transcript_tokens. Reserve clamped so visible transcript is never zero for a non-zero window.
  • Error class semantics: unchanged; the probe is advisory and never fails a run.
  • Sandbox/native/host names: N/A.

Database Impact

None. resolved_context_budget is per-run, in-memory only — deliberately not persisted (design note §5).

Blast Radius

ironclaw_loop_contracts (type + run-context field), ironclaw_agent_loop (compaction reads the run-context budget; replay digests), ironclaw_loop_host (gateway method + wrapper threading), ironclaw_turn_runner (per-run resolution in host build), ironclaw_llm (TokenRefreshingProvider::model_metadata no longer refreshes). Behavior changes only for runs whose provider populates context_length (today: Gemini's static table, and any provider later wired). Everything else is on the identical default, pinned by the None tests at each tier.

What could break: a provider that advertises a window smaller than the transcript the product expects (e.g. a stale table entry) would compact earlier than before; the 90% margin and reserve clamp bound how aggressive that gets, and the route-identity check prevents borrowing the wrong model's window.

Rollback Plan

Revert the branch (16 commits, no schema, no persisted state). A partial rollback is also safe: reverting only b963204bc (turn-runner resolution) returns every run to the compiled-in default while leaving the plumbing inert.

Review Follow-Through

  • Whether a tracking issue should be filed and linked (Linked Issue above).
  • The ponytail: in loop_driver_host.rs records the one deliberate shortcut: the window probe awaits on the host-build critical path, justified by the new CONTRACT rule that model_metadata() is I/O-free; if that rule is ever relaxed, promote it to a tokio::spawn beside user_profile_fetch.
  • Follow-on slice named in the design note: populating context_length for more providers (static table + free catalog fields), which is what makes this budget bite for them.

Review track: B (feature; touches loop runtime behavior but no security, DB, or CI surface)

🤖 Generated with Claude Code

henrypark133 and others added 16 commits September 1, 2026 22:51
…ed window

Adds PromptContextTokenBudget::from_advertised_window, which turns a
provider-advertised total context window into a usable budget: 90% of the
window to absorb chars/4 estimate error, with the flat response reserve
clamped so a small-window model still has room for transcript.

None (or zero) reproduces the compiled-in 128k/20k exactly, so a provider
that advertises nothing behaves as it does today. No caller yet.

Also adds serde::Deserialize, which the LoopRunContext field needs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
Adds resolved_context_budget: Option<PromptContextTokenBudget>, mirroring
the existing resolved_model_route field and builder exactly: serde-default
so runs recorded before this change still replay, and one builder for the
single writer that will set it.

No producer or consumer yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
Structural only. can_evaluate and trigger_at take the budget instead of
reading self.prompt_context_budget, and all 7 call sites pass exactly the
value the field read produced, so every compaction decision is identical.

Zero test edits: no test calls either helper directly (the test named
can_evaluate_skips_when_visible_threshold_equals_preserve_tail asserts on
should_compact). That zero-churn result is the proof this is internal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
should_compact now reads ctx.resolved_context_budget via effective_budget,
falling back to the strategy's own budget when the run resolved none. Both
CompactionStrategy implementors go through the same helper.

The compaction ceiling is no longer a fixed property of the family, so the
four replay fingerprints say context_limit=run_context and all four
ComponentDigest constants are recomputed from failing-test output. The
digest now stays stable across models rather than encoding one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
… gateway

Adds HostManagedModelGateway::advertised_context_window_tokens, a defaulted
async trait method copying diagnostic_effective_model's shape, so every
other implementor keeps today's behavior unchanged.

LlmProviderModelGateway overrides it by reading ModelMetadata.context_length
-- a field that has had no consumer outside ironclaw_llm until now. It only
trusts the window when the model the provider describes is the model this
run will actually be served: request_model_override lets a route override
the model, and borrowing a larger model's window is exactly the provider
rejection this work exists to prevent. A mismatch falls back to the default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
Structural only. ThreadResolvingLoopModelGatewayParts gains a thirteenth
field, carried onto the gateway and applied to ThreadBackedLoopModelPort
via the builder that until now had only test callers.

That port's resolve_model_messages is what calls
select_prompt_context_messages -- the call deciding which transcript
messages actually reach the provider. It was the one budget consumer the
earlier draft of this work left on the compiled-in default, which would
have shipped a loop that compacts against one ceiling and sizes requests
against another.

Both construction arms pass self.config.prompt_context_budget, which is
PromptContextTokenBudget::default() today -- byte-identical to the port's
own default -- so request sizing is unchanged. Task 7 changes the value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
…odel

build_text_only_host_with_capabilities now asks the run's gateway for the
model's advertised context window, derives a budget from it, and carries it
on LoopRunContext. All four consumers read the one resolved local: the
compaction strategy (via the run context), the prompt context port, both
model-gateway construction arms, and structured finalization.

Two details worth keeping:

- The gateway is resolved once via resolve_for_scope and that same object
  answers the window query and serves the run. Asking self.model_gateway
  while a scope override is active would let the budget describe a
  different gateway than the one issuing the request.
- The await sits after the three advisory prefetch kickoffs, not before
  them, so it cannot serialize work the surrounding code deliberately runs
  in parallel.

A gateway that advertises nothing leaves the run on the configured default,
which is exactly today's behavior.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
TraceLlm -- the SDK-seam fake the in-process integration tier mocks at --
had no model_metadata impl, so it inherited the trait default and reported
no window. That made the model-derived budget unreachable from any
integration scenario.

Adds an advertised_context_window field with a with_* setter and a
model_metadata override, plus an advertised_context_window(tokens) method
on the harness thread builder that applies it. Unset stays None, so every
existing scenario is unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
…ext budget

ThreadResolvingLoopModelGateway is what the driver host hands the loop, so a
budget that stops at the wrapper never sizes the outbound request. The
port-level test already proved the port honors a budget; this pins that the
wrapper forwards its own budget instead of the port's compiled-in default.
Fails with all five seeded messages forwarded when the builder call in
stream_model_inner is removed.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
…und request

The existing tests proved a resolved budget lands on the run context; this
one drives an empty model request through the built host and asserts the
gateway receives only the transcript tail the derived budget admits. Fails
with all five seeded messages forwarded when the model port is handed the
config default instead of the run-derived budget.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
…eal turn

The scripted provider advertises a 40k context window; through the real
LlmProviderModelGateway (route-identity check included) the run derives a
36k/9k budget and the model is sent fewer transcript messages than an
unadvertised run, which keeps the compiled-in 128k/20k ceiling. Fails with
'13 vs 13' when the production gateway stops reporting a window.

Compaction and request sizing share one threshold on this path, so this tier
does not separate them; each link of the request-sizing chain is pinned by
mutation-verified crate tests in ironclaw_loop_host and ironclaw_turn_runner.
Adds the harness builder knob and captured_request_message_count, and the
tests/AGENTS.md coverage row.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
- ironclaw_llm CONTRACT: model_metadata().context_length now drives the
  per-run prompt budget via LlmProviderModelGateway::advertised_context_window_tokens.
- model_gateway: tag the advisory .ok()? with the silent-ok marker the
  error-handling rule greps for.
- integration support: restore script()'s doc comment, displaced by the
  advertised_context_window builder.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
Matches the crate's host/run_context/tests.rs and runtime_context/tests.rs
layout. No behavior change; same eight tests.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
…13773

Genuine contract growth from the model-derived prompt budget: the
from_advertised_window derivation on the type this crate owns and the
optional resolved budget on LoopRunContext (~41 lines). main already sat
3 lines under the effective ceiling. Count read from the test's own failure
message; reason recorded in the ladder comment.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
TokenRefreshingProvider::model_metadata ran ensure_fresh_token first — a
possible HTTP POST to the OAuth token endpoint under the renewal lock shared
with in-flight complete() calls — before delegating to an inner that builds
a static struct needing no credential. With model_metadata() now awaited
once per turn-run host build (advertised_context_window_tokens), every turn
on the OpenAI Codex chain could pay a network round trip, or wait behind an
in-flight refresh, to learn nothing.

Delete the refresh; pin it with a test whose counting local endpoint saw one
connection before the fix and none after. Make the rule explicit in the
ironclaw_llm CONTRACT (model_metadata is a static, I/O-free description;
decorators delegate it unchanged) and point the driver host's critical-path
comment at that contract instead of at an observation.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
Copilot AI lite review requested due to automatic review settings September 3, 2026 15:56
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-8053 September 3, 2026 15:56 Destroyed
@railway-app

railway-app Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-8053 environment in ironclaw-ci-preview

Service Status Web Updated
ironclaw ✅ Success (View Logs) Web Sep 3, 2026 at 11:11 pm UTC

@github-actions github-actions Bot added scope: docs Documentation scope: dependencies Dependency updates labels Sep 3, 2026
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-03T16:01:09.102864Z 120402e PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions github-actions Bot added size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules labels Sep 3, 2026
@github-actions github-actions Bot added the contributor: core 20+ merged PRs label Sep 3, 2026
@coderabbitai

coderabbitai Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Team

Run ID: 3f4b595a-2921-4e03-808c-9a47ad5af395

📥 Commits

Reviewing files that changed from the base of the PR and between f0cdace and 16c51cc.

📒 Files selected for processing (2)
  • crates/loop/ironclaw_loop_host/src/lib.rs
  • crates/loop/ironclaw_loop_host/src/model_gateway.rs

Included review availability: Your plan provides up to 10 included reviews per hour; 5 remain after this review.


📝 Summary

Summary by CodeRabbit

  • New Features

    • Prompt budgets can be derived from a model’s advertised context window.
    • Transcript sizing and compaction adapt to the resolved budget.
    • Models without advertised limits retain the existing default budget.
    • Multi-model routing and failover use the smallest known shared context window.
    • Integration tools support configuring and validating advertised windows.
  • Bug Fixes

    • Metadata lookup no longer refreshes tokens or performs network activity.
    • Unknown model context limits are no longer estimated inaccurately.
  • Documentation

    • Added guidance for context-window handling and provider metadata.

Walkthrough

The change derives per-run prompt budgets from provider-advertised context windows. It stores budgets on the run context and applies them to transcript sizing and compaction. It preserves compiled-in defaults when metadata is absent or mismatched.

Changes

Model-derived prompt context budget

Layer / File(s) Summary
Budget and run-context contracts
crates/contracts/ironclaw_loop_contracts/src/context_budget*, crates/contracts/ironclaw_loop_contracts/src/host/run_context*
Adds advertised-window budget derivation, reserve clamping, deserialization, and an optional serialized run-context budget.
Provider metadata gateway
crates/domains/ironclaw_llm/*, crates/loop/ironclaw_loop_host/src/{lib.rs,model_gateway.rs}
Adds advertised-window metadata handling, minimum-window reporting, model identity checks, and no-refresh metadata access.
Run-specific compaction
crates/loop/ironclaw_agent_loop/src/strategies/*, crates/loop/ironclaw_agent_loop/src/families/*
Compaction uses the resolved budget, clamps preserved tails, and updates replay fingerprints and digests.
Turn-runner propagation
crates/loop/ironclaw_turn_runner/src/loop_driver_host*, crates/loop/ironclaw_loop_host/src/thread_resolving_model_gateway.rs
Resolves one budget per run and passes it to prompt context and model request sizing.
Validation and support
tests/integration/*, tests/support/*, tests/trace_llm_tests.rs, Cargo.toml, crates/app/ironclaw_architecture_tests/*
Adds provider fixtures, harness configuration, request assertions, integration coverage, test registration, and architecture-size accounting.
Documentation and CI
docs/internal/*, tests/AGENTS.md, scripts/ci/*, .config/nextest.toml
Documents the design, updates test inventories, registers root test partitions, and adjusts selected slow-test thresholds.
Notification setup test correction
crates/product/ironclaw_webui/frontend/src/lib/api-boundary.test.ts
Uses the web-app extension ID in the notification setup rejection test.

Estimated code review effort: 4 (Complex) | ~45 minutes

Merge Risk: ⚪ Minimal · up to 16c51

This change adapts agent prompt budgets to matching provider-advertised model windows while preserving existing defaults when metadata is unavailable. No current merge-blocking risk remains.

Sequence Diagram(s)

sequenceDiagram
  participant TurnRunner
  participant LlmProviderModelGateway
  participant LlmProvider
  participant LoopRunContext
  participant CompactionStrategy
  participant ThreadBackedLoopModelPort
  TurnRunner->>LlmProviderModelGateway: query advertised context window
  LlmProviderModelGateway->>LlmProvider: model_metadata()
  LlmProvider-->>LlmProviderModelGateway: context_length and model id
  LlmProviderModelGateway-->>TurnRunner: matching window or None
  TurnRunner->>LoopRunContext: store resolved_context_budget
  TurnRunner->>CompactionStrategy: evaluate with effective budget
  TurnRunner->>ThreadBackedLoopModelPort: apply prompt context token budget
Loading

Suggested reviewers: serrrfirat

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description covers the required sections and provides detailed validation evidence. However, the new feature has no linked approved issue, despite the template requiring one. The description also … Link an approved issue for this new feature, or obtain and document an explicit maintainer exception. Run pr-shepherd before requesting review and update the Validation section with the result.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title follows Conventional Commits style and accurately summarizes the main change: deriving the prompt context budget from the model's advertised window.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Description check

Explanation

The description covers the required sections and provides detailed validation evidence. However, the new feature has no linked approved issue, despite the template requiring one. The description also states that pr-shepherd has not yet been run after a coding-agent change.

  • Fix all pre-merge checks with AI

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 120402e9ac

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/loop/ironclaw_loop_host/src/model_gateway.rs
Comment thread crates/loop/ironclaw_loop_host/src/model_gateway.rs

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new integration test relies on a fixed captured-request index that can shift when compaction adds extra model calls, making the assertion potentially brittle/flaky.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

This PR makes the agent loop’s prompt context budget model-aware by deriving it from the provider-advertised context window (when available) and threading the resolved budget through the turn-runner host build into all relevant loop consumers (compaction + message selection), while preserving today’s 128k/20k behavior when no window is advertised.

Changes:

  • Add PromptContextTokenBudget::from_advertised_window(..) and carry an optional resolved budget on LoopRunContext.
  • Extend the loop-host gateway surface to report an advertised context window (with route/model-id verification) and wire the resolved budget through turn-runner host construction and model/message selection.
  • Update compaction strategies and replay fingerprints/digests to treat the compaction ceiling as run-scoped, and add/adjust contract + unit + integration tests (including SDK-seam fakes) to pin the end-to-end behavior.
File summaries
File Description
Cargo.toml Registers new integration test target for context-budget scenario.
crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rs Re-pins ironclaw_loop_contracts size ceiling with rationale for the new lines.
crates/contracts/ironclaw_loop_contracts/src/context_budget.rs Adds Deserialize + from_advertised_window derivation; moves tests out-of-line.
crates/contracts/ironclaw_loop_contracts/src/context_budget/tests.rs Adds coverage for advertised-window budget derivation behavior.
crates/contracts/ironclaw_loop_contracts/src/host/run_context.rs Adds resolved_context_budget: Option<_> + builder on LoopRunContext.
crates/contracts/ironclaw_loop_contracts/src/host/run_context/tests.rs Adds serde/default and builder/round-trip tests for the new run-context field.
crates/domains/ironclaw_llm/CONTRACT.md Documents model_metadata() invariants + the new runtime consumption of context_length.
crates/domains/ironclaw_llm/src/token_refreshing.rs Removes token refresh from model_metadata() and adds a regression test to pin I/O-free behavior.
crates/loop/ironclaw_agent_loop/src/families/mod.rs Updates default family fingerprint/digest for run-scoped compaction budget.
crates/loop/ironclaw_agent_loop/src/families/subagent.rs Updates subagent family fingerprint/digest for run-scoped compaction budget.
crates/loop/ironclaw_agent_loop/src/families/unbound.rs Updates unbound family fingerprints/digests for run-scoped compaction budget.
crates/loop/ironclaw_agent_loop/src/strategies/active_task_compaction.rs Uses run-context resolved budget for compaction; adds tests for override/fallback behavior.
crates/loop/ironclaw_agent_loop/src/strategies/compaction.rs Uses run-context resolved budget for compaction; refactors helper signatures; adds tests.
crates/loop/ironclaw_loop_host/src/lib.rs Adds defaulted advertised_context_window_tokens to HostManagedModelGateway.
crates/loop/ironclaw_loop_host/src/model_gateway.rs Implements advertised-window lookup on provider-backed gateway with route/model-id check; adds tests.
crates/loop/ironclaw_loop_host/src/thread_resolving_model_gateway.rs Threads a prompt-context budget into the wrapper so outbound request sizing uses the resolved budget.
crates/loop/ironclaw_loop_host/tests/thread_loop_host_contract.rs Adds contract test proving the wrapper applies its injected budget to message selection.
crates/loop/ironclaw_turn_runner/src/loop_driver_host.rs Resolves advertised window once per host build, sets run-context budget, and wires the effective budget to ports/gateway.
crates/loop/ironclaw_turn_runner/src/loop_driver_host/context_budget_tests.rs Adds driver-host tests pinning run-context budget resolution and request sizing behavior.
docs/internal/reborn/design/model-derived-context-budget.md Adds design/spec doc describing the end-to-end shape and invariants.
docs/internal/superpowers/plans/2026-09-01-model-derived-context-budget.md Adds detailed implementation plan/playbook for the change.
tests/AGENTS.md Updates scenario coverage map to include the new integration test.
tests/integration/context_budget.rs Adds integration scenario asserting advertised-window runs send fewer messages than unadvertised runs.
tests/integration/support/assertions.rs Adds harness helper for inspecting captured request message counts.
tests/integration/support/builder.rs Plumbs advertised context window through integration harness builder to scripted provider.
tests/integration/support/group.rs Plumbs advertised context window through grouped-thread builder path.
tests/support/trace_llm.rs Extends TraceLlm to optionally advertise a context window via model_metadata().
tests/trace_llm_tests.rs Adds tests verifying TraceLlm’s default vs configured advertised context window.
Review details
  • Files reviewed: 28/28 changed files
  • Comments generated: 3
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread tests/integration/context_budget.rs Outdated
Comment thread crates/loop/ironclaw_loop_host/src/lib.rs Outdated
Comment thread tests/integration/support/assertions.rs Outdated
@ironloopai

ironloopai Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Review · Status

🟩 Completed

IronLoop completed the review and posted it to GitHub.

Result

Open submitted review →

Run details
  • Run: c4b32b31-1edd-4b73-b25a-e9248fff5913
  • Base: main at d3e62cf
  • Head: context-length at 120402e
  • Created: 2026-09-03 16:01 UTC
  • Updated: 2026-09-03 16:23 UTC

Automatic trigger · attempt 1 of 3 · completed in 21m 55s

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It changes core loop budgeting behavior across multiple crates and relies on cross-component invariants (provider metadata + host build + compaction + request sizing) that warrant final human review despite strong test coverage.

Review details
  • Files reviewed: 34/34 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Derive per-run prompt context budgets from the served model’s advertised window while preserving defaults and removing unnecessary metadata token-refresh I/O.

Stats: 3 findings (from 6 raw, 6 after filter, 3 after dedup) across 3 files. Reviewers run: correctness, security, performance, design, coverage. Reviewers failed: none. Body-only: 0

Unresolved review threads at emission: 0

Mechanical

  1. Medium New plan file exceeds the repository’s 1,000-line file-size ceiling (docs/internal/superpowers/plans/2026-09-01-model-derived-context-budget.md:1-1332, confidence 90) — anchor: docs/internal/superpowers/plans/2026-09-01-model-derived-context-budget.md:1
    This change adds a 1,332-line plan document, crossing the repository’s stated 1,000-line file-size ceiling.

Correctness / bugs

  1. High Budget ignores smaller failover provider windows (crates/loop/ironclaw_loop_host/src/model_gateway.rs:431-451, confidence 95) — anchor: crates/loop/ironclaw_loop_host/src/model_gateway.rs:433
    When the gateway wraps FailoverProvider, model_metadata() describes only its current last_used provider, initially the primary. If a large-window primary fails and fallback advances to a smaller-window provider, the run retains the primary-derived budget and can send a prompt larger than the fallback accepts, causing repeated context-overflow failure instead of recovery.
    Also flagged by: coverage/High, design/Medium

Coverage / tests

  1. Medium Token-refresh regression is not tested through its production caller (crates/domains/ironclaw_llm/src/token_refreshing.rs:620-632, confidence 91) — anchor: crates/domains/ironclaw_llm/src/token_refreshing.rs:620
    model_metadata_does_not_refresh_the_token calls TokenRefreshingProvider::model_metadata() directly. The side effect matters because LlmProviderModelGateway::advertised_context_window_tokens invokes that decorator during host construction; a future gateway/decorator wiring change could reintroduce the HTTP call while the leaf test remains green.
    Also flagged by: coverage/Medium

Comment thread crates/loop/ironclaw_loop_host/src/model_gateway.rs
Comment thread crates/domains/ironclaw_llm/src/token_refreshing.rs
henrypark133 and others added 2 commits September 3, 2026 21:06
…chain

model_metadata() returned providers[last_used]'s metadata. The gateway reads
it once per run at host build, while last_used is still the primary, and the
derived budget lives on the run context for the whole run; a later in-run
failure (or the gateway's fallback_index walk) serves a smaller-window
fallback with the primary's budget and overflows instead of recovering.
Composition builds the chain from a primary and a fallback_model it expects
to differ.

Advertise min(context_length) across every chain member, None if any is
unknown, keeping id on last_used so the gateway's identity check stays
consistent with active_model_name(). Same shape as SmartRoutingProvider;
CONTRACT names failover chains alongside routing. Swappable delegates live
to its current inner and is operator-swapped, not failure-driven — unchanged.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
…h the gateway

The token-refresh regression test called TokenRefreshingProvider directly;
the production caller is LlmProviderModelGateway::advertised_context_window_tokens
during host build, and a decorator re-wiring could reintroduce the HTTP call
with the leaf test still green. Drive the real gateway over the real
decorator with a session loaded from an expired on-disk session file (no new
seam) and a counting local endpoint: 1 connection with the refresh
reinserted, 0 with the fix.

Lives in tests/llm_gateway.rs beside the other real-gateway tests: the
reborn_product_api_crates_do_not_bind_http_ingress gate scans this crate's
src/ (test modules included) for TcpListener::bind, and the loopback counting
endpoint is exactly that.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
Copilot AI review requested due to automatic review settings September 3, 2026 21:50
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-8053 September 3, 2026 21:50 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

There are still correctness/perf issues on the host-build-critical model_metadata() path (redundant metadata awaits in FailoverProvider) and a semantic mismatch where “unknown” advertised windows can incorrectly populate resolved_context_budget instead of leaving it None.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review details

Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

crates/loop/ironclaw_turn_runner/src/loop_driver_host.rs:1687

  • PromptContextTokenBudget::from_advertised_window(Some(window)) intentionally treats None/0/too-small windows as “unknown” by returning the compiled-in default. But the host-build currently stores that default back into run_context.resolved_context_budget whenever the gateway returns Some(window) (including 0 or other values that derive to the default), which contradicts the LoopRunContext field docs that say None represents “provider advertises no window / fall back to the default”. Consider only setting resolved_context_budget when the derived budget actually differs from the configured fallback, and filter out window == 0 up front.
  • Files reviewed: 36/36 changed files
  • Comments generated: 1
  • Review effort level: Lite

Comment thread crates/domains/ironclaw_llm/src/failover.rs Outdated
…indow

Review nit: model_metadata() awaited the last_used member twice and kept
querying after the aggregate window had already become None. One call per
member, id captured for last_used, early exit once the aggregate is None.
Same result. Pinned by two tests counting model_metadata() calls per member:
on the old body the last_used member is queried twice (left: 2, right: 1).

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
Copilot AI review requested due to automatic review settings September 3, 2026 22:17
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-8053 September 3, 2026 22:17 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It changes core loop runtime budgeting/compaction semantics across multiple crates and relies on subtle critical-path contract guarantees that merit a final human review despite strong test coverage.

Review details

Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

crates/loop/ironclaw_turn_runner/src/loop_driver_host.rs:1664

  • The comment here says the await resolves immediately because model_metadata() is I/O-free, but the awaited call is advertised_context_window_tokens(), which is a gateway method that can be overridden independently of LlmProvider::model_metadata(). As written, this reads like a stronger guarantee than the trait provides and could mislead future implementors into adding a slow/I/O implementation without noticing it lands on the host-build critical path.

Consider rewording to explicitly tie the “fast/no-I/O” assumption to the provider-backed gateway’s implementation (which delegates to model_metadata()), and keep the ponytail note as the fallback if that assumption changes.

  • Files reviewed: 36/36 changed files
  • Comments generated: 1
  • Review effort level: Lite

Comment thread crates/loop/ironclaw_loop_host/src/lib.rs
…room

The three reborn_*_location_scan binaries each carry one scan of the entire
crates/ tree. On the last green run they finished at 176.8s against the ci
profile's 60s x 3 = 180s hard kill; the next run was terminated at 180.008s
with no code change that touches their runtime. reborn_sealed_evidence_mint_ratchet
rides the same ladder to 136-141s. One override for that family at 60s x 6
keeps the SLOW cadence and stops every PR's architecture bucket being a
coin flip; the default stays as is for every other binary.

Match verified locally: with the override's period set to 150s no SLOW marker
fires (tests take ~110s); with 60s the markers return at 60s.

Rollback: delete the override block.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer
Copilot AI review requested due to automatic review settings September 3, 2026 22:24
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-8053 September 3, 2026 22:24 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It changes critical-path prompt sizing/compaction behavior across multiple crates and provider decorators, so it needs a final human review despite strong test coverage.

Review details

Suppressed comments (2)

Previously missed (1) — in code that hasn't changed since the last review.

tests/integration/support/assertions.rs:591

  • captured_last_request_message_count() calls TraceLlm::captured_requests() (which clones every captured request and message) and then reads .last(). With large transcripts (as in the new context-budget integration tests), this does unnecessary O(total captured size) cloning just to get the last request length and can significantly slow tests / increase memory.

Prefer a dedicated TraceLlm accessor that reads the last captured call under the lock (e.g. captured_last_request_len() / captured_last_request_messages()), cloning only what’s needed (or just returning the length), and use that here (and in captured_last_request_contents()).

This issue also appears on line 599 of the same file.

tests/integration/support/assertions.rs:603

  • captured_last_request_contents() also calls TraceLlm::captured_requests() and then takes .last(), which clones all captured requests/messages even though only the last request is needed. This is especially expensive when messages contain large content strings.

Consider adding a TraceLlm API that returns a clone of only the last request’s messages (or iterates them under the lock) so this helper doesn’t duplicate the entire capture history.

  • Files reviewed: 37/37 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Derive prompt context budgets from the model’s advertised context window and remove token-refresh I/O.

Stats: 3 findings (from 3 raw, 3 after filter, 3 after dedup) across 2 files. Reviewers run: correctness, security, performance, design, coverage. Reviewers failed: none. Body-only: 0

Mechanical

  1. Medium New plan document exceeds the 1,000-line file-size ceiling (docs/internal/superpowers/plans/2026-09-01-model-derived-context-budget.md:1-1332, confidence 90) — anchor: docs/internal/superpowers/plans/2026-09-01-model-derived-context-budget.md:1
    This change adds a 1,332-line plan document. The repository change contract caps touched source files at roughly 1,000 lines; large planning artifacts are harder to review and maintain.

Regression escape

  1. Medium Route-override model identity is not regression-tested (crates/loop/ironclaw_loop_host/src/model_gateway.rs:443-450, confidence 94) — anchor: crates/loop/ironclaw_loop_host/src/model_gateway.rs:443
    The metadata path is tested only with resolved_model_route = None and a policy-level override. Production runs can carry a HostManagedModelRouteSnapshot; a regression could derive a budget from the wrong model.

Tests

  1. Medium Metadata-error fallback is not regression-tested (crates/loop/ironclaw_loop_host/src/model_gateway.rs:436, confidence 87) — anchor: crates/loop/ironclaw_loop_host/src/model_gateway.rs:436
    The advisory fallback when provider.model_metadata() returns an error is untested. A provider/decorator failure must leave the run on the compiled-in budget rather than fail host construction or use partial metadata.

Comment thread crates/loop/ironclaw_loop_host/src/model_gateway.rs
Comment thread crates/loop/ironclaw_loop_host/src/model_gateway.rs
…a-error paths

Review round three: the resolved-route-snapshot branch of the identity check
and the model_metadata() error fallback had no tests. Snapshot naming another
model -> None (fails with Some(200000) when the identity check is removed);
snapshot naming the served model -> the window; metadata error -> None (the
double's error flag is what flips the outcome). The trait doc now states the
probe is awaited on the host-build critical path and must be cheap.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X6JWfzEzWKThikJ9wsHAer

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It changes critical-path turn-run host construction and cross-crate loop budgeting behavior, so it warrants final human review despite strong test coverage.

Review details
  • Files reviewed: 37/37 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The derivation is consistently wired end-to-end with strong multi-tier regression coverage, and the model_metadata() critical-path I/O removal is enforced by both code and contract tests.

Review details
  • Files reviewed: 37/37 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

This branch was successfully deployed

1 active deployment
ironclaw-ci-preview / ironclaw-pr-8053 — 16c51cca Deployed Sep 3, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: dependencies Dependency updates scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants