Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
4bede65
feat(loop-contracts): derive a prompt context budget from an advertis…
henrypark133 Sep 1, 2026
198974f
feat(loop-contracts): carry a resolved context budget on LoopRunContext
henrypark133 Sep 1, 2026
328ffaa
refactor(agent-loop): pass the compaction budget as an argument
henrypark133 Sep 1, 2026
267b0b5
feat(agent-loop): honor a run-resolved context budget when compacting
henrypark133 Sep 1, 2026
c98383f
feat(loop-host): expose the provider-advertised context window on the…
henrypark133 Sep 1, 2026
2100238
refactor(loop-host): thread a prompt context budget to the model port
henrypark133 Sep 1, 2026
b963204
feat(turn-runner): resolve the prompt context budget from the run's m…
henrypark133 Sep 1, 2026
1c714d2
test(support): let the scripted model advertise a context window
henrypark133 Sep 1, 2026
2888785
test(loop-host): pin that the gateway wrapper applies its prompt cont…
henrypark133 Sep 2, 2026
55ff1c9
test(turn-runner): pin that the derived budget sizes the host's outbo…
henrypark133 Sep 2, 2026
d293fe1
test(integration): prove the model-derived context budget through a r…
henrypark133 Sep 2, 2026
adfc802
docs: record the context_length consumer and tidy review leftovers
henrypark133 Sep 2, 2026
c842ed0
refactor(loop-contracts): move context_budget tests out of line
henrypark133 Sep 2, 2026
b687623
chore(architecture-tests): re-pin ironclaw_loop_contracts ceiling at …
henrypark133 Sep 2, 2026
ca68ac4
docs(internal): design note and plan for the model-derived context bu…
henrypark133 Sep 2, 2026
120402e
fix(llm): keep model_metadata() free of token refresh I/O
henrypark133 Sep 2, 2026
ec9ead5
refactor(agent-loop): move compaction strategy tests out of line
henrypark133 Sep 3, 2026
3ad5c72
ci(reborn-plan): register tests/trace_llm_tests.rs as a root test par…
henrypark133 Sep 3, 2026
2ced021
fix(loop-contracts): treat a window too small for any transcript as u…
henrypark133 Sep 3, 2026
2b35d1e
test(integration): read the last captured model request, not index 5
henrypark133 Sep 3, 2026
b1f172f
fix(llm): never advertise a guessed or too-large context window
henrypark133 Sep 3, 2026
0eb4701
fix(agent-loop): bound the compaction preserve tail by the run's visi…
henrypark133 Sep 3, 2026
03bf75c
Merge origin/main into context-length
henrypark133 Sep 3, 2026
a2699f9
test(webui): use the live extension id in the notification-setup boun…
henrypark133 Sep 3, 2026
cd277c0
fix(llm): FailoverProvider advertises the smallest window across its …
henrypark133 Sep 3, 2026
c7af1f1
test(loop-host): pin the window probe never refreshes a token, throug…
henrypark133 Sep 3, 2026
1f870b7
refactor(llm): query each failover member once when advertising the w…
henrypark133 Sep 3, 2026
f0cdace
ci(nextest): give the whole-tree architecture scans real timeout head…
henrypark133 Sep 3, 2026
16c51cc
test(loop-host): pin the window probe's snapshot-identity and metadat…
henrypark133 Sep 3, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions .config/nextest.toml
Original file line number Diff line number Diff line change
Expand Up @@ -64,3 +64,15 @@ slow-timeout = { period = "300s", terminate-after = 2 }
[[profile.ci.overrides]]
filter = "binary(~reborn_group_)"
slow-timeout = { period = "28m", terminate-after = 1 }

# ironclaw_architecture_tests: reborn_{extension_contract,loop_port,product_contract}_location_scan
# and reborn_sealed_evidence_mint_ratchet each contain one whole-crates/-tree
# scan test. On the last green run the location-scan tests finished at
# 176.8s against the default 60s/3 = 180s hard kill; the very next run was
# terminated at 180.008s. reborn_sealed_evidence_mint_ratchet rides the same
# ladder to 136-141s. Locally these take ~110s. 60s/6 = 360s keeps the same
# SLOW cadence but gives real headroom instead of raising the default for
# every binary in the crate.
[[profile.ci.overrides]]
filter = 'binary(/_location_scan$/) | binary(~reborn_sealed_evidence_mint_ratchet)'
slow-timeout = { period = "60s", terminate-after = 6 }
4 changes: 4 additions & 0 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -332,6 +332,10 @@ path = "tests/integration/backend_matrix.rs"
name = "reborn_integration_budget"
path = "tests/integration/budget.rs"

[[test]]
name = "reborn_integration_context_budget"
path = "tests/integration/context_budget.rs"
Comment thread
henrypark133 marked this conversation as resolved.

[[test]]
name = "reborn_integration_cancel"
path = "tests/integration/cancel.rs"
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1039,7 +1039,15 @@ fn reborn_contracts_crates_carry_a_checked_size_ceiling() {
// production prompt validation checks structural limits and control
// characters only; decoded Basic-auth samples remain test-only. Count
// read from this test's own failure message after merging #7416 and #6985.
("ironclaw_loop_contracts", 13_608),
// 13_608 -> 13_773 (2026-09-02, model-derived prompt context budget):
// `PromptContextTokenBudget::from_advertised_window` derives the
// per-run budget on the type this crate owns, and `LoopRunContext`
// carries the optional resolved budget; derivation stays a pure
// function of the DTO, resolution and consumption stay in
// `ironclaw_loop_host` / `ironclaw_turn_runner` / `ironclaw_agent_loop`.
// `main` already sat 3 lines under the effective ceiling. Count read
// from this test's own failure message.
("ironclaw_loop_contracts", 13_773),
// Raised 15_685 -> 15_758 by #7220 (operator inspector API): the growth
// is bounded, output-only read-view descriptors. Capture, retention,
// authorization, and transport behavior remain in their owning
Expand Down
62 changes: 37 additions & 25 deletions crates/contracts/ironclaw_loop_contracts/src/context_budget.rs
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@
/// Storage still scans transcript context by message count. Host adapters use
/// this budget after that scan, and compaction strategies use the same budget
/// shape to decide when the observed prompt is near its context ceiling.
#[derive(Debug, Clone, Copy, PartialEq, Eq, serde::Serialize)]
#[derive(Debug, Clone, Copy, PartialEq, Eq, serde::Serialize, serde::Deserialize)]
pub struct PromptContextTokenBudget {
pub context_limit_tokens: u64,
pub reserve_tokens: u64,
Expand All @@ -16,6 +16,13 @@ impl PromptContextTokenBudget {
pub const DEFAULT_RESERVE_TOKENS: u64 = 20_000;
pub const DEFAULT_MAIN_LOOP_MAX_OUTPUT_TOKENS: u64 = 0;

/// Fraction of a provider-advertised window we are willing to fill.
///
/// This margin exists to absorb error in the chars/4 token estimate
Comment thread
henrypark133 marked this conversation as resolved.
/// (`estimate_tokens_from_chars`), which is the only reason for it. Room
/// for the model's *response* is a separate axis — `reserve_tokens`.
pub const DEFAULT_USABLE_FRACTION_PERCENT: u64 = 90;

pub const fn new(
context_limit_tokens: u64,
reserve_tokens: u64,
Expand All @@ -32,6 +39,34 @@ impl PromptContextTokenBudget {
self.context_limit_tokens
.saturating_sub(self.reserve_tokens.max(self.main_loop_max_output_tokens))
}

/// Derive a budget from a provider-advertised total context window.
///
/// `None` (or a nonsense zero) reproduces the compiled-in default
/// exactly, so a provider that advertises nothing behaves as it always
/// has. Never guess a window for an unknown model: guessing high
/// produces the provider rejection this mechanism exists to avoid. A
/// window too small to leave any visible transcript after the reserve
/// (today: below 2 tokens) is treated as unknown, like `None` and `0`.
Comment thread
henrypark133 marked this conversation as resolved.
pub fn from_advertised_window(advertised_tokens: Option<u64>) -> Self {
let Some(advertised) = advertised_tokens.filter(|tokens| *tokens > 0) else {
return Self::default();
};
let context_limit_tokens =
advertised.saturating_mul(Self::DEFAULT_USABLE_FRACTION_PERCENT) / 100;
// A small-window model would otherwise have its whole budget consumed
// by the flat response reserve, leaving zero visible transcript.
let reserve_tokens = Self::DEFAULT_RESERVE_TOKENS.min(context_limit_tokens / 4);
let candidate = Self {
context_limit_tokens,
reserve_tokens,
main_loop_max_output_tokens: Self::DEFAULT_MAIN_LOOP_MAX_OUTPUT_TOKENS,
};
if candidate.visible_transcript_tokens() == 0 {
return Self::default();
}
candidate
}
}

impl Default for PromptContextTokenBudget {
Expand All @@ -45,27 +80,4 @@ impl Default for PromptContextTokenBudget {
}

#[cfg(test)]
mod tests {
use super::PromptContextTokenBudget;

#[test]
fn visible_transcript_tokens_reserves_larger_output_buffer() {
let budget = PromptContextTokenBudget::new(100, 10, 30);

assert_eq!(budget.visible_transcript_tokens(), 70);
}

#[test]
fn visible_transcript_tokens_saturates_when_reserve_exceeds_limit() {
let budget = PromptContextTokenBudget::new(10, 20, 0);

assert_eq!(budget.visible_transcript_tokens(), 0);
}

#[test]
fn visible_transcript_tokens_uses_reserve_when_larger_than_output_budget() {
let budget = PromptContextTokenBudget::new(100, 30, 10);

assert_eq!(budget.visible_transcript_tokens(), 70);
}
}
mod tests;
112 changes: 112 additions & 0 deletions crates/contracts/ironclaw_loop_contracts/src/context_budget/tests.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,112 @@
use super::*;

#[test]
fn visible_transcript_tokens_reserves_larger_output_buffer() {
let budget = PromptContextTokenBudget::new(100, 10, 30);

assert_eq!(budget.visible_transcript_tokens(), 70);
}

#[test]
fn visible_transcript_tokens_saturates_when_reserve_exceeds_limit() {
let budget = PromptContextTokenBudget::new(10, 20, 0);

assert_eq!(budget.visible_transcript_tokens(), 0);
}

#[test]
fn visible_transcript_tokens_uses_reserve_when_larger_than_output_budget() {
let budget = PromptContextTokenBudget::new(100, 30, 10);

assert_eq!(budget.visible_transcript_tokens(), 70);
}

#[test]
fn advertised_window_of_none_reproduces_the_compiled_in_default() {
// A provider that reports nothing must behave exactly as it does
// today. This is the compatibility guarantee of the whole change.
assert_eq!(
PromptContextTokenBudget::from_advertised_window(None),
PromptContextTokenBudget::default()
);
}

#[test]
fn advertised_window_of_zero_is_treated_as_unknown() {
assert_eq!(
PromptContextTokenBudget::from_advertised_window(Some(0)),
PromptContextTokenBudget::default()
);
}

#[test]
fn large_advertised_window_keeps_the_flat_response_reserve() {
let budget = PromptContextTokenBudget::from_advertised_window(Some(2_000_000));

assert_eq!(budget.context_limit_tokens, 1_800_000);
assert_eq!(
budget.reserve_tokens,
PromptContextTokenBudget::DEFAULT_RESERVE_TOKENS
);
assert_eq!(budget.visible_transcript_tokens(), 1_780_000);
}

#[test]
fn small_advertised_window_clamps_the_reserve_and_keeps_budget_usable() {
// An 8k model would otherwise have its entire budget consumed by the
// flat 20k response reserve, leaving zero visible transcript and a
// loop that cannot run at all.
let budget = PromptContextTokenBudget::from_advertised_window(Some(8_000));

assert_eq!(budget.context_limit_tokens, 7_200);
assert_eq!(budget.reserve_tokens, 1_800);
assert!(
budget.visible_transcript_tokens() > 0,
"a small-window model must still have room for transcript"
);
}

#[test]
fn smallest_positive_window_is_treated_as_unknown() {
// Some(1) survives the `> 0` filter but derives a zero visible
// transcript, which must fall back to the default exactly like
// None and Some(0).
assert_eq!(
PromptContextTokenBudget::from_advertised_window(Some(1)),
PromptContextTokenBudget::default()
);
}

#[test]
fn smallest_usable_window_keeps_a_nonzero_visible_transcript() {
// Find the smallest advertised window whose derivation does NOT
// fall back to the default, and prove it still leaves visible
// transcript room rather than trusting the arithmetic.
let smallest_non_default = (1..=16)
.find(|&candidate| {
PromptContextTokenBudget::from_advertised_window(Some(candidate))
!= PromptContextTokenBudget::default()
})
.expect("some small window must derive a non-default budget");

let budget = PromptContextTokenBudget::from_advertised_window(Some(smallest_non_default));

assert_ne!(budget, PromptContextTokenBudget::default());
assert!(
budget.visible_transcript_tokens() > 0,
"the smallest non-default derived budget must still leave visible transcript room"
);
}

#[test]
fn advertised_window_matching_todays_constant_is_reduced_by_the_margin() {
// 128k advertised is NOT the same as the 128k fallback: the fallback
// is a guess, an advertised value gets the estimate-error margin.
let budget = PromptContextTokenBudget::from_advertised_window(Some(128_000));

assert_eq!(budget.context_limit_tokens, 115_200);
assert_eq!(
budget.reserve_tokens,
PromptContextTokenBudget::DEFAULT_RESERVE_TOKENS
);
}
12 changes: 12 additions & 0 deletions crates/contracts/ironclaw_loop_contracts/src/host/run_context.rs
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ use ironclaw_host_api::{
};
use serde::{Deserialize, Serialize};

use crate::context_budget::PromptContextTokenBudget;
use crate::refs::{CheckpointSchemaId, LoopDriverId};
use crate::snapshot::ResolvedRunProfile;
use ironclaw_host_api::turn::{
Expand Down Expand Up @@ -245,6 +246,11 @@ pub struct LoopRunContext {
pub resolved_run_profile: ResolvedRunProfile,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub resolved_model_route: Option<LoopModelRouteSnapshot>,
/// Prompt context budget resolved from this run's model at host
/// construction. `None` — an older serialized context, or a provider that
/// advertises no window — means the compiled-in default.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub resolved_context_budget: Option<PromptContextTokenBudget>,
pub loop_driver_id: LoopDriverId,
pub loop_driver_version: RunProfileVersion,
pub checkpoint_schema_id: CheckpointSchemaId,
Expand Down Expand Up @@ -279,6 +285,7 @@ impl LoopRunContext {
run_id,
resolved_run_profile,
resolved_model_route: None,
resolved_context_budget: None,
loop_driver_id,
loop_driver_version,
checkpoint_schema_id,
Expand Down Expand Up @@ -347,6 +354,11 @@ impl LoopRunContext {
self
}

pub fn with_resolved_context_budget(mut self, budget: PromptContextTokenBudget) -> Self {
self.resolved_context_budget = Some(budget);
self
}

pub fn with_product_context(mut self, product_context: ProductTurnContext) -> Self {
self.product_context = Some(product_context);
self
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -249,3 +249,59 @@ mod acting_identity_ladder {
);
}
}

fn sample_run_context() -> LoopRunContext {
let scope = TurnScope::new(
ironclaw_host_api::ids::TenantId::new("tenant-context-budget").expect("tenant"),
None,
None,
ThreadId::new("thread-context-budget").expect("thread"),
);
let profile = ResolvedRunProfile::legacy_compatibility(
ironclaw_host_api::turn::RunProfileId::default_profile(),
ironclaw_host_api::turn::RunProfileVersion::new(1),
true,
);
LoopRunContext::new(scope, TurnId::new(), TurnRunId::new(), profile)
}

#[test]
fn run_context_defaults_to_no_resolved_context_budget() {
let context = sample_run_context();

assert_eq!(context.resolved_context_budget, None);
}

#[test]
fn run_context_carries_a_resolved_context_budget() {
let budget = PromptContextTokenBudget::from_advertised_window(Some(200_000));
let context = sample_run_context().with_resolved_context_budget(budget);

assert_eq!(context.resolved_context_budget, Some(budget));
}

#[test]
fn run_context_without_a_budget_field_still_deserializes() {
// Runs recorded before this change must replay, landing on the
// compiled-in default rather than failing to deserialize.
let context = sample_run_context();
let mut wire = serde_json::to_value(&context).expect("serialize");
wire.as_object_mut()
.expect("object")
.remove("resolved_context_budget");

let restored: LoopRunContext = serde_json::from_value(wire).expect("deserialize");

assert_eq!(restored.resolved_context_budget, None);
}

#[test]
fn resolved_context_budget_round_trips_through_the_wire() {
let budget = PromptContextTokenBudget::from_advertised_window(Some(1_000_000));
let context = sample_run_context().with_resolved_context_budget(budget);

let wire = serde_json::to_string(&context).expect("serialize");
let restored: LoopRunContext = serde_json::from_str(&wire).expect("deserialize");

assert_eq!(restored.resolved_context_budget, Some(budget));
}
2 changes: 2 additions & 0 deletions crates/domains/ironclaw_llm/CONTRACT.md
Original file line number Diff line number Diff line change
Expand Up @@ -236,6 +236,8 @@ Key notes:
- `cost_per_token()` returns `(Decimal, Decimal)` using `rust_decimal`. Look up via `costs::model_cost()` in your constructor; fall back to `costs::default_cost()` for unknowns.
- `RigAdapter` forwards per-request model overrides through rig-core's typed request model field. Do not put `model` in flattened `additional_params`, which would serialize a duplicate top-level JSON key.
- `complete_with_tools()` is never cached (tool calls can have side effects) — `CachedProvider` always passes them through.
- `model_metadata().context_length` is now consumed at runtime: `ironclaw_loop_host`'s `LlmProviderModelGateway::advertised_context_window_tokens` reads it — only when the returned `ModelMetadata::id` matches the model the request will actually be served by — to derive the per-run prompt context budget. A provider that populates `context_length` therefore changes prompt sizing and compaction thresholds for runs served by that model; a provider that leaves it `None` keeps the compiled-in default. A provider that may serve a call from more than one model (routing, fan-out, failover chains — `SmartRoutingProvider`, `FailoverProvider`) advertises the smallest window among them, and a table lookup returns `None` for a model it does not know — a guessed window is worse than none.
- `model_metadata()` must be a static description of the configured model: no network or credential I/O (no token refresh, no discovery call) and no lock shared with an in-flight call — it is awaited on the turn-run host-build critical path (`ironclaw_turn_runner::loop_driver_host`), so any I/O there stalls or serializes behind that path. Decorators (`TokenRefreshingProvider` and peers) must delegate it unchanged rather than wrapping it with pre-call work they add to the other trait methods.

To add a new provider:
1. Create `crates/domains/ironclaw_llm/src/myprovider.rs` implementing `LlmProvider` <!-- check-guidance: path-ok --> (prescriptive: the file you are about to add, not one that exists)
Expand Down
Loading
Loading