Align Codex cache affinity header - #57012
Draft
ildunari wants to merge 1 commit into
Draft
Conversation
Contributor
|
Thanks for the focused cache-affinity fix. Current main still derives Automated hermes-sweeper review. |
JoaoMarcos44
added a commit
to JoaoMarcos44/hermes-agent
that referenced
this pull request
Aug 5, 2026
…des, pin Grok cron affinity Fixes NousResearch#79013, NousResearch#79014, NousResearch#79015. - session_id header now carries the raw physical session id (NousResearch#57012 contract); x-client-request-id mirrors the body's effective prompt_cache_key instead of both diverging to a bare scope string. - extra_body.prompt_cache_key for xAI Responses now reads back a caller's top-level request_overrides={"prompt_cache_key": ...} instead of always using the auto-derived hash, so an explicit override actually governs the field xAI reads. - x-grok-conv-id (native xAI Responses transport and Grok-via-OpenRouter profile) is now scoped through _cache_scope_from_session_id(), so cron re-fires of the same job pin to the same backend instead of a new one every fire. - Fallback cache_key (when instructions/tools are empty) now falls back to the normalized scope instead of the raw session_id, same class of fix.
kshitijk4poor
added a commit
that referenced
this pull request
Aug 15, 2026
…ation Legacy compaction mode (compression.in_place: false) rotates the physical session_id mid-conversation. The prompt-cache scope introduced in #79161 was derived from that physical id, so every rotation moved the same conversation into a fresh cache bucket - the prompt cache went cold at every rotation boundary (#79017). Fix: resolve a rotation-stable logical scope - the compression-lineage ROOT of the current session (SessionDB.get_compression_lineage, fork-aware post-#79193) - once per turn, memoized per transcript segment, and prefer it over the physical session_id at every prompt_cache_key derivation site: - agent/prompt_cache_scope.py (new): resolve_prompt_cache_scope(agent) - lineage-root walk with per-segment memo; falls back to the physical id when no DB is attached or the walk fails, degrading to pre-fix behavior. - transports/codex.py: build_kwargs accepts cache_scope_id and prefers it for the body prompt_cache_key, the xAI x-grok-conv-id header, and the Codex x-client-request-id routing header. The Codex session_id header keeps the raw physical id (transcript identity, #57012 contract). - transports/chat_completions.py: _add_prompt_cache_key accepts cache_scope_id with the same precedence. - chat_completion_helpers.py: build_api_kwargs threads the resolved scope into all three build_kwargs call sites (codex, profile, legacy). - auxiliary_client.py: set_runtime_main carries cache_scope; the aux Responses cache-key site prefers it over the physical session_id. - turn_context.py: resolves the scope once per turn and threads it through set_runtime_main (no DB walk on the per-API-call hot path). Scope semantics preserved from #79161: /new starts a fresh scope (new lineage), /branch children, delegate subagents, and tool children stay isolated (explicit-fork exclusion in get_compression_lineage), unrelated sessions keep distinct buckets, and cron per-fire timestamps still normalize via _cache_scope_from_session_id. Default installs compact in place (session_id never rotates), so they hit the memo and produce byte-identical keys to before. Fixes #79017
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fix Codex prompt-cache affinity by aligning the cache-routing
x-client-request-idheader with the finalprompt_cache_keyfor the ChatGPT/Codex backend.Hermes already builds a stable, content-addressed
prompt_cache_keyfrom the static prefix (instructions + tools). That keeps recurring jobs and fresh Hermes sessions from going cache-cold when the visible Hermessession_idchanges. But the Codex backend path still sentx-client-request-idas the raw Hermes session id, so the body cache key and request-affinity header could point at different cache namespaces.This keeps the raw
session_idheader unchanged for transcript/session identity, and only changesx-client-request-idto follow the finalprompt_cache_key.Fixes #47126.
What changed
prompt_cache_keybehavior.session_idheader as the real Hermes session id.x-client-request-idto the finalprompt_cache_keyvalue.request_overrides["prompt_cache_key"]by making the header follow that final overridden key.Validation
Ran against a clean branch from
origin/main:Result:
169 passed.Also verified locally in an isolated disposable
HERMES_HOMEsandbox that resumedopenai-codex / gpt-5.5turns still reach high per-call cache reuse after warmup, while this patch keeps the existing content-addressed key behavior needed for timestamped recurring jobs.