Skip to content

[WRONG BRANCH] release: promote dev to main (v2.32.0) - #2479

Merged
lidge-jun merged 111 commits into
mainfrom
dev
Aug 24, 2026
Merged

lidge-jun merged 111 commits into
mainfrom
dev

Conversation

@lidge-jun

Copy link
Copy Markdown
Owner

Summary

Promote dev (c44e43f) to main for the v2.32.0 stable release.

Verification

Cross-platform CI green on c44e43f; full local suite re-run before release dispatch.

Checklist

  • CI green on promoted head
  • No merge conflicts

lidge-jun and others added 30 commits August 22, 2026 00:16
… diagnostics/GC-relief/smol-worker plans, macmini measurement protocol
…nt thread

The serving-identity record was keyed only on `x-codex-parent-thread-id`.
Without that header there was no scope at all, so the record could never be
written or compared: every turn stayed permanently cold, the deterministic
pre-flight never fired, and each turn fell through to the opaque-blob
recovery — one extra full upload of the transcript, every turn.

Measured on live traffic. Across 95 xAI conversations, 70 recoveries occurred
and 67 of them were in two conversations:

  f4be51de   86 requests  55 recoveries
  c14e85a7   66 requests  12 recoveries
  e925d065  165 requests   1 recovery     <- healthy: one cold first turn

Both outliers are conversations where the backend was switched mid-session, so
their transcripts permanently carry foreign-minted reasoning blobs replayed on
every later turn. An instrumented build showed those requests carrying no
client thread id, which is why the record never warmed up. Those turns were
~150k input tokens each, sent twice.

The recovery was working as designed — without it the turns would fail
outright. The defect is that the deterministic path was structurally
unavailable to them, so the recovery paid full price every turn instead of
once.

`conversationIdFromResponsesRequest` already resolves a conversation identity
for the request log through a four-level fallback, so reuse it as the replay
scope key when the header is absent. `_clientThreadId` is untouched: it
remains the routing and continuation identity, and the header path is
byte-for-byte unchanged.

The scope is shared with the process-local raw-reasoning replay and the
durable thought-signature replay. Widening is safe for both because they key
additionally by provider, destination, adapter, model and credential, so a
conversation namespace only narrows what they already isolate — and a fallback
that yields no identity still produces no scope, preserving today's keep-the-
blobs behaviour.

Pinned by a three-turn headerless regression asserting sendCount [2, 1, 1]:
recover once, then strip pre-flight. That sequence is the entire point.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 22375cf980ee7990f36f6d6c9d231966ff84102a)
The serving-identity record tracks which destination served the previous turn.
That is the right signal for detecting a switch and the wrong one for what
actually costs money, because foreign blobs stay in the client transcript
forever while the switch happens only once.

Measured on the deployed build — three consecutive headerless turns replaying
a grok-minted blob to gpt-5.6-sol:

  turn 1  sends=1  recovery=[]                       pre-flight strips, one send
  turn 2  sends=2  recovery=[opaque-blob-rejection]   record now says sol == sol,
  turn 3  sends=2  recovery=[opaque-blob-rejection]   no strip, upstream rejects

After the first turn commits the new destination every later comparison
returns "same identity", so the pre-flight stops stripping while the
grok-minted blob is still in the replayed history. Each of those turns paid a
full extra upload. This is the production pathology: 86 requests / 55
recoveries and 66 / 12 in the two conversations where the backend was switched
mid-session, against 165 / 1 for a healthy one, at ~150k input tokens a send.

When a recovery succeeds the upstream has just proven this conversation's
replayed opaque state is unusable for that destination. Remember it and
pre-strip instead of rediscovering it once per turn.

The memo is keyed by conversation **and** durable serving identity. Keyed by
conversation alone it would strip the original destination's own valid blobs
the moment the user switched back — a silent, permanent quality regression with
no error to notice. It is recorded only when the blobless resend actually
succeeded, so a resend that also failed teaches nothing.

TTL is five minutes against the serving record's hour, and the asymmetry is
deliberate: a stale memo silently degrades reasoning, while an expired one
costs a single visible recovery round trip that re-establishes it.

An earlier attempt at this test alternated destinations between turns, which
passes for the wrong reason — the identity changes every turn, so the ordinary
switch detection fires and the memo is never exercised. The regression now
holds the destination constant and asserts sendCount [2, 1, 1], plus the
switch-back case, a failed resend recording nothing, and expiry rechecking
once before settling.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit fe8be1ac4d00d855e98393642e1f2e796b21ca2f)
Do not reuse the hashed request-log conversation id. Mixed parent-thread
and session_id headers that carry the same conversation must hit one
serving record, and a shared or synthetic session_id must not coalesce
distinct thread or Cursor conversations.
Restore the architecture note for the conversation-and-serving-identity
memo: five-minute TTL, successful blobless-retry admission, and later
pre-flight stripping.
Verified by an independent read-only review lane plus focused suites on the merged tree: tsc exit 0, 339 pass / 0 fail across the nine suites covering every changed file.
Verified by an independent read-only review lane plus focused suites on the merged tree: tsc exit 0, 339 pass / 0 fail across the nine suites covering every changed file.
Verified by an independent read-only review lane plus focused suites on the merged tree: tsc exit 0, 339 pass / 0 fail across the nine suites covering every changed file.
Verified by an independent read-only review lane plus focused suites on the merged tree: tsc exit 0, 339 pass / 0 fail across the nine suites covering every changed file.
Opens devlog/_plan/260822_backlog_disposition_program/ as the planning unit for
clearing the open PR/issue backlog by explicit per-item disposition.

000  objective, 45-PR inventory captured at unit open, disposition classes, and the
     dependency-ordered wp0-wp9 map
001  baseline verifier evidence actually run at unit open (tool-argument-integers
     24 pass, tsc exit 0), remote host state, and repository authority
002  A-phase audit synthesis: round 1 returned FAIL with 7 blockers, all accepted
     with zero rebuttals, each re-verified against the tree before disposition
003  live drift at the A gate (45 -> 50 open PRs) and the disposition of competitor
     PR #2360, which fixes the same issue as wp3
010  wp1 green-and-ready merges, with per-PR verified change maps
020  wp2 changes-requested rebuilds, including the full 16-PR roster the audit
     found missing
030  wp3 #2316, re-scoped by the audit to a single file after the bare-name alias
     was shown unreachable behind the bridge authorization guard
040  wp4 #2292 Windows picker, with a bounded subprocess seam
050  wp5 #2221 native main token refresh, with external-writer CAS promoted into
     acceptance criteria
060  wp6 #1049, recorded as deferred: it needs a crash-safe publisher phase first
070  wp7 Bun 1.4 memory stack retarget, preserving the recorded FAIL verdicts
080  wp8 conflicting and remaining PR disposition

Docs only: no production file is touched by this commit, and nothing in the build,
typecheck, or test path reads from devlog/. privacy:scan passes.
devlog: backlog disposition program roadmap (work-phase 0, docs-only)
Codex advertises multi_agent_v1__wait_agent's timeout_ms as a JSON Schema
number, but its Rust runtime deserializes the field as u64. Grok serializes
the integer through a float, so a wait of 120000 arrives as 120000.0 and
Codex rejects the call before the tool runs, with an invalid-type error
naming a floating point value where u64 was expected.

The #1611 repair already existed but declined here, because it only fires on
a declared integer. The schema lookup was never the problem: the error text
comes from Codex's own deserializer, which only sees the call after the
bridge emitted it.

Treat a known native u64 field as integer-declared when the schema declares a
numeric type, so the existing re-stringify path emits 120000. A fractional
value is a real disagreement and still fails upstream.

The allowlist is one field wide on purpose. It names only what has a captured
u64 rejection, because the repair is unambiguous only for a field that cannot
hold a fraction; a generic name like start, priority, or port would silently
rewrite a third-party tool's legitimate fractional value. Cursor's sibling
yield_time_ms is also declared number and is deliberately not included: it
gets its own change when it gets its own reproduction.

Array items are judged by their own schema rather than inheriting the key, so
an array named like the allowlist is not rewritten.

Closes #2316
Ingwannu and others added 22 commits August 23, 2026 17:38
fix(sidecars): keep caller cancellations account-neutral
* feat(compatibility): add fixture-backed OpenAI contract manifest

* fix(compatibility): bind manifest to canonical route

* docs: sync compatibility guidance across locales

---------

Co-authored-by: Ingwannu <ingwannu@users.noreply.github.com>
Co-authored-by: Ingwannu <ingwannu@users.noreply.github.com>
* refactor(responses): isolate fetch helper imports

* test(responses): reject dynamic import bypasses

---------

Co-authored-by: Ingwannu <ingwannu@users.noreply.github.com>
Co-authored-by: Ingwannu <ingwannu@users.noreply.github.com>
Co-authored-by: Ingwannu <ingwannu@users.noreply.github.com>
…rminal once (#2449)

* fix(combos): fail over zero-output stream failures

* fix(combos): preserve preflight ownership boundaries

* fix(responses): record streamed terminal outcome once

---------

Co-authored-by: Ingwannu <ingwannu@users.noreply.github.com>
#2448 scoped the wait repair correctly but allowlisted the hyphenated
yield-time_ms name. Live Grok 4.6 Codex Desktop calls still emit
yield_time_ms: 20000.0 and Codex rejects them as u64 before wait runs.
Keep the repair wait-scoped so namespaced Cursor calls stay untouched.
A combined payload re-stringifies ignored number fields, so an isolated
priority:2.0 case is the actual guard that wait did not absorb it.
…liation (#2456)

* devlog: wp11 — combo quota badges opened as PR #2454

* devlog: wp6-wp11 records and the closing reconciliation
…ult preset (#2463 #2464 #2465) (#2466)

* devlog: model/provider UX design unit — aliases, new-models-off, default preset

Design-only unit 260824_model_ux_aliases_and_defaults: 000 research, 010 provider+model aliases, 020 new-models-arrive-off baseline, 030 latest-only default preset. Basis for three feature issues.

* devlog: link filed issues #2463 #2464 #2465
`ocx models` built `reasoningEfforts` from a bare per-model lookup falling back
to the provider-wide list. The catalog (`provider-fetch`) and the effort cap
(`effort-policy`) both go through `configuredReasoningEfforts`, which does three
more things: it returns `[]` for a `noReasoningModels` match, drops levels Codex
does not declare, and re-adds tiers a wire map proves the model emits.

Restating two of its five lines meant the command reported a ladder the proxy
strips and echoed junk as a supported level:

  noReasoningModels: ["model-b"]        ocx models ["low","medium","high"]
                                        runtime   []
  modelReasoningEfforts:
    model-c: ["high","bogus","low"]     ocx models ["high","bogus","low"]
                                        runtime   ["low","high"]

`ocx models` is what an operator reads to check what a config actually did, so a
row that disagrees with the proxy is the one thing it must not print.

This is the sibling of the modality fix in #2086, which routed the three maps
through `modelRecordValue` on the lines above but left this one a partial
re-implementation.
fix(cli): resolve the effort ladder the way the runtime resolves it
…core

fix(tools): coerce wait.yield_time_ms as an integral float
@lidge-jun
lidge-jun requested a review from Ingwannu as a code owner August 24, 2026 09:58
@github-actions

Copy link
Copy Markdown
Contributor

✅ Deterministic PR hygiene checks passed.

@github-actions github-actions Bot changed the title release: promote dev to main (v2.32.0) [WRONG BRANCH] release: promote dev to main (v2.32.0) Aug 24, 2026
@github-actions

github-actions Bot commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

⏳ DRAFT

  • wrong target branch (main); retarget to dev. UI screenshot required.

What to do

  • Retarget this PR to dev — all contributions go to dev.
  • Add a screenshot of the UI change to the PR description.

Its title has been prefixed with [WRONG BRANCH].
Automatic draft conversion failed (token cannot change draft status). Please convert this pull request to a draft manually. The required enforce-target check will keep failing until every issue above is resolved.

@github-actions
github-actions Bot marked this pull request as draft August 24, 2026 09:59
@lidge-jun
lidge-jun marked this pull request as ready for review August 24, 2026 09:59
@lidge-jun
lidge-jun merged commit 8bc7aa8 into main Aug 24, 2026
68 of 72 checks passed
agentHits pushed a commit to agentHits/opencodex that referenced this pull request Sep 17, 2026
[WRONG BRANCH] release: promote dev to main (v2.32.0)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants