Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions docs-site/src/content/docs/reference/proxy-formats.md
Original file line number Diff line number Diff line change
Expand Up @@ -688,3 +688,17 @@ that repair, it becomes a normal user message. If a current v2 task remains genu
but the selected routed target cannot read native ChatGPT ciphertext, opencodex fails with
`unreadable_encrypted_agent_task` instead of sending unreadable bytes to that provider. See
[Sub-agent Surface](/guides/sub-agent-surface/) for the client behavior around worker tasks.

History is handled too, and differently, because losing a replayed message should not end a
conversation. A replayed `agent_message` that mixes readable text with backend ciphertext cannot
be lowered to a public message, so a routed Responses destination would otherwise receive the
ciphertext along with an item type only the ChatGPT backend declares. Before dispatch, opencodex
replaces that ciphertext with `[encrypted content omitted]` — the same marker it already
substitutes after an upstream decrypt failure — which leaves the item lowerable and the readable
text intact. The provider never sees the ciphertext or the private item, and the conversation
continues. Combo targets are repaired individually, since each receives its own copy of the
request. The canonical ChatGPT Codex backend is exempt because it is the destination that minted
and can read those bytes; a `forward` provider pointed at any other origin is not exempt.
Explicitly trusted `allowEncryptedV2AgentTasks` routes and translated Chat or Anthropic wires are
unaffected, as are other item types such as reasoning and tool-output blobs, which keep their
existing decrypt-failure recovery.
2 changes: 1 addition & 1 deletion src/server/responses.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ import { requestPacingOverloadResponse } from "./responses/pacing-overload";

export { buildToolBridgeMaps, isV1CollabSurface, collabSurface, multiAgentGuidanceText, V2_GUIDANCE_CHAR_BUDGET, injectDeveloperMessage } from "./responses/collaboration";
export type { MultiAgentGuidanceOptions, MultiAgentGuidanceDeps } from "./responses/collaboration";
export { hasUnreadableEncryptedAgentTask, sanitizeEncryptedContentInPlace } from "./responses/encrypted-payload";
export { hasUnreadableEncryptedAgentTask, sanitizeEncryptedContentInPlace, stripAgentMessageCiphertextInPlace } from "./responses/encrypted-payload";
export { COMPACT_RESPONSE_MAX_BYTES, bufferCompactResponse } from "./responses/compact";
export { disableResponsesRequestTimeout, safeHostLabel, fetchWithHeaderTimeout } from "./responses/fetch-helpers";
export { sidecarOutcomeRecorder, isShadowSourceModel, codexLogAccountId, usesCodexForwardPoolAuth, codexForwardTerminalOutcomeRecorder, decodeRequestErrorResponse, buildComboChildHeaders, linkAbortSignal } from "./responses/core";
Expand Down
45 changes: 44 additions & 1 deletion src/server/responses/core.ts
Original file line number Diff line number Diff line change
Expand Up @@ -394,7 +394,7 @@ import type { EffectiveSubagentRoster, SpawnAgentSurface } from "../../codex/cat

import { buildToolBridgeMaps, collabSurface, injectDeveloperMessage, multiAgentGuidanceText } from "./collaboration";
import { mapCodexAuthContextErrorToResponse, nativeMainRefreshFailureResponse } from "./codex-auth-error";
import { hasUnreadableEncryptedAgentTask, looksLikeBackendCiphertext, sanitizeEncryptedContentInPlace } from "./encrypted-payload";
import { hasUnreadableEncryptedAgentTask, looksLikeBackendCiphertext, sanitizeEncryptedContentInPlace, stripAgentMessageCiphertextInPlace } from "./encrypted-payload";
import { fetchWithHeaderTimeout, providerFetch, safeHostLabel, safeOriginLabel, storedPoolReplayDispatchNotifier, type ProviderFetchOptions } from "./fetch-helpers";
import { classifyTransportFailureKind, transportErrorCode } from "../../lib/upstream-reachability";
import {
Expand Down Expand Up @@ -4024,6 +4024,49 @@ async function handleResponsesInner(
return unreadableEncryptedAgentTaskResponse(recoveryFailureReason);
}

// The guard above asks whether the CURRENT worker task is readable, and it only inspects the
// tail item. An `agent_message` that mixes readable text with backend ciphertext answers
// "readable" to that question at every position, so it passed -- and then
// `normalizeRoutedAgentMessages` refused to lower it, because lowering requires every part to
// be representable. The raw Responses passthrough serialized the private item as it stood, so
// backend ciphertext and an item type only the Codex backend declares reached a third-party
// provider, which answered `422 unknown item type "agent_message"` (#4454).
//
// The opaque-blob path already knows the repair: replace the undecryptable part with an
// omission marker, which leaves the item lowerable. It applied that repair only AFTER an
// upstream rejection. For a destination that cannot accept the private item under any
// circumstances, that round trip was never going to succeed and sent the ciphertext to find
// out, so do the repair here instead. Recovery above has already had its chance to turn the
// same bytes into real plaintext; only what it could not rescue reaches this.
if (inboundWire === "responses" && !finalRouteCanPassThroughEncryptedTask) {
// Only the raw Responses passthrough puts input items on the wire verbatim, so that is the
// only wire this has to repair: translated wires rebuild the body from parsed messages, where
// `inputContentParts` drops an encrypted part instead of forwarding it. The exemption is the
// canonical Codex backend alone, because it is the one destination that minted these bytes and
// can read them. `authMode: "forward"` is NOT that test -- a noncanonical forward gateway is
// somebody else's server that happens to be configured for passthrough, and it receives the
// ciphertext like any other third party.
//
// Combo children run this too. Each child carries its own `structuredClone` of the body
// (`concreteComboRequestBody`) and its own concrete route, so a sibling's repair is invisible
// here and a target that resolves to a routed Responses wire would otherwise send the
// ciphertext that the parent's own dispatch no longer does.
const wireProvider = resolveWireProtocolOverride(
route.providerName,
route.modelId,
route.provider,
inboundWire,
);
if (wireProvider.adapter === "openai-responses" && !isCanonicalOpenAiForwardProvider(wireProvider)) {
const repaired = stripAgentMessageCiphertextInPlace((body as { input?: unknown } | undefined)?.input);
if (repaired > 0) {
console.warn(
`[opencodex] replaced ciphertext in ${repaired} replayed agent message(s) with an omission marker; the selected provider cannot read native ChatGPT ciphertext`,
);
}
}
}

// The canonical ChatGPT backend rejects previous_response_id, so a local replay miss leaves no
// safe way to recover the omitted history. Fail before auth, adapter construction, or upstream
// I/O instead of stripping the id and silently forwarding a context-free delta (#702).
Expand Down
161 changes: 161 additions & 0 deletions src/server/responses/encrypted-payload.ts
Original file line number Diff line number Diff line change
Expand Up @@ -319,6 +319,167 @@ export function hasEncryptedContentPart(content: unknown): boolean {
}


/** The marker `prepareOpaqueBlobRecovery` already substitutes for an undecryptable part. */
export const OMITTED_ENCRYPTED_CONTENT_TEXT = "[encrypted content omitted]";

/**
* The Fernet WIRE shape, without validating the body: version prefix, base64url alphabet, and a
* canonical encoded length. Free text is judged by this rather than by
* `looksLikeBackendCiphertext`, which is length >= 64 over a character class that a SHA-256 hex
* digest matches exactly at 64 characters -- as do a SHA-512 digest, a long key, and adjacent
* short encoded fragments. An `encrypted_content` slot carries ciphertext by definition and is
* stripped whatever it holds; a text part does not, and replacing a digest a child deliberately
* printed would destroy readable content to protect bytes that were never secret.
*/
const FERNET_SHAPED = /^g[A-Za-z0-9_-]+={0,2}$/;

function looksLikeFernetToken(text: string): boolean {
return text.length >= 100 && text.length % 4 === 0 && FERNET_SHAPED.test(text);
}

function textWithRunsOmitted(payload: string, runs: readonly FernetTokenRun[]): string {
let last = 0;
let out = "";
for (const run of runs) {
out += payload.slice(last, run.index) + OMITTED_ENCRYPTED_CONTENT_TEXT;
last = run.index + run.token.length;
}
return out + payload.slice(last);
}

/**
* Replace ciphertext inside `agent_message` items with an omission marker, so
* `normalizeRoutedAgentMessages` can lower them onto public messages. Returns how many items
* were repaired. Items are replaced rather than mutated, and every other item type is left
* alone: reasoning and function-output blobs keep their own reactive recovery.
*
* `hasUnreadableEncryptedAgentTask` above answers a different question: can the CURRENT worker
* task be read at all? It inspects only the tail item and reports false the moment any plaintext
* survives the envelope. `normalizeRoutedAgentMessages` asks the opposite question -- is EVERY
* part lowerable? -- and forwards the private item verbatim when one is not. A mixed
* `input_text` + `encrypted_content` item answers "readable" to the first and "not lowerable"
* to the second, so it fell between them: the guard never fired, the adapter refused to lower it,
* and the raw Responses passthrough put a private item and backend ciphertext on the wire
* (#4454). Position was never the discriminator -- a replayed child result lands mid-history and
* the tail-only scan cannot see it -- but the tail is equally exposed when it is mixed.
*
* This is the repair `prepareOpaqueBlobRecovery` performs after an upstream rejection, applied
* before dispatch for a destination that cannot accept the private item under any circumstances.
* The round trip it replaces was never going to succeed, and it sent ciphertext to a third party
* to find that out. Nothing is decrypted, and nothing readable is lost: the parent could not read
* these bytes either.
*
* The two kinds of slot are judged differently, because they carry different guarantees. An
* `encrypted_content` slot holds ciphertext by definition, so it is stripped whatever it holds:
* demanding a well-formed token there would reopen this defect one payload later, since a
* truncated token, a standard-base64 blob carrying `+` or `/`, an unexpected version byte, or a
* run past the recovery size limits would each keep the item and forward the bytes.
*
* A text part carries no such guarantee, so it is matched strictly: embedded runs that validate
* as Fernet, or a whole slot with the Fernet wire shape. A loose character-class test would be
* worse than the defect for that half -- a SHA-256 digest is exactly 64 characters of
* `[A-Za-z0-9]` and would be replaced with a marker, silently deleting something a child
* deliberately printed.
*/
export function stripAgentMessageCiphertextInPlace(input: unknown): number {
if (!Array.isArray(input)) return 0;
let repaired = 0;
for (let index = 0; index < input.length; index += 1) {
const item = input[index];
if (!item || typeof item !== "object" || Array.isArray(item)) continue;
const record = item as Record<string, unknown>;
if (record.type !== "agent_message") continue;
const content = record.content;
if (typeof content === "string") {
const replaced = textWithoutCiphertext(content);
if (replaced === content) continue;
input[index] = { ...record, content: replaced };
repaired += 1;
continue;
}
if (!Array.isArray(content)) continue;
const parts = contentWithoutCiphertext(content);
if (parts === content) continue;
input[index] = { ...record, content: parts };
repaired += 1;
}
return repaired;
}

/** Free text: drop embedded token runs, and replace a slot that is nothing but a token. */
function textWithoutCiphertext(text: string): string {
const runs = fernetTokenRuns(text);
if (runs.length > 0) return textWithRunsOmitted(text, runs);
return looksLikeFernetToken(text.trim()) ? OMITTED_ENCRYPTED_CONTENT_TEXT : text;
}

function ciphertextTextOfPart(part: unknown): string | undefined {
if (!part || typeof part !== "object") return undefined;
const record = part as { type?: unknown; text?: unknown };
return (record.type === "input_text" || record.type === "text") && typeof record.text === "string"
? record.text
: undefined;
}

/**
* Adjacent text slots that are one token between them. Each fragment can be too short to judge on
* its own, which is the text-side twin of the split `encrypted_content` run. The join must still
* be Fernet-shaped, so two ordinary encoded fragments do not become a marker by being adjacent.
*/
function joinedCiphertextTextParts(content: readonly unknown[]): Set<object> {
const flagged = new Set<object>();
let run: Array<{ part: object; text: string }> = [];
const finish = (): void => {
if (run.length > 1 && looksLikeFernetToken(run.map(entry => entry.text).join(""))) {
for (const entry of run) flagged.add(entry.part);
}
run = [];
};
for (const part of content) {
const text = ciphertextTextOfPart(part);
if (text === undefined || text.trim().length === 0 || !/^[A-Za-z0-9_-]+={0,2}$/.test(text)) {
finish();
continue;
}
run.push({ part: part as object, text });
}
finish();
return flagged;
}

function contentWithoutCiphertext(content: unknown[]): unknown[] {
let changed = false;
const joined = joinedCiphertextTextParts(content);
const parts = content.map((part: unknown) => {
if (!part || typeof part !== "object") return part;
if (joined.has(part)) {
changed = true;
return { type: "input_text", text: OMITTED_ENCRYPTED_CONTENT_TEXT };
}
const record = part as { type?: unknown; text?: unknown; encrypted_content?: unknown };
if (record.type === "encrypted_content" && typeof record.encrypted_content === "string") {
changed = true;
// Keep whatever plaintext a recognizable slot carries around its token; a slot this
// cannot parse is replaced whole rather than forwarded on the chance that it is benign.
const runs = fernetTokenRuns(record.encrypted_content);
return {
type: "input_text",
text: runs.length > 0
? textWithRunsOmitted(record.encrypted_content, runs)
: OMITTED_ENCRYPTED_CONTENT_TEXT,
};
}
const text = ciphertextTextOfPart(part);
if (text === undefined) return part;
const replaced = textWithoutCiphertext(text);
if (replaced === text) return part;
changed = true;
return { ...record, text: replaced };
});
return changed ? parts : content;
}



export function sanitizeEncryptedContentInPlace(input: unknown): number {
if (!Array.isArray(input)) return 0;
Expand Down
60 changes: 60 additions & 0 deletions structure/subagents.md
Original file line number Diff line number Diff line change
Expand Up @@ -148,6 +148,66 @@ unreadable split-token shapes. The sanitizer preserves just those fragment objec
normalizing independent plaintext slots. Detection never authorizes reconstruction or recovery;
other fragment layouts and mixed readable content retain their documented residual boundaries.

## Routed agent-message ciphertext egress

Two questions about an `agent_message` were asked in two places, and the gap between them was
open. `hasUnreadableEncryptedAgentTask` asks whether the CURRENT worker task can be read and
inspects only the tail item; `normalizeRoutedAgentMessages` asks whether EVERY part can be lowered
onto a public message and forwards the private item verbatim when one cannot. An item mixing
`input_text` with `encrypted_content` is readable by the first measure and unlowerable by the
second, so it passed the guard, kept its private type through the raw Responses passthrough, and
left the process as backend ciphertext plus an item type only the Codex backend declares. The
destination answered `422 unknown item type "agent_message"` after the bytes were already sent.
Position was incidental: a replayed child result sits mid-history, where a tail-only scan cannot
see it, and the tail is exposed the same way once it is mixed.

The repair already existed reactively. `prepareOpaqueBlobRecovery` replaces an undecryptable part
with `[encrypted content omitted]`, which leaves the item lowerable, and it ran after an upstream
rejection. A destination that cannot accept the private item under any circumstances was never
going to answer that request, so the round trip only served to send the ciphertext.
`stripAgentMessageCiphertextInPlace` in `src/server/responses/encrypted-payload.ts` applies the
same repair before dispatch, and `src/server/responses/core.ts` runs it against the final route,
after `expandPreviousResponseInput`, after the sanitizer has rewritten plaintext parked in
encrypted slots, and after encrypted-task recovery has had its chance to produce real plaintext
instead of a marker.

The two kinds of slot are judged differently, because they carry different guarantees. An
`encrypted_content` slot holds ciphertext by definition, so it is stripped whatever it holds:
demanding a well-formed token there would reopen the same defect one payload later, since a
truncated token, a standard-base64 blob carrying `+` or `/`, an unexpected version byte, or a run
past the recovery size limits would each keep the item and forward the bytes. A text part carries
no such guarantee, so it is matched strictly -- embedded runs that validate as Fernet, or a whole
slot with the Fernet wire shape, which is the version prefix, the base64url alphabet and a
canonical length of at least 100 divisible by four. Adjacent text fragments are joined before that
test, so a token split across slots is still caught. `looksLikeBackendCiphertext` is deliberately
NOT used on text: it is length >= 64 over a character class that a SHA-256 digest matches exactly
at 64 characters, and replacing a digest a child deliberately printed would delete readable content
to protect bytes that were never secret. Other item types are untouched: reasoning and
function-output blobs keep the reactive opaque-blob recovery, which still rescues a destination
that merely failed to decrypt something it was entitled to read, and which stays reachable for the
canonical backend and for explicitly trusted routes.

The repair resolves the same wire override the adapter is built from rather than restating routing
policy, and runs for `openai-responses` whenever the destination is not the canonical Codex
backend. `authMode: "forward"` is deliberately not that test: it describes how this proxy treats
credentials, not who answers, and a forward-configured gateway at another origin receives the
ciphertext like any third party. Only `isCanonicalOpenAiForwardProvider` is exempt, because it
alone minted these bytes and can read them. The wire override matters for the reported destination,
where the provider row names the Chat wire and a registry model default moves the model onto
Responses. Translated wires are untouched because `inputContentParts` drops an encrypted part
instead of forwarding it, and `canPassThroughEncryptedV2AgentTask` keeps an explicitly trusted
route exempt. Combo children run the repair themselves: `concreteComboRequestBody` gives each
target its own `structuredClone` and its own concrete route, so a sibling's repair is invisible to
them and a target resolving to a routed Responses wire would otherwise send what the parent's own
dispatch no longer does.

Nothing here decrypts, and the tail NEW_TASK envelope keeps `unreadable_encrypted_agent_task` and
its opt-in recovery unchanged: an unreadable current task still fails closed rather than reaching a
child with a marker where its assignment should be. An `agent_message` carrying unknown parts but
no ciphertext still reaches the wire unchanged and still draws the destination's own 422, which is
a compatibility gap rather than an egress one. Covered by
`tests/server/v2-agent-message-failfast.test.ts`.

## Subagents

New non-OAuth provider registrations carry `initialModelSelection` with a unique
Expand Down
Loading
Loading