Skip to content
This repository was archived by the owner on Aug 25, 2026. It is now read-only.

feat: add native OpenAI Responses compaction - #3

Merged
YaseenHQ merged 6 commits into
mainfrom
codex/pi-server-compaction
Jul 25, 2026
Merged

YaseenHQ merged 6 commits into
mainfrom
codex/pi-server-compaction

Conversation

@YaseenHQ

@YaseenHQ YaseenHQ commented Jul 23, 2026 •

Copy link
Copy Markdown
Owner

Related Issue

None. This PR adds native OpenAI Responses compaction while retaining Kimi existing local compaction as the portable fallback.

Problem

Kimi previously owned compaction entirely on the client. Responses-compatible providers can instead return an opaque compaction item from /responses/compact, which must be persisted and replayed on later requests. Dropping it, rendering it as normal content, or forwarding it to another model or endpoint is incorrect.

What changed

  • Add an optional provider-native compaction capability to current Kosong and agent-core v2.
  • Automatically try /responses/compact for Responses providers during normal full compaction.
  • Preserve local summary compaction for custom instructions, unsupported endpoints, and ordinary provider failures.
  • Persist provider-owned replacement windows atomically across context replay and resume.
  • Record provider-reported compaction usage and lifecycle telemetry.
  • Stamp opaque compaction state with its source model and endpoint, replaying it only to that same source.
  • Suppress repeated attempts after known unsupported responses while still surfacing authentication failures.
  • Keep TUI, VS Code, and daemon transcripts human-readable by hiding opaque state from presentation adapters.
  • Preserve upstream reasoning-field dialect detection while skipping opaque placeholders on non-Responses wires.
  • Document the automatic behavior in the English and Chinese session guides.

The branch is synced with current main after the upstream 0.29.1 integration. Conflict resolution combines both required behaviors: endpoint-specific reasoning-key selection and opaque-message filtering.

Validation

  • full monorepo package build/typecheck completed through the Kimi CLI; integration then exposed and fixed exhaustive opaque-part handling in the CLI and VS Code adapters
  • Kimi CLI, VS Code, Kimi Web, vis server, and vis web typechecks pass after those fixes
  • focused provider and compaction suite: 370 passed, 1 skipped
  • Kimi TUI message-flow suite: 163 passed
  • current synced main full suites: Kimi CLI 2,411 passed; agent-core-v2 4,084 passed; kap-server 797 passed; pi-tui 745 passed
  • git diff --check

Checklist

  • I have read the CONTRIBUTING document.
  • I have linked a related issue, or explained the problem above.
  • I have added tests that prove the feature and integration behavior.
  • Ran the gen-changesets skill; the feature changeset remains in the branch.
  • Updated relevant documentation.

Summary by CodeRabbit

  • New Features
    • Added native context compaction for compatible OpenAI Responses providers, preserving opaque replacement state and replaying the provider’s replacement window.
    • Introduced provider “compact” support with safer remote compaction attempts (telemetry/usage tracking) and automatic fallback to local summarization when unavailable.
  • Documentation
    • Updated session guides (English/Chinese) to explain native context compression behavior and expectations.
  • Bug Fixes
    • Prevented compaction-only/opaque content from being emitted or forwarded as normal chat content, including improved truncation and token counting behavior.
  • Tests
    • Expanded coverage for compaction replay and OpenAI Responses compact/streaming behavior.

@coderabbitai

coderabbitai Bot commented Jul 23, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: d9f1a32d-a975-486d-bffc-49728c5fe43c

📥 Commits

Reviewing files that changed from the base of the PR and between 68b24b3 and 93f77be.

📒 Files selected for processing (12)
  • apps/kimi-code/src/tui/utils/message-replay.ts
  • apps/vscode/src/runtime/replay-adapter.ts
  • apps/vscode/src/utils/session-context.ts
  • packages/agent-core-v2/docs/wire-manifest.d.ts
  • packages/agent-core-v2/src/kosong/provider/bases/openai/openai-legacy.ts
  • packages/agent-core-v2/test/kosong/provider/composition.test.ts
  • packages/agent-core/src/agent/context/index.ts
  • packages/agent-core/src/agent/index.ts
  • packages/kosong/src/providers/kimi.ts
  • packages/kosong/src/providers/openai-legacy.ts
  • packages/kosong/test/kimi.test.ts
  • packages/kosong/test/openai-legacy.test.ts

📝 Walkthrough

Walkthrough

Adds provider-native OpenAI Responses compaction with opaque state preservation, replacement-window context installation, fallback to local summarization, cross-provider filtering, and tests and documentation.

Changes

OpenAI Compaction State

Layer / File(s) Summary
Opaque content and provider contracts
packages/{kosong,agent-core-v2}/src/..., packages/agent-core-v2/docs/...
Adds openai_compaction, opaque assistant detection, provider compaction contracts, requester methods, token estimation, and wire payload support.
OpenAI Responses wire flow
packages/{kosong,agent-core-v2}/src/kosong/provider/...
Encodes and decodes compaction items, validates compact responses, tracks compaction sources, and exposes native compaction calls.
Replacement-window orchestration
packages/{agent-core,agent-core-v2}/src/agent/...
Persists replacement messages, installs provider-owned context windows, records usage, and falls back to local summarization.
Cross-provider projection behavior
packages/{kosong,agent-core,agent-core-v2}/src/...
Filters opaque assistant shells and omits compaction parts from incompatible protocol projections.
Validation and release support
packages/*/test/..., apps/..., docs/{en,zh}/guides/sessions.md, .changeset/*
Adds context, requester, provider, streaming, transcript, replay, and type-safety tests plus release metadata and documentation.

Estimated code review effort: 5 (Critical) | ~120 minutes

Sequence Diagram(s)

sequenceDiagram
  participant FullCompaction
  participant AgentLLMRequesterService
  participant OpenAIResponsesChatProvider
  participant ContextMemory
  FullCompaction->>AgentLLMRequesterService: request provider compaction
  AgentLLMRequesterService->>OpenAIResponsesChatProvider: call compact with history
  OpenAIResponsesChatProvider-->>AgentLLMRequesterService: replacement messages and usage
  AgentLLMRequesterService-->>FullCompaction: ProviderCompactionResult
  FullCompaction->>ContextMemory: apply replacement window
  ContextMemory-->>FullCompaction: compaction complete
Loading

Suggested reviewers: wbxl2000

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 26.19% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title is concise and accurately summarizes the main change: native OpenAI Responses compaction.
Description check ✅ Passed The description follows the template well, with related issue, problem, what changed, validation, and checklist sections filled in.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch codex/pi-server-compaction

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@YaseenHQ YaseenHQ changed the title feat: preserve OpenAI Responses compaction state feat: add native OpenAI Responses compaction Jul 23, 2026
@YaseenHQ
YaseenHQ marked this pull request as ready for review July 23, 2026 14:10

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
packages/agent-core/src/agent/index.ts (1)

323-346: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Native compact requests bypass request logging/recording.

compactProvider doesn't call llmRequestLogger.logRequest/llmRequestRecorder.record the way generate's run closure does, so native compaction requests won't appear in the same request log/record trail as ordinary generate calls (including the local-summarizer fallback, which does log via the normal generate path). Telemetry usage is still captured separately, so this is an observability nice-to-have rather than a functional gap.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/agent-core/src/agent/index.ts` around lines 323 - 346, Update
compactProvider to wrap native provider.compact calls with the same
llmRequestLogger.logRequest and llmRequestRecorder.record flow used by
generate’s run closure, including both authenticated and unauthenticated paths.
Preserve the existing auth resolution and compaction behavior while ensuring
every native compaction request is logged and recorded.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/agent-core/src/agent/compaction/full.ts`:
- Around line 454-479: The remote compaction path in the visible branch of full
compaction must preserve real user messages appended during the provider call.
In the `remote` handling around `applyCompaction`, append the live history tail
after `originalHistory.length` to `remote.messages` when constructing
`replacementMessages`; apply the same change to the corresponding v2
remote-compaction branch, while retaining the existing validation and
cancellation behavior.

In `@packages/agent-core/src/agent/context/index.ts`:
- Around line 315-347: The replacementMessages branch in the compaction flow
must add a compaction boundary marker compatible with undo() before or alongside
the retained provider messages. Ensure native/remote compaction creates a
compaction_summary-origin checkpoint so undo() stops at the boundary and does
not remove both the proxy checkpoint and retained user prompt; preserve the
existing result bookkeeping and history replacement behavior.

---

Nitpick comments:
In `@packages/agent-core/src/agent/index.ts`:
- Around line 323-346: Update compactProvider to wrap native provider.compact
calls with the same llmRequestLogger.logRequest and llmRequestRecorder.record
flow used by generate’s run closure, including both authenticated and
unauthenticated paths. Preserve the existing auth resolution and compaction
behavior while ensuring every native compaction request is logged and recorded.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 5e824882-4791-4e16-84f1-14698093fa6b

📥 Commits

Reviewing files that changed from the base of the PR and between 02dd07e and a158738.

📒 Files selected for processing (48)
  • .changeset/preserve-openai-compaction-state.md
  • docs/en/guides/sessions.md
  • docs/zh/guides/sessions.md
  • packages/agent-core-v2/docs/wire-manifest.d.ts
  • packages/agent-core-v2/src/agent/contextMemory/compactionHandoff.ts
  • packages/agent-core-v2/src/agent/contextMemory/contextMemory.ts
  • packages/agent-core-v2/src/agent/contextMemory/contextMemoryService.ts
  • packages/agent-core-v2/src/agent/contextMemory/contextOps.ts
  • packages/agent-core-v2/src/agent/contextMemory/messageProjection.ts
  • packages/agent-core-v2/src/agent/fullCompaction/fullCompactionService.ts
  • packages/agent-core-v2/src/agent/llmRequester/llmRequester.ts
  • packages/agent-core-v2/src/agent/llmRequester/llmRequesterService.ts
  • packages/agent-core-v2/src/agent/loop/loopService.ts
  • packages/agent-core-v2/src/agent/mcp/output.ts
  • packages/agent-core-v2/src/kosong/contract/message.ts
  • packages/agent-core-v2/src/kosong/contract/provider.ts
  • packages/agent-core-v2/src/kosong/contract/tokens.ts
  • packages/agent-core-v2/src/kosong/model/modelRequester.ts
  • packages/agent-core-v2/src/kosong/model/modelRequesterImpl.ts
  • packages/agent-core-v2/src/kosong/provider/bases/google-genai/google-genai.ts
  • packages/agent-core-v2/src/kosong/provider/bases/openai/openai-common.ts
  • packages/agent-core-v2/src/kosong/provider/bases/openai/openai-legacy.ts
  • packages/agent-core-v2/src/kosong/provider/bases/openai/openai-responses.ts
  • packages/agent-core-v2/test/agent/contextMemory/context.test.ts
  • packages/agent-core-v2/test/agent/fullCompaction/fullCompaction.test.ts
  • packages/agent-core-v2/test/agent/llmRequester/llmRequesterService.test.ts
  • packages/agent-core-v2/test/kosong/provider/composition.test.ts
  • packages/agent-core/src/agent/compaction/full.ts
  • packages/agent-core/src/agent/compaction/types.ts
  • packages/agent-core/src/agent/context/index.ts
  • packages/agent-core/src/agent/index.ts
  • packages/agent-core/src/mcp/output.ts
  • packages/agent-core/src/services/message/message.ts
  • packages/agent-core/src/services/message/transcript.ts
  • packages/agent-core/src/utils/tokens.ts
  • packages/agent-core/test/agent/context.test.ts
  • packages/agent-core/test/services/message-transcript.test.ts
  • packages/kosong/src/message.ts
  • packages/kosong/src/provider.ts
  • packages/kosong/src/providers/google-genai.ts
  • packages/kosong/src/providers/kimi.ts
  • packages/kosong/src/providers/openai-common.ts
  • packages/kosong/src/providers/openai-legacy.ts
  • packages/kosong/src/providers/openai-responses.ts
  • packages/kosong/test/kimi.test.ts
  • packages/kosong/test/openai-legacy.test.ts
  • packages/kosong/test/openai-responses.test.ts
  • packages/kosong/test/type-safety.test.ts

Comment on lines +454 to +479
if (remote !== undefined) {
const newHistory = this.agent.context.history;
for (let i = 0; i < originalHistory.length; i++) {
if (newHistory[i] !== originalHistory[i]) {
this.cancel();
return undefined;
}
}
if (
newHistory
.slice(originalHistory.length)
.some((message) => !isRealUserInput(message))
) {
this.cancel();
return undefined;
}
if (remote.usage !== undefined) {
this.agent.usage.record(model, remote.usage);
}
const result = this.agent.context.applyCompaction({
summary: REMOTE_COMPACTION_SUMMARY,
contextSummary: REMOTE_COMPACTION_SUMMARY,
replacementMessages: remote.messages,
compactedCount: originalHistory.length,
tokensBefore,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🔴 Critical | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
rg -n -B5 -A20 "replacementMessages: remote.messages" packages/agent-core-v2/src/agent/fullCompaction/fullCompactionService.ts

Repository: YaseenHQ/kimi

Length of output: 1380


🏁 Script executed:

#!/bin/bash
set -eu

echo "## locate files"
fd -a 'full.ts|fullCompactionService.ts|context/index.ts|interface.ts' packages | sed 's#^\./##'

echo
echo "## inspect full.ts relevant sections"
sed -n '400,540p' packages/agent-core/src/agent/compaction/full.ts
echo
echo "## inspect fullCompactionService.ts remote section"
sed -n '520,590p' packages/agent-core-v2/src/agent/fullCompaction/fullCompactionService.py 2>/dev/null || sed -n '520,590p' packages/agent-core-v2/src/agent/fullCompaction/fullCompactionService.ts
echo
echo "## inspect applyCompaction implementation candidates"
fd -a 'context/index.ts|context.ts|interface.ts' packages/agent-core packages/agent-core-v2 | while read -r f; do
  echo "--- $f"
  rg -n -B5 -A40 "applyCompaction|replacementMessages|_history" "$f"
done
echo
echo "## behavioral probe of replacementMessages semantics from source"
python3 - <<'PY'
import re, pathlib
for root in ["packages/agent-core","packages/agent-core-v2"]:
    files=list(pathlib.Path(root).rglob("*"))
    for f in files:
        if f.is_file() and "context" in str(f).split("/")[-2:][-1].lower():
            p=root+"/"+str(f).replace(str(root+"/"),"")
            text=f.read_text(errors="ignore")
            if "applyCompaction" in text and "replacementMessages" in text:
                print(f"--- {p}")
                for line_no, line in enumerate(text.splitlines(),1):
                    if "applyCompaction" in line or "replacementMessages" in line or "_history" in line:
                        print(f"{line_no}: {line}")
PY

Repository: YaseenHQ/kimi

Length of output: 9668


🏁 Script executed:

#!/bin/bash
set -eu

echo "## v2 applyCompaction definitions/usages"
rg -n -B10 -A40 "applyCompaction|replacementMessages" packages/agent-core-v2/src packages/agent-core/src | head -n 220

echo
echo "## agent-core context compact implementation"
path="$(fd 'context.ts|Context.ts|context/index.ts' packages/agent-core/src | head -n 5 | tr '\n' ' ')"
for f in $path; do
  echo "--- $f"
  wc -l "$f"
  rg -n -B10 -A50 "applyCompaction|replacementMessages|_history|this\.history" "$f"
done

echo
echo "## local compaction tail handling around local path"
line="$(rg -n "collectCompactableUserMessages|replacementMessages:" packages/agent-core/src/agent/compaction/full.ts | tail -n 20 | awk -F: '{print $1}')"
for l in $line; do
  n=$((l-35)); [ "$n" -lt 1 ] && n=1
  e=$((l+45))
  echo "--- lines $n-$e"
  sed -n "${n},${e}p" packages/agent-core/src/agent/compaction/full.ts
done

Repository: YaseenHQ/kimi

Length of output: 24015


🏁 Script executed:

#!/bin/bash
set -eu

echo "## agent-core context applyCompaction"
rg -n -B20 -A80 "applyCompaction|replacementMessages|_history|context" packages/agent-core/src/profile/context.ts

echo
echo "## agent-core-v2 applyCompaction definitions"
rg -n -B25 -A65 "applyCompaction|replacementMessages|set\(|this\([^)]*history" packages/agent-core-v2/src/agent | head -n 300

echo
echo "## read-only source-based consistency check"
python3 - <<'PY'
from pathlib import Path
import re

for root in ["packages/agent-core/src/agent/compaction/full.ts",
             "packages/agent-core-v2/src/agent/fullCompaction/fullCompactionService.ts"]:
    text = Path(root).read_text()
    print(f"\n--- {root}")
    m = re.search(r"replacementMessages:\s*(?:remote\.messages|[^\n]+)\s*,", text)
    print("replacement assignment contains remote.messages:", "remote.messages" in (m.group(0) or ""))
    m2 = re.search(r"newHistory\s*=\s*.*\.history|this\.context\.get\(\)", text)
    print("uses live history snapshot:", bool(m2))
PY

Repository: YaseenHQ/kimi

Length of output: 34404


Remote compaction silently drops user messages appended mid-flight.

Both v1 and v2 remote-compaction branches pass only remote.messages to applyCompaction after confirming the live tail changed. Since replacementMessages replaces the compacted window, real user input added while the provider call is in flight is neither summarized nor preserved. Include the grown tail in the replacement array before applying compaction.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/agent-core/src/agent/compaction/full.ts` around lines 454 - 479, The
remote compaction path in the visible branch of full compaction must preserve
real user messages appended during the provider call. In the `remote` handling
around `applyCompaction`, append the live history tail after
`originalHistory.length` to `remote.messages` when constructing
`replacementMessages`; apply the same change to the corresponding v2
remote-compaction branch, while retaining the existing validation and
cancellation behavior.

Comment thread packages/agent-core/src/agent/context/index.ts
YaseenHQ added 4 commits July 23, 2026 22:25
The compact() request omitted include: ['reasoning.encrypted_content'],
so compaction preserved message continuity but dropped the model's
accumulated reasoning chain. Codex sets this unconditionally on every
Responses request (client.rs:888); match that behavior here.
# Conflicts:
#	packages/agent-core-v2/src/kosong/provider/bases/openai/openai-legacy.ts
#	packages/kosong/src/providers/kimi.ts
@YaseenHQ
YaseenHQ merged commit 5a65517 into main Jul 25, 2026
1 check was pending
@github-actions github-actions Bot mentioned this pull request Aug 11, 2026
@YaseenHQ
YaseenHQ deleted the codex/pi-server-compaction branch August 24, 2026 23:05
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant