Skip to content
Merged
Show file tree
Hide file tree
Changes from 8 commits
Commits
Show all changes
30 commits
Select commit Hold shift + click to select a range
8c959b4
fix(agent-core-v2): honor request-level toolMessageConversion in open…
7Sageer Sep 8, 2026
1874f4c
refactor(agent-core-v2): split the protocol trait into per-protocol d…
7Sageer Sep 8, 2026
1833731
refactor(agent-core-v2): compose format stages and dialect hooks in t…
7Sageer Sep 8, 2026
54ce16e
refactor(agent-core-v2): rename protocol dialects to traits
7Sageer Sep 9, 2026
4844b9a
fix(agent-core-v2): include the system message in openai mergeHistory…
7Sageer Sep 9, 2026
fef7f8d
refactor(agent-core-v2): enforce the format/trait contract boundary
7Sageer Sep 9, 2026
38598cf
refactor(agent-core-v2): thread the request config by spread and drop…
7Sageer Sep 9, 2026
e330f50
refactor(agent-core-v2): keep protocol format modules internal to the…
7Sageer Sep 9, 2026
3effd9b
refactor(agent-core-v2): plug llm credentials in through request config
7Sageer Sep 9, 2026
fe5b4b9
fix(agent-core-v2): surface request-actor failures as llm.failed.remote
7Sageer Sep 9, 2026
0f899d2
docs(agent-core-v2): attribute retry/recovery to the turn machine in …
7Sageer Sep 9, 2026
8bb2f04
refactor(agent-core-v2): unify credentials in human/credentials, turn…
7Sageer Sep 9, 2026
d7ffce0
fix(agent-core-v2): credential recovery for direct paths, abort guard…
7Sageer Sep 9, 2026
7bb6114
feat(agent-core-v2): IModelCatalog.generate with stream credential re…
7Sageer Sep 9, 2026
2e870d1
refactor(agent-core-v2): inline credential recovery at call sites
7Sageer Sep 9, 2026
fc59aec
fix(agent-core-v2): settle the queued turn when a machine turn settle…
7Sageer Sep 9, 2026
426e0ba
fix(agent-core-v2): bind and end pre-gate failures through the normal…
7Sageer Sep 9, 2026
b9be6d3
fix(agent-core-v2): settle seeded turns on pre-gate failure, preserve…
7Sageer Sep 9, 2026
bc38daa
fix(kap-server): add generate to IModelCatalog test fakes
7Sageer Sep 9, 2026
bb569c8
fix(klient): migrate examples to the credentials API
7Sageer Sep 9, 2026
0be003f
test(agent-core-v2): drop credential recovery coverage
7Sageer Sep 9, 2026
eba30ea
refactor(agent-core-v2): fold kimi-oauth credential adapter into cred…
7Sageer Sep 10, 2026
17b256b
refactor(agent-core-v2): pair credential refresh with invalidate, opt…
7Sageer Sep 10, 2026
d78c485
refactor(agent-core-v2): credentials recovery strategy chain, shared …
7Sageer Sep 10, 2026
7086d06
fix(agent-core-v2): settle message-less notifications on pre-gate fai…
7Sageer Sep 10, 2026
812ab19
Merge branch 'main' into refactor/llm-protocol-dialect
7Sageer Sep 10, 2026
4227356
Merge branch 'main' into refactor/human-connection-credentials
7Sageer Sep 10, 2026
22cc391
Merge branch 'main' into refactor/human-connection-credentials
7Sageer Sep 10, 2026
124c383
test(kap-server): add the IModelCatalog.generate stub to the history …
7Sageer Sep 10, 2026
a0ec899
Merge branch 'main' into refactor/human-connection-credentials
7Sageer Sep 10, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 11 additions & 7 deletions packages/agent-core-v2/docs/en/llm.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ llm is a standalone LLM request library inside the human layer (`src/human/llm/`
2. **Streaming-native; events are the contract**. The only outward surface is a single, purely serializable event stream (requester level: `llm.sent / streaming.headers / streaming.part / streaming.usage / streaming.finish / streaming.message_id / failed.syntax / failed.remote / done`; the machine level adds `llm.retrying / llm.recovering`, and `llm.sent` carries the most recent recovery record). Streaming and non-streaming are isomorphic (non-streaming also accumulates over the stream, just without deltas). Events are emitted as they arrive — no caching, no fallback.
3. **format masks inter-protocol differences; trait expresses provider dialects**. format lives at the protocol layer and handles encoding/decoding of requests, responses, errors, usage, and finish. trait is a bundle of hooks a provider attaches (endpoint, headers, convertMessage, buildParams, withThinking, etc.). Protocol differences must not leak into the machine or into requester decorators.
4. **Two-layer error model**. Internally, code throws the SDK's native errors; local request validation throws the shared `SyntaxRequestFormatError` (`llm/syntax-errors.ts`), which the requester converts uniformly via `toLlmSyntaxErrorMessage`, with no intermediate layer. Externally there are only `llm.failed.syntax` (local message syntax errors, never retried) and `llm.failed.remote` (remote streaming errors, subdivided into connection / timeout / rate_limit / quota_exhausted / context_overflow / request_structure, etc.), converted by format at the boundary.
5. **Stateless core + state machine shell**. `generate(config, content, control)` is a stateless function; errors are delivered via onEvent, never thrown. The llm machine wraps a single request (messageResolvers, abort scope, event forwarding) and drives retry and recovery through the pure policy functions in retry.ts / recovery.ts: recovery re-sends with replacement messages produced by the pure `propose` function (attempt resets to 1), retry backs off in the `retrying` state (honoring Retry-After), and the machine emits `llm.recovering / llm.retrying` for each. Empty response is judged by `withEmptyResponseGuard` at the requester boundary and raised as `llm.failed.remote`, entering the same retry path. Abort is carried by an AbortController owned by the turn: the controller is passed into the machine and the request actor via `LlmInput.signal`, and the turn aborts it directly on `turn.abort`, with the request ending as `llm.failed.remote`; the request actor neither creates its own controller nor touches any signal on teardown, so a finished request can never abort a shared signal. The accumulator is held by the turn and fed by the event stream; on `llm.retrying / llm.recovering` the turn rolls it back and recreates it, so every attempt accumulates from zero while as much interrupted state as possible is preserved (the turn finishes the complete message out of the accumulator at `llm.done`).
5. **Stateless core + state machine shell**. `generate(config, content, control)` is a stateless function; errors are delivered via onEvent, never thrown. The llm machine wraps a single request (messageResolvers, event forwarding). The turn machine drives retry and recovery through the pure policy functions in retry.ts / recovery.ts: recovery re-sends with replacement messages produced by the pure `propose` function (attempt resets to 1), retry backs off in the turn's `retrying` state (honoring Retry-After), and the turn emits `llm.recovering / llm.retrying` for each. Empty response is judged by the turn at `llm.done` via the pure `emptyResponseError` and re-raised as `llm.failed.remote`, entering the same failure cascade. Abort is carried by an AbortController owned by the turn: the controller is passed into the machine and the request actor via `LlmInput.signal`, and the turn aborts it directly on `turn.abort`, with the request ending as `llm.failed.remote`; the request actor neither creates its own controller nor touches any signal on teardown, so a finished request can never abort a shared signal. The accumulator is held by the turn and fed by the event stream; on `llm.retrying / llm.recovering` the turn rolls it back and recreates it, so every attempt accumulates from zero while as much interrupted state as possible is preserved (the turn finishes the complete message out of the accumulator at `llm.done`).
6. **No silent fallback**. Configuration is taken exactly as given. For beta features, thinking, empty response, and similar scenarios, define explicit error conditions first, fail at request time, and guide the user to fix the configuration — never fall back silently.
7. **Every variable capability is a contribution point**. Providers, media upload/degradation, usage, traceId, and error recovery (compaction / media degradation) all plug in through extension points; the llm core contains none of these concepts.
8. **Data is data**. A model is pure, function-free data (endpoint url + model uniquely identifies a model), serializable and directly usable as generate input. The catalog is a derived `provider -> models` cache; the dependency direction only goes from models-dev into llm internals, never the reverse.
Expand All @@ -32,11 +32,15 @@ llm/
│
├── requester/
│ ├── requester.ts LlmRequester.generate(config, content, control);
│ │ ExtraParams typed per protocol {openai?, responses?, anthropic?, googleGenai?}
│ ├── machine.ts llm state machine (single request + retry/recovery + empty response
│ │ judgment; emits llm.retrying / llm.recovering)
│ ├── retry.ts / recovery.ts pure retry/recovery policy functions (driven by the llm machine; propose is pure)
│ ├── empty-response.ts withEmptyResponseGuard: judges empty responses at finish and raises llm.failed.remote
│ │ ExtraParams typed per protocol {openai?, responses?, anthropic?, googleGenai?};
│ │ LlmRequestConfig.credentials: credential contribution point
│ │ (resolve/canRecover/invalidate), resolved per attempt by the caller;
│ │ factories live in human/credentials (staticCredentials / oauthCredentials;
│ │ kimiOAuthCredentialProvider adapts Kimi OAuth tokens)
│ ├── machine.ts llm state machine (single request: messageResolvers +
│ │ event forwarding)
│ ├── retry.ts / recovery.ts pure retry/recovery policy functions (driven by the turn machine; propose is pure)
│ ├── empty-response.ts emptyResponseError: pure empty-response judgment; the turn raises it as llm.failed.remote at llm.done
│ └── bases/ four protocol bases: openai / openai-responses / anthropic / google-genai
│ each with format / lower / patterns / capability / extra-params / requester
│
Expand All @@ -51,7 +55,7 @@ llm/
└── media/ media contribution points: cache / degrade / ref / resolver / store / upload
```

Request lifecycle: `generate` receives (config, content, control) → format lowers the generic Message[] through the Pattern Rewriter into protocol requestParams → internalGenerate calls the official SDK → streaming chunks are converted by the stateless parser callbacks into `llm.streaming.part / streaming.usage / streaming.finish / streaming.message_id` events → errors are converted by format into `llm.failed.*`; on success the requester emits `llm.done`, on failure it ends with `llm.failed.syntax / llm.failed.remote` and never emits `llm.done`. `withEmptyResponseGuard` judges empty responses at finish and raises `llm.failed.remote`; the llm machine first tries recovery on `llm.failed.remote` (replacement messages from the pure `propose` function, emitting `llm.recovering`), then retries with backoff (honoring Retry-After, emitting `llm.retrying`), and only lands in the failed final state once attempts are exhausted. The upper-layer turn holds the HistoryAccumulator, fed by the event stream, rolls it back and recreates it on `llm.retrying / llm.recovering`, and finishes the complete message at `llm.done`; usage accounting, tracing, compaction, and media degradation all attach to the event stream as plugins/contribution points.
Request lifecycle: `generate` receives (config, content, control) → the caller resolves `config.credentials` into a fully-credentialed model before each attempt (the llm machine's request actor for the machine path), so requests always carry fresh credentials and a credential-refresh recovery (recoverable 401 → `credentials.invalidate()`, emitted as `llm.recovering` with strategy `credentials`) naturally re-resolves on the re-send (direct callers outside the state machines — ping, generate, full compaction, media upload — inline the same single-retry recovery at the call site) → format lowers the generic Message[] through the Pattern Rewriter into protocol requestParams → internalGenerate calls the official SDK → streaming chunks are converted by the stateless parser callbacks into `llm.streaming.part / streaming.usage / streaming.finish / streaming.message_id` events → errors are converted by format into `llm.failed.*`; on success the requester emits `llm.done`, on failure it ends with `llm.failed.syntax / llm.failed.remote` and never emits `llm.done`. at `llm.done` the turn judges empty responses via `emptyResponseError` and re-raises them as `llm.failed.remote`; the turn machine first tries recovery on `llm.failed.remote` (credential refresh on a recoverable 401, and replacement messages from the pure `propose` function, emitting `llm.recovering`), then retries with backoff (honoring Retry-After, emitting `llm.retrying`), and only lands in the failed final state once attempts are exhausted. The upper-layer turn holds the HistoryAccumulator, fed by the event stream, rolls it back and recreates it on `llm.retrying / llm.recovering`, and finishes the complete message at `llm.done`; usage accounting, tracing, compaction, and media degradation all attach to the event stream as plugins/contribution points.

## Rejected Schemes (do not reintroduce)

Expand Down
17 changes: 10 additions & 7 deletions packages/agent-core-v2/docs/zh/llm.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ llm 是 human 层内一个独立的 LLM 请求库(`src/human/llm/`),提供
2. **流式原生、事件即契约**。对外只暴露一条纯可序列化的事件流(requester 层:`llm.sent / streaming.headers / streaming.part / streaming.usage / streaming.finish / streaming.message_id / failed.syntax / failed.remote / done`;machine 层补充 `llm.retrying / llm.recovering`,`llm.sent` 携带最近一次 recovery 记录),流式与非流式同构(非流式也走流式累积,只是不发 delta);事件收到即发,不缓存、不兜底。
3. **format 屏蔽协议间差异,trait 表达 provider 方言**。format 位于 protocol 层,负责请求、响应、错误、usage 和 finish 的编解码;trait 是 provider 附加的一包 hooks(endpoint、headers、convertMessage、buildParams、withThinking 等)。协议差异不允许泄漏到 machine 或 requester 的装饰层。
4. **错误两层模型**。内部 throw SDK 原生错误;本地请求校验抛共享的 `SyntaxRequestFormatError`(`llm/syntax-errors.ts`),由 requester 经 `toLlmSyntaxErrorMessage` 统一转换,不加中间层。对外只有 `llm.failed.syntax`(本地消息语法错误,不重试)与 `llm.failed.remote`(远程流式错误,细分为 connection/timeout/rate_limit/quota_exhausted/context_overflow/request_structure 等),由 format 在边界完成转换。
5. **无状态内核 + 状态机外壳**。`generate(config, content, control)` 是无状态函数,错误走 onEvent 不 throw;llm machine 包装单次请求(messageResolvers、abort 作用域、事件转发),并借助 retry.ts / recovery.ts 的纯策略函数驱动重试与 recovery:recovery 由纯函数 `propose` 产出替换消息直接重发(attempt 重置为 1),重试走 `retrying` 状态的 backoff(尊重 Retry-After),两者分别对外补发 `llm.recovering / llm.retrying` 事件;empty response 由 `withEmptyResponseGuard` 在 requester 边界判定并转为 `llm.failed.remote`,进入同一重试路径;abort 由 turn 持有的 AbortController 承载:controller 经 `LlmInput.signal` 传入 machine 与 request actor,turn 在 `turn.abort` 时直接 abort 它,请求随即以 `llm.failed.remote` 收尾;request actor 不自建 controller、回收时不触碰任何 signal,正常完成的请求绝不可能误 abort 共享 signal。累积器由 turn 持有并随事件流喂入,在 `llm.retrying / llm.recovering` 时 rollback 并重建,每次 attempt 从零累积,从而尽可能保留中断现场(turn 在 `llm.done` 时从累加器 finish 出完整消息)。
5. **无状态内核 + 状态机外壳**。`generate(config, content, control)` 是无状态函数,错误走 onEvent 不 throw;llm machine 包装单次请求(messageResolvers、事件转发);重试与 recovery 由 turn machine 借助 retry.ts / recovery.ts 的纯策略函数驱动:recovery 由纯函数 `propose` 产出替换消息直接重发(attempt 重置为 1),重试走 turn 的 `retrying` 状态的 backoff(尊重 Retry-After),两者分别由 turn 对外补发 `llm.recovering / llm.retrying` 事件;empty response 由 turn 在 `llm.done` 时经纯函数 `emptyResponseError` 判定并重新转为 `llm.failed.remote`,进入同一失败级联;abort 由 turn 持有的 AbortController 承载:controller 经 `LlmInput.signal` 传入 machine 与 request actor,turn 在 `turn.abort` 时直接 abort 它,请求随即以 `llm.failed.remote` 收尾;request actor 不自建 controller、回收时不触碰任何 signal,正常完成的请求绝不可能误 abort 共享 signal。累积器由 turn 持有并随事件流喂入,在 `llm.retrying / llm.recovering` 时 rollback 并重建,每次 attempt 从零累积,从而尽可能保留中断现场(turn 在 `llm.done` 时从累加器 finish 出完整消息)。
6. **不兜底**。配置是什么就是什么;beta 特性、thinking、empty response 等场景先定义明确报错条件,在请求阶段报错并引导用户修正,而不是静默兜底。
7. **一切可变能力都是贡献点**。provider、媒体上传/降级、usage、traceId、错误恢复(compaction/媒体降级)都通过扩展点接入,llm 内核不含这些概念。
8. **数据即数据**。model 是无函数的纯数据(endpoint url + model 唯一标识一个模型),可序列化、可直接作为 generate 输入;catalog 是 `provider -> models` 的派生缓存,依赖方向只能从 models-dev 指向 llm 内部,不能反向依赖。
Expand All @@ -32,11 +32,14 @@ llm/
│
├── requester/
│ ├── requester.ts LlmRequester.generate(config, content, control);
│ │ ExtraParams 按协议带类型 {openai?, responses?, anthropic?, googleGenai?}
│ ├── machine.ts llm 状态机(单次请求 + 重试/恢复 + empty response 判定;
│ │ 对外补发 llm.retrying / llm.recovering)
│ ├── retry.ts / recovery.ts 重试/恢复策略纯函数(由 llm machine 驱动;propose 为纯函数)
│ ├── empty-response.ts withEmptyResponseGuard:finish 时判定空响应并转为 llm.failed.remote
│ │ ExtraParams 按协议带类型 {openai?, responses?, anthropic?, googleGenai?};
│ │ LlmRequestConfig.credentials:凭证贡献点
│ │ (resolve/canRecover/invalidate),由调用方在每次 attempt 前解析;
│ │ 工厂位于 human/credentials(staticCredentials / oauthCredentials;
│ │ kimiOAuthCredentialProvider 适配 Kimi OAuth token)
│ ├── machine.ts llm 状态机(单次请求:messageResolvers + 事件转发)
│ ├── retry.ts / recovery.ts 重试/恢复策略纯函数(由 turn machine 驱动;propose 为纯函数)
│ ├── empty-response.ts emptyResponseError:空响应判定纯函数,由 turn 在 llm.done 时转为 llm.failed.remote
│ └── bases/ 四个协议基座:openai / openai-responses / anthropic / google-genai
│ 各自含 format / lower / patterns / capability / extra-params / requester
│
Expand All @@ -51,7 +54,7 @@ llm/
└── media/ 媒体贡献点:cache / degrade / ref / resolver / store / upload
```

请求生命周期:`generate` 收到 (config, content, control) → format 将通用 Message[] 经 Pattern Rewriter 降低为协议 requestParam → internalGenerate 调用官方 SDK → 流式 chunk 经无状态 parser 回调转换为 `llm.streaming.part / streaming.usage / streaming.finish / streaming.message_id` 事件 → 错误由 format 转换为 `llm.failed.*`;成功时 requester 发出 `llm.done`,失败时以 `llm.failed.syntax / llm.failed.remote` 收尾、不再发 `llm.done`。`withEmptyResponseGuard` 在 finish 时判定空响应并转为 `llm.failed.remote`;llm machine 对 `llm.failed.remote` 先尝试 recovery(纯函数 `propose` 产出替换消息,发 `llm.recovering`),再按策略 backoff 重试(尊重 Retry-After,发 `llm.retrying`),耗尽后才以 failed 终态收尾。上层的 turn 持有 HistoryAccumulator 随事件流累积,在 `llm.retrying / llm.recovering` 时 rollback 并重建累加器,`llm.done` 时 finish 出完整消息;usage 统计、trace、compaction、媒体降级均以插件/贡献点身份挂接在事件流上。
请求生命周期:`generate` 收到 (config, content, control) → 调用方在每次 attempt 前把 `config.credentials` 解析成带完整凭证的 model(machine 路径由 llm machine 的 request actor 完成),请求因此始终携带新鲜凭证,而凭证刷新恢复(可恢复的 401 → `credentials.invalidate()`,以 `llm.recovering`(strategy 为 `credentials`)发出)在重发时自然重新解析(不经状态机的 direct 调用方——ping、generate、full compaction、媒体上传——在调用点内联同样的单次重试恢复) → format 将通用 Message[] 经 Pattern Rewriter 降低为协议 requestParam → internalGenerate 调用官方 SDK → 流式 chunk 经无状态 parser 回调转换为 `llm.streaming.part / streaming.usage / streaming.finish / streaming.message_id` 事件 → 错误由 format 转换为 `llm.failed.*`;成功时 requester 发出 `llm.done`,失败时以 `llm.failed.syntax / llm.failed.remote` 收尾、不再发 `llm.done`。turn 在 `llm.done` 时经 `emptyResponseError` 判定空响应并重新转为 `llm.failed.remote`;turn machine 对 `llm.failed.remote` 先尝试恢复(可恢复 401 的凭证刷新,以及纯函数 `propose` 产出的替换消息,发 `llm.recovering`),再按策略 backoff 重试(尊重 Retry-After,发 `llm.retrying`),耗尽后才以 failed 终态收尾。上层的 turn 持有 HistoryAccumulator 随事件流累积,在 `llm.retrying / llm.recovering` 时 rollback 并重建累加器,`llm.done` 时 finish 出完整消息;usage 统计、trace、compaction、媒体降级均以插件/贡献点身份挂接在事件流上。

## 已被否决的方案(不要再引入)

Expand Down
Loading
Loading