Skip to content
Merged
29 changes: 29 additions & 0 deletions changelogs/v2.2.0.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
## What's new

### Features

- **Thinking / reasoning stream**: AI chat and the Cairn Agent now stream model reasoning text (Claude `thinking_delta`, OpenAI `delta.reasoning`) as a separate channel alongside content — never merged into note content or tool-call JSON. A collapsible **Thinking** panel renders above the assistant's reply, expanded by default while reasoning unfolds, auto-collapsing the instant the first content token arrives, and re-expandable via chevron at any time. Reasoning text is persisted to the `chat_messages` and `pi_agent_messages` tables so past messages retain their thinking panel across app restarts. For models that don't expose reasoning (e.g. standard GPT-4o), the panel simply never appears — zero overhead.
- **Backend**: `runToolLoop` in `electron/ipc/chat.ts` and `runAgentLoop` in `electron/lib/pi-agent-loop.ts` both parse `delta.reasoning` from SSE chunks, accumulate it through the tool-call loop, and emit new `chat:thought` / `pi-agent:thought` IPC channels mirroring the existing token channels. Non-streaming localllm path also reads `message.reasoning` from the completion response.
- **Usage**: `completion_tokens_details.reasoning_tokens` is now parsed from the OpenAI usage object and threaded through `onUsage` callbacks on both paths. The ContextRing popover (top-right of chat) now shows an **Output** section with **Answer** / **Thinking** / **Total** token breakdowns when a turn completes with reasoning token data.
- **IPC**: New preload listeners `window.electron.chat.onThought` and `window.electron.piAgent.onThought` added; `chat:usage` / `pi-agent:usage` payloads extended with `reasoningTokens`; `chat:done` payload now carries `reasoning` text for persistence.
- **Store**: `ChatMessage` and `PiAgentMessage` types gain `reasoning?: string`; `ChatThread.lastUsage` and `TerminalSession.lastUsage` gain `reasoningTokens?: number`. New `appendPiThought` Zustand action mirrors `appendPiToken` for live agent reasoning accumulation.
- **DB migration v20**: `ALTER TABLE chat_messages ADD COLUMN reasoning TEXT` and the same for `pi_agent_messages`. Idempotent — checks `PRAGMA table_info` before adding.
- **UI components**: New shared `ThinkingPanel.tsx` (`src/components/chat/chat-panel/`) handles both streaming (auto-collapse on first content token, re-expand on new reasoning, honour user override) and persisted (collapsed by default with chevron toggle) states. Rendered in `ToolCallIndicator` (live chat), `ChatMessageBubble` (persisted chat), and `AgentMessageBubble` (both live and persisted agent messages + subagent rows).

- **GitHub-style raw HTML in markdown**: All three markdown pipelines (`NoteMarkdownPreview`, `note-editor` read mode, `MarkdownContent` for chat) now install `rehype-raw` as the first rehype plugin, preserving raw HTML like `<p align="center">`, `<img>`, `<details>`/`<summary>`, and `<br>` instead of stripping them. Matches GitHub's rendering behaviour for README files pasted into notes or chat.

### Fixes

- **New chat thread not selected**: `SessionPane.handleNewChatThread` created a new thread but didn't call `setActiveChatThreadId` with the new thread's ID, so the chat panel kept showing the old session instead of switching to the freshly-created one. Now immediately switches.

- **Chat thread not switching on project change**: The `ChatPanel` init effect short-circuited whenever `activeChatThreadId` was non-null, so switching projects kept displaying the old project's chat thread. The effect now detects when `activeProjectId` changes and switches to a thread scoped to the new project via `getOrCreateThread`.

- **ESLint clean-install failure**: `eslint.config.mjs` imports from `eslint/config` (`@eslint/core`) which was only available transitively through `eslint@9` — clean installs (e.g. CI, PR review bots) couldn't resolve it. `@eslint/core` is now an explicit devDependency.

### Changes

- **Reasoning is intentionally stripped from compaction**: The chat compaction flow (`compactChatThread` in `src/store/slices/chat.ts`) sends only `{ role, content }` when summarising — reasoning text is never embedded into summary content, per the guidance that reasoning shouldn't leak into user-visible notes as text. Token counts (`reasoningTokens`) remain tracked for metrics/billing.

- **MCP / LLM history excludes reasoning**: `pi_agent_llm_history` (the raw model replay buffer for agent sessions) stores only role + content, so reasoning blocks don't interfere with downstream model turns.

- **`max_tokens` unchanged**: No request-side changes to `max_tokens` / `max_completion_tokens` were made. Models that split reasoning from content (e.g. Gemini-3.5-flash) produce `finish_reason: "length"` if the limit is too low — the existing `max_tokens: 4096` in `runToolLoop` is sufficient for most turns.
20 changes: 12 additions & 8 deletions electron/db/queries.ts
Original file line number Diff line number Diff line change
Expand Up @@ -534,13 +534,13 @@ export function getChatMessages(db: Database.Database, threadId: string) {
}

export function addChatMessage(db: Database.Database, m: {
id: string; threadId: string; role: string; content: string; contextRefs?: unknown; toolCalls?: unknown;
id: string; threadId: string; role: string; content: string; contextRefs?: unknown; toolCalls?: unknown; reasoning?: string;
}) {
const now = ts();
db.prepare(`
INSERT INTO chat_messages (id, thread_id, role, content, context_refs, tool_calls, created_at)
VALUES (?, ?, ?, ?, ?, ?, ?)
`).run(m.id, m.threadId, m.role, m.content, m.contextRefs ? JSON.stringify(m.contextRefs) : null, m.toolCalls ? JSON.stringify(m.toolCalls) : null, now);
INSERT INTO chat_messages (id, thread_id, role, content, context_refs, tool_calls, reasoning, created_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?)
`).run(m.id, m.threadId, m.role, m.content, m.contextRefs ? JSON.stringify(m.contextRefs) : null, m.toolCalls ? JSON.stringify(m.toolCalls) : null, m.reasoning ?? null, now);
return toChatMessage(db.prepare("SELECT * FROM chat_messages WHERE id = ?").get(m.id));
}

Expand Down Expand Up @@ -952,6 +952,7 @@ export interface PiMessageRow {
sessionId: string;
role: "user" | "assistant" | "error" | "system";
content: string;
reasoning: string | null;
toolCalls: unknown[] | null;
subagents: unknown[] | null;
timestamp: string;
Expand All @@ -965,6 +966,7 @@ function toPiMessage(row: any): PiMessageRow {
sessionId: row.session_id as string,
role: row.role as "user" | "assistant" | "error" | "system",
content: row.content as string,
reasoning: (row.reasoning as string | null) ?? null,
toolCalls: row.tool_calls ? JSON.parse(row.tool_calls as string) : null,
subagents: row.subagents ? JSON.parse(row.subagents as string) : null,
timestamp: row.timestamp as string,
Expand All @@ -974,17 +976,19 @@ function toPiMessage(row: any): PiMessageRow {

export function upsertPiMessage(
db: Database.Database,
msg: { id: string; sessionId: string; role: "user" | "assistant" | "error" | "system"; content: string; toolCalls?: unknown[] | null; subagents?: unknown[] | null; timestamp: string; order: number },
msg: { id: string; sessionId: string; role: "user" | "assistant" | "error" | "system"; content: string; reasoning?: string | null; toolCalls?: unknown[] | null; subagents?: unknown[] | null; timestamp: string; order: number },
) {
db.prepare(`
INSERT INTO pi_agent_messages (id, session_id, role, content, tool_calls, subagents, timestamp, "order")
VALUES (?, ?, ?, ?, ?, ?, ?, ?)
INSERT INTO pi_agent_messages (id, session_id, role, content, reasoning, tool_calls, subagents, timestamp, "order")
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(id) DO UPDATE SET
content = excluded.content,
reasoning = excluded.reasoning,
tool_calls = excluded.tool_calls,
subagents = excluded.subagents
`).run(
msg.id, msg.sessionId, msg.role, msg.content,
msg.reasoning ?? null,
msg.toolCalls ? JSON.stringify(msg.toolCalls) : null,
msg.subagents ? JSON.stringify(msg.subagents) : null,
msg.timestamp, msg.order,
Expand All @@ -994,7 +998,7 @@ export function upsertPiMessage(
export function savePiMessages(
db: Database.Database,
sessionId: string,
messages: Array<{ id: string; role: "user" | "assistant" | "error" | "system"; content: string; toolCalls?: unknown[] | null; subagents?: unknown[] | null; timestamp: string }>,
messages: Array<{ id: string; role: "user" | "assistant" | "error" | "system"; content: string; reasoning?: string | null; toolCalls?: unknown[] | null; subagents?: unknown[] | null; timestamp: string }>,
) {
const save = db.transaction(() => {
db.prepare("DELETE FROM pi_agent_messages WHERE session_id = ?").run(sessionId);
Expand Down
15 changes: 15 additions & 0 deletions electron/db/schema.ts
Original file line number Diff line number Diff line change
Expand Up @@ -480,6 +480,21 @@ const MIGRATIONS: Migration[] = [
db.exec("ALTER TABLE relationship_cache ADD COLUMN target_section_title TEXT");
}
},

// v20: Persist reasoning / thinking text on chat messages so the
// collapsible "Thinking" panel survives app restarts. Mirrors v14's
// approach for tool_calls. Also covers pi_agent_messages so terminal
// agent sessions retain their thinking across restarts too.
(db) => {
const chatCols = db.prepare("PRAGMA table_info(chat_messages)").all() as { name: string }[];
if (!chatCols.some((c) => c.name === "reasoning")) {
db.exec("ALTER TABLE chat_messages ADD COLUMN reasoning TEXT");
}
const piCols = db.prepare("PRAGMA table_info(pi_agent_messages)").all() as { name: string }[];
if (!piCols.some((c) => c.name === "reasoning")) {
db.exec("ALTER TABLE pi_agent_messages ADD COLUMN reasoning TEXT");
}
},
];

export function applySchema(db: Database.Database): void {
Expand Down
85 changes: 62 additions & 23 deletions electron/ipc/chat.ts
Original file line number Diff line number Diff line change
Expand Up @@ -83,29 +83,43 @@ async function runToolLoop(
signal?: AbortSignal,
getWin?: () => BrowserWindow | null,
provider?: string,
onUsage?: (pt: number, ct: number) => void,
onUsage?: (pt: number, ct: number, rt?: number) => void,
emitToolCallDone?: (e: { tool: string; cairnRef?: { type: "note" | "task"; id: string; title: string } }) => void,
onToken?: (delta: string) => void,
): Promise<{ exhausted: true; content: string } | { exhausted: false; content: string }> {
onThought?: (delta: string) => void,
): Promise<{ exhausted: true; content: string; reasoning: string } | { exhausted: false; content: string; reasoning: string }> {
const maxSteps = req.config?.maxSteps ?? 20;
const temperature = req.config?.temperature ?? 0.3;
let accumulatedContent = "";
let accumulatedReasoning = "";

for (let round = 0; round < maxSteps; round++) {
if (signal?.aborted) return { exhausted: true, content: "" };
let assistantMsg: OpenAIMessage;
if (signal?.aborted) return { exhausted: true, content: "", reasoning: accumulatedReasoning };

let assistantMsg: OpenAIMessage & { reasoning?: string };

if (provider === "localllm") {
try {
const { callLocalLLMChat } = await import("../lib/local-llm");
const res = await callLocalLLMChat(messages, TOOLS);
const choice = res.choices?.[0];
if (!choice) return { exhausted: true, content: "No response from local Llama on-device model." };
if (!choice) return { exhausted: true, content: "No response from local Llama on-device model.", reasoning: "" };
if (res.usage && onUsage) {
onUsage(res.usage.prompt_tokens ?? 0, res.usage.completion_tokens ?? 0);
onUsage(
res.usage.prompt_tokens ?? 0,
res.usage.completion_tokens ?? 0,
res.usage.completion_tokens_details?.reasoning_tokens ?? 0,
);
}
assistantMsg = choice.message as OpenAIMessage;
const rawMsg = choice.message as OpenAIMessage & { reasoning?: string };
if (rawMsg.reasoning) {
accumulatedReasoning += rawMsg.reasoning;
if (onThought) onThought(rawMsg.reasoning);
}
// Strip reasoning before assigning — it must not enter the messages
// array that gets re-sent to the API on subsequent rounds.
const { reasoning: _r, ...msgWithoutReasoning } = rawMsg;
assistantMsg = msgWithoutReasoning;

// Self-Healing Parser for On-Device XML-style tool calls and tokenizers
if (assistantMsg.content && assistantMsg.content.includes("<|tool_call>call:")) {
Expand Down Expand Up @@ -138,7 +152,7 @@ async function runToolLoop(
if (onToken) onToken(assistantMsg.content);
}
} catch (err) {
return { exhausted: true, content: `Local LLM Engine error: ${String(err)}` };
return { exhausted: true, content: `Local LLM Engine error: ${String(err)}`, reasoning: accumulatedReasoning };
}
} else {
let response: Response;
Expand All @@ -161,27 +175,31 @@ async function runToolLoop(
}),
});
} catch (_err) {
if (signal?.aborted) return { exhausted: true, content: "" };
return { exhausted: true, content: `Could not reach the AI endpoint at \`${baseUrl}\`. Check your endpoint URL and make sure the server is running.` };
if (signal?.aborted) return { exhausted: true, content: "", reasoning: accumulatedReasoning };
return { exhausted: true, content: `Could not reach the AI endpoint at \`${baseUrl}\`. Check your endpoint URL and make sure the server is running.`, reasoning: "" };
}

if (!response.ok) {
const errText = await response.text().catch(() => response.statusText);
return { exhausted: true, content: `AI endpoint error (${response.status}): ${errText.slice(0, 300)}` };
return { exhausted: true, content: `AI endpoint error (${response.status}): ${errText.slice(0, 300)}`, reasoning: "" };
}

const reader = response.body?.getReader();
if (!reader) return { exhausted: true, content: "No response stream" };
if (!reader) return { exhausted: true, content: "No response stream", reasoning: "" };

let contentBuffer = "";
const toolCallBuffers: Map<number, { id: string; name: string; args: string }> = new Map();
const toolCallBuffers: Map<number, { id: string; name: string; args: string; thought_signature?: string }> = new Map();

for await (const jsonStr of iterSseData(reader, signal ?? undefined)) {
try {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const chunk = JSON.parse(jsonStr) as any;
if (chunk.usage && onUsage) {
onUsage(chunk.usage.prompt_tokens ?? 0, chunk.usage.completion_tokens ?? 0);
onUsage(
chunk.usage.prompt_tokens ?? 0,
chunk.usage.completion_tokens ?? 0,
chunk.usage.completion_tokens_details?.reasoning_tokens ?? 0,
);
}
const delta = chunk.choices?.[0]?.delta;
if (!delta) continue;
Expand All @@ -192,6 +210,14 @@ async function runToolLoop(
if (onToken) onToken(delta.content);
}

// Reasoning / thinking stream (Claude thinking_delta, OpenAI delta.reasoning).
// Models that don't expose reasoning text simply never emit this field —
// the panel stays hidden. Reasoning is NOT merged into content/tool JSON.
if (delta.reasoning) {
accumulatedReasoning += delta.reasoning;
if (onThought) onThought(delta.reasoning);
}

if (delta.tool_calls) {
for (const tc of delta.tool_calls) {
const idx: number = tc.index ?? 0;
Expand All @@ -202,6 +228,8 @@ async function runToolLoop(
if (tc.id) buf.id = tc.id;
if (tc.function?.name) buf.name = tc.function.name;
if (tc.function?.arguments) buf.args += tc.function.arguments;
// Gemini 3.x thought signature — opaque blob to round-trip back.
if (tc.thought_signature) buf.thought_signature = tc.thought_signature;
}
}
} catch { /* skip malformed SSE JSON line */ }
Expand All @@ -216,7 +244,7 @@ async function runToolLoop(
});
}

if (signal?.aborted) return { exhausted: true, content: "" };
if (signal?.aborted) return { exhausted: true, content: "", reasoning: accumulatedReasoning };

const toolCalls = toolCallBuffers.size > 0
? Array.from(toolCallBuffers.entries())
Expand All @@ -225,20 +253,25 @@ async function runToolLoop(
id: buf.id,
type: "function" as const,
function: { name: buf.name, arguments: buf.args },
...(buf.thought_signature ? { thought_signature: buf.thought_signature } : {}),
}))
: undefined;

assistantMsg = {
role: "assistant",
role: "assistant" as const,
content: contentBuffer || null,
// Note: reasoning is intentionally NOT included here. It is
// accumulated separately in `accumulatedReasoning` and returned
// to the caller for UI/persistence. Sending it back to the API
// would violate both OpenAI and Anthropic message schemas.
tool_calls: toolCalls,
};
}

// No tool calls — model is ready to produce its final reply
if (!assistantMsg.tool_calls?.length) {
messages.push(assistantMsg);
return { exhausted: false, content: accumulatedContent };
return { exhausted: false, content: accumulatedContent, reasoning: accumulatedReasoning };
Comment thread
coderabbitai[bot] marked this conversation as resolved.
}

messages.push(assistantMsg);
Expand Down Expand Up @@ -285,6 +318,7 @@ async function runToolLoop(
return {
exhausted: true,
content: "I reached the maximum number of steps for this request. Any actions taken so far have been saved — check your board and notes. Try breaking the request into smaller steps.",
reasoning: accumulatedReasoning,
};
}

Expand Down Expand Up @@ -366,17 +400,19 @@ export function registerChatHandler(db: Database.Database, workspacePath: string

let promptTokens = 0;
let completionTokens = 0;
let reasoningTokens = 0;
let lastBreakdown: TokenBreakdown | undefined = undefined;
const addUsage = (pt: number, ct: number) => {
const addUsage = (pt: number, ct: number, rt?: number) => {
promptTokens += pt;
completionTokens += ct;
if (typeof rt === "number") reasoningTokens += rt;
try {
const rawBreakdown = calculatePromptBreakdown(buildSystemPrompt(req), messages, TOOLS);
lastBreakdown = scaleBreakdown(rawBreakdown, promptTokens);
} catch (err) {
console.error("[chat] failed to calculate breakdown:", err);
}
send("chat:usage", { promptTokens, completionTokens, breakdown: lastBreakdown });
send("chat:usage", { promptTokens, completionTokens, reasoningTokens, breakdown: lastBreakdown });
};

const loopResult = await runToolLoop(
Expand All @@ -385,7 +421,10 @@ export function registerChatHandler(db: Database.Database, workspacePath: string
emitToolCallDone,
(delta) => {
send("chat:token", { delta });
}
},
(delta) => {
send("chat:thought", { delta });
},
);

abortControllers.delete(event.sender.id);
Expand All @@ -399,10 +438,10 @@ export function registerChatHandler(db: Database.Database, workspacePath: string
}

if (abortCtrl.signal.aborted) {
send("chat:done", { content: "", contextRefs: [], usage: promptTokens > 0 ? { promptTokens, completionTokens, breakdown: lastBreakdown } : undefined });
send("chat:done", { content: "", reasoning: loopResult.reasoning, contextRefs: [], usage: promptTokens > 0 ? { promptTokens, completionTokens, reasoningTokens, breakdown: lastBreakdown } : undefined });
return;
Comment thread
coderabbitai[bot] marked this conversation as resolved.
}

send("chat:done", { content: loopResult.content, contextRefs: [], usage: promptTokens > 0 ? { promptTokens, completionTokens, breakdown: lastBreakdown } : undefined });
send("chat:done", { content: loopResult.content, reasoning: loopResult.reasoning, contextRefs: [], usage: promptTokens > 0 ? { promptTokens, completionTokens, reasoningTokens, breakdown: lastBreakdown } : undefined });
});
}
Loading