chore: merge upstream main (1896d1a1) into fork - #20
Merged
Conversation
…oonshotAI#2362) isImageFormatError missed the production phrasing "Unsupported image. Please try another one." — the existing pattern requires a url/format/ type suffix, so the deterministic image rejection never triggered the media-stripped resend, and the session failed on every later request. Add a standalone-sentence pattern (punctuation- or end-terminated) to both the kosong and agent-core-v2 classifiers, keeping the deliberate boundary that count/size phrasings must not match. Co-authored-by: fengchenchen <fengchenchen@moonshot.ai>
…s via the manifest systemPrompt field (MoonshotAI#2314) * feat(agent-core-v2): let plugins contribute system prompt instructions via the manifest systemPrompt field * feat(agent-core-v2): add systemPromptPath to load plugin system prompt from a file * docs: explain plugin system prompt templates * fix(agent-core-v2): refresh plugin system prompts after changes * fix(agent-core-v2): freeze restored profile bindings and converge plugin contributions at session scope - restore no longer re-renders or re-persists prompts: a resumed agent keeps its replayed profile binding (prompt and tool set) as persisted - a new Session-level convergence point reloads plugin skills into the session skill catalog before fanning out to every live agent prompt, and every catalog-kind plugin mutation awaits the whole pipeline; MCP-only toggles carry a distinct change kind and skip it - live refreshes after a restart re-resolve the bound profile by name and rebind the full slice (prompt, disallowed tools, active tools) atomically, warning and keeping the persisted state when the profile is gone; renders reuse the first-render timestamp and unchanged prompts are not re-persisted, so convergence never churns the wire - cap plugin system-prompt contributions (32 KB per field/file, 64 KB aggregate per prompt build) with manifest diagnostics and warnings - bump the changeset to minor: this is a new user-facing capability * fix(agent-core-v2): register the new session domain and dedupe the missing-profile warning - add sessionPluginContribution to the domain-layer registry so lint:domain stays green - emit system-prompt-refresh-profile-missing once per profile name, matching the service's other deduped warnings - document the convergence timeout escape hatch and the klient exclusion of enabledSystemPrompts * fix(agent-core-v2): dedupe the plugin budget warning and surface section read failures - emit plugin-sections-oversized once per skipped-plugin signature - let enabledSystemPrompts failures propagate to the refresh catch (keeps the current prompt and warns) instead of silently rendering and persisting a prompt without plugin instructions - cover the convergence timeout cut-off with a fake-timers test - clarify that the first-render timestamp anchors per process * fix(agent-core-v2): serialize session convergence and restore onDidReload timing - run at most one convergence per session and bound each change's wait by the timeout, so a fan-out emitter never interleaves deliveries after a timed-out convergence - fire onDidReload as soon as the reload commits again, keeping hook reloads independent of prompt convergence - sign the plugin budget warning with an unambiguous key * docs(agent-core-v2): align convergence wording with the serialized semantics - the timeout retry promise only holds once stalled work clears - note the per-session serial delivery cost model on the plugin change contract and the dual-queue invariant on the service * fix(agent-core-v2): keep empty plugin sections byte-neutral in the prompt template - place ${plugin_sections} on the same template line as ${skills_section} so prompts without either block render exactly as before this feature - note on the change contract that waitUntil work must not call back into plugin mutations, and spell out the per-session convergence order in the user docs * fix(agent-core-v2): pin a fork's profile so refresh triggers never rebind it - applyBindingSnapshot left the fork with no pinned profile, which routed in-process forks into the post-restart catalog rebind and could reset an inherited tool set; forks now inherit the source agent's pinned profile object - pin the first-render timestamp reuse with a ${now}-embedding test and document the anchored ${now} semantics - tighten the plugin docs budget and resume-refresh wording * fix(agent-core-v2): join in-flight convergence during agent bootstrap - an agent created while a plugin convergence is in flight now waits for it, and a restored agent refreshes once after it, so a plugin mutation never straddles an agent's bootstrap - warn on a non-string systemPrompt field and strip a UTF-8 BOM from systemPromptPath files before trimming - correct the consumption-surface wording (every CLI surface on the experimental flag, not just kimi -p), the per-session queueing note, and the single-plugin combined budget clause * fix(agent-core-v2): bound the bootstrap convergence join by the timeout A permanently wedged convergence kept convergeTail pending forever, and the unconditional settled() wait in bindBootstrap would have blocked every later agent creation in that session; the join now races the shared convergence timeout and continues (a restored agent still refreshes once, which never touches the tail), and the timeout constant moves to the contract for reuse * fix(agent-core-v2): close the convergence race against in-progress restores - a convergence fan-out could land while an agent's wire log is still replaying, dispatching a replay-visible config record whose effect the rest of the replay then overwrites; refreshSystemPrompt now skips while the wire restore is in progress - convergence completion is tracked by a generation counter; bootstrap compares it (after a bounded join) and refreshes a restored agent exactly once when a round completed after its creation began, replacing the wasConverging flag that could miss both windows * fix(agent-core-v2): bound each convergence so a wedged participant cannot stop the pipeline - the fan-out now races the convergence timeout, so convergeTail always settles: a permanently hung refresh delays its round (blocked entries drain oldest-first on later changes) instead of killing the session's convergence for good - warn when agent bootstrap stops waiting on a stalled convergence - diagnose a blank systemPromptPath and pin the plugin-root escape guard with traversal, absolute-path, and symlink tests * fix(agent-core-v2): bound the skill reload, preserve user-tool overlays, roll the prompt clock daily - the convergence's skill-reload segment now races the same timeout as the fan-out, so no segment of the pipeline can wedge a session for good; it continues with the previous catalog and retries next change - a cold rebind that resets the tool set replays session-added user tools onto the new base instead of dropping them for the rest of the process - the rendered timestamp re-anchors when the UTC date rolls over, so long-lived processes keep a fresh clock while steady-state renders stay byte-stable within a day - the plugin budget warning dedupes per plugin id, and the docs note that systemPromptPath content is frozen until the next reload * feat(agent-core-v2): converge cold plugin changes on resume through a drift-free gate - restore replays the persisted binding untouched, then bootstrap refreshes only when drift-free inputs changed while the session was cold: the catalog profile's tool set/denylist, or the plugin-sections baseline persisted alongside the prompt on the existing bind/update payloads; directory-listing and date drift wait for live triggers, so quiet resumes append no replay-visible records - the rendered timestamp is day-precision (UTC date at 00:00, re-anchored on rollover), keeping steady-state renders byte-stable across resumes and sessions on the same day - consolidate both timeout helpers onto a shared raceOutcome, and drop the generation counter the gate supersedes - align the plugin-sections precedence prose with the AGENTS.md disclaimer (no self-granted authority, system instructions win on conflict) * fix(agent-core-v2): bound the restored-prompt gate and land the sections baseline - the gate's plugin-sections read now races the convergence timeout, so agent creation never blocks behind an unrelated plugin mutation - refreshes serialize per agent through a tail, so overlapping triggers cannot write prompts out of order - when plugin sections change but a plugin-free custom prompt does not, the new baseline lands as a sections-only update instead of making every later resume re-render in vain - align the system prompt's Date and Time paragraph with the day-precision anchored timestamp * Update plugin system-prompt instructions in changeset Live sessions pick up plugin changes, while the default TUI and `kimi -p` paths ignore these fields. Signed-off-by: 7Sageer <sag77r@hotmail.com> * refactor(agent-core-v2): keep plugin skill reload user-driven Plugin mutations still converge live agent prompts, but the session skill catalog goes back to refreshing only on explicit plugin reload, as before: the prompt feature does not need skill convergence, and the pre-existing manual-reload semantics stay uniform across all plugin contributions. Removes the convergence-driven skill reload, the reloadSource de-privatization, and their tests; restores the PluginSkillSource onDidReload forwarding and its catalog tests. * refactor(agent-core-v2): apply plugin system-prompt changes only on explicit reload Drop the live convergence machinery (the plugin onDidChange barrier, the sessionPluginContribution fan-out, the restored-prompt drift gate, and the day-precision render clock) so plugin system-prompt sections take effect at the same point as every other plugin contribution: /plugins reload or a new session. The profile now refreshes when the session skill catalog re-pulls its plugin source on reload, reading both the skill list and the prompt sections fresh. * feat(agent-core): let plugins contribute system prompt instructions via the manifest systemPrompt field * chore(agent-core-v2): remove inline implementation comment * docs: clarify plugin prompt refresh semantics --------- Signed-off-by: 7Sageer <sag77r@hotmail.com>
…ection (…" (MoonshotAI#2368) This reverts commit dbb69a2.
zicochaos
force-pushed
the
merge/upstream-2026-07-30
branch
from
July 29, 2026 13:36
bb11a2c to
cb4b611
Compare
zicochaos
force-pushed
the
merge/upstream-2026-07-30
branch
from
July 29, 2026 13:40
cb4b611 to
4f6057b
Compare
6 tasks
zicochaos
pushed a commit
that referenced
this pull request
Aug 5, 2026
…oonshotAI#2604) * perf(minidb): skip idle everysec fsyncs and add lifecycle stats - everysec WAL now fsyncs on the timer only while dirty (tracked by a write/sync generation watermark); close() keeps its unconditional final sync, and background sync failures surface via walFsyncErrors plus a sticky lastWalFsyncError instead of being silently swallowed - add WAL queue/group-commit counters (walQueuedBytes, walMaxQueuedBytes, walGroupCommits, walGroupCommitFrames) and lifecycle phase stats: recovery bytes/frames/duration, index/text rebuild durations, compaction total/snapshot/rotation/postings durations, rotation pause, and query candidates/decoded/sorted rows; add a syncIntervalMs open option threaded through compaction WAL rotation - rewrite the bench on fixed-seed synthetic data with a stable machine-readable JSON report (cold open 10k/50k/100k, word/ngram search, idle-fsync acceptance, 100k compaction, event-loop delay, peak heap/RSS per scenario) and pin the schema in test/bench-json.test.ts; add app-side baselines with loose complexity budgets in sessionIndex and searchService tests - fix ClusterDb lock-pool closeAll() leaking in-flight shard opens and drain the query store's async close on server shutdown, eliminating the ENOTEMPTY directory-teardown race * perf(minidb): bound startup rebuild and steady-state hot paths - rebuild all derived indexes in one shared store walk: a single decode per record fans out to staged builders, dt rebuild reads record metadata only, and index-less opens no longer decode at all - rank full-text results with a bounded min-heap plus a stable key tie-break instead of sorting every candidate - remove/overwrite text docs via a docID -> delta-terms reverse map instead of scanning the whole delta vocabulary - validate unique batches incrementally against touched postings instead of copying the full per-index owner map - reap due TTL entries from the expiry heap on the write path instead of a full-store sweep per write * feat(session-index): add minidb read model with keyset pagination - add ISessionIndex read-model lifecycle (prepare/status, ready/degraded states) behind the persistence_minidb_readmodel experimental flag - add ISessionIndexMirror write side recording fresh summaries into a bounded, coalescing queue after the authoritative document is durable - replace the offset cursor with before/after keyset pagination; rename list/countActive to listRecent/count - extend IQueryStore with ordered columns and pageByColumn, plus getMany/listKeys/dropCollection - wire the read model through kap-server routes and start, and update the klient sessions contract - index every session for global search instead of the 500 most recent * feat(kap-server): bound search sync lifecycle, pagination, and query budgets - split search requests from sync work: searchIndex() no longer awaits runSync/reopen/reindex; a single-flight sync coordinator with debounce and backpressure runs in the background, stale generations keep serving with explicit stale/degraded state, and refresh/sync/reindex failures surface via lastRefreshError instead of being swallowed - scope file-meta keys by session id (\0meta\file\<sessionId>\<hash>) with lazy + one-shot background migration from the legacy hash-only keys, so one session sync only touches its own meta rows - make authoritative scans incremental (mtime/ino/size rescan conditions, unchanged files no longer rewrite meta) and read wire deltas in 1 MiB chunks instead of whole-file buffer + split - replace offset pagination with versioned v2 keyset page tokens (fingerprint + index generation + sort boundary); generation changes fail old tokens with invalid_page_token, legacy v1 offset tokens are served once and upgraded, and pages collect via bounded top-K instead of full sort + offset skip - add query budgets enforced at the postings/score stage: max query terms, literal length cap, postings visit budget (minidb searchBounded/maxVisits with prefix decoding that never fabricates hits and skips the postings LRU), candidate caps, deadline and text budget; truncation is reported via incomplete reasons candidate_cap/postings_budget/deadline - reopen read-only dbs by opening the next handle before closing the previous one so a failed refresh keeps the old generation serving; failed opens now self-heal through search traffic 100k-message bench: first page p95 < 300ms and page-100 cost on par with page 1; event-loop delay during queries stays sub-millisecond. * fix(minidb): poison, roll back, and recover the WAL on write failures - give WAL writes a commit point: a failed flushBatch poisons the WAL (WAL_POISONED, tracked separately as walWriteErrors vs walFsyncErrors), rejects queued frames in reverse enqueue order, and stops scheduling further batches; everysec background sync failures stay non-rejecting per stage-1 semantics - recover in place to a known-safe point: a serialized recovery chain truncates the WAL back to the first un-acked frame, rebuilds size/nextOffset, and clears the poison; writes queue behind the recovery gate (zero-cost when idle), a failed truncate flips the instance into an explicit writeDisabled state, and a stale truncate offset (WAL file replaced by a rotation) skips the truncate - roll failed flush groups back as a unit: frames are stamped with their batchId, MiniDb keeps per-group earliest pre-state, and the first rejection restores every key of the group (rejected writes no longer reappear after reopen, and in-memory state matches reopen for any failure interleaving); the per-op seq guard remains for cross-group and rotation-retry races - wrap applyOp and the following in-memory mutations so a contract violation poisons the WAL and rolls the group back instead of escaping as a half-commit; frames never enqueued (seal race) roll back per-op without poisoning - tag errors past the commit point with ambiguous: true so callers can distinguish "definitely not applied" from "maybe applied but revoked" - close() waits for the recovery chain to go idle and backup() fences behind in-flight recovery before copying files Controlled A/B bench (22 alternating iterations, 100k concurrent sets): write-path throughput regression is within the 2% budget. * fix(minidb): turn the file lock into an instance-owned serialized lease - distinguish lock ownership by instance instead of pid: every acquire mints a pid:uuid token carried by lock/bid/watch files, inspect().mine compares tokens, liveness still follows pid, tokenless legacy files keep the old stale-takeover path, and hasLiveForeignWatch excludes self by token so same-process contenders see each other (closing the double-win takeover and the cross-instance release); a live same-pid lock is still respected, and re-acquiring a held lock is idempotent - serialize acquire/renew/release through a per-instance promise-chain mutex: renew re-checks held inside the chain and release waits for an in-flight renew, eliminating the renew/rename-after-unlink ghost lock - make MiniDb.close() a state machine (open/closing/closed) with a shared closePromise: cleanup runs per-resource try/catch in dependency order (text indexes, store, valueReader, WAL, lock), aggregates every cleanup error into an AggregateError, stays in 'closing' on failure so a retry finishes the cleanup, and no longer leaks the lock when the WAL close fails; a rejected in-flight compaction no longer escapes the cleanup pass * fix(minidb): keep readers on one consistent file generation - add an internal persistent-files module as the single source of truth for the persisted file set (snapshot, WAL, sidecars, postings pattern, fingerprint subset); lock-pool fingerprints, persistentFiles, open stale-tmp cleanup, and backup/restore filtering all derive from it, and fingerprints upgrade to dev:ino:size:mtimeMs so compound sidecar changes can no longer hide from cluster readers - pair snapshot and WAL generations during recovery (transitional stat-pairing until stage-5 manifests): each pass anchors the fds it scans, re-stats afterwards, tolerates append-only WAL growth, retries bounded times on generation churn with a clean store reset, and throws RECOVERY_GENERATION_CHURN when churn exceeds the budget; the disk-mode ValueReader attach re-validates inodes so stale offsets never read a replaced file - make the rotation directory fsyncs strict: failures abort the rotation through the existing rollback path instead of being swallowed, while platforms without directory fsync degrade once with a warn and stats.dirFsyncUnsupported * fix(minidb): serialize index-definition sidecar mutations, persist before publish - extract the promise-chain mutex into a shared createSerializer() and give each sidecar family (secondary/compound/text) its own chain: create/drop run uninterruptibly (memory change + rebuild + persist), different families stay independent, and the data write path never shares these chains - reverse the publication order to staged -> persist -> publish: a create stages the definition, rebuilds via the staged builder, persists the sidecar including the new definition, then publishes atomically; any failure discards the staged state leaving live and sidecar untouched (no phantom indexes, retry-safe); a drop persists the sidecar without the definition before removing it live; text index create/drop adopt the same pattern, replacing the hand-rolled unwind, and a dropping marker keeps compaction postings rebuilds out of the persist window - feed staged indexes from the incremental write path (add/remove/ checkUnique/checkUniqueBatch visit live+staged) so writes landing in the persist window are not lost at publish; queries still see live only - harden writeFileAtomic: instance-unique tmp names (.tmp-pid-seq), a strict fsyncDir after rename so a successful persist is crash durable, and whitelist-based stale-tmp cleanup that never touches lock tmp files * fix(minidb): validate writes before any side effect, canonicalize values once - canonical value at the write boundary: the json codec re-parses the encoded bytes once and every downstream consumer (unique checks, secondary/compound/text indexes, dt extraction) sees exactly the persisted representation, so getter/toJSON/Proxy documents can no longer diverge between the index view and the storage view - reorder the set/batch pipeline so every fallible check happens before any visible side effect: prepare (key/ttl checks, encoding, canonical decode, index field extraction, tokenization) -> unique checks -> ensureMemoryFor eviction -> commit; a constraint failure now leaves the database untouched (no more evicted victims on rejected inserts), and applyOp is structurally pure against pre-validated data - tokenize at the prepare boundary: TextIndex gains prepareAdd/ addPrepared and the buildQueue carries validated key+tokens mutations instead of raw docs, so a throwing custom tokenizer can no longer poison the live view or the queue, and custom-tokenizer output is rejected per token over 0xffff bytes before it can permanently break postings rebuilds; prepared tokens are keyed by index instance so a same-name drop+create mid-write re-tokenizes instead of crossing tokenizers - strict batch structure validation: scanBatchOpRefs/decodeBatchOps reject unknown op types, out-of-bounds lengths, and trailing bytes (offset must equal body length), so a valid-CRC but malformed batch is skipped as a unit and counted via RecoveryInfo.corruptBatches instead of being partially applied Bench vs the stage-1 baseline: json write throughput regression is within the 5% budget (median ~2-4% depending on the measurement). * feat(minidb): add OpTracker drain primitive and atomic backup, harden tests - introduce the internal OpTracker (close gate + in-flight counter with enter/leave/close/whenIdle and reference-counted pause/resume) and drive every shutdown/drain path from it: WAL background syncs are tracked so close() waits out an in-flight sync before closing the fd, cluster lock-pool closeAll() closes the gates and drains busy callbacks before closing handles, and MiniDb writes pass a write gate - make backup() atomic with a defined linearization point: pause the write gate, drain in-flight writes (every acknowledged write is now included), copy to a sibling temp dir with per-file fsyncs, write the manifest last as the commit marker, and rename into place; failures clean up and leave no partial backup, and concurrent writes are rejected with BACKUP_IN_PROGRESS - reap emptied compound-index groups on remove (the groups map no longer grows monotonically), move the open-time mkdir behind the readOnly check so a read-only open of a missing directory fails with ENOENT instead of creating it, and never run a destructive rebuild for a read-only open failure (explicit or onLockFail fallback) - consolidate every review fault-injection repro into the formal suite behind deterministic barrier helpers (programmable writev/sync/ rename/tokenize hooks) and convert the six timing-based tests to barrier/tick-driven assertions; the .tmp repro scripts are removed The converted timing tests and the full suite pass 50 repeat runs (including under CPU load injection) with zero flakes. * feat(minidb): persist derived indexes as atomic generations, open from WAL delta - checkpoint the store, dt/secondary/compound indexes, and text dictionary/postings/docs into immutable generations under generations/g-NNNNNN published atomically (tmp build, per-file checksums and fsyncs, dir rename, CURRENT swap, strict dir fsyncs); the manifest records the format version, WAL/snapshot checkpoint anchors, per-index definition hashes, and codec/value-mode compatibility - open now loads the published generation and replays only the WAL delta after its checkpoint: no full value decode, corpus tokenization, or postings rewrite on a normal reopen (warm opens are 3.5-13.8x faster at 100k/1M records); a definition change rebuilds only the affected index, and corrupt generation files fall back to the previous generation or the legacy full recovery without ever touching the authoritative snapshot/WAL - build generations transactionally with compaction (rotation plus derived state publish as one unit, replacing the synchronous rebuildTextPostings tail), capture concurrent writes through a sealed op queue with byte/op caps, hard-link clean postings and the snapshot into the new generation, and repoint every live text base into the CURRENT generation after publish - cluster/read-only refresh watches CURRENT and the WAL watermark: pure generation publishes keep readers on incremental catch-up while rotations reopen onto the new generation; writers building the next generation never disturb readers of the current one - legacy databases open through the old path unchanged and gain their first generation in the background; OpenOptions.indexGenerations: false fully restores the pre-generation behavior * feat(minidb): workerize text-index builds and split MiniDb into facets - split the monolithic src/index.ts into facet modules (mini-db, types, value-codec, memory-guard, backup, query-engine, text-registry, wal-group, generation-builder/loader, write-path, read-path, index-admin, lifecycle, stats) and move text-index.ts to text-index/ - run corpus-scale text-index builds off the main thread via the bounded worker engine (src/worker/), exported through the new worker-runtime subpath, with inline fallback for small corpora and rollback switches - defer the open-time fallback text rebuild into a maintenance task; searches on a not-yet-committed base raise TextIndexBuildingError - add the unified maintenance scheduler, bounded async read surface, and a maintenance bench - kap-server search: switch to searchBoundedAsync and serve the building page while the index base rebuilds after fallback recovery - kimi-code: install the SEA-bundled minidb text-build worker at startup, bundle it via the native asset scripts, and add the startup-trace util plus the KIMI_TUI_INPUT_LATENCY debug probe * fix(minidb): treat win32 EPERM as unsupported directory fsync - extract isUnsupportedDirectoryFsyncError and cover win32 EPERM - drop the one-shot console.warn; stats.dirFsyncUnsupported carries the degraded state * fix(kap-server): harden search-index dispose and drain lifecycle - dispose() now closes an OpTracker gate and drains in-flight sync/refresh passes before closing the db, so no background write can hit a closed handle; the deleteSessionDocs loop and trailing stats write skip once the gate closes (review #20) - drainGlobalSearchDisposals loops to a fixpoint so disposals registered while a drain is in flight are also awaited (review #21) - pin the post-open failure semantics with a regression test: a failed text-index setup closes the handle and the next open reacquires the writer lock instead of self-locking read-only (review #19) - export OpTracker from the minidb root for the search service's drain * chore: fix oxlint type-aware lint errors
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Conflict resolutions
Verification