Skip to content

perf(core): speed up and harden snapshot capture - #51996

Merged
nexxeln merged 6 commits into
v2from
snapshot-lab
Oct 8, 2026
Merged

nexxeln merged 6 commits into
v2from
snapshot-lab

Conversation

@nexxeln

@nexxeln nexxeln commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Why the change

Snapshots added 153 to 363 ms of Git work to every session step, could break for good after a crash, when two processes captured at once, or under some Git settings, and grew the store one loose file per object forever. This roughly halves the per-step cost, makes capture self-healing, and keeps memory flat and the store packed, without changing what a snapshot records.

Special things to note

  • Output matches v2 across 32 edge-case scenarios under 7 Git configurations (script/snapshot-parity-run.sh). The only differences are cases where v2 failed or lost data: names starting with : (v2 fails the capture and undo deletes the file), a global core.splitIndex=true (v2 writes empty trees; this PR writes none), a corrupt index, a stale index.lock, concurrent processes, and ..scope directories.
  • Retention (deleting old snapshots) and capturing on file mutation are left out, because both change what history users keep or what a step records.

Memory and disk

Every step writes new objects into the store, and v2 never packed them, so stores grew one loose file per object (the largest real one had 23,924 loose objects, 198 MiB) and lookups slowed as they grew. Deleting old snapshots is a separate product call, so lossless packing is the only disk fix that keeps every snapshot. The memory work keeps the new design from costing anything that grows over a long session.

  • No copies. The temporary index is a hard link, so isolating each capture costs no bytes. On a 100k-file repo the index is 8.4 MB, which a copy would duplicate every step.
  • Nothing grows with session length. The only in-process state is one remembered tree per store. Over 50 captures on a 50k-file repo, RSS went 158 MB → 171 MB peak → 129 MB, open file descriptors stayed at 8, and no temporary files were left behind.
  • Fewer processes. 6.8 Git processes per step instead of 10.6, and 2 instead of 200 to restore 100 files.
  • Linear path handling. Paths stream to update-index on stdin and are filtered with sets, replacing git add with a path list (quadratic in paths) and an includes scan over arrays.
  • Bounded packing. pack-objects runs in the background with 2 threads and a 64 MiB window (about 67 MB peak on the largest real store), at most every 10 minutes, only past 2048 loose objects or 16 packs, and never blocks a capture.
  • Never prunes. Snapshot trees have no refs, so git gc would treat every snapshot as garbage. Loose objects and old packs are removed only after the new pack index lists every one of them. On copies of real stores every object survived, and the largest went from 198 MiB to 88 MiB with 0 loose files left.

API changes for direct @opencode/core importers

Schema, protocol, the server API, session events, and the generated client are unchanged, so SDK users see nothing. Code importing @opencode/core modules directly sees three changes:

  • Git.Interface gains objects.pack. Hand-written Git.Service fakes need it; the one in this repo spreads the real service.
  • The createLLMEventPublisher input snapshot?: Snapshot.ID becomes pendingSnapshot?: Effect<Snapshot.ID | undefined>.
  • Git.OperationError["operation"] gains "pack".

Git tree functions now treat paths literally; v2 read *, ?, [ and a leading : as pathspec magic, which is how undo deleted :colon.txt. Snapshot stores on disk stay readable by older releases.

Change outline

Every write goes through a temporary index that is scanned, updated, and then renamed over the store's index, so processes never share index.lock and a killed writer never leaves a broken index.

 capture(scopes)
+  sweep temporary indexes older than an hour (age from the name); fix store config (once per store)
+  hard-link the store index -> temporary index (GIT_INDEX_FILE)
-  diff-files + ls-files -o          (against the shared index, under an in-process lock only)
+  diff-files + ls-files -o          (against the temporary index, :(literal) scope)
+  if the temporary index is the one this process installed last, nothing tracked changed,
+  and the source ignores every untracked path
+    return the remembered tree      (2-3 git calls instead of 5)
-  rm --cached / add --pathspec-from-file   (on the shared index; add is quadratic in paths)
+  update-index --force-remove / --add --remove --replace --stdin   (linear; second pass for gitlinks)
   write-tree                        (tree ID validated)
+  rename the temporary index over the store index (atomic; last writer wins)
+  on failure: only if Git reports the index unreadable, install the source's index and retry once

Restore batches work when the paths are independent of each other, and otherwise falls back to the original per-path order.

 restore(files)
-  for each path: ls-tree, then checkout or remove        (2 processes per file)
+  if paths are canonical ASCII and none contains another
+    one ls-tree + one checkout per source tree, then removals in parallel
+  else
+    original per-path order, with :(literal) pathspecs

The start snapshot overlaps the model request. It is still awaited before Step.Started, so it always exists before any local tool runs, and the wait stays cancellable.

 SessionStep.attempt
-  startSnapshot = capture()
-  llm.stream(...)
+  pendingStartSnapshot = fork(capture())
+  llm.stream(...)
+    per event, before Step.Started: join(pendingStartSnapshot)   (interruptible)
+    startAssistant -> publish Step.Started with the start snapshot
+  cancelled before the snapshot exists -> record nothing, like v2
   ...settle tools...
+  startSnapshot = join(pendingStartSnapshot)
   endSnapshot = capture()

Files and responsibilities:

 packages/core/src/
 ├── git.ts               # temporary-index capture, clean-capture memo, index recovery,
 │                        # atomic store creation, batched restore, objects.pack
 ├── snapshot.ts          # ..scope fix, rate-limited background packing
 └── session/runner/
     ├── step.ts          # start capture overlaps the provider request
     └── publish-llm-event.ts  # Step.Started awaits the pending snapshot
 packages/core/script/
+├── snapshot-parity*.ts / snapshot-parity-run.sh   # differential corpus vs a baseline worktree
+├── benchmark-snapshot*.ts                         # engine, session, and packing benchmarks

Measured against origin/v2, back to back on the same machine with the same scripts (benchmark-snapshot.ts, benchmark-snapshot-session.ts):

                                         v2          this PR
snapshot overhead per step, opencode     153 ms      75 ms     (10.6 -> 6.8 git calls)
snapshot overhead per step, 100k files   363 ms      171 ms
capture with nothing changed, opencode   58 ms       41 ms
restore 100 files, opencode              1.8 s       41 ms
restore 100 files, 100k files            2.4 s       46 ms
robustness probes (zeroed index, stale
  index.lock, ..scope, 4 processes x 25)  4 fail      4 pass

- Scan the temporary index that the capture writes, never the shared one, so a concurrent process cannot roll entries back into a snapshot.
- Sweep abandoned temporary indexes by the creation time in their name; a hard link keeps the store index's old mtime while in use.
- Rebuild the index only when Git itself reports it unreadable, and install the source index atomically instead of deleting first.
- Record a tracked file replaced by an embedded repository as a gitlink.
- Drop exclude mirroring, which changed how source negations apply.
- Keep step cancellation responsive while the start snapshot is captured.
- Restore checkouts before removals so one failed removal cannot block the rest.
- Rename internals: objects.pack (not compact), withTemporaryIndex, lastCaptures, attemptCapture, canBatchRestore, pendingSnapshot; fold the capture's seed into ignores.
…ntracked paths

- Without exclude mirroring, a path ignored only by the source's info/exclude is listed as untracked on every scan, which defeated the clean-capture memo and cost check-ignore, update-index, and write-tree on every step. When the index is unchanged and nothing tracked changed, one check-ignore confirms every untracked path is ignored and returns the remembered tree; the result is reused on a miss.
- Only tracked paths that became ignored need --force-remove; untracked paths are never in the index.
…he Windows command-line limit

- Record an embedded repository's gitlink by repeating its path in the same update-index call instead of a second pass.
- Read pack index object IDs in-process; packing runs a single pack-objects (show-index only for v1 or >2 GiB packs).
- Restore runs one ls-tree per source tree: paths as arguments while they fit in 24,000 characters, otherwise the whole tree, so Windows' 32,767-character limit is never hit.
- Diff passes --literal-pathspecs once instead of prefixing every path, keeping the command line as long as before.
Revert diffs every restored path, passing them as git diff arguments. Past Windows' 32,767-character command line the spawn failed, after the files were already restored. On Windows, a selection over 24,000 characters is now diffed in byte-sorted groups that share the patch cap, so output matches a single call. Other platforms keep a single call.
@nexxeln
nexxeln merged commit 801c152 into v2 Oct 8, 2026
21 of 22 checks passed
@nexxeln
nexxeln deleted the snapshot-lab branch October 8, 2026 12:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant