Skip to content

Commit the Mac LongMemEval-S harness and preregistration byte-exact (#274) - #380

Merged
seathatflowsinourveins merged 4 commits into
mainfrom
claude/memory-stack-longmemeval-harness-20260925
Sep 27, 2026
Merged

seathatflowsinourveins merged 4 commits into
mainfrom
claude/memory-stack-longmemeval-harness-20260925

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Scope

  • Commits the Mac LongMemEval-S retrieval harness, preregistration (A1-A16), eligible-question manifest and the agentmemory 0.9.29 lock into blueprints/memory-stack/longmemeval/, byte-exact, so the workstation can draft amendment A17 and run the confirmatory C3/C4 rerun against files in Git (host request [mac-coordinator] other: commit the memory-stack LongMemEval harness and preregistration for the NativeStack rerun #274).
  • Base commit: 5af288d968b39a1a083a908ddad0ff74c4681534
  • Lane: lane:foundation
  • Owned paths touched: blueprints/memory-stack/longmemeval/**, docs/decisions/2026-09-25-longmemeval-frozen-npm-lock.md, manifests/evidence.json (hot-file protocol re-registration only, rebased onto current main).

SOTA sources

  • LongMemEval (MIT, Di Wu) at 9e0b455f4ef0e2ab8f2e582289761153549043fc: the official repository lme_harness.py imports at run time (not vendored). README.md in the added directory maps each file to this pin.
  • Dataset: Hugging Face xiaowu0162/longmemeval-cleaned at revision 98d7416c24c778c2fee6e6f3006e7a073259d48f, longmemeval_s_cleaned.json, sha256 d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442, 277,383,467 bytes; fetched by hash at run time, never committed (MIT on its dataset card).
  • docs/lanes.md "Hot-file protocol": followed to rebase onto main and re-register this branch's files in manifests/evidence.json without hand-merging it.
  • Source of record for PREREGISTRATION.md only: the private agent-ecosystem repository at 2c3e09b, evals/longmemeval/PREREGISTRATION.md (blob f5c475db), the owner's own prior work. Every other copied file comes from the Mac run directory or its parent bench directory, as the directory's README.md "Files" table records per file; the two reconstructed/ files have no live source, and the README says how they were rebuilt.

Evidence-class table

Claim Evidence class Command / receipt
Each of the 13 copied files is byte-identical to its source source_review sha256 and cmp of PREREGISTRATION.md against agent-ecosystem@2c3e09b, and of the other 12 against the Mac run and bench directories; sources per file in the directory's README.md "Files" table. README.md, .gitignore, SHA256SUMS and the decision record were written for this PR, and reconstructed/* has no live source
SHA256SUMS matches the files in the directory source_review shasum -a 256 -c SHA256SUMS (exit 0)
The harness and summarizer run --help from a clean clone with no Mac-only path local_integration clean git clone + isolated HOME, official LongMemEval checked out at 9e0b455f4, .venv-official (Python 3.11.16); all 6 files exit 0
manifests/evidence.json re-registration is internally consistent local_integration python3 scripts/validate.py, python3 scripts/component_matrix.py --check-equivalent --write (no diff), python3 scripts/new_host_grand_list.py --write (no diff)
No secret or personal path was introduced local_integration Gitleaks via adoption/tools/gitleaks-guarded-macos dir . (exit 0, no leaks found) for secrets; the private-content scan in python3 scripts/validate.py (its PRIVATE_CONTENT patterns, passed) for home paths and similar private strings
No arm, model or embedding server ran for this commit none new n/a — this PR transfers files only; the Mac's descriptive C3 results are unchanged in convergence.json

Local commands run

$ git fetch origin && git rebase origin/main   # hot-file protocol; conflict in manifests/evidence.json resolved
  by taking origin/main's copy and re-registering this branch's 19 files with host_receipts.register_file
$ python3 scripts/component_matrix.py --write   -> {"flip_rule_violations": 0, "rows": 32, "status": "written"}, exit 0
$ python3 scripts/new_host_grand_list.py --write -> {"status": "written", "layers": 32, "winners": 66}, exit 0
$ python3 scripts/validate.py                   -> {"components": 69, "hashed_files": 7160, "profiles": 4, "receipts": 159, "status": "passed"}, exit 0
$ python3 scripts/validate_catalogs.py          -> {"repository_entries": 154, "unique_catalog_repositories": 148, ...}, exit 0
$ git diff --check origin/main...HEAD           -> exit 0, no output
$ (cd blueprints/memory-stack/longmemeval && shasum -a 256 -c SHA256SUMS)  -> 17 files, all OK, exit 0
$ adoption/tools/gitleaks-guarded-macos dir . --config .gitleaks.toml --max-target-megabytes 2 --redact --no-banner
  -> "no leaks found", exit 0
$ python3 -m unittest                           -> full suite as CI's "Test validation failure modes" step runs it;
  Ran 6453 tests, OK (skipped=861), 557.85 s, exit 0
$ env -i HOME=<isolated> PATH=/usr/bin:/bin <.venv-official python> {lme_harness.py,lme_harness.v1.py,lme_harness.v2.py,
  reconstructed/lme_harness.v1a.py,reconstructed/lme_harness.v3pre.py,lme_summarize.py} --help
  -> all 6 exit 0, from a clean clone of this branch with an isolated HOME holding only LongMemEval@9e0b455f4

Decision record

docs/decisions/2026-09-25-longmemeval-frozen-npm-lock.md — why the agentmemory lock and package.json are stored as *.frozen rather than under their npm names.

Not transferred

Copied faithfully from the directory's README.md ("Not committed, and why"):

  • The dataset and the official oracle file (data/): fetched by sha256 at run time, as [mac-coordinator] other: commit the memory-stack LongMemEval harness and preregistration for the NativeStack rerun #274 asks.
  • Run outputs: result rows (results*/, 470 rows per complete arm), the official runner's logs (about 291 MB each, holding dataset text), run logs (logs/; one holds an absolute home path), embedding caches (cache/, about 660 MB), model weights (models/) and the Python venv. [mac-coordinator] other: commit the memory-stack LongMemEval harness and preregistration for the NativeStack rerun #274 does not ask for them.
  • The two reports and their JSON: not requested. convergence.json carries their figures, and the summarizer here reproduces both byte for byte.
  • The two Codex reviews the preregistration cites, prereg-review-codex.md (A1-A7) and harness-review-codex.md (A10), and their event-stream logs: they fail this repository's private-content scan (a home path; the logs also hold session identifiers), so they cannot be committed byte-exact.
  • lme_summarize.v1.py (e01b78bc885e9458d0b95f5d851c45d1b63d2b897eec8b15884c1ed1b594a91f) and lme_summarize.v2.py (29f2f0b7fa5e7c426c8b14f37ebd910c9b19e9a82aa8a15e18de5feef9caa200), the copies A13 and A14 keep: neither produced a recorded report.
  • run_tail.sh (9e30e42ec09832d336d8a9d5b1a736fbfe466c9675e9602659a76f609980423f): it was superseded at 22:38 while still waiting and never ran an arm.
  • The inline launch commands: the A0 and A1 official runs, the arm chains started at 18:55, 18:57, 20:32 and 21:37, the embed server started at 21:35 and the proxy started at 22:38. They exist only in the coordinator session's transcript, with home paths; the directory's README arm table gives each one's time and arms.
  • Later preregistration revisions: the Mac's current file adds A16.1-A16.3 (sha256 a9b1db335eee1ff99d1883e048bcd8e443ca2afee3505a34fd874dd1ef412d5b, 37,885 bytes), and its first 33,527 bytes are PREREGISTRATION.md here. [mac-coordinator] other: commit the memory-stack LongMemEval harness and preregistration for the NativeStack rerun #274 asks for the A1-A16 revision. The workstation should draft amendment A17 knowing that these three further amendments already exist, written before any VelaNext run, on the Mac and in the private agent-ecosystem repository, but not in this repository.
  • Binaries: the ai-memory build, the iii engine and node_modules/, pinned above (by hash, in the directory's README).

Host evidence

Not applicable: this PR does not add or change files under evidence/hosts/. No host receipt accompanies it because no component ran: the PR transfers files and checks --help, and scripts/host_receipts.py record refuses a use receipt whose commands are only help or version calls (docs/contributing-evidence.md section 2). The mac-coordinator's receipts for component runs belong to #276 and #379.

Checklist

  • New/changed GitHub Actions are pinned to a full commit SHA with a version comment (no floating tags). (none added/changed)
  • New/changed workflows declare top-level permissions: contents: read (or a narrower, explicitly justified addition). (none added/changed)
  • No secrets are printed, logged or committed; no new required secret was added without a documented owner.
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved (not deleted, moved or overwritten).

Review requested

  • An Opus review ran at head 291f7daa: a read-only Opus subagent of the mac-coordinator session, on the same host and account, in a fresh clone. It is not cross-family. Its verdict and findings are in a PR comment.
  • The GPT-6 cross-family review ran from the workstation at 291f7daa. Its verdict was needs_changes, with one medium and one low finding, both README-only: the frozen harness's forced cleanup kills listeners it does not own, and a byte-identity claim also covered the reconstructed files. Commit 87e120f4 adds a shared-host warning and narrows the claim. It also folds in the Opus decision-record nit. No frozen file changed.
  • Requested: a GPT-6 re-check at 87e120f4. This PR merges only after that verdict, once the branch is updated from main under the hot-file protocol.

Host request: #274

🤖 Generated with Claude Code

…274)

Add blueprints/memory-stack/longmemeval/ for host request #274, so the
workstation can draft amendment A17 and run the confirmatory rerun against
files in Git:

- PREREGISTRATION.md with amendments A1-A16, sha256 1f3ca9d5..., from
  agent-ecosystem commit 2c3e09b;
- harness v3 (lme_harness.py), the kept v1 and v2 copies, summarizer v3,
  the A13 embedding cache proxy, the A1 oracle shim, and the three Mac
  launch scripts (mac-drivers/, records only);
- reconstructed/: harness v1a (B1, B2, D1) and v3-pre (C3 rows 1-31, D3),
  the two revisions that no longer exist as files on the Mac;
- eligible-manifest.json (871f5da1...);
- the agentmemory 0.9.29 lock and package.json the D arms used, stored as
  *.frozen because their npm names would fail the required osv-scanner and
  dependency-review checks (GHSA-45rx-2jwx-cxfr high, GHSA-8988-4f7v-96qf
  medium, via iii-sdk 0.11.2); docs/decisions/2026-09-25-longmemeval-
  frozen-npm-lock.md records why and what would overturn it.

Every copied file is byte-identical to its source; SHA256SUMS lists all
files in the directory. The README maps each Mac arm to the code that ran
it, gives the run-time pins (official code 9e0b455, dataset d6f21ea9...),
the --help results from a clean clone and what was not transferred.
No arm or benchmark ran for this commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Register the 18 files under blueprints/memory-stack/longmemeval/ and the
frozen-lock decision record with scripts/host_receipts.register_file
(hot-file protocol, last commit). component_matrix.py --write and
new_host_grand_list.py --write left their reports unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 27, 2026
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Opus review of head 291f7daa: approve with nits

The mac-coordinator session ran this as a read-only Opus subagent, on the same host and account, in a fresh clone of the branch. It is not independent across model families. The GPT-6 cross-family review is still requested (see the end of this comment).

#274 acceptance holds.

  • All four items are present.
  • The SHA256SUMS line for PREREGISTRATION.md is 1f3ca9d5…c11c2. shasum -a 256 -c and sha256sum --check --strict both report 17/17 OK.
  • Byte identity was reproduced on this Mac with cmp (0 for each):
    • PREREGISTRATION.md against agent-ecosystem@2c3e09b;
    • the 7 run-directory files and the 3 drivers;
    • the agentmemory lock pair.
  • The README's quoted hashes all match their files: the Mac's current a9b1db33… preregistration (whose first 33,527 bytes equal the committed file), summarizer v1/v2, run_tail.sh and the dataset (d6f21ea9…, 277,383,467 bytes).
  • --help from a clean clone, with a fake HOME holding only LongMemEval@9e0b455f: all six harness and summarizer entry points exit 0. The controls fail as documented.
  • manifests/evidence.json adds 19 entries (7141 → 7160), none removed or changed.
  • validate.py, validate_catalogs.py, component_matrix.py --check and new_host_grand_list.py --check all exit 0.
  • osv-scanner 2.6.0 --no-resolve reproduces the decision record: under npm names, exactly the two advisories; under .frozen names, no package sources found.

Findings and what was done

Severity Finding Action
should-fix PR body, SOTA sources: 2c3e09b was named as the source for "the preregistration and harness". Only 6 of the 15 files are in that tree, and some there are different revisions. Fixed in the body: 2c3e09b is the source of PREREGISTRATION.md only. The other files are per the README Files table.
should-fix PR body, evidence row 1 claimed "every committed file" is byte-identical, "see table below", and no table followed. Fixed: "each of the 13 copied files". It points to the README table and names the files with no live source.
nit The secrets row cited only Gitleaks. Fixed: it also cites the PRIVATE_CONTENT scan in scripts/validate.py.
nit The unittest line had no total. Fixed: 6453 tests, OK (skipped=861).
nit docs/decisions/2026-09-25-longmemeval-frozen-npm-lock.md lines 65–66 say the inventory's excluded list admits only deliberately vulnerable fixtures. It also holds three captured dvc.lock copies. The deciding reason is that dependency-review (and Socket and Scorecard) read package-lock.json by name whatever the inventory says. The Socket app is also missing from the overturn list. Pending. It will go in one follow-up commit (with re-registration), batched with the next item.
nit README.md publishes private-repository details beyond the four items: a PR number, two commit ids, the existence of a v4 harness and the session's edit history. It is accurate and quotes no file contents. Pending the owner's decision: keep, or drop the PR number and commit ids. README.md is listed in SHA256SUMS, so an edit also updates that line and its registration.

The coordinator also corrected one line the review didn't flag. The "Not transferred" item said A16.1–A16.3 were "uncommitted". They are committed in the private agent-ecosystem repository, just not in this repository.

Requested: the GPT-6 cross-family review from the workstation, at head 291f7daa. If the pending nit commit lands, its delta covers only the decision-record wording, possibly the README, and the matching SHA256SUMS and evidence.json lines. No copied file changes. This PR merges only after the GPT-6 verdict, and after a fresh hot-file check against main.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

GPT-6 cross-family review at 291f7daa: VERDICT: needs_changes (1 medium, 1 low)

Setup: gpt-6-astra, effort max, read-only, run from the workstation (session mac-memory-layer-plan). It tested behaviour rather than only reading.

[medium] blueprints/memory-stack/longmemeval/lme_harness.py:451: slot cleanup kills listeners without checking ownership. Starting an agentmemory arm can terminate unrelated services on the host.

  • The pre-start cleanup at :479 is unconditional, and :457 sends SIGKILL.
  • Probe: the original function with mocked lsof and os.kill returned kill calls=[(424242, 9)] for an unrelated listener. No real process was signalled.
  • Fix direction: keep the frozen bytes, since this PR is a byte-exact transfer. Add a README warning not to run lme_harness.py directly on a shared host, and put the ownership guard or isolated execution path in the workstation rerun's own code (amendment A17), not in the frozen file.

[low] README.md:301: the blanket "The files match the Mac originals" claim also covers the two reconstructed files (reconstructed/lme_harness.v1a.py, reconstructed/lme_harness.v3pre.py). Narrow it to the 13 copied files.

Verified:

  • sha256sum -c SHA256SUMS exits 0. All 17 entries match their file contents, their sizes and manifests/evidence.json. The PREREGISTRATION.md digest is 1f3ca9d554cff24dc785c35f275335027a831dd386f460890f8a3082a07c11c2, and amendments A1 to A16 are present.
  • The manifests hold exactly 470 full-track and 419 official-track ids. The dataset digest d6f21ea9…, recall_all@5 and the embedder identities agree with the records. The metric calls the pinned official evaluator.
  • Arm provenance:
    • B1, B2 and D1 map to v1a. Removing the six documented A8 lines from v1 reproduces v1a byte for byte.
    • C1, C2 and D2 map to v2.
    • C3 maps to v3-pre followed by v3.
    • Both reconstructions are labelled.
  • The agentmemory 0.9.29 and iii-sdk 0.11.2 lock entries match their registry integrity fields. The decision record's advisories match their published records.
  • No /Users/ paths, emails or credential patterns appear. Gitleaks: no leaks. validate.py passes with 7,160 files hashed, and git diff --check is clean.

Not established here: python3 lme_harness.py --help exits 1 in a bare environment (env -i, no numpy). The dependencies are documented, but a passing --help was not independently reproduced, and there is no dry-run mode.

For the workstation rerun (A17), not blockers for this PR:

  • The runner selects the old 19b6429 build; the rerun needs the official v2.4.0 control.
  • The summarizer implements A1 to A14 at α 0.05, not A15.2's 0.0333.
  • The Mac drivers are non-portable.
  • The MiniLM downloads were not revision-pinned.

…claims (#380 review)

- README: the agentmemory arms clean up by force and kill processes they do
  not own. v3 and v3-pre SIGKILL every listener on a slot's ports; v1, v1a
  and v2 pkill the install's index path. Run them only where nothing else
  uses those ports or that prefix. The ownership guard belongs in the A17
  rerun's own code (GPT-6 review, medium).
- README: the byte-identity claim now covers the 13 copied files, not the
  two reconstructed ones (GPT-6 review, low).
- Decision record: reject the inventory's excluded list for the real
  reason. dependency-review reads package-lock.json by name and fails on
  high severity. Note that the lock is the first file hidden by renaming
  that is not a test fixture, and add the Socket app to the overturn list
  (Opus review, nit).
- Update the README line in SHA256SUMS and re-register three
  manifests/evidence.json entries. No frozen file changed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Fix commit 87e120f4 for the GPT-6 review at 291f7daa. The delta covers only README.md, SHA256SUMS (the README line), the decision record and 3 manifests/evidence.json entries. No frozen file changed, and shasum -a 256 -c SHA256SUMS gives 17/17 OK.

  • [medium] listener kills: added a README warning under the top bullets, verified against the frozen code. lme_harness.py and reconstructed/lme_harness.v3pre.py send SIGKILL to every listener on a slot's REST, stream and engine ports, before the slot starts and at teardown. With the default four slots those are 3611–3612, 3711–3712, 3811–3812, 3911–3912 and 49634/49734/49834/49934; the engine ports are in the dynamic range. lme_harness.v1.py, lme_harness.v2.py and reconstructed/lme_harness.v1a.py run pkill -f and then pkill -9 -f on the agentmemory index path. The warning says not to run them on a shared host, and that the ownership guard or isolated path belongs in the A17 rerun's own code.
  • [low] README claim: the Verifying table now claims byte identity for the 13 copied files only. The two reconstructed/ files are pinned by SHA256SUMS with their provenance.
  • Opus nit, decision record: the excluded-list alternative is now rejected for the real reason. dependency-review reads package-lock.json by name, whatever the inventory says, and fails on high severity. The record also notes this is the first non-fixture file hidden by renaming, and adds the Socket app to the overturn list.

Checks at 87e120f4, all exit 0:

  • validate.py and validate_catalogs.py;
  • component_matrix.py --check (32 rows) and new_host_grand_list.py --check;
  • git diff --check;
  • 220 targeted tests (lockfile coverage, adoption docs consistency, verdict review gate, workflow security coverage): OK, 1 skipped;
  • Gitleaks on the changed files: no leaks.

Still open, pending the owner: whether the README keeps the private-repository PR number and commit ids. Before merging, the branch will be updated from main under the hot-file protocol, since main moved one commit and touched manifests/evidence.json. The trial merge is clean.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

GPT-6 cross-family re-check at 87e120f4: VERDICT: accept. No new findings.

  • [medium] Resolved. README.md:21 warns against running the harness on a shared host, and it assigns the ownership guard or isolated execution path to A17's own code. The frozen code confirms both claims:
    • v3 and v3-pre am_teardown SIGKILL every listener on the stated ports;
    • v1, v1a and v2 am_reap run pkill -f and then pkill -9 -f on the agentmemory index path.
  • [low] Resolved. README.md:315 limits byte identity to the 13 copied files and names the two reconstructions separately.
  • Decision-record edit: accurate and sourced. The OSV exclusion needs a fixture reason and existing evidence, and dependency-review keeps fail-on-severity: high.

Verified at 87e120f4:

  • The diff from 291f7daa changes 4 files. All 16 non-README checksum-listed files are byte-identical, and only the README line of SHA256SUMS changed.
  • sha256sum -c SHA256SUMS: 17/17 OK.
  • All 19 relevant manifests/evidence.json entries match.
  • validate.py passes (7,160 files).

It merges in a main-free window of the token lane's queue, which the workstation will relay. The README owner question (private agent-ecosystem ids) stays with the user.

Take main's manifests/evidence.json and re-register this branch's 19 files
with host_receipts.register_file. component_matrix and new_host_grand_list
were rewritten with no change, and validate.py passed with 7200 hashed
files. No transferred file changed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins enabled auto-merge (squash) September 27, 2026 05:29
@seathatflowsinourveins
seathatflowsinourveins merged commit 1aa9776 into main Sep 27, 2026
24 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/memory-stack-longmemeval-harness-20260925 branch September 27, 2026 05:47
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
Replayed onto main: the bootstrap.md note merges after #366's OTEL tool-details
sentence, manifests/evidence.json is re-registered, and the matrix and grand
list are regenerated. validate FAILS=0; 438 related tests OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
…be, measured RTK rule) (#378)

* SubagentStart token-lanes carrier for non-blind subagents

Adds a stdlib, fail-open SubagentStart hook that returns
adoption/hooks/claude/token-lanes-block.md as additionalContext to every
subagent except agent types starting "blind-", registers it in the Claude
settings template beside ai-memory, and installs both files through the
existing sha256-checked guard step.

docs/token-session-handbook.md "Token lanes carried into subagents" is the
source of truth the block copies line for line. Its RTK line is the measured
RTK 0.50.0 rule (rtk hook check and the rtk hook claude PreToolUse hook;
rewrite_multiline_block, rtk-ai/rtk#3319): && chains and multi-line blocks
are rewritten segment by segment; $(...), backticks, <(...), file redirects
and heredocs are never rewritten; gh pipelines are not rewritten.

Native probe (Claude Code 2.1.283, Agent-tool Haiku children): with the hook
the child found and quoted the block's first 60 characters exactly; without
it, and for a blind-judge child with it, the child found nothing.

Decision record: docs/decisions/2026-09-27-token-lanes-subagent-start.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Native probe and RTK hook-check receipt for the token-lanes carrier

evidence/artifacts/token-lanes-subagent-start-20260927/ retains, sanitized,
the five native probe runs (Claude Code 2.1.283, Agent-tool Haiku children:
without the hook, with it, and a blind-judge child with it, plus the failed
v1 question kept as a failed design) and 17 RTK 0.50.0 inputs through both
`rtk hook check` and the `rtk hook claude` PreToolUse processor. Raw CLI
output and child transcripts stay private and are identified by SHA-256 and
counts. Evidence class local_integration; not native_proven for Workflow
children, which have not run with the hook.

The handbook's RTK sources paragraph and the decision record now cite the
receipt instead of unretained observations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Count the carrier block's budget in UTF-8 bytes (cross-family review)

The GPT-6 review of #378 at 85dd27c found that the text test checked
len(block), a character count, although the decision record states a
3,500-byte budget: a 3,280-character, 3,520-byte block would pass. The test
now measures the file's bytes through fits_budget() and adds that case as a
regression test.

Controls: a character-counting fits_budget fails the regression test, and a
block padded to 3,560 bytes (3,360 characters) fails the byte check but
passed the old character check. The committed block is ASCII, 3,159 bytes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Register the carrier's files on main 1aa9776 (after #366 and #380)

Replayed onto main: the bootstrap.md note merges after #366's OTEL tool-details
sentence, manifests/evidence.json is re-registered, and the matrix and grand
list are regenerated. validate FAILS=0; 438 related tests OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
Verified each finding against the frozen code and the merged #380 state before
editing; only README.md, the decision record and evidence.json change, no
frozen file touched.

[should_fix] The "outside D2h/C4" line for the three still-excluded Modelfiles
was misleading: run_velanext.sh's `gates` command is mandatory before any arm
(require_gates, called by cmd_cpu/cmd_arms/cmd_a16) and smoke-tests all four
reranker-LLM tags plus K1 before writing logs/gates/passed.json, so no D2h, C4
or K1 run passes gates from this directory without all three files. README
title now matches the PR body's scope (D2h, C4, the A16 arms).

[should_fix] #380 has merged to `main` since this PR opened; removed the two
"not yet merged"/"lives only on #380's branch" statements.

[should_fix] Added a dated 2026-09-27 addendum to
docs/decisions/2026-09-25-longmemeval-frozen-npm-lock.md, extending its scope
from the two agentmemory/*.frozen files to this driver's six additional
.frozen files. Evidence gathered directly: osv-scanner 2.6.0 against plainly
named scratch copies (official.lock.txt: 114 unique advisories across 9 of 64
packages, dominated by nltk/transformers/torch/pillow; mempalace.lock.txt's
chromadb: 4, including a pre-auth code-injection advisory; embed.lock.txt: 2;
the npm lock repeats the original record's 2 OpenTelemetry advisories;
hindsight.lock.txt and build.lock.txt: none), plus the same discriminating
control the original record used (.frozen names: "No package sources found",
exit 128). Also corrected README.md's claim that the .in files are "unpinned"
(official.in alone pins 25 of 28 lines) and that TRACKED "does not match"
them by design (it is a basename/directory-name technicality, not a
deliberate exemption).

[should_fix] Added a "Before anything runs" section: eligible-manifest.json,
the agentmemory install files, embed_gates.json, reference/v3-check-ids.txt
and the six renamed locks are all read by setup_velanext.sh/run_velanext.sh,
but are not available from this directory alone. The table gives each
file:line that reads it and where to get it (two are byte-identical to files
#380 now carries on main; two are not present anywhere in this repository and
need the private source; the locks need isolated-prefix restore commands).

[nit] Split the evidence-class row that credited "reproducing... with inert
probes" to source_review; that reproduction is the GPT-6 review's own
independent observation, now cited by its actual PR-comment URL (confirmed
via `gh api .../issues/386/comments`, which also confirmed the review's 1
high / 2 medium / 1 low was carried through in full, nothing dropped or
narrowed).
[nit] Qualified two over-broad safety claims: stop_bg's group-mode signals
(the CPU stream) are unconditional, not same_process-guarded; setup_velanext.sh's
`kill "$oll_pid"` can hit a reused PID in the failed-bind case, the same gap
attributed to mp_kill_mines.
[nit] Corrected the LME_HOME/Path.home() claim: LME_HOME itself defaults to
this directory (lme_harness.py:87), not Path.home(); only the AE-derived
paths do.
[nit] Added .gitignore to the SHA256SUMS provenance exception, and noted the
plain-name/`.frozen`-name distinction for git-show provenance commands.
[nit] Added README notes for the three shared-host/portability gaps that are
frozen-code findings, not README claims: two vendor harnesses launch without
env -i (run_velanext.sh:563,596-597); lane_tools.py's v3-shim/fetch-hf
--pin-main mutate shared state outside their lane-owned wrapper paths; the X
arm's kalm-small reranker loads with trust_remote_code=True (mitigated by a
pinned revision, local_files_only and HF_HUB_OFFLINE).
[hygiene] Deleted the blob:none clone of the private agent-ecosystem
repository left in the review session's scratchpad; nothing referenced it.

Re-merged origin/main (three more commits landed) under the hot-file
protocol before this commit: took main's manifests/evidence.json and
re-registered this branch's 28 v4 files plus the two edited entries above.

Checks after this round: shasum -a 256 -c SHA256SUMS (27/27 OK); python3
scripts/validate.py ({"hashed_files": 7278, "status": "passed"}); python3
scripts/validate.py --scan-file on the three changed files (passed); gitleaks
over blueprints/memory-stack/longmemeval/v4 and docs/decisions (no leaks);
python3 scripts/host_receipts.py validate (passed); python3
scripts/component_matrix.py --write and new_host_grand_list.py --write (no
unexpected diff); git diff --check; the four unittest modules the PR body
cites plus tests.test_validate.ScanFileForPrivateContentTests (all OK,
221+9 tests) -- all exit 0.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
…4) (#386)

* Commit the LongMemEval v4 driver byte-exact for the A17 rerun (D2h, C4)

blueprints/memory-stack/longmemeval/v4/: the harness v4 driver (lme_harness.py,
lane_tools.py, lme_summarize.py, pins.json), the v3 equivalence reference
(reference/lme_harness_v3.py), the X-arm and embedding-server tools
(rerank_stage.py, embed_server_st.py, gguf_embed_front.py, mteb_lmeb.py), the
VelaNext drivers (run_velanext.sh, setup_velanext.sh), PREREGISTRATION.md
(hash-verified against a9b1db33...), the k1-subset.json question-id subset,
the five requirements/*.in + *.lock.txt(.frozen) pairs, and
vendor/agentmemory-repo-package-lock.json.frozen. Extracted byte-exact from
the private agent-ecosystem repository, branch
claude/longmemeval-velanext-lane-20260925, commit 576689a, per host request
#274 and #384, following the pattern of #380 (not yet merged).

The five requirements/*.lock.txt files match this repository's TRACKED
lockfile pattern and are stored as .frozen for the same reason #380 freezes
its npm lock (docs/decisions/2026-09-25-longmemeval-frozen-npm-lock.md,
still on #380's branch); official.lock.txt additionally pins nltk==3.9.1,
which would also trip the repo-wide ignore-scope test if scanned live.
vendor/agentmemory-repo-package-lock.json.frozen does not itself match that
pattern but is frozen under the same rationale per this task's directive.

README.md corrects the task's premise that C4 runs through rerank_stage.py:
C4 (aimem-qwen3-rerank) is implemented in lme_harness.py itself via its LLM
reranker path (LLM_URL, A12); rerank_stage.py implements the separate,
non-decided X arms. It also documents each teardown function's actual
ownership checks (am_teardown, mp_kill_mines, run_velanext.sh's
stop_bg/same_process) against the code, not its docstrings, and lists what
was left out of transfer scope and why (eligible-manifest.json, the
agentmemory install lock and run_official_oracle.py are byte-identical to
#380's copies and are not duplicated; embed_cache_proxy.py at 576689a is a
changed v4 revision not in this task's requested file list).

Checks: private-content scan (24/24 new source files plus README.md and
SHA256SUMS, individually and combined), gitleaks-guarded-macos (no leaks),
SHA256SUMS self-check, component_matrix.py --check, new_host_grand_list.py
--check, validate_catalogs.py, git diff --check, scripts/validate.py (7173
hashed files), and the targeted unittest modules (221 tests, 1 skipped, all
passing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* README fixes for the GPT-6 review at 97697b8 (#386)

Verified each finding against the frozen code before editing; only README.md,
SHA256SUMS and one new file change, nothing frozen is touched.

[high] lme_harness.py's teardown does not establish ownership. _ours()
(:698-701) accepts any command merely containing the substring "agentmemory",
or the configured iii path/basename; am_teardown() (:749) also treats an
empty `ps` result as "ours"; am_start() (:797-800) calls that same teardown
with proc=None before it has started anything, so it can SIGKILL a foreign
process before owning anything. The README wrongly called the empty-command
case "the one gap" and said the functions check ownership; it now describes
the actual name/liveness matching and ends with a plain warning: do not run
lme_harness.py, run_velanext.sh or setup_velanext.sh on a shared host. A17's
runner is where isolated ports/stores and identity-checked cleanup belong.

[medium] Two warnings added, both verified by reading the code directly:
- setup_velanext.sh:220's readiness loop accepts any Ollama server that
  answers on port 11438, with no check that it is the oll_pid process just
  started; :222-224 then pulls models, and :228-229 creates aliases, against
  whichever server answered.
- lme_harness.py's mp_kill_mines() (:1149) trusts mp_mine_pids() (:1045),
  which only checks _pid_alive() (:1035, kill(pid, 0)) with no identity
  check, then killpg()s the whole process group; a reused PID's group can be
  hit.

[low] pins.json:146 and run_velanext.sh:654 confirm qwen3.5-9b-64k.Modelfile
is a C4 prerequisite (also used by the A15.2 F arms and the pooled
aimem-qwen3-rerank stage), not outside D2h/C4 as the README said. Transferred
it byte-exact from the private agent-ecosystem repository at 576689a
(evals/longmemeval/ollama/qwen3.5-9b-64k.Modelfile, blob de951c9, git
hash-object verified against the source blob id), added to SHA256SUMS and
the Files table. The other three Modelfiles stay excluded and are now named
individually with size/sha256/blob: lfm2.5-2.6b-64k.Modelfile (H3),
nemotron-3.5-lightning-30b-a3b-64k.Modelfile (H2), qwen3.6-35b-a3b-64k.Modelfile
(H1, and K1's LLM via H1_BUILD at run_velanext.sh:68) -- verified against
lme_harness.py's H_RERANK_MODELS and run_velanext.sh's matching loop, not
asserted from memory.

Checks: shasum -a 256 -c SHA256SUMS (27/27 OK), python3 scripts/validate.py
--scan-file on README.md/SHA256SUMS/the new Modelfile, and
adoption/tools/gitleaks-guarded-macos over the whole directory (no leaks).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* README: fix the evidence-class table's own source list (#386)

Line ~289's "teardown behavior described above" row still named only the
pre-fix set of functions (am_teardown, stop, mp_kill_mines, stop_bg/
same_process). Updated it to list every function and file the rewritten
shared-host warning now cites (_ours, am_teardown, am_start,
mp_mine_pids/_pid_alive/mp_kill_mines, stop_bg/same_process,
setup_velanext.sh's readiness loop), so the table matches the prose above it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* README/decision-record fixes for the second review round (#386)

Verified each finding against the frozen code and the merged #380 state before
editing; only README.md, the decision record and evidence.json change, no
frozen file touched.

[should_fix] The "outside D2h/C4" line for the three still-excluded Modelfiles
was misleading: run_velanext.sh's `gates` command is mandatory before any arm
(require_gates, called by cmd_cpu/cmd_arms/cmd_a16) and smoke-tests all four
reranker-LLM tags plus K1 before writing logs/gates/passed.json, so no D2h, C4
or K1 run passes gates from this directory without all three files. README
title now matches the PR body's scope (D2h, C4, the A16 arms).

[should_fix] #380 has merged to `main` since this PR opened; removed the two
"not yet merged"/"lives only on #380's branch" statements.

[should_fix] Added a dated 2026-09-27 addendum to
docs/decisions/2026-09-25-longmemeval-frozen-npm-lock.md, extending its scope
from the two agentmemory/*.frozen files to this driver's six additional
.frozen files. Evidence gathered directly: osv-scanner 2.6.0 against plainly
named scratch copies (official.lock.txt: 114 unique advisories across 9 of 64
packages, dominated by nltk/transformers/torch/pillow; mempalace.lock.txt's
chromadb: 4, including a pre-auth code-injection advisory; embed.lock.txt: 2;
the npm lock repeats the original record's 2 OpenTelemetry advisories;
hindsight.lock.txt and build.lock.txt: none), plus the same discriminating
control the original record used (.frozen names: "No package sources found",
exit 128). Also corrected README.md's claim that the .in files are "unpinned"
(official.in alone pins 25 of 28 lines) and that TRACKED "does not match"
them by design (it is a basename/directory-name technicality, not a
deliberate exemption).

[should_fix] Added a "Before anything runs" section: eligible-manifest.json,
the agentmemory install files, embed_gates.json, reference/v3-check-ids.txt
and the six renamed locks are all read by setup_velanext.sh/run_velanext.sh,
but are not available from this directory alone. The table gives each
file:line that reads it and where to get it (two are byte-identical to files
#380 now carries on main; two are not present anywhere in this repository and
need the private source; the locks need isolated-prefix restore commands).

[nit] Split the evidence-class row that credited "reproducing... with inert
probes" to source_review; that reproduction is the GPT-6 review's own
independent observation, now cited by its actual PR-comment URL (confirmed
via `gh api .../issues/386/comments`, which also confirmed the review's 1
high / 2 medium / 1 low was carried through in full, nothing dropped or
narrowed).
[nit] Qualified two over-broad safety claims: stop_bg's group-mode signals
(the CPU stream) are unconditional, not same_process-guarded; setup_velanext.sh's
`kill "$oll_pid"` can hit a reused PID in the failed-bind case, the same gap
attributed to mp_kill_mines.
[nit] Corrected the LME_HOME/Path.home() claim: LME_HOME itself defaults to
this directory (lme_harness.py:87), not Path.home(); only the AE-derived
paths do.
[nit] Added .gitignore to the SHA256SUMS provenance exception, and noted the
plain-name/`.frozen`-name distinction for git-show provenance commands.
[nit] Added README notes for the three shared-host/portability gaps that are
frozen-code findings, not README claims: two vendor harnesses launch without
env -i (run_velanext.sh:563,596-597); lane_tools.py's v3-shim/fetch-hf
--pin-main mutate shared state outside their lane-owned wrapper paths; the X
arm's kalm-small reranker loads with trust_remote_code=True (mitigated by a
pinned revision, local_files_only and HF_HUB_OFFLINE).
[hygiene] Deleted the blob:none clone of the private agent-ecosystem
repository left in the review session's scratchpad; nothing referenced it.

Re-merged origin/main (three more commits landed) under the hot-file
protocol before this commit: took main's manifests/evidence.json and
re-registered this branch's 28 v4 files plus the two edited entries above.

Checks after this round: shasum -a 256 -c SHA256SUMS (27/27 OK); python3
scripts/validate.py ({"hashed_files": 7278, "status": "passed"}); python3
scripts/validate.py --scan-file on the three changed files (passed); gitleaks
over blueprints/memory-stack/longmemeval/v4 and docs/decisions (no leaks);
python3 scripts/host_receipts.py validate (passed); python3
scripts/component_matrix.py --write and new_host_grand_list.py --write (no
unexpected diff); git diff --check; the four unittest modules the PR body
cites plus tests.test_validate.ScanFileForPrivateContentTests (all OK,
221+9 tests) -- all exit 0.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Correct the addendum's dependency-review claim and a citation (#386)

Self-review of the previous commit found two accuracy issues before they
reached a second reviewer:

- The 2026-09-25 decision record's new addendum claimed that under plain
  names, the required `dependency-review` check would fail this pull request
  on specific advisories. That does not follow: `dependency-review` reads
  GitHub's native dependency graph, which recognizes manifests by fixed
  per-ecosystem names (`requirements.txt`, `package-lock.json`, ...), not by
  this repository's `TRACKED` pattern. None of the six files' plain names
  (`official.lock.txt`, `mempalace.lock.txt`, `agentmemory-repo-package-lock.json`,
  and so on) is one of those recognized names, so `dependency-review` would
  not see any of them regardless of the `.frozen` suffix -- unlike the
  original record's own file, which was named exactly `package-lock.json`.
  The gate these five Python locks actually escape by staying `.frozen` is
  `osv-scanner`, via the `TRACKED`-forced inventory entry. Reworded the
  addendum's "Under their plain names" paragraph and the PR body's matching
  evidence row accordingly.
- README.md's new "Before anything runs" table cited "`:130` area's sibling
  installs" for where `setup_velanext.sh` reads the five lock files under
  their plain names; `:130` is the unrelated `eligible-manifest.json` hash
  check. The actual citation is `venv()` and its five calls at `:94-109`.

Also fixed the pr-body.md line describing this round's `origin/main` merge:
it was a clean `git merge` with no conflict (git's 3-way merge resolved
`manifests/evidence.json` on its own), not a manual `checkout --theirs` +
re-registration as the first merge round used; reworded to say so, and
pointed the decision-record link at a resolvable full GitHub URL instead of
a relative path that would not resolve in a PR body.

Rechecked after these edits: shasum -a 256 -c SHA256SUMS (27/27 OK); python3
scripts/validate.py ({"hashed_files": 7278, "status": "passed"}, unchanged
file count); python3 scripts/validate.py --scan-file on the three changed
files (passed); gitleaks over both changed directories (no leaks); python3
scripts/host_receipts.py validate, component_matrix.py --write and
new_host_grand_list.py --write (no further diff beyond the two hash/byte
updates); git diff --check -- all exit 0.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Correct the OSV ignore claim: the repository's ignores are global (#386)

OSV-Scanner 2.6.0 matches an [[IgnoredVulns]] entry by advisory ID and
expiry only (internal/config/config.go:104-112, ShouldIgnore), so the two
entries written for the Lumibot lock also filter the v4 locks:
GHSA-8mgp-746c-j5xp on nltk 3.9.1 (official.lock.txt) and
GHSA-h35f-9h28-mq5c on setuptools 81.0.0 (embed.lock.txt).

Re-scanned with an empty config: official 115 unique advisories (114
with the config) and embed 3 (2). The addendum table and README now give
both counts and name the filtered IDs. No frozen file changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: seathatflowsinourveins <234074349+seathatflowsinourveins@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant