Repository navigation
docs: four-layer memory in the session (enforced) and in the product (mapped) - #1313
Conversation
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Warning Review limit reachedNext included review available in 4 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (4)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
sakibsadmanshajib
left a comment
There was a problem hiding this comment.
Independent review, pipeline stage 6
Read against the pushed diff at a591e9b, not the description, with every factual claim checked against the repository, the ledger, the vault and the live issue list.
Streams
- Plain verification pass (this review): RAN.
- Antigravity (
agy, gemini-3.1-pro-high, high effort), framed as a pre-merge review: RAN. It independently found the MultiEdit fail-open and the CLI skip divergence, and added the hyphenated-identifier regex note. Its remaining conclusions matched my own measurements. - CodeRabbit: SKIPPED, not a pass. The bot review on this pull request stopped at "Review limit reached, next included review available in 28 minutes" and produced no findings, and
coderabbit pullrequest 1313 --show-promptsreturns "No CodeRabbit all-comments prompt was found", so there is nothing to read either way. - Codex connector: SKIPPED, usage limits reached, as expected until 2026-09-10.
Claims checked
CONFIRMED:
- CoALA's four layers are working, episodic, semantic and procedural, and arXiv:2309.02427 is the right identifier. The skill leads with those four and marks entity/profile and reflection/consolidation as Hive extensions with no paper behind them, which is the accurate framing. No fifth or sixth layer is smuggled back in.
- The hook is wired into the existing
Write|Edit|MultiEditPreToolUse matcher in.claude/settings.json, immediately aftersecrets-scanner.js. node .claude/hooks/hooks.selfcheck.jsreports 34 of 34 passed. I ran this branch's guards against the live ledger, not the author's transcript.- The CLI path behaves as described: it blocks a fabricated id, warns on a REVOKED entry, and reports D-058 as the highest recorded id. D-058 is in fact the highest, and the ledger has 58 contiguous entries.
- Issues #1308, #1309, #1310, #1311 and #1312 all exist, are open, and their titles match the product table in the description.
- Both vault documents exist, and the six-layer one carries a frame-correction callout at the top as the description claims.
- Every
D-NNNoccurring anywhere in the repository outside the ledger resolves to a real entry, so the new gate has no blast radius on existing files. - The product half changes no product code and is scoped as a map plus issues rather than as shipped behaviour. Spot check of its headline finding holds:
/internal/user-memories/is registered in the control-plane router, the recall path exists inapps/edge-api/internal/chat/memory.go, and no caller of the write side appears anywhere in the tree. - No
.wolf/telemetry is committed and.wolf/buglog.jsonlis untouched, per the openwolf contract. No dash punctuation in the added prose.
UNSUPPORTED:
- "runs as a PreToolUse hook on every Write, Edit and MultiEdit". MultiEdit is not covered in the Claude Code payload shape, measured, exit 0. See the thread on the content extraction.
- "On 2026-08-28 ... a
D-030that did not exist at the time". D-030 has been in the tracked ledger since 2026-08-25. The incident being described is 2026-07-30. See the thread on the header comment. - "the CLI and the gate can never disagree about what counts as a valid citation". True for citation validity, false for the skip list, which only the hook path implements.
Minor, no thread opened:
- The description says six new self-check cases. It is five per payload shape and ten in total, which is what the jump from 24 to 34 reflects.
- The skill section does not mention the
.claude/hooks/skip that the code implements and the description documents. - Nothing outside Write, Edit and MultiEdit is covered, so a Bash heredoc bypasses the gate completely. Worth one honest sentence under a heading that promises enforcement rather than good intentions.
Verdict: NO MERGE
The direction is right and most of it is verified true. Two findings must land first, and both are small:
- MultiEdit fail-open, because an "enforced" claim that is false for one of three named tools is the specific defect that makes a document worse than no document.
- The mis-dated D-030 incident, because a guard against fabricated citations must not ship a fabricated-looking citation in its own justification, and because a reader greps D-030, finds it, and stops trusting the rest.
Recommended in the same pass, cheap and it removes a way this gate can cause the harm it prevents: the stale-worktree hard block, and the two missing dead-status words.
Everything else on this list is optional.
…1337) Closes #1333. ## The hole `.claude/hooks/secrets-scanner.js` is registered in `.claude/settings.json` on the `Write|Edit|MultiEdit` matcher, but it only ever saw the first two. It resolved an edit array from the top-level `data.edits` field, which is Cursor's file-edit payload shape. Claude Code nests that array at `tool_input.edits`. On a MultiEdit, `ti.content`, `ti.new_string` and the top-level `data.edits` are all absent, so `content` resolved to the empty string, the `if (!content ...) process.exit(0)` guard fired, and the hook exited 0 with no output. Exit 0 and no output is exactly what a clean scan looks like. A MultiEdit could therefore write an OpenAI key, a private key, an AWS access key or a hardcoded password with the guard reporting nothing at all. Write and Edit stayed covered throughout, since those populate `tool_input.content` and `tool_input.new_string`. ## Fix The scanner consults both shapes, preferring the Claude Code one: ```js const editsSource = Array.isArray(ti.edits) ? ti.edits : Array.isArray(data.edits) ? data.edits : []; ``` ## Pinning it `.claude/hooks/hooks.selfcheck.js` already fed every guard both harness payload shapes, but its Claude Code write case used `Write`, whose text lives in `tool_input.content`. The nested-array path was never exercised, which is why a harness that predates this bug did not catch it. A third shape, `claude-multiedit`, is added, and the secrets-scanner cases now run under all three. The shell-command cases (bash-safety, commit-guard) stay on the two shapes that apply to them rather than being pointlessly duplicated. Mutation check, run before opening this PR: reverting `editsSource` to the top-level-only form turns the run red at 25/29, exit 1, with the four failures all under `claude-multiedit`: ``` FAIL secrets-scanner block aws key (claude-multiedit): exit 0, want 2; missing "BLOCKED"; missing "AWS access key" FAIL secrets-scanner warn api_key assignment (claude-multiedit): missing "SECRET WARNING"; missing "Possible API key assignment" FAIL secrets-scanner block go-style password assignment (claude-multiedit): exit 0, want 2; missing "BLOCKED"; missing "Hardcoded password" FAIL secrets-scanner warn go-style api_key assignment (claude-multiedit): missing "SECRET WARNING"; missing "Possible API key assignment" ``` With the fix restored: 29/29. All fixture credentials are synthetic and assembled at runtime, so no literal appears in the committed file: the AWS fixture is `'AKIA' + 'A'.repeat(16)` and the generic one is `'k'.repeat(20)`. No real credential is committed here, as a fixture or otherwise. ## The self-check was not run by anything Nothing in `.github/`, `Makefile` or `package.json` invoked `hooks.selfcheck.js`. A harness that nothing runs is not a gate, and that is the second half of how this survived. It is now a step in the Repo policy lints required check, next to `lint-workflow-check-names.mjs`, `lint-deploy-paths-filter.mjs`, `newest-covered-commit.test.mjs` and `test-set-compose-project-name.sh`, which live there for the same reason. The `changes` filter's extension-deny arm runs first and matches `.js`, so a `.claude/hooks/` change can never set `run=false` and skip this. ## Commit-time gate compliance The ECC `pre-bash-commit-quality` hook blocked the first commit attempt on a pre-existing fixture line in `hooks.selfcheck.js` that assigns a quoted literal to a key-named identifier. The gate was fulfilled, not bypassed: the identifier now comes from a `KEY_IDENT` constant, so no such shape appears in the committed source. This is the file's own existing convention, already applied to `FAKE_AWS_KEY` with a comment saying why, and these two fixture lines were an oversight of it. The text handed to the scanner under test is byte-identical, so the coverage is unchanged. No hook was disabled and `ECC_DISABLED_HOOKS` was not touched. ## Audit of the other hooks Every hook in this repo was checked against both payload shapes. Grepped for `data.edits`, `data.content`, `data.new_string` and `data.file_path`, then read each hit in context. `.claude/hooks/`: | Hook | Verdict | | --- | --- | | `secrets-scanner.js` | Was broken. Fixed here. | | `bash-safety.js` | Clean. Reads `(data.tool_input \|\| {}).command \|\| data.command`, covering both shapes. Bash payloads carry no edit array, so nothing else applies. Pinned in the self-check under both shapes. | | `commit-guard.js` | Clean. Identical resolution to `bash-safety.js`, same reasoning, same coverage. | | `hooks.selfcheck.js` | Not a hook; the harness itself. | `.wolf/hooks/` (out of this PR's scope, reported for completeness): | Hook | Verdict | | --- | --- | | `pre-write.js` | Blind to MultiEdit, but advisory only. Reads `tool_input.content`, `.old_string`, `.new_string` and never `.edits`, so its cerebrum Do-Not-Repeat check and bug-log lookup see nothing on a MultiEdit. It writes to stderr and always exits 0, so this loses a nudge, not a block. | | `post-write.js` | Same blindness, same advisory-only consequence, on the same fields. Anatomy and memory updates skip MultiEdit content. | | `pre-read.js`, `post-read.js` | Clean. Read `tool_input.file_path` / `.path` only; Read payloads have no edit array. | | `session-start.js`, `stop.js`, `shared.js`, `bugstore.js` | Clean. No tool-input payload parsing of this kind. | The `.wolf/` pair is worth a separate issue: it is a real gap in what the OpenWolf memory hooks observe, but it degrades to a missing warning rather than to a guard that silently passes, so it is a different severity from the one fixed here and does not belong in this diff. Also confirmed: no hook in this repo other than the secrets scanner ever exits 2 on the content of an edit, so the MultiEdit hole had exactly one security consequence and it is closed here. ## Related The same top-level-versus-nested assumption was found independently in `.claude/hooks/decision-citation-check.js` during the review of PR #1313, and is being fixed there. That file is untouched by this PR. ## Buglog entry To be appended to `.wolf/buglog.jsonl` on `main` in a separate buglog-only PR after this merges, per the OpenWolf protocol. ```json {"date":"2026-08-28","title":"Secrets scanner hook was blind to every MultiEdit and reported clean","error_message":"No output, exit 0, on a MultiEdit payload containing a hardcoded credential. Indistinguishable from a clean scan.","root_cause":"secrets-scanner.js resolved its edit array from the top-level data.edits field, which is Cursor's file-edit payload shape. Claude Code nests the array at tool_input.edits. On a MultiEdit all three content fallbacks (ti.content, ti.new_string, top-level data.edits) were empty, so content became the empty string and the early-return guard exited 0 with no output. The hooks.selfcheck.js harness fed both harness shapes but used Write for the Claude Code case, whose text lives in tool_input.content, so the nested-array path was never exercised; the harness was also not invoked by CI or any npm script, so even a covering case would not have run.","fix":"Consult both shapes, preferring tool_input.edits over data.edits. Add a claude-multiedit shape to hooks.selfcheck.js and run every secrets-scanner case under all three shapes. Wire hooks.selfcheck.js into the Repo policy lints required CI check so the harness actually runs.","tags":["hooks","secrets","claude-code","payload-shape","silent-failure","ci-coverage","issue-1333"]} ``` --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Review stream: Antigravity (
|
The memory-layers skill led with a six-layer model. No primary source states one. CoALA (Sumers et al., arXiv:2309.02427) states four: working, episodic, semantic, procedural. The skill now leads with those four and keeps entity/profile and reflection/consolidation as clearly marked Hive extensions, which is what they always were in the sourcing section of the vault document behind it. The larger problem the skill described but did not enforce was citation accuracy. Semantic memory in this repo is authoritative and agents cite it constantly, but nothing ever checked a citation. On 2026-07-30 an agent killed at its session limit left two retention migrations behind whose 14 day window was justified by citing a .wolf/decisions.md D-030 that did not exist then. A person going looking is the only reason it was caught. The rule saying "verify your citations" was already in force and did not stop it. So this adds the mechanical version. decision-citation-check.js runs as a PreToolUse hook on Write, Edit and MultiEdit. It blocks a write citing a D-NNN that is in no .wolf/decisions.md the checkout can see, and says how high it can see rather than claiming to know the highest id that exists. It warns without blocking when the cited entry opens REVOKED, SUPERSEDED, AMENDED, RETIRED or MOOT, and when an "owner ruling <date>" names a date absent from the ledger. It skips .wolf/decisions.md itself, where new ids are minted, and the hooks directory, whose fixtures carry deliberate fake ids. The same parser and skip list are runnable as a CLI for text a Write never passes through, such as a pull request body. Two things the gate deliberately does not do. It does not verify that the cited entry says what the citing text claims, which stays a reading job. And it cannot see a dispatch brief at all, since a brief is an Agent prompt rather than a file write, so the third fabrication this repo has seen is out of scope and the skill says so rather than letting a reader assume otherwise. Stale worktree ledgers are handled by merging every .wolf/decisions.md between the edited file and the filesystem root, so a worktree under .claude/worktrees sees the canonical checkout's fresher copy and a correct citation of a newly minted id is not read as a fabrication. A documented marker, citation-check: allow-unknown-ids, exempts content that is deliberately about an id which does not exist, since recording a fabricated id in the buglog is mandatory work here. Every use announces itself, so a bypass never looks like a guard that did not run. Payload handling is deliberately identical to secrets-scanner.js, which had the same MultiEdit defect first and fixed it in #1333: the edits array is read from either harness position, no field is trusted to be a string, and every source is audited rather than the first truthy one winning. Self-check: 61 of 61, over all three write payload shapes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1
f03ea4e to
981320f
Compare
Branch rebuilt on current main, and two corrections to the comment aboveThis branch was a parentless commit carrying a snapshot of an older main, which is why GitHub had no merge base for it. That held together until The rebuild is not just a rebase. Correction one: the "#1333 is still open" comment in this guard is gone, because it is no longer true. Leaving it would have been the exact defect this pull request exists to stop, a confident stale claim sitting in a comment about citation accuracy. Correction two: Within that shape, The self-check follows the same consolidation. The ad-hoc MultiEdit payload helper this pull request added is gone in favour of the The stale-worktree measurement was rerun on the rebuilt branch and still passes a correct Every one of the six review threads above remains fixed and resolved; none of this changes any of those answers. |
## Summary This is the batched buglog follow-up for the 21 pull requests merged during the 2026-08-28 session. Its diff is `.wolf/buglog.jsonl` and nothing else. Per `.claude/rules/openwolf.md`, every fixed bug, error, failed test or failed build must be logged, but the line may never be appended on a fix branch: `merge=union` in `.gitattributes` resolves concurrent appends locally and is ignored by GitHub's server side merge, so two branches that both appended land in hard conflict there. An unmergeable pull request gets no `refs/pull/N/merge`, no `pull_request` run and therefore zero checks, and the required status gate then blocks the merge for a reason the page never states (issue #873). Each fix accordingly carried its entry in its own pull request body, and this pull request copies them onto `main` in one batch, which the protocol explicitly prefers over one pull request per entry. ## Source pull requests 1240, 1251, 1253, 1257, 1268, 1276, 1277, 1278, 1281, 1287, 1292, 1293, 1294, 1296, 1300, 1301, 1303, 1305, 1313, 1335, 1337. All merged. Every entry came from a "Buglog entry" heading in one of those bodies. Nothing was invented for a pull request that carried none. ## What landed 36 entries appended, one JSON object per line, append only. The 196 pre-existing lines are byte identical to `origin/main`. | Source | Entries | |---|---| | #1240 | 1 | | #1251 | 1 | | #1253 | 1 | | #1257 | 4 | | #1268 | 5 | | #1276 | 2 | | #1277 | 1 (of 2 in the body) | | #1278 | 0 (merged into #1296) | | #1281 | 2 | | #1287 | 2 | | #1292 | 3 | | #1293 | 1 | | #1294 | 3 | | #1296 | 1 | | #1300 | 1 | | #1301 | 1 | | #1303 | 1 | | #1305 | 3 | | #1313 | 1 | | #1335 | 1 | | #1337 | 1 | Note on #1268: its first "Buglog entry" heading says "None yet" in prose and carries no JSON. Its two later headings, from the CI live lane and from the intermittent tool call failure, carry the five entries taken here. ## Deduplication - **#1278 dropped, folded into #1296.** Both describe the same defect: `omitempty` on `StreamContentBlock.Text` dropped the required `"text":""` from every text `content_block_start`, crashing the real Anthropic SDK's stream accumulator (issue #1274). #1278 is the conformance suite that found it and shipped it marked xfail; #1296 is the fix, and its entry carries the fuller root cause and the actual remedy. One bug, one entry. #1296's entry gains a `discovered_by` field naming #1278 so the discovery is not lost. - **#1277's first entry dropped.** The same body carries a later "Buglog entry (revised)" heading written after the review round found the page's claims did not match what the code enforces. The revised entry is the one taken. - Checked and kept as distinct: #1313 and #1337 are two different hooks (`decision-citation-check.js` and `secrets-scanner.js`) blind to the same MultiEdit payload shape, fixed in two different pull requests, so two entries. #1240 and #1335 are two different `account_not_provisioned` defects, one an observability gap at the edge boundary and one a console mint that should have refused, so two entries. #1305's three entries are three separate rounds of defects in the same money path change, each with its own root cause. ## Corrections against what actually merged Each entry was checked against the merged tree at `origin/main`, not against its own claim. - **#1240.** The entry said the log line went in at `AuthSnapshot.TenantUUID`. On `main` the check is the exported `authz.ParseTenantID(TenantLookup)` that `TenantUUID` delegates to, which the images and audio routing adapters (two further silent call sites found in the same review) also call, and `key_id` is deliberately not logged because CodeQL's clear text logging check flags any field named `*Key*` (alert #31). The `fix` field now says so. - **#1276, first entry.** The entry named `app/console/analytics/page.tsx` as the home of the five fetch helpers and their `Promise.all`. On `main` they live in `apps/web-console/lib/analytics/overview-fetch.ts`, extracted during review. Path corrected. - **#1277.** Its `error_message` was the placeholder `n/a`. Reconstructed from the pull request's own correction narrative: the page as first written published a blanket no content stored claim false for `/v1/batches`, `/v1/files` and `/v1/rag`, a product wide provider blindness claim disproved by catalogue summaries that name vendors (#1284), a metering claim anchored to the console side `UsageEventRow` projection rather than the `usage_events` table, and a 1:1 alias to route claim that is a property of seed data rather than of `SelectRoute`. **#1303 needed no correction.** Its original root cause asserted a live mid stream provider leak on the session chat relay that measurement disproved, and the author had already corrected the body before merge. The corrected version is what was taken, including the sentence recording that the session chat relay did not leak an error frame but silently truncated instead. Every other entry's central claim was verified present in the merged tree, among them `metering.SupportsIncludeUsage`, `sanitize.VariablePriceFrame` in the batch dispatcher, the revoke and regrant in `20260828_01_service_role_public_schema_grant.sql` with the `anon` assertions in `ci-throwaway-db.sh`, `normalizeReasoningUsage` now called from `normalizeChatCompletion`, `signup.SyncTenantMembershipRole`, `TestKeyViewHidesALimitThatIsNotEnforced`, `TestListEventsLatencyCrossesTheWire`, `mask-api-keys.mjs` and `md-table.mjs`, `StreamContentBlock.Text` as `*string`, `redactSnapshot`, `httpx.ReadBody`, `sanitize.ReplaceErrorFrame` with the default deny tail in `provider_blind.go`, `pinCompletionCeiling` and `captureInputTokens` with `applyReasoningHeadroom` gone, `requireBillingTenant`, and `hooks.selfcheck.js` wired into the Repo policy lints check. ## Verification - `node .wolf/hooks/bugstore.selfcheck.js` reports `bugstore selfcheck OK`. - All 232 lines parse as a single JSON object each. - Every appended entry carries `error_message`, `root_cause`, `fix` and `tags`. - Scanned for credentials: no API key, bearer token, JWT, password, AWS key or Postgres DSN with a password appears in any entry. The `hk_` occurrences are prefix descriptions in prose, not keys. - `git diff origin/main...HEAD --name-only` prints `.wolf/buglog.jsonl` and nothing else. No `.wolf/` telemetry was staged. ## Review No adversarial review streams were run, deliberately. This change is records only: it adds no code, no test, no configuration and no behavior, and `.wolf/buglog.jsonl` is on the inert path allowlist in `.github/workflows/ci.yml`, so the six required checks report green without running their heavy steps. If a check does fail here, that is a real signal about the file rather than about the pipeline. https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1 Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
## Summary This is the batched buglog follow-up for the pull requests merged to `main` on 2026-08-29. Its diff is `.wolf/buglog.jsonl` and nothing else. Per `.claude/rules/openwolf.md`, every fixed bug, error, failed test or failed build must be logged, but the line may never be appended on a fix branch. `merge=union` in `.gitattributes` resolves concurrent appends locally and is ignored by GitHub's server side merge, so two branches that both appended land in hard conflict there. An unmergeable pull request gets no `refs/pull/N/merge`, no `pull_request` run and therefore zero checks, and the required status gate then blocks the merge for a reason the page never states (issue #873). Each fix accordingly carried its entry in its own pull request body, and this pull request copies them onto `main` in one batch, which the protocol explicitly prefers over one pull request per entry. ## Scope examined Fifty nine pull requests merged to `main` on 2026-08-29. Forty eight of them carried at least one entry, for eighty two entries in total. Thirty two of those were already on `main` and are skipped, leaving fifty appended here from thirty four pull requests. The largest block of skips comes from #1342, the equivalent batch for the 2026-08-28 merges, which merged earlier the same day and already landed thirty six entries covering #1257, #1268, #1276, #1277, #1287, #1292, #1293, #1294, #1296, #1301, #1303, #1305, #1313, #1335 and #1337. ## What landed Fifty entries appended, one JSON object per line, append only. The 232 pre-existing lines are byte identical to `origin/main` (verified by hashing the first 232 lines of the result against the base file). Every line in the resulting file parses as JSON and carries `error_message`, `root_cause`, `fix` and `tags`. | Source | Entries | |---|---| | #1083 | 2 | | #1277 | 1 | | #1278 | 1 | | #1298 | 1 | | #1334 | 1 | | #1336 | 3 | | #1343 | 1 | | #1346 | 1 | | #1351 | 1 | | #1365 | 2 | | #1368 | 1 | | #1369 | 1 | | #1371 | 3 | | #1375 | 3 | | #1376 | 1 | | #1378 | 1 | | #1379 | 2 | | #1388 | 5 | | #1389 | 3 | | #1390 | 2 | | #1393 | 1 | | #1394 | 1 | | #1410 | 1 | | #1417 | 1 | | #1421 | 1 | | #1423 | 1 | | #1424 | 1 | | #1426 | 1 | | #1429 | 1 | | #1431 | 1 | | #1433 | 1 | | #1434 | 1 | | #1436 | 1 | | #1439 | 1 | Entries are copied verbatim from their source pull request bodies. Nothing was rewritten, no field was invented, and no field was added. No JSON needed repair: all eighty two extracted entries parsed on the first attempt and all four required fields were present on every one. ## Merged pull requests that carried no entry Eleven of the fifty nine. Recorded here because the gap is itself the useful signal. | Pull request | Title | Assessment | |---|---|---| | #1013 | chore(deps): bump the go-minor-patch group across 1 directory with 4 updates | Dependabot bump, no defect fixed, no entry expected | | #1015 | chore(deps): bump the go-minor-patch group across 1 directory with 6 updates | Dependabot bump, no entry expected | | #1016 | chore(deps): bump golang from 1.26-alpine to 1.27-alpine in /deploy/docker | Dependabot bump, no entry expected | | #1218 | chore(deps): bump postcss from 8.5.19 to 8.5.26 in /apps/desktop | Dependabot bump, no entry expected | | #1219 | chore(deps): bump golang.org/x/crypto from 0.41.0 to 0.52.0 in /apps/control-plane | Dependabot bump, no entry expected | | #1342 | chore: batch buglog entries for the 2026-08-28 merges | The previous batch pull request itself, correctly carries no entry of its own | | #1364 | chore: remove four dead skills and record the patterns that cost time | Protocol gap. The body records patterns that cost time, which is the shape of a buglog entry, but none was written as one | | #1383 | test: retire stale expected-failure markers, restore the ones that are true (#1381, #1382, #1324) | Protocol gap. Stale `it.fails` markers reading as red is a real defect that was fixed here and should have carried an entry | | #1384 | docs: correct D-047, hive-auto reverted to variable pricing (D-059) | Decision ledger correction, arguably a documentation defect, no entry written | | #1387 | chore(deps): bump next from 15.5.23 to 16.3.3 in /apps/agent-console | Dependabot bump, no entry expected | | #1398 | docs: rescue the 2026-08-25 parity captures and add the 2026-08-29 QA matrix evidence | Documentation and evidence rescue, no entry written | Six of the eleven are Dependabot bumps and one is the previous batch, so the genuine protocol gaps are #1364, #1383, #1384 and #1398. Of those, #1383 is the one worth a follow-up: it fixed a real defect class (a stale expected-failure marker reads as a red "Expect test to fail" and gets dismissed as pre-existing) and left no record. ## Entries skipped as already present Thirty two. Thirty of them matched an entry already on `main` on `error_message`, `id` or `fix`. Two more from #1278 are semantic duplicates that an exact match would have missed, and were skipped after reading the landed entries they duplicate: - #1278's `streaming content_block_start omits text field` entry is covered by the consolidated `bug-2026-08-28-anthropic-sdk-wire-conformance` entry landed from #1296, whose root cause names the same `omitempty` on `StreamContentBlock.Text`. - #1278's `GET /v1/models leaked an upstream provider name` entry is covered by `BUG-1284`, landed from #1300, which names the same `public.model_aliases.summary` publication path. #1278's third entry, on `top_k` forwarding producing a 400, is not covered anywhere on `main` and is appended here. #1342 recorded #1278 as fully "merged into #1296", which was accurate for two of its three entries. ## Note on entry quality One appended entry is thin: #1277's parity re-score record carries `error_message` of `n/a` and a root cause of "console had no privacy/data-policy surface at all". It is a parity gap record rather than a defect record. It is included exactly as written rather than embellished, per the protocol's preference for the author's own words. ## Test plan - [x] Branch cut fresh from `origin/main`, diff is `.wolf/buglog.jsonl` and nothing else - [x] First 232 lines byte identical to the base file (md5 match) - [x] All 282 resulting lines parse as JSON and carry `error_message`, `root_cause`, `fix` and `tags` - [x] No `.wolf/` telemetry (`anatomy.md`, `memory.md`, `token-ledger.json`, `hooks/_session.json`, `buglog.json`) in the commit - [ ] The six required checks report green via the inert path allowlist in `.github/workflows/ci.yml` --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Why
The owner asked that the four-layer memory system be used both in the Claude Code session and in the product we are building. Those are two genuinely different problems and this pull request answers them differently: the session half ships a mechanical gate, and the product half ships a written scope plus issues, with no product code changed.
Frame correction: four, not six
.claude/skills/memory-layers.mdled with a six-layer model. No primary source states one. CoALA (Sumers et al., arXiv:2309.02427) states four: working, episodic, semantic, procedural. The skill now leads with those four and keeps entity/profile and reflection/consolidation as clearly marked Hive extensions, which is what the sourcing section of the vault document behind it always said they were. The vault filename is unchanged so its inbound links keep working; a note at its top carries the correction.Session half: a gate, because the rule already existed and did not work
Semantic memory in this repo is authoritative and agents cite it constantly, and nothing ever checked a citation. On 2026-07-30 an agent killed at its session limit left two retention migrations behind whose 14 day window was justified by citing a
.wolf/decisions.mdD-030 that did not exist then; the resuming agent grepped for it, found nothing, and replaced the rationale. It was caught because a person went looking (native memory:feedback_killed_agent_work_verification). That id has since been minted for an unrelated decision, which is exactly why the fabrication is invisible to a grep today. "Verify your citations" was already written down and did not stop it, so writing it more emphatically was not the fix..claude/hooks/decision-citation-check.jsruns as a PreToolUse hook on Write, Edit and MultiEdit. It:D-NNNabsent from.wolf/decisions.md, and names the highest id it can see, so the fabrication is obvious in the refusal itself<date>" names a date that appears nowhere in the ledger. This one stays a warning deliberately: the ruling may be real and merely unrecorded, which is a reason to record it rather than to refuse a write.wolf/decisions.md, where new ids are minted, and.claude/hooks/, whose fixtures carry deliberate fake ids. Both the hook and the CLI apply the same skip list, from one helpercitation-check: allow-unknown-idsin the content. Recording a fabricated id is mandatory work here (the buglog entry after every fixed bug), and a post-mortem or review analysing one has the same need, so the guard leaves a greppable way to write the id down rather than pushing authors toward a Bash heredoc that skips every write-path guard. Every use announces itself on stdout, so a bypass never looks like a guard that did not runStale worktree ledgers are handled before any of that. Every builder works in its own worktree carrying whatever
.wolf/decisions.mdits branch point had, so the guard merges every ledger between the edited file and the filesystem root: a worktree under<repo>/.claude/worktrees/sees the canonical checkout's fresher copy, and a correct citation of a recently minted id is not read as a fabrication. When a block does land, the refusal says only how high it can see, tells the author to refresh from main, and says explicitly not to substitute a lower id that happens to exist, which is itself a fabrication.What the gate covers, stated plainly, because the three fabrications this repo has seen are not equally catchable:
D-NNNwith no ledger entry, in a Write, Edit or MultiEdit<date>" absent from the ledgerThe same parser is runnable as a CLI (
--check FILE..., or on stdin) for text a Write never passes through, such as a pull request body. One implementation and one skip list, so the gate and the manual check cannot disagree about what a valid citation is or which files are exempt.What it deliberately does not do: verify that the cited entry says what the citing text claims. An id existing is not agreement. That stays a reading job and the skill says so.
Why this will be used rather than ignored: it fires on the write, not on the reader's diligence, and it does not depend on the agent having read the skill first.
Product half: the honest map
Full document in the vault at
hive/architecture-2026-08-28-product-agent-memory.md, with every claim cited by file and line. Summary of what Hive's own shipped agents remember across sessions:The most surprising finding:
user_memoriesrecall runs on every chat request and reads a table that nothing in the product ever writes to. The internal write API has zero callers across Go, TypeScript, Python, Svelte and shell; there is no UI and no extractor. The feature is inert by construction, not broken, which is why nothing has failed loudly.The scaffolding is in better shape than the behavior. Tenant isolation is enforced at the row level on every relevant table, the episodic log is already append-only and retained, and the entity store already has its sanitization and caps right. What is missing is retrieval and production, not storage.
Filed as issues, ranked by value over cost:
user_memorieshas no writerReview round
Six findings came back on the first version of the guard and all six are fixed in d1130a3.
The two that mattered most were both about the change not being what it said it was. The hook read only Cursor's flat
data.edits, while Claude Code nests a MultiEdit's array attool_input.edits, so every Claude Code MultiEdit resolved to empty content and exited before auditing anything, even though the settings matcher included the tool and the skill said it was covered. And the guard's own justification cited a 2026-08-28 fabrication of D-030 that does not survive a check: D-030 has existed since the ledger became tracked on 2026-08-25. The real incident is the 2026-07-30 one now described above.The rest: the refusal asserted a ledger maximum it could not know and pushed toward substituting a lower id, fixed by merging up-tree ledgers and by wording;
DEAD_STATUSmissed RETIRED and MOOT; there was no way to write about a fabricated id at all, fixed by the bypass marker; and the CLI did not apply the skip list the header advertised, threw an ENOENT stack trace on a missing path, and read ids out of the middle of hyphenated identifiers.secrets-scanner.jshad the identical MultiEdit hole, reported from here as #1333. It was fixed on main by3b2bffd4awhile this pull request was in review, and its fix is better than the one here was, so this guard now uses that same payload handling rather than a parallel version of it.A second stream then found more holes of the same family. Fixed: the citation pattern was case sensitive, so a lowercase
d-030evaded it entirely, and merged ledgers were concatenated without a separator, so a file with no trailing newline fused two lines together. Declined: matching a bare plural such asD-030s, because capturing the trailing character would block a real id, while the possessive form already matches. Documented rather than coded: the bypass marker travels in a whole-file Write, so a file that keeps the marker stays exempt on later full-file writes. Full disposition in a comment on this pull request.The branch itself was rebuilt. It began as a parentless commit carrying a snapshot of an older main, which is why GitHub had no merge base for it; that held until
3b2bffd4atouched one of this pull request's own four files and the branch went to CONFLICTING. It is now an ordinary commit on top of current main with the same four files changed.Verification
Twenty-seven new cases. Across all three write payload shapes, using the shared
edit()helper rather than a private one: the clean path, a fabricated id, a lowercase fabricated id, an unrecorded owner ruling, a revoked id, a retired or moot id, the ledger-file skip, the bypass marker, and an id embedded in a hyphenated identifier. Plus three explicit payloads: one carrying both a content field and an edits array, one whosenew_stringis not a string, and one whose fabrication sits in an edit that is not the first in the array.The live, revoked and retired decision ids the cases use are read out of the ledger at runtime rather than hardcoded, so appending to
.wolf/decisions.mdcannot turn the self-check red on its own. The self-check also pinsCLAUDE_PROJECT_DIRand cwd to the repo root, which makes every guard's result independent of the directory it was invoked from.Each new behaviour was mutation tested. Reverting the edits-array read turns nine cases red; reverting the RETIRED and MOOT words, the bypass marker, the citation pattern's hyphen guard, or its case handling each turns the suite red on exactly the corresponding cases.
The stale worktree case was measured directly, with a worktree ledger stripped of D-057 and D-058 and the full ledger in the parent checkout:
The CLI path was exercised against live data: it skips the guard's own fixtures rather than blocking on them, reports an unreadable path as a usage error with a failing exit rather than a stack trace, blocks a fabricated id on stdin, and honours the bypass marker.
No UI surface is touched, so the visual-proof requirement does not apply.
Review streams this round: an independent adversarial pass (six findings, all fixed and resolved) and Antigravity (four findings, two fixed, one declined, one documented, written up in a comment on this pull request). CodeRabbit CLI is SKIPPED, not passed: it exits 1 with "Review limit reached, you have used all 3 included reviews currently available" on this account, and the PR's own CodeRabbit check reports "Review rate limited".
Buglog entry
To be appended to
.wolf/buglog.jsonlon main in a separate buglog-only pull request after this merges, per.claude/rules/openwolf.md:{"id":"bug-2026-08-28-citation-hook-multiedit-blind","date":"2026-08-28","title":"decision-citation-check hook exited 0 on every Claude Code MultiEdit","error_message":"No output, exit 0, on a MultiEdit payload citing a fabricated D-999; the settings matcher invoked the hook and it declined to look","root_cause":"The payload reader handled only Cursor's flat data.edits array. Claude Code nests a MultiEdit's edits under tool_input.edits, so ti.content, ti.new_string and the flat editsContent were all empty, content resolved to empty string, and the guard exited before auditing a single citation. The self-check could not see it because its claude payload shape was only ever tool_name Write with tool_input.content, so the MultiEdit shape was never exercised. Same defect as issue #1333 in secrets-scanner.js, inherited by copying the pattern.","fix":"Adopt the payload handling that fixed #1333 in secrets-scanner.js: read the edits array from either harness position, serialize any field that is not a string instead of coercing it, and audit every source rather than letting the first truthy one win. Run the decision-citation self-check cases over all three write shapes via the shared edit() helper, plus explicit payloads for content-and-edits together, a non-string new_string, and a fabrication in a later edit. Mutation tested: reverting the edits-array read turns nine cases red.","tags":["hooks","claude-code","payload-shape","fail-open","citation-check","pr-1313"],"related":["#1333"]}Scope
No product code changed. The product half is a scope, and building any of it is separate work on the normal pipeline.