Skip to content

docs: four-layer memory in the session (enforced) and in the product (mapped) - #1313

Merged
sakibsadmanshajib merged 1 commit into
mainfrom
docs/four-layer-memory-session-and-product
Aug 29, 2026
Merged

sakibsadmanshajib merged 1 commit into
mainfrom
docs/four-layer-memory-session-and-product

Conversation

@sakibsadmanshajib

@sakibsadmanshajib sakibsadmanshajib commented Aug 28, 2026 •

Copy link
Copy Markdown
Owner

Why

The owner asked that the four-layer memory system be used both in the Claude Code session and in the product we are building. Those are two genuinely different problems and this pull request answers them differently: the session half ships a mechanical gate, and the product half ships a written scope plus issues, with no product code changed.

Frame correction: four, not six

.claude/skills/memory-layers.md led with a six-layer model. No primary source states one. CoALA (Sumers et al., arXiv:2309.02427) states four: working, episodic, semantic, procedural. The skill now leads with those four and keeps entity/profile and reflection/consolidation as clearly marked Hive extensions, which is what the sourcing section of the vault document behind it always said they were. The vault filename is unchanged so its inbound links keep working; a note at its top carries the correction.

Session half: a gate, because the rule already existed and did not work

Semantic memory in this repo is authoritative and agents cite it constantly, and nothing ever checked a citation. On 2026-07-30 an agent killed at its session limit left two retention migrations behind whose 14 day window was justified by citing a .wolf/decisions.md D-030 that did not exist then; the resuming agent grepped for it, found nothing, and replaced the rationale. It was caught because a person went looking (native memory: feedback_killed_agent_work_verification). That id has since been minted for an unrelated decision, which is exactly why the fabrication is invisible to a grep today. "Verify your citations" was already written down and did not stop it, so writing it more emphatically was not the fix.

.claude/hooks/decision-citation-check.js runs as a PreToolUse hook on Write, Edit and MultiEdit. It:

  • blocks a write citing a D-NNN absent from .wolf/decisions.md, and names the highest id it can see, so the fabrication is obvious in the refusal itself
  • warns without blocking when the cited entry opens REVOKED, SUPERSEDED, AMENDED, RETIRED or MOOT, since citing one as history is legitimate and citing one as a current rule is not. The last two are the words this ledger actually uses for D-028 and D-029
  • warns when an "owner ruling <date>" names a date that appears nowhere in the ledger. This one stays a warning deliberately: the ruling may be real and merely unrecorded, which is a reason to record it rather than to refuse a write
  • skips .wolf/decisions.md, where new ids are minted, and .claude/hooks/, whose fixtures carry deliberate fake ids. Both the hook and the CLI apply the same skip list, from one helper
  • can be bypassed deliberately with the literal citation-check: allow-unknown-ids in the content. Recording a fabricated id is mandatory work here (the buglog entry after every fixed bug), and a post-mortem or review analysing one has the same need, so the guard leaves a greppable way to write the id down rather than pushing authors toward a Bash heredoc that skips every write-path guard. Every use announces itself on stdout, so a bypass never looks like a guard that did not run

Stale worktree ledgers are handled before any of that. Every builder works in its own worktree carrying whatever .wolf/decisions.md its branch point had, so the guard merges every ledger between the edited file and the filesystem root: a worktree under <repo>/.claude/worktrees/ sees the canonical checkout's fresher copy, and a correct citation of a recently minted id is not read as a fabrication. When a block does land, the refusal says only how high it can see, tells the author to refresh from main, and says explicitly not to substitute a lower id that happens to exist, which is itself a fabrication.

What the gate covers, stated plainly, because the three fabrications this repo has seen are not equally catchable:

Class Treatment
A D-NNN with no ledger entry, in a Write, Edit or MultiEdit Blocked, against the ledger state of the day it happens
An "owner ruling <date>" absent from the ledger Warned, never blocked, by design
A dispatch brief misstating the ledger Out of scope: a brief is an Agent prompt, not a file write, so no PreToolUse hook on a write ever sees one

The same parser is runnable as a CLI (--check FILE..., or on stdin) for text a Write never passes through, such as a pull request body. One implementation and one skip list, so the gate and the manual check cannot disagree about what a valid citation is or which files are exempt.

What it deliberately does not do: verify that the cited entry says what the citing text claims. An id existing is not agreement. That stays a reading job and the skill says so.

Why this will be used rather than ignored: it fires on the write, not on the reader's diligence, and it does not depend on the agent having read the skill first.

Product half: the honest map

Full document in the vault at hive/architecture-2026-08-28-product-agent-memory.md, with every claim cited by file and line. Summary of what Hive's own shipped agents remember across sessions:

Layer State
Working Real, per run, destroyed with the sandbox
Episodic Recorded faithfully, never read back into a later run
Semantic Exists for chat, unreachable from an agent run
Procedural Does not exist
Entity/profile (extension) Recall wired, no writer anywhere in the product
Reflection/consolidation (extension) Does not exist

The most surprising finding: user_memories recall runs on every chat request and reads a table that nothing in the product ever writes to. The internal write API has zero callers across Go, TypeScript, Python, Svelte and shell; there is no UI and no extractor. The feature is inert by construction, not broken, which is why nothing has failed loudly.

The scaffolding is in better shape than the behavior. Tenant isolation is enforced at the row level on every relevant table, the episodic log is already append-only and retained, and the entity store already has its sanitization and caps right. What is missing is retrieval and production, not storage.

Filed as issues, ranked by value over cost:

Review round

Six findings came back on the first version of the guard and all six are fixed in d1130a3.

The two that mattered most were both about the change not being what it said it was. The hook read only Cursor's flat data.edits, while Claude Code nests a MultiEdit's array at tool_input.edits, so every Claude Code MultiEdit resolved to empty content and exited before auditing anything, even though the settings matcher included the tool and the skill said it was covered. And the guard's own justification cited a 2026-08-28 fabrication of D-030 that does not survive a check: D-030 has existed since the ledger became tracked on 2026-08-25. The real incident is the 2026-07-30 one now described above.

The rest: the refusal asserted a ledger maximum it could not know and pushed toward substituting a lower id, fixed by merging up-tree ledgers and by wording; DEAD_STATUS missed RETIRED and MOOT; there was no way to write about a fabricated id at all, fixed by the bypass marker; and the CLI did not apply the skip list the header advertised, threw an ENOENT stack trace on a missing path, and read ids out of the middle of hyphenated identifiers.

secrets-scanner.js had the identical MultiEdit hole, reported from here as #1333. It was fixed on main by 3b2bffd4a while this pull request was in review, and its fix is better than the one here was, so this guard now uses that same payload handling rather than a parallel version of it.

A second stream then found more holes of the same family. Fixed: the citation pattern was case sensitive, so a lowercase d-030 evaded it entirely, and merged ledgers were concatenated without a separator, so a file with no trailing newline fused two lines together. Declined: matching a bare plural such as D-030s, because capturing the trailing character would block a real id, while the possessive form already matches. Documented rather than coded: the bypass marker travels in a whole-file Write, so a file that keeps the marker stays exempt on later full-file writes. Full disposition in a comment on this pull request.

The branch itself was rebuilt. It began as a parentless commit carrying a snapshot of an older main, which is why GitHub had no merge base for it; that held until 3b2bffd4a touched one of this pull request's own four files and the branch went to CONFLICTING. It is now an ordinary commit on top of current main with the same four files changed.

Verification

$ node .claude/hooks/hooks.selfcheck.js
61/61 passed

Twenty-seven new cases. Across all three write payload shapes, using the shared edit() helper rather than a private one: the clean path, a fabricated id, a lowercase fabricated id, an unrecorded owner ruling, a revoked id, a retired or moot id, the ledger-file skip, the bypass marker, and an id embedded in a hyphenated identifier. Plus three explicit payloads: one carrying both a content field and an edits array, one whose new_string is not a string, and one whose fabrication sits in an edit that is not the first in the array.

The live, revoked and retired decision ids the cases use are read out of the ledger at runtime rather than hardcoded, so appending to .wolf/decisions.md cannot turn the self-check red on its own. The self-check also pins CLAUDE_PROJECT_DIR and cwd to the repo root, which makes every guard's result independent of the directory it was invoked from.

Each new behaviour was mutation tested. Reverting the edits-array read turns nine cases red; reverting the RETIRED and MOOT words, the bypass marker, the citation pattern's hyphen guard, or its case handling each turns the suite red on exactly the corresponding cases.

The stale worktree case was measured directly, with a worktree ledger stripped of D-057 and D-058 and the full ledger in the parent checkout:

Per D-058, this holds.   EXIT=0   (before: BLOCKED ... highest recorded: D-056)
Per D-999, this holds.   EXIT=2   BLOCKED ... highest id visible here: D-058

The CLI path was exercised against live data: it skips the guard's own fixtures rather than blocking on them, reports an unreadable path as a usage error with a failing exit rather than a stack trace, blocks a fabricated id on stdin, and honours the bypass marker.

No UI surface is touched, so the visual-proof requirement does not apply.

Review streams this round: an independent adversarial pass (six findings, all fixed and resolved) and Antigravity (four findings, two fixed, one declined, one documented, written up in a comment on this pull request). CodeRabbit CLI is SKIPPED, not passed: it exits 1 with "Review limit reached, you have used all 3 included reviews currently available" on this account, and the PR's own CodeRabbit check reports "Review rate limited".

Buglog entry

To be appended to .wolf/buglog.jsonl on main in a separate buglog-only pull request after this merges, per .claude/rules/openwolf.md:

{"id":"bug-2026-08-28-citation-hook-multiedit-blind","date":"2026-08-28","title":"decision-citation-check hook exited 0 on every Claude Code MultiEdit","error_message":"No output, exit 0, on a MultiEdit payload citing a fabricated D-999; the settings matcher invoked the hook and it declined to look","root_cause":"The payload reader handled only Cursor's flat data.edits array. Claude Code nests a MultiEdit's edits under tool_input.edits, so ti.content, ti.new_string and the flat editsContent were all empty, content resolved to empty string, and the guard exited before auditing a single citation. The self-check could not see it because its claude payload shape was only ever tool_name Write with tool_input.content, so the MultiEdit shape was never exercised. Same defect as issue #1333 in secrets-scanner.js, inherited by copying the pattern.","fix":"Adopt the payload handling that fixed #1333 in secrets-scanner.js: read the edits array from either harness position, serialize any field that is not a string instead of coercing it, and audit every source rather than letting the first truthy one win. Run the decision-citation self-check cases over all three write shapes via the shared edit() helper, plus explicit payloads for content-and-edits together, a non-string new_string, and a fabrication in a later edit. Mutation tested: reverting the edits-array read turns nine cases red.","tags":["hooks","claude-code","payload-shape","fail-open","citation-check","pr-1313"],"related":["#1333"]}

Scope

No product code changed. The product half is a scope, and building any of it is separate work on the normal pipeline.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 28, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 4 minutes.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 27160fce-81d5-4c3d-9464-2e245026138c

📥 Commits

Reviewing files that changed from the base of the PR and between 8d32fb1 and 981320f.

📒 Files selected for processing (4)
  • .claude/hooks/decision-citation-check.js
  • .claude/hooks/hooks.selfcheck.js
  • .claude/settings.json
  • .claude/skills/memory-layers.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sakibsadmanshajib sakibsadmanshajib left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review, pipeline stage 6

Read against the pushed diff at a591e9b, not the description, with every factual claim checked against the repository, the ledger, the vault and the live issue list.

Streams

  • Plain verification pass (this review): RAN.
  • Antigravity (agy, gemini-3.1-pro-high, high effort), framed as a pre-merge review: RAN. It independently found the MultiEdit fail-open and the CLI skip divergence, and added the hyphenated-identifier regex note. Its remaining conclusions matched my own measurements.
  • CodeRabbit: SKIPPED, not a pass. The bot review on this pull request stopped at "Review limit reached, next included review available in 28 minutes" and produced no findings, and coderabbit pullrequest 1313 --show-prompts returns "No CodeRabbit all-comments prompt was found", so there is nothing to read either way.
  • Codex connector: SKIPPED, usage limits reached, as expected until 2026-09-10.

Claims checked

CONFIRMED:

  • CoALA's four layers are working, episodic, semantic and procedural, and arXiv:2309.02427 is the right identifier. The skill leads with those four and marks entity/profile and reflection/consolidation as Hive extensions with no paper behind them, which is the accurate framing. No fifth or sixth layer is smuggled back in.
  • The hook is wired into the existing Write|Edit|MultiEdit PreToolUse matcher in .claude/settings.json, immediately after secrets-scanner.js.
  • node .claude/hooks/hooks.selfcheck.js reports 34 of 34 passed. I ran this branch's guards against the live ledger, not the author's transcript.
  • The CLI path behaves as described: it blocks a fabricated id, warns on a REVOKED entry, and reports D-058 as the highest recorded id. D-058 is in fact the highest, and the ledger has 58 contiguous entries.
  • Issues #1308, #1309, #1310, #1311 and #1312 all exist, are open, and their titles match the product table in the description.
  • Both vault documents exist, and the six-layer one carries a frame-correction callout at the top as the description claims.
  • Every D-NNN occurring anywhere in the repository outside the ledger resolves to a real entry, so the new gate has no blast radius on existing files.
  • The product half changes no product code and is scoped as a map plus issues rather than as shipped behaviour. Spot check of its headline finding holds: /internal/user-memories/ is registered in the control-plane router, the recall path exists in apps/edge-api/internal/chat/memory.go, and no caller of the write side appears anywhere in the tree.
  • No .wolf/ telemetry is committed and .wolf/buglog.jsonl is untouched, per the openwolf contract. No dash punctuation in the added prose.

UNSUPPORTED:

  • "runs as a PreToolUse hook on every Write, Edit and MultiEdit". MultiEdit is not covered in the Claude Code payload shape, measured, exit 0. See the thread on the content extraction.
  • "On 2026-08-28 ... a D-030 that did not exist at the time". D-030 has been in the tracked ledger since 2026-08-25. The incident being described is 2026-07-30. See the thread on the header comment.
  • "the CLI and the gate can never disagree about what counts as a valid citation". True for citation validity, false for the skip list, which only the hook path implements.

Minor, no thread opened:

  • The description says six new self-check cases. It is five per payload shape and ten in total, which is what the jump from 24 to 34 reflects.
  • The skill section does not mention the .claude/hooks/ skip that the code implements and the description documents.
  • Nothing outside Write, Edit and MultiEdit is covered, so a Bash heredoc bypasses the gate completely. Worth one honest sentence under a heading that promises enforcement rather than good intentions.

Verdict: NO MERGE

The direction is right and most of it is verified true. Two findings must land first, and both are small:

  1. MultiEdit fail-open, because an "enforced" claim that is false for one of three named tools is the specific defect that makes a document worse than no document.
  2. The mis-dated D-030 incident, because a guard against fabricated citations must not ship a fabricated-looking citation in its own justification, and because a reader greps D-030, finds it, and stops trusting the rest.

Recommended in the same pass, cheap and it removes a way this gate can cause the harm it prevents: the stale-worktree hard block, and the two missing dead-status words.

Everything else on this list is optional.

Comment thread .claude/hooks/decision-citation-check.js Outdated
Comment thread .claude/hooks/decision-citation-check.js Outdated
Comment thread .claude/hooks/decision-citation-check.js
Comment thread .claude/hooks/decision-citation-check.js Outdated
Comment thread .claude/hooks/decision-citation-check.js Outdated
Comment thread .claude/hooks/decision-citation-check.js
sakibsadmanshajib added a commit that referenced this pull request Aug 29, 2026
…1337)

Closes #1333.

## The hole

`.claude/hooks/secrets-scanner.js` is registered in
`.claude/settings.json` on the `Write|Edit|MultiEdit` matcher, but it
only ever saw the first two.

It resolved an edit array from the top-level `data.edits` field, which
is Cursor's file-edit payload shape. Claude Code nests that array at
`tool_input.edits`. On a MultiEdit, `ti.content`, `ti.new_string` and
the top-level `data.edits` are all absent, so `content` resolved to the
empty string, the `if (!content ...) process.exit(0)` guard fired, and
the hook exited 0 with no output.

Exit 0 and no output is exactly what a clean scan looks like. A
MultiEdit could therefore write an OpenAI key, a private key, an AWS
access key or a hardcoded password with the guard reporting nothing at
all. Write and Edit stayed covered throughout, since those populate
`tool_input.content` and `tool_input.new_string`.

## Fix

The scanner consults both shapes, preferring the Claude Code one:

```js
const editsSource = Array.isArray(ti.edits) ? ti.edits
  : Array.isArray(data.edits) ? data.edits
  : [];
```

## Pinning it

`.claude/hooks/hooks.selfcheck.js` already fed every guard both harness
payload shapes, but its Claude Code write case used `Write`, whose text
lives in `tool_input.content`. The nested-array path was never
exercised, which is why a harness that predates this bug did not catch
it.

A third shape, `claude-multiedit`, is added, and the secrets-scanner
cases now run under all three. The shell-command cases (bash-safety,
commit-guard) stay on the two shapes that apply to them rather than
being pointlessly duplicated.

Mutation check, run before opening this PR: reverting `editsSource` to
the top-level-only form turns the run red at 25/29, exit 1, with the
four failures all under `claude-multiedit`:

```
FAIL  secrets-scanner block aws key (claude-multiedit): exit 0, want 2; missing "BLOCKED"; missing "AWS access key"
FAIL  secrets-scanner warn api_key assignment (claude-multiedit): missing "SECRET WARNING"; missing "Possible API key assignment"
FAIL  secrets-scanner block go-style password assignment (claude-multiedit): exit 0, want 2; missing "BLOCKED"; missing "Hardcoded password"
FAIL  secrets-scanner warn go-style api_key assignment (claude-multiedit): missing "SECRET WARNING"; missing "Possible API key assignment"
```

With the fix restored: 29/29.

All fixture credentials are synthetic and assembled at runtime, so no
literal appears in the committed file: the AWS fixture is `'AKIA' +
'A'.repeat(16)` and the generic one is `'k'.repeat(20)`. No real
credential is committed here, as a fixture or otherwise.

## The self-check was not run by anything

Nothing in `.github/`, `Makefile` or `package.json` invoked
`hooks.selfcheck.js`. A harness that nothing runs is not a gate, and
that is the second half of how this survived. It is now a step in the
Repo policy lints required check, next to
`lint-workflow-check-names.mjs`, `lint-deploy-paths-filter.mjs`,
`newest-covered-commit.test.mjs` and `test-set-compose-project-name.sh`,
which live there for the same reason. The `changes` filter's
extension-deny arm runs first and matches `.js`, so a `.claude/hooks/`
change can never set `run=false` and skip this.

## Commit-time gate compliance

The ECC `pre-bash-commit-quality` hook blocked the first commit attempt
on a pre-existing fixture line in `hooks.selfcheck.js` that assigns a
quoted literal to a key-named identifier. The gate was fulfilled, not
bypassed: the identifier now comes from a `KEY_IDENT` constant, so no
such shape appears in the committed source. This is the file's own
existing convention, already applied to `FAKE_AWS_KEY` with a comment
saying why, and these two fixture lines were an oversight of it. The
text handed to the scanner under test is byte-identical, so the coverage
is unchanged. No hook was disabled and `ECC_DISABLED_HOOKS` was not
touched.

## Audit of the other hooks

Every hook in this repo was checked against both payload shapes. Grepped
for `data.edits`, `data.content`, `data.new_string` and
`data.file_path`, then read each hit in context.

`.claude/hooks/`:

| Hook | Verdict |
| --- | --- |
| `secrets-scanner.js` | Was broken. Fixed here. |
| `bash-safety.js` | Clean. Reads `(data.tool_input \|\| {}).command
\|\| data.command`, covering both shapes. Bash payloads carry no edit
array, so nothing else applies. Pinned in the self-check under both
shapes. |
| `commit-guard.js` | Clean. Identical resolution to `bash-safety.js`,
same reasoning, same coverage. |
| `hooks.selfcheck.js` | Not a hook; the harness itself. |

`.wolf/hooks/` (out of this PR's scope, reported for completeness):

| Hook | Verdict |
| --- | --- |
| `pre-write.js` | Blind to MultiEdit, but advisory only. Reads
`tool_input.content`, `.old_string`, `.new_string` and never `.edits`,
so its cerebrum Do-Not-Repeat check and bug-log lookup see nothing on a
MultiEdit. It writes to stderr and always exits 0, so this loses a
nudge, not a block. |
| `post-write.js` | Same blindness, same advisory-only consequence, on
the same fields. Anatomy and memory updates skip MultiEdit content. |
| `pre-read.js`, `post-read.js` | Clean. Read `tool_input.file_path` /
`.path` only; Read payloads have no edit array. |
| `session-start.js`, `stop.js`, `shared.js`, `bugstore.js` | Clean. No
tool-input payload parsing of this kind. |

The `.wolf/` pair is worth a separate issue: it is a real gap in what
the OpenWolf memory hooks observe, but it degrades to a missing warning
rather than to a guard that silently passes, so it is a different
severity from the one fixed here and does not belong in this diff.

Also confirmed: no hook in this repo other than the secrets scanner ever
exits 2 on the content of an edit, so the MultiEdit hole had exactly one
security consequence and it is closed here.

## Related

The same top-level-versus-nested assumption was found independently in
`.claude/hooks/decision-citation-check.js` during the review of PR
#1313, and is being fixed there. That file is untouched by this PR.

## Buglog entry

To be appended to `.wolf/buglog.jsonl` on `main` in a separate
buglog-only PR after this merges, per the OpenWolf protocol.

```json
{"date":"2026-08-28","title":"Secrets scanner hook was blind to every MultiEdit and reported clean","error_message":"No output, exit 0, on a MultiEdit payload containing a hardcoded credential. Indistinguishable from a clean scan.","root_cause":"secrets-scanner.js resolved its edit array from the top-level data.edits field, which is Cursor's file-edit payload shape. Claude Code nests the array at tool_input.edits. On a MultiEdit all three content fallbacks (ti.content, ti.new_string, top-level data.edits) were empty, so content became the empty string and the early-return guard exited 0 with no output. The hooks.selfcheck.js harness fed both harness shapes but used Write for the Claude Code case, whose text lives in tool_input.content, so the nested-array path was never exercised; the harness was also not invoked by CI or any npm script, so even a covering case would not have run.","fix":"Consult both shapes, preferring tool_input.edits over data.edits. Add a claude-multiedit shape to hooks.selfcheck.js and run every secrets-scanner case under all three shapes. Wire hooks.selfcheck.js into the Repo policy lints required CI check so the harness actually runs.","tags":["hooks","secrets","claude-code","payload-shape","silent-failure","ci-coverage","issue-1333"]}
```

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Review stream: Antigravity (agy, gemini-3.1-pro-high, effort high)

Four findings on the guard as it stood at d1130a3. Three are fixed in f03ea4e; the fourth is answered rather than coded. The first attempt at this stream timed out with no output after 20 minutes and was rerun with a narrower prompt, so the stream ran, and this is its result rather than an assumed pass.

Fixed: the citation pattern was case sensitive (Medium). A lowercase d-030 evaded the check entirely. CITATION now matches either case and the id is normalised to upper before the ledger lookup, so a lowercase citation of a real id resolves and a lowercase citation of a fabricated one blocks. Measured before allowing it, because a widened pattern is how a gate starts refusing correct writes: the repository contains no lowercase d-NNN token anywhere, so this adds no new class of false refusal. Self-check case added in both payload shapes, and reverting the pattern turns them red.

Fixed: merged ledgers concatenated without a separator (Low). loadLedger did text += raw, so a ledger with no trailing newline fused its last line onto the next ledger's first. That string is exactly what the owner-ruling date lookup searches. Now text += raw + '\n'.

Fixed: tool_input.new_str was not read (reported High). The severity is overstated, since nothing in the current Write|Edit|MultiEdit matcher sends that field; it is the Anthropic text-editor tool's name for it. Taken anyway, because the defect this pull request exists to fix was precisely an unrecognised field name resolving to empty content and exiting clean, and reading one more field can only add text to audit, never refuse something it would otherwise have passed. Self-check case added for that payload shape.

Not taken: match a bare plural such as D-030s. Capturing the trailing s would make the looked-up id D-030s, which is absent from the ledger, so a real id written that way would be blocked. The possessive form D-030's already matches, since an apostrophe is neither a word character nor a hyphen. The suggestion trades a typo case for a new class of false refusal, which is the failure this pull request just spent three findings removing.

Answered, not coded: the bypass marker is carried in a whole-file Write. The observation is right and it now has a paragraph in memory-layers.md instead of a code change. On an Edit or MultiEdit the audited content is only the new strings, so the marker scopes to that change. A Write carries the whole file by construction and there is no smaller unit to check, so a file that keeps the marker stays exempt on every later full-file write. The mitigation is visibility, not a narrower check: every exempt write announces itself on stdout, so such a file says so each time.

Self-check after these: 45 of 45, with both new behaviours mutation tested.

Stream status for the record: adversarial pass ran (six findings, all fixed and resolved above). Antigravity ran (this comment). CodeRabbit CLI is SKIPPED, not passed: coderabbit review --base-commit a591e9b9b exits 1 with "Review limit reached, you have used all 3 included reviews currently available", and the PR's own CodeRabbit check reports "Review rate limited".

The memory-layers skill led with a six-layer model. No primary source states
one. CoALA (Sumers et al., arXiv:2309.02427) states four: working, episodic,
semantic, procedural. The skill now leads with those four and keeps
entity/profile and reflection/consolidation as clearly marked Hive extensions,
which is what they always were in the sourcing section of the vault document
behind it.

The larger problem the skill described but did not enforce was citation
accuracy. Semantic memory in this repo is authoritative and agents cite it
constantly, but nothing ever checked a citation. On 2026-07-30 an agent killed
at its session limit left two retention migrations behind whose 14 day window
was justified by citing a .wolf/decisions.md D-030 that did not exist then. A
person going looking is the only reason it was caught. The rule saying "verify
your citations" was already in force and did not stop it.

So this adds the mechanical version. decision-citation-check.js runs as a
PreToolUse hook on Write, Edit and MultiEdit. It blocks a write citing a D-NNN
that is in no .wolf/decisions.md the checkout can see, and says how high it can
see rather than claiming to know the highest id that exists. It warns without
blocking when the cited entry opens REVOKED, SUPERSEDED, AMENDED, RETIRED or
MOOT, and when an "owner ruling <date>" names a date absent from the ledger.
It skips .wolf/decisions.md itself, where new ids are minted, and the hooks
directory, whose fixtures carry deliberate fake ids. The same parser and skip
list are runnable as a CLI for text a Write never passes through, such as a
pull request body.

Two things the gate deliberately does not do. It does not verify that the
cited entry says what the citing text claims, which stays a reading job. And
it cannot see a dispatch brief at all, since a brief is an Agent prompt rather
than a file write, so the third fabrication this repo has seen is out of scope
and the skill says so rather than letting a reader assume otherwise.

Stale worktree ledgers are handled by merging every .wolf/decisions.md between
the edited file and the filesystem root, so a worktree under .claude/worktrees
sees the canonical checkout's fresher copy and a correct citation of a newly
minted id is not read as a fabrication. A documented marker,
citation-check: allow-unknown-ids, exempts content that is deliberately about
an id which does not exist, since recording a fabricated id in the buglog is
mandatory work here. Every use announces itself, so a bypass never looks like
a guard that did not run.

Payload handling is deliberately identical to secrets-scanner.js, which had
the same MultiEdit defect first and fixed it in #1333: the edits array is read
from either harness position, no field is trusted to be a string, and every
source is audited rather than the first truthy one winning.

Self-check: 61 of 61, over all three write payload shapes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1
@sakibsadmanshajib
sakibsadmanshajib force-pushed the docs/four-layer-memory-session-and-product branch from f03ea4e to 981320f Compare August 29, 2026 03:00
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Branch rebuilt on current main, and two corrections to the comment above

This branch was a parentless commit carrying a snapshot of an older main, which is why GitHub had no merge base for it. That held together until 3b2bffd4a landed, because that commit touches .claude/hooks/hooks.selfcheck.js, one of this pull request's own four files, and the branch went to CONFLICTING. It is now rebuilt as an ordinary commit on top of current main, so it has a real merge base, and the diff is the same four files: 981320f9d.

The rebuild is not just a rebase. 3b2bffd4a is the fix for #1333, the identical MultiEdit hole in secrets-scanner.js that this pull request reported and deliberately left alone. It merged while this branch was in review, and it did the job better than this guard did, so two things change.

Correction one: the "#1333 is still open" comment in this guard is gone, because it is no longer true. Leaving it would have been the exact defect this pull request exists to stop, a confident stale claim sitting in a comment about citation accuracy.

Correction two: tool_input.new_str is reverted, and the comment above claiming it as fixed is wrong. secrets-scanner.js now resolves payload content in a way strictly better than the fallback chain that finding was aimed at: it reads the edits array from either harness position, serializes any field that is not a string rather than letting [object Object] swallow it, and audits every source instead of letting the first truthy one win. This guard now uses that same code, so a payload carrying both a content field and an edits array can no longer hide the array, and a non-string new_string can no longer hide a citation.

Within that shape, new_str is dead in both guards: no tool in the Write|Edit|MultiEdit matcher sends it. Adding it to one guard and not the other would reintroduce exactly the divergence between these two files that produced #1333 in the first place, so it is not carried here. If that field ever becomes real it belongs in both guards in one change.

The self-check follows the same consolidation. The ad-hoc MultiEdit payload helper this pull request added is gone in favour of the claude-multiedit shape 3b2bffd4a added to the shared edit() helper, so the decision-citation cases now run over all three write shapes rather than two plus a special case, and they pick up main's two edge shapes as well: a payload with both content and edits, and a non-string new_string. One case is kept that the shared helper cannot express, a fabrication sitting in an edit that is not the first in the array.

$ node .claude/hooks/hooks.selfcheck.js
61/61 passed

The stale-worktree measurement was rerun on the rebuilt branch and still passes a correct D-058 from a worktree ledger that lacks it. Mutation testing was rerun too: reverting the edits-array read now turns nine cases red rather than one, since the guard is exercised over three shapes instead of two.

Every one of the six review threads above remains fixed and resolved; none of this changes any of those answers.

@sakibsadmanshajib
sakibsadmanshajib merged commit 9bbbaf9 into main Aug 29, 2026
29 checks passed
@sakibsadmanshajib
sakibsadmanshajib deleted the docs/four-layer-memory-session-and-product branch August 29, 2026 03:05
sakibsadmanshajib added a commit that referenced this pull request Aug 29, 2026
## Summary

This is the batched buglog follow-up for the 21 pull requests merged
during the 2026-08-28 session. Its diff is `.wolf/buglog.jsonl` and
nothing else.

Per `.claude/rules/openwolf.md`, every fixed bug, error, failed test or
failed build must be logged, but the line may never be appended on a fix
branch: `merge=union` in `.gitattributes` resolves concurrent appends
locally and is ignored by GitHub's server side merge, so two branches
that both appended land in hard conflict there. An unmergeable pull
request gets no `refs/pull/N/merge`, no `pull_request` run and therefore
zero checks, and the required status gate then blocks the merge for a
reason the page never states (issue #873). Each fix accordingly carried
its entry in its own pull request body, and this pull request copies
them onto `main` in one batch, which the protocol explicitly prefers
over one pull request per entry.

## Source pull requests

1240, 1251, 1253, 1257, 1268, 1276, 1277, 1278, 1281, 1287, 1292, 1293,
1294, 1296, 1300, 1301, 1303, 1305, 1313, 1335, 1337. All merged.

Every entry came from a "Buglog entry" heading in one of those bodies.
Nothing was invented for a pull request that carried none.

## What landed

36 entries appended, one JSON object per line, append only. The 196
pre-existing lines are byte identical to `origin/main`.

| Source | Entries |
|---|---|
| #1240 | 1 |
| #1251 | 1 |
| #1253 | 1 |
| #1257 | 4 |
| #1268 | 5 |
| #1276 | 2 |
| #1277 | 1 (of 2 in the body) |
| #1278 | 0 (merged into #1296) |
| #1281 | 2 |
| #1287 | 2 |
| #1292 | 3 |
| #1293 | 1 |
| #1294 | 3 |
| #1296 | 1 |
| #1300 | 1 |
| #1301 | 1 |
| #1303 | 1 |
| #1305 | 3 |
| #1313 | 1 |
| #1335 | 1 |
| #1337 | 1 |

Note on #1268: its first "Buglog entry" heading says "None yet" in prose
and carries no JSON. Its two later headings, from the CI live lane and
from the intermittent tool call failure, carry the five entries taken
here.

## Deduplication

- **#1278 dropped, folded into #1296.** Both describe the same defect:
`omitempty` on `StreamContentBlock.Text` dropped the required
`"text":""` from every text `content_block_start`, crashing the real
Anthropic SDK's stream accumulator (issue #1274). #1278 is the
conformance suite that found it and shipped it marked xfail; #1296 is
the fix, and its entry carries the fuller root cause and the actual
remedy. One bug, one entry. #1296's entry gains a `discovered_by` field
naming #1278 so the discovery is not lost.
- **#1277's first entry dropped.** The same body carries a later "Buglog
entry (revised)" heading written after the review round found the page's
claims did not match what the code enforces. The revised entry is the
one taken.
- Checked and kept as distinct: #1313 and #1337 are two different hooks
(`decision-citation-check.js` and `secrets-scanner.js`) blind to the
same MultiEdit payload shape, fixed in two different pull requests, so
two entries. #1240 and #1335 are two different `account_not_provisioned`
defects, one an observability gap at the edge boundary and one a console
mint that should have refused, so two entries. #1305's three entries are
three separate rounds of defects in the same money path change, each
with its own root cause.

## Corrections against what actually merged

Each entry was checked against the merged tree at `origin/main`, not
against its own claim.

- **#1240.** The entry said the log line went in at
`AuthSnapshot.TenantUUID`. On `main` the check is the exported
`authz.ParseTenantID(TenantLookup)` that `TenantUUID` delegates to,
which the images and audio routing adapters (two further silent call
sites found in the same review) also call, and `key_id` is deliberately
not logged because CodeQL's clear text logging check flags any field
named `*Key*` (alert #31). The `fix` field now says so.
- **#1276, first entry.** The entry named
`app/console/analytics/page.tsx` as the home of the five fetch helpers
and their `Promise.all`. On `main` they live in
`apps/web-console/lib/analytics/overview-fetch.ts`, extracted during
review. Path corrected.
- **#1277.** Its `error_message` was the placeholder `n/a`.
Reconstructed from the pull request's own correction narrative: the page
as first written published a blanket no content stored claim false for
`/v1/batches`, `/v1/files` and `/v1/rag`, a product wide provider
blindness claim disproved by catalogue summaries that name vendors
(#1284), a metering claim anchored to the console side `UsageEventRow`
projection rather than the `usage_events` table, and a 1:1 alias to
route claim that is a property of seed data rather than of
`SelectRoute`.

**#1303 needed no correction.** Its original root cause asserted a live
mid stream provider leak on the session chat relay that measurement
disproved, and the author had already corrected the body before merge.
The corrected version is what was taken, including the sentence
recording that the session chat relay did not leak an error frame but
silently truncated instead.

Every other entry's central claim was verified present in the merged
tree, among them `metering.SupportsIncludeUsage`,
`sanitize.VariablePriceFrame` in the batch dispatcher, the revoke and
regrant in `20260828_01_service_role_public_schema_grant.sql` with the
`anon` assertions in `ci-throwaway-db.sh`, `normalizeReasoningUsage` now
called from `normalizeChatCompletion`,
`signup.SyncTenantMembershipRole`,
`TestKeyViewHidesALimitThatIsNotEnforced`,
`TestListEventsLatencyCrossesTheWire`, `mask-api-keys.mjs` and
`md-table.mjs`, `StreamContentBlock.Text` as `*string`,
`redactSnapshot`, `httpx.ReadBody`, `sanitize.ReplaceErrorFrame` with
the default deny tail in `provider_blind.go`, `pinCompletionCeiling` and
`captureInputTokens` with `applyReasoningHeadroom` gone,
`requireBillingTenant`, and `hooks.selfcheck.js` wired into the Repo
policy lints check.

## Verification

- `node .wolf/hooks/bugstore.selfcheck.js` reports `bugstore selfcheck
OK`.
- All 232 lines parse as a single JSON object each.
- Every appended entry carries `error_message`, `root_cause`, `fix` and
`tags`.
- Scanned for credentials: no API key, bearer token, JWT, password, AWS
key or Postgres DSN with a password appears in any entry. The `hk_`
occurrences are prefix descriptions in prose, not keys.
- `git diff origin/main...HEAD --name-only` prints `.wolf/buglog.jsonl`
and nothing else. No `.wolf/` telemetry was staged.

## Review

No adversarial review streams were run, deliberately. This change is
records only: it adds no code, no test, no configuration and no
behavior, and `.wolf/buglog.jsonl` is on the inert path allowlist in
`.github/workflows/ci.yml`, so the six required checks report green
without running their heavy steps. If a check does fail here, that is a
real signal about the file rather than about the pipeline.

https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
sakibsadmanshajib added a commit that referenced this pull request Aug 29, 2026
## Summary

This is the batched buglog follow-up for the pull requests merged to
`main` on 2026-08-29. Its diff is `.wolf/buglog.jsonl` and nothing else.

Per `.claude/rules/openwolf.md`, every fixed bug, error, failed test or
failed build must be logged, but the line may never be appended on a fix
branch. `merge=union` in `.gitattributes` resolves concurrent appends
locally and is ignored by GitHub's server side merge, so two branches
that both appended land in hard conflict there. An unmergeable pull
request gets no `refs/pull/N/merge`, no `pull_request` run and therefore
zero checks, and the required status gate then blocks the merge for a
reason the page never states (issue #873). Each fix accordingly carried
its entry in its own pull request body, and this pull request copies
them onto `main` in one batch, which the protocol explicitly prefers
over one pull request per entry.

## Scope examined

Fifty nine pull requests merged to `main` on 2026-08-29. Forty eight of
them carried at least one entry, for eighty two entries in total. Thirty
two of those were already on `main` and are skipped, leaving fifty
appended here from thirty four pull requests.

The largest block of skips comes from #1342, the equivalent batch for
the 2026-08-28 merges, which merged earlier the same day and already
landed thirty six entries covering #1257, #1268, #1276, #1277, #1287,
#1292, #1293, #1294, #1296, #1301, #1303, #1305, #1313, #1335 and #1337.

## What landed

Fifty entries appended, one JSON object per line, append only. The 232
pre-existing lines are byte identical to `origin/main` (verified by
hashing the first 232 lines of the result against the base file). Every
line in the resulting file parses as JSON and carries `error_message`,
`root_cause`, `fix` and `tags`.

| Source | Entries |
|---|---|
| #1083 | 2 |
| #1277 | 1 |
| #1278 | 1 |
| #1298 | 1 |
| #1334 | 1 |
| #1336 | 3 |
| #1343 | 1 |
| #1346 | 1 |
| #1351 | 1 |
| #1365 | 2 |
| #1368 | 1 |
| #1369 | 1 |
| #1371 | 3 |
| #1375 | 3 |
| #1376 | 1 |
| #1378 | 1 |
| #1379 | 2 |
| #1388 | 5 |
| #1389 | 3 |
| #1390 | 2 |
| #1393 | 1 |
| #1394 | 1 |
| #1410 | 1 |
| #1417 | 1 |
| #1421 | 1 |
| #1423 | 1 |
| #1424 | 1 |
| #1426 | 1 |
| #1429 | 1 |
| #1431 | 1 |
| #1433 | 1 |
| #1434 | 1 |
| #1436 | 1 |
| #1439 | 1 |

Entries are copied verbatim from their source pull request bodies.
Nothing was rewritten, no field was invented, and no field was added. No
JSON needed repair: all eighty two extracted entries parsed on the first
attempt and all four required fields were present on every one.

## Merged pull requests that carried no entry

Eleven of the fifty nine. Recorded here because the gap is itself the
useful signal.

| Pull request | Title | Assessment |
|---|---|---|
| #1013 | chore(deps): bump the go-minor-patch group across 1 directory
with 4 updates | Dependabot bump, no defect fixed, no entry expected |
| #1015 | chore(deps): bump the go-minor-patch group across 1 directory
with 6 updates | Dependabot bump, no entry expected |
| #1016 | chore(deps): bump golang from 1.26-alpine to 1.27-alpine in
/deploy/docker | Dependabot bump, no entry expected |
| #1218 | chore(deps): bump postcss from 8.5.19 to 8.5.26 in
/apps/desktop | Dependabot bump, no entry expected |
| #1219 | chore(deps): bump golang.org/x/crypto from 0.41.0 to 0.52.0 in
/apps/control-plane | Dependabot bump, no entry expected |
| #1342 | chore: batch buglog entries for the 2026-08-28 merges | The
previous batch pull request itself, correctly carries no entry of its
own |
| #1364 | chore: remove four dead skills and record the patterns that
cost time | Protocol gap. The body records patterns that cost time,
which is the shape of a buglog entry, but none was written as one |
| #1383 | test: retire stale expected-failure markers, restore the ones
that are true (#1381, #1382, #1324) | Protocol gap. Stale `it.fails`
markers reading as red is a real defect that was fixed here and should
have carried an entry |
| #1384 | docs: correct D-047, hive-auto reverted to variable pricing
(D-059) | Decision ledger correction, arguably a documentation defect,
no entry written |
| #1387 | chore(deps): bump next from 15.5.23 to 16.3.3 in
/apps/agent-console | Dependabot bump, no entry expected |
| #1398 | docs: rescue the 2026-08-25 parity captures and add the
2026-08-29 QA matrix evidence | Documentation and evidence rescue, no
entry written |

Six of the eleven are Dependabot bumps and one is the previous batch, so
the genuine protocol gaps are #1364, #1383, #1384 and #1398. Of those,
#1383 is the one worth a follow-up: it fixed a real defect class (a
stale expected-failure marker reads as a red "Expect test to fail" and
gets dismissed as pre-existing) and left no record.

## Entries skipped as already present

Thirty two. Thirty of them matched an entry already on `main` on
`error_message`, `id` or `fix`. Two more from #1278 are semantic
duplicates that an exact match would have missed, and were skipped after
reading the landed entries they duplicate:

- #1278's `streaming content_block_start omits text field` entry is
covered by the consolidated
`bug-2026-08-28-anthropic-sdk-wire-conformance` entry landed from #1296,
whose root cause names the same `omitempty` on
`StreamContentBlock.Text`.
- #1278's `GET /v1/models leaked an upstream provider name` entry is
covered by `BUG-1284`, landed from #1300, which names the same
`public.model_aliases.summary` publication path.

#1278's third entry, on `top_k` forwarding producing a 400, is not
covered anywhere on `main` and is appended here. #1342 recorded #1278 as
fully "merged into #1296", which was accurate for two of its three
entries.

## Note on entry quality

One appended entry is thin: #1277's parity re-score record carries
`error_message` of `n/a` and a root cause of "console had no
privacy/data-policy surface at all". It is a parity gap record rather
than a defect record. It is included exactly as written rather than
embellished, per the protocol's preference for the author's own words.

## Test plan

- [x] Branch cut fresh from `origin/main`, diff is `.wolf/buglog.jsonl`
and nothing else
- [x] First 232 lines byte identical to the base file (md5 match)
- [x] All 282 resulting lines parse as JSON and carry `error_message`,
`root_cause`, `fix` and `tags`
- [x] No `.wolf/` telemetry (`anatomy.md`, `memory.md`,
`token-ledger.json`, `hooks/_session.json`, `buglog.json`) in the commit
- [ ] The six required checks report green via the inert path allowlist
in `.github/workflows/ci.yml`

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant