From 981320f9d1caf740fea8441b8e175f2190de433e Mon Sep 17 00:00:00 2001 From: Sakib Sadman Shajib Date: Fri, 28 Aug 2026 23:00:23 -0400 Subject: [PATCH] docs: reframe memory layers on CoALA's four and gate decision citations The memory-layers skill led with a six-layer model. No primary source states one. CoALA (Sumers et al., arXiv:2309.02427) states four: working, episodic, semantic, procedural. The skill now leads with those four and keeps entity/profile and reflection/consolidation as clearly marked Hive extensions, which is what they always were in the sourcing section of the vault document behind it. The larger problem the skill described but did not enforce was citation accuracy. Semantic memory in this repo is authoritative and agents cite it constantly, but nothing ever checked a citation. On 2026-07-30 an agent killed at its session limit left two retention migrations behind whose 14 day window was justified by citing a .wolf/decisions.md D-030 that did not exist then. A person going looking is the only reason it was caught. The rule saying "verify your citations" was already in force and did not stop it. So this adds the mechanical version. decision-citation-check.js runs as a PreToolUse hook on Write, Edit and MultiEdit. It blocks a write citing a D-NNN that is in no .wolf/decisions.md the checkout can see, and says how high it can see rather than claiming to know the highest id that exists. It warns without blocking when the cited entry opens REVOKED, SUPERSEDED, AMENDED, RETIRED or MOOT, and when an "owner ruling " names a date absent from the ledger. It skips .wolf/decisions.md itself, where new ids are minted, and the hooks directory, whose fixtures carry deliberate fake ids. The same parser and skip list are runnable as a CLI for text a Write never passes through, such as a pull request body. Two things the gate deliberately does not do. It does not verify that the cited entry says what the citing text claims, which stays a reading job. And it cannot see a dispatch brief at all, since a brief is an Agent prompt rather than a file write, so the third fabrication this repo has seen is out of scope and the skill says so rather than letting a reader assume otherwise. Stale worktree ledgers are handled by merging every .wolf/decisions.md between the edited file and the filesystem root, so a worktree under .claude/worktrees sees the canonical checkout's fresher copy and a correct citation of a newly minted id is not read as a fabrication. A documented marker, citation-check: allow-unknown-ids, exempts content that is deliberately about an id which does not exist, since recording a fabricated id in the buglog is mandatory work here. Every use announces itself, so a bypass never looks like a guard that did not run. Payload handling is deliberately identical to secrets-scanner.js, which had the same MultiEdit defect first and fixed it in #1333: the edits array is read from either harness position, no field is trusted to be a string, and every source is audited rather than the first truthy one winning. Self-check: 61 of 61, over all three write payload shapes. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1 --- .claude/hooks/decision-citation-check.js | 329 +++++++++++++++++++++++ .claude/hooks/hooks.selfcheck.js | 128 +++++++++ .claude/settings.json | 5 + .claude/skills/memory-layers.md | 117 ++++++-- 4 files changed, 562 insertions(+), 17 deletions(-) create mode 100755 .claude/hooks/decision-citation-check.js diff --git a/.claude/hooks/decision-citation-check.js b/.claude/hooks/decision-citation-check.js new file mode 100755 index 000000000..1a8285809 --- /dev/null +++ b/.claude/hooks/decision-citation-check.js @@ -0,0 +1,329 @@ +#!/usr/bin/env node +// PreToolUse hook: refuses a write that cites a decision id which is not in +// `.wolf/decisions.md`. +// +// Why this exists. Semantic memory in this repo is authoritative +// (`.claude/skills/memory-layers.md`), and agents cite it constantly, but +// nothing ever checked a citation. On 2026-07-30 an agent killed at its +// session limit left behind two retention migrations whose 14 day window was +// justified by citing a `.wolf/decisions.md` D-030 that did not exist then. +// The resuming agent grepped, found nothing, and replaced the rationale; a +// person looking is the only reason it was caught (native memory: +// `feedback_killed_agent_work_verification`). D-030 has since been minted for +// an unrelated decision, which is precisely why a fabricated id is so hard to +// spot after the fact: grep the same string today and it resolves. A written +// rule saying "verify your citations" was already in force and did not stop +// it, so this is the mechanical version: the check runs on the write itself, +// whether or not the agent read the rule. +// +// What this actually stops, stated plainly, because the fabrications this repo +// has seen are not equally catchable and a reader will otherwise assume all of +// them are covered: +// +// - A `D-NNN` with no ledger entry, inside a Write, Edit or MultiEdit: +// BLOCKED. That is the 2026-07-30 case, judged against the ledger state of +// the day it happens. +// - An "owner ruling " naming a date that appears nowhere in the ledger: +// WARNS, never blocks, deliberately. The ruling may be real and merely +// unrecorded, which is a reason to record it, not a reason to refuse a write. +// - A dispatch brief that misstates the ledger (claiming it tops out ten +// entries below where it does, say): OUT OF SCOPE entirely. A brief is an +// Agent prompt, not a file write, so a PreToolUse hook on Write, Edit and +// MultiEdit never sees it. Nothing here fires on that class at all. +// +// SKIPS `.wolf/decisions.md` itself (where new ids are minted) and +// `.claude/hooks/` (this guard and its self-check carry deliberate fake ids as +// fixtures). Both entry points below apply the same skip list. +// +// BYPASS: content carrying the literal `citation-check: allow-unknown-ids` is +// not audited. Recording a fabricated id is mandatory work here (a buglog +// entry after every fixed bug, `.claude/rules/openwolf.md`), and a post-mortem +// or review analysing one has the same need, so the guard has to leave a +// deliberate, greppable way to write the id down. Without one the reachable +// workaround is a Bash heredoc, which skips every write-path guard and teaches +// the routing-around this gate exists to stop. The marker also covers a real +// entry that landed on main after this checkout was cut and cannot be +// refreshed in place. +// +// Staleness. Every builder here works in its own worktree, carrying whatever +// `.wolf/decisions.md` its branch point had, so a correct citation of an entry +// minted since can look fabricated. The ledger is therefore merged from every +// `.wolf/decisions.md` on the path from the edited file up to the filesystem +// root, which picks up the canonical checkout's fresher copy whenever the +// worktree sits under it (`/.claude/worktrees/`, the layout this +// repo uses). What that still cannot reach is handled by wording: a refusal +// never claims to know the highest id that exists, only the highest it can +// see, and it says to refresh from main rather than substitute a lower id that +// happens to exist. Substitution is what a confident wrong number invites, and +// it produces exactly the fabrication this guard exists to stop. +// +// Also runnable directly, for text a Write never passes through (a PR body, +// an existing file, a report): +// +// node .claude/hooks/decision-citation-check.js --check FILE... +// ... | node .claude/hooks/decision-citation-check.js --check +// +// Exit 2 means at least one citation is unverifiable. The parser and the skip +// list live here once and both entry points call them, so the CLI and the gate +// cannot disagree about what counts as a valid citation or an exempt file. + +const fs = require('fs'); +const path = require('path'); + +// The ledger's format is `- D-NNN | decision | source | date` (see the header +// of .wolf/decisions.md). Ids are always three digits there, so requiring +// three digits is what keeps stray matches like a "D-1" in prose out. +const LEDGER_ENTRY = /^-\s*(D-\d{3})\s*\|(.*)$/; +// A bare \b treats a hyphen as a word boundary, so an identifier like +// `part-D-123-abc` would read as a citation and be blocked. Excluding word +// characters and hyphens on both sides keeps the match to a standalone id. +// The `d` is accepted in either case and normalised to upper before the ledger +// lookup, so a lowercase `d-030` cannot slip a fabricated id past the check. +// Measured before allowing it: the repo contains no lowercase `d-NNN` token at +// all, so this adds no new class of false refusal. +const CITATION = /(?, text, sources }. +// The nearest ledger wins on a conflict, because an entry amended on this +// branch describes this branch's state better than an outer checkout's copy. +function loadLedger(startDir) { + const entries = new Map(); + const sources = []; + let text = ''; + for (const p of findLedgers(startDir)) { + let raw; + try { + raw = fs.readFileSync(p, 'utf8'); + } catch (e) { + continue; + } + sources.push(p); + // The newline matters: without it a ledger with no trailing newline fuses + // its last line onto the next ledger's first, and `text` is what the + // owner-ruling date lookup searches. + text += raw + '\n'; + for (const line of raw.split('\n')) { + const m = LEDGER_ENTRY.exec(line); + if (m && !entries.has(m[1])) entries.set(m[1], m[2]); + } + } + return { entries, text, sources }; +} + +function highestId(entries) { + let max = 0; + for (const id of entries.keys()) max = Math.max(max, Number(id.slice(2))); + return max ? `D-${String(max).padStart(3, '0')}` : '(none)'; +} + +// Returns { missing: [id], dead: [{id, status}], unrecordedRulings: [date] }. +function auditCitations(text, ledger) { + const missing = new Set(); + const dead = []; + const seenDead = new Set(); + + for (const raw of text.match(CITATION) || []) { + const cited = raw.toUpperCase(); + if (!ledger.entries.has(cited)) { + missing.add(cited); + continue; + } + if (seenDead.has(cited)) continue; + const status = DEAD_STATUS.exec(ledger.entries.get(cited).slice(0, STATUS_WINDOW)); + if (status) { + dead.push({ id: cited, status: status[1] }); + seenDead.add(cited); + } + } + + const unrecordedRulings = new Set(); + let m; + OWNER_RULING.lastIndex = 0; + while ((m = OWNER_RULING.exec(text)) !== null) { + if (!ledger.text.includes(m[1])) unrecordedRulings.add(m[1]); + } + + return { missing: [...missing], dead, unrecordedRulings: [...unrecordedRulings] }; +} + +// Emits the human-facing lines. Returns the exit code. +function report(label, audit, entries) { + for (const { id, status } of audit.dead) { + console.log( + `CITATION WARNING in ${label}: ${id} is marked ${status} in .wolf/decisions.md. ` + + `Read its entry before relying on it; a retired decision is history, not a current rule.` + ); + } + for (const date of audit.unrecordedRulings) { + console.log( + `CITATION WARNING in ${label}: cites an owner ruling dated ${date}, but that date ` + + `appears nowhere in .wolf/decisions.md. Record the ruling as a new decision first, ` + + `or cite what the owner actually said instead of a ledger entry that does not exist.` + ); + } + if (audit.missing.length === 0) return 0; + console.log( + `BLOCKED: ${label} cites ${audit.missing.join(', ')}, which ${audit.missing.length > 1 ? 'are' : 'is'} ` + + `not in any .wolf/decisions.md this checkout can see (highest id visible here: ${highestId(entries)}, ` + + `which is not necessarily the highest that exists). ` + + `If the entry is real and landed on main after this checkout was cut, refresh .wolf/decisions.md ` + + `from main and re-read the entry. Do NOT substitute a lower id that happens to exist: that is the ` + + `fabrication this check exists to stop. If you are deliberately writing about an id that does not ` + + `exist here, such as the buglog entry or post-mortem recording a fabricated one, put the literal ` + + `"${BYPASS_MARKER}" in the text. Fabricated decision citations have already shipped in this repo.` + ); + return 2; +} + +function runCli(argv) { + const files = argv.slice(argv.indexOf('--check') + 1); + const ledger = loadLedger(process.cwd()); + if (ledger.entries.size === 0) { + console.log('No .wolf/decisions.md found; nothing to check against.'); + process.exit(0); + } + + const readStdin = () => new Promise(resolve => { + let buf = ''; + process.stdin.setEncoding('utf8'); + process.stdin.on('data', c => buf += c); + process.stdin.on('end', () => resolve(buf)); + }); + + const check = (label, text) => { + if (text.includes(BYPASS_MARKER)) { + console.log(`Skipped ${label}: carries the ${BYPASS_MARKER} marker.`); + return 0; + } + return report(label, auditCitations(text, ledger), ledger.entries); + }; + + if (files.length === 0) { + readStdin().then(text => process.exit(check('stdin', text))); + return; + } + let worst = 0; + for (const f of files) { + if (isExemptPath(f)) { + console.log(`Skipped ${f}: exempt path (the ledger itself, or a hooks fixture).`); + continue; + } + let text; + try { + text = fs.readFileSync(f, 'utf8'); + } catch (e) { + // A path that cannot be read is a usage error, not a crash, and not a + // pass either: reporting it green would be the quiet absence this repo + // keeps getting burned by. + console.log(`Cannot read ${f}: ${e.message}`); + worst = 2; + continue; + } + worst = Math.max(worst, check(f, text)); + } + console.log( + `Checked ${files.length} file(s) against ${ledger.sources.length} ledger(s) ` + + `(highest id visible: ${highestId(ledger.entries)}).` + ); + process.exit(worst); +} + +function runHook() { + let input = ''; + const stdinTimeout = setTimeout(() => process.exit(0), 5000); + process.stdin.setEncoding('utf8'); + process.stdin.on('data', chunk => input += chunk); + process.stdin.on('end', () => { + clearTimeout(stdinTimeout); + try { + const data = JSON.parse(input); + const ti = data.tool_input || {}; + const filePath = ti.file_path || data.file_path || ''; + // Payload handling is deliberately identical to secrets-scanner.js, + // which had this defect first and fixed it in #1333. Claude Code nests a + // MultiEdit's array at `tool_input.edits`; Cursor's flat file-edit + // payload carries `edits` at the top level. Reading only the flat shape + // meant every Claude Code MultiEdit resolved to empty content and exited + // before auditing a single citation, while settings.json matched + // MultiEdit and the skill claimed it was covered. + const editsSource = Array.isArray(ti.edits) ? ti.edits + : Array.isArray(data.edits) ? data.edits + : []; + // Nothing here trusts a payload field to be a string, because the + // default coercion of an object is "[object Object]", which would hide + // any citation nested inside it. + const asText = v => (typeof v === 'string' ? v : v == null ? '' : JSON.stringify(v)); + const editsContent = editsSource.map(e => asText((e || {}).new_string)).join('\n'); + // Every source is audited rather than the first truthy one winning. A + // fallback chain lets a payload carrying both a content field and an + // edits array hide the array behind the field. + const content = [asText(ti.content), asText(ti.new_string), editsContent] + .filter(Boolean) + .join('\n'); + if (!content) process.exit(0); + if (isExemptPath(filePath)) process.exit(0); + if (content.includes(BYPASS_MARKER)) { + // Announced rather than silent. A bypass nobody can see in the + // transcript is indistinguishable from a guard that never ran, and + // any file documenting the literal (the memory-layers skill, say) + // trips it too, which a reader should be told rather than left to + // discover. + console.log( + `Citation audit skipped for ${filePath ? path.basename(filePath) : 'this write'}: ` + + `content carries the ${BYPASS_MARKER} marker.` + ); + process.exit(0); + } + + const ledger = loadLedger( + process.env.CLAUDE_PROJECT_DIR || (filePath ? path.dirname(filePath) : process.cwd()) + ); + if (ledger.entries.size === 0) process.exit(0); + + const label = filePath ? path.basename(filePath) : 'this write'; + process.exit(report(label, auditCitations(content, ledger), ledger.entries)); + } catch (e) { + // A guard that cannot parse its input must not block the write. + process.exit(0); + } + }); +} + +if (process.argv.includes('--check')) runCli(process.argv); +else runHook(); diff --git a/.claude/hooks/hooks.selfcheck.js b/.claude/hooks/hooks.selfcheck.js index 943cb8df3..b9485cbb9 100755 --- a/.claude/hooks/hooks.selfcheck.js +++ b/.claude/hooks/hooks.selfcheck.js @@ -19,6 +19,11 @@ const { spawnSync } = require('child_process'); const path = require('path'); const HOOKS_DIR = __dirname; +// decision-citation-check.js resolves .wolf/decisions.md relative to +// CLAUDE_PROJECT_DIR or the edited file's directory, so the guards are run +// from the repo root with that variable pinned. Without this the result +// depends on the directory the self-check happened to be invoked from. +const REPO_ROOT = path.resolve(HOOKS_DIR, '..', '..'); // Synthetic, never a real credential. Assembled at runtime so the literal // never appears in this file, which the secrets scanner reads on write. @@ -33,10 +38,35 @@ const FAKE_GENERIC_KEY = 'k'.repeat(20); // receives stays identical. const KEY_IDENT = 'apiKey'; +// Real ids pulled from the live ledger: one that exists, and ones whose entry +// opens with a dead-status marker. Reading them instead of hardcoding them +// keeps the decision-citation cases correct as the ledger grows. All are null +// on a checkout with no ledger, and the cases that need them are skipped +// rather than failing. +const LEDGER_LINES = (() => { + try { + return require('fs') + .readFileSync(path.join(REPO_ROOT, '.wolf', 'decisions.md'), 'utf8') + .split('\n') + .map(l => /^-\s*(D-\d{3})\s*\|(.*)$/.exec(l)) + .filter(Boolean); + } catch (e) { + return []; + } +})(); +const DEAD_WORDS = /\b(REVOKED|SUPERSEDED|AMENDED|RETIRED|MOOT)\b/; +const LIVE_DECISION_ID = (LEDGER_LINES.find(m => !DEAD_WORDS.test(m[2].slice(0, 80))) || [])[1] || null; +const REVOKED_DECISION_ID = (LEDGER_LINES.find(m => /\bREVOKED\b/.test(m[2].slice(0, 80))) || [])[1] || null; +// This ledger also retires entries with RETIRED (D-028) and MOOT (D-029). +// Both were invisible to the guard's warn arm until the vocabulary grew. +const OTHERWISE_DEAD_ID = (LEDGER_LINES.find(m => /\b(RETIRED|MOOT)\b/.test(m[2].slice(0, 80))) || [])[1] || null; + function run(hook, payload) { const res = spawnSync(process.execPath, [path.join(HOOKS_DIR, hook)], { input: JSON.stringify(payload), encoding: 'utf8', + cwd: REPO_ROOT, + env: { ...process.env, CLAUDE_PROJECT_DIR: REPO_ROOT }, }); return { code: res.status, out: (res.stdout || '').trim() }; } @@ -113,6 +143,61 @@ for (const shape of ['claude', 'claude-multiedit', 'cursor']) { ); } +// decision-citation-check: a real citation passes, a fabricated one is +// blocked. The ids the cases use are read out of the ledger at runtime rather +// than hardcoded, so appending to .wolf/decisions.md can never turn this +// self-check red on its own. This runs over all three write shapes, including +// claude-multiedit, because the guard read only the top-level `edits` array +// and so exited clean on every Claude Code MultiEdit, which is the same +// defect the secrets scanner had in #1333. +for (const shape of ['claude', 'claude-multiedit', 'cursor']) { + if (LIVE_DECISION_ID) { + // Clean first, so a block below can never be a pre-existing failure in + // disguise. + cases.push({ name: `decision-citation clean (${shape})`, hook: 'decision-citation-check.js', + payload: edit(shape, 'docs/note.md', `Per ${LIVE_DECISION_ID}, the ledger is authoritative.\n`), + code: 0, absent: ['BLOCKED', 'CITATION WARNING'] }); + } + cases.push( + { name: `decision-citation block fabricated id (${shape})`, hook: 'decision-citation-check.js', + payload: edit(shape, 'docs/note.md', 'Per D-999, we already ruled on this.\n'), + code: 2, present: ['BLOCKED', 'D-999', 'highest id visible here'] }, + // Case is normalised before the ledger lookup, so lowercasing an id is + // not a way around the block. + { name: `decision-citation blocks a lowercase fabricated id (${shape})`, hook: 'decision-citation-check.js', + payload: edit(shape, 'docs/note.md', 'Per d-999, we already ruled on this.\n'), + code: 2, present: ['BLOCKED', 'D-999'] }, + { name: `decision-citation warn unrecorded owner ruling (${shape})`, hook: 'decision-citation-check.js', + payload: edit(shape, 'docs/note.md', 'This follows an owner ruling on 1999-01-01.\n'), + code: 0, present: ['CITATION WARNING', '1999-01-01'], absent: ['BLOCKED'] }, + // The ledger is where new ids are minted, so a not-yet-known id in that + // file is the normal case, not a fabrication. + { name: `decision-citation skips the ledger itself (${shape})`, hook: 'decision-citation-check.js', + payload: edit(shape, '.wolf/decisions.md', '- D-999 | a brand new decision | owner | 2026-08-28\n'), + code: 0, absent: ['BLOCKED'] }, + // Recording a fabricated id is mandatory work here (the buglog entry after + // every fixed bug), so the documented marker has to let that text through. + { name: `decision-citation honours the bypass marker (${shape})`, hook: 'decision-citation-check.js', + payload: edit(shape, 'docs/note.md', 'citation-check: allow-unknown-ids\nThe agent cited D-999, which never existed.\n'), + code: 0, present: ['Citation audit skipped'], absent: ['BLOCKED'] }, + // A hyphen is a word boundary, so a bare \b read the middle of an + // identifier as a citation and blocked on it. + { name: `decision-citation ignores an id inside an identifier (${shape})`, hook: 'decision-citation-check.js', + payload: edit(shape, 'docs/note.md', 'The fixture is named part-D-999-abc in the harness.\n'), + code: 0, absent: ['BLOCKED'] }, + ); + if (REVOKED_DECISION_ID) { + cases.push({ name: `decision-citation warn revoked id (${shape})`, hook: 'decision-citation-check.js', + payload: edit(shape, 'docs/note.md', `Still bound by ${REVOKED_DECISION_ID}.\n`), + code: 0, present: ['CITATION WARNING', REVOKED_DECISION_ID], absent: ['BLOCKED'] }); + } + if (OTHERWISE_DEAD_ID) { + cases.push({ name: `decision-citation warn retired/moot id (${shape})`, hook: 'decision-citation-check.js', + payload: edit(shape, 'docs/note.md', `Still bound by ${OTHERWISE_DEAD_ID}.\n`), + code: 0, present: ['CITATION WARNING', OTHERWISE_DEAD_ID], absent: ['BLOCKED'] }); + } +} + // Two payload shapes that no harness is documented to send, pinned anyway // because this guard fails silently when it fails at all: it exits 0 with no // output, which is byte-identical to a clean scan. Both were raised in the @@ -147,6 +232,49 @@ cases.push( code: 2, present: ['BLOCKED', 'AWS access key'] }, ); +// The same three shapes against decision-citation-check, which resolves its +// content the same way and would fail the same way. The third pins the case +// the single-edit helper above cannot reach: a fabrication in an edit that is +// not the first one in the array. +cases.push( + { name: 'decision-citation scans edits even when content is also present', hook: 'decision-citation-check.js', + payload: { + hook_event_name: 'PreToolUse', + tool_name: 'MultiEdit', + tool_input: { + file_path: 'docs/note.md', + content: 'A line with no citation in it.\n', + edits: [{ old_string: 'placeholder', new_string: 'Per D-999, we already ruled on this.\n' }], + }, + }, + code: 2, present: ['BLOCKED', 'D-999'] }, + + { name: 'decision-citation scans a non-string new_string', hook: 'decision-citation-check.js', + payload: { + hook_event_name: 'PreToolUse', + tool_name: 'MultiEdit', + tool_input: { + file_path: 'docs/note.md', + edits: [{ old_string: 'placeholder', new_string: { note: 'Per D-999, we already ruled on this.' } }], + }, + }, + code: 2, present: ['BLOCKED', 'D-999'] }, + + { name: 'decision-citation scans every edit, not just the first', hook: 'decision-citation-check.js', + payload: { + hook_event_name: 'PreToolUse', + tool_name: 'MultiEdit', + tool_input: { + file_path: 'docs/note.md', + edits: [ + { old_string: 'a', new_string: 'An edit with no citation in it.\n' }, + { old_string: 'b', new_string: 'Per D-999, we already ruled on this.\n' }, + ], + }, + }, + code: 2, present: ['BLOCKED', 'D-999'] }, +); + let failed = 0; for (const c of cases) { const { code, out } = run(c.hook, c.payload); diff --git a/.claude/settings.json b/.claude/settings.json index ef9350610..1f4cb18b8 100644 --- a/.claude/settings.json +++ b/.claude/settings.json @@ -58,6 +58,11 @@ "type": "command", "command": "node \"$CLAUDE_PROJECT_DIR/.claude/hooks/secrets-scanner.js\"", "timeout": 5 + }, + { + "type": "command", + "command": "node \"$CLAUDE_PROJECT_DIR/.claude/hooks/decision-citation-check.js\"", + "timeout": 5 } ] }, diff --git a/.claude/skills/memory-layers.md b/.claude/skills/memory-layers.md index 4aa205839..02193163e 100644 --- a/.claude/skills/memory-layers.md +++ b/.claude/skills/memory-layers.md @@ -1,20 +1,25 @@ --- -name: Memory Layers (Six-Layer Model) -description: Map a memory question to the right layer/store, retrieve in the right order, verify staleness before trusting a hit, and consolidate (promote/retire) between layers. Companion to memory-tools.md, which only covers tool detection. +name: Memory Layers (Four-Layer Model) +description: Map a memory question to the right layer/store, retrieve in the right order, verify a citation and check staleness before trusting a hit, and consolidate (promote/retire) between layers. Companion to memory-tools.md, which only covers tool detection. --- ## Memory Layers -Six layers, built on CoALA's four layers (Sumers et al. 2024, arXiv:2309.02427): -working, episodic, semantic, procedural, plus two practitioner extensions -this project specifically needs: entity/profile memory and -reflection/consolidation (grounded in Park et al. 2023, arXiv:2304.03442). -"Six" is not an established standard; no primary source states it. Full -sourcing, the one source that does use the word "six" and why it is not -treated as authoritative, and the gap analysis: vault -`hive/architecture-2026-08-28-six-layer-agent-memory.md`. +**Four layers, from CoALA** (Sumers et al. 2024, [arXiv:2309.02427](https://arxiv.org/abs/2309.02427)): +working, episodic, semantic, procedural. That is the frame, and it is the only +one with a primary source behind it. -### The six layers +Two further rows below are **Hive extensions, not CoALA and not a standard**: +entity/profile memory and reflection/consolidation. They are kept because this +project needs both jobs done and nothing else does them, not because any paper +says a memory model has six layers. No primary source states a six-layer model; +a doc or brief that presents one as established is wrong. Full sourcing, the +secondary sources that use the words five and six, why they are not treated as +authoritative, and the gap analysis: vault +`hive/architecture-2026-08-28-six-layer-agent-memory.md` (filename kept for its +inbound links; read its research section, not its title, for the frame). + +### The four layers | # | Layer | Store/source | Location | Lifespan | |---|-------|---------------|----------|----------| @@ -22,8 +27,13 @@ treated as authoritative, and the gap analysis: vault | 2 | Episodic | `.wolf/buglog.jsonl`; claude-mem sqlite | repo (tracked, append-only); external (`~/.claude-mem/`, read-only) | permanent | | 3 | Semantic | `.wolf/decisions.md` (authoritative), MEMORY.md `project_*`/`reference_*`, vault specs | repo; native (`~/.claude/projects/.../memory/`); external (Obsidian vault) | permanent, edit in place | | 4 | Procedural | `.wolf/cerebrum.md` Do-Not-Repeat, MEMORY.md `feedback_*`, `.claude/skills/*` | repo; native; repo | permanent, edit in place | -| 5 | Entity/Profile | MEMORY.md `user_*.md` | native | permanent, edit in place | -| 6 | Reflection/Consolidation | no store (this skill) | n/a | promotion + retirement process | + +### Two Hive extensions (not CoALA, not a standard) + +| # | Layer | Store/source | Location | Lifespan | +|---|-------|---------------|----------|----------| +| E1 | Entity/Profile | MEMORY.md `user_*.md` | native | permanent, edit in place | +| E2 | Reflection/Consolidation | no store (this skill) | n/a | promotion + retirement process | `project_*.md` lives under Semantic (row 3), not Entity/Profile: in this repo's actual MEMORY.md convention every `project_*.md` file is a durable @@ -34,7 +44,7 @@ instance, `user_sakib.md`; the layer stays because the category is real (a future second tenant/customer/named sub-project would land here), not because this repo currently has many of them. -Why layers 5 and 6 exist beyond CoALA's four: **entity/profile** is needed +Why the two extensions exist beyond CoALA's four: **entity/profile** is needed because a fact about one specific named person (preferences, role) is a different shape than a general project fact, and folding it into semantic memory loses the "whose fact is this" anchor the moment there is more than @@ -64,6 +74,79 @@ that check runs. 5. Entity: MEMORY.md `user_*.md`. Anything specific to this person? 6. Vault: only when a terse layer points at a doc and full detail is needed. +### Citation check (enforced by a hook, not by good intentions) + +Semantic memory is only worth having if a citation to it is true. On +2026-07-30 an agent killed at its session limit left two retention migrations +behind whose 14 day window was justified by citing a `.wolf/decisions.md` +D-030 that did not exist then; a person grepping for it is what caught it +(native memory: `feedback_killed_agent_work_verification`). That id has since +been minted for an unrelated decision, which is exactly why a fabricated +citation is so hard to catch after the fact: grep the same string today and it +resolves. Writing "verify your citations" here again would have changed +nothing, so the rule is mechanical instead: + +`.claude/hooks/decision-citation-check.js` runs as a PreToolUse hook on every +Write, Edit and MultiEdit. It **blocks** the write when the content cites a +`D-NNN` with no entry in `.wolf/decisions.md`, and names the highest id it can +see so the fabrication is obvious. It **warns**, without blocking, when the +cited entry opens REVOKED / SUPERSEDED / AMENDED / RETIRED / MOOT (citing one +as history is legitimate; citing one as a current rule is not) and when an +"owner ruling ``" names a date that appears nowhere in the ledger. That +second one stays a warning deliberately: the ruling may be real and merely +unrecorded, which is a reason to record it, not to refuse a write. + +What it does **not** cover, so nobody reads more into it than is there: a +dispatch brief that misstates the ledger (claiming it tops out ten entries +below where it does, say). A brief is an Agent prompt, not a file write, so no +PreToolUse hook on Write, Edit or MultiEdit ever sees one. That class is still +entirely on you. + +It skips `.wolf/decisions.md` itself, where new ids are minted, and +`.claude/hooks/`, whose fixtures carry deliberate fake ids. Both the hook and +the CLI below apply the same skip list. + +Writing **about** an id that does not exist is mandatory work here (the buglog +entry after every fixed bug, and the post-mortem or review analysing one), so +there is a deliberate, greppable bypass: put the literal +`citation-check: allow-unknown-ids` in the content and the audit is skipped. +It also covers a real entry that landed on main after this worktree was cut +and cannot be refreshed in place. It is not for getting past a block you have +not read. Every use is announced on stdout, so a bypass is visible in the +transcript rather than looking like a guard that never ran. + +Know the reach of the marker before using it. On an Edit or MultiEdit only the +new strings are audited, so the marker exempts that change alone. A Write +carries the whole file, so a marker anywhere in it exempts the entire write, +and a file that keeps the marker stays exempt on every later full-file write. +This section is itself an example: it contains the literal, so writes to this +file are exempt. That is the price of a greppable marker rather than a bug, +and the per-write announcement is what keeps it visible. + +Stale worktree ledgers are handled before it comes to that. Every builder +works in its own worktree carrying whatever `.wolf/decisions.md` its branch +point had, so the guard merges every ledger from the edited file up to the +filesystem root: a worktree under `/.claude/worktrees/` sees the +canonical checkout's fresher copy, and a correct citation of a recently minted +id is not read as a fabrication. When a block does land, refresh +`.wolf/decisions.md` from main and re-read the entry. Never substitute a lower +id that happens to exist: that turns a stale ledger into a real fabrication, +which is the failure this whole check exists to stop. + +For text a Write never passes through, a PR body or an existing file, run the +same parser directly: + +``` +node .claude/hooks/decision-citation-check.js --check FILE... +gh pr view --json body -q .body | node .claude/hooks/decision-citation-check.js --check +``` + +Self-check for the guard itself: `node .claude/hooks/hooks.selfcheck.js`. + +What the hook cannot check, and stays your job: whether the entry you cited +says what you claimed it says. An id existing is not agreement. Read the +line before you lean on it. + ### Staleness check (mandatory before trusting any hit above working) - Grep the hit's own text for `REVOKED`, `SUPERSEDED`, or `STALE` first. @@ -90,13 +173,13 @@ that type. ### Retirement: a memory turns out wrong -Applies to the mutable layers (semantic, procedural, entity/profile: layers -3, 4, 5). Never delete. Edit in place: +Applies to the mutable layers: semantic and procedural (layers 3 and 4) plus +the entity/profile extension (E1). Never delete. Edit in place: 1. Prefix `REVOKED` or `SUPERSEDED ` on both the frontmatter/index line and the opening sentence of the body. 2. Point at what supersedes it (decision id, PR, or new fact). 3. Leave the original wrong text below so a future grep for the old term finds the correction, not silence. -4. `.wolf/decisions.md` specifically is append-only within this pattern: add a new `D-0xx` row rather than editing the old one; the old row gets the `REVOKED` prefix pointing at the new row's id (pattern: D-013, D-029, D-035, D-036). +4. `.wolf/decisions.md` specifically is append-only within this pattern: add a new `D-0xx` row rather than editing the old one; the old row gets a dead-status prefix pointing at the new row's id. `REVOKED` is the usual word (pattern: D-013, D-035, D-036), and the ledger also uses `RETIRED` (D-028) and `MOOT` (D-029) where those read truer. The citation hook treats all five words as dead, so any of them earns the warning. Episodic (layer 2) does not get this treatment at all: `.wolf/buglog.jsonl` is append-only history by contract (`.claude/rules/openwolf.md`: "never