Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 56 additions & 0 deletions FORK_CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -605,6 +605,62 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
### Fixed


- **Curated hits order above the transcripts that quote them; diary hits say they have no source** ([`HEAD`](https://github.com/techempower-org/mempalace/commit/HEAD))
A paraphrased question ranked session transcripts above the curated card
that answers it, so a reader at the default limit never reached the
correction. Measured on production (wing ``2g``, limit 20): the
``CLAUDE.md`` chunk carrying the REFUTED banner came back at rank 13 on
``bm25-fast`` and 14 on hybrid while a transcript quoting the same claim sat
at rank 1 — the class the reporting librarian named **FAIL-D**: right file,
right chunk, right markers, ranked below the cut. Its sibling: palace diary
summaries ranked first with no citable source at all.

New ``mempalace/result_ordering.py`` raises each hit to the earliest
position it has EARNED — the topmost hit it directly duplicates — searching
back no further than the nearest earlier hit it does not outrank by kind
(curated document, then transcript, then diary). Near-duplicate is the
overlap coefficient over word 3-grams at 0.35, chosen by measurement over
134 cross-kind pairs: bare token overlap is 0.63 even for merely same-topic
chunks, 3-gram Jaccard peaks at 0.16 for TRUE quoting pairs, and 3-gram
overlap gives 0.373/0.473/0.491 for the genuine buried-card pairs against
<= 0.323 for same-topic non-quoting pairs. A transcript quotes a card and
surrounds it with conversation, so the two differ wildly in length: overlap
asks "is most of the shorter inside the longer", which is the real question,
while Jaccard is punished by exactly the conversation that makes a quote a
quote.

Two bounds, both found by review rather than by design, and each needing a
case the original tests did not contain. Grouping by connected component
makes near-duplicate transitive, so a chain A~B~C let A clear C with no
shared 3-gram — that needs three hits, and every test used two. Testing rank
only against the hit earned FROM let a hit sail past one it duplicates and
does NOT outrank on the way, demoting a curated card below a transcript
quoting it — that needs one hit duplicating two others of different kinds,
which nothing covered. Both are now tests, alongside property tests over
4,000 randomised three-kind inputs each for idempotence and for the bound as
a pairwise invariant. Passes repeat until the order settles because a single
pass is not idempotent; exhausting the pass cap is logged rather than
swallowed, since an under-settled order is otherwise indistinguishable from
a settled one.

Reordering can only promote a hit the ranker actually returned, so when
nothing curated came back at all exactly one wider fetch is made (2x the
limit, capped at 40) before reordering and truncating to the requested
limit. The common case makes no extra call: this wave is cutting daemon
load, so the round trip is spent only where the search is broken.

Applied to the CLI's bm25-fast, hybrid and MCP-envelope routes and to
``searcher.search_memories``. Verified drift-free on production — one fetch,
ranks recorded, the reorder applied to a copy of that same response, because
a re-mine moved curated ranks repeatedly during the session and a
two-process before/after credited the corpus's work to the code. Diary hits
now render "palace diary summary — no source file" and sort below curated
hits.

*Tests:* 47 new (test_result_ordering x35 incl. the transitivity triple, the pass-what-you-outrank repro, the stabilise-cap pin and property tests for idempotence and the bound over 4,000 randomised three-kind inputs each; test_cli_daemon route-ordering + conditional widen + diary tag x10; test_provenance diary note x2)
*Files:* `mempalace/result_ordering.py`, `mempalace/cli.py`, `mempalace/searcher.py`, `mempalace/provenance.py`


- **Postgres backend embeds through get_embedding_function() instead of rebuilding an ONNX session per call** ([`4975b0a`](https://github.com/techempower-org/mempalace/commit/4975b0a))
``backends/postgres.py::_embed`` constructed its own
``DefaultEmbeddingFunction`` and never called
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -290,6 +290,7 @@ The full enumeration of fork-ahead changes. The canonical source is [`docs/fork-

| Description | Upstream PR | Fork commit |
|---|---|---|
| Curated hits order above the transcripts that quote them; diary hits say they have no source | — | [`HEAD`](https://github.com/techempower-org/mempalace/commit/HEAD) |
| Wake-up L1 stops leading with harness prompt-echo, diff fragments and tool-call receipts | — | [`HEAD`](https://github.com/techempower-org/mempalace/commit/HEAD) |
| mempalace tunnels --rebuild --wing W — refresh the derived graph on purpose | — | [`HEAD`](https://github.com/techempower-org/mempalace/commit/HEAD) |
| Batch tunnel persistence — a 2,000-tunnel rebuild goes from 37.9 s to 0.08 s | — | [`HEAD`](https://github.com/techempower-org/mempalace/commit/HEAD) |
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
seq: 145
id: search-curated-over-transcript-ordering
date: '2026-09-10'
bucket: Fixed
commit: HEAD
area: Search
summary: Curated hits order above the transcripts that quote them; diary hits say they have no source
body: |
A paraphrased question ranked session transcripts above the curated card
that answers it, so a reader at the default limit never reached the
correction. Measured on production (wing ``2g``, limit 20): the
``CLAUDE.md`` chunk carrying the REFUTED banner came back at rank 13 on
``bm25-fast`` and 14 on hybrid while a transcript quoting the same claim sat
at rank 1 — the class the reporting librarian named **FAIL-D**: right file,
right chunk, right markers, ranked below the cut. Its sibling: palace diary
summaries ranked first with no citable source at all.

New ``mempalace/result_ordering.py`` raises each hit to the earliest
position it has EARNED — the topmost hit it directly duplicates — searching
back no further than the nearest earlier hit it does not outrank by kind
(curated document, then transcript, then diary). Near-duplicate is the
overlap coefficient over word 3-grams at 0.35, chosen by measurement over
134 cross-kind pairs: bare token overlap is 0.63 even for merely same-topic
chunks, 3-gram Jaccard peaks at 0.16 for TRUE quoting pairs, and 3-gram
overlap gives 0.373/0.473/0.491 for the genuine buried-card pairs against
<= 0.323 for same-topic non-quoting pairs. A transcript quotes a card and
surrounds it with conversation, so the two differ wildly in length: overlap
asks "is most of the shorter inside the longer", which is the real question,
while Jaccard is punished by exactly the conversation that makes a quote a
quote.

Two bounds, both found by review rather than by design, and each needing a
case the original tests did not contain. Grouping by connected component
makes near-duplicate transitive, so a chain A~B~C let A clear C with no
shared 3-gram — that needs three hits, and every test used two. Testing rank
only against the hit earned FROM let a hit sail past one it duplicates and
does NOT outrank on the way, demoting a curated card below a transcript
quoting it — that needs one hit duplicating two others of different kinds,
which nothing covered. Both are now tests, alongside property tests over
4,000 randomised three-kind inputs each for idempotence and for the bound as
a pairwise invariant. Passes repeat until the order settles because a single
pass is not idempotent; exhausting the pass cap is logged rather than
swallowed, since an under-settled order is otherwise indistinguishable from
a settled one.

Reordering can only promote a hit the ranker actually returned, so when
nothing curated came back at all exactly one wider fetch is made (2x the
limit, capped at 40) before reordering and truncating to the requested
limit. The common case makes no extra call: this wave is cutting daemon
load, so the round trip is spent only where the search is broken.

Applied to the CLI's bm25-fast, hybrid and MCP-envelope routes and to
``searcher.search_memories``. Verified drift-free on production — one fetch,
ranks recorded, the reorder applied to a copy of that same response, because
a re-mine moved curated ranks repeatedly during the session and a
two-process before/after credited the corpus's work to the code. Diary hits
now render "palace diary summary — no source file" and sort below curated
hits.
tests: "47 new (test_result_ordering x35 incl. the transitivity triple, the pass-what-you-outrank repro, the stabilise-cap pin and property tests for idempotence and the bound over 4,000 randomised three-kind inputs each; test_cli_daemon route-ordering + conditional widen + diary tag x10; test_provenance diary note x2)"
fork_pr: 477
files: [mempalace/result_ordering.py, mempalace/cli.py, mempalace/searcher.py, mempalace/provenance.py]
75 changes: 69 additions & 6 deletions mempalace/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,7 @@
provenance_note,
source_kind,
)
from .result_ordering import prefer_curated
from .version import __version__


Expand Down Expand Up @@ -685,6 +686,7 @@ def _provenance_tag(hit: dict) -> str:
"""Short source-shape tag for ``--format compact`` (empty when unremarkable).

``⟨transcript⟩`` — a quoted copy from a session transcript.
``⟨diary⟩`` — a palace-written summary with no file behind it.
``⟨stale⟩`` — the file on disk has been modified since it was indexed.
Never both: a growing session transcript is expected, so ``⟨stale⟩`` is
suppressed for transcripts exactly as ``provenance_note`` suppresses the
Expand All @@ -700,6 +702,8 @@ def _provenance_tag(hit: dict) -> str:
flags = []
if kind == "transcript":
flags.append("transcript")
elif kind == "diary":
flags.append("diary")
elif hit.get("source_stale") is True:
flags.append("stale")
return f" ⟨{','.join(flags)}⟩" if flags else ""
Expand Down Expand Up @@ -2609,9 +2613,43 @@ def _resolve_search_limit(args) -> int:
return getattr(args, "results", 5)


def _daemon_search_fast(query: str, n_results: int, wing: str = None) -> dict | None:
"""BM25 fast path via GET /search/fast. Returns normalised data dict or None."""
rest_params = {"q": query, "limit": n_results}
# Reordering can only promote a hit the ranker actually returned. On the
# measured #451 case the curated card sat at rank 13 of 20 while the reader
# asked for 10, so it was never in the window to promote. When NOTHING
# curated came back, one wider fetch is made and the result is reordered and
# truncated; when a curated hit is already present — the common case — no
# extra call happens at all. Bounded deliberately: this wave is cutting
# daemon load, not adding to it.
_WIDEN_FACTOR = 2
_WIDEN_CAP = 40


def _widen_when_nothing_curated(hits: list, n_results: int, fetch) -> list:
"""One wider fetch when no curated document matched; else ``hits`` as-is.

``fetch(limit)`` returns a fresh, already-normalised hit list (or None).
The widened list is annotated before it is returned so the caller can
rank it. A failed or unhelpful second call degrades to the original
hits — the extra fetch is an optimisation, never a dependency.
"""
if not no_curated_source(hits):
return hits
widened_limit = min(n_results * _WIDEN_FACTOR, _WIDEN_CAP)
if widened_limit <= n_results:
return hits
try:
wider = fetch(widened_limit)
except DaemonError:
return hits
if not isinstance(wider, list) or len(wider) <= len(hits):
return hits
annotate(wider)
return wider


def _fast_hits(query: str, limit: int, wing: str = None) -> list | None:
"""GET /search/fast, normalised to the shared hit shape. None if unusable."""
rest_params = {"q": query, "limit": limit}
if wing:
rest_params["wing"] = wing
raw = _call_daemon_rest("/search/fast", rest_params)
Expand All @@ -2634,8 +2672,20 @@ def _daemon_search_fast(query: str, n_results: int, wing: str = None) -> dict |
hit["bm25_score"] = round(hit.pop("rank"), 3)
if hit.get("source_file"):
hit["source"] = hit["source_file"]
return hits


def _daemon_search_fast(query: str, n_results: int, wing: str = None) -> dict | None:
"""BM25 fast path via GET /search/fast. Returns normalised data dict or None."""
hits = _fast_hits(query, n_results, wing)
if hits is None:
return None
annotate(hits)
return {"results": hits, "query": query, "source": "bm25-fast"}
hits = _widen_when_nothing_curated(
hits, n_results, lambda limit: _fast_hits(query, limit, wing)
)
prefer_curated(hits)
return {"results": hits[:n_results], "query": query, "source": "bm25-fast"}


def _daemon_search_hybrid(
Expand All @@ -2651,7 +2701,18 @@ def _daemon_search_hybrid(
if data is None:
return None
data.setdefault("source", "hybrid")
annotate(data.get("results"))
hits = data.get("results")
annotate(hits)
if isinstance(hits, list):

def _refetch(limit):
wider_body = dict(body, limit=limit)
wider = _post_daemon_rest("/search/hybrid", wider_body)
return (wider or {}).get("results")

hits = _widen_when_nothing_curated(hits, n_results, _refetch)
prefer_curated(hits)
data["results"] = hits[:n_results]
return data


Expand Down Expand Up @@ -2883,7 +2944,9 @@ def cmd_search(args):

if data is None:
data = _call_daemon_tool("mempalace_search", arguments)
annotate(data.get("results") if isinstance(data, dict) else None)
results = data.get("results") if isinstance(data, dict) else None
annotate(results)
prefer_curated(results)
except DaemonError as e:
if want_json:
_emit_json({"error": str(e), "source": "daemon", "query": args.query})
Expand Down
6 changes: 6 additions & 0 deletions mempalace/provenance.py
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,7 @@
_STALE_GRACE_SECONDS = 60

_TRANSCRIPT_NOTE = "quoted copy from a session transcript — verify at the curated source"
_DIARY_NOTE = "palace diary summary — no source file"


def _paths(hit) -> list:
Expand Down Expand Up @@ -228,6 +229,11 @@ def provenance_note(hit) -> Optional[str]:
notes = []
if kind == "transcript":
notes.append(_TRANSCRIPT_NOTE)
elif kind == "diary":
# A palace-written summary: legitimate memory, but there is no file a
# reader can open to check it, and these rank first on literal-term
# queries (techempower-org/mempalace#451 item G).
notes.append(_DIARY_NOTE)
if stale is True and kind != "transcript":
indexed = _indexed_at(hit)
when = indexed.strftime("%Y-%m-%d") if indexed else "unknown date"
Expand Down
Loading
Loading