Skip to content

feat: preserve expand-query evidence provenance - #484

Open
100yenadmin wants to merge 2 commits into
stephenschoettler:mainfrom
100yenadmin:feat/expand-query-evidence-provenance
Open

100yenadmin wants to merge 2 commits into
stephenschoettler:mainfrom
100yenadmin:feat/expand-query-evidence-provenance

Conversation

@100yenadmin

@100yenadmin 100yenadmin commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

Preserve bounded evidence provenance in lcm_expand_query answers

Summary

lcm_expand_query already builds rich recursive context before asking the auxiliary model to synthesize an answer. That context can contain:

  • matched root summaries;
  • recursively reached child summaries;
  • leaf messages;
  • raw-query windows;
  • hydrated externalized payload content;
  • durable transcript placeholders/content;
  • node/store/session identity and recursive source paths.

Today, most of that trace disappears after synthesis. The successful response returns the answer plus top-level matches, but not necessarily the exact descendants or excerpts the model saw. A fluent answer therefore cannot be traced back to the complete bounded synthesis input from the response alone.

This PR adds an additive, bounded evidence_provenance object to successful, no-match, and structured degraded lcm_expand_query responses. The object is extracted by the tool from the exact context_blocks supplied to synthesis; it is not authored by the auxiliary model.

No ranking, retrieval selection, recursive expansion, prompt construction, model routing, storage schema, durable state, or existing top-level response field changes.

Why

PR #266 correctly made lcm_expand_query recursively descend summary DAGs and include leaf evidence before synthesis. That improved answer quality, but it also created a response-level observability gap:

  1. the model can see a recursively reached child or leaf;
  2. that source can materially influence the answer;
  3. the final response may expose only the selected root summary;
  4. an agent/operator cannot tell which exact bounded excerpts were available to synthesis.

The same gap exists for raw FTS windows and hydrated externalized payloads: the model sees a bounded context representation, while the caller receives only search metadata or a top-level match.

This PR preserves that boundary without pretending to solve replay verification, answer verification, or authorization. It explicitly separates:

  • locator coverage — are bounded locator arguments present for each unique synthesis-context identity?;
  • locator replay — do those arguments still resolve to the quoted representation?;
  • semantic entailment — does every answer claim follow from those inputs?;
  • authority — does possession of an identifier authenticate or authorize the caller?

Only bounded locator coverage and context representation are implemented here. Locator replay remains explicitly unverified.

Response contract

Every successful answer, no-match result, and currently structured degraded lcm_expand_query response receives:

{
  "evidence_provenance": {
    "retrieval_scope": {
      "kind": "current_session",
      "session_ids": ["..."]
    },
    "identifiers_are_authority": false,
    "locator_replay_safety": "not_guaranteed",
    "synthesis_status": "completed",
    "locator_coverage": "complete",
    "semantic_entailment": "not_verified",
    "quote_origin": "tool_extracted_from_synthesis_context",
    "context_truncated": false,
    "context_unique_item_count": 3,
    "context_occurrence_count": 4,
    "items_truncated": false,
    "serialized_char_limit": 10000,
    "serialization": {
      "scope": "evidence_provenance_only",
      "json_ensure_ascii": true
    },
    "quote_max_chars": 500,
    "quotes_truncated_by_provenance_cap": 0,
    "metadata_truncated": false,
    "items": []
  }
}

Evidence representations

Items use three source types:

  • summary — root or recursively reached summary text;
  • raw_message — stored, recursive, or raw-search message content;
  • externalized_payload — hydrated externalized content with direct externalized_ref expansion.

Raw-message and externalized items additionally use content_source to distinguish the exact representation supplied to synthesis, including raw_search_hit, externalized_payload, and transcript_content when the durable transcript representation differs from hydrated content.

Items retain available:

  • bounded node_id, store_id, externalized_ref, and session_id metadata;
  • bounded source and role metadata;
  • content_source and content_offset;
  • occurrence counts and up to eight distinct source-path records, each preserving up to eight production-shaped {node_id, source_index} hops plus original depth/truncation metadata;
  • a bounded synthesis-context quote, its pre-provenance-cap length, and provenance-cap truncation state;
  • direct lcm_expand arguments when a locator is present;
  • locator_replay_status: unverified for present locators, never an exact-replay claim.

When hydrated payload content and durable transcript_content differ, both are preserved because both representations are present in the context block supplied to synthesis. If an externalized reference cannot be emitted intact under the metadata bound, the hydrated quote is retained without substituting a misleading transcript store_id locator.

Bounds

The response is intentionally bounded:

  • at most 24 unique evidence identities;
  • duplicate identities aggregate occurrence count and bounded distinct paths;
  • at most 500 quote characters per item;
  • at most 10,000 characters for the nested evidence_provenance object under default json.dumps serialization, enforced by dropping whole trailing items;
  • string metadata capped at 256 characters with original character count and SHA-256 disclosure when shortened;
  • a compact fixed-envelope fallback for pathological metadata;
  • whole-item truncation, never partial item removal;
  • deterministic first-seen context order;
  • explicit unique/occurrence counts, item truncation, metadata truncation, and quote-cap metadata.

The 10,000-character ceiling does not claim to bound the complete lcm_expand_query response. quote_chars_before_provenance_cap describes the synthesis-context representation before this layer's 500-character cap, not necessarily the unsliced durable source length.

locator_coverage is complete only when each unique bounded context identity has locator arguments present. It does not claim those locators still resolve. Item truncation or a selected identity without a locator makes coverage partial; no evidence yields none. context_truncated remains orthogonal: it reports that the synthesis budget omitted additional retrieval context.

Synthesis states

Path synthesis_status semantic_entailment Evidence behavior
successful answer completed not_verified preserves bounded synthesis inputs
no retrieval matches not_run not_applicable explicit empty bundle
timeout failed not_applicable preserves the already-built evidence bundle
blank synthesis result failed not_applicable preserves the already-built evidence bundle

Unexpected internal exceptions retain existing behavior and are not broadly swallowed by this PR.

Why evidence_provenance, not evidence

Open PR #452 introduces cross-session DAG expansion and an output="evidence" mode whose top-level evidence field contains bounded raw context blocks without running synthesis.

This PR deliberately uses a separate name and shape:

  • evidence remains suitable for evidence-only raw context inspection;
  • evidence_provenance is the compact answer-side binding to the context that synthesis saw.

That keeps response types stable and allows the two features to compose rather than making one field change meaning between answer and evidence-only modes.

If #452 lands first, integration requires a semantic review—not merely a mechanical rebase—to:

  • use its explicit selected-session set in retrieval_scope;
  • retain each item's originating session identity without implying that a selector authorizes access;
  • reconcile accepted session-aware expansion arguments and locator bounds;
  • verify evidence and evidence_provenance coexist without shape, privacy, or response-size collisions;
  • rerun cross-session authorization and Unicode/metadata cap probes on the combined behavior.

Caller-supplied session IDs should be described as retrieval scope/filter selectors, not authentication or authorization tokens.

Security and privacy boundary

This PR does not broaden the retrieval performed by lcm_expand_query; it serializes only bounded data already present in the context_blocks supplied to the auxiliary model for this tool call. It does make recursively selected excerpts newly visible in the returned tool payload, so clients, gateways, and logs that retain complete responses will retain these additional bounded excerpts.

This PR adds no authorization mechanism. identifiers_are_authority is always false: IDs locate evidence but do not authenticate a principal, grant access, or establish replay safety. Current expansion behavior is asymmetric and remains so:

  • node_id and externalized_ref expansion retain current-session checks;
  • store_id is an intentional cross-session locator in the existing core API;
  • a host/policy layer must authorize expansion before invoking it and must not treat any returned identifier as a capability.

locator_replay_safety is not_guaranteed, and each emitted locator is unverified: current scalar locators do not detect database replacement, ID reuse, deletion, or content revision. This is intentionally compatible with the future principal-scoped authorization work in issue #473 and provenance-bound handle contract in #476 without claiming either contract early.

Relationship to adjacent work

Read-only GitHub API inspection on 2026-08-03 confirmed #266 is closed/merged, #436/#443/#452/#461 are open pull requests, and #417/#473/#476 are open issues. No direct existing contribution was found that preserves answer-side lcm_expand_query synthesis-context provenance.

Non-goals

This PR does not add:

  • claim-level citations;
  • semantic entailment verification;
  • contradiction resolution;
  • model-authored source selection;
  • a new provider/model call;
  • ranking or retrieval behavior changes;
  • cross-session retrieval by itself;
  • principal identity/authorization architecture;
  • storage/schema migrations;
  • durable state or configuration;
  • changes to lcm_recall.

Implementation guide

  • tools.py
    • snapshots the retrieval session once per lcm_expand_query request and threads it through retrieval, recursive expansion, externalized hydration, and provenance serialization;
    • builds the bounded provenance object from the exact synthesis context_blocks;
    • preserves summary/message/externalized/transcript source identities and appends each selected DAG item's final {node_id, source_index} edge before path normalization and dedupe;
    • attaches provenance derived from the same bounded context blocks to success, no-match, timeout, and blank-result paths.
  • schemas.py
    • documents locator coverage, unverified replay/semantic entailment, and the actual asymmetric non-authoritative identifier boundary.
  • docs/retrieval-tools.md
    • documents nested-only default-JSON bounds, metadata truncation/digests, occurrence/path aggregation, response semantics, compatibility, and non-claims.
  • tests/test_expand_query_provenance.py
    • covers success/no-match/degraded paths, request-scoped session identity across rebind, complete root/recursive/message DAG paths, raw-window offsets, externalized/transcript dual representation, Unicode escaping, oversized metadata/session/ref inputs, unverified locators, occurrence/path aggregation, deterministic bounds/order/dedupe, contradictory sources, context truncation, and schema wording.

Validation

Environment: Python 3.11.15, pytest 9.1.1, pydantic 2.13.4.

Current-main semantic rebase verification (b6288eb0c339e05608b3ed933f0f88988e3470a7 → db9e03ba729784c911636716bc2447e6fbfee98b):

  • 17 passed in tests/test_expand_query_provenance.py;
  • 28 passed, 724 deselected for neighboring expand_query / exact-reference behavior;
  • 4 passed in the tool-contract gate;
  • 1118 passed, 1 skipped across the isolated provenance/core/engine/packaging/tool-contract matrix;
  • Ruff, compilation, base-range diff checks, final tree re-pin, and schema raw-byte checks passed.

The ordinary/low-FD full-suite and release-validator results below were collected on the original pre-rebase candidate and are retained as historical baseline-classification evidence. Fresh upstream CI on the rebased head is the authoritative current-main cross-version gate.

  • Focused validation: pytest tests/test_expand_query_provenance.py -q -o addopts= -> 17 passed
  • Neighboring behavior: pytest tests/test_lcm_engine.py tests/test_host_capability.py tests/test_packaging_install.py -q -o addopts= -k 'expand_query' -> 28 passed, 758 deselected
  • Default focused suite: pytest tests/test_lcm_core.py tests/test_lcm_engine.py tests/test_packaging_install.py tests/test_tool_contracts.py -q -o addopts= -> 1071 passed, 1 skipped
  • pytest -q -o addopts= -> 2323 passed, 1 skipped, 12 xfailed, 4 failed (all four independently reproduced on untouched base; details below)
  • exact-final low-FD full suite (RLIMIT_NOFILE=256) -> the same 2323 passed, 1 skipped, 12 xfailed, 4 failed; this rules out descriptor exhaustion as the cause and reproduces the two macOS /var → /private/var path-alias assertions unchanged
  • python -m compileall -q .
  • python -m py_compile scripts/import_lossless_claw.py
  • bash -n scripts/install.sh scripts/update.sh
  • python -m ruff check schemas.py tools.py tests/test_expand_query_provenance.py
  • git diff --check upstream/main...HEAD && git diff --check && git diff --cached --check (upstream/main is the canonical repository remote; origin is the contributor fork)
  • exact-final scripts/validate_release.sh --full --keep-going -> overall exit 1 only because its full and low-FD pytest gates each report the same four base failures at 2323 passed, 1 skipped, 12 xfailed; focused pytest, benchmark smoke, stress smoke, stress release, compilation, shell syntax, and all diff checks passed
  • Workflow validation: not applicable; no workflow files changed

The four full-suite failures were independently reproduced on an untouched worktree at the exact upstream base commit (854e8699e129774fa6320cfdae00d79cb08556c9):

  • test_import_lossless_claw_externalizes_legacy_data_uri_content
  • test_import_lossless_claw_respects_externalization_path_env
  • test_path_containment_within_allowed_base
  • test_configured_externalization_path_inside_allowed_base_accepted

The first two are existing import/externalization expectation failures; the latter two are the existing macOS /var/... versus /private/var/... path-alias failures. No candidate-caused full-suite or release-validation failure was found.

Fresh disposable adversarial verification reproduced the pre-publication reviewer's attacks against the corrected helper. Twenty-four Unicode-heavy items serialized under default json.dumps to 9,717 characters after retaining the first two whole items in deterministic order. A separate 20,008-character session ID plus 10,004-character Unicode externalized ref and 2,000-character role/source metadata serialized to 5,879 characters; emitted scope metadata was 256 characters with original length and SHA-256 disclosure, and the shortened externalized ref correctly omitted locator arguments instead of substituting an inexact store_id locator. A production-shaped path generated by _bounded_source_path_payload() preserved its last eight {node_id, source_index} hops, original depth 10, and truncated: true marker in a 1,414-character provenance object. The disposable verifiers were removed after execution.

The available GitNexus index predates this base and line-maps the current diff onto unrelated lcm_inspect/lcm_recent symbols, so no structural blast-radius claim is made from it. Current-source tests and diff review are the acceptance authority.

Backward compatibility and rollback

The response field is additive and existing top-level keys/values remain unchanged. Unknown-field-tolerant callers continue to work. Strict response validators, fixed-size gateway envelopes, complete-response logs, and clients that assume the old exact payload shape must be updated and should review the additional bounded excerpt visibility.

Rollback is source-only: remove the helper, response wiring, tests, schema sentence, and documentation. There is no migration, stored-data conversion, config rollback, or cleanup operation.

Maintainer review order

  1. Review the nested serialization bound and pathological metadata fallback in tools.py.
  2. Review locator presence/replay/authorization non-claims in this description, schema, and docs.
  3. Verify success/no-match/timeout/blank wiring all use the same context-derived helper.
  4. Review recursive occurrence/path aggregation plus Unicode, oversized-metadata, raw-window, and externalized fidelity tests.
  5. Decide merge order relative to fix: guard expand-query LLM response access; degrade instead of raising #443 and feat: expand summary DAGs across sessions #452; feat: expand summary DAGs across sessions #452 requires semantic integration review if it lands first.

Notes

  • Built on the recursive pre-synthesis evidence path introduced by #266 by @Tosko4.
  • Related open pull requests: #436, #443, #452, and #461.
  • Related open issues: #417, #473, and #476.
  • Reviewer credit remains in this PR description; no Co-authored-by trailer is used because prior work and review informed the design but did not author this patch.
  • The canonical base-range diff check uses upstream/main because origin is the contributor fork.

Refs #266

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 342bb1ca2b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tools.py Outdated
Comment thread tools.py
Eva added 2 commits August 4, 2026 05:37
Build bounded tool-extracted provenance from the exact context supplied to lcm_expand_query synthesis, expose it on successful/no-match/degraded responses, document traceability and authority limits, and cover recursive/raw/externalized paths plus bounds and degradation.
@100yenadmin

Copy link
Copy Markdown
Contributor Author

Keeping this open. We reviewed it against current main to check whether it had been overtaken while it sat, and it has not been.

Verdict: not superseded. Every externally observable capability it adds is still absent upstream, and nothing upstream implements the same thing by another route — so closing it would drop working functionality rather than tidy a duplicate.

The two findings worth stating, both checkable:

  • Upstream's expand-query still returns no provenance. Its response is prompt / query / answer / model / max_tokens / context_max_tokens / context_truncated / context_pagination / node_ids / matches / raw_matches — the blocks that fed synthesis are used and then dropped. This PR is what attaches them to the answer, on the no-match and degraded paths too, not only the happy one.
  • lcm_evidence_pack is adjacent, not a replacement. Its own description says it "returns no prose answer" — it builds an evidence packet from baseline refs. This PR couples evidence to a synthesized answer. They solve different problems and neither makes the other redundant.

We also deliberately did not fold this into the maintenance roll-up (#505). That roll-up is #486→#487→#489, a genuine linear stack around retrieval references and FTS bootstrap concurrency; this touches expand-query provenance in a different area. Bundling it would be convenience rather than coherence, and would make a reviewable change harder to review.

No changes pushed — flagging the review outcome so it is not mistaken for stale.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant