feat: preserve expand-query evidence provenance - #484
100yenadmin wants to merge 2 commits into
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 342bb1ca2b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Build bounded tool-extracted provenance from the exact context supplied to lcm_expand_query synthesis, expose it on successful/no-match/degraded responses, document traceability and authority limits, and cover recursive/raw/externalized paths plus bounds and degradation.
f1909a5 to
db9e03b
Compare
|
Keeping this open. We reviewed it against current Verdict: not superseded. Every externally observable capability it adds is still absent upstream, and nothing upstream implements the same thing by another route — so closing it would drop working functionality rather than tidy a duplicate. The two findings worth stating, both checkable:
We also deliberately did not fold this into the maintenance roll-up (#505). That roll-up is #486→#487→#489, a genuine linear stack around retrieval references and FTS bootstrap concurrency; this touches expand-query provenance in a different area. Bundling it would be convenience rather than coherence, and would make a reviewable change harder to review. No changes pushed — flagging the review outcome so it is not mistaken for stale. |
Preserve bounded evidence provenance in
lcm_expand_queryanswersSummary
lcm_expand_queryalready builds rich recursive context before asking the auxiliary model to synthesize an answer. That context can contain:Today, most of that trace disappears after synthesis. The successful response returns the answer plus top-level matches, but not necessarily the exact descendants or excerpts the model saw. A fluent answer therefore cannot be traced back to the complete bounded synthesis input from the response alone.
This PR adds an additive, bounded
evidence_provenanceobject to successful, no-match, and structured degradedlcm_expand_queryresponses. The object is extracted by the tool from the exactcontext_blockssupplied to synthesis; it is not authored by the auxiliary model.No ranking, retrieval selection, recursive expansion, prompt construction, model routing, storage schema, durable state, or existing top-level response field changes.
Why
PR #266 correctly made
lcm_expand_queryrecursively descend summary DAGs and include leaf evidence before synthesis. That improved answer quality, but it also created a response-level observability gap:The same gap exists for raw FTS windows and hydrated externalized payloads: the model sees a bounded context representation, while the caller receives only search metadata or a top-level match.
This PR preserves that boundary without pretending to solve replay verification, answer verification, or authorization. It explicitly separates:
Only bounded locator coverage and context representation are implemented here. Locator replay remains explicitly unverified.
Response contract
Every successful answer, no-match result, and currently structured degraded
lcm_expand_queryresponse receives:{ "evidence_provenance": { "retrieval_scope": { "kind": "current_session", "session_ids": ["..."] }, "identifiers_are_authority": false, "locator_replay_safety": "not_guaranteed", "synthesis_status": "completed", "locator_coverage": "complete", "semantic_entailment": "not_verified", "quote_origin": "tool_extracted_from_synthesis_context", "context_truncated": false, "context_unique_item_count": 3, "context_occurrence_count": 4, "items_truncated": false, "serialized_char_limit": 10000, "serialization": { "scope": "evidence_provenance_only", "json_ensure_ascii": true }, "quote_max_chars": 500, "quotes_truncated_by_provenance_cap": 0, "metadata_truncated": false, "items": [] } }Evidence representations
Items use three source types:
summary— root or recursively reached summary text;raw_message— stored, recursive, or raw-search message content;externalized_payload— hydrated externalized content with directexternalized_refexpansion.Raw-message and externalized items additionally use
content_sourceto distinguish the exact representation supplied to synthesis, includingraw_search_hit,externalized_payload, andtranscript_contentwhen the durable transcript representation differs from hydrated content.Items retain available:
node_id,store_id,externalized_ref, andsession_idmetadata;content_sourceandcontent_offset;{node_id, source_index}hops plus original depth/truncation metadata;lcm_expandarguments when a locator is present;locator_replay_status: unverifiedfor present locators, never an exact-replay claim.When hydrated payload content and durable
transcript_contentdiffer, both are preserved because both representations are present in the context block supplied to synthesis. If an externalized reference cannot be emitted intact under the metadata bound, the hydrated quote is retained without substituting a misleading transcriptstore_idlocator.Bounds
The response is intentionally bounded:
evidence_provenanceobject under defaultjson.dumpsserialization, enforced by dropping whole trailing items;The 10,000-character ceiling does not claim to bound the complete
lcm_expand_queryresponse.quote_chars_before_provenance_capdescribes the synthesis-context representation before this layer's 500-character cap, not necessarily the unsliced durable source length.locator_coverageiscompleteonly when each unique bounded context identity has locator arguments present. It does not claim those locators still resolve. Item truncation or a selected identity without a locator makes coveragepartial; no evidence yieldsnone.context_truncatedremains orthogonal: it reports that the synthesis budget omitted additional retrieval context.Synthesis states
synthesis_statussemantic_entailmentcompletednot_verifiednot_runnot_applicablefailednot_applicablefailednot_applicableUnexpected internal exceptions retain existing behavior and are not broadly swallowed by this PR.
Why
evidence_provenance, notevidenceOpen PR #452 introduces cross-session DAG expansion and an
output="evidence"mode whose top-levelevidencefield contains bounded raw context blocks without running synthesis.This PR deliberately uses a separate name and shape:
evidenceremains suitable for evidence-only raw context inspection;evidence_provenanceis the compact answer-side binding to the context that synthesis saw.That keeps response types stable and allows the two features to compose rather than making one field change meaning between answer and evidence-only modes.
If #452 lands first, integration requires a semantic review—not merely a mechanical rebase—to:
retrieval_scope;evidenceandevidence_provenancecoexist without shape, privacy, or response-size collisions;Caller-supplied session IDs should be described as retrieval scope/filter selectors, not authentication or authorization tokens.
Security and privacy boundary
This PR does not broaden the retrieval performed by
lcm_expand_query; it serializes only bounded data already present in thecontext_blockssupplied to the auxiliary model for this tool call. It does make recursively selected excerpts newly visible in the returned tool payload, so clients, gateways, and logs that retain complete responses will retain these additional bounded excerpts.This PR adds no authorization mechanism.
identifiers_are_authorityis alwaysfalse: IDs locate evidence but do not authenticate a principal, grant access, or establish replay safety. Current expansion behavior is asymmetric and remains so:node_idandexternalized_refexpansion retain current-session checks;store_idis an intentional cross-session locator in the existing core API;locator_replay_safetyisnot_guaranteed, and each emitted locator isunverified: current scalar locators do not detect database replacement, ID reuse, deletion, or content revision. This is intentionally compatible with the future principal-scoped authorization work in issue #473 and provenance-bound handle contract in #476 without claiming either contract early.Relationship to adjacent work
lcm_expand_query. This PR preserves that selected evidence's identity/excerpt information after synthesis.output="evidence". This PR reservesevidence_provenancefor the answer-side compact binding so both contracts can coexist.lcm_recall, with retrieval: add exhaustive citable recall mode #461 deliberately extracting the narrower retrieval-only lane from [wave 1 · R2] Trajectory/experience-memory subsystem + scaling fixes + citable delivery + benchmark evidence #436. This PR targets a different tool and makes a weaker claim: source traceability forlcm_expand_query, not answer-ready citations or semantic support.lcm_recall; this PR reuses the bounded/detail/expandability direction forlcm_expand_querywithout changing recall.Read-only GitHub API inspection on 2026-08-03 confirmed #266 is closed/merged, #436/#443/#452/#461 are open pull requests, and #417/#473/#476 are open issues. No direct existing contribution was found that preserves answer-side
lcm_expand_querysynthesis-context provenance.Non-goals
This PR does not add:
lcm_recall.Implementation guide
tools.pylcm_expand_queryrequest and threads it through retrieval, recursive expansion, externalized hydration, and provenance serialization;context_blocks;{node_id, source_index}edge before path normalization and dedupe;schemas.pydocs/retrieval-tools.mdtests/test_expand_query_provenance.pyValidation
Environment: Python 3.11.15, pytest 9.1.1, pydantic 2.13.4.
Current-main semantic rebase verification (
b6288eb0c339e05608b3ed933f0f88988e3470a7→db9e03ba729784c911636716bc2447e6fbfee98b):17 passedintests/test_expand_query_provenance.py;28 passed, 724 deselectedfor neighboringexpand_query/ exact-reference behavior;4 passedin the tool-contract gate;1118 passed, 1 skippedacross the isolated provenance/core/engine/packaging/tool-contract matrix;The ordinary/low-FD full-suite and release-validator results below were collected on the original pre-rebase candidate and are retained as historical baseline-classification evidence. Fresh upstream CI on the rebased head is the authoritative current-main cross-version gate.
pytest tests/test_expand_query_provenance.py -q -o addopts=->17 passedpytest tests/test_lcm_engine.py tests/test_host_capability.py tests/test_packaging_install.py -q -o addopts= -k 'expand_query'->28 passed, 758 deselectedpytest tests/test_lcm_core.py tests/test_lcm_engine.py tests/test_packaging_install.py tests/test_tool_contracts.py -q -o addopts=->1071 passed, 1 skippedpytest -q -o addopts=->2323 passed, 1 skipped, 12 xfailed, 4 failed(all four independently reproduced on untouched base; details below)RLIMIT_NOFILE=256) -> the same2323 passed, 1 skipped, 12 xfailed, 4 failed; this rules out descriptor exhaustion as the cause and reproduces the two macOS/var→/private/varpath-alias assertions unchangedpython -m compileall -q .python -m py_compile scripts/import_lossless_claw.pybash -n scripts/install.sh scripts/update.shpython -m ruff check schemas.py tools.py tests/test_expand_query_provenance.pygit diff --check upstream/main...HEAD && git diff --check && git diff --cached --check(upstream/mainis the canonical repository remote;originis the contributor fork)scripts/validate_release.sh --full --keep-going-> overall exit 1 only because its full and low-FD pytest gates each report the same four base failures at2323 passed, 1 skipped, 12 xfailed; focused pytest, benchmark smoke, stress smoke, stress release, compilation, shell syntax, and all diff checks passedThe four full-suite failures were independently reproduced on an untouched worktree at the exact upstream base commit (
854e8699e129774fa6320cfdae00d79cb08556c9):test_import_lossless_claw_externalizes_legacy_data_uri_contenttest_import_lossless_claw_respects_externalization_path_envtest_path_containment_within_allowed_basetest_configured_externalization_path_inside_allowed_base_acceptedThe first two are existing import/externalization expectation failures; the latter two are the existing macOS
/var/...versus/private/var/...path-alias failures. No candidate-caused full-suite or release-validation failure was found.Fresh disposable adversarial verification reproduced the pre-publication reviewer's attacks against the corrected helper. Twenty-four Unicode-heavy items serialized under default
json.dumpsto 9,717 characters after retaining the first two whole items in deterministic order. A separate 20,008-character session ID plus 10,004-character Unicode externalized ref and 2,000-character role/source metadata serialized to 5,879 characters; emitted scope metadata was 256 characters with original length and SHA-256 disclosure, and the shortened externalized ref correctly omitted locator arguments instead of substituting an inexactstore_idlocator. A production-shaped path generated by_bounded_source_path_payload()preserved its last eight{node_id, source_index}hops, original depth 10, andtruncated: truemarker in a 1,414-character provenance object. The disposable verifiers were removed after execution.The available GitNexus index predates this base and line-maps the current diff onto unrelated
lcm_inspect/lcm_recentsymbols, so no structural blast-radius claim is made from it. Current-source tests and diff review are the acceptance authority.Backward compatibility and rollback
The response field is additive and existing top-level keys/values remain unchanged. Unknown-field-tolerant callers continue to work. Strict response validators, fixed-size gateway envelopes, complete-response logs, and clients that assume the old exact payload shape must be updated and should review the additional bounded excerpt visibility.
Rollback is source-only: remove the helper, response wiring, tests, schema sentence, and documentation. There is no migration, stored-data conversion, config rollback, or cleanup operation.
Maintainer review order
tools.py.Notes
Co-authored-bytrailer is used because prior work and review informed the design but did not author this patch.upstream/mainbecauseoriginis the contributor fork.Refs #266