Skip to content

feat(workspace): add tiered context summaries - #1566

Closed
G7CNF wants to merge 3 commits into
nearai:mainfrom
G7CNF:codex/issue-1473-tiered-context-summaries
Closed

G7CNF wants to merge 3 commits into
nearai:mainfrom
G7CNF:codex/issue-1473-tiered-context-summaries

Conversation

@G7CNF

@G7CNF G7CNF commented Mar 22, 2026

Copy link
Copy Markdown
Contributor

Implements the first slice of #1473.

What changed:

  • add L0/L1 summary fields to workspace documents
  • generate and persist summaries after writes for sufficiently large docs
  • default workspace search to L1 content, with explicit L0/L2 detail levels
  • expose summaries in memory tool output
  • add schema migrations for postgres and libsql
  • add unit and integration coverage for summary generation and search detail levels

Validation:

  • cargo fmt --all -- --check
  • cargo check -q
  • cargo clippy -p ironclaw --lib --all-features -- -D warnings
  • cargo test -p ironclaw --lib workspace::tests::test_generate_document_summaries_parses_json_response -- --nocapture
  • cargo test -p ironclaw --lib workspace::search::tests::test_search_detail_levels_choose_content -- --nocapture
  • cargo test -p ironclaw --test workspace_integration test_workspace_tiered_summaries_default_search_returns_l1 -- --nocapture

Note:

  • rust-pr-preflight still reports existing repository baseline boundary issues unrelated to this branch.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, you can upgrade your account or add credits to your account and enable them for code reviews in your settings.

@github-actions github-actions Bot added scope: tool/builtin Built-in tools scope: db Database trait / abstraction scope: db/postgres PostgreSQL backend scope: db/libsql libSQL / Turso backend scope: workspace Persistent memory / workspace scope: docs Documentation size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Mar 22, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces tiered context summaries (L0/L1) to workspace documents, enhancing search capabilities and memory tool output. It includes automatic summary generation, schema migrations for database compatibility, and extensive testing to ensure functionality and reliability. The changes aim to improve search result relevance and provide more structured context within the workspace.

Highlights

  • Tiered Context Summaries: Introduces L0 and L1 summary fields to workspace documents, enhancing search and memory tool output.
  • Summary Generation and Persistence: Automatically generates and persists summaries after document writes for sufficiently large documents, optimizing search detail levels.
  • Default Workspace Search: Configures workspace search to default to L1 content, with options for explicit L0 and L2 detail levels, improving search result relevance.
  • Schema Migrations: Includes schema migrations for both PostgreSQL and LibSQL to support the new summary fields, ensuring database compatibility.
  • Testing and Validation: Adds comprehensive unit and integration tests for summary generation and search detail levels, ensuring functionality and reliability.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces tiered context summaries (L0/L1) for workspace documents, enhancing search functionality and document previews. The database schema is updated to include summary_l0 (one-line abstract) and summary_l1 (structured overview) columns. Document updates now trigger asynchronous background generation of these summaries using an LLM, with a backfill process initiated on application startup. Search results can now specify a detail level (L0, L1, or raw content) for the returned content, defaulting to L1. Review comments suggest improving error handling for prompt injection checks in summary generation, considering a retry mechanism for LLM failures, logging warnings when summaries are unexpectedly missing, and potentially triggering synchronous summary updates in the database layer for immediate consistency.

Comment thread src/workspace/mod.rs Outdated
Comment thread migrations/V1__initial.sql
Comment thread src/db/libsql/workspace.rs
Comment thread src/workspace/mod.rs Outdated
Comment thread src/workspace/search.rs

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review: feat(workspace): add tiered context summaries

Verdict: REQUEST_CHANGES

Solid feature design with clean L0/L1/L2 tiering model. Dual-backend support is properly implemented (PostgreSQL + libSQL migrations, trait method, both implementations). Prompt injection defense via reject_if_injected is good. Three issues need addressing.

Database Dual-Backend Compliance: PASS

Both backends properly handled -- V14 migration for both, update_document_summaries in trait and both impls, column indices shifted correctly in libSQL, summaries NULLed on content update.

Must Fix

1. No content truncation before LLM call

generate_document_summaries passes entire document content to the LLM with no size limit. A 500KB document will blow the context window or incur massive token costs. with_max_tokens(1024) only limits output tokens.

Fix: Truncate content to a reasonable limit (e.g., first ~32KB) before sending to the LLM.

2. Unbounded backfill on startup

backfill_summaries loads ALL documents into memory and iterates sequentially with LLM calls. For workspaces with hundreds of documents, this hammers the LLM endpoint at startup with no concurrency limit, rate limiting, or batch cap.

Fix: Add a configurable batch limit (e.g., max N documents per startup) and/or a concurrency semaphore.

3. Error variant misuse

Summary generation/persistence/JSON parse failures all map to WorkspaceError::SearchFailed. This produces confusing error messages and logs.

Fix: Add WorkspaceError::SummaryGenerationFailed { reason: String } in src/error.rs following the thiserror convention.

Suggestions (non-blocking)

  1. Redundant data in tool output -- memory_search returns content (tier-selected) plus summary_l0 and summary_l1 as separate fields. Consider omitting raw summary fields for the selected tier.

  2. Double word-count check -- schedule_summary_refresh and refresh_document_summaries both check SUMMARY_MIN_WORDS. Defensive but redundant.

  3. extract_json_object fragility -- outer-brace matching (find('{')/rfind('}')) could pick wrong boundaries with nested braces. Acceptable for now given prompt design and serde_json validation.

  4. No retry on LLM failure -- summary silently lost until next content update. Consider marking documents as "needs summary" for the next backfill.

Comment thread src/workspace/mod.rs
Comment thread src/workspace/mod.rs
Comment thread src/tools/builtin/memory.rs Outdated
Comment thread src/workspace/repository.rs Outdated
@github-actions github-actions Bot added the scope: agent Agent core (agent loop, router, scheduler) label Apr 6, 2026
@G7CNF
G7CNF force-pushed the codex/issue-1473-tiered-context-summaries branch from 30af8d6 to 252d141 Compare April 6, 2026 20:31
@henrypark133
henrypark133 changed the base branch from staging to main May 1, 2026 06:18
@ilblackdragon

Copy link
Copy Markdown
Member

Code Review — feat(workspace): add tiered context summaries

Overview

Adds L0 (one-sentence abstract) and L1 (≤500-word structured overview) summary fields to memory_documents. Summaries are generated by a (preferably cheap) LLM after each write and during a startup backfill. Search now defaults to returning L1 content, with l0/l1/l2 selectable via the memory_search tool and SearchConfig.detail. Both Postgres and libSQL/Turso backends are covered with a migration (V14) and an updated fresh-init schema.

What's good

  • Dual-backend coverage: schema added in V1__initial.sql (fresh) + V14__... (existing) + libsql_migrations.rs. All search read paths in repository.rs and db/libsql/workspace.rs updated consistently.
  • Defense in depth on LLM output: reject_if_injected scans both summary_l0 and summary_l1 before persisting (src/workspace/mod.rs:1039-1048).
  • Stale-write guard: re-reads current_content after the LLM call and skips persistence if the document changed mid-generation (src/workspace/mod.rs:1025-1037).
  • Retry with backoff for transient LLM errors via is_retryable/retry_backoff_delay.
  • UTF-8 safety: truncate_summary_input uses floor_char_boundary — covered by a dedicated test.
  • Startup backfill is spawned, not awaited, so it doesn't block boot (src/app.rs:937).

Correctness issues

  • Workspace::search hardcodes SearchDetailLevel::L1, overriding search_defaults.detail (src/workspace/mod.rs:1935):
    ```rust
    self.search_with_config(query, self.search_defaults.clone().with_limit(limit).with_detail(L1))
    ```
    This silently discards any env-configured default detail for the simple search() path while search_with_config honors it. Either drop the override or remove detail from SearchConfig defaults — current state is inconsistent.

  • refresh_summaries_after_write is awaited inline on every write. update_document clears summary_l0/l1 = NULL and then the caller (e.g., memory_write tool) blocks on a full LLM round-trip + retry budget before returning. For append_daily_log and any high-frequency write path this stacks up latency for the agent and the user. Consider:
    ```rust
    tokio::spawn(async move { refresh_document_summaries(...).await });
    ```
    The integration test (test_workspace_tiered_summaries_default_search_returns_l1) already uses a 5-second polling loop, which suggests the author considered async-ness — clarifying which model is intended would help.

  • extract_json_object is naive (src/workspace/mod.rs:952): first { to last } works for clean output but won't survive an LLM that returns a fenced code block plus a trailing JSON example. Low-risk but worth a regression test for fenced output.

  • Length cap not enforced on parsed summaries. The prompt asks for ≤30 / ≤500 words but the LLM can ignore that. Consider clamping summary_l1 to a hard char count after parse — otherwise a misbehaving model can blow up directory listings (list_workspace_files returns summary_l0 as the preview).

Style / convention

  • SUMMARY_BACKFILL_LIMIT is read directly from std::env::var in src/workspace/mod.rs:1111. Project convention (CLAUDE.md, src/config/) is that env-driven settings live in src/config/*.rs. This should be in WorkspaceConfig or a sibling, so it's discoverable via Config and testable.

  • Scope creep: src/agent/routine_engine.rs un-cfg(test)s sanitize_summary / strip_html_tags and wires them into complete_dispatched_run (lines 606-741). That's a reasonable hardening, but it's unrelated to tiered summaries and should be a separate PR — the title doesn't mention it and reviewers won't expect it.

  • reject_if_injected warns can spam. choose_search_content emits a warn! whenever an L0/L1 search hits a chunk without summaries (src/workspace/search.rs:1432). Short documents (<200 words) never get summaries by design, so every search over a fresh workspace will log a warning per result. Downgrade to debug! or only warn once per document_id per process.

  • with_llm fallback to the main LLM: in src/app.rs, cheap_llm.cloned().unwrap_or_else(|| Arc::clone(llm)) means a deployment without CHEAP_LLM_* configured silently uses the main (potentially expensive) model for summaries. Worth a tracing::warn! at init telling the operator they're paying main-model rates for summaries.

  • crate::llm::* imports (src/workspace/mod.rs:649-651): per CLAUDE.md, LLM types live in ironclaw_llm and should be imported as use ironclaw_llm::{...}. The crate::llm::* shim still works but new code should use the extracted-crate path.

Performance

  • Backfill walks newest 50 docs, ordered by updated_at DESC. Old docs without summaries never get backfilled across restarts. A partial index on summary_l0 IS NULL and a WHERE summary_l0 IS NULL clause would amortize backfill over many startups.
  • Each write does: index → null summaries → LLM call → write summaries. The intermediate NULL window means a search between write and summary refresh falls back to raw content. With the inline await this window is small but nonzero; with a spawned refresh, it widens. Acceptable but document the semantics.

Test coverage

Solid baseline: JSON parser test, char-boundary test, L0/L1/L2 selection test, integration test with a StubLlm. Gaps:

  • No test for the stale-write guard (write A → write B → A's summary should be rejected).
  • No test for update_document clearing summaries to NULL.
  • No test for SUMMARY_BACKFILL_LIMIT=0 disabling backfill or invalid-value fallback to default.
  • The integration test's tokio::time::timeout(5s) poll loop is a smell — if the refresh is meant to be synchronous, a single search() should suffice; if it's meant to be async, this should be tightened or made deterministic via a notification channel.

Security

  • LLM-generated summaries are sanitizer-scanned ✓
  • Summaries are persisted but not re-scanned on read — fine, since write-time scan is the gate.
  • LLM is fed raw user document content; a hostile document could engineer the L1 summary to contain instructions, then any later context including the L1 summary could carry those instructions. The injection scanner catches obvious patterns but not subtle ones. Worth a callout in the workspace README.

Recommendation

Request changes on:

  1. Drop the with_detail(L1) override in search() or document why it diverges from configured defaults.
  2. Move SUMMARY_BACKFILL_LIMIT into src/config/.
  3. Decide sync vs. async refresh, then either remove the polling in the integration test or document the async semantics.
  4. Pull the routine_engine.rs sanitizer changes into a separate PR.
  5. Downgrade the missing-summary log to debug!.

Nice-to-have: length-clamp parsed summaries; warn at init if summary_llm is the main LLM; add tests for stale-write guard, NULL-on-update, and BACKFILL_LIMIT=0.

@serrrfirat

Copy link
Copy Markdown
Collaborator

Thank you for exploring tiered L0/L1 summaries to reduce memory-search context cost.\n\nWe are closing this legacy implementation because Reborn now retrieves bounded, scope-checked, sanitized memory snippets through the MemoryService/native provider boundary. It does not currently persist L0/L1 summaries, and we do not want to carry over the v1 database schema and summary-generation lifecycle without benchmark evidence that the existing Reborn retrieval path needs another persisted tier.\n\nReferences:\n- https://github.com/nearai/ironclaw/blob/main/crates/ironclaw_memory/src/service.rs\n- https://github.com/nearai/ironclaw/blob/main/crates/ironclaw_memory_native/src/service.rs\n- https://github.com/nearai/ironclaw/pull/5327\n\nIf Reborn memory benchmarks later show context-quality or token-cost problems, tiered summaries can be reconsidered as a provider-owned design with explicit invalidation and cost policy. This is an architecture-transition closure, not a reflection on the quality of your work.

@serrrfirat serrrfirat closed this Jul 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: agent Agent core (agent loop, router, scheduler) scope: db/libsql libSQL / Turso backend scope: db/postgres PostgreSQL backend scope: db Database trait / abstraction scope: docs Documentation scope: tool/builtin Built-in tools scope: workspace Persistent memory / workspace size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants