Skip to content

feat(memory): schedule STMO jobs after conversation persistence (NoOp writer — follow-up for Oracle wiring) - #1346

Closed
spalimpaaces-star wants to merge 5 commits into
smg-project:mainfrom
spalimpaaces-star:feat/stmo-scheduling
Closed

spalimpaaces-star wants to merge 5 commits into
smg-project:mainfrom
spalimpaaces-star:feat/stmo-scheduling

Conversation

@spalimpaaces-star

@spalimpaaces-star spalimpaaces-star commented Apr 23, 2026 •

Copy link
Copy Markdown
Contributor

Description

Problem

  • Needs to support Short-Term Memory Optimization (STMO) scheduling for responses api.
  • STMO jobs should be enqueued after conversation items are persisted, closer to where the conversation state lives.

Solution

  • Thread stm_enabled and stm_condenser_model_id from the x-conversation-memory-config header through MemoryHeaderView and MemoryExecutionContext
  • Wire ConversationMemoryWriter and MemoryExecutionContext into persist_conversation_items; gRPC path uses a NoOp writer
  • Turn count is computed at route time (after history is assembled) and threaded through to persistence — by persistence time only the current request payload is available, not the full conversation history
  • Add enqueue_stmo_if_needed — fires at turns 4, 7, 10, 13... matching the upstream worker trigger contract
  • Compute turn count from the fully assembled conversation input after history is loaded, not from the current request payload alone
  • Gate stm_enabled by the runtime memory switch — header alone is not sufficient to activate STMO scheduling
  • Best-effort: enqueue failures are logged and swallowed, never fail the request
  • Real writer implementation is a follow-up PR

Changes

model_gateway/src/memory/context.rs

  • Add stm_enabled and stm_condenser_model_id to MemoryExecutionContext
  • Gate stm_enabled by runtime.enabled — same pattern as LTM store_ltm and recall

model_gateway/src/routers/common/header_utils.rs

  • Add stm_enabled and stm_condenser_model_id fields to MemoryHeaderView
  • Parse both from x-conversation-memory-config JSON header

model_gateway/src/routers/openai/responses/route.rs

  • Remove redundant second parse of x-conversation-memory-config
  • After load_input_history, compute conversation_user_turn_count (gated on stm_enabled) and store in ResponsesPayloadState

model_gateway/src/routers/openai/context.rs

  • Add conversation_user_turn_count: Option<usize> to ResponsesPayloadState and StorageHandles
  • Wire through into_streaming_context

model_gateway/src/routers/openai/responses/history.rs

  • Add count_conversation_user_turns(input: &ResponseInput) -> usize — counts from typed ResponseInput variants, handles both Text and Items

model_gateway/src/routers/common/persistence_utils.rs

  • Add conversation_memory_writer, memory_execution_context, and conversation_user_turn_count params to persist_conversation_items
  • Add enqueue_stmo_if_needed — fires at turns 4, 7, 10, 13...
  • Turn count uses precomputed value from assembled history, not local JSON counting
  • Condenser model is optional — enqueues without it for worker fallback parity
  • Best-effort: failures logged and swallowed, never fail the request

model_gateway/src/routers/openai/responses/non_streaming.rs

  • Pass conversation_memory_writer, memory_execution_context, and conversation_user_turn_count to persist_conversation_items

model_gateway/src/routers/openai/responses/streaming.rs

  • Pass conversation_memory_writer, memory_execution_context, and conversation_user_turn_count at both persist_conversation_items call sites

model_gateway/src/routers/grpc/common/responses/utils.rs

  • Pass NoOpConversationMemoryWriter and MemoryExecutionContext::default() — gRPC path never has STMO enabled

Test Plan

  • Added unit test Turn boundary unit test: fires at 4, 7, 10, 13; skips 1–3, 5–6, 8–9
  • Enqueue test: verifies condenser_model, last_index, target_item_end in job config
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • New Features
    • Short-term memory (STM) is now configurable per request with optional condenser model selection.
    • Conversation turn tracking implemented to enable intelligent memory management.
    • Deterministic memory condensation scheduling: first triggered at conversation turn 4, then every 3 turns thereafter.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@coderabbitai

coderabbitai Bot commented Apr 23, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR extends the conversation memory system with Short-Term Memory (STM) configuration and enqueuing logic. It adds STM enablement flags and condenser model IDs to context and header objects, implements deterministic user-turn counting, and conditionally creates STMO enqueue tasks during conversation persistence (triggered at turn 4, then every 3 turns thereafter).

Changes

Cohort / File(s) Summary
Memory Configuration & Headers
model_gateway/src/memory/context.rs, model_gateway/src/routers/common/header_utils.rs
Added stm_enabled boolean and optional stm_condenser_model_id fields to MemoryExecutionContext and MemoryHeaderView. Header extraction gates STM enablement against runtime capability; condenser model ID is read from headers only when STM is enabled.
Conversation Persistence & STM Enqueue
model_gateway/src/routers/common/persistence_utils.rs
Extended persist_conversation_items to accept ConversationMemoryWriter and MemoryExecutionContext. Implements conditional STMO enqueue with deterministic trigger (eligible at user turn 4, then every 3 turns). Builds memory_config JSON with last_index, target_item_end, and optional condenser_model, creating NewConversationMemory rows of type Stmo with status Ready. Enqueue failures are logged as warnings and do not block persistence.
Response Context & User Turn Threading
model_gateway/src/routers/openai/context.rs, model_gateway/src/routers/grpc/common/responses/utils.rs
Added conversation_user_turn_count: Option<usize> field to ResponsesPayloadState and StorageHandles. Response flow now threads turn count alongside memory context through streaming/storage layers.
User Turn Counting & Route Logic
model_gateway/src/routers/openai/responses/history.rs, model_gateway/src/routers/openai/responses/route.rs
Replaced no-op inject_memory_context with deterministic count_conversation_user_turns function. Counts ResponseInput::Text as one turn and scans ResponseInput::Items for user-role Message and SimpleInputMessage entries. Route conditionally computes and stores turn count in payload state when STM is enabled.
Response Handlers
model_gateway/src/routers/openai/responses/non_streaming.rs, model_gateway/src/routers/openai/responses/streaming.rs
Updated response persistence calls to supply conversation memory writer, execution context, and user turn count to persist_conversation_items function.

Sequence Diagram(s)

sequenceDiagram
    participant Client
    participant Route as Route Handler
    participant Context as Memory Context
    participant Persister as persist_conversation_items
    participant Writer as ConversationMemoryWriter
    participant DB as Database

    Client->>Route: ResponsesRequest (with STM enabled)
    Route->>Route: count_conversation_user_turns(input)
    Route->>Context: Store turn_count in ResponsesPayloadState
    Route->>Persister: Call with memory_writer, memory_context, turn_count
    
    Persister->>Persister: Link input/output items to conversation
    Persister->>Persister: Check: stm_enabled && eligible_turn?<br/>(turn 4, then every 3)
    
    alt Turn is Eligible for STMO Enqueue
        Persister->>Persister: Compute target_item_end<br/>Build memory_config JSON<br/>(condenser_model, last_index, target_item_end)
        Persister->>Writer: Create NewConversationMemory<br/>(type: Stmo, status: Ready)
        Writer->>DB: INSERT memory record
        DB-->>Writer: Success/Failure
        Writer-->>Persister: Result (warn on failure)
    else Turn Not Eligible
        Persister->>Persister: Skip STMO enqueue
    end
    
    Persister-->>Route: Persistence complete
    Route-->>Client: Response
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~27 minutes

Possibly related PRs

Suggested labels

data-connector, tests

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237

Poem

🐰 A fluffy tail wiggles with delight,
STM memories now queue just right—
Four turns first, then threes—hooray!
The condenser hops, data to weigh! ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title clearly and specifically summarizes the main change: scheduling STMO (Short-Term Memory Optimization) jobs after conversation persistence, with a note about the NoOp writer being a follow-up. It directly matches the core objectives outlined in the PR description.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@mergify

mergify Bot commented Apr 23, 2026

Copy link
Copy Markdown
Contributor

Hi @spalimpaaces-star, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes openai OpenAI router changes labels Apr 23, 2026
spalimpaaces-star added 3 commits April 22, 2026 23:09
…cutionContext

x-conversation-memory-config was parsed twice per request — once in
middleware to build MemoryExecutionContext, and again in route_responses
for the inject_memory_context no-op stub. Remove the second parse and
the stub entirely.
Signed-off-by: Saikiran Palimpati <saikiran.palimpati@oracle.com>

Signed-off-by: spalimpaaces-star <saikiran.palimpati@oracle.com>
… into persist_conversation_items

Signed-off-by: saikiranpalimpati <34260562+saikiranpalimpati@users.noreply.github.com>
Signed-off-by: spalimpaaces-star <saikiran.palimpati@oracle.com>
…n flow

Signed-off-by: saikiranpalimpati <34260562+saikiranpalimpati@users.noreply.github.com>
Signed-off-by: spalimpaaces-star <saikiran.palimpati@oracle.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a3b0237ffb

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/routers/common/persistence_utils.rs Outdated
Comment thread model_gateway/src/memory/context.rs Outdated
@spalimpaaces-star
spalimpaaces-star marked this pull request as draft April 23, 2026 07:38
spalimpaaces-star added 2 commits April 23, 2026 02:45
Previously enqueue_stmo_if_needed derived the user turn count from
original_body.input at persistence time, which only contains the current
request payload. STMO trigger boundaries (4/7/10/...) were never reached
for multi-turn conversations.

Fix: count_conversation_user_turns runs over request_body.input after
load_input_history assembles the full conversation context. The count is
threaded through ResponsesPayloadState and StorageHandles to both the
streaming and non-streaming persistence paths.

Signed-off-by: spalimpaaces-star <saikiran.palimpati@oracle.com>
stm_enabled was set directly from the request header without checking
runtime.enabled, allowing STMO scheduling to be activated even when the
master memory feature gate is off. Apply the same runtime guard used by
store_ltm and recall.

Signed-off-by: spalimpaaces-star <saikiran.palimpati@oracle.com>
@spalimpaaces-star
spalimpaaces-star marked this pull request as ready for review April 23, 2026 10:50

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1c910fe89d

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +182 to +185
let conversation_user_turn_count = if ctx.memory_execution_context.stm_enabled {
Some(super::history::count_conversation_user_turns(
&request_body.input,
))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use full conversation length for STMO turn count

conversation_user_turn_count is derived from request_body.input, but in conversation-id flows that input is assembled by load_input_history, which still limits history to MAX_CONVERSATION_HISTORY_ITEMS = 100 (history.rs) before this count is taken. Fresh evidence: the loader fetches a capped slice, so after long conversations the count stops reflecting total user turns, and STMO boundary checks (4/7/10/...) will fire at the wrong times or not at all. This directly affects production threads once they exceed the history cap.

Useful? React with 👍 / 👎.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
model_gateway/src/routers/grpc/common/responses/utils.rs (1)

154-171: 🧹 Nitpick | 🔵 Trivial

LGTM — double-gated STMO inhibition on the gRPC path.

Passing NoOpConversationMemoryWriter, a default MemoryExecutionContext (stm_enabled=false), and None for the turn count ensures the gRPC persistence path never enqueues STMO jobs, matching the PR's "NoOp writer — follow-up for Oracle wiring" scope.

Minor nit (optional): Arc::new(NoOpConversationMemoryWriter::new()) allocates on every persisted response. If this path is hot, consider a OnceLock<Arc<dyn ConversationMemoryWriter>> or storing the Arc on a shared components struct to avoid per-call allocation. Not required for this PR.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/grpc/common/responses/utils.rs` around lines 154 -
171, The current code calls Arc::new(NoOpConversationMemoryWriter::new()) each
time before calling persist_conversation_items which may allocate on every hot
path; change this to a shared, lazily-initialized Arc so the
NoOpConversationMemoryWriter is created once and reused (e.g., store Arc<dyn
ConversationMemoryWriter> in a static OnceLock or on the shared components
struct) and pass that shared Arc into persist_conversation_items instead of
creating a new one per call.
model_gateway/src/routers/common/header_utils.rs (1)

496-511: 🧹 Nitpick | 🔵 Trivial

Add STM assertions to this test to lock in the independence of STM and LTM gates.

The input header already carries "short_term_memory":{"enabled":true,"condenser_model_id":"cond-1"}, but no assertion covers the STM fields. Per the new from_http_headers logic, STM should remain enabled regardless of LTM state — adding the assertions below exercises that contract and guards against an accidental future coupling.

✏️ Proposed test additions
         let view = MemoryHeaderView::from_http_headers(&headers);

         assert_eq!(view.policy, None);
         assert_eq!(view.subject_id, None);
         assert_eq!(view.embedding_model, None);
         assert_eq!(view.extraction_model, None);
+        // STM is independent of the LTM enabled flag.
+        assert!(view.stm_enabled);
+        assert_eq!(view.stm_condenser_model_id.as_deref(), Some("cond-1"));
     }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/common/header_utils.rs` around lines 496 - 511, In
test_memory_header_view_defaults_when_ltm_disabled, add assertions after
creating view = MemoryHeaderView::from_http_headers(&headers) to lock the
STM/LTM independence: assert that the short-term memory flag on MemoryHeaderView
is enabled (e.g., view.short_term_enabled or view.short_term_memory_enabled is
true) and that the condenser model id is preserved (e.g.,
view.condenser_model_id or view.condenser_model equals "cond-1"), so the test
explicitly verifies STM remains active and its condenser model is parsed even
when LTM is disabled.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/memory/context.rs`:
- Around line 89-91: The MemoryExecutionContext currently sets
stm_condenser_model_id unconditionally which can leave a model id present when
stm_enabled is false; update the constructor/initializer that builds
MemoryExecutionContext so stm_condenser_model_id is only assigned when both
headers.stm_condenser_model_id.is_some() and runtime.enabled are true (otherwise
set it to None) to keep it consistent with stm_enabled, and add a unit test
similar to stm_enabled_gated_off_when_runtime_disabled that asserts
ctx.stm_condenser_model_id.is_none() when runtime.enabled is false but the
header provided a condenser id; refer to MemoryExecutionContext, stm_enabled,
stm_condenser_model_id, headers and runtime to locate affected code.

In `@model_gateway/src/routers/common/persistence_utils.rs`:
- Around line 605-793: Tests cover STMO boundaries and config shapes but miss
the stm_enabled=false short-circuit; add a tokio::test that constructs a
MemoryExecutionContext with stm_enabled: false (keep other fields like
stm_condenser_model_id as needed), use a RecordingConversationMemoryWriter and
call enqueue_stmo_if_needed with a turn that would normally enqueue (e.g.,
Some(4)), then assert that writer.rows remains empty; reference the existing
RecordingConversationMemoryWriter, MemoryExecutionContext, and
enqueue_stmo_if_needed to implement the test consistent with the other test
patterns.
- Around line 540-603: The STMO enqueue behavior is undocumented and currently
sets scope_id = None which can produce duplicate jobs on retries; update the
code in enqueue_stmo_if_needed to include an explicit inline comment documenting
the chosen semantic (either "best-effort / duplicates allowed; consumer must be
idempotent" OR "at-most-once / deduplicate by scope_id") and follow the same
pattern used in enqueue_conversation_memory_rows (PR `#1343`): if you decide
duplicates are acceptable, leave scope_id = None and state that clearly next to
NewConversationMemory creation and the create_memory call; if you decide to
enforce uniqueness, compute and set a deterministic scope_id (e.g., based on
conversation_id + user_turns + response_id or target_item_end) and document that
producers/writers must use scope_id to dedupe when implementing Postgres/Oracle
writers; ensure the comment references enqueue_stmo_if_needed,
NewConversationMemory.scope_id, and ConversationMemoryWriter.create_memory so
future writers know how to implement deduplication.

---

Outside diff comments:
In `@model_gateway/src/routers/common/header_utils.rs`:
- Around line 496-511: In test_memory_header_view_defaults_when_ltm_disabled,
add assertions after creating view =
MemoryHeaderView::from_http_headers(&headers) to lock the STM/LTM independence:
assert that the short-term memory flag on MemoryHeaderView is enabled (e.g.,
view.short_term_enabled or view.short_term_memory_enabled is true) and that the
condenser model id is preserved (e.g., view.condenser_model_id or
view.condenser_model equals "cond-1"), so the test explicitly verifies STM
remains active and its condenser model is parsed even when LTM is disabled.

In `@model_gateway/src/routers/grpc/common/responses/utils.rs`:
- Around line 154-171: The current code calls
Arc::new(NoOpConversationMemoryWriter::new()) each time before calling
persist_conversation_items which may allocate on every hot path; change this to
a shared, lazily-initialized Arc so the NoOpConversationMemoryWriter is created
once and reused (e.g., store Arc<dyn ConversationMemoryWriter> in a static
OnceLock or on the shared components struct) and pass that shared Arc into
persist_conversation_items instead of creating a new one per call.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 41805c2b-8b75-420e-8174-7da0b7b92356

📥 Commits

Reviewing files that changed from the base of the PR and between c9f37a0 and 1c910fe.

📒 Files selected for processing (9)
  • model_gateway/src/memory/context.rs
  • model_gateway/src/routers/common/header_utils.rs
  • model_gateway/src/routers/common/persistence_utils.rs
  • model_gateway/src/routers/grpc/common/responses/utils.rs
  • model_gateway/src/routers/openai/context.rs
  • model_gateway/src/routers/openai/responses/history.rs
  • model_gateway/src/routers/openai/responses/non_streaming.rs
  • model_gateway/src/routers/openai/responses/route.rs
  • model_gateway/src/routers/openai/responses/streaming.rs

Comment on lines +89 to 91
stm_enabled: headers.stm_enabled && runtime.enabled,
stm_condenser_model_id: headers.stm_condenser_model_id.clone(),
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Minor inconsistency: stm_condenser_model_id is not gated by runtime.enabled.

When headers.stm_enabled=true but runtime.enabled=false, the resulting MemoryExecutionContext has stm_enabled=false yet stm_condenser_model_id=Some(...). This is fine as long as every downstream consumer gates on stm_enabled (not on stm_condenser_model_id.is_some()). To remove the footgun entirely and keep the two STM fields consistent, consider clearing the model id when the gate trips:

♻️ Proposed diff
-            stm_enabled: headers.stm_enabled && runtime.enabled,
-            stm_condenser_model_id: headers.stm_condenser_model_id.clone(),
+            stm_enabled: headers.stm_enabled && runtime.enabled,
+            stm_condenser_model_id: if headers.stm_enabled && runtime.enabled {
+                headers.stm_condenser_model_id.clone()
+            } else {
+                None
+            },

Adding a test asserting ctx.stm_condenser_model_id.is_none() when runtime is disabled but the header set a condenser would lock this in alongside the existing stm_enabled_gated_off_when_runtime_disabled test.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
stm_enabled: headers.stm_enabled && runtime.enabled,
stm_condenser_model_id: headers.stm_condenser_model_id.clone(),
}
stm_enabled: headers.stm_enabled && runtime.enabled,
stm_condenser_model_id: if headers.stm_enabled && runtime.enabled {
headers.stm_condenser_model_id.clone()
} else {
None
},
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/memory/context.rs` around lines 89 - 91, The
MemoryExecutionContext currently sets stm_condenser_model_id unconditionally
which can leave a model id present when stm_enabled is false; update the
constructor/initializer that builds MemoryExecutionContext so
stm_condenser_model_id is only assigned when both
headers.stm_condenser_model_id.is_some() and runtime.enabled are true (otherwise
set it to None) to keep it consistent with stm_enabled, and add a unit test
similar to stm_enabled_gated_off_when_runtime_disabled that asserts
ctx.stm_condenser_model_id.is_none() when runtime.enabled is false but the
header provided a condenser id; refer to MemoryExecutionContext, stm_enabled,
stm_condenser_model_id, headers and runtime to locate affected code.

Comment on lines +540 to +603
async fn enqueue_stmo_if_needed(
conversation_memory_writer: &Arc<dyn ConversationMemoryWriter>,
memory_execution_context: &MemoryExecutionContext,
conversation_user_turn_count: Option<usize>,
conversation_id: &ConversationId,
response_id: &ResponseId,
output_items: &[Value],
input_item_count: usize,
) {
if !memory_execution_context.stm_enabled {
return;
}

let Some(user_turns) = conversation_user_turn_count else {
return;
};

if !should_enqueue_stmo_for_current_turn(user_turns) {
return;
}

let target_item_end = input_item_count + output_items.len();
// STMO worker config semantics:
// - `last_index`: latest observed user-turn count at enqueue time.
// - `target_item_end`: exclusive end index for items included in this run.
let mut job_config = Map::new();
if let Some(condenser_model) = memory_execution_context.stm_condenser_model_id.as_deref() {
job_config.insert(
STMO_CFG_KEY_CONDENSER_MODEL.to_string(),
Value::String(condenser_model.to_string()),
);
}
job_config.insert(STMO_CFG_KEY_LAST_INDEX.to_string(), json!(user_turns));
job_config.insert(
STMO_CFG_KEY_TARGET_ITEM_END.to_string(),
json!(target_item_end),
);
let job_config = Value::Object(job_config).to_string();

let row = NewConversationMemory {
conversation_id: conversation_id.clone(),
conversation_version: None,
response_id: Some(response_id.clone()),
memory_type: ConversationMemoryType::Stmo,
status: ConversationMemoryStatus::Ready,
attempt: 0,
owner_id: None,
next_run_at: Utc::now(),
lease_until: None,
content: None,
memory_config: Some(job_config),
scope_id: None,
error_msg: None,
};

if let Err(err) = conversation_memory_writer.create_memory(row).await {
warn!(
conversation_id = %conversation_id.0,
response_id = %response_id.0,
error = %err,
"Failed to enqueue STMO job (best-effort; request flow continues)"
);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Does the writer (or its Oracle impl) enforce uniqueness on scope_id / some key?
rg -nP -C8 'create_memory\s*\(|scope_id'
ast-grep --pattern $'impl ConversationMemoryWriter for $_ {
  $$$
}'

Repository: lightseekorg/smg

Length of output: 14198


🏁 Script executed:

# Check Postgres and Oracle create_memory implementations
fd -e rs '\.(postgres|oracle)\.rs' | xargs rg -l 'create_memory|conversation_memory'

Repository: lightseekorg/smg

Length of output: 903


🏁 Script executed:

# Look for STMO worker implementation to understand duplicate handling
fd -e rs | xargs rg -l 'stmo|STMO' | head -20

Repository: lightseekorg/smg

Length of output: 147


🏁 Script executed:

# Check for uniqueness constraints in schema or implementation
rg -n 'UNIQUE|unique|PRIMARY' crates/data_connector/src/schema.rs | grep -i 'conversation_memory\|scope'

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Find Postgres and Oracle data connector implementations
fd -e rs 'postgres|oracle' crates/data_connector/src/

Repository: lightseekorg/smg

Length of output: 229


🏁 Script executed:

# Search for STMO worker implementation
rg -n 'stmo_worker|StmoWorker|STMO.*worker' --type rs

Repository: lightseekorg/smg

Length of output: 86


🏁 Script executed:

# Check if scope_id is used for uniqueness enforcement anywhere
rg -n 'scope_id' crates/data_connector/src/postgres.rs crates/data_connector/src/oracle.rs 2>/dev/null | head -30

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Look at schema definition for conversation_memory table
rg -A 20 'conversation_memory.*=>' crates/data_connector/src/schema.rs

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Search for create_memory in postgres and oracle files
rg 'create_memory' crates/data_connector/src/postgres.rs crates/data_connector/src/oracle.rs

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Look for STMO worker
rg -n 'struct.*Stmo|impl.*Stmo' --max-count 20

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Check conversation_memory table schema in postgres migrations
rg -A 30 'conversation_memory' crates/data_connector/src/postgres_migrations.rs | head -60

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Check for idempotency handling in STMO
rg -n 'idempotent|dedupe|duplicate' crates/workers/ 2>/dev/null || rg -n 'idempotent|dedupe|duplicate' --max-count 10

Repository: lightseekorg/smg

Length of output: 16525


🏁 Script executed:

# Get postgres create_memory implementation
rg -A 30 'impl ConversationMemoryWriter' crates/data_connector/src/postgres.rs

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Get oracle create_memory implementation
rg -A 30 'impl ConversationMemoryWriter' crates/data_connector/src/oracle.rs

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Check if there's a unique constraint on scope_id in migrations
rg -n 'scope_id|UNIQUE|unique' crates/data_connector/src/postgres_migrations.rs | head -40

Repository: lightseekorg/smg

Length of output: 173


🏁 Script executed:

# Search for STMO worker or condensing logic
find . -name '*.rs' | xargs rg -l 'condense|Condense' | grep -v test | head -10

Repository: lightseekorg/smg

Length of output: 204


🏁 Script executed:

# Search for ConversationMemoryWriter implementations across the codebase
rg 'impl.*ConversationMemoryWriter' --type rs

Repository: lightseekorg/smg

Length of output: 86


🏁 Script executed:

# Check if there's any INSERT logic for conversation_memory in postgres/oracle
rg -n 'INSERT.*conversation_memory|conversation_memory.*INSERT' crates/data_connector/src/

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Look at the data_connector factory to see which writers are used
rg -A 20 'fn.*conversation_memory_writer|create_conversation_memory' crates/data_connector/src/

Repository: lightseekorg/smg

Length of output: 1759


🏁 Script executed:

# Check for any dedup or uniqueness logic in the conversation memory handling
rg -B 5 -A 10 'conversation_memory' crates/data_connector/src/postgres.rs | head -60

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Try without -type flag
rg 'impl.*ConversationMemoryWriter'

Repository: lightseekorg/smg

Length of output: 632


🏁 Script executed:

# List all files in data_connector src
ls -la crates/data_connector/src/

Repository: lightseekorg/smg

Length of output: 1281


🏁 Script executed:

# Check if Postgres/Oracle writers exist in the factory or elsewhere
rg 'PostgresConversationMemoryWriter|OracleConversationMemoryWriter'

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Look at how conversation_memory rows are actually inserted
rg -n 'NewConversationMemory' crates/data_connector/src/postgres.rs crates/data_connector/src/oracle.rs 2>/dev/null

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Search for conversation_memory table handling in postgres.rs
head -100 crates/data_connector/src/postgres.rs

Repository: lightseekorg/smg

Length of output: 3681


🏁 Script executed:

# Look for the actual conversation_memory table structure
rg -B 5 -A 15 'conversation_memory' crates/data_connector/src/postgres.rs | head -100

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Check factory to see which writer is actually instantiated
cat crates/data_connector/src/factory.rs | head -100

Repository: lightseekorg/smg

Length of output: 4229


🏁 Script executed:

# Look for STMO-related code and understand the worker flow
rg -n 'STMO\|stmo' model_gateway/src/routers/common/persistence_utils.rs

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Get the complete factory to see the Oracle/Postgres backend handling
grep -A 50 'HistoryBackend::Postgres =>' crates/data_connector/src/factory.rs | head -60

Repository: lightseekorg/smg

Length of output: 1833


🏁 Script executed:

# Look at how enqueue_conversation_memory_rows is called and documented
rg -B 5 -A 15 'enqueue_conversation_memory_rows' model_gateway/src/routers/common/persistence_utils.rs | head -80

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Check where STMO enqueue is called from
rg -B 10 'enqueue_stmo_if_needed' model_gateway/src/routers/common/persistence_utils.rs

Repository: lightseekorg/smg

Length of output: 3068


🏁 Script executed:

# Look at the complete context of both enqueuing functions
cat model_gateway/src/routers/common/persistence_utils.rs | head -300 | tail -200

Repository: lightseekorg/smg

Length of output: 8351


🏁 Script executed:

# Look at enqueue_conversation_memory_rows to understand the documented pattern
rg -B 10 -A 40 'fn enqueue_conversation_memory_rows' model_gateway/src/routers/common/persistence_utils.rs

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Check if there's any documentation or inline comment about LTM/OnDemand semantics
rg -B 5 -A 20 'LTM.*OnDemand|best-effort' model_gateway/src/routers/common/persistence_utils.rs

Repository: lightseekorg/smg

Length of output: 1660


🏁 Script executed:

# Check if Postgres/Oracle backends create a real memory writer
rg -A 5 'create_postgres_storage\|create_oracle_storage' crates/data_connector/src/factory.rs

Repository: lightseekorg/smg

Length of output: 42


🏁 Script executed:

# Check what create_postgres_storage returns
rg -A 30 'fn create_postgres_storage' crates/data_connector/src/factory.rs

Repository: lightseekorg/smg

Length of output: 1493


🏁 Script executed:

# Check what create_oracle_storage returns  
rg -A 30 'fn create_oracle_storage' crates/data_connector/src/factory.rs

Repository: lightseekorg/smg

Length of output: 1499


🏁 Script executed:

# Search for all ConversationMemoryWriter implementations to confirm which ones exist
rg 'ConversationMemoryWriter' crates/data_connector/src/ | grep -v test | grep -v '//'

Repository: lightseekorg/smg

Length of output: 2070


Document STMO enqueue as best-effort, matching LTM/OnDemand semantics.

enqueue_stmo_if_needed currently sets scope_id = None and logs failures as best-effort. When a persistent ConversationMemoryWriter is implemented for Postgres/Oracle (currently only NoOp backends exist), duplicate STMO jobs can be enqueued on client/router retry at the same user_turns boundary, causing the worker to condense twice. Following the pattern established in PR #1343 for enqueue_conversation_memory_rows, clarify whether STMO enqueue should accept at-most-once semantics (duplicates allowed, idempotent consumer) or require dedup via scope_id. Document the chosen semantics inline so the future Postgres/Oracle writer implementations know whether to enforce uniqueness.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/common/persistence_utils.rs` around lines 540 -
603, The STMO enqueue behavior is undocumented and currently sets scope_id =
None which can produce duplicate jobs on retries; update the code in
enqueue_stmo_if_needed to include an explicit inline comment documenting the
chosen semantic (either "best-effort / duplicates allowed; consumer must be
idempotent" OR "at-most-once / deduplicate by scope_id") and follow the same
pattern used in enqueue_conversation_memory_rows (PR `#1343`): if you decide
duplicates are acceptable, leave scope_id = None and state that clearly next to
NewConversationMemory creation and the create_memory call; if you decide to
enforce uniqueness, compute and set a deterministic scope_id (e.g., based on
conversation_id + user_turns + response_id or target_item_end) and document that
producers/writers must use scope_id to dedupe when implementing Postgres/Oracle
writers; ensure the comment references enqueue_stmo_if_needed,
NewConversationMemory.scope_id, and ConversationMemoryWriter.create_memory so
future writers know how to implement deduplication.

Comment on lines +605 to +793
#[cfg(test)]
mod tests {
use std::sync::{Arc, Mutex};

use serde_json::Value;
use smg_data_connector::{ConversationMemoryId, ConversationMemoryResult};

use super::*;

struct RecordingConversationMemoryWriter {
rows: Mutex<Vec<NewConversationMemory>>,
}

impl RecordingConversationMemoryWriter {
fn new() -> Self {
Self {
rows: Mutex::new(Vec::new()),
}
}
}

#[async_trait::async_trait]
impl ConversationMemoryWriter for RecordingConversationMemoryWriter {
async fn create_memory(
&self,
input: NewConversationMemory,
) -> ConversationMemoryResult<ConversationMemoryId> {
self.rows.lock().expect("rows mutex poisoned").push(input);
Ok(ConversationMemoryId::from("mem_test"))
}
}

#[test]
fn stmo_turn_boundary_matches_expected_sequence() {
let cases = [
(1, false),
(2, false),
(3, false),
(4, true),
(5, false),
(6, false),
(7, true),
(8, false),
(9, false),
(10, true),
(11, false),
(12, false),
(13, true),
];

for (turn, expected) in cases {
assert_eq!(
should_enqueue_stmo_for_current_turn(turn),
expected,
"turn={turn}"
);
}
}

#[tokio::test]
async fn enqueue_stmo_if_needed_enqueues_expected_row_on_boundary() {
let writer = Arc::new(RecordingConversationMemoryWriter::new());
let writer_dyn: Arc<dyn ConversationMemoryWriter> = writer.clone();

let memory_execution_context = MemoryExecutionContext {
stm_enabled: true,
stm_condenser_model_id: Some("condense-1".to_string()),
..MemoryExecutionContext::default()
};

let output_items = vec![json!({ "type": "message", "role": "assistant" })];
let conversation_id = ConversationId::from("conv_test");
let response_id = ResponseId::from("resp_test");

enqueue_stmo_if_needed(
&writer_dyn,
&memory_execution_context,
Some(4),
&conversation_id,
&response_id,
&output_items,
4,
)
.await;

let rows = writer.rows.lock().expect("rows mutex poisoned");
assert_eq!(rows.len(), 1, "should enqueue exactly one STMO row");

let row = &rows[0];
assert_eq!(row.conversation_id, conversation_id);
assert_eq!(row.response_id, Some(response_id));
assert_eq!(row.memory_type, ConversationMemoryType::Stmo);
assert_eq!(row.status, ConversationMemoryStatus::Ready);

let config = row
.memory_config
.as_deref()
.expect("memory_config should be set");
let config_json: Value =
serde_json::from_str(config).expect("memory_config must be valid JSON");

assert_eq!(
config_json
.get(STMO_CFG_KEY_CONDENSER_MODEL)
.and_then(Value::as_str),
Some("condense-1")
);
assert_eq!(
config_json
.get(STMO_CFG_KEY_LAST_INDEX)
.and_then(Value::as_u64),
Some(4)
);
assert_eq!(
config_json
.get(STMO_CFG_KEY_TARGET_ITEM_END)
.and_then(Value::as_u64),
Some(5)
);
}

#[tokio::test]
async fn enqueue_stmo_if_needed_skips_when_turn_count_missing() {
let writer = Arc::new(RecordingConversationMemoryWriter::new());
let writer_dyn: Arc<dyn ConversationMemoryWriter> = writer.clone();

let memory_execution_context = MemoryExecutionContext {
stm_enabled: true,
stm_condenser_model_id: Some("condense-1".to_string()),
..MemoryExecutionContext::default()
};

enqueue_stmo_if_needed(
&writer_dyn,
&memory_execution_context,
None,
&ConversationId::from("conv_test"),
&ResponseId::from("resp_test"),
&[],
0,
)
.await;

let rows = writer.rows.lock().expect("rows mutex poisoned");
assert!(
rows.is_empty(),
"should not enqueue when turn count is absent"
);
}

#[tokio::test]
async fn enqueue_stmo_if_needed_enqueues_without_condenser_model() {
let writer = Arc::new(RecordingConversationMemoryWriter::new());
let writer_dyn: Arc<dyn ConversationMemoryWriter> = writer.clone();

let memory_execution_context = MemoryExecutionContext {
stm_enabled: true,
stm_condenser_model_id: None,
..MemoryExecutionContext::default()
};

enqueue_stmo_if_needed(
&writer_dyn,
&memory_execution_context,
Some(4),
&ConversationId::from("conv_test"),
&ResponseId::from("resp_test"),
&[json!({ "type": "message", "role": "assistant" })],
4,
)
.await;

let rows = writer.rows.lock().expect("rows mutex poisoned");
assert_eq!(rows.len(), 1, "should enqueue without condenser model");

let config = rows[0]
.memory_config
.as_deref()
.expect("memory_config should be set");
let config_json: Value =
serde_json::from_str(config).expect("memory_config must be valid JSON");
assert!(config_json.get(STMO_CFG_KEY_CONDENSER_MODEL).is_none());
assert_eq!(
config_json
.get(STMO_CFG_KEY_LAST_INDEX)
.and_then(Value::as_u64),
Some(4)
);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Good test coverage for boundary + config shape.

The boundary table (1..=13) pins the 4/7/10/13 contract; the condenser-present / condenser-absent / missing-turn-count cases cover the enqueue decision tree well. One small gap: no test for the stm_enabled: false short-circuit — worth adding since the gate is the primary runtime switch that determines whether STMO ever fires.

♻️ Suggested additional test
+    #[tokio::test]
+    async fn enqueue_stmo_if_needed_skips_when_stm_disabled() {
+        let writer = Arc::new(RecordingConversationMemoryWriter::new());
+        let writer_dyn: Arc<dyn ConversationMemoryWriter> = writer.clone();
+
+        let memory_execution_context = MemoryExecutionContext {
+            stm_enabled: false,
+            stm_condenser_model_id: Some("condense-1".to_string()),
+            ..MemoryExecutionContext::default()
+        };
+
+        enqueue_stmo_if_needed(
+            &writer_dyn,
+            &memory_execution_context,
+            Some(4),
+            &ConversationId::from("conv_test"),
+            &ResponseId::from("resp_test"),
+            &[json!({ "type": "message", "role": "assistant" })],
+            4,
+        )
+        .await;
+
+        assert!(writer.rows.lock().expect("rows mutex poisoned").is_empty());
+    }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/common/persistence_utils.rs` around lines 605 -
793, Tests cover STMO boundaries and config shapes but miss the
stm_enabled=false short-circuit; add a tokio::test that constructs a
MemoryExecutionContext with stm_enabled: false (keep other fields like
stm_condenser_model_id as needed), use a RecordingConversationMemoryWriter and
call enqueue_stmo_if_needed with a turn that would normally enqueue (e.g.,
Some(4)), then assert that writer.rows remains empty; reference the existing
RecordingConversationMemoryWriter, MemoryExecutionContext, and
enqueue_stmo_if_needed to implement the test consistent with the other test
patterns.

@spalimpaaces-star
spalimpaaces-star marked this pull request as draft April 23, 2026 14:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes openai OpenAI router changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant