Skip to content

GitHub issue #3451: [Reborn] Add direct DB operations for loop checkpoint mappings - #3468

Merged
serrrfirat merged 10 commits into
reborn-integrationfrom
kb/kb-002
May 11, 2026
Merged

serrrfirat merged 10 commits into
reborn-integrationfrom
kb/kb-002

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Refs #3451

KB task: KB-002

GitHub issue #3451: [Reborn] Add direct DB operations for loop checkpoint mappings

URL: #3451
Repo: nearai/ironclaw
Labels: risk: medium, scope: db, reborn
Assignees: none

Scope: IronClaw Reborn issue imported from GitHub.
Before coding: read full issue body/comments, check linked PRs/blockers, work from current origin/reborn-integration unless issue says otherwise, keep PR tightly scoped.
Acceptance: satisfy GitHub issue acceptance criteria; include tests/verification evidence; never merge without explicit user approval.

Auto-opened by kb when task reached In Review. Auto-merge remains disabled; do not merge without explicit operator approval.

@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) scope: channel/cli TUI / CLI channel scope: channel/web Web gateway channel size: XL 500+ changed lines scope: tool Tool infrastructure scope: tool/builtin Built-in tools scope: worker Container worker scope: docs Documentation scope: dependencies Dependency updates risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels May 11, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a significant architectural expansion to the IronClaw Reborn system, adding several new crates for conversation management, event projections, runtime policy, and durable event storage. It implements LibSql and PostgreSQL backends for capability leases, run states, and conversation bindings, alongside a new runtime planner and production wiring validation. Feedback highlights several database-related improvements, including the need to replace inefficient wipe-and-reload persistence patterns and full-state memory loading with targeted read/write paths to prevent memory exhaustion. Additionally, the reviewer recommends decomposing the packed owner_key JSON column into individual columns for better queryability and optimizing lease lookups by filtering at the database level. Finally, a documentation discrepancy was noted in the ironclaw_conversations guardrails regarding the inclusion of durable adapters.

Comment on lines +342 to +515
async fn load_state_from_conn(
conn: &::libsql::Connection,
) -> Result<PersistedConversationState, InboundTurnError> {
let revision = load_revision(conn).await?;
let mut state = InMemoryState::default();

let mut rows = conn
.query(
"SELECT key_payload, user_id FROM reborn_conversation_actor_pairings",
(),
)
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
let key: ActorKey = from_json(&row.get::<String>(0).map_err(db_error)?)?;
let user_id = ironclaw_host_api::UserId::new(row.get::<String>(1).map_err(db_error)?)
.map_err(|error| InboundTurnError::DurableState {
reason: error.to_string(),
})?;
state.pairings.insert(key, user_id);
}

let mut rows = conn
.query(
"SELECT key_payload, payload FROM reborn_conversation_bindings",
(),
)
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
let key: BindingKey = from_json(&row.get::<String>(0).map_err(db_error)?)?;
let binding: BindingRecord = from_json(&row.get::<String>(1).map_err(db_error)?)?;
state.source_bindings.insert(
binding.source_binding_ref.as_str().to_string(),
binding.clone(),
);
state.bindings.insert(key, binding);
}

let mut rows = conn
.query("SELECT payload FROM reborn_conversation_reply_targets", ())
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
let reply_target: ReplyTargetRecord = from_json(&row.get::<String>(0).map_err(db_error)?)?;
state.reply_targets.insert(
reply_target.reply_target_binding_ref.as_str().to_string(),
reply_target,
);
}

let mut rows = conn
.query("SELECT payload FROM reborn_conversation_threads", ())
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
let thread_key_and_record: (ThreadKey, ThreadRecord) =
from_json(&row.get::<String>(0).map_err(db_error)?)?;
state
.threads
.insert(thread_key_and_record.0, thread_key_and_record.1);
}

let mut rows = conn
.query(
"SELECT tenant_id, thread_id, user_id FROM reborn_conversation_thread_participants",
(),
)
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
let tenant_id = ironclaw_host_api::TenantId::new(row.get::<String>(0).map_err(db_error)?)
.map_err(|error| InboundTurnError::DurableState {
reason: error.to_string(),
})?;
let thread_id = ironclaw_host_api::ThreadId::new(row.get::<String>(1).map_err(db_error)?)
.map_err(|error| InboundTurnError::DurableState {
reason: error.to_string(),
})?;
let user_id = ironclaw_host_api::UserId::new(row.get::<String>(2).map_err(db_error)?)
.map_err(|error| InboundTurnError::DurableState {
reason: error.to_string(),
})?;
if let Some(thread) = state
.threads
.get_mut(&ThreadKey::new(&tenant_id, &thread_id))
{
thread.participants.insert(user_id);
}
}

let mut rows = conn
.query(
"SELECT key_payload, identity_payload FROM reborn_conversation_external_event_routes",
(),
)
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
state.external_event_routes.insert(
from_json(&row.get::<String>(0).map_err(db_error)?)?,
from_json(&row.get::<String>(1).map_err(db_error)?)?,
);
}

let mut rows = conn
.query(
"SELECT payload FROM reborn_conversation_accepted_messages",
(),
)
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
let message: ThreadMessageRecord = from_json(&row.get::<String>(0).map_err(db_error)?)?;
let idempotency_key = MessageIdempotencyKey {
tenant_id: message.accepted.tenant_id.clone(),
source_binding_ref: message.accepted.source_binding_ref.as_str().to_string(),
external_event_id: message.external_event_id.clone(),
};
state
.message_idempotency
.insert(idempotency_key, message.accepted.clone());
state.messages.push(message);
}

let mut rows = conn
.query(
"SELECT key_payload, payload FROM reborn_conversation_message_replays",
(),
)
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
state.message_replays.insert(
from_json(&row.get::<String>(0).map_err(db_error)?)?,
from_json(&row.get::<String>(1).map_err(db_error)?)?,
);
}

let mut rows = conn
.query(
"SELECT message_ref, idempotency_key FROM reborn_conversation_submission_keys",
(),
)
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
let message_ref =
ironclaw_turns::AcceptedMessageRef::new(row.get::<String>(0).map_err(db_error)?)
.map_err(|reason| InboundTurnError::InvalidCanonicalRef { reason })?;
let key = ironclaw_turns::IdempotencyKey::new(row.get::<String>(1).map_err(db_error)?)
.map_err(|reason| InboundTurnError::InvalidCanonicalRef { reason })?;
state.submission_keys.insert(message_ref, key);
}

let mut rows = conn
.query(
"SELECT message_ref, payload FROM reborn_conversation_submit_responses",
(),
)
.await
.map_err(db_error)?;
while let Some(row) = rows.next().await.map_err(db_error)? {
let message_ref =
ironclaw_turns::AcceptedMessageRef::new(row.get::<String>(0).map_err(db_error)?)
.map_err(|reason| InboundTurnError::InvalidCanonicalRef { reason })?;
state.submitted_message_responses.insert(
message_ref,
from_json(&row.get::<String>(1).map_err(db_error)?)?,
);
}

Ok(PersistedConversationState { state, revision })
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The load_state function fetches the entire database content into memory. While this aligns with the current snapshot-adapter design pattern, it poses a risk of memory exhaustion as data grows. Per project guidelines, we should defer this optimization for now, but you must document the requirement for targeted read paths as a follow-up task.

References
  1. If a design pattern like snapshot-adapter necessitates loading full state, defer optimization and document the requirement for targeted read paths as a follow-up task.

Comment on lines +536 to +539
conn.execute(&format!("DELETE FROM {table}"), ())
.await
.map_err(db_error)?;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The save_state implementation wipes all tables and re-inserts the entire state. While this is part of the snapshot-adapter pattern, it is inefficient. As per our rules, we should defer the optimization to granular operations for now, but please document the need for a more efficient persistence strategy as a follow-up.

References
  1. If a design pattern like snapshot-adapter necessitates an inefficient implementation, defer optimization and document the requirement for targeted paths as a follow-up task.

Comment on lines +338 to +511
async fn load_state_from_txn(
txn: &deadpool_postgres::Transaction<'_>,
) -> Result<PersistedConversationState, InboundTurnError> {
let revision = load_revision(txn).await?;
let mut state = InMemoryState::default();

for row in txn
.query(
"SELECT key_payload, user_id FROM reborn_conversation_actor_pairings",
&[],
)
.await
.map_err(pg_error)?
{
let key: ActorKey = from_json(row.get::<_, &str>(0))?;
let user_id = ironclaw_host_api::UserId::new(row.get::<_, String>(1)).map_err(|error| {
InboundTurnError::DurableState {
reason: error.to_string(),
}
})?;
state.pairings.insert(key, user_id);
}

for row in txn
.query(
"SELECT key_payload, payload FROM reborn_conversation_bindings",
&[],
)
.await
.map_err(pg_error)?
{
let key: BindingKey = from_json(row.get::<_, &str>(0))?;
let binding: BindingRecord = from_json(row.get::<_, &str>(1))?;
state.source_bindings.insert(
binding.source_binding_ref.as_str().to_string(),
binding.clone(),
);
state.bindings.insert(key, binding);
}

for row in txn
.query("SELECT payload FROM reborn_conversation_reply_targets", &[])
.await
.map_err(pg_error)?
{
let reply_target: ReplyTargetRecord = from_json(row.get::<_, &str>(0))?;
state.reply_targets.insert(
reply_target.reply_target_binding_ref.as_str().to_string(),
reply_target,
);
}

for row in txn
.query("SELECT payload FROM reborn_conversation_threads", &[])
.await
.map_err(pg_error)?
{
let (key, record): (ThreadKey, ThreadRecord) = from_json(row.get::<_, &str>(0))?;
state.threads.insert(key, record);
}

for row in txn
.query(
"SELECT tenant_id, thread_id, user_id FROM reborn_conversation_thread_participants",
&[],
)
.await
.map_err(pg_error)?
{
let tenant_id =
ironclaw_host_api::TenantId::new(row.get::<_, String>(0)).map_err(|error| {
InboundTurnError::DurableState {
reason: error.to_string(),
}
})?;
let thread_id =
ironclaw_host_api::ThreadId::new(row.get::<_, String>(1)).map_err(|error| {
InboundTurnError::DurableState {
reason: error.to_string(),
}
})?;
let user_id = ironclaw_host_api::UserId::new(row.get::<_, String>(2)).map_err(|error| {
InboundTurnError::DurableState {
reason: error.to_string(),
}
})?;
if let Some(thread) = state
.threads
.get_mut(&ThreadKey::new(&tenant_id, &thread_id))
{
thread.participants.insert(user_id);
}
}

for row in txn
.query(
"SELECT key_payload, identity_payload FROM reborn_conversation_external_event_routes",
&[],
)
.await
.map_err(pg_error)?
{
state.external_event_routes.insert(
from_json(row.get::<_, &str>(0))?,
from_json(row.get::<_, &str>(1))?,
);
}

for row in txn
.query(
"SELECT payload FROM reborn_conversation_accepted_messages",
&[],
)
.await
.map_err(pg_error)?
{
let message: ThreadMessageRecord = from_json(row.get::<_, &str>(0))?;
let idempotency_key = MessageIdempotencyKey {
tenant_id: message.accepted.tenant_id.clone(),
source_binding_ref: message.accepted.source_binding_ref.as_str().to_string(),
external_event_id: message.external_event_id.clone(),
};
state
.message_idempotency
.insert(idempotency_key, message.accepted.clone());
state.messages.push(message);
}

for row in txn
.query(
"SELECT key_payload, payload FROM reborn_conversation_message_replays",
&[],
)
.await
.map_err(pg_error)?
{
state.message_replays.insert(
from_json(row.get::<_, &str>(0))?,
from_json(row.get::<_, &str>(1))?,
);
}

for row in txn
.query(
"SELECT message_ref, idempotency_key FROM reborn_conversation_submission_keys",
&[],
)
.await
.map_err(pg_error)?
{
let message_ref = ironclaw_turns::AcceptedMessageRef::new(row.get::<_, String>(0))
.map_err(|reason| InboundTurnError::InvalidCanonicalRef { reason })?;
let key = ironclaw_turns::IdempotencyKey::new(row.get::<_, String>(1))
.map_err(|reason| InboundTurnError::InvalidCanonicalRef { reason })?;
state.submission_keys.insert(message_ref, key);
}

for row in txn
.query(
"SELECT message_ref, payload FROM reborn_conversation_submit_responses",
&[],
)
.await
.map_err(pg_error)?
{
let message_ref = ironclaw_turns::AcceptedMessageRef::new(row.get::<_, String>(0))
.map_err(|reason| InboundTurnError::InvalidCanonicalRef { reason })?;
state
.submitted_message_responses
.insert(message_ref, from_json(row.get::<_, &str>(1))?);
}

Ok(PersistedConversationState { state, revision })
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The load_state implementation for Postgres fetches the entire state into memory. Following the snapshot-adapter pattern, we will defer optimization for now, but please ensure the requirement for targeted read paths is documented as a follow-up task to mitigate OOM risks.

References
  1. If a design pattern like snapshot-adapter necessitates loading full state, defer optimization and document the requirement for targeted read paths as a follow-up task.

Comment on lines +532 to +535
txn.execute(&format!("DELETE FROM {table}"), &[])
.await
.map_err(pg_error)?;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The save_state implementation for Postgres uses the destructive wipe-and-reload pattern. Since this follows the snapshot-adapter design, we will defer refactoring to granular updates, but this requirement must be documented as a follow-up task.

References
  1. If a design pattern like snapshot-adapter necessitates an inefficient implementation, defer optimization and document the requirement for targeted paths as a follow-up task.

Comment on lines +15 to +22
CREATE TABLE IF NOT EXISTS reborn_capability_lease_records (
owner_key TEXT NOT NULL,
invocation_id TEXT NOT NULL,
lease_id TEXT NOT NULL,
status TEXT NOT NULL,
payload TEXT NOT NULL,
PRIMARY KEY (owner_key, invocation_id, lease_id)
);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Packing multiple identifiers into a single JSON string column (owner_key) and using it as a primary key component is a poor database design. It prevents efficient indexing and querying of individual fields (like user_id or project_id) for analytics or cross-invocation management. Consider using individual columns for each identifier in the ResourceScope to allow for better query performance and data integrity.

References
  1. To prevent key collisions and improve queryability, use structured keys or individual columns instead of string concatenation or packed strings.

conn: &libsql::Connection,
scope: &ResourceScope,
) -> Result<Vec<CapabilityLease>, CapabilityLeaseError> {
let mut rows = conn.query("SELECT invocation_id, lease_id, status, payload FROM reborn_capability_lease_records WHERE owner_key = ?1 ORDER BY lease_id", libsql::params![owner_key(scope)?]).await.map_err(db_error)?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The query in libsql_leases_for_scope filters only by owner_key, but the ResourceScope passed in contains a specific invocation_id. Since the table has an invocation_id column and the primary consumer (active_leases_for_context) filters by it in memory, it would be significantly more efficient to perform this filtering at the database level to reduce data transfer and processing overhead.

    let mut rows = conn.query("SELECT invocation_id, lease_id, status, payload FROM reborn_capability_lease_records WHERE owner_key = ?1 AND invocation_id = ?2 ORDER BY lease_id", libsql::params![owner_key(scope)?, scope.invocation_id.to_string()]).await.map_err(db_error)?;
References
  1. Use targeted database queries to fetch specific records instead of loading all records and filtering in the application to prevent performance bottlenecks.

client: &impl deadpool_postgres::GenericClient,
scope: &ResourceScope,
) -> Result<Vec<CapabilityLease>, CapabilityLeaseError> {
let rows = client.query("SELECT invocation_id, lease_id, status, payload::text FROM reborn_capability_lease_records WHERE owner_key = $1 ORDER BY lease_id", &[&owner_key(scope)?]).await.map_err(db_error)?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Similar to the LibSql implementation, postgres_leases_for_scope should filter by invocation_id in the SQL query to avoid fetching unnecessary rows and filtering them in memory.

    let rows = client.query("SELECT invocation_id, lease_id, status, payload::text FROM reborn_capability_lease_records WHERE owner_key = $1 AND invocation_id = $2 ORDER BY lease_id", &[&owner_key(scope)?, &scope.invocation_id.to_string()]).await.map_err(db_error)?;
References
  1. Use targeted database queries to fetch specific records instead of loading all records and filtering in the application to prevent performance bottlenecks.

- Preserve typed `ironclaw_turns::TurnError` values across this boundary instead of flattening turn failures to strings.
- Automatic first-contact binding must not trust raw adapter-supplied agent/project scope hints; use explicit linking or a future trusted thread-creation seam for scoped thread selection.
- Explicit links are idempotent only for the same target thread; never silently retarget an already-bound external conversation to another thread.
- Keep durable PostgreSQL/libSQL adapters out of this crate until the transcript/thread storage boundary has a scoped implementation plan with parity tests.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The addition of durable PostgreSQL/libSQL adapters in this PR directly contradicts the guardrail stating they should be kept out of this crate until a scoped implementation plan exists. If this PR constitutes that plan, the guardrail should be updated to reflect the new status.

References
  1. Documentation for complex logic and architectural state must precisely match the code implementation and distinguish between current and future states.

@github-actions github-actions Bot added size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules and removed size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules labels May 11, 2026
@zmanian

zmanian commented May 11, 2026

Copy link
Copy Markdown
Collaborator

Review

Summary: Replaces wipe-and-reload snapshot persistence for LoopCheckpointStore with targeted INSERT/SELECT against turn_loop_checkpoints, on both libSQL and Postgres, plus adds scope_key/kind columns to turn_checkpoints with idempotent migrations.

Findings:

  1. .unwrap() in production code (CLAUDE.md violation). crates/ironclaw_turns/src/memory.rs:499-500 and src/store.rs:530 call LoopCheckpointStateRef::new("checkpoint:...").unwrap() as placeholder/serde defaults. The safety comment handwaves it, but the project rule is no .unwrap() in prod — use OnceLock/LazyLock or return Result. These encode "placeholder will be threaded later" — landing risks shipping the placeholder permanently.

  2. Non-transactional INSERT+SELECT race. libsql_insert_loop_checkpoint_record (db.rs:209) uses INSERT OR IGNORE then a separate SELECT outside any transaction — a concurrent UPDATE/DELETE between the two statements can misclassify a true conflict as idempotent or yield "conflicted but row not readable" spuriously. Postgres version (db.rs:273) has the same shape on a pooled client without BEGIN. Previous code held SHARE ROW EXCLUSIVE — that protection is gone. Wrap in a transaction or use INSERT ... ON CONFLICT DO NOTHING RETURNING + a single SELECT FOR UPDATE follow-up.

  3. JSON re-deserialization as truth, then compared to request fields (db.rs:108-114, :190-196). WHERE clause already constrains by checkpoint_id/scope_key/turn_id/run_id. If persisted JSON drifts from indexed columns (partial migration, manual repair), this returns Ok(None) and silently hides a real row. Either trust the WHERE clause or treat mismatch as an error.

  4. libSQL migration string-matches "duplicate column name" (db.rs:45). Fragile across libsql versions/locales. Use PRAGMA table_info(turn_checkpoints) for presence check. Also: new NOT NULL DEFAULT '' columns mean existing rows get empty scope_key — no backfill plan documented.

  5. checkpoints.is_empty() tests (lines 698-700, 729-732) presume loop checkpoints never land in turn_checkpoints, but libsql_replace_snapshot (db.rs:251-262) still writes TurnCheckpointRecord checkpoints there with the new columns. Test naming is correct for loop checkpoints but the assertion is brittle if any block checkpoint is later added to setup.

Minor: test helpers use .unwrap() (fine in #[cfg(test)]). created_at.to_rfc3339() → Postgres timestamptz via cast loses microsecond precision vs. binding chrono::DateTime<Utc> directly.

Verdict: Request changes — (1) and (2) are blocking. Tests and dual-backend coverage otherwise solid.

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

@zmanian addressed in a37f794d:

  1. Removed production placeholder .unwrap() path by threading LoopCheckpointStateRef through BlockRunRequest, LoopBlocked, TurnRunnerOutcome::Blocked, and checkpoint record persistence. Legacy serde default now uses an internal sentinel constructor, not .unwrap().
  2. Replaced insert-then-select loop-checkpoint idempotency with atomic conflict handling: libSQL uses one upsert statement; Postgres wraps the write in a transaction and only returns success for identical payload conflicts.
  3. get_loop_checkpoint no longer treats payload/index drift as Ok(None); it now returns a conflict if the selected row payload disagrees with constrained metadata.
  4. libSQL migration now checks PRAGMA table_info(turn_checkpoints) instead of matching duplicate column name, and turn-persistence.md documents legacy empty scope_key semantics/backfill requirement.
  5. Loop checkpoint tests now assert loop mappings do not land in turn_checkpoints by state ref, instead of assuming turn_checkpoints is globally empty.
  6. Minor timestamp note addressed by binding chrono::DateTime<Utc> directly for Postgres TIMESTAMPTZ writes.

Verification:

  • cargo fmt --check
  • cargo test -p ironclaw_turns --all-features
  • cargo clippy -p ironclaw_turns --all-features --tests --examples --benches -- -D warnings
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings
  • bash scripts/pre-commit-safety.sh
  • GitHub checks are green; PR merge state is clean.

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: approved

All five blocking findings resolved at a37f794d:

  1. .unwrap() in production — RESOLVED. store.rs:113-120 replaces both sites with serde defaults: default_checkpoint_kind() returns LoopCheckpointKind::BeforeBlock and default_checkpoint_state_ref() returns a new infallible LoopCheckpointStateRef::legacy_unknown(). state_ref is now a real plumbed field on BlockRunRequest / TurnRunnerOutcome::Blocked / record_checkpoint rather than a sentinel.
  2. Non-transactional INSERT+SELECT race — RESOLVED. Both backends do a single atomic UPSERT with a guard. libSQL: INSERT ... ON CONFLICT(checkpoint_id) DO UPDATE SET checkpoint_id = ... WHERE payload = excluded.payload returning row count (rows == 1 ⇒ success, 0 ⇒ Conflict). Postgres mirrors with RETURNING payload::text inside client.transaction().
  3. JSON re-deserialization drift — RESOLVED. ensure_loop_checkpoint_record_matches_request compares persisted scope/turn_id/run_id/checkpoint_id against the request and raises TurnError::Conflict on drift — silent miss replaced by loud error.
  4. libSQL migration robustness — RESOLVED. libsql_column_exists via PRAGMA table_info with identifier whitelist replaces the string match; Postgres uses ADD COLUMN IF NOT EXISTS. Legacy NOT NULL DEFAULT '' rows documented in docs/reborn/contracts/turn-persistence.md with explicit future-read-path requirements.
  5. Brittle is_empty() tests — RESOLVED. New contract tests assert exact loop-checkpoint cardinality + absence of state_ref in turn_checkpoints, plus cross-scope/cross-run miss and drift conflict for both backends.

Substantive fixes, not annotations — particularly the state_ref plumbing in #1.

@serrrfirat
serrrfirat merged commit b682584 into reborn-integration May 11, 2026
14 checks passed
@serrrfirat
serrrfirat deleted the kb/kb-002 branch May 11, 2026 22:05
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
GitHub issue nearai#3451: [Reborn] Add direct DB operations for loop checkpoint mappings
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: agent Agent core (agent loop, router, scheduler) scope: channel/cli TUI / CLI channel scope: channel/web Web gateway channel scope: dependencies Dependency updates scope: docs Documentation scope: tool/builtin Built-in tools scope: tool Tool infrastructure scope: worker Container worker size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants