Skip to content

feat(reborn): add process lifecycle substrate - #3017

Merged
serrrfirat merged 7 commits into
reborn-integrationfrom
reborn-land-04a-processes
Apr 29, 2026
Merged

serrrfirat merged 7 commits into
reborn-integrationfrom
reborn-land-04a-processes

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

Carves the process lifecycle/result/output substrate from the Reborn stack.

Adds crates/ironclaw_processes with:

  • process records, starts, statuses, exits, and subscriptions
  • ProcessStore, ProcessResultStore, ProcessManager, ProcessHost, and ProcessExecutor
  • in-memory and filesystem-backed process/result stores
  • BackgroundProcessManager
  • cooperative cancellation registry/tokens
  • eventing process store wrapper for redacted process events
  • resource-managed process store that reserves/reconciles/releases process-owned reservations
  • output refs for large/binary process outputs

Stacking note

This PR is draft/stacked because it is intended to land after the current Reborn substrate/control stack:

After those merge into reborn-integration, rebase this branch so the diff narrows to only:

crates/ironclaw_processes
Cargo.toml / Cargo.lock membership for that crate

Scope boundary

This PR intentionally does not include:

  • dispatcher runtime routing
  • WASM/Script/MCP runtime lanes
  • ironclaw_capabilities / CapabilityHost
  • host runtime composition
  • built-in obligation handling
  • secrets/network substrates
  • production app wiring or user-visible behavior changes

Exposure checklist

Does this affect existing src/ runtime behavior? no
Does this alter app startup/config defaults? no
Does this expose new routes/CLI/user-visible APIs? no
Does this create a second writer for existing production state? no
Are all Reborn paths feature-gated or unused by default? yes, unused internal crate substrate
Are docs/status labels accurate: implemented slice vs product complete? yes, process substrate only

TDD note

Copied the process contract tests first against a RED stub and confirmed cargo test -p ironclaw_processes failed due missing process types before porting the implementation.

Verification

Passed:

cargo test -p ironclaw_processes
cargo clippy -p ironclaw_processes --all-targets -- -D warnings
cargo fmt --check
git diff --check
cargo tree -p ironclaw_processes -e normal | grep -E 'ironclaw_(authorization|approvals|capabilities|dispatcher|extensions|host_runtime|mcp|network|run_state|scripts|secrets|wasm)' || true

Boundary grep returned no forbidden normal Reborn dependencies.

Refs #2987.

@github-actions github-actions Bot added scope: db/postgres PostgreSQL backend scope: docs Documentation scope: dependencies Dependency updates DB MIGRATION PR adds or modifies PostgreSQL or libSQL migration definitions size: XL 500+ changed lines labels Apr 28, 2026
@github-actions github-actions Bot added risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Apr 28, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces several core infrastructure crates for the IronClaw Reborn project, covering approval resolution, authorization, event logging, extension management, scoped filesystem access, and process lifecycle tracking. Feedback identifies critical performance bottlenecks in the database-backed filesystem and authorization service due to inefficient record scanning. Reliability issues were noted regarding the process manager's reliance on volatile in-memory state for durable resource reservations and the use of synchronous I/O in async contexts. Additionally, better error reporting for background task failures was recommended to improve system observability.

Comment thread crates/ironclaw_filesystem/src/lib.rs Outdated
}

async fn all_paths(&self) -> Result<Vec<(VirtualPath, u64, FileType)>, FilesystemError> {
let client = self.client().await?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The all_paths helper (and its counterpart in LibSqlRootFilesystem at line 1397) performs a full table scan of root_filesystem_entries. This is used by list_dir and stat to resolve directory contents and metadata. As the number of files grows, this will lead to severe performance degradation and high memory usage. The database queries should be optimized to fetch only relevant rows using prefix matching (e.g., WHERE path LIKE '/dir/%') to move the logic to the database layer and prevent performance bottlenecks.

References
  1. Use targeted database queries to fetch specific records instead of loading all records and filtering in the application to prevent performance bottlenecks.
  2. When application logic becomes complex or inefficient, consider moving it to the database layer (e.g., a dedicated SQL query) to improve performance.

Comment thread crates/ironclaw_processes/src/lib.rs Outdated
Comment on lines +844 to +849
.ok_or(ProcessError::ResourceReservationNotOwned {
process_id,
reservation_id: record_reservation_id,
})?;
if Some(reservation_id) != record_reservation_id {
self.owned_reservations

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

ResourceManagedProcessStore relies on an in-memory owned_reservations map to track active resource reservations. This map is lost when the application restarts. If the underlying ProcessStore is durable (like FilesystemProcessStore), processes that were running before the restart will still exist in the store with a valid resource_reservation_id, but any attempt to complete, fail, or kill them will result in a ResourceReservationNotOwned error because the entry is missing from the in-memory map. The store should trust the resource_reservation_id present in the persisted ProcessRecord as a fallback.

        let reservation_id = self
            .owned_reservations
            .lock()
            .unwrap()
            .remove(&ProcessKey::new(scope, process_id))
            .or(record_reservation_id)
            .ok_or(ProcessError::ResourceReservationNotOwned {
                process_id,
                reservation_id: record_reservation_id,
            })?;

Comment on lines +483 to +489
async fn active_leases_for_context(&self, context: &ExecutionContext) -> Vec<CapabilityLease> {
self.leases_for_scope(&context.resource_scope)
.await
.into_iter()
.filter(|lease| lease_is_authorizing(lease, context))
.collect()
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

active_leases_for_context is implemented by calling leases_for_scope, which scans all historical invocation directories for a user to find active leases. This operation is performed during every capability dispatch. As the number of invocations grows, this will become a major performance bottleneck. Since lease_is_authorizing only accepts leases matching the current invocation_id, the implementation should directly access the specific invocation's directory instead of loading all records and filtering in the application.

    async fn active_leases_for_context(&self, context: &ExecutionContext) -> Vec<CapabilityLease> {
        let Ok(root) = lease_invocation_root(&context.resource_scope) else {
            return Vec::new();
        };
        let Ok(files) = self.list_lease_files(&root).await else {
            return Vec::new();
        };
        let mut leases = Vec::new();
        for path in files {
            if let Ok(lease) = self.read_lease_file(&path).await {
                if lease_is_authorizing(&lease, context) {
                    leases.push(lease);
                }
            }
        }
        leases
    }
References
  1. Use targeted database queries or direct resource access to fetch specific records instead of loading all records and filtering in the application to prevent performance bottlenecks.


#[async_trait]
impl RootFilesystem for LocalFilesystem {
async fn read_file(&self, path: &VirtualPath) -> Result<Vec<u8>, FilesystemError> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

LocalFilesystem implements the RootFilesystem trait using synchronous std::fs operations (e.g., std::fs::read, std::fs::write, std::fs::read_dir). Since these methods are called within an async context, they will block the Tokio executor threads. This can lead to thread pool starvation and poor performance under high I/O load. These operations should be performed using tokio::fs or wrapped in tokio::task::spawn_blocking.

Comment thread crates/ironclaw_processes/src/lib.rs Outdated
Comment on lines +1199 to +1208
if let Ok(record) = store.complete(&scope, process_id).await
&& let Some(result_store) = &result_store
{
let _ = result_store
.complete(&record.scope, record.process_id, result.output)
.await;
}
}
Err(error) => {
if let Ok(record) = store.fail(&scope, process_id, error.kind).await

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The background task spawned in spawn silently ignores failures when updating the process status in the ProcessStore. If the store update fails, the process will remain in the Running state indefinitely. These failures should be surfaced via tracing::warn! to ensure system observability, and if spawn_blocking is used, the JoinError should be logged to distinguish between panics and cancellations.

References
  1. Errors in background persistence or status updates should be surfaced via tracing::warn! to ensure reliability and observability.
  2. When handling errors from tokio::task::spawn_blocking, log the JoinError to capture debugging information and distinguish between panics and cancellations.

@serrrfirat
serrrfirat force-pushed the reborn-land-04a-processes branch from ce8173b to d471de7 Compare April 28, 2026 20:01
@serrrfirat
serrrfirat marked this pull request as ready for review April 28, 2026 20:05
@serrrfirat serrrfirat added the reborn IronClaw Reborn architecture and landing work label Apr 29, 2026
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Addressed all 9 findings from the paranoid review in fc96d6596. 44 tests pass, clippy clean.

ID Where Fix
M1 services.rs Added BackgroundFailure / BackgroundFailureStage + BackgroundProcessManager::with_error_handler(..). Errors from store.complete/fail and result_store.complete/fail inside the spawned task are now reported via the handler instead of silently dropped.
M2 services.rs Inverted write order in BackgroundProcessManager::spawn: result store is persisted before lifecycle status flips to terminal. Contract is now "if status is terminal, result record is on disk" — await_result no longer needs the single-retry race mitigation as a primary correctness mechanism. Added a store.get pre-check so the inverted ordering does not overwrite a Killed result if host.kill already terminalized the process (caught by an existing test).
M3 services.rs Removed the blanket impl<T> ProcessManager for T where T: ProcessStore + ?Sized so raw stores can no longer silently impersonate a ProcessManager. No callers depend on it.
M4 services.rs Doc note on BackgroundProcessManager::spawn describing the detached-task orphan risk on runtime shutdown + TODO for built-in startup reconciliation of stuck Running records.
L1 wrappers.rs New ReservationDropGuard RAII wrapper around inner.start in ResourceManagedProcessStore::start. If inner.start panics, the just-acquired reservation is released via Drop rather than leaked.
L2 filesystem_store.rs Doc on FilesystemProcessResultStore::complete describing the two-step write (write_output → write_result) and the orphan-blob cleanup expectation if the second write fails.
L3 filesystem_store.rs Doc on FilesystemProcessStore::new and from_arc calling out the single-instance invariant (transition lock only serializes within one instance; share via Arc, do not construct multiple instances against the same root).
L4 tests/process_store_contract.rs Added FailingProcessResultStore + background_process_manager_reports_result_store_complete_failure_and_keeps_running_status — drives a result-store that returns Err and asserts (a) the error handler receives ResultStoreComplete for the right process_id and (b) the lifecycle status stays Running (M2 ordering: status does not promote when result write fails).
N1 cancellation.rs Doc on ProcessCancellationToken::cancelled explaining the notify-then-check ordering: notified() is created before the flag re-check, so any notify_waiters racing the check is captured rather than lost.

Diff: 6 files, +342 / −49.

@serrrfirat serrrfirat added the skip-regression-check Bypass regression test CI gate (tests exist but not in tests/ dir) label Apr 29, 2026
Carves the 1953-line lib.rs into 7 focused modules so each file fits
in one tool-read and an agent can grep-jump straight to the right
concern:

- types.rs            data types, errors, traits, shared helpers
- cancellation.rs     token + registry
- host.rs             ProcessHost + Subscription
- memory_store.rs     InMemory{Store,ResultStore}
- filesystem_store.rs Filesystem* + path/serde helpers
- wrappers.rs         Eventing + ResourceManaged decorators
- services.rs         ProcessServices + BackgroundProcessManager

lib.rs is now a thin module-decl + re-export hub so the public API
surface is visible at a glance.

No behavior change. Verified:
- cargo fmt
- cargo check -p ironclaw_processes
- cargo clippy -p ironclaw_processes --all-targets -- -D warnings
- cargo test -p ironclaw_processes (43/43 pass)
- M1 + L4: BackgroundFailure/Stage + with_error_handler so swallowed
  store/result-store errors in spawned tasks are now observable; covered
  by new FailingProcessResultStore test.
- M2: invert write order in BackgroundProcessManager::spawn (result
  store first, then lifecycle status) so terminal status implies result
  is persisted. Added pre-check on store.get to skip writes when the
  process was already terminalized externally (e.g. by host.kill),
  avoiding overwriting a kill record.
- M3: drop blanket impl<T: ProcessStore> ProcessManager for T so raw
  stores no longer impersonate a manager.
- M4: doc note on spawn re detached-task orphan risk + TODO for
  startup reconciliation.
- L1: ReservationDropGuard RAII guard around inner.start in
  ResourceManagedProcessStore::start; reservations released on panic.
- L2: doc orphan-blob expectation on FilesystemProcessResultStore
  ::complete.
- L3: doc single-instance invariant on FilesystemProcessStore::new
  and from_arc.
- N1: doc the notify-then-check ordering in
  ProcessCancellationToken::cancelled.
[skip-regression-check]

The actual fix commit (fc96d65) added a new integration test in
crates/ironclaw_processes/tests/process_store_contract.rs, but the
regression-check script only matches `^tests/` (repo-root) and misses
crate-level integration tests. Skip explicitly via marker.
@serrrfirat
serrrfirat force-pushed the reborn-land-04a-processes branch from 2506385 to be69726 Compare April 29, 2026 10:47
@serrrfirat
serrrfirat merged commit e1880c3 into reborn-integration Apr 29, 2026
18 checks passed
@serrrfirat
serrrfirat deleted the reborn-land-04a-processes branch April 29, 2026 11:36
serrrfirat added a commit that referenced this pull request May 1, 2026
Refs: #3145, #3080, #3017, #3087.

Decisions: host-runtime now wraps process stores with ProcessObligationLifecycleStore so spawn-phase resource reservations are reconciled on success or released on failure/kill, and staged network/secret handoffs are discarded at terminal lifecycle. Process-start failure remains CapabilityHost abort-owned.

Files changed: ironclaw_host_runtime lib exports, obligations lifecycle/store cleanup, HostRuntimeServices process graph wiring, host_runtime_services_contract tests.

Notes: full host-runtime tests need CARGO_BUILD_JOBS=1 in this environment to avoid linker OOM.
serrrfirat added a commit that referenced this pull request May 1, 2026
* RALPH: complete issue 3145 background obligation lifecycle

Refs: #3145, #3080, #3017, #3087.

Decisions: host-runtime now wraps process stores with ProcessObligationLifecycleStore so spawn-phase resource reservations are reconciled on success or released on failure/kill, and staged network/secret handoffs are discarded at terminal lifecycle. Process-start failure remains CapabilityHost abort-owned.

Files changed: ironclaw_host_runtime lib exports, obligations lifecycle/store cleanup, HostRuntimeServices process graph wiring, host_runtime_services_contract tests.

Notes: full host-runtime tests need CARGO_BUILD_JOBS=1 in this environment to avoid linker OOM.

* Fix process obligation cleanup lifecycle

* Fix stale reservation lifecycle cleanup

* Fix background obligation cleanup failures

* fix(reborn): harden process obligation cleanup

* fix(host-runtime): preserve cancel side effects on cleanup failure

* fix(reborn): enforce single active process handoff

* fix: address review findings (iteration 1)
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
Carves out the Reborn process lifecycle substrate in ironclaw_processes with scoped process/result stores, ProcessHost, BackgroundProcessManager, cooperative cancellation, event/resource decorators, and contract tests.
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…#3161)

* RALPH: complete issue 3145 background obligation lifecycle

Refs: nearai#3145, nearai#3080, nearai#3017, nearai#3087.

Decisions: host-runtime now wraps process stores with ProcessObligationLifecycleStore so spawn-phase resource reservations are reconciled on success or released on failure/kill, and staged network/secret handoffs are discarded at terminal lifecycle. Process-start failure remains CapabilityHost abort-owned.

Files changed: ironclaw_host_runtime lib exports, obligations lifecycle/store cleanup, HostRuntimeServices process graph wiring, host_runtime_services_contract tests.

Notes: full host-runtime tests need CARGO_BUILD_JOBS=1 in this environment to avoid linker OOM.

* Fix process obligation cleanup lifecycle

* Fix stale reservation lifecycle cleanup

* Fix background obligation cleanup failures

* fix(reborn): harden process obligation cleanup

* fix(host-runtime): preserve cancel side effects on cleanup failure

* fix(reborn): enforce single active process handoff

* fix: address review findings (iteration 1)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs DB MIGRATION PR adds or modifies PostgreSQL or libSQL migration definitions reborn IronClaw Reborn architecture and landing work risk: medium Business logic, config, or moderate-risk modules scope: db/postgres PostgreSQL backend scope: dependencies Dependency updates scope: docs Documentation size: XL 500+ changed lines skip-regression-check Bypass regression test CI gate (tests exist but not in tests/ dir)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant