Skip to content

feat(worker): harden reasoning streams and sanitize job event payloads - #461

Closed
panosAthDBX wants to merge 5 commits into
nearai:stagingfrom
panosAthDBX:split/worker-orchestrator-streaming-v2
Closed

panosAthDBX wants to merge 5 commits into
nearai:stagingfrom
panosAthDBX:split/worker-orchestrator-streaming-v2

Conversation

@panosAthDBX

Copy link
Copy Markdown
Contributor

Summary

  • harden worker/orchestrator job event streaming by sanitizing and redacting tool payloads before persistence/broadcast
  • add structured worker reasoning event emission (narrative + per-tool decisions), with stable parallel-group numbering and rationale fallback behavior
  • enrich job_result events with success/message metadata and update web event typings/SSE routing accordingly

Why this split

Reviewer feedback on #361 asked for decomposition into focused PRs. This PR isolates the worker/orchestrator streaming and safety hardening layer.

Scope

  • worker runtime: reasoning/job event emission + sanitization behavior (src/worker/runtime.rs)
  • orchestrator event bridge: sanitize/redact and map reasoning/result payloads (src/orchestrator/api.rs)
  • event contracts: web SSE event type/routing updates (src/channels/web/types.rs, src/channels/web/sse.rs)
  • integration glue and tests: src/agent/worker.rs, src/agent/job_monitor.rs, src/main.rs, src/safety/mod.rs

Validation

Ran locally on this branch:

  • cargo fmt --check
  • cargo clippy --all --all-features
  • cargo test

All passed.

@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) scope: channel/web Web gateway channel scope: tool Tool infrastructure scope: safety Prompt injection defense scope: llm LLM integration scope: orchestrator Container orchestrator scope: worker Container worker size: XL 500+ changed lines risk: high Safety, secrets, auth, or critical infrastructure contributor: new First-time contributor labels Mar 2, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly improves the worker and orchestrator's event handling and safety mechanisms. It introduces a comprehensive system for sanitizing and redacting sensitive information within job event payloads, ensuring data privacy and compliance. Additionally, it enhances the observability of worker operations by providing structured reasoning events and more detailed job result metadata. The changes also bolster the stability of LLM integrations by gracefully handling provider-level panics.

Highlights

  • Enhanced Job Event Streaming and Safety: Implemented robust sanitization and redaction for tool payloads (inputs and outputs) before they are persisted or broadcast, significantly hardening the worker/orchestrator job event streaming layer against sensitive data leakage and policy violations.
  • Structured Worker Reasoning Events: Introduced new structured reasoning event emissions, including narrative and per-tool decisions, with stable parallel-group numbering and fallback behavior for rationales. This provides richer insights into the worker's decision-making process.
  • Enriched Job Result Events: Updated job_result events to include success and message metadata, providing more detailed outcomes for job completions. Corresponding web event typings and SSE routing have been updated to support these new fields.
  • Improved LLM Provider Stability: Added panic suppression and error mapping for rig-core based LLM providers, converting potential panics into structured LlmError::RequestFailed errors to prevent process crashes and improve overall system stability.
  • New JSON Redaction Module: A new module src/tools/redaction.rs was added to intelligently identify and redact sensitive keys within JSON payloads, supporting exact matches, camel case, and contextual analysis.
Changelog
  • src/channels/web/types.rs
    • Added ToolDecisionSsePayload struct for detailed tool decision information.
    • Introduced ReasoningUpdate and JobReasoning SSE event types for structured worker reasoning.
    • Enriched JobResult SSE event with success and message fields.
  • src/llm/provider.rs
    • Added reasoning field to the ToolCall struct.
    • Introduced normalize_tool_reasoning function and DEFAULT_TOOL_RATIONALE constant for consistent reasoning handling.
  • src/llm/rig_adapter.rs
    • Integrated panic suppression logic to convert rig-core panics into LlmError::RequestFailed.
  • src/orchestrator/api.rs
    • Implemented sanitize_job_event_data to redact sensitive tool_use inputs and sanitize tool_result outputs.
    • Updated job_event_handler to process and broadcast new reasoning events and enriched result events.
  • src/tools/redaction.rs
    • Added new module for sensitive JSON key redaction, supporting various key naming conventions.
  • src/worker/runtime.rs
    • Modified job execution logic to emit structured reasoning events with parallel group numbering.
    • Applied sensitive JSON redaction to tool_use inputs and safety sanitization to tool_result outputs.
    • Implemented sanitize_worker_narrative and sanitize_worker_rationale for safety policy enforcement on worker-generated text.
Activity
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request significantly hardens the worker and orchestrator by introducing sanitization and redaction for job event payloads, which is a great security improvement. The addition of structured reasoning events with stable parallel-group numbering is also a valuable enhancement for observability. The code is well-structured and the changes are well-tested. I've found one minor issue regarding redundant sanitization which could be optimized.

Comment thread src/orchestrator/api.rs Outdated
Comment on lines +344 to +349
input: crate::tools::redaction::redact_sensitive_json(
payload
.data
.get("input")
.unwrap_or(&serde_json::Value::Null),
),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The payload.data is already sanitized on line 303 by sanitize_job_event_data, which calls redact_sensitive_json for tool_use events. This second call to redact_sensitive_json is redundant and introduces unnecessary overhead from cloning and traversing the JSON value again.

You can simplify this by directly using the already-sanitized value from payload.data.

Suggested change
input: crate::tools::redaction::redact_sensitive_json(
payload
.data
.get("input")
.unwrap_or(&serde_json::Value::Null),
),
input: payload
.data
.get("input")
.cloned()
.unwrap_or(serde_json::Value::Null),

@henrypark133
henrypark133 changed the base branch from main to staging March 10, 2026 02:31

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: PR #461 -- feat(worker): harden reasoning streams and sanitize job event payloads

Reviewed against IronClaw standards. This PR delivers security-critical sanitization of job event payloads flowing through the orchestrator, plus structured reasoning event emission from workers.

Positive observations

  • No .unwrap()/.expect() in production code -- all instances are in test modules only. Clean.
  • Job event sanitization (sanitize_job_event_data) is correctly placed: it runs before both persistence and broadcast, covering both vectors. The tool_use path redacts sensitive JSON keys in input, and the tool_result path runs safety.sanitize_tool_output() on output text. This is the right approach.
  • Worker-side sanitization: sanitize_worker_narrative() and sanitize_worker_rationale() mirror the agent-side patterns from PR #460, ensuring consistency. Both fall back safely (None for narratives, DEFAULT_TOOL_RATIONALE for rationales) when blocked.
  • Worker tool output now runs through safety.sanitize_tool_output() instead of raw truncate() -- significant security improvement for job event streams.
  • Worker tool input now uses redact_sensitive_json() instead of truncate(&tc.arguments.to_string(), 500) -- prevents credential leakage in tool_use events.
  • SseEvent::JobResult now carries success and message fields -- good structured enrichment that avoids ambiguous status strings.
  • Rig adapter panic guard (run_completion_guarded) is a solid defensive addition -- catches known rig-core usage underflow panics, marks the adapter as poisoned, and converts panics to typed LlmErrors. The AssertUnwindSafe usage is documented with a safety comment.
  • Test coverage: job_event_redacts_tool_use_input, job_event_sanitizes_tool_result_output, job_event_handles_reasoning_event, job_event_handles_result_event, worker parallel group monotonicity -- all the security-critical paths are tested.

Minor notes (non-blocking)

  1. The OrchestratorState now requires a safety: Arc<SafetyLayer> field -- the main.rs change correctly wires this. All test constructors also updated. Clean.
  2. The worker completion state machine change (removing mark_completed logic from the Ok(Ok(())) branch in worker.rs) is a behavioral change worth noting -- the worker no longer marks itself completed on success. Ensure the caller/job monitor handles this transition.
  3. The normalize_tool_reasoning("").into_owned() pattern appears frequently in test code -- could be a helper, but not worth changing.

LGTM. Approving.

@zmanian
zmanian enabled auto-merge (squash) March 12, 2026 19:24
@zmanian

zmanian commented Mar 12, 2026

Copy link
Copy Markdown
Collaborator

Closing — the source branch has been deleted and can no longer be merged. If this feature is still needed, please open a fresh PR from a new branch.

@zmanian zmanian closed this Mar 12, 2026
auto-merge was automatically disabled March 12, 2026 22:07

Pull request was closed

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: new First-time contributor risk: high Safety, secrets, auth, or critical infrastructure scope: agent Agent core (agent loop, router, scheduler) scope: channel/web Web gateway channel scope: llm LLM integration scope: orchestrator Container orchestrator scope: safety Prompt injection defense scope: tool Tool infrastructure scope: worker Container worker size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants