Skip to content

chore: promote staging to main (2026-03-10 02:35 UTC) - #807

Merged
henrypark133 merged 9 commits into
mainfrom
staging-promote/83950d11-22884429853
Mar 10, 2026
Merged

henrypark133 merged 9 commits into
mainfrom
staging-promote/83950d11-22884429853

Conversation

@ironclaw-ci

@ironclaw-ci ironclaw-ci Bot commented Mar 10, 2026

Copy link
Copy Markdown
Contributor

Auto-promotion from staging CI

Batch range: a5f88b32fd0e716df19eaaba4238efc99aa81cc3..83950d11a43ded67b8762b5cf76043d3880f3d7a
Promotion branch: staging-promote/83950d11-22884429853
Base: main
Triggered by: Staging CI batch at 2026-03-10 02:35 UTC

Waiting for gates:

  • Tests: pending
  • E2E: pending
  • Claude Code review: pending (will post comments on this PR)

Auto-created by staging-ci workflow

henrypark133 and others added 7 commits March 9, 2026 16:41
…-check]

GitHub Actions step-level `if:` doesn't have access to `secrets` context.
Replace `if: secrets.X != ''` with `continue-on-error: true` and let
the Set token step handle the fallback.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…798)

Cherry-pick of #794: remove continue-on-error hack, skip redundant checks
on staging PRs, allow ironclaw-ci[bot] in Claude Code review.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* fix: destructive actions from ambiguous user prompts

* review fixes

* review fixes
…799)

When users authenticate via NEAR AI Cloud API key (option 4) during
onboarding, the key is stored as an env var but fetch_nearai_models()
was hardcoding api_key: None. This caused resolve_bearer_token() to
re-trigger the interactive auth prompt at step 4 (model selection).

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
…eck] (#803)

Cherry-pick of #802: run fmt + clippy on staging PRs, skip Windows clippy,
simplify claude-review trigger to labeled-only.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
#788)

* feat: persist user_id in save_job and expose job_id on routine runs (#709)

* feat: persist worker events to DB and fix activity tab rendering

In-process Worker (used by Scheduler::dispatch_job) now persists events
via save_job_event at key execution points: plan creation, LLM
responses, tool_use, tool_result, and job completion/failure/stuck.
Event data shapes match the container worker format so the gateway
activity tab renders them correctly.

Frontend: tool_result errors now show a red X icon with danger styling
instead of a silent empty output. The result event falls back to the
error field when message is absent.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: wire RoutineEngine into gateway for direct manual trigger firing

Replace the message-channel hack in routines_trigger_handler with a
direct call to RoutineEngine::fire_manual(), ensuring FullJob routines
dispatch correctly when triggered from the web UI. Inject the engine
into GatewayState from Agent::run after construction.

Also persists user_id in save_job for both PG and libSQL backends,
removes the source='sandbox' filter so all jobs are visible, and
exposes job_id on RoutineRunInfo for the frontend job link.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: remove stale gateway_state argument from Agent::new test call sites

The gateway_state parameter was removed from Agent::new during rebase
(replaced by post-construction set_routine_engine_slot), but three test
call sites still passed the extra None argument.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR review — restore sandbox source filter, remove blank lines

- Revert removal of `source = 'sandbox'` filter in all SandboxStore
  queries (8 sites across PG and libSQL). Sandbox-specific APIs should
  stay scoped to sandbox jobs; unified job listing for the Jobs tab
  should use a separate query path.
- Remove extra blank lines in agent_loop.rs and worker.rs that caused
  formatting CI failure.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address review — regenerate Cargo.lock, add user_id regression test

- Regenerate Cargo.lock from main's lockfile to eliminate dependency
  version downgrades (anyhow, syn, etc.) that were churn from rebase.
- Add regression test verifying user_id round-trips through save_job
  and get_job in the libSQL backend.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: remove trailing blank line in libsql jobs.rs

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test: add Postgres-side regression test for user_id persistence in save_job

Mirrors the existing libSQL test (test_save_job_persists_user_id) for the
Postgres backend. Gated behind #[cfg(feature = "postgres")] + #[ignore]
since it requires a running PostgreSQL instance (integration tier).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* fix: add job token budget, change iteration cap to Failed, fix web cancel (#698)

Jobs could enter infinite retry loops because: (1) no token budget was
enforced, (2) iteration cap marked jobs as Stuck (allowing self-repair to
restart them), and (3) the web UI cancel button only updated the DB without
stopping the running worker.

- Add `max_tokens_per_job` config (settings.json + AGENT_MAX_TOKENS_PER_JOB
  env var, default 0 = unlimited) with per-job metadata override
- Track token usage after respond_with_tools() and fail the job on budget
  exceeded
- Change iteration cap and persistent rate limiting from mark_stuck to
  mark_failed, preventing self-repair restart loops
- Fix web cancel handler to call scheduler.stop() which updates in-memory
  state AND aborts the worker task, falling back to DB-only update

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR review — always persist cancel to DB, simplify token check

- Cancel handler now always persists Cancelled to DB regardless of whether
  scheduler.stop() ran, fixing the edge case where stop() returns Ok(())
  for jobs not in the scheduler map
- Collapse nested ifs per clippy (let-chains)
- Add NOTE comment about select_tools() not exposing TokenUsage

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: rustfmt formatting in wizard.rs (pre-existing)

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) scope: channel/web Web gateway channel scope: tool/builtin Built-in tools scope: config Configuration scope: setup Onboarding / setup scope: ci CI/CD workflows size: XL 500+ changed lines risk: high Safety, secrets, auth, or critical infrastructure contributor: new First-time contributor labels Mar 10, 2026
@claude

claude Bot commented Mar 10, 2026

Copy link
Copy Markdown

Code review

Found 12 issues:

  1. [CRITICAL:100] Error handling block unreachable: .await? propagates errors before let Err(msg) pattern can match. Token budget overflow is not logged or marked as failed; job fails silently without diagnostics.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/agent/worker.rs#L494-L502

    if total_tokens > 0
        && let Err(msg) = self
            .context_manager()
            .update_context(self.job_id, |ctx| ctx.add_tokens(total_tokens))
            .await?  // <-- ? propagates error before pattern match
    {
        self.mark_failed(&msg).await?;  // unreachable
  1. [CRITICAL:75] TOCTOU race between get_job() and update_job_status() without transaction. Job state can change between check and update; cancellation silently succeeds even if job is no longer active.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/channels/web/handlers/jobs.rs#L283-L305

    let job = store.get_job(job_id).await?;
    if !job.state.is_active() {
        return Err(ApiError::InvalidState(...));
    }
    // <-- Job state can change here
    store.update_job_status(job_id, JobState::Cancelled).await?;
  1. [CRITICAL:50] Non-transactional multi-step context updates between metadata/token setup and DB persist. Concurrent worker can modify context between final get_context() and save_job(), causing in-memory/DB divergence. Crash recovery exposes the bug.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/agent/scheduler.rs#L170-L195

    self.context_manager
        .update_context(job_id, |ctx| { ctx.metadata = meta; })
        .await?;
    // ... separate update for token budget ...
    let ctx = self.context_manager.get_context(job_id).await?;
    // <-- Worker can modify ctx here
    store.save_job(&ctx).await?;  // Persists stale snapshot
  1. [HIGH:100] Token budget not persisted to database. PR adds max_tokens to JobContext but neither PostgreSQL nor libSQL schemas include the column. Jobs lose token budgets on restart, violating CLAUDE.md dual-backend requirement.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/db/libsql/jobs.rs#L28-L66

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/history/store.rs#L572-L596

    // Neither INSERT nor ON CONFLICT includes max_tokens column
    INSERT INTO agent_jobs (... user_id, ... ) VALUES (... ?16, ... )
  1. [HIGH:85] User-supplied metadata bypasses configured token budget without validation. Job metadata can inject {"max_tokens": 9999999} to exceed intended limits.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/agent/scheduler.rs#L164-L168

    let max_tokens = metadata
        .as_ref()
        .and_then(|m| m.get("max_tokens"))
        .and_then(|v| v.as_u64())
        .unwrap_or(self.config.max_tokens_per_job);  // <-- No validation
  1. [HIGH:85] Token budget enforcement incomplete: select_tools() LLM calls bypass tracking while respond_with_tools() calls are counted. Jobs can exceed budgets via the planning loop without being caught.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/agent/worker.rs#L489-L501

    // Line 429: select_tools() LLM call — NOT tracked
    let selections = self.select_tools(&mut state, &request).await?;
    // ...
    // Line 460+: Only respond_with_tools() tokens tracked
    let total_tokens = respond_output.usage.total() as u64;
    if total_tokens > 0 && ... ctx.add_tokens(total_tokens) ...
  1. [HIGH:75] Unsafe environment variable mutation without synchronization in src/setup/wizard.rs. unsafe { std::env::set_var() } and remove_var() mutate global state; concurrent tests cause data races.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/setup/wizard.rs#L3661-L3676

    unsafe {
        std::env::set_var(key, value);  // <-- No locking, test parallel race
    }
  1. [HIGH:75] Removed error handling guard in .github/workflows/staging-ci.yml. Changed from if: ${{ secrets.GH_RELEASES_MANAGER_APP_ID != '' }} to always attempting token generation. Workflow fails hard instead of gracefully degrading to fallback token.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/.github/workflows/staging-ci.yml#L69

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/.github/workflows/staging-ci.yml#L230

    - name: Create GitHub App token
      id: app-token
      uses: actions/create-github-app-token@v1
      // <-- Previously had: if: ${{ secrets.GH_RELEASES_MANAGER_APP_ID != '' }}
  1. [HIGH:75] Breaking UX change: tool_remove and skill_remove downgraded from auto-approvable to always-require-approval without release notes, affecting workflows where users previously auto-approved these tools.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/tools/builtin/extension_tools.rs#L494-L495

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/tools/builtin/skill_tools.rs#L911-L912

    approval: ApprovalRequirement::Always,  // Changed from UnlessAutoApproved
  1. [MEDIUM:75] Stringly-typed token budget with semantic magic number 0 = unlimited violates stated code style preference for strong types. Requires defensive > 0 checks in 3+ places; error-prone for future maintainers.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/config/agent.rs#L485

    max_tokens_per_job: u64,  // Semantic: 0 = unlimited, documented only in comments
    // Better: enum TokenBudget { Unlimited | Limited(u64) }
  1. [MEDIUM:75] Missing error context mapping in token budget enforcement. Error is bare String instead of structured JobError::ContextError. Loss of context about which LLM phase (respond vs select) caused overflow makes debugging harder.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/agent/worker.rs#L499-L500

    self.mark_failed(&msg).await?;  // msg is bare String, no phase context
  1. [LOW:50] Configuration precedence for max_tokens_per_job not documented in one place. Resolution order spans job metadata, env var, settings.json, and default—correct implementation but non-obvious to maintainers and operators.

https://github.com/anthropics/ironclaw/blob/577e26eff4960cd31932e22ef017ae1599f3fbc6/src/agent/scheduler.rs#L163-L168

@github-actions github-actions Bot added size: L 200-499 changed lines and removed size: XL 500+ changed lines labels Mar 10, 2026
@henrypark133
henrypark133 enabled auto-merge March 10, 2026 18:27
@henrypark133
henrypark133 disabled auto-merge March 10, 2026 18:40
@henrypark133
henrypark133 merged commit b442a1f into main Mar 10, 2026
27 checks passed
@henrypark133
henrypark133 deleted the staging-promote/83950d11-22884429853 branch March 10, 2026 18:40
bkutasi pushed a commit to bkutasi/ironclaw that referenced this pull request Mar 28, 2026
…884429853

chore: promote staging to main (2026-03-10 02:35 UTC)
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
…884429853

chore: promote staging to main (2026-03-10 02:35 UTC)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: new First-time contributor risk: high Safety, secrets, auth, or critical infrastructure scope: agent Agent core (agent loop, router, scheduler) scope: channel/web Web gateway channel scope: ci CI/CD workflows scope: config Configuration scope: setup Onboarding / setup scope: tool/builtin Built-in tools size: L 200-499 changed lines staging-promotion

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants