Skip to content

chore: promote staging to staging-promote/fdb093fb-24118285581 (2026-04-08 08:28 UTC) - #2137

Merged
henrypark133 merged 15 commits into
staging-promote/fdb093fb-24118285581from
staging-promote/10d970d4-24125728989
Apr 9, 2026
Merged

henrypark133 merged 15 commits into
staging-promote/fdb093fb-24118285581from
staging-promote/10d970d4-24125728989

Conversation

@ironclaw-ci

@ironclaw-ci ironclaw-ci Bot commented Apr 8, 2026 •

Copy link
Copy Markdown
Contributor

Auto-promotion from staging CI

Batch range: a55aff980a4e235590c3af57ded2542512e2f9f6..10d970d456497699d2a4f2c9f41bcd29f4189f28
Promotion branch: staging-promote/10d970d4-24125728989
Base: staging-promote/fdb093fb-24118285581
Triggered by: Staging CI batch at 2026-04-08 08:28 UTC

Commits in this batch (56):

Current commits in this promotion (9)

Current base: staging-promote/fdb093fb-24118285581
Current head: staging-promote/10d970d4-24125728989
Current range: origin/staging-promote/fdb093fb-24118285581..origin/staging-promote/10d970d4-24125728989

Auto-updated by staging promotion metadata workflow

Waiting for gates:

  • Tests: pending
  • E2E: pending
  • Claude Code review: pending (will post comments on this PR)

Auto-created by staging-ci workflow

serrrfirat and others added 2 commits April 8, 2026 17:03
…ner (#2042)

* test: add Slack E2E tests, Rust integration tests, and smoke runner

Replicate the Telegram test infrastructure for the Slack WASM channel:
- Add Slack URL rewriting in wrapper.rs for test API redirection
- Create fake_slack_api.py mock server for E2E tests
- Add 12 Python E2E tests covering setup, DM, mentions, auth, threads, files
- Add 12 Rust integration tests for WASM channel behavior
- Add conftest.py fixtures for isolated Slack test instances
- Add local smoke test runner for pre-release validation with real Slack

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: wrap env::set_var/remove_var in unsafe blocks for Rust 1.83+

CI uses Rust 1.94 which requires unsafe blocks for std::env::set_var
and std::env::remove_var. Wrap the test-only calls in unsafe blocks
with safety comments.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR review feedback

- Replace fragile time.time()-1 fallback with explicit SmokeError in
  run_smoke.py attachment case (reviewer finding #1)
- Add OnceLock<Mutex> guard around env var mutation in wrapper.rs unit
  test to prevent parallel test races (reviewer finding #2)
- Extract duplicated git-worktree discovery into find_project_file()
  helper in slack_auth_integration.rs (reviewer finding #3)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test(channels): generalize WASM HTTP test rewrites

* fix(channels): gate Slack test URL rewrites from release builds

* fix(ci): update wrapper test pairing store ctor

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) scope: channel/wasm WASM channel runtime scope: docs Documentation size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Apr 8, 2026
@claude

claude Bot commented Apr 8, 2026

Copy link
Copy Markdown

Code review

Found 9 issues:

  1. [HIGH:90] Missing error handling on HTTP POST in reset_fake_slack() — the response status is not checked before continuing

https://github.com/anthropics/ironclaw/blob/10d970d456497699d2a4f2c9f41bcd29f4189f28/tests/e2e/conftest.py#L1147-L1149

Network failures will silently corrupt test state when calling await c.post(f"{fake_slack_url}/__mock/reset"). Add .raise_for_status() or check the response status.

  1. [HIGH:85] Synchronous blocking sleep in polling loop

https://github.com/anthropics/ironclaw/blob/10d970d456497699d2a4f2c9f41bcd29f4189f28/scripts/slack_smoke/run_smoke.py#L138

time.sleep(poll_interval_secs) is a blocking synchronous call. If called from async context, it will block the event loop. Use asyncio.sleep() or document the threading context.

  1. [HIGH:80] Race condition with environment variable mutations in tests

https://github.com/anthropics/ironclaw/blob/10d970d456497699d2a4f2c9f41bcd29f4189f28/tests/slack_auth_integration.rs#L767-L802

unsafe { std::env::set_var() } is called before acquiring the ENV_MUTEX lock. The lock is only held after the set, not during. Parallel tests can observe stale env vars. Acquire lock before calling set_var().

  1. [MEDIUM:75] .expect() calls in test fixture setup for WASM runtime

https://github.com/anthropics/ironclaw/blob/10d970d456497699d2a4f2c9f41bcd29f4189f28/tests/slack_auth_integration.rs#L1747

WasmChannelRuntime::new(config).expect("Failed to create runtime") panics on misconfiguration. Use proper error propagation instead of panics in test setup.

  1. [MEDIUM:75] Potential race condition accessing private socket implementation detail

https://github.com/anthropics/ironclaw/blob/10d970d456497699d2a4f2c9f41bcd29f4189f28/tests/e2e/fake_slack_api.py#L164-L167

Accessing site._server.sockets[0] directly is fragile and could raise IndexError. Use aiohttp public API or add bounds checking.

  1. [MEDIUM:75] Missing error handling on HTTP requests in E2E fixture

https://github.com/anthropics/ironclaw/blob/10d970d456497699d2a4f2c9f41bcd29f4189f28/tests/e2e/conftest.py#L1339-L1348

await c.get(...).json() without checking status code. A 500 error would silently fail during JSON parsing. Add .raise_for_status() before calling .json().

  1. [MEDIUM:70] Unbounded polling loop with no minimum interval enforcement

https://github.com/anthropics/ironclaw/blob/10d970d456497699d2a4f2c9f41bcd29f4189f28/scripts/slack_smoke/run_smoke.py#L102-L141

SLACK_SMOKE_POLL_INTERVAL_SECS can be 0 from environment, causing excessive polling. Add minimum validation like max(poll_interval_secs, 0.1).

  1. [MEDIUM:65] File mutation without rollback in test fixture

https://github.com/anthropics/ironclaw/blob/10d970d456497699d2a4f2c9f41bcd29f4189f28/tests/slack_auth_integration.rs#L1776-L1788

Test modifies Slack WASM capabilities file in-place without backup. If test fails after mutation, corruption persists for subsequent tests. Create temporary copy or implement cleanup.

  1. [MEDIUM:60] Hard-coded test credentials scattered across fixtures

https://github.com/anthropics/ironclaw/blob/10d970d456497699d2a4f2c9f41bcd29f4189f28/tests/slack_auth_integration.rs#L1800

Test tokens hardcoded in multiple locations. Centralize in a constants module or use environment-based injection.

* chore(ci): add Dependabot and pin GitHub Actions by SHA

Add automated dependency vulnerability scanning via Dependabot for both
Cargo crates (weekly) and GitHub Actions (weekly). Pin all 101 external
action references across 14 workflow files to full commit SHAs to prevent
supply-chain attacks via compromised tags.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): harden workflows — persist-credentials, permissions, template injection

Address zizmor security audit findings:
- Add persist-credentials: false to all checkout steps (artipacked)
- Add explicit minimal permissions to all workflows (excessive-permissions)
- Move workflow-level write permissions to job level where possible
- Fix template injection in regression-test-check.yml by using env vars

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): address PR review comments

- Group Dependabot updates by ecosystem to reduce PR noise (gemini)
- Add persist-credentials: false to docker.yml checkout (Copilot)
- Move inputs.tag and other expansions to env vars in docker.yml to
  eliminate template injection from workflow_dispatch user input

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): restore git push auth and harden git fetch

Address PR #2043 review comments:
- staging-ci create-promotion-pr: generate App token before checkout
  and pass it to checkout so 'git push origin "$BRANCH"' works
- staging-ci update-tag: re-enable credential persistence so the
  'staging-tested' tag force-push succeeds (job is internal-only)
- release update-registry-checksums: re-enable credential persistence
  so the checksum-update branch push succeeds
- regression-test-check: add '--' to git fetch to prevent refs that
  start with '-' from being interpreted as options

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): address serrrfirat PR review comments

- release.yml: move github.ref_name, needs.plan.outputs.tag, and
  needs.plan.outputs.tag-flag to env vars across plan, build-local-artifacts,
  build-global-artifacts, and host jobs. The tag pattern
  '[0-9]+.[0-9]+.[0-9]+*' has a trailing glob, so a tag like
  '1.2.3\$(curl evil)' could match and be shell-expanded.
- dependabot.yml: split Cargo groups into tokio-ecosystem, serialization,
  wasm, and everything-else to make regression bisection easier when
  CI fails on a Dependabot PR.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): add missing job-level permissions for gh CLI calls

- resolve-promotion-base: add pull-requests: read for 'gh pr list'
- gate: add checks: read for 'gh api .../commits/{sha}/check-runs'

Both were dropped when workflow-level permissions moved to job level.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Zaki Manian <zaki@iqlusion.io>
serrrfirat and others added 12 commits April 8, 2026 23:23
* feat: port ratatui tui onto staging

* Add TUI model picker for /model

* Fix TUI CI lint failures

* Format /tools output as vertical list

* Restore TUI approval modal on thread switch

* Re-emit pending approval events on follow-up messages

* Improve TUI thread handling and activity UI

* Sort TUI resume conversations by activity

* fix(tui): address PR review feedback

* Add TUI thread detail modal for activity sidebar

* feat(tui): improve conversation scrolling UX

- Mouse wheel: 1-line increments (was 3-line jumps)
- PageUp/PageDown: full-page scroll based on viewport height (was 5 lines)
- Add scrollbar widget on conversation right edge (track │, thumb ┃)
- Add "↓ N more ↓ End to return" indicator when scrolled up
- Add auto-follow (pinned_to_bottom) that disengages on scroll-up
  and re-engages when reaching bottom or pressing End
- Clamp scroll offset to valid range (can't scroll past content)
- Add End key binding to jump to bottom

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tui): use engine context pressure data for status bar

The context bar was using cumulative session tokens (total_input +
total_output) which grow unboundedly across turns, making the bar
always show 100% after a few exchanges. Now uses the actual context
window usage from ContextPressure events when available, falling back
to cumulative tokens only before the first engine update arrives.

Also syncs context_window from the engine's max_tokens so the limit
reflects the real model capability instead of name-based heuristics.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tui): render markdown in thread detail modal

The thread detail modal was displaying raw markdown text (plain
line splitting). Now uses render_markdown() for proper formatting
of headers, lists, bold, code blocks, etc.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tui): hydrate sidebar with engine threads and routines at startup

The TUI sidebar was empty until the first user message because
EngineThreadList and RoutineUpdate events were only sent after
processing a message. Now sends initial data right before the
message loop so the activity panel shows existing threads and
routines immediately on startup.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tui): use owner_id for engine thread hydration at startup

list_engine_threads filters by user_id, so passing "" matched no
threads. Now uses self.owner_id() which matches the TUI channel's
user_id, so threads are visible in the sidebar immediately.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tui): fix CI — type errors and formatting in TUI tests

Wrap `started_at` and `updated_at` in `Some(...)` to match
`Option<DateTime<Utc>>` after upstream struct change, and run
`cargo fmt` on files with formatting drift.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): resolve clippy warnings — collapsible ifs and needless borrow

Collapse three nested `if` blocks into `if && let` chains and remove
a needless `&` on the `process_list_threads` call, all in agent_loop.rs.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): add live_harness.rs with updated StatusUpdate patterns

The live_harness.rs file was added to staging after this branch diverged.
When CI merges the PR into staging, the file uses old StatusUpdate patterns
that don't account for the new `detail` and `call_id` fields added by this
branch. Add the file with `..` rest patterns to fix the merge-time compile
errors.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(engine): add skill repair learning loop

* fix(engine): guard skill repair mission updates

* fix(engine): persist skill repair provenance

* fix(engine): address skill-repair PR review feedback

- Fix hex formatting: iterate GenericArray bytes individually instead of
  relying on Display impl which produces debug-like output
- Always recompute content hash from actual doc.content when archiving a
  revision to prevent drift from out-of-band writes
- Prune repair history on rollback to remove records for versions newer
  than the one being restored
- Combine collect_error_messages + collect_observed_actions into a single
  pass (collect_errors_and_actions) to avoid redundant event iteration
- Document bounded revision eviction policy (cap at 10)
- Add comment clarifying concurrent skill-repair / error-diagnosis triggers

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): constrain skill repair updates

* fix(engine): keep insights on completed threads

* style(engine): satisfy fmt and clippy on mission.rs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
* allow private local llm endpoints

* Fix private endpoint review issues

* Fix link-local clippy warning

* Tighten base URL validation follow-ups

* fix(config): non-blocking DNS validation and admin-only LLM key filtering

Address PR #1955 review feedback:

- Wrap to_socket_addrs() in tokio::task::block_in_place when called from
  a multi-threaded async runtime so the LLM utility handlers
  (/api/llm/test_connection, /api/llm/list_models) can no longer stall a
  worker thread on slow DNS. Synchronous callers (env config, CLI) are
  unaffected.

- Add ADMIN_ONLY_LLM_SETTING_KEYS + strip_admin_only_llm_keys helper as
  defense-in-depth: Config::from_db_with_toml and re_resolve_llm_with_secrets
  now take an is_operator flag and strip admin-only base-URL-bearing keys
  (llm_builtin_overrides, llm_custom_providers, ollama_base_url,
  openai_compatible_base_url) from the DB merge for non-operator users.
  This guards future per-user resolve paths and any pre-existing legacy
  rows from reactivating a private/loopback endpoint via the operator
  validation policy. Existing call sites pass true (owner_id is the
  operator scope).

Adds regression tests covering:
  * strip_admin_only_llm_keys removes all four keys, leaves others
  * validate_base_url is callable from a multi-thread tokio runtime
    without panicking on the strict short-circuit path
  * validate_operator_base_url remains callable from async handlers
  * re_resolve_llm filters admin-only keys when is_operator=false and
    keeps them when is_operator=true

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(settings): de-dup admin-only LLM key list (#1955 review)

`handlers/settings.rs::is_admin_only_setting_key` now delegates to
`crate::config::helpers::ADMIN_ONLY_LLM_SETTING_KEYS` so the write-side
gate cannot drift from the read-side `strip_admin_only_llm_keys` filter.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(workspace): admin system prompt shared with all users (#2088)

Introduce SYSTEM.md in a well-known __admin__ scope so admins can set a
system prompt that all tenants receive. Gated behind multi-tenant mode
(WorkspacePool sets admin_prompt_enabled on each workspace; owner
workspace in app.rs also gets the flag when has_any_users() is true).

New endpoints:
- GET  /api/admin/system-prompt — read admin system prompt
- PUT  /api/admin/system-prompt — set admin system prompt (64 KB limit)

Safety:
- SYSTEM.md added to injection scan list
- is_reserved_scope() guard on user creation (defense-in-depth)
- Multi-tenancy gate on both API and prompt assembly layers

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: remove review audit file from tracked files

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: add 64 KB size limit to admin system prompt PUT handler

Addresses PR review feedback:
- Enforce 64 KB limit on system prompt content to prevent token budget
  exhaustion (the content is injected into every user's system prompt)
- Add regression tests for the size limit (413 for oversized, not-413
  for at-limit)
- Document that is_multi_tenant is evaluated once at startup and the
  owner workspace requires a restart after the first user is created

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address remaining review feedback on admin system prompt

- Restore rustdoc comments stripped from document.rs (DocumentMetadata,
  HygieneMetadata, DocumentVersion, VersionSummary, PatchResult, etc.)
  to keep the diff focused on feature additions only
- Replace silent error swallowing (if let Ok) with discriminated match
  in admin prompt read — only DocumentNotFound is silent, other errors
  logged at debug! level
- Cache admin system prompt on WorkspacePool to avoid an extra DB read
  on every turn; invalidated on PUT via invalidate_admin_prompt()
- Add cache invalidation integration test

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(workspace): tighten reserved-scope check and admin-prompt body limit

- is_reserved_scope: case-insensitive, whitespace-tolerant, and reserves
  the entire `__*__` namespace so future system scopes (alongside
  `__admin__`) cannot be impersonated by hand-crafted user IDs
- admin system-prompt route: layer-level DefaultBodyLimit of 128 KB
  rejects oversized payloads before JSON parse, complementing the
  in-handler 64 KB content cap
- system_prompt put_handler: clarify that the in-handler size check is
  a clearer-error fallback for the layer cap
- users_create_handler: drop the dead is_reserved_scope check on a
  freshly-minted UUID; the guard belongs at a code path that actually
  accepts user-supplied IDs
- expand is_reserved_scope tests for case, whitespace, and the wider
  `__*__` namespace

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
* Fix skill installs for invalid catalog names

* Fix clippy test module ordering

* fix: address PR review feedback

* fix: use PairingStore::new_noop() in SSRF test after merge with staging

The staging branch introduced a new test (test_http_request_rejects_private_ip_targets)
that calls PairingStore::new(), but this branch changed the signature to require
db and cache arguments. Use new_noop() since this is a test context.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR #2040 review — remove expect() and dead strip_prefix

- Restructure download_key flow in skills_install_handler to use the
  value directly instead of round-tripping through Option + expect(),
  satisfying the no-expect-in-production-code rule.
- Remove dead strip_prefix("---\n") in render_skill_md — serde_yml does
  not emit a leading document marker for structs.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(skills): preserve unknown frontmatter and tighten install matching

- rewrite install-recovery to mutate the `name` field via raw YAML
  Value rather than re-serializing the typed SkillManifest, so unknown
  frontmatter keys (vendor extensions, future fields) survive the
  install rewrite
- catalog_entry_is_installed: case-insensitive comparison for the
  display-name and normalized-slug branches, matching the slug branch
- normalize_skill_identifier: document non-ASCII handling
- normalizing-invalid-name log: warn -> debug (REPL/TUI rule)
- add round-trip test asserting unknown top-level keys, nested
  mappings, and sequences survive install recovery

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
The test used a hyphenated channel name ("test-failing-channel") but
canonicalize_extension_name() converts hyphens to underscores. This
caused configure() to look for "test_failing_channel.capabilities.json"
which didn't exist, returning an early Err before reaching the
activation code path the test was designed to exercise.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…1770

chore: promote staging to staging-promote/bb2c3e1d-24154330911 (2026-04-08 21:41 UTC)
…0911

chore: promote staging to staging-promote/315c4cf8-24151502580 (2026-04-08 19:29 UTC)
…2580

chore: promote staging to staging-promote/8aa09412-24146405548 (2026-04-08 18:23 UTC)
…5548

chore: promote staging to staging-promote/482ee57c-24140982267 (2026-04-08 16:25 UTC)
…2267

chore: promote staging to staging-promote/3df1bf38-24132614443 (2026-04-08 14:32 UTC)
…4443

chore: promote staging to staging-promote/10d970d4-24125728989 (2026-04-08 11:19 UTC)
@github-actions github-actions Bot added scope: channel Channel infrastructure scope: channel/web Web gateway channel scope: tool/builtin Built-in tools scope: workspace Persistent memory / workspace scope: extensions Extension management scope: ci CI/CD workflows labels Apr 9, 2026
@github-actions github-actions Bot added the scope: dependencies Dependency updates label Apr 9, 2026
@henrypark133
henrypark133 merged commit 2c6aedf into staging-promote/fdb093fb-24118285581 Apr 9, 2026
14 checks passed
@henrypark133
henrypark133 deleted the staging-promote/10d970d4-24125728989 branch April 9, 2026 04:15
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…4125728989

chore: promote staging to staging-promote/ac420970-24118285581 (2026-04-08 08:28 UTC)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: agent Agent core (agent loop, router, scheduler) scope: channel/wasm WASM channel runtime scope: channel/web Web gateway channel scope: channel Channel infrastructure scope: ci CI/CD workflows scope: dependencies Dependency updates scope: docs Documentation scope: extensions Extension management scope: tool/builtin Built-in tools scope: workspace Persistent memory / workspace size: XL 500+ changed lines staging-promotion

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants