Skip to content

fix(tests): eliminate env mutex poison cascade - #1558

Merged
ilblackdragon merged 4 commits into
stagingfrom
fix/env-mutex-poison-cascade
Mar 22, 2026
Merged

ilblackdragon merged 4 commits into
stagingfrom
fix/env-mutex-poison-cascade

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

Summary

  • Poison-recovering lock_env() helper: Added config::helpers::lock_env() that acquires ENV_MUTEX with .unwrap_or_else(|e| e.into_inner()) instead of .unwrap()/.expect(). A single test panic no longer cascades ~68 tests into PoisonError failures.
  • Consolidated rogue module-local locks: workspace.rs, orchestrator/mod.rs, and bootstrap.rs each had their own ENV_LOCK/ENV_MUTEX that didn't synchronize with the crate-wide mutex, causing cross-module env var races. All now use the shared lock_env().
  • Fixed gateway user_id fallback: ChannelsConfig::resolve() hardcoded the gateway user_id fallback to "default" instead of using the owner_id parameter (HTTP channel already did this correctly).
  • Fixed test_ironclaw_env_path: Was calling ironclaw_env_path() which reads from a LazyLock whose cached value depends on test execution order. Now uses compute_ironclaw_base_dir() directly, matching the pattern of all other bootstrap tests.

18 files changed across src/config/, src/cli/, src/llm/, src/db/, src/extensions/, src/orchestrator/, src/setup/, and src/bootstrap.rs.

Test plan

  • cargo fmt clean
  • cargo clippy --all --benches --tests --examples --all-features zero warnings
  • cargo test --lib --all-features passes 3585 tests, 0 failures
  • Config tests that previously cascaded (config::llm, config::embeddings, config::search, config::safety, config::sandbox, config::wasm, config::builder) all pass
  • Extension manager gateway callback tests pass (71/71)
  • Bootstrap tests pass including previously-failing test_ironclaw_env_path
  • Channels resolve_uses_settings_channel_values_with_owner_scope_user_ids now passes

🤖 Generated with Claude Code

The shared ENV_MUTEX used by ~68 config tests would cascade a single
test panic into failures across every module. Replace all .unwrap() /
.expect() lock acquisitions with a poison-recovering lock_env() helper.
Consolidate rogue module-local ENV_LOCK instances (workspace, orchestrator,
bootstrap) onto the shared global mutex to prevent cross-module races.

Also fixes:
- gateway user_id fallback was hardcoded to "default" instead of owner_id
- test_ironclaw_env_path used LazyLock which is order-dependent

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings March 22, 2026 07:10
@github-actions github-actions Bot added scope: channel/cli TUI / CLI channel scope: llm LLM integration scope: orchestrator Container orchestrator scope: extensions Extension management scope: setup Onboarding / setup size: L 200-499 changed lines risk: high Safety, secrets, auth, or critical infrastructure contributor: core 20+ merged PRs labels Mar 22, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the stability and reliability of the test suite by addressing environment variable handling. It introduces a robust mechanism to recover from poisoned mutexes during tests, preventing widespread failures. Additionally, it centralizes environment variable locking across various modules, ensuring consistent and safe access. Several specific configuration fallbacks and test logic were also refined to improve correctness and prevent unexpected behavior.

Highlights

  • Test Stability Improvement: Introduced a poison-recovering lock_env() helper that acquires ENV_MUTEX using unwrap_or_else(|e| e.into_inner()), preventing test panics from cascading into PoisonError failures across ~68 tests.
  • Environment Variable Synchronization: Consolidated module-local ENV_LOCK/ENV_MUTEX instances in workspace.rs, orchestrator/mod.rs, and bootstrap.rs to use the new shared lock_env() helper, eliminating cross-module environment variable race conditions.
  • Gateway User ID Fallback Correction: Fixed an issue in ChannelsConfig::resolve() where the gateway user_id fallback was hardcoded to 'default' instead of correctly utilizing the owner_id parameter.
  • Bootstrap Test Reliability: Corrected test_ironclaw_env_path to directly use compute_ironclaw_base_dir() instead of ironclaw_env_path(), resolving issues caused by LazyLock caching and test execution order dependencies.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This is a great pull request that significantly improves the test suite's robustness and fixes a couple of bugs. The introduction of the poison-recovering lock_env() helper and the consolidation of environment variable mutexes are excellent changes that will prevent cascading test failures and eliminate potential race conditions. The bug fixes for the gateway user_id fallback and the test_ironclaw_env_path are also correct and well-implemented. Overall, this is a high-quality contribution.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR hardens Rust unit tests that touch process environment variables by centralizing env locking with poison recovery, preventing one test panic from cascading into widespread PoisonError failures. It also fixes a config bug where the gateway channel’s user_id fallback ignored the provided owner_id.

Changes:

  • Introduce config::helpers::lock_env() (test-only) to acquire the shared env mutex while recovering from poison.
  • Replace module-local env mutexes / direct ENV_MUTEX.lock() usage across many test modules with lock_env().
  • Fix ChannelsConfig::resolve() to default gateway user_id to the passed owner_id (matching HTTP behavior) and adjust a bootstrap path test to avoid LazyLock ordering sensitivity.

Reviewed changes

Copilot reviewed 18 out of 18 changed files in this pull request and generated no comments.

Show a summary per file
File Description
src/config/helpers.rs Adds lock_env() helper that recovers from poisoned env mutex in tests.
src/setup/wizard.rs Updates tests to use lock_env() for env serialization.
src/orchestrator/mod.rs Removes module-local env lock in tests; uses shared lock_env().
src/llm/oauth_helpers.rs Updates OAuth helper tests to use lock_env().
src/extensions/manager.rs Updates extension manager tests/guards to use lock_env().
src/db/libsql/workspace.rs Updates workspace dimension resolution tests to use lock_env().
src/config/workspace.rs Removes module-local env lock in tests; uses lock_env().
src/config/wasm.rs Updates WASM config tests to use lock_env().
src/config/search.rs Updates search config tests to use lock_env().
src/config/sandbox.rs Updates sandbox config tests to use lock_env().
src/config/safety.rs Updates safety config tests to use lock_env().
src/config/llm.rs Updates LLM config tests to use lock_env().
src/config/embeddings.rs Updates embeddings config tests to use lock_env().
src/config/channels.rs Fixes gateway user_id fallback to owner_id; updates tests to use lock_env().
src/config/builder.rs Updates builder config tests to use lock_env().
src/cli/oauth_defaults.rs Updates CLI OAuth defaults tests to use lock_env().
src/cli/doctor.rs Updates doctor CLI tests to use lock_env().
src/bootstrap.rs Removes module-local env mutex in tests; updates env-path test to avoid LazyLock ordering issues and uses lock_env().

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

ilblackdragon and others added 2 commits March 22, 2026 00:28
Satisfies the regression-test-check CI gate by adding a test that
intentionally poisons ENV_MUTEX and verifies lock_env() recovers.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The regression test check relied on git diff -W to expand context to
function boundaries, but git doesn't recognize Rust `mod tests {}` as a
function boundary. Changes to imports, helpers, or lock calls inside
test modules were invisible to the check.

Add a line-level fallback: for each changed .rs file, find where
#[cfg(test)] starts and check if any diff hunk targets a line at or
after that boundary. This catches edits anywhere inside test modules
regardless of git's language awareness.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings March 22, 2026 07:33
@github-actions github-actions Bot added the scope: ci CI/CD workflows label Mar 22, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 19 out of 19 changed files in this pull request and generated 2 comments.


💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/config/helpers.rs

// The mutex is now poisoned. lock_env() should recover, not cascade.
assert!(ENV_MUTEX.lock().is_err(), "mutex should be poisoned");
let _guard = lock_env(); // must not panic

Copilot AI Mar 22, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lock_env_recovers_from_poisoned_mutex intentionally poisons the global ENV_MUTEX but never clears the poison flag, leaving the mutex permanently poisoned for the remainder of the test suite. Even though lock_env() recovers, this introduces persistent global state and can break any future code/tests that still call ENV_MUTEX.lock() directly. Consider calling ENV_MUTEX.clear_poison() after the assertions (or otherwise resetting the poison state) so this test doesn't affect unrelated tests.

Suggested change
let _guard = lock_env(); // must not panic
let _guard = lock_env(); // must not panic
// Reset global state so this test does not leave ENV_MUTEX permanently poisoned.
ENV_MUTEX.clear_poison();

Copilot uses AI. Check for mistakes.
Comment on lines +136 to +157
# Line-level check: detect changes inside #[cfg(test)] regions.
# git -W relies on function boundary detection which misses Rust mod blocks,
# so this fallback checks whether changed line numbers fall within test modules.
CHANGED_RS=$(echo "$CHANGED_FILES" | grep '\.rs$' || true)
if [ -n "$CHANGED_RS" ]; then
while IFS= read -r rs_file; do
[ -f "$rs_file" ] || continue

# Find the line number where #[cfg(test)] appears (start of test module).
TEST_MOD_START=$(grep -n '#\[cfg(test)\]' "$rs_file" | head -1 | cut -d: -f1 || true)
[ -n "$TEST_MOD_START" ] || continue

# Get changed line numbers in this file from the diff hunk headers.
# Each @@ line looks like: @@ -old,count +new,count @@
while IFS= read -r hunk_line; do
line_no=$(echo "$hunk_line" | sed -E 's/^@@ -[0-9,]+ \+([0-9]+).*/\1/')
[ -n "$line_no" ] || continue
if [ "$line_no" -ge "$TEST_MOD_START" ]; then
echo "Test changes found: $rs_file has changes at line $line_no inside #[cfg(test)] region (starts at line $TEST_MOD_START)."
exit 0
fi
done < <(git diff "${BASE_REF}...${HEAD_REF}" -U0 -- "$rs_file" | grep -E '^@@')

Copilot AI Mar 22, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fallback 'Line-level check' treats the first #[cfg(test)] occurrence in a file as the start of a test-only region. This is not reliable: some files define test-only items near the top (e.g., src/config/helpers.rs has #[cfg(test)] for ENV_MUTEX near the top), so any later non-test code edits would be misclassified as "test changes" and incorrectly satisfy this enforcement check. Consider detecting #[cfg(test)] mod tests blocks specifically (and ideally their end), or scanning diff context for mod tests/#[test] markers rather than comparing against the first #[cfg(test)] line number.

Suggested change
# Line-level check: detect changes inside #[cfg(test)] regions.
# git -W relies on function boundary detection which misses Rust mod blocks,
# so this fallback checks whether changed line numbers fall within test modules.
CHANGED_RS=$(echo "$CHANGED_FILES" | grep '\.rs$' || true)
if [ -n "$CHANGED_RS" ]; then
while IFS= read -r rs_file; do
[ -f "$rs_file" ] || continue
# Find the line number where #[cfg(test)] appears (start of test module).
TEST_MOD_START=$(grep -n '#\[cfg(test)\]' "$rs_file" | head -1 | cut -d: -f1 || true)
[ -n "$TEST_MOD_START" ] || continue
# Get changed line numbers in this file from the diff hunk headers.
# Each @@ line looks like: @@ -old,count +new,count @@
while IFS= read -r hunk_line; do
line_no=$(echo "$hunk_line" | sed -E 's/^@@ -[0-9,]+ \+([0-9]+).*/\1/')
[ -n "$line_no" ] || continue
if [ "$line_no" -ge "$TEST_MOD_START" ]; then
echo "Test changes found: $rs_file has changes at line $line_no inside #[cfg(test)] region (starts at line $TEST_MOD_START)."
exit 0
fi
done < <(git diff "${BASE_REF}...${HEAD_REF}" -U0 -- "$rs_file" | grep -E '^@@')
# Hunk-level check: detect changes in or near Rust test code.
# Instead of assuming the first #[cfg(test)] marks a test-only region,
# scan the diff hunks for common test markers like `mod tests`, `#[test]`,
# or `#[cfg(test)]`. This avoids misclassifying non-test code below
# early #[cfg(test)] items as test changes.
CHANGED_RS=$(echo "$CHANGED_FILES" | grep '\.rs$' || true)
if [ -n "$CHANGED_RS" ]; then
while IFS= read -r rs_file; do
[ -f "$rs_file" ] || continue
# Look for test markers in the diff hunks for this file.
if git diff "${BASE_REF}...${HEAD_REF}" -- "$rs_file" \
| grep -E 'mod tests|#\[test\]|#\[cfg\(test\)\]' >/dev/null 2>&1; then
echo "Test changes found in Rust file: $rs_file (diff contains test markers)."
exit 0
fi

Copilot uses AI. Check for mistakes.
- Clear ENV_MUTEX poison after regression test so it doesn't leave
  global state dirty for subsequent tests.
- Fix CI regression-test-check to match #[cfg(test)] only when followed
  by `mod` (the test module pattern), avoiding false positives from
  standalone #[cfg(test)] items like statics or functions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@ilblackdragon

Copy link
Copy Markdown
Member Author

Addressed both review comments in 7e2ce9e:

  1. Poison cleanup (helpers.rs): Added ENV_MUTEX.clear_poison() after assertions so the regression test doesn't leave the mutex permanently poisoned for subsequent tests.

  2. CI false positive (regression-test-check.yml): The #[cfg(test)] line-level check now uses an awk pattern that matches only when followed by mod on the same or next line. This avoids false positives from standalone #[cfg(test)] items (like the ENV_MUTEX static at line 14 of helpers.rs).

bkutasi pushed a commit to bkutasi/ironclaw that referenced this pull request Mar 28, 2026
* fix(tests): eliminate env mutex poison cascade and fix test flakiness

The shared ENV_MUTEX used by ~68 config tests would cascade a single
test panic into failures across every module. Replace all .unwrap() /
.expect() lock acquisitions with a poison-recovering lock_env() helper.
Consolidate rogue module-local ENV_LOCK instances (workspace, orchestrator,
bootstrap) onto the shared global mutex to prevent cross-module races.

Also fixes:
- gateway user_id fallback was hardcoded to "default" instead of owner_id
- test_ironclaw_env_path used LazyLock which is order-dependent

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(helpers): add regression test for lock_env poison recovery

Satisfies the regression-test-check CI gate by adding a test that
intentionally poisons ENV_MUTEX and verifies lock_env() recovers.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): detect test changes inside #[cfg(test)] regions

The regression test check relied on git diff -W to expand context to
function boundaries, but git doesn't recognize Rust `mod tests {}` as a
function boundary. Changes to imports, helpers, or lock calls inside
test modules were invisible to the check.

Add a line-level fallback: for each changed .rs file, find where
#[cfg(test)] starts and check if any diff hunk targets a line at or
after that boundary. This catches edits anywhere inside test modules
regardless of git's language awareness.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review feedback

- Clear ENV_MUTEX poison after regression test so it doesn't leave
  global state dirty for subsequent tests.
- Fix CI regression-test-check to match #[cfg(test)] only when followed
  by `mod` (the test module pattern), avoiding false positives from
  standalone #[cfg(test)] items like statics or functions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
* fix(tests): eliminate env mutex poison cascade and fix test flakiness

The shared ENV_MUTEX used by ~68 config tests would cascade a single
test panic into failures across every module. Replace all .unwrap() /
.expect() lock acquisitions with a poison-recovering lock_env() helper.
Consolidate rogue module-local ENV_LOCK instances (workspace, orchestrator,
bootstrap) onto the shared global mutex to prevent cross-module races.

Also fixes:
- gateway user_id fallback was hardcoded to "default" instead of owner_id
- test_ironclaw_env_path used LazyLock which is order-dependent

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(helpers): add regression test for lock_env poison recovery

Satisfies the regression-test-check CI gate by adding a test that
intentionally poisons ENV_MUTEX and verifies lock_env() recovers.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): detect test changes inside #[cfg(test)] regions

The regression test check relied on git diff -W to expand context to
function boundaries, but git doesn't recognize Rust `mod tests {}` as a
function boundary. Changes to imports, helpers, or lock calls inside
test modules were invisible to the check.

Add a line-level fallback: for each changed .rs file, find where
#[cfg(test)] starts and check if any diff hunk targets a line at or
after that boundary. This catches edits anywhere inside test modules
regardless of git's language awareness.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review feedback

- Clear ENV_MUTEX poison after regression test so it doesn't leave
  global state dirty for subsequent tests.
- Fix CI regression-test-check to match #[cfg(test)] only when followed
  by `mod` (the test module pattern), avoiding false positives from
  standalone #[cfg(test)] items like statics or functions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: high Safety, secrets, auth, or critical infrastructure scope: channel/cli TUI / CLI channel scope: ci CI/CD workflows scope: extensions Extension management scope: llm LLM integration scope: orchestrator Container orchestrator scope: setup Onboarding / setup size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants