Skip to content

fix(security): tool boundary checks - #4869

Merged
think-in-universe merged 6 commits into
mainfrom
codex/fix-security-tool-boundaries
Jun 16, 2026
Merged

think-in-universe merged 6 commits into
mainfrom
codex/fix-security-tool-boundaries

Conversation

@think-in-universe

@think-in-universe think-in-universe commented Jun 14, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Fixes the validated security issues around built-in filesystem and shell tool boundaries:

  • reject dangling final symlinks during sandbox path validation so write_file cannot follow them outside base_dir
  • classify newline-separated shell commands as chained commands
  • inspect transparent shell wrappers such as env ... sh -c, direct shell -c, and time ... before reusing session-level shell auto-approval
  • force sort --compress-program through the high-risk approval floor
  • add regression coverage through both helper and caller-path tests

Root Cause

write_file validated non-existent paths by checking the nearest existing ancestor while the final write sink followed symlinks. For shell execution, risk classification only inspected shallow command segments, so dangerous payloads hidden behind newlines or wrappers could be downgraded to UnlessAutoApproved and inherit prior session approval.

Validation

  • cargo fmt
  • cargo test --test shell_risk_regression
  • cargo test --lib tools::builtin::path_utils::tests
  • cargo test --lib tools::builtin::file::tests::test_write_file_rejects_dangling_final_symlink

Fixes #4797.
Fixes #4861.
Fixes #4862.
Fixes #4863.
Fixes #4864.
Fixes #4865.

Summary by CodeRabbit

  • Bug Fixes
    • File writes and path validation now reject attempts to use dangling symlinks that resolve outside the sandbox.
  • New Features
    • Shell command risk detection is more robust, using quote/escape-aware parsing and recognizing common wrapper/delegation patterns, including newline and CRLF handling.
  • Tests
    • Added Unix-only coverage for dangling-symlink write rejection.
    • Added shell-risk regression cases for newline/CRLF chaining and embedded high-risk payloads (including env and sort --compress-program).

@coderabbitai

coderabbitai Bot commented Jun 14, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Patches three CVEs in sandbox isolation and shell approval gates. validate_path now detects dangling symlinks via symlink_metadata and rejects before the unsafe ancestor-walk fallback branch. classify_command_risk adds newline/CRLF as segment delimiters and a wrapper-aware recursive classifier that unwraps sh/bash/env/time/sort --compress-program delegation patterns before applying existing risk patterns. Regression tests verify both fixes handle end-to-end attack chains.

Changes

Dangling Symlink Sandbox Escape Fix

Layer / File(s) Summary
validate_path dangling symlink rejection + tests
src/tools/builtin/path_utils.rs, src/tools/builtin/file.rs
validate_path invokes symlink_metadata() in the non-existent-path branch: if the resolved path is a symlink (dangling), returns ToolError::NotAuthorized("Path is a dangling symlink") immediately, bypassing the unsafe nearest-existing-ancestor canonicalization. If metadata fails unexpectedly, returns ToolError::ExecutionFailed. Unix unit test creates a sandbox symlink pointing outside to a temp directory and asserts rejection. WriteFileTool integration test verifies the write is rejected and the external target file is never created.

Shell Risk Classifier Security Hardening

Layer / File(s) Summary
Wrapper detection and recursive segment risk classifier
src/tools/builtin/shell.rs (lines 340–547)
Adds quote/escape-aware tokenization, command basename extraction, env-assignment detection, and wrapper pattern detection for sh/bash/zsh/dash -c, env --split-string, time output/format delegation, and sort --compress-program. classify_segment_risk recursively unwraps delegations and applies existing never-auto-approve/low-risk/medium-risk pattern lists with unknown commands defaulting to Medium. Wrapper-invoked nested commands classified as High elevate the wrapper's risk to High.
classify_command_risk wiring and newline segment splitting
src/tools/builtin/shell.rs (lines 554–576, 788, 801–825)
classify_command_risk delegates to classify_command_risk_inner, which splits the command via the updated split_shell_segments, maps classify_segment_risk over each segment, and returns the maximum risk. split_shell_segments extends operator detection to treat \n and \r as delimiters (with \r\n collapsed to single separator) alongside existing `
Shell risk regression tests
tests/shell_risk_regression.rs (lines 173–295)
Regression tests cover newline-chained, CRLF-chained, and background-chained commands asserting RiskLevel::High + ApprovalRequirement::Always. Transparent wrappers (bash -lc, env ... bash -lc, bash --rcfile, env /bin/sh -c, /usr/bin/time -p) reveal High-risk inner payloads and require Always approval. Non-shell wrapper prefixes and env option-argument forms tested. sort --compress-program both = and space forms assert High + Always. Empty env --split-string conservatively classifies as Medium + UnlessAutoApproved.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Poem

A dangling symlink tried to escape the box,
But symlink_metadata picked the locks. 🔗
Newlines and wrappers once hid deadly rm -rf,
Now the tokenizer sees through every shell riff.
Three CVEs patched, three approval bypasses sealed —
ApprovalRequirement::Always remains revealed. 🛡️

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed Title follows Conventional Commits style (fix(security): tool boundary checks) and accurately describes the main security-focused changes across path validation and shell risk classification.
Description check ✅ Passed Description includes summary with 5 bullets, change type (Security), linked issues (6 fixes), validation evidence (4 cargo commands), security impact detail, and explicit trust-boundary checklist items addressing sandbox escape, policy boundaries, and approval mechanisms.
Linked Issues check ✅ Passed All six linked security issues (#4797, #4861, #4862, #4863, #4864, #4865) have corresponding code changes: dangling symlink rejection in path_utils.rs [#4797], newline separator handling in shell.rs [#4861], sort --compress-program detection [#4862], transparent wrapper unwrapping for env/shell/-c/time [#4863-#4865].
Out of Scope Changes check ✅ Passed All changes directly address documented security objectives: symlink metadata checks reject dangling finals, split_shell_segments now treats newline/CRLF as separators, classify_segment_risk unwraps delegated commands (env, -c, time, sort --compress-program), and regression tests cover all variants. No unrelated refactoring or scope creep detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added scope: tool/builtin Built-in tools size: L 200-499 changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Jun 14, 2026
@think-in-universe think-in-universe changed the title [codex] Fix security tool boundary checks fix(security): tool boundary checks Jun 14, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request enhances path validation to reject dangling symlinks and significantly improves shell command risk classification by introducing robust tokenization and unwrapping of transparent command wrappers (like env, time, and shells). Feedback on these changes highlights several critical security and compatibility improvements: ensuring the shell script argument parser excludes long options starting with -- to prevent bypasses, handling wrapper command options that take separate arguments (e.g., env -u), and supporting Windows path separators (\\) in both command_basename and is_env_assignment.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread src/tools/builtin/shell.rs Outdated
Comment thread src/tools/builtin/shell.rs Outdated
Comment thread src/tools/builtin/shell.rs
Comment thread src/tools/builtin/shell.rs
@think-in-universe
think-in-universe marked this pull request as ready for review June 15, 2026 17:16
Copilot AI review requested due to automatic review settings June 15, 2026 17:16

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/tools/builtin/path_utils.rs`:
- Around line 100-107: Add a `// silent-ok:` annotation to justify the silent
error handling on the symlink_metadata call. This annotation should be placed
before or on the line containing `.unwrap_or(false)` and should explain why
silently converting errors to false is acceptable in this context (e.g., that
NotFound errors are expected for new files and other IO errors will surface at
write time). This satisfies the "Fail loud" invariant by making the intentional
error suppression explicit and documented.

In `@src/tools/builtin/shell.rs`:
- Around line 402-450: The delegated_env_command function currently skips the
argument following -S or --split-string without inspecting its content for
dangerous shell metacharacters or injection patterns. This allows malicious
payloads to hide in that argument. Modify the function to extract and inspect
the argument at idx + 1 when encountering -S or --split-string, checking for
shell metacharacters and dangerous patterns before deciding whether to skip or
process it further. Either return the dangerous content so it gets properly
classified by the risk classifier, or integrate an inline risk check to detect
shell injection attempts within the -S argument itself rather than blindly
skipping it with idx += 2.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 51395c0b-8da0-4c0e-9a50-ec4dd7b3948f

📥 Commits

Reviewing files that changed from the base of the PR and between 9b4fa63 and 2963f61.

📒 Files selected for processing (4)
  • src/tools/builtin/file.rs
  • src/tools/builtin/path_utils.rs
  • src/tools/builtin/shell.rs
  • tests/shell_risk_regression.rs

Comment thread src/tools/builtin/path_utils.rs Outdated
Comment thread src/tools/builtin/shell.rs

@think-in-universe think-in-universe left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review skill pass for current head 2963f61250d6eecc0bd6a006507a2c1978ea1754.

Findings: 4 total, including 2 high-severity approval-classification regressions. GitHub would not allow this account to request changes on this PR, so I am posting as review comments.

Comment thread src/tools/builtin/shell.rs
Comment thread src/tools/builtin/shell.rs Outdated
Comment thread src/tools/builtin/path_utils.rs Outdated
Comment thread src/tools/builtin/shell.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
src/tools/builtin/path_utils.rs (1)

266-278: 🧹 Nitpick | 🔵 Trivial | 💤 Low value

Consider asserting the specific error variant.

The test confirms rejection but doesn't verify it's NotAuthorized with the "dangling symlink" message. A more specific assertion would catch accidental regressions where the path is rejected for a different reason.

Suggested tightening
     let result = validate_path("jump", Some(sandbox.path()));
-    assert!(result.is_err());
+    let err = result.unwrap_err();
+    assert!(
+        matches!(&err, ToolError::NotAuthorized(msg) if msg.contains("dangling symlink")),
+        "expected NotAuthorized(dangling symlink), got {err:?}"
+    );
 }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/tools/builtin/path_utils.rs` around lines 266 - 278, The test
`test_validate_path_rejects_dangling_final_symlink` uses a generic assertion
that only checks if the result is an error without verifying the specific error
type or message. Replace the `assert!(result.is_err());` statement with a more
specific assertion that verifies the error is the `NotAuthorized` variant and
that the error message contains the text "dangling symlink" to ensure the path
is being rejected for the correct reason and catch regressions where rejection
occurs for a different cause.
src/tools/builtin/shell.rs (1)

421-454: ⚠️ Potential issue | 🔴 Critical | ⚡ Quick win

env -S<adjacent> form bypasses detection; destructive payload is classified Low.

delegated_env_command checks exact "-S" / "--split-string" matches (line 425) and the --split-string= prefix (line 428), but not the POSIX-style -S<value> (no space). When the attacker writes:

env -S'rm -rf /tmp/marker'

…the token -S'rm -rf /tmp/marker' falls through to line 443 (generic -* skip), delegated_env_command returns None, and the whole command is classified against env in LOW_RISK_PATTERNS → Low.

Add handling analogous to --split-string=:

 if token == "-S" || token == "--split-string" {
     return tokens.get(idx + 1).map(|script| shell_tokens(script));
 }
 if let Some(script) = token.strip_prefix("--split-string=") {
     return Some(shell_tokens(script));
 }
+if let Some(script) = token.strip_prefix("-S") {
+    if !script.is_empty() {
+        return Some(shell_tokens(script));
+    }
+}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/tools/builtin/shell.rs` around lines 421 - 454, The delegated_env_command
function handles the `-S` flag only when it has a space before its argument, and
the `--split-string=` prefix form, but it does not handle the POSIX-style
`-S<value>` form where the value is adjacent to the flag with no space (e.g.,
`-S'rm -rf /tmp/marker'`). Add a check using token.strip_prefix("-S") after the
existing `--split-string=` check to detect this form, extract the script value,
and return Some(shell_tokens(script)), treating it the same way as the spaced
and equals-sign variants.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/shell_risk_regression.rs`:
- Around line 248-260: Add a regression test case for the `-S<adjacent>` form
(where the flag is directly adjacent to the argument with no space) to the
env_split_string_payloads_are_classified test function. Insert the string "env
-S'rm -rf /tmp/env-adjacent-marker'" into the cmds array alongside the existing
test cases. This ensures that once the delegated_env_command bypass is fixed,
future refactors cannot reintroduce the gap by missing this command form.

---

Outside diff comments:
In `@src/tools/builtin/path_utils.rs`:
- Around line 266-278: The test
`test_validate_path_rejects_dangling_final_symlink` uses a generic assertion
that only checks if the result is an error without verifying the specific error
type or message. Replace the `assert!(result.is_err());` statement with a more
specific assertion that verifies the error is the `NotAuthorized` variant and
that the error message contains the text "dangling symlink" to ensure the path
is being rejected for the correct reason and catch regressions where rejection
occurs for a different cause.

In `@src/tools/builtin/shell.rs`:
- Around line 421-454: The delegated_env_command function handles the `-S` flag
only when it has a space before its argument, and the `--split-string=` prefix
form, but it does not handle the POSIX-style `-S<value>` form where the value is
adjacent to the flag with no space (e.g., `-S'rm -rf /tmp/marker'`). Add a check
using token.strip_prefix("-S") after the existing `--split-string=` check to
detect this form, extract the script value, and return
Some(shell_tokens(script)), treating it the same way as the spaced and
equals-sign variants.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 14dc5a89-6a43-4bd7-bb05-03d94939b37e

📥 Commits

Reviewing files that changed from the base of the PR and between 2963f61 and 369ff3d.

📒 Files selected for processing (3)
  • src/tools/builtin/path_utils.rs
  • src/tools/builtin/shell.rs
  • tests/shell_risk_regression.rs

Comment thread tests/shell_risk_regression.rs
Copilot AI review requested due to automatic review settings June 16, 2026 07:25

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copy link
Copy Markdown
Collaborator Author

Human final review guidance: focus on the security-sensitive tool boundary changes. Please verify shell risk classification still unwraps delegated commands through shell/env/time/sort wrappers, CRLF and single-& separators raise approval correctly, and dangling symlink metadata failures fail closed without weakening normal new-file writes. CI is green at head 931ef930 and review threads are addressed.

Comment thread src/tools/builtin/shell.rs
Comment thread src/tools/builtin/shell.rs Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/tools/builtin/shell.rs (1)

438-451: ⚠️ Potential issue | 🔴 Critical

These three signal options do not consume a separate token; remove them from the idx += 2 branch.

GNU env accepts --block-signal, --default-signal, and --ignore-signal with optional signal values attached via = (e.g., --block-signal=TERM). The signal is not a separate token, so env --block-signal TERM rm -rf / will fail because TERM is misinterpreted as an environment variable assignment, not a signal argument.

The code currently skips these three flags with idx += 2, which is incorrect. They should either be removed from this match arm (so idx += 1 applies) or checked for --flag=value prefix variants separately. Move them to the generic token.starts_with('-') fallback to skip only the current token.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/tools/builtin/shell.rs` around lines 438 - 451, The three signal options
`--block-signal`, `--default-signal`, and `--ignore-signal` are incorrectly
grouped in the match arm that increments idx by 2, but these flags accept their
signal value as an optional `=value` suffix (not a separate token), so they
should only skip one token like other flag prefixes. Remove these three options
from the current match branch and allow them to fall through to the generic
`token.starts_with('-')` fallback that applies `idx += 1` instead.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@src/tools/builtin/shell.rs`:
- Around line 438-451: The three signal options `--block-signal`,
`--default-signal`, and `--ignore-signal` are incorrectly grouped in the match
arm that increments idx by 2, but these flags accept their signal value as an
optional `=value` suffix (not a separate token), so they should only skip one
token like other flag prefixes. Remove these three options from the current
match branch and allow them to fall through to the generic
`token.starts_with('-')` fallback that applies `idx += 1` instead.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 1fa05b44-3d7c-4c06-b9ed-8cacf2905922

📥 Commits

Reviewing files that changed from the base of the PR and between 931ef93 and 3e3b702.

📒 Files selected for processing (2)
  • src/tools/builtin/shell.rs
  • tests/shell_risk_regression.rs

@think-in-universe
think-in-universe added this pull request to the merge queue Jun 16, 2026

Copy link
Copy Markdown
Collaborator Author

I found one still-actionable issue in the current head 3e3b70247887ce30fbefcaf4e73e356caeeb9518: env --block-signal, --default-signal, and --ignore-signal are no-argument options, but they were still handled with the two-token skip path, which can hide the delegated command from shell-risk classification.

I prepared and locally validated a fix in detached commit 3c52b894e:

  • remove those three signal flags from the env-wrapper options that consume the following token
  • add env_signal_options_do_not_consume_command_tokens regression coverage

Local verification: cargo test --test shell_risk_regression passed all 22 tests.

I could not push the fix because GitHub rejected updates to codex/fix-security-tool-boundaries while the PR is in the merge queue:

GH006: Protected branch update failed ... A pull request for this branch has been added to a merge queue. Branches that are queued for merging cannot be updated.

Please dequeue this PR or apply the same patch before merge; otherwise this review comment remains unresolved.

Merged via the queue into main with commit a1d7c3b Jun 16, 2026
41 checks passed
@think-in-universe
think-in-universe deleted the codex/fix-security-tool-boundaries branch June 16, 2026 11:07
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
* Fix security tool boundary checks

* Address shell wrapper review feedback

* Fix shell clippy lifetime lint

* Fix shell wrapper risk regressions

* test: cover adjacent env split-string payloads

* fix: harden env wrapper risk parsing

---------

Co-authored-by: Codex <codex@openai.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: tool/builtin Built-in tools size: L 200-499 changed lines

Projects

None yet

4 participants