Skip to content

fix(safety): add credential patterns and sensitive path blocklist - #1675

Merged
ilblackdragon merged 19 commits into
nearai:stagingfrom
j-bloggs:security/leak-detector-improvements
Apr 7, 2026
Merged

ilblackdragon merged 19 commits into
nearai:stagingfrom
j-bloggs:security/leak-detector-improvements

Conversation

@j-bloggs

@j-bloggs j-bloggs commented Mar 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Move is_sensitive_path to ironclaw_safety crate as shared module (sensitive_paths.rs)
  • Guard ListDirTool with sensitive path checks at top-level AND recursive traversal
  • Annotate blocked directories with [sensitive - access blocked] in listings
  • Add missing sensitive paths: ~/.config/gh/hosts.yml, /etc/shadow, ~/.terraform.d/credentials.tfrc.json, ~/.azure/
  • Add path traversal regression test

Change Type

  • Bug fix (non-breaking change which fixes an issue)

Linked Issue

Addresses review feedback from zmanian, gemini-code-assist, and copilot on this PR.

Validation

  • cargo fmt --all -- --check: pass
  • cargo clippy (all-features, default, libsql): pass (zero warnings)
  • cargo test -p ironclaw_safety -- sensitive_paths: 9 tests pass
  • cargo test --lib -- tools::builtin::file: 21 tests pass

Security Impact

  • Closes ListDirTool enumeration gap (attacker could enumerate ~/.ssh/, ~/.aws/ contents)
  • Shell tool bypass via cat ~/.ssh/id_rsa is a known limitation. Shell DANGEROUS_PATTERNS provide partial coverage but are bypassable. Full mitigation requires filesystem-level sandboxing.

Database Impact

None

Blast Radius

  • ListDirTool may now block listing of directories it previously allowed (sensitive credential dirs)
  • File tools now import is_sensitive_path from ironclaw_safety instead of local definition

Rollback Plan

Revert to staging (no sensitive path blocking). Credential patterns are additive and safe to keep.

Review Track

Track C - security changes in crates/ironclaw_safety/ and src/tools/builtin/

Feature Parity

No FEATURE_PARITY.md changes needed.

@github-actions github-actions Bot added the scope: tool/builtin Built-in tools label Mar 26, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the security posture of the system by expanding its ability to detect and prevent the leakage of sensitive credentials. It introduces new patterns for various API keys and tokens within the leak detector and establishes a comprehensive blocklist for file system access to common credential storage locations. These changes are crucial for hardening the application against credential exfiltration and unauthorized file manipulation, discovered through proactive security testing.

Highlights

  • Leak Detector Enhancements: Added four new critical credential patterns to the leak detector, including OpenRouter, Anthropic OAuth, Telegram bot tokens, and Groq API keys, to proactively identify and block sensitive information.
  • Sensitive Path Blocklist: Implemented a robust sensitive path blocklist for ReadFileTool, WriteFileTool, and ApplyPatchTool to prevent access to credential-bearing files such as .env, .ssh/, and .aws/ configurations.
  • .env File Handling: The .env file blocking mechanism now intelligently covers common variants like .env.local and .env.production while explicitly allowing safe suffixes such as .env.example and .env.template.
  • Symlink Safety: Ensured the sensitive path checks are symlink-safe by independently canonicalizing paths within is_sensitive_path() to prevent bypasses through symbolic links.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions github-actions Bot added size: L 200-499 changed lines risk: medium Business logic, config, or moderate-risk modules contributor: new First-time contributor labels Mar 26, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request enhances security by adding new leak detection patterns for various API keys (OpenRouter, Anthropic OAuth, Telegram, Groq) and implementing a sensitive file path blocklist for file access tools. The is_sensitive_path function was introduced to prevent ReadFileTool, WriteFileTool, and ApplyPatchTool from interacting with credential-bearing files like .env or .ssh configurations. Feedback suggests that the sensitive path check needs to be made platform-agnostic by normalizing path separators, and the new leak detection regexes should incorporate word boundaries to prevent false positives and ensure consistency.

Comment thread src/tools/builtin/file.rs Outdated
Comment thread crates/ironclaw_safety/src/leak_detector.rs Outdated
Comment thread crates/ironclaw_safety/src/leak_detector.rs Outdated
Comment thread crates/ironclaw_safety/src/leak_detector.rs Outdated

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review: credential patterns and sensitive path blocklist

Solid security hardening. Good credential patterns for OpenRouter/Anthropic OAuth/Telegram/Groq. However:

Critical

  1. ListDirTool not guarded: ReadFileTool, WriteFileTool, ApplyPatchTool get is_sensitive_path checks but ListDirTool does not. Attacker can enumerate ~/.ssh/, ~/.aws/, ~/.gnupg/ contents.

High

  1. Shell tool bypass: cat ~/.ssh/id_rsa via the shell tool trivially reads any sensitive file. At minimum document as known limitation with a tracking issue.

Medium

  1. Missing sensitive paths: ~/.config/gh/hosts.yml (GitHub CLI), /etc/shadow, ~/.terraform.d/credentials.tfrc.json.
  2. No test for path traversal bypass (../../.ssh/id_rsa). canonicalize() should handle it but an explicit test would strengthen confidence.

Low

  1. Telegram pattern \d{8,12}:AA could false-match log lines with timestamps. Low risk but document the :AA requirement.

Positive

  • Credential patterns are well-crafted with appropriate minimum lengths
  • Symlink-through-canonicalize defense is correct
  • .env.example/.env.local safe suffixes are thoughtful
  • 6 tests covering credential detection

Fix the ListDirTool gap and track the shell bypass before merge.

j-bloggs and others added 7 commits March 27, 2026 23:03
Addresses critical credential leakage found by security testing
(~/.ironclaw/tests/SECURITY_REPORT.md, test ce-02).

Leak detector (crates/ironclaw_safety/src/leak_detector.rs):
- Add OpenRouter API key pattern (sk-or-v1-<hex>)
- Add Anthropic OAuth token pattern (sk-ant-oat<NN>-<base64url>)
- Add Telegram bot token pattern (word-bounded, 8-12 digit bot ID)
- Add Groq API key pattern (gsk_<alphanumeric>)
- 12 new tests with synthetic keys (positive, false-positive, integration)

File tools (src/tools/builtin/file.rs):
- Add sensitive path blocklist to ReadFileTool, WriteFileTool, and
  ApplyPatchTool (defense-in-depth for all file access vectors)
- Blocks: .env (and .env.local/.env.production/etc.), .ssh/, .aws/,
  .netrc, .pgpass, .npmrc, .pypirc, .docker/config.json, .kube/config,
  .git-credentials, .gcloud/, .config/gcloud/, .gnupg/, .vault-token,
  .ironclaw/secrets/
- Allows .env.example, .env.template, .env.sample (safe suffixes)
- Case-insensitive; resolves symlinks via canonicalize() before check
- 8 new tests covering blocking, safe suffixes, .env variants, case

Known gap: shell tool can still `cat ~/.env` — different security
domain (denylist-gated in autonomous mode, user-initiated in
interactive mode). Tracked for follow-up.

Note: new patterns use .unwrap() on Regex::new() matching the
established convention of the 16 existing patterns in this file
(all use // safety: hardcoded literal). Follow-up to address the
existing .unwrap() debt across all patterns.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
…blocklist

- Move is_sensitive_path to ironclaw_safety crate for shared use
- Guard ListDirTool with sensitive path checks (including recursive traversal)
- Add missing sensitive paths: ~/.config/gh/hosts.yml, /etc/shadow,
  ~/.terraform.d/credentials.tfrc.json, ~/.azure/
- Add path traversal regression test and ListDirTool blocking test

Addresses review feedback from zmanian, gemini-code-assist, and copilot.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When ListDirTool's recursive traversal encounters a sensitive directory,
annotate it with [sensitive - access blocked] so users understand why
its contents are suppressed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@j-bloggs
j-bloggs force-pushed the security/leak-detector-improvements branch from 84f574b to d9ba583 Compare March 27, 2026 12:07

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: credential patterns and sensitive path blocklist

3 of 4 original findings fixed. Good progress.

Finding Status
CRITICAL: ListDirTool not guarded Fixed -- blocks sensitive dirs, annotates as [sensitive - access blocked]
HIGH: Shell tool bypass Still open -- cat/head/less on protected paths still works via shell tool
MEDIUM: Missing sensitive paths Fixed -- added gh/hosts.yml, /etc/shadow, terraform, azure
MEDIUM: No path traversal test Fixed -- canonicalize + raw string matching tests

The shell tool bypass remains the main gap. Consider synchronizing DANGEROUS_PATTERNS with SENSITIVE_PATH_PATTERNS, or integrating is_sensitive_path() into shell argument scanning.

…hecking

Add defense-in-depth check for shell commands that read sensitive
credential files (cat, head, tail, less, cp, etc.). Extracts file
path arguments from known file-reading commands and checks them
against the shared is_sensitive_path function from ironclaw_safety.

This is best-effort — shell-level bypass via aliases, variable
expansion, or encoding is still possible. Full mitigation requires
filesystem-level sandboxing (seccomp/landlock).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added size: XL 500+ changed lines and removed size: L 200-499 changed lines labels Mar 28, 2026
@j-bloggs

Copy link
Copy Markdown
Contributor Author

Addressed in 46e2846: integrated is_sensitive_path() from the safety crate into the shell tool's validation pipeline. File-reading commands (cat, head, tail, less, cp, mv, etc.) now have their path arguments checked against the shared sensitive path list.

What's covered:

  • cat ~/.ssh/id_rsa, head /home/user/.env, cp ~/.aws/credentials /tmp/
  • Piped chains: cat ~/.env | grep KEY
  • Input redirection: < ~/.ssh/id_rsa
  • Full-path commands: /usr/bin/cat ~/.ssh/id_rsa
  • 8 regression tests

Known limitation (documented): Shell-level bypass via aliases, variable expansion, encoding, or non-listed commands remains possible. This is defense-in-depth, not a complete fix. Full mitigation requires filesystem-level sandboxing (seccomp/landlock).

Validation: cargo fmt, clippy x3, cargo test --lib -- sensitive_file_access (8 pass).

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review of commit 46e2846. See detailed findings below.

@zmanian

zmanian commented Mar 28, 2026

Copy link
Copy Markdown
Collaborator

Re-review of 46e2846 - posting detailed review

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: shell tool integration (commit 46e2846)

The new commit addresses the HIGH-priority shell tool bypass. Good progress.

What was done well

  • Correct placement: check_sensitive_file_access runs after injection detection, before execution
  • Shared is_sensitive_path reuse from ironclaw_safety -- no duplicated logic
  • Tilde expansion, pipe/semicolon splitting, input redirection, full-path command stripping all handled
  • Honest documentation: clearly states this is defense-in-depth, not a complete sandbox
  • 8 regression tests covering key scenarios

Previous findings status

All 4 previous findings (ListDirTool, shell bypass, missing paths, traversal test) are fixed.

New findings

Important (should fix):

  1. Output redirection bypass: The function checks read commands but not write targets via > or >>. An attacker could write to sensitive credential paths via redirection. Consider scanning redirect targets against is_sensitive_path. This mirrors the existing comment on line 902 about redirect-aware parsing.

  2. Subshell / command substitution bypass: $() and backtick substitution not handled. A sensitive read nested in command substitution would not be caught. Exploitability is lower since detect_command_injection partially covers this upstream. Document the gap.

  3. Ampersand splitting fragility: Splitting on single & incorrectly splits && into segments. Works by accident (empty segment is harmless) but is fragile. Consider splitting on ["&&", "||", "|", ";"] explicitly.

Verdict

Meaningful security improvement that closes the most obvious bypass. Known limitations honestly documented. Items 1-3 worth addressing before merge -- item 1 (output redirection) is most actionable as the write-path equivalent of the read-path protection this commit adds.

Overall: good security work. Substantially better than at the start of the review cycle.

Address zmanian's re-review findings:
- Check > and >> redirection targets against is_sensitive_path
  (write-path equivalent of read-path protection)
- Replace single-char & splitting with proper &&/|| aware parser
  to avoid fragmenting double operators
- Document subshell/command-substitution gap (partially covered by
  detect_command_injection upstream)
- Extract helpers: split_shell_segments, check_segment_file_commands,
  check_redirect_target, expand_tilde
- Add tests for output redirection, chained commands, segment splitting

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@j-bloggs

Copy link
Copy Markdown
Contributor Author

Addressed all 3 findings from re-review in 8175f93:

  1. Output redirection bypass -- now checks > and >> targets against is_sensitive_path. Blocks echo pwned > ~/.ssh/authorized_keys and echo extra >> ~/.env. Test: sensitive_file_access_blocks_output_redirection.

  2. Subshell/command-substitution gap -- documented in function doc comment. Partially mitigated by detect_command_injection upstream which flags $( patterns.

  3. Ampersand splitting fragility -- replaced single-char & split with a byte-level state machine that correctly handles && and || as 2-char operators. Single | and ; still split as before. Test: split_shell_segments_handles_operators.

Also refactored into smaller helpers: split_shell_segments, check_segment_file_commands, check_redirect_target, expand_tilde.

Validation: cargo fmt, clippy x3, 11 tests pass.

@github-actions github-actions Bot added contributor: regular 2-5 merged PRs and removed contributor: new First-time contributor labels Mar 28, 2026

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Request Changes

Credential patterns are well-crafted (all ReDoS-safe). Putting is_sensitive_path in ironclaw_safety is architecturally correct. However, there's a critical conflict.

Must-fix

  1. Direct conflict with PR #1713 -- both implement sensitive path protection, touching the exact same insertion points in file.rs and shell.rs. Recommend: use this PR as base (safety crate is the right home), absorb #1713's more complete path list, close #1713 as superseded.

  2. Telegram pattern needs trailing \b -- \b\d{8,12}:AA[A-Za-z0-9_-]{30,} could false-positive on structured data.

Should-fix

  1. Missing patterns from #1713: .bash_history, .zsh_history, .histfile, /etc/gshadow, key extensions (.pem, .key, .p12, .pfx), individual key files (id_ed25519, id_ecdsa, id_dsa)
  2. grep missing from FILE_READ_COMMANDS -- also consider awk, sed, python, curl
  3. check_redirect_target only finds first > or < per segment
  4. .env error message should point to specific IronClaw secrets management commands

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: commit 339977a (fix 3 items from previous review)

All 3 medium findings from my last review are addressed:

Finding Status Notes
Blocking canonicalize() in async context Addressed Documented trade-off with guidance to make async if needed. Pragmatic -- local FS is sub-ms.
Overly broad standalone SSH key patterns Fixed Moved to SENSITIVE_FILENAMES with exact file_name() match. Test confirms grid_rsa_data no longer false-positives.
Duplicate test suites Fixed Removed duplicate is_sensitive_path unit tests from file.rs. Only integration-level execute() test remains.

No new issues in this commit. LGTM.

@j-bloggs

j-bloggs commented Apr 4, 2026

Copy link
Copy Markdown
Contributor Author

Hey! 👋 Just checking in — this one's still good to go (no conflicts with staging). Would appreciate a re-review when you have a moment! @ilblackdragon @serrrfirat

Copy link
Copy Markdown
Collaborator

Reviewing the full worktree relative to origin/main surfaced a few correctness issues outside the narrow PR-vs-staging diff. Posting here for visibility because they are present in this branch checkout, but some may already be inherited from the base branch.

Findings:

  1. High: Legacy conversations will stop accepting approvals after a restart.
    src/agent/thread_ops.rs now treats a missing source_channel as fail-closed, but the migration only adds a nullable column and does not backfill existing rows (migrations/V15__conversation_source_channel.sql). Any pre-migration conversation rehydrated from DB will therefore reject approvals forever, even from its original channel. This needs either a migration backfill (source_channel = channel) or a legacy fallback when the stored value is null.

  2. Medium: Sandbox job restart silently drops the new mcp_servers and max_iterations inputs.
    src/tools/builtin/job.rs parses both fields, but the persisted sandbox job record only stores credential_grants_json (src/history/store.rs). The restart handler then recreates the job with JobCreationParams { credential_grants, ..Default::default() } (src/channels/web/handlers/jobs.rs). A restarted job will therefore run with the default worker iteration cap and without its original MCP filter.

  3. Medium: The new per-job MCP mounting logic points at the wrong source of truth.
    src/orchestrator/job_manager.rs hardcodes /opt/ironclaw/config/worker/mcp-servers.json, but the normal MCP config path is ~/.ironclaw/mcp-servers.json (src/tools/mcp/config.rs), and bootstrap migrates that file into DB anyway (src/bootstrap.rs). The worker image also never creates /opt/ironclaw/config/worker/... (Dockerfile.worker). So with MCP_PER_JOB_ENABLED=true, most installs will mount no MCP config at all, and the feature will not behave as advertised.

Residual risk:

  • Job-level approval allowlists still do not survive DB restore (src/db/libsql/jobs.rs, src/history/store.rs both hydrate approval_context: None).
  • The shell sensitive-path protection is still heuristic by design, not a hard guarantee (src/tools/builtin/shell.rs).
  • I did not get a completed result from cargo test --test tool_approval_context; it spent the turn in compile/link work, so these findings are from code inspection rather than a finished test run.

@j-bloggs

j-bloggs commented Apr 4, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough review @serrrfirat! 🙏

I've looked at all three findings — they're all pre-existing issues inherited from the staging base branch, not introduced by this PR's sensitive-path protection changes:

  1. source_channel backfill — This is from fix(security): block cross-channel approval thread hijacking #1590's V15__conversation_source_channel.sql migration (already merged to staging). Our PR doesn't touch conversations or approvals.
  2. Sandbox mcp_servers/max_iterations persistence — This is from feat(jobs): per-job MCP server filtering and max_iterations cap #1243's per-job MCP feature (already on staging). Our PR only touches ironclaw_safety and file tool paths.
  3. MCP config path mismatch — Same as above, from feat(jobs): per-job MCP server filtering and max_iterations cap #1243.

Happy to file separate issues for these if the maintainers want them tracked. Our diff is limited to crates/ironclaw_safety/src/ (leak detector patterns) and src/tools/builtin/file.rs (path validation calls).

@ilblackdragon
ilblackdragon merged commit e7fa167 into nearai:staging Apr 7, 2026
14 checks passed
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
…arai#1675)

* fix(safety): add credential patterns and sensitive path blocklist

Addresses critical credential leakage found by security testing
(~/.ironclaw/tests/SECURITY_REPORT.md, test ce-02).

Leak detector (crates/ironclaw_safety/src/leak_detector.rs):
- Add OpenRouter API key pattern (sk-or-v1-<hex>)
- Add Anthropic OAuth token pattern (sk-ant-oat<NN>-<base64url>)
- Add Telegram bot token pattern (word-bounded, 8-12 digit bot ID)
- Add Groq API key pattern (gsk_<alphanumeric>)
- 12 new tests with synthetic keys (positive, false-positive, integration)

File tools (src/tools/builtin/file.rs):
- Add sensitive path blocklist to ReadFileTool, WriteFileTool, and
  ApplyPatchTool (defense-in-depth for all file access vectors)
- Blocks: .env (and .env.local/.env.production/etc.), .ssh/, .aws/,
  .netrc, .pgpass, .npmrc, .pypirc, .docker/config.json, .kube/config,
  .git-credentials, .gcloud/, .config/gcloud/, .gnupg/, .vault-token,
  .ironclaw/secrets/
- Allows .env.example, .env.template, .env.sample (safe suffixes)
- Case-insensitive; resolves symlinks via canonicalize() before check
- 8 new tests covering blocking, safe suffixes, .env variants, case

Known gap: shell tool can still `cat ~/.env` — different security
domain (denylist-gated in autonomous mode, user-initiated in
interactive mode). Tracked for follow-up.

Note: new patterns use .unwrap() on Regex::new() matching the
established convention of the 16 existing patterns in this file
(all use // safety: hardcoded literal). Follow-up to address the
existing .unwrap() debt across all patterns.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Update src/tools/builtin/file.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update crates/ironclaw_safety/src/leak_detector.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update crates/ironclaw_safety/src/leak_detector.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update crates/ironclaw_safety/src/leak_detector.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* fix(safety): address review feedback on credential patterns and path blocklist

- Move is_sensitive_path to ironclaw_safety crate for shared use
- Guard ListDirTool with sensitive path checks (including recursive traversal)
- Add missing sensitive paths: ~/.config/gh/hosts.yml, /etc/shadow,
  ~/.terraform.d/credentials.tfrc.json, ~/.azure/
- Add path traversal regression test and ListDirTool blocking test

Addresses review feedback from zmanian, gemini-code-assist, and copilot.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): annotate sensitive dirs as blocked in recursive listing

When ListDirTool's recursive traversal encounters a sensitive directory,
annotate it with [sensitive - access blocked] so users understand why
its contents are suppressed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): integrate is_sensitive_path into shell tool file-access checking

Add defense-in-depth check for shell commands that read sensitive
credential files (cat, head, tail, less, cp, etc.). Extracts file
path arguments from known file-reading commands and checks them
against the shared is_sensitive_path function from ironclaw_safety.

This is best-effort — shell-level bypass via aliases, variable
expansion, or encoding is still possible. Full mitigation requires
filesystem-level sandboxing (seccomp/landlock).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): add output redirection checks, fix segment splitting

Address zmanian's re-review findings:
- Check > and >> redirection targets against is_sensitive_path
  (write-path equivalent of read-path protection)
- Replace single-char & splitting with proper &&/|| aware parser
  to avoid fragmenting double operators
- Document subshell/command-substitution gap (partially covered by
  detect_command_injection upstream)
- Extract helpers: split_shell_segments, check_segment_file_commands,
  check_redirect_target, expand_tilde
- Add tests for output redirection, chained commands, segment splitting

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(safety): address PR nearai#1675 review feedback and absorb nearai#1713 patterns

Absorb nearai#1713's sensitive path patterns into ironclaw_safety crate:
- Add shell history files, SSH key types, /etc/gshadow
- Add sensitive file extensions (.pem, .key, .p12, .pfx, .jks, .keystore)
- Add .dist safe suffix; smart .env matching (excludes .envrc, .environment)
- Directory-level blocking for .aws/, .docker/, .kube/ (not just specific files)
- Trailing-slash matching so bare directory paths trigger detection

Shell tool hardening:
- Strip surrounding quotes from tokens before sensitive path check
- Add grep, awk, sed to FILE_READ_COMMANDS
- Scan ALL redirect operators in a segment, not just the first

Other fixes:
- Add trailing \b to Telegram bot token regex to prevent over-matching
- Update error messages to reference secret_list/secret_create
- Strengthen ListDirTool test with tempfile-based .ssh directory

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(safety): close 4 adversarial bypass vectors in leak detector

- Fix .env suffix check: use exact remainder matching instead of
  ends_with, so .env.production.dist is no longer allowed through
- Detect process substitution <(...) in redirect checks and scan
  inner tokens for sensitive paths
- Add missing /id_rsa to SENSITIVE_PATH_PATTERNS (other SSH key
  types were already present)
- Check --flag=value tokens for sensitive paths instead of skipping
  all tokens starting with -

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Fix 3 items from zmanian re-review on leak detector

1. Document blocking canonicalize() in async context: added comment
   explaining the trade-off (sub-ms on local FS, could block on NFS)
   with guidance to make async if needed.

2. Fix overly broad standalone key patterns: moved id_rsa, id_ed25519,
   id_ecdsa, id_dsa, authorized_keys, known_hosts from substring-based
   SENSITIVE_PATH_PATTERNS to exact filename matching via SENSITIVE_FILENAMES.
   This prevents false positives on paths like /project/grid_rsa_data while
   still blocking /project/test_fixtures/id_rsa.

3. Remove duplicate is_sensitive_path unit tests from file.rs: these
   belong in sensitive_paths.rs which already has comprehensive coverage.
   Kept integration-level tests that exercise tool execute() methods.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
@ironclaw-ci ironclaw-ci Bot mentioned this pull request Apr 10, 2026
JZKK720 pushed a commit to JZKK720/ironclaw that referenced this pull request Apr 13, 2026
…arai#1675)

* fix(safety): add credential patterns and sensitive path blocklist

Addresses critical credential leakage found by security testing
(~/.ironclaw/tests/SECURITY_REPORT.md, test ce-02).

Leak detector (crates/ironclaw_safety/src/leak_detector.rs):
- Add OpenRouter API key pattern (sk-or-v1-<hex>)
- Add Anthropic OAuth token pattern (sk-ant-oat<NN>-<base64url>)
- Add Telegram bot token pattern (word-bounded, 8-12 digit bot ID)
- Add Groq API key pattern (gsk_<alphanumeric>)
- 12 new tests with synthetic keys (positive, false-positive, integration)

File tools (src/tools/builtin/file.rs):
- Add sensitive path blocklist to ReadFileTool, WriteFileTool, and
  ApplyPatchTool (defense-in-depth for all file access vectors)
- Blocks: .env (and .env.local/.env.production/etc.), .ssh/, .aws/,
  .netrc, .pgpass, .npmrc, .pypirc, .docker/config.json, .kube/config,
  .git-credentials, .gcloud/, .config/gcloud/, .gnupg/, .vault-token,
  .ironclaw/secrets/
- Allows .env.example, .env.template, .env.sample (safe suffixes)
- Case-insensitive; resolves symlinks via canonicalize() before check
- 8 new tests covering blocking, safe suffixes, .env variants, case

Known gap: shell tool can still `cat ~/.env` — different security
domain (denylist-gated in autonomous mode, user-initiated in
interactive mode). Tracked for follow-up.

Note: new patterns use .unwrap() on Regex::new() matching the
established convention of the 16 existing patterns in this file
(all use // safety: hardcoded literal). Follow-up to address the
existing .unwrap() debt across all patterns.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Update src/tools/builtin/file.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update crates/ironclaw_safety/src/leak_detector.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update crates/ironclaw_safety/src/leak_detector.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update crates/ironclaw_safety/src/leak_detector.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* fix(safety): address review feedback on credential patterns and path blocklist

- Move is_sensitive_path to ironclaw_safety crate for shared use
- Guard ListDirTool with sensitive path checks (including recursive traversal)
- Add missing sensitive paths: ~/.config/gh/hosts.yml, /etc/shadow,
  ~/.terraform.d/credentials.tfrc.json, ~/.azure/
- Add path traversal regression test and ListDirTool blocking test

Addresses review feedback from zmanian, gemini-code-assist, and copilot.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): annotate sensitive dirs as blocked in recursive listing

When ListDirTool's recursive traversal encounters a sensitive directory,
annotate it with [sensitive - access blocked] so users understand why
its contents are suppressed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): integrate is_sensitive_path into shell tool file-access checking

Add defense-in-depth check for shell commands that read sensitive
credential files (cat, head, tail, less, cp, etc.). Extracts file
path arguments from known file-reading commands and checks them
against the shared is_sensitive_path function from ironclaw_safety.

This is best-effort — shell-level bypass via aliases, variable
expansion, or encoding is still possible. Full mitigation requires
filesystem-level sandboxing (seccomp/landlock).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): add output redirection checks, fix segment splitting

Address zmanian's re-review findings:
- Check > and >> redirection targets against is_sensitive_path
  (write-path equivalent of read-path protection)
- Replace single-char & splitting with proper &&/|| aware parser
  to avoid fragmenting double operators
- Document subshell/command-substitution gap (partially covered by
  detect_command_injection upstream)
- Extract helpers: split_shell_segments, check_segment_file_commands,
  check_redirect_target, expand_tilde
- Add tests for output redirection, chained commands, segment splitting

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(safety): address PR nearai#1675 review feedback and absorb nearai#1713 patterns

Absorb nearai#1713's sensitive path patterns into ironclaw_safety crate:
- Add shell history files, SSH key types, /etc/gshadow
- Add sensitive file extensions (.pem, .key, .p12, .pfx, .jks, .keystore)
- Add .dist safe suffix; smart .env matching (excludes .envrc, .environment)
- Directory-level blocking for .aws/, .docker/, .kube/ (not just specific files)
- Trailing-slash matching so bare directory paths trigger detection

Shell tool hardening:
- Strip surrounding quotes from tokens before sensitive path check
- Add grep, awk, sed to FILE_READ_COMMANDS
- Scan ALL redirect operators in a segment, not just the first

Other fixes:
- Add trailing \b to Telegram bot token regex to prevent over-matching
- Update error messages to reference secret_list/secret_create
- Strengthen ListDirTool test with tempfile-based .ssh directory

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(safety): close 4 adversarial bypass vectors in leak detector

- Fix .env suffix check: use exact remainder matching instead of
  ends_with, so .env.production.dist is no longer allowed through
- Detect process substitution <(...) in redirect checks and scan
  inner tokens for sensitive paths
- Add missing /id_rsa to SENSITIVE_PATH_PATTERNS (other SSH key
  types were already present)
- Check --flag=value tokens for sensitive paths instead of skipping
  all tokens starting with -

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Fix 3 items from zmanian re-review on leak detector

1. Document blocking canonicalize() in async context: added comment
   explaining the trade-off (sub-ms on local FS, could block on NFS)
   with guidance to make async if needed.

2. Fix overly broad standalone key patterns: moved id_rsa, id_ed25519,
   id_ecdsa, id_dsa, authorized_keys, known_hosts from substring-based
   SENSITIVE_PATH_PATTERNS to exact filename matching via SENSITIVE_FILENAMES.
   This prevents false positives on paths like /project/grid_rsa_data while
   still blocking /project/test_fixtures/id_rsa.

3. Remove duplicate is_sensitive_path unit tests from file.rs: these
   belong in sensitive_paths.rs which already has comprehensive coverage.
   Kept integration-level tests that exercise tool execute() methods.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
(cherry picked from commit e7fa167)
@ironclaw-ci ironclaw-ci Bot mentioned this pull request Apr 18, 2026
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…arai#1675)

* fix(safety): add credential patterns and sensitive path blocklist

Addresses critical credential leakage found by security testing
(~/.ironclaw/tests/SECURITY_REPORT.md, test ce-02).

Leak detector (crates/ironclaw_safety/src/leak_detector.rs):
- Add OpenRouter API key pattern (sk-or-v1-<hex>)
- Add Anthropic OAuth token pattern (sk-ant-oat<NN>-<base64url>)
- Add Telegram bot token pattern (word-bounded, 8-12 digit bot ID)
- Add Groq API key pattern (gsk_<alphanumeric>)
- 12 new tests with synthetic keys (positive, false-positive, integration)

File tools (src/tools/builtin/file.rs):
- Add sensitive path blocklist to ReadFileTool, WriteFileTool, and
  ApplyPatchTool (defense-in-depth for all file access vectors)
- Blocks: .env (and .env.local/.env.production/etc.), .ssh/, .aws/,
  .netrc, .pgpass, .npmrc, .pypirc, .docker/config.json, .kube/config,
  .git-credentials, .gcloud/, .config/gcloud/, .gnupg/, .vault-token,
  .ironclaw/secrets/
- Allows .env.example, .env.template, .env.sample (safe suffixes)
- Case-insensitive; resolves symlinks via canonicalize() before check
- 8 new tests covering blocking, safe suffixes, .env variants, case

Known gap: shell tool can still `cat ~/.env` — different security
domain (denylist-gated in autonomous mode, user-initiated in
interactive mode). Tracked for follow-up.

Note: new patterns use .unwrap() on Regex::new() matching the
established convention of the 16 existing patterns in this file
(all use // safety: hardcoded literal). Follow-up to address the
existing .unwrap() debt across all patterns.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Update src/tools/builtin/file.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update crates/ironclaw_safety/src/leak_detector.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update crates/ironclaw_safety/src/leak_detector.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update crates/ironclaw_safety/src/leak_detector.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* fix(safety): address review feedback on credential patterns and path blocklist

- Move is_sensitive_path to ironclaw_safety crate for shared use
- Guard ListDirTool with sensitive path checks (including recursive traversal)
- Add missing sensitive paths: ~/.config/gh/hosts.yml, /etc/shadow,
  ~/.terraform.d/credentials.tfrc.json, ~/.azure/
- Add path traversal regression test and ListDirTool blocking test

Addresses review feedback from zmanian, gemini-code-assist, and copilot.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): annotate sensitive dirs as blocked in recursive listing

When ListDirTool's recursive traversal encounters a sensitive directory,
annotate it with [sensitive - access blocked] so users understand why
its contents are suppressed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): integrate is_sensitive_path into shell tool file-access checking

Add defense-in-depth check for shell commands that read sensitive
credential files (cat, head, tail, less, cp, etc.). Extracts file
path arguments from known file-reading commands and checks them
against the shared is_sensitive_path function from ironclaw_safety.

This is best-effort — shell-level bypass via aliases, variable
expansion, or encoding is still possible. Full mitigation requires
filesystem-level sandboxing (seccomp/landlock).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): add output redirection checks, fix segment splitting

Address zmanian's re-review findings:
- Check > and >> redirection targets against is_sensitive_path
  (write-path equivalent of read-path protection)
- Replace single-char & splitting with proper &&/|| aware parser
  to avoid fragmenting double operators
- Document subshell/command-substitution gap (partially covered by
  detect_command_injection upstream)
- Extract helpers: split_shell_segments, check_segment_file_commands,
  check_redirect_target, expand_tilde
- Add tests for output redirection, chained commands, segment splitting

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(safety): address PR nearai#1675 review feedback and absorb nearai#1713 patterns

Absorb nearai#1713's sensitive path patterns into ironclaw_safety crate:
- Add shell history files, SSH key types, /etc/gshadow
- Add sensitive file extensions (.pem, .key, .p12, .pfx, .jks, .keystore)
- Add .dist safe suffix; smart .env matching (excludes .envrc, .environment)
- Directory-level blocking for .aws/, .docker/, .kube/ (not just specific files)
- Trailing-slash matching so bare directory paths trigger detection

Shell tool hardening:
- Strip surrounding quotes from tokens before sensitive path check
- Add grep, awk, sed to FILE_READ_COMMANDS
- Scan ALL redirect operators in a segment, not just the first

Other fixes:
- Add trailing \b to Telegram bot token regex to prevent over-matching
- Update error messages to reference secret_list/secret_create
- Strengthen ListDirTool test with tempfile-based .ssh directory

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(safety): close 4 adversarial bypass vectors in leak detector

- Fix .env suffix check: use exact remainder matching instead of
  ends_with, so .env.production.dist is no longer allowed through
- Detect process substitution <(...) in redirect checks and scan
  inner tokens for sensitive paths
- Add missing /id_rsa to SENSITIVE_PATH_PATTERNS (other SSH key
  types were already present)
- Check --flag=value tokens for sensitive paths instead of skipping
  all tokens starting with -

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Fix 3 items from zmanian re-review on leak detector

1. Document blocking canonicalize() in async context: added comment
   explaining the trade-off (sub-ms on local FS, could block on NFS)
   with guidance to make async if needed.

2. Fix overly broad standalone key patterns: moved id_rsa, id_ed25519,
   id_ecdsa, id_dsa, authorized_keys, known_hosts from substring-based
   SENSITIVE_PATH_PATTERNS to exact filename matching via SENSITIVE_FILENAMES.
   This prevents false positives on paths like /project/grid_rsa_data while
   still blocking /project/test_fixtures/id_rsa.

3. Remove duplicate is_sensitive_path unit tests from file.rs: these
   belong in sensitive_paths.rs which already has comprehensive coverage.
   Kept integration-level tests that exercise tool execute() methods.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: regular 2-5 merged PRs risk: medium Business logic, config, or moderate-risk modules scope: tool/builtin Built-in tools size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants