Skip to content

shell: glob a demoted brace group as literal text - #39634

Closed
robobun wants to merge 1 commit into
mainfrom
farm/96bd8276/shell-literal-brace-glob
Closed

robobun wants to merge 1 commit into
mainfrom
farm/96bd8276/shell-literal-brace-glob

Conversation

@robobun

@robobun robobun commented Aug 19, 2026

Copy link
Copy Markdown
Collaborator

Problem

  • With files {x}.a.txt and x.a.txt, Bun.$\echo {x}..txt`printsx.a.txt. bash prints {x}.a.txt. {x},.txt, {a,{x}}..txtand{a,b` fail the same way.
  • Since shell: treat comma-less brace groups as literal instead of truncating #34856 the brace lexer (src/shell_parser/braces.rs) treats a {...} group with no comma as text. The glob matcher still reads every {...} as a group, and neutralize_glob_metachars (src/runtime/shell/states/Expansion.rs:402) passes every template brace byte through to it.

Fix

  • The brace lexer returns LexerOutput::kept_as_syntax: one bool per unescaped {, ,, }, read off the tokens after demotion and rollback.
  • retain_brace_syntax drops the demoted brace bytes from meta_offsets before the word goes to the walker. A word with no brace hint drops all of them. The neutralizer then wraps them as [{], [,], [}], as it already does for interpolated braces.
  • Correct because the shell decided what the braces mean one step earlier. The walker must receive that decision. Kept groups pass through unchanged. Every new fixture matches bash 5.2.
  • Verified: test/js/bun/shell/brace.test.ts (9 new tests, 8 fail on main) and a unit test in braces.rs. Also bunshell.test.ts, parse.test.ts, lex.test.ts, miri, clippy.

Background

  • Expansion.rs expands one shell word. meta_offsets lists the bytes written by template *, **, {, ,, } atoms. Only these may act as pattern syntax. neutralize_glob_metachars wraps every other metacharacter in a [c] class.
  • The parser sets the brace hint only when a word has {, } and ,. Without it the word goes to the walker directly.
  • With it, do_brace_expand runs the brace lexer, pushes the variants, and also hands the original pattern to the walker, which expands the groups again (shell: pathname-expand each brace variant instead of appending the patterns #33423 changes that part). This PR edits that pattern.
Notes

Before and after, with the fixtures from the tests:

word main this PR (= bash 5.2)
echo {x}.*.txt x.a.txt {x}.a.txt {x}.b.txt
echo *.{x} a.x a.{x}
echo {x}/*.txt x/a.txt {x}/a.txt
echo {a,b* no matches found {a,b1.txt
echo {a,b}x{c* no matches matches ax{c1 bx{c2
echo {x},*.txt matches x,a.txt matches {x},a.txt
echo {a,{x}}.*.txt matches a.1.txt x.1.txt matches a.1.txt {x}.1.txt

The last three words also emit the un-expanded variants as extra argv words on main. That is the pre-existing behavior #33423 removes, so those tests assert on the matches only. When #33423 lands, its test "a comma-less brace group still globs a literal *" needs {x}.a.txt style fixtures.

Alignment between the verdicts and meta_offsets: do_brace_expand backslash-escapes every {, ,, } and \ that is not in meta_offsets, so the unescaped brace characters the lexer sees are exactly the recorded ones, in order. retain_brace_syntax zips the two and debug_asserts that both run out together. The {a,b},*.txt test guards the pairing: the stray comma is dropped, the group is kept.

, and } outside any group were already literal in the matcher, so dropping them changes nothing. An unclosed { used to make the matcher fail the whole word. Now it is literal, as in bash.

#32902 takes the other route and changes Bun.Glob itself. This PR keeps the shell correct either way: if #32902 lands, [{] and { in a walker pattern both mean a literal brace.

Ran: bun bd test on test/js/bun/shell/brace.test.ts, parse.test.ts, lex.test.ts, bunshell.test.ts, the glob-using commands/ tests (three ls failures are environmental: root user, no registry access) and test/internal/source-lints/. cargo test -p bun_shell_parser, bun run rust:miri -p bun_shell_parser, cargo clippy --no-deps on bun_shell_parser and bun_runtime. All new fixtures were also run through bash 5.2.37.

The brace lexer now reports, for each unescaped `{`, `,` and `}`, whether
it survived as brace syntax. Before a word reaches the glob walker, the
shell drops the brace bytes the lexer demoted to text from `meta_offsets`,
so `neutralize_glob_metachars` wraps them like any other data byte. A word
with no brace hint at all drops every brace byte.

`echo {x}.*.txt` matched `x.a.txt` because the glob matcher reads every
`{...}` as a group. It now matches `{x}.a.txt`, as bash does.
@robobun

robobun commented Aug 19, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: closed, superseded by #32902 (matcher fix, covers every case here) and #33423 (removes the shell path the lexer change feeds). Details in the comment below.

Reproduced on bun 1.4.0 and on main:

mkdir t && cd t && touch '{x}.a.txt' 'x.a.txt'
bun -e 'await Bun.$`echo {x}.*.txt`'   # prints x.a.txt, bash prints {x}.a.txt

@coderabbitai

coderabbitai Bot commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

Important

Review skipped

This review includes 3 billable files. This on-demand review is free during your promotion.

Your included review limit has been reached. Run @coderabbitai review --use-credits to review the latest changes using usage credits.

  • Run review — free
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 334af8fa-97b0-4171-81b2-ca523b69b4a5

📥 Commits

Reviewing files that changed from the base of the PR and between a356964 and 6759e13.

📒 Files selected for processing (3)
  • src/runtime/shell/states/Expansion.rs
  • src/shell_parser/braces.rs
  • test/js/bun/shell/brace.test.ts

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.


Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. The alignment invariant between kept_as_syntax and the brace bytes in meta_offsets checks out, and the test coverage is thorough. Because it's a correctness change to the shell's brace/glob interaction with a subtle ordering invariant, a human look would still be worthwhile.

What was reviewed:

  • Traced the kept_as_syntax ↔ meta_offsets zip: do_brace_expand's escaping loop leaves exactly the meta-offset brace bytes unescaped, and meta_offsets is monotone, so the pairing is one-to-one (guarded by the debug_assert).
  • Verified kept_as_syntax is read off tokens after demotion and rollback_braces but before flatten_tokens reorders indices.
  • Confirmed the no-brace-hint path (retain_brace_syntax(me, &[])) is sound: brace_expansion_hint at parse.rs:1671 requires {, } and ,, so any word the lexer would expand already takes the do_brace_expand path.
Extended reasoning...

Overview

The PR fixes a bash-compat bug where a comma-less brace group ({x}) that the brace lexer demotes to literal text was still handed to the glob walker as an active brace group, so {x}.* matched x.a instead of {x}.a. Three files change: src/shell_parser/braces.rs gains a kept_as_syntax: Vec<bool> on LexerOutput (one verdict per unescaped {/,/} in input order), src/runtime/shell/states/Expansion.rs gains retain_brace_syntax which drops the demoted entries from meta_offsets before neutralize_glob_metachars runs, and test/js/bun/shell/brace.test.ts gains 9 integration tests plus a 12-case Rust unit test for the new field.

Security risks

None. This is a correctness fix that narrows what the glob walker treats as pattern syntax; it cannot broaden a match. The neutralization path ([c] wrapping) is pre-existing and only receives more bytes to wrap.

Level of scrutiny

Medium-high. The fix rests on an ordering invariant: the sequence of brace bytes in meta_offsets (offsets into current_out) must line up one-to-one with the lexer's kept_as_syntax (verdicts over the escaped string). I traced this through: expand_simple_no_io pushes meta_offsets in append order (monotone), the tilde prefix adjusts them uniformly, and do_brace_expand's escaping loop backslash-escapes every {/,/}/\\ not at a meta offset — so the unescaped brace bytes the lexer sees are exactly the meta-offset brace bytes, in order. The * entries in meta_offsets are skipped by the match in retain_brace_syntax and don't consume a verdict. The debug_assert!(kept.next().is_none()) catches any drift.

I also checked the lexer side: brace_chars records Some(tok_idx) for each syntax token and None for immediate text, then re-reads the token variant after rollback_braces (which uses in-place replace_token_with_string, so indices stay valid) but before flatten_tokens (which swaps tokens). Walked the {a,{b}}, {a,{b,c},d and {a\\,b} cases by hand and they match the unit-test expectations. The }-with-empty-stack and stray-, arms correctly push None and fall through to append_char (no continue), preserving pre-existing token output.

For the no-brace-hint path (has_glob_expansion() without has_brace_expansion()), retain_brace_syntax(me, &[]) drops every recorded brace byte. This is safe because brace_expansion_hint at parse.rs:1671 is has_brace_open && has_brace_close && has_comma — any word the brace lexer could expand sets the hint, so on this path no brace can be legitimate glob-walker syntax.

Other factors

The tests are strong: each fixture creates both the {x}.* and x.* filenames so a regression to the old behavior fails observably, the {a,b},* case guards the zip pairing specifically, and the Rust unit test covers rollback, nesting, escaped chars and multi-byte text on both encodings. The PR description documents the interaction with #33423 and #32902. This is well-executed but not mechanical — the invariant spans two subsystems and someone who owns the shell expansion pipeline should confirm it.

@robobun

robobun commented Aug 19, 2026

Copy link
Copy Markdown
Collaborator Author

Closing after self-review. The bug is real, but this is the wrong place to fix it, and the change has a regression of its own.

Wrong layer:

Regression (verified with this branch):

  • This PR wraps a template , or } as [,] or [}]. That turns a path segment such as a,b or c}d from a literal segment into a pattern segment. The walker does not follow symlinks for pattern segments, so with a,b -> real, echo a,b/*.txt prints a,b/f1.txt on main and fails with no matches found on this branch. bash prints a,b/f1.txt. No test here covered that.

The nine shell cases in test/js/bun/shell/brace.test.ts at 6759e13 are still useful as shell level regression coverage for #32902. I left a note there.

@robobun

robobun commented Aug 19, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded by #32902 and #33423, see the comment above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant