Repository navigation
Conversation
WalkthroughA control-flow change in the Zig shell lexer prevents digits that are part of an in-progress word from being parsed as file-descriptor redirections. New regression tests validate parsing of command names with trailing digits versus standalone FD redirects. Changes
🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Comment |
|
Updated 8:53 PM PT - May 1st, 2026
❌ @robobun, your commit 58599ca has 3 failures in
DetailsDetails🧪 To try this PR locally: bunx bun-pr 27257That installs a local version of the PR into your bun-27257 --bun |
|
Newest first ✅ 0a2a7 — Looks good! Reviewed 2 files across Previous reviews✅ 08afc — Looks good! Reviewed 2 files across Previous reviews✅ d9f72 — Looks good! Reviewed 2 files across |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@test/regression/issue/12602.test.ts`:
- Around line 15-20: Capture and assert the exit code of the setup chmod command
and add one negative redirect case: assign the chmod invocation to a variable
(e.g., chmodRes = await $`chmod +x ${dir}/script1`.quiet()) and assert
chmodRes.exitCode === 0; keep the existing successful run (result = await $`cd
${dir} && ./script1<input.txt`.quiet()) and its output assertion, then add a
failing redirect test by running the command against a missing file with
.nothrow() (e.g., bad = await $`cd ${dir} && ./script1<missing.txt`.nothrow())
and assert bad.exitCode !== 0 to cover the error path.
The lexer treated any digit 0-2 followed by > or < as an fd-prefix redirect, even when the digit was part of a larger word. This caused `echo abc1>file` to lex as `echo abc` + `1>file` (writing "abc\n" instead of "abc1\n"), and `echo test100>out` to lex as `echo test10` + `0>out`. POSIX only recognizes an IO number when the digits form a separate word. Add a word-boundary check before entering the fd-redirect path: the digit must not extend an in-progress text fragment (word_start==j) and must not immediately follow a word-continuation token (Text, Var, quoted text, CmdSubstEnd, etc.). Digits that begin a word, including after operators/semicolons or at start of input, are still parsed as fd prefixes.
0a2a7ca to
2934757
Compare
|
Addressed review feedback in 7ac36fe + 58599ca:
All shell tests pass on every platform (lex 31/31, parse 17/17, 12602 3/3 on darwin and 1+2skip on Windows). CI reds are unrelated flakes also hitting other open PRs: |
…redirect - Treat a digit following `**` as part of the compound word, same as `*`, so `**2>file` lexes as glob `**2` + `>file` instead of `**` + `2>file`. - In eat_redirect's '<' arm, eat the '<' before checking for a second one so `N<file` no longer sets append=true and `N<<file` lexes as one token.
…sterisk Whitespace after `**` didn't emit a Delimit (unlike `*`), so adding `.DoubleAsterisk` to the fd-prefix guard made `echo ** 2>file` glue the `2` onto the glob word and parse `>` as a stdout redirect instead of `2>` as stderr. Move `.DoubleAsterisk` into the `=> true` arm of break_word_impl so it delimits like `.Asterisk` does — this also keeps `echo ** foo` from gluing into a single compound word. Also drop the unnecessary Windows skip on the `cat 0<input.txt` regression test since it only uses shell builtins.
What does this PR do?
Fixes the shell lexer treating a digit at the end of a word as an fd-prefix when followed by
>/<.Repro
POSIX only recognizes an IO number when the digits form their own word. bash writes
abc1\n,test100\n,abc1\n, and runs./script1respectively.Root cause
In the
'0'...'9'arm of the lexer (src/shell/shell.zig),eat_redirect()was attempted unconditionally whenever a digit appeared in Normal state, thenbreak_word()was called after the redirect was accepted — silently splitting the trailing digit off whatever word preceded it.Fix
Before entering the fd-redirect path, require that the digit begins a new word:
self.word_start == self.j— no text currently being accumulated, andText,Var,SingleQuotedText,DoubleQuotedText,CmdSubstEnd,Asterisk, brace tokens) — same classificationbreak_word_implalready uses.Digits that genuinely begin a word — after whitespace (
echo 2>file), at start of input (2>file), or after an operator (echo foo;2>file) — still parse as fd prefixes.How did you verify your code works?
USE_SYSTEM_BUN=1 bun test test/js/bun/shell/lex.test.ts -t "op_redirect digit"→ failbun bd test test/js/bun/shell/lex.test.ts→ 30/30 passbun bd test test/js/bun/shell/parse.test.ts→ 17/17 passbun bd test test/js/bun/shell/file-io.test.ts→ 25/25 passbun bd test test/js/bun/shell/bunshell.test.ts -t redirect→ 28/28 passbun bd test test/regression/issue/12602.test.ts→ 3/3 passCloses #12602