Skip to content

watcher(windows): keep one ReadDirectoryChangesW outstanding - #42938

Closed
robobun wants to merge 1 commit into
mainfrom
robobun/8b7d3578/windows-watcher-one-read
Closed

robobun wants to merge 1 commit into
mainfrom
robobun/8b7d3578/windows-watcher-one-read

Conversation

@robobun

@robobun robobun commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • WindowsWatcher::next (src/watcher/WindowsWatcher.rs:296) issues a ReadDirectoryChangesW before every wait, also while an earlier read is still queued. Each cycle leaves one more read queued: 200 watcher cycles add about 100 KiB of nonpaged pool.
  • All reads share one buffer. A read that finds recorded changes fills it inside the call. A later completion fills it when its packet is dequeued. So an older packet overwrites newer records, and the watcher never reports them.
  • The lost change is observed on Windows debug builds only, as hangs in "should hot reload when a file is renamed() into place" and "should work with sourcemap loading". On release builds it is modelled, not observed.

Fix

Background

  • One watcher thread reads the changes under the project root with overlapped ReadDirectoryChangesW calls. An I/O completion port queues one packet per completed read.
  • The watch loop waits for the first packet, then polls with a zero timeout for more. A poll that times out leaves its read queued.
  • Between two reads, Windows records the changes for the next read.
Notes

Where this came from. A Windows debug build never finishes test/cli/hot/hot.test.ts (its timeout is Infinity on debug builds). In "should work with sourcemap loading" the --hot side is fine. The bun build --watch child never reports the write to bundle_in.ts, so the outfile is never rebuilt. bun build --watch and --watch restart on every change, so each generation has a new watcher with one read, which is the state that loses a change. Log of the child that hung (BUN_DEBUG_watcher=1, read numbers added):

read #1
event: hot-runner-root.js Modified        packet of read #1
read #2                                   poll, times out, read #2 stays queued
calling onFileUpdate                      about 20 ms on a debug build
                                          meanwhile: the outfile's second change completes read #2,
                                          then the test writes bundle_in.ts twice, with no read queued
read #3                                   completes inside the call: buffer = [hot-runner-root.js, bundle_in.ts]
event: hot-runner-root.js Modified        packet of read #2 is dequeued first and overwrites the buffer
read #4
event: hot-runner-root.js Modified        packet of read #3: the buffer still holds the records of read #2
read #5, read #6                          nothing else arrives, the test waits forever

The renamed() test loses the Removed record of the entry point the same way. BUN_WATCHER_TRACE of its --hot child shows {dir: []} {dir: [write]} and nothing more without the fix, and {dir: [delete, rename, write], entry: [delete, write]} with it. The same analysis is in a comment on #40017.

Counts. Windows Server 2019 x64, 16 vCPUs, no extra load. "Without the fix" is a debug build of main, or the 1.4.3 canary where named. The hot.test.ts rows are from main at f937cf4, the rows of the new test from main at c6b7fcb, which this branch is now based on.

without the fix with the fix
new test, 1.4.3 canary against a release build of this branch fails 8 of 8 (+101 KiB) passes 10 of 10 (+0)
new test, debug build fails 8 of 8 (+98 to +115 KiB) passes 10 of 10 (+0)
hot.test.ts, debug build hangs after 7 tests 12 pass, 47 s (50 s on c6b7fcb)
"should work with sourcemap loading" alone, debug build hangs in 12 of 12 runs hangs in 0 of 25
"renamed() into place" alone, debug build hangs in 4 of 4 runs hangs in 0 of 9

Nonpaged pool of the --hot process over 200 watcher cycles (Get-Process, NonpagedSystemMemorySize64): canary 7,240 to 104,520 bytes, release build with the fix 7,240 to 7,240, debug build with the fix 11,456 to 11,456. The counter does not move for the first 10 to 25 queued reads, then it grows by 500 to 1,000 bytes for each one. So a short run proves nothing: 30 reloads measured +0 bytes on the canary, 100 reloads +43 to +57 KiB.

Probe. A 100 line program with raw kernel32 calls and bun's setup: an overlapped directory handle on a completion port, one buffer, one OVERLAPPED. One read is queued. File A is written, then file B (a write is two records). Then a second read, then two dequeues:

buffer before read #2: <empty>                         (read #1 completed, packet not dequeued yet)
buffer right after read #2 returned: A + B + B         (filled inside the call)
dequeue 1 (packet of read #1): nbytes=48  buffer: A    (filled at dequeue)
dequeue 2 (packet of read #2): nbytes=120 buffer: A    (stale, B is lost)
Source of the first probe (rustc, no crates)
// Probe 2: a ReadDirectoryChangesW that finds changes already recorded completes inside the call
// and writes its records into the caller's buffer at once. If an older completion packet is still
// queued on the port, dequeuing that packet copies ITS records over the buffer. The newer packet
// then finds stale records.
#![allow(non_snake_case, non_camel_case_types, clippy::all)]
use std::ffi::c_void;
use std::os::windows::ffi::OsStrExt;
use std::ptr;

type HANDLE = *mut c_void;
type DWORD = u32;
type BOOL = i32;

#[repr(C)]
struct OVERLAPPED { Internal: usize, InternalHigh: usize, Offset: u32, OffsetHigh: u32, hEvent: HANDLE }

#[link(name = "kernel32")]
unsafe extern "system" {
    fn CreateFileW(name: *const u16, access: DWORD, share: DWORD, sa: *mut c_void, disp: DWORD, flags: DWORD, templ: HANDLE) -> HANDLE;
    fn CreateIoCompletionPort(file: HANDLE, existing: HANDLE, key: usize, threads: DWORD) -> HANDLE;
    fn ReadDirectoryChangesW(dir: HANDLE, buf: *mut c_void, len: DWORD, subtree: BOOL, filter: DWORD, ret: *mut DWORD, ov: *mut OVERLAPPED, cb: *mut c_void) -> BOOL;
    fn GetQueuedCompletionStatus(port: HANDLE, nbytes: *mut DWORD, key: *mut usize, ov: *mut *mut OVERLAPPED, ms: DWORD) -> BOOL;
    fn GetLastError() -> DWORD;
}
const FILTER: DWORD = 0x1 | 0x2 | 0x10 | 0x40;

#[repr(C, align(8))]
struct Buf([u8; 64 * 1024]);

fn records(b: &[u8]) -> String {
    let mut off = 0usize;
    let mut out = Vec::new();
    loop {
        let next = u32::from_ne_bytes(b[off..off + 4].try_into().unwrap()) as usize;
        let action = u32::from_ne_bytes(b[off + 4..off + 8].try_into().unwrap());
        let len = u32::from_ne_bytes(b[off + 8..off + 12].try_into().unwrap()) as usize;
        if len == 0 { out.push("<empty>".to_string()); break; }
        let name: Vec<u16> = b[off + 12..off + 12 + len.min(200)].chunks_exact(2).map(|c| u16::from_ne_bytes([c[0], c[1]])).collect();
        out.push(format!("{}(action {})", String::from_utf16_lossy(&name), action));
        if next == 0 || out.len() > 8 { break; }
        off += next;
    }
    out.join(" + ")
}

fn main() {
    let root = std::env::temp_dir().join(format!("rdcw-probe2-{}", std::process::id()));
    std::fs::create_dir_all(&root).unwrap();
    let root = std::fs::canonicalize(&root).unwrap();
    let a = root.join("hot-runner-root.js");
    let b = root.join("bundle_in.ts");
    std::fs::write(&a, "0").unwrap();
    std::fs::write(&b, "0").unwrap();

    unsafe {
        let w: Vec<u16> = root.as_os_str().encode_wide().chain(Some(0)).collect();
        let dir = CreateFileW(w.as_ptr(), 1, 7, ptr::null_mut(), 3, 0x0200_0000 | 0x4000_0000, ptr::null_mut());
        assert!(dir as isize != -1);
        let iocp = CreateIoCompletionPort(dir, ptr::null_mut(), 0, 1);
        let mut ov: Box<OVERLAPPED> = Box::new(std::mem::zeroed());
        let mut buf: Box<Buf> = Box::new(Buf([0; 64 * 1024]));
        let mut arm = |buf: &mut Buf, ov: &mut OVERLAPPED, label: &str| {
            let ok = ReadDirectoryChangesW(dir, buf.0.as_mut_ptr().cast(), buf.0.len() as u32, 1, FILTER, ptr::null_mut(), ov, ptr::null_mut());
            println!("{label}: ReadDirectoryChangesW ok={ok} lasterr={}", GetLastError());
        };
        let mut dequeue = |buf: &Buf, label: &str| {
            let mut nbytes: DWORD = 0; let mut key = 0usize; let mut pov: *mut OVERLAPPED = ptr::null_mut();
            let rc = GetQueuedCompletionStatus(iocp, &mut nbytes, &mut key, &mut pov, 500);
            if rc == 0 && pov.is_null() { println!("{label}: no packet (lasterr={})", GetLastError()); return; }
            println!("{label}: packet rc={rc} nbytes={nbytes} -> buffer holds: {}", records(&buf.0));
        };

        arm(&mut buf, &mut ov, "read #1 (nothing recorded yet, stays pending)");
        std::fs::write(&a, "1").unwrap();
        println!("wrote hot-runner-root.js: its first change completes read #1, the packet waits on the port");
        std::fs::write(&b, "1").unwrap();
        println!("wrote bundle_in.ts: no read is outstanding, so the kernel records it for the next read");
        std::thread::sleep(std::time::Duration::from_millis(50));
        println!("buffer before read #2: {}", records(&buf.0));
        arm(&mut buf, &mut ov, "read #2 (changes are recorded, completes inside the call)");
        println!("buffer right after read #2 returned, nothing dequeued yet: {}", records(&buf.0));
        dequeue(&buf, "dequeue 1 (packet of read #1)");
        dequeue(&buf, "dequeue 2 (packet of read #2)");
        dequeue(&buf, "dequeue 3");
    }
    let _ = std::fs::remove_dir_all(&root);
}

A second probe runs bun's loop shape against bursts of writes, with a new watcher for each trial, and counts the trials in which the one write to the watched file is never reported. "Dispatch" is a busy wait after each batch that stands in for on_file_update:

dispatch read before every wait one read outstanding
20 us 0 / 200 0 / 200
100 us 2 / 200 0 / 200
1 ms 176 / 200 0 / 200
5 ms 176 / 200 0 / 200
Source of the second probe. The table rows are rdcw_burst3.exe every|once 200 2 4 200 <dispatch_us>
// Burst probe 2: like bun build --watch, every trial uses a FRESH watcher (new directory handle, one
// read outstanding). Noise files are rewritten by several threads at once and the watched file Y is
// written once during the burst. Count the trials in which the watcher loop never reports Y.
// usage: rdcw_burst2.exe <every|once> <trials> <writer_threads> <noise_writes_per_thread> <y_delay_us>
#![allow(non_snake_case, non_camel_case_types, clippy::all)]
use std::ffi::c_void;
use std::os::windows::ffi::OsStrExt;
use std::ptr;
use std::sync::atomic::{AtomicU64, AtomicUsize, Ordering};
use std::sync::Arc;
use std::time::{Duration, Instant};

type HANDLE = *mut c_void;
type DWORD = u32;
type BOOL = i32;

#[repr(C)]
struct OVERLAPPED { Internal: usize, InternalHigh: usize, Offset: u32, OffsetHigh: u32, hEvent: HANDLE }

#[link(name = "kernel32")]
unsafe extern "system" {
    fn CreateFileW(name: *const u16, access: DWORD, share: DWORD, sa: *mut c_void, disp: DWORD, flags: DWORD, templ: HANDLE) -> HANDLE;
    fn CreateIoCompletionPort(file: HANDLE, existing: HANDLE, key: usize, threads: DWORD) -> HANDLE;
    fn ReadDirectoryChangesW(dir: HANDLE, buf: *mut c_void, len: DWORD, subtree: BOOL, filter: DWORD, ret: *mut DWORD, ov: *mut OVERLAPPED, cb: *mut c_void) -> BOOL;
    fn GetQueuedCompletionStatus(port: HANDLE, nbytes: *mut DWORD, key: *mut usize, ov: *mut *mut OVERLAPPED, ms: DWORD) -> BOOL;
    fn GetLastError() -> DWORD;
    fn CloseHandle(h: HANDLE) -> BOOL;
}
const FILTER: DWORD = 0x1 | 0x2 | 0x10 | 0x40;

#[repr(C, align(8))]
struct Buf([u8; 64 * 1024]);

struct Shared { seen_y: AtomicUsize, records: AtomicU64, packets: AtomicU64, zero: AtomicU64, issued: AtomicU64 }

fn start_watcher(arm_every: bool, dispatch_us: u64, shared: Arc<Shared>, root: std::path::PathBuf) -> usize {
    let (tx, rx) = std::sync::mpsc::channel::<usize>();
    std::thread::spawn(move || unsafe {
        let w: Vec<u16> = root.as_os_str().encode_wide().chain(Some(0)).collect();
        let dir = CreateFileW(w.as_ptr(), 1, 7, ptr::null_mut(), 3, 0x0200_0000 | 0x4000_0000, ptr::null_mut());
        assert!(dir as isize != -1);
        let iocp = CreateIoCompletionPort(dir, ptr::null_mut(), 0, 1);
        let mut ov: Box<OVERLAPPED> = Box::new(std::mem::zeroed());
        let mut buf: Box<Buf> = Box::new(Buf([0; 64 * 1024]));
        let mut pending = false;
        let mut told = false;
        loop {
            let mut timeout: DWORD = 0xFFFF_FFFF;
            loop {
                if arm_every || !pending {
                    let ok = ReadDirectoryChangesW(dir, buf.0.as_mut_ptr().cast(), buf.0.len() as u32, 1, FILTER, ptr::null_mut(), &mut *ov, ptr::null_mut());
                    if ok == 0 { CloseHandle(iocp); return; }
                    pending = true;
                    shared.issued.fetch_add(1, Ordering::Relaxed);
                }
                if !told { told = true; tx.send(dir as usize).unwrap(); }
                let mut nbytes: DWORD = 0; let mut key = 0usize; let mut pov: *mut OVERLAPPED = ptr::null_mut();
                let rc = GetQueuedCompletionStatus(iocp, &mut nbytes, &mut key, &mut pov, timeout);
                if rc == 0 {
                    let err = GetLastError();
                    if pov.is_null() && (err == 258 || err == 1460) { break; }
                    CloseHandle(iocp);
                    return;
                }
                pending = false;
                shared.packets.fetch_add(1, Ordering::Relaxed);
                if nbytes == 0 { shared.zero.fetch_add(1, Ordering::Relaxed); continue; }
                let b = &buf.0;
                let mut off = 0usize;
                loop {
                    let next = u32::from_ne_bytes(b[off..off + 4].try_into().unwrap()) as usize;
                    let len = u32::from_ne_bytes(b[off + 8..off + 12].try_into().unwrap()) as usize;
                    let name: Vec<u16> = b[off + 12..off + 12 + len].chunks_exact(2).map(|c| u16::from_ne_bytes([c[0], c[1]])).collect();
                    shared.records.fetch_add(1, Ordering::Relaxed);
                    if String::from_utf16_lossy(&name) == "entry.js" { shared.seen_y.fetch_add(1, Ordering::SeqCst); }
                    if next == 0 { break; }
                    off += next;
                }
                timeout = 0;
            }
            let until = Instant::now() + Duration::from_micros(dispatch_us);
            while Instant::now() < until { std::hint::spin_loop(); }
        }
    });
    rx.recv().unwrap()
}

fn main() {
    let args: Vec<String> = std::env::args().collect();
    let arm_every = args.get(1).map(|s| s == "every").unwrap_or(true);
    let trials: usize = args.get(2).and_then(|s| s.parse().ok()).unwrap_or(50);
    let writers: usize = args.get(3).and_then(|s| s.parse().ok()).unwrap_or(4);
    let per_thread: usize = args.get(4).and_then(|s| s.parse().ok()).unwrap_or(16);
    let y_delay: u64 = args.get(5).and_then(|s| s.parse().ok()).unwrap_or(300);
    let dispatch_us: u64 = args.get(6).and_then(|s| s.parse().ok()).unwrap_or(0);

    let root = std::env::temp_dir().join(format!("rdcw-burst2-{}", std::process::id()));
    std::fs::create_dir_all(&root).unwrap();
    let root = std::fs::canonicalize(&root).unwrap();
    let y = root.join("entry.js");
    std::fs::write(&y, "0").unwrap();
    for t in 0..writers { for n in 0..per_thread { std::fs::write(root.join(format!("noise-{t}-{n}.txt")), "0").unwrap(); } }

    let shared = Arc::new(Shared { seen_y: AtomicUsize::new(0), records: AtomicU64::new(0), packets: AtomicU64::new(0), zero: AtomicU64::new(0), issued: AtomicU64::new(0) });

    let mut lost = 0;
    for trial in 0..trials {
        let dir_handle = start_watcher(arm_every, dispatch_us, shared.clone(), root.clone());
        std::thread::sleep(Duration::from_millis(5));
        let before = shared.seen_y.load(Ordering::SeqCst);
        let barrier = Arc::new(std::sync::Barrier::new(writers + 1));
        let mut hs = Vec::new();
        for t in 0..writers {
            let root = root.clone();
            let barrier = barrier.clone();
            hs.push(std::thread::spawn(move || {
                barrier.wait();
                for n in 0..per_thread { std::fs::write(root.join(format!("noise-{t}-{n}.txt")), format!("{trial}")).unwrap(); }
            }));
        }
        barrier.wait();
        let until = Instant::now() + Duration::from_micros(y_delay);
        while Instant::now() < until { std::hint::spin_loop(); }
        std::fs::write(&y, format!("{trial}")).unwrap();
        for h in hs { h.join().unwrap(); }
        let deadline = Instant::now() + Duration::from_millis(1000);
        while shared.seen_y.load(Ordering::SeqCst) == before && Instant::now() < deadline { std::thread::sleep(Duration::from_millis(1)); }
        if shared.seen_y.load(Ordering::SeqCst) == before { lost += 1; }
        unsafe { CloseHandle(dir_handle as HANDLE); }
        std::thread::sleep(Duration::from_millis(10));
    }
    println!("mode={} trials={} writers={} noise/thread={} y_delay={}us dispatch={}us => Y never reported in {} trials; reads issued={} packets={} zero-byte packets={} records parsed={}",
        if arm_every { "every" } else { "once" }, trials, writers, per_thread, y_delay, dispatch_us, lost,
        shared.issued.load(Ordering::Relaxed), shared.packets.load(Ordering::Relaxed), shared.zero.load(Ordering::Relaxed), shared.records.load(Ordering::Relaxed));
    std::process::exit(0);
}

This is the model for release builds. I could not make the 1.4.3 canary lose a change: about 2,000 generations of --watch with bursts of up to 64 writes around the watched write, and 130 more with 3,000 to 10,000 modules in the watchlist, all restarted. A long lived process is also protected by the defect itself: each cycle leaves one more read queued, so a burst rarely finds the handle without a read.

The test. bun --hot entry.js with BUN_WATCHER_TRACE. A write of a file that nothing imports is one watcher cycle: the watcher traces a batch for the directory and reloads nothing. The test makes 200 such writes and waits for the trace to grow after each one, which takes about 17 ms in total. It reads NonpagedSystemMemorySize64 of the child before and after from one PowerShell process (test/bundler/compile-windows-metadata.test.ts also asks PowerShell). Windows charges each queued read to that quota, so the old loop adds about 100 KiB and the new one nothing. The limit is 16 KiB. The whole test takes 1.0 to 1.4 s on a release build and on a debug build, most of it the start of PowerShell, and it has no timeout of its own. Before the first sample the test writes until the trace shows a batch, at most 20 times: a new watcher records nothing before its thread issues the first read, which is the separate problem of #40017. Two earlier drafts are gone. The first checked the lost change directly (a write of the watched file inside a burst of other writes, 10 of 10 lost on a debug build), but the release lanes of CI pass it with or without the fix. The second counted the pool over 100 reloads with two PowerShell runs, which review found too slow.

Not changed. When Windows reports that its own buffer overflowed (nbytes == 0), the watcher still drops those changes and tells nobody. That is #42936.

Overlap. Five open PRs carry an equivalent gate inside a larger change: #30644, #35596, #39488, #39820 and #40017 (the hunk in #40017 is the same code). None of them has merged in four weeks to four months. This PR is the gate alone with the cause and a test, and is meant to land first. #40017 then keeps only the start-up arming and the stop drain, and #39820 can drop its pending flag commit.

Suites run with the fix. Windows debug and Windows release: test/cli/hot/hot.test.ts, test/cli/hot/watch-many-dirs.test.ts, test/cli/hot/watch.test.ts, test/cli/watch/watch.test.ts, test/cli/watch/watcher-trace.test.ts. Windows debug only: test/bake/dev/hot.test.ts, test/bake/deinitialization.test.ts, test/js/bun/resolve/bun-main-entry-point.test.ts, test/cli/test/test-filter-lifecycle-snapshot.test.ts, test/js/bun/http/bun-serve-html-hot-reload-drop.test.ts. Linux debug: test/cli/hot/watch-many-dirs.test.ts (the new test is skipped). cargo check -p bun_watcher --target x86_64-pc-windows-msvc passes.


no test proof · iteration 1 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/cli/hot/watch-many-dirs.test.ts

@coderabbitai

coderabbitai Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 1f7d201f-8aeb-4816-94c6-cdeefb7e13b9

📥 Commits

Reviewing files that changed from the base of the PR and between c6b7fcb and a9a3b37.

📒 Files selected for processing (2)
  • src/watcher/WindowsWatcher.rs
  • test/cli/hot/watch-many-dirs.test.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 3 remain after this review.


Walkthrough

Changes

Windows watcher lifecycle

Layer / File(s) Summary
Pending read control
src/watcher/WindowsWatcher.rs
WindowsWatcher tracks pending directory reads. Guarded arming prevents overlapping reads, and completion dequeue clears the pending state. Overflow recovery uses the guarded path.
Windows regression coverage
test/cli/hot/watch-many-dirs.test.ts
A Windows-only hot-watch test performs 200 directory-write cycles and checks that nonpaged pool growth remains below 16 KiB.

Suggested reviewers: jarred-sumner

Priority: ⬇️ Low

Merge Risk: ⚪ Minimal · up to a9a3b

No merge-blocking issue was identified in the Windows watcher lifecycle or its regression coverage.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the primary Windows watcher change: keeping one ReadDirectoryChangesW operation outstanding.
Description check ✅ Passed The description clearly explains the problem, fix, background, test coverage, measured results, and verification commands. It does not use the exact template headings, but it provides the required inf…

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: closed in favor of #44344. Its first change is the same gate.

How to reproduce, on Windows x64:

# release build (1.4.3 canary): the new test fails 8 of 8 runs, +101 KiB of nonpaged pool over 200 watcher cycles
bun test test/cli/hot/watch-many-dirs.test.ts -t "one directory read"

# debug build of main: both hang, 12 of 12 and 4 of 4 runs. Bound the run from outside, the file's timeout is Infinity on debug builds.
bun bd test test/cli/hot/hot.test.ts -t "sourcemap loading$"
bun bd test test/cli/hot/hot.test.ts -t "into place"

With this change the new test passes 10 of 10 runs on a release build and on a debug build (+0 bytes), and the debug build finishes test/cli/hot/hot.test.ts (12 of 12).

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline findings (both on the new test), I also checked the Rust side of the change: read_pending is cleared only when the dequeued overlapped is the watcher's own, so a foreign/spurious completion at src/watcher/WindowsWatcher.rs:353 correctly leaves the flag set (the real read is still queued), and the rc == 0 error and nbytes == 0 re-arm paths both keep the flag consistent with what is on the port. Closing the directory handle in stop() while a read is outstanding is unchanged from before the PR (the old code had more reads queued, not fewer), so it is not introduced here.

Extended reasoning...

The native change is a single read_pending flag with three touch points (arm(), the post-GetQueuedCompletionStatus clear keyed on pointer identity, and the two prepare() call sites replaced by arm()). I traced each exit of next(): timeout returns leave the flag set (read still queued, which is the whole fix); the matching-overlapped branch clears it before either the error return or the event return; the non-matching-overlapped continue leaves it set, which is correct since a foreign packet does not consume the watcher's read. The stop()-with-pending-read concern is pre-existing behaviour and is not made worse by this PR. The two confirmed findings are on the test's runtime budget and its undeadlined setup loop; nothing further on the Rust side needed to be added.

Comment thread test/cli/hot/watch-many-dirs.test.ts Outdated
Comment thread test/cli/hot/watch-many-dirs.test.ts Outdated
@robobun

robobun commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 8:50 AM PT - Sep 16th, 2026

✅ @robobun, your commit a9a3b370799436101986c920ad4ef644c48696e3 passed in Build #116583! 🎉


🧪   To try this PR locally:

bunx bun-pr 42938

That installs a local version of the PR into your bun-42938 executable, so you can run:

bun-42938 --bun

@robobun
robobun force-pushed the robobun/8b7d3578/windows-watcher-one-read branch from f1f574d to 272d0df Compare September 16, 2026 15:23
Comment thread src/watcher/WindowsWatcher.rs Outdated
Comment thread src/watcher/WindowsWatcher.rs Outdated
WindowsWatcher::next issued a new read before every wait, also while
the read of an earlier, timed out wait was still queued. Every watcher
cycle left one more read queued on the directory handle.

All reads share one buffer. A read that finds changes already recorded
completes inside the call and fills the buffer at once. A read that
completes later fills it only when its packet is dequeued. With two
reads in flight, the dequeue of the older packet overwrote the records
of the newer read. The watcher parsed the older records twice and never
reported the newer changes.

Issue a read only when none is outstanding.
@robobun
robobun force-pushed the robobun/8b7d3578/windows-watcher-one-read branch from 272d0df to a9a3b37 Compare September 16, 2026 15:27

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed the updated push and found no bugs; since this changes the Windows watcher's I/O-completion handshake and the test can only run on Windows, a human look on the Windows CI result would still be worthwhile.

What was reviewed:

  • read_pending bookkeeping in WindowsWatcher::next: checked GetQueuedCompletionStatus out-param semantics — lpOverlapped is NULL on timeout and set to the failed OVERLAPPED on an aborted read, so clearing before the rc == 0 branch is correct; stop() is the only lifecycle exit and is terminal (no re-init), so no stale-flag path.
  • Spurious-packet continue and the nbytes==0 re-arm path keep the flag consistent with what is actually queued.
  • Test: the setup loop is now bounded and the explicit timeout is gone. Not verifiable from Linux: whether 200 watcher round-trips plus PowerShell startup fit the default per-test timeout on the Windows debug lane (the PR description still cites 100 reloads, so the numbers there look stale relative to the code).
Extended reasoning...

Overview

The diff (+104/-6) touches src/watcher/WindowsWatcher.rs and test/cli/hot/watch-many-dirs.test.ts. The Rust change adds a read_pending: bool to WindowsWatcher, an arm() helper that only calls DirWatcher::prepare() (the ReadDirectoryChangesW issue) when no read is outstanding, and clears the flag in next() whenever GetQueuedCompletionStatus hands back the watcher's own OVERLAPPED. Both former prepare() call sites (top of next() and the nbytes==0 overflow re-arm) go through arm(). The test adds a Windows-only case that drives ~200 watcher cycles by writing an unimported file and compares the process's nonpaged-pool counter before and after via a single long-lived PowerShell.

Security risks

None identified. The change does not touch input parsing, paths, or privileges; it only changes when an already-issued kernel read is re-issued. The test spawns PowerShell with a fixed command string and the child's PID interpolated, which is a number.

Level of scrutiny

Moderate. The logic is small but sits on a kernel-level handshake where a wrong assumption about out-param semantics would silently re-introduce the double-read or, worse, never re-arm. I checked: on timeout GetQueuedCompletionStatus sets *lpOverlapped to NULL (so the equality check does not fire and the flag stays true, which matches reality — the read is still queued); on a failed I/O such as ERROR_OPERATION_ABORTED after stop() closes the directory handle it sets the failed OVERLAPPED, so the flag is cleared and Err is returned. stop() has a single caller in Watcher::thread_body on the error exit and nothing re-initializes the same WindowsWatcher afterward, so there is no lifecycle path where a stale read_pending = true would suppress a needed prepare(). The spurious-packet continue leaves the flag untouched, which is correct because the watcher's own packet is still on the port.

Other factors

The exit reason was dry_streak with no findings. The prior review's two test concerns (unbounded setup loop, explicit 30 s timeout / large workload) are addressed in this push: the setup loop is capped at 20 attempts of 100 ms with a named failure, and the explicit timeout is gone. The workload changed shape rather than shrinking — 200 cycles of a non-reloading write instead of ~100 full reloads — and the single PowerShell now answers both samples. I cannot execute the test from this Linux checkout, so whether it stays under the default per-test timeout on the Windows debug lane is something the CI run has to settle; the PR description's "100 reloads" figures appear to predate this version of the test. I am not approving outright because the correctness of the fix is observable only on Windows and the claimed measurements cannot be reproduced here.

@robobun

robobun commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Closing in favor of #44344. Its first change (read_pending) is the gate of this PR: next() starts a ReadDirectoryChangesW only when no read is outstanding.

If #44344 does not land, reopen this PR.

@robobun robobun closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant