Repository navigation
fix(file_tools): refuse plain-text writes that corrupt binary documents - #86362
Merged
kshitijk4poor merged 3 commits intoAug 14, 2026
Merged
kshitijk4poor merged 3 commits into
kshitijk4poor merged 3 commits into
Conversation
Port from nearai/ironclaw#7109: read_file auto-extracts .docx/.xlsx/.pptx (and PDF via anydoc) to readable text, so a model plausibly believes it holds the file's contents and writes the edited text back with write_file/patch — silently destroying the document container. Proven live on main: write_file over a valid .docx left a non-zip corpse, and a text write over an existing .pdf clobbered the %PDF header. - tools/binary_extensions.py: OPAQUE_DOCUMENT_EXTENSIONS + has_opaque_document_extension() + is_pdf_path() (pure string checks) - tools/file_tools.py: _check_binary_document_write() — opaque container formats (doc/docx/xls/xlsx/ppt/pptx/odt/ods/odp) always rejected; .pdf rejected only when overwriting an existing regular file (new-PDF creation stays allowed, matching the upstream split guard). Wired into write_file_tool and patch_tool (replace + V4A Update/Add headers; Delete/Move skip the guard since they write no text). - tests/tools/test_binary_document_write_guard.py: guard unit tests + end-to-end write_file/patch coverage incl. bytes-untouched assertions.
OPAQUE_DOCUMENT_EXTENSIONS was missing 10 extensions that read_file auto-extracts via anydoc: .docm, .xlsm, .xlsb, .pptm, .ppsx, .ppsm, .pps, .pot, .rtf, .epub. Each has the same corruption path: read_file shows extracted text, model writes it back, container is destroyed. Flagged by @egilewski on PR NousResearch#82818 — proven live for .docm (text write left a non-zip corpse). Added bytes-untouched regression test for .docm.
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
kshitijk4poor
enabled auto-merge (rebase)
August 14, 2026 21:16
Lazy import inside _check_binary_document_write was unnecessary — binary_extensions is a leaf module already imported at line 15. Hoisted has_opaque_document_extension and is_pdf_path to the existing module-level import. /simplify-code finding.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
write_file/patchnow refuse plain-text writes that would corrupt binary documents — always for opaque container formats (.doc/.docx/.docm/.xls/.xlsx/.xlsm/.xlsb/.ppt/.pps/.pot/.pptx/.pptm/.ppsx/.ppsm/.odt/.ods/.odp/.rtf/.epub), and for.pdfonly when overwriting an existing file (new-PDF creation stays allowed).Port of nearai/ironclaw#7109. Based on #82818 by @teknium1.
Root cause:
read_fileauto-extracts these formats to readable text, so a model plausibly believes it holds the file's contents and writes the edited text back — silently destroying the document container. Proven live on current main before fixing:write_file_toolover a valid.docxleft a non-zip corpse, and a text write over an existing.pdfclobbered the%PDFheader.Changes
tools/binary_extensions.py:OPAQUE_DOCUMENT_EXTENSIONS(19 extensions covering all anydoc-extracted container formats) +has_opaque_document_extension()+is_pdf_path()(pure string checks, no I/O)tools/file_tools.py:_check_binary_document_write()gate, wired intowrite_file_toolandpatch_tool(replace mode + V4AUpdate/Addheaders;Delete/Moveskip the guard since they write no text)tests/tools/test_binary_document_write_guard.py: guard unit tests + end-to-end write_file/patch coverage with bytes-untouched assertionsSalvage notes
Cherry-picked from #82818 by @teknium1 (authorship preserved). Follow-up commit adds 10 missing extensions (
.docm,.xlsm,.xlsb,.pptm,.ppsx,.ppsm,.pps,.pot,.rtf,.epub) thatread_fileauto-extracts via anydoc but were absent from the originalOPAQUE_DOCUMENT_EXTENSIONSset. Flagged by @egilewski — proven live for.docm(text write left a non-zip corpse). Added bytes-untouched regression test for.docm.Validation
write_fileover existing valid.docxwrite_fileover existing.docmwrite_fileto newreport.docxwrite_fileover existing.pdf%PDFheader clobberedwrite_fileto NEW.pdfpatch(replace + V4A Update) on.docx.docxTargeted tests: 17/17 passed. E2E probe with real imports + isolated
HERMES_HOMEconfirmed all scenarios.Closes #82818