Repository navigation
Conversation
The WebLoader, the legacy HTMLChunker and the factory html chunker turned HTML into text with a chain of global regex replacements (script, style, comment, block end, generic tag). Each removal could join the text on either side of it into a new tag for the next replacement, the end-tag match missed `</script >`, and the lazy bodies rescanned the whole rest of the input from every unclosed `<script`, `<style` or `<!--`, which is quadratic. One shared scanner (utils/htmlText.ts) now reads the input once, left to right. Text is only copied from outside removed spans, script and style end tags may carry whitespace or attributes, and every search either consumes what it scans or is answered from an index computed up front. The llms-txt build script gets the same treatment for its tag strip. This is text conversion, not output sanitization. Entity decoding is unchanged and still happens exactly once, after markup removal.
✅ Single Commit Policy - COMPLIANTStatus: Policy requirements met • 1 commit • Valid format • Ready for merge 📊 View validation details📝 Commit Details
✅ Validation Results
🤖 Automated validation by NeuroLink Single Commit Enforcement |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info
📝 Walkthrough
Merge Risk | ⚪ Minimal · up to
|
Documentation Validation Results🚀 Documentation validation passed!
📦 Build artifact uploaded successfully. Ready for deployment preview. Commit: |
Summary
Replaces the chained regex replacements that turn HTML into text in three RAG paths, and in one docs build script, with a single left-to-right scanner. CodeQL alerts addressed: #157-#160 and #166 (
src/lib/rag/chunking/htmlChunker.ts), #161-#164 (src/lib/rag/document/loaders.ts), #150 and #156 (src/lib/rag/chunkers/HTMLChunker.ts), #206 (docs-site/scripts/build-llms-txt.ts). Open alerts #151 (htmlChunker.ts) and #152 (loaders.ts), bothjs/bad-tag-filteron the same</script>expressions, were not in my list; they are on the code this removes.This is text conversion, not output sanitization. The result is plain text for retrieval and chunking. It is not meant to be written back into a page, and the existing entity decoding still runs afterwards and is allowed to turn
<b>into<b>. Nothing here makes the output safe to render as HTML.Whether the alerts close is decided by CodeQL after this PR runs, not by me. CodeQL was not available locally. What I checked is that no
.replace(/<...>/...)tag chain is left in the four flagged files (grep), not what CodeQL will say about the new code.What was wrong (behaviour, not only the static-analysis shape)
Each caller chained global replacements (script, style, comment, block end, tag). That had three real defects:
<scr<!---->ipt>.</script >and</SCRIPT\t>, so a script body could be indexed as page text.[\s\S]*?bodies rescan the rest of the input from every unclosed<script,<styleor<!--, which is quadratic. In the reversal run below the old web loader took 36,138 ms and 22,027 ms on two 400 KB documents.What changes for a caller
No public API or type change;
removeHtmlMarkupinsrc/lib/utils/htmlText.tsis imported by the three callers only and is not exported from any entry point (grep), sodocs/apiis not regenerated. The WebLoader, the legacyHTMLChunker(rag/chunking/htmlChunker.ts) and thecreateChunker("html")chunker (rag/chunkers/HTMLChunker.ts) call it with their own replacement text and line-break tag lists, so each keeps its previous spacing.The scanner reads the input once. Text is copied only from outside removed spans, so a removal cannot assemble a tag, and the output never matches
<[^>]+>. A<survives only when no>follows it, or as the empty<>.Behaviour that differs, all on malformed or hostile input:
</script >,</SCRIPT\t>and</style x="y">now close their element.<scripty>, is an ordinary tag, not a script element (<scripty>x</script>ynow givesxy, wasy).a <scr<!---->ipt>BODY</script> zgivesa ipt>BODY z, wasa BODY z.createChunker("html")chunker removes a whole comment even when it holds>(<!-- a > b -->); it used to leave residue such as-->behind.<that is not a tag now ends at the first>after it. The loader used to turn</p>into a newline first and then let the bare<run on to a later>:5 < 6 is true</p>after<b>x</b>gave5 x, now5 afterx. Both lose the<...>span; the new one keeps more text.docs-site/scripts/build-llms-txt.tsgets a localremoveTagSpanswith the same result asreplace(/<[^>]+>/g, ""). Its output is unchanged. The remaining JSX-component regexes in that function (<Tabs,<TabItem,<[A-Z]...) are untouched.Tests
New suite
test/continuous-test-suite-rag-html-markup.ts(10 cases), wired aspnpm run test:rag-html-markupand into the credential-free CI step. It takes everything fromdist/index.js: WebLoader on an owned local HTTP document, the legacyHTMLChunker, andcreateChunker("html"). Every case runs all three paths and names the paths that are wrong; assertion messages carry path names only, never converted text.>.Per CLAUDE.md rule 15 there is no import from
src/lib/.test/continuous-test-suite-rag-entity-decoding.tsis unchanged.What I ran
All on the committed tree; machine load average was between about 40 and 240 during these runs.
pnpm run test:rag-html-markup: 10 of 10 pass, 0.08 to 0.19 s.pnpm exec tsx test/continuous-test-suite-rag-entity-decoding.ts: 13 of 13 pass.dist/for the pre-change sources (parent commit, transpiled as ES2022 modules), ran the suite, then restored them from backups and verified sha256 on all three (all matched). This is a swap of compiled modules, not a rebuild. Result: exit 1, 0 passed, 10 failed, each reported as a failure, not a skip. Per path: assembled script pair, spaced end tags, unclosed-then-closed, name boundary and the ordinary page failed on web loader, legacy chunker andcreateChunker("html"); the assembled comment case on web loader and legacy chunker; the comment holding>oncreateChunker("html")only; two hostile cases exceeded 2,000 ms on all three paths. The third hostile case (unclosed<script>start tags) failed withfetch failedwhile the old conversion blocked the process; I did not isolate that cause. After restoring, 10 of 10 pass again. Run time with the old code was 104.9 s.<br>,</p>), each through three option sets (900,000 conversions): 0 outputs contained a<...>span. With empty replacements, on the same 300,000 strings, 0 outputs were changed by a second pass and 0 were not a subsequence of the input. A further 300,000 strings from a tag-only alphabet (no-,!, so no comments) matchedreplace(/<[^>]+>/g, "")exactly.indexOf,lastIndexOf,slice,startsWith,toLowerCase,charCodeAt,charAt) and ran 41 hostile families (unclosed or mismatched script, style and comment openers,</scripty,<brplus spaces, assembled tags, long names, lone<and>) under two option sets, with the repeat count doubling from 12,500 to 200,000. Work per input character was constant across the sizes in every row, the highest being 9.00 (the<>family), and the largest growth for one doubling of the input was 2.00. The regex whitespace loop in the<brprobe is not counted, but it only steps over characters the removed tag then consumes. This is the evidence for linear time; it does not depend on machine load.<...>span, the slowest conversion was 8.2 s (the<script>x n plus<style>x n document, about 12 million characters, on the loaded machine). Timings were noisy under load; one row's 16x ratio read 109 because its 50,000 point was a warm-up outlier (698 ms at 100,000 and 4,054 ms at 800,000 for the same row), so I did not use wall-clock ratios as proof.>inside) gave 0 differences on all three paths. Comments holding>: 0 of 20,000 differ on the loader and legacy chunker, 9,209 of 20,000 differ oncreateChunker("html")(the residue case above). Fragments with a bare<: 1,497, 859 and 750 of 20,000 differ (loader, legacy, factory); I read the example above, not every case. 100,000 malformed mixes differ in 16,112, 9,333 and 10,563 cases, as intended; I read one sample, not all; no output contained a<...>span in the bare-<and malformed sets.removeTagSpansagainst the regex: 0 differences over all 4,450 Markdown files underdocs/(14,196,910 characters) and over 300,000 random strings.check(svelte-check 0 errors,tsc --noEmit --strict) andvalidate:all. svelte-check printed a config load error forlanding/svelte.config.js(missing@sveltejs/adapter-vercelin this worktree) and reported 0 errors.Not covered
pnpm test, the rest of the suites, andtest/continuous-test-suite-rag.ts(it importsdotenv/config, and I was told not to read.env). Hosted CI has not run.pnpm run check:docs-apiwas not run by me; no exported type changed, and the pre-push hook runs it.WebLoader.extractMainContent, the legacy chunker's tag splitting) are untouched; I did not measure them and no open alert points at them.build-llms-txt.ts, because its output is identical; the evidence is the corpus and fuzz comparison above. This touchesdocs-site/, so the non-required Docs-site Artifacts check will run;search-index.jsonis not produced by this script and I did not regenerate it.test:unitaggregate (one line per file was the brief); it runs in the CI step.htmlToMarkdown.tsandmarkupSniff.tsare unchanged.Summary by CodeRabbit
Bug Fixes
<>sequences when removing markup from generated text.Tests