Skip to content

fix: make maxPrefixPenalty actually cap the prefix penalty [patch] - #83

Merged
matt-edmondson merged 1 commit into
mainfrom
fix/prefix-penalty-cap
Sep 22, 2026
Merged

matt-edmondson merged 1 commit into
mainfrom
fix/prefix-penalty-cap

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #82

The bug

The unmatched prefix was charged twice, and only one of the two charges was capped:

  1. PenalizeNonPatternCharacters applies Math.Max(n * unmatchedPrefixLetterPenalty, maxPrefixPenalty) when the first pattern codepoint matches — correctly capped at -5.
  2. The generic unmatchedLetterPenalty in the scoring loop had already charged every one of those same prefix codepoints, uncapped.

So the cap bounded one source while the other kept running. Matching pattern "y" against ever-longer junk prefixes:

subject before after
y 10 10
xy -2 -1
xxxy -6 -3
xxxxxy — -5
xxxxxxy -11 -5
xxxxxxxxxxxxy -17 -5
x×100 + y -106 -5

The score now falls 1 per prefix character and flattens at the documented -5, instead of falling without bound. (The 10 at zero prefix is the separator bonus a match at the start of the string earns — a separate effect, not prefix penalty.)

For real use — ranking matches in long strings or file paths — a match behind a longer irrelevant prefix is now deprioritized rather than punished indefinitely, which is what the constant's name, its doc comment, and PenalizeNonPatternCharacters' own Math.Max all intended.

The fix

This takes the first of the two options the issue offers: make PenalizeNonPatternCharacters the sole, capped source of prefix cost.

Rather than skipping the generic penalty in the prefix region outright, the loop tracks what it charged before the first match (prefixPenaltyCharged) and refunds it at the moment that match lands, just before the capped penalty is applied. This keeps one behaviour that a plain skip would have lost: a subject that never matches the pattern at all keeps its per-codepoint penalties, so its score still reflects how much was skipped, as Contains' outScore documents ("the score reflects how close the match was").

PenalizeNonPatternCharacters itself is unchanged — it was always correct in isolation, which is why its own unit tests never caught this. The double-counting was in the caller.

The doc comment on maxPrefixPenalty now states what the cap is for.

Tests

Four new tests in FuzzyTests.cs, exercising the full CalculateScore pipeline rather than PenalizeNonPatternCharacters in isolation:

  • CalculateScore_LongUnmatchedPrefix_PenaltyFlattensAtTheCap — prefixes of 5, 12 and 100 characters all score identically. This is the issue's acceptance criterion.
  • CalculateScore_PrefixPenalty_NeverExceedsTheDocumentedCap — across prefix lengths 1-20, the score never drops below maxPrefixPenalty.
  • CalculateScore_LongerPrefix_NeverScoresHigherThanAShorterOne — the penalty is monotonic, so flattening never becomes a reward.
  • CalculateScore_NonMatchingSubject_StillReflectsHowMuchWasSkipped — pins the behaviour the refund deliberately preserves.

Verified by reverting only the Fuzzy.cs change and re-running: the two cap tests fail, 49 pass. With the fix: 51 passed, 0 failed — including all 49 pre-existing tests, untouched.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FyZyutu7Xna2FUQK6o8KAC


Generated by Claude Code

The prefix had two independent penalty sources. PenalizeNonPatternCharacters
applied a capped cost when the first pattern codepoint matched, but the
generic per-codepoint unmatchedLetterPenalty in the scoring loop had
already charged every prefix codepoint, uncapped, and the two stacked.
maxPrefixPenalty bounded only one of them, so the score for pattern "y"
kept falling roughly 1 per extra prefix character without limit rather
than flattening at -5.

Track what the loop charged before the first match and refund it when
that match lands, leaving PenalizeNonPatternCharacters as the sole,
capped source of prefix cost. A subject that never matches keeps its
per-codepoint penalties, so its score still reflects how much was
skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FyZyutu7Xna2FUQK6o8KAC
@sonarqubecloud

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Documented prefix-penalty cap doesn't actually cap anything — unbounded score drop for long prefixes

2 participants