Skip to content

Improve saved skill search discoverability - #126

Merged
kentcdodds merged 2 commits into
mainfrom
cursor/skill-search-discoverability-9a44
Apr 1, 2026
Merged

kentcdodds merged 2 commits into
mainfrom
cursor/skill-search-discoverability-9a44

Conversation

@kentcdodds

@kentcdodds kentcdodds commented Mar 31, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • improve saved-skill search relevance by indexing canonical skill names alongside titles and descriptions
  • strengthen saved-skill lexical ranking and cross-entity tie-breaking for natural-language queries
  • add regressions covering launch-cursor-cloud-agent discoverability without skill_collection

Testing

  • npm run test -- packages/worker/src/mcp/skills/skill-embed-and-flags.node.test.ts packages/worker/src/mcp/capabilities/unified-search.workers.test.ts
  • npm run test:mcp -- packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts
  • Validation artifact: /opt/cursor/artifacts/search-ranking-validation.txt
Open in Web Open in Cursor 

Summary by CodeRabbit

  • Improvements

    • Enhanced unified search ranking with phrase normalization, per-entity lexical scoring, and refined tie-breaking for more relevant results
    • Improved cross-entity result ordering and bounded candidate selection for consistent rankings
    • Skill names are now included in embed text and factored into search relevance
  • Tests

    • Added unit and e2e tests validating strong skill matching across unified search
    • Updated skill embed text validation tests

Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
@coderabbitai

coderabbitai Bot commented Mar 31, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Introduces normalized phrase matching and per-entity lexical scoring into unified search; precomputes per-skill lexical scores; enforces a bounded candidateLimit for fetches; reworks unified ranking to combine fused RRF scores, unified lexical scores, and entity scores. Skill embed text now optionally includes a normalized skill name and a humanized variant.

Changes

Cohort / File(s) Summary
Unified Search Core
packages/worker/src/mcp/capabilities/unified-search.ts
Added phrase normalization and lexical scoring helpers; precompute per-skill lexicalScoreById in searchSkillsForUser; introduce bounded candidateLimit (clamped 25–100) for fetches; replace prior sort with a custom comparator that orders by fused cross-entity RRF, unified lexical score (entity-specific), entity fused score, then key; added getEntityScore and getUnifiedLexicalScore helpers.
Skill Embed / Persistence
packages/worker/src/mcp/skills/skill-embed-and-flags.ts, packages/worker/src/mcp/skills/skill-mutation.ts
buildSkillEmbedText gained optional `skillName?: string
Tests / E2E
packages/worker/src/mcp/capabilities/unified-search.workers.test.ts, packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts, packages/worker/src/mcp/skills/skill-embed-and-flags.node.test.ts
Added unit test ensuring skill name/description matches survive cross-entity ranking; added e2e test that saved skills surface without collection filters; updated embed test to assert new skill name lines.

Sequence Diagram

sequenceDiagram
    participant Client
    participant UnifiedSearch as Unified Search
    participant SkillSearch as Skill Search
    participant Embedding as Skill Embedding
    participant Ranker as Multi‑Signal Ranker

    Client->>UnifiedSearch: searchUnified(query, limit)
    UnifiedSearch->>Ranker: normalizeSearchPhrase(query)
    UnifiedSearch->>SkillSearch: fetch top candidates (candidateLimit)
    SkillSearch->>Embedding: buildSkillEmbedText(skillName?, title, desc, ...)
    Embedding-->>SkillSearch: embed text (includes normalized/humanized name)
    SkillSearch->>Ranker: compute lexicalScoreById (lexical matcher + phrase bonuses)
    Ranker->>Ranker: compute fused RRF scores (cross-entity)
    Ranker->>Ranker: compute unified lexical score per entity type
    Ranker-->>UnifiedSearch: ranked skill results (by fused → lexical → vector)
    UnifiedSearch->>Ranker: unify results from skills/capabilities/secrets/ui_artifacts
    Ranker->>Ranker: final sort by fused RRF → unified lexical → entity score → key
    Ranker-->>UnifiedSearch: unified ranked list
    UnifiedSearch-->>Client: return final results (bounded by original limit)
Loading

Estimated Code Review Effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly Related PRs

Poem

🐰 I hopped through phrases, trimmed and neat,
I nudged each score where matches meet,
A normalized name, a human line,
Candidates rank, the best align,
Hooray — the search now hops to beat!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Improve saved skill search discoverability' directly summarizes the main change: enhancing how saved skills are discovered through search by indexing skill names and strengthening lexical ranking.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch cursor/skill-search-discoverability-9a44

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@kentcdodds
kentcdodds marked this pull request as ready for review March 31, 2026 23:57
@github-actions

github-actions Bot commented Mar 31, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Preview deployed: https://kody-pr-126.kentcdodds.workers.dev

Worker: kody-pr-126
D1: kody-pr-126-db
KV: kody-pr-126-oauth-kv

Mocks:

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
packages/worker/src/mcp/capabilities/unified-search.ts (1)

548-584: Consider caching the normalized query to avoid repeated computation.

normalizeSearchPhrase(input.query) is called multiple times per skill hit within getUnifiedLexicalScore (lines 564, 566, 568-570). This could be computed once and passed as a parameter or cached at the outer scope.

🔧 Cache normalized query at outer scope
+	const normalizedQuery = normalizeSearchPhrase(input.query)
 	function getUnifiedLexicalScore(key: string): number {
 		if (key.startsWith('c:')) {
 			const hit = capByName.get(key.slice(2))
 			return hit ? scoreCapabilityLexicalMatch(input.query, hit) : 0
 		}
 		if (key.startsWith('s:')) {
 			const hit = skillByName.get(key.slice(2))
 			if (!hit) return 0
 			return (
 				lexicalScore(input.query, [
 					hit.skillName,
 					hit.title,
 					hit.description,
 					hit.collection ?? '',
 					hit.keywords.join(' '),
 				].join('\n')) +
-				scoreSkillPhraseMatch(normalizeSearchPhrase(input.query), hit.skillName) *
+				scoreSkillPhraseMatch(normalizedQuery, hit.skillName) *
 					2 +
-				scoreSkillPhraseMatch(normalizeSearchPhrase(input.query), hit.title) *
+				scoreSkillPhraseMatch(normalizedQuery, hit.title) *
 					1.5 +
 				scoreSkillPhraseMatch(
-					normalizeSearchPhrase(input.query),
+					normalizedQuery,
 					hit.description,
 				) *
 					1.25
 			)
 		}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/worker/src/mcp/capabilities/unified-search.ts` around lines 548 -
584, In getUnifiedLexicalScore compute normalizeSearchPhrase(input.query) once
and reuse it instead of calling normalizeSearchPhrase(input.query) multiple
times; e.g., create a local const (e.g., normalizedQuery) in the scope where
getUnifiedLexicalScore can access input.query (or at the start of
getUnifiedLexicalScore) and pass that to scoreSkillPhraseMatch calls for
hit.skillName, hit.title, and hit.description, leaving lexicalScore(...) as-is
to avoid repeated normalization work.
packages/worker/src/mcp/skills/skill-embed-and-flags.ts (1)

42-50: Redundant inner conditional for humanizedSkillName.

When normalizedSkillName is truthy (line 45 condition is met), humanizedSkillName will always be a non-empty string since it's derived from normalizedSkillName via .replace(/[-_]+/g, ' '). The inner conditional on line 48 is always true in this context.

🔧 Simplify by removing redundant conditional
 	const baseParts = [
 		...(normalizedSkillName
 			? [
 					`name ${normalizedSkillName}`,
-					...(humanizedSkillName ? [humanizedSkillName] : []),
+					humanizedSkillName,
 				]
 			: []),
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/worker/src/mcp/skills/skill-embed-and-flags.ts` around lines 42 -
50, The inner conditional checking humanizedSkillName inside the baseParts
spread is redundant because humanizedSkillName is derived from
normalizedSkillName and will be non-empty whenever normalizedSkillName is
truthy; update the baseParts construction (around normalizedSkillName,
humanizedSkillName, baseParts) to always include humanizedSkillName in the array
when normalizedSkillName is present instead of the nested conditional,
simplifying the branch and removing the unnecessary check.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@packages/worker/src/mcp/capabilities/unified-search.ts`:
- Around line 548-584: In getUnifiedLexicalScore compute
normalizeSearchPhrase(input.query) once and reuse it instead of calling
normalizeSearchPhrase(input.query) multiple times; e.g., create a local const
(e.g., normalizedQuery) in the scope where getUnifiedLexicalScore can access
input.query (or at the start of getUnifiedLexicalScore) and pass that to
scoreSkillPhraseMatch calls for hit.skillName, hit.title, and hit.description,
leaving lexicalScore(...) as-is to avoid repeated normalization work.

In `@packages/worker/src/mcp/skills/skill-embed-and-flags.ts`:
- Around line 42-50: The inner conditional checking humanizedSkillName inside
the baseParts spread is redundant because humanizedSkillName is derived from
normalizedSkillName and will be non-empty whenever normalizedSkillName is
truthy; update the baseParts construction (around normalizedSkillName,
humanizedSkillName, baseParts) to always include humanizedSkillName in the array
when normalizedSkillName is present instead of the nested conditional,
simplifying the branch and removing the unnecessary check.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 9ad0f958-713b-495b-bc9a-648d6207b0a7

📥 Commits

Reviewing files that changed from the base of the PR and between 74e1e37 and 5169596.

📒 Files selected for processing (6)
  • packages/worker/src/mcp/capabilities/unified-search.ts
  • packages/worker/src/mcp/capabilities/unified-search.workers.test.ts
  • packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts
  • packages/worker/src/mcp/skills/skill-embed-and-flags.node.test.ts
  • packages/worker/src/mcp/skills/skill-embed-and-flags.ts
  • packages/worker/src/mcp/skills/skill-mutation.ts

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

Bugbot Autofix prepared a fix for the issue found in the latest run.

  • ✅ Fixed: Lexical scores recomputed inside sort comparator instead of precomputed
    • Cached unified lexical scores in a map before sorting to avoid repeated expensive recomputation during comparisons.

You can send follow-ups to this agent here.

Comment thread packages/worker/src/mcp/capabilities/unified-search.ts
Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
packages/worker/src/mcp/capabilities/unified-search.ts (1)

548-584: Consider extracting shared lexical scoring logic to reduce duplication.

The skill scoring logic in getUnifiedLexicalScore (lines 553-573) duplicates the weighted phrase-match pattern from scoreSkillLexicalMatch. Since they operate on different types (SkillSearchHit vs McpSkillRow), some duplication is inevitable, but extracting a shared helper that accepts the relevant fields could improve maintainability.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/worker/src/mcp/capabilities/unified-search.ts` around lines 548 -
584, getUnifiedLexicalScore duplicates the weighted phrase-match logic used in
scoreSkillLexicalMatch for skills; extract a small shared helper (e.g.,
buildSkillTextScore or scoreSkillLike) that takes the search text (input.query)
and a normalized/adapter object or explicit fields (skillName, title,
description, collection, keywords) so both getUnifiedLexicalScore and
scoreSkillLexicalMatch can call it; use existing helpers lexicalScore,
scoreSkillPhraseMatch, and normalizeSearchPhrase inside the new helper and
update getUnifiedLexicalScore to call it when handling 's:' keys (operating on
SkillSearchHit) and update scoreSkillLexicalMatch to call it for McpSkillRow,
mapping fields as needed.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@packages/worker/src/mcp/capabilities/unified-search.ts`:
- Around line 105-118: The doc string construction in
scoreUiArtifactLexicalMatch can include literal "undefined" when hit.runtime,
hit.description, or parameter.description are missing; update the joins to skip
falsy parts (e.g., filter(Boolean)) so only defined strings are concatenated and
also filter/map parameterText entries (or default missing parameter.description
to '') before joining; keep the existing behavior of scoring but ensure the
array passed to join for doc and parameterText excludes undefined/null values to
avoid injecting "undefined" into the document used by lexicalScore.
- Around line 84-94: In scoreCapabilityLexicalMatch, the doc string is built by
joining [hit.name, hit.domain, hit.description] which can introduce the literal
"undefined" when hit.domain or hit.description are undefined; update the
construction of doc inside scoreCapabilityLexicalMatch to only include
defined/nonnull fields (e.g., build an array [hit.name, hit.domain,
hit.description] and filter out undefined/null/empty values before joining with
'\n') so lexicalScore receives a clean document string.

---

Nitpick comments:
In `@packages/worker/src/mcp/capabilities/unified-search.ts`:
- Around line 548-584: getUnifiedLexicalScore duplicates the weighted
phrase-match logic used in scoreSkillLexicalMatch for skills; extract a small
shared helper (e.g., buildSkillTextScore or scoreSkillLike) that takes the
search text (input.query) and a normalized/adapter object or explicit fields
(skillName, title, description, collection, keywords) so both
getUnifiedLexicalScore and scoreSkillLexicalMatch can call it; use existing
helpers lexicalScore, scoreSkillPhraseMatch, and normalizeSearchPhrase inside
the new helper and update getUnifiedLexicalScore to call it when handling 's:'
keys (operating on SkillSearchHit) and update scoreSkillLexicalMatch to call it
for McpSkillRow, mapping fields as needed.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 8e2951c7-91d4-40f5-921b-1dfdf98af7b4

📥 Commits

Reviewing files that changed from the base of the PR and between 5169596 and 2c62cf6.

📒 Files selected for processing (1)
  • packages/worker/src/mcp/capabilities/unified-search.ts

Comment on lines +84 to +94
function scoreCapabilityLexicalMatch(
query: string,
hit: CapabilitySearchHit,
): number {
const normalizedQuery = normalizeSearchPhrase(query)
const doc = [hit.name, hit.domain, hit.description].join('\n')
let bonus = 0
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.name) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
return lexicalScore(query, doc) + bonus
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Filter out undefined values before joining to avoid "undefined" string in doc.

If hit.domain or hit.description is undefined, Array.join will convert it to the literal string "undefined", which could affect lexical matching.

Proposed fix
 function scoreCapabilityLexicalMatch(
 	query: string,
 	hit: CapabilitySearchHit,
 ): number {
 	const normalizedQuery = normalizeSearchPhrase(query)
-	const doc = [hit.name, hit.domain, hit.description].join('\n')
+	const doc = [hit.name, hit.domain, hit.description].filter(Boolean).join('\n')
 	let bonus = 0
 	bonus += scoreSkillPhraseMatch(normalizedQuery, hit.name) * 1.5
 	bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
 	return lexicalScore(query, doc) + bonus
 }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/worker/src/mcp/capabilities/unified-search.ts` around lines 84 - 94,
In scoreCapabilityLexicalMatch, the doc string is built by joining [hit.name,
hit.domain, hit.description] which can introduce the literal "undefined" when
hit.domain or hit.description are undefined; update the construction of doc
inside scoreCapabilityLexicalMatch to only include defined/nonnull fields (e.g.,
build an array [hit.name, hit.domain, hit.description] and filter out
undefined/null/empty values before joining with '\n') so lexicalScore receives a
clean document string.

Comment on lines +105 to +118
function scoreUiArtifactLexicalMatch(
query: string,
hit: UiArtifactSearchHit,
): number {
const normalizedQuery = normalizeSearchPhrase(query)
const parameterText = (hit.parameters ?? [])
.map((parameter) => `${parameter.name} ${parameter.description}`)
.join('\n')
const doc = [hit.title, hit.description, hit.runtime, parameterText].join('\n')
let bonus = 0
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.title) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
return lexicalScore(query, doc) + bonus
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Same filter(Boolean) recommendation applies here.

hit.runtime or hit.description being undefined would inject literal "undefined" into the doc string.

Proposed fix
 function scoreUiArtifactLexicalMatch(
 	query: string,
 	hit: UiArtifactSearchHit,
 ): number {
 	const normalizedQuery = normalizeSearchPhrase(query)
 	const parameterText = (hit.parameters ?? [])
 		.map((parameter) => `${parameter.name} ${parameter.description}`)
 		.join('\n')
-	const doc = [hit.title, hit.description, hit.runtime, parameterText].join('\n')
+	const doc = [hit.title, hit.description, hit.runtime, parameterText].filter(Boolean).join('\n')
 	let bonus = 0
 	bonus += scoreSkillPhraseMatch(normalizedQuery, hit.title) * 1.5
 	bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
 	return lexicalScore(query, doc) + bonus
 }
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
function scoreUiArtifactLexicalMatch(
query: string,
hit: UiArtifactSearchHit,
): number {
const normalizedQuery = normalizeSearchPhrase(query)
const parameterText = (hit.parameters ?? [])
.map((parameter) => `${parameter.name} ${parameter.description}`)
.join('\n')
const doc = [hit.title, hit.description, hit.runtime, parameterText].join('\n')
let bonus = 0
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.title) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
return lexicalScore(query, doc) + bonus
}
function scoreUiArtifactLexicalMatch(
query: string,
hit: UiArtifactSearchHit,
): number {
const normalizedQuery = normalizeSearchPhrase(query)
const parameterText = (hit.parameters ?? [])
.map((parameter) => `${parameter.name} ${parameter.description}`)
.join('\n')
const doc = [hit.title, hit.description, hit.runtime, parameterText].filter(Boolean).join('\n')
let bonus = 0
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.title) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
return lexicalScore(query, doc) + bonus
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/worker/src/mcp/capabilities/unified-search.ts` around lines 105 -
118, The doc string construction in scoreUiArtifactLexicalMatch can include
literal "undefined" when hit.runtime, hit.description, or parameter.description
are missing; update the joins to skip falsy parts (e.g., filter(Boolean)) so
only defined strings are concatenated and also filter/map parameterText entries
(or default missing parameter.description to '') before joining; keep the
existing behavior of scoring but ensure the array passed to join for doc and
parameterText excludes undefined/null values to avoid injecting "undefined" into
the document used by lexicalScore.

@kentcdodds
kentcdodds merged commit c951968 into main Apr 1, 2026
9 checks passed
@kentcdodds
kentcdodds deleted the cursor/skill-search-discoverability-9a44 branch April 17, 2026 01:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants