Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
207 changes: 184 additions & 23 deletions packages/worker/src/mcp/capabilities/unified-search.ts
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,83 @@ function buildSkillUsage(skillName: string): string {
return `Run with meta_run_skill: ${runArgs}. Optionally include "params": { ... }. To inspect code, call meta_get_skill then execute.`
}

function normalizeSearchPhrase(value: string | null | undefined): string {
return (value ?? '')
.toLowerCase()
.replace(/[-_]+/g, ' ')
.replace(/[^a-z0-9\s]+/g, ' ')
.replace(/\s+/g, ' ')
.trim()
}

function scoreSkillPhraseMatch(
normalizedQuery: string,
value: string | null | undefined,
): number {
const normalizedValue = normalizeSearchPhrase(value)
if (!normalizedQuery || !normalizedValue) return 0
if (normalizedValue === normalizedQuery) return 1.5
if (normalizedValue.includes(normalizedQuery)) return 1
return 0
}

function scoreSkillLexicalMatch(
query: string,
row: McpSkillRow,
doc: string,
keywords: ReadonlyArray<string>,
): number {
const normalizedQuery = normalizeSearchPhrase(query)
let bonus = 0

bonus += scoreSkillPhraseMatch(normalizedQuery, row.name) * 2
bonus += scoreSkillPhraseMatch(normalizedQuery, row.title) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, row.description) * 1.25
bonus += scoreSkillPhraseMatch(normalizedQuery, row.search_text) * 1
bonus += scoreSkillPhraseMatch(normalizedQuery, row.collection_name) * 0.25
for (const keyword of keywords) {
bonus += scoreSkillPhraseMatch(normalizedQuery, keyword) * 0.5
}

return lexicalScore(query, doc) + bonus
}

function scoreCapabilityLexicalMatch(
query: string,
hit: CapabilitySearchHit,
): number {
const normalizedQuery = normalizeSearchPhrase(query)
const doc = [hit.name, hit.domain, hit.description].join('\n')
let bonus = 0
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.name) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
return lexicalScore(query, doc) + bonus
}
Comment on lines +84 to +94

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Filter out undefined values before joining to avoid "undefined" string in doc.

If hit.domain or hit.description is undefined, Array.join will convert it to the literal string "undefined", which could affect lexical matching.

Proposed fix
 function scoreCapabilityLexicalMatch(
 	query: string,
 	hit: CapabilitySearchHit,
 ): number {
 	const normalizedQuery = normalizeSearchPhrase(query)
-	const doc = [hit.name, hit.domain, hit.description].join('\n')
+	const doc = [hit.name, hit.domain, hit.description].filter(Boolean).join('\n')
 	let bonus = 0
 	bonus += scoreSkillPhraseMatch(normalizedQuery, hit.name) * 1.5
 	bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
 	return lexicalScore(query, doc) + bonus
 }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/worker/src/mcp/capabilities/unified-search.ts` around lines 84 - 94,
In scoreCapabilityLexicalMatch, the doc string is built by joining [hit.name,
hit.domain, hit.description] which can introduce the literal "undefined" when
hit.domain or hit.description are undefined; update the construction of doc
inside scoreCapabilityLexicalMatch to only include defined/nonnull fields (e.g.,
build an array [hit.name, hit.domain, hit.description] and filter out
undefined/null/empty values before joining with '\n') so lexicalScore receives a
clean document string.


function scoreSecretLexicalMatch(query: string, hit: SecretSearchHit): number {
const normalizedQuery = normalizeSearchPhrase(query)
const doc = [hit.name, hit.description].join('\n')
let bonus = 0
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.name) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
return lexicalScore(query, doc) + bonus
}

function scoreUiArtifactLexicalMatch(
query: string,
hit: UiArtifactSearchHit,
): number {
const normalizedQuery = normalizeSearchPhrase(query)
const parameterText = (hit.parameters ?? [])
.map((parameter) => `${parameter.name} ${parameter.description}`)
.join('\n')
const doc = [hit.title, hit.description, hit.runtime, parameterText].join('\n')
let bonus = 0
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.title) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
return lexicalScore(query, doc) + bonus
}
Comment on lines +105 to +118

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Same filter(Boolean) recommendation applies here.

hit.runtime or hit.description being undefined would inject literal "undefined" into the doc string.

Proposed fix
 function scoreUiArtifactLexicalMatch(
 	query: string,
 	hit: UiArtifactSearchHit,
 ): number {
 	const normalizedQuery = normalizeSearchPhrase(query)
 	const parameterText = (hit.parameters ?? [])
 		.map((parameter) => `${parameter.name} ${parameter.description}`)
 		.join('\n')
-	const doc = [hit.title, hit.description, hit.runtime, parameterText].join('\n')
+	const doc = [hit.title, hit.description, hit.runtime, parameterText].filter(Boolean).join('\n')
 	let bonus = 0
 	bonus += scoreSkillPhraseMatch(normalizedQuery, hit.title) * 1.5
 	bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
 	return lexicalScore(query, doc) + bonus
 }
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
function scoreUiArtifactLexicalMatch(
query: string,
hit: UiArtifactSearchHit,
): number {
const normalizedQuery = normalizeSearchPhrase(query)
const parameterText = (hit.parameters ?? [])
.map((parameter) => `${parameter.name} ${parameter.description}`)
.join('\n')
const doc = [hit.title, hit.description, hit.runtime, parameterText].join('\n')
let bonus = 0
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.title) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
return lexicalScore(query, doc) + bonus
}
function scoreUiArtifactLexicalMatch(
query: string,
hit: UiArtifactSearchHit,
): number {
const normalizedQuery = normalizeSearchPhrase(query)
const parameterText = (hit.parameters ?? [])
.map((parameter) => `${parameter.name} ${parameter.description}`)
.join('\n')
const doc = [hit.title, hit.description, hit.runtime, parameterText].filter(Boolean).join('\n')
let bonus = 0
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.title) * 1.5
bonus += scoreSkillPhraseMatch(normalizedQuery, hit.description) * 1
return lexicalScore(query, doc) + bonus
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/worker/src/mcp/capabilities/unified-search.ts` around lines 105 -
118, The doc string construction in scoreUiArtifactLexicalMatch can include
literal "undefined" when hit.runtime, hit.description, or parameter.description
are missing; update the joins to skip falsy parts (e.g., filter(Boolean)) so
only defined strings are concatenated and also filter/map parameterText entries
(or default missing parameter.description to '') before joining; keep the
existing behavior of scoring but ensure the array passed to join for doc and
parameterText excludes undefined/null values to avoid injecting "undefined" into
the document used by lexicalScore.


function skillRowEmbedDoc(
row: McpSkillRow,
specs: Record<string, CapabilitySpec>,
Expand All @@ -56,6 +133,7 @@ function skillRowEmbedDoc(
inferred = []
}
return buildSkillEmbedText({
skillName: row.name,
title: row.title,
description: row.description,
collectionName: row.collection_name,
Expand Down Expand Up @@ -188,9 +266,21 @@ async function searchSkillsForUser(input: {
(row) => [row.id, skillRowEmbedDoc(row, input.specs)] as const,
),
)
const keywordsById = Object.fromEntries(
filteredRows.map((row) => [row.id, parseJsonStringArray(row.keywords)] as const),
)
const lexicalScoreById = Object.fromEntries(
ids.map((id) => {
const row = idSet.get(id)!
return [
id,
scoreSkillLexicalMatch(q, row, docsById[id]!, keywordsById[id] ?? []),
] as const
}),
)

const lexicalOrder = sortIdsByScore(ids, (id) =>
lexicalScore(q, docsById[id]!),
lexicalScoreById[id]!,
)

let vectorOrder: Array<string>
Expand Down Expand Up @@ -262,10 +352,18 @@ async function searchSkillsForUser(input: {
[lexicalOrder, vectorOrder],
CAPABILITY_SEARCH_RRF_K,
)
const ordered = sortIdsByScore(ids, (id) => fused.get(id) ?? 0).slice(
0,
Math.max(1, Math.min(input.limit, ids.length)),
)
const ordered = [...ids]
.sort((a, b) => {
const fusedDiff = (fused.get(b) ?? 0) - (fused.get(a) ?? 0)
if (fusedDiff !== 0) return fusedDiff
const lexicalDiff = lexicalScoreById[b]! - lexicalScoreById[a]!
if (lexicalDiff !== 0) return lexicalDiff
return (
(vectorRankById.get(a) ?? Number.MAX_SAFE_INTEGER) -
(vectorRankById.get(b) ?? Number.MAX_SAFE_INTEGER)
)
})
.slice(0, Math.max(1, Math.min(input.limit, ids.length)))

const matches = ordered.map((id) => {
const row = idSet.get(id)!
Expand Down Expand Up @@ -355,10 +453,11 @@ export async function searchUnified(input: {
const builtinFilter: VectorizeVectorMetadataFilter = {
kind: { $eq: 'builtin' },
}
const candidateLimit = Math.min(100, Math.max(input.limit * 3, 25))
const capResult = await searchCapabilities({
env: input.env,
query: input.query,
limit: input.limit,
limit: candidateLimit,
detail: false,
specs: input.specs,
vectorMetadataFilter: isCapabilitySearchOffline(input.env)
Expand Down Expand Up @@ -387,13 +486,13 @@ export async function searchUnified(input: {
if (input.userId) {
secretResult = await searchSecretsForUser({
query: input.query,
limit: input.limit,
limit: candidateLimit,
rows: input.userSecretRows,
})
skillResult = await searchSkillsForUser({
env: input.env,
query: input.query,
limit: input.limit,
limit: candidateLimit,
specs: input.specs,
userId: input.userId,
collectionSlug: input.skillCollectionSlug,
Expand All @@ -403,13 +502,23 @@ export async function searchUnified(input: {
baseUrl: input.baseUrl,
env: input.env,
query: input.query,
limit: input.limit,
limit: candidateLimit,
userId: input.userId,
rows: input.uiArtifactRows,
appSecretsByAppId: input.appSecretsByAppId,
})
}

const capByName = new Map(capResult.matches.map((m) => [m.name, m] as const))
const skillByName = new Map(
skillResult.matches.map((m) => [m.skillName, m] as const),
)
const secretByName = new Map(
secretResult.matches.map((m) => [m.name, m] as const),
)
const uiArtifactById = new Map(
uiArtifactResult.matches.map((m) => [m.appId, m] as const),
)
const capKeys = capResult.matches.map((m) => `c:${m.name}`)
const skillKeys = skillResult.matches.map((m) => `s:${m.skillName}`)
const secretKeys = secretResult.matches.map((m) => `u:${m.name}`)
Expand All @@ -421,21 +530,73 @@ export async function searchUnified(input: {
const allKeys = [
...new Set([...capKeys, ...skillKeys, ...secretKeys, ...uiArtifactKeys]),
]
const sortedKeys = sortIdsByScore(
allKeys,
(k) => fusedCross.get(k) ?? 0,
).slice(0, Math.max(1, input.limit))

const capByName = new Map(capResult.matches.map((m) => [m.name, m] as const))
const skillByName = new Map(
skillResult.matches.map((m) => [m.skillName, m] as const),
)
const secretByName = new Map(
secretResult.matches.map((m) => [m.name, m] as const),
)
const uiArtifactById = new Map(
uiArtifactResult.matches.map((m) => [m.appId, m] as const),
function getEntityScore(key: string): number {
if (key.startsWith('c:')) {
return capByName.get(key.slice(2))?.fusedScore ?? 0
}
if (key.startsWith('s:')) {
return skillByName.get(key.slice(2))?.fusedScore ?? 0
}
if (key.startsWith('u:')) {
return secretByName.get(key.slice(2))?.fusedScore ?? 0
}
if (key.startsWith('a:')) {
return uiArtifactById.get(key.slice(2))?.fusedScore ?? 0
}
return 0
}
function getUnifiedLexicalScore(key: string): number {
if (key.startsWith('c:')) {
const hit = capByName.get(key.slice(2))
return hit ? scoreCapabilityLexicalMatch(input.query, hit) : 0
}
if (key.startsWith('s:')) {
const hit = skillByName.get(key.slice(2))
if (!hit) return 0
return (
lexicalScore(input.query, [
hit.skillName,
hit.title,
hit.description,
hit.collection ?? '',
hit.keywords.join(' '),
].join('\n')) +
scoreSkillPhraseMatch(normalizeSearchPhrase(input.query), hit.skillName) *
2 +
scoreSkillPhraseMatch(normalizeSearchPhrase(input.query), hit.title) *
1.5 +
scoreSkillPhraseMatch(
normalizeSearchPhrase(input.query),
hit.description,
) *
1.25
)
}
if (key.startsWith('u:')) {
const hit = secretByName.get(key.slice(2))
return hit ? scoreSecretLexicalMatch(input.query, hit) : 0
}
if (key.startsWith('a:')) {
const hit = uiArtifactById.get(key.slice(2))
return hit ? scoreUiArtifactLexicalMatch(input.query, hit) : 0
}
return 0
}
const lexicalScoreByKey = new Map(
allKeys.map((key) => [key, getUnifiedLexicalScore(key)] as const),
)
const sortedKeys = [...allKeys]
.sort((a, b) => {
const fusedDiff = (fusedCross.get(b) ?? 0) - (fusedCross.get(a) ?? 0)
if (fusedDiff !== 0) return fusedDiff
const lexicalDiff =
(lexicalScoreByKey.get(b) ?? 0) - (lexicalScoreByKey.get(a) ?? 0)
if (lexicalDiff !== 0) return lexicalDiff
const entityDiff = getEntityScore(b) - getEntityScore(a)
if (entityDiff !== 0) return entityDiff
return a.localeCompare(b)
})
Comment thread
cursor[bot] marked this conversation as resolved.
.slice(0, Math.max(1, input.limit))

const matches: Array<UnifiedSearchMatch> = []
for (const key of sortedKeys) {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -120,6 +120,80 @@ test('skill collection filter narrows saved skill matches', async () => {
expect(skill.collection).toBe('GitHub Workflows')
})

test('skill name and description matches survive cross-entity ranking', async () => {
const env = { SENTRY_ENVIRONMENT: 'test' } as Env
const specs = {
launch_agent: {
name: 'launch_agent',
domain: 'coding',
description: 'Launch an agent for generic coding work.',
keywords: ['launch', 'agent', 'coding'],
readOnly: true,
idempotent: true,
destructive: false,
inputFields: ['task'],
requiredInputFields: ['task'],
outputFields: ['agentId'],
inputSchema: {},
},
cursor_docs: {
name: 'cursor_docs',
domain: 'coding',
description: 'Look up Cursor product documentation.',
keywords: ['cursor', 'docs'],
readOnly: true,
idempotent: true,
destructive: false,
inputFields: ['query'],
requiredInputFields: ['query'],
outputFields: ['results'],
inputSchema: {},
},
cursor_agent_status: {
name: 'cursor_agent_status',
domain: 'coding',
description: 'Inspect the status of an existing Cursor agent.',
keywords: ['cursor', 'agent', 'status'],
readOnly: true,
idempotent: true,
destructive: false,
inputFields: ['agentId'],
requiredInputFields: ['agentId'],
outputFields: ['status'],
inputSchema: {},
},
} satisfies Record<string, CapabilitySpec>
const skillRow: McpSkillRow = {
...createSkillRow('launch-cursor-cloud-agent'),
name: 'launch-cursor-cloud-agent',
title: 'Launch Cursor Cloud Agent',
description: 'Launch a Cursor Cloud Agent for an autonomous coding task.',
collection_name: 'Cursor',
collection_slug: 'cursor',
keywords: JSON.stringify(['cursor', 'cloud', 'agent', 'launch']),
search_text: 'launch cursor cloud agent autonomous coding task',
}

const result = await searchUnified({
env,
baseUrl: 'http://localhost',
query: 'launch Cursor Cloud agent',
limit: 5,
specs,
userId: 'user-123',
skillRows: [skillRow],
uiArtifactRows: [],
userSecretRows: [],
appSecretsByAppId: new Map(),
})

expect(result.matches[0]).toMatchObject({
type: 'skill',
skillName: 'launch-cursor-cloud-agent',
title: 'Launch Cursor Cloud Agent',
})
})

test('search can return standalone user secrets and nest app secrets on apps', async () => {
const env = { SENTRY_ENVIRONMENT: 'test' } as Env
const specs = {} as Record<string, CapabilitySpec>
Expand Down
Loading
Loading