calls+watcher: route guardian copy and watch handlers through call-site IDs#26105
Conversation
| undefined, | ||
| undefined, | ||
| { signal, config: { modelIntent: "latency-optimized" } }, | ||
| { signal, config: { callSite: "guardianQuestionCopy" } }, |
There was a problem hiding this comment.
🔴 callSite not passed to resolveConfiguredProvider(), causing potential provider/model mismatch
The PR migrates guardian-question-copy.ts from modelIntent to callSite in the sendMessage config (line 85), but the resolveConfiguredProvider() call at assistant/src/calls/guardian-question-copy.ts:55 is not updated to pass the callSite. resolveConfiguredProvider(callSite?) uses the call-site config to select the correct provider when callSite is given (assistant/src/providers/provider-send-message.ts:67-68), but falls back to the legacy services.inference.provider path when it's omitted. If a user configures llm.callSites.guardianQuestionCopy.provider = "openai", the code selects the default provider (e.g. Anthropic) at line 55 but RetryProvider.normalizeViaCallSite resolves an OpenAI model from the callSite config — sending an incompatible model name to the wrong provider, causing an API error and forcing a fallback.
Prompt for agents
The callSite "guardianQuestionCopy" is passed in the sendMessage config at line 85, but the provider is resolved at line 55 without the callSite argument. resolveConfiguredProvider() supports an optional callSite parameter that routes provider selection through the llm.callSites config (see provider-send-message.ts:48-69). To fix, pass the callSite to the provider resolution: change line 55 from resolveConfiguredProvider() to resolveConfiguredProvider("guardianQuestionCopy"). This ensures the provider selected matches the model/settings that the RetryProvider resolves via the callSite config.
Was this helpful? React with 👍 or 👎 to provide feedback.
| { | ||
| config: { | ||
| modelIntent: "latency-optimized", | ||
| callSite: "watchCommentary", |
There was a problem hiding this comment.
🔴 callSite not passed to getConfiguredProvider() for watchCommentary, causing potential provider/model mismatch
Same pattern as in guardian-question-copy: callSite: "watchCommentary" is passed in the sendMessage config, but getConfiguredProvider() at assistant/src/daemon/watch-handler.ts:116 is called without the callSite argument. getConfiguredProvider(callSite?) delegates to resolveConfiguredProvider(callSite) which selects the provider from llm.callSites[id].provider when given a callSite (assistant/src/providers/provider-send-message.ts:94-97). Without it, the legacy services.inference.provider is used, risking a provider/model mismatch if the user configured a different provider for watchCommentary.
Prompt for agents
The callSite "watchCommentary" is passed in the sendMessage config at line 167, but the provider is resolved at line 116 via getConfiguredProvider() without a callSite argument. getConfiguredProvider() supports an optional callSite parameter (see provider-send-message.ts:93-97). To fix, change line 116 from getConfiguredProvider() to getConfiguredProvider("watchCommentary"). This ensures the provider selected matches the model/settings resolved from the callSite config by the RetryProvider.
Was this helpful? React with 👍 or 👎 to provide feedback.
| { | ||
| config: { | ||
| modelIntent: "quality-optimized", | ||
| callSite: "watchSummary", |
There was a problem hiding this comment.
🔴 callSite not passed to getConfiguredProvider() for watchSummary, causing potential provider/model mismatch
Same pattern as the other two call sites: callSite: "watchSummary" is passed in the sendMessage config, but getConfiguredProvider() at assistant/src/daemon/watch-handler.ts:228 is called without the callSite argument. If a user configures llm.callSites.watchSummary.provider to a different provider than the default, the wrong provider is used for the actual API call while the RetryProvider resolves model/settings from the callSite config, causing a mismatch.
Prompt for agents
The callSite "watchSummary" is passed in the sendMessage config at line 332, but the provider is resolved at line 228 via getConfiguredProvider() without a callSite argument. getConfiguredProvider() supports an optional callSite parameter (see provider-send-message.ts:93-97). To fix, change line 228 from getConfiguredProvider() to getConfiguredProvider("watchSummary"). This ensures the provider selected matches the model/settings resolved from the callSite config by the RetryProvider.
Was this helpful? React with 👍 or 👎 to provide feedback.
| undefined, | ||
| undefined, | ||
| { signal, config: { modelIntent: "latency-optimized" } }, | ||
| { signal, config: { callSite: "guardianQuestionCopy" } }, |
There was a problem hiding this comment.
🚩 AGENTS.md rule about modelIntent may need updating for callSite migration
The AGENTS.md states: "Use modelIntent ('latency-optimized', 'quality-optimized', 'vision-optimized') instead of hardcoded model IDs." This PR migrates away from modelIntent to the newer callSite system which supersedes it. The codebase clearly treats modelIntent as legacy (assistant/src/providers/retry.ts:91 comments "Legacy modelIntent path (preserved)") while callSite is the new unified path. The AGENTS.md rule should be updated to reflect the current callSite pattern as the preferred approach, with modelIntent documented as legacy. Not flagged as a bug since the spirit of the rule (avoid hardcoded model IDs) is preserved and the codebase has already established the callSite pattern.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: aea4da25af
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| { | ||
| config: { | ||
| modelIntent: "latency-optimized", | ||
| callSite: "watchCommentary", |
There was a problem hiding this comment.
Restore latency routing for live watch commentary
Replacing modelIntent: "latency-optimized" with callSite: "watchCommentary" changes default model selection because resolveCallSiteConfig falls back to llm.default when no llm.callSites.watchCommentary entry exists (assistant/src/config/llm-resolver.ts:31-37), and this repo does not seed that call-site override anywhere. On default config, that means commentary now runs on llm.default.model (claude-opus-4-6), not the prior latency model (claude-haiku-4-5-20251001), which can noticeably increase latency/cost for real-time commentary.
Useful? React with 👍 / 👎.
| undefined, | ||
| undefined, | ||
| { signal, config: { modelIntent: "latency-optimized" } }, | ||
| { signal, config: { callSite: "guardianQuestionCopy" } }, |
There was a problem hiding this comment.
Keep guardian copy on latency-optimized model
This migration drops the previous latency-optimized intent for guardian copy without providing an equivalent per-call-site override. Since unresolved call sites inherit llm.default (assistant/src/config/llm-resolver.ts:31-37), guardian copy now uses the default model (assistant/src/config/schemas/llm.ts:223-225) rather than the old latency intent mapping (assistant/src/providers/model-intents.ts:15), increasing the chance that the 5s generation window falls back to deterministic copy.
Useful? React with 👍 / 👎.
| { | ||
| config: { | ||
| modelIntent: "latency-optimized", | ||
| callSite: "watchCommentary", |
There was a problem hiding this comment.
Pass callSite into provider selection for watch calls
The request now declares a call-site ID, but watch generation still obtains its transport via getConfiguredProvider() with no call-site argument, so provider choice stays on services.inference.provider (assistant/src/providers/provider-send-message.ts:66-69). At the same time, RetryProvider applies model values from the call-site config (assistant/src/providers/retry.ts:178-194), which can produce cross-provider mismatches (e.g., OpenAI model resolved for a call-site but request sent through Anthropic) and runtime failures when users set llm.callSites.watchCommentary/watchSummary.provider overrides.
Useful? React with 👍 / 👎.
8ce9500
into
siddseethepalli/unify-llm-callsites
…es} (#26159) * config(llm): add unified llm schema with call-site enum and profile refines (#26089) * config(llm): add unified llm schema with call-site enum and profile refines * fix(llm-schema): replace deepPartialObject helper with explicit .partial().extend() Zod 4's readonly shape typing tripped TS2542 in the LSP for the generic walker. Inline the one-level expansion for ContextWindowSchema and switch the superRefine issue code to the string literal (Zod 4 deprecated ZodIssueCode). * config(llm): add resolveCallSiteConfig resolver with deep merge (#26094) * config(llm): add resolveCallSiteConfig resolver with deep merge * fix(llm-resolver): deep-clone nested objects so resolved configs are isolated snapshots Codex flagged that the merge helper aliased nested objects from llm.default when no override touched them, so a caller mutating the returned config would silently corrupt the source. Recurse into plain-object sources unconditionally and add a regression test. * config(llm): add llm field to AssistantConfigSchema (no behavior change) (#26095) * config(llm): add llm field to AssistantConfigSchema (no behavior change) * fix(llm-schema): add field-level defaults so partial llm configs don't trigger full config reset Codex flagged that requiring all LLMConfigBase fields meant the loader's leaf-deletion recovery couldn't repair partial/invalid llm blocks — falling through to cloneDefaultConfig() and discarding the user's other valid settings. Add .default(...) to every leaf so LLMSchema.parse({}) returns a fully-defaulted object, matching the pattern used by sibling config schemas. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * providers: accept callSite in per-call config; resolve via resolveCallSiteConfig (#26102) * workspace: migrate scattered LLM config keys into unified llm structure (#26101) * workspace: migrate scattered LLM config keys into unified llm structure * fix(migration): preserve existing llm subtree; map notification intent to both call sites Codex flagged two issues: - The migration assignment replaced config.llm wholesale, destroying any pre-existing llm.callSites/profiles when llm.default was absent. Now merges into existing config.llm, preserving non-conflicting entries. - notifications.decisionModelIntent drives both notification classification and preference extraction, but the migration only seeded notificationDecision. Now seeds both call sites. * memory: route extraction/consolidation/retrieval through call-site IDs (#26106) * memory: route narrative/pattern/summarization/starters through call-site IDs (#26107) * notifications: route decision and preference extraction through call-site IDs (#26109) * calls+watcher: route guardian copy and watch handlers through call-site IDs (#26105) * utility: route classifier and analyzer LLM calls through call-site IDs (#26111) * macos(settings): migrate InferenceServiceCard reads/writes to llm.default.* (#26113) * workspace+conversation: route commit message and title through call-site IDs (#26112) * ui: route identity intro and empty-state greeting through call-site IDs (#26108) * daemon: thread callSite through processMessage options and adapter callbacks (#26115) * daemon: thread callSite through processMessage options and adapter callbacks * fix(callsite-threading): complete interface contract and server.ts symmetry Devin flagged two gaps in PR #26115: - ProcessConversationContext interface missing callSite in its runAgentLoop options type (works via structural typing but contract was incomplete; mocks would silently drop the field). - DaemonServer.persistAndProcessMessage didn't thread callSite to conversation.runAgentLoop, while DaemonServer.processMessage did. Aligned. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(callsite): don't default unspecified callers to 'mainAgent' Codex flagged that defaulting to mainAgent for every turn routes them through the new RetryProvider call-site resolver, which reads from llm.default — but config-model.setModel still writes to services.inference without syncing llm.default. Result: stale/incompatible model IDs after a model switch. Defer the cutover. agent-loop turns now keep using the legacy modelIntent path (turnCallSite = options?.callSite, no fallback). PRs 7-11 still explicitly pass callSite and route through the new resolver as intended. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * heartbeat: pass callSite: 'heartbeatAgent' instead of speed kwarg (#26125) * filing: pass callSite: 'filingAgent' instead of speed kwarg (#26124) * runtime/analyze-conversation: route through callSite: 'analyzeConversation' (#26126) * subagent: pass callSite: 'subagentSpawn' when spawning isolated agents (#26122) * calls: route the call agent loop through callSite: 'callAgent' (#26123) * macos(settings): add SettingsStore APIs for per-call-site overrides (#26128) * macos(settings): add SettingsStore APIs for per-call-site overrides * fix(callsite-overrides): harden setCallSiteOverrides against dup-id crash and batch divergence Devin and Codex flagged two issues: - Dictionary(uniqueKeysWithValues:) crashes if callers pass duplicate CallSiteOverride.id values (external input — must be tolerant). Switch to Dictionary(_:uniquingKeysWith:) with last-write-wins. - Batch updates locally cleared entries omitted from the input but only PATCHed entries that were present, so omitted entries appeared cleared in the UI but reappeared on next sync. Now the PATCH payload includes NSNull clears for every catalog entry not in the batch, aligning remote with local. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(callsite-overrides): null entire entry on clear so non-UI leaves get cleared too Codex P2 (PR #26128 cycle 2): clearCallSiteOverride only nulled provider/model/profile, but call-site config supports additional leaves (maxTokens, effort, speed, thinking, contextWindow). If those were set via manual edits, the UI would report cleared while the daemon kept applying hidden overrides. Switch the PATCH payload from { provider: null, model: null, profile: null } to a single null on the entry itself. The Zod fragment treats null as absent, so the resolver falls back to llm.default. Same fix applies to the omitted-catalog-entry clears in setCallSiteOverrides batch. Tests updated to assert the new shape. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * macos(settings): confirm default-provider switch when call-site overrides exist (#26133) * macos(settings): show 'N call-site overrides' badge with read-only list sheet (#26135) * macos(settings): show 'N call-site overrides' badge with read-only list sheet * fix(comments): drop PR-number breadcrumbs in callsite override files Devin flagged that comments referencing PR 22/23/24 violate clients/AGENTS.md 'Comment Quality' rule (no breadcrumbs). Replaced with timeless descriptions of code intent. * macos(settings): make per-task override sheet editable with provider/model pickers (#26136) * macos(settings): make per-task override sheet editable with provider/model pickers * fix(callsite-sheet): preserve external updates and seed override from active default provider Codex flagged two P1s: - syncDraftsFromStore compared drafts against the NEW persisted value to decide 'touched', so external store updates were treated as user edits and got overwritten by Save All. Track the previously-persisted value in lastSyncedFromStore and consider a row touched only when the draft differs from that baseline. - Toggling 'Override default' on initialized provider from providerIds.first instead of the user's actual default provider, which could pin the wrong provider on save. Pass the user's default provider into CallSiteOverrideRow and seed from it. * fix(callsite-sheet): use entry-level null path for cleared rows in saveAll/resetAll Devin flagged that saveAll() and resetAll() were passing all-nil entries to setCallSiteOverrides, which routed them through the field-level null path (provider/model/profile = null). That left advanced leaves (maxTokens, effort, temperature, contextWindow) untouched on the daemon. Fix: - saveAll(): filter to entries with hasOverride == true; toggled-off rows fall through to the entry-level null path. - resetAll(): pass an empty list so every catalog entry hits the entry-level null path. * config(llm): remove deprecated scattered LLM keys (#26140) * fix(config-loader): treat JSON null as key deletion in deepMergeOverwrite (#26153) * fix(agent-loop): default user-initiated turns to callSite: 'mainAgent' (#26154) * fix(meet-join): migrate consent-monitor + session-manager to callSite contract (#26155) * fix(macos): atomic provider+model save via single PATCH (#26156) * fix(cleanup): remove dead code, refresh comments, add migration test, update docs (#26157) * fix(r2): catalog test count, skill self-knowledge doc, AGENTS.md, loader docstring (#26158) * fix(llm-callsite): refresh stale docstring, restore overflow budget, restore SettingsStore fallback (#26252) * fix(llm-callsite): route provider transport and field precedence through callSite (#26254) * fix(llm-callsite): pass CI + address subagent/thinking/temperature review comments (#26258) * test(extension-id-guard): allow CWS URL matches; mirrors main PR #26263 (#26270) * fix(llm-callsite): UI override state divergence, null-as-delete, migration gaps (#26271) * Fix Chrome extension allowlist ID and clarify README dev setup (#26259) Update the canonical allowlist to use the correct published CWS extension ID (hphbdmpffeigpcdjkckleobjmhhokpne). Restructure the Chrome extension README to clearly explain the allowlist merge strategy, separate the macOS app (automatic) path from the manual native messaging setup, and show how dev + prod extensions work side-by-side. Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(clients): enable non-contiguous glyph layout for NSTextView-backed code views (#26242) TextKit 1 defaults NSLayoutManager.allowsNonContiguousLayout to false, which forces full-document glyph layout from character 0 on the main thread whenever a glyph range is queried. Attaching an NSTextView to its scroll view (setDocumentView: -> _setSuperview: -> setNeedsDisplayInRect: -> _glyphRangeForBoundingRect:) triggers that query during makeNSView, producing multi-second hangs on large code blocks. Opt into non-contiguous layout on every TextKit 1 stack we build via NSViewRepresentable so glyph generation is confined to the requested bounding rect. Also replace NSLayoutManager.ensureLayout(for:) in the code-view sizeThatFits paths with direct lineCount * fixedLineHeight math: the text container is unbounded horizontally (no wrapping) and paragraph style pins minimumLineHeight == maximumLineHeight, so the geometry is exact and avoids a second O(glyph count) main-thread path. Fixes VELLUM-ASSISTANT-MACOS-J2. Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: ashlee@vellum.ai <ashlee@vellum.ai> * fix(contacts): show Assistant badge for assistant-type contacts (LUM-1009) (#26239) * fix(contacts): show Assistant badge for assistant-type contacts (LUM-1009) * Move role/contactType derivation onto Kind for valid initializer --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(llm-callsite): UI override state divergence, null-as-delete, migration gaps - deepMergeOverwrite: null on scalar/null targets assigns null (preserves nullable config fields like activeHoursStart); null on object targets still deletes (call-site clearing). Fixes regression where PATCH with null for nullable fields was deleted then re-defaulted. - InferenceServiceCard: override confirmation dialog only fires when the resolved provider ID actually changes, not on mode-only toggles where both old and new resolve to the same provider. - CallSiteOverridesSheet: per-row Save uses replaceCallSiteOverride (clear-then-set) so stale daemon-side leaves are removed. The partial-update setCallSiteOverride would retain fields the draft nil'd. - CallSiteOverrideRow: merge consecutive .padding modifiers into single EdgeInsets call per macOS AGENTS.md layout rule. - SettingsStore: add replaceCallSiteOverride for full-entry replacement. --------- Co-authored-by: Noa Flaherty <noa@vellum.ai> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: ashlee@vellum.ai <ashlee@vellum.ai> * fix(llm-callsite): seed latency-optimized defaults and fix guardian provider routing (#26275) * fix(meet-bot): address review feedback — Docker build, scraper races, audio capture, storage writer (#26264) * fix(meet): chat concurrency, dispose teardown, and wake adapter fidelity (#26265) * fix: heartbeat dual-emit, analysis dedup, test hermiticity, credential executor discovery (#26266) * fix: model default fallback, empty-response nudge scan (#26268) - Update FALLBACK_DEFAULT_MODEL to claude-opus-4-7 + test - Fix resolveModel to check Anthropic catalog (not just current default) so stale persisted defaults (e.g. claude-opus-4-6) don't get sent to non-Anthropic providers - Fix priorAssistantHadVisibleText backward scan to check ALL prior assistant messages, not just the most recent one Addresses review feedback from PRs #26247, #26164. * fix(meet): TTS stream races, barge-in tracking, ffmpeg error classification (#26267) * Fix extension-id-sync-guard test after canonical ID update (#26263) The guard test asserts that canonical extension IDs appear only in the allowlist config file. After updating the canonical ID to match the published CWS extension, it now collides with CWS URLs in README and browser-execution.ts. Fix by stripping CWS URLs before checking for bare ID occurrences, and ignore .codex-worktrees (repo copies). Also remove hardcoded CWS ID from README in favor of reading from the canonical config. Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(llm-callsite): seed latency-optimized defaults, fix guardian provider routing, clean stale comments - Add LATENCY_OPTIMIZED_CALLSITE_DEFAULTS to schema for new installs - Create migration 040 to seed latency-optimized call-site entries for existing workspaces - Fix guardian-action-generators to use getConfiguredProvider() instead of bypassing call-site resolution - Restore commitMessage maxTokens: 120 and temperature: 0.2 via call-site defaults - Remove stale PR-reference comments from analyze-conversation.ts and voice-session-bridge.ts Addresses consolidated review feedback from PRs #26101-#26140. --------- Co-authored-by: Noa Flaherty <noa@vellum.ai> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(retry): stop forwarding contextWindow/provider to provider request body (#26280) * chore(skills): regenerate catalog.json --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Noa Flaherty <noa@vellum.ai> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: ashlee@vellum.ai <ashlee@vellum.ai>
Summary
calls/guardian-question-copy.ts->callSite: 'guardianQuestionCopy'.daemon/watch-handler.ts(commentary path) ->callSite: 'watchCommentary'.daemon/watch-handler.ts(summary path) ->callSite: 'watchSummary'.Part of plan: unify-llm-callsites.md (PR 17 of 24)