Skip to content

feat(voice): hand a sentence naming no offered label to Clark's turn with the focused widget's actions - #600

Merged
mrgoonie merged 6 commits into
mainfrom
feat/444-label-free-spoken-action
Oct 7, 2026
Merged

mrgoonie merged 6 commits into
mainfrom
feat/444-label-free-spoken-action

Conversation

@mrgoonie

@mrgoonie mrgoonie commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Addresses #444

Summary

Implements the decided option 3 for label-free spoken matching.

Before this change, a sentence spoken while a widget was focused that matched none of its offered labels was refused by the voice session ("Tôi chưa rõ bạn muốn làm gì với widget đang mở…"). It never reached Clark's turn, so "format this as a percentage" with the spreadsheet focused did nothing.

Now, when the focused widget offers actions.perform@1 actions and the session's page sent widgetPerform: 1:

  • the unmatched sentence goes to Clark's normal voice turn;
  • the turn's data section (not the note) lists the focused widget and each offered action: binding id, label, description and input schema. These are the package's words, quoted inertly on one line under a heading that says they are data, not instructions;
  • a short host-authored note tells the model it may perform one through perform_widget_action;
  • the perform goes through the existing tool, so the declared schema, execution policy, host approval card, ledger and origin are the typed path's. Neither the turn nor the widget can approve it.

The host adds no matching heuristics and packages declare no phrasings. A sentence that says a label keeps the direct spoken press. Without a page that can perform, or when the widget offers no performable actions, the sentence is refused by naming what the widget offers, as before.

Code:

  • apps/runtime/src/widget-perform-tool.ts: focusedWidgetActionsContext and FOCUSED_WIDGET_NOTE.
  • apps/runtime/src/voice-session.ts: new injected focusedWidgetContext option. An unmatched sentence is handed to answer with widgetContext when the conditions above hold.
  • apps/runtime/src/bootstrap/voice-bootstrap.ts: the inline answer is extracted as answerSpokenSentence (behaviour unchanged), which adds the note and the data. The new option is wired.

Tests

  • apps/runtime/test/widget-perform.spec.ts, new describe "a spoken sentence that names none of the focused widget's offered actions". It uses a real voice socket and the node's own wiring; only the provider and the model are scripted, and the scripted model calls the real perform_widget_action tool:
    • the turn gets the offered actions as data and the frame is asked once, with a confirmed effect;
    • under an asking policy the host card is placed, nothing is sent, and a spoken "đồng ý" runs it exactly once;
    • a page without widgetPerform is refused and reaches no turn;
    • a labelled sentence keeps the direct press and reaches no turn;
    • package words are made inert, and only the focused instance in this conversation is listed.
  • apps/web/e2e/widget-perform.spec.ts, "a range is formatted when the person says so out loud without naming the action's label". It fails without the wiring (B2 stays 0.25) and passes with it.
  • corepack pnpm verify and corepack pnpm invariants.

Docs impact

  • docs/widget-development.md and .vi.md: the voice paragraph now describes the label-free handoff in place of "not matched yet".
  • docs/widgets-and-extensions.md and .vi.md: the spreadsheet bullet.
  • docs/conformance-traceability.md: the evidence.
  • docs/manifest.json: digests.
  • Official clarkcant-web docs (docs/cli.html and vi/docs/cli.html, the text-editor note) still say label-free sentences are not matched yet and need a follow-up.

🤖 Generated with Claude Code

…with the focused widget's actions

A sentence spoken while a widget that offers actions is focused, and that names none of their labels, now goes to Clark's voice turn instead of being refused. The turn's data lists the focused widget's offered actions (binding id, label, description, input schema) as inert package words, and Clark can perform one only through perform_widget_action, under the same schema, execution policy and host card as a typed request. A labelled sentence keeps the direct spoken press; a page that cannot perform keeps the refusal.
…arts, and hand every unmatched sentence to Clark

- A spoken sentence that names none of the focused widget's labels now goes to Clark's turn whatever the widget offers; it is refused only when no agent is wired.
- The widget's offered actions are read when the turn starts (dataAtStart, resolved after the turn claims the conversation), and left out if the widget was closed or another one focused meanwhile.
- A description is listed only from a running definition with the instance's version and digest that declares the action with the label and schema its binding recorded.
- The package's label, description and input schema share one quoted, inert renderer with the perform tool's list.
- The spoken approval question and the widget refusals follow the person's language.
- The fixture model takes a spoken turn's binding from the turn's data, so the browser journey fails if the data is dropped.
…he widget id in offered-action listings

- A dataAtStart hook that throws gives no data instead of leaving the conversation's claimed turn unreleased.
- The widget id is quoted and made inert in the perform tool's list and in the focused widget's context.
- The voice runner's no-focus and action-gone refusals follow the person's language, and with no agent wired an unmatched sentence is refused whether or not a widget is focused.
- A closed voice session's turn gets no widget data.
- The renderer and the fixture model share one binding key constant.
- Tests: a throwing hook frees the conversation, typed words carrying start-time data are never steered, event-driven waits in place of sleeps, a hostile widget id, the no-focus refusal and the closed session.
@mrgoonie

mrgoonie commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Review attestation: ready to merge at 3f9d83c41f22c7e908aebf3a7c238ac6bd0482fd, reviewed by agent:code-reviewer.

A push to this PR makes this attestation stale; the new head needs its own review.

@mrgoonie

mrgoonie commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Review attestation: ready to merge at 79d12d84553e5371fe46f0de5cdc76b15625315d, reviewed by agent:code-reviewer.

A push to this PR makes this attestation stale; the new head needs its own review.

…poken-action

# Conflicts:
#	apps/runtime/src/bootstrap/voice-bootstrap.ts
#	docs/manifest.json
@mrgoonie

mrgoonie commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Review attestation: ready to merge at 7b54e4891d8fe754700d028a4d86a59a4400ae14, reviewed by agent:code-reviewer.

A push to this PR makes this attestation stale; the new head needs its own review.

@mrgoonie

mrgoonie commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Review attestation: ready to merge at 7b54e4891d8fe754700d028a4d86a59a4400ae14, reviewed by agent:code-reviewer.

A push to this PR makes this attestation stale; the new head needs its own review.

@mrgoonie
mrgoonie merged commit 24c7afe into main Oct 7, 2026
22 checks passed
@mrgoonie
mrgoonie deleted the feat/444-label-free-spoken-action branch October 7, 2026 21:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant