feat(studio): let an agent author motion - #3520
Conversation
Adds `studio_select` and `studio_seek`, so an agent and the human are looking at the same element and the same instant. Selecting reveals the inspector, exactly as a click does, which is what makes the agent's move visible. Selection is shared state, not a per-call argument, and that is forced rather than chosen. Most of Studio's edit handlers read the ambient React selection, and `applyDomSelection` only schedules a state update, so selecting and committing inside ONE call would write to whatever was selected before. Two tool calls are separated by a render, so the contract is select first, then act. That is also how a human works: click, then type. `studio_seek` uses `requestSeek`, not `setCurrentTime`. The latter only moves the timeline's displayed number and leaves the composition where it was. Two things the tools refuse to fake: Seek does not clamp. `seek()` already clamps against the adapter's duration, which can differ from the store's, and clamping again would give that invariant two owners that can disagree. The tool reports where the playhead actually landed instead, read back afterwards. `requestSeek` is fire-and-forget, so it cannot report that no adapter was mounted to receive it. The tool compares the playhead before and after and fails rather than claiming a seek that never happened. Select separates three failures that a single message would have merged: the preview is not mounted yet (wait), no element matches the handle (re-read), and the element cannot be selected (try a neighbour). The agent's next move differs for each, so collapsing them would cost it a round trip or a retry loop.
Renders the composition to a PNG at a given time and returns the URL. This is what turns the tool set from a remote control into a loop: author a change, capture the instant it affects, look, adjust. No agent can judge motion from source, because "what does this look like at 2.4 seconds" is not a question a file answers. Reuses Studio's existing capture endpoint via `buildFrameCaptureUrl` rather than inventing a second one. Two things this does not fake: It reports the time the playhead LANDED on, not the time requested. The player clamps, so those differ at the ends, and attaching the wrong time to a frame is how an agent draws a confident wrong conclusion about motion. It waits before capturing, by default 150ms. The frame is rendered from the file on disk, and the render cache is cleared by a file watcher with a 40ms write-stability threshold, so a capture that beats the watcher renders the PRE-edit composition. That exact staleness was a real bug here once. An agent reading a stale frame as "my edit failed" would thrash, so the wait is on by default, `settleMs` makes it tunable, and the tool description names the failure rather than leaving it to be rediscovered. It probes with HEAD before returning, so a URL that 404s comes back as a failure with a hint instead of as a link the agent cannot render.
Everything about one element in one call: resolved styles, text fields, box, data attributes, GSAP animations, and what the element will and will not accept. The point is to prevent a failed write rather than to satisfy curiosity. `can.reasonIfDisabled` is passed through verbatim from Studio's own capabilities, so an agent that reads first should never attempt an edit the element would refuse. Three things it refuses to get wrong: Animations are reported ONLY for the current selection, because that is the only element Studio parses them for. Attributing them to any other element would be reporting the wrong element's motion, which is worse than reporting none. When a handle names something else the field is empty and `animationEditingBlocked` says why. `animationEditingBlocked` also carries the two states where animation editing is off entirely, multiple timelines and an unsupported timeline pattern. Both live on the selection context. Learning them from a read costs one call; learning them from a failed write costs a retry loop. Inspecting a handle does NOT change what is selected. It is a read, and stealing the human's selection would be a side effect they did not ask for. There is a test asserting `applySelection` is never called. Nothing selected and no handle given is a failure, not an empty result. An empty result would assert "this element has nothing", which is a different and false claim.
The first tools that change the composition. Both act on the current selection and take no handle, which is forced rather than chosen: the handlers read the ambient React selection, and `applyDomSelection` only schedules a state update, so selecting and committing inside one call would write to whatever was selected before. Select first, then edit. Also plumbs the write-blocked state, which was the blocker for shipping any write at all. `domEditSaveQueuePaused` and the external-file conflict both lived on App and were unreachable from the tool surface, so `canWrite` was optimistic and a comment said so. They now derive into a single `writeBlockedReason` on the shell context: one field, one owner, conflict taking precedence because resolving it is what unblocks the queue. That guard matters more than it looks. Both states are BANNERS in Studio with no lock behind them, so nothing else was stopping a programmatic write from landing on top of a conflict the user had been asked to adjudicate. Three things the tools refuse to fake: They check the outcome, not the absence of a throw. Studio has several paths where a failed commit resolves anyway, so awaiting the handler proves nothing. The tagged outcome added earlier is what proves the write landed. A partial style result is reported as partial. `handleDomStyleCommit` is one property per call, so N properties are N commits; the result carries `applied` and `rejected` maps rather than a single boolean that would have to pick a side. Style commits run sequentially, never concurrently. Two commits racing through Studio's client-side read-modify-write can record undo entries that both claim the same starting content. There is a test that measures concurrency rather than trusting the loop. Every decline reason maps to a hint naming what to do instead, so a refusal routes the agent rather than just stopping it.
`studio_transform` does what a drag does, and then checks. The box in the
result is READ BACK after the write, never echoed from the request, and
`applied` lists what actually took effect.
That is not belt-and-braces. The plan for this unit said to re-derive the
geometry handlers' behaviour rather than trust any description of them, and
doing that turned up three different behaviours behind one interface.
The handlers on `DomEditActionsValue` are the GSAP-AWARE wrappers, aliased in
`useDomEditSession.ts:534-538`, not the CSS ones in `useDomGeometryCommits.ts`
that an earlier note in this workstream described.
`handleGsapAwarePathOffsetCommit` and `handleGsapAwareRotationCommit` are
`if (gsapCommitMutation) { ...intercept... }` with no else branch. Their own
comments say the absence is deliberate: position and rotation are written as
GSAP code and there is no CSS fallback to write to. So they can return having
done nothing.
`handleGsapAwareBoxSizeCommit` is not like the other two. It runs through
`runGestureTransaction` with separate scale and width/height routes, so resize
works more generally.
Reading back is what turns that middle case from a silent lie into a reported
one. A move that did nothing comes back in `unchanged` with a reason.
Three smaller decisions:
Operations re-read between each other, so a move is judged against the box
AFTER a resize in the same call. Comparing against the original would credit
the resize's change to the move.
Rotation is reported as dispatched, not verified. `rotate` is an individual
transform property and does not appear in the computed transform, so there is
no honest box-derived signal, and claiming one would be worse than saying so.
x pairs with y and width pairs with height. Accepting one alone would mean
inventing the other from the current value, which moves the element somewhere
the caller did not ask for. The pairing rule and its minimum live in one
`parsePair` helper rather than as four separate branches.
Four tools: add an animation, change its duration/ease/position, add a keyframe, delete it. This is the capability that makes the tool set worth having, because motion is the one thing an agent cannot judge or author from source. These are deliberately less confident than the rest of the set, and the reason is the handlers underneath them: `handleGsapAddAnimation(method)` takes only a method. Its insert position comes from the live playhead, not the caller, and the call is `void ...catch()` so it returns nothing. `handleGsapAddKeyframeBatch` returns a promise but catches its own failure, so awaiting proves the call finished, not that it landed. `handleGsapDeleteAnimation` discards its promise entirely. `handleGsapUpdateMeta` is the one honest signal. It returns a boolean. U8 handled the same problem by reading the result back. That does not work here: the animation list comes from React state that only refreshes on a render, and no render happens inside one tool call. Rather than fake a verification with a frame-timer, these report what was DISPATCHED and the descriptions tell the agent to call studio_inspect to see the result. Saying "I asked for this" is honest; saying "this happened" would not be. Three consequences worth stating: `studio_add_animation` takes no position. The handler reads the playhead, so accepting one would report a number that had no effect. It reports where the playhead actually was and tells the agent to seek first. `studio_update_animation` rules out the no-selection case BEFORE dispatch. The handler answers `false` for both "nothing selected" and "the write failed", so eliminating one is what makes the other legible. Keyframe percent and properties are validated in the tool, because nothing in the platform checks input against the declared schema.
96cebf2 to
b4123be
Compare
721c271 to
90cef4f
Compare
…3517) Everything about one element in one call: resolved styles, text fields, box, data attributes, GSAP animations, and what the element will and will not accept. The point is to prevent a failed write rather than to satisfy curiosity. `can.reasonIfDisabled` is passed through verbatim from Studio's own capabilities, so an agent that reads first should never attempt an edit the element would refuse. Three things it refuses to get wrong: Animations are reported ONLY for the current selection, because that is the only element Studio parses them for. Attributing them to any other element would be reporting the wrong element's motion, which is worse than reporting none. When a handle names something else the field is empty and `animationEditingBlocked` says why. `animationEditingBlocked` also carries the two states where animation editing is off entirely, multiple timelines and an unsupported timeline pattern. Both live on the selection context. Learning them from a read costs one call; learning them from a failed write costs a retry loop. Inspecting a handle does NOT change what is selected. It is a read, and stealing the human's selection would be a side effect they did not ask for. There is a test asserting `applySelection` is never called. Nothing selected and no handle given is a failure, not an empty result. An empty result would assert "this element has nothing", which is a different and false claim.
Resolve conflicts from squash-merged #3515 (selection tools). Branch retains frame + inspect tools from the stack. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Resolve conflicts from squash-merged #3515 (selection tools). Branch retains the full WebMCP tool stack from the PR chain. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
`studio_transform` does what a drag does, and then checks. The box in the
result is READ BACK after the write, never echoed from the request, and
`applied` lists what actually took effect.
That is not belt-and-braces. The plan for this unit said to re-derive the
geometry handlers' behaviour rather than trust any description of them, and
doing that turned up three different behaviours behind one interface.
The handlers on `DomEditActionsValue` are the GSAP-AWARE wrappers, aliased in
`useDomEditSession.ts:534-538`, not the CSS ones in `useDomGeometryCommits.ts`
that an earlier note in this workstream described.
`handleGsapAwarePathOffsetCommit` and `handleGsapAwareRotationCommit` are
`if (gsapCommitMutation) { ...intercept... }` with no else branch. Their own
comments say the absence is deliberate: position and rotation are written as
GSAP code and there is no CSS fallback to write to. So they can return having
done nothing.
`handleGsapAwareBoxSizeCommit` is not like the other two. It runs through
`runGestureTransaction` with separate scale and width/height routes, so resize
works more generally.
Reading back is what turns that middle case from a silent lie into a reported
one. A move that did nothing comes back in `unchanged` with a reason.
Three smaller decisions:
Operations re-read between each other, so a move is judged against the box
AFTER a resize in the same call. Comparing against the original would credit
the resize's change to the move.
Rotation is reported as dispatched, not verified. `rotate` is an individual
transform property and does not appear in the computed transform, so there is
no honest box-derived signal, and claiming one would be worse than saying so.
x pairs with y and width pairs with height. Accepting one alone would mean
inventing the other from the current value, which moves the element somewhere
the caller did not ask for. The pairing rule and its minimum live in one
`parsePair` helper rather than as four separate branches.
…owser (#3521) * docs: document Studio's WebMCP agent tools Adds `guides/webmcp`, under Developers > Agent setup. Its first job is to defuse a name collision. `guides/mcp` already exists and covers HeyGen's HOSTED MCP connector, which builds a video from a chat. This page is about an agent working inside Studio on a composition already open in front of you. Different feature, confusingly similar name, so the page says what it is not before it says what it is. Written to DOCS_GUIDELINES: one-sentence intro, outcome before implementation, real values rather than placeholders, and three callouts. The three things a reader most needs are the ones easiest to get wrong: The API is `document.modelContext`, not `navigator.modelContext`. Most published examples use the second, which is a polyfill compatibility shim rather than a spec member, so feature-detecting it misleads. Select first, then edit. Most editing tools act on the current selection, and an agent that skips it gets an error rather than a wrong-element write. Leave Studio visible. Some of Studio's write paths report failure through a toast rather than a return value, so the human is the one who sees it. That is a real property of the co-pilot design, not a nicety, so the page says it plainly. Verified with `npx mint validate` and `npx mint broken-links --check-redirects`, both passing. * fix(studio): target the text field that exists, not one named self Found by running the tools end to end in a browser, which is the only way it could have been found: the unit tests mock `setText`, so they never crossed the boundary where this breaks. An element's text usually lives in a CHILD field, keyed like `self:0:h1` or `child:0:h1`. `studio_set_text` passed no field key, so `buildNextDomTextFields` planned zero operations, the request went out with an empty patch, and the server answered: POST /api/projects/<id>/file-mutations/patch-element -> 400 {"error":"target and operations required"} Which surfaced as `persist-failed`. The tool was telling the truth, so the reporting work in the earlier PRs did its job, but the failure looked like a server problem and was not. The tool now resolves the field: the one the caller named, or the element's single field when it has exactly one. An element with several fields is asked to name one; an element with none is reported blocked. Naming a field the element does not have is rejected with the list of the ones it does have, rather than silently writing nowhere. Four regression tests, including the exact `child:0:h1` shape that failed. One existing assertion changed: it expected the field to be `undefined`, which is precisely the bug, so it now expects the resolved key. Also documents two things the browser run surfaced, both real and neither a defect: registration is asynchronous, so a caller reading `getTools()` too early sees a partial list; and the tools that act on the current selection need a render between the select and the edit, which a real agent gets for free because its calls arrive as separate messages. * docs: give the agent-tools kill switch instructions that work The page told readers to set agentToolsEnabled in Studio's preferences. Nothing writes that flag: it is read in useStudioAgentTools and parsed in studioUiPreferences, but there is no settings UI and no toggle, so the instruction could not be followed. Replace it with the localStorage write that actually flips it, and spell out the merge, since overwriting the key drops every other stored preference. * docs: do not promise a per-call permission prompt we have not verified The page said the browser asks before any agent calls a tool. Prompt granularity is browser-specific and unsettled during the origin trial, and we have not observed it on the native path. Say what holds, that access is gated, and name the part that is still moving.
terencecho
left a comment
There was a problem hiding this comment.
APPROVE on the motion commit 90cef4f251393a8b9fe3c96188fa2c6ae36c74c3 (the true delta this PR owns). Reading the head SHA d7a692c9 (which folds in the already-merged #3517/#3519/#3521 layers due to the pending rebase); scope of this stamp is the motion commit only. Two design calls are worth stating out loud, plus one silent-coercion smell that's the only real nit.
1. Dispatched-not-landed is the right shape here — and it's honest about which handler is different.
The docstring's four-way handler table is the point:
| Handler | Signal actually available |
|---|---|
handleGsapAddAnimation(method) |
none: void ...catch() |
handleGsapAddKeyframeBatch |
promise that catches its own failure — arrival ≠ landing |
handleGsapDeleteAnimation |
promise discarded entirely |
handleGsapUpdateMeta |
boolean — the one honest one |
So three tools return dispatched: true and one returns a real success/failure. That asymmetry is faithfully reflected: STUDIO_UPDATE_ANIMATION_DESCRIPTION calls out "the one animation tool that CONFIRMS its write", and the other three descriptions carry the DISPATCH_CAVEAT verbatim. An agent reading tools/list gets the honest contract. Sibling reviewers on #3515 flagged the class-of-bug for read-then-write races; this PR takes the read-back option off the table (React state won't refresh inside one tool call) and picks "tell the model to call studio_inspect to see the result" — the right substitute given the constraint.
2. studio_update_animation's "rule out no-selection BEFORE dispatch" is the load-bearing design decision.
The underlying handleGsapUpdateMeta returns false for both "nothing was selected" and "the write failed on a real target". If both paths could reach the boolean, false would be ambiguous and the tool couldn't return a "failed" kind vs an "invalid" kind honestly. guard(deps) before dispatch drops the no-selection case, so a false afterward is unambiguous → toolFailure("failed", ..., "The animation id may be stale. studio_inspect lists the current ones."). That's what makes kind: "failed" legible to the agent, and the hint points at the right recovery. Test-pinned at line 208-223 (test_hf_component_batch_flag_lookup_is_retained_only_for_pre_retirement_histories equivalent — checks no-selection is caught before dispatch, so updateAnimation mock is never called).
3. studio_add_animation refuses position — and the description explains why.
Right call. handleGsapAddAnimation takes ONLY a method and reads the playhead itself. Accepting a caller-supplied position would let the tool echo back a number that had no effect on the actual insertion point. Instead it reports insertedAtSeconds: currentTime (read-back from the playhead) and the description tells the agent to studio_seek first if it wants a specific time. That's a design decision that would be easy to get wrong by mirroring the schema of a lookalike tool. ✓
4. Silent-drop of wrong-type fields on studio_update_animation misleads the agent's next call. (nit, non-blocker)
studioUpdateAnimation builds updates by dropping any field that fails its type/finite/bound guard:
if (typeof input.duration === "number" && Number.isFinite(input.duration)) {
if (input.duration < 0) return toolFailure("invalid", "duration must not be negative");
updates.duration = input.duration;
}
if (typeof input.ease === "string" && input.ease.trim()) updates.ease = input.ease;
if (typeof input.position === "number" && Number.isFinite(input.position)) {
updates.position = input.position;
}
if (Object.keys(updates).length === 0) {
return toolFailure("invalid", "give at least one of duration, ease, position");
}An agent that sends {animationId: "x", duration: "1.5"} (string instead of number) trips the third-branch check → updates stays empty → "give at least one of duration, ease, position". The agent thinks nothing was given; the agent actually gave a duration in the wrong shape. Next call likely re-sends with the same shape and loops.
The keyframe path has the same pattern for individual property values (non-string/non-number silently dropped, only "all dropped" surfaces as an error), though the properties must contain at least one number or string value reason is slightly more actionable there because the property vocabulary is open.
Concrete fix that would eliminate the loop-inducing UX: when input.duration !== undefined && (typeof input.duration !== "number" || !Number.isFinite(input.duration)), surface toolFailure("invalid", "duration must be a finite number"). Same for ease and position. Then wrong-type is distinguishable from missing, and the agent's next call fixes the actual shape mismatch.
Not a blocker — the sibling read tool (studio_inspect) gives the agent the actual schema by example, so a well-guided agent recovers. But it's the "silent coercion masks typos" pattern in miniature, and the fix is a couple of lines.
5. STUDIO_UPDATE_ANIMATION_INPUT_SCHEMA position: number description — GSAP semantics ambiguous. (low-severity)
Description reads "Start position in seconds". GSAP positions are relative-to-previous offsets or absolute times depending on parent context, and are traditionally allowed to be tokens like "<" or ">". This tool declares type: "number" and reads only numeric positions, so token positions aren't reachable — that's fine. But an agent reading the description will believe "position" is absolute seconds; the actual handler behavior depends on Studio's timeline setup, which isn't exposed. Worth a description tweak like "Start position in the timeline (numeric only; use studio_add_animation + studio_seek for absolute placement)" — makes the constraint explicit rather than implicit.
6. Confirmed:
studioDeleteAnimationreturns{animationId, dispatched: true}but its test only assertsresult.ok === true. Addingexpect(...dispatched).toBe(true)would pin the field-shape contract; the current test still catches a regression where the tool returns something completely different. Minor coverage nit. ✓ (functionally covered)- Method whitelist
METHODS = ["to", "from", "set", "fromTo"]— matches GSAP's four core methods; anything else routes totoolFailure("invalid"). Test at line 149 asserts the whitelist and that the dispatch mock is never called. ✓ - Percent test at 0 and 100 — inclusive bounds asserted at line 296-305. Boundary-covered. ✓
- String "50" rejected as percent (line 274) — schema-vs-actual mismatch caught, exactly the "platform doesn't validate inputSchema" defense. ✓
- Idempotent registration across re-renders (lines 646-655) — count assertion 8 → 12 pinned. ✓
- Keyframe one-commit-one-undo (line 400-405) —
addKeyframecalled ONCE with all properties, not per-property; test assertstoHaveBeenCalledTimes(1). Batch-safe undo model preserved. ✓ - StudioAgentTools deps array updated to include all four new handlers — line 76-80. useMemo would otherwise return stale closures. ✓
Rebase mechanics: PR head is d7a692c9 which folds in the just-merged #3521 docs commit — mergeStateStatus: DIRTY reflects the pending rebase onto master (which now has #3515/#3517/#3519/#3521 landed). Only Preflight (lint + format) and WIP visible in the check rollup so far — the full test/tsc/lint matrix hasn't run at this head. Post-rebase, the head SHA changes and require_last_push_approval=true on the public repo will dismiss this approval; a fresh re-approve will be needed at the rebased head. Signaling my read at this SHA so re-approval on the rebased head is a fast concur (assuming no code drift beyond the automatic rebase).
Review by tai (pr-review)
Resolve conflicts from parallel squash merges in the stack. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Resolve conflicts from parallel squash merges in the stack. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…ntent # Conflicts: # packages/studio/src/webmcp/StudioAgentTools.tsx # packages/studio/src/webmcp/useStudioAgentTools.test.tsx # packages/studio/src/webmcp/useStudioAgentTools.ts
… feat/studio-webmcp-animation
…imation # Conflicts: # packages/studio/src/webmcp/StudioAgentTools.tsx # packages/studio/src/webmcp/tools/contentTools.test.ts # packages/studio/src/webmcp/tools/contentTools.ts # packages/studio/src/webmcp/useStudioAgentTools.test.tsx # packages/studio/src/webmcp/useStudioAgentTools.ts
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Workflows to automatically generate PRs for you. |
terencecho
left a comment
There was a problem hiding this comment.
Concur — re-approving at 2f2ca158 after the rebase.
Verified the animation-tool delta this PR owns is byte-identical to the previously-reviewed head: the feat(studio): let an agent author motion commit is still 90cef4f2513..., unchanged in place. The 36-commit compare vs my prior head d7a692c9 is entirely (a) main commits pulled in by the rebase (0.8.17 → 0.8.20 releases + assorted CLI/core/producer fixes + the hw-write-title registry addition) and (b) the sibling WebMCP stack fold-ins now that #3515 / #3516 / #3517 / #3518 / #3519 have landed. Nothing touches this PR's own animation-tool surface.
CI still in progress at head — approval doesn't force merge; auto-merge (already armed) gates on green as expected.
Prior review 5060096243 carries.
Review by tai (pr-review)
What
Four tools that author motion:
studio_add_animation,studio_update_animation,studio_add_keyframe,studio_delete_animation.Stacked on #3519. This completes the tool set; only the docs page and the browser proof remain.
Why
Motion is the one thing an agent cannot judge or author from source. Reading GSAP code tells you the tween exists; it does not tell you the title lands a beat late or the ease feels mushy. Paired with
studio_seekandstudio_framefrom earlier in the stack, these close that loop.How, and why these are less confident than the rest
The handlers underneath mostly cannot report back:
handleGsapAddAnimation(method)void ...catch()handleGsapAddKeyframeBatchhandleGsapDeleteAnimationhandleGsapUpdateMeta#3519 solved the same problem by reading the result back. That does not work here: the animation list comes from React state that only refreshes on a render, and no render happens inside one tool call.
So rather than fake a verification with a frame-timer, these report what was dispatched, and every description tells the agent to call
studio_inspectto see the result. Saying "I asked for this" is honest; saying "this happened" would not be.Three consequences worth stating:
studio_add_animationtakes no position. The handler reads the playhead itself, so accepting one would report a number that had no effect. It reports where the playhead actually was and tells the agent tostudio_seekfirst.studio_update_animationrules out the no-selection case before dispatch. The handler answersfalsefor both "nothing selected" and "the write failed". Eliminating one beforehand is what makes the other legible, and it is why afalsefrom this tool is reported as a real failure with a stale-id hint.Keyframe percent and properties are validated in the tool, because nothing in the platform checks input against the declared
inputSchema.Test plan
15 tests in
animationTools.test.ts:dispatched, not claimed as landed.falsebecomes a real failure with a stale-id hint.falsecan never be reached that way."50"; properties rejected for{},{y: null},[]and a string; 0 and 100 accepted as the ends.Full package suite 4602 passing across 414 files.
bunx tsc --noEmit,bunx oxlintandbunx fallow audit --fail-on-issuesall clean.Still owed
Nothing in this stack has been exercised in a real browser. That is the last unit, and it needs
chrome://flags/#enable-webmcp-testing.