From eb6201f09274d41bb4a569cf25e45e5e70b2d0cd Mon Sep 17 00:00:00 2001 From: amo-tech-ai Date: Wed, 16 Sep 2026 16:35:35 -0500 Subject: [PATCH 1/2] feat(skills): extract canonical MDE skill foundation --- .claude/skills/cloudinary/SKILL.md | 33 + .claude/skills/cloudinary/evals/evals.json | 32 + .claude/skills/copilotkit | 1 - .claude/skills/copilotkit/SKILL.md | 57 + .claude/skills/copilotkit/evals/evals.json | 48 + .../copilotkit/references/ag-ui-and-tools.md | 7 + .../copilotkit/references/mastra-bridge.md | 11 + .../official/copilotkit-cli/SKILL.md | 141 ++ .../references/official/copilotkit/SKILL.md | 98 ++ .../references/runtime-and-react.md | 12 + .claude/skills/copilotkit/upstream.yaml | 26 + .claude/skills/events/SKILL.md | 29 + .claude/skills/events/evals/evals.json | 32 + .claude/skills/gemini | 1 - .claude/skills/gemini/SKILL.md | 56 + .claude/skills/gemini/evals/evals.json | 36 + .../official/gemini-api-dev/SKILL.md | 435 +++++++ .../gemini-api-dev/references/migration.md | 122 ++ .../official/gemini-live-api-dev/SKILL.md | 399 ++++++ .../official/gemini-omni-flash-api/SKILL.md | 450 +++++++ .../scripts/upload_file.py | 321 +++++ .../scripts/video/generate_video.py | 734 +++++++++++ .../scripts/video/inspect_video.py | 197 +++ .../scripts/video/prep_video.py | 255 ++++ .claude/skills/gemini/upstream.yaml | 21 + .claude/skills/maps/SKILL.md | 284 ++++ .../maps/references/gmaps-cli-behavior.md | 16 + .../maps/references/react-vis-gl/README.md | 19 + .../references/react-vis-gl/components-api.md | 272 ++++ .../react-vis-gl/geometry-components.md | 495 +++++++ .../maps/references/react-vis-gl/hooks-api.md | 305 +++++ .../maps/references/react-vis-gl/patterns.md | 577 +++++++++ .../react-vis-gl/places-autocomplete.md | 476 +++++++ .../references/security-and-optimization.md | 266 ++++ .../vendor/google-maps-platform/SKILL.md | 454 +++++++ .../vendor/google-maps-platform/UPSTREAM.md | 9 + .claude/skills/maps/scripts/gmaps.py | 983 ++++++++++++++ .claude/skills/mastra/SKILL.md | 323 +---- .claude/skills/mastra/links.md | 4 +- .claude/skills/mastra/references/README.md | 2 +- .claude/skills/mastra/references/browser.md | 2 +- .../skills/mastra/references/copilotkit.md | 2 +- .../skills/mastra/references/examples-v0.md | 2 +- .../mastra/references/mcp-docs-lookup.md | 2 +- .../references/official/mastra/SKILL.md | 133 ++ .../mastra/references/common-errors.md | 537 ++++++++ .../mastra/references/core-concepts.md | 17 + .../mastra/references/create-mastra.md | 222 ++++ .../mastra/references/embedded-docs.md | 103 ++ .../official/mastra/references/mastra-api.md | 201 +++ .../mastra/references/migration-guide.md | 180 +++ .../mastra/references/model-selection.md | 24 + .../official/mastra/references/remote-docs.md | 193 +++ .../mastra/references/trace-intelligence.md | 220 ++++ .../mastra/scripts/provider-registry.mjs | 180 +++ .../skills/mastra/references/rag-pgvector.md | 2 +- .claude/skills/mastra/references/streaming.md | 2 +- .../skills/mastra/references/topic-routing.md | 2 +- .claude/skills/mastra/references/workflows.md | 2 +- .../mastra/references/workspace-skills.md | 4 +- .claude/skills/mastra/references/workspace.md | 2 +- .claude/skills/mastra/upstream.yaml | 16 + .claude/skills/nextjs/SKILL.md | 34 + .claude/skills/nextjs/evals/evals.json | 32 + .claude/skills/real-estate/SKILL.md | 34 + .claude/skills/real-estate/gemini | 1 + .../skills/real-estate/industry-context.md | 781 +++++++++++ .claude/skills/real-estate/marketplace-v1.md | 132 ++ .claude/skills/real-estate/mls-v2.md | 89 ++ .../real-estate/sub-agents/lead-qualifier.md | 257 ++++ .../sub-agents/neighborhood-guide.md | 339 +++++ .../sub-agents/property-description.md | 241 ++++ .claude/skills/stripe/SKILL.md | 36 + .claude/skills/stripe/evals/evals.json | 32 + .claude/skills/supabase/SKILL.md | 60 + .claude/skills/supabase/better-auth.md | 141 ++ .claude/skills/supabase/client-and-auth.md | 113 ++ .claude/skills/supabase/edge-functions.md | 1150 +++++++++++++++++ .claude/skills/supabase/evals/evals.json | 36 + .claude/skills/supabase/postgres.md | 65 + .claude/skills/supabase/realtime.md | 269 ++++ .../supabase/references/ai-edge-functions.md | 47 + .../references/docs/rls-row-level-security.md | 416 ++++++ .../supabase/references/docs/testing.md | 57 + .../references/edge-functions-inventory.md | 79 ++ .../CHANGELOG.md | 73 ++ .../supabase-postgres-best-practices/SKILL.md | 64 + .../references/_contributing.md | 170 +++ .../references/_sections.md | 39 + .../references/_template.md | 34 + .../references/advanced-full-text-search.md | 55 + .../references/advanced-jsonb-indexing.md | 49 + .../references/conn-idle-timeout.md | 46 + .../references/conn-limits.md | 44 + .../references/conn-pooling.md | 41 + .../references/conn-prepared-statements.md | 46 + .../references/data-batch-inserts.md | 54 + .../references/data-n-plus-one.md | 53 + .../references/data-pagination.md | 50 + .../references/data-upsert.md | 50 + .../references/lock-advisory.md | 56 + .../references/lock-deadlock-prevention.md | 68 + .../references/lock-short-transactions.md | 50 + .../references/lock-skip-locked.md | 54 + .../references/monitor-explain-analyze.md | 45 + .../references/monitor-pg-stat-statements.md | 55 + .../references/monitor-vacuum-analyze.md | 55 + .../references/query-composite-indexes.md | 44 + .../references/query-covering-indexes.md | 40 + .../references/query-index-types.md | 48 + .../references/query-missing-indexes.md | 43 + .../references/query-partial-indexes.md | 45 + .../references/schema-constraints.md | 80 ++ .../references/schema-data-types.md | 46 + .../references/schema-foreign-key-indexes.md | 59 + .../schema-lowercase-identifiers.md | 55 + .../references/schema-partitioning.md | 55 + .../references/schema-primary-keys.md | 61 + .../references/security-privileges.md | 54 + .../references/security-rls-basics.md | 50 + .../references/security-rls-performance.md | 63 + .../references/official/supabase/CHANGELOG.md | 78 ++ .../references/official/supabase/SKILL.md | 149 +++ .../assets/feedback-issue-template.md | 17 + .../supabase/references/skill-feedback.md | 17 + .../references/postgres/_contributing.md | 177 +++ .../supabase/references/postgres/_sections.md | 46 + .../supabase/references/postgres/_template.md | 34 + .../postgres/advanced-full-text-search.md | 55 + .../postgres/advanced-jsonb-indexing.md | 49 + .../references/postgres/conn-idle-timeout.md | 46 + .../references/postgres/conn-limits.md | 44 + .../references/postgres/conn-pooling.md | 41 + .../postgres/conn-prepared-statements.md | 46 + .../references/postgres/data-batch-inserts.md | 54 + .../references/postgres/data-n-plus-one.md | 53 + .../references/postgres/data-pagination.md | 50 + .../references/postgres/data-upsert.md | 50 + .../references/postgres/lock-advisory.md | 56 + .../postgres/lock-deadlock-prevention.md | 68 + .../postgres/lock-short-transactions.md | 50 + .../references/postgres/lock-skip-locked.md | 54 + .../postgres/monitor-explain-analyze.md | 45 + .../postgres/monitor-pg-stat-statements.md | 55 + .../postgres/monitor-vacuum-analyze.md | 55 + .../postgres/query-composite-indexes.md | 44 + .../postgres/query-covering-indexes.md | 40 + .../references/postgres/query-index-types.md | 48 + .../postgres/query-missing-indexes.md | 43 + .../postgres/query-partial-indexes.md | 45 + .../references/postgres/schema-constraints.md | 80 ++ .../references/postgres/schema-data-types.md | 46 + .../postgres/schema-foreign-key-indexes.md | 59 + .../postgres/schema-lowercase-identifiers.md | 55 + .../postgres/schema-partitioning.md | 55 + .../postgres/schema-primary-keys.md | 61 + .../postgres/security-privileges.md | 54 + .../postgres/security-rls-basics.md | 50 + .../postgres/security-rls-performance.md | 57 + .../supabase-database-functions.md | 130 ++ .../supabase-declarative-schema.md | 78 ++ .../project-rules/supabase-edge-functions.md | 110 ++ .../project-rules/supabase-migrations.md | 49 + .../project-rules/supabase-patterns.md | 50 + .../project-rules/supabase-realtime.md | 463 +++++++ .../project-rules/supabase-rls-policies.md | 251 ++++ .../project-rules/supabase-sql-style.md | 133 ++ .../migration-from-postgres-changes.md | 154 +++ .../realtime/rls-policy-cookbook.md | 199 +++ .../references/storage/api-cheatsheet.md | 120 ++ .../references/storage/rls-policies.md | 138 ++ .../assets/feedback-issue-template.md | 24 + .../references/supabase/skill-feedback.md | 25 + .../supabase/references/tables-overview.md | 28 + .../supabase/scripts/verify-edge-inventory.sh | 50 + .claude/skills/supabase/storage.md | 308 +++++ .claude/skills/supabase/upstream.yaml | 17 + 177 files changed, 20680 insertions(+), 295 deletions(-) create mode 100644 .claude/skills/cloudinary/SKILL.md create mode 100644 .claude/skills/cloudinary/evals/evals.json delete mode 120000 .claude/skills/copilotkit create mode 100644 .claude/skills/copilotkit/SKILL.md create mode 100644 .claude/skills/copilotkit/evals/evals.json create mode 100644 .claude/skills/copilotkit/references/ag-ui-and-tools.md create mode 100644 .claude/skills/copilotkit/references/mastra-bridge.md create mode 100644 .claude/skills/copilotkit/references/official/copilotkit-cli/SKILL.md create mode 100644 .claude/skills/copilotkit/references/official/copilotkit/SKILL.md create mode 100644 .claude/skills/copilotkit/references/runtime-and-react.md create mode 100644 .claude/skills/copilotkit/upstream.yaml create mode 100644 .claude/skills/events/SKILL.md create mode 100644 .claude/skills/events/evals/evals.json delete mode 120000 .claude/skills/gemini create mode 100644 .claude/skills/gemini/SKILL.md create mode 100644 .claude/skills/gemini/evals/evals.json create mode 100644 .claude/skills/gemini/references/official/gemini-api-dev/SKILL.md create mode 100644 .claude/skills/gemini/references/official/gemini-api-dev/references/migration.md create mode 100644 .claude/skills/gemini/references/official/gemini-live-api-dev/SKILL.md create mode 100644 .claude/skills/gemini/references/official/gemini-omni-flash-api/SKILL.md create mode 100755 .claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/upload_file.py create mode 100755 .claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/generate_video.py create mode 100755 .claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/inspect_video.py create mode 100755 .claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/prep_video.py create mode 100644 .claude/skills/gemini/upstream.yaml create mode 100644 .claude/skills/maps/SKILL.md create mode 100644 .claude/skills/maps/references/gmaps-cli-behavior.md create mode 100644 .claude/skills/maps/references/react-vis-gl/README.md create mode 100644 .claude/skills/maps/references/react-vis-gl/components-api.md create mode 100644 .claude/skills/maps/references/react-vis-gl/geometry-components.md create mode 100644 .claude/skills/maps/references/react-vis-gl/hooks-api.md create mode 100644 .claude/skills/maps/references/react-vis-gl/patterns.md create mode 100644 .claude/skills/maps/references/react-vis-gl/places-autocomplete.md create mode 100644 .claude/skills/maps/references/security-and-optimization.md create mode 100644 .claude/skills/maps/references/vendor/google-maps-platform/SKILL.md create mode 100644 .claude/skills/maps/references/vendor/google-maps-platform/UPSTREAM.md create mode 100755 .claude/skills/maps/scripts/gmaps.py create mode 100644 .claude/skills/mastra/references/official/mastra/SKILL.md create mode 100644 .claude/skills/mastra/references/official/mastra/references/common-errors.md create mode 100644 .claude/skills/mastra/references/official/mastra/references/core-concepts.md create mode 100644 .claude/skills/mastra/references/official/mastra/references/create-mastra.md create mode 100644 .claude/skills/mastra/references/official/mastra/references/embedded-docs.md create mode 100644 .claude/skills/mastra/references/official/mastra/references/mastra-api.md create mode 100644 .claude/skills/mastra/references/official/mastra/references/migration-guide.md create mode 100644 .claude/skills/mastra/references/official/mastra/references/model-selection.md create mode 100644 .claude/skills/mastra/references/official/mastra/references/remote-docs.md create mode 100644 .claude/skills/mastra/references/official/mastra/references/trace-intelligence.md create mode 100755 .claude/skills/mastra/references/official/mastra/scripts/provider-registry.mjs create mode 100644 .claude/skills/mastra/upstream.yaml create mode 100644 .claude/skills/nextjs/SKILL.md create mode 100644 .claude/skills/nextjs/evals/evals.json create mode 100644 .claude/skills/real-estate/SKILL.md create mode 120000 .claude/skills/real-estate/gemini create mode 100644 .claude/skills/real-estate/industry-context.md create mode 100644 .claude/skills/real-estate/marketplace-v1.md create mode 100644 .claude/skills/real-estate/mls-v2.md create mode 100644 .claude/skills/real-estate/sub-agents/lead-qualifier.md create mode 100644 .claude/skills/real-estate/sub-agents/neighborhood-guide.md create mode 100644 .claude/skills/real-estate/sub-agents/property-description.md create mode 100644 .claude/skills/stripe/SKILL.md create mode 100644 .claude/skills/stripe/evals/evals.json create mode 100644 .claude/skills/supabase/SKILL.md create mode 100644 .claude/skills/supabase/better-auth.md create mode 100644 .claude/skills/supabase/client-and-auth.md create mode 100644 .claude/skills/supabase/edge-functions.md create mode 100644 .claude/skills/supabase/evals/evals.json create mode 100644 .claude/skills/supabase/postgres.md create mode 100644 .claude/skills/supabase/realtime.md create mode 100644 .claude/skills/supabase/references/ai-edge-functions.md create mode 100644 .claude/skills/supabase/references/docs/rls-row-level-security.md create mode 100644 .claude/skills/supabase/references/docs/testing.md create mode 100644 .claude/skills/supabase/references/edge-functions-inventory.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/CHANGELOG.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/SKILL.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/_contributing.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/_sections.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/_template.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/advanced-full-text-search.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/advanced-jsonb-indexing.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/conn-idle-timeout.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/conn-limits.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/conn-pooling.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/conn-prepared-statements.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/data-batch-inserts.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/data-n-plus-one.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/data-pagination.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/data-upsert.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/lock-advisory.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/lock-deadlock-prevention.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/lock-short-transactions.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/lock-skip-locked.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/monitor-explain-analyze.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/monitor-pg-stat-statements.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/monitor-vacuum-analyze.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/query-composite-indexes.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/query-covering-indexes.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/query-index-types.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/query-missing-indexes.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/query-partial-indexes.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/schema-constraints.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/schema-data-types.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/schema-foreign-key-indexes.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/schema-lowercase-identifiers.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/schema-partitioning.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/schema-primary-keys.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/security-privileges.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/security-rls-basics.md create mode 100644 .claude/skills/supabase/references/official/supabase-postgres-best-practices/references/security-rls-performance.md create mode 100644 .claude/skills/supabase/references/official/supabase/CHANGELOG.md create mode 100644 .claude/skills/supabase/references/official/supabase/SKILL.md create mode 100644 .claude/skills/supabase/references/official/supabase/assets/feedback-issue-template.md create mode 100644 .claude/skills/supabase/references/official/supabase/references/skill-feedback.md create mode 100644 .claude/skills/supabase/references/postgres/_contributing.md create mode 100644 .claude/skills/supabase/references/postgres/_sections.md create mode 100644 .claude/skills/supabase/references/postgres/_template.md create mode 100644 .claude/skills/supabase/references/postgres/advanced-full-text-search.md create mode 100644 .claude/skills/supabase/references/postgres/advanced-jsonb-indexing.md create mode 100644 .claude/skills/supabase/references/postgres/conn-idle-timeout.md create mode 100644 .claude/skills/supabase/references/postgres/conn-limits.md create mode 100644 .claude/skills/supabase/references/postgres/conn-pooling.md create mode 100644 .claude/skills/supabase/references/postgres/conn-prepared-statements.md create mode 100644 .claude/skills/supabase/references/postgres/data-batch-inserts.md create mode 100644 .claude/skills/supabase/references/postgres/data-n-plus-one.md create mode 100644 .claude/skills/supabase/references/postgres/data-pagination.md create mode 100644 .claude/skills/supabase/references/postgres/data-upsert.md create mode 100644 .claude/skills/supabase/references/postgres/lock-advisory.md create mode 100644 .claude/skills/supabase/references/postgres/lock-deadlock-prevention.md create mode 100644 .claude/skills/supabase/references/postgres/lock-short-transactions.md create mode 100644 .claude/skills/supabase/references/postgres/lock-skip-locked.md create mode 100644 .claude/skills/supabase/references/postgres/monitor-explain-analyze.md create mode 100644 .claude/skills/supabase/references/postgres/monitor-pg-stat-statements.md create mode 100644 .claude/skills/supabase/references/postgres/monitor-vacuum-analyze.md create mode 100644 .claude/skills/supabase/references/postgres/query-composite-indexes.md create mode 100644 .claude/skills/supabase/references/postgres/query-covering-indexes.md create mode 100644 .claude/skills/supabase/references/postgres/query-index-types.md create mode 100644 .claude/skills/supabase/references/postgres/query-missing-indexes.md create mode 100644 .claude/skills/supabase/references/postgres/query-partial-indexes.md create mode 100644 .claude/skills/supabase/references/postgres/schema-constraints.md create mode 100644 .claude/skills/supabase/references/postgres/schema-data-types.md create mode 100644 .claude/skills/supabase/references/postgres/schema-foreign-key-indexes.md create mode 100644 .claude/skills/supabase/references/postgres/schema-lowercase-identifiers.md create mode 100644 .claude/skills/supabase/references/postgres/schema-partitioning.md create mode 100644 .claude/skills/supabase/references/postgres/schema-primary-keys.md create mode 100644 .claude/skills/supabase/references/postgres/security-privileges.md create mode 100644 .claude/skills/supabase/references/postgres/security-rls-basics.md create mode 100644 .claude/skills/supabase/references/postgres/security-rls-performance.md create mode 100644 .claude/skills/supabase/references/project-rules/supabase-database-functions.md create mode 100644 .claude/skills/supabase/references/project-rules/supabase-declarative-schema.md create mode 100644 .claude/skills/supabase/references/project-rules/supabase-edge-functions.md create mode 100644 .claude/skills/supabase/references/project-rules/supabase-migrations.md create mode 100644 .claude/skills/supabase/references/project-rules/supabase-patterns.md create mode 100644 .claude/skills/supabase/references/project-rules/supabase-realtime.md create mode 100644 .claude/skills/supabase/references/project-rules/supabase-rls-policies.md create mode 100644 .claude/skills/supabase/references/project-rules/supabase-sql-style.md create mode 100644 .claude/skills/supabase/references/realtime/migration-from-postgres-changes.md create mode 100644 .claude/skills/supabase/references/realtime/rls-policy-cookbook.md create mode 100644 .claude/skills/supabase/references/storage/api-cheatsheet.md create mode 100644 .claude/skills/supabase/references/storage/rls-policies.md create mode 100644 .claude/skills/supabase/references/supabase/assets/feedback-issue-template.md create mode 100644 .claude/skills/supabase/references/supabase/skill-feedback.md create mode 100644 .claude/skills/supabase/references/tables-overview.md create mode 100755 .claude/skills/supabase/scripts/verify-edge-inventory.sh create mode 100644 .claude/skills/supabase/storage.md create mode 100644 .claude/skills/supabase/upstream.yaml diff --git a/.claude/skills/cloudinary/SKILL.md b/.claude/skills/cloudinary/SKILL.md new file mode 100644 index 00000000..16d76e37 --- /dev/null +++ b/.claude/skills/cloudinary/SKILL.md @@ -0,0 +1,33 @@ +--- +name: cloudinary +description: >- + Use when MDE work explicitly introduces or changes Cloudinary uploads, assets, transformations, delivery URLs, signatures, webhooks, or media lifecycle behavior. +--- + +# Cloudinary + +Own Cloudinary-specific media behavior. Do not assume Cloudinary is active merely because the project may use media assets. + +## Current-state rule + +The current repo has no direct Cloudinary dependency or active source reference in the audited application paths. First verify that the task actually uses Cloudinary. If not, do not introduce it as architecture by default. + +## Invariants + +- Keep API secrets and signing secrets server-side. +- Prefer signed uploads/transformations when the operation requires trust. +- Validate resource ownership before destructive asset changes. +- Keep durable application metadata in the application database; Cloudinary is media infrastructure, not the business source of truth. +- Verify webhook signatures before applying side effects. + +## Workflow + +1. Confirm Cloudinary is in scope and inspect any existing integration. +2. Verify the current official API/SDK contract. +3. Define upload, transformation, deletion, and failure behavior. +4. Implement only the required integration surface. +5. Test invalid signatures, missing assets, and retry behavior when applicable. + +## Handoff + +Use `nextjs` for framework upload routes, `supabase` for application metadata/authorization, and the owning domain skill for product behavior. diff --git a/.claude/skills/cloudinary/evals/evals.json b/.claude/skills/cloudinary/evals/evals.json new file mode 100644 index 00000000..d84692fa --- /dev/null +++ b/.claude/skills/cloudinary/evals/evals.json @@ -0,0 +1,32 @@ +{ + "version": 1, + "evals": [ + { + "prompt": "Handle a focused cloudinary change in MDE and identify the correct implementation boundary.", + "expected_output": "Uses cloudinary as the primary specialist and keeps unrelated domains out.", + "expectations": [ + "inspects current repo state", + "uses current official docs for version-sensitive behavior", + "makes the smallest safe change" + ] + }, + { + "prompt": "A bug appears near cloudinary, but the root cause is not known. What owns diagnosis?", + "expected_output": "Routes diagnosis to systematic-debugging while retaining the specialist for domain evidence.", + "expectations": [ + "does not guess root cause", + "uses systematic-debugging for diagnosis", + "keeps specialist scope bounded" + ] + }, + { + "prompt": "A production-critical cloudinary change is ready to ship. What proof is required?", + "expected_output": "Requires targeted tests and independent task verification before Done.", + "expectations": [ + "requires current evidence", + "does not self-certify", + "uses task-verifier for completion proof" + ] + } + ] +} diff --git a/.claude/skills/copilotkit b/.claude/skills/copilotkit deleted file mode 120000 index 37d9aa42..00000000 --- a/.claude/skills/copilotkit +++ /dev/null @@ -1 +0,0 @@ -../../../.agents/skills/copilotkit \ No newline at end of file diff --git a/.claude/skills/copilotkit/SKILL.md b/.claude/skills/copilotkit/SKILL.md new file mode 100644 index 00000000..7d22c623 --- /dev/null +++ b/.claude/skills/copilotkit/SKILL.md @@ -0,0 +1,57 @@ +--- +name: copilotkit +description: >- + Use when MDE work involves CopilotKit, /api/copilotkit, CopilotKit v2 React hooks, generative UI, frontend tools or actions, shared agent state, AG-UI transport, runtime wiring, CLI verification, HITL, or the CopilotKit-to-Mastra bridge. +metadata: + mde-version: "2.0.0" + upstream-commit: "8a7446186cd3e0d368ec885e61c5913f0918ef5d" + verified-package: "@copilotkit/react-core 1.55.2 / @copilotkit/runtime 1.55.2" +--- + +# CopilotKit — official upstream + MDE overlay + +## Source order + +1. Inspect the installed MDE CopilotKit packages and current runtime/provider code. +2. Read the pinned official CopilotKit core skill in `references/official/copilotkit/SKILL.md`. +3. For wiring/debugging, also read `references/official/copilotkit-cli/SKILL.md` and run `npx copilotkit@latest verify --json` when safe and applicable. +4. Use current official CopilotKit/AG-UI docs or source when the pinned skill directs you there. +5. Apply the MDE-specific invariants below. + +Do not answer volatile CopilotKit API questions from memory. Do not silently replace installed-version behavior with latest-main examples. + +## Ownership + +Own the browser-facing agent bridge: provider/hooks, same-origin runtime, AG-UI events, frontend tool/action registration, shared state, generative UI, and CopilotKit-visible failures. `mastra` owns agents/tools/workflows/memory/storage/HITL semantics. Product-domain skills own business invariants. +## Current MDE invariants + +- MDE uses CopilotKit v2 React APIs from `@copilotkit/react-core/v2`. +- Browser traffic uses the same-origin `/api/copilotkit` runtime; do not switch to hosted Intelligence implicitly. +- Tool render/action names must match the Mastra tool-map key, not a `createTool()` id. +- Preserve auth, distributed rate limits, request context, telemetry, and agent allowlists on the runtime route. +- Treat AG-UI messages/events as the frontend-agent transport contract; do not create parallel ad-hoc chat state. +- Keep provider props stable across renders. + +## Workflow + +1. Classify the issue: wiring/runtime, React/provider, AG-UI/tool rendering, shared state, CLI verification, or Mastra bridge. +2. Use the official core/CLI skill as the default procedure and lookup guide. +3. Load only the matching MDE reference: `runtime-and-react.md`, `ag-ui-and-tools.md`, or `mastra-bridge.md`. +4. Reproduce before editing when debugging. +5. Keep the smallest safe contract change and hand Mastra-internal changes to `mastra`. +6. Prove the affected contract with targeted tests plus browser/stream evidence when user-visible behavior changed. + +## Verification + +At minimum verify the affected runtime URL, agent identity, version surface, tool/action mapping, auth/rate-limit behavior, AG-UI result/state handling, and rendered user path. A passing CopilotKit CLI wiring check does not prove tool execution, streaming order, state synchronization, or rendered generative UI; test those separately. + +## References + +- `upstream.yaml` — immutable source provenance and update policy +- `references/official/copilotkit/SKILL.md` — official vendor skill, pinned and read-only +- `references/official/copilotkit-cli/SKILL.md` — official CLI skill, pinned and read-only +- `references/runtime-and-react.md` +- `references/ag-ui-and-tools.md` +- `references/mastra-bridge.md` +- https://docs.copilotkit.ai/ +- https://github.com/CopilotKit/CopilotKit diff --git a/.claude/skills/copilotkit/evals/evals.json b/.claude/skills/copilotkit/evals/evals.json new file mode 100644 index 00000000..2cf86b8a --- /dev/null +++ b/.claude/skills/copilotkit/evals/evals.json @@ -0,0 +1,48 @@ +{ + "skill_name": "copilotkit", + "evals": [ + { + "id": 1, + "prompt": "The /api/copilotkit stream is returning 500 after a runtime change. Diagnose the CopilotKit side without rewriting Mastra tools.", + "expected_output": "Routes to CopilotKit plus systematic debugging; verifies installed v2/runtime wiring and current official sources before proposing a change.", + "files": [], + "expectations": [ + "Uses CopilotKit for the browser/runtime boundary", + "Verifies installed/current source rather than answering volatile API details from memory", + "Does not rewrite Mastra tool business logic" + ] + }, + { + "id": 2, + "prompt": "A Mastra search tool executes successfully but the CopilotKit v2 card never renders. What should we verify?", + "expected_output": "Checks frontend action/tool registration against the Mastra tool-map key, AG-UI result flow, and rendered browser path.", + "files": [], + "expectations": [ + "Checks the CopilotKit action/tool-name contract", + "Checks AG-UI/tool result handling", + "Requires browser-visible proof, not only a tool unit test" + ] + }, + { + "id": 3, + "prompt": "The event host publishing rule needs to change from draft-only to draft-or-reviewed.", + "expected_output": "Does not select CopilotKit as the owner because this is a product-domain rule with no CopilotKit dependency.", + "files": [], + "expectations": [ + "Does not treat CopilotKit as the primary owner", + "Routes to the relevant product/domain workflow instead" + ] + }, + { + "id": 4, + "prompt": "Can we copy a CopilotKit main-branch example into our app? We are on @copilotkit/react-core 1.55.2.", + "expected_output": "Checks installed package/version and pinned/current official source first; preserves v2 import/runtime contract and refuses blind latest-main copying.", + "files": [], + "expectations": [ + "Checks installed version before using examples", + "Does not blindly substitute latest-main API behavior", + "Preserves the v2 surface when applicable" + ] + } + ] +} diff --git a/.claude/skills/copilotkit/references/ag-ui-and-tools.md b/.claude/skills/copilotkit/references/ag-ui-and-tools.md new file mode 100644 index 00000000..82ded0b4 --- /dev/null +++ b/.claude/skills/copilotkit/references/ag-ui-and-tools.md @@ -0,0 +1,7 @@ +# AG-UI and frontend tools + +CopilotKit v2 uses AG-UI as the agent/user interaction transport. Treat text streaming, tool calls/results, and state snapshots/deltas as protocol contracts. + +MDE-specific rule: CopilotKit tool/action names align to Mastra tool map keys. Verify `src/platform/copilot/mastra-tool-action-names.ts` and its tests before changing render/action registration. + +When tool result envelopes or state change, test normalization plus the rendered browser path; protocol-shape tests alone do not prove UI behavior. diff --git a/.claude/skills/copilotkit/references/mastra-bridge.md b/.claude/skills/copilotkit/references/mastra-bridge.md new file mode 100644 index 00000000..eb5f12da --- /dev/null +++ b/.claude/skills/copilotkit/references/mastra-bridge.md @@ -0,0 +1,11 @@ +# CopilotKit ↔ Mastra bridge + +MDE uses `@ag-ui/mastra` adapters behind the CopilotKit runtime. CopilotKit owns transport/UI registration; `mastra` owns agent definitions, tool implementations, workflows, memory, persistence, and HITL semantics. + +Before changing the bridge inspect: +- `src/mastra/copilotkit/logging-mastra-agent.ts` +- `src/platform/copilot/mastra-tool-action-names.ts` +- the `/api/copilotkit` route +- current Mastra agent registry/allowlist + +Do not duplicate Mastra business/tool logic in React actions. Preserve request/tenant context and telemetry across the bridge. diff --git a/.claude/skills/copilotkit/references/official/copilotkit-cli/SKILL.md b/.claude/skills/copilotkit/references/official/copilotkit-cli/SKILL.md new file mode 100644 index 00000000..fb44453f --- /dev/null +++ b/.claude/skills/copilotkit/references/official/copilotkit-cli/SKILL.md @@ -0,0 +1,141 @@ +--- +name: copilotkit-cli +description: "Use for the CopilotKit CLI — `npx copilotkit@latest`. Covers proving a project's wiring with `verify` before debugging anything by hand, scaffolding with `create`, signing in and selecting a hosted Intelligence project, agent-assisted onboarding of an existing app, generating type-safe agent ids, and importing thread history. Reach for `verify` first whenever a CopilotKit app is not working." +version: 1.0.0 +--- + +# CopilotKit CLI + +```bash +npx copilotkit@latest +``` + +`--help` on any command prints its flags. The commands below are the ones worth knowing +before you start reading someone's project by hand. + +## `verify` — do this before debugging + +```bash +npx copilotkit@latest verify --json +``` + +One command replaces the manual survey. It settles up to eleven things: a hosted project is +selected; the project API key is present, loadable by the app, and authenticates; the runtime +responds, declares an agent, actually consumes the credential, and serves the thread routes; +the frontend serves its own assets; the runtime accepts the browser's origin; and the +installed CopilotKit packages match the version the runtime reports. It also reports the +runtime version, the agent framework in use, whether transcription is wired, the realtime +gateway wiring, and the license state. + +Eleven is the ceiling, not a promise. The last three are omitted when there was nothing to +check them against — no frontend origin was found, or no installed packages were. Count +`checks[]` rather than assuming a fixed set. With `--expect-runtime oss` the hosted-project +and credential checks do not apply at all, so that run is a smaller set. + +Crucially it names **which URL it probed and where that URL came from** — the project's +`runtimeUrl`, an environment variable, the app's own dev configuration, or an assumed +default. A survey done by hand cannot tell you that, and the provenance changes the verdict: +nothing answering at a URL **the project named** is a FAIL, while nothing answering at an +**assumed** default is UNKNOWN, because an app on a port the command never learned is not a +wiring failure. + +Useful flags: + +- `--frontend-url ` — adds the browser-facing checks, including a + real CORS preflight when that origin differs from the runtime's +- `--round-trip` — also runs the agent and reads its answer back. Costs a model call, so it + is opt-in +- `--expect-runtime oss` — for a self-hosted runtime with no Intelligence. It exits zero + only when `/info` declares the named agent, reports no Intelligence entitlement, and + `--round-trip --agent ` passes +- `--agent ` — which declared agent to run, when several are registered +- `--runtime-url ` — probe this endpoint instead of the one read from the project +- `--header ': '` — repeatable. Needed when the project's `identifyUser` reads + a session the CLI does not carry +- `--timeout ` — how long to wait for an answer, default 90 + +Read `checks[]` and fix in the order given: + +- The checks **chain**. A later check that could not run says so and names the earlier one to + fix first, so the first failure is the real one. +- `UNKNOWN` means the check could not run. It never means the check passed, and the command + exits non-zero unless every check passed. + +### What `verify` does not cover + +Reach past it only once it is clean. + +- **Tool execution.** `--round-trip` deliberately asks a question that needs no tools and + sends no context, so a passing round trip says nothing about whether your tools work. +- **Event ordering and streaming.** It reports pass or fail on a run, not the sequence inside + it. A run that starts and never finishes, or stalls mid-stream, is a job for the Inspector. +- **State synchronisation.** Snapshot-versus-delta divergence is agent behaviour, not wiring. + +## Starting a project + +```bash +npx copilotkit@latest init # `create` is an alias for it +``` + +Prompts for a name and framework, scaffolds a starter, signs you in when needed, and connects +the app to a cloud-hosted Intelligence project. The name it asks for names the new directory, +so this is the path for a project that does not exist yet. For an app you already have, use +`onboard start` below. + +To add CopilotKit to an existing app, either follow the [quickstart](/quickstart), or hand +the job to your coding agent: + +```bash +npx copilotkit@latest onboard start +``` + +That runs an agent-guided flow over the repository you are already in, with checkpoints and +proof steps rather than a scaffold. `onboard start --intent ` targets one feature on +an app that already has CopilotKit. + +## Signing in and picking a project + +```bash +npx copilotkit@latest login --json # agent-readable JSON lines, no browser launch +npx copilotkit@latest login # interactive: opens a browser +npx copilotkit@latest whoami # who is signed in, and the active organization +npx copilotkit@latest project select # pick or create a hosted project for this directory +npx copilotkit@latest project list --json +``` + +Use `login --json` when you are driving the CLI. Bare `login` tries to open a browser, which +is not something you can complete. + +There is no `auth` command. It is `login`. + +`project select` records the choice in `.copilotkit/project.json` and provisions a +project-scoped runtime key into `.env`: + +``` +CPK_INTELLIGENCE_API_KEY=cpk_... +``` + +`CPK_INTELLIGENCE_API_KEY` is the canonical name and the only one the CLI writes. Keep it +server-side — it is a runtime key, not a frontend token, so it takes **no** `NEXT_PUBLIC_` or +`VITE_` prefix. Do not set the platform URLs: they default to the managed hosts, so any value +you supply can only replace a correct default with a worse one. + +Without a TTY — which is what a coding agent has — use `project list --json` to see the +choices and `project select --project ` or `--create ` to name the answer up front. + +## Other commands + +| Command | What it does | +| ------------------------------------------ | ------------------------------------------------------------------------------ | +| `skills install` | Installs these skills into a project (`skills onboard` also starts onboarding) | +| `typegen` | Generates type-safe agent ids from a running runtime | +| `import --source adk\|langgraph --dry-run` | Previews importing historical threads into Intelligence | +| `license create` / `license list` | Issues and lists license tokens | +| `channels` | Sets up managed Intelligence Channels for Slack or Microsoft Teams | +| `framework list` | The agent frameworks `create` accepts, and their flags | +| `logs` | The CLI log path, or recent lines | +| `telemetry` | Shows or changes the CLI telemetry preference | +| `docs` | Opens the documentation | +| `version` | Version, build, and commit | + +The CLI collects usage data. `DO_NOT_TRACK=1` or `COPILOTKIT_TELEMETRY_DISABLED=1` opts out. diff --git a/.claude/skills/copilotkit/references/official/copilotkit/SKILL.md b/.claude/skills/copilotkit/references/official/copilotkit/SKILL.md new file mode 100644 index 00000000..4ff65aad --- /dev/null +++ b/.claude/skills/copilotkit/references/official/copilotkit/SKILL.md @@ -0,0 +1,98 @@ +--- +name: copilotkit +description: "Use for any CopilotKit question — adding it to an app, chat UI, frontend or server tools, generative UI, shared state, human-in-the-loop, agent frameworks (LangGraph, CrewAI, Mastra, ADK, PydanticAI, and others), the runtime, Intelligence, threads, voice, or diagnosing something that is not working. Do not answer from memory: this skill exists to point you at the current documentation and source, both of which are searchable." +version: 3.0.0 +--- + +# CopilotKit + +CopilotKit's APIs move. Anything written down in a skill file is a copy that starts drifting +the day it is written, so this skill carries almost no API detail on purpose. It tells you +where the current answer lives and how to get it. + +**Look it up before you write code.** Not because the docs are more convenient, but because +they are regenerated from the source and a recollection is not. + +## The search tools + +An MCP server, `copilotkit-docs`, is bundled with this plugin. It exposes four search tools +and two exploration tools over four separate corpora. Picking the wrong one is the most +common way to come up empty: + +| Tool | Corpus | Use it for | +| ------------------------------- | ------------------------- | ----------------------------------------------------------------------------------------------------- | +| `search-docs` | docs.copilotkit.ai | Usage, configuration, guides, quickstarts, the generated API reference | +| `search-code` | CopilotKit library source | How something is implemented, and exact signatures. Library packages only — not examples or showcases | +| `search-ag-ui-docs` | AG-UI protocol docs | The protocol itself: event types, transports, the SDKs | +| `search-ag-ui-code` | AG-UI protocol SDK source | Protocol implementation detail | +| `explore-docs` / `explore-code` | either tree | Browsing structure when you do not yet know what to search for | + +CopilotKit questions go to the first two. AG-UI protocol questions go to the second two — +they are a different repository, and `search-docs` will not find them. + +**These are semantic searches.** Several short, differently-phrased queries beat one long +one. If a query returns something off-target, rephrase rather than widen — and if you get a +plausible-looking page that does not actually contain the term you need, say so instead of +reasoning from the title. + +## Setup + +The server is registered once per tool, and installing the skills does not register it. +Check first — if the search tools above are already available, skip this. + +**Claude Code** — the CopilotKit plugin declares the server in its `.mcp.json`, so a +plugin install needs nothing. A skills-only install (`npx skills add`) does not carry +that file, so register it: + +```bash +claude mcp add --transport sse copilotkit-mcp https://mcp.copilotkit.ai/sse +``` + +**Codex**: + +```bash +codex mcp add copilotkit --url https://mcp.copilotkit.ai/mcp +``` + +**Anything else** — `https://mcp.copilotkit.ai/mcp` for streamable HTTP, +`https://mcp.copilotkit.ai/sse` for SSE. Keep the path: the bare host returns 404. +[Per-tool instructions](/build-with-agents) cover Cursor, Windsurf, Cline, GitHub +Copilot and VS Code. + +Without the server, [the documentation](https://docs.copilotkit.ai) works as a plain +site, and appending `.md` to any docs URL returns that page as Markdown. + +## Where the answers are + +Worth knowing so a search has somewhere to land: + +- **Getting started** — [quickstart](/quickstart), and [the CLI](/cli) for the + CLI-driven path +- **Frontend** — [frontend tools](/frontend-tools), [human-in-the-loop](/human-in-the-loop), + [prebuilt components](/prebuilt-components), [styling](/custom-look-and-feel/css), + [attachments](/multimodal-attachments), [voice](/voice) +- **Runtime** — [the runtime](/backend/copilot-runtime), + [HTTP endpoints](/backend/runtime-endpoints), [runners](/backend/agent-runner), + [factory mode](/backend/custom-agent), [server adapters](/runtime-server-adapter), + [auth](/auth) +- **Agent frameworks** — one quickstart per framework under `/integrations/` +- **Intelligence** — [overview](/intelligence/overview) and the pages under it +- **Not working** — [common issues](/troubleshooting/common-issues), + [error reference](/troubleshooting/error-reference), and the generated + [`CopilotKitCoreErrorCode`](/reference/core/enums/CopilotKitCoreErrorCode) for a code the + app actually reported +- **Protocol** — [AG-UI](/backend/ag-ui), and the AG-UI docs for the protocol itself + +## Before you debug anything + +Run the CLI's wiring check first — `npx copilotkit@latest verify --json`. It settles up to +eleven things in one command and is almost always faster than reading the project. See the +`copilotkit-cli` skill. + +## Two versions exist + +v2 is current. Import from the `/v2` subpath — `@copilotkit/react-core/v2`, +`@copilotkit/runtime/v2`. The package root is the deprecated v1 surface and still resolves, +so mixing the two raises nothing at import time and surfaces later as a runtime mismatch. +Check which subpath a project imports before trusting anything else about it. v1 is +deprecated but supported; [the migration guide](/migrate/v2) covers moving off it. diff --git a/.claude/skills/copilotkit/references/runtime-and-react.md b/.claude/skills/copilotkit/references/runtime-and-react.md new file mode 100644 index 00000000..c59ee6b6 --- /dev/null +++ b/.claude/skills/copilotkit/references/runtime-and-react.md @@ -0,0 +1,12 @@ +# Runtime and React + +Current MDE baseline: `@copilotkit/react-core` 1.55.2 and `@copilotkit/runtime` 1.55.2. The app imports v2 React APIs and styles and uses `/api/copilotkit` as the same-origin runtime. + +Inspect before editing: +- `src/components/copilot/copilot-kit-provider*` +- `src/lib/copilotkit-client-props.ts` +- `src/app/api/copilotkit/[[...path]]/route*` +- `src/lib/hooks/use-concierge-chat.ts` +- `src/lib/hooks/use-host-ops-chat.ts` + +Protect auth/rate limiting and stable provider props. Verify current installed source when an API signature is uncertain; do not copy latest-main examples blindly into 1.55.2. diff --git a/.claude/skills/copilotkit/upstream.yaml b/.claude/skills/copilotkit/upstream.yaml new file mode 100644 index 00000000..7ead608a --- /dev/null +++ b/.claude/skills/copilotkit/upstream.yaml @@ -0,0 +1,26 @@ +vendor: CopilotKit +repository: https://github.com/CopilotKit/CopilotKit +source_path: skills +reviewed_commit: 8a7446186cd3e0d368ec885e61c5913f0918ef5d +reviewed_at: 2026-09-15 +policy: + active_skill: copilotkit + upstream_files_are_read_only: true + auto_update: false + update_flow: detect-diff-review-evals-pin +sources: + core: + path: skills/copilotkit/SKILL.md + local: references/official/copilotkit/SKILL.md + sha256: 8f910089af420640f5863506963f1aaf54314b5316b36d24789680bc680e14b1 + cli: + path: skills/copilotkit-cli/SKILL.md + local: references/official/copilotkit-cli/SKILL.md + sha256: 3a3708ff10e185c589c6d65b1be6c339268b3a2f5d6ed760bd66fd3138713fc9 +optional_upstream_skills: + - channels-setup + - copilotkit-channels + - inspector-docs + - inspector-workbench + - intelligence-docs + - setup-slack-channel diff --git a/.claude/skills/events/SKILL.md b/.claude/skills/events/SKILL.md new file mode 100644 index 00000000..4be9598d --- /dev/null +++ b/.claude/skills/events/SKILL.md @@ -0,0 +1,29 @@ +--- +name: events +description: >- + Use when MDE work changes event creation, host workflows, event publishing, ticket configuration, event discovery, attendee flows, or event-domain rules. +--- + +# Events + +Own event-domain behavior. Stack implementation details remain with `copilotkit`, `mastra`, `supabase`, `stripe`, `nextjs`, and other stack skills. + +## Core domain boundaries + +- Host creation/editing and publishing must preserve explicit approval where the flow requires HITL. +- Ticket tiers, event capacity, dates, venue, publish state, and buyer-facing availability must stay internally consistent. +- Payment success is owned by verified payment state, never by UI optimism. +- Domain writes must remain organization/user scoped. +- Do not duplicate event state across UI, agent state, and database without an explicit source-of-truth contract. + +## Workflow + +1. Identify the persona and surface: host creation, discovery, ticket purchase, saved event, or admin operations. +2. Identify the canonical event state and transition being changed. +3. Load only the affected stack skills. +4. Define success/failure/empty/retry states. +5. Verify the user journey and persistence path end to end. + +## Handoff + +Use `stripe` for payment mechanics, `supabase` for persistence/RLS, `mastra` for agent logic, `copilotkit` for AI UI/runtime, `maps` for venue/place behavior, and `nextjs` for framework-specific behavior. diff --git a/.claude/skills/events/evals/evals.json b/.claude/skills/events/evals/evals.json new file mode 100644 index 00000000..8d4abaca --- /dev/null +++ b/.claude/skills/events/evals/evals.json @@ -0,0 +1,32 @@ +{ + "version": 1, + "evals": [ + { + "prompt": "Handle a focused events change in MDE and identify the correct implementation boundary.", + "expected_output": "Uses events as the primary specialist and keeps unrelated domains out.", + "expectations": [ + "inspects current repo state", + "uses current official docs for version-sensitive behavior", + "makes the smallest safe change" + ] + }, + { + "prompt": "A bug appears near events, but the root cause is not known. What owns diagnosis?", + "expected_output": "Routes diagnosis to systematic-debugging while retaining the specialist for domain evidence.", + "expectations": [ + "does not guess root cause", + "uses systematic-debugging for diagnosis", + "keeps specialist scope bounded" + ] + }, + { + "prompt": "A production-critical events change is ready to ship. What proof is required?", + "expected_output": "Requires targeted tests and independent task verification before Done.", + "expectations": [ + "requires current evidence", + "does not self-certify", + "uses task-verifier for completion proof" + ] + } + ] +} diff --git a/.claude/skills/gemini b/.claude/skills/gemini deleted file mode 120000 index 977702c6..00000000 --- a/.claude/skills/gemini +++ /dev/null @@ -1 +0,0 @@ -../../.agents/skills/gemini \ No newline at end of file diff --git a/.claude/skills/gemini/SKILL.md b/.claude/skills/gemini/SKILL.md new file mode 100644 index 00000000..c2175c28 --- /dev/null +++ b/.claude/skills/gemini/SKILL.md @@ -0,0 +1,56 @@ +--- +name: gemini +description: >- + Use when MDE work changes or verifies Gemini models, Google AI SDK/provider usage, function calling, structured output, multimodal input/output, grounding, embeddings, research agents, streaming, or Gemini-specific failures. Do not use for generic Mastra orchestration, generic AI prompting, or Google Maps behavior unless Gemini is part of the implementation. +metadata: + mde-version: "1.0.0" + upstream-commit: "80dd31dda25bbe1410207df0adb3e0d591c2c634" + mde-provider-baseline: "@ai-sdk/google 2.0.74" +--- + +# Gemini — official upstream + MDE overlay + +## Source order + +1. Inspect the current MDE provider package, model strings, and call sites. +2. Read `references/official/gemini-api-dev/SKILL.md` for current Gemini API/model guidance. +3. Fetch the current official Gemini documentation page required by that skill before changing version-sensitive code. +4. Use the Live API or Omni Flash official skills only when the task actually uses those runtimes. +5. Apply the MDE-specific integration rules below. + +Do not trust remembered Gemini model names, SDK signatures, quotas, or deprecation status. Do not migrate MDE from `@ai-sdk/google` to `@google/genai` unless the task explicitly requires that architectural change. + +## Ownership + +Own Gemini model/provider selection, Gemini-specific tool/function contracts, structured output, multimodal/grounding/embedding behavior, and Gemini-specific errors. `mastra` owns agent/workflow orchestration. `maps` owns map/place behavior. Domain skills own business meaning. +## Current MDE invariants + +- Current MDE integration uses `@ai-sdk/google` 2.0.74; preserve that provider boundary unless migration is explicitly in scope. +- Verify the exact model identifier and provider support before changing model strings. +- Keep structured-output schemas, tool contracts, and fallback behavior explicit and tested. +- Grounding/search or Maps behavior must preserve citations/attribution requirements where applicable. +- Treat model changes as behavior changes: rerun representative prompts and structured-output/tool tests, not only typecheck. +- Do not enable Live API or Omni-specific paths merely because the official references exist. + +## Workflow + +1. Classify: model/provider, structured output, function/tool calling, multimodal, grounding/search, embeddings, research agent, streaming, Live API, or Omni. +2. Read the matching official Gemini skill/reference first. +3. Verify the installed MDE provider and current call-site contract. +4. Check current official docs/model availability before editing. +5. Make the smallest provider-compatible change. +6. Verify representative success, schema/tool correctness, and failure/fallback behavior. + +## Verification + +For model/provider changes, prove the requested model exists and the installed integration supports it. For structured output/tools, validate schema and actual runtime output. For grounding/search, verify source metadata/attribution. For S3/S4 cross-system changes, finish with independent `code-review` and `task-verifier`. + +## References + +- `upstream.yaml` +- `references/official/gemini-api-dev/SKILL.md` +- `references/official/gemini-api-dev/references/migration.md` +- `references/official/gemini-live-api-dev/SKILL.md` — only for Live API work +- `references/official/gemini-omni-flash-api/SKILL.md` — only for Omni Flash work +- https://ai.google.dev/gemini-api/docs +- https://github.com/google-gemini/gemini-skills diff --git a/.claude/skills/gemini/evals/evals.json b/.claude/skills/gemini/evals/evals.json new file mode 100644 index 00000000..3387db71 --- /dev/null +++ b/.claude/skills/gemini/evals/evals.json @@ -0,0 +1,36 @@ +{ + "skill_name": "gemini", + "evals": [ + { + "id": 1, + "prompt": "Switch the MDE Gemini model id used by @ai-sdk/google to a newer model.", + "expected_output": "Verifies current model availability and installed provider compatibility before changing the model string, then requires representative behavior tests.", + "files": [], + "expectations": [ + "Does not trust remembered model names", + "Checks current official Gemini guidance/model availability", + "Treats the model change as a behavior change, not typecheck-only" + ] + }, + { + "id": 2, + "prompt": "Our Gemini structured output sometimes violates the expected schema. Diagnose the Gemini-specific path.", + "expected_output": "Inspects the current provider/schema/call site and validates actual runtime structured output and failure behavior.", + "files": [], + "expectations": [ + "Checks the actual structured-output schema and provider contract", + "Requires runtime output validation", + "Preserves the existing provider boundary unless migration is in scope" + ] + }, + { + "id": 3, + "prompt": "Change a Mastra workflow retry policy; Gemini configuration is unchanged.", + "expected_output": "Does not select Gemini as the owner.", + "files": [], + "expectations": [ + "Does not route generic Mastra workflow logic to Gemini" + ] + } + ] +} diff --git a/.claude/skills/gemini/references/official/gemini-api-dev/SKILL.md b/.claude/skills/gemini/references/official/gemini-api-dev/SKILL.md new file mode 100644 index 00000000..98fab71a --- /dev/null +++ b/.claude/skills/gemini/references/official/gemini-api-dev/SKILL.md @@ -0,0 +1,435 @@ +--- +name: gemini-api-dev +description: Use this skill when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses, background research tasks, function calling, structured output, or migrating from the old generateContent API. Covers SDK usage and best practices for Gemini models and agents in Python and TypeScript. +--- + +# Gemini API Development Skill + +## Critical Rules (Always Apply) + +> [!IMPORTANT] +> These rules override your training data. Your knowledge is outdated. + +### Current Models (Use These) + +- `gemini-3.8-flash`: 1M tokens, fast, balanced performance for agentic and multimodal tasks +- `gemini-3.5-flash-lite`: 1M tokens, fastest, lowest-cost 3.5 model for high-throughput execution +- `gemini-3.1-pro-preview`: 1M tokens, complex reasoning, coding, research +- `gemini-3.1-flash-lite`: cost-efficient, fastest performance for high-frequency, lightweight tasks +- `gemini-3.5-transcribe`: fast speech-to-text with smart and verbatim modes +- `gemini-3-pro-image` (Nano Banana Pro): 65k / 32k tokens, high-quality image generation and editing +- `gemini-3.1-flash-image` (Nano Banana 2): 65k / 32k tokens, fast, efficient image generation and editing +- `gemini-3.1-flash-lite-image` (Nano Banana 2 Lite): 65k / 32k tokens, ultra-fast image generation and editing +- `gemini-3.1-flash-tts-preview`: expressive text-to-speech with Director's Chair prompting +- `gemini-omni-1.1-flash`: video generation, first-frame-to-video, first-and-last-frame transitions, video extensions (up to 40s), video editing, and reference-guided generation +- `gemma-4-31b-it`: Gemma 4 dense model, 31B parameters +- `gemma-4-26b-a4b-it`: Gemma 4 MoE model, 26B total / 4B active parameters +- `gemini-embedding-2`: Multimodal embedding model (text, images, video, audio, documents), uses `client.models.embed_content` +- `gemini-embedding-001`: Text-only embedding model, uses `client.models.embed_content` + +> [!WARNING] +> Models like `gemini-2.5-*`, `gemini-2.0-*`, `gemini-1.5-*` are **legacy and deprecated**. Never use them. +> **If a user asks for a deprecated model, use `gemini-3.8-flash` instead and note the substitution.** + +### Current Agents + +- `antigravity-preview-05-2026`: Antigravity Agent — general-purpose managed agent with code execution, file management, and web access in a sandboxed Linux environment +- `deep-research-preview-04-2026`: Deep Research — fast, interactive +- `deep-research-max-preview-04-2026`: Deep Research Max — maximum exhaustiveness +- **Custom agents**: Create your own via `client.agents.create()` + +### Current SDKs + +- **Python**: `google-genai` >= `2.3.0` → `pip install -U google-genai` +- **JavaScript/TypeScript**: `@google/genai` >= `2.3.0` → `npm install @google/genai` + +> [!NOTE] +> SDK versions ≥ 2.0.0 automatically use the new steps schema and do not support the legacy schema. +> Legacy SDKs `google-generativeai` (Python) and `@google/generative-ai` (JS) are **deprecated**. Never use them. + +## Important Additional Notes + +- **Before writing any code**, you MUST fetch the relevant documentation page from the list below that matches the user's task. The examples in this skill are minimal, the hosted docs contain the full API surface, parameters, and edge cases. +- Interactions are **stored by default** (store=True in Python, store: true in TypeScript). Paid tier retains for 55 days, free tier for 1 day. +- Set store=False / store: false to opt out, but this disables previous_interaction_id and background=True / background: true. +- `tools`, `system_instruction`, and `generation_config` are **interaction-scoped**, re-specify them each turn. +- **Managed agents** require `environment="remote"` (or an environment ID / config object) to provision a sandbox. +- **Migrating from `generateContent`**: Read `references/migration.md` for the scoping, checklist, and before/after code examples. Always confirm scope with the user before editing. +- **Model upgrades**: Drop-in, swap the model string. Deprecated models (`gemini-2.0-*`, `gemini-1.5-*`) must be replaced, see `references/migration.md`. +- **Migrating to Gemini 3.8 Flash or Gemini 3.5 Flash-Lite**: Read `references/migration.md` for the scoping and checklist. + +## Quick Start + +### Python +```python +from google import genai + +client = genai.Client() + +interaction = client.interactions.create( + model="gemini-3.8-flash", + input="Tell me a short joke about programming." +) +print(interaction.output_text) +``` + +### JavaScript/TypeScript +```typescript +import { GoogleGenAI } from "@google/genai"; + +const client = new GoogleGenAI({}); + +const interaction = await client.interactions.create({ + model: "gemini-3.8-flash", + input: "Tell me a short joke about programming.", +}); +console.log(interaction.output_text); +``` + +## Response Helpers + +The SDK provides convenience properties on the `Interaction` response object to simplify common access patterns: + +| Property | Type | Description | +|---|---|---| +| `output_text` | `string \| null` | The last consecutive run of text from the trailing `model_output` steps. Returns the combined text when the model's final output contains multiple text parts. | +| `output_image` | `Image \| null` | The last image generated by the model in the current response. Returns an object with `data` (base64) and `mime_type`. | +| `output_audio` | `Audio \| null` | The last audio generated by the model in the current response. Returns an object with `data` (base64) and `mime_type`. | + +## Stateful Conversation + +### Python +```python +interaction1 = client.interactions.create( + model="gemini-3.8-flash", + input="Hi, my name is Phil." +) +# Second turn — server remembers context +interaction2 = client.interactions.create( + model="gemini-3.8-flash", + input="What is my name?", + previous_interaction_id=interaction1.id +) +print(interaction2.output_text) +``` + +### JavaScript/TypeScript +```typescript +const interaction1 = await client.interactions.create({ + model: "gemini-3.8-flash", + input: "Hi, my name is Phil.", +}); +const interaction2 = await client.interactions.create({ + model: "gemini-3.8-flash", + input: "What is my name?", + previous_interaction_id: interaction1.id, +}); +console.log(interaction2.output_text); +``` + +## Deep Research Agent + +Use `deep-research-preview-04-2026` for fast research or `deep-research-max-preview-04-2026` for maximum exhaustiveness. Agents require `background=True`. + +### Python +```python +import time + +interaction = client.interactions.create( + agent="deep-research-preview-04-2026", + input="Research the history of Google TPUs.", + background=True +) +while True: + interaction = client.interactions.get(interaction.id) + if interaction.status == "completed": + print(interaction.output_text) + break + elif interaction.status == "failed": + print(f"Failed: {interaction.error}") + break + time.sleep(10) +``` + +### JavaScript/TypeScript +```typescript +import { GoogleGenAI } from "@google/genai"; + +const client = new GoogleGenAI({}); + +// Start background research +const initialInteraction = await client.interactions.create({ + agent: "deep-research-preview-04-2026", + input: "Research the history of Google TPUs.", + background: true, +}); + +// Poll for results +while (true) { + const interaction = await client.interactions.get(initialInteraction.id); + if (interaction.status === "completed") { + console.log(interaction.output_text); + break; + } else if (["failed", "cancelled"].includes(interaction.status)) { + console.log(`Failed: ${interaction.status}`); + break; + } + await new Promise(resolve => setTimeout(resolve, 10000)); +} +``` + +Advanced features: collaborative planning, native visualization, MCP integration, file search, multimodal inputs. See [Deep Research docs](https://ai.google.dev/gemini-api/docs/deep-research.md.txt). + +## Managed Agents + +Managed agents run inside a sandboxed Linux environment hosted by Google. Fetch the [Managed Agents Quickstart](https://ai.google.dev/gemini-api/docs/managed-agents-quickstart.md.txt) before writing agent code. + +### Antigravity Agent + +The Antigravity agent (`antigravity-preview-05-2026`) is the general-purpose managed agent. It can execute code (Bash, Python, Node.js), manage files, browse the web, and use Google Search. See [Antigravity Agent docs](https://ai.google.dev/gemini-api/docs/antigravity-agent.md.txt) for capabilities, tools, multimodal input, and pricing. + +#### Python +```python +from google import genai + +client = genai.Client() + +interaction = client.interactions.create( + agent="antigravity-preview-05-2026", + input="Write a Python script that generates the first 20 Fibonacci numbers and saves them to fibonacci.txt. Then read the file and print its contents.", + environment="remote", +) + +print(f"Environment ID: {interaction.environment_id}") +print(interaction.output_text) +``` + +#### JavaScript/TypeScript +```typescript +import { GoogleGenAI } from "@google/genai"; + +const client = new GoogleGenAI({}); + +const interaction = await client.interactions.create({ + agent: "antigravity-preview-05-2026", + input: "Write a Python script that generates the first 20 Fibonacci numbers and saves them to fibonacci.txt. Then read the file and print its contents.", + environment: "remote", +}); + +console.log(`Environment ID: ${interaction.environment_id}`); +console.log(interaction.output_text); +``` + +### Custom Agents + +See [Building Custom Agents docs](https://ai.google.dev/gemini-api/docs/custom-agents.md.txt). + +#### Python +```python +agent = client.agents.create( + id="code-reviewer", + base_agent="antigravity-preview-05-2026", + system_instruction="You are a senior code reviewer. Check every file for bugs, style issues, and security vulnerabilities.", + base_environment={ + "type": "remote", + "sources": [ + { + "type": "repository", + "source": "https://github.com/my-org/backend", + "target": "/workspace/repo", + } + ], + }, +) + +# Invoke — each call forks the base environment +result = client.interactions.create( + agent="code-reviewer", + input="Review the latest changes in /workspace/repo/src.", + environment="remote", +) +print(result.output_text) +``` + +#### JavaScript/TypeScript +```typescript +const agent = await client.agents.create({ + id: "code-reviewer", + base_agent: "antigravity-preview-05-2026", + system_instruction: "You are a senior code reviewer. Check every file for bugs, style issues, and security vulnerabilities.", + base_environment: { + type: "remote", + sources: [ + { + type: "repository", + source: "https://github.com/my-org/backend", + target: "/workspace/repo", + } + ], + }, +}); + +const result = await client.interactions.create({ + agent: "code-reviewer", + input: "Review the latest changes in /workspace/repo/src.", + environment: "remote", +}); +console.log(result.output_text); +``` + +Manage agents with `client.agents.list()`, `client.agents.get(id=...)`, and `client.agents.delete(id=...)`. + +## Streaming + +Set `stream=True` to receive incremental server-sent events. Each stream follows: `interaction.created` → (`step.start` → `step.delta`(s) → `step.stop`)+ → `interaction.completed`. + +### Python +```python +for event in client.interactions.create( + model="gemini-3.8-flash", + input="Explain quantum entanglement in simple terms.", + stream=True, +): + if event.event_type == "step.delta": + if event.delta.type == "text": + print(event.delta.text, end="", flush=True) + elif event.event_type == "interaction.completed": + print(f"\n\nTotal Tokens: {event.interaction.usage.total_tokens}") +``` + +### JavaScript/TypeScript +```typescript +const stream = await client.interactions.create({ + model: "gemini-3.8-flash", + input: "Explain quantum entanglement in simple terms.", + stream: true, +}); +for await (const event of stream) { + if (event.event_type === "step.delta") { + if (event.delta.type === "text") { + process.stdout.write(event.delta.text); + } + } else if (event.event_type === "interaction.completed") { + console.log(`\n\nTotal Tokens: ${event.interaction?.usage?.total_tokens}`); + } +} +``` + +For streaming with tools, thinking, agents, and image generation see the full [Streaming guide](https://ai.google.dev/gemini-api/docs/streaming.md.txt). + + + +## Documentation Pages + +**You MUST fetch the matching page below before writing code.** These hosted docs are the source of truth for parameters, types, and edge cases — do not rely solely on the examples above. + +**Core Documentation:** +- [Interactions API Overview](https://ai.google.dev/gemini-api/docs/interactions.md.txt) +- [Quickstart](https://ai.google.dev/gemini-api/docs/quickstart.md.txt) +- [Text Generation](https://ai.google.dev/gemini-api/docs/text-generation.md.txt) +- [Streaming](https://ai.google.dev/gemini-api/docs/streaming.md.txt) +- [Tokens](https://ai.google.dev/gemini-api/docs/tokens.md.txt) +- [API Keys](https://ai.google.dev/gemini-api/docs/api-key.md.txt) + +**Tools & Function Calling:** +- [Function Calling](https://ai.google.dev/gemini-api/docs/function-calling.md.txt) +- [Google Search](https://ai.google.dev/gemini-api/docs/google-search.md.txt) +- [Code Execution](https://ai.google.dev/gemini-api/docs/code-execution.md.txt) +- [URL Context](https://ai.google.dev/gemini-api/docs/url-context.md.txt) +- [File Search](https://ai.google.dev/gemini-api/docs/file-search.md.txt) +- [Tool Combination](https://ai.google.dev/gemini-api/docs/tool-combination.md.txt) +- [Computer Use](https://ai.google.dev/gemini-api/docs/computer-use.md.txt) +- [Maps Grounding](https://ai.google.dev/gemini-api/docs/maps-grounding.md.txt) + +**Generation & Output:** +- [Structured Output](https://ai.google.dev/gemini-api/docs/structured-output.md.txt) +- [Thinking](https://ai.google.dev/gemini-api/docs/thinking.md.txt) +- [Thought Signatures](https://ai.google.dev/gemini-api/docs/thought-signatures.md.txt) +- [Image Generation](https://ai.google.dev/gemini-api/docs/image-generation.md.txt) +- [Image Understanding](https://ai.google.dev/gemini-api/docs/image-understanding.md.txt) +- [Video Generation & Editing (Omni Flash)](https://ai.google.dev/gemini-api/docs/omni.md.txt) +- [Speech Generation](https://ai.google.dev/gemini-api/docs/speech-generation.md.txt) +- [Music Generation](https://ai.google.dev/gemini-api/docs/music-generation.md.txt) +- [Embeddings](https://ai.google.dev/gemini-api/docs/embeddings.md.txt) + +**Multimodal Understanding:** +- [Audio](https://ai.google.dev/gemini-api/docs/audio.md.txt) +- [Audio Transcription](https://ai.google.dev/gemini-api/docs/transcribe.md.txt) +- [Video Understanding](https://ai.google.dev/gemini-api/docs/video-understanding.md.txt) +- [Document Processing](https://ai.google.dev/gemini-api/docs/document-processing.md.txt) + +**Files & Context:** +- [Files](https://ai.google.dev/gemini-api/docs/files.md.txt) +- [File Input Methods](https://ai.google.dev/gemini-api/docs/file-input-methods.md.txt) +- [Caching](https://ai.google.dev/gemini-api/docs/caching.md.txt) +- [Media Resolution](https://ai.google.dev/gemini-api/docs/media-resolution.md.txt) + +**Agents:** +- [Agents Overview](https://ai.google.dev/gemini-api/docs/agents.md.txt) +- [Managed Agents Quickstart](https://ai.google.dev/gemini-api/docs/managed-agents-quickstart.md.txt) +- [Antigravity Agent](https://ai.google.dev/gemini-api/docs/antigravity-agent.md.txt) +- [Agent Environments](https://ai.google.dev/gemini-api/docs/agent-environment.md.txt) +- [Agent Hooks](https://ai.google.dev/gemini-api/docs/agent-hooks.md.txt) +- [Building Custom Agents](https://ai.google.dev/gemini-api/docs/custom-agents.md.txt) +- [Deep Research](https://ai.google.dev/gemini-api/docs/deep-research.md.txt) + +**Advanced Features:** +- [Latest Models (3.8 Flash & 3.5 Flash-Lite)](https://ai.google.dev/gemini-api/docs/latest-model.md.txt) +- [Flex Inference](https://ai.google.dev/gemini-api/docs/flex-inference.md.txt) +- [Priority Inference](https://ai.google.dev/gemini-api/docs/priority-inference.md.txt) + +**API Reference:** +- [API Reference](https://ai.google.dev/static/api/interactions.md.txt) +- [OpenAPI Spec](https://ai.google.dev/static/api/interactions.openapi.json) +- [May 2026 Breaking Changes Migration Guide](https://ai.google.dev/gemini-api/docs/interactions-breaking-changes-may-2026.md.txt) + +## Data Model + +An `Interaction` response contains `steps`, an array of typed step objects representing a structured timeline of the interaction turn. + +### Step Types + +**User steps:** +- `user_input`: User input (text, audio, multimodal). Contains `content` array. + +**Model/server steps:** +- `model_output`: Final model generation. Contains `content` array with `text`, `image`, `audio`, etc. +- `thought`: Model reasoning/Chain of Thought. Has `signature` field (required) and optional `summary`. +- `function_call`: Tool call request (`id`, `name`, `arguments`). +- `function_result`: Tool result you send back (`call_id`, `name`, `result`). +- `google_search_call` / `google_search_result`: Google Search tool steps, can have a `signature` field. +- `code_execution_call` / `code_execution_result`: Code execution tool steps, can have a `signature` field. +- `url_context_call` / `url_context_result`: URL context tool steps, can have a `signature` field. +- `mcp_server_tool_call` / `mcp_server_tool_result`: Remote MCP tool steps. +- `file_search_call` / `file_search_result`: File search tool steps, can have a `signature` field. + +### Content types (inside `content` array on `model_output` and `user_input` steps) +- `text`: Text content (`text` field) +- `image` / `audio` / `document` / `video`: Content with `data`, `mime_type`, or `uri` + +### Streaming Event Types + +| Event | Description | +|---|---| +| `interaction.created` | Interaction created; includes metadata. | +| `interaction.status_update` | Interaction-level status change. | +| `step.start` | A new step begins. Contains step `type` and initial metadata. | +| `step.delta` | Incremental data for the current step. Contains a typed `delta` object. | +| `step.stop` | The step is complete. Contains `index`. | +| `interaction.completed` | Interaction finished. Contains final `usage`. | + +### Delta Types + +| Delta Type | Parent Step | Description | +|---|---|---| +| `text` | `model_output` | Incremental text token. | +| `audio` | `model_output` | audio chunk (base64). | +| `image` | `model_output` | image chunk (base64). | +| `thought_summary` | `thought` | thinking summary text. | +| `thought_signature` | `thought` | Opaque signature for thought verification. | + +**Status values:** `completed`, `in_progress`, `requires_action`, `failed`, `cancelled` + +## Gemini Live API + +For real-time, bidirectional audio/video/text streaming with the Gemini Live API, install the **`google-gemini/gemini-live-api-dev`** skill. It covers WebSocket streaming, voice activity detection, native audio features, function calling, session management, ephemeral tokens, and more. diff --git a/.claude/skills/gemini/references/official/gemini-api-dev/references/migration.md b/.claude/skills/gemini/references/official/gemini-api-dev/references/migration.md new file mode 100644 index 00000000..83f9a995 --- /dev/null +++ b/.claude/skills/gemini/references/official/gemini-api-dev/references/migration.md @@ -0,0 +1,122 @@ +# Migration Reference + +How to migrate existing Gemini API code to the Interactions API and/or upgrade between model generations. Covers the agent workflow for performing migrations safely. + +For detailed before/after code examples across all feature areas (text generation, multi-turn, streaming, function calling, structured output, grounding, multimodal), fetch the full migration guide: https://ai.google.dev/gemini-api/docs/migrate-to-interactions.md.txt + +## Confirm the Migration Scope + +**Before any edits, confirm the scope.** If the user's request does not explicitly name a single file, a specific directory, or an explicit file list, ask first and do not start editing. + +Even imperative requests like "migrate my code", "upgrade to gemini 3", "migrate my app to gemini", or "switch to the Interactions API" leave the scope ambiguous. Ask: + +> Before I start editing, can you confirm the scope? +> 1. Entire project +> 2. Specific subdirectory (e.g. `src/`, `api/`) +> 3. Specific file or list of files + +**Sizing the scope (large repos).** Before asking, get a per-directory count: + +```sh +rg -l "generate_content|generateContent|gemini-1\.5|gemini-2\.0|gemini-2\.5|gemini-3\.1-flash-lite|gemini-3\.5|gemini-3\.6|gemini-3\.7|gemini-3-flash|thinking_budget|temperature" --type-not md | cut -d/ -f1 | sort | uniq -c | sort -rn +``` + +Present the breakdown in your question (e.g. *"Found 42 references across 3 directories: src/ (28), tests/ (10), scripts/ (4). Which to migrate?"*). + +**Proceed without asking** only when the scope is already unambiguous, the user named an exact file ("migrate `app.py`"), pointed at a directory ("migrate everything under `src/`"), or already confirmed scope in an earlier turn. + +## API Migration: `generateContent` → `Interactions` + +The core changes when migrating from `generateContent` to the Interactions API: + +| What | `generateContent` | Interactions API | +|------|----------------|-----------------| +| **SDK method** | `client.models.generate_content()` | `client.interactions.create()` | +| **Response text** | `response.text` | `interaction.steps[-1].content[0].text` | +| **Multi-turn** | Manual history array or `client.chats.create()` | `previous_interaction_id=interaction.id` | +| **Streaming** | `generate_content_stream()` / `:streamGenerateContent` | `stream=True` + `step.delta` events | +| **Structured output** | `config.response_format` inside `GenerateContentConfig` | Top-level `response_format` array | +| **Function calling** | `candidates[0].content.parts[0].function_call` | `function_call` step in `interaction.steps` | +| **Search grounding** | `groundingMetadata` on candidates | `google_search_call`/`google_search_result` steps + inline `annotations` | +| **Config/types** | `types.GenerateContentConfig(...)`, `types.Tool(...)`, `types.Content(...)`, `types.Part.*` | Not used. Interactions API uses plain Python dicts and direct params. Check the feature docs for exact format. | +| **REST endpoint** | `POST /v1beta/models/{model}:generateContent` | `POST /v1beta/interactions` | +| **SDK package** | `google-genai` ≥ 1.x or legacy `google-generativeai` | `google-genai` ≥ 2.0.0 | + +For full before/after code examples, fetch the [Migration Guide](https://ai.google.dev/gemini-api/docs/migrate-to-interactions.md.txt) or read the Interactions API documentation pages for each feature. + +## Model Migration + +### Deprecated Models + +| Model | Status | Drop-in Replacement | +|-------|--------|-------------------| +| `gemini-2.0-flash` | Deprecated | `gemini-3.8-flash` | +| `gemini-2.0-flash-lite` | Deprecated | `gemini-3.5-flash-lite` | +| `gemini-1.5-pro` | Deprecated | `gemini-3.8-flash` | +| `gemini-1.5-flash` | Deprecated | `gemini-3.8-flash` | + +### Active Legacy Models (migration recommended) + +| Current Model | Recommended Target | Why | +|--------------|-------------------|-----| +| `gemini-3.7-flash`, `gemini-3.6-flash`, `gemini-3.5-flash`, or `gemini-3-flash-preview` | `gemini-3.8-flash` | Latest Flash: stronger agentic/multimodal performance, reduced token usage/loop spiraling | +| `gemini-2.5-flash` | `gemini-3.8-flash` or `gemini-3.5-flash-lite` | Latest Flash with Interactions API support, or latest Flash-Lite for cheaper/simpler tasks. | +| `gemini-2.5-flash-lite` or `gemini-3.1-flash-lite` | `gemini-3.5-flash-lite` | Latest Flash-lite with Interactions API support | +| `gemini-2.5-pro` | `gemini-3.1-pro-preview` | Latest Pro with 1M context, complex reasoning | + +> **Note:** Within the Interactions API, model upgrades are generally drop-in — change the model string and verify. The breaking changes are at the **API level** (generateContent → Interactions) and parameter deprecations (`temperature`, `top_p`, `top_k`). + +## Migration Checklist + +Every item is tagged: **`[BLOCKS]`** items cause errors or broken behavior if missed. **`[TUNE]`** items are quality/performance adjustments. + +### API Migration (generateContent → Interactions) + +- [ ] Updated SDK: `google-genai` ≥ 2.0.0 (Python) / `@google/genai` ≥ 2.0.0 (JS) +- [ ] Replaced `client.models.generate_content()` → `client.interactions.create()` +- [ ] Replaced `response.text` → `interaction.steps[-1].content[0].text` +- [ ] Replaced `response.candidates[0].content.parts` → iterate `interaction.steps` +- [ ] Replaced `client.chats.create()` / manual history → `previous_interaction_id` +- [ ] Removed all `types.*` wrappers (`GenerateContentConfig`, `Tool`, `Content`, `Part`) — Interactions API uses plain dicts. Check feature docs for exact format. +- [ ] Moved `response_format` from `GenerateContentConfig` to top-level parameter +- [ ] Replaced `generate_content_stream()` → `stream=True` + step-based event handling +- [ ] Updated function calling: candidates-based → step-based tool lifecycle +- [ ] REST: Changed endpoint to `/v1beta/interactions` +- [ ] REST: Add `Api-Revision: 2026-05-20` header (SDK ≥ 2.0.0 sets it automatically) +- [ ] Replaced `google-generativeai` (Python) → `google-genai` ≥ 2.0.0 +- [ ] Replaced `@google/generative-ai` (JS) → `@google/genai` ≥ 2.0.0 +- [ ] Updated all import statements to match new package names + +### Model String Updates + +- [ ] Replaced `gemini-2.0-*` model strings with current equivalents +- [ ] Replaced `gemini-1.5-*` model strings with current equivalents +- [ ] Consider upgrading `gemini-3.7-flash` → `gemini-3.8-flash` +- [ ] Consider upgrading `gemini-3.6-flash` → `gemini-3.8-flash` +- [ ] Consider upgrading `gemini-3.5-flash` → `gemini-3.8-flash` +- [ ] Consider upgrading `gemini-3-flash-preview` → `gemini-3.8-flash` +- [ ] Consider upgrading `gemini-2.5-flash` → `gemini-3.8-flash` +- [ ] Consider upgrading `gemini-3.1-flash-lite` → `gemini-3.5-flash-lite` +- [ ] Consider upgrading `gemini-2.5-flash-lite` → `gemini-3.5-flash-lite` or `gemini-3.1-flash-lite` +- [ ] Consider upgrading `gemini-2.5-pro` → `gemini-3.1-pro-preview` + +### Migrate to Gemini 3.8 Flash or Gemini 3.5 Flash-Lite + +Use this checklist if the user requests to migrate to Gemini 3.8 Flash or Gemini 3.5 Flash-Lite. For full documentation of the changes, fetch the [Latest Gemini models guide](https://ai.google.dev/gemini-api/docs/latest-model.md.txt) and look for the migration section. + +- [ ] Updated model name to `gemini-3.8-flash` or `gemini-3.5-flash-lite` (depending on user request) +- [ ] Removed `temperature`, `top_p`, `top_k` from config +- [ ] Replaced `thinking_budget` with `thinking_level` (`minimal`, `low`, `medium`, `high`) + +--- + +## Verify the Migration + +After updating, run a spot-check to confirm the Interactions API is working: + +1. Make a single `client.interactions.create()` call with a simple input +2. Assert `interaction.steps` is not empty +3. Assert at least one step has `type == "model_output"` with non-empty text +4. For multi-turn, verify `previous_interaction_id` preserves context across turns + +For verification code snippets, fetch the [Migration Guide](https://ai.google.dev/gemini-api/docs/migrate-to-interactions.md.txt). diff --git a/.claude/skills/gemini/references/official/gemini-live-api-dev/SKILL.md b/.claude/skills/gemini/references/official/gemini-live-api-dev/SKILL.md new file mode 100644 index 00000000..69393dc4 --- /dev/null +++ b/.claude/skills/gemini/references/official/gemini-live-api-dev/SKILL.md @@ -0,0 +1,399 @@ +--- +name: gemini-live-api-dev +description: Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), native audio features, function calling, session management, ephemeral tokens for client-side auth, live translation, and all Live API configuration options. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript). +--- + +# Gemini Live API Development Skill + +## Overview + +The Live API enables **low-latency, real-time voice and video interactions** with Gemini over WebSockets. It processes continuous streams of audio, video, or text to deliver immediate, human-like spoken responses. + +Key capabilities: +- **Bidirectional audio streaming** — real-time mic-to-speaker conversations +- **Video streaming** — send camera/screen frames alongside audio +- **Text input/output** — send and receive text within a live session +- **Audio transcriptions** — get text transcripts of both input and output audio +- **Voice Activity Detection (VAD)** — automatic interruption handling +- **Native audio** — thinking (with configurable `thinkingLevel`) +- **Function calling** — synchronous tool use +- **Google Search grounding** — ground responses in real-time search results +- **Session management** — context compression, session resumption, GoAway signals +- **Ephemeral tokens** — secure client-side authentication + +> [!NOTE] +> The Live API currently **only supports WebSockets**. For WebRTC support or simplified integration, use a [partner integration](#partner-integrations). + +## Models + +- `gemini-3.1-flash-live-preview` — Optimized for low-latency, real-time dialogue. Native audio output, thinking (via `thinkingLevel`). 128k context window. **This is the recommended model for all Live API use cases.** +- `gemini-3.5-transcribe-live` — Real-time streaming speech-to-text with interim hypotheses, finalized transcripts, smart formatting, and Hybrid VAD. +- `gemini-3.5-live-translate-preview` — Real-time streaming translation model. + +> [!WARNING] +> The following Live API models are **deprecated** and will be shut down. Migrate to `gemini-3.1-flash-live-preview`. +> - `gemini-2.5-flash-native-audio-preview-12-2025` — Migrate to `gemini-3.1-flash-live-preview`. +> - `gemini-live-2.5-flash-preview` — Released June 17, 2025. Shutdown: December 9, 2025. +> - `gemini-2.0-flash-live-001` — Released April 9, 2025. Shutdown: December 9, 2025. + +## SDKs + +- **Python**: `google-genai` — `pip install google-genai` +- **JavaScript/TypeScript**: `@google/genai` — `npm install @google/genai` + +> [!WARNING] +> Legacy SDKs `google-generativeai` (Python) and `@google/generative-ai` (JS) are deprecated. Use the new SDKs above. + +## Partner Integrations + +To streamline real-time audio/video app development, use a third-party integration supporting the Gemini Live API over **WebRTC** or **WebSockets**: + +- [LiveKit](https://docs.livekit.io/agents/models/realtime/plugins/gemini/) — Use the Gemini Live API with LiveKit Agents. +- [Pipecat by Daily](https://docs.pipecat.ai/guides/features/gemini-live) — Create a real-time AI chatbot using Gemini Live and Pipecat. +- [Fishjam by Software Mansion](https://docs.fishjam.io/tutorials/gemini-live-integration) — Create live video and audio streaming applications with Fishjam. +- [Vision Agents by Stream](https://visionagents.ai/integrations/gemini) — Build real-time voice and video AI applications with Vision Agents. +- [Voximplant](https://voximplant.com/products/gemini-client) — Connect inbound and outbound calls to Live API with Voximplant. +- [Firebase AI SDK](https://firebase.google.com/docs/ai-logic/live-api?api=dev) — Get started with the Gemini Live API using Firebase AI Logic. + +## Audio Formats + +- **Input**: Raw PCM, little-endian, 16-bit, mono. 16kHz native (will resample others). MIME type: `audio/pcm;rate=16000` +- **Output**: Raw PCM, little-endian, 16-bit, mono. 24kHz sample rate. + +> [!IMPORTANT] +> Use `send_realtime_input` / `sendRealtimeInput` for all real-time user input (audio, video, **and text**). `send_client_content` / `sendClientContent` is **only** supported for seeding initial context history (requires setting `initial_history_in_client_content` in `history_config`). Do **not** use it to send new user messages during the conversation. + +> [!WARNING] +> Do **not** use `media` in `sendRealtimeInput`. Use the specific keys: `audio` for audio data, `video` for images/video frames, and `text` for text input. + +--- + +## Quick Start + +### Authentication + +#### Python + +```python +from google import genai + +client = genai.Client(api_key="YOUR_API_KEY") +``` + +#### JavaScript + +```js +import { GoogleGenAI } from '@google/genai'; + +const ai = new GoogleGenAI({ apiKey: 'YOUR_API_KEY' }); +``` + +### Connecting to the Live API + +#### Python +```python +from google.genai import types + +config = types.LiveConnectConfig( + response_modalities=[types.Modality.AUDIO], + system_instruction=types.Content( + parts=[types.Part(text="You are a helpful assistant.")] + ) +) + +async with client.aio.live.connect(model="gemini-3.1-flash-live-preview", config=config) as session: + pass # Session is active +``` + +#### JavaScript +```js +const session = await ai.live.connect({ + model: 'gemini-3.1-flash-live-preview', + config: { + responseModalities: ['audio'], + systemInstruction: { parts: [{ text: 'You are a helpful assistant.' }] } + }, + callbacks: { + onopen: () => console.log('Connected'), + onmessage: (response) => console.log('Message:', response), + onerror: (error) => console.error('Error:', error), + onclose: () => console.log('Closed') + } +}); +``` + +### Sending Text + +#### Python +```python +await session.send_realtime_input(text="Hello, how are you?") +``` + +#### JavaScript +```js +session.sendRealtimeInput({ text: 'Hello, how are you?' }); +``` + +### Sending Audio + +#### Python +```python +await session.send_realtime_input( + audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000") +) +``` + +#### JavaScript +```js +session.sendRealtimeInput({ + audio: { data: chunk.toString('base64'), mimeType: 'audio/pcm;rate=16000' } +}); +``` + +### Sending Video + +#### Python +```python +# frame: raw JPEG-encoded bytes +await session.send_realtime_input( + video=types.Blob(data=frame, mime_type="image/jpeg") +) +``` + +#### JavaScript +```js +session.sendRealtimeInput({ + video: { data: frame.toString('base64'), mimeType: 'image/jpeg' } +}); +``` + +### Receiving Audio and Text + +> [!IMPORTANT] +> A single server event can contain **multiple content parts simultaneously** (e.g., audio chunks and transcript). Always process **all** parts in each event to avoid missing content. + +#### Python +```python +async for response in session.receive(): + content = response.server_content + if content: + # Audio — process ALL parts in each event + if content.model_turn: + for part in content.model_turn.parts: + if part.inline_data: + audio_data = part.inline_data.data + # Transcription + if content.input_transcription: + print(f"User: {content.input_transcription.text}") + if content.output_transcription: + print(f"Gemini: {content.output_transcription.text}") + # Interruption + if content.interrupted is True: + pass # Stop playback, clear audio queue +``` + +#### JavaScript +```js +// Inside the onmessage callback +const content = response.serverContent; +if (content?.modelTurn?.parts) { + for (const part of content.modelTurn.parts) { + if (part.inlineData) { + const audioData = part.inlineData.data; // Base64 encoded + } + } +} +if (content?.inputTranscription) console.log('User:', content.inputTranscription.text); +if (content?.outputTranscription) console.log('Gemini:', content.outputTranscription.text); +if (content?.interrupted) { /* Stop playback, clear audio queue */ } +``` + +--- + +## Live Translation (Gemini Live Translate) + +The Live API supports real-time, low-latency streaming translation of speech (audio) across 70+ languages. For full details on options and capabilities, see the [Live Translate Guide](https://ai.google.dev/gemini-api/docs/live-api/live-translate.md.txt). + +### Model +- `gemini-3.5-live-translate-preview` — The recommended translation model for all Live Translate use cases. + +### Configuration (`TranslationConfig`) + +To enable translation, specify a `TranslationConfig` object inside your live session setup: + +- **Python SDK**: Configure the connection using `translation_config` on `LiveConnectConfig`: + ```python + config = types.LiveConnectConfig( + response_modalities=[types.Modality.AUDIO], + translation_config=types.TranslationConfig( + target_language_code="es", # Target language code (e.g. es, fr, pl) + echo_target_language=True, + ), + input_audio_transcription=types.AudioTranscriptionConfig(), + output_audio_transcription=types.AudioTranscriptionConfig(), + ) + ``` +- **Raw WebSockets**: Place `translationConfig` inside `generationConfig`: + ```json + { + "setup": { + "model": "models/gemini-3.5-live-translate-preview", + "generationConfig": { + "responseModalities": ["AUDIO"], + "translationConfig": { + "targetLanguageCode": "es", + "echoTargetLanguage": true + } + } + } + } + ``` + +--- + +## Live Streaming Transcription (Gemini Live Transcribe) + +The Live API supports real-time streaming speech-to-text over WebSockets with low-latency interim hypotheses, finalized transcripts, and Hybrid VAD. For full details, see the [Live Transcription Guide](https://ai.google.dev/gemini-api/docs/live-api/live-transcribe.md.txt) and [Colab Cookbook](https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/Get_started_transcribe.ipynb). + +### Model +- `gemini-3.5-transcribe-live` + +### Modes +- `smart`: cleans up filler words, resolves inline self-corrections, and structures formatting. +- `verbatim` (default): exact word-for-word transcript. + +### Python +```python +config = types.LiveConnectConfig( + response_modalities=["TEXT"], + input_audio_transcription=types.AudioTranscriptionConfig(), +) + +async with client.aio.live.connect(model="gemini-3.5-transcribe-live", config=config) as session: + # Stream audio + await session.send_realtime_input(audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000")) + # Hybrid VAD: notify turn end on client-detected silence for zero latency + await session.send_realtime_input(audio_stream_end=True) +``` + +### JavaScript +```javascript +const session = await ai.live.connect({ + model: 'gemini-3.5-transcribe-live', + config: { + responseModalities: ['text'], + inputAudioTranscription: { mode: 'smart' } + }, + callbacks: { + onmessage: (msg) => { + if (msg.serverContent?.interimInputTranscription) { + console.log('Interim:', msg.serverContent.interimInputTranscription.text); + } + if (msg.serverContent?.inputTranscription) { + console.log('Final:', msg.serverContent.inputTranscription.text); + } + } + } +}); + +session.sendRealtimeInput({ audio: { data: chunkBase64, mimeType: 'audio/pcm;rate=16000' } }); +session.sendRealtimeInput({ audioStreamEnd: true }); // Hybrid VAD +``` + +### Raw WebSockets +```json +{ + "setup": { + "model": "models/gemini-3.5-transcribe-live", + "generationConfig": { + "responseModalities": ["TEXT"], + "speechConfig": { + "voiceConfig": {} + } + }, + "inputAudioTranscription": { + "mode": "smart" + } + } +} +``` + +--- + +## Limitations + +- **Response modality** — Only `TEXT` **or** `AUDIO` per session, not both. Native audio models only support audio. +- **Audio-only session** — 15 min without compression +- **Audio+video session** — 2 min without compression +- **Connection lifetime** — ~10 min (use session resumption) +- **Context window** — 128k tokens (native audio) / 32k tokens (standard) +- **Async function calling** — Not yet supported; function calling is synchronous only. The model will not start responding until you've sent the tool response. +- **Proactive audio** — Not yet supported in Gemini 3.1 Flash Live. Remove any configuration for this feature. +- **Affective dialogue** — Not yet supported in Gemini 3.1 Flash Live. Remove any configuration for this feature. +- **Code execution** — Not supported +- **URL context** — Not supported + +## Migrating from Gemini 2.5 Flash Live + +When migrating from `gemini-2.5-flash-native-audio-preview-12-2025` to `gemini-3.1-flash-live-preview`: + +1. **Model string** — Update from `gemini-2.5-flash-native-audio-preview-12-2025` to `gemini-3.1-flash-live-preview`. +2. **Thinking configuration** — Use `thinkingLevel` (`minimal`, `low`, `medium`, `high`) instead of `thinkingBudget`. Default is `minimal` for lowest latency. +3. **Server events** — A single event can contain multiple content parts simultaneously (audio + transcript). Process **all** parts in each event. +4. **Client content** — `send_client_content` is only for seeding initial context history (set `initial_history_in_client_content` in `history_config`). Use `send_realtime_input` for text during conversation. +5. **Turn coverage** — Defaults to `TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO` instead of `TURN_INCLUDES_ONLY_ACTIVITY`. If sending constant video frames, consider sending only during audio activity to reduce costs. +6. **Async function calling** — Not yet supported. Function calling is synchronous only. +7. **Proactive audio & affective dialogue** — Not yet supported. Remove any configuration for these features. + +## Best Practices + +1. **Use headphones** when testing mic audio to prevent echo/self-interruption +2. **Enable context window compression** for sessions longer than 15 minutes +3. **Implement session resumption** to handle connection resets gracefully +4. **Use ephemeral tokens** for client-side deployments — never expose API keys in browsers +5. **Use `send_realtime_input`** for all real-time user input (audio, video, text). Reserve `send_client_content` only for seeding initial context history +6. **Send `audioStreamEnd`** when the mic is paused to flush cached audio +7. **Clear audio playback queues** on interruption signals +8. **Process all parts** in each server event — events can contain multiple content parts + +## Documentation Lookup + +### When MCP is Installed (Preferred) + +If the **`search_docs`** tool (from the Google MCP server) is available, use it as your **only** documentation source: + +1. Call `search_docs` with your query +2. Read the returned documentation +3. **Trust MCP results** as source of truth for API details — they are always up-to-date. + +> [!IMPORTANT] +> When MCP tools are present, **never** fetch URLs manually. MCP provides up-to-date, indexed documentation that is more accurate and token-efficient than URL fetching. + +### When MCP is NOT Installed (Fallback Only) + +If no MCP documentation tools are available, fetch from the official docs index: + +**llms.txt URL**: `https://ai.google.dev/gemini-api/docs/llms.txt` + +This index contains links to all documentation pages in `.md.txt` format. Use web fetch tools to: + +1. Fetch `llms.txt` to discover available documentation pages +2. Fetch specific pages (e.g., `https://ai.google.dev/gemini-api/docs/live-session.md.txt`) + +### Key Documentation Pages + +> [!IMPORTANT] +> Those are not all the documentation pages. Use the `llms.txt` index to discover available documentation pages + +- [Live API Overview](https://ai.google.dev/gemini-api/docs/live.md.txt) — getting started, raw WebSocket usage +- [Live Transcription](https://ai.google.dev/gemini-api/docs/live-api/live-transcribe.md.txt) — real-time speech-to-text, interim hypotheses, smart formatting, and Hybrid VAD +- [Live Translate](https://ai.google.dev/gemini-api/docs/live-api/live-translate.md.txt) — configuration options and capabilities for translation +- [Live API Capabilities Guide](https://ai.google.dev/gemini-api/docs/live-guide.md.txt) — voice config, transcription config, native audio (thinking), VAD configuration, media resolution +- [Live API Tool Use](https://ai.google.dev/gemini-api/docs/live-tools.md.txt) — function calling (sync and async), Google Search grounding +- [Session Management](https://ai.google.dev/gemini-api/docs/live-session.md.txt) — context window compression, session resumption, GoAway signals +- [Ephemeral Tokens](https://ai.google.dev/gemini-api/docs/ephemeral-tokens.md.txt) — secure client-side authentication for browser/mobile +- [WebSockets API Reference](https://ai.google.dev/api/live.md.txt) — raw WebSocket protocol details + +## Supported Languages + +The Live API supports 70 languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Russian, and many more. Native audio models automatically detect and switch languages. diff --git a/.claude/skills/gemini/references/official/gemini-omni-flash-api/SKILL.md b/.claude/skills/gemini/references/official/gemini-omni-flash-api/SKILL.md new file mode 100644 index 00000000..844fefcb --- /dev/null +++ b/.claude/skills/gemini/references/official/gemini-omni-flash-api/SKILL.md @@ -0,0 +1,450 @@ +--- +name: gemini-omni-flash-api +description: Use this skill for generative video editing, text-to-video, image-referenced video generation, first-frame-to-video, first-and-last-frame transitions, and video extensions using Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) via the official google-genai SDK. Includes workflows for pre-processing/optimizing high-resolution or long source videos with ffmpeg, stripping audio for full sound regeneration, and handling turn-by-turn video editing and parallel execution. +--- + +# Gemini Omni Flash Skill + +This skill uses the Gemini Omni 1.1 Flash model (`gemini-omni-1.1-flash`) to perform text to video generation, image to video generation (first frame and last frame transitions), video extensions (up to 40s), and video editing. + +> [!WARNING] +> **Important Regional Restrictions**: Uploading videos to use for video edits or extensions is **NOT** available in the EEA, Switzerland, the United Kingdom, and some US states. If a video-to-video edit completes quickly with empty outputs (`total_output_tokens: 0` or no video content), it is likely due to this restriction. + +## Core capabilities + +1. **Text to video**: Generating videos from a text prompt. +2. **First frame to video**: Generating videos from a starting image (`--first-frame`). +3. **First and last frame transition**: Generating videos interpolating between a starting image and a final image (`--first-frame` and `--last-frame`; note: `--last-frame` must be used with `--first-frame`). +4. **Video extensions**: Extending existing videos by up to 10 seconds per turn, up to a total length of 40 seconds (`--extend` or `--previous-interaction-id`). +5. **Video editing and refinement**: Editing existing videos (maximum duration 10 seconds), applying stylistic changes, or performing inpainting/outpainting. +6. **Image and video referenced generation**: Using style, character, or object references from images or videos to guide video generation. + +## Workflow + +1. **Analyze request**: Determine the target task (e.g., first-frame-to-video, first-and-last-frame transition, video extension, reference-guided editing) and identify any input media assets. +2. **Run SDK scripts**: + + * Directly run the appropriate utility (`scripts/video/generate_video.py` or `scripts/upload_file.py`). + * Configure settings like `--aspect-ratio` (e.g. `16:9`, `9:16`), `--resolution` (`360p`, `720p`, `1080p`, `4k`; default: `720p`), and `--duration` (any integer between `3` and `10` seconds, e.g. `3`, `5`, `10`). *Note: `4k` requests take longer to generate.* + +3. **Retrieve and process output**: Outputs are saved to the local filesystem (e.g. `media/`). Report back the completed media path to the user. + +## Reference Documentation + +* **Interactions API**: All operations and state management for the Gemini Omni 1.1 Flash model (`gemini-omni-1.1-flash`) are handled via the [Interactions API](https://ai.google.dev/gemini-api/docs/interactions-overview). +* **Files API**: Input media files (such as reference images and videos) must be uploaded via the [Files API](https://ai.google.dev/gemini-api/docs/files) first before being referenced in generations. The uploaded file URI and MIME type are then included in the `interactions.create` input parts array. +* **[Gemini API Skill Reference](https://github.com/google-gemini/gemini-skills/blob/main/skills/gemini-api-dev/SKILL.md)**: Platform-wide guidelines, current model specifications, and SDK usage rules for the Gemini API. + +## Dependencies and Prerequisites + +* **Python SDK (`google-genai`)**: Requires `google-genai >= 2.19.0` (Python) to support the `interactions` client and full video output resolution configuration (`360p`, `720p`, `1080p`, `4k`). Install or upgrade using: + ```bash + pip install -U google-genai + ``` +* **Python Runtime**: Requires **Python >= 3.10** (for compatibility with modern `google-genai` SDK types and methods). +* **ffmpeg & ffprobe**: `prep_video.py`, `inspect_video.py`, and `generate_video.py` (when stripping audio via `--strip-audio`) require `ffmpeg` and `ffprobe` binaries installed and available in your system `PATH`. +* **API Key**: Set the `GEMINI_API_KEY` environment variable: + ```bash + export GEMINI_API_KEY="your-api-key" + ``` + +## Available scripts + +Use the following Python scripts to upload media with the Files API, prepare input videos with ffmpeg, and generate video outputs using the Interactions API. + +1. **[upload_file.py](scripts/upload_file.py)**: Uploads local media (images and videos) to the Files API and polls until `ACTIVE`. If uploading a video larger than 25MB, it prints an informative warning/tip highlighting that Gemini Omni Flash is optimized for editing 10s videos at 720p/24fps, and recommends pre-processing with `prep_video.py` first to speed up the upload. + + ```bash + ./scripts/upload_file.py path/to/image.png + ``` + +2. **[generate_video.py](scripts/video/generate_video.py)**: Performs end-to-end video generation and downloads the output video. It detects and uploads local media references (images or videos) before calling the Interactions API. Large video assets (>25MB) will trigger informative pre-processing recommendations without blocking the upload. + + * **Text to video**: + + ```bash + ./scripts/video/generate_video.py "A close-up of a cat drinking tea" --output media/cat_tea.mp4 + ``` + + * **Output resolution options (`--resolution`)**: + + Gemini Omni 1.1 Flash natively supports four output resolutions across both landscape (`16:9`) and portrait (`9:16`) aspect ratios: + - `360p`: `640x360` (16:9) or `360x640` (9:16) + - `720p`: `1280x720` (16:9) or `720x1280` (9:16) — *(default)* + - `1080p`: `1920x1080` (16:9) or `1080x1920` (9:16) + - `4k`: `3840x2160` (16:9) or `2160x3840` (9:16) + + ```bash + # High-definition (1080p) + ./scripts/video/generate_video.py "A cinematic drone shot over misty mountains at sunrise" --resolution 1080p --output media/mountains_1080p.mp4 + + # Ultra-high-definition 4K (Note: 4K requests take longer to generate; pass --timeout if needed) + ./scripts/video/generate_video.py "A macro shot of a dewdrop on a flower petal in golden sunlight" --resolution 4k --timeout 900 --output media/flower_4k.mp4 + ``` + + * **Configurable request timeouts (`--timeout`)**: + + Default HTTP timeout is `600` seconds (10 minutes). For computationally intensive requests — such as extending a 30s video in 4K by 10s (up to the maximum 40s total video length) — generation can take several minutes. Use `--timeout 900` (or `1200`) to provide an extended execution budget. + + * **First frame to video**: + + ```bash + ./scripts/video/generate_video.py "The waves crash against the shore." --first-frame start.png --output media/waves.mp4 + ``` + + * **First and last frame transition**: + + Provide a starting frame and an ending frame to generate a smooth transition between them (note: `--last-frame` **must** be used together with `--first-frame`): + + ```bash + ./scripts/video/generate_video.py "A smooth timelapse from sunrise to sunset" --first-frame start.png --last-frame end.png --output media/interpolation.mp4 + ``` + + * **Looping video (identical start and end frame)**: + + ```bash + ./scripts/video/generate_video.py "A crystal orb spinning continuously in place" --first-frame orb.png --last-frame orb.png --output media/loop.mp4 + ``` + + * **Image-referenced video generation**: + + ```bash + ./scripts/video/generate_video.py "A cybernetic warrior in the style of " --image reference.png --output media/warrior.mp4 + ``` + + * **Video-referenced video generation**: + + Provide one or more reference videos (`--video-reference` / `-vr`) to guide character, object, or motion style (ideal duration is ~3s, up to 3 reference videos recommended): + + ```bash + ./scripts/video/generate_video.py "A musician playing cello in the style of " --video-reference ref_dance.mp4 --output media/cello.mp4 + ``` + + * **Video extension (extend an existing video)**: + + Extend an existing video by up to 10 seconds (total duration up to 40 seconds): + + ```bash + ./scripts/video/generate_video.py "The scene continues as the sun sets over the horizon" --extend media/sunset.mp4 --output media/sunset_extended.mp4 + ``` + + * **Video extension with reference images and reference videos**: + + Prompt-based extension allows passing reference images and reference videos simultaneously: + + ```bash + ./scripts/video/generate_video.py "Extend this video. The character in enters dancing like the dancer in ." --extend media/sunset.mp4 --image character.png --video-reference dance_ref.mp4 --output media/sunset_extended_with_refs.mp4 + ``` + + * **Video editing (keep original audio)**: + + ```bash + ./scripts/video/generate_video.py "Transform the style to Japanese anime" --video input.mp4 --output media/anime_style.mp4 + ``` + + * **Video editing (regenerate all audio from scratch)**: + + ```bash + ./scripts/video/generate_video.py "Transform the style to Japanese anime" --video input.mp4 --strip-audio --output media/anime_style_new_audio.mp4 + ``` + + * **Turn-by-turn video editing (edit previous interaction)**: + + Edit a prior video generation without re-uploading assets by passing the interaction ID: + + ```bash + ./scripts/video/generate_video.py "Change the setting to a snowy winter wonderland." --previous-interaction-id "v1_..." --output media/winter_wonderland.mp4 + ``` + + * **Turn-by-turn video extension (extend previous interaction)**: + + Extend a prior video generation by passing the previous interaction ID: + + ```bash + ./scripts/video/generate_video.py "Extend this video. The character turns around and begins to run." --previous-interaction-id "v1_..." --output media/extended_turn.mp4 + ``` + + * **Parallel batch execution (prompts file)**: Run multiple prompts from a line-by-line text file concurrently: + + ```bash + ./scripts/video/generate_video.py --prompts-file prompts.txt --concurrency 3 + ``` + + * **Parallel batch execution (JSON config)**: Execute fully configured, distinct generation and editing jobs in parallel: + + ```bash + ./scripts/video/generate_video.py --batch jobs.json --concurrency 3 + ``` + + *Example `jobs.json`:* + + ```json + [ + { + "prompt": "A smooth timelapse from sunrise to sunset.", + "first_frame": "start.png", + "last_frame": "end.png", + "resolution": "1080p", + "output": "media/interpolation.mp4" + }, + { + "prompt": "Extend this video. The scene continues with the character in dancing like .", + "extend": "media/sunset.mp4", + "image": "character.png", + "video_reference": "dance_ref.mp4", + "output": "media/extended_with_refs.mp4" + }, + { + "prompt": "A macro shot of a crystal orb refracting cosmic nebula colors.", + "resolution": "4k", + "output": "media/nebula_orb_4k.mp4" + }, + { + "prompt": "A musician playing cello in the style of .", + "video_reference": "cello_ref.mp4", + "output": "media/cello.mp4" + }, + { + "prompt": "Transform the style to Japanese anime.", + "video": "input.mp4", + "output": "media/anime_style.mp4", + "strip_audio": false, + "aspect_ratio": "16:9" + } + ] + ``` + +3. **[inspect_video.py](scripts/video/inspect_video.py)**: Inspects a local video file (using `ffprobe`) to check its duration, resolution, frame rate (FPS), audio stream presence, and format details. + + ```bash + ./scripts/video/inspect_video.py media/output.mp4 + ``` + + * To get a pre-parsed, structured JSON summary: + + ```bash + ./scripts/video/inspect_video.py media/output.mp4 --json + ``` + + * To get the complete, unmodified `ffprobe` raw JSON dump: + + ```bash + ./scripts/video/inspect_video.py media/output.mp4 --raw + ``` + +4. **[prep_video.py](scripts/video/prep_video.py)**: Normalizes, trims, and formats any video file to fit standard Gemini Omni Flash generation and editing limits. It handles timecode-based trimming, optional frame rate conversion, and proportional scaling of large videos (max 1280x720 for landscape, 720x1280 for portrait) to optimize upload times without stretching. If the video is longer than 10 seconds and the script is run interactively (in a TTY), it prompts the user to select the first 10s, last 10s, or enter a custom timecode (defaulting to the first 10s). + + * **Trim first 10s (default)**: + + ```bash + ./scripts/video/prep_video.py path/to/source.mp4 + ``` + + or explicitly specify the start and duration: + + ```bash + ./scripts/video/prep_video.py path/to/source.mp4 --start 0 --duration 10 + ``` + + * **Trim last 10s** (automatically calculates starting point based on source length): + + ```bash + ./scripts/video/prep_video.py path/to/source.mp4 --start last + ``` + + * **Trim 10s starting at specific timecode** (MM:SS or HH:MM:SS): + + ```bash + ./scripts/video/prep_video.py path/to/source.mp4 --start 00:03 --output media/custom.mp4 + ``` + + * **Custom frame rate and resolution**: + + ```bash + ./scripts/video/prep_video.py path/to/source.mp4 --fps 30 --resolution 1920x1080 + ``` + + * **Strip audio for audio regeneration**: + + ```bash + ./scripts/video/prep_video.py path/to/source.mp4 --strip-audio --output media/video_with_no_audio.mp4 + ``` + +## Audio handling in video editing + +When editing a source video that contains audio, you must choose between keeping the original audio or regenerating all audio from scratch. + +* **Keep original audio**: By default, Gemini Omni Flash preserves the existing audio layer (though it may modify or adapt it slightly during generation). Use this when the original background music, dialogue, or sound effects are desired. +* **Regenerate all audio from scratch**: If you want Gemini Omni Flash to re-create a brand-new audio layer tailored to the new visual style or prompt, you **must** upload the video with its audio stream stripped out. If any audio stream is present, Gemini Omni Flash will attempt to preserve/modify it instead of starting from scratch. + + * Use `--strip-audio` (or `-a`) when pre-processing with `scripts/video/prep_video.py` or executing `scripts/video/generate_video.py`. + * This forces Gemini Omni Flash to perform full audio generation. + +## Prompting Gemini Omni Flash + +### Single scene + +By default Gemini Omni Flash will try to create a video with a few different shots. It'll attempt to craft an interesting narrative based on the prompt. + +If you need the output video to contain a single scene, you must prompt for that: + +* In a single unbroken scene +* In a single continuous shot +* No scene cuts + +For example: + +``` +Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air. Sound design: Gentle breeze, distant bird chirps. No dialogue. +``` + +### Removing unwanted elements + +If the generated video contains things you don't want, include simple negative prompts to avoid them: + +* No dialogue +* No embellishments +* No extra sound effects + +### Prompts for editing + +Simple prompts work best for video editing. Overly descriptive prompts can lead to unintended changes. + +The following are more examples of simple editing prompts: + +* Make this video anime +* Put a fashionable hat on this person +* Change the lighting to be more dramatic +* Change the text on the sign to say "Omni Flash" + +When editing a specific aspect of the video, include `"Keep everything else the same"` to maintain visual consistency. + +The following are some examples to show how to apply this technique: + +* **Avoid:** `In the video of the man sitting on the sofa, please add a small black cat that runs from the right side of the screen, jumps onto his lap, and then he starts to stroke its head while looking down.` + * **Simplify:** `Add a cat that jumps onto his lap, he begins to pet it. Keep everything else the same.` +* **Avoid:** `Please remove the cell phone that the person is holding in their hand and fill in the background so it looks like they are just holding their hand empty.` + * **Simplify:** `Make the phone invisible. Keep everything else the same.` + +### Prompting the audio + +By default the model will try to generate an appropriate audio track for a video. This might not always be what you want. You can use your prompt to describe the type of audio you want. This is especially important if you want music in your video: + +* Include calm background music +* The video has a high energy techno beat +* The audio is a low tinny radio broadcast in the background, playing a song + +### Timing events + +You can prompt for things to happen at specific times in the video, there is no precise syntax needed and you can use natural language. This is especially useful in creating your own scene cuts, rhythm or rapid fire sequences. See the following for examples: + +* After 3 seconds, a woman enters the scene. +* At 5s the chorus starts in the background audio. +* Every 2s cut to a new frame. +* In a rapid fire sequence, every half a second (12 frames at 24fps) change the scene to a new location. + +You can also use a timecode syntax: + +``` +[0-3s] A person is walking +[3-6s] They stop and turn around +[6-10s] They start running +``` + +### Meta prompting + +You can ask Gemini Omni Flash to pay attention to general qualities or principles of video generation: + +* Consider micro-detail, expression and timing to create a very rich, detailed but entirely natural scene. +* Be extremely detailed in your descriptions of characters and environments. Apply costume design principles to characters. Be very specific about the people, items and objects in the scene. +* Include plenty of appropriate detail in the background elements to make the scene feel realistic and natural. +* Make a rapid fire video that shows a different rare `[thing]` every 1s, upbeat music, include text to label the thing. + +### Text in videos + +You can prompt to include text in your video and Gemini Omni will render in a way that is correct and readable. If there will be naturally occurring text in your video, even in background elements, it can help to define what it should say. + +* One word on the screen at a time: "did, you, know, that, Omni, can, do, awesome, text?" Each word appears for 1s with a different animated style. No dialogue. +* There is a street sign that says: "This is an AI generation by Omni", there is a storefront that says: "All you need AI", there's a car with the number plate: "OMNI1.1" + +### Prompts for extending a video + +With Gemini Omni 1.1 Flash you can extend videos with prompts like, `"Extend this video"` or `"The scene continues"`. You can extend videos by 10s, up to a total length of 40s. + +Omni creates an extension that keeps video, motion, characters and audio coherent by using the last 10s of your original video as context. Some of the final frames in your input video will be edited to make the transition seamless. + +> [!TIP] +> **Extending with References**: Video extensions can be done with a prompt (e.g., `"Extend this video"`, `"The scene continues"`) without setting the API's `task="extend"` parameter. Omitting the `task` parameter allows passing in reference images (`--image`) and reference videos (`--video-reference`) during video extensions to introduce new characters, objects, or styles seamlessly into the extended scene. If `task="extend"` is explicitly set, multimodal references cannot be passed. + +When extending, all of this guide's Omni prompting tips still apply: + +* Describe the audio in your extended scene, especially if you need it to change, `"The music continues into the chorus"` +* Describe if the scene continues, or if there is a shot cut to a new scene (perhaps with the same characters), `"Show the same characters in the next scene"` +* Include images and videos as references when extending to help keep your outputs accurate, or to introduce new characters, `"The person shown in the reference image enters the scene"`, `"The dog in the reference video jumps onto the sofa"` +* If using timestamps or a timecode syntax, 0s refers to the beginning of the extended part of the video. If extending a 10s video, the scene cut in this prompt will happen after 12s: `"After 2s cut to a new scene with the same characters"` + +### Video extension constraints and guidelines + +* **Duration limit**: Input videos for extension must be 10 seconds or less in length when uploading (unless using multi-turn). +* **Spoken dialogue on uploaded videos**: Currently, you cannot extend an uploaded video where someone is talking to add additional dialogue (it is supported if the character remains silent or if the prompt does not add dialogue). +* **Multi-turn voice extension**: Generating spoken dialogue or speech is supported when extending previously generated videos via multi-turn (`previous_interaction_id`). +* **Task parameter recommendation**: We recommend relying primarily on prompting and using the `task="extend"` parameter only when prompting alone does not work and you need to help the model understand which mode it should use, as setting the `task` field adds constraints (such as disabling multimodal reference inputs). + +### Using tags in prompts to set image and video roles + +You can use tags to bind uploaded media to specific generation roles. This lets you specify whether each image or video is a starting frame, a final frame, or a reference. + +#### Simple tags (recommended) + +For simple cases where media roles are clear from the prompt, you can bind images and videos to roles directly: + +* **``**: use the image as the starting frame of the video, for example: ` a woman is walking` +* **``**: use the image as the final frame of the video to transition to. Must be used with ``, for example: ` a woman is walking` +* **``**: use the image as a reference, for example: `in the style of a woman is walking` (combines style reference from the first image and subject reference from the second image). Image references start from 0. +* **``**: use the video as a character or object reference, for example: `the person in is playing the violin`. Video references also start from 0. + +> [!NOTE] +> **Reference Video Guidelines**: +> * **Duration**: Ideal reference videos are **~3 seconds**, though longer videos are fine. Use `prep_video.py --duration 3` if you want to trim longer source files down to reference length. +> * **Quantity**: Up to **3 reference videos** is ideal, though more can also be used. + +The following is an example with 6 reference images: + +``` +[0-3s] A studio fashion sequence. Starting with woman , she is holding +[3-6s] Then we see the man holding +[6-10s] And finally another woman who is holding while walking. +``` + +#### Declaring sources and references + +For more complex cases with multiple media inputs and multiple roles, you can use explicit prefix tags paired with natural language instructions. You should declare these sources and references at the start of your prompt. + + * `[# Sources @Image1]` will use the first image as the starting frame. + * `[# Sources @Image1 @Image2]` will use the first image as the starting frame and the second image as the final frame. + * `[# Sources @Image1 @Image1]` will use the first image as both the first frame and the last frame, creating a video that loops. + * `[# Sources @Image1] [# References @Image2]` will use the first image as the starting frame and the second image as a reference. + * `[# Sources @Video1]` will use the video as the primary source video to edit or modify. + * `[# Sources @Video1]` will use the video from the previous turn to extend. + * `[# References @Image1]` will use the first image as a reference. + * `[# References @Image2]` will use the second image as a reference. + * `[# References @Image1 @Image2]` will use both images as references. + * `[# References @Video1]` will use the first video as a reference. + * `[# References @Image1 @Video1]` will use both an image and a video as a reference. + +Add guiding instructions at the end of your prompt: + + * For a starting frame: `"Use this image as the starting frame."` + * For a looping video via start and end frames: `"Use this image as the first frame and the last frame."` + * For reference images: `"Use the given image(s) as references for video generation. The images should not be used as literal initial frames."` + * For reference videos: `"Use the given video(s) as references. Do not use them as a source for video editing."` + +Some examples of prompts with source and reference declarations: + +``` +[# Sources @Image1] [# References @Image2] a woman is walking. Use Image1 as the starting frame. Use Image2 as a reference for the video generation. +``` + +``` +[# References @Image1 @Video1] The woman in is playing the violin shown in . Use Video1 as a character reference and Image1 as an object reference. +``` diff --git a/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/upload_file.py b/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/upload_file.py new file mode 100755 index 00000000..1ecb3ade --- /dev/null +++ b/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/upload_file.py @@ -0,0 +1,321 @@ +#!/usr/bin/env python3 +""" +Uploads a file to the Gemini Files API and waits for it to become ACTIVE. +Uses the official google-genai SDK. +""" + +import argparse +import json +import mimetypes +import os +import re +import sys +import time +import urllib.parse +from google import genai +from google.genai import types, errors + +def get_api_key(): + return os.environ.get("GEMINI_API_KEY") + +def strip_query_params(url): + """Strips query parameters from URL for clean logging and security.""" + if not url: + return "" + parsed = urllib.parse.urlparse(url) + return urllib.parse.urlunparse(parsed._replace(query="")) + +def sanitize_error(err): + """Sanitizes error messages by redacting API keys, tokens, query parameters, internal provider URLs, and raw response bodies.""" + if not err: + return "" + + # If it is an SDK APIError, prefer clean code + status + message over raw dictionary dump + if isinstance(err, errors.APIError): + parts = [] + if getattr(err, "code", None): + parts.append(str(err.code)) + if getattr(err, "status", None): + parts.append(f"({err.status})") + if getattr(err, "message", None): + parts.append(f": {err.message}") + elif getattr(err, "details", None): + parts.append(f": {err.details}") + err_str = " ".join(parts) if parts else str(err) + else: + err_str = str(err) + + # 1. Directly redact the active API key value if present + key = get_api_key() + if key and len(key) >= 8: + err_str = err_str.replace(key, "[REDACTED_KEY]") + + # 2. Redact known Google API key & OAuth token patterns: + # Classic Google API key: AIza... (30-45 chars) + err_str = re.sub(r'AIza[0-9A-Za-z_\-]{30,}', '[REDACTED_KEY]', err_str) + # Newer Google Cloud / Gemini API key: AQ.... + err_str = re.sub(r'AQ\.[0-9A-Za-z_\-]{20,}', '[REDACTED_KEY]', err_str) + # Google OAuth access token: ya29.... + err_str = re.sub(r'ya29\.[0-9A-Za-z_\-]+', '[REDACTED_TOKEN]', err_str) + # Bearer tokens + err_str = re.sub(r'Bearer\s+[A-Za-z0-9_\-\.]+', 'Bearer [REDACTED_TOKEN]', err_str, flags=re.IGNORECASE) + + # 3. Strip sensitive query parameters from any URLs (key, token, secret, signature) + err_str = re.sub(r'([?&](?:key|api_key|apiKey|access_token|auth|signature)=)[^&\s"\'<>()]+', r'\1[REDACTED]', err_str) + + # 4. Strip internal Google / provider infrastructure URLs & hosts + err_str = re.sub(r'https?://[a-zA-Z0-9.\-_]*\.corp\.goog[^\s"\'<>)]*', '[INTERNAL_HOST]', err_str) + err_str = re.sub(r'https?://[a-zA-Z0-9.\-_]*\.sandbox\.googleapis\.com[^\s"\'<>)]*', '[SANDBOX_ENDPOINT]', err_str) + err_str = re.sub(r'https?://(?:generativelanguage|aiplatform)\.googleapis\.com/v[0-9a-z_]+/', '[API_ENDPOINT]/', err_str) + + # 5. Redact raw response bodies (once a body= marker is found, redact the remainder) + err_str = re.sub(r'\bbody\s*=.*\Z', 'body=[REDACTED_BODY]', err_str, flags=re.IGNORECASE | re.DOTALL) + + # 6. Strip HTML error pages / raw HTML response bodies + if "(.*?)', err_str, re.IGNORECASE) + if title_match: + err_str = f"HTTP Error Response: {title_match.group(1).strip()}" + else: + err_str = re.sub(r'<[^>]+>', ' ', err_str) + + # 7. Collapse excessive whitespace and cap oversized error dumps + err_str = re.sub(r'\s+', ' ', err_str).strip() + if len(err_str) > 500: + err_str = err_str[:497] + "..." + + return err_str + +def detect_mime_type(file_path): + """Determines MIME type based on file extension, falling back to standard mimetypes module.""" + ext = os.path.splitext(file_path)[1].lower() + mime_map = { + ".png": "image/png", + ".jpg": "image/jpeg", + ".jpeg": "image/jpeg", + ".webp": "image/webp", + ".mp4": "video/mp4", + ".mp3": "audio/mpeg", + ".wav": "audio/wav", + ".pdf": "application/pdf", + ".txt": "text/plain", + } + + if ext in mime_map: + return mime_map[ext] + + mime_type, _ = mimetypes.guess_type(file_path) + if mime_type: + return mime_type + + return "application/octet-stream" + +def upload_file(file_path, display_name=None): + """Performs an upload using google-genai SDK, with automatic pre-processing for large videos.""" + if not get_api_key(): + raise RuntimeError("Error: GEMINI_API_KEY environment variable is not set.") + + file_size = os.path.getsize(file_path) + mime_type = detect_mime_type(file_path) + + # Large video file size check (>25MB) + is_video = mime_type.startswith("video/") + if is_video and file_size > 25 * 1024 * 1024: + size_mb = file_size / (1024 * 1024) + print(f"\nWARNING: Video file '{file_path}' is very large ({size_mb:.2f} MB)!") + print("Note: Gemini Omni Flash is optimized for 10s videos at 720p and 24fps. Uploading very large or") + print("high-resolution videos will significantly increase upload times and may cause Out-Of-Memory (OOM) errors.") + + # Determine if terminal is interactive + if sys.stdin.isatty(): + print("\nWould you like to automatically pre-process this video first using prep_video.py?") + print("This will trim, scale, and optimize the video to ensure a fast, OOM-safe upload.") + try: + choice = input("Pre-process video? [Y/n]: ").strip().lower() + if choice in ("", "y", "yes"): + prepped_output_path = os.path.join("media", f"prepped_{os.path.basename(file_path)}") + os.makedirs("media", exist_ok=True) + + # Resolve prep_video.py script path + import subprocess + prep_script = os.path.join(os.path.dirname(os.path.abspath(__file__)), "video", "prep_video.py") + if not os.path.exists(prep_script): + prep_script = os.path.join(os.path.dirname(os.path.abspath(__file__)), "prep_video.py") + + cmd = [sys.executable, prep_script, file_path, "--output", prepped_output_path] + print(f"Running: {' '.join(cmd)}") + + try: + result = subprocess.run(cmd) + if result.returncode == 0 and os.path.exists(prepped_output_path): + file_path = prepped_output_path + file_size = os.path.getsize(file_path) + print(f"\nPre-processing completed successfully! Proceeding with upload of prepped video ({file_size / (1024*1024):.2f} MB)...") + else: + raise RuntimeError("Error: Video pre-processing failed. Proceeding with original file upload is not recommended.") + except Exception as e: + raise RuntimeError(f"Error executing prep_video.py: {sanitize_error(e)}") + else: + proceed_choice = input("Do you want to proceed with uploading the original large video anyway? [y/N]: ").strip().lower() + if proceed_choice not in ("y", "yes"): + raise RuntimeError("Upload cancelled by user. Please pre-process the video manually first.") + except (KeyboardInterrupt, EOFError): + raise RuntimeError("\nNo input received. Upload cancelled to prevent OOM.") + else: + # Non-interactive mode + if file_size > 100 * 1024 * 1024: # Block files larger than 100MB in non-interactive mode + err_msg = ( + f"Error: Video file is extremely large ({size_mb:.2f} MB) and script is running in non-interactive mode.\n" + "To prevent Out-Of-Memory (OOM) errors, upload has been blocked.\n" + "Please pre-process the video first using prep_video.py." + ) + raise RuntimeError(err_msg) + else: + print("Proceeding with upload in non-interactive mode...", file=sys.stderr) + + if not display_name: + display_name = os.path.basename(file_path) + + print(f"Preparing upload of '{file_path}' ({file_size} bytes, type: {mime_type})...") + + # Step 1: Initialize Client with explicit request timeout bounds (300s = 300,000ms) + client = genai.Client(http_options=types.HttpOptions(timeout=300 * 1000)) + + # Step 2: Upload file using SDK + print("Uploading file bytes using google-genai SDK...") + try: + config = types.UploadFileConfig( + display_name=display_name, + mime_type=mime_type, + ) + file_obj = client.files.upload(file=file_path, config=config) + # Convert Pydantic File model to dictionary with both camelCase and snake_case keys for compatibility + file_dict = json.loads(file_obj.model_dump_json()) + # Ensure URI is stripped of any sensitive query parameters + if "uri" in file_dict and file_dict["uri"]: + file_dict["uri"] = strip_query_params(file_dict["uri"]) + # Add camelCase field for mimeType + if "mime_type" in file_dict: + file_dict["mimeType"] = file_dict["mime_type"] + return file_dict + except errors.APIError as e: + raise RuntimeError(f"API Error uploading file via SDK: {sanitize_error(e)}") + except Exception as e: + raise RuntimeError(f"Error uploading file via SDK: {sanitize_error(e)}") + +def wait_for_active(file_name, poll_interval=3, max_attempts=60, backoff_factor=1.2, max_interval=15, max_timeout=600): + """Polls the file status until state is ACTIVE or FAILED using exponential backoff with finite request bounds.""" + if not get_api_key(): + raise RuntimeError("Error: GEMINI_API_KEY environment variable is not set.") + + print(f"Waiting for file {file_name} to finish processing...") + + client = genai.Client(http_options=types.HttpOptions(timeout=30 * 1000)) + attempt = 0 + current_interval = poll_interval + consecutive_errors = 0 + max_consecutive_errors = 5 + start_time = time.time() + + while attempt < max_attempts: + if time.time() - start_time > max_timeout: + raise TimeoutError(f"Error: Maximum timeout ({max_timeout}s) reached waiting for file {file_name} to become ACTIVE.") + + try: + file_obj = client.files.get(name=file_name) + except errors.APIError as e: + consecutive_errors += 1 + if consecutive_errors >= max_consecutive_errors: + raise RuntimeError(f"Error: Too many consecutive API errors checking status ({sanitize_error(e)}). Exiting.") + + print(f"Warning: API error checking status ({sanitize_error(e)}). Retrying in {current_interval:.1f}s...") + time.sleep(current_interval) + current_interval = min(current_interval * backoff_factor, max_interval) + attempt += 1 + continue + except Exception as e: + consecutive_errors += 1 + if consecutive_errors >= max_consecutive_errors: + raise RuntimeError(f"Error: Too many consecutive errors checking status ({sanitize_error(e)}). Exiting.") + + print(f"Warning: Error checking status ({sanitize_error(e)}). Retrying in {current_interval:.1f}s...") + time.sleep(current_interval) + current_interval = min(current_interval * backoff_factor, max_interval) + attempt += 1 + continue + + # Reset consecutive errors on successful API response + consecutive_errors = 0 + state = file_obj.state + state_str = state.name if hasattr(state, "name") else str(state) + + if state == types.FileState.ACTIVE or state_str == "ACTIVE": + print("File is ACTIVE and ready for generations!") + file_dict = json.loads(file_obj.model_dump_json()) + if "uri" in file_dict and file_dict["uri"]: + file_dict["uri"] = strip_query_params(file_dict["uri"]) + if "mime_type" in file_dict: + file_dict["mimeType"] = file_dict["mime_type"] + return file_dict + elif state == types.FileState.FAILED or state_str == "FAILED": + # Terminal FAILED state: abort immediately without retrying + err_details = getattr(file_obj, "error", None) + err_msg = f"Error: File processing failed on backend for '{file_name}' (State: FAILED)" + if err_details: + err_msg += f": {sanitize_error(err_details)}" + raise RuntimeError(err_msg) + + print(f"Current state: {state_str}. Retrying in {current_interval:.1f}s...") + time.sleep(current_interval) + + # Increase interval for the next poll (backoff) + current_interval = min(current_interval * backoff_factor, max_interval) + attempt += 1 + + raise RuntimeError(f"Error: Maximum polling attempts ({max_attempts}) reached. File is still not ACTIVE.") + +class SanitizedArgumentParser(argparse.ArgumentParser): + """Custom ArgumentParser that ensures any error output is sanitized to prevent secret leaks.""" + def error(self, message): + sanitized_msg = sanitize_error(message) + self.print_usage(sys.stderr) + self.exit(2, f"{self.prog}: error: {sanitized_msg}\n") + +def main(): + parser = SanitizedArgumentParser(description="Upload files to Gemini Files API using google-genai SDK.") + parser.add_argument("file", help="Path to the file to upload") + parser.add_argument("--name", help="Custom display name for the file") + parser.add_argument("--no-wait", action="store_true", help="Don't wait for ACTIVE status") + + args = parser.parse_args() + + api_key = get_api_key() + if not api_key: + print("Error: GEMINI_API_KEY environment variable is not set.", file=sys.stderr) + sys.exit(1) + + if not os.path.exists(args.file): + print(f"Error: File '{args.file}' not found.", file=sys.stderr) + sys.exit(1) + + try: + file_meta = upload_file(args.file, args.name) + file_name = file_meta.get("name") + + print(f"File metadata created:") + print(f" Name: {file_name}") + print(f" URI: {strip_query_params(file_meta.get('uri'))}") + print(f" Type: {file_meta.get('mimeType')}") + + if not args.no_wait: + file_meta = wait_for_active(file_name) + + print("\nFile upload successfully completed! JSON Output:") + print(json.dumps(file_meta, indent=2)) + sys.exit(0) + except Exception as e: + print(f"Error: {sanitize_error(e)}", file=sys.stderr) + sys.exit(1) + +if __name__ == "__main__": + main() diff --git a/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/generate_video.py b/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/generate_video.py new file mode 100755 index 00000000..80254562 --- /dev/null +++ b/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/generate_video.py @@ -0,0 +1,734 @@ +#!/usr/bin/env python3 +""" +Generates, extends, and edits videos using the Gemini Omni 1.1 Flash model via the google-genai Interactions API. +Can automatically upload local media references using the Files API. +Supports first frame and first+last frame transitions, video extensions (up to 40s), +video references (, etc.), image references, and parallel batch execution. +Uses the official google-genai SDK. +""" + +import argparse +from concurrent.futures import ThreadPoolExecutor, as_completed +import json +import os +import re +import subprocess +import sys +import time +import urllib.parse +import uuid +from google import genai +from google.genai import types, errors + +# Ensure stdout is unbuffered/line-buffered for real-time progress logs +if hasattr(sys.stdout, "reconfigure"): + sys.stdout.reconfigure(line_buffering=True) + +# Load local upload helper logic inline to prevent dependency issues +sys.path.append(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) +from upload_file import upload_file, wait_for_active + +DEFAULT_MODEL = "gemini-omni-1.1-flash" + +def get_api_key(): + """Retrieves API key strictly from GEMINI_API_KEY environment variable.""" + return os.environ.get("GEMINI_API_KEY") + +def strip_query_params(url): + """Strips query parameters from URL for clean logging and security.""" + if not url: + return "" + parsed = urllib.parse.urlparse(url) + return urllib.parse.urlunparse(parsed._replace(query="")) + +def sanitize_error(err): + """Sanitizes error messages by redacting API keys, tokens, query parameters, internal provider URLs, and raw response bodies.""" + if not err: + return "" + + # If it is an SDK APIError, prefer clean code + status + message over raw dictionary dump + if isinstance(err, errors.APIError): + parts = [] + if getattr(err, "code", None): + parts.append(str(err.code)) + if getattr(err, "status", None): + parts.append(f"({err.status})") + if getattr(err, "message", None): + parts.append(f": {err.message}") + elif getattr(err, "details", None): + parts.append(f": {err.details}") + err_str = " ".join(parts) if parts else str(err) + else: + err_str = str(err) + + # 1. Directly redact the active API key value if present + key = get_api_key() + if key and len(key) >= 8: + err_str = err_str.replace(key, "[REDACTED_KEY]") + + # 2. Redact known Google API key & OAuth token patterns: + # Classic Google API key: AIza... (30-45 chars) + err_str = re.sub(r'AIza[0-9A-Za-z_\-]{30,}', '[REDACTED_KEY]', err_str) + # Newer Google Cloud / Gemini API key: AQ.... + err_str = re.sub(r'AQ\.[0-9A-Za-z_\-]{20,}', '[REDACTED_KEY]', err_str) + # Google OAuth access token: ya29.... + err_str = re.sub(r'ya29\.[0-9A-Za-z_\-]+', '[REDACTED_TOKEN]', err_str) + # Bearer tokens + err_str = re.sub(r'Bearer\s+[A-Za-z0-9_\-\.]+', 'Bearer [REDACTED_TOKEN]', err_str, flags=re.IGNORECASE) + + # 3. Strip sensitive query parameters from any URLs (key, token, secret, signature) + err_str = re.sub(r'([?&](?:key|api_key|apiKey|access_token|auth|signature)=)[^&\s"\'<>()]+', r'\1[REDACTED]', err_str) + + # 4. Strip internal Google / provider infrastructure URLs & hosts + err_str = re.sub(r'https?://[a-zA-Z0-9.\-_]*\.corp\.goog[^\s"\'<>)]*', '[INTERNAL_HOST]', err_str) + err_str = re.sub(r'https?://[a-zA-Z0-9.\-_]*\.sandbox\.googleapis\.com[^\s"\'<>)]*', '[SANDBOX_ENDPOINT]', err_str) + err_str = re.sub(r'https?://(?:generativelanguage|aiplatform)\.googleapis\.com/v[0-9a-z_]+/', '[API_ENDPOINT]/', err_str) + + # 5. Redact raw response bodies (once a body= marker is found, redact the remainder) + err_str = re.sub(r'\bbody\s*=.*\Z', 'body=[REDACTED_BODY]', err_str, flags=re.IGNORECASE | re.DOTALL) + + # 6. Strip HTML error pages / raw HTML response bodies + if "(.*?)', err_str, re.IGNORECASE) + if title_match: + err_str = f"HTTP Error Response: {title_match.group(1).strip()}" + else: + err_str = re.sub(r'<[^>]+>', ' ', err_str) + + # 7. Collapse excessive whitespace and cap oversized error dumps + err_str = re.sub(r'\s+', ' ', err_str).strip() + if len(err_str) > 500: + err_str = err_str[:497] + "..." + + return err_str + +def is_file_uri(uri): + """Returns True if the string is a standard Gemini File URI.""" + if not uri: + return False + return "files/" in uri and ("generativelanguage.googleapis.com" in uri or uri.startswith("files/")) + +def normalize_file_uri(uri): + """Normalizes any File API URI/reference to the standard format with query parameters stripped.""" + if not uri: + return None + match = re.search(r'files/([a-zA-Z0-9]+)', uri) + if match: + file_id = match.group(1) + return f"https://generativelanguage.googleapis.com/files/{file_id}" + return strip_query_params(uri) + +def slugify(text): + """Converts a text prompt into a safe, descriptive filename slug.""" + text = text.lower() + text = re.sub(r'[^a-z0-9]+', '_', text) + return text.strip('_')[:50] + +def parse_and_validate_duration(value): + """Parses and formats a duration integer between 3 and 10 with optional 's' suffix.""" + if value is None: + return None + if isinstance(value, (int, float)): + val = float(value) + else: + clean_value = str(value).strip().lower() + if clean_value in ('none', ''): + return None + if clean_value.endswith('s'): + clean_value = clean_value[:-1] + try: + val = float(clean_value) + except ValueError: + raise ValueError(f"Invalid duration value: '{value}'. Must be an integer (e.g., 5, 10).") + + if not val.is_integer(): + raise ValueError(f"Duration must be an integer, not a float (e.g., got {value}).") + + val_int = int(val) + if val_int < 3 or val_int > 10: + raise ValueError(f"Duration must be between 3 (inclusive) and 10 (inclusive) seconds. Got {val_int}.") + + return f"{val_int}s" + +def argparse_duration_type(value): + """argparse type converter for validating duration.""" + if value is None or str(value).strip().lower() in ('none', ''): + return None + try: + return parse_and_validate_duration(value) + except ValueError as e: + raise argparse.ArgumentTypeError(str(e)) + +def parse_and_validate_resolution(value): + """Validates and normalizes video output resolution to '360p', '720p', '1080p', or '4k'.""" + if value is None: + return None + val = str(value).strip().lower() + if val in ('none', ''): + return None + valid_resolutions = {"360p", "720p", "1080p", "4k"} + if val in valid_resolutions: + return val + raise ValueError( + f"Invalid resolution '{value}'. Supported resolutions for Gemini Omni Flash are: 360p, 720p, 1080p, 4k." + ) + +def argparse_resolution_type(value): + """argparse type converter for validating video resolution.""" + if value is None or str(value).strip().lower() in ('none', ''): + return None + try: + return parse_and_validate_resolution(value) + except ValueError as e: + raise argparse.ArgumentTypeError(str(e)) + +def resolve_or_upload_asset(asset_path, mime_type, strip_audio=False): + """ + If asset_path is a File API URI, returns it directly (normalized). + If it is a local file path, uploads it and returns its File API URI (normalized). + """ + if not asset_path: + return None, None + + if not get_api_key(): + raise RuntimeError("Error: GEMINI_API_KEY environment variable is not set.") + + if is_file_uri(asset_path): + if strip_audio: + raise ValueError( + "Error: --strip-audio cannot be applied to an existing remote File API URI. " + "Please provide a local video file so audio can be stripped before upload." + ) + normalized = normalize_file_uri(asset_path) + print(f"Using existing File URI: {strip_query_params(normalized)}") + return normalized, mime_type + + if os.path.exists(asset_path): + upload_path = asset_path + temp_stripped_path = None + + try: + if strip_audio: + print(f"Detected local asset path '{asset_path}'. Stripping audio before upload...") + + # Check if ffmpeg is available + try: + subprocess.run(["ffmpeg", "-version"], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, check=True) + except (subprocess.SubprocessError, FileNotFoundError): + raise RuntimeError( + "Error: ffmpeg is not installed or not found in system PATH. " + "ffmpeg is required to strip audio from local videos when --strip-audio is specified." + ) + + os.makedirs("media", exist_ok=True) + base_name = os.path.basename(asset_path) + name, ext = os.path.splitext(base_name) + temp_stripped_path = os.path.join("media", f"temp_stripped_{name}_{uuid.uuid4().hex}{ext}") + + # Fast stream-copy audio stripping + cmd = ["ffmpeg", "-y", "-i", asset_path, "-c:v", "copy", "-an", temp_stripped_path] + proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True) + if proc.returncode != 0: + err_msg = sanitize_error(proc.stderr.strip()) if proc.stderr else "Unknown ffmpeg error" + raise RuntimeError(f"ffmpeg failed to strip audio from '{asset_path}': {err_msg}") + + print(f"Successfully stripped audio. Temporary video file created at: {temp_stripped_path}") + upload_path = temp_stripped_path + + print(f"Uploading asset '{upload_path}'...") + file_meta = upload_file(upload_path) + file_name = file_meta.get("name") + # Wait for file to become active + file_meta = wait_for_active(file_name) + normalized = normalize_file_uri(file_meta.get("uri")) + + # Handle both mimeType and mime_type key formats returned from upload_file + returned_mime = file_meta.get("mimeType") or file_meta.get("mime_type") + return normalized, returned_mime + finally: + # Clean up temporary stripped file if we created one, guaranteed in finally block + if temp_stripped_path and os.path.exists(temp_stripped_path): + try: + os.remove(temp_stripped_path) + print(f"Cleaned up temporary video file: {temp_stripped_path}") + except Exception as e: + print(f"Warning: Failed to remove temporary file {temp_stripped_path}: {sanitize_error(e)}", file=sys.stderr) + else: + raise FileNotFoundError(f"Asset path '{asset_path}' is neither a valid File API URI nor a local file path.") + +def download_video_file(file_or_uri, output_path, client=None): + """Downloads generated video file using the official google-genai SDK files.download method.""" + if not file_or_uri: + raise ValueError("Download error: file reference or URI is required.") + + if not get_api_key(): + raise RuntimeError("Error: GEMINI_API_KEY environment variable is not set.") + + if client is None: + client = genai.Client() + + display_uri = getattr(file_or_uri, "uri", str(file_or_uri)) + print(f"Downloading video from {strip_query_params(display_uri)} to {output_path} via SDK...") + try: + download_target = getattr(file_or_uri, "uri", file_or_uri) + video_bytes = client.files.download(file=download_target) + parent_dir = os.path.dirname(output_path) + if parent_dir: + os.makedirs(parent_dir, exist_ok=True) + + with open(output_path, "wb") as f: + f.write(video_bytes) + print(f"Video successfully saved to: {output_path}") + except Exception as e: + raise RuntimeError(f"Error downloading video file via SDK: {sanitize_error(e)}") + +def generate_video( + prompt, + model=DEFAULT_MODEL, + aspect_ratio="16:9", + duration=None, + resolution=None, + first_frame=None, + last_frame=None, + image_path=None, + image_reference=None, + video_path=None, + extend_video=None, + video_reference=None, + task=None, + output_path="output.mp4", + strip_audio=False, + previous_interaction_id=None, + timeout=600 +): + """Creates an interaction with the video model and downloads the resulting video using the official google-genai SDK.""" + if not get_api_key(): + raise RuntimeError("Error: GEMINI_API_KEY environment variable is not set.") + + # Validation: last_frame requires first_frame + if last_frame and not first_frame: + raise ValueError("--last-frame must be used with --first-frame (the model requires a starting frame when specifying a final frame).") + + duration = parse_and_validate_duration(duration) + input_parts = [] + + # Collect reference images (supporting image_path and image_reference aliases) + ref_images = [] + if image_path: + if isinstance(image_path, list): + ref_images.extend(image_path) + else: + ref_images.append(image_path) + if image_reference: + if isinstance(image_reference, list): + ref_images.extend(image_reference) + else: + ref_images.append(image_reference) + + # Collect reference videos + ref_videos = [] + if video_reference: + if isinstance(video_reference, list): + ref_videos.extend(video_reference) + else: + ref_videos.append(video_reference) + + if len(ref_videos) > 3: + print(f"Note: {len(ref_videos)} reference videos provided. The ideal number of reference videos is up to 3 (though more are supported).") + + # 1. Resolve and add first frame image + f1_uri = None + f1_mime = None + if first_frame: + f1_uri, f1_mime = resolve_or_upload_asset(first_frame, "image/png") + input_parts.append({ + "type": "image", + "uri": f1_uri, + "mime_type": f1_mime + }) + + # 2. Resolve and add last frame image (must follow first frame) + if last_frame: + if last_frame == first_frame and f1_uri: + # Re-use already uploaded asset for looping video + input_parts.append({ + "type": "image", + "uri": f1_uri, + "mime_type": f1_mime + }) + else: + f2_uri, f2_mime = resolve_or_upload_asset(last_frame, "image/png") + input_parts.append({ + "type": "image", + "uri": f2_uri, + "mime_type": f2_mime + }) + + # 3. Resolve and add general reference image inputs (, ...) + for img in ref_images: + img_uri, img_mime = resolve_or_upload_asset(img, "image/png") + input_parts.append({ + "type": "image", + "uri": img_uri, + "mime_type": img_mime + }) + + # 4. Resolve and add source video for extension or editing + if extend_video: + ext_uri, ext_mime = resolve_or_upload_asset(extend_video, "video/mp4", strip_audio=strip_audio) + input_parts.append({ + "type": "video", + "uri": ext_uri, + "mime_type": ext_mime + }) + + if video_path: + if isinstance(video_path, list): + for path in video_path: + vid_uri, vid_mime = resolve_or_upload_asset(path, "video/mp4", strip_audio=strip_audio) + input_parts.append({ + "type": "video", + "uri": vid_uri, + "mime_type": vid_mime + }) + else: + vid_uri, vid_mime = resolve_or_upload_asset(video_path, "video/mp4", strip_audio=strip_audio) + input_parts.append({ + "type": "video", + "uri": vid_uri, + "mime_type": vid_mime + }) + + # 5. Resolve and add reference videos (, , ...) + for ref_vid in ref_videos: + rv_uri, rv_mime = resolve_or_upload_asset(ref_vid, "video/mp4", strip_audio=strip_audio) + input_parts.append({ + "type": "video", + "uri": rv_uri, + "mime_type": rv_mime + }) + + # 6. Format and add text prompt with appropriate role tags if not explicitly declared + prompt_text = prompt + + if "[# Sources" not in prompt and "[# References" not in prompt: + if extend_video: + sources_part = "[# Sources @Video1]" + ref_parts = [] + img_start_num = (2 if last_frame and last_frame != first_frame else 1) if first_frame else 0 + for idx, _ in enumerate(ref_images): + ref_parts.append(f"@Image{img_start_num + idx + 1}") + vid_start_num = 1 + if video_path: + if isinstance(video_path, list): + vid_start_num += len(video_path) + else: + vid_start_num += 1 + for idx, _ in enumerate(ref_videos): + ref_parts.append(f"@Video{vid_start_num + idx + 1}") + + if ref_parts: + prompt_text = f"{sources_part} [# References {' '.join(ref_parts)}] {prompt_text}" + else: + prompt_text = f"{sources_part} {prompt_text}" + elif first_frame and last_frame: + if "" not in prompt and "" not in prompt: + prompt_text = f" {prompt_text}" + elif first_frame: + if "" not in prompt: + prompt_text = f" {prompt_text}" + + input_parts.append({ + "type": "text", + "text": prompt_text + }) + + # Construct response_format video configuration + video_config = { + "type": "video", + "delivery": "uri" + } + + # Only omit aspect_ratio if task is explicitly 'extend' + if task != "extend": + video_config["aspect_ratio"] = aspect_ratio + + if duration: + video_config["duration"] = duration + + if resolution: + video_config["resolution"] = parse_and_validate_resolution(resolution) + + # Construct generation_config if task is explicitly specified and no previous_interaction_id is used + generation_config = None + if task and not previous_interaction_id: + generation_config = { + "video_config": { + "task": task + } + } + + res_display = video_config.get("resolution", "default (720p)") + print(f"\nSending generation request using official google-genai SDK and model '{model}'...") + print(f"Prompt: '{prompt_text}' | Aspect Ratio: {video_config.get('aspect_ratio', 'inherited')} | Resolution: {res_display} | Duration: {duration or 'default'}") + if res_display == "4k": + print("Note: 4K video generation selected. Processing high-resolution video may take longer to complete.") + if task: + print(f"Task Mode: {task}") + + # Initialize the client with explicit per-request timeout bounds in milliseconds + timeout_s = int(timeout) if timeout else 600 + if timeout_s <= 0: + raise ValueError(f"Timeout must be greater than 0 (got {timeout}).") + client = genai.Client(http_options=types.HttpOptions(timeout=timeout_s * 1000)) + create_kwargs = { + "model": model, + "input": input_parts, + "response_format": video_config, + } + if generation_config: + create_kwargs["generation_config"] = generation_config + if previous_interaction_id: + create_kwargs["previous_interaction_id"] = previous_interaction_id + + try: + interaction = client.interactions.create(**create_kwargs) + except errors.APIError as e: + raise RuntimeError(f"API Error generating video via SDK: {sanitize_error(e)}") + except Exception as e: + raise RuntimeError(f"Error generating video via SDK: {sanitize_error(e)}") + + print(f"Generation complete for '{prompt}'! Processing response...") + + interaction_id = interaction.id + if interaction_id: + print(f"Interaction ID: {interaction_id}") + + output_video = interaction.output_video + if not output_video or not output_video.uri: + err_msg = f"No video content found in response for '{prompt}'." + if video_path or extend_video or ref_videos: + err_msg += ( + "\nWARNING: IMPORTANT REGIONAL RESTRICTION: Uploading videos to use for video edits, extensions, or references is " + "not available in the EEA, Switzerland, United Kingdom, and some US states." + ) + raise RuntimeError(err_msg) + + video_uri = output_video.uri + print(f"Generated video URI for '{prompt}': {strip_query_params(video_uri)}") + + # Download the final video using the official google-genai SDK + download_video_file(output_video or video_uri, output_path, client=client) + +def run_job(job): + """Runs a single generation job inside a thread pool, catching exceptions.""" + prompt = job.get("prompt") + if not prompt: + print("Warning: Skipping job with empty prompt.", file=sys.stderr) + return {"job": job, "status": "SKIPPED", "error": "Empty prompt"} + + first_frame = job.get("first_frame") + last_frame = job.get("last_frame") + if last_frame and not first_frame: + err = "last_frame must be used with first_frame (the model requires a starting frame when specifying a final frame)." + print(f"[Parallel] Failed: '{prompt}' - Error: {err}", file=sys.stderr) + return {"job": job, "status": "FAILED", "error": err} + + aspect_ratio = job.get("aspect_ratio", "16:9") + duration = job.get("duration") + resolution = job.get("resolution") + image_path = job.get("image") or job.get("image_reference") or job.get("ref_image") + video_path = job.get("video") + extend_video = job.get("extend") or job.get("extend_video") + video_reference = job.get("video_reference") or job.get("ref_video") or job.get("video_ref") + task = job.get("task") + output_path = job.get("output") + model = job.get("model", DEFAULT_MODEL) + strip_audio = job.get("strip_audio", False) + previous_interaction_id = job.get("previous_interaction_id") + timeout = job.get("timeout", 600) + + if not output_path: + output_path = f"media/output_{slugify(prompt)}.mp4" + + print(f"[Parallel] Dispatching: '{prompt}' (Output: {output_path})") + + try: + generate_video( + prompt=prompt, + model=model, + aspect_ratio=aspect_ratio, + duration=duration, + resolution=resolution, + first_frame=first_frame, + last_frame=last_frame, + image_path=image_path, + video_path=video_path, + extend_video=extend_video, + video_reference=video_reference, + task=task, + output_path=output_path, + strip_audio=strip_audio, + previous_interaction_id=previous_interaction_id, + timeout=timeout + ) + return {"job": job, "status": "SUCCESS", "output_path": output_path} + except Exception as e: + cleaned_err = sanitize_error(e) + print(f"[Parallel] Failed: '{prompt}' - Error: {cleaned_err}", file=sys.stderr) + return {"job": job, "status": "FAILED", "error": cleaned_err} + +class SanitizedArgumentParser(argparse.ArgumentParser): + """Custom ArgumentParser that ensures any error output is sanitized to prevent secret leaks.""" + def error(self, message): + sanitized_msg = sanitize_error(message) + self.print_usage(sys.stderr) + self.exit(2, f"{self.prog}: error: {sanitized_msg}\n") + +def main(): + parser = SanitizedArgumentParser(description="Generate, extend, and edit videos using Gemini Omni 1.1 Flash model via google-genai SDK (supports parallel batch execution).") + parser.add_argument("prompt", nargs="?", help="Text prompt / instruction for a single video generation") + parser.add_argument("--first-frame", "-f", help="Local image path or File API URI for starting frame ()") + parser.add_argument("--last-frame", "-l", help="Local image path or File API URI for final transition frame (). Must be used with --first-frame.") + parser.add_argument("--image", "--image-reference", "--ref-image", "-i", action="append", dest="image", help="Optional local image path or File API URI for referencing / image-to-video (can be specified multiple times)") + parser.add_argument("--video", "-v", action="append", help="Optional local video path or File API URI for editing / source video (can be specified multiple times)") + parser.add_argument("--extend", "-e", help="Optional local video path or File API URI to extend (by up to 10s per turn, up to 40s total)") + parser.add_argument("--video-reference", "--ref-video", "-vr", action="append", dest="video_reference", help="Optional local video path or File API URI for reference video(s) (, , ...). Ideal duration is ~3s (up to 3 recommended). Can be specified multiple times.") + parser.add_argument("--task", choices=["text_to_video", "image_to_video", "reference_to_video", "edit", "extend"], default=None, help="Explicit video task mode. Note: omitting task allows combining extensions with reference images and videos.") + parser.add_argument("--aspect-ratio", default="16:9", choices=["16:9", "9:16"], help="Aspect ratio (default: 16:9, omitted automatically if explicit task='extend' is set)") + parser.add_argument("--duration", type=argparse_duration_type, default=None, help="Video duration as an integer between 3 and 10 seconds (e.g., 5, 10). Default: None (API/Model decides, typically 10s or matches source)") + parser.add_argument("--resolution", "-res", type=argparse_resolution_type, default=None, help="Video output resolution: 360p, 720p, 1080p, or 4k (default: 720p). Note: 4k requests take longer to generate.") + parser.add_argument("--model", default=DEFAULT_MODEL, help=f"Gemini Omni Flash video model ID (default: {DEFAULT_MODEL})") + parser.add_argument("--output", "-o", help="Local output file path for single generation (default: media/output.mp4)") + parser.add_argument("--strip-audio", "-a", action="store_true", help="Completely strip/disable audio stream from the input video(s) before uploading so Gemini Omni Flash can regenerate new audio from scratch") + parser.add_argument("--previous-interaction-id", help="Optional Interaction ID of a previous generation for turn-by-turn editing or extending") + parser.add_argument("--timeout", type=int, default=600, help="Per-request HTTP timeout in seconds (default: 600). Recommend 900+ for multi-turn 4K extensions up to 40s.") + + # Parallel batch configuration options + parser.add_argument("--batch", help="Path to a JSON file containing an array of generation jobs") + parser.add_argument("--prompts-file", help="Path to a text file containing one prompt per line to run in parallel") + parser.add_argument("--concurrency", type=int, default=3, help="Maximum number of concurrent executions (default: 3)") + + args = parser.parse_args() + + if args.last_frame and not args.first_frame: + parser.error("--last-frame must be used with --first-frame (the model requires a starting frame when specifying a final frame).") + + if not get_api_key(): + print("Error: GEMINI_API_KEY environment variable is not set.", file=sys.stderr) + sys.exit(1) + + # 1. Handle Batch JSON execution + if args.batch: + if not os.path.exists(args.batch): + print(f"Error: Batch JSON file '{args.batch}' not found.", file=sys.stderr) + sys.exit(1) + try: + with open(args.batch, "r", encoding="utf-8") as f: + jobs = json.load(f) + if not isinstance(jobs, list): + print("Error: Batch JSON file must contain a list/array of job objects.", file=sys.stderr) + sys.exit(1) + except Exception as e: + print(f"Error parsing Batch JSON: {sanitize_error(e)}", file=sys.stderr) + sys.exit(1) + + print(f"Loaded {len(jobs)} jobs from batch JSON. Running with concurrency={args.concurrency}...") + + # 2. Handle Prompts File execution + elif args.prompts_file: + if not os.path.exists(args.prompts_file): + print(f"Error: Prompts file '{args.prompts_file}' not found.", file=sys.stderr) + sys.exit(1) + + jobs = [] + with open(args.prompts_file, "r", encoding="utf-8") as f: + for line in f: + line = line.strip() + if line and not line.startswith("#"): + jobs.append({ + "prompt": line, + "aspect_ratio": args.aspect_ratio, + "duration": args.duration, + "resolution": args.resolution, + "first_frame": args.first_frame, + "last_frame": args.last_frame, + "image": args.image, + "video": args.video, + "extend": args.extend, + "video_reference": args.video_reference, + "task": args.task, + "model": args.model, + "strip_audio": args.strip_audio, + "previous_interaction_id": args.previous_interaction_id, + "timeout": args.timeout + }) + print(f"Loaded {len(jobs)} prompts from text file. Running with concurrency={args.concurrency}...") + + # 3. Handle standard single prompt execution + else: + if not args.prompt: + parser.print_help() + sys.exit(1) + + output_path = args.output if args.output else "media/output.mp4" + try: + generate_video( + prompt=args.prompt, + model=args.model, + aspect_ratio=args.aspect_ratio, + duration=args.duration, + resolution=args.resolution, + first_frame=args.first_frame, + last_frame=args.last_frame, + image_path=args.image, + video_path=args.video, + extend_video=args.extend, + video_reference=args.video_reference, + task=args.task, + output_path=output_path, + strip_audio=args.strip_audio, + previous_interaction_id=args.previous_interaction_id, + timeout=args.timeout + ) + sys.exit(0) + except Exception as e: + print(f"Error: Generation failed: {sanitize_error(e)}", file=sys.stderr) + sys.exit(1) + + # Parallel Execution Loop + if not jobs: + print("Warning: No valid jobs found to execute.") + sys.exit(0) + + results = [] + with ThreadPoolExecutor(max_workers=args.concurrency) as executor: + futures = {executor.submit(run_job, job): job for job in jobs} + for future in as_completed(futures): + results.append(future.result()) + + # Print Batch Results Summary + print("\n" + "="*50) + print("BATCH PARALLEL EXECUTION SUMMARY") + print("="*50) + success_count = sum(1 for r in results if r["status"] == "SUCCESS") + failed_count = sum(1 for r in results if r["status"] == "FAILED") + skipped_count = sum(1 for r in results if r["status"] == "SKIPPED") + + print(f"Total: {len(results)} | Success: {success_count} | Failed: {failed_count} | Skipped: {skipped_count}\n") + for r in results: + status_str = r["status"] + prompt = r["job"].get("prompt") + if r["status"] == "SUCCESS": + print(f" [{status_str}] '{prompt}' -> {r['output_path']}") + else: + print(f" [{status_str}] '{prompt}' -> Error: {r.get('error')}") + print("="*50) + + if failed_count > 0: + sys.exit(1) + sys.exit(0) + +if __name__ == "__main__": + main() + diff --git a/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/inspect_video.py b/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/inspect_video.py new file mode 100755 index 00000000..032c899b --- /dev/null +++ b/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/inspect_video.py @@ -0,0 +1,197 @@ +#!/usr/bin/env python3 +import argparse +import json +import os +import subprocess +import sys + +def format_size(size_bytes): + """Formats file size in bytes to a human-readable string.""" + try: + size_bytes = int(size_bytes) + except (ValueError, TypeError): + return "Unknown size" + + for unit in ['B', 'KB', 'MB', 'GB']: + if size_bytes < 1024.0: + return f"{size_bytes:.2f} {unit}" + size_bytes /= 1024.0 + return f"{size_bytes:.2f} TB" + +def parse_fps(fps_str): + """Parses fractional frame rates like '30/1' or '24000/1001' into floats.""" + if not fps_str: + return "Unknown" + if "/" in fps_str: + try: + num, den = map(float, fps_str.split("/")) + if den != 0: + val = num / den + if val.is_integer(): + return f"{int(val)} fps" + return f"{val:.2f} fps" + except (ValueError, ZeroDivisionError): + pass + try: + val = float(fps_str) + if val.is_integer(): + return f"{int(val)} fps" + return f"{val:.2f} fps" + except ValueError: + return fps_str + +def inspect_video(file_path, raw=False): + """Runs ffprobe on the video file and returns parsed metadata dictionary.""" + if not os.path.exists(file_path): + raise FileNotFoundError(f"File not found: {file_path}") + + # Check if ffprobe is available + try: + subprocess.run(["ffprobe", "-version"], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, check=True) + except (subprocess.SubprocessError, FileNotFoundError): + raise RuntimeError("ffprobe is not installed or not found in system PATH.") + + cmd = [ + "ffprobe", + "-v", "error", + "-show_format", + "-show_streams", + "-of", "json", + file_path + ] + + result = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True, check=True) + data = json.loads(result.stdout) + + if raw: + return data + + # Extract format level details + fmt = data.get("format", {}) + duration = fmt.get("duration") + size_bytes = fmt.get("size") + bitrate = fmt.get("bit_rate") + + # Format files size + size_str = format_size(size_bytes) if size_bytes else "Unknown" + + # Parse duration + try: + duration_val = float(duration) if duration else 0.0 + duration_str = f"{duration_val:.2f}s" + except ValueError: + duration_str = "Unknown" + duration_val = None + + # Parse bitrate + try: + bitrate_kbps = f"{int(float(bitrate) / 1000)} kbps" if bitrate else "Unknown" + except ValueError: + bitrate_kbps = "Unknown" + + video_streams = [s for s in data.get("streams", []) if s.get("codec_type") == "video"] + audio_streams = [s for s in data.get("streams", []) if s.get("codec_type") == "audio"] + + has_video = len(video_streams) > 0 + has_audio = len(audio_streams) > 0 + + video_info = {} + if has_video: + v = video_streams[0] + width = v.get("width") + height = v.get("height") + codec = v.get("codec_name", "Unknown").upper() + r_fps = parse_fps(v.get("r_frame_rate")) + avg_fps = parse_fps(v.get("avg_frame_rate")) + + # Prefer r_frame_rate but fallback to avg + fps = r_fps if r_fps != "0 fps" and r_fps != "Unknown" else avg_fps + + video_info = { + "resolution": f"{width}x{height}" if width and height else "Unknown", + "width": width, + "height": height, + "fps": fps, + "codec": codec, + "duration": v.get("duration") + } + + audio_info = {} + if has_audio: + a = audio_streams[0] + codec = a.get("codec_name", "Unknown").upper() + channels = a.get("channels", "Unknown") + sample_rate = a.get("sample_rate") + sample_rate_khz = f"{float(sample_rate)/1000:.1f} kHz" if sample_rate else "Unknown" + + audio_info = { + "codec": codec, + "channels": channels, + "sample_rate": sample_rate_khz + } + + return { + "file_name": os.path.basename(file_path), + "file_size": size_str, + "size_bytes": size_bytes, + "duration": duration_str, + "duration_seconds": duration_val, + "bitrate": bitrate_kbps, + "has_video": has_video, + "video": video_info, + "has_audio": has_audio, + "audio": audio_info + } + +def print_terminal_report(info): + """Prints an aligned terminal report.""" + print(f"\nVideo Inspection Report: {info['file_name']}") + print("=" * 50) + print(f"File Size : {info['file_size']}") + print(f"Duration : {info['duration']}") + print(f"Bitrate : {info['bitrate']}") + + print("\nVideo Stream Details:") + if info["has_video"]: + v = info["video"] + print(f" * Resolution : {v['resolution']}") + print(f" * Frame Rate : {v['fps']}") + print(f" * Codec : {v['codec']}") + else: + print(" * No Video Stream Found.") + + print("\nAudio Stream Details:") + if info["has_audio"]: + a = info["audio"] + print(" * Status : Audio Present") + print(f" * Codec : {a['codec']}") + print(f" * Channels : {a['channels']}") + print(f" * Sample Rate: {a['sample_rate']}") + else: + print(" * Status : No Audio Stream Present") + print() + +def main(): + parser = argparse.ArgumentParser(description="Inspect video details (duration, frame rate, resolution, audio presence) using ffprobe.") + parser.add_argument("file", help="Path to the video file to inspect") + parser.add_argument("--json", action="store_true", help="Output parsed summary in JSON format") + parser.add_argument("--raw", action="store_true", help="Output raw unmodified ffprobe JSON data") + + args = parser.parse_args() + + try: + if args.raw: + info = inspect_video(args.file, raw=True) + print(json.dumps(info, indent=2)) + else: + info = inspect_video(args.file, raw=False) + if args.json: + print(json.dumps(info, indent=2)) + else: + print_terminal_report(info) + except Exception as e: + print(f"Error inspecting video: {e}", file=sys.stderr) + sys.exit(1) + +if __name__ == "__main__": + main() diff --git a/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/prep_video.py b/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/prep_video.py new file mode 100755 index 00000000..dbe83922 --- /dev/null +++ b/.claude/skills/gemini/references/official/gemini-omni-flash-api/scripts/video/prep_video.py @@ -0,0 +1,255 @@ +#!/usr/bin/env python3 +import argparse +import os +import subprocess +import sys +from inspect_video import inspect_video, format_size + +def parse_timecode(time_str, total_duration=None): + """Parses a time string (seconds, MM:SS, HH:MM:SS, or 'last') into float seconds.""" + if not time_str: + return 0.0 + time_str = time_str.strip().lower() + + if time_str == "last": + if total_duration is None: + raise ValueError("Total duration is required to calculate 'last' starting point.") + target_dur = 10.0 + if total_duration <= target_dur: + return 0.0 + return total_duration - target_dur + + if ":" in time_str: + parts = time_str.split(":") + if len(parts) == 2: # MM:SS + m, s = map(float, parts) + return m * 60.0 + s + elif len(parts) == 3: # HH:MM:SS + h, m, s = map(float, parts) + return h * 3600.0 + m * 60.0 + s + else: + raise ValueError(f"Invalid timecode format: '{time_str}'. Use HH:MM:SS or MM:SS.") + + try: + return float(time_str) + except ValueError: + raise ValueError(f"Invalid timecode: '{time_str}'. Must be float seconds, HH:MM:SS, or 'last'.") + +def prep_video(input_path, output_path, start_time_str=None, duration=10, fps=None, resolution=None, strip_audio=False): + """Preps a video file by trimming, optionally re-encoding to target fps and resolution.""" + if not os.path.exists(input_path): + raise FileNotFoundError(f"Input file not found: {input_path}") + + # Check if ffmpeg is available + try: + subprocess.run(["ffmpeg", "-version"], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, check=True) + except (subprocess.SubprocessError, FileNotFoundError): + raise RuntimeError("ffmpeg is not installed or not found in system PATH.") + + # Inspect input video first + print(f"Analyzing source video: {os.path.basename(input_path)}...") + source_info = inspect_video(input_path) + total_duration = source_info.get("duration_seconds", 0.0) + + # Resolve start time + if start_time_str is None: + if total_duration and total_duration > 10.0 and sys.stdin.isatty(): + print(f"\nThe input video is longer than 10s ({total_duration:.2f}s).") + print("Please choose a 10s segment to trim:") + print(" 1) First 10 seconds [default]") + print(" 2) Last 10 seconds") + print(" 3) Custom starting timecode (e.g., MM:SS, HH:MM:SS, or seconds)") + try: + choice = input("Your choice [1/2/3, default 1]: ").strip() + if choice == "2": + start_time_str = "last" + elif choice == "3": + custom_start = input("Enter starting timecode (e.g., 00:03 or 15): ").strip() + start_time_str = custom_start if custom_start else "0" + else: + start_time_str = "0" + except (KeyboardInterrupt, EOFError): + print("\nNo input received. Defaulting to first 10 seconds.") + start_time_str = "0" + else: + start_time_str = "0" + + try: + start_seconds = parse_timecode(start_time_str, total_duration) + except Exception as e: + raise ValueError(f"Timecode parsing failed: {e}") + + + if start_seconds < 0 or (total_duration and start_seconds >= total_duration): + raise ValueError(f"Start time {start_seconds}s is out of bounds for video of length {total_duration}s.") + + # Construct output path if not specified + if not output_path: + os.makedirs("media", exist_ok=True) + base_name = os.path.basename(input_path) + name, ext = os.path.splitext(base_name) + output_path = os.path.join("media", f"prepped_{name}.mp4") + + # Check if the source video file is large (>25MB) + size_bytes_str = source_info.get("size_bytes") + is_large = False + try: + if size_bytes_str and int(size_bytes_str) > 25 * 1024 * 1024: + is_large = True + except (ValueError, TypeError): + pass + + # Target resolution parsing + scale_filter = None + orig_width = None + orig_height = None + if "video" in source_info: + try: + orig_width = int(source_info["video"].get("width")) + orig_height = int(source_info["video"].get("height")) + except (ValueError, TypeError): + pass + + if resolution: + try: + target_w, target_h = map(int, resolution.lower().split("x")) + if orig_width and orig_height: + # Scale to fit target_w and target_h while preserving aspect ratio + scale_factor = min(target_w / orig_width, target_h / orig_height) + width = int(orig_width * scale_factor) + height = int(orig_height * scale_factor) + else: + width, height = target_w, target_h + # Ensure divisible by 2 for standard decoders/encoders + width = (width // 2) * 2 + height = (height // 2) * 2 + scale_filter = f"scale={width}:{height}" + resolution = f"{width}x{height}" + except ValueError: + raise ValueError(f"Invalid resolution: '{resolution}'. Format must be WIDTHxHEIGHT (e.g. 1280x720).") + elif is_large: + if orig_width and orig_height: + # Scale down large videos proportionally (max 1280x720 for landscape, 720x1280 for portrait) + if orig_width >= orig_height: + max_w, max_h = 1280, 720 + else: + max_w, max_h = 720, 1280 + scale_factor = min(max_w / orig_width, max_h / orig_height) + if scale_factor < 1.0: + width = int(orig_width * scale_factor) + height = int(orig_height * scale_factor) + else: + width, height = orig_width, orig_height + else: + width, height = 1280, 720 + + # Ensure divisible by 2 + width = (width // 2) * 2 + height = (height // 2) * 2 + resolution = f"{width}x{height}" + print(f"\nRecommendation: Source video is very large ({source_info.get('file_size')}).") + print(" Automatically scaling to optimize upload times for Gemini Omni Flash.") + scale_filter = f"scale={width}:{height}" + + fps_spec = f"{fps} fps" if fps else "Original frame rate" + print(f"\nPreparing Video Processing:") + print(f" * Source Duration: {total_duration:.2f}s") + print(f" * Trim Range : Start at {start_seconds:.2f}s | Length {duration:.2f}s") + if resolution: + print(f" * Encoding Specs : {width}x{height} @ {fps_spec}") + else: + print(f" * Encoding Specs : Original Resolution @ {fps_spec}") + print(f" * Target Path : {output_path}") + print("=" * 50) + + # ffmpeg command construction + cmd = [ + "ffmpeg", + "-y", # Overwrite output + "-ss", str(start_seconds), # Seek start + "-i", input_path, # Input file + "-t", str(duration), # Duration to copy + ] + if scale_filter: + cmd.extend(["-vf", scale_filter]) + + cmd.extend([ + "-c:v", "libx264", # Standard H264 video codec + "-pix_fmt", "yuv420p", # Standard pixel format for web/Gemini compatibility + ]) + + if fps: + cmd.extend(["-r", str(fps)]) # Output frame rate if requested + + if strip_audio or not source_info.get("has_audio", False): + if not source_info.get("has_audio", False) and not strip_audio: + print("No audio stream detected in source video. Disabling audio output.") + else: + print("Stripping audio stream from video as requested.") + cmd.append("-an") # Disable audio streams completely + else: + cmd.extend([ + "-c:a", "aac", # Convert audio to standard AAC + "-b:a", "128k", # Standard audio bitrate + "-ac", "2", # Convert to stereo + ]) + + cmd.append(output_path) + + print("Running ffmpeg encoding...") + process = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True) + + if process.returncode != 0: + print("Error: ffmpeg failed. Stderr output follows:", file=sys.stderr) + print(process.stderr, file=sys.stderr) + raise RuntimeError("ffmpeg execution failed.") + + print("Video preparation completed successfully!") + print("=" * 50) + + # Call inspection tool on output to print clean specs + output_info = inspect_video(output_path) + return output_info + +def main(): + parser = argparse.ArgumentParser(description="Prep videos for editing (trimming, re-encoding to target fps and resolution).") + parser.add_argument("file", help="Path to the source video file to prep") + parser.add_argument("--start", "-s", default=None, help="Start timecode (seconds, MM:SS, HH:MM:SS, or 'last' for last 10s). Default: 0 (or prompted if > 10s)") + parser.add_argument("--duration", "-d", type=int, default=10, help="Duration of trimmed segment in seconds. Default: 10") + parser.add_argument("--fps", "-r", type=int, default=None, help="Target frame rate. Default: None (keep original frame rate)") + parser.add_argument("--resolution", "-g", default=None, help="Target resolution (e.g., 1280x720). Default: None (keep original resolution)") + parser.add_argument("--output", "-o", help="Custom output path. Defaults to media/prepped_.mp4") + parser.add_argument("--strip-audio", "-a", action="store_true", help="Completely strip/disable audio stream so the model can generate new audio") + + args = parser.parse_args() + + try: + info = prep_video( + input_path=args.file, + output_path=args.output, + start_time_str=args.start, + duration=args.duration, + fps=args.fps, + resolution=args.resolution, + strip_audio=args.strip_audio + ) + + # Display final output report + print(f"\nPrepped Video Specifications: {info['file_name']}") + print("=" * 50) + print(f"File Size : {info['file_size']}") + print(f"Duration : {info['duration']}") + print(f"Bitrate : {info['bitrate']}") + print(f"Resolution : {info['video']['resolution']}") + print(f"Frame Rate : {info['video']['fps']}") + print(f"Video Codec : {info['video']['codec']}") + if info["has_audio"]: + print(f"Audio Spec : {info['audio']['codec']} | {info['audio']['channels']} ch | {info['audio']['sample_rate']}") + print() + + except Exception as e: + print(f"Error prepping video: {e}", file=sys.stderr) + sys.exit(1) + +if __name__ == "__main__": + main() diff --git a/.claude/skills/gemini/upstream.yaml b/.claude/skills/gemini/upstream.yaml new file mode 100644 index 00000000..7af19830 --- /dev/null +++ b/.claude/skills/gemini/upstream.yaml @@ -0,0 +1,21 @@ +vendor: Google Gemini +repository: https://github.com/google-gemini/gemini-skills +source_path: skills +reviewed_commit: 80dd31dda25bbe1410207df0adb3e0d591c2c634 +reviewed_at: 2026-09-15 +policy: + active_skill: gemini + upstream_files_are_read_only: true + auto_update: false + update_flow: detect-diff-review-evals-pin +sources: + api: + local: references/official/gemini-api-dev + live_api: + local: references/official/gemini-live-api-dev + activation: reference-only-until-adopted + omni_flash: + local: references/official/gemini-omni-flash-api + activation: reference-only-until-adopted +combined_tree_sha256: 9cb960898d2a22bde524ac27ea08f2f5188ae5f54814ee99d5e7e20596d8d3e9 +hash_algorithm: sha256(sorted base-name/relative-path + NUL + file-bytes + NUL) diff --git a/.claude/skills/maps/SKILL.md b/.claude/skills/maps/SKILL.md new file mode 100644 index 00000000..4986f330 --- /dev/null +++ b/.claude/skills/maps/SKILL.md @@ -0,0 +1,284 @@ +--- +name: maps +description: >- + Use when MDE work changes or diagnoses Google Maps, Places, map state, markers, routes, location search, Maps grounding, or map-related keys. +title: maps — Google Maps Platform (comprehensive) +impact: HIGH +impactDescription: Places enrichment, Maps grounding, ChatMap, batch APIs, security, AI code assist +tags: google-maps, places-api, maps-links, gemini-grounding, mdeai, mastra, security, cli, mcp +paths: + - "src/**/*Map*" + - "src/**/*map*" + - "src/mastra/tools/*place*" + - "supabase/functions/*maps*/**" + - "supabase/functions/*places*/**" +--- + +# maps — Google Maps Platform + +## When NOT to use + +- Generic GIS / spatial math with no Google Maps Platform APIs +- **Mapbox-only** or **Leaflet/OpenStreetMap-only** stacks (no GMP) +- Unrelated mapping tutorials or homework off the mdeai repo +- **Non-mdeAI** products—still read-only here; prefer not to expand scope in this skill + +## Load order (keep context small) + +1. This **`SKILL.md`** — Quick routing table + Consolidated sibling note. +2. **GMP doc questions** — read the pinned official Google Maps skill first, then one MDE reference relevant to the task. +3. **One** MDE `references/*.md` file for implementation; do not bulk-load unrelated references. +4. **`scripts/gmaps.py` + `references/gmaps-cli-behavior.md`** only when running or editing batch CLI work. + +Historical Cursor/MCP files are not active local dependencies. Verify current Maps tooling before relying on an MCP integration. + +--- + +## Official Google Maps upstream + MDE overlay + +**Decision:** Google’s official `google-maps-platform` skill is the pinned upstream/reference layer. This `maps` skill remains the only active MDE Maps skill. Do not install a second active top-level Maps skill in this repo. + +- Official docs: https://developers.google.com/maps/ai/agent-skills +- Official source: https://github.com/googlemaps/agent-skills +- Pinned reviewed copy: [`references/vendor/google-maps-platform/SKILL.md`](references/vendor/google-maps-platform/SKILL.md) +- Reviewed upstream commit: `84f0e9a2527403a408a61b8705bea0c3900b76a8` + +For Google Maps API/SDK implementation, read the pinned official skill first, then apply the MDE rules here. For changing facts such as API availability, deprecations, pricing, and regional coverage, verify current official Google documentation or Code Assist rather than historical MDE notes. + +MDE-specific ownership remains: Supabase owns inventory truth; Mastra owns orchestration; Maps/Places own geo truth; Gemini must not invent coordinates, place IDs, hours, or routes. + +--- + +## Quick routing + +| Task | Go to | +|------|-------| +| **PRD / audit** — Places API (New) v2.1 feature matrix + score (PLACES-002–081) | Repo: `tasks/maps/maps-prd-v2.md`, `tasks/maps/places-api-new-audit.md` | +| **Interactive** — search_places, get_directions, show_on_map in Claude session | [§ Interactive MCP tools below](#interactive-mcp-tools) | +| **CLI batch** — use the maintained batch helper and behavior notes | [`scripts/gmaps.py`](scripts/gmaps.py) + [`references/gmaps-cli-behavior.md`](references/gmaps-cli-behavior.md) | +| **Security** — API key architecture, HTML pages, embed iframes | [`references/security-and-optimization.md`](references/security-and-optimization.md) | +| **Former `google-maps` skill** — removed 2026-05-14 (last stub copy in `_archive/2026-05-14/google-maps-stub/`) | § [Interactive MCP tools](#interactive-mcp-tools) below | +| **Batch Maps helper** — `gmaps.py` + operator rules | [`scripts/gmaps.py`](scripts/gmaps.py) + [`references/gmaps-cli-behavior.md`](references/gmaps-cli-behavior.md) | +| **Former `react-google-maps` skill** — `@vis.gl/react-google-maps` | [`references/react-vis-gl/README.md`](references/react-vis-gl/README.md) | + +## mdeAI environment + +``` +NEXT_PUBLIC_GOOGLE_MAPS_API_KEY — Frontend (browser) — Maps JS API, AdvancedMarkerElement +GOOGLE_PLACES_API_KEY — Server-side only — Places API (New), enrichment scripts +GOOGLE_MAPS_API_KEY — Edge functions — Directions, Routes +GOOGLE_ROUTES_API_KEY — Edge functions — Routes API +``` + +**Medellín anchor:** `{ latitude: 6.2442, longitude: -75.5812 }` — default `locationBias` center and Maps grounding `latLng`. + +**Never expose `GOOGLE_PLACES_API_KEY` through a `NEXT_PUBLIC_*` variable** — it is server-side only. + +--- + +## Interactive MCP tools + +Use these when answering location questions **in a Claude session** (not for mdeAI production code). Tools call Google Maps APIs live. + +### Tools available + +``` +search_places(query, location?, radius?, type?, open_now?, language?) + query — text query ("restaurants in Laureles") + location — "lat,lng" center (optional) + radius — meters, max 50000 (optional) + type — place type filter ("restaurant", "tourist_attraction", "hotel") + open_now — boolean, default false + language — language code, default "en" + +search_nearby_places(location, radius, keyword?, type?, rank_by?, open_now?, language?) + location — "lat,lng" (required) + radius — meters (required, max 50000) + rank_by — "prominence" (default) or "distance" + +get_place_details(place_id, language?, reviews_sort?) + place_id — from search results + reviews_sort — "most_relevant" (default) or "newest" + +get_directions(origin, destination, mode?, alternatives?, avoid?, language?) + mode — "driving" (default), "walking", "bicycling", "transit" + avoid — "tolls", "highways", or "ferries" + +geocode_address(address, language?, region?) + region — country code for regional bias + +reverse_geocode(latlng, language?) + latlng — "lat,lng" + +show_on_map(map_type, markers?, directions?, center?, zoom?) + map_type — "markers", "directions", or "area" + markers — array of {lat, lng} objects +``` + +### Response pattern — Text → Map → Text + +**Always follow this sequence. Never call `show_on_map` in parallel with other calls.** + +1. **Text** — introduce what you'll show ("Here are top restaurants near Poblado:") +2. **Map** — call `show_on_map` to render results +3. **Text** — explain results in plain language (names, ratings, notes) + +**Multiple categories:** sequential maps — events then restaurants, not parallel. + +**Never echo raw map_data JSON** (coordinates, markers, zoom) in your text response. The map renders visually; describe places by name and quality only. + +### Intent → tool mapping + +| User says | Tool to use | +|-----------|-------------| +| "Where is X?" | `geocode_address` | +| "Find restaurants near..." | `search_places` or `search_nearby_places` | +| "What are the hours for...?" | `get_place_details` | +| "How do I get from A to B?" | `get_directions` | +| "What address is at these coords?" | `reverse_geocode` | +| "Show me these places on a map" | `show_on_map` | + +**Preserve `place_id`** from search results for use in `get_place_details`. + +--- + +## Places API (New) — field masks for mdeAI + +Every request uses `X-Goog-FieldMask` header. Only request fields you need — billing is field-mask driven. + +### Enrichment mask (PLACES-005-010) + +``` +places.id,places.displayName,places.googleMapsLinks,places.location,places.generativeSummary +``` + +| Field | Returns | mdeAI use | +|-------|---------|-----------| +| `places.id` | Place ID (`ChIJ...`) | Store as `place_id` | +| `places.googleMapsLinks.placeUri` | Canonical `https://maps.app.goo.gl/...` URL | Store as `maps_url` | +| `places.googleMapsLinks.directionsUri` | Directions link | Optional card button | +| `places.googleMapsLinks.photosUri` | Google Maps photos link | Optional "see photos" | +| `places.location` | `{ latitude, longitude }` | Backfill lat/lng | +| `places.generativeSummary` | `{ text, disclosureText }` | Store as `ai_summary`; show `disclosureText` | + +### generativeSummary constraints + +- **Coverage:** English only; US and India only currently +- **Attribution required:** Display `disclosureText` ("Summarized with Gemini") wherever `ai_summary` appears — ToS requirement +- **Cache in DB:** Fetch once at seeding time. Never call per chat turn. + +### googleMapsLinks — currently free + +`googleMapsLinks` is in preview and **free** as of 2026-05. Use `placeUri` (not lat/lng-constructed URLs) — it's stable and canonical. + +--- + +## Node.js client — enrichment script pattern + +```typescript +import { PlacesClient } from '@googlemaps/places'; + +const client = new PlacesClient({ apiKey: process.env.GOOGLE_PLACES_API_KEY }); + +const [response] = await client.searchText( + { + textQuery: `${venueName} ${neighborhood} Medellín Colombia`, + locationBias: { + circle: { center: { latitude: 6.2442, longitude: -75.5812 }, radius: 30000 }, + }, + }, + { otherArgs: { headers: { 'X-Goog-FieldMask': 'places.id,places.displayName,places.googleMapsLinks,places.location,places.generativeSummary' } } }, +); +``` + +--- + +## Gemini Maps grounding — summary + +Use current official Google Maps grounding documentation; do not rely on retired offline mirrors. + +| Mode | Free tier | Cost | Enable | +|------|-----------|------|--------| +| Grounding with Google Maps (Gemini API) | 500/day | $25/1K | `tools: [{ googleMaps: {} }]` | +| Maps Grounding Lite (MCP) — **GA** | pay-as-you-go | per SKU | `mapstools.googleapis.com/mcp` | + +**Kill switch:** `MAPS_GROUNDING_DAILY_LIMIT=0` → fall back to Supabase immediately. + +**Sequential calls for structured output:** grounded call (no `responseMimeType`) → structured output call (no grounding). Maps + custom function declarations CAN be combined in one call (March 2026 update). + +--- + +## Maps JavaScript API — ChatMap.tsx summary + +Use [`references/react-vis-gl/README.md`](references/react-vis-gl/README.md) plus current app source for Maps JavaScript implementation. + +- Loader: `@googlemaps/js-api-loader` with `libraries: ['marker']` +- `mapId` required for `AdvancedMarkerElement` +- `data-testid="map-pin"` on every pin content element (MASTRA-045 smoke spec) +- Per-category pin merge: `setPins(prev => [...prev.filter(p => p.category !== cat), ...newPins])` +- Frontend key restricted to HTTP referrers + Maps JS API only + +--- + +## Session tokens — autocomplete billing + +Use UUID v4 session tokens to group autocomplete keystrokes + final Place Details into one billing event: + +```typescript +import { v4 as uuidv4 } from 'uuid'; +const sessionToken = uuidv4(); // new UUID per search session +// Pass as sessionToken on each Autocomplete call +// Generate fresh UUID after user selects a place +``` + +--- + +## GCP key setup — quick reference + +| Key | Restrictions | APIs enabled | +|-----|-------------|-------------| +| `NEXT_PUBLIC_GOOGLE_MAPS_API_KEY` | HTTP referrers for approved MDE origins | Maps JavaScript API only | +| `GOOGLE_PLACES_API_KEY` | Server IP | Places API (New) only | +| `GOOGLE_MAPS_API_KEY` | Server IP | Directions API, Maps Static | +| `GOOGLE_ROUTES_API_KEY` | Server IP | Routes API | + +> Full 2-key security architecture → [`references/security-and-optimization.md`](references/security-and-optimization.md) + +**Enable "Places API (New)"** in GCP Console — NOT "Places API" (legacy). Different billing, different endpoints, different field names. + +--- + +## Event discovery — Maps / Places / ADK (plan 10 §11) + +| Layer | mdeai use | Skill / task | +|-------|-----------|----------------| +| **Places API (New)** | Batch venue enrich → `place_id`, `maps_url`, lat/lng | **EVD-06** → EVP-024 (historical) | +| **Maps JS** | Camila’s event pins (`mapId` + `AdvancedMarker`) | EVP-016 (historical) | +| **ADK sidecar** | Freshness / `search_grounded_places` — not event inventory | EVP-023 (historical) | +| **Web grounding** | C-004 citations — Google Search, not Places catalog | EVP-021 (historical) | + +Historical event-discovery task links were retired; resolve current work through Linear and the canonical `events` skill. + +**Golden rule:** Places enriches DB once; grounding answers live geo questions — never invent event listings from Maps. + +--- + +## Mastra handoff + +For Maps-related Mastra work, use the canonical `mastra` skill plus current source and the live Linear task. Retired `tasks/mastra/maps/**` paths are not active instructions. + +--- + +## Common gotchas + +| Gotcha | Fix | +|--------|-----| +| `generativeSummary` null | Handle gracefully — not all places have summaries | +| No `disclosureText` shown | Required by ToS — show "Summarized with Gemini" | +| Legacy Places API | Switch to Places API (New) — different billing, different endpoints | +| `googleMapsLinks` missing | Must be in field mask explicitly | +| Constructing Maps URLs from lat/lng | Use `placeUri` from `googleMapsLinks` — it's canonical and stable | +| `AdvancedMarkerElement` not found | Add `'marker'` to `libraries` in js-api-loader | +| Missing `mapId` | Required for AdvancedMarkerElement — set in Map constructor | +| Frontend key 403 | Verify HTTP referrer restriction includes current origin | +| Places server key in `NEXT_PUBLIC_*` | Server-side keys must never be browser-exposed | diff --git a/.claude/skills/maps/references/gmaps-cli-behavior.md b/.claude/skills/maps/references/gmaps-cli-behavior.md new file mode 100644 index 00000000..e93a5580 --- /dev/null +++ b/.claude/skills/maps/references/gmaps-cli-behavior.md @@ -0,0 +1,16 @@ +--- +title: gmaps.py — operator behavior (from legacy google-maps-api skill) +canonical_skill: maps +--- + +# `gmaps.py` — critical operator rules + +These rules were carried over when **`google-maps-api`** was folded into **`maps`**. Command reference: official Google Maps API documentation. Script path: **`.claude/skills/maps/scripts/gmaps.py`** (moved from legacy `google-maps-api/scripts/` on 2026-05-14). + +1. **Communicate blockers immediately.** On API failure (`403`, `REQUEST_DENIED`, API not enabled, etc.), **stop** and explain in plain language. Do not silently substitute web search for a failed Maps call unless the user explicitly asked for a fallback. + +2. **Ask before generating HTML.** Do not start a full interactive HTML page without confirming the user wants a visual artifact vs text/JSON. + +3. **Ask before choosing output format** when the request is ambiguous (summary vs table vs downloadable artifact). + +4. **Enablement:** If an API is disabled in GCP, say which API to enable and link Console paths — see official Google Maps API documentation API list table. diff --git a/.claude/skills/maps/references/react-vis-gl/README.md b/.claude/skills/maps/references/react-vis-gl/README.md new file mode 100644 index 00000000..3f5bd306 --- /dev/null +++ b/.claude/skills/maps/references/react-vis-gl/README.md @@ -0,0 +1,19 @@ +--- +title: "@vis.gl/react-google-maps — reference pack" +canonical_skill: maps +source: Consolidated from the former `react-google-maps` skill (symlink stub only). +--- + +# @vis.gl/react-google-maps — deep references + +Use this folder for **React + Maps JS API** work beyond the summary in [Google Maps JavaScript API documentation](https://developers.google.com/maps/documentation/javascript). + +| Task | File | +|------|------| +| Map, Marker, InfoWindow, Pin API | [components-api.md](components-api.md) | +| `useMap`, `useMapsLibrary`, hooks | [hooks-api.md](hooks-api.md) | +| Circle, Polygon, Polyline (manual overlays) | [geometry-components.md](geometry-components.md) | +| Places autocomplete / geocoding UI | [places-autocomplete.md](places-autocomplete.md) | +| Draggable markers, controlled maps | [patterns.md](patterns.md) | + +**mdeai rule:** extend **`maps`** only — do not add new content to the `react-google-maps` symlink stub. diff --git a/.claude/skills/maps/references/react-vis-gl/components-api.md b/.claude/skills/maps/references/react-vis-gl/components-api.md new file mode 100644 index 00000000..f13ac264 --- /dev/null +++ b/.claude/skills/maps/references/react-vis-gl/components-api.md @@ -0,0 +1,272 @@ +# Components API Reference + +## APIProvider + +Loads the Google Maps JavaScript API. Must wrap all map components. + +```tsx + + {children} + +``` + +**Props:** +- `apiKey` (required): Google Maps API key +- `libraries`: Array of libraries to preload (`'places'`, `'geocoding'`, `'drawing'`, `'geometry'`, `'visualization'`) +- `version`: API version (`'weekly'`, `'quarterly'`, `'beta'`) +- `region`: Region code for localized behavior +- `language`: Language code + +--- + +## Map + +The map container component. + +```tsx + console.log(e.detail.latLng)} + onCameraChanged={(e) => console.log(e.detail)} +> + {/* Markers, InfoWindows, etc. */} + +``` + +### Controlled vs Uncontrolled + +**Uncontrolled (default):** Use `defaultCenter`, `defaultZoom` - map manages its own state after init. + +**Controlled:** Use `center`, `zoom` - React controls the viewport, map always syncs to props. + +```tsx +// Uncontrolled - user can freely pan/zoom + + +// Controlled - viewport locked to React state +const [camera, setCamera] = useState({ center, zoom: 10 }); + setCamera(e.detail)} +/> +``` + +### Key Props + +| Prop | Type | Description | +|------|------|-------------| +| `mapId` | string | Required for AdvancedMarker, enables cloud styling | +| `defaultCenter` / `center` | LatLngLiteral | Initial/controlled center | +| `defaultZoom` / `zoom` | number | Initial/controlled zoom (0-22) | +| `gestureHandling` | `'cooperative'` \| `'greedy'` \| `'none'` \| `'auto'` | How map handles gestures | +| `disableDefaultUI` | boolean | Hide all default controls | +| `mapTypeId` | `'roadmap'` \| `'satellite'` \| `'hybrid'` \| `'terrain'` | Map type | +| `clickableIcons` | boolean | Whether POI icons are clickable | +| `style` / `className` | CSSProperties / string | Container styling | + +### Events + +| Event | Type | Description | +|-------|------|-------------| +| `onClick` | `(e: MapMouseEvent) => void` | Map clicked | +| `onDblClick` | `(e: MapMouseEvent) => void` | Map double-clicked | +| `onContextMenu` | `(e: MapMouseEvent) => void` | Right-click | +| `onCameraChanged` | `(e: CameraChangedEvent) => void` | Camera changed | +| `onBoundsChanged` | `() => void` | Bounds changed | +| `onIdle` | `() => void` | Map idle after pan/zoom | + +--- + +## AdvancedMarker + +Modern marker component. **Requires `mapId` on the Map component.** + +```tsx + console.log('clicked')} + onDrag={(e) => console.log(e.latLng)} + onDragEnd={(e) => console.log('final position:', e.latLng)} +> + {/* Optional: custom content instead of default pin */} + + +``` + +### Custom HTML Marker + +```tsx + +
+ marker + Label +
+
+``` + +### Props + +| Prop | Type | Description | +|------|------|-------------| +| `position` | LatLngLiteral \| LatLngAltitudeLiteral | Marker position | +| `title` | string | Accessibility title | +| `draggable` | boolean | Enable dragging | +| `clickable` | boolean | Enable click events | +| `zIndex` | number | Stacking order | +| `collisionBehavior` | CollisionBehavior | Collision handling | +| `className` / `style` | string / CSSProperties | Content element styling | + +### Events + +| Event | Type | Description | +|-------|------|-------------| +| `onClick` | `(e: AdvancedMarkerClickEvent) => void` | Marker clicked | +| `onDragStart` | `(e: MapMouseEvent) => void` | Drag started | +| `onDrag` | `(e: MapMouseEvent) => void` | **During drag** (for real-time updates) | +| `onDragEnd` | `(e: MapMouseEvent) => void` | Drag ended | +| `onMouseEnter` | `(e: MouseEvent) => void` | Mouse entered | +| `onMouseLeave` | `(e: MouseEvent) => void` | Mouse left | + +### Anchor Point + +Control which part of the marker aligns with the position: + +```tsx +import { AdvancedMarkerAnchorPoint } from '@vis.gl/react-google-maps'; + + + + +``` + +--- + +## Pin + +Customizable pin for AdvancedMarker. + +```tsx + + + +``` + +### Props + +| Prop | Type | Description | +|------|------|-------------| +| `background` | string | Pin background color | +| `borderColor` | string | Pin border color | +| `glyphColor` | string | Glyph (icon/text) color | +| `glyph` | string \| Element | Content inside pin | +| `scale` | number | Size multiplier | + +--- + +## InfoWindow + +Popup window attached to a marker or position. + +```tsx +const [markerRef, marker] = useAdvancedMarkerRef(); +const [open, setOpen] = useState(false); + + setOpen(true)} /> + +{open && ( + setOpen(false)} + headerContent={

Title

} + > +

Window content

+
+)} +``` + +### Props + +| Prop | Type | Description | +|------|------|-------------| +| `anchor` | AdvancedMarkerElement \| Marker | Anchor to marker | +| `position` | LatLngLiteral | Position if no anchor | +| `onClose` | `() => void` | Close callback (required for sync) | +| `onCloseClick` | `() => void` | Close button clicked | +| `headerContent` | ReactNode | Header content | +| `headerDisabled` | boolean | Hide header | +| `minWidth` / `maxWidth` | number | Width constraints | +| `disableAutoPan` | boolean | Don't pan map to show window | +| `pixelOffset` | [number, number] | Offset from anchor | + +**Important:** Always provide `onClose` to keep React state in sync when the map closes the InfoWindow. + +--- + +## MapControl + +Add custom UI elements to the map. + +```tsx +import { ControlPosition } from '@vis.gl/react-google-maps'; + + + + + + + +
Legend content
+
+
+``` + +### ControlPosition Values + +``` +TOP_LEFT TOP_CENTER TOP_RIGHT +LEFT_TOP RIGHT_TOP +LEFT_CENTER RIGHT_CENTER +LEFT_BOTTOM RIGHT_BOTTOM +BOTTOM_LEFT BOTTOM_CENTER BOTTOM_RIGHT +``` + +--- + +## Marker (Deprecated) + +Legacy marker. Use AdvancedMarker for new projects. + +```tsx +import { useMarkerRef } from '@vis.gl/react-google-maps'; + +const [markerRef, marker] = useMarkerRef(); + + +``` diff --git a/.claude/skills/maps/references/react-vis-gl/geometry-components.md b/.claude/skills/maps/references/react-vis-gl/geometry-components.md new file mode 100644 index 00000000..95ef218c --- /dev/null +++ b/.claude/skills/maps/references/react-vis-gl/geometry-components.md @@ -0,0 +1,495 @@ +# Geometry Components (Circle, Polygon, Polyline) + +**Important:** These components are NOT exported by `@vis.gl/react-google-maps`. Copy the implementations below into your project. + +## Circle Component + +```tsx +// components/circle.tsx +import { forwardRef, useContext, useEffect, useImperativeHandle, useRef } from 'react'; +import { GoogleMapsContext } from '@vis.gl/react-google-maps'; +import type { Ref } from 'react'; + +type CircleEventProps = { + onClick?: (e: google.maps.MapMouseEvent) => void; + onDrag?: (e: google.maps.MapMouseEvent) => void; + onDragStart?: (e: google.maps.MapMouseEvent) => void; + onDragEnd?: (e: google.maps.MapMouseEvent) => void; + onMouseOver?: (e: google.maps.MapMouseEvent) => void; + onMouseOut?: (e: google.maps.MapMouseEvent) => void; + onRadiusChanged?: (radius: number) => void; + onCenterChanged?: (center: google.maps.LatLng) => void; +}; + +export type CircleProps = google.maps.CircleOptions & CircleEventProps; +export type CircleRef = Ref; + +function useCircle(props: CircleProps) { + const { + onClick, + onDrag, + onDragStart, + onDragEnd, + onMouseOver, + onMouseOut, + onRadiusChanged, + onCenterChanged, + center, + radius, + ...circleOptions + } = props; + + const circleRef = useRef(null); + const map = useContext(GoogleMapsContext)?.map; + + // Create circle instance + useEffect(() => { + if (!map) return; + + const circle = new google.maps.Circle(); + circle.setMap(map); + circleRef.current = circle; + + return () => { + circle.setMap(null); + circleRef.current = null; + }; + }, [map]); + + // Update options + useEffect(() => { + if (!circleRef.current) return; + circleRef.current.setOptions(circleOptions); + }, [circleOptions]); + + // Update center + useEffect(() => { + if (!circleRef.current || !center) return; + circleRef.current.setCenter(center); + }, [center]); + + // Update radius + useEffect(() => { + if (!circleRef.current || radius === undefined) return; + circleRef.current.setRadius(radius); + }, [radius]); + + // Event listeners + useEffect(() => { + if (!circleRef.current) return; + const circle = circleRef.current; + const listeners: google.maps.MapsEventListener[] = []; + + if (onClick) listeners.push(circle.addListener('click', onClick)); + if (onDrag) listeners.push(circle.addListener('drag', onDrag)); + if (onDragStart) listeners.push(circle.addListener('dragstart', onDragStart)); + if (onDragEnd) listeners.push(circle.addListener('dragend', onDragEnd)); + if (onMouseOver) listeners.push(circle.addListener('mouseover', onMouseOver)); + if (onMouseOut) listeners.push(circle.addListener('mouseout', onMouseOut)); + if (onRadiusChanged) { + listeners.push(circle.addListener('radius_changed', () => { + onRadiusChanged(circle.getRadius() ?? 0); + })); + } + if (onCenterChanged) { + listeners.push(circle.addListener('center_changed', () => { + const c = circle.getCenter(); + if (c) onCenterChanged(c); + })); + } + + return () => listeners.forEach((l) => l.remove()); + }, [onClick, onDrag, onDragStart, onDragEnd, onMouseOver, onMouseOut, onRadiusChanged, onCenterChanged]); + + return circleRef; +} + +export const Circle = forwardRef((props, ref) => { + const circleRef = useCircle(props); + useImperativeHandle(ref, () => circleRef.current, []); + return null; +}); + +Circle.displayName = 'Circle'; +``` + +### Circle Usage + +```tsx +import { Circle } from '@/components/circle'; + + console.log('Clicked', e.latLng)} +/> +``` + +--- + +## Polygon Component + +```tsx +// components/polygon.tsx +import { forwardRef, useContext, useEffect, useImperativeHandle, useRef } from 'react'; +import { GoogleMapsContext, useMapsLibrary } from '@vis.gl/react-google-maps'; +import type { Ref } from 'react'; + +type PolygonEventProps = { + onClick?: (e: google.maps.MapMouseEvent) => void; + onDrag?: (e: google.maps.MapMouseEvent) => void; + onDragStart?: (e: google.maps.MapMouseEvent) => void; + onDragEnd?: (e: google.maps.MapMouseEvent) => void; + onMouseOver?: (e: google.maps.MapMouseEvent) => void; + onMouseOut?: (e: google.maps.MapMouseEvent) => void; +}; + +type PolygonCustomProps = { + encodedPaths?: string[]; // Encoded polyline paths +}; + +export type PolygonProps = google.maps.PolygonOptions & PolygonEventProps & PolygonCustomProps; +export type PolygonRef = Ref; + +function usePolygon(props: PolygonProps) { + const { + onClick, + onDrag, + onDragStart, + onDragEnd, + onMouseOver, + onMouseOut, + encodedPaths, + paths, + ...polygonOptions + } = props; + + const polygonRef = useRef(null); + const map = useContext(GoogleMapsContext)?.map; + const geometryLib = useMapsLibrary('geometry'); + + // Create polygon instance + useEffect(() => { + if (!map) return; + + const polygon = new google.maps.Polygon(); + polygon.setMap(map); + polygonRef.current = polygon; + + return () => { + polygon.setMap(null); + polygonRef.current = null; + }; + }, [map]); + + // Update options + useEffect(() => { + if (!polygonRef.current) return; + polygonRef.current.setOptions(polygonOptions); + }, [polygonOptions]); + + // Update paths + useEffect(() => { + if (!polygonRef.current) return; + + if (encodedPaths && geometryLib) { + const decodedPaths = encodedPaths.map((p) => + geometryLib.encoding.decodePath(p) + ); + polygonRef.current.setPaths(decodedPaths); + } else if (paths) { + polygonRef.current.setPaths(paths); + } + }, [paths, encodedPaths, geometryLib]); + + // Event listeners + useEffect(() => { + if (!polygonRef.current) return; + const polygon = polygonRef.current; + const listeners: google.maps.MapsEventListener[] = []; + + if (onClick) listeners.push(polygon.addListener('click', onClick)); + if (onDrag) listeners.push(polygon.addListener('drag', onDrag)); + if (onDragStart) listeners.push(polygon.addListener('dragstart', onDragStart)); + if (onDragEnd) listeners.push(polygon.addListener('dragend', onDragEnd)); + if (onMouseOver) listeners.push(polygon.addListener('mouseover', onMouseOver)); + if (onMouseOut) listeners.push(polygon.addListener('mouseout', onMouseOut)); + + return () => listeners.forEach((l) => l.remove()); + }, [onClick, onDrag, onDragStart, onDragEnd, onMouseOver, onMouseOut]); + + return polygonRef; +} + +export const Polygon = forwardRef((props, ref) => { + const polygonRef = usePolygon(props); + useImperativeHandle(ref, () => polygonRef.current, []); + return null; +}); + +Polygon.displayName = 'Polygon'; +``` + +### Polygon Usage + +```tsx +import { Polygon } from '@/components/polygon'; + +const trianglePaths = [ + { lat: 25.774, lng: -80.19 }, + { lat: 18.466, lng: -66.118 }, + { lat: 32.321, lng: -64.757 }, +]; + + console.log('Dragging polygon')} +/> +``` + +--- + +## Polyline Component + +```tsx +// components/polyline.tsx +import { forwardRef, useContext, useEffect, useImperativeHandle, useRef } from 'react'; +import { GoogleMapsContext, useMapsLibrary } from '@vis.gl/react-google-maps'; +import type { Ref } from 'react'; + +type PolylineEventProps = { + onClick?: (e: google.maps.MapMouseEvent) => void; + onDrag?: (e: google.maps.MapMouseEvent) => void; + onDragStart?: (e: google.maps.MapMouseEvent) => void; + onDragEnd?: (e: google.maps.MapMouseEvent) => void; + onMouseOver?: (e: google.maps.MapMouseEvent) => void; + onMouseOut?: (e: google.maps.MapMouseEvent) => void; +}; + +type PolylineCustomProps = { + encodedPath?: string; // Encoded polyline path +}; + +export type PolylineProps = google.maps.PolylineOptions & PolylineEventProps & PolylineCustomProps; +export type PolylineRef = Ref; + +function usePolyline(props: PolylineProps) { + const { + onClick, + onDrag, + onDragStart, + onDragEnd, + onMouseOver, + onMouseOut, + encodedPath, + path, + ...polylineOptions + } = props; + + const polylineRef = useRef(null); + const map = useContext(GoogleMapsContext)?.map; + const geometryLib = useMapsLibrary('geometry'); + + // Create polyline instance + useEffect(() => { + if (!map) return; + + const polyline = new google.maps.Polyline(); + polyline.setMap(map); + polylineRef.current = polyline; + + return () => { + polyline.setMap(null); + polylineRef.current = null; + }; + }, [map]); + + // Update options + useEffect(() => { + if (!polylineRef.current) return; + polylineRef.current.setOptions(polylineOptions); + }, [polylineOptions]); + + // Update path + useEffect(() => { + if (!polylineRef.current) return; + + if (encodedPath && geometryLib) { + const decodedPath = geometryLib.encoding.decodePath(encodedPath); + polylineRef.current.setPath(decodedPath); + } else if (path) { + polylineRef.current.setPath(path); + } + }, [path, encodedPath, geometryLib]); + + // Event listeners + useEffect(() => { + if (!polylineRef.current) return; + const polyline = polylineRef.current; + const listeners: google.maps.MapsEventListener[] = []; + + if (onClick) listeners.push(polyline.addListener('click', onClick)); + if (onDrag) listeners.push(polyline.addListener('drag', onDrag)); + if (onDragStart) listeners.push(polyline.addListener('dragstart', onDragStart)); + if (onDragEnd) listeners.push(polyline.addListener('dragend', onDragEnd)); + if (onMouseOver) listeners.push(polyline.addListener('mouseover', onMouseOver)); + if (onMouseOut) listeners.push(polyline.addListener('mouseout', onMouseOut)); + + return () => listeners.forEach((l) => l.remove()); + }, [onClick, onDrag, onDragStart, onDragEnd, onMouseOver, onMouseOut]); + + return polylineRef; +} + +export const Polyline = forwardRef((props, ref) => { + const polylineRef = usePolyline(props); + useImperativeHandle(ref, () => polylineRef.current, []); + return null; +}); + +Polyline.displayName = 'Polyline'; +``` + +### Polyline Usage + +```tsx +import { Polyline } from '@/components/polyline'; + +const routePath = [ + { lat: 37.772, lng: -122.214 }, + { lat: 21.291, lng: -157.821 }, + { lat: -18.142, lng: 178.431 }, + { lat: -27.467, lng: 153.027 }, +]; + + +``` + +--- + +## Real-Time Sync: Marker + Circle + +Key pattern for making shapes follow a draggable marker during drag (not just on drop): + +```tsx +function DraggableMarkerWithCircle() { + const [position, setPosition] = useState({ lat: 40.7128, lng: -74.006 }); + + return ( + + { + // Update position DURING drag - this is the key! + if (e.latLng) { + setPosition({ + lat: e.latLng.lat(), + lng: e.latLng.lng(), + }); + } + }} + /> + + {/* Circle follows marker in real-time */} + + + ); +} +``` + +### With Multiple Shapes + +```tsx +function MarkerWithMultipleShapes() { + const [center, setCenter] = useState({ lat: 40.7128, lng: -74.006 }); + + // Calculate polygon points based on center + const polygonPaths = useMemo(() => { + const offset = 0.01; + return [ + { lat: center.lat + offset, lng: center.lng }, + { lat: center.lat, lng: center.lng + offset }, + { lat: center.lat - offset, lng: center.lng }, + { lat: center.lat, lng: center.lng - offset }, + ]; + }, [center]); + + return ( + <> + { + if (e.latLng) { + setCenter({ lat: e.latLng.lat(), lng: e.latLng.lng() }); + } + }} + /> + + + + + ); +} +``` + +--- + +## Common Options Reference + +### Shared Shape Options + +| Option | Type | Description | +|--------|------|-------------| +| `strokeColor` | string | Stroke color (CSS color) | +| `strokeOpacity` | number | Stroke opacity (0-1) | +| `strokeWeight` | number | Stroke width in pixels | +| `fillColor` | string | Fill color (CSS color) | +| `fillOpacity` | number | Fill opacity (0-1) | +| `draggable` | boolean | Can be dragged | +| `editable` | boolean | Can be edited (resize/reshape) | +| `visible` | boolean | Is visible | +| `zIndex` | number | Stacking order | +| `clickable` | boolean | Responds to click events | + +### Circle-Specific + +| Option | Type | Description | +|--------|------|-------------| +| `center` | LatLngLiteral | Circle center | +| `radius` | number | Radius in meters | + +### Polygon/Polyline-Specific + +| Option | Type | Description | +|--------|------|-------------| +| `paths` / `path` | LatLngLiteral[] | Array of coordinates | +| `geodesic` | boolean | Geodesic segments (follow Earth curve) | diff --git a/.claude/skills/maps/references/react-vis-gl/hooks-api.md b/.claude/skills/maps/references/react-vis-gl/hooks-api.md new file mode 100644 index 00000000..e017f028 --- /dev/null +++ b/.claude/skills/maps/references/react-vis-gl/hooks-api.md @@ -0,0 +1,305 @@ +# Hooks API Reference + +## useMap + +Access the `google.maps.Map` instance. + +```tsx +import { useMap } from '@vis.gl/react-google-maps'; + +function MapControls() { + const map = useMap(); + + useEffect(() => { + if (!map) return; + + // Access all google.maps.Map methods + map.panTo({ lat: 40.7128, lng: -74.006 }); + map.setZoom(15); + map.fitBounds(bounds); + }, [map]); + + return ; +} +``` + +### With Multiple Maps + +```tsx +// Give maps explicit IDs + + + +// Access specific map +const mainMap = useMap('main-map'); +const miniMap = useMap('mini-map'); +``` + +### Common Map Methods + +```tsx +const map = useMap(); + +// Viewport control +map.panTo(latLng); +map.setCenter(latLng); +map.setZoom(level); +map.fitBounds(bounds, padding?); +map.panToBounds(bounds); + +// Get current state +map.getCenter(); +map.getZoom(); +map.getBounds(); + +// Other +map.setMapTypeId('satellite'); +map.setOptions({ gestureHandling: 'none' }); +``` + +--- + +## useMapsLibrary + +Dynamically load Google Maps libraries. Returns the library object or `null` while loading. + +```tsx +import { useMapsLibrary } from '@vis.gl/react-google-maps'; + +function GeocodingComponent() { + const geocodingLib = useMapsLibrary('geocoding'); + const [geocoder, setGeocoder] = useState(null); + + useEffect(() => { + if (!geocodingLib) return; + setGeocoder(new geocodingLib.Geocoder()); + }, [geocodingLib]); + + // Use geocoder... +} +``` + +### Available Libraries + +| Library | Use Case | Key Classes | +|---------|----------|-------------| +| `'places'` | Places search, autocomplete | `AutocompleteService`, `PlacesService` | +| `'geocoding'` | Address ↔ coordinates | `Geocoder` | +| `'drawing'` | Drawing tools | `DrawingManager` | +| `'geometry'` | Distance, area calculations | `spherical`, `poly`, `encoding` | +| `'visualization'` | Heatmaps | `HeatmapLayer` | +| `'marker'` | Marker utilities | `PinElement` | +| `'routes'` | Directions | `DirectionsService`, `DirectionsRenderer` | +| `'maps3d'` | 3D maps | `Map3DElement` | + +### Example: Geocoding + +```tsx +function useGeocoder() { + const geocodingLib = useMapsLibrary('geocoding'); + + const geocode = useCallback(async (address: string) => { + if (!geocodingLib) return null; + + const geocoder = new geocodingLib.Geocoder(); + const response = await geocoder.geocode({ address }); + + if (response.results.length > 0) { + const { lat, lng } = response.results[0].geometry.location; + return { lat: lat(), lng: lng() }; + } + return null; + }, [geocodingLib]); + + return { geocode, isReady: !!geocodingLib }; +} +``` + +### Example: Directions + +```tsx +function useDirections() { + const routesLib = useMapsLibrary('routes'); + const map = useMap(); + const [renderer, setRenderer] = useState(null); + + useEffect(() => { + if (!routesLib || !map) return; + + const directionsRenderer = new routesLib.DirectionsRenderer({ map }); + setRenderer(directionsRenderer); + + return () => directionsRenderer.setMap(null); + }, [routesLib, map]); + + const getRoute = useCallback(async (origin: string, destination: string) => { + if (!routesLib || !renderer) return; + + const service = new routesLib.DirectionsService(); + const result = await service.route({ + origin, + destination, + travelMode: google.maps.TravelMode.DRIVING, + }); + + renderer.setDirections(result); + }, [routesLib, renderer]); + + return { getRoute, isReady: !!renderer }; +} +``` + +--- + +## useAdvancedMarkerRef + +Connect an AdvancedMarker to an InfoWindow. + +```tsx +import { useAdvancedMarkerRef } from '@vis.gl/react-google-maps'; + +function MarkerWithInfoWindow({ position }: { position: google.maps.LatLngLiteral }) { + const [markerRef, marker] = useAdvancedMarkerRef(); + const [infoOpen, setInfoOpen] = useState(false); + + return ( + <> + setInfoOpen(true)} + /> + + {infoOpen && ( + setInfoOpen(false)}> +
Info content
+
+ )} + + ); +} +``` + +### Return Value + +```tsx +const [markerRef, marker] = useAdvancedMarkerRef(); +// markerRef: RefCallback to pass to AdvancedMarker's ref prop +// marker: The AdvancedMarkerElement instance (or null) +``` + +--- + +## useMarkerRef + +Same as `useAdvancedMarkerRef` but for the legacy `Marker` component. + +```tsx +const [markerRef, marker] = useMarkerRef(); + + +``` + +--- + +## useApiIsLoaded + +Check if the Maps JavaScript API has finished loading. + +```tsx +import { useApiIsLoaded } from '@vis.gl/react-google-maps'; + +function MapComponent() { + const isLoaded = useApiIsLoaded(); + + if (!isLoaded) { + return
Loading map...
; + } + + return ; +} +``` + +--- + +## useApiLoadingStatus + +Get detailed loading status. + +```tsx +import { useApiLoadingStatus, APILoadingStatus } from '@vis.gl/react-google-maps'; + +function MapComponent() { + const status = useApiLoadingStatus(); + + switch (status) { + case APILoadingStatus.NOT_LOADED: + return
Not started
; + case APILoadingStatus.LOADING: + return
Loading...
; + case APILoadingStatus.LOADED: + return ; + case APILoadingStatus.FAILED: + return
Failed to load
; + } +} +``` + +--- + +## Custom Hook Patterns + +### useMapBounds + +```tsx +function useMapBounds() { + const map = useMap(); + const [bounds, setBounds] = useState(null); + + useEffect(() => { + if (!map) return; + + const listener = map.addListener('bounds_changed', () => { + setBounds(map.getBounds() ?? null); + }); + + return () => listener.remove(); + }, [map]); + + return bounds; +} +``` + +### useFitBounds + +```tsx +function useFitBounds(positions: google.maps.LatLngLiteral[]) { + const map = useMap(); + + useEffect(() => { + if (!map || positions.length === 0) return; + + const bounds = new google.maps.LatLngBounds(); + positions.forEach((pos) => bounds.extend(pos)); + map.fitBounds(bounds, { padding: 50 }); + }, [map, positions]); +} +``` + +### useMapClick + +```tsx +function useMapClick(callback: (latLng: google.maps.LatLng) => void) { + const map = useMap(); + + useEffect(() => { + if (!map) return; + + const listener = map.addListener('click', (e: google.maps.MapMouseEvent) => { + if (e.latLng) callback(e.latLng); + }); + + return () => listener.remove(); + }, [map, callback]); +} +``` diff --git a/.claude/skills/maps/references/react-vis-gl/patterns.md b/.claude/skills/maps/references/react-vis-gl/patterns.md new file mode 100644 index 00000000..aaf04e9b --- /dev/null +++ b/.claude/skills/maps/references/react-vis-gl/patterns.md @@ -0,0 +1,577 @@ +# Advanced Patterns + +## Controlled vs Uncontrolled Maps + +### Uncontrolled (Default) + +Map manages its own viewport. Use `default*` props for initial values: + +```tsx + +``` + +User can freely pan/zoom. React doesn't track viewport changes. + +### Controlled + +React controls the viewport. Map always syncs to props: + +```tsx +function ControlledMap() { + const [camera, setCamera] = useState({ + center: { lat: 40.7128, lng: -74.006 }, + zoom: 12, + }); + + return ( + setCamera(e.detail)} + /> + ); +} +``` + +### Hybrid: Track without Lock + +Track viewport in state but don't force it: + +```tsx +function TrackedMap() { + const [viewport, setViewport] = useState({ + center: { lat: 40.7128, lng: -74.006 }, + zoom: 12, + }); + + return ( + <> + setViewport(e.detail)} + /> +

Current zoom: {viewport.zoom}

+ + ); +} +``` + +--- + +## Real-Time Drag Synchronization + +The key to making shapes follow markers during drag is using `onDrag` (not just `onDragEnd`): + +### Single Shape Following Marker + +```tsx +function MarkerWithCircle() { + const [position, setPosition] = useState({ lat: 40.7128, lng: -74.006 }); + + return ( + <> + { + // This fires continuously during drag + if (e.latLng) { + setPosition({ + lat: e.latLng.lat(), + lng: e.latLng.lng(), + }); + } + }} + /> + + + ); +} +``` + +### Multiple Dependent Shapes + +```tsx +function MarkerWithDependentShapes() { + const [center, setCenter] = useState({ lat: 40.7128, lng: -74.006 }); + + // Derived positions for polygon vertices + const polygonPath = useMemo(() => { + const d = 0.005; // offset in degrees + return [ + { lat: center.lat + d, lng: center.lng }, + { lat: center.lat, lng: center.lng + d }, + { lat: center.lat - d, lng: center.lng }, + { lat: center.lat, lng: center.lng - d }, + ]; + }, [center]); + + // Polyline from center to a fixed point + const polylinePath = useMemo(() => [ + center, + { lat: 40.72, lng: -73.99 }, // fixed endpoint + ], [center]); + + return ( + <> + { + if (e.latLng) { + setCenter({ lat: e.latLng.lat(), lng: e.latLng.lng() }); + } + }} + /> + + + + + ); +} +``` + +### Performance: Throttled Updates + +For complex shapes, throttle updates: + +```tsx +import { useCallback, useRef } from 'react'; + +function useThrottledDrag(callback: (pos: google.maps.LatLngLiteral) => void, delay = 16) { + const lastCall = useRef(0); + + return useCallback((e: google.maps.MapMouseEvent) => { + const now = Date.now(); + if (now - lastCall.current >= delay && e.latLng) { + lastCall.current = now; + callback({ lat: e.latLng.lat(), lng: e.latLng.lng() }); + } + }, [callback, delay]); +} + +// Usage +function OptimizedDrag() { + const [position, setPosition] = useState({ lat: 40.7128, lng: -74.006 }); + const handleDrag = useThrottledDrag(setPosition, 32); // ~30fps + + return ( + + ); +} +``` + +--- + +## Drawing Manager + +Allow users to draw shapes on the map: + +```tsx +import { useMap, useMapsLibrary } from '@vis.gl/react-google-maps'; + +function useDrawingManager() { + const map = useMap(); + const drawingLib = useMapsLibrary('drawing'); + const [drawingManager, setDrawingManager] = useState(null); + + useEffect(() => { + if (!map || !drawingLib) return; + + const manager = new drawingLib.DrawingManager({ + map, + drawingMode: null, // Start with no drawing mode + drawingControl: true, + drawingControlOptions: { + position: google.maps.ControlPosition.TOP_CENTER, + drawingModes: [ + google.maps.drawing.OverlayType.MARKER, + google.maps.drawing.OverlayType.CIRCLE, + google.maps.drawing.OverlayType.POLYGON, + google.maps.drawing.OverlayType.POLYLINE, + google.maps.drawing.OverlayType.RECTANGLE, + ], + }, + circleOptions: { + fillColor: '#FF0000', + fillOpacity: 0.3, + strokeWeight: 2, + editable: true, + draggable: true, + }, + polygonOptions: { + fillColor: '#00FF00', + fillOpacity: 0.3, + strokeWeight: 2, + editable: true, + draggable: true, + }, + }); + + setDrawingManager(manager); + + return () => manager.setMap(null); + }, [map, drawingLib]); + + return drawingManager; +} + +// Handle drawn shapes +function DrawableMap() { + const drawingManager = useDrawingManager(); + const [shapes, setShapes] = useState([]); + + useEffect(() => { + if (!drawingManager) return; + + const listeners = [ + google.maps.event.addListener(drawingManager, 'circlecomplete', (circle: google.maps.Circle) => { + setShapes((prev) => [...prev, circle]); + console.log('Circle:', circle.getCenter()?.toJSON(), circle.getRadius()); + }), + google.maps.event.addListener(drawingManager, 'polygoncomplete', (polygon: google.maps.Polygon) => { + setShapes((prev) => [...prev, polygon]); + const path = polygon.getPath().getArray().map(p => p.toJSON()); + console.log('Polygon:', path); + }), + ]; + + return () => listeners.forEach((l) => l.remove()); + }, [drawingManager]); + + return ; +} +``` + +--- + +## Synchronized Maps + +Keep multiple maps in sync: + +```tsx +function SyncedMaps() { + const [camera, setCamera] = useState({ + center: { lat: 40.7128, lng: -74.006 }, + zoom: 12, + }); + + return ( +
+ setCamera(e.detail)} + mapId="MAP_ID_1" + > + + + + setCamera(e.detail)} + mapTypeId="satellite" + mapId="MAP_ID_2" + /> +
+ ); +} +``` + +--- + +## Marker Clustering + +Use `@googlemaps/markerclusterer`: + +```bash +npm install @googlemaps/markerclusterer +``` + +```tsx +import { MarkerClusterer } from '@googlemaps/markerclusterer'; +import { useMap, AdvancedMarker } from '@vis.gl/react-google-maps'; + +function ClusteredMarkers({ points }: { points: google.maps.LatLngLiteral[] }) { + const map = useMap(); + const [markers, setMarkers] = useState([]); + const clusterer = useRef(null); + + // Initialize clusterer + useEffect(() => { + if (!map) return; + + clusterer.current = new MarkerClusterer({ map }); + + return () => { + clusterer.current?.clearMarkers(); + }; + }, [map]); + + // Update markers in clusterer + useEffect(() => { + if (!clusterer.current) return; + + clusterer.current.clearMarkers(); + clusterer.current.addMarkers(markers); + }, [markers]); + + // Collect marker refs + const setMarkerRef = useCallback((marker: google.maps.marker.AdvancedMarkerElement | null, key: string) => { + setMarkers((prev) => { + if (marker) { + return [...prev.filter((m) => m.title !== key), marker]; + } + return prev.filter((m) => m.title !== key); + }); + }, []); + + return ( + <> + {points.map((point, i) => ( + setMarkerRef(marker, `marker-${i}`)} + /> + ))} + + ); +} +``` + +--- + +## Heatmap Layer + +```tsx +function HeatmapLayer({ data }: { data: google.maps.LatLngLiteral[] }) { + const map = useMap(); + const visualizationLib = useMapsLibrary('visualization'); + + useEffect(() => { + if (!map || !visualizationLib || !data.length) return; + + const heatmap = new visualizationLib.HeatmapLayer({ + data: data.map((d) => new google.maps.LatLng(d.lat, d.lng)), + map, + radius: 20, + opacity: 0.7, + }); + + return () => heatmap.setMap(null); + }, [map, visualizationLib, data]); + + return null; +} +``` + +--- + +## Fit Bounds to Markers + +```tsx +function useFitBounds(positions: google.maps.LatLngLiteral[], padding = 50) { + const map = useMap(); + + useEffect(() => { + if (!map || positions.length === 0) return; + + const bounds = new google.maps.LatLngBounds(); + positions.forEach((pos) => bounds.extend(pos)); + + map.fitBounds(bounds, padding); + }, [map, positions, padding]); +} + +// Usage +function MapWithMarkers({ markers }: { markers: google.maps.LatLngLiteral[] }) { + useFitBounds(markers); + + return ( + <> + {markers.map((pos, i) => ( + + ))} + + ); +} +``` + +--- + +## Custom Overlay (Advanced) + +For completely custom rendering: + +```tsx +import { useMap } from '@vis.gl/react-google-maps'; +import { createPortal } from 'react-dom'; + +function useCustomOverlay(position: google.maps.LatLngLiteral) { + const map = useMap(); + const [container, setContainer] = useState(null); + + useEffect(() => { + if (!map) return; + + class CustomOverlay extends google.maps.OverlayView { + private div: HTMLDivElement | null = null; + private position: google.maps.LatLng; + + constructor(position: google.maps.LatLngLiteral) { + super(); + this.position = new google.maps.LatLng(position.lat, position.lng); + } + + onAdd() { + this.div = document.createElement('div'); + this.div.style.position = 'absolute'; + const panes = this.getPanes(); + panes?.overlayMouseTarget.appendChild(this.div); + setContainer(this.div); + } + + draw() { + if (!this.div) return; + const projection = this.getProjection(); + const point = projection.fromLatLngToDivPixel(this.position); + if (point) { + this.div.style.left = `${point.x}px`; + this.div.style.top = `${point.y}px`; + } + } + + onRemove() { + this.div?.remove(); + setContainer(null); + } + } + + const overlay = new CustomOverlay(position); + overlay.setMap(map); + + return () => overlay.setMap(null); + }, [map, position]); + + return container; +} + +// Usage +function CustomMarker({ position, children }: { position: google.maps.LatLngLiteral; children: React.ReactNode }) { + const container = useCustomOverlay(position); + + if (!container) return null; + + return createPortal(children, container); +} +``` + +--- + +## 3D Maps (Experimental) + +```tsx +function Map3D() { + useMapsLibrary('maps3d'); + + return ( + + ); +} +``` + +Requires `@types/google.maps` and proper API configuration. + +--- + +## Error Boundaries + +Wrap map components to handle errors gracefully: + +```tsx +import { ErrorBoundary } from 'react-error-boundary'; + +function MapErrorFallback({ error }: { error: Error }) { + return ( +
+

Map failed to load

+

{error.message}

+
+ ); +} + +function App() { + return ( + + + + + + ); +} +``` + +--- + +## Next.js Considerations + +### Client Component + +Maps must be client components: + +```tsx +'use client'; + +import { APIProvider, Map } from '@vis.gl/react-google-maps'; + +export function MapComponent() { + return ( + + + + ); +} +``` + +### Dynamic Import (Avoid SSR Issues) + +```tsx +import dynamic from 'next/dynamic'; + +const MapComponent = dynamic( + () => import('@/components/map').then((mod) => mod.MapComponent), + { ssr: false, loading: () =>
Loading map...
} +); +``` + +### Environment Variables + +Use `NEXT_PUBLIC_` prefix for client-side access: + +```env +NEXT_PUBLIC_GOOGLE_MAPS_API_KEY=your_api_key +NEXT_PUBLIC_GOOGLE_MAPS_MAP_ID=your_map_id +``` diff --git a/.claude/skills/maps/references/react-vis-gl/places-autocomplete.md b/.claude/skills/maps/references/react-vis-gl/places-autocomplete.md new file mode 100644 index 00000000..a85b619d --- /dev/null +++ b/.claude/skills/maps/references/react-vis-gl/places-autocomplete.md @@ -0,0 +1,476 @@ +# Places Autocomplete & Geocoding + +## Important: Deprecation Notice + +As of **March 1st, 2025**, `google.maps.places.Autocomplete` (the widget) is **not available to new customers**. Use `AutocompleteService` with a custom UI instead. + +--- + +## Custom Places Autocomplete (Recommended) + +Build your own autocomplete UI using `AutocompleteService`: + +```tsx +// components/place-autocomplete.tsx +import { useState, useEffect, useCallback, useRef } from 'react'; +import { useMapsLibrary } from '@vis.gl/react-google-maps'; + +interface PlaceAutocompleteProps { + onPlaceSelect: (place: google.maps.places.PlaceResult | null) => void; + placeholder?: string; +} + +export function PlaceAutocomplete({ onPlaceSelect, placeholder = 'Search places...' }: PlaceAutocompleteProps) { + const placesLib = useMapsLibrary('places'); + const [inputValue, setInputValue] = useState(''); + const [suggestions, setSuggestions] = useState([]); + const [isOpen, setIsOpen] = useState(false); + + const autocompleteService = useRef(null); + const placesService = useRef(null); + const sessionToken = useRef(null); + + // Initialize services + useEffect(() => { + if (!placesLib) return; + + autocompleteService.current = new placesLib.AutocompleteService(); + // PlacesService needs an element or map - use a dummy div + const div = document.createElement('div'); + placesService.current = new placesLib.PlacesService(div); + sessionToken.current = new placesLib.AutocompleteSessionToken(); + }, [placesLib]); + + // Fetch suggestions + const fetchSuggestions = useCallback(async (input: string) => { + if (!autocompleteService.current || !input.trim()) { + setSuggestions([]); + return; + } + + try { + const response = await autocompleteService.current.getPlacePredictions({ + input, + sessionToken: sessionToken.current!, + // Optional: bias results + // componentRestrictions: { country: 'us' }, + // types: ['address'], // or 'establishment', 'geocode', etc. + }); + + setSuggestions(response.predictions || []); + setIsOpen(true); + } catch (error) { + console.error('Autocomplete error:', error); + setSuggestions([]); + } + }, []); + + // Debounce input + useEffect(() => { + const timer = setTimeout(() => { + fetchSuggestions(inputValue); + }, 300); + + return () => clearTimeout(timer); + }, [inputValue, fetchSuggestions]); + + // Handle suggestion selection + const handleSelect = useCallback((prediction: google.maps.places.AutocompletePrediction) => { + if (!placesService.current || !placesLib) return; + + // Get full place details + placesService.current.getDetails( + { + placeId: prediction.place_id, + fields: ['geometry', 'name', 'formatted_address', 'place_id'], + sessionToken: sessionToken.current!, + }, + (place, status) => { + if (status === google.maps.places.PlacesServiceStatus.OK && place) { + onPlaceSelect(place); + setInputValue(place.formatted_address || place.name || ''); + setIsOpen(false); + // Reset session token after selection + sessionToken.current = new placesLib.AutocompleteSessionToken(); + } + } + ); + }, [placesLib, onPlaceSelect]); + + return ( +
+ setInputValue(e.target.value)} + onFocus={() => suggestions.length > 0 && setIsOpen(true)} + onBlur={() => setTimeout(() => setIsOpen(false), 200)} + placeholder={placeholder} + className="w-full px-4 py-2 border rounded-lg" + /> + + {isOpen && suggestions.length > 0 && ( +
    + {suggestions.map((suggestion) => ( +
  • handleSelect(suggestion)} + className="px-4 py-2 cursor-pointer hover:bg-gray-100" + > + + {suggestion.structured_formatting.main_text} + + + {suggestion.structured_formatting.secondary_text} + +
  • + ))} +
+ )} +
+ ); +} +``` + +### Usage with Map + +```tsx +function MapWithSearch() { + const [selectedPlace, setSelectedPlace] = useState(null); + const [markerRef, marker] = useAdvancedMarkerRef(); + + const position = selectedPlace?.geometry?.location + ? { lat: selectedPlace.geometry.location.lat(), lng: selectedPlace.geometry.location.lng() } + : null; + + return ( + +
+ + + + {position && ( + + + + )} + +
+
+ ); +} +``` + +--- + +## Autocomplete with MapControl + +Place the autocomplete inside the map: + +```tsx +import { MapControl, ControlPosition } from '@vis.gl/react-google-maps'; + +function MapWithEmbeddedSearch() { + const map = useMap(); + const [selectedPlace, setSelectedPlace] = useState(null); + + // Pan to selected place + useEffect(() => { + if (!map || !selectedPlace?.geometry?.location) return; + map.panTo(selectedPlace.geometry.location); + map.setZoom(15); + }, [map, selectedPlace]); + + return ( + + +
+ +
+
+ + {selectedPlace?.geometry?.location && ( + + )} +
+ ); +} +``` + +--- + +## Geocoding + +Convert addresses to coordinates and vice versa. + +### Address to Coordinates (Geocoding) + +```tsx +import { useMapsLibrary } from '@vis.gl/react-google-maps'; + +function useGeocoder() { + const geocodingLib = useMapsLibrary('geocoding'); + const [geocoder, setGeocoder] = useState(null); + + useEffect(() => { + if (!geocodingLib) return; + setGeocoder(new geocodingLib.Geocoder()); + }, [geocodingLib]); + + const geocode = useCallback(async (address: string): Promise => { + if (!geocoder) return null; + + try { + const response = await geocoder.geocode({ address }); + + if (response.results.length > 0) { + const location = response.results[0].geometry.location; + return { lat: location.lat(), lng: location.lng() }; + } + } catch (error) { + console.error('Geocoding error:', error); + } + + return null; + }, [geocoder]); + + return { geocode, isReady: !!geocoder }; +} + +// Usage +function AddressSearch() { + const { geocode, isReady } = useGeocoder(); + const [address, setAddress] = useState(''); + const [result, setResult] = useState(null); + + const handleSearch = async () => { + const coords = await geocode(address); + if (coords) setResult(coords); + }; + + return ( +
+ setAddress(e.target.value)} /> + + {result &&

Lat: {result.lat}, Lng: {result.lng}

} +
+ ); +} +``` + +### Coordinates to Address (Reverse Geocoding) + +```tsx +function useReverseGeocoder() { + const geocodingLib = useMapsLibrary('geocoding'); + const [geocoder, setGeocoder] = useState(null); + + useEffect(() => { + if (!geocodingLib) return; + setGeocoder(new geocodingLib.Geocoder()); + }, [geocodingLib]); + + const reverseGeocode = useCallback(async ( + latLng: google.maps.LatLngLiteral + ): Promise => { + if (!geocoder) return null; + + try { + const response = await geocoder.geocode({ location: latLng }); + + if (response.results.length > 0) { + return response.results[0].formatted_address; + } + } catch (error) { + console.error('Reverse geocoding error:', error); + } + + return null; + }, [geocoder]); + + return { reverseGeocode, isReady: !!geocoder }; +} + +// Usage: Get address when clicking on map +function ClickToAddress() { + const { reverseGeocode, isReady } = useReverseGeocoder(); + const [address, setAddress] = useState(null); + + const handleMapClick = async (e: google.maps.MapMouseEvent) => { + if (!e.latLng || !isReady) return; + + const result = await reverseGeocode({ + lat: e.latLng.lat(), + lng: e.latLng.lng(), + }); + + setAddress(result); + }; + + return ( + <> + + {address &&

Address: {address}

} + + ); +} +``` + +--- + +## Places Details + +Get detailed information about a place: + +```tsx +function usePlaceDetails() { + const placesLib = useMapsLibrary('places'); + const serviceRef = useRef(null); + + useEffect(() => { + if (!placesLib) return; + const div = document.createElement('div'); + serviceRef.current = new placesLib.PlacesService(div); + }, [placesLib]); + + const getPlaceDetails = useCallback(async ( + placeId: string, + fields: string[] = ['name', 'formatted_address', 'geometry', 'photos', 'rating', 'reviews'] + ): Promise => { + if (!serviceRef.current) return null; + + return new Promise((resolve) => { + serviceRef.current!.getDetails( + { placeId, fields }, + (place, status) => { + if (status === google.maps.places.PlacesServiceStatus.OK) { + resolve(place); + } else { + resolve(null); + } + } + ); + }); + }, []); + + return { getPlaceDetails, isReady: !!serviceRef.current }; +} +``` + +--- + +## Nearby Search + +Find places near a location: + +```tsx +function useNearbySearch() { + const placesLib = useMapsLibrary('places'); + const map = useMap(); + const serviceRef = useRef(null); + + useEffect(() => { + if (!placesLib || !map) return; + serviceRef.current = new placesLib.PlacesService(map); + }, [placesLib, map]); + + const searchNearby = useCallback(async ( + location: google.maps.LatLngLiteral, + radius: number, + type?: string + ): Promise => { + if (!serviceRef.current) return []; + + return new Promise((resolve) => { + serviceRef.current!.nearbySearch( + { + location, + radius, + type: type as string, + }, + (results, status) => { + if (status === google.maps.places.PlacesServiceStatus.OK && results) { + resolve(results); + } else { + resolve([]); + } + } + ); + }); + }, []); + + return { searchNearby, isReady: !!serviceRef.current }; +} + +// Usage +function NearbyRestaurants() { + const { searchNearby, isReady } = useNearbySearch(); + const [restaurants, setRestaurants] = useState([]); + + const findRestaurants = async () => { + const results = await searchNearby( + { lat: 40.7128, lng: -74.006 }, + 1500, // 1.5km radius + 'restaurant' + ); + setRestaurants(results); + }; + + return ( + + ); +} +``` + +--- + +## Autocomplete Options Reference + +```tsx +const autocompleteOptions: google.maps.places.AutocompletionRequest = { + input: 'pizza', + + // Bias to a location + locationBias: { + center: { lat: 40.7128, lng: -74.006 }, + radius: 5000, + }, + + // Or restrict to bounds + locationRestriction: { + east: -73.9, + west: -74.1, + north: 40.8, + south: 40.6, + }, + + // Restrict to countries + componentRestrictions: { country: ['us', 'ca'] }, + + // Filter by type + types: ['address'], // 'establishment', 'geocode', '(cities)', '(regions)' + + // Session token (for billing) + sessionToken: new google.maps.places.AutocompleteSessionToken(), +}; +``` + +### Place Types + +- `'address'` - Street addresses +- `'establishment'` - Businesses +- `'geocode'` - Geographic areas +- `'(cities)'` - Cities only +- `'(regions)'` - Administrative regions diff --git a/.claude/skills/maps/references/security-and-optimization.md b/.claude/skills/maps/references/security-and-optimization.md new file mode 100644 index 00000000..c7ff3e11 --- /dev/null +++ b/.claude/skills/maps/references/security-and-optimization.md @@ -0,0 +1,266 @@ +--- +title: Maps Platform — API key security + HTML pages + optimization +--- + +# Maps Platform — Security, HTML Pages & Optimization + +Official: https://developers.google.com/maps/api-security-best-practices +Optimization: https://developers.google.com/maps/optimization-guide +Coverage: https://developers.google.com/maps/coverage + +--- + +## API key architecture — three deployment modes + +### Mode 1: Personal / CLI use + +API key in `.env`, used locally. Key never leaves the machine. + +``` +.env (GOOGLE_MAPS_API_KEY=...) +gmaps.py → calls Google APIs directly +HTML pages → zero-key embed iframes (no key in HTML) +``` + +**Risk:** Low — it's the user's own key on their own machine. +**Best practice:** Even personally, prefer zero-key `output=embed` iframes for HTML. Only use Maps JS API when advanced features (custom markers, polylines, clustering) are needed. + +### Mode 2: Users bring their own key (BYOK) + +Users configure their own Google API key. They control their own billing. + +- Store key in user profile (encrypted at rest) +- Use server-side for data queries +- For shareable/exported HTML: generate static exports with **zero keys in HTML** +- Never embed a user's key in a downloadable file + +### Mode 3: Platform key — you pay, users must never see it + +You provide the Google API key. Architecture requires two separate keys. + +``` +Browser +├─ Interactive map rendering ← Frontend Key (HTTP referrer restricted) +└─ All data (weather, places, directions) + ↑ pre-rendered from backend — NO API key in browser + +Your Backend Server +├─ Backend Key (env var, IP restricted — never sent to browser) +├─ Geocoding, Directions, Places, Weather, etc. +└─ All data APIs proxied server-side +``` + +**Frontend key** (client-side, in the HTML): +- Enables: **Maps JavaScript API only** +- Restricted by: **HTTP referrer** → `https://app.yourdomain.com/*` +- If copied by someone: only works from your domain, useless elsewhere + +**Backend key** (server-side, hidden): +- Enables: all data APIs (geocoding, directions, places, weather, etc.) +- Restricted by: **Server IP address** +- Never sent to the browser + +--- + +## Setting up two keys in GCP Console + +### Frontend key + +1. **APIs & Services → Credentials → Create Credentials → API Key** +2. Edit key → Application restrictions: **HTTP referrers** +3. Add: `https://app.yourdomain.com/*` +4. API restrictions: **Restrict key** → enable **Maps JavaScript API only** +5. Save + +### Backend key + +1. Create another API key +2. Application restrictions: **IP addresses** → add your server IP(s) +3. API restrictions: enable all data APIs you use: + - Geocoding, Routes, Places (New), Weather, Air Quality, Pollen + - Solar, Elevation, Time Zone, Address Validation, Roads + - Street View Static, Geolocation, Aerial View, Route Optimization +4. Save + +```bash +# Your server environment +GOOGLE_MAPS_BACKEND_KEY=AIzaSy...xxx # IP-restricted, server only +GOOGLE_MAPS_FRONTEND_KEY=AIzaSy...yyy # Domain-restricted, in HTML +``` + +--- + +## HTML pages — zero-key embed iframes (default) + +**Always default to zero-key embed iframes for HTML maps.** No API key in HTML, free, unlimited. + +```html + + + + + +``` + +**Parameters:** +- `q` — place name or address (URL-encoded, `+` for spaces) +- `saddr` / `daddr` — origin/destination for directions +- `z` — zoom level 1–20 +- `output=embed` — required +- `ll` — optional center coordinates + +**WARNING:** Never use `loading="lazy"` on Google Maps embed iframes — it causes maps below the fold to appear permanently blank. + +Only use `