Skip to content

SAN-1272 — Make MDE AI coding agents choose the right skills, workflow, and checks - #45

Closed
amoai-tech wants to merge 9 commits into
mainfrom
chore/san-1272-canonical-skills
Closed

amoai-tech wants to merge 9 commits into
mainfrom
chore/san-1272-canonical-skills

Conversation

@amoai-tech

@amoai-tech amoai-tech commented Sep 15, 2026 •

Copy link
Copy Markdown
Owner

What this PR does

This PR standardizes how coding agents work on MDE AI.

Real-world example: if a developer asks, “Fix a Supabase RLS policy used by a Mastra workflow and update the approval UI,” the agent should not load every skill or guess the owner. It should route the work to the smallest correct set of skills, run the right tests, review the diff, and require stronger proof for high-risk changes.

Linear task: https://linear.app/amo100/issue/SAN-1272/mde-skills-001-build-canonical-skills-subagent-orchestration-for-mde

Evidence index: https://linear.app/amo100/document/san-1272-skills-orchestration-evidence-index-45e1d463c91a

Status at a glance

Forensic verdict: NOT READY TO MERGE yet.

The skill architecture is strong and the exact PR head passes the MDE skill/vendor validators, but the PR still has merge blockers: Floor CI is red, committed bootstrap files still contain stale pre-rename skill names, real model routing/behavior tests have not executed, and CodeRabbit skipped the PR because it contains 401 files.

Area Current status
Linear implementation tracker 88% — 28/32 core items
Exact PR head 398bf626daf35635c6b8c0ff7891e8b960497e23
MDE skill validator ✅ PASS — 27 core skills, 42 routing eval definitions validated
Vendor wrapper validator ✅ PASS — 4 wrappers, pinned local hashes verified
Broken active top-level skill symlinks ✅ 0
Lint / TypeScript / Next build ✅ PASS in Floor CI
Vitest ✅ 1244 passed / 12 skipped
Mastra gate ✅ PASS
Floor CI 🔴 FAIL — npm audit --audit-level=critical
Real routing/behavior executions 🟡 Pending — definitions are not execution proof
Independent automated review 🟠 CodeRabbit skipped — 401 files exceeds its 100-file limit

Developer workflow / user journey

flowchart TD
    U[Developer request] --> Q{Single clear owner?}
    Q -->|Yes| D[Use the directly relevant skill]
    Q -->|No / multi-system| R[using-mde-skills]
    R --> C[Classify S0-S4 and choose one primary owner]
    C -->|S0-S1| D
    C -->|S2-S4 substantial work| T[tasks orchestrator]
    T --> S[Load only affected stack/domain skills]
    D --> V[Implement smallest safe change]
    S --> V
    V --> TEST[tdd / testing]
    TEST --> CR[code-review]
    CR --> H{High risk or Done claim?}
    H -->|Yes| TV[task-verifier independent gate]
    H -->|No| PR[PR / handoff]
    TV -->|Pass| PR
    TV -->|Fail| FIX[Return to owning workflow]
    FIX --> V
Loading

Portable agent setup

flowchart LR
    REPO[MDE AI repo] --> AG[AGENTS.md\nportable project rules]
    REPO --> CL[CLAUDE.md\nClaude-specific rules]
    REPO --> SK[.claude/skills/\ncanonical skill library]

    CL --> CC[Claude Code]
    SK --> CC

    AG --> OC[OpenCode]
    SK --> OC

    AG --> CU[Cursor]
    SK --> CU

    AG --> CX[Codex / other agents]

    SK --> ROUTE[using-mde-skills]
    ROUTE --> TASKS[tasks]
    TASKS --> OWNERS[stack + domain owners]
Loading

Claude Code uses .claude/skills/ natively. OpenCode and Cursor also support Claude-compatible project skills. AGENTS.md provides the small tool-neutral bootstrap instead of maintaining four copies of the skill library.

What changed

Core workflow

  • modernized tasks, task-verifier, wireframe, and mermaid-diagrams
  • added using-mde-skills as a router only
  • kept tasks as the substantial implementation orchestrator
  • added shared prompting, outcome, subagent, and skill-authoring standards
  • added deterministic validators for skill structure, routing references, task trackers, vendor pins, and upstream drift
  • added AGENTS.md for portable repo-level guidance
  • corrected the Claude session-start architecture and removed broken active skill aliases

Canonical stack skills

  • copilotkit
  • mastra
  • supabase
  • gemini
  • stripe
  • nextjs
  • maps
  • cloudinary

Canonical domain skills currently in scope

  • events
  • real-estate

Do not create placeholder domain skills merely to fill a list. Add restaurants, venues, trips, partners, or ecommerce only when there is real project-specific workflow knowledge worth preserving.

Skills and MCPs/tools to use

Need Best owner
Ambiguous or multi-system request using-mde-skills
Substantial implementation tasks
Unknown failure/root cause systematic-debugging
Test-first regression fix tdd
Verification strategy testing
Current official API/framework research research + Context7 / official docs
Existing PR/diff review code-review
Done / merge / production claim task-verifier
Architecture or journey diagram mermaid-diagrams
GitHub PR, exact diff, CI GitHub connector
Linear scope/progress/evidence Linear connector
Current library documentation Context7
Exact local files/commands/commit proof Remote Desktop Commander

Efficiency rule: simple task → direct skill. Use the router only when ownership is unclear or multiple systems are involved. Subagents are optional accelerators, not mandatory ceremony.

Application tech stack context

This PR changes the developer-agent control plane, not the customer application runtime.

Layer Current repo
Frontend Next.js 16.2.6, React 19.2.1, CopilotKit 1.55.2, shadcn, Tailwind CSS 4, Google Maps
AI/runtime Mastra beta, AG-UI, Gemini through @ai-sdk/google 2.0.74
Data/backend Supabase JS ^2.106.1, Supabase SSR, Postgres, RLS, Next.js Route Handlers
Maps @vis.gl/react-google-maps, Google Places, MarkerClusterer
Commerce context Medusa SDK present; Stripe skill owns payment contracts but there is no direct Stripe npm dependency in this PR head
Quality TypeScript, ESLint, Vitest 4.1.6, Playwright 1.60.0, GitHub Actions

Frontend / backend / screens

No customer-facing screen implementation is changed by this PR. It changes repository instructions, skill packages, references, validators, and session bootstrap.

Current app surfaces that must continue to work after merge include /, /chat, /events, /rentals, /restaurants, /cafes, /nightlife, /venues, /host/*, /me/tickets, /trips, and /partners/*. The existing CI production build successfully discovers these routes.

Forensic audit findings

Priority Finding Why it matters Required action
🔴 Blocker Floor CI is red at npm audit --audit-level=critical A merge-ready PR cannot claim green production gates Resolve or formally isolate the critical dependency finding, then rerun Floor
🔴 Blocker AGENTS.md and CLAUDE.md at the PR head still reference mde-maps / mde-real-estate; CLAUDE.md also says events and stripe are not active even though the canonical skills now exist Fresh sessions can receive contradictory routing instructions Update both bootstraps to the canonical names and rerun validators/fresh-session proof
🔴 Blocker Real routing and vendor behavior evals have not executed JSON definitions prove test design, not agent behavior Run 10–15 real routing prompts + one real CopilotKit/Mastra/Supabase/Gemini behavior case
🟠 High CodeRabbit skipped the PR because 401 files exceeds its review limit Independent automated review coverage is incomplete Split review units or perform an independent exact-head review before merge
🟠 High Canonical maps/SKILL.md still contains stale/broken project links and legacy assumptions; validator reports active Maps warnings The active skill can direct an agent to missing paths Repair the active Maps references first; leave deep historical warning cleanup incremental
🟡 Medium 84 legacy/reference warnings remain Noise can hide new reference regressions Reduce gradually; do not block the entire core on historical mirrors
🟡 Medium mde-worktree-pr-flow/SKILL.md is 519 lines Above Anthropic’s ~500-line progressive-disclosure target Move low-frequency detail into references when next touched
🟡 Existing repo debt Floor audit reports 51 vulnerabilities, including 1 critical; PR does not modify package.json/lockfile Mostly pre-existing, but still a repository merge gate Track/fix separately if not safe to solve in this PR; do not hide the red gate

Current dependency-security blocker

Floor CI reaches lint, typecheck, production build, Vitest, and Mastra successfully, then fails on the audit gate. The audit currently reports 51 vulnerabilities: 17 low, 18 moderate, 15 high, 1 critical. The critical set includes the installed Next.js 16.2.6; npm reports a newer Next.js release outside the current declared range as a possible fix.

This PR does not modify dependency manifests, so the dependency issue should not be silently folded into the skill refactor. The efficient path is to fix it in a focused dependency/security change or explicitly establish the correct merge policy, then rerun the exact PR head.

Scores

Area Score
Architecture / ownership model 92/100
Skill organization / progressive disclosure 88/100
Vendor provenance / pinned-reference design 92/100
Portability design 88/100
Routing definition quality 88/100
Bootstrap consistency at current PR head 58/100
Real executed eval proof 35/100
Reviewability 45/100
CI / security readiness 45/100
Current merge readiness 58/100 — blocked

The Linear checklist can still be 88% implementation-complete while merge readiness is lower: implementation progress and release evidence are different measurements.

Fastest safe path to finish

flowchart TD
    A[PR #45 current head] --> B[Fix CLAUDE.md + AGENTS.md canonical names]
    B --> C[Repair active Maps broken/stale references]
    C --> D[Run exact skill + vendor validators]
    D --> E[Run 10-15 real routing prompts]
    E --> F[Run 4 vendor behavior cases]
    F --> G[Run S0-S4 fresh-session certification]
    G --> H{Independent review available?}
    H -->|PR still too large| I[Split review units or perform exact-head manual review]
    H -->|Yes| J[Resolve red Floor security gate]
    I --> J
    J --> K[All required checks green]
    K --> L[Merge]
Loading

Do not add more orchestration features before these gates pass.

Pre-merge production-readiness checklist

Already verified

  • Exact PR head identified: 398bf626daf35635c6b8c0ff7891e8b960497e23
  • validate-skills.py passes: 27 core skills / 42 routing definitions
  • validate-vendor-skills.py passes: 4 pinned vendor wrappers / local hashes verified
  • Broken active top-level skill symlinks = 0
  • Routing validator rejects missing active skill identifiers
  • Structural evals are correctly labeled definitions validated, not executed
  • Lint passes in CI
  • TypeScript passes in CI
  • Next.js production build passes in CI
  • Vitest passes: 1244 passed / 12 skipped
  • Mastra gate passes

Must pass before merge

  • Replace stale skill names in committed CLAUDE.md and AGENTS.md
  • Repair active maps skill links/legacy path assumptions that can misroute current work
  • Run 10–15 real routing prompts and record the actual result, not just the expected JSON
  • Run one real behavior case each for CopilotKit, Mastra, Supabase, and Gemini
  • Run fresh-session S0–S4 certification
  • Prove S0 avoids unnecessary orchestration
  • Prove S4 cannot self-certify and reaches task-verifier
  • Obtain independent review despite the current 401-file CodeRabbit limit
  • Resolve/triage the critical npm audit finding and make Floor CI green
  • Rerun all required checks against the final exact head SHA
  • Confirm no unrelated docs migration is included

Success criteria

The skill platform is ready when:

  1. every active routing identifier resolves to a real skill;
  2. active broken skill symlinks = 0;
  3. CLAUDE.md, AGENTS.md, routing, registry, and skill folders all use the same canonical names;
  4. pinned vendor bytes still match their reviewed hashes and upstream changes never auto-update trusted instructions;
  5. at least 10 representative real routing prompts select the expected owner/complexity without unnecessary orchestration;
  6. CopilotKit, Mastra, Supabase, and Gemini each pass one real behavior case;
  7. S0 stays direct and S4 requires independent verification;
  8. required CI is green at the final head;
  9. an independent reviewer has actually reviewed the change set;
  10. Linear evidence reflects the exact merged SHA rather than an earlier local checkpoint.

Post-merge actions

  • Fast-forward local main to the exact merge SHA
  • Rerun skill + vendor validators on main
  • Start a fresh Claude Code session and verify canonical skill discovery/routing
  • Smoke the same skill library with OpenCode and Cursor where available
  • Re-run the representative routing/behavior cases on main
  • Keep adaptive routing memory disabled until real observations have been reviewed
  • Record merge SHA + post-merge evidence in SAN-1272
  • Remove the merged feature branch only after the above proof is recorded
  • Continue reducing the 84 warning backlog separately; do not mix historical cleanup into unrelated feature work

Official references used for this review

Skill authoring / agent behavior

Official vendor skill sources pinned by this PR

MDE source of truth

Commits in this PR

  • 674c61d — modernize MDE task design and verification skills
  • ca80963 — checkpoint portable MDE skill foundation
  • 398bf62 — normalize canonical stack and domain skills

Review note on PR size

The existing commits are useful checkpoints, but simply splitting by commit is not enough: the three slices touch roughly 48, 244, and 130 files. If automated review coverage is required, split by review responsibility instead of blindly by commit — core router/bootstrap/validators, vendor snapshots/wrappers, and large domain/reference migrations. That is faster to review and safer than rewriting the whole implementation.

@coderabbitai

coderabbitai Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: cc2a7a0e-5ae8-48e7-a42c-7311d43c407d


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

Failed to generate code suggestions for PR

@amoai-tech amoai-tech changed the title SAN-1272 · MDE-SKILLS-001 — Canonical skills and portable orchestration foundation SAN-1272 — One shared MDE skill system for Claude Code, Cursor, OpenCode & Codex Sep 15, 2026
@amoai-tech amoai-tech changed the title SAN-1272 — One shared MDE skill system for Claude Code, Cursor, OpenCode & Codex SAN-1272 — Make MDE AI coding agents choose the right skills, workflow, and checks Sep 15, 2026
@codacy-production

Copy link
Copy Markdown

Not up to standards ⛔

🔴 Issues 5 critical · 38 high · 51 medium · 6 minor

Alerts:
⚠ 100 issues (≤ 0 issues of at least minor severity)

Results:
100 new issues

Category Results
UnusedCode 2 medium
BestPractice 20 medium
ErrorProne 4 high
Security 6 minor
34 high
5 critical
14 medium
Complexity 15 medium

View in Codacy

🟢 Metrics 421 complexity · 6 duplication

Metric Results
Complexity 421
Duplication 6

View in Codacy

AI Reviewer: first review requested successfully. AI can make mistakes. Always validate suggestions.

Run reviewer

TIP This summary will be updated as you push new changes.

Copy link
Copy Markdown
Owner Author

Superseded by the staged SAN-1274/SAN-1273 landing sequence: PRs #47, #48, #49, #50 and #52.

The useful canonical skills, bootstrap, security, routing, and live-certification work has been extracted into smaller independently verified PRs.

PR #45 should not be merged because it contains the older large orchestration design and would reintroduce superseded/conflicting files.

The PR #45 branch/history is being retained temporarily as forensic/recovery evidence; local historical branches are not being deleted as part of this closure.

@amoai-tech amoai-tech closed this Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants