Skip to content

feat: add fleet capacity dashboard - #5

Merged
purple-phoenix merged 21 commits into
mainfrom
fm/capacity-build-capacity-pipeline-dashboard-skill-ce
Jul 28, 2026
Merged

purple-phoenix merged 21 commits into
mainfrom
fm/capacity-build-capacity-pipeline-dashboard-skill-ce

Conversation

@purple-phoenix

Copy link
Copy Markdown
Owner

Intent

Create a discoverable captain-invocable /capacity skill for Firstmate that gathers a fresh bounded provenance-aware fleet snapshot across the main and registered secondmate homes using authoritative state owners; classifies meaningful ready-work supply, independent starts, delivery-stage flow, blockers, scope alignment, dispatch/runtime/auth availability, backlog definition health, and aging without inventing quota, busywork, utilization targets, self-scheduling, or unsafe concurrency; and renders a polished accessible responsive self-contained offline HTML dashboard under gitignored data/ with stable CAP action IDs and copyable captain prompts. Keep one complete procedure owner, minimal AGENTS.md trigger and safety stub, README discovery, and exact mechanics in script help or headers. The invocation is read-mostly: no Lavish, secrets, PHI, services, dispatch, merge, teardown, task mutation, or speculative work, and approved actions must re-enter normal Firstmate lifecycles. Follow existing bearings, fleet-state, backlog, secondmate, dispatch-profile, harness/backend, approval, overlap, project-write, merge, security, and teardown contracts; add focused deterministic tests and browser validation; preserve tracked Markdown style; push only to purple-phoenix/firstmate and never merge. Preserve the approved dashboard-output symlink hardening while allowing legitimate system symlink ancestors, bounded secondmate probes, per-home credential/runtime/dispatch attribution for ready and active work, cross-home overlap and unavailable-state semantics, project-registry resolution evidence, ShellCheck-clean tests, and every prior pipeline fix commit. Integrate origin/main d225b87 containing the pre-existing main-suite test fixes. The captain verified the xoxb-shaped privacy fixture is fabricated test data and explicitly approved GitHub push-protection allowance as used-in-tests; do not rewrite history.

What Changed

  • Add the captain-invocable /capacity skill and producer for provenance-aware fleet classification, stable CAP-NN recommendations, JSON output, and a private responsive offline dashboard.
  • Extend fleet snapshots with project-resolution, delivery-mode, secondmate scope, runtime, credential, dispatch, overlap, and aging evidence while preserving bounded probes and hardened output paths.
  • Document the new workflow and add deterministic coverage for classification, unavailable-state handling, cross-home attribution, dashboard rendering, accessibility, and security boundaries.

Risk Assessment

✅ Low: The follow-up removes the premature aggregate snapshot timeout and documents the required bounds, probe mechanics, action meanings, and priority order without introducing a material new risk.

Testing

The previously successful full baseline was supplemented by focused capacity and fleet-snapshot suites, a fresh end-to-end /capacity production run, desktop visual capture, mobile responsive/offline checks, and copy-prompt interaction testing; everything passed and no actionable issues remain.

  • Evidence: Rendered capacity dashboard (local file: /var/folders/g0/z_4x96f92cgfpm7940cqvtt00000gn/T/no-mistakes-evidence/01KYK7YWRRN3X6T8VQZJG9A395/capacity-dashboard-desktop.png)
Evidence: Self-contained offline dashboard
<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width, initial-scale=1">
  <meta name="color-scheme" content="dark light">
  <title>Firstmate capacity dashboard</title>
  <style>
    :root{--ink:#e9eef8;--muted:#9cabbe;--panel:#111923;--panel2:#172231;--line:#2b3b50;--accent:#61d7c2;--accent2:#f1bd66;--danger:#ff8f8f;--bg:#081019;--shadow:0 16px 40px rgba(0,0,0,.24);font-family:Inter,ui-sans-serif,system-ui,-apple-system,BlinkMacSystemFont,"Segoe UI",sans-serif;color-scheme:dark}
    *{box-sizing:border-box}html{background:var(--bg);scroll-behavior:smooth}body{margin:0;color:var(--ink);background:radial-gradient(circle at 15% 0%,#12303d 0,transparent 34rem),var(--bg);line-height:1.5;overflow-wrap:anywhere}a{color:#8ae6d6;text-underline-offset:.18em}button,a{outline-offset:3px}button:focus-visible,a:focus-visible{outline:3px solid var(--accent2)}.skip{position:absolute;left:-9999px}.skip:focus{left:1rem;top:1rem;z-index:10;background:#fff;color:#000;padding:.7rem 1rem;border-radius:.5rem}.shell{width:min(1560px,100%);margin:auto;padding:clamp(1rem,3vw,3rem)}.hero{display:grid;grid-template-columns:minmax(0,1.6fr) minmax(18rem,.7fr);gap:1.5rem;align-items:end;padding:clamp(1.5rem,4vw,3.5rem) 0 2rem}.eyebrow{margin:0 0 .5rem;color:var(--accent);font-weight:800;letter-spacing:.12em;text-transform:uppercase;font-size:.78rem}.hero h1{font-size:clamp(2.2rem,6vw,5.4rem);line-height:.95;letter-spacing:-.055em;margin:0;max-width:12ch}.hero p{color:var(--muted);max-width:70ch}.freshness{border:1px solid var(--line);background:rgba(17,25,35,.82);padding:1rem;border-radius:1rem;box-shadow:var(--shadow)}.freshness strong,.freshness span{display:block}.freshness span{color:var(--muted);font-size:.9rem;margin-top:.35rem}.metrics{display:grid;grid-template-columns:repeat(5,minmax(0,1fr));gap:.85rem;margin:1rem 0 2.5rem}.metric{min-width:0;background:linear-gradient(145deg,var(--panel2),var(--panel));border:1px solid var(--line);border-radius:1rem;padding:1rem}.metric span,.metric small{display:block;color:var(--muted)}.metric strong{display:block;font-size:clamp(1.4rem,2.6vw,2.4rem);line-height:1.1;margin:.4rem 0}.section-title{display:flex;justify-content:space-between;gap:1rem;align-items:end;margin:3.5rem 0 1rem}.section-title h2{font-size:clamp(1.6rem,3vw,2.5rem);margin:0}.section-title p{color:var(--muted);max-width:65ch;margin:0}.pipeline{display:grid;grid-template-columns:repeat(4,minmax(0,1fr));gap:1rem;align-items:start}.stage{min-width:0;border:1px solid var(--line);background:rgba(12,20,29,.78);border-radius:1rem;padding:.85rem}.stage>header{display:flex;align-items:center;justify-content:space-between;gap:.75rem;border-bottom:1px solid var(--line);padding:.25rem .2rem .75rem}.stage>header h2{font-size:1rem;margin:0}.stage>header span,.action-id{background:#21443f;color:#a9f7e8;border-radius:999px;padding:.16rem .58rem;font-weight:800;font-size:.78rem}.stage-cards{display:grid;gap:.7rem;margin-top:.7rem}.work-card,.lane-card,.recommendation{min-width:0;background:var(--panel);border:1px solid var(--line);border-radius:.8rem;padding:.85rem}.work-card h3,.lane-card h3,.recommendation h3{font-size:1rem;margin:.55rem 0}.work-card p,.lane-card p{font-size:.88rem;color:var(--muted);margin:.4rem 0}.card-top,.rec-head{display:flex;align-items:center;justify-content:space-between;gap:.6rem}.item-id{font:700 .76rem ui-monospace,SFMono-Regular,Menlo,monospace;color:var(--accent)}.owner,.status{font-size:.72rem;color:var(--muted);text-align:right}.status{border:1px solid var(--line);padding:.14rem .45rem;border-radius:999px}.status.healthy{color:var(--accent);border-color:#2c675d}.status.warning{color:var(--danger);border-color:#70414a}.work-card dl,.recommendation dl{margin:.6rem 0 0}.work-card dl div,.recommendation dl div{margin-top:.45rem}.work-card dt,.recommendation dt{color:var(--muted);font-size:.72rem;text-transform:uppercase;letter-spacing:.06em}.work-card dd,.recommendation dd{margin:0;font-size:.86rem}.artifact,.path{font:500 .75rem ui-monospace,SFMono-Regular,Menlo,monospace;margin-top:.65rem;max-width:100%;overflow-wrap:anywhere}.empty{color:var(--muted);font-style:italic}.lanes{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:1rem}.lane-group{min-width:0;border:1px solid var(--line);border-radius:1rem;background:rgba(12,20,29,.75);padding:1rem}.lane-group>h3{margin:.2rem 0 1rem}.lane-grid{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:.75rem}.recommendations{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:1rem}.recommendation{border-left:4px solid var(--accent2);padding:1.1rem}.rec-head{justify-content:flex-start}.rec-head h3{font-size:1.2rem}.recommendation dl div{display:grid;grid-template-columns:minmax(9rem,.34fr) minmax(0,1fr);gap:1rem;border-top:1px solid var(--line);padding-top:.65rem}.prompt{display:grid;grid-template-columns:minmax(0,1fr) auto;gap:.6rem;align-items:center;margin-top:1rem;background:#09121b;border:1px solid var(--line);border-radius:.65rem;padding:.65rem}.prompt code{font-size:.78rem;white-space:normal}.prompt button{border:0;border-radius:.5rem;background:var(--accent);color:#06241f;font-weight:800;padding:.55rem .75rem;cursor:pointer}.health-grid{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:1rem}.list-panel{background:var(--panel);border:1px solid var(--line);border-radius:1rem;padding:1rem;min-width:0}.list-panel h3{margin-top:0}.clean-list{list-style:none;margin:0;padding:0;display:grid;gap:.7rem}.clean-list li{display:grid;grid-template-columns:minmax(8rem,.35fr) minmax(0,1fr);gap:.75rem;border-top:1px solid var(--line);padding-top:.65rem;min-width:0}.clean-list li>*{min-width:0}.clean-list span{color:var(--muted)}footer{margin-top:4rem;padding:1.5rem 0;color:var(--muted);border-top:1px solid var(--line);font-size:.85rem}
    @media(max-width:1100px){.metrics{grid-template-columns:repeat(3,minmax(0,1fr))}.pipeline{grid-template-columns:repeat(2,minmax(0,1fr))}.recommendations{grid-template-columns:1fr}}
    @media(max-width:760px){.shell{padding:1rem}.hero{grid-template-columns:1fr}.metrics,.pipeline,.lanes,.health-grid,.lane-grid{grid-template-columns:1fr}.section-title{display:block}.section-title p{margin-top:.4rem}.recommendation dl div,.clean-list li{grid-template-columns:1fr;gap:.15rem}.prompt{grid-template-columns:1fr}.prompt button{width:100%}.hero h1{font-size:clamp(2.4rem,15vw,4rem)}}
    @media(prefers-reduced-motion:reduce){html{scroll-behavior:auto}}
    @media print{body{background:#fff;color:#111}.shell{width:100%;padding:0}.work-card,.lane-card,.recommendation,.stage,.lane-group,.list-panel,.metric,.freshness{box-shadow:none;background:#fff;color:#111;border-color:#bbb}.pipeline{grid-template-columns:repeat(2,minmax(0,1fr))}.prompt button{display:none}a{color:#0645ad}}
  </style>
</head>
<body>
  <a class="skip" href="#main">Skip to dashboard</a>
  <div class="shell">
    <header class="hero">
      <div><p class="eyebrow">Firstmate / meaningful throughput</p><h1>Capacity, without busywork.</h1><p>This view identifies what can safely flow now, what is actually holding delivery back, and where healthy idle should stay idle.</p></div>
      <aside class="freshness" aria-label="Snapshot freshness"><strong>Generated 2026-07-28T02:41:11Z</strong><span>Fresh command observation on each normal invocation; status-log tails are historical only and never current-state authority.</span><span>Validated structured-home summaries with registered-table route metadata; fallback parent events never override readable home state.</span></aside>
    </header>
    <main id="main">
      <section aria-label="Top capacity measures" class="metrics"><article class="metric"><span>Useful ready supply</span><strong>0</strong><small>Grounded and independently startable</small></article><article class="metric"><span>Active flow</span><strong>0</strong><small>Ephemeral and secondmate child work</small></article><article class="metric"><span>Waiting work</span><strong>0</strong><small>Queued, gated, or at delivery gates</small></article><article class="metric"><span>Captain actions</span><strong>0</strong><small>Structured decisions and ready approvals</small></article><article class="metric"><span>Primary bottleneck</span><strong>unavailable state</strong><small>CAP-03</small></article></section>
      <div class="section-title"><h2>Delivery pipeline</h2><p>One current card per item, classified from authoritative current state and structured backlog evidence.</p></div>
      <div class="pipeline"><section class="stage" aria-labelledby="stage-queued">
    <header><h2 id="stage-queued">Queued</h2><span>0</span></header>
    <div class="stage-cards"><p class="empty">No current items.</p></div>
  </section><section class="stage" aria-labelledby="stage-ready">
    <header><h2 id="stage-ready">Ready</h2><span>0</span></header>
    <div class="stage-cards"><p class="empty">No current items.</p></div>
  </section><section class="stage" aria-labelledby="stage-building">
    <header><h2 id="stage-building">Building</h2><span>0</span></header>
    <div class="stage-cards"><p class="empty">No current items.</p></div>
  </section><section class="stage" aria-labelledby="stage-validating_fixing">
    <header><h2 id="stage-validating_fixing">Validating / fixing</h2><span>0</span></header>
    <div class="stage-cards"><p class="empty">No current items.</p></div>
  </section><section class="stage" aria-labelledby="stage-pr_ci_approval">
    <header><h2 id="stage-pr_ci_approval">PR / CI / approval</h2><span>0</span></header>
    <div class="stage-cards"><p class="empty">No current items.</p></div>
  </section><section class="stage" aria-labelledby="stage-blocked">
    <header><h2 id="stage-blocked">Blocked</h2><span>0</span></header>
    <div class="stage-cards"><p class="empty">No current items.</p></div>
  </section><section class="stage" aria-labelledby="stage-recently_landed">
    <header><h2 id="stage-recently_landed">Recently landed</h2><span>0</span></header>
    <div class="stage-cards"><p class="empty">No current items.</p></div>
  </section></div>
      <div class="section-title"><h2>Lane and scope alignment</h2><p>Ephemeral workers are created on demand. Persistent secondmates are healthy when idle unless grounded in-scope work is already ready.</p></div>
      <section class="lanes" aria-label="Lane utilization">
        <div class="lane-group"><h3>Ephemeral dispatch lanes</h3><p>unbounded on demand - Firstmate has no fixed ephemeral pool or concurrency target</p><p>Runtime backend: <strong>tmux</strong> - required runtime surface available. GitHub auth: available.</p><div class="lane-grid"><article class="lane-card">
    <div class="card-top"><span class="item-id">codex</span><span class="status healthy">available</span></div>
    <h3>configured dispatch rule</h3>
    <p>configured executable present</p>
    <small>not observed - capacity never guesses quota</small>
  </article></div></div>
        <div class="lane-group"><h3>Persistent secondmates</h3><div class="lane-grid"><p class="empty">No persistent secondmates are registered.</p></div></div>
      </section>
      <div class="section-title"><h2>Capacity recommendations</h2><p>Stable action IDs are discussion handles. Copying or approving a prompt re-enters normal Firstmate lifecycles in chat; this page executes nothing.</p></div>
      <section class="recommendations" aria-label="Prioritized capacity recommendations"><article class="recommendation" id="CAP-03">
    <div class="rec-head"><span class="action-id">CAP-03</span><h3>unavailable state</h3></div>
    <dl>
      <div><dt>Evidence</dt><dd>1 current-state surface is unavailable: main/main-backlog.</dd></div>
      <div><dt>Throughput consequence</dt><dd>Unknown state prevents safe overlap and dispatch decisions, reducing knowable throughput.</dd></div>
      <div><dt>Safety and authority</dt><dd>Recovery must use the normal backend and stuck-crewmate or secondmate lifecycle; no endpoint is restarted here.</dd></div>
      <div><dt>Next action</dt><dd>Reconcile the unavailable state owners before treating apparent idle capacity as real.</dd></div>
    </dl>
    <div class="prompt"><code>Approve CAP-03: reconcile the unavailable fleet state through the normal recovery lifecycle, then rerun /capacity.</code><button type="button" data-copy="Approve CAP-03: reconcile the unavailable fleet state through the normal recovery lifecycle, then rerun /capacity." aria-label="Copy CAP-03 follow-up prompt">Copy prompt</button></div>
  </article></section>
      <div class="section-title"><h2>Definition health and landed context</h2><p>Nominal queue depth is separated from dispatch-grade supply, with recent outcomes retained for context.</p></div>
      <section class="health-grid">
        <div class="list-panel"><h3>Backlog definition gaps</h3><ul class="clean-list"><li><strong>None</strong><span>No definition gaps detected in the bounded queue.</span></li></ul></div>
        <div class="list-panel"><h3>Recently landed (observed; incomplete)</h3><ul class="clean-list"><li><strong>None</strong><span>No recent completions are in the bounded baseline.</span></li></ul></div>
      </section>
    </main>
    <footer><p>Provenance: bin/fm-fleet-snapshot.sh fm-fleet-snapshot.v1. Structured backlog captain holds and keyed open-decision folds only; scout reports and visual artifacts are not scraped. Authoritative backend functions, configured dispatch profiles, executable presence, and bootstrap-equivalent GitHub auth status; quota is not observed or guessed.</p><p>This private dashboard contains bounded operational metadata only. It uses no CDN, remote asset, analytics, network service, or Lavish integration.</p></footer>
  </div>
  <script>
    document.querySelectorAll('[data-copy]').forEach((button) => button.addEventListener('click', async () => {
      try { await navigator.clipboard.writeText(button.dataset.copy); button.textContent = 'Copied'; }
      catch { button.textContent = 'Select prompt'; }
    }));
  </script>
</body>
</html>
Evidence: Fresh fm-capacity.v1 model
{
  "schema": "fm-capacity.v1",
  "generated": "2026-07-28T02:41:11Z",
  "dashboard_path": "private data/capacity-dashboard.html",
  "provenance": {
    "fleet": "bin/fm-fleet-snapshot.sh fm-fleet-snapshot.v1",
    "freshness": "Fresh command observation on each normal invocation; status-log tails are historical only and never current-state authority.",
    "secondmates": "Validated structured-home summaries with registered-table route metadata; fallback parent events never override readable home state.",
    "decisions": "Structured backlog captain holds and keyed open-decision folds only; scout reports and visual artifacts are not scraped.",
    "environment": "Authoritative backend functions, configured dispatch profiles, executable presence, and bootstrap-equivalent GitHub auth status; quota is not observed or guessed.",
    "recent_landings": "Main backlog completion evidence is incomplete; the displayed count is an observed lower bound."
  },
  "measures": {
    "useful_ready_work": 0,
    "independent_tasks_safe_to_start_now": 0,
    "active_independent_work": 0,
    "waiting_work": 0,
    "open_captain_actions": 0,
    "recently_landed": 0,
    "recently_landed_complete": false
  },
  "primary_bottleneck": {
    "id": "CAP-03",
    "classification": "unavailable state",
    "evidence": "1 current-state surface is unavailable: main/main-backlog."
  },
  "pipeline": {
    "queued": [],
    "ready": [],
    "building": [],
    "validating_fixing": [],
    "pr_ci_approval": [],
    "blocked": [],
    "recently_landed": []
  },
  "lanes": {
    "ephemeral_workers": {
      "active": 0,
      "pool": "unbounded on demand - Firstmate has no fixed ephemeral pool or concurrency target",
      "backend": {
        "name": "tmux",
        "available": true,
        "evidence": "required runtime surface available",
        "owner": "bin/fm-backend.sh"
      },
      "github_auth": {
        "status": "available",
        "evidence": "credential check passed",
        "owner": "bin/fm-bootstrap.sh"
      },
      "configured_dispatch": {
        "config_present": false,
        "valid": true,
        "reason": null,
        "lanes": [
          {
            "harness": "codex",
            "model": null,
            "effort": null,
            "when": "configured dispatch rule",
            "available": true,
            "availability_evidence": "configured executable present",
            "quota": "not observed - capacity never guesses quota"
          }
        ]
      }
    },
    "persistent_secondmates": []
  },
  "readiness": {
    "queued_considered": 0,
    "grounded_candidates": 0,
    "independent_start_count": 0,
    "available": false,
    "definition_gaps": [],
    "explicit_gates": [],
    "conservative_overlap_gates": []
  },
  "aging": [],
  "recommendations": [
    {
      "id": "CAP-03",
      "classification": "unavailable state",
      "priority": 20,
      "evidence": "1 current-state surface is unavailable: main/main-backlog.",
      "expected_throughput_consequence": "Unknown state prevents safe overlap and dispatch decisions, reducing knowable throughput.",
      "safety_authority_boundary": "Recovery must use the normal backend and stuck-crewmate or secondmate lifecycle; no endpoint is restarted here.",
      "recommended_next_action": "Reconcile the unavailable state owners before treating apparent idle capacity as real.",
      "prompt": "Approve CAP-03: reconcile the unavailable fleet state through the normal recovery lifecycle, then rerun /capacity."
    }
  ],
  "omissions": [
    "No quota inference or utilization target is computed.",
    "Natural-language secondmate scopes and project names are withheld and not machine-guessed against main-home work.",
    "Backlog bodies are used only for bounded definition checks and are never rendered.",
    "Status tails, terminal chat, scout report contents, and visual artifacts are not consulted.",
    "Main backlog completion evidence is incomplete; the displayed count is an observed lower bound."
  ]
}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed ✅
  • 🚨 bin/fm-capacity.mjs:106 - The required “bounded secondmate probes” must gather state “across the main and registered secondmate homes,” but this new 45-second wrapper timeout aborts the entire snapshot. The canonical snapshot probes up to 20 homes sequentially with an 8-second per-home bound, so six slow homes can make /capacity fail without producing unavailable-state records or a dashboard. Use an aggregate strategy that preserves bounded results instead of killing the producer.
  • ⚠️ bin/fm-capacity.mjs:61 - The intent requires “exact mechanics in script help or headers,” and the skill explicitly assigns the model schema, probe bounds, bottleneck priority, and stable action identifiers to this help/header. The current help only describes inputs and output; it does not document CAP-01–CAP-10 semantics or priority order, and the header merely claims ownership. Add the promised mechanics to the header or help.

🔧 Fix: Captain, harden capacity snapshot bounds and help contract
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"
  • Configured baseline: command -v tmux &gt;/dev/null || { echo &#34;tmux is required for e2e tests&#34; &gt;&amp;2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo &#34;== $t ==&#34;; bash &#34;$t&#34; || rc=1; done; exit &#34;$rc&#34; (reported successful)
  • bash tests/fm-capacity.test.sh
  • bash tests/fm-bearings-snapshot.test.sh
  • bin/fm-capacity.mjs --json > /var/folders/g0/z_4x96f92cgfpm7940cqvtt00000gn/T/no-mistakes-evidence/01KYK7YWRRN3X6T8VQZJG9A395/live-capacity-model.json
  • Opened data/capacity-dashboard.html directly from disk in Chrome at 1440×1100 and captured the rendered dashboard
  • Emulated a 390×844 mobile viewport in Chrome; verified scrollWidth == innerWidth, single-column metric/pipeline layouts, and zero remote resources
  • Clicked the rendered CAP-03 copy control and verified it copied the stable lifecycle prompt and changed to Copied
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…ity-build-capacity-pipeline-dashboard-skill-ce
@purple-phoenix
purple-phoenix merged commit d8403b6 into main Jul 28, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant