Skip to content

docs: gsd-method revision 2 from the Blockvalley Phase 1 retro - #92

Merged
withally merged 2 commits into
mainfrom
fm/fm-gsd-method-doc-r2
Aug 30, 2026
Merged

withally merged 2 commits into
mainfrom
fm/fm-gsd-method-doc-r2

Conversation

@withally

@withally withally commented Aug 30, 2026 •

Copy link
Copy Markdown
Owner

Intent

Revise docs/gsd-method.md as revision 2 from section 3 of the Blockvalley Phase 1 retro. Fold in five concrete rules, each with a one-line why: a one-page phase-plan slice contract with bounded task and acceptance-criterion counts plus a mandatory deletion pass before CHECK; a signed kept/changed/parked scope-delta artifact acknowledged by every live worker; a code-level Don’t-Hand-Roll audit plus independent critic before local-main landing of a proof slice and never deferred; an on-phone capture manifest with element IDs, camera bounds, deterministic frame, and per-criterion evidence links; and check-in cost accounting for model, rounds, and wall time. Keep the existing provenance-tag, CONTEXT, Don’t-Hand-Roll, and plan-check structure without restyling it. Touch docs/gsd-method.md only unless a pointer is strictly needed, use one sentence per Markdown line and plain dashes, keep net growth under about 60 lines, copy nothing private, and cite the source only as Blockvalley Phase 1 retro. Validate with bin/fm-lint.sh and bin/fm-doc-audience-check.sh. Attest and update existing PR 92 titled docs: gsd-method revision 2 from the Blockvalley Phase 1 retro; do not open a second PR.

What Changed

  • Added a one-page phase-plan slice contract with numeric task and acceptance-criterion caps, a mandatory pre-CHECK deletion pass, and signed kept/changed/parked scope-delta reconciliation with live-worker acknowledgments.
  • Added a code-level Don't-Hand-Roll audit and independent critic gate before local-main proof-slice landing, plus an on-phone capture manifest covering element IDs, camera bounds, deterministic state and frame, and per-criterion evidence links.
  • Added check-in accounting for model, rounds, and wall time, and marked the document as revision 2 with a dated Blockvalley Phase 1 retro log entry and stable-name/path evidence citations.

Risk Assessment

✅ Low: The documentation-only change satisfies the stated requirements and preserves the corrected heading and named-proof success signal.

Testing

Focused documentation-audience behavior, Markdown rendering, visual presentation, whitespace, and worktree custody all passed. Evidence was written to the permitted evidence directory. fm-lint.sh and other lint/static-analysis phases were intentionally left to the outer executor, while the audience checker was exercised through its focused test.

Evidence: Rendered overview

Source: Rendered overview

Rendered target Markdown and visually reviewed the revised document surface.
Evidence: Rendered rules

Source: Rendered rules

Shows phase-plan budget and scope-cut reconciliation rules.
Evidence: Rendered proof and cost

Source: Rendered proof and cost

Shows phone-capture and check-in cost rules.
Evidence: Rendered HTML

Source: Rendered HTML

Portable HTML render of the target document.

<!DOCTYPE html>
<html xmlns="http://www.w3.org/1999/xhtml">
<head>
  <meta charset="utf-8" />
  <meta name="generator" content="pandoc 3.11" />
  <meta name="viewport" content="width=device-width, initial-scale=1.0, user-scalable=yes" />
  <title>gsd-method</title>
  <style>
    /* Default styles provided by pandoc.
    ** See https://pandoc.org/MANUAL.html#variables-for-html for config info.
    */
    html {
      color: #1a1a1a;
      background-color: #fdfdfd;
    }
    body {
      margin: 0 auto;
      max-width: 36em;
      padding-left: 50px;
      padding-right: 50px;
      padding-top: 50px;
      padding-bottom: 50px;
      hyphens: auto;
      overflow-wrap: break-word;
      text-rendering: optimizeLegibility;
      font-kerning: normal;
    }
    @media (max-width: 600px) {
      body {
        font-size: 0.9em;
        padding: 12px;
      }
      h1 {
        font-size: 1.8em;
      }
    }
    @media print {
      html {
        background-color: white;
      }
      body {
        background-color: transparent;
        color: black;
      }
      p, h2, h3 {
        orphans: 3;
        widows: 3;
      }
      h2, h3, h4 {
        page-break-after: avoid;
      }
    }
    p {
      margin: 1em 0;
    }
    a {
      color: #1a1a1a;
    }
    a:visited {
      color: #1a1a1a;
    }
    img {
      max-width: 100%;
    }
    svg {
      height: auto;
      max-width: 100%;
    }
    h1, h2, h3, h4, h5, h6 {
      margin-top: 1.4em;
    }
    h5, h6 {
      font-size: 1em;
      font-style: italic;
    }
    h6 {
      font-weight: normal;
    }
    ol, ul {
      padding-left: 1.7em;
      margin-top: 1em;
    }
    li > ol, li > ul {
      margin-top: 0;
    }
    blockquote {
      margin: 1em 0 1em 1.7em;
      padding-left: 1em;
      border-left: 2px solid #e6e6e6;
      color: #606060;
    }
    code {
      white-space: pre-wrap;
      font-family: Menlo, Monaco, Consolas, 'Lucida Console', monospace;
      font-size: 85%;
      margin: 0;
      hyphens: manual;
    }
    pre {
      margin: 1em 0;
      overflow: auto;
    }
    pre code {
      padding: 0;
      overflow: visible;
      overflow-wrap: normal;
    }
    .sourceCode {
     background-color: transparent;
     overflow: visible;
    }
    hr {
      border: none;
      border-top: 1px solid #1a1a1a;
      height: 1px;
      margin: 1em 0;
    }
    table {
      margin: 1em 0;
      border-collapse: collapse;
      width: 100%;
      overflow-x: auto;
      display: block;
      font-variant-numeric: lining-nums tabular-nums;
    }
    table caption {
      margin-bottom: 0.75em;
    }
    tbody {
      margin-top: 0.5em;
      border-top: 1px solid #1a1a1a;
      border-bottom: 1px solid #1a1a1a;
    }
    th {
      border-top: 1px solid #1a1a1a;
      padding: 0.25em 0.5em 0.25em 0.5em;
    }
    td {
      padding: 0.125em 0.5em 0.25em 0.5em;
    }
    header {
      margin-bottom: 4em;
      text-align: center;
    }
    #TOC li {
      list-style: none;
    }
    #TOC ul {
      padding-left: 1.3em;
    }
    #TOC > ul {
      padding-left: 0;
    }
    #TOC a:not(:hover) {
      text-decoration: none;
    }
    span.smallcaps{font-variant: small-caps;}
    div.columns{display: flex; gap: 1.5em;}
    div.column{flex: auto;}
    @media screen {
    div.columns{gap: min(4vw, 1.5em);}
    div.column{overflow-x: auto;}
    }
    div.hanging-indent{margin-left: 1.5em; text-indent: -1.5em;}
    /* The extra [class] is a hack that increases specificity enough to
       override a similar rule in reveal.js */
    ul.task-list[class]{list-style: none;}
    ul.task-list li input[type="checkbox"] {
      font-size: inherit;
      width: 0.8em;
      margin: 0 0.8em 0.2em -1.6em;
      vertical-align: middle;
    }
  </style>
</head>
<body>
<h1 id="gsd-method-in-firstmate-revision-2">GSD method in Firstmate,
revision 2</h1>
<p>This document is Firstmate's adaptation of the GSD method from <a
href="https://github.com/open-gsd/gsd-core">open-gsd/gsd-core</a>: the
approach without the tool. It is a running document: the rules above the
log are the current contract, and the dated log at the bottom is the
evidence that shaped them.</p>
<h2 id="1-purpose">1. Purpose</h2>
<p>Phase-sized builds fail in one recurring way: a fact is assumed,
never checked, and a subsystem that an engine already provides gets
hand-rolled on top of that assumption. The GSD method's research-first
discipline catches that failure before code exists, and Blockvalley
Phase 0 proved it does (see the log).</p>
<p>Why not the tool: gsd-core ships an installer, <code>/gsd-*</code>
slash commands, <code>STATE.md</code> and milestone/phase file
bookkeeping, git hooks that prompt the user, and an executor framework
that runs plans in parallel fresh-context waves. Firstmate already
provides each of those needs differently - the backlog, generated
briefs, scouts and ship workers in isolated worktrees, keyed check-ins,
captain-held tasks, and guarded merges - so installing the tool would
add a second orchestrator beside the one this repo is. The captain's
ruling is explicit: "GSD is a method here, not a tool: no gsd-core, no
gsd-pi, no extra hooks." Section 6 names, for every dropped piece, the
Firstmate piece that covers it.</p>
<h2 id="2-when-it-applies">2. When it applies</h2>
<ul>
<li>Use the method for phase-sized builds: a new subsystem, an engine or
runtime choice, or a multi-week slice whose shape is not yet known.</li>
<li>Do not use it for bug fixes, one-file features, or clearly specified
changes; those take the ordinary ship path in <code>AGENTS.md</code>
section 7.</li>
<li>The test is uncertainty that could materially change whether or what
to build; that is the same threshold section 7 uses to justify a scout
over a ship.</li>
</ul>
<h2 id="3-the-six-kept-parts">3. The six kept parts</h2>
<p>Each part is stated as the rule, the Firstmate mechanism that
enforces it, and the artifact with its owner. The mechanisms are
referenced, not restated; their owners are <code>AGENTS.md</code>
sections 7 and 10, <code>bin/fm-brief.sh --help</code>, and
<code>.agents/skills/captain-hold-lifecycle/SKILL.md</code>.</p>
<h3 id="31-research-before-planning">3.1 Research before planning</h3>
<ul>
<li>Rule: every phase starts with research in fresh-context researchers,
one per stream, and one synthesizer that reads the researcher reports
plus the authoritative CONTEXT/rulings document, nothing else.</li>
<li>Mechanism: each researcher and the synthesizer is a scout spawned
from a <code>bin/fm-brief.sh --scout</code> brief in its own worktree,
so no researcher inherits another's context or the planner's sunk
cost.</li>
<li>Artifact and owner: each researcher leaves one
<code>data/&lt;task&gt;/report.md</code>, and the synthesizer scout's
single <code>data/&lt;task&gt;/report.md</code> is the durable synthesis
artifact; the Blockvalley <code>RESEARCH-0.md</code> was a verbatim copy
of that report, and the governing firstmate or secondmate owns the phase
plan that names the streams.</li>
</ul>
<h3 id="32-claim-provenance">3.2 Claim provenance</h3>
<ul>
<li>Rule: every factual claim in research and plans carries
<code>[VERIFIED: source]</code>, <code>[CITED: url]</code>, or
<code>[ASSUMED]</code>, and an ASSUMED claim never becomes a locked
decision without the captain's word.</li>
<li>Rule detail: an API or engine capability is VERIFIED only by
official documentation read that session or a working probe; a blog post
or training memory is ASSUMED.</li>
<li>Mechanism: the plan-check fails a load-bearing claim whose tag is
false, and every ASSUMED item that gates the build is registered as a
captain-held task through <code>captain-hold-lifecycle</code> before the
research or plan is treated as complete.</li>
<li>Artifact and owner: the provenance tags live in the report or plan
that makes the claim; the held task in the owning home's backlog carries
the captain's answer, closed only with his actual words.</li>
</ul>
<h3 id="33-the-context-document">3.3 The CONTEXT document</h3>
<ul>
<li>Rule: research is bounded by one document of locked decisions,
discretion areas, and deferred ideas, and researchers inves

... [4632 bytes truncated] ...

search and
the CONTEXT document.</li>
<li>CHECK: one adversarial plan-check returns BLOCKER and WARNING
findings with a concrete fix per finding.</li>
<li>PATCH: one scoped patch applies exactly the named fixes and nothing
else, and the checker re-verifies the changed lines, the new claims, and
every untouched section that consumes a changed contract.</li>
<li>If blockers remain after the patch, the governing secondmate stops
and posts one keyed check-in to the parent firstmate with the remaining
blockers and two options.</li>
<li>No second rewrite happens on the worker's or the secondmate's own
authority; the parent firstmate decides between the options or escalates
to the captain.</li>
<li>A gate that depends only on already-closed findings may be
dispatched while an unrelated blocker is escalated, when the checker's
dispatch judgment says so explicitly.</li>
</ul>
<p>Roles:</p>
<ul>
<li>Researcher -&gt; report.</li>
<li>Synthesizer -&gt; plan.</li>
<li>Checker -&gt; verdict, BLOCKER or WARNING, on the single question
"is it executable".</li>
<li>Parent firstmate -&gt; the escalation decision when the cap is
reached, and quota confirmation for the Fable seat.</li>
<li>Captain -&gt; locked decisions only, reached through captain-held
tasks.</li>
</ul>
<h2 id="5-model-seats">5. Model seats</h2>
<p>This is the current setting as the captain set it on 2026-08-29; the
captain's preference file (<code>data/captain.md</code>, or the
governing secondmate's copy) wins whenever it differs from this
section.</p>
<ul>
<li>Round 1 synthesis and plan writing: Claude Fable 5 at HIGH effort,
never xhigh or max; this is the only Fable seat, and the parent
firstmate confirms Fable quota before each Round 1 dispatch.</li>
<li>If the Fable seat is unavailable, Round 1 waits; the parent
firstmate reports the seat unavailable and the captain decides whether
that round runs on Codex gpt-5.6-sol HIGH instead.</li>
<li>Round 1 never downgrades on a worker or secondmate authority.</li>
<li>Plan-check and every later patch: Codex gpt-5.6-sol HIGH; Fable does
not re-enter to rewrite.</li>
<li>Researchers: Codex gpt-5.6-sol HIGH.</li>
<li>Execution workers: Codex gpt-5.6-sol medium by default, high when
the task is genuinely hard, and every task names the slice element it
unblocks.</li>
</ul>
<h2 id="6-explicitly-out">6. Explicitly out</h2>
<ul>
<li>Installer: Firstmate's scaffolded briefs and <code>bin/</code>
scripts are already installed in every home and secondmate home.</li>
<li>Slash commands (<code>/gsd-new-project</code>,
<code>/gsd-onboard</code>, and the phase commands): the governing
firstmate dispatches each phase step as a scout or ship task from the
backlog.</li>
<li>Milestone and phase file bookkeeping (<code>STATE.md</code>, phase
archives): <code>data/backlog.md</code> and the dated rulings files
under <code>data/</code> are the durable state, and this document's log
is the method's own record.</li>
<li>Hook-driven user prompts: keyed <code>needs-decision</code>
check-ins and captain-held tasks carry every question to the parent
firstmate and, only when genuinely his, to the captain.</li>
<li>Executor framework: <code>bin/fm-spawn.sh</code> launches each
execution task in its own isolated worktree with a fresh context;
<code>bin/fm-pr-merge.sh</code> owns PR merges, and
<code>bin/fm-merge-local.sh</code> owns approved local-only
landing.</li>
</ul>
<h2 id="7-success-signals">7. Success signals</h2>
<ul>
<li>The plan survives one check with a bounded blocker count and reaches
dispatch after the single patch.</li>
<li>The build hits the phase's named proof, for example "the slice on
the captain's phone", without a re-plan.</li>
<li>Fable spend per phase is at most one HIGH-effort Round 1 seat; zero
is valid only under the captain-approved fallback in section 5.</li>
</ul>
<h3 id="71-capture-coverage-contract">7.1 Capture-coverage contract</h3>
<ul>
<li>Rule: an "on the phone" proof passes only with a capture manifest
that names every required element ID, the camera bounds, one
deterministic state and frame, and an evidence link for every acceptance
criterion.</li>
<li>Why: the Blockvalley Phase 1 retro found that phone captures could
look successful while omitting required elements and interactions.</li>
<li>Artifact and owner: the proof-slice worker owns the capture
manifest, and the independent critic verifies its criterion coverage
before local-main landing.</li>
</ul>
<h2 id="8-how-we-review-our-use-of-it">8. How we review our use of
it</h2>
<ul>
<li>After each phase's named proof lands, the governing secondmate posts
one keyed check-in as a usage retrospective: what the method caught,
what it cost, and what this document got wrong.</li>
<li>The parent firstmate folds accepted changes into the sections above
and records the retrospective as a log entry below.</li>
<li>The log is the record; it is evidence-backed, dated, and pruned
rather than appended forever, under the same contract as
<code>data/learnings.md</code>.</li>
</ul>
<h3 id="81-check-in-cost-accounting">8.1 Check-in cost accounting</h3>
<ul>
<li>Rule: every phase check-in records the model used, the number of
rounds, and elapsed wall time, with unavailable values marked unknown
rather than reconstructed later.</li>
<li>Why: the Blockvalley Phase 1 retro could not recover complete
worker, model, and wall-time totals after task cleanup.</li>
</ul>
<h2 id="log">Log</h2>
<p>Entries are dated, cite their evidence by stable name or path, and
are rewritten or removed when superseded.</p>
<ul>
<li>2026-08-28 - Blockvalley Phase 0 ran the method by hand from a plan
with a CONTEXT section, three researchers, and one synthesizer
(<code>data/blockvalley-research-plan-phase0-2026-08-28.md</code>,
<code>data/blockvalley-p0-synth-s/RESEARCH-0.md</code>, both in the
Blockvalley home). It caught the real failure: the earlier substitution
program had assumed a renderer could be hand-rolled from repainted
photos, and the research named y-sort, autotile, and shadows as
Don't-Hand-Roll items with provenance. The synthesis was prescriptive
(Flutter + Flame, runner-up Godot 4) and its open ASSUMED items became
one captain-held engine decision, which the captain answered the same
day. The learnings entry that seeded this document is the GSD METHOD
line in <code>data/learnings.md</code>.</li>
<li>2026-08-28 to 2026-08-29 - Blockvalley Phase 1 planning ran four
plan rounds before a cap existed. The Round 1 plan was written on Claude
Fable 5 at xhigh (<code>data/blockvalley-p1-plan/report.md</code>) and
failed its first check on nine dimensions, including a false
<code>[VERIFIED]</code> Tiled command and locked UX behaviour reduced to
Phase 2 (<code>data/blockvalley-p1-plancheck/report.md</code>). Revision
4 on Fable still carried one blocker after the round-4 check, an
argument-count mismatch in one task's Tiled contract, while Gate 0 was
judged dispatchable
(<code>data/blockvalley-p1-planrev3/report.md</code>,
<code>data/blockvalley-p1-plancheck-r4/report.md</code>). The cost was
roughly a week of Fable quota with no Fable or Opus seats available
until the 2026-08-30 reset
(<code>data/blockvalley-p1-scope-cut-2026-08-29.md</code>). Consequence
recorded in section 4 and section 5: write once, check once, patch once
on Codex, escalate on the third failure, Fable only at HIGH and only for
Round 1.</li>
<li>2026-08-29 - The captain set the planning and model rules that
sections 4 and 5 state, in the Blockvalley steering protocol
(<code>data/steering-protocol.md</code>, "Planning and model rules");
the Phase 1 scope cut restated the plan-round cap and the slice-element
rule for every task
(<code>data/blockvalley-p1-scope-cut-2026-08-29.md</code>). The
captain's UX rulings file is the worked example of a CONTEXT document
that bounds behaviour without ruling on looks
(<code>data/blockvalley-ux-v1/UX-RULINGS-v1.md</code>).</li>
<li>2026-08-30 - The Blockvalley Phase 1 retro tightened revision 2 with
a phase-plan budget, a scope-cut reconciliation gate, a pre-landing
code-level Don't-Hand-Roll audit and independent critic, a
capture-coverage contract, and minimal check-in cost accounting.</li>
</ul>
</body>
</html>

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 3 issues found → auto-fixed ✅
  • ⚠️ docs/gsd-method.md:46 - The intent says to keep the existing Don't-Hand-Roll structure without restyling it, but this renames ### 3.4 The Don&#39;t-Hand-Roll list to ... audit, changing its public heading and anchor. Retain the existing heading or confirm the intentional rename.
  • 🚨 docs/gsd-method.md:123 - The existing success signal requiring the build to hit its named proof without a re-plan was removed. The new capture rule only constrains an on-phone proof if one is attempted, so a phase can now land without any named proof and bypass the proof-triggered check-in; restore the existing signal alongside the new contract.
  • ⚠️ docs/gsd-method.md:137 - The cost rule uses unscoped singular fields while its rationale identifies missing complete worker, model, and wall-time totals. A multi-worker phase can satisfy this wording with only the check-in poster's values; define phase-wide totals or require per-worker records plus totals.

🔧 Fix: Restored heading and named-proof signal; focused diff check passed
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • tests/fm-documentation-audiences.test.sh
  • git diff --check 28d7049036bee2a92c7dc82b38501c1fae731680 372e974b4e727f9d81116c1c4c3fd6792a95cab5
  • pandoc docs/gsd-method.md --from=gfm --to=html5 --standalone
  • Isolated Chrome render and screenshots at 1440x1200
  • Manual scope review of the target diff
  • git status --short --untracked-files=all and git diff --quiet HEAD --
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@withally
withally force-pushed the fm/fm-gsd-method-doc-r2 branch from aef8147 to 372e974 Compare August 30, 2026 13:16
@withally
withally merged commit e2399c6 into main Aug 30, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant