Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .agents/skills/firstmate-coding-guidelines/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -118,6 +118,8 @@ Never configure a deterministic suite-walk `commands.test` in any repository's n
Targeted validation belongs to the no-mistakes evidence path, while CI owns broad deterministic regression coverage.
Firstmate PR #3644 demonstrated the cost: pinning a 75-162-script walk took 32.7 minutes per validation, while removing it restored the 3.6-minute targeted-validation posture.

When the gate's test step refuses a scenario the repo's own suites already execute, use `test.instructions` in `.no-mistakes.yaml`, never a `commands.test`; that config is the single owner of the repository's live-evidence boundary.

## Repo style rules

- Put one full sentence per line in tracked Markdown.
Expand Down
112 changes: 112 additions & 0 deletions .no-mistakes.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,118 @@ commands:

# Publish each run's test evidence to the orphan no-mistakes/evidence branch linked from the PR.
# The evidence is not committed to the feature or default branch.
#
# Runbook for the gate's live-evidence test agent. Without it the analyzer
# derives scenarios, proves them with this repo's own suites, and is then
# refused for reporting a "pass" it never recognized as a live drive
# (`scenario N result "pass" requires live validation`). The runbook tells the
# analyzer what this product and these suites actually are, so a genuinely
# executed scenario can be reported live and an un-executed one still cannot.
# It narrows and clarifies the gate's live-evidence policy; it never weakens it.
# Trusted-only: the analyzer reads this runbook from the default-branch copy of
# this file, so a feature branch cannot alter the active live-evidence policy.
test:
evidence:
store_in_repo: true
instructions: |
WHAT THE PRODUCT IS HERE.
Firstmate is not a service you stand up. Its product entry points are the
executable files under bin/ that the fleet runs as processes, whatever
language each is written in - the bin/fm-*.sh commands and equally
bin/fm-voice-client.py, bin/fm-voice-relay.py, bin/fm-mail.py,
bin/fm-arm-command-policy.mjs, bin/fm-cd-command-policy.mjs and
bin/fm-extension-launch-barrier.mjs - plus the .pi/extensions TypeScript
editor extensions. Being something the fleet runs is the test, not
the language it happens to be written in. A module under bin/ that only
gets imported, such as bin/fm_voice_frame.py, is library code those entry
points use, not an entry point. Its end user is an operator, or a
supervising agent, who runs those entry points against a firstmate home
(FM_HOME) holding data/, state/, config/ and projects/. So driving a
scenario end-to-end here means running those real entry points as real
processes against a throwaway FM_HOME and reading what they actually did to
that home - not importing a function, and not reasoning about the source.

THE SUITES IN tests/ ARE THAT DRIVE, NOT UNIT TESTS.
Most tests/<subject>.test.sh scripts build a disposable state root and git
fixture and then run the real bin/ entry points as real subprocesses,
asserting only what is observable through their public interfaces;
asserting implementation source bytes is forbidden in this repo. A case
drove the product when it exercised the shipped artefact through the real
runtime that actually runs it, in a process the case launched. Launching
the entry point itself qualifies, directly or through a child contract
executable the case launches - tests/fm-inbox-conversation.test.sh drives
bin/fm-inbox.sh that way, through tests/fm-inbox-conversation-cases.py.
Importing a module into an interpreter the test controls and calling its
functions does NOT qualify, even for a file the product ships, even when
that interpreter is a child process, and even when real packages are
symlinked beside it: the product's runtime never started. Read the case
behind the transcript line you are about to cite and confirm what it did:
the transcript proves the case ran and its named assertion held, never that
it launched anything, and a script that drives entry points in its other
cases does not make this one live. Run the few scripts that cover your
scenarios, by path:

bin/fm-test-run.sh tests/<subject>.test.sh [more scripts...]

bin/fm-test-run.sh --help owns the selection and timing mechanics. Do not
run the whole suite.

When such a case executes a scenario, that IS driving the product live in
this run: report the scenario "pass" with "live": true and cite the script
and the transcript lines that case printed. An "ok - <behavior>" line, or a
"PASS: <behavior>" line a child contract executable prints through its
owning .test.sh wrapper, proves only that the named assertion held in a
case that ran; whether that case drove the product is settled by the
principle above, never by the line itself. The closing
"FM_TEST_END ... exit=0 ... gate_skip=false" proves the script as a whole
ran, not that every case inside it did, so cite the transcript line that
names your scenario, never the marker alone, and cite it only once you have
read the case behind it. This is the standing rule that an existing
automated test which drives a scenario end-to-end may be run as the
scenario and cited as its evidence. It is not an exemption from live
evidence, so apply it only to a case that actually executed here.

THE BOUNDARY THAT STAYS REFUSED.
Ask of every scenario: did a script execute the real surface as a process
in this run, or did it replay a recorded fixture standing in for something
this host cannot produce? Only the first is live. A scenario whose verdict
depends on a real vendor agent harness rendering in a real terminal pane, a
real network or forge call, a real remote host over SSH, or a real
credential is NOT proven by the portable suite, which substitutes a fixture
or a fake for exactly that. Neither is a scenario covered only by a case
that sources a library and feeds it recorded strings: that proves the
classifier's logic, not the live surface it was recorded from. Neither is a
scenario covered only by a case that parses a declarative artifact - a
YAML, JSON or config file - without running its consumer: that proves the
artifact's shape, not that the consumer reads it, so a scenario whose
consumer is the gate itself or any other process this run never launched
stays untested. tests/fm-nm-test-contract.test.sh is exactly such a case:
its "ok -" lines prove this file's shape, never that the gate read it.
Report every such scenario "untested" with "live": false
and a reason naming what was missing and how to supply it - never "pass",
and never "live": true.

The runner hands you that reason. bin/fm-test-run.sh classifies a script as
a gate skip only when its FIRST meaningful line is "skip: <reason>", and
only such a script ends with "gate_skip=true". An opt-in skip names the
variable that would turn the guard on; a capability skip names the tool,
harness or host requirement that was unavailable. A script may also print
"skip: <reason>" mid-transcript when one case cannot run here while its
other cases do; that leaves gate_skip=false for the script, so read every
"skip:" line in the transcript and treat the case it names as untested. A
skipped guard or case is not a met guarantee and never licenses a live
"pass", and a fixture replaying a harness's rendered output proves only the
assumption recorded in the fixture.

DO NOT MANUFACTURE LIVE EVIDENCE.
Do not set FM_LIVE, or any guard's own control variable, to force a skipped
live guard on: those guards submit prompts to real agent harnesses and
spend metered operator accounts without explicit authorization. Do not
launch a real agent harness, a real second mate, or a real
fleet worker for evidence either. bin/fm-spawn.sh, bin/fm-send.sh and
bin/fm-teardown.sh deliberately refuse to run from a gate worktree
(bin/fm-gate-refuse-lib.sh), and the suites reach them only for their own
throwaway fixture homes. Never hand-set FM_GATE_REFUSE_BYPASS, and never
point a scenario at the operator's real firstmate home. An honest
"untested" is listed on the pull request and costs nothing; a forced or
guessed "pass" costs everything.
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,7 +113,7 @@ Those sleeps look like recoverable overhead - `fm-watch-triage.test.sh` alone is
Sampling less often does not remove that wait, it only delays detection: raising the interval to 0.5s and charging each sample proportionally measured `fm-watch-triage.test.sh` at 435s and 440s against 390s and 393s for the unchanged script, back to back on 2026-09-03, because each of its ~40 poll-cycle waits and ~73 process-exit waits paid up to half a second more.
Some of those loops are also catching a transient rather than waiting for a settled condition, so a coarser sample can step over the state they assert on.
Discover tests by listing `tests/*.test.sh`: each is a self-contained bash script named `<subject>.test.sh`, and its header comment describes what it covers, so pass one to `bin/fm-test-run.sh` to focus on a subject with canonical timing output.
Shared test helpers live in `tests/lib.sh` (reporters, temp roots, git fixtures), `tests/fixtures.sh` (fake toolchain and spawn-world builders), `tests/wake-helpers.sh`, and `tests/secondmate-helpers.sh`.
Shared test helpers live in `tests/lib.sh` (reporters, temp roots, git fixtures, and `fm_yaml_to_json` for parsing a tracked YAML file with whichever parser the host has), `tests/fixtures.sh` (fake toolchain and spawn-world builders), `tests/wake-helpers.sh`, and `tests/secondmate-helpers.sh`.
Source those instead of copying a fake toolchain into a new suite.
A fixture may shorten a production timeout to keep a failure path prompt, but never below what the real work inside that window costs on a loaded machine: a fork, an exec, a lock acquisition, a beacon publication, or a first-poll check.
Where a case's assertion is not about the timeout itself, give that window headroom over the measured loaded cost, and bound the test's own waiting with iteration-counted poll loops, which stretch under load where a wall-clock budget does not.
Expand Down
1 change: 1 addition & 0 deletions bin/fm-test-run.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1429,6 +1429,7 @@ families_for_changed_path() {
.github/workflows/ci.yml|.no-mistakes.yaml)
printf '%s\n' pure-contract-unit
printf '%s\n' real-herdr-gated
printf '%s\n' __script__:fm-nm-test-contract.test.sh
;;
docs/fm-test-portable-shards.md|docs/fm-test-isolation-proof.md|\
docs/fm-test-isolation-proof.json)
Expand Down
4 changes: 3 additions & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -213,7 +213,9 @@ The flag is a home-local supervision-noise preference and is not inherited by se
The tracked `.no-mistakes.yaml` sets `test.evidence.store_in_repo: true` and pins `commands.lint` to `bin/fm-lint.sh`, the same owner CI invokes.
Storing evidence in the repo publishes each run's test artifacts to the orphan `no-mistakes/evidence` branch and links them from the PR body, instead of keeping them on local disk under the no-mistakes home.
That branch shares no history with code branches, so evidence never enters a pushed feature branch or the default branch; the worktree's `.no-mistakes/` stays local and CI rejects tracked entries under that path.
The [`firstmate-coding-guidelines` skill](../.agents/skills/firstmate-coding-guidelines/SKILL.md#no-mistakes-test-configuration) owns why `commands.test` stays absent and targeted validation belongs to the evidence path.
It also sets `test.instructions`, the runbook that tells the gate's test analyzer what this repository's entry points and `tests/` suites are and which scenarios stay untested.
That runbook is trusted like `commands.test`: the analyzer reads it only from the default-branch copy of `.no-mistakes.yaml`, so a pushed feature branch cannot alter the live-evidence policy its own run is judged by.
The [`firstmate-coding-guidelines` skill](../.agents/skills/firstmate-coding-guidelines/SKILL.md#no-mistakes-test-configuration) owns why `commands.test` stays absent, why `test.instructions` is the single owner of the live-evidence boundary, and why targeted validation belongs to the evidence path.
`commands.test` executes code, so no-mistakes honors it only from the default-branch copy of `.no-mistakes.yaml`; a pushed branch cannot change what the gate runs.
See [CONTRIBUTING.md](../CONTRIBUTING.md) for the firstmate-specific local test policy and entry points.
Portable shard evidence and coverage rules are in [fm-test-portable-shards.md](fm-test-portable-shards.md); [herdr-backend.md](herdr-backend.md#destructive-lab-safety) owns the real-Herdr lane's isolation boundary, and [runtime-backends.md](verification/runtime-backends.md#herdr) owns active evidence.
Expand Down
51 changes: 41 additions & 10 deletions tests/fm-nm-test-contract.test.sh
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
#!/usr/bin/env bash
# Contract: parsed .no-mistakes.yaml must leave commands.test absent or empty.
# Contract: parsed .no-mistakes.yaml must leave commands.test absent or empty
# and must carry a non-empty test.instructions runbook.
set -u

# shellcheck source=tests/lib.sh
Expand All @@ -8,19 +9,49 @@ set -u
NM="$ROOT/.no-mistakes.yaml"

test_nm_has_no_deterministic_test_command() {
command -v ruby >/dev/null 2>&1 \
|| fail "ruby is required to parse .no-mistakes.yaml for this contract"
local val
val=$(ruby -ryaml -e '
doc = YAML.load_file(ARGV[0]) || {}
cmds = doc["commands"] || {}
val = cmds.is_a?(Hash) ? cmds["test"] : nil
puts (val.nil? || val == false || val == "") ? "" : val.inspect
' "$NM") || fail "failed to parse .no-mistakes.yaml as YAML"
local json val
json=$(fm_yaml_to_json "$NM") \
|| fail "could not parse .no-mistakes.yaml as YAML (needs python3 with PyYAML, or ruby with psych)"
val=$(printf '%s' "$json" | python3 -c '
import json, sys

doc = json.load(sys.stdin) or {}
commands = doc.get("commands")
val = commands.get("test") if isinstance(commands, dict) else None
empty = val is None or val is False or (isinstance(val, str) and not val.strip())
print("" if empty else repr(val))
') || fail "failed to read commands.test from the parsed .no-mistakes.yaml"
if [ -n "$val" ]; then
fail "commands.test must be absent or empty so Test stays intent-targeted; got: $val"
fi
pass "no-mistakes does not configure commands.test"
}

test_nm_carries_a_test_instructions_runbook() {
local json status
json=$(fm_yaml_to_json "$NM") \
|| fail "could not parse .no-mistakes.yaml as YAML (needs python3 with PyYAML, or ruby with psych)"
status=$(printf '%s' "$json" | python3 -c '
import json, sys

doc = json.load(sys.stdin) or {}
test = doc.get("test")
if not isinstance(test, dict):
print("test is not a mapping")
else:
val = test.get("instructions")
if not isinstance(val, str):
print("test.instructions is %s rather than a string" % type(val).__name__)
elif not val.strip():
print("test.instructions is empty")
else:
print("ok")
') || fail "failed to read test.instructions from the parsed .no-mistakes.yaml"
if [ "$status" != ok ]; then
fail "test.instructions must stay present and non-empty so the test analyzer keeps its live-evidence runbook; $status"
fi
pass "no-mistakes carries a non-empty test.instructions runbook"
}

test_nm_has_no_deterministic_test_command
test_nm_carries_a_test_instructions_runbook
59 changes: 39 additions & 20 deletions tests/fm-test-run.test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -99,6 +99,7 @@ init_changed_fixture_repo() {
fm-documentation-audiences.test.sh \
fm-test-isolation-proof.test.sh \
fm-test-run.test.sh \
fm-nm-test-contract.test.sh \
fm-cd-pretool-check.test.sh \
fm-daemon.test.sh \
fm-harness-adapter-instructions-live-e2e.test.sh \
Expand Down Expand Up @@ -157,6 +158,7 @@ init_changed_fixture_repo() {
: >"$repo/.pi/extensions/lib/fm-operational-input.ts"
: >"$repo/docs/fm-test-isolation-proof.md"
: >"$repo/CONTRIBUTING.md"
: >"$repo/.no-mistakes.yaml"
: >"$repo/src/unmapped.ts"
git -C "$repo" init -q
git -C "$repo" add .
Expand Down Expand Up @@ -285,6 +287,19 @@ test_shell_line_ending_policy_selects_runner_contract() {
pass "shell line-ending policy selects runner coverage"
}

test_nm_config_change_selects_its_contract_test() {
local tmp repo listed
tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-nmconfig.XXXXXX")
repo="$tmp/repo"
init_changed_fixture_repo "$repo"
printf 'test:\n instructions: |\n runbook\n' >>"$repo/.no-mistakes.yaml"
listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD)
assert_contains "$listed" "tests/fm-nm-test-contract.test.sh" \
"a .no-mistakes.yaml change selects the contract that guards it"
rm -rf "$tmp"
pass "no-mistakes config change selects its contract test"
}

test_changed_dependency_selection_and_unmapped_failure() {
local tmp repo listed rc
tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-changed.XXXXXX")
Expand Down Expand Up @@ -1508,27 +1523,30 @@ test_herdr_ci_family_run_has_a_step_timeout() {
# The required Herdr lane's hang tripwire is the family-run *step* bound, not
# the 75-minute job cap. Parse the workflow as YAML so nested `with.name`
# artifact keys cannot masquerade as the step contract.
command -v ruby >/dev/null 2>&1 \
|| fail "ruby is required to parse .github/workflows/ci.yml as YAML"
local json job_timeout step_timeout
json=$(ruby -ryaml -rjson -e '
doc = YAML.load_file(ARGV[0])
job = doc.fetch("jobs").fetch("tests-herdr")
step = job.fetch("steps").find { |s|
s.is_a?(Hash) && s["name"] == "Run real-Herdr family (serial, required)"
}
raise "missing family-run step" if step.nil?
raise "family-run step has no timeout-minutes" unless step.key?("timeout-minutes")
puts JSON.generate(
"job_timeout" => job.fetch("timeout-minutes"),
"step_timeout" => step.fetch("timeout-minutes")
local json timeouts job_timeout step_timeout
json=$(fm_yaml_to_json "$ROOT/.github/workflows/ci.yml") \
|| fail "could not parse .github/workflows/ci.yml as YAML (needs python3 with PyYAML, or ruby with psych)"
timeouts=$(printf '%s' "$json" | python3 -c '
import json, sys

doc = json.load(sys.stdin)
job = doc["jobs"]["tests-herdr"]
step = next(
(
s
for s in job["steps"]
if isinstance(s, dict)
and s.get("name") == "Run real-Herdr family (serial, required)"
),
None,
)
' "$ROOT/.github/workflows/ci.yml") \
|| fail "could not parse tests-herdr timeouts from ci.yml"
job_timeout=$(python3 -c 'import json,sys; print(json.load(sys.stdin)["job_timeout"])' <<<"$json") \
|| fail "could not read job timeout from parsed workflow"
step_timeout=$(python3 -c 'import json,sys; print(json.load(sys.stdin)["step_timeout"])' <<<"$json") \
|| fail "could not read step timeout from parsed workflow"
if step is None:
raise SystemExit("missing family-run step")
if "timeout-minutes" not in step:
raise SystemExit("family-run step has no timeout-minutes")
print(job["timeout-minutes"], step["timeout-minutes"])
') || fail "could not parse tests-herdr timeouts from ci.yml"
read -r job_timeout step_timeout <<<"$timeouts"
[ "$job_timeout" = 75 ] \
|| fail "tests-herdr job backstop must stay 75 minutes, got $job_timeout"
[ "$step_timeout" = 20 ] \
Expand Down Expand Up @@ -1587,6 +1605,7 @@ test_changed_file_selection_is_conservative
test_task_marker_refuses_the_primary_checkout
test_changed_runner_surfaces_select_their_family
test_shell_line_ending_policy_selects_runner_contract
test_nm_config_change_selects_its_contract_test
test_changed_dependency_selection_and_unmapped_failure
test_changed_bin_reference_selects_per_script_not_per_family
test_changed_uses_bounded_automatic_concurrency
Expand Down
Loading
Loading