Skip to content

fix(agy): supply the ladder with current quota so the 25% Opus floor holds - #22

Merged
prajwal-395 merged 5 commits into
mainfrom
fm/fm-agy-floor-fixes
Aug 19, 2026
Merged

prajwal-395 merged 5 commits into
mainfrom
fm/fm-agy-floor-fixes

Conversation

@prajwal-395

@prajwal-395 prajwal-395 commented Aug 19, 2026 •

Copy link
Copy Markdown
Owner

Why

The 25% Opus 4.6 reserve floor shipped in c648692, is live, and its arithmetic is correct - the floor, the descent, the climb-back and the captain's override all do exactly what was asked. It still was not holding.

The defect was entirely in the evidence feeding it. The gate refuses only on positive, fresh quota evidence, while every way of losing that evidence resolved to "allow":

  • Nothing read quota at dispatch time. The only writer was bin/fm-watch.sh, which recorded a reading when a live agy pane happened to redraw with its footer in view. With no agy work running, nothing ever refreshed.
  • A reading stayed authoritative for its whole reset window, with no decay. Quota only falls during that window, so an early reading systematically overstated what was left, for hours. One number authorised five hours of unrestricted dispatch.
  • No headroom was reserved for launches in flight. At a fresh 26%, an unbounded burst of concurrent workers all cleared the floor and drove the true value under it before anything redrew.
  • Four parser hazards each silently orphaned or discarded a valid reading, every one of them a floor breach.

The captain's ruling shaped the fix

Asked whether missing evidence should refuse dispatch, the captain ruled on 2026-08-19 that it must not:

"we should not allow the fleet to stall but there has to be a way to quickly poll the quota before going in and launching a crewmate in the agy ladder to see if we are above the 25% quota or if we should drop rungs in the ladder."

So this change does not make missing evidence refuse. It stops the evidence from going missing. The intake poll is the centrepiece; unknown still allows a rung-1 launch, and the gate says out loud when it launched unchecked. An unproven descent is still refused, because that direction was already correct.

What changed

1. A live quota poll on the dispatch path (fm_agy_quota_poll). agy --print /quota --output-format json is answered by agy's own local command runner and runs no model turn, so it costs none of the budget the floor protects. One call answers every model at once, which lets a single poll serve both ladder decisions - whether rung 1 is above its floor, and which rung to drop to if it is not. fm_agy_ladder_gate polls before deciding, only for models the ladder actually ranks, so an off-ladder launch pays no network call while spawn locks are held.

Verified and pinned in docs/verification/agy-quota-poll.md, including what agy does not offer, so a future reader does not repeat the search.

2. A max-age ceiling (FM_AGY_QUOTA_MAX_AGE, default 300s), independent of the reset window. A reading older than the ceiling is unknown no matter how much window is left. This is what converts "one number authorises five hours" into "one number authorises minutes".

3. In-flight headroom reservation. An append-only per-rung ledger records each authorized launch the moment it is authorized, and the floor compares remaining - (in_flight * margin) rather than the bare reading, so a burst at 26% cannot cross the reserve before any refresh.

4. Parser hardening - ANSI stripped before the model key is derived (this one wrote a reading under an escape-laden key that the gate could never look up, silently orphaning it); tolerant whitespace around the separators; a non-numeric percentage reads unknown instead of 0; and an unparseable or 0m reset window no longer throws away an otherwise valid percentage, which is safe now that the ceiling bounds it.

5. Regression tests for all of it, none of which existed: ANSI-wrapped footer, padded/tight/multi-line footers, a reading aged past the ceiling but still inside its window, unparseable and 0m windows, non-numeric percentages, the poll's fail-soft paths, and a burst of concurrent launches just above the floor.

Evidence this was real, from production

Firstmate ran the discovered command live in the working home on 2026-08-19. Two things came back at once:

  • num_turns 0 and total_tokens 0 - the poll is genuinely free, confirming the load-bearing claim.
  • It reported Claude Opus 4.6 at 93% remaining while the recorded state file still said 33.8%, because that reading was hours old and its window had already reset.

That is the stale-evidence defect reproduced in production. Here the gap happened to be the harmless direction - the stale number was pessimistic and merely blocked launches the account could afford - but the gap is symmetric, and the dangerous direction is precisely how automatic dispatch ran below the reserved 25%.

Out of scope, deliberately

The floor value, the descent rule, the climb-back rule, and the FM_AGY_LADDER_OVERRIDE escape that reserves the last quarter for the captain are unchanged. They were correct.

Local config change, not in this PR

config/ is gitignored, so this is recorded rather than shipped. The _ladder_note in config/crew-dispatch.json still claimed enforcement had not landed and told firstmate to apply the floor by hand at intake - which firstmate reads at every intake, making a wrong note a live defect. It has been corrected locally to state that enforcement is live, point at bin/fm-agy-ladder-lib.sh, and describe the intake poll actually built. Old and new text are quoted in the task report.

Testing

CI is green on every check, including all four serial shards and the Herdr lane:

summary: "12 passed, 0 failed, 12 total"

An earlier green on this branch was thinner than it looked: serial shard 1 ran 14.1 minutes against a 15-minute cap, about 54 seconds of headroom, so that green was not evidence the shard timing was healthy. This branch is now rebased onto the shard rebalance from #21, and the run above is from that base:

Shard Scripts Duration Headroom vs 15m cap
serial 1 30 10.4m 4.6m
serial 2 31 11.7m 3.3m
serial 3 30 10.0m 5.0m
serial 4 31 9.9m 5.1m

Worst-case headroom is now 3.3 minutes rather than 54 seconds, and shard 1 in particular dropped from 14.1m to 10.4m.

bin/fm-lint.sh is ShellCheck-clean.

Before and after, against the pre-fix library

True remaining 20%, below the 25% floor - every reading should read 20 and refuse:

=== BEFORE (base 7427976) ===          === AFTER (this branch) ===
  plain          -> 20  enforced         plain          -> 20  enforced
  ANSI-wrapped   -> unknown  BREACH      ANSI-wrapped   -> 20  enforced
  padded spacing -> unknown  BREACH      padded spacing -> 20  enforced
  reset 0m       -> unknown  BREACH      reset 0m       -> 20  enforced
  reset unparsed -> unknown  BREACH      reset unparsed -> 20  enforced

One reading (95% remaining, 5h window) taken at T+0 and never refreshed:

=== BEFORE ===        === AFTER (300s ceiling) ===
  T+0h -> 95            T+0h -> 95
  T+1h -> 95            T+1h -> unknown
  T+2h -> 95            T+2h -> unknown
  T+3h -> 95            T+3h -> unknown
  T+4h -> 95            T+4h -> unknown

That is the headline defect: one number authorising five hours, now authorising minutes.

No agy quota was spent testing this: the poll costs none by construction, and every parser and ladder path is driven from synthetic fixtures and a stubbed agy.

Verdict on the two local test failures

Both were investigated against the base tree (7427976, exported with git archive so no worktree was created). Both test files are byte-identical at base and at this branch, and this branch touches no watcher, lock, wake, or classify code.

1. fm-wake-queue.test.sh - "the oversized unread status line was truncated or omitted". Pre-existing, proven, and a genuine macOS portability bug in that test.

It fails identically at base, where this change does not exist. The mechanism is exact: the assertion is grep -Fx "$expected", where $expected embeds the whole 20,006-byte huge.status line, so the pattern is oversized, not the input. macOS's system grep cannot build it:

$ /usr/bin/grep --version
grep (BSD grep, GNU compatible) 2.6.0-FreeBSD

pattern bytes: 20006
/usr/bin/grep -Fx : rc=2 grep: out of memory      <- byte-identical to the test failure
ugrep         -Fx : rc=0 (no stderr)

CI passes it because GNU grep on Linux handles a 20KB fixed pattern. This is worth a separate fix - the test is not portable to macOS - but it is not this PR's.

2. fm-watcher-lock.test.sh - "arm did not exit with HUP status (got 124)". Not attributable to this change, but I could not reproduce it.

124 is a timeout. Since the single observed failure I have run it 11 times with zero failures: 3x idle at base, 3x idle at this branch, 1x at base under 30 CPU hogs on 10 cores, and 4x concurrently at base. I could not make it fail on demand, including under deliberate load.

So I will not claim it proven pre-existing the way I can for the grep one. What is established: the test is byte-identical at base, this branch changes nothing it exercises, and it passes consistently at both revisions. The one failure occurred while several full lanes ran concurrently. Treating it as a load-sensitive flake is the reasonable reading, but it is an unreproduced flake and it deserves its own look, not a dismissal.

A note on running the suite locally

While this branch was in development, the local suite could destroy live worker panes: sixteen tests invoke bin/fm-bootstrap.sh, which reaches bin/fm-herdr-orphan-reaper.sh --close, and a shell carrying a live herdr identity pointed that reaper at the real session and workspace. Local lanes here were therefore run with HERDR_ENV, HERDR_PANE_ID, HERDR_TAB_ID, HERDR_WORKSPACE_ID, HERDR_SOCKET_PATH and HERDR_SESSION unset.

That workaround is no longer needed: bin/fm-test-run.sh now scrubs those variables itself, and this branch is rebased onto that fix. The suite verification above was re-run on the new base with the variables still present in the shell, relying on the runner's own scrub.

None of this is related to the change itself - nothing here touches bootstrap, the tangle guard, or the reaper.

…holds

The ladder's arithmetic was already correct - the floor, the descent, the
climb-back and the captain's override all decided exactly as asked. What failed
was the evidence they decided on.

Readings were written by one opportunistic writer: bin/fm-watch.sh, only while a
live agy pane happened to redraw with its footer in view. A home running no agy
pane refreshed nothing at all, and a reading counted as authoritative for the
whole remainder of its own reset window - hours, during which quota only falls.
One number recorded once therefore authorised five hours of unrestricted Opus
4.6 dispatch.

That did not only fail the floor open. It also blocked every DESCENT, because a
descent is refused unless the rungs above are proven exhausted. Rung 1 launched
unchecked below its floor while rung 2 was refused for lack of proof, so the
ladder became unusable in both directions and dispatch escaped it upward to a
costlier model outside the policy entirely.

The captain ruled that missing evidence must not stall the fleet, so this does
not make absence refuse. It stops the evidence going missing:

- An intake poll. `agy --print /quota --output-format json` answers the live
  quota for every model at once and runs no turn (agy 1.1.15 reports
  "num_turns":0 and "total_tokens":0), so reading the floor costs none of the
  budget the floor protects. fm_agy_ladder_gate runs it once at dispatch time,
  ahead of both decisions, which is why one call serves the floor AND the
  descent. Evidence in docs/verification/agy-quota-poll.md.
- A max-age ceiling on any reading, independent of its reset window. This is
  what turns "one number authorises five hours" into "minutes".
- Headroom reserved for launches already in flight, so a burst at 26% cannot all
  clear a 25% floor on the same pre-burst reading.
- Four parser hazards closed, each a silent floor breach: an ANSI-wrapped footer
  stored a reading under a key no reader could look up; whitespace variation
  around the separators dropped the reading entirely; a non-numeric percentage
  was recorded as a value rather than as an absence; and an unparseable or "0m"
  reset window discarded an otherwise valid percentage.

fm_agy_bounded_output gains a perl fallback. A stock macOS has neither `timeout`
nor `gtimeout`, so without it every bounded agy probe - the model catalogue as
well as this poll - silently declined to run on the platform the fleet runs on,
which would have left this fix inert where it is needed.

When the poll cannot answer, the gate still decides and says which fallback it
took: rung 1 launches, an unproven descent does not, and the refusal names the
failed live read and the override rather than reporting a bare absence.

The floor value, the descent rule, the climb-back rule, and the
FM_AGY_LADDER_OVERRIDE escape are unchanged.
…ange

[ -\/] parsed as the range space-to-backslash plus a literal slash, which
over-matched inside an escape sequence. The intended class is the CSI
intermediate range 0x20-0x2F, which [ -/] states exactly.
A caller that reached the message builder with something other than a count
must add no clause rather than abort the reason it was assembling.
…e clock

The reading ages these cases assert are load-bearing - the defect being fixed
was an age going unchecked - so recording on the wall clock and then asking how
old the reading was left a one-second window in which the answer changed.
fm_agy_quota_observe now accepts the same optional stamp fm_agy_quota_record
already took, and the suite pins every reading to it. The two-argument case
keeps the wall clock, because that default is what it exercises.
@prajwal-395
prajwal-395 force-pushed the fm/fm-agy-floor-fixes branch from 3c0237e to 2a1f6c8 Compare August 19, 2026 21:42
@prajwal-395
prajwal-395 merged commit 699b14a into main Aug 19, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant