chore(test): guard a plan against a single-exercise monoculture, and close #89 - #152
Merged
Merged
Conversation
…lose #89 Issue #89 tracked "~36 exercises carry ~80% of every plan" as a library concentration defect. Measured over 144 plans, that metric is not one: no exercise exceeds 15% of a plan's blocks or 14% of its minutes, 80% of the blocks sit on 55% of the rows a plan uses where an even spread reads 80%, and the head count is mostly set by the eligible pool — 94 rows with full gear against 38 on a bare bouldering wall — so it moves on a purchase rather than on a defect. Ruling 47 had already priced the one route to moving it. Closed as a non-defect by ruling 55, which keeps the part that IS owed: PR #120 fixed a real monoculture at its cause and nothing pinned the result, so a library or ranking edit could have walked it back silently. The guard asserts a per-plan ceiling on the single largest exercise over the existing twelve-row sweep at six session counts and three weaknesses, 216 plans: 17% of a plan's blocks and 23% of its seconds. Both halves, because #120 stated its result in minutes while #89's metric was blocks, and a ceiling on one leaves the other free to regress. `OPEN_CLIMBING_KEYS` is out of both numerator and denominator — those blocks arrive as ruling 27's length fill by ruling 29's decision, so counting them would measure a ruling, not a defect. Shown to fail, not assumed to: narrowing `prescribable()` to one row per cell — the generalised pre-#120 condition — puts `limit_boulders` at 17.50% of blocks and 30.78% of minutes, over both ceilings, on all twelve rows. Freezing the rotation instead stays green, which is the useful negative: pool narrowness makes a monoculture, rotation order does not. The worst case in both eras is the narrow-equipment column, so a full-vocabulary-only sweep would have been green for the wrong reason. Tests only — no product code changes, and no new sweep, plan builder or minutes helper: both arms read `generate()` output through the file's own `_input` and `_block_seconds`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #89 as a non-defect, and keeps the part that is genuinely owed.
Why #89 closed
It tracked "~36 exercises carry ~80% of every plan" as a library concentration defect. Measured over 144 plans (both disciplines × three grade bands × 1–7 days/week × three weakness settings, crossed with a full-vocabulary and a bouldering-wall-only equipment column):
There is no concentrated head: an even spread over the rows a plan uses reads 0.80 and it reads 0.55. The head count is mostly set by the eligible pool, which the climber's equipment fixes, so it moves on a purchase rather than on a defect. Ruling 47 had already measured the one route to moving it — re-weighting the rotation — as moving presence, not proportion, because a 2-session week reads only 2 of 16 ring positions.
A hypothesis worth recording as dead: the head is not floored by progression arithmetic.
progressed()takes no history and no plan argument and the dose is a pure function of the week's ordinal, so 68.3% of (block, exercise) pairs appear in exactly one week of their block. Nothing requires a row to repeat.What ruling 55 keeps
PR #120 fixed a real monoculture at its cause, and nothing pinned the result — no test in the repo asserted a per-plan single-exercise share, so a library or ranking edit could have walked it back silently.
test_no_single_exercise_DOMINATES_a_planasserts a per-plan ceiling on the single largest exercise over the existing twelve-row sweep × six session counts × three weaknesses = 216 plans: 17% of blocks, 23% of seconds. Both halves, because #120 stated its result in minutes while #89's metric was blocks.OPEN_CLIMBING_KEYSis out of numerator and denominator — those blocks arrive as ruling 27's length fill by ruling 29's decision.Shown to fail
Narrowing
prescribable()to one row per cell — the generalised pre-#120 condition — putslimit_bouldersat 17.50% of blocks and 30.78% of minutes, over both ceilings, on all twelve rows. Freezing the rotation instead stays green, which is the useful negative: pool narrowness makes a monoculture, rotation order does not. The worst case in both eras is the narrow-equipment column, so a full-vocabulary-only sweep would have been green for the wrong reason.Known limit, accepted deliberately: on the blocks half the window is narrow (14.85% today, 17.50% under total collapse, ceiling 17%), so that half fires only on near-total collapse and the detection headroom lives in the minutes half. Tightening it would turn red on a benign re-dose, which is why the measured maxima are narrated in the docstring rather than asserted.
Scope
Tests only. No product code, and no new sweep, plan builder or minutes helper — both arms read
generate()output through the file's own_inputand_block_seconds.npm run check:servergreen, 1451 passed.🤖 Generated with Claude Code