docs(auto-combo): complete the mode pack table and gate what it claims - #12316
Merged
diegosouzapw merged 1 commit intoSep 1, 2026
Merged
diegosouzapw merged 1 commit into
diegosouzapw merged 1 commit into
Conversation
maxmad64bis
force-pushed
the
docs/mode-packs-table-gate
branch
from
September 1, 2026 14:15
b6a89e7 to
48e2d48
Compare
This was referenced Sep 1, 2026
maxmad64bis
force-pushed
the
docs/mode-packs-table-gate
branch
from
September 1, 2026 14:25
48e2d48 to
c783546
Compare
diegosouzapw
added a commit
to maxmad64bis/OmniRoute
that referenced
this pull request
Sep 1, 2026
…iers-confidence-claim diegosouzapw#12316 landed the docs-count gate extension this PR builds on, so scripts/check/check-docs-counts-sync.mjs and its test take the tip's side plus this PR's own required-claim additions.
diegosouzapw
pushed a commit
that referenced
this pull request
Sep 1, 2026
) docs/reference/FREE_TIERS.md said its numbers were "gathered by web research (confidence tagged per row)". No entry carries one: grep -c confidence on the catalog returns 0, the type does not declare the field, and the API serves nothing of the sort. A reader looking for "how much can I trust this figure" was pointed at a per-row signal that never existed. Replaced with the two counts the data actually supports — 7 of 446 entries carry hardStopGuaranteed (the field with the strictest sourcing rule in the repo: set only when the provider's own terms document that exceeding the free allowance refuses the request, source in a comment, never defaulted to true) and 13 carry a prompt-training disclosure. check:docs-counts reads both from the catalog at runtime, as required claims, so a reworded or deleted sentence fails rather than passing as "no claim in this file". The PR deliberately does not add a confidence field — curating one is a product call, and it says so instead of inventing it. Reconciled on merge: #12316 landed the gate extension underneath, so scripts/check/check-docs-counts-sync.mjs and its test took the tip's side plus this PR's own required-claim additions. Verified afterwards: check:docs-counts green, 48/48 across check-docs-counts-sync and free-catalog-no-confidence-field. Thanks @maxmad64bis — checking the gate against a number it should reject (7 swapped for 99) is the right way to prove a gate works.
7 tasks done
diegosouzapw
pushed a commit
that referenced
this pull request
Sep 1, 2026
…12317) The weight table said stability accounts for "low latency stdDev / error rate". Grep errorRate in scoring.ts and you find it declared on ProviderCandidate and read nowhere — while combo.ts pulls 24 hours of usage history behind a ten-sample floor, falls back to real-time metrics, and hands every candidate an errorRate the scorer ignores. Two candidates, one failing 1% of calls and one failing 99%, scored identically at 0.459486. This declares reliability as a sixteenth factor: 1 - failureRate, using the same formula, field precedence and rate-bounding speedRanking.ts already applies, so a corrupt reading means "nothing observed" rather than "fails every call". It ships at weight 0, leaving the ranking unchanged to the digit — the honest default, since which weight this deserves is a product call backed by traffic the author does not have. Two declared-but-silent factors already ship (cacheAffinity, resetWindowAffinity), so the pattern is not new. The stability row now describes what that factor actually computes: latency variance. The rest is the mechanical 15 → 16 across nineteen documents and the forty-two llm.txt mirrors — sourced from check:docs-counts rather than a grep, the first real use of the gate #12316 extended. Protected-surface note: this PR touches AGENTS.md, llm.txt and its 42 mirrors, and skills/omni-combos-routing/SKILL.md. Every changed line in those 45 files is a digit substitution and nothing else — masking all digits makes the removed and added lines identical, with no sentence added, removed or reworded. Reviewed and approved on that basis before merging. Verified on the author's rebased head: check:docs-counts green (the gate that now enforces the count this PR moves), typecheck:core clean, and 71/71 focused tests across scoring-reliability-factor, combo-scoring-weights-schema-coverage, check-docs-counts-sync, lkgp-enabled-context, intelligent-routing-options and the combo-matrix auto integration suite. Thanks @maxmad64bis — shipping the factor at weight 0 and saying plainly that the weight is someone else's call is the right way to land this.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
diegosouzapw#12316) docs/routing/AUTO-COMBO.md documented four mode packs; six ship. It printed 0.14 where modePacks.ts says 0.1333. And nothing was watching: four documents stated a scoring-factor count and check:docs-counts covered none of them. Wiring them up turned the gate red on seven real drifts — ARCHITECTURE.md and REPOSITORY_MAP.md at "9-factor" (code: 15) and "4 mode packs" (code: 6), RESILIENCE_GUIDE.md and SKILL.md at 13, AUTO-COMBO-GUIDE.md at both 5 and 13. ARCHITECTURE.md did not merely have the wrong number: it named nine factors that are not the engine's, and its four "mode packs" were the auto/* request prefixes. A product fact fell out of writing the table: no pack sets quality, and applying a pack replaces the weight map wholesale (weights = pack in engine.ts, not a merge), so quality carries 0.03 by default and normalizes to 0 under any pack — pick a mode pack and the observed-quality signal stops voting. Documented, not changed. The gate reads pack names from the module through the tsx subprocess that already reads every other code-derived count, matching the three spellings the docs actually use; on the reference document a missing claim now fails rather than passing. The dashboard was behind too (four of six packs offered); the count is dropped from the strategy label rather than corrected, since nothing reads selector labels and a right-today number goes stale unnoticed. Verified in a combined batch worktree: 174/174 focused tests across all 11 PRs of this batch, typecheck:core clean, and check:docs-counts green with the four newly-wired documents. Thanks @maxmad64bis — finding the quality-under-a-pack behaviour while writing a docs table is the kind of thing a table is for.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…gosouzapw#12318) docs/reference/FREE_TIERS.md said its numbers were "gathered by web research (confidence tagged per row)". No entry carries one: grep -c confidence on the catalog returns 0, the type does not declare the field, and the API serves nothing of the sort. A reader looking for "how much can I trust this figure" was pointed at a per-row signal that never existed. Replaced with the two counts the data actually supports — 7 of 446 entries carry hardStopGuaranteed (the field with the strictest sourcing rule in the repo: set only when the provider's own terms document that exceeding the free allowance refuses the request, source in a comment, never defaulted to true) and 13 carry a prompt-training disclosure. check:docs-counts reads both from the catalog at runtime, as required claims, so a reworded or deleted sentence fails rather than passing as "no claim in this file". The PR deliberately does not add a confidence field — curating one is a product call, and it says so instead of inventing it. Reconciled on merge: diegosouzapw#12316 landed the gate extension underneath, so scripts/check/check-docs-counts-sync.mjs and its test took the tip's side plus this PR's own required-claim additions. Verified afterwards: check:docs-counts green, 48/48 across check-docs-counts-sync and free-catalog-no-confidence-field. Thanks @maxmad64bis — checking the gate against a number it should reject (7 swapped for 99) is the right way to prove a gate works.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…iegosouzapw#12317) The weight table said stability accounts for "low latency stdDev / error rate". Grep errorRate in scoring.ts and you find it declared on ProviderCandidate and read nowhere — while combo.ts pulls 24 hours of usage history behind a ten-sample floor, falls back to real-time metrics, and hands every candidate an errorRate the scorer ignores. Two candidates, one failing 1% of calls and one failing 99%, scored identically at 0.459486. This declares reliability as a sixteenth factor: 1 - failureRate, using the same formula, field precedence and rate-bounding speedRanking.ts already applies, so a corrupt reading means "nothing observed" rather than "fails every call". It ships at weight 0, leaving the ranking unchanged to the digit — the honest default, since which weight this deserves is a product call backed by traffic the author does not have. Two declared-but-silent factors already ship (cacheAffinity, resetWindowAffinity), so the pattern is not new. The stability row now describes what that factor actually computes: latency variance. The rest is the mechanical 15 → 16 across nineteen documents and the forty-two llm.txt mirrors — sourced from check:docs-counts rather than a grep, the first real use of the gate diegosouzapw#12316 extended. Protected-surface note: this PR touches AGENTS.md, llm.txt and its 42 mirrors, and skills/omni-combos-routing/SKILL.md. Every changed line in those 45 files is a digit substitution and nothing else — masking all digits makes the removed and added lines identical, with no sentence added, removed or reworded. Reviewed and approved on that basis before merging. Verified on the author's rebased head: check:docs-counts green (the gate that now enforces the count this PR moves), typecheck:core clean, and 71/71 focused tests across scoring-reliability-factor, combo-scoring-weights-schema-coverage, check-docs-counts-sync, lkgp-enabled-context, intelligent-routing-options and the combo-matrix auto integration suite. Thanks @maxmad64bis — shipping the factor at weight 0 and saying plainly that the weight is someone else's call is the right way to land this.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
docs/routing/AUTO-COMBO.mddocuments four mode packs. Six ship. Its numbers are wrong too — it prints0.14wheremodePacks.tssays0.1333.The other half is that nothing was watching. Four documents state a scoring-factor count and
check:docs-countscovered none of them. Wiring them up turned the gate red on seven real drifts:ARCHITECTURE.mddidn't just have the wrong number — it named nine factors that aren't the engine's, and its four "mode packs" (coding, fast, cheap, smart) are theauto/*request prefixes, not packs.A product fact fell out of writing the table. No pack sets
quality, and applying a pack replaces the weight map wholesale (weights = packinengine.ts, not a merge). Soqualitycarries0.03by default and normalizes to0under any pack — pick a mode pack and the observed-quality signal stops voting. Documented, not changed.The gate reads pack names from the module via the
tsxsubprocess that already reads every other code-derived count, so there's no hand-rolled TypeScript parser to go quietly wrong. Its count pattern matches the three spellings the docs actually use (6 curated **mode packs**,6 pre-defined weight profiles,4 weight profiles) — matching only the first left the other two unguarded, which is how the stale4 weight profilesline inAUTO-COMBO.mditself surfaced. On the reference document a missing claim now fails rather than passing, so rewording a sentence past the pattern can't silence it.The dashboard was behind too: four of six packs offered, and the default strategy labelled
Rules (6-Factor Scoring). The count is gone from the label rather than corrected — nothing reads selector labels, so a number that's right today goes stale unnoticed.chaos-modeis now offered asChaos Mode (fault injection — testing); next to "Ship Fast" it would otherwise read as one more routing preference.Related Issues
Validation
npm run lintThe gate output above is the pre-fix run with the new checks wired in; after the corrections
check:docs-countsexits 0. The dashboard test failed on two counts before the change — the missing packs, and the factor count in the label.npm run lintexits 2 here and on an untouched checkout of the base alike (suppressions left that do not occur anymore, zero reported errors).Tests Added Or Updated
tests/unit/check-docs-counts-sync.test.ts(extended: pack-name validator, token-boundary matching, the three count spellings, required-claim behaviour)tests/unit/dashboard/intelligent-routing-options.test.ts(new: every shipped pack is offered, no strategy label states a factor count, the fault-injection pack says so)Coverage Notes
Nothing in
src/,open-sse/,electron/orbin/changes behaviour: the only code touched is a selector option list and two labels, both covered by the new dashboard test. The gate script's added functions are pure and unit-tested.Reviewer Notes
includes.ship-fastis a substring ofship-fast-v2, and a doc shouldn't pass by naming a pack that doesn't ship.REPOSITORY_MAP.mdshows 18 changed lines for a one-cell edit: the cell grew by a character andlint-stagedreformatted the table. Same forAUTO-COMBO-GUIDE.md, which wasn't prettier-clean on the base.0.9999at four decimals, not exactly1.0. The text says so rather than rounding the claim.chaos-modeandreliability-firstwere already reachable throughauto/chaosand the engine; this only makes them selectable from the panel that claims to list the packs.