Skip to content

docs(auto-combo): complete the mode pack table and gate what it claims - #12316

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.51from
maxmad64bis:docs/mode-packs-table-gate
Sep 1, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.51from
maxmad64bis:docs/mode-packs-table-gate

Conversation

@maxmad64bis

@maxmad64bis maxmad64bis commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

docs/routing/AUTO-COMBO.md documents four mode packs. Six ship. Its numbers are wrong too — it prints 0.14 where modePacks.ts says 0.1333.

The other half is that nothing was watching. Four documents state a scoring-factor count and check:docs-counts covered none of them. Wiring them up turned the gate red on seven real drifts:

✗ ARCHITECTURE.md        — "9-factor"  — code has 15      ✗ ARCHITECTURE.md   — "4 mode packs" — code has 6
✗ REPOSITORY_MAP.md      — "9-factor"  — code has 15      ✗ REPOSITORY_MAP.md — "4 mode packs" — code has 6
✗ RESILIENCE_GUIDE.md    — "13-factor" — code has 15
✗ AUTO-COMBO-GUIDE.md    — "5 factors", "13-factor"       ✗ SKILL.md — "13 factors" — code has 15

ARCHITECTURE.md didn't just have the wrong number — it named nine factors that aren't the engine's, and its four "mode packs" (coding, fast, cheap, smart) are the auto/* request prefixes, not packs.

A product fact fell out of writing the table. No pack sets quality, and applying a pack replaces the weight map wholesale (weights = pack in engine.ts, not a merge). So quality carries 0.03 by default and normalizes to 0 under any pack — pick a mode pack and the observed-quality signal stops voting. Documented, not changed.

The gate reads pack names from the module via the tsx subprocess that already reads every other code-derived count, so there's no hand-rolled TypeScript parser to go quietly wrong. Its count pattern matches the three spellings the docs actually use (6 curated **mode packs**, 6 pre-defined weight profiles, 4 weight profiles) — matching only the first left the other two unguarded, which is how the stale 4 weight profiles line in AUTO-COMBO.md itself surfaced. On the reference document a missing claim now fails rather than passing, so rewording a sentence past the pattern can't silence it.

The dashboard was behind too: four of six packs offered, and the default strategy labelled Rules (6-Factor Scoring). The count is gone from the label rather than corrected — nothing reads selector labels, so a number that's right today goes stale unnoticed. chaos-mode is now offered as Chaos Mode (fault injection — testing); next to "Ship Fast" it would otherwise read as one more routing preference.

Related Issues

Validation

  • Change type: other (docs + CI gate) / UI
  • Focused tests and category gates from the golden path
  • npm run lint
  • Reconciled with the current active release base; focused checks rerun afterward
  • Production-code changes include a new or updated automated test in this PR

The gate output above is the pre-fix run with the new checks wired in; after the corrections check:docs-counts exits 0. The dashboard test failed on two counts before the change — the missing packs, and the factor count in the label.

npm run lint exits 2 here and on an untouched checkout of the base alike (suppressions left that do not occur anymore, zero reported errors).

Tests Added Or Updated

  • tests/unit/check-docs-counts-sync.test.ts (extended: pack-name validator, token-boundary matching, the three count spellings, required-claim behaviour)
  • tests/unit/dashboard/intelligent-routing-options.test.ts (new: every shipped pack is offered, no strategy label states a factor count, the fault-injection pack says so)

Coverage Notes

Nothing in src/, open-sse/, electron/ or bin/ changes behaviour: the only code touched is a selector option list and two labels, both covered by the new dashboard test. The gate script's added functions are pure and unit-tested.

Reviewer Notes

  • The pack-name check uses a token boundary, not includes. ship-fast is a substring of ship-fast-v2, and a doc shouldn't pass by naming a pack that doesn't ship.
  • REPOSITORY_MAP.md shows 18 changed lines for a one-cell edit: the cell grew by a character and lint-staged reformatted the table. Same for AUTO-COMBO-GUIDE.md, which wasn't prettier-clean on the base.
  • Each pack sums to 0.9999 at four decimals, not exactly 1.0. The text says so rather than rounding the claim.
  • chaos-mode and reliability-first were already reachable through auto/chaos and the engine; this only makes them selectable from the panel that claims to list the packs.

@maxmad64bis
maxmad64bis force-pushed the docs/mode-packs-table-gate branch from 48e2d48 to c783546 Compare September 1, 2026 14:25
@diegosouzapw
diegosouzapw merged commit 3b82d85 into diegosouzapw:release/v3.8.51 Sep 1, 2026
16 checks passed
diegosouzapw added a commit to maxmad64bis/OmniRoute that referenced this pull request Sep 1, 2026
…iers-confidence-claim

diegosouzapw#12316 landed the docs-count gate extension this PR builds on, so scripts/check/check-docs-counts-sync.mjs and its test take the tip's side plus this PR's own required-claim additions.
diegosouzapw pushed a commit that referenced this pull request Sep 1, 2026
)

docs/reference/FREE_TIERS.md said its numbers were "gathered by web research (confidence tagged per row)". No entry carries one: grep -c confidence on the catalog returns 0, the type does not declare the field, and the API serves nothing of the sort. A reader looking for "how much can I trust this figure" was pointed at a per-row signal that never existed.

Replaced with the two counts the data actually supports — 7 of 446 entries carry hardStopGuaranteed (the field with the strictest sourcing rule in the repo: set only when the provider's own terms document that exceeding the free allowance refuses the request, source in a comment, never defaulted to true) and 13 carry a prompt-training disclosure. check:docs-counts reads both from the catalog at runtime, as required claims, so a reworded or deleted sentence fails rather than passing as "no claim in this file".

The PR deliberately does not add a confidence field — curating one is a product call, and it says so instead of inventing it.

Reconciled on merge: #12316 landed the gate extension underneath, so scripts/check/check-docs-counts-sync.mjs and its test took the tip's side plus this PR's own required-claim additions. Verified afterwards: check:docs-counts green, 48/48 across check-docs-counts-sync and free-catalog-no-confidence-field.

Thanks @maxmad64bis — checking the gate against a number it should reject (7 swapped for 99) is the right way to prove a gate works.
diegosouzapw pushed a commit that referenced this pull request Sep 1, 2026
…12317)

The weight table said stability accounts for "low latency stdDev / error rate". Grep errorRate in scoring.ts and you find it declared on ProviderCandidate and read nowhere — while combo.ts pulls 24 hours of usage history behind a ten-sample floor, falls back to real-time metrics, and hands every candidate an errorRate the scorer ignores. Two candidates, one failing 1% of calls and one failing 99%, scored identically at 0.459486.

This declares reliability as a sixteenth factor: 1 - failureRate, using the same formula, field precedence and rate-bounding speedRanking.ts already applies, so a corrupt reading means "nothing observed" rather than "fails every call". It ships at weight 0, leaving the ranking unchanged to the digit — the honest default, since which weight this deserves is a product call backed by traffic the author does not have. Two declared-but-silent factors already ship (cacheAffinity, resetWindowAffinity), so the pattern is not new. The stability row now describes what that factor actually computes: latency variance.

The rest is the mechanical 15 → 16 across nineteen documents and the forty-two llm.txt mirrors — sourced from check:docs-counts rather than a grep, the first real use of the gate #12316 extended.

Protected-surface note: this PR touches AGENTS.md, llm.txt and its 42 mirrors, and skills/omni-combos-routing/SKILL.md. Every changed line in those 45 files is a digit substitution and nothing else — masking all digits makes the removed and added lines identical, with no sentence added, removed or reworded. Reviewed and approved on that basis before merging.

Verified on the author's rebased head: check:docs-counts green (the gate that now enforces the count this PR moves), typecheck:core clean, and 71/71 focused tests across scoring-reliability-factor, combo-scoring-weights-schema-coverage, check-docs-counts-sync, lkgp-enabled-context, intelligent-routing-options and the combo-matrix auto integration suite.

Thanks @maxmad64bis — shipping the factor at weight 0 and saying plainly that the weight is someone else's call is the right way to land this.
@maxmad64bis
maxmad64bis deleted the docs/mode-packs-table-gate branch September 24, 2026 21:11
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
diegosouzapw#12316)

docs/routing/AUTO-COMBO.md documented four mode packs; six ship. It printed 0.14 where modePacks.ts says 0.1333. And nothing was watching: four documents stated a scoring-factor count and check:docs-counts covered none of them. Wiring them up turned the gate red on seven real drifts — ARCHITECTURE.md and REPOSITORY_MAP.md at "9-factor" (code: 15) and "4 mode packs" (code: 6), RESILIENCE_GUIDE.md and SKILL.md at 13, AUTO-COMBO-GUIDE.md at both 5 and 13. ARCHITECTURE.md did not merely have the wrong number: it named nine factors that are not the engine's, and its four "mode packs" were the auto/* request prefixes.

A product fact fell out of writing the table: no pack sets quality, and applying a pack replaces the weight map wholesale (weights = pack in engine.ts, not a merge), so quality carries 0.03 by default and normalizes to 0 under any pack — pick a mode pack and the observed-quality signal stops voting. Documented, not changed.

The gate reads pack names from the module through the tsx subprocess that already reads every other code-derived count, matching the three spellings the docs actually use; on the reference document a missing claim now fails rather than passing. The dashboard was behind too (four of six packs offered); the count is dropped from the strategy label rather than corrected, since nothing reads selector labels and a right-today number goes stale unnoticed.

Verified in a combined batch worktree: 174/174 focused tests across all 11 PRs of this batch, typecheck:core clean, and check:docs-counts green with the four newly-wired documents.

Thanks @maxmad64bis — finding the quality-under-a-pack behaviour while writing a docs table is the kind of thing a table is for.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…gosouzapw#12318)

docs/reference/FREE_TIERS.md said its numbers were "gathered by web research (confidence tagged per row)". No entry carries one: grep -c confidence on the catalog returns 0, the type does not declare the field, and the API serves nothing of the sort. A reader looking for "how much can I trust this figure" was pointed at a per-row signal that never existed.

Replaced with the two counts the data actually supports — 7 of 446 entries carry hardStopGuaranteed (the field with the strictest sourcing rule in the repo: set only when the provider's own terms document that exceeding the free allowance refuses the request, source in a comment, never defaulted to true) and 13 carry a prompt-training disclosure. check:docs-counts reads both from the catalog at runtime, as required claims, so a reworded or deleted sentence fails rather than passing as "no claim in this file".

The PR deliberately does not add a confidence field — curating one is a product call, and it says so instead of inventing it.

Reconciled on merge: diegosouzapw#12316 landed the gate extension underneath, so scripts/check/check-docs-counts-sync.mjs and its test took the tip's side plus this PR's own required-claim additions. Verified afterwards: check:docs-counts green, 48/48 across check-docs-counts-sync and free-catalog-no-confidence-field.

Thanks @maxmad64bis — checking the gate against a number it should reject (7 swapped for 99) is the right way to prove a gate works.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…iegosouzapw#12317)

The weight table said stability accounts for "low latency stdDev / error rate". Grep errorRate in scoring.ts and you find it declared on ProviderCandidate and read nowhere — while combo.ts pulls 24 hours of usage history behind a ten-sample floor, falls back to real-time metrics, and hands every candidate an errorRate the scorer ignores. Two candidates, one failing 1% of calls and one failing 99%, scored identically at 0.459486.

This declares reliability as a sixteenth factor: 1 - failureRate, using the same formula, field precedence and rate-bounding speedRanking.ts already applies, so a corrupt reading means "nothing observed" rather than "fails every call". It ships at weight 0, leaving the ranking unchanged to the digit — the honest default, since which weight this deserves is a product call backed by traffic the author does not have. Two declared-but-silent factors already ship (cacheAffinity, resetWindowAffinity), so the pattern is not new. The stability row now describes what that factor actually computes: latency variance.

The rest is the mechanical 15 → 16 across nineteen documents and the forty-two llm.txt mirrors — sourced from check:docs-counts rather than a grep, the first real use of the gate diegosouzapw#12316 extended.

Protected-surface note: this PR touches AGENTS.md, llm.txt and its 42 mirrors, and skills/omni-combos-routing/SKILL.md. Every changed line in those 45 files is a digit substitution and nothing else — masking all digits makes the removed and added lines identical, with no sentence added, removed or reworded. Reviewed and approved on that basis before merging.

Verified on the author's rebased head: check:docs-counts green (the gate that now enforces the count this PR moves), typecheck:core clean, and 71/71 focused tests across scoring-reliability-factor, combo-scoring-weights-schema-coverage, check-docs-counts-sync, lkgp-enabled-context, intelligent-routing-options and the combo-matrix auto integration suite.

Thanks @maxmad64bis — shipping the factor at weight 0 and saying plainly that the weight is someone else's call is the right way to land this.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants