fix(routing): remove heuristic batch and embedding decisions - #1000
fix(routing): remove heuristic batch and embedding decisions#1000seonghobae wants to merge 106 commits into
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
2026-09-02 no-heuristics continuation — exact-head RCA found two additional NIM decision defects on the canonical branch: (1) |
|
@jules Fresh exact-head inventory shows this Draft has advanced to Repair the current branch directly through ordinary source/test commits; do not wait for these workflows to self-publish and do not add another source-fix successor. Acceptance for the next exact head:
Do not force-push, weaken gates, transfer predecessor GREEN, or treat a queued repair workflow as production evidence. |
|
Fresh exact-head/source revalidation found additional production policy defaults in |
| @@ -0,0 +1,93 @@ | |||
| name: Source fix PR1000 optimizer executable provenance | |||
| jobs: | ||
| repair: | ||
| permissions: | ||
| contents: write |
| @@ -0,0 +1,84 @@ | |||
| name: Source fix PR1000 optimizer selection | |||
| jobs: | ||
| repair: | ||
| permissions: | ||
| contents: write |
| @@ -0,0 +1,88 @@ | |||
| name: Source fix PR1000 reasoning effort | |||
| jobs: | ||
| repair: | ||
| permissions: | ||
| contents: write |
|
Adjudication evidence (host 1 session, 2026-09-06 KST; full report with commands in #1080). Nothing here closes, flips, or retargets anything — the decision is the opener's. Stray CI machinery in the tree: three branch-scoped |
Exact-log RCA and non-force source-fix repair — 2026-09-07The stalled reasoning-effort source-fix run Root cause: Commit Fresh source-fix run |
Follow-up exact-log RCA — run 34066033532The anchor correction worked: RED-before-change and the production transformation both completed, and Two causal classes remain:
Those integration failures must be repaired without reinstating priority/discovery-order/name fallback and without merely weakening the tests. The correct owner work is to define explicit eligibility/selection evidence in the fixtures and production virtual-provider contract, or fail closed where no such authority exists. Until that integrated RED→GREEN repair lands, the reasoning-effort one-shot must remain unexecuted and #1000 remains Draft/non-merge-ready. |
|
Exact-head RCA and bounded repair for Security and Quality run 34066307901:
This does not claim that the other 345 full-suite failures are fixed, does not weaken the production selector, and does not transfer the predecessor run as evidence. The branch remains Draft pending fresh exact-head hosted results and integration repair. |
|
CO #1067/#1074 integration notice: read the full existing 715f24a delta and freshly fetched this branch at 37cf3f6. The measurement_evidence_only/null-threshold classification repair already exists here; no duplicate estimator or replacement cutoff is needed. We intend to preserve that decision-authority slice on the psychometric parent, then normally integrate the child locked-cohort and failure-inclusive pairing changes. The separate explicit token-allocation slice remains owned here; this is not a claim of complete 715f24a or PR1000 supersession. Counterexample reproduced on child54b2b809: the same 30 successful route_once/conduct_bounded task pairs receive production_candidate_review with completion1.0 for declared locked_task_count30 and30000 alike. This is a constructed helper-contract case, not evidence of live contamination. The current owner helper still lacks the child locked-cell restriction, so integration must retain both. We will not move this PR branch, rerun its failed suite blindly, or treat its current integration failures as resolved. |
Current exact-head fuzz integration repair — 2026-09-07
37cf3f6bb0810ff3bcf0132ac08f09352380966734066307901on predecessor3f4a895…reached the full suite and failed 346 tests; its fuzz shard isolated one independent fixture regression.test_orchestration_on_arbitrary_promptconstructed three eligible mock agents and therefore hit the intended ambiguous-selection fail-closed contract before exercising prompt/trace/SSE invariants.Root cause
Several production routing, benchmark, and compute-allocation seams made substantive decisions without an identified model or authoritative measurement contract:
RoutingPolicyinferred sync-vs-batch fromlatency_tolerant,priority, andbatch_min_tokens;cheapest_upstreamassumed a representative 1000-prompt/1000-completion-token request shape;ModelGroupRoutersynthesizedP(success)/EWMA latencyas route authority;Under the no-heuristics contract, missing routing/accounting/evaluation evidence must fail closed rather than be replaced by a hand-authored selector.
Implemented production boundary
The current branch already changes the core routing seams so that:
channel=sync|batch, with compatibility latency/priority/token hints non-authoritative;TaskOrchestratorto the evidence-only model-selection bridge: multiple general model candidates require complete, unique, exact-context fast-mlsirm psychometric evidence or explicit eligible model/agent selection;The earlier NIM token/accounting repair has executable RED→GREEN evidence recorded on its source-fix lineage (
tests/test_nim_benchmark_no_heuristic_tokens.pyplustests/test_nim_benchmark.py: 102 passed on the repaired worktree before that one-shot was removed). That predecessor evidence is not transferred as hosted evidence for the current head.Active exact-head follow-up repairs
Exact current head, freshly re-fetched on 2026-09-02:
512ee46b42d71300e5bdc9170027915d6ddb1c2a.The changed-file set still contains temporary source-fix machinery for optimizer provenance, optimizer selection, and reasoning effort. Their presence is not completion evidence; do not overlap those source writers with a competing mutation.
A fresh exact-head source read also proves that
contextual_orchestrator/nim_benchmark.pystill constructsOrchestrationPolicywithroute_p95_seconds=2.5andworkflow_planning="template", withmax_workflow_steps=MAX_WORKFLOW_DEPTH. Protectedmaincontains the sameroute_p95_seconds=2.5policy. These values remain open no-heuristics findings until the branch either (a) binds each decision to an executable research/experimental/model authority or explicit caller-supplied governed evaluation design, or (b) fails closed instead of selecting a route/plan/compute allocation. Do not replace them with different repository-authored constants.MAX_WORKFLOW_DEPTH=5must likewise be treated as decision authority only if its cited issue/experiment is verified to govern this exact evaluation; a Conductor/TRINITY name in a comment is not by itself evidence that those papers authorize five steps.Current PR-triggered exact-head Tests, Security, Fuzz, SAST Semgrep, and Security Scan runs are queued. Queued state is non-passing and does not transfer predecessor evidence.
Research conformance refresh
docs/doctoring/routing-literature-refresh-2026-09.mdrefreshes the research boundary against TRINITY, the Conductor, the 2026 Sakana Fugu technical report, Select-then-Solve, and TwinRouterBench. The refreshed conclusion is deliberately narrower than those learned systems: their results support learned/validated per-task coordination and execution-grounded evaluation, but do not authorize hand-authored thresholds, weights, provider preferences, static workflow choices, or surrogate scores in this repository. Until a learned router or other explicit decision model has deployment-valid training/evaluation evidence, ambiguous selection remains fail-closed or requires exact-context fast-mlsirm evidence/explicit unique selection.Current merge boundary
The PR is mechanically mergeable but remains Draft and is not merge-ready. Temporary one-shot repair artifacts are still in the changed-file set and current exact-head required evidence is non-terminal. Landing requires the one-shot artifacts to be absent from one unchanged exact head, their intended production source changes to be present and verified, the newly re-confirmed
route_p95_seconds/workflow-planning/compute-allocation findings above to be resolved without another heuristic, fresh hosted Tests/Security/Fuzz/Semgrep/Security Scan and required central review evidence to be terminal-success, all valid exact-head review findings to be resolved, and ordinary repository protection/independent review to be satisfied. No force push, self-approval, administrative bypass, gate weakening, or fabricated/predecessor evidence is authorized.