Repository navigation
feat: decouple FPM collection limits and validate AgentX coverage - #300
Merged
Arsene12358 merged 1 commit intoSep 19, 2026
Merged
Conversation
Signed-off-by: Yiming Liu <yimingl@nvidia.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueComment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why and what changed
FPM onboarding currently derives collection limits from one synthetic validation workload. This change makes context, scheduler token/sequence limits and prefill CUDA graph capture independent, reviewable collection settings. Dynamo self-benchmark generates the grid. Users can change validation traffic and regenerate evaluation inputs while preserving verified collection data and checkpoints.
Adds
aisimulate onboard validate-fpmfor a local Weka/AgentX trace through existing cold aggregated replay. The canonical Rust performance model records measured, interpolated and unsupported direct-FPM lookups without changing timings. The command saves an ordinary prediction config, input/data hashes and replay/coverage reports. Coverage passes only when every selected request and play completes; missing timing preserves partial evidence and returns nonzero. The public guide and agent instructions describe this staged workflow.Review map
support/schema.py,support/plan.pyandcollector/fpm_forward/{config,planner,runner}.pyfor independent limits and safe plan refresh; thenperfmodel/fpm/coverage.rs,crates/core/src/python.rsandsupport/validation.pyfor lookup ownership and replay completion.estimator_config.fpm_interpolation.collect_coverageand returned-model coverage/raw-timing methods; additive collector runtime flags; onboarding request/plan v2 and coverage reports. Coverage requires explicit direct FPM with denied fallback and is disabled by default.Evidence
Local source is commit
10186e316be0a39e444ab1ca9c8cbb36b6288376. The installed wheel was built from the same frozen production files; the final changes before committing affected tests only.cargo test --workspace --no-default-features— 1,573 passed, one existing ignored fixture; embedded-Python boundary — 33 passed.python -m pytest -p no:timeout -c python/aisimulate/pytest.ini crates/core/parity_tests/perfmodel/test_engine_step_parity.py crates/core/parity_tests/perfmodel/test_compile_engine_parity.py— 386 passed; no golden changes.PYTHONPATH.python -m pytest -p no:timeout -p no:cacheprovider -c /dev/null -n 4 -q testsagainst the installed wheel — 2,676 passed, five skipped, 132 subtests; one existing import-isolation failure reproduced on the base branch (test_supervisor_argument_and_output_setup_do_not_import_runtime). Both base and this branch import lightweightaisimulate_core, neither imports the native runtime through that path.onboard init,plan, collection preview,validate-fpm, ordinarypredictandrecommendfrom outside the checkout withoutPYTHONPATH. Source/wheel/installed hashes match for every changed production Python file and the native stub.10186e316be0a39e444ab1ca9c8cbb36b6288376; CODEOWNERS and DCO also passed.10186e316be0a39e444ab1ca9c8cbb36b6288376— independent whole-branch and Standards reviews are clean, with no deferred findings. All five task reviews and both repair reviews completed separately; reviewers independently reran tests and behavioral controls.Modeling or data provenance
Functional replay checks use unchanged published FPM data for MiniMax-M2.7/H200 TP4 (vLLM 0.25.1) and GLM-5.2-NVFP4/GB200 DEP16 (vLLM 0.28.0), plus complete plays selected from the pinned Weka AgentX subset. Resource profiles are illustrative declared bounds, not qualified serving capacity. No trace rows, derived datasets or additional silicon measurements are committed.
Both deployments complete the selected seven-request play with zero unsupported lookups. A different complete play stops at decode batch 1 / 134,719 total past-KV tokens and correctly reports incomplete coverage. The full 393-play corpus is rejected by the existing host-memory preflight before any timing queries, so corpus-wide coverage is not claimed. This validates workflow behavior and timing availability, not predictive accuracy. No GPU collection was run.
Current scope is cold aggregated replay, one client lane, HBM-only cache and no speculation. Seeded snapshots, explicit warmup and expanded AgentX replay remain follow-up work after #207/#235; neither PR is a dependency here.
Tracking
feat/fpm-config-onboarding).