Skip to content

feat: decouple FPM collection limits and validate AgentX coverage - #300

Merged
Arsene12358 merged 1 commit into
feat/fpm-config-onboardingfrom
feat/fpm-agentx-collection-coverage
Sep 19, 2026
Merged

Arsene12358 merged 1 commit into
feat/fpm-config-onboardingfrom
feat/fpm-agentx-collection-coverage

Conversation

@Arsene12358

@Arsene12358 Arsene12358 commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

These changes are now included in #248 on feat/fpm-config-onboarding at commit 10186e316be0a39e444ab1ca9c8cbb36b6288376. Continue review and follow-up work in #248. GitHub marked this PR merged when its base branch advanced to that exact commit; #248 remains a draft stacked on #238.

Why and what changed

FPM onboarding currently derives collection limits from one synthetic validation workload. This change makes context, scheduler token/sequence limits and prefill CUDA graph capture independent, reviewable collection settings. Dynamo self-benchmark generates the grid. Users can change validation traffic and regenerate evaluation inputs while preserving verified collection data and checkpoints.

Adds aisimulate onboard validate-fpm for a local Weka/AgentX trace through existing cold aggregated replay. The canonical Rust performance model records measured, interpolated and unsupported direct-FPM lookups without changing timings. The command saves an ordinary prediction config, input/data hashes and replay/coverage reports. Coverage passes only when every selected request and play completes; missing timing preserves partial evidence and returns nonzero. The public guide and agent instructions describe this staged workflow.

Review map

  • Risk level: medium; spans onboarding, collection and the Python/Rust replay boundary.
  • Start with support/schema.py, support/plan.py and collector/fpm_forward/{config,planner,runner}.py for independent limits and safe plan refresh; then perfmodel/fpm/coverage.rs, crates/core/src/python.rs and support/validation.py for lookup ownership and replay completion.
  • Public or serialized contract changed: opt-in estimator_config.fpm_interpolation.collect_coverage and returned-model coverage/raw-timing methods; additive collector runtime flags; onboarding request/plan v2 and coverage reports. Coverage requires explicit direct FPM with denied fallback and is disabled by default.
  • Compatibility: saved v1 requests preserve declared limits during migration. Legacy collection plans require a new directory because their limits depended on validation traffic. Verified v2 plans support validation-only refresh. Interpolation math and frozen engine timing goldens are unchanged.

Evidence

Local source is commit 10186e316be0a39e444ab1ca9c8cbb36b6288376. The installed wheel was built from the same frozen production files; the final changes before committing affected tests only.

  • Native: cargo test --workspace --no-default-features — 1,573 passed, one existing ignored fixture; embedded-Python boundary — 33 passed.
  • Frozen parity: python -m pytest -p no:timeout -c python/aisimulate/pytest.ini crates/core/parity_tests/perfmodel/test_engine_step_parity.py crates/core/parity_tests/perfmodel/test_compile_engine_parity.py — 386 passed; no golden changes.
  • Package unit suite: 7,630 passed across the full run and two corrected source-import reruns; 11 skipped. macOS timeout disabled and the documented torch-dependent test omitted. Source-only legacy collector tests ran from the Python project; subprocesses that change directory used its absolute PYTHONPATH.
  • Repository suite: python -m pytest -p no:timeout -p no:cacheprovider -c /dev/null -n 4 -q tests against the installed wheel — 2,676 passed, five skipped, 132 subtests; one existing import-isolation failure reproduced on the base branch (test_supervisor_argument_and_output_setup_do_not_import_runtime). Both base and this branch import lightweight aisimulate_core, neither imports the native runtime through that path.
  • Packaged CLI: a fresh wheel installation passed onboard init, plan, collection preview, validate-fpm, ordinary predict and recommend from outside the checkout without PYTHONPATH. Source/wheel/installed hashes match for every changed production Python file and the native stub.
  • Policy: full Ruff lint/format, Rust format, copyright, packaged legal files, documentation links, generated CODEOWNERS and single-oracle checks passed.
  • Negative/boundary cases: missing timings, failed/empty/interrupted replay, modified inputs, unsafe output paths, legacy migration, invalid scheduler bounds and validation-only plan refresh. Native tests use literal measured fixture values and direct midpoint expectations; coverage-on/off tests preserve returned timings.
  • Fast CI: passed for 10186e316be0a39e444ab1ca9c8cbb36b6288376; CODEOWNERS and DCO also passed.
  • Full CI: not run. NVIDIA's runner gate currently requires vetting.
  • CodeRabbit reviewed commit: none; the bot skipped this draft PR.
  • Codex reviewed commit: 10186e316be0a39e444ab1ca9c8cbb36b6288376 — independent whole-branch and Standards reviews are clean, with no deferred findings. All five task reviews and both repair reviews completed separately; reviewers independently reran tests and behavioral controls.

Modeling or data provenance

Functional replay checks use unchanged published FPM data for MiniMax-M2.7/H200 TP4 (vLLM 0.25.1) and GLM-5.2-NVFP4/GB200 DEP16 (vLLM 0.28.0), plus complete plays selected from the pinned Weka AgentX subset. Resource profiles are illustrative declared bounds, not qualified serving capacity. No trace rows, derived datasets or additional silicon measurements are committed.

Both deployments complete the selected seven-request play with zero unsupported lookups. A different complete play stops at decode batch 1 / 134,719 total past-KV tokens and correctly reports incomplete coverage. The full 393-play corpus is rejected by the existing host-memory preflight before any timing queries, so corpus-wide coverage is not claimed. This validates workflow behavior and timing availability, not predictive accuracy. No GPU collection was run.

Current scope is cold aggregated replay, one client lane, HBM-only cache and no speculation. Seeded snapshots, explicit warmup and expanded AgentX replay remain follow-up work after #207/#235; neither PR is a dependency here.

Tracking

Signed-off-by: Yiming Liu <yimingl@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Sep 19, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Sep 19, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

@Arsene12358
Arsene12358 merged commit 10186e3 into feat/fpm-config-onboarding Sep 19, 2026
7 checks passed
@Arsene12358
Arsene12358 deleted the feat/fpm-agentx-collection-coverage branch September 19, 2026 14:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant