Repository navigation
feat: onboard models for FPM simulation on target hardware #196
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
20 commits
Select commit
Hold shift + click to select a range
4cb33d1
feat: support local systems roots in predict and recommend
Arsene12358 dcb73f9
feat: add independent guided FPM onboarding workflow
Arsene12358 3cb22fd
Merge main into guided FPM support and preserve local hardware lookups
Arsene12358 ecb6ee5
fix: guard guided FPM collection and recover interrupted plans
Arsene12358 b4af59d
fix: report unexpected FPM collection failures cleanly
Arsene12358 d6650fd
test: assert local FPM timing for worker overrides
Arsene12358 29a19ae
fix: preserve installed-wheel imports in recommendation tests
Arsene12358 dffb80d
feat: name model and hardware FPM onboarding explicitly
Arsene12358 949c0df
Merge main to align CI and release contracts
Arsene12358 24486a9
fix: use onboard as the sole model onboarding command
Arsene12358 396af2d
fix: reject nested symlinks in onboarding collection outputs
Arsene12358 661fd97
docs: rename onboarding guide to FPM self-service
Arsene12358 f5a3ac0
docs: align FPM onboarding with planned model decoupling
Arsene12358 1e92af9
Merge canonical estimator APIs into guided FPM onboarding
Arsene12358 05bf614
Merge main into guided FPM onboarding
Arsene12358 7f7cd55
Merge main into guided FPM onboarding
Arsene12358 45fa257
fix: address FPM onboarding review feedback
Arsene12358 bd96d2f
Merge main into feat/guided-fpm-cli
simone-chen 6236b65
fix: preserve worker data roots in performance metadata
simone-chen a7c1ccd
docs: fix renamed FPM collection guide link
simone-chen File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,145 @@ | ||
| <!-- | ||
| SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. | ||
| SPDX-License-Identifier: Apache-2.0 | ||
| --> | ||
|
|
||
| # FPM self-service | ||
|
|
||
| `aisimulate onboard` guides onboarding a new model for FPM simulation on your designated hardware platform. It records the model, runtime, target GPU system and allocation, plans one pure tensor-parallel worker, and produces ordinary `predict` and `recommend` configurations that read your collected FPM data. It builds on AISimulate's existing per-worker FPM support and packaged collector. | ||
|
|
||
| Planning works before the model has an AISimulate model class or measured FPM timings. A valid request records your choices; model metadata and execution readiness, runtime compatibility, and data readiness remain **unchecked**, and accuracy is **not assessed**. This setup does not run target preflight checks, provision GPU resources, or establish measured accuracy. | ||
|
|
||
| In this revision, ordinary FPM `predict` and `recommend` still construct a registered analytical model for resource accounting and SOL-based timing transfer. A separate decoupling change will add a route that uses model/resource metadata and direct interpolation of measured timings without an op-level model class. The existing registered-model/SOL route will remain available. See [Choose the model execution route](#choose-the-model-execution-route) before handing off a new architecture. | ||
|
|
||
| The Inkling pilot is planned for NVIDIA GB200 with TP, DEP, and TEP configurations through the class-independent route. This guided planner currently generates pure-TP configurations only. The checkpoint, runtime, allocation, and strategy degrees must be pinned before that pilot; this guide does not establish Inkling readiness or GB200 accuracy. | ||
|
|
||
| ## Create the request | ||
|
|
||
| Use an environment installed from this checkout; see [development setup](../DEVELOPMENT.md). Guided setup requires a terminal and starts only when explicitly requested: | ||
|
|
||
| ```bash | ||
| aisimulate onboard init --interactive --output support-request.yaml | ||
| ``` | ||
|
|
||
| Enter the actual model identifier or checkpoint path, pinned model revision, dense/MoE kind, pinned vLLM version, GPU system, allocation, and interconnect. Then choose a TP size and pilot workload. Supplied options skip their prompts. Enter accepts displayed defaults; invalid values can be corrected; Ctrl-C or end-of-input cancels without saving. Existing files require `--overwrite`. | ||
|
|
||
| Both guided and scripted setup use onboarding. `--profile onboarding` is an optional spelling of the same behavior. For automation, supply the identity flags directly: | ||
|
|
||
| ```bash | ||
| aisimulate onboard init \ | ||
| --model /models/your-pinned-checkpoint \ | ||
| --model-revision YOUR_IMMUTABLE_REVISION \ | ||
| --model-kind dense \ | ||
| --framework-version YOUR_PINNED_VLLM_VERSION \ | ||
| --gpu h200_sxm --gpu-count 4 --interconnect nvswitch \ | ||
| --tensor-parallel 2 \ | ||
| --output support-request.yaml | ||
| ``` | ||
|
|
||
| Replace the model and runtime placeholders with the actual inputs. The example does not identify an Inkling checkpoint or claim that four H200s can run your model. Framework support currently selects vLLM. Optional tokenizer, chat-template, and AISimulate revisions are recorded only when supplied. | ||
|
|
||
| The default pilot uses 1,024 input tokens, 128 output tokens, concurrency 1, four requests, a 16,384-token context limit, TTFT target 1,000 ms, and TPOT target 100 ms. Scripted setup defaults to TP1 unless `--tensor-parallel` is supplied. These are planning defaults, not measured model capacity or latency. Change them with the corresponding flags shown by `aisimulate onboard init --help`. | ||
|
|
||
| One node is the default; GPUs per node then equals `--gpu-count`. For multiple nodes, supply both `--node-count` and `--gpus-per-node`; their product must equal the total allocation. Each pure-TP worker must fit on one node. Advanced flags include `--request-count`, `--max-candidates`, `--objective`, `--seed`, and `--sm`. | ||
|
|
||
| ## Plan, preview, and explicitly execute | ||
|
|
||
| ```bash | ||
| aisimulate onboard plan \ | ||
| --config support-request.yaml --output-dir ./aisimulate-support | ||
|
|
||
| aisimulate onboard collect-fpm \ | ||
| --config ./aisimulate-support/request.yaml \ | ||
| --output-dir ./aisimulate-support | ||
| ``` | ||
|
|
||
| The first command saves the request, `support-plan.json`, `commands.json`, `predict/pilot.yaml`, `recommend/pilot.yaml`, and a local `systems/` directory. The second prints the collector command without launching it. Generated command vectors and printed next commands use absolute output paths and preserve spaces or shell punctuation. Use a separate output directory for each request. `--overwrite` can repair missing generated files for the identical saved request, including an interrupted write before `support-plan.json` was created. Repair requires an intact matching `request.yaml`, rejects changed inputs, and preserves existing collected data. | ||
|
|
||
| If a software update changes generated guidance, recreate the plan in a new output directory: repair compares generated files byte for byte and does not migrate older plans. Guidance changes do not change the saved-request identity checks used by collection. | ||
|
|
||
| The search uses one selected TP size. By default it evaluates a single worker, so recommendation is not a broad deployment search. `--max-candidates 2` additionally considers the largest count of identical workers that fits the allocation, when that differs from one worker. Each choice gets an independent recommendation config pinned to that replica count with a one-trial budget. The single worker keeps `recommend/pilot.yaml`; the second choice uses `recommend/replicas-N.yaml`, where `N` is its replica count. The plan reports the actual candidate count and lists both config and result paths. Dense collection uses the `tp` preset; MoE uses `pure_tp`; the selected TP size remains exact. | ||
|
|
||
| Before execution, prepare the real checkpoint and the pinned runtime using the existing [FPM collection guide](../python/aisimulate/docs/fpm/self-benchmarking-and-onboarding.md). The packaged collector invokes a Generator-resolved Dynamo/vLLM deployment and needs the corresponding GPU resources, deployment configuration, permissions, and model access. Invoking its command locally does not create that environment. `commands.json` publishes the guarded `aisimulate onboard collect-fpm --execute` command for collection, alongside a read-only collector planning command. | ||
|
|
||
| After those prerequisites are ready, explicitly launch collection from that environment: | ||
|
|
||
| ```bash | ||
| aisimulate onboard collect-fpm \ | ||
| --config ./aisimulate-support/request.yaml \ | ||
| --output-dir ./aisimulate-support --execute | ||
| ``` | ||
|
|
||
| Execution requires a matching saved plan. The request records model and runtime revisions; this setup does not download a pinned checkpoint or check the installed runtime before execution. Keep the actual checkpoint and runtime consistent with the request before collecting or predicting. Successful formal collection and resume verify that the published data's pod-reported runtime version exactly matches `framework_version`, including suffixes such as `+cu128`. A mismatch exits 1 and preserves collection artifacts. Use the declared runtime or create a new request and plan for the observed version; generated configs are never silently retargeted. Diagnostic smoke runs do not publish or verify formal data. | ||
|
|
||
| Set deployment options directly on `onboard collect-fpm`: `--dynamo-version VERSION`, `--image IMAGE`, `--namespace NAME`, `--model-cache NAME[:MOUNT[:SUBPATH]]`, `--transport nvlink|ib|efa`, and `--image-pull-secret NAME`. The mount, when supplied, is an absolute container path. Prefer an immutable image digest. Supply the same options when previewing, executing, and resuming; deployment settings are part of the collector's frozen-plan identity, so changed settings require a new output directory. Arbitrary collector arguments and engine overrides are not accepted by this command. | ||
|
|
||
| For a diagnostic run, add `--execute --smoke`; `--limit N` also requires `--smoke`. Diagnostic smoke and limited runs do not publish formal FPM data. Existing campaign data, raw artifacts, or checkpoints require explicit `--resume` and a readable matching collector checkpoint; otherwise choose a new output directory. A custom `--checkpoint-dir`, if needed, must remain inside the plan's `fpm-checkpoint/` directory. Selecting an empty checkpoint directory does not allow reuse of existing campaign artifacts. Smoke and formal campaigns have separate checkpoints and artifact directories, so an existing smoke run does not prevent the first formal run, or vice versa. The collector verifies the resumed checkpoint's frozen-plan identity. | ||
|
|
||
| Planning and collection reject concurrent onboarding operations. The persistent `.support.lock` file uses an OS advisory lock; ownership is released when the process exits, including after an abrupt termination. Leave the file in place. Request validation, saved onboarding-plan checks, and collector input resolution exit 2. Failures after collector execution starts, including a frozen checkpoint identity mismatch, exit 1 with a concise message. Interruption exits 130. | ||
|
|
||
| The collector narrows initial prefill sampling with the pilot's input-token and concurrency bounds. Decode uses the collector's existing profile; a four-request synthetic pilot does not imply four timing samples or a short decode campaign. Inspect the generated command and collector plan before committing GPU time. Successful formal collection publishes the FPM Parquet file and metadata pair into the plan's local systems data directory; diagnostic success alone does not provide that pair. | ||
|
|
||
| ## Run the generated ordinary configurations | ||
|
|
||
| After the selected model execution route is available and formal data collection is complete: | ||
|
|
||
| ```bash | ||
| aisimulate predict \ | ||
| --config ./aisimulate-support/predict/pilot.yaml \ | ||
| --output-dir ./aisimulate-support/predict-results/pilot | ||
|
|
||
| aisimulate recommend \ | ||
| --config ./aisimulate-support/recommend/pilot.yaml \ | ||
| --output-dir ./aisimulate-support/recommend-results/pilot | ||
| ``` | ||
|
|
||
| With two candidates, run both replica counts. This loop derives the candidate names from the validated saved request and constructs fixed `aisimulate recommend` commands: | ||
|
|
||
| ```bash | ||
| python3 - <<'PY' | ||
| import shlex | ||
| import subprocess | ||
| from pathlib import Path | ||
|
|
||
| from aisimulate.support.plan import check_plan | ||
| from aisimulate.support.schema import SupportRequest | ||
|
|
||
| root = Path("./aisimulate-support").resolve() | ||
| request = SupportRequest.from_yaml(root / "request.yaml") | ||
| check_plan(request, root) | ||
| max_replicas = request.identity.node_count * (request.identity.gpus_per_node // request.search.tensor_parallel) | ||
| replicas = list(dict.fromkeys((1, max_replicas)))[:request.search.max_candidates] | ||
| statuses = [] | ||
| for count in replicas: | ||
| name = "pilot" if count == 1 else f"replicas-{count}" | ||
| command = [ | ||
| "aisimulate", "recommend", | ||
| "--config", str(root / f"recommend/{name}.yaml"), | ||
| "--output-dir", str(root / f"recommend-results/{name}"), | ||
| "--format", "json", | ||
| ] | ||
| result = subprocess.run(command, check=False) | ||
|
simone-chen marked this conversation as resolved.
|
||
| statuses.append(result.returncode) | ||
| print(f"exit {result.returncode}: {shlex.join(command)}", flush=True) | ||
| raise SystemExit(1 if any(statuses) else 0) | ||
| PY | ||
| ``` | ||
|
|
||
| The loop reports each command's exit status and attempts all commands, including when the pilot finds no feasible candidate. It exits 1 after all attempts if any command returned a nonzero status, otherwise 0. Successful outputs remain usable even when the loop exits 1; inspect each result before comparing candidates. | ||
|
|
||
| Run either the individual recommendation command or the loop against fresh result directories. The loop also handles a one-candidate plan and does not execute entries from `commands.json`. Results remain separate under `recommend-results/pilot` and, when present, `recommend-results/replicas-N`; compare their objective and latency results for the same workload. This plan does not produce a combined ranking or search additional TP sizes or scheduler settings. | ||
|
|
||
| The generated configurations select `engine.workers.aggregated.timing.estimation_mode: fpm_interpolation` with `fallback_policy: deny` and set `engine.systems_paths` to a list containing the plan's absolute local systems directory. The same root supplies hardware and collected FPM data. Recommendation preserves it in exported prediction configs. Moving the plan to another machine requires updating absolute paths or regenerating it there. The single-root `engine.systems_path` input is an alias added by onboarding; saved configurations use the canonical `engine.systems_paths` list. Both spellings resolve relative local paths against the working directory when the configuration is loaded, so saving and reloading from another directory preserves the selected roots. The canonical list's `default` entry continues to select packaged data. A worker’s `timing.systems_paths` overrides engine-level roots; performance metadata records those effective roots. Empty or whitespace-only roots are rejected by the shared root type in both prediction and recommendation configurations. | ||
|
|
||
| These configurations target AISimulate's standalone `predict` and `recommend` commands. Dynamo's replay Planner AIC session adapter does not currently forward custom systems roots; using that downstream path requires a separate adapter update. | ||
|
|
||
| Ordinary `predict` and `recommend` commands retain their existing behavior and defaults. Their generated configs can be loaded and edited through the public configuration schema. In this revision, missing model registration or data may still prevent execution; a successful simulation is not an accuracy result. Compare its output with an independent run of the same model, runtime, topology, and workload to assess accuracy. | ||
|
|
||
| ## Choose the model execution route | ||
|
|
||
| Hand off the saved request, pinned model configuration, and plan. Both routes need a canonical checkpoint identity, effective precision and topology, correct weight and KV-cache accounting, and matching whole-forward FPM measurements. Collected timings alone do not establish memory fit. | ||
|
|
||
| - **Registered-model/SOL route, available in this revision:** reuse a compatible analytical class or follow [How to Add a New Model](../python/aisimulate/docs/add_a_new_model.md) when choosing to add one. Verify its operation graph, memory and cache accounting, and native FPM SOL execution. A dedicated class is an option for this route, not the intended prerequisite for every FPM onboarding. | ||
| - **Class-independent route, planned in the separate decoupling change:** resolve model and resource metadata without constructing an operation graph. Use direct measured-time interpolation, including wider two-sided KV brackets at the same batch size when both neighboring prompt curves cover the query. Missing brackets or unsupported metadata must produce explicit errors. This route is not implemented by the guided foundation; 2D interpolation remains experimental. | ||
|
|
||
| Per-operation silicon profiling described in the model guide is not required by either FPM route. The intended self-service workflow collects whole-forward timings, then verifies prediction and recommendation for the exact target deployment. Collection bootstrap can already resolve some unregistered model configurations; that does not establish that this revision's ordinary FPM prediction path can construct them. | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.