Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
4cb33d1
feat: support local systems roots in predict and recommend
Arsene12358 Sep 14, 2026
dcb73f9
feat: add independent guided FPM onboarding workflow
Arsene12358 Sep 14, 2026
3cb22fd
Merge main into guided FPM support and preserve local hardware lookups
Arsene12358 Sep 16, 2026
ecb6ee5
fix: guard guided FPM collection and recover interrupted plans
Arsene12358 Sep 16, 2026
b4af59d
fix: report unexpected FPM collection failures cleanly
Arsene12358 Sep 16, 2026
d6650fd
test: assert local FPM timing for worker overrides
Arsene12358 Sep 16, 2026
29a19ae
fix: preserve installed-wheel imports in recommendation tests
Arsene12358 Sep 16, 2026
dffb80d
feat: name model and hardware FPM onboarding explicitly
Arsene12358 Sep 16, 2026
949c0df
Merge main to align CI and release contracts
Arsene12358 Sep 16, 2026
24486a9
fix: use onboard as the sole model onboarding command
Arsene12358 Sep 16, 2026
396af2d
fix: reject nested symlinks in onboarding collection outputs
Arsene12358 Sep 16, 2026
661fd97
docs: rename onboarding guide to FPM self-service
Arsene12358 Sep 16, 2026
f5a3ac0
docs: align FPM onboarding with planned model decoupling
Arsene12358 Sep 16, 2026
1e92af9
Merge canonical estimator APIs into guided FPM onboarding
Arsene12358 Sep 18, 2026
05bf614
Merge main into guided FPM onboarding
Arsene12358 Sep 19, 2026
7f7cd55
Merge main into guided FPM onboarding
Arsene12358 Sep 22, 2026
45fa257
fix: address FPM onboarding review feedback
Arsene12358 Sep 28, 2026
bd96d2f
Merge main into feat/guided-fpm-cli
simone-chen Sep 29, 2026
6236b65
fix: preserve worker data roots in performance metadata
simone-chen Sep 29, 2026
a7c1ccd
docs: fix renamed FPM collection guide link
simone-chen Sep 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -264,6 +264,7 @@ Use the focused SDK documentation instead of treating CLI internals as public
APIs:

- [Estimator/FPE Python and Rust SDK](docs/core-api.md)
- [FPM self-service: onboard a model on target hardware](docs/fpm-self-service.md)
- [AIC-compatible modeled-power contract (semantics only)](docs/power-model.md)
- [Self-benchmarking and FPM onboarding guide](python/aisimulate/docs/fpm/self-benchmarking-and-onboarding.md)
- [Replay SDK and artifact contract](crates/core/src/replay/README.md)
Expand Down
145 changes: 145 additions & 0 deletions docs/fpm-self-service.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
<!--
SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: Apache-2.0
-->

# FPM self-service

`aisimulate onboard` guides onboarding a new model for FPM simulation on your designated hardware platform. It records the model, runtime, target GPU system and allocation, plans one pure tensor-parallel worker, and produces ordinary `predict` and `recommend` configurations that read your collected FPM data. It builds on AISimulate's existing per-worker FPM support and packaged collector.

Planning works before the model has an AISimulate model class or measured FPM timings. A valid request records your choices; model metadata and execution readiness, runtime compatibility, and data readiness remain **unchecked**, and accuracy is **not assessed**. This setup does not run target preflight checks, provision GPU resources, or establish measured accuracy.

In this revision, ordinary FPM `predict` and `recommend` still construct a registered analytical model for resource accounting and SOL-based timing transfer. A separate decoupling change will add a route that uses model/resource metadata and direct interpolation of measured timings without an op-level model class. The existing registered-model/SOL route will remain available. See [Choose the model execution route](#choose-the-model-execution-route) before handing off a new architecture.

The Inkling pilot is planned for NVIDIA GB200 with TP, DEP, and TEP configurations through the class-independent route. This guided planner currently generates pure-TP configurations only. The checkpoint, runtime, allocation, and strategy degrees must be pinned before that pilot; this guide does not establish Inkling readiness or GB200 accuracy.

## Create the request

Use an environment installed from this checkout; see [development setup](../DEVELOPMENT.md). Guided setup requires a terminal and starts only when explicitly requested:

```bash
aisimulate onboard init --interactive --output support-request.yaml
```

Enter the actual model identifier or checkpoint path, pinned model revision, dense/MoE kind, pinned vLLM version, GPU system, allocation, and interconnect. Then choose a TP size and pilot workload. Supplied options skip their prompts. Enter accepts displayed defaults; invalid values can be corrected; Ctrl-C or end-of-input cancels without saving. Existing files require `--overwrite`.

Both guided and scripted setup use onboarding. `--profile onboarding` is an optional spelling of the same behavior. For automation, supply the identity flags directly:

```bash
aisimulate onboard init \
--model /models/your-pinned-checkpoint \
--model-revision YOUR_IMMUTABLE_REVISION \
--model-kind dense \
--framework-version YOUR_PINNED_VLLM_VERSION \
--gpu h200_sxm --gpu-count 4 --interconnect nvswitch \
--tensor-parallel 2 \
--output support-request.yaml
```

Replace the model and runtime placeholders with the actual inputs. The example does not identify an Inkling checkpoint or claim that four H200s can run your model. Framework support currently selects vLLM. Optional tokenizer, chat-template, and AISimulate revisions are recorded only when supplied.

The default pilot uses 1,024 input tokens, 128 output tokens, concurrency 1, four requests, a 16,384-token context limit, TTFT target 1,000 ms, and TPOT target 100 ms. Scripted setup defaults to TP1 unless `--tensor-parallel` is supplied. These are planning defaults, not measured model capacity or latency. Change them with the corresponding flags shown by `aisimulate onboard init --help`.

One node is the default; GPUs per node then equals `--gpu-count`. For multiple nodes, supply both `--node-count` and `--gpus-per-node`; their product must equal the total allocation. Each pure-TP worker must fit on one node. Advanced flags include `--request-count`, `--max-candidates`, `--objective`, `--seed`, and `--sm`.

## Plan, preview, and explicitly execute

```bash
aisimulate onboard plan \
--config support-request.yaml --output-dir ./aisimulate-support

aisimulate onboard collect-fpm \
--config ./aisimulate-support/request.yaml \
--output-dir ./aisimulate-support
```

The first command saves the request, `support-plan.json`, `commands.json`, `predict/pilot.yaml`, `recommend/pilot.yaml`, and a local `systems/` directory. The second prints the collector command without launching it. Generated command vectors and printed next commands use absolute output paths and preserve spaces or shell punctuation. Use a separate output directory for each request. `--overwrite` can repair missing generated files for the identical saved request, including an interrupted write before `support-plan.json` was created. Repair requires an intact matching `request.yaml`, rejects changed inputs, and preserves existing collected data.

If a software update changes generated guidance, recreate the plan in a new output directory: repair compares generated files byte for byte and does not migrate older plans. Guidance changes do not change the saved-request identity checks used by collection.

The search uses one selected TP size. By default it evaluates a single worker, so recommendation is not a broad deployment search. `--max-candidates 2` additionally considers the largest count of identical workers that fits the allocation, when that differs from one worker. Each choice gets an independent recommendation config pinned to that replica count with a one-trial budget. The single worker keeps `recommend/pilot.yaml`; the second choice uses `recommend/replicas-N.yaml`, where `N` is its replica count. The plan reports the actual candidate count and lists both config and result paths. Dense collection uses the `tp` preset; MoE uses `pure_tp`; the selected TP size remains exact.

Before execution, prepare the real checkpoint and the pinned runtime using the existing [FPM collection guide](../python/aisimulate/docs/fpm/self-benchmarking-and-onboarding.md). The packaged collector invokes a Generator-resolved Dynamo/vLLM deployment and needs the corresponding GPU resources, deployment configuration, permissions, and model access. Invoking its command locally does not create that environment. `commands.json` publishes the guarded `aisimulate onboard collect-fpm --execute` command for collection, alongside a read-only collector planning command.

After those prerequisites are ready, explicitly launch collection from that environment:

```bash
aisimulate onboard collect-fpm \
--config ./aisimulate-support/request.yaml \
--output-dir ./aisimulate-support --execute
```

Execution requires a matching saved plan. The request records model and runtime revisions; this setup does not download a pinned checkpoint or check the installed runtime before execution. Keep the actual checkpoint and runtime consistent with the request before collecting or predicting. Successful formal collection and resume verify that the published data's pod-reported runtime version exactly matches `framework_version`, including suffixes such as `+cu128`. A mismatch exits 1 and preserves collection artifacts. Use the declared runtime or create a new request and plan for the observed version; generated configs are never silently retargeted. Diagnostic smoke runs do not publish or verify formal data.

Set deployment options directly on `onboard collect-fpm`: `--dynamo-version VERSION`, `--image IMAGE`, `--namespace NAME`, `--model-cache NAME[:MOUNT[:SUBPATH]]`, `--transport nvlink|ib|efa`, and `--image-pull-secret NAME`. The mount, when supplied, is an absolute container path. Prefer an immutable image digest. Supply the same options when previewing, executing, and resuming; deployment settings are part of the collector's frozen-plan identity, so changed settings require a new output directory. Arbitrary collector arguments and engine overrides are not accepted by this command.

For a diagnostic run, add `--execute --smoke`; `--limit N` also requires `--smoke`. Diagnostic smoke and limited runs do not publish formal FPM data. Existing campaign data, raw artifacts, or checkpoints require explicit `--resume` and a readable matching collector checkpoint; otherwise choose a new output directory. A custom `--checkpoint-dir`, if needed, must remain inside the plan's `fpm-checkpoint/` directory. Selecting an empty checkpoint directory does not allow reuse of existing campaign artifacts. Smoke and formal campaigns have separate checkpoints and artifact directories, so an existing smoke run does not prevent the first formal run, or vice versa. The collector verifies the resumed checkpoint's frozen-plan identity.

Planning and collection reject concurrent onboarding operations. The persistent `.support.lock` file uses an OS advisory lock; ownership is released when the process exits, including after an abrupt termination. Leave the file in place. Request validation, saved onboarding-plan checks, and collector input resolution exit 2. Failures after collector execution starts, including a frozen checkpoint identity mismatch, exit 1 with a concise message. Interruption exits 130.

The collector narrows initial prefill sampling with the pilot's input-token and concurrency bounds. Decode uses the collector's existing profile; a four-request synthetic pilot does not imply four timing samples or a short decode campaign. Inspect the generated command and collector plan before committing GPU time. Successful formal collection publishes the FPM Parquet file and metadata pair into the plan's local systems data directory; diagnostic success alone does not provide that pair.

## Run the generated ordinary configurations

After the selected model execution route is available and formal data collection is complete:

```bash
aisimulate predict \
--config ./aisimulate-support/predict/pilot.yaml \
--output-dir ./aisimulate-support/predict-results/pilot

aisimulate recommend \
--config ./aisimulate-support/recommend/pilot.yaml \
--output-dir ./aisimulate-support/recommend-results/pilot
```

With two candidates, run both replica counts. This loop derives the candidate names from the validated saved request and constructs fixed `aisimulate recommend` commands:

```bash
python3 - <<'PY'
import shlex
import subprocess
from pathlib import Path

from aisimulate.support.plan import check_plan
from aisimulate.support.schema import SupportRequest

root = Path("./aisimulate-support").resolve()
request = SupportRequest.from_yaml(root / "request.yaml")
check_plan(request, root)
max_replicas = request.identity.node_count * (request.identity.gpus_per_node // request.search.tensor_parallel)
replicas = list(dict.fromkeys((1, max_replicas)))[:request.search.max_candidates]
statuses = []
for count in replicas:
name = "pilot" if count == 1 else f"replicas-{count}"
command = [
"aisimulate", "recommend",
"--config", str(root / f"recommend/{name}.yaml"),
"--output-dir", str(root / f"recommend-results/{name}"),
"--format", "json",
]
result = subprocess.run(command, check=False)
Comment thread
coderabbitai[bot] marked this conversation as resolved.
Comment thread
simone-chen marked this conversation as resolved.
statuses.append(result.returncode)
print(f"exit {result.returncode}: {shlex.join(command)}", flush=True)
raise SystemExit(1 if any(statuses) else 0)
PY
```

The loop reports each command's exit status and attempts all commands, including when the pilot finds no feasible candidate. It exits 1 after all attempts if any command returned a nonzero status, otherwise 0. Successful outputs remain usable even when the loop exits 1; inspect each result before comparing candidates.

Run either the individual recommendation command or the loop against fresh result directories. The loop also handles a one-candidate plan and does not execute entries from `commands.json`. Results remain separate under `recommend-results/pilot` and, when present, `recommend-results/replicas-N`; compare their objective and latency results for the same workload. This plan does not produce a combined ranking or search additional TP sizes or scheduler settings.

The generated configurations select `engine.workers.aggregated.timing.estimation_mode: fpm_interpolation` with `fallback_policy: deny` and set `engine.systems_paths` to a list containing the plan's absolute local systems directory. The same root supplies hardware and collected FPM data. Recommendation preserves it in exported prediction configs. Moving the plan to another machine requires updating absolute paths or regenerating it there. The single-root `engine.systems_path` input is an alias added by onboarding; saved configurations use the canonical `engine.systems_paths` list. Both spellings resolve relative local paths against the working directory when the configuration is loaded, so saving and reloading from another directory preserves the selected roots. The canonical list's `default` entry continues to select packaged data. A worker’s `timing.systems_paths` overrides engine-level roots; performance metadata records those effective roots. Empty or whitespace-only roots are rejected by the shared root type in both prediction and recommendation configurations.

These configurations target AISimulate's standalone `predict` and `recommend` commands. Dynamo's replay Planner AIC session adapter does not currently forward custom systems roots; using that downstream path requires a separate adapter update.

Ordinary `predict` and `recommend` commands retain their existing behavior and defaults. Their generated configs can be loaded and edited through the public configuration schema. In this revision, missing model registration or data may still prevent execution; a successful simulation is not an accuracy result. Compare its output with an independent run of the same model, runtime, topology, and workload to assess accuracy.

## Choose the model execution route

Hand off the saved request, pinned model configuration, and plan. Both routes need a canonical checkpoint identity, effective precision and topology, correct weight and KV-cache accounting, and matching whole-forward FPM measurements. Collected timings alone do not establish memory fit.

- **Registered-model/SOL route, available in this revision:** reuse a compatible analytical class or follow [How to Add a New Model](../python/aisimulate/docs/add_a_new_model.md) when choosing to add one. Verify its operation graph, memory and cache accounting, and native FPM SOL execution. A dedicated class is an option for this route, not the intended prerequisite for every FPM onboarding.
- **Class-independent route, planned in the separate decoupling change:** resolve model and resource metadata without constructing an operation graph. Use direct measured-time interpolation, including wider two-sided KV brackets at the same batch size when both neighboring prompt curves cover the query. Missing brackets or unsupported metadata must produce explicit errors. This route is not implemented by the guided foundation; 2D interpolation remains experimental.

Per-operation silicon profiling described in the model guide is not required by either FPM route. The intended self-service workflow collects whole-forward timings, then verifies prediction and recommendation for the exact target deployment. Collection bootstrap can already resolve some unregistered model configurations; that does not establish that this revision's ordinary FPM prediction path can construct them.
2 changes: 1 addition & 1 deletion python/aisimulate/collector/fpm_forward/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@
from .config import add_fpm_arguments, add_fpm_generator_arguments
from .entry import resolve_inputs, resolve_run_inputs, run_resolved

_INPUT_ERRORS = (FileNotFoundError, RuntimeError, TypeError, ValueError)
_INPUT_ERRORS = (OSError, RuntimeError, TypeError, ValueError)


def _parser() -> argparse.ArgumentParser:
Expand Down
2 changes: 1 addition & 1 deletion python/aisimulate/src/aisimulate/capacity.py
Original file line number Diff line number Diff line change
Expand Up @@ -252,7 +252,7 @@ def estimate_num_gpu_blocks(
attention_backend: str | None = None,
enable_eplb: bool = False,
wideep_num_slots: int | None = None,
systems_path: str | None = None,
systems_path: str | list[str] | None = None,
cuda_graph_reserved_bytes: int = 0,
diagnostics: dict[str, Any] | None = None,
) -> int:
Expand Down
2 changes: 2 additions & 0 deletions python/aisimulate/src/aisimulate/cli_args.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@
import yaml

from .detail import parse_detail_sections
from .support.cli import add_support_parser


class _CliConfigError(ValueError):
Expand Down Expand Up @@ -75,6 +76,7 @@ def build_parser() -> argparse.ArgumentParser:
action="store_true",
help="pace prediction against the real wall clock instead of virtual time",
)
add_support_parser(subparsers)
return parser


Expand Down
8 changes: 8 additions & 0 deletions python/aisimulate/src/aisimulate/compiler.py
Original file line number Diff line number Diff line change
Expand Up @@ -420,6 +420,9 @@ def _worker_performance_model_metadata(
if value is not None:
config[field] = value
config["database_mode"] = worker.timing.database_mode or engine.database_mode
systems_paths = worker.timing.systems_paths or engine.systems_paths
if systems_paths is not None:
config["systems_paths"] = systems_paths
if engine.speculation is not None:
config["speculation"] = engine.speculation.cost_config()
if worker.timing.fpm_parquet_path is not None:
Expand Down Expand Up @@ -488,6 +491,10 @@ def _worker_engine_args(
payload["speculation"] = engine.speculation.model_dump(mode="json")
if engine.backend_version is not None:
payload["aic_backend_version"] = engine.backend_version
if engine.systems_paths is not None and worker.timing.type != "default":
from .sweeper.forward_pass_estimator import resolve_systems_paths

payload["systems_path"] = list(resolve_systems_paths(engine.systems_paths))
if engine.decoder_replay:
payload["aic_decoder_replay"] = True
for field in ("database_mode", "enable_shared_layer", "strict_provenance"):
Expand Down Expand Up @@ -543,6 +550,7 @@ def _worker_engine_args(
if capacity.type == "default" and cache.state_cache is None:
payload = materialize_aic_num_gpu_blocks(payload)
for name in (
"systems_path",
"aic_backend_version",
"aic_system",
"aic_model_path",
Expand Down
22 changes: 21 additions & 1 deletion python/aisimulate/src/aisimulate/config/common.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,13 +11,33 @@
from typing import Annotated, Any, Generic, Literal, TypeVar

import yaml
from pydantic import BaseModel, ConfigDict, Field, field_validator, model_validator
from pydantic import AfterValidator, BaseModel, ConfigDict, Field, field_validator, model_validator


class StrictModel(BaseModel):
model_config = ConfigDict(extra="forbid")


def _absolute_systems_path(value: str) -> str:
if not value.strip() or "," in value:
raise ValueError("systems_path must name one nonempty local directory, not a comma-separated list")
return str(Path(value).expanduser().resolve())


SystemsPath = Annotated[str, Field(strict=True, min_length=1), AfterValidator(_absolute_systems_path)]


def _normalize_systems_root(value: str) -> str:
if not value.strip():
raise ValueError("systems_paths entries must be nonempty")
if value.lower() == "default":
return "default"
return str(Path(value).expanduser().resolve())


SystemsRoot = Annotated[str, Field(strict=True, min_length=1), AfterValidator(_normalize_systems_root)]


T = TypeVar("T")
PositiveFiniteFloat = Annotated[float, Field(strict=True, gt=0, allow_inf_nan=False)]
PositiveStrictInt = Annotated[int, Field(strict=True, gt=0)]
Expand Down
Loading
Loading