Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 18 additions & 5 deletions THIRD_PARTY_NOTICES.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,19 @@ This material is licensed under the Apache License 2.0. The upstream license
at the identified revision is available at:
https://github.com/ai-dynamo/aiconfigurator/blob/915f590680d8a79fe9c39f6f3a9ff13bc267fcce/LICENSE

## Dynamo V4.1 FPM collection adapter

`collector/fpm_forward/runtime/dsv41/dsv41_scheduler.py` is modified code
adapted from `components/src/dynamo/vllm/instrumented_scheduler.py` in
https://github.com/ai-dynamo/dynamo/tree/54960177085413259859c88bd34ed0734d4c2ea9.
It adds bounded same-request real-KV collection while preserving the native
benchmark and FPM contracts. Copyright (c) 2025-2026 NVIDIA CORPORATION &
AFFILIATES. All rights reserved. Licensed under Apache-2.0; the upstream
license is preserved in the adapter's adjacent `LICENSE`. The adjacent README
records the inspected vLLM API revision and immutable runtime image/source
hashes. vLLM implementation files are not vendored. The text fixture and
lifecycle tests are original work for this change, with no external corpus.

## NVIDIA AIConfigurator speculative decoding

The speculation SDK, compatibility exports, CLI/task integration, attention and whole-forward FPM operation changes, native bindings, and their tests are adapted and modified from AIConfigurator PR #1563, pinned at commit `6290c161a354da5250c391bd43372b2e9c6f4a51`. Original paths are under `aic-core/src/aiconfigurator_core/sdk/`, `src/aiconfigurator/`, `aic-core/rust/aiconfigurator-core/`, `aic-core/rust/tests/public-api/`, and `tests/`.
Expand Down Expand Up @@ -535,7 +548,7 @@ SOFTWARE.

## Meta Muse Glimmer model configuration

`src/aiconfigurator_core/model_configs/meta-models--Muse-Glimmer-30B_config.json`
`src/aisimulate_core/model_configs/meta-models--Muse-Glimmer-30B_config.json`
is an unmodified copy of `config.json` from the Meta Muse Glimmer model
repository at immutable revision
`f84ecc3a0ea984a4c04542a84269e3d065350a6e`:
Expand All @@ -557,8 +570,8 @@ named Qwen model repositories at the immutable revisions shown:

| Packaged file | Upstream revision |
| --- | --- |
| `aiconfigurator_core/model_configs/Qwen--Qwen3.8-2.4T-A95B_config.json` | `Qwen/Qwen3.8-2.4T-A95B@207bd685a7e3696cfaff12ded7c6a7ea0f88c996` |
| `aiconfigurator_core/model_configs/Qwen--Qwen3.8-2.4T-A95B-FP8_config.json` | `Qwen/Qwen3.8-2.4T-A95B-FP8@d2dc35658bcf77e66643428cb52e774cc3b5bd29` |
| `aisimulate_core/model_configs/Qwen--Qwen3.8-2.4T-A95B_config.json` | `Qwen/Qwen3.8-2.4T-A95B@207bd685a7e3696cfaff12ded7c6a7ea0f88c996` |
| `aisimulate_core/model_configs/Qwen--Qwen3.8-2.4T-A95B-FP8_config.json` | `Qwen/Qwen3.8-2.4T-A95B-FP8@d2dc35658bcf77e66643428cb52e774cc3b5bd29` |

Upstream repositories:
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B and
Expand Down Expand Up @@ -592,10 +605,10 @@ For any questions regarding this license, please contact model-business@notice.q
The following bundled model configs are modified copies of Meta Llama 4
checkpoint configuration files:

- `src/aiconfigurator_core/model_configs/meta-llama--Llama-4-Scout-17B-16E-Instruct_config.json`
- `src/aisimulate_core/model_configs/meta-llama--Llama-4-Scout-17B-16E-Instruct_config.json`
from revision `92f3b1597a195b523d8d9e5700e57e4fbb8f20d3`:
https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct/blob/92f3b1597a195b523d8d9e5700e57e4fbb8f20d3/config.json
- `src/aiconfigurator_core/model_configs/meta-llama--Llama-4-Maverick-17B-128E-Instruct_config.json`
- `src/aisimulate_core/model_configs/meta-llama--Llama-4-Maverick-17B-128E-Instruct_config.json`
from revision `73d14711bcc77c16df3470856949c3764056b617`:
https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct/blob/73d14711bcc77c16df3470856949c3764056b617/config.json

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -2553,7 +2553,9 @@ def test_fpm_spec_tags(self, fpm_systems_root, monkeypatch):
ctx_op = spec["context_ops"][0]["FpmForward"]
assert ctx_op["phase"] == "prefill"
assert spec["generation_ops"][0]["FpmForward"]["phase"] == "decode"
assert len(ctx_op["match_identity"]) == 15
assert len(ctx_op["match_identity"]) == 19
assert ctx_op["match_identity"][-4:] == ["", "full", "none", "text"]
assert spec["generation_ops"][0]["FpmForward"]["match_identity"] == ctx_op["match_identity"]
assert ctx_op["sol_ops"], "sol_ops must carry the original granular list"

@pytest.mark.parametrize(
Expand Down
6 changes: 4 additions & 2 deletions crates/core/src/engine/common/perf_model.rs
Original file line number Diff line number Diff line change
Expand Up @@ -42,10 +42,12 @@ impl std::fmt::Debug for PerfModel {
}

impl PerfModel {
pub(crate) fn prefill_batch_validation_can_fail(&self) -> bool {
// External prediction and duration conversion can fail even when geometry
// validation accepts every batch. Keep admission transactional for all providers.
pub(crate) fn prefill_pass_can_fail(&self) -> bool {
match self {
Self::Polynomial => false,
Self::External { timing } => timing.prefill_batch_validation_can_fail(),
Self::External { .. } => true,
}
}

Expand Down
8 changes: 4 additions & 4 deletions crates/core/src/engine/scheduler/sglang/core.rs
Original file line number Diff line number Diff line change
Expand Up @@ -778,11 +778,11 @@ impl SglangCore {
}
}
}
// Only providers with fallible geometry validation need to preserve the
// admission state. Normal polynomial and unrestricted AIC passes avoid
// copying radix metadata. Lease checkpoints never become independent owners.
// Preserve admission around validation, prediction and duration conversion.
// The built-in polynomial remains infallible; external providers may fail
// after accepting geometry. Lease checkpoints are not independent owners.
let admission_checkpoint = (!self.waiting.is_empty()
&& self.config.perf_model.prefill_batch_validation_can_fail())
&& self.config.perf_model.prefill_pass_can_fail())
.then(|| {
let waiting = self
.waiting
Expand Down
4 changes: 4 additions & 0 deletions crates/core/src/engine/scheduler/sglang/tests.rs
Original file line number Diff line number Diff line change
Expand Up @@ -2658,6 +2658,10 @@ mod admission_validation_rollback {
}

impl crate::engine::TimingModel for FallibleTiming {
fn prefill_batch_validation_can_fail(&self) -> bool {
!self.fail_in_prediction
}

fn validate_prefill_batch(&self, _: &[(usize, usize)]) -> anyhow::Result<()> {
anyhow::ensure!(
self.fail_in_prediction || !self.fail.load(Ordering::Relaxed),
Expand Down
7 changes: 4 additions & 3 deletions crates/core/src/engine/timing.rs
Original file line number Diff line number Diff line change
Expand Up @@ -447,9 +447,10 @@ pub enum TimingModelConfig {
/// Implementations may call AIC, interpolate profiler data, or use another
/// provider without adding that dependency to `aisimulate-core`.
pub trait TimingModel: Send + Sync {
/// Whether admission needs a checkpoint around `validate_prefill_batch`.
/// Existing custom providers default to the safe, fallible contract. Providers
/// opting out must accept every batch geometry in that validation hook.
/// Whether `validate_prefill_batch` may reject geometry. A false return
/// guarantees only that validation accepts every batch, not that prediction
/// or duration conversion is infallible. Admission must still protect those
/// later operations for external providers.
fn prefill_batch_validation_can_fail(&self) -> bool {
true
}
Expand Down
8 changes: 7 additions & 1 deletion crates/core/src/perfmodel/config.rs
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,10 @@ pub const ENGINE_CONFIG_SCHEMA_VERSION: u32 = 1;
// - 19 (DeepSeek-V4.1 review): Dsv41AttentionOp gained kv_cache_layout,
// separating physical backend KV payload from attention arithmetic precision.
// Its appended enum changes positional bincode layout; old JSON defaults only.
pub const ENGINE_SPEC_SCHEMA_VERSION: u32 = 19;
// - 20 (DeepSeek-V4.1 FPM): FpmForwardOp gained original_fmha_quant_mode
// for selector diagnostics. This appends a positional field after the schema-19
// release; serde defaults support legacy JSON, not legacy bincode.
pub const ENGINE_SPEC_SCHEMA_VERSION: u32 = 20;

/// Static engine identity and setup information carried by an
/// [`crate::perfmodel::engine::spec::EngineSpec`].
Expand Down Expand Up @@ -254,6 +257,9 @@ pub struct QuantizationConfig {
#[serde(default)]
pub moe_dtype: Option<DataType>,
pub activation_dtype: Option<DataType>,
/// FPM cell selector only; does not override model arithmetic or memory.
#[serde(default)]
pub fpm_fmha_dtype: Option<DataType>,
pub kv_cache_dtype: Option<DataType>,
}

Expand Down
Loading
Loading