Skip to content

[Perf][Cosmos3] Add SeaCache support for Cosmos3 - #6922

Merged
linyueqian merged 21 commits into
vllm-project:mainfrom
yzhautouskay:yzhautouskay/cosmos3_diffusion_caching
Sep 15, 2026
Merged

linyueqian merged 21 commits into
vllm-project:mainfrom
yzhautouskay:yzhautouskay/cosmos3_diffusion_caching

Conversation

@yzhautouskay

@yzhautouskay yzhautouskay commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add a native SeaCache backend for Cosmos3. Paper: SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion Models
  • Use spectral filtering to compare the most informative latent features when deciding whether to skip and approximate a denoising step.
  • Unlike TeaCache, it requires no model-specific coefficient fitting
  • Keep SeaCache opt-in initially, with the goal of enabling it by default after broader quality validation.

Output videos

I2V

vllm_omni_i2v_baseline_seacache_metrics_synced.mp4

T2V

vllm_omni_t2v_baseline_seacache_metrics_synced.mp4

Performance and fidelity

Both runs use Cosmos3-Nano with 35 steps, guidance 6, shift 10, 1280×720 output, 189 frames at 24 FPS, and regional torch.compile. SeaCache uses threshold 0.25, residual order 1, maximum 2 consecutive cached calls, and power 3.

Mode Baseline E2E SeaCache E2E Speedup LPIPS ↓
I2V 92.655 s 51.533 s 1.798× 0.11570
T2V 92.906 s 51.091 s 1.818× 0.23509

LPIPS is computed across all 189 frames against the corresponding uncached output.

Reproduction commands

Both benchmarks used one GB200 with regional torch.compile enabled.

Baseline server

CUDA_VISIBLE_DEVICES=0 vllm serve nvidia/Cosmos3-Nano \
  --omni --host 0.0.0.0 --port 8000 --init-timeout 1800 \
  --cache-backend none --no-guardrails

SeaCache server

CUDA_VISIBLE_DEVICES=0 vllm serve nvidia/Cosmos3-Nano \
  --omni --host 0.0.0.0 --port 8000 --init-timeout 1800 \
  --cache-backend sea_cache \
  --cache-config '{"sea_threshold":0.25,"sea_residual_order":1,"sea_max_consecutive_cached":2,"sea_power_exp":3.0}' \
  --no-guardrails

The same request was executed once against each server.

T2V request

curl -sS -X POST http://localhost:8000/v1/videos/sync \
  -H "Accept: video/mp4" \
  -F "model=nvidia/Cosmos3-Nano" \
  -F "prompt=A robot arm is cleaning a plate in the kitchen" \
  -F "negative_prompt=blurry, distorted, low quality, jittery, deformed" \
  -F "size=1280x720" \
  -F "num_frames=189" \
  -F "fps=24" \
  -F "num_inference_steps=35" \
  -F "guidance_scale=6.0" \
  -F "max_sequence_length=4096" \
  -F "flow_shift=10.0" \
  -F "seed=123" \
  -F 'extra_params={"use_resolution_template":false,"use_duration_template":false}' \
  -o t2v.mp4

I2V request

curl -sS -X POST http://localhost:8000/v1/videos/sync \
  -H "Accept: video/mp4" \
  -F "model=nvidia/Cosmos3-Nano" \
  -F "prompt=The scene comes to life with smooth, natural motion." \
  -F "negative_prompt=blurry, distorted, low quality" \
  -F "size=1280x720" \
  -F "num_frames=189" \
  -F "fps=24" \
  -F "num_inference_steps=35" \
  -F "guidance_scale=6.0" \
  -F "max_sequence_length=4096" \
  -F "flow_shift=10.0" \
  -F "seed=1111" \
  -F 'extra_params={"use_resolution_template":false,"use_duration_template":false}' \
  -F "input_reference=@Cosmos3-Nano/assets/example_i2v_input.jpg;type=image/jpeg" \
  -o i2v.mp4

Test Plan

vLLM Version: 0.28.0
vLLM-Omni Commit: 3db6064e2561b19d449e84f2e5d7af24bd54f8bb

  • Run scoped pre-commit checks.
  • Run Cosmos3 and diffusion-cache tests.
  • Compare cache-disabled and SeaCache T2V/I2V generation with regional torch.compile.
  • Measure end-to-end latency and exact-frame PSNR, SSIM, and LPIPS.

Test Result

python -m pytest -sv tests/diffusion/models/cosmos3/ tests/diffusion/cache/ -m "core_model and cpu"
  • Pre-commit checks passed.
  • 309 tests passed; SeaCache suite rerun: 10 passed.

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/cache_management.md.

Module owners: @Isotr0py @princepride @SamitHuang

@yzhautouskay, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 1, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Resolved as of cad97dc31d44: the high-priority or low-quality signal noted on an earlier commit no longer applies.

@yzhautouskay yzhautouskay changed the title Add SeaCache support for Cosmos3 [Perf][Cosmos3] Add SeaCache support for Cosmos3 Sep 1, 2026
@hsliuustc0106 hsliuustc0106 added diffusion codes related to diffusion models Kernel optimization Codes related to optimize kernel execution to improve hardware utilization labels Sep 2, 2026

@MaciejBalaNV MaciejBalaNV left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall the changes look good. However, there is some improvements to be made about Cosmos3 specific comments and logs - I didn't mark all of these. Can you also update the Cosmos3 docs mentioning this is the recommended cache for these models?

Comment thread vllm_omni/diffusion/cache/seacache/backend.py Outdated
Comment thread vllm_omni/diffusion/cache/seacache/backend.py Outdated
Comment thread vllm_omni/diffusion/cache/seacache/backend.py Outdated
Comment thread vllm_omni/diffusion/cache/seacache/config.py Outdated
Comment thread vllm_omni/diffusion/cache/seacache/hook.py Outdated
Comment thread vllm_omni/diffusion/cache/seacache/hook.py Outdated
Comment thread vllm_omni/diffusion/worker/diffusion_model_runner.py Outdated
@yzhautouskay

yzhautouskay commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Self-review / current state

This PR adds a native SeaCache backend. Mainly targeting Cosmos3, but potentially possible to be used by other models. It is currently opt-in before broader quality validation, with the explicit goal of enabling SeaCache by default

Remaining action items

@alex-jw-brooks

Copy link
Copy Markdown
Collaborator

Hi, thanks for the contribution!

I think we need to be very careful about adding more diffusion caching backends unless there is a clear reason to do so, especially when the caching needs potential changes in the model. Adding too many approaches is also confusing to users since it's not clear what to pick, and it can also cause model support to be more sparse since contributions become diluted over the different backends.

It would also be easier to tell whether or not a new caching backend should be added based on how it benchmarks (both for quality and accuracy) against the existing diffusion caching approaches, and not just the baseline. Then we can potentially provide better guidance to users on when to use what.

Also cc @hsliuustc0106 @wtomin @NickCao @RuixiangMa in case any of you have thoughts as well

@MaciejBalaNV

Copy link
Copy Markdown
Contributor

Hey @alex-jw-brooks , in our internal tests for Cosmos3 models SeaCache is a clear winner and the only solution in a production-ready state.

I understand the worry about confusing users with too many options, but at the same time we should make sure that vLLM-Omni stays on SOTA level. In the fast moving field it's inevitable that new options will show up, as is the case here for e.g. SeaCache (first paper Feb 2026) vs TeaCache (first paper Nov 2024). IMO we should take advantage of a modular design with easy to swap caches and extend the framework with new options, rather than limit ourselves because the amount of choices can be confusing.

We can provide more examples and benchmarks for Cosmos3 model with different cache backends. We will also update the docs for Cosmos3 models explaining which cache backend should be used.

Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
@yzhautouskay
yzhautouskay force-pushed the yzhautouskay/cosmos3_diffusion_caching branch from ea2ac43 to 183a380 Compare September 3, 2026 10:54
@yzhautouskay

Copy link
Copy Markdown
Contributor Author

@alex-jw-brooks @MaciejBalaNV Here's diffusion-cache comparison to demonstrate SeaCache's optimal quality-speedup balance for Cosmos3

Settings

Single-seed comparison on 1× GB200 per mode: SeaCache, Cache-DiT, and TeaCache via vLLM-Omni PR #4389. vLLM-Omni cache-acceleration post for Cache-DiT background.

Time is server-side Cosmos3 pipeline time.

Video: 1280×720, 189 frames, 35 steps, CFG 6.0. T2I: 1024×1024, 50 steps, CFG 7.0.

  • SeaCache: threshold 0.25, residual order 1, max consecutive cached 2.
  • Cache-DiT default: F1B0.
  • Cache-DiT quality: F8B8 with threshold 0.12.
  • Cache-DiT blog-hybrid: F1B0 with TaylorSeer.
  • TeaCache: rel_l1_thresh=0.2, num_warmup_steps=12, matching PR [Feature] Add TeaCache support for Cosmos3 #4389.
  • MagCache: excluded; Cosmos3 is not wired and requires magnitude calibration in offline run on the representative dataset.
  • StepCache: excluded; not wired to Cosmos3 and specialized for DreamZero.

Environments

  • SeaCache and Cache-DiT: commit ea2ac433, vLLM/vLLM-Omni 0.28.0, Cache-DiT 1.5.0, PyTorch 2.13.0.
  • TeaCache: PR commit 016fb722, vLLM/vLLM-Omni 0.24.0, Cache-DiT 1.3.0, PyTorch 2.11.0.

Warning: TeaCache was run on its PR branch with the v0.24 stack and compared only against the no-cache baseline from that same branch and environment. SeaCache and Cache-DiT use the v0.28 baseline; do not compare absolute values across the two stacks.

Cosmos3-Nano

T2V

Mode Time Speedup LPIPS PSNR SSIM
No cache (v0.28) 97.10s — — — —
No cache (v0.24) 95.30s — — — —
SeaCache 53.39s 1.82× 0.2351 22.22 0.7844
Cache-DiT default 53.09s 1.83× 0.3851 19.15 0.7105
Cache-DiT quality 83.46s 1.16× 0.2800 20.68 0.7544
Cache-DiT blog-hybrid 60.34s 1.61× 0.3625 17.77 0.6811
TeaCache (v0.24) 63.70s 1.50× 0.2080 21.83 0.8056
cosmos3_nano_t2v_comparison_pair1.mp4
cosmos3_nano_t2v_comparison_pair2.mp4
cosmos3_nano_t2v_comparison_pair3.mp4

I2V

Mode Time Speedup LPIPS PSNR SSIM
No cache (v0.28) 96.99s — — — —
No cache (v0.24) 94.42s — — — —
SeaCache 55.70s 1.74× 0.1169 27.92 0.8004
Cache-DiT default 52.95s 1.83× 0.1731 25.22 0.7564
Cache-DiT quality 81.78s 1.19× 0.1453 26.16 0.7863
Cache-DiT blog-hybrid 58.61s 1.65× 0.1760 24.86 0.7527
TeaCache (v0.24) 62.70s 1.51× 0.1596 26.33 0.7923
cosmos3_nano_i2v_comparison_pair1.mp4
cosmos3_nano_i2v_comparison_pair2.mp4
cosmos3_nano_i2v_comparison_pair3.mp4

T2I

Mode Time Speedup LPIPS PSNR SSIM
No cache (v0.28) 2.75s — — — —
No cache (v0.24) 2.53s — — — —
SeaCache 1.63s 1.69× 0.1921 20.40 0.8019
Cache-DiT default 2.68s 1.03× 0.1606 22.21 0.8394
Cache-DiT quality 4.44s 0.62× 0.0652 28.67 0.9245
Cache-DiT blog-hybrid 2.68s 1.08× 0.1829 21.52 0.7966
TeaCache (v0.24) 1.78s 1.42× 0.1829 20.81 0.8238
cosmos3_nano_t2i_comparison_3072x2208

Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
@yzhautouskay

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106
Posted a comprehensive quality-validation report here:
https://yzhautouskay.github.io/vllm-omni/seacache-validation/

It includes:

  • Aggregated PaiBench-G I2V (1,040 prompts × 5 seeds) quality scores to establish acceptable video quality beyond small-scale manual inspection and PSNR/SSIM/LPIPS metrics
  • PSNR/SSIM/LPIPS metrics across 102 matched no-cache/SeaCache pairs + side-by-side visualizations
  • Coverage across prompts, seeds, modes, and parallelism configurations
  • Per-task worst-case fidelity examples

I also merged the latest upstream main and addressed the previous review feedback. Could you please review the updated PR?

Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
@linyueqian linyueqian added the ready label to trigger buildkite CI label Sep 14, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The SeaCache implementation itself is careful and I did not find a defect in it. The indicator does latent[batch].movedim(0, -1) before apply_sea_filter, so the separable Wiener gain really runs over (T, H, W); extrapolate_residual is a correct in-place Newton divided-difference table with a zero-denominator guard; the gate forces full compute on the first and last step, on max_consecutive_cached, and whenever history or the indicator is missing; history is bounded to residual_order + 1 entries; and _synchronize_compute all-reduces the skip decision with MAX across the FS, SP and layerwise-offload groups so ranks cannot disagree and hang a collective. The get_extractor change to walk __mro__ is what lets FSDP-wrapped transformers resolve, and the two-key enabler map (Cosmos3OmniDiffusersPipeline, Cosmos3OmniPipeline) matches the alias cache_dit already uses in model_specific.py:885. Cosmos3EdgeVFMTransformer exists as a subclass in transformer_cosmos3_edge.py, so both extractor keys are live.

What holds this back is not the cache; it is two changes to the default path that the description presents as an opt-in feature. Both are inline. The first is that sampling_dtype = torch.float32 is unconditional: initial noise, sound latents, I2V image latents and the velocity mask are now created in float32 and every prediction is cast back to it, whether or not --cache-backend sea_cache is set. That is very likely the right numerical choice, but it changes what every Cosmos3 user gets on upgrade, including the noise draw at a fixed seed, and the baseline column in the PR's own table was measured on this branch, so the 1.8x is fp32-vs-fp32 and says nothing about main-vs-PR. The second is _dit_any_rank_failed, which on main imports a get_dit_group that does not exist, catches the ImportError, and silently returns the local flag; the PR points it at the world group and so turns on a cross-rank all-reduce that has never actually run. That is a bug fix and probably a good one, but it belongs in the description with a sentence on why every rank is guaranteed to reach it in lockstep.

Two smaller things worth noting rather than acting on. The transformer forward and the new extractor both raise TypeError on any unexpected kwarg where the old forward silently accepted **kwargs; CI is green so every in-tree caller is clean, but it is a tightening third-party wrappers could trip on. And the PR carries two unrelated hardenings in extract_qwen_context (torch.as_tensor for a possibly non-tensor timestep) and extract_flux2_klein_context (a None guard); harmless, but they are not SeaCache and a reader of the history will not find them here.

Validation: tests/diffusion/cache/test_seacache.py is core_model/cpu so it gates per-PR, and buildkite/vllm-omni is green at this head (build 15156). Reviewed statically; the head is on a fork and no PR code was executed. Verdict is comment rather than approve only because the float32 sampling change needs to be either stated and defended for the default path or gated behind the backend, and that is the author's call to make, not mine.

Comment thread vllm_omni/diffusion/models/cosmos3/pipeline_cosmos3.py Outdated
Comment thread vllm_omni/diffusion/worker/diffusion_model_runner.py Outdated
Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 15, 2026
Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 15, 2026
noisy_frame_mask = extra_states.get("sea_cache_noisy_frame_mask")
if isinstance(noisy_frame_mask, torch.Tensor) and not bool(torch.any(noisy_frame_mask != 0).item()):
self._warn_once("SeaCache requires noisy vision; conditioning-only calls run in full.")
return self._run_uncached(ctx)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Move uncached execution outside this try block; model failures currently trigger a second forward.

gain = reshaped_gain if gain is None else gain * reshaped_gain

assert gain is not None
mean_gain = gain.mean()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Normalize by mean_gain directly; validated gains are finite and positive.

@linyueqian linyueqian left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving at 7adab72c. Both items from round one have been taken out of this PR rather than argued inside it: the unconditional fp32 sampling state now lives in #7592 with its own motivation and no-cache numbers, and the _dit_any_rank_failed group change is reverted so that function is byte-for-byte main again, with a clean follow-up promised. What is left is exactly the opt-in cache and nothing that alters the default Cosmos3 path.

The cache itself did not change since my first read and was clean then: the separable Wiener filter runs over (T, H, W) after the movedim, the Newton residual extrapolation is a correct divided-difference table with a zero-denominator guard, first and last steps are forced, history is bounded to residual_order + 1, the skip decision is all-reduced with MAX across the FS, SP and layerwise-offload groups so ranks cannot diverge, and both extractor keys resolve to live classes. The runner delta against main is the SeaCache registration plus the mypy-driven tidy that came with it: a local pipeline alias asserted non-None instead of repeated self.pipeline reads, matched_state instead of shadowing state, an isinstance(self.cache_backend, CacheDiTBackend) check in place of the string compare, and an explicit error when a stepwise request reaches the scheduler with no latents; none of it changes control flow.

Validation: lane 15246 on this head is green on the general and AMD lanes (Intel is red on every main commit today, NPU is informational), tests/diffusion/cache/test_seacache.py is core_model/cpu and gates per-PR, and a cross-model panel on this head found nothing beyond the one statistics nit inline. Reviewed statically; the head is on a fork and no PR code was executed.

One request before this is merged rather than before it is approved: the branch is 56 commits behind main and two of them touch files in this diff (13b85c56 from #7427 in pipeline_cosmos3.py and its test, 68294066 from #4167 in serve.py), and the green lane ran on the head as-is, so it has never seen those changes. Please rebase onto current main; I will cycle ready after the push so the lane fires, and merge on green.

Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
…3_diffusion_caching

Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 15, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-approving at 2dd3f786, which is the rebase I asked for plus one small follow-up. The base is now current main (78934753), and the merge commit is byte-identical to the automatic merge of 6ba87dfb with main, so the #7427 and #4167 changes that the previous lane never saw are in this tree. The follow-up commit does two things and both check out: the finite-and-positive guard on gain.mean() in sea_filter.py is gone, which is safe because every axis gain is signal_scale * clean_power / (signal_scale**2 * clean_power + noise_scale**2 + 1e-16) with clean_power > 0 and signal_scale clamped into (1e-6, 1 - 1e-6), so the product is strictly positive and finite for any shape and the guard was dead; and the conditioning-only early return in hook.py now happens after the scheduler metadata is validated instead of before, which only changes which one-time warning fires when both conditions hold and adds a test pinning that a failing forward on a conditioning-only call runs exactly once.

Everything I approved at 7adab72c is otherwise unchanged: both round-one items are out of this PR (#7592 and a follow-up), the default Cosmos3 path is main's, and the cache itself was clean on my read and on a three-model panel; the one statistics nit from that panel is inline as a suggestion and does not hold the merge.

Validation: the general lane on this head has been re-fired and I will merge on green; tests/diffusion/cache/test_seacache.py is core_model/cpu and gates per-PR. Reviewed statically; the head is on a fork and no PR code was executed.

if residual.device != ctx.hidden_states.device:
residual = residual.to(ctx.hidden_states.device)
state.consecutive_cached += 1
self.skip_count += 1

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] skip_count and consecutive_cached are bumped before the can_reuse check, so the defensive fallback a few lines down (shape, device or dtype mismatch) runs the full stack and _record_execution, which resets consecutive_cached but never takes the skip back out of skip_count. The step is counted as both skipped and executed, which only corrupts the summary statistics and only on a path that should never fire, so this is not holding the merge; moving the two increments under if can_reuse: keeps the numbers honest. Credit to the cross-model panel for spotting it.

@linyueqian
linyueqian merged commit 6040bb9 into vllm-project:main Sep 15, 2026
8 of 9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: Yuliya Zhautouskaya <yzhautouskay@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner diffusion codes related to diffusion models high priority high priority issue, needs to be done asap Kernel optimization Codes related to optimize kernel execution to improve hardware utilization ready label to trigger buildkite CI world model

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants