Skip to content

[Multimodal] Optimize sparse frames decoding for PyAV video backend - #44598

Draft
Isotr0py wants to merge 1 commit into
vllm-project:mainfrom
Isotr0py:pyav-segment-demux
Draft

Isotr0py wants to merge 1 commit into
vllm-project:mainfrom
Isotr0py:pyav-segment-demux

Conversation

@Isotr0py

@Isotr0py Isotr0py commented Jun 5, 2026 •

Copy link
Copy Markdown
Member

Purpose

Test Plan

  • Test on a video with 900 frames (30s, fps=30):
image

Test Result


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
@mergify mergify Bot added the multi-modality Related to multi-modality (#4194) label Jun 5, 2026
@Isotr0py Isotr0py changed the title [Multimodal] Optimize sparse frames decoding for PyAV backend [Multimodal] Optimize sparse frames decoding for PyAV video backend Jun 5, 2026
@WindChimeRan

Copy link
Copy Markdown
Contributor

I've opened #45203, which adds an opt-in lossy keyframe-only sampling policy (pyav_keyframes).

I think it's complementary to this PR. This PR makes decoding the same exact frames faster, while mine is keyframes-only for prefilling-heavy video classification scenarios, and pays for it in accuracy on motion-sensitive tasks.

@NicolasHug

NicolasHug commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

As requested by @Isotr0py I benchmarked this against TorchCodec and against the previous pyav backend in #46609 (see results there). Here are my 2 cents:

  • This new pyav backend is much faster than the previous pyav backend on sparse sampling, but slower on dense sampling (with high number of threads, which is the current vllm default).
  • TorchCodec is faster than both pyav backends on both sampling scenarios: TC is much faster on dense sampling , and a tiny bit faster on sparse sampling.

The reason TC is faster is that it's smart about when it should seek vs just decode frames forwards. I recently implemented a heuristic for that (see meta-pytorch/torchcodec#1488). You could implement the same heuristic for pyav, although of course it's up to you to decide whether you want to own that complexity within vllm, or let TorchCodec handle it.

My note also explains why your new pyav implementation is slower than the previous one on dense sampling: sometimes, it's just better to NOT seek! And this is especially true with large number of threads in FRAME threading (what this PR is doing).


What this PR is doing via _demux_keyframes() is what TorchCodec's does by default: it scans the file to find the keyframes, which allows for exact and fast seeking behavior. It's worth noting that this scan strategy can be detrimental:

TorchCodec solves all this already, and allows to not scan via seek_mode="approximate". If you want to allow that with the pyav backend, you'd need yet another implementation, potentially many.


On keyframe sampling: TorchCodec exposes a private VideoDecoder(...)._get_key_frame_indices() method. I'm happy to expose it publicly if you need it.

@mergify

mergify Bot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @Isotr0py.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

multi-modality Related to multi-modality (#4194) needs-rebase

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants