Repository navigation
Conversation
|
I've opened #45203, which adds an opt-in lossy keyframe-only sampling policy (pyav_keyframes). I think it's complementary to this PR. This PR makes decoding the same exact frames faster, while mine is keyframes-only for prefilling-heavy video classification scenarios, and pays for it in accuracy on motion-sensitive tasks. |
|
As requested by @Isotr0py I benchmarked this against TorchCodec and against the previous
The reason TC is faster is that it's smart about when it should seek vs just decode frames forwards. I recently implemented a heuristic for that (see meta-pytorch/torchcodec#1488). You could implement the same heuristic for pyav, although of course it's up to you to decide whether you want to own that complexity within vllm, or let TorchCodec handle it. My note also explains why your new pyav implementation is slower than the previous one on dense sampling: sometimes, it's just better to NOT seek! And this is especially true with large number of threads in What this PR is doing via
TorchCodec solves all this already, and allows to not scan via On keyframe sampling: TorchCodec exposes a private |
|
This pull request has merge conflicts that must be resolved before it can be |
Purpose
Test Plan
Test Result
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.