dataset link on Hugging dataset
https://huggingface.co/datasets/suimu/InsAVE-80K
Arxiv link
https://arxiv.org/abs/2605.18467
Description of the dataset
InsAVE-80K is a dataset introduced with InstructAV2AV for instruction-conditioned audio-visual/video retrieval and editing.
The relevant retrieval setting for MOEB is: video + text → video
A query combines a source video with a natural-language instruction or modification, and the goal is to retrieve the corresponding target video.
This dataset would provide additional coverage for instruction-conditioned composed video retrieval.
dataset link on Hugging dataset
https://huggingface.co/datasets/suimu/InsAVE-80K
Arxiv link
https://arxiv.org/abs/2605.18467
Description of the dataset
InsAVE-80K is a dataset introduced with InstructAV2AV for instruction-conditioned audio-visual/video retrieval and editing.
The relevant retrieval setting for MOEB is: video + text → video
A query combines a source video with a natural-language instruction or modification, and the goal is to retrieve the corresponding target video.
This dataset would provide additional coverage for instruction-conditioned composed video retrieval.