Model link on Hugging Face
https://huggingface.co/YunzeLiu/OmniRetriever-7B
Arxiv link
https://arxiv.org/abs/2605.26641
Description of the model
OmniRetriever-7B is a multimodal retrieval model designed to produce shared embeddings across text, audio, and video modalities.
The model supports cross-modal retrieval settings involving text, audio, and video, making it relevant for MOEB tasks that evaluate multimodal representation alignment and retrieval performance across these modalities.
Model link on Hugging Face
https://huggingface.co/YunzeLiu/OmniRetriever-7B
Arxiv link
https://arxiv.org/abs/2605.26641
Description of the model
OmniRetriever-7B is a multimodal retrieval model designed to produce shared embeddings across text, audio, and video modalities.
The model supports cross-modal retrieval settings involving text, audio, and video, making it relevant for MOEB tasks that evaluate multimodal representation alignment and retrieval performance across these modalities.