dataset link on Hugging dataset
https://huggingface.co/datasets/RyanWW/XModBench
Arxiv link
https://arxiv.org/pdf/2510.15148
Description of the dataset
I’d like to propose adding XModBench to MOEB.
It covers text, audio, images, and video. The benchmark contains tri-modal QA tasks, and has several directions:
audio + text → image
audio + text → video
video + text → audio
image + text → audio
text → image
text → video
image + text → text
video + text → text
The different directions are built from the same underlying aligned examples, with the context and candidate modalities swapped. So it is one benchmark suite rather than a collection of unrelated datasets.
We could either add each direction as an individual task, or combine them into one XModBench task covering all four modalities
dataset link on Hugging dataset
https://huggingface.co/datasets/RyanWW/XModBench
Arxiv link
https://arxiv.org/pdf/2510.15148
Description of the dataset
I’d like to propose adding XModBench to MOEB.
It covers text, audio, images, and video. The benchmark contains tri-modal QA tasks, and has several directions:
audio + text → image
audio + text → video
video + text → audio
image + text → audio
text → image
text → video
image + text → text
video + text → text
The different directions are built from the same underlying aligned examples, with the context and candidate modalities swapped. So it is one benchmark suite rather than a collection of unrelated datasets.
We could either add each direction as an individual task, or combine them into one XModBench task covering all four modalities