MOAT(Multimodal model Of All Trades)是一个用于评估大型多模态模型(LMMs)的挑战性基准测试数据集。它包含视觉语言任务,要求LMM整合多种视觉语言能力,并参与类似人类的通用视觉问题解决。MOAT的许多任务还关注LMMs在复杂文本和视觉指令上的基础能力,这对于LMMs在野外应用至关重要。
MMCBench是由Sea AI Lab创建的一个全面基准,用于评估超过100种流行的大型多模态模型(LMMs)在面对文本、图像和语音交互中的四种基本生成任务时的自我一致性。该数据集包含超过150个模型检查点,旨在通过彻底的评估,促进对尖端LMMs可靠性的更好理解。MMCBench特别关注于测量模型输出在遭受常见损坏时的自我一致性,为多模态模型的实际部署提供了关键的评估工具。
We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meti
Introduction Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating ar