iWISDM
收藏arXiv2025-09-30 收录
数据链接:
官方服务:
资源简介:
该数据集是在iWISDM环境中构建的,旨在生成各种复杂度的视觉-语言任务,以评估在多模态环境下遵循指令的能力。该数据集包含三个对应不同复杂度任务的基准,并已用于评估多个大型多模态模型与人类表现的对标。这些任务的复杂度分为低、中、高三个级别,总计包含150项试验,任务内容主要是关于多模态任务中的指令遵循。
This dataset was constructed in the iWISDM environment, aiming to generate visual-language tasks of varying complexities to evaluate instruction-following capabilities in multimodal settings. It contains three benchmarks corresponding to tasks of distinct complexity levels, and has been used to assess multiple large multimodal models against human performance. The complexity of these tasks is categorized into three tiers: low, medium, and high, totaling 150 trials. The tasks primarily focus on instruction following in multimodal scenarios.
提供机构:
BashivanLab搜集汇总
数据集介绍

背景与挑战
背景概述
iWISDM是一个用于评估多模态模型指令遵循能力的虚拟环境数据集,能够生成多样化的视觉语言任务,涵盖执行功能如工作记忆和任务切换。它是一个可扩展框架,支持用户自定义任务,并基于ShapeNet数据集进行渲染,提供基准测试和工具包以便于使用和研究。
以上内容由遇见数据集搜集并总结生成



