4D-Bench
收藏资源简介:
4D-Bench是一个专为评估多模态大型语言模型在4D对象理解能力上的新基准。该数据集包含多样化的4D对象类别,高质量注释,并设计有需要多视角时空理解的4D对象问答和4D对象字幕任务。数据集通过渲染Objaverse-XL中的动态3D对象来构建,并经过精心设计的数据清洗流程以确保数据质量。它为多模态大型语言模型在4D对象理解方面的评估提供了新的挑战,并可作为图像/视频MLLMs的泛化评估基准。
4D-Bench is a novel benchmark specifically designed to evaluate the 4D object understanding capabilities of multimodal large language models. This dataset encompasses diverse 4D object categories and high-quality annotations, and features 4D object question answering and 4D object captioning tasks that demand multi-view spatio-temporal comprehension. Constructed by rendering dynamic 3D objects from Objaverse-XL, the dataset has undergone a meticulously designed data cleaning pipeline to ensure data quality. It presents new challenges for evaluating the 4D object understanding performance of multimodal large language models, and can serve as a generalizable evaluation benchmark for image/video-based MLLMs.




