LiveXiv
收藏资源简介:
LiveXiv是一个基于arXiv论文内容的多模态实时基准数据集,由美国密歇根大学统计系等机构创建。该数据集通过自动生成视觉问答对(VQA)来测试大型多模态模型(LMMs)的能力,避免了传统数据集的污染问题。数据集的内容包括从arXiv论文中提取的图表、表格等多模态数据,并通过GPT-4o等模型生成问答对。创建过程中,数据集通过结构化文档解析和多模态模型过滤,确保了数据的质量和多样性。LiveXiv主要应用于评估和提升LMMs在科学领域的性能,旨在解决现有基准数据集的污染和更新问题。
LiveXiv is a real-time multimodal benchmark dataset built on arXiv paper content, developed by institutions including the Department of Statistics at the University of Michigan, USA, and other affiliated organizations. This dataset automatically generates visual question-answer pairs (VQA) to evaluate the capabilities of Large Multimodal Models (LMMs), thereby avoiding the dataset contamination problem plaguing traditional datasets. The dataset comprises multimodal data such as charts and tables extracted from arXiv papers, with question-answer pairs generated via models like GPT-4o. During its construction, the dataset underwent structured document parsing and multimodal model-based filtering to guarantee data quality and diversity. Primarily designed to assess and improve the performance of LMMs in scientific domains, LiveXiv aims to resolve the contamination and staleness issues associated with current benchmark datasets.




