OmniParsingBench
收藏资源简介:
OmniParsingBench是由阿里巴巴集团提出的一个多模态解析基准数据集,旨在支持文档、图像和视听流的统一解析。该数据集包含丰富的知识密集型图像样本和优化的视频注释,用于细粒度分析和长教育内容理解。数据集构建过程采用了三阶段渐进式解析框架,从整体检测到细粒度识别再到多级解释,最终输出标准化的JSON格式数据。该数据集主要应用于多模态大语言模型的训练和评估,旨在解决复杂视听信号到结构化知识的转换问题,提升模型在检索增强生成、问答等下游任务中的可靠性。
OmniParsingBench is a multimodal parsing benchmark dataset proposed by Alibaba Group, aiming to support unified parsing of documents, images and audiovisual streams. This dataset contains rich knowledge-intensive image samples and optimized video annotations, which are designed for fine-grained analysis and long educational content understanding. The dataset construction adopts a three-stage progressive parsing framework, ranging from holistic detection, fine-grained recognition to multi-level interpretation, and finally outputs standardized JSON-formatted data. This dataset is mainly applied to the training and evaluation of multimodal large language models (LLMs), with the goal of solving the conversion problem from complex audiovisual signals to structured knowledge and improving the reliability of models in downstream tasks such as retrieval-augmented generation and question answering.




