infly/Infinity-Doc2-5M
收藏资源简介:
Infinity-Doc2-5M 是一个大规模、高质量的文档解析训练数据集,专门用于文档解析任务。该数据集包含500万个样本,覆盖多种文档类型,如学术论文、研究报告、财务报告、报纸、教科书、考试试卷、杂志等,支持中文和英文语言以及多种布局类型(如单栏、多栏、垂直文本)。数据集提供丰富的注释,包括块级类别(标题、文本段落、表格、公式、页眉、页脚等)、文档元素定位信息、每个元素区域的识别结果(文本字符串、表格HTML、公式LaTeX、化学SMILES、图表)以及文档的整体阅读顺序。此外,数据集设计了多样化的提示词以增强生成模型的训练多样性和泛化能力。数据质量高,通过人工过滤、智能标注和数据合成相结合,并经过专家质量检查,确保准确性。部分数据从原始语料库合成,无敏感信息,符合版权规定,适用于学术和非商业用途。该数据集为文档布局分析、元素检测与识别、公式解析和文档理解等任务提供了坚实的数据基础。
Infinity-Doc2-5M is a large-scale, high-quality document parsing training dataset specifically designed for document parsing tasks. This dataset contains 5 million samples, covering diverse document types including academic papers, research reports, financial reports, newspapers, textbooks, examination papers, magazines and more, and supports both Chinese and English languages as well as various layout types such as single-column, multi-column, and vertical text. The dataset provides rich annotations, including block-level categories (headings, text paragraphs, tables, formulas, headers, footers, etc.), positioning information of document elements, recognition results for each element region (text strings, table HTML, formula LaTeX, chemical SMILES, and charts), as well as the overall reading order of the document. Additionally, the dataset incorporates diverse prompt designs to enhance the training diversity and generalization capability of generative models. It guarantees high data quality through a combination of manual filtering, intelligent annotation, and data synthesis, alongside expert quality inspections to ensure accuracy. Some of the data is synthesized from original corpora, contains no sensitive information, complies with copyright regulations, and is suitable for academic and non-commercial use. This dataset provides a solid data foundation for tasks such as document layout analysis, element detection and recognition, formula parsing, and document understanding.




