相关数据集
HenriqueGodoy/extract-0
Extract-0文档信息提取数据集是一个包含280,128个合成训练示例的数据集,用于训练Extract-0语言模型,该模型在文档信息提取任务上表现优于GPT-4和其他大型模型。数据集由来自arXiv论文、PubMed Central文章、Wikipedia内容和FDA监管文档的文本片段构成,每个示例都配有一个基于模式的提取任务和相应的结构化输出。
Hugging Face2025-09-30 更新130
Artifacts for Keyword Extraction From Specification Documents for Planning Security Mechanisms
This dataset contains the data used for evaluating VDocScan - a keyword extraction based security vulnerability prediction method. The repository includes an extensive list of Products and Vulnerabili
Zenodo2023-01-28 更新60
DUDE competition train and validation splits ground truth
This JSON file contains the ground truth annotations for the train and validation set of the DUDE competition (https://rrc.cvc.uab.es/?ch=23&com=tasks) of ICDAR 2023 (https://icdar2023.org/). <str
Zenodo2023-03-23 更新20
tinixai/ocr_annual_financials
TiniX越南OCR年度财务报表数据集是一个大规模越南语财务文档数据集,包含2015年至2025年越南上市公司的年度财务报表原始PDF文件和相应OCR提取的文本。数据集涵盖18,231份财务报告,涉及1,491个不同的股票代码,总数据量约为194 GB。OCR文本内容完整保留了文档结构,包括合并财务报表、母公司财务报表、年度审计报告、财务报表附注及相关附录和表格,在数字和表格数据上的准确率达到95
Hugging Face2026-05-26 更新80



