DocILE
收藏资源简介:
DocILE数据集是由Rossum.ai创建的,是目前最大的商业文档信息定位与提取(KILE)及行项目识别(LIR)数据集。该数据集包含6700个标注的商业文档、10万个合成文档以及近100万个用于无监督预训练的未标注文档。DocILE数据集的特点包括:(i) 55个类别的标注,远超以往发布的KILE数据集的粒度;(ii) LIR任务代表了一个高度实用的信息提取任务,需要将关键信息分配到表格中的项目;(iii) 文档来自多种布局,测试集包括零样本和少样本情况,以及在训练集中常见的布局。
The DocILE dataset, created by Rossum.ai, is currently the largest dataset for commercial document information location and extraction (KILE) and line item recognition (LIR). This dataset includes 6,700 annotated commercial documents, 100,000 synthetic documents, and nearly 1 million unannotated documents for unsupervised pre-training. The characteristics of the DocILE dataset are as follows: (i) Annotations covering 55 categories, which far exceeds the annotation granularity of previously released KILE datasets; (ii) The LIR task represents a highly practical information extraction task that requires assigning key information to items in tables; (iii) The documents originate from diverse layouts, and the test set encompasses zero-shot and few-shot scenarios, as well as layouts commonly encountered in the training set.




