jordanparker6/publaynet
收藏资源简介:
--- title: PubLayNet license: other annotations_creators: [] language: - en size_categories: - 100B<n<1T source_datasets: [] task_categories: - image-to-text task_ids: [] --- # PubLayNet PubLayNet is a large dataset of document images, of which the layout is annotated with both bounding boxes and polygonal segmentations. The source of the documents is [PubMed Central Open Access Subset (commercial use collection)](https://www.ncbi.nlm.nih.gov/pmc/tools/openftlist/). The annotations are automatically generated by matching the PDF format and the XML format of the articles in the PubMed Central Open Access Subset. More details are available in our paper ["PubLayNet: largest dataset ever for document layout analysis."](https://arxiv.org/abs/1908.07836). The public dataset is in tar.gz format which doesn't fit nicely with huggingface streaming. Modifications have been made to optimise the delivery of the dataset for the hugginface datset api. The original files can be found [here](https://developer.ibm.com/exchanges/data/all/publaynet/). Licence: [Community Data License Agreement – Permissive – Version 1.0 License](https://cdla.dev/permissive-1-0/) Author: IBM GitHub: https://github.com/ibm-aur-nlp/PubLayNet @article{ zhong2019publaynet, title = { PubLayNet: largest dataset ever for document layout analysis }, author = { Zhong, Xu and Tang, Jianbin and Yepes, Antonio Jimeno }, journal = { arXiv preprint arXiv:1908.07836}, year. = { 2019 } }
--- 标题:PubLayNet 许可证:其他 标注创作者:无 语言:英语 数据规模:100B < 样本量 < 1T 源数据集:无 任务类别:图像到文本 任务子项:无 --- # PubLayNet PubLayNet是一款大规模文档图像数据集,其文档布局同时通过边界框(bounding box)与多边形分割(polygonal segmentation)进行标注。该数据集的文档源自[PubMed Central 开放获取子集(商业使用合集)](https://www.ncbi.nlm.nih.gov/pmc/tools/openftlist/)。标注通过匹配PubMed Central开放获取子集内文章的PDF格式与XML格式自动生成。更多细节可参阅我们的论文《PubLayNet:用于文档布局分析的超大规模数据集》(https://arxiv.org/abs/1908.07836)。 原始公开数据集采用tar.gz格式,与Hugging Face流式加载机制适配性不佳。本版本已针对Hugging Face数据集API的加载需求进行了优化调整。原始数据集文件可在此处获取:https://developer.ibm.com/exchanges/data/all/publaynet/。 许可证:[社区数据许可协议——宽松版——1.0版](https://cdla.dev/permissive-1-0/) 开发方:IBM GitHub仓库:https://github.com/ibm-aur-nlp/PubLayNet 参考文献: bibtex @article{zhong2019publaynet, title = {PubLayNet:用于文档布局分析的超大规模数据集}, author = {Zhong, Xu and Tang, Jianbin and Yepes, Antonio Jimeno}, journal = {arXiv预印本 arXiv:1908.07836}, year = {2019} }
PubLayNet 数据集概述
基本信息
- 标题: PubLayNet
- 许可证: Community Data License Agreement – Permissive – Version 1.0
- 语言: 英语 (en)
- 大小分类: 100B<n<1T
- 任务类别: 图像到文本 (image-to-text)
数据集描述
PubLayNet 是一个大型文档图像数据集,其布局通过边界框和多边形分割进行标注。数据来源于 PubMed Central Open Access Subset (商业用途集合)。标注是通过匹配 PubMed Central Open Access Subset 中的文章的 PDF 格式和 XML 格式自动生成的。
数据集来源
相关文献
- 论文: "PubLayNet: largest dataset ever for document layout analysis."
- 作者: Zhong, Xu; Tang, Jianbin; Yepes, Antonio Jimeno
- 发表年份: 2019




