遇见数据集

DocLayNet

收藏
Opencsg2023-01-25 更新2025-05-03 收录
官方服务:

资源简介:

DocLayNet提供页面布局分割能力,它基于来自金融、科学、专利、招标、法律文本和手册等6个文档类别中80863个独特页面的数据集,数据集中每个页面都使用边界框进行标注,总共包含11个不同的类别标签。DocLayNet的数据规模属于1万到10万之间,标注信息由众包完成,采用COCO格式。数据集包含PNG图像、COCO格式的边界框标注、单页PDF文件以及包含坐标和内容的JSON文件四种类型的数据资产,并预定义了训练集、验证集和测试集。该数据集支持目标检测和图像分割等任务,并采用CDLA-Permissive-1.0许可协议。

DocLayNet is a dataset designed for page layout segmentation. It consists of 80,863 unique pages sourced from 6 document categories: finance, scientific literature, patents, tender documents, legal texts, and manuals. Every page in the dataset is annotated with bounding boxes, resulting in a total of 11 distinct category labels. The scale of the DocLayNet dataset ranges from 10,000 to 100,000 samples. All annotation information is generated via crowdsourcing and follows the COCO data format. The dataset contains four types of data assets: PNG images, COCO-formatted bounding box annotations, single-page PDF files, and JSON files that store coordinate information and content. Predefined training, validation, and test splits are provided. This dataset supports downstream tasks including object detection and image segmentation, and is released under the CDLA-Permissive-1.0 license.

创建时间:
2024-07-19
搜集汇总
数据集介绍
DocLayNet 数据集图片
背景与挑战
背景概述
DocLayNet是一个用于文档布局分割的大型数据集,包含80,863个来自金融、科学、专利等6个类别的独特页面,使用边界框标注11个类别标签,支持目标检测和图像分割任务。数据集提供PNG图像、COCO格式标注、PDF和JSON文件,并预定义了训练、验证和测试集,采用CDLA-Permissive-1.0许可协议。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务