遇见数据集

youssefkhalil320/urdu_images_doc_tags_all_v12

收藏
Hugging Face2025-10-24 更新2025-10-25 收录
官方服务:

资源简介:

该数据集包含了文档的多种信息,如文档的ID、源PDF文件路径、图片、图片预览、HTML和OTS标签格式文档、Markdown格式文档、文档语言及其置信度、难度评分、页面文本长度、是否缺失边界框、是否含有非全宽文本、页面宽高、渲染宽高等信息。数据集分为训练集,共有6741个样本,大小为约2.1GB。

The dataset includes various information of documents, such as document ID, source PDF file path, images, image preview, HTML and OTSL tagged documents, Markdown document, document language and its confidence, difficulty score, page text length, whether there is missing bounding box, whether there is non-full width text, page width and height, rendered width and height, etc. The dataset is split into a training set with a total of 6741 samples, with a size of approximately 2.1GB.

提供机构:
youssefkhalil320
二维码
社区交流群
二维码
科研交流群
商业服务