遇见数据集

youssefkhalil320/urdu_images_doc_tags_all_v9

收藏
Hugging Face2025-10-22 更新2025-10-25 收录
官方服务:

资源简介:

这是一个包含了文档相关特征的训练数据集,其中包括文档的唯一标识符、源PDF路径、图像、图像预览、文档标签的HTML和OTS格式、Markdown文本、语言类型、语言识别置信度、难度评分、页面文本长度、是否缺失边界框、是否含有非全宽文本、页面宽度和高度等信息。数据集被划分为训练集,可用于文档处理相关的机器学习任务。

This is a training dataset that includes various document-related features such as unique identifiers for documents, source PDF paths, images, image previews, document tags in HTML and OTSL formats, Markdown text, language type, language detection confidence, difficulty score, page text length, whether missing bounding boxes, whether containing non-full width text, page width, and height. The dataset is split into a training set and can be used for machine learning tasks related to document processing.

提供机构:
youssefkhalil320
二维码
社区交流群
二维码
科研交流群
商业服务