遇见数据集

youssefkhalil320/urdu_images_doc_tags_all_v8

收藏
Hugging Face2025-10-21 更新2025-10-25 收录
官方服务:

资源简介:

该数据集包含文档的多种表示形式,包括PDF源文件、图片、预览图、HTML和OTS标签格式、Markdown格式等。数据集还包含了文档的语言类型及置信度评分、难度评分、页面文本长度、是否缺失边界框、是否含有非全宽文本等元信息。数据集被划分为训练集,共有6741个样本,总大小约为1.06GB。

The dataset includes various representations of documents, such as source PDFs, images, preview images, HTML and OTSL tags, Markdown format, etc. The dataset also contains metadata like the documents language type and confidence score, difficulty score, page text length, whether there are missing bounding boxes, and whether there is non-full width text. The dataset is split into a training set with a total of 6741 samples, with an overall size of approximately 1.06GB.

提供机构:
youssefkhalil320
二维码
社区交流群
二维码
科研交流群
商业服务