遇见数据集

youssefkhalil320/arabic_images_doc_tags_all_v13

收藏
Hugging Face2025-10-26 更新2025-11-15 收录
官方服务:

资源简介:

该数据集包含文档的多种格式信息,如PDF,图片,HTML标记,以及文本内容的Markdown格式。此外,还提供了语言类型和置信度,文本难度评分,页面文本长度,以及页面尺寸等信息。这些特征表明数据集可能是用于文档分析,自然语言处理,或者光学字符识别等任务。训练集包含了大量的示例,用于模型的训练。

The dataset includes various formats of document information such as PDF, images, HTML tags, and Markdown format of text content. It also provides language type and confidence, text difficulty score, page text length, and page dimensions, etc. These features suggest that the dataset might be used for tasks like document analysis, natural language processing, or optical character recognition. The training set contains a large number of examples for model training.

提供机构:
youssefkhalil320
二维码
社区交流群
二维码
科研交流群
商业服务