遇见数据集

youssefkhalil320/hebrew_images_doc_tags_all_v6

收藏
Hugging Face2025-10-10 更新2025-10-25 收录
官方服务:

资源简介:

该数据集包含了图片、HTML文档标签、文本元素数量等多种类型的数据。数据集中的每个样本可能包含缺失的边界框引用和标签,同时记录了缺失边界框的数量和是否存在缺失的边界框。数据还包含了原始PDF来源、页码、是否包含非全宽文本、是否包含图形以及页面文本的长度。数据集分为训练集,其大小为830MB,共有7652个示例。

The dataset includes various types of data such as images, HTML document tags, and the number of text elements. Each sample in the dataset may contain missing bounding box references and labels, and records the count of missing bounding boxes and whether there are missing bounding boxes. The data also includes the original PDF source, page number, whether it contains non-full width text, whether it contains figures, and the length of the page text. The dataset is split into a training set, which is 830MB in size and contains 7652 examples.

提供机构:
youssefkhalil320
二维码
社区交流群
二维码
科研交流群
商业服务