raphael0202/ingredient-detection-layout-dataset
收藏Hugging Face2023-11-01 更新2024-03-04 收录
下载链接:
https://hf-mirror.com/datasets/raphael0202/ingredient-detection-layout-dataset
下载链接
链接失效反馈官方服务:
资源简介:
该数据集名为ingredient-detection-layout-dataset,主要用于成分检测和布局分析。数据集包含多个特征,如命名实体识别标签(ner_tags)、单词序列(words)、边界框序列(bboxes)、图像(image)、文本(text)、偏移量序列(offsets)和元数据(meta,包括条形码、图像ID、URL、ID和测试分割标志)。数据集分为训练集和测试集,训练集包含5065个样本,测试集包含556个样本。数据集的下载大小为2271205424字节,总大小为2304124809.875字节。
This dataset is named ingredient-detection-layout-dataset, which is primarily designed for ingredient detection and layout analysis. It contains multiple features including named entity recognition tags (ner_tags), word sequences, bounding box sequences, images, text, offset sequences, and metadata (meta, which covers barcode, image ID, URL, ID and test split flag). The dataset is split into training and test sets, with 5065 samples in the training set and 556 samples in the test set. The download size of the dataset is 2271205424 bytes, and the total storage size is 2304124809.875 bytes.
提供机构:
raphael0202原始信息汇总
数据集概述
数据集信息
特征
- ner_tags: 序列特征,包含类别标签,标签名称包括:
- 0: O
- 1: B-ING
- 2: I-ING
- words: 序列特征,字符串类型
- bboxes: 序列特征,整数类型
- image: 图像类型
- text: 字符串类型
- offsets: 序列特征,整数类型
- meta: 结构化特征,包含以下字段:
- barcode: 字符串类型
- image_id: 字符串类型
- url: 字符串类型
- id: 字符串类型
- in_test_split: 布尔类型
数据分割
- train: 训练集,包含5065个样本,大小为2059533770.875字节
- test: 测试集,包含556个样本,大小为244591039.0字节
数据集大小
- 下载大小: 2271205424字节
- 数据集大小: 2304124809.875字节




