five

raphael0202/ingredient-detection-layout-dataset

收藏
Hugging Face2023-11-01 更新2024-03-04 收录
下载链接:
https://hf-mirror.com/datasets/raphael0202/ingredient-detection-layout-dataset
下载链接
链接失效反馈
官方服务:
资源简介:
该数据集名为ingredient-detection-layout-dataset,主要用于成分检测和布局分析。数据集包含多个特征,如命名实体识别标签(ner_tags)、单词序列(words)、边界框序列(bboxes)、图像(image)、文本(text)、偏移量序列(offsets)和元数据(meta,包括条形码、图像ID、URL、ID和测试分割标志)。数据集分为训练集和测试集,训练集包含5065个样本,测试集包含556个样本。数据集的下载大小为2271205424字节,总大小为2304124809.875字节。

This dataset is named ingredient-detection-layout-dataset, which is primarily designed for ingredient detection and layout analysis. It contains multiple features including named entity recognition tags (ner_tags), word sequences, bounding box sequences, images, text, offset sequences, and metadata (meta, which covers barcode, image ID, URL, ID and test split flag). The dataset is split into training and test sets, with 5065 samples in the training set and 556 samples in the test set. The download size of the dataset is 2271205424 bytes, and the total storage size is 2304124809.875 bytes.
提供机构:
raphael0202
原始信息汇总

数据集概述

数据集信息

特征

  • ner_tags: 序列特征,包含类别标签,标签名称包括:
    • 0: O
    • 1: B-ING
    • 2: I-ING
  • words: 序列特征,字符串类型
  • bboxes: 序列特征,整数类型
  • image: 图像类型
  • text: 字符串类型
  • offsets: 序列特征,整数类型
  • meta: 结构化特征,包含以下字段:
    • barcode: 字符串类型
    • image_id: 字符串类型
    • url: 字符串类型
    • id: 字符串类型
    • in_test_split: 布尔类型

数据分割

  • train: 训练集,包含5065个样本,大小为2059533770.875字节
  • test: 测试集,包含556个样本,大小为244591039.0字节

数据集大小

  • 下载大小: 2271205424字节
  • 数据集大小: 2304124809.875字节
二维码
社区交流群
二维码
科研交流群
商业服务