ncar-ocr-dataset10-split
收藏资源简介:
该数据集是一个多模态数据集,包含图像和文本两种数据类型。数据集结构包含以下字段:'image'(图像类型)和'text'(字符串类型)。数据被划分为三个部分:训练集(5,171个样本,约1.6GB)、验证集(273个样本,约84.7MB)和测试集(287个样本,约89MB)。数据文件按默认配置存储在指定路径下,其中训练集文件路径为'data/train-*',验证集为'data/validation-*',测试集为'data/test-*'。数据集总下载大小约1.77GB,解压后总大小约1.78GB。
This is a multimodal dataset containing two data modalities: image and text. The dataset structure includes the following fields: 'image' (image-type data) and 'text' (string-type data). The data is split into three subsets: training set (5,171 samples, ~1.6 GB), validation set (273 samples, ~84.7 MB), and test set (287 samples, ~89 MB). The data files are stored at the designated path according to the default configuration, where the training set files are located at 'data/train-*', the validation set files at 'data/validation-*', and the test set files at 'data/test-*'. The total download size of the dataset is approximately 1.77 GB, and the total uncompressed size is approximately 1.78 GB.
NCAR-OCR-Dataset10-Split 数据集概述
数据集基本信息
- 数据集名称:NCAR-OCR-Dataset10-Split
- 数据集地址:https://huggingface.co/datasets/Abdalrahmankamel/ncar-ocr-dataset10-split
- 下载大小:1,768,223,040 字节
- 数据集大小:1,778,170,890 字节
数据集特征
- 特征字段:
image:图像数据,数据类型为 imagetext:文本数据,数据类型为 string
数据划分
- 训练集:
- 样本数量:5,171 个
- 数据大小:1,604,418,369 字节
- 文件路径:data/train-*
- 验证集:
- 样本数量:273 个
- 数据大小:84,704,354 字节
- 文件路径:data/validation-*
- 测试集:
- 样本数量:287 个
- 数据大小:89,048,167 字节
- 文件路径:data/test-*
配置信息
- 默认配置:
- 配置名称:default
- 数据文件按照训练集、验证集和测试集划分,分别对应指定路径模式




