遇见数据集

jacobcohen/mechaptcha

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个文本图像数据集,包含图像、文本标识符、文本内容以及多种视觉扰动或样式属性。具体特征包括:图像数据、唯一ID、文本字符串,以及多个布尔标记(如模糊、加粗、点状噪声、简单线条、复杂线条、斜体、旋转、椒盐噪声、双线、波浪效果、波浪线、字符抖动和间距抖动)。数据集分为训练集(1,200,000个样本)、验证集(150,000个样本)和测试集(150,000个样本),总大小约3.5GB。该数据集可能用于光学字符识别(OCR)或文本图像处理任务,以评估模型在受扰动文本图像上的性能。

This dataset is a text image dataset containing images, text identifiers, text content, and various visual perturbations or style attributes. Specific features include: image data, unique ID, text string, and multiple boolean flags (such as blur, bold, dots, easy_line, hard_line, italic, rotation, salt_pepper, two_lines, wave, wavy_line, char_jitter, and spacing_jitter). The dataset is split into training set (1,200,000 examples), validation set (150,000 examples), and test set (150,000 examples), with a total size of approximately 3.5 GB. It is likely designed for optical character recognition (OCR) or text image processing tasks to evaluate model performance on perturbed text images.

提供机构:
jacobcohen
二维码
社区交流群
二维码
科研交流群
商业服务