synthetic-captchas-library
收藏资源简介:
该数据集是一个多语言合成的4字符CAPTCHA图像库,旨在支持OCR、多语言视觉模型和文字识别研究。数据集覆盖了44种世界文字系统,特别适用于低资源语言的OCR训练。每种文字系统包含10万张独特的CAPTCHA图像,总计440万张独特图像(包括重复格式则为880万张)。数据集以两种格式提供:CSV(标准机器学习工作流)和Parquet(大规模管道快速加载)。每个CAPTCHA包含4个字符,图像文件名与正确文本标签相对应。数据集适用于多语言OCR训练、视觉语言模型预训练、文字识别研究以及在扭曲条件下的鲁棒文本识别,但不应用于绕过真实世界的CAPTCHA安全系统。
This dataset is a multilingual synthetic 4-character CAPTCHA image library designed to support research on OCR, multilingual vision models and text recognition. The dataset covers 44 global writing systems, and is particularly suitable for OCR training for low-resource languages. Each writing system contains 100,000 unique CAPTCHA images, with a total of 4.4 million unique images across all systems (8.8 million if duplicate formats are included). The dataset is provided in two formats: CSV (for standard machine learning workflows) and Parquet (for fast loading in large-scale pipelines). Each CAPTCHA contains 4 characters, and the image filename corresponds to the correct text label. This dataset is applicable to multilingual OCR training, visual-language model pre-training, text recognition research and robust text recognition under distorted conditions, but it should not be used to bypass real-world CAPTCHA security systems.




