OCR-Synthetic-Multilingual-v1
收藏资源简介:
OCR-Synthetic-Multilingual-v1 是一个大规模合成的多语言OCR训练数据集,专为文本检测和识别任务设计。该数据集由NVIDIA Corporation使用经过大量修改和扩展的SynthDoG(合成文档生成器)工具生成,支持英语、日语、韩语、俄语、简体中文和繁体中文六种语言。数据集采用HDF5格式存储,每个文件包含图像数据、文本标注、图像尺寸、完整文本标签、JPEG质量参数和样本ID等信息。标注信息采用JSON格式,包含单词级、行级和段落级的边界框标注以及阅读顺序图。数据集总样本量超过1200万,存储容量达5.45TB,按语言分为不同子目录,每个语言又分为训练集、测试集和验证集。该数据集被用于训练Nemotron OCR v2模型,适用于机器学习研究人员和AI工程师进行OCR相关研究,采用CC BY 4.0许可协议,允许商业和非商业用途。
OCR-Synthetic-Multilingual-v1 is a large-scale synthetic multilingual OCR training dataset specifically tailored for text detection and recognition tasks. This dataset was generated by NVIDIA Corporation using a heavily modified and expanded SynthDoG (Synthetic Document Generator) tool, supporting six languages: English, Japanese, Korean, Russian, Simplified Chinese and Traditional Chinese. The dataset is stored in HDF5 format, with each file containing image data, text annotations, image dimensions, complete text labels, JPEG quality parameters, sample ID and other relevant information. The annotation information follows the JSON format, including word-level, line-level and paragraph-level bounding box annotations as well as reading order maps. The total number of samples in the dataset exceeds 12 million, with a total storage capacity of 5.45 TB. It is organized into separate subdirectories by language, and each language's subset is further split into training, test and validation sets. This dataset has been utilized to train the Nemotron OCR v2 model, and is suitable for machine learning researchers and AI engineers engaged in OCR-related research. It is distributed under the CC BY 4.0 license, permitting both commercial and non-commercial usage.




