<b>M</b>ulti-<b>T</b>ype <b>A</b>ncient <b>C</b>hinese <b>C</b>haracter <b>R</b>ecognition (MTACCR) dataset
收藏资源简介:
<b>The Multi-Type Ancient Chinese Character Recognition (MTACCR) dataset</b> is a large-scale resource designed to advance research in ancient Chinese script analysis. It is constructed based on the <i>Table of General Standard Chinese Characters</i> (通用规范汉字表), covering <b>7,874 Chinese characters</b> across three levels (3,500 Level-1, 3,000 Level-2, and 1,605 Level-3) and including standard, traditional, and variant forms. With <b>over 9 million samples</b>, the dataset features diverse ancient character images—ranging from original scanned manuscripts to segmented glyphs—collected from calligraphic databases, open-source datasets, and data augmentation. MTACCR significantly surpasses existing datasets in character coverage, typological diversity, and scale, providing a comprehensive benchmark for recognition and historical linguistics studies.
多类型古汉字识别(Multi-Type Ancient Chinese Character Recognition, MTACCR)数据集是专为推进古汉字文字分析研究而构建的大规模学术资源。该数据集以《通用规范汉字表》为基础搭建,涵盖三级字表共7874个汉字,其中一级字表3500个、二级字表3000个、三级字表1605个,同时包含规范字、繁体字与异体字三类字形。数据集拥有逾900万张样本,涵盖类型丰富的古汉字图像——从原始扫描手稿到经分割处理的字形均有覆盖——数据来源包括书法数据库、开源数据集,并通过数据增强技术扩充样本。MTACCR在汉字覆盖范围、字形类型多样性与数据集规模三方面均显著优于现有同类数据集,可为古汉字识别与历史语言学研究提供全面的基准测试支撑。



