<b>M</b>ulti-<b>T</b>ype <b>A</b>ncient <b>C</b>hinese <b>C</b>haracter <b>R</b>ecognition (MTACCR) dataset
收藏资源简介:
<b>The Multi-Type Ancient Chinese Character Recognition (MTACCR) dataset</b> is a large-scale resource designed to advance research in ancient Chinese script analysis. It is constructed based on the <i>Table of General Standard Chinese Characters</i> (通用规范汉字表), covering <b>7,874 Chinese characters</b> across three levels (3,500 Level-1, 3,000 Level-2, and 1,605 Level-3) and including standard, traditional, and variant forms. With <b>over 9 million samples</b>, the dataset features diverse ancient character images—ranging from original scanned manuscripts to segmented glyphs—collected from calligraphic databases, open-source datasets, and data augmentation. MTACCR significantly surpasses existing datasets in character coverage, typological diversity, and scale, providing a comprehensive benchmark for recognition and historical linguistics studies.



