遇见数据集

<b>M</b>ulti-<b>T</b>ype <b>A</b>ncient <b>C</b>hinese <b>C</b>haracter <b>R</b>ecognition (MTACCR) dataset

收藏
Figshare2025-06-08 更新2026-04-08 收录
官方服务:

资源简介:

​<b>​The Multi-Type Ancient Chinese Character Recognition (MTACCR) dataset​</b>​ is a large-scale resource designed to advance research in ancient Chinese script analysis. It is constructed based on the <i>Table of General Standard Chinese Characters</i> (通用规范汉字表), covering ​<b>​7,874 Chinese characters​</b>​ across three levels (3,500 Level-1, 3,000 Level-2, and 1,605 Level-3) and including standard, traditional, and variant forms. With ​<b>​over 9 million samples​</b>​, the dataset features diverse ancient character images—ranging from original scanned manuscripts to segmented glyphs—collected from calligraphic databases, open-source datasets, and data augmentation. MTACCR significantly surpasses existing datasets in character coverage, typological diversity, and scale, providing a comprehensive benchmark for recognition and historical linguistics studies.

多类型古汉字识别(Multi-Type Ancient Chinese Character Recognition, MTACCR)数据集是专为推进古汉字文字分析研究而构建的大规模学术资源。该数据集以《通用规范汉字表》为基础搭建,涵盖三级字表共7874个汉字,其中一级字表3500个、二级字表3000个、三级字表1605个,同时包含规范字、繁体字与异体字三类字形。数据集拥有逾900万张样本,涵盖类型丰富的古汉字图像——从原始扫描手稿到经分割处理的字形均有覆盖——数据来源包括书法数据库、开源数据集,并通过数据增强技术扩充样本。MTACCR在汉字覆盖范围、字形类型多样性与数据集规模三方面均显著优于现有同类数据集,可为古汉字识别与历史语言学研究提供全面的基准测试支撑。

创建时间:
2025-06-08
二维码
社区交流群
二维码
科研交流群
商业服务