five

<b>M</b>ulti-<b>T</b>ype <b>A</b>ncient <b>C</b>hinese <b>C</b>haracter <b>R</b>ecognition (MTACCR) dataset

收藏
DataCite Commons2025-06-08 更新2025-09-08 收录
下载链接:
https://figshare.com/articles/dataset/_b_M_b_ulti-_b_T_b_ype_b_A_b_ncient_b_C_b_hinese_b_C_b_haracter_b_R_b_ecognition_MTACCR_dataset/29263991
下载链接
链接失效反馈
官方服务:
资源简介:
​<b>​The Multi-Type Ancient Chinese Character Recognition (MTACCR) dataset​</b>​ is a large-scale resource designed to advance research in ancient Chinese script analysis. It is constructed based on the <i>Table of General Standard Chinese Characters</i> (通用规范汉字表), covering ​<b>​7,874 Chinese characters​</b>​ across three levels (3,500 Level-1, 3,000 Level-2, and 1,605 Level-3) and including standard, traditional, and variant forms. With ​<b>​over 9 million samples​</b>​, the dataset features diverse ancient character images—ranging from original scanned manuscripts to segmented glyphs—collected from calligraphic databases, open-source datasets, and data augmentation. MTACCR significantly surpasses existing datasets in character coverage, typological diversity, and scale, providing a comprehensive benchmark for recognition and historical linguistics studies.
提供机构:
figshare
创建时间:
2025-06-08
二维码
社区交流群
二维码
科研交流群
商业服务