MTHv2
收藏资源简介:
本项目共享的数据集是高丽藏汉文(TKH)数据集和多版本高丽藏汉文(MTH)数据集。为了促进对中国历史文献的研究,我们扩展了原始数据集的规模,增加了布局、字符和文本行的标注。从互联网上添加了更多具有挑战性的文档图像到MTH数据集,其图像数量现已达到2200张,TKH和MTH的合并数据集被命名为MTHv2。
The dataset shared in this project comprises the Tripitaka Koreana in Chinese (TKH) dataset and the Multi-version Tripitaka Koreana in Chinese (MTH) dataset. To facilitate research on Chinese historical documents, we have expanded the scale of the original datasets by adding annotations for layout, characters, and text lines. More challenging document images from the internet have been added to the MTH dataset, increasing the number of images to 2,200. The combined dataset of TKH and MTH is now named MTHv2.
数据集概述
数据集名称
- Tripitaka Koreana in Han (TKH) Dataset
- Multiple Tripitaka in Han (MTH) Dataset
- MTHv2 (TKH和MTH的合并数据集)
数据集内容
- 原始数据集扩展,增加了布局、字符和文本行的标注。
- 新增了来自互联网的更具挑战性的文档图像,总数达到2200张。
数据集结构
- 包含三种类型的标注:
- 行级标注:文本行位置及其转录,按阅读顺序保存。
- 字符级标注:包括类别和边界框坐标。
- 边界线:由线段的起始和结束点表示。
数据集划分
- 随机分为训练集和测试集,比例为3:1。
数据集下载
- Google Drive链接:Google Drive
- Baidu Drive链接:Baidu Drive 提取码: eweb
引用信息
@article{ title={Joint Layout Analysis, Character Detection and Recognition for Historical Document Digitization}, author={Weihong Ma, Hesuo Zhang, Lianwen Jin, Sihang Wu, Jiapeng Wang, Yongpan Wang}, journal={ICFHR 2020}, year={2020} }




