SynthMap+ (English) Synthetic Train Data for ICDAR'25 MapText Competition
收藏资源简介:
Dataset of synthetic map images in English for the ICDAR'25 Competition on Historical Map Text Detection, Recognition, and Linking. Annotations and images follow the format described at the competition website. Please refer to [1] for the generation process and usage. We extend [1] to provide grouping labels for location phrases. Train Annotations en25synth_train.json Images train.zip Files en25synth/train/*.jpg Tiles 35,000 Map Sheets - Words 348,494 Label Groups 157,483 Label Groups (Group Size > 1) 133,955 Illegible Words 0 Truncated Words 0 Valid Words 348,494 [1] Lin, Y., & Chiang, Y. -Y. (2024). Hyper-local deformable transformers for text spotting on historical maps. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5387-5397).
本数据集为面向ICDAR'25历史地图文本检测、识别与关联竞赛的英文合成地图图像数据集。 数据集的标注与图像均遵循竞赛官网公布的格式规范。 有关数据集的生成流程与使用方式,请参见参考文献[1];本数据集在[1]的基础上新增了位置短语的分组标注。 训练集 标注文件:en25synth_train.json 图像文件:train.zip 文件路径:en25synth/train/*.jpg 图像瓦片数:35,000 地图幅数:无 文本总词数:348,494 标注组总数:157,483 规模大于1的标注组数量:133,955 无法识别的文本词数:0 截断文本词数:0 有效文本词数:348,494 [1] Lin, Y. 与 Chiang, Y.-Y. (2024). 面向历史地图文本定位的超局部可变形Transformer(Transformer)方法. 见:第30届ACM SIGKDD知识发现与数据挖掘大会论文集, 第5387-5397页。



