遇见数据集

IGN Synthetic Train Data for ICDAR'25 MapText Competition

收藏
Zenodo2025-01-20 更新2026-05-26 收录
官方服务:

资源简介:

Data set of 2Kx2K synthetic image tiles for the ICDAR'25 Competition on Historical Map Text Detection, Recognition, and Linking. Annotations and images follow the format described at the competition website and can be evaluated using the official evaluation repository script. This synthetic dataset is supplementary to the dataset of real tiles IGN Train and Validation Data for ICDAR'25 MapText Competition. This synthetic training set mimics the style (background and fonts) of the original maps, and leverages the actual, modern land use database from the French government to generate realistic geometries and names from similar geographic areas (both in terms of vocabulary and urban density). This synthetic data is meant to be used as a supplementary training set, and is organized as such. We also provide a sample for fast download and code testing, containing only the images and ground truth for the first 10 images of the dataset. Synthetic Train Sample (in sample.zip) Annotations ign25synth_train.json (same) Images synthtrain.zip (same) Files ign25synth/train/*.jpg (same) Tiles 18,072 10 Map Sheets a dozen of different styles 1 style Words 1,614,631 111 Label Groups 1,485,245 90 Illegible Words 34 0 Truncated Words 74,846 2 Valid Words 1,539,785 109 All data used to generate this dataset is public domain. ℹ️ This version 2 contains a fix in the ign25synth_train.json file from which very small regions (<1 square pixel) were removed to mitigate evaluation issues. This results in a smaller number of total words and groups, but the number of valid words remains the same compared to version 1. The sample in sample.zip and the images in ign25synth_train.zip were not changed and are identical to version 1.

本数据集为ICDAR'25历史地图文本检测、识别与关联竞赛所用的2K×2K合成图像瓦片数据集。 标注与图像均遵循竞赛官网规定的格式,可通过官方评估仓库脚本进行评估。 本合成数据集是ICDAR'25 MapText竞赛的真实瓦片数据集IGN(Institut Géographique National,法国国家地理研究院)训练与验证数据集的补充数据集。 该合成训练集还原了原始地图的风格(包括背景与字体),并依托法国政府公开的现代土地利用数据库,从词汇与城市密度相近的地理区域生成具备真实感的几何形状与地名。本合成数据旨在作为补充训练集,并已按该用途完成组织。 我们还提供了用于快速下载与代码测试的示例集,仅包含数据集前10张图像及其真值标注。 ### 合成训练集 示例(存于sample.zip) ### 标注文件:ign25synth_train.json(格式一致) ### 图像文件:synthtrain.zip(格式一致) ### 文件结构:ign25synth/train/*.jpg(格式一致) 瓦片数量 18,072 10 地图图幅样式 十余种不同风格 1种 词汇总数 1,614,631 111 标注组总数 1,485,245 90 难以识别词汇 34 0 截断词汇 74,846 2 有效词汇 1,539,785 109 本数据集生成所用的全部数据均属于公有领域。 ℹ️ 本V2版本修复了ign25synth_train.json文件,移除了其中面积小于1平方像素的极小区域以缓解评估问题。这导致总词汇数与标注组数有所减少,但有效词汇数量与V1版本保持一致。sample.zip中的示例集与ign25synth_train.zip中的图像均未做修改,与V1版本完全一致。

提供机构:
Zenodo
创建时间:
2025-01-07
二维码
社区交流群
二维码
科研交流群
商业服务