IGN Synthetic Train Data for ICDAR'25 MapText Competition
收藏资源简介:
Data set of 2Kx2K synthetic image tiles for the ICDAR'25 Competition on Historical Map Text Detection, Recognition, and Linking. Annotations and images follow the format described at the competition website and can be evaluated using the official evaluation repository script. This synthetic dataset is supplementary to the dataset of real tiles IGN Train and Validation Data for ICDAR'25 MapText Competition. This synthetic training set mimics the style (background and fonts) of the original maps, and leverages the actual, modern land use database from the French government to generate realistic geometries and names from similar geographic areas (both in terms of vocabulary and urban density). This synthetic data is meant to be used as a supplementary training set, and is organized as such. We also provide a sample for fast download and code testing, containing only the images and ground truth for the first 10 images of the dataset. Synthetic Train Sample (in sample.zip) Annotations ign25synth_train.json (same) Images synthtrain.zip (same) Files ign25synth/train/*.jpg (same) Tiles 18,072 10 Map Sheets a dozen of different styles 1 style Words 1,615,354 111 Label Groups 1,485,898 90 Illegible Words 34 0 Truncated Words 75,569 2 Valid Words 1,539,785 109 All data used to generate this dataset is public domain.



