遇见数据集

TANVOM v1: Dataset for Orienteering Map Vectorization

收藏
Zenodo2026-04-06 更新2026-05-26 收录
官方服务:

资源简介:

This dataset accompanies research presented at ICAI 2026 on topology-aware neural vectorization of printed orienteering maps. The broader research framework targets multiple ISOM symbol classes, but the present release is primarily designed to support reproducible experiments on processed raster tiles and derived supervision masks for map digitization, contour extraction, thin-line segmentation, and downstream raster-to-vector evaluation. The dataset is derived from high-resolution raster exports of printed orienteering maps following the International Specification for Orienteering Maps (ISOM). Ground-truth supervision was generated from symbol-specific exports and converted into task-specific raster masks using deterministic preprocessing rules. In the associated research pipeline, neural models predict raster outputs, which are then converted into editable vector objects using deterministic post-processing and vectorization. The main methodological motivation is that practical map reconstruction quality cannot be assessed reliably from raster overlap alone; downstream vector structure, topology, and editability must also be considered. Important note on data release scope:For privacy and data-protection reasons, this Zenodo release contains only processed data products. The original full map sheets, original OCAD files, and full-sheet source exports are not redistributed here. Instead, the release provides cropped tile-level RGB inputs and the corresponding derived masks and metadata required for reproducible machine learning experiments. This means that the dataset is suitable for training, validation, benchmarking, and methodological comparison, but it is not intended to recreate or redistribute the original source maps in their full form. Data organizationThe shared dataset is organized at tile level. Full map rasters were partitioned into 512×512 pixel tiles with 128-pixel stride, and each retained base tile was expanded into four augmented variants. The train/validation split is performed at map-sheet level rather than tile level in order to reduce spatial leakage and better reflect generalization to unseen maps. In the current processed snapshot, the dataset contains 21,544 samples in total, with 18,880 training samples and 2,664 validation samples. Tasks and labelsDepending on the included subset, the release may contain one or more of the following targets: 1. Contours:A binary contour mask in which ISOM contour symbols 101 and 102 are merged into a single geometry target. This design favors continuity and downstream vectorizability over subtype discrimination. 2. Roads-black:A binary union mask for black road symbols (ISOM 503–508), intended for geometry-first learning. In the associated pipeline, semantic road typing is performed later, after vectorization. 3. Rare line symbols:Optional per-code binary masks for rare symbols such as 107 and 108, stored as separate channels/files. 4. Region targets:Optional region annotations consisting of:- a base landcover label map (multi-class, one class per pixel), and/or- overlay masks (multi-label binary layers) for additional overprinted region symbols. Preprocessing conventionsThe data were prepared using deterministic preprocessing. Masks are binarized using a fixed convention (“alpha > 0, otherwise non-white / threshold-based foreground detection”), and tile alignment for derived targets is handled in a reproducible way. The pipeline uses ISOM symbol identifiers consistently across mask naming, metadata, manifests, and downstream export logic. Appearance variability such as contrast changes, blur, compression artifacts, and mild scan-like degradation is modeled through augmentation in order to improve robustness to printed/scanned map conditions. What is includedA typical release may include:- processed RGB tile images,- corresponding task-specific mask files,- split information,- metadata and/or manifest files describing paths and class mappings,- optional class mapping files for region or road-type tasks. What is not includedThis release does not include:- original full map images,- original editable cartographic source files,- raw OCAD projects,- full-sheet exports that would allow direct redistribution of the original maps. How to use this datasetThis dataset is intended for research use in:- semantic segmentation of cartographic symbols,- thin-line extraction from printed maps,- contour reconstruction,- map digitization and raster-to-vector studies,- topology-aware evaluation,- benchmarking of deterministic post-processing pipelines. Typical usage:1. Load RGB tile images as model inputs.2. Load the corresponding binary or multi-class masks as supervision targets.3. Respect the provided train/validation split at map level.4. Train segmentation models on the raster prediction task.5. Optionally apply your own skeletonization, tracing, polygonization, or graph-based vectorization methods to compare downstream vector quality. Because only processed tiles are released, methods that require full-map contextual reconstruction should treat this dataset as a tile-based benchmark rather than as a full-sheet archival source. Users interested in end-to-end vectorization can still use the dataset to compare raster predictions, topology-aware metrics, and vectorization behavior under controlled, reproducible conditions. Recommended citationIf you use this dataset, please cite the Zenodo record and also cite the associated ICAI 2026 abstract or any subsequent full paper when available. Supported by the EKOP-25-I-1 scholarship program

提供机构:
Zenodo
创建时间:
2026-04-06
二维码
社区交流群
二维码
科研交流群
商业服务