Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground–Aerial Robotic Teams
收藏资源简介:
We have created a planetary rover/aerial cross-view dataset from images and motion data captured with in the Extraterrestrial Environment Simulation (Exterres) laboratory at the University of Adelaide. This is complemented by a higher-volume synthetic dataset designed to appear similar to the laboratory environment, generated using PANGU. This dataset comprises the data used in the paper “Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground–Aerial Robotic Teams” by Lachlan Holden, Feras Dayoub, Alberto Candela, David Harvey, and Tat-Jun Chin, in the International Conference on Space Robotics 2025 in Sendai, Japan. Raw data The archive calibration-images.tar.gz contains a set of checkerboard calibration images for the camera CG1 used for all laboratory ("real") data collection. The checkerboard is a 7x10 checkerboard comprising 25mm squares. The configuration-*.tar.gz archives contain the laboratory collected data for each run. Each archive contains the data for all runs for a specific rock configuration. The configurations are numbered from 002 due to a prior, internal test data collection attempt being given the index 001. The filenames within these archives can be interpreted as follows: /> date / /> rock configuration (globally indexed across dates) / / /> run ID (indexed per configuration) / / / /> filetype code (see below), with opt. index / / / / /> camera code (if relevant) / / / / / /> file extension as normal 2024-001-R01-GDV_01-CG1.mp4 All data in this dataset uses camera code CG1. The filetype codes are as follows: GDV: ground video AES: aerial still for these, no suffix is raw image, -calibrated is with distortion removed, and -rectified is lined up with OptiTrack 3D motion tracking data according to sidecar JSON file OPT: OptiTrack (motion tracking) take OPC: OptiTrack calibration data SYNC: Manually-processed synchronisation data The specific runs that correspond to trajectories A through F in the paper are as follows: 004-R01 A 005-R01 B 007-R05 C 004-R02 D 005-R03 E 007-R04 F Image pairs The various pairs-*.tar.gz archives contain aerial/ground image pairs that were directly used for training and validation of our neural network backbone. The filenames correspond to their entries in Table III of the paper. The pairs-synthetic-rgb-{train,val}.tar.gz archives contain PANGU-generated synthetic image pairs, and -mask their corresponding foundation-model-generated rock masks. The pairs-real-rgb-{train,val}.tar.gz archives contain image pairs sampled from the laboratory data. The validation set comprises images from runs 004-R01, 004-R02, 005-R01, 005-R07, 007-R01, and 007-R05. The training set comprises images from runs 002-R02, 002-R02, 003-R01, 003-R07, 006-R01, and 006-R02. Crucially, as mentioned in the paper, the training set does not contain data from any of the rock configurations used for trajectories A through F, above. The pairs-real-mask-val.tar.gz archive contains manually-guided masked image pairs based on real laboratory data. There is no corresponding training data, as this was just used to evaluate the foundation model segmentation performance as well as the ability for a network trained on only synthetic masks to generalise.



