oxford-spires-multimodal
收藏资源简介:
# Dataset Card for Oxford Spires Multimodal (MCAP)  A FiftyOne build of the **Oxford Spires Dataset**, the large-scale LiDAR-visual localisation, reconstruction and radiance-field benchmark from the Oxford Robotics Institute (ORI). This build repackages 6 of the 24 source sequences — one per historic Oxford landmark — as time-synchronised [MCAP](https://mcap.dev/) recordings for FiftyOne's native [multimodal dataset support](https://docs.voxel51.com/user_guide/multimodal.html) (FiftyOne 1.19+). Each sample is one episode, viewable in FiftyOne's tiled multimodal viewer with synchronised three-camera fisheye imagery, motion-undistorted LiDAR point clouds, IMU telemetry, and a live 6-DoF pose track from LiDAR-inertial SLAM. Alongside the raw sensor streams, each episode also carries the three per-camera products of the source devkit's own `generate_depth.py` pipeline — 16-bit depth maps, HSV depth overlays on the camera image, and surface-normal maps — logged as additional streams so the devkit's canonical visualisation is reproducible inside the App. Oxford Spires is a raw multi-sensor dataset for benchmarking SLAM, Structure-from-Motion, Multi-View Stereo, NeRF and 3D Gaussian Splatting methods; it carries **no object-level annotations**. "Ground truth" in the source dataset means millimetre-accurate Terrestrial LiDAR Scanner (TLS) 3D models and the centimetre-accurate trajectories registered against them. This repackaging does not add or alter any ground truth; see [Dataset Creation](#dataset-creation) for exactly what was kept, converted, and left out. This is a [FiftyOne](https://github.com/voxel51/fiftyone) dataset with 6 samples. ## Installation If you haven't already, install FiftyOne: ```bash pip install -U fiftyone ``` ## Usage ```python import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/oxford-spires-multimodal") # Launch the App session = fo.launch_app(dataset) ``` ## Dataset Details ### Dataset Description The Oxford Spires Dataset was captured in and around six well-known historic landmarks in Oxford, UK, using a custom handheld multi-sensor perception unit called **Frontier**, carried in a backpack at walking pace. The unit comprises three synchronised global-shutter colour fisheye cameras (forward-, left- and right-facing), a 64-beam automotive 3D LiDAR, and an inertial sensor, all precisely calibrated. Each site is additionally covered by a millimetre-accurate reference 3D model captured with a Terrestrial LiDAR Scanner, which the authors use both as reconstruction ground truth and — via ICP registration of the mobile LiDAR scans — as the source of centimetre-accurate ground-truth trajectories. In total the source dataset contains 24 sequences across the six sites, covering more than 125,000 m² (about the size of a small town), with the average distance travelled per sequence exceeding 400 metres. The three forward/left/right camera configuration is a distinguishing feature: it widens the field of view for texture mapping and supplies the extra view constraints that vision-only methods need to infer 3D structure from a single linear pass through an environment. The authors establish three benchmarks on this data — localisation, 3D reconstruction, and novel-view synthesis — and use them to show that state-of-the-art radiance field methods overfit to training poses and generalise poorly to out-of-sequence viewpoints. This FiftyOne build covers 6 full-length episodes, one per site (see [Curation Rationale](#curation-rationale)). - **Curated by:** Oxford Robotics Institute, Department of Engineering Science, University of Oxford, in collaboration with the Group of Automation, Robotics and Computer Vision (AUROVA), University of Alicante — original data collection, sensor calibration, TLS reference models, and ground-truth trajectory post-processing. This MCAP/FiftyOne multimodal repackaging (episode authoring, dataset card) was prepared independently by Harpreet Sahota. - **Funded by:** Partly funded by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT), No. RS-2024-00461409. Miguel Ángel Muñoz-Bañón is supported by the Valencian Community Government and the European Union through the CIBEST/2023/44 fellowship and the PROMETEO/2021/075 project. - **Shared by:** Harpreet Sahota (this repackaging); the original Oxford Spires Dataset is shared by the Oxford Robotics Institute via https://dynamic.robots.ox.ac.uk/datasets/oxford-spires/ and the [`ori-drs/oxford_spires_dataset`](https://huggingface.co/datasets/ori-drs/oxford_spires_dataset) Hugging Face dataset repository. - **Language(s):** N/A (sensor data — camera, LiDAR, IMU, pose; no text). - **License:** [CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/) — Copyright (c) 2024, University of Oxford; intended for non-commercial academic use. Commercial use requires contacting the original authors at oxfordspiresdataset@robots.ox.ac.uk. ### Dataset Sources - **Repository:** https://github.com/ori-drs/oxford_spires_dataset (devkit, `oxspires_tools`); data at https://huggingface.co/datasets/ori-drs/oxford_spires_dataset - **Paper:** Tao, Y., Muñoz-Bañón, M. Á., Zhang, L., Wang, J., Fu, L. F. T., & Fallon, M. (2025). *The Oxford Spires Dataset: Benchmarking Large-Scale LiDAR-Visual Localisation, Reconstruction and Radiance Field Methods*. The International Journal of Robotics Research. arXiv:[2411.10546](https://arxiv.org/abs/2411.10546); [SAGE](https://journals.sagepub.com/doi/10.1177/02783649251369905) - **Demo:** https://dynamic.robots.ox.ac.uk/datasets/oxford-spires/ (official project site) and https://www.youtube.com/watch?v=AKZ-YrOob_4 (dataset video) ## Uses ### Direct Use - Exercising/demoing FiftyOne's multimodal MCAP support: synchronised playback of three fisheye camera streams, motion-undistorted LiDAR point clouds, IMU telemetry, and a moving 6-DoF sensor pose across full-length, real handheld recordings (3.5–13.7 minutes each). - Inspecting LiDAR-camera calibration quality interactively — the depth overlay streams (`/overlay_cam_*`) reproduce the devkit's own LiDAR-on-image visualisation, which the paper uses (Figs. 4 and 5) to demonstrate calibration and motion-undistortion accuracy. - Prototyping analyses that need per-frame LiDAR-to-camera projection: 16-bit euclidean depth maps and surface-normal maps are logged per camera at every image-synchronised keyframe. - Qualitative review of LiDAR-inertial SLAM trajectory behaviour, with the `world` → `base` transform track driving the 3D tile's camera and point-cloud placement. - Browsing site-to-site variation (limestone vs. Keble's alternating red and white brick; open squares vs. narrow passages; indoor/outdoor lighting transitions) via the `location` sample field. ### Out-of-Scope Use - Reproducing the paper's localisation, reconstruction, or novel-view synthesis benchmark numbers. Those benchmarks require artifacts that are **not** in this build: the TLS ground-truth 3D models, the TLS-registered ground-truth trajectories (`gt-tum.txt`), the HBA and COLMAP trajectory variants, and the COLMAP SfM outputs. See [Parsing decisions](#parsing-decisions). - Any task needing object-level labels (detection, segmentation, classification). No such annotations exist anywhere in the source dataset. - Metric evaluation against the poses in this build. The pose track here is the **VILENS-SLAM estimate** shipped in `slam-poses.csv`, which is itself one of the systems *evaluated* in the paper's localisation benchmark (ATE 0.03–1.11 m depending on sequence), not the TLS-registered ground truth (≈1–2 cm accuracy). - Photometrically consistent colour reconstruction from the merged `/lidar_rgb` clouds. The paper explicitly flags this as an open problem for this data: camera auto-exposure was enabled, so the same 3D structure observed from different viewpoints has inconsistent pixel intensity, and merging colourised LiDAR clouds "would lead to a mixture of different colours in the reconstruction." - Map-based workflows. There is no GNSS/GPS stream anywhere in the source dataset, so the App's Map tile is empty for every episode by design. ## Dataset Structure This is a flat (ungrouped) FiftyOne dataset with `media_type: "multimodal"` and **6 samples**. Each sample is one **episode**, stored as one `.mcap` file; FiftyOne infers the multimodal media type automatically from the `.mcap` extension. There are no separate per-frame image or point-cloud samples — the episode is the sample unit, and every stream inside it (cameras, LiDAR, IMU, transforms, and the derived depth/overlay/normal products) is decoded live by FiftyOne's multimodal viewer. The dataset carries no per-sample tags, no temporal tags, and `dataset.info` is empty — there is no extra dataset-level metadata beyond the per-sample fields below. Sensor calibration is not stored in `dataset.info`; it lives inside each MCAP as `foxglove.CameraCalibration` and `foxglove.FrameTransform` messages, so the viewer can use it directly. The built-in `metadata` field is unpopulated (`None`) because `compute_metadata()` was not run. Totals across the 6 episodes: 1,091,709 MCAP messages, 2,481.3 seconds (41.4 minutes) of recording, 37 GB of MCAP on disk. ### Episodes in this dataset | `sequence_id` | `location` | `recording_date` | `duration_s` | `message_count` | `channel_count` | `/lidar` clouds | `has_rgb_lidar` | |---|---|---|---|---|---|---|---| | `2024-03-12-keble-college-02` | `keble-college` | 2024-03-12 | 300.1 | 148,939 | 18 | 805 | False | | `2024-03-13-observatory-quarter-02` | `observatory-quarter` | 2024-03-13 | 275.1 | 131,263 | 18 | 361 | False | | `2024-03-14-blenheim-palace-05` | `blenheim-palace` | 2024-03-14 | 372.9 | 160,664 | 19 | 361 | True | | `2024-03-20-christ-church-05` | `christ-church` | 2024-03-20 | 822.1 | 336,685 | 19 | 713 | True | | `2024-05-20-bodleian-library-02` | `bodleian-library` | 2024-05-20 | 503.0 | 219,393 | 19 | 521 | True | | `2024-07-09-new-college-01` | `new-college` | 2024-07-09 | 208.1 | 94,765 | 19 | 192 | True | The two 18-channel episodes lack `/lidar_rgb` because their source point clouds already ship with colour baked in — see [Parsing decisions](#parsing-decisions). The effective image rate varies markedly between episodes (from ≈4.8 Hz for `christ-church-05` to ≈20 Hz for `keble-college-02`, computed as image count divided by episode duration). The paper documents the raw camera streams as 20 Hz, so the `raw/images.zip` archives for some sequences appear to be a decimated subset rather than the full stream; the counts logged here match those archives exactly. ### Fields | Field | FiftyOne type | Description | |-------|---------------|-------------| | `filepath` | `StringField` | Absolute path to the episode's `.mcap` file — the sample's multimodal media | | `sequence_id` | `StringField` | Source sequence name (`YYYY-MM-DD-<site>-<NN>`), verbatim from the source repository's folder name | | `location` | `StringField` | Site slug parsed from `sequence_id` (e.g. `keble-college`); one of 6 | | `recording_date` | `StringField` | Recording date parsed from `sequence_id`, `YYYY-MM-DD` | | `run_number` | `IntField` | Run index at that site, parsed from the `sequence_id` suffix | | `duration_s` | `FloatField` | Episode duration in seconds, computed from the MCAP's message-time span | | `message_count` | `IntField` | Total MCAP message count across all channels | | `channel_count` | `IntField` | Total MCAP channel (topic) count — 18 or 19 | | `topics` | `ListField(StringField)` | Every MCAP topic present (see [MCAP topics](#mcap-topics-inside-each-episode)) | | `schemas` | `ListField(StringField)` | Every distinct message schema present in the episode | | `has_image` | `BooleanField` | Has an Image-tile-decodable stream (`foxglove.CompressedImage` or `foxglove.RawImage`) — `True` for all | | `has_pointcloud` | `BooleanField` | Has a 3D-tile-decodable point-cloud stream (`foxglove.PointCloud`) — `True` for all | | `has_gps` | `BooleanField` | Has a Map-tile GPS fix stream (`foxglove.LocationFix`) — `False` for all; no GNSS exists in the source | | `has_imu` | `BooleanField` | Has IMU telemetry — `True` for all; hardcoded, because the stream is logged as generic JSON rather than a recognised IMU schema | | `has_logs` | `BooleanField` | Has a Logs-tile stream (`foxglove.Log`) — `False` for all | | `has_trajectory` | `BooleanField` | Has a 6-DoF pose track (the `world` → `base` `FrameTransform` series) — `True` for all | | `has_camera_calibration` | `BooleanField` | Has `foxglove.CameraCalibration` streams — `True` for all | | `has_depth` | `BooleanField` | Has per-camera 16-bit depth maps (`/depth_*`) — `True` for all | | `has_depth_overlay` | `BooleanField` | Has per-camera depth overlays on the camera image (`/overlay_*`) — `True` for all | | `has_surface_normals` | `BooleanField` | Has per-camera surface-normal maps (`/normal_*`) — `True` for all | | `has_rgb_lidar` | `BooleanField` | Has the derived camera-colourised full-sweep cloud (`/lidar_rgb`) — `True` for 4 of 6 episodes | | `num_cameras` | `IntField` | Number of cameras in the rig — 3 for every episode | Standard FiftyOne bookkeeping fields (`id`, `tags`, `metadata`, `created_at`, `last_modified_at`) are also present but not source-specific. ### MCAP topics (inside each episode) Cameras use the devkit's own labels from `configs/sensor.yaml`: `cam_front`, `cam_left`, `cam_right` (source directories `cam0`, `cam1`, `cam2` respectively). | Topic(s) | Schema | Tile | Notes | |----------|--------|------|-------| | `/cam_front/image_raw`, `/cam_left/image_raw`, `/cam_right/image_raw` | `foxglove.CompressedImage` (jpeg) | Image | Raw, still-distorted fisheye JPEGs (1440×1080), byte-for-byte from the source `raw/cam{0,1,2}` folders | | `/cam_front/calibration`, `/cam_left/calibration`, `/cam_right/calibration` | `foxglove.CameraCalibration` | 3D (frustums) | Static, logged once at episode start; `equidistant` distortion model with `K`/`D` from `configs/sensor.yaml` | | `/lidar` | `foxglove.PointCloud` | 3D | VILENS-SLAM motion-undistorted keyframe clouds (≈1 Hz, one per SLAM pose-graph node), in frame `base`. Fields `x,y,z,intensity,normal_x,normal_y,normal_z,curvature`, or `x,y,z,red,green,blue,normal_*,curvature` for the two episodes whose source clouds ship colour | | `/lidar_rgb` (4 of 6 episodes) | `foxglove.PointCloud` | 3D | **Derived, not a devkit product.** The full sweep with RGB sampled by projecting every point into all three cameras; points no camera sees keep an intensity-derived grey. Fields `x,y,z,red,green,blue` | | `/depth_cam_front`, `/depth_cam_left`, `/depth_cam_right` | `foxglove.RawImage` (`16UC1`) | Image | Euclidean depth × 256 as uint16, equivalent to the devkit's `depths_euc_accum_0` output | | `/overlay_cam_front`, `/overlay_cam_left`, `/overlay_cam_right` | `foxglove.CompressedImage` (jpeg) | Image | Depth-coloured LiDAR points drawn over the camera image (HSV colormap, radius-2 circles), equivalent to the devkit's `*_overlay` output | | `/normal_cam_front`, `/normal_cam_left`, `/normal_cam_right` | `foxglove.CompressedImage` (png) | Image | Surface normals from the PCD `normal_*` fields, rotated into the camera frame and encoded `(n+1)/2·255`, equivalent to the devkit's `normals_euc_accum_0` output | | `/tf` | `foxglove.FrameTransform` | 3D | One static `base` → `lidar` transform at episode start, plus per SLAM keyframe: `world` → `base` and `world` → {`cam_front`, `cam_left`, `cam_right`} | | `/imu` | generic JSON (auto-named `jsonschema`, e.g. `schema-f5wrpOiP`) | Plot / Message | `acc_x/y/z` and `ang_vel_x/y/z` at the IMU's native rate, from `raw/imu.csv` | ### Label types and why **No FiftyOne sample-level label fields (`Detections`, `Keypoints`, `Classification`, etc.) are attached.** Two reasons: 1. The source dataset has no object-level annotations at all — it is a raw multi-sensor SLAM/reconstruction benchmark. Its "ground truth" is TLS 3D models and 6-DoF trajectories, used to compute benchmark metrics, not per-sample labels. 2. Each sample is a continuous 3.5–13.7 minute recording rather than a single frame, so there is no single fixed-length list that a sample-level label field could hold. The one label-like quantity that does exist per episode — the 6-DoF trajectory — is therefore embedded as an MCAP transform stream (`/tf`) inside the same timeline as the sensor data, decoded live by the multimodal viewer, exactly like the sensor topics themselves. This also lets the viewer place point clouds and camera frustums correctly in the world frame during playback. It was logged as `foxglove.FrameTransform` rather than as a static pose list because the viewer consumes transforms to resolve frames at each timestamp. All remaining per-sample fields are primitives, not labels: identifiers (`sequence_id`, `location`, `recording_date`, `run_number`), MCAP inventory (`duration_s`, `message_count`, `channel_count`, `topics`, `schemas`), and boolean capability flags (`has_*`). The flags exist so episodes can be filtered in the grid without opening every MCAP first, e.g. `dataset.match(F("has_rgb_lidar"))`. They are derived from **schema names, not topic substrings**, except `has_imu` (hardcoded `True`, since the IMU rides on a generic JSON channel) and the `has_depth`/`has_depth_overlay`/ `has_surface_normals`/`has_rgb_lidar` flags (derived from topic prefixes, since those products share the generic image/point-cloud schemas). ### Parsing decisions - **One sample = one full episode.** The source Hugging Face repository is roughly 1.3 TB, so 6 of the 24 sequences were selected — one per site, choosing the smallest sequence at each site by download size — and each was authored as a complete recording rather than a short window. - **Built from the per-file raw artifacts, not the ROS bags.** The source ships every sequence as both ROS 1 `.bag` and ROS 2 `.db3` (2–15 GB each) *and* as individual files. This build reads the individual files (`raw/images.zip`, `raw/imu.csv`, `processed/vilens-slam/undist-clouds.zip`, `processed/vilens-slam/slam-poses.csv`), which avoids the bag conversion entirely and gives the motion-undistorted, image-synchronised clouds the devkit's own tooling expects. - **Devkit frame naming.** Frames are `base`, `lidar`, `cam_front`, `cam_left`, `cam_right`, matching the labels in `configs/sensor.yaml` rather than inventing ROS-style names. - **Cloud points are in the body frame, so they are logged as `base`.** The undistorted PCD files' `VIEWPOINT` header equals the matching `slam-poses.csv` row, and the devkit's own `undistort_sequence.py` writes an identity viewpoint — i.e. the point data itself is already expressed in the sensor/body frame at capture time. Verified directly: a mid-sequence cloud whose pose is 65 m from the origin has its point centroid at the origin. - **Camera extrinsics follow the devkit's own composition.** `T_base_cam = T_base_lidar @ inv(T_cam_lidar)`, the same expression used in the devkit's `scripts/reconstruction_benchmark/main.py`, with the per-camera `T_cam_lidar` and rig `T_base_lidar` read from `configs/sensor.yaml`. Cameras are logged as direct time-varying `world` → `cam_*` transforms at each SLAM keyframe, so the viewer never has to chain through a single-sample static transform to resolve `world`. - **Depth/overlay/normal maps are ports of the devkit's own functions** (`oxspires_tools.depth.projection.encode_points_as_depthmap`, `depth.utils.get_overlay`, `depth.surface_normal.compute_normalmap`): euclidean depth scaled by 256 into uint16, nearest point winning per pixel, HSV colormap overlay with radius-2 circles, and normal maps that flip normals toward the camera and reserve `(128,128,128)` for empty pixels. They were reimplemented in numpy/OpenCV rather than called directly because the devkit's projection module imports `open3d`, which has no wheel for this machine's platform. - **Hidden-point removal is skipped.** It is the one step of the devkit's depth pipeline that genuinely requires `open3d`; the devkit itself exposes this as a supported `--skip_hpr` flag. For a single LiDAR keyframe sweep this mainly affects points on the far side of thin occluders. - **The projection field-of-view cone is 160°**, taken from the devkit's `camera_fov: 160.0` in `configs/sensor.yaml`, which is the value its own projection code uses to filter points before `cv2.fisheye.projectPoints`. Note this differs from the per-camera field of view quoted in the paper (126° × 92.4°); the devkit value was kept for fidelity to the devkit's output. - **Image-to-cloud pairing uses a 25 ms threshold** (`max_time_diff_camera_and_pose: 0.025` from `configs/sensor.yaml`), matching each cloud to the nearest image per camera. Cameras that fall outside the threshold for a given keyframe simply get no derived frame there, which is why per-camera depth/overlay/normal counts differ slightly within an episode. - **`/lidar_rgb` is only built where the source clouds lack colour.** The `keble-college-02` and `observatory-quarter-02` undistorted clouds carry a packed PCL `rgb` field instead of `intensity` (colourised upstream by the dataset authors); for those, the colour is decoded into `red`/`green`/ `blue` on `/lidar` itself and the derived `/lidar_rgb` is skipped as a strictly worse duplicate — camera projection can only colour the 56–84% of points that fall inside a camera's view, whereas the upstream colour covers the whole sweep. - **The IMU rides on a generic JSON channel**, not `foxglove.Imu`, because `imu.csv` provides only linear acceleration and angular velocity — no orientation or covariance. Consequence: it is inspectable in the Plot and Message tiles but is not reported under `schemas` as a recognised IMU schema, hence the hardcoded `has_imu`. - **Real capture timestamps throughout.** Every message is logged at its source timestamp (Unix epoch nanoseconds) with no rebasing, so `duration_s` reflects the true recording span and all streams share one coherent clock. - **Clouds without a matching SLAM pose are still logged.** `blenheim-palace-05` has 361 undistorted clouds but only 340 poses in `slam-poses.csv`; all 361 clouds appear on `/lidar`, while `/tf` carries 340 keyframes, so the surplus clouds render at the last resolved transform. - **Not included in this build:** the raw 10 Hz LiDAR sweeps (`raw/lidar-clouds.zip`), the COLMAP SfM outputs (`processed/colmap/`), the TLS ground-truth 3D models (`ground_truth_map/`), the TLS-registered ground-truth and refined trajectory variants (`gt-tum.txt`, `hba-tum.txt`, `colmap-tum.txt`), the TLS-rendered ground-truth depth images, and the reconstruction and novel-view-synthesis benchmark artifacts. Note also that the source repository does not currently ship `processed/trajectory/` or `processed/colmap/` for the New College sequences at all. ## Dataset Creation ### Curation Rationale The full source repository is roughly 1.3 TB — every sequence ships raw images, raw LiDAR sweeps, ROS 1 and ROS 2 bags, VILENS-SLAM outputs and COLMAP outputs, on top of per-site TLS reference maps (about 76 GB) and benchmark result artifacts. Exhaustive coverage is impractical for a lightweight FiftyOne showcase, so this build optimises for **site diversity at minimum download**: one sequence per landmark, choosing the smallest available sequence at each site, giving all six architectural settings the paper describes (Bodleian Library ≈37,000 m²; Christ Church College ≈26,000 m²; Keble College ≈18,000 m²; New College ≈18,000 m²; Blenheim Palace ≈14,000 m²; Radcliffe Observatory Quarter ≈12,000 m²). Within each episode the priority was the opposite of trimming: episodes are complete recordings, and the streams were chosen to reproduce what the devkit itself visualises — hence the inclusion of all three `generate_depth.py` products rather than depth maps alone. ### Source Data #### Data Collection and Processing Per the paper: each sequence was collected by walking with the **Frontier** handheld perception unit mounted in a backpack. The unit carries three colour fisheye cameras facing forward, left and right — a customised Alphasense Core Development Kit from Sevensense Robotics AG — each 1440×1080 (1.6 MP) global shutter with a 126° × 92.4° field of view and roughly 36° of overlap between adjacent cameras, running at 20 Hz with auto-exposure enabled. A cellphone-grade IMU inside the Alphasense Core runs at 400 Hz and is hardware-synchronised to the three cameras by a Sevensense FPGA. A 64-channel Hesai QT64 LiDAR (10 Hz, 104° field of view, 60 m maximum range, ±3 cm typical accuracy) is mounted on top of the cameras. Synchronisation is both hardware and software: the Alphasense device clock and the Hesai LiDAR are synchronised to the unit's host computer using Precision Time Protocol (sub-microsecond accuracy); the cameras' exposure intervals are aligned about their midpoints so the image triplets share one timestamp; and because the QT64 scans continuously, each point cloud is motion-corrected with IMU preintegration using VILENS and undistorted to the time of the next camera frame. The result is that every node in the SLAM pose graph has three camera images and one undistorted LiDAR cloud at an identical timestamp. Calibration used the equidistant (Kannala-Brandt) model for the fisheye lenses: camera intrinsics and inter-camera extrinsics with Kalibr (sub-pixel reprojection residuals, 0.22–0.23 px mean), IMU noise from an eight-hour Allan variance sequence, per-camera camera-IMU extrinsics with Kalibr, and a single SE(3) camera-bundle-to-LiDAR transform with DiffCal. Reference data: a Leica RTC360 TLS (360° × 300° field of view, 130 m range, 1.9 mm point accuracy at 10 m and 5.3 mm at 40 m, colourised from 432 megapixel imagery) scanned each site; scans were registered with Leica Cyclone REGISTER 360 Plus to 3–7 mm average cloud-to-cloud error and merged into a 1 cm colourised map. Ground-truth trajectories were then produced by ICP-registering each undistorted LiDAR cloud to that merged map with an offline version of VILENS, reaching approximately 1–2 cm accuracy. Processed outputs released alongside the raw data include the VILENS-SLAM trajectory and undistorted clouds, and COLMAP SfM results computed over images spaced 1 m apart (about 1 Hz at walking pace). For this repackaging: the six sequences' individual-file artifacts were downloaded from the public Hugging Face repository, extracted, and packed into one `.mcap` file per episode with the `foxglove-sdk`, computing the depth/overlay/normal products with ports of the devkit's own projection code. Every episode was then verified against its own source files — image counts per camera, IMU row count, cloud count, and expected transform count all had to match exactly, with per-channel timestamp monotonicity checked — before ingest. No sensor data was synthesized, relabeled, or altered beyond the conversions documented in [Parsing decisions](#parsing-decisions). #### Who are the source data producers? The Oxford Robotics Institute, Department of Engineering Science, University of Oxford, with the Group of Automation, Robotics and Computer Vision (AUROVA) at the University of Alicante — specifically Yifu Tao, Miguel Ángel Muñoz-Bañón, Lintong Zhang, Jiahao Wang, Lanke Frank Tarimo Fu, and Maurice Fallon. The paper additionally acknowledges Tobit Flatscher, Ayoung Kim, Matias Mattamala, Christina Kassab, Haedam Oh, Jianeng Wang, and Dongjae Lee for help with sensing, calibration, collection, post-processing and proofreading. ### Annotations #### Annotation process There is no annotation process: the Oxford Spires Dataset contains no manually labelled, object-level annotations. What the source dataset calls ground truth is produced automatically — reconstruction ground truth is the registered and merged Leica RTC360 TLS point cloud per site, and localisation ground truth is the trajectory obtained by ICP-registering each motion-undistorted LiDAR cloud against that TLS map using an offline version of VILENS. Neither of those artifacts is included in this build; the pose track logged here is the VILENS-SLAM *estimate* from `slam-poses.csv` (see [Out-of-Scope Use](#out-of-scope-use)). #### Who are the annotators? No human annotators were involved. All ground-truth quantities in the source dataset are the output of automated survey-grade scanning, registration, and SLAM/ICP post-processing performed by the dataset authors. #### Personal and Sensitive Information The recordings were made in and around public and semi-public university and palace grounds in Oxford, so incidental pedestrians can appear in the camera imagery. The ROS bag filenames in the source repository contain `blurred_filtered` (e.g. `..._blurred_filtered_compressed.db3`), which indicates that a privacy filtering step — presumably face and licence-plate blurring — was applied upstream to the bag streams before public release. [More Information Needed] on the specifics: neither the paper nor the devkit documents an anonymisation procedure or tool, and it is not documented whether the `raw/images.zip` JPEG archives used to build these episodes carry the same blurring as the bag streams. This repackaging performs no additional processing, re-identification, or redaction beyond what the Oxford Robotics Institute already released publicly. ## Citation **BibTeX:** ```bibtex @article{tao2025spires, title = {The Oxford Spires Dataset: Benchmarking Large-Scale LiDAR-Visual Localisation, Reconstruction and Radiance Field Methods}, author = {Tao, Yifu and Mu{\~n}oz-Ba{\~n}{\'o}n, Miguel {\'A}ngel and Zhang, Lintong and Wang, Jiahao and Fu, Lanke Frank Tarimo and Fallon, Maurice}, journal = {The International Journal of Robotics Research}, year = {2025}, doi = {10.1177/02783649251369905} } ``` **APA:** Tao, Y., Muñoz-Bañón, M. Á., Zhang, L., Wang, J., Fu, L. F. T., & Fallon, M. (2025). The Oxford Spires Dataset: Benchmarking large-scale LiDAR-visual localisation, reconstruction and radiance field methods. *The International Journal of Robotics Research*. ## More Information This repository is an independently-curated, derived subset of the official Oxford Spires Dataset, repackaged as MCAP for FiftyOne's multimodal support. It is not an official Oxford Robotics Institute artifact, and it is subject to the source dataset's non-commercial CC BY-NC-SA 4.0 licence. For the full dataset (all 24 sequences, raw LiDAR sweeps, ROS 1/ROS 2 bags, COLMAP outputs, per-site TLS reference maps, TLS-registered ground-truth trajectories, rendered ground-truth depth images, and the localisation / reconstruction / novel-view-synthesis benchmark artifacts and evaluation code), see: - https://dynamic.robots.ox.ac.uk/datasets/oxford-spires/ - https://huggingface.co/datasets/ori-drs/oxford_spires_dataset - https://github.com/ori-drs/oxford_spires_dataset and its [wiki](https://github.com/ori-drs/oxford_spires_dataset/wiki) Viewing these episodes requires FiftyOne 1.19 or newer for multimodal media support. The 6 MCAP files total 37 GB. ## Dataset Card Authors Harpreet Sahota ([@harpreetsahota](https://huggingface.co/harpreetsahota)) — MCAP repackaging and this card. Original dataset producers are listed under [Dataset Description](#dataset-description). ## Dataset Card Contact Harpreet Sahota — https://huggingface.co/harpreetsahota



