YEG3D: A Canadian Benchmark LiDAR Dataset for Urban Infrastructure and 3D Scene Understanding
收藏资源简介:
Overview YEG3D is a large-scale, point-wise annotated mobile laser scanning (MLS) dataset designed to support research in 3D scene understanding, semantic segmentation, autonomous driving, intelligent transportation systems, geomatics, and urban analytics. The dataset was collected in Edmonton, Alberta, Canada, using a vehicle-mounted MLS platform equipped with a HESAI XT32M1X LiDAR sensor, RTK-enabled GNSS, and a high-frequency inertial measurement unit (IMU). The current release contains approximately 682 million annotated LiDAR points covering 14 km of urban roadway corridors and neighborhoods. The dataset is organized into seven spatial zones representing diverse roadway geometries, land-use contexts, and infrastructure characteristics. In total, YEG3D includes 18 semantic classes with a particular emphasis on fine-grained roadway and pedestrian infrastructure that is often underrepresented in existing public LiDAR benchmarks. The semantic taxonomy includes: Class Description Road Paved roadway surfaces, including travel lanes, bus lanes, and on-street parking areas Other Road Driveways, off-street parking lots, and other paved vehicular areas not part of the main roadway Median Raised medians separating opposing or adjacent traffic flows Sidewalk Designated pedestrian sidewalks located adjacent to roadways Shared Pathway Multi-use paths intended for shared pedestrian and cyclist use, typically separated from the road Walkway Walkable spaces not classified as sidewalks, such as building entrances, plazas, and connector paths Bike Lane Marked or separated lanes designated for cyclists Crosswalk Marked pedestrian crossing areas across road surfaces Pole Vertical pole-like objects such as streetlights, utility poles, bollards, and signposts without attached signage Sign Traffic signs, including STOP, YIELD, regulatory, warning, and informational signs Signal Traffic signal heads and their supporting poles; when a pole contains both a signal and a sign, the sign portion is classified as Sign Building Exterior facades and structural surfaces of buildings Fence Fences, guardrails, and other narrow vertical barriers High Vegetation Trees, tall shrubs, and dense vertical vegetation Low Vegetation Grass, small plants, and other low-height vegetation Marking Road markings including lane lines, arrows, symbols, and other painted elements Vehicle Motor vehicles such as cars, vans, trucks, and buses Not Classified Points outside the defined taxonomy or areas that were intentionally left unlabeled The dataset was generated through a multi-stage workflow involving mobile LiDAR acquisition, SLAM-based point cloud reconstruction, georeferencing, point cloud cleaning, segmentation, and manual semantic annotation. Annotation was performed using CloudCompare and required approximately 600 hours of manual labeling effort across 295 segmented files. Each spatial zone contains scene-wise annotated point clouds divided into approximately 50-m roadway segments and provided in both `.txt` and `.ply` formats. These files include global coordinates (x, y, z), LiDAR intensity values, and semantic labels for all classes within a scene, making them suitable for semantic segmentation and scene understanding tasks. In addition, class-wise annotation files are provided to isolate individual infrastructure categories for visualization and analysis. Each zone also includes raw LiDAR packet capture files (`.pcap`), georeferenced vehicle trajectory files (`.las`), GNSS observations (`.csv`), and IMU measurements (`.csv`), enabling users to reconstruct, analyze, and process the complete mobile mapping workflow. To establish benchmark performance, five state-of-the-art semantic segmentation models were evaluated using a 7-fold spatial cross-validation framework: PointNet++, DGCNN, KPConv, KPConvX, and Point Transformer V3. Geographic Coverage Edmonton, Alberta, Canada Data Collection Year 2025 Dataset Size 682 million annotated points Approximately 14 km of annotated roadway corridors Seven spatial zones Eighteen semantic classes Applications 3D semantic segmentation Autonomous driving perception HD mapping Intelligent transportation systems Digital twins Urban infrastructure inventory Geomatics Remote sensing Computer vision Machine learning Citation Please cite both the associated publication and this Zenodo dataset when using YEG3D in your work.



