Voxel51/kitscenes-multimodal
收藏资源简介:
KITScenes Multimodal(KIT-MRT)是一个高保真的欧洲城市自动驾驶多模态数据集,已整合为分组的FiftyOne数据集。它包含4个验证场景,每个场景由多个帧组成,总共有680组(帧)。每组对应一个时间戳帧,包含10个样本:来自九个全局快门摄像头的360度覆盖图像(形成环视和长距离立体视觉),以及一个融合的3D激光雷达/雷达点云(融合了七个激光雷达和三个雷达的数据)。数据集还提供了丰富的标注信息,包括投影的激光雷达深度热图、Lanelet2高精地图折线、自我轨迹路径点以及基于Mapillary-Vistas分类的实例预测(如车辆、行人、交通标志等)。此外,每个样本都包含场景上下文、时间戳、自我姿态(位置和方向)、GNSS数据(如经纬度、海拔)等字段。该数据集适用于多模态浏览与标注、高精地图感知、长距离深度估计、轨迹分析和2D对象分析等任务,但请注意,这是一个预览子集,不包含3D边界框或动态代理的跟踪,且实例预测为模型输出而非人工标注。数据采集自德国法兰克福等城市,使用CC-BY-NC-4.0许可证,仅限非商业用途。
KITScenes Multimodal (KIT-MRT) is a high-fidelity European urban autonomous-driving multimodal dataset, ingested into a grouped FiftyOne dataset. It includes 4 validation scenes, with a total of 680 groups (frames). Each group corresponds to a timestamped frame and contains 10 samples: synchronized captures from nine global-shutter cameras providing 360° coverage (forming surround view and long-range stereo vision), plus a fused 3D lidar/radar point cloud (combining seven lidars and three radars). The dataset is enriched with annotations such as projected lidar-depth heatmaps, Lanelet2 HD-map polylines, ego trajectory waypoints, and instance predictions based on the Mapillary-Vistas taxonomy (e.g., cars, pedestrians, traffic signs). Each sample also includes contextual fields like scene ID, timestamp, ego pose (translation and orientation), GNSS data (e.g., longitude, latitude, altitude), and more. This dataset is suitable for multimodal browsing and curation, HD-map perception, long-range depth estimation, trajectory/motion analysis, and 2D object analysis. Note that it is an early-release preview subset and does not include 3D bounding boxes or tracking for dynamic agents; the instance predictions are model outputs, not human annotations. Data was recorded in cities like Frankfurt, Germany, under the CC-BY-NC-4.0 license for non-commercial use only.




