Multimodal Dataset for Evaluating Driving Distractors
收藏资源简介:
Driver safety has never been as important as it is now, due to the changes our society is facing. New technologies such as AI have been successfully employed in various technological fields. It is natural to want to replicate that success in the field of driving. The dataset is multimodal, consisting of synchronized audio, video, and annotation streams. Audio was recorded using four microphones, while video data was obtained from multiple cameras, including a dashboard camera and eye trackers. Two distinct scenarios are available, with durations of 41:12 (S1) and 39:55 (S2), respectively. This dataset was developed to quantify the influence of auditory distractions on driving performance. For this reason, we focused on auditory elements — both on the different classes identified as distractions and on the quantity of audio events present in the dataset. Objectives of the Dataset This dataset provides samples that allow users to: Train models based on noises and sound events. Train models to detect distracting elements inside the vehicle. Train models to identify distracting elements outside the vehicle. Data Collection Methodology The dataset was formed from videos and audio recordings obtained in two scenarios. In these, participants are observed during normal driving, which is interrupted by different sound effects. Equipment Used Neon - Pupil Labs. ZOOM h6 Recorder. Behringer ECM8000. Samsung Galaxy Watch 5. Data Format and Structure The dataset is organized into two distinct scenarios, identified using the abbreviations S1 and S2, respectively. The information is divided into raw data and labels. Audio_Elements Folder Divided into data and labels. Data Contains 8 audio elements in .wav format: 3 from Dash cameras.(Sx_audio_D_25fps,Sx_audio_F_25fps,Sx_audio_M_25fps) 1 recorded by Pupil Labs.(SX_audio_Eyetracker) 4 from the Zoom recorder microphones.(Sx_audio_Zoom_Left,Sx_audio_Zoom_Right,Sx_audio_Zoom_Tr1,Sx_audio_Zoom_Tr3) Labels A .txt file containing: Start and end offset of each sound event. Class name. This table presents information on the number of occurrences of each class across different scenarios (Instances_Sx), as well as the total accumulated duration of those instances(Duration_Sx). ID Name Instances_S1 Instances_S2 Duration_S1 Duration_S2 0 Manipulating Objects 5 1 3.2 5.2 1 Speech 268 155 1514.71 1830.18 2 Laughter 39 43 51.27 82.39 3 Music 8 12 80.11 212.8 4 Warning Beeps 5 4 25.72 18.7 5 Ring Tone 12 4 69.39 5.37 6 Notifications 14 15 14.6 12.41 7 Cough 3 3 2.1 3.09 8 Cry 3 2 63.24 20.64 9 Hit 1 2 0.24 1.1 10 Yell 1 0 1.5 0 11 Dog 6 2 2.99 2.3 12 Car Horn 4 2 4.5 2.14 13 Scream 0 1 0 1.39 14 Microphone Noise 0 1 0 1.9 15 Stone 0 10 0 9.34 16 Engine 0 2 0 19.22 Additional metadata for the audio labels is also provided. Videos_Raw Folder Videos are separated based on their file size and stored in .mp4 format. There are 4 video sources. 3 from the Dash camera.(Sx_Video_D_25fps,Sx_Video_D_25fps,Sx_Video_D_25fps) 1 from the Pupil Labs.(Sx_Video_Eyetracker_25fps) Additional information about each video is included here. Video_Labels Folder Includes the following subfolders: Coordinates_Eye_Tracker Contains gaze data recorded with Pupil Eye Tracker for each scenario, stored in .csv files. First CSV includes: Unix timestamp. Video time in seconds. Video frame number X and Y gaze coordinates. Fixation ID (if any). Labeled coordinates CSV includes additionally: Object being viewed (refers to interior_classes; left empty if none) Sound event. Fixation indicator. Fixation ID. Fixation coordinates (X, Y). Pixel variation between the current gaze and the fixation point. Object_Segmentation_labels Video frames were extracted using FFmpeg with the following command: ffmpeg -i video_name.mp4 -vf fps=1 -start_number 1 frames/test_%06d.png One frame per second was extracted (instead of 25 fps), which is reflected in the frame naming. The .txt file contains: Class ID. Object position coordinates within the frame. Additional information about the segmentation labels is also included. The first table refers to exterior objects and distractors, while the second one corresponds to interior objects and distractions. These tables summarize the instances in which each class appears. Interior_Classes ID Name Instances_S1 Instances_S2 0 Car 9719 8426 1 Traffic light 1148 1547 2 Traffic signal 4297 3619 3 Person 678 573 4 Road 4991 3605 5 Pedestrian crossing 1205 958 Exterior_Classes ID Name Instances_S1 Instances_S2 0 Interior rear-view mirror 2127 2145 1 Outside left 2388 1946 2 Outside Central 2479 2400 3 Screen 2212 2209 4 Speedometer 2298 2237 5 Left rear-view visor 2035 1536 6 Right rear-view mirror 682 713 7 Outside Right 447 630 8 Co-driver 2189 1901 9 Radio 2157 1528 10 Steering wheel 2378 2016 Extra Folder Contains supplementary information not directly related to the audio or video elements. It includes JSON files for both scenarios and heart rate data. Each file stores: Average, maximum, and minimum heart rate. Time period in Unix timestamp format Warning The heart rate data may be inaccurate or unreliable. We recommend exercising caution when using or interpreting this information. Authors C. Castorena, A. Roche, J. Lopez-Ballester, J. A. De Rus, J. J. Perez-Solano, S. Roger, C. Botella-Mascarell, J. J. Lopez, J. M. Mossi, F. J. Ferri and M. Cobos Acknolegments Dataset collected within the AEOLIAN-CONNECT (TED2021-131003B-C21) project, funded by MICIU/AEI /10.13039/501100011033 and by European Union NextGenerationEU/ PRTR.



