音频-视觉拥挤场景分类数据集
收藏资源简介:
音频-视觉拥挤场景分类数据集是由奥地利技术研究所收集的一个包含341个视频的数据集,总时长近29.06小时,涵盖五种真实生活中的拥挤场景:‘暴乱’、‘嘈杂街道’、‘烟花事件’、‘音乐事件’和‘体育氛围’。数据集通过从YouTube收集的野外场景视频构建,每个视频被分割成10秒的片段,并标注相应的场景类别。该数据集旨在通过深度学习框架分析音频和视觉输入,以提高对特定拥挤场景的分类准确性,特别是在预测和检测潜在的暴乱事件方面具有重要应用。
The Audio-Visual Crowded Scene Classification Dataset is a collection curated by the Austrian Institute of Technology, consisting of 341 videos with a total runtime of nearly 29.06 hours. It covers five real-life crowded scenarios: "riots", "busy streets", "firework events", "music events" and "sports atmospheres". Constructed from in-the-wild scene videos sourced from YouTube, each original video is segmented into 10-second clips and annotated with the corresponding scene category. This dataset is designed to analyze audio and visual inputs via deep learning frameworks to improve the classification accuracy of specific crowded scenes, and holds significant applications particularly in predicting and detecting potential riot incidents.




