遇见数据集

Event-Shifted Acoustic Scene (ESAS) Dataset

收藏
Zenodo2026-06-10 更新2026-06-12 收录
官方服务:

资源简介:

The ESAS dataset is a new benchmark for evaluating the robustness of Acoustic Scene Classification (ASC) systems against unknown sound events in real-world environments. Overview Existing ASC datasets mostly use clean background recordings. However, real-world scenes are often disrupted by unexpected sound events (alarms, voices, traffic, etc.). ESAS simulates this variability by injecting foreground events from FSD50K into background scenes from CochlScene, with semantic consistency ensured by Large Language Models (LLM). Key Features 13 scene classes, 96 event classes (27 known + 69 unknown) 76,115 clips (~ 211 hours of audio) Mix types: Background Only, Known Events, Unknown Events Rich metadata: SNR (-15 ~ +15 dB), timestamps, event counts, etc. Unknown events appear only in the test set Construction Background scenes come from CochlScene. Foreground events are filtered and mixed using randomized overlapping, time-stretching, pitch-shifting, and SNR control. The ESAS dataset contains only the mixed samples (Known + Unknown Events). The background-only samples are not included and must be downloaded separately from CochlScene. Benchmark ASC models suffer significant performance drops when encountering unknown events, highlighting the need for event-robust methods. Resources Paper: Towards Event-Robust Acoustic Scene Classification (Interspeech 2026) Code & Generation Pipeline: GitHub Repository

提供机构:
Zenodo
创建时间:
2026-06-06
二维码
社区交流群
二维码
科研交流群
商业服务