Extended-IS3
收藏资源简介:
Extended-IS3数据集是由韩国科学技术院的研究团队创建的,该数据集在IS3数据集的基础上增加了spoken utterances,用于评估模型在同时定位和识别视觉场景中的混合音频类型(包括重叠的口语和非口语声音)的性能。数据集包含了音频和对应的视觉信息,用于训练模型在视觉场景中同时定位和识别不同的音频类型。该数据集的应用领域主要在于提高音频-视觉定位模型的性能,解决实际场景中音频源混合的问题。
The Extended-IS3 dataset was created by a research team from the Korea Advanced Institute of Science and Technology (KAIST). This dataset expands the original IS3 dataset by adding spoken utterances, and is designed to evaluate models' performance in simultaneously localizing and identifying mixed audio types (including overlapping speech and non-speech sounds) within visual scenes. The dataset includes audio and its corresponding visual information, which is used to train models to simultaneously localize and recognize various audio types in visual scenarios. Its primary application areas are to improve the performance of audio-visual localization models and address the issue of mixed audio sources in real-world scenarios.

- 1Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes韩国科学技术院 · 2025年



