AISHELL8-RealScene
收藏资源简介:
AISHELL8-RealScene是一个在真实世界环境中录制的公开多场景、多视角视听对话语料库,专注于包含环境噪声、视角变化、说话人干扰、背景活动、室外场景和对话互动等现实条件。该语料库由AISHELL使用同步多设备录音设置采集。数据内容涵盖普通话对话,采集自五个真实地点(建筑、大厅、酒店、公园、街道),包含室内和室外场景。数据由171位前景说话人参与录制,总时长约102.19小时。数据集提供了多模态数据:音频方面,包含用于转录和对齐的说话人佩戴近场麦克风采集的单通道音频,以及由圆形麦克风阵列(原始16个,公开发布版均匀选取为8个通道)采集的16kHz、16位远场音频;视频方面,包含由三个水平间隔约30度放置的同步高清RGB摄像头(左、中、右视图)采集的、与音频同步的256x256分辨率、25帧/秒的面部视频。数据集已划分为训练集(79.84小时,133个会话/说话人,56组)、开发集(10.70小时,18个会话/说话人,7组)和评估集(11.65小时,20个会话/说话人,7组),且不同划分间的说话人不重叠,所有划分均包含全部五个地点的录音。该数据集适用于音频-视觉语音识别(AVSR)、多模态语音处理、远场语音处理、鲁棒性语音识别、说话人识别以及对话系统等研究领域,尤其适合用于在复杂真实环境下开发和评估相关模型的性能。数据集采用CC BY-NC-SA 4.0许可证发布,仅限于学术和非商业研究使用。
AISHELL8-RealScene is a publicly available multi-scene, multi-view audio-visual dialogue corpus recorded in real-world environments, focusing on realistic conditions such as environmental noise, viewpoint changes, speaker interference, background activities, outdoor scenes, and conversational interactions. The corpus was collected by AISHELL using a synchronized multi-device recording setup. The data content includes Mandarin conversations collected from five real locations (building, hall, hotel, park, street), covering both indoor and outdoor scenes. The data involves 171 foreground speakers, with a total duration of approximately 102.19 hours. The dataset provides multimodal data: in terms of audio, it includes single-channel audio from near-field microphones worn by speakers for transcription and alignment, as well as far-field audio at 16kHz and 16-bit from a circular microphone array (originally 16 channels, uniformly selected as 8 channels in the publicly released version); in terms of video, it includes synchronized facial videos at 256x256 resolution and 25 frames per second, captured by three horizontally spaced (approximately 30 degrees apart) synchronized high-definition RGB cameras (left, center, right views). The dataset is divided into a training set (79.84 hours, 133 sessions/speakers, 56 groups), a development set (10.70 hours, 18 sessions/speakers, 7 groups), and an evaluation set (11.65 hours, 20 sessions/speakers, 7 groups), with no overlap of speakers between different splits, and all splits include recordings from all five locations. This dataset is suitable for research areas such as audio-visual speech recognition (AVSR), multimodal speech processing, far-field speech processing, robust speech recognition, speaker recognition, and dialogue systems, particularly for developing and evaluating related models in complex real-world environments. The dataset is released under the CC BY-NC-SA 4.0 license and is limited to academic and non-commercial research use.
数据集概述:AISHELL8-RealScene
AISHELL8-RealScene 是一个公开的多场景、多视角音视频会话语料库,专注于真实世界环境下的语音与视觉数据。
核心属性
| 属性 | 数值 |
|---|---|
| 语言 | 中文(普通话) |
| 许可证 | CC BY-NC-SA 4.0 |
| 总时长 | 102.19 小时 |
| 说话人 | 171 位前景说话人 |
| 场景 | 5 个真实世界地点 |
| 录制风格 | 会话式语音 |
| 音频 | 近场 + 远场 |
| 视频 | 多视角 RGB 面部视频 |
| 采样率 | 16 kHz |
| 远场音频 | 8 通道 |
| 视频分辨率 | 256×256 @ 25 fps |
录制环境
语料库包含来自五个真实世界地点的录音:
- L1: 建筑物(室外)
- L2: 大厅(室内)
- L3: 酒店(室内)
- L4: 公园(室外)
- L5: 街道(室外)
录音涵盖室内和室外场景。
音频规格
- 近场音频: 每位前景说话人佩戴近讲麦克风,用于转录和对齐。
- 远场音频: 使用 16 麦克风环形阵列录制。公开版本提供8通道远场音频(均匀选取阵列中的麦克风)。
- 格式:
.wav文件,采样率 16 kHz。近场为单声道,远场为 8 声道。
视频规格
- 使用三台同步的 HD RGB 摄像机以约 30° 的水平角度间隔(左 D0、中心 D1、右 D2)捕获。
- 参数: 分辨率 256×256,帧率 25 fps。
- 释放的面部视频与音频同步。
数据集统计与划分
数据集被划分为训练集、开发集和评估集,且三者之间的说话人不重叠。所有划分均包含来自全部五个地点的录音。
| 划分 | 时长 (小时) | 对话轮次 | 组数 | 说话人数 |
|---|---|---|---|---|
| 训练集 | 79.84 | 133 | 56 | 133 |
| 开发集 | 10.70 | 18 | 7 | 18 |
| 评估集 | 11.65 | 20 | 7 | 20 |
| 总计 | 102.19 | 171 | 70 | 171 |
相关资源
- 论文: arXiv 2606.05763 (https://arxiv.org/abs/2606.05763)
- 项目页面: 即将上线




