CHiME-Home
收藏资源简介:
CHiME-Home 是用于家庭环境中声源识别的数据集。它使用大约 6.8 小时的家庭环境录音。这些录音是从 CHiME 项目中获得的——多源环境中的计算听力——其中录音设备位于英国维多利亚时代的半独立式房屋内。录音选自 22 个会议,总计 19.5 小时,每个会议在早上 7:30 和晚上 20:00 之间进行。在所考虑的录音中,设备被放置在靠近通往走廊的门的休息室(起居室)中,走廊通向没有门的厨房。由于休息室的门通常是打开的,因此突出的声音可能来自休息室和厨房的来源。允许标签的选择是由考虑的声学环境中存在的来源推动的:人类说话者(c,m,f);人类活动(p);电视 (v);家用电器 (b)。进一步的标签 o、S、U 分别与任何其他可识别的声音、沉默、无法识别的声音相关。标签 S、U 只能分别单独分配。获取注释器以将至少一个标签分配给块,因此注释器可以从集合 {c,m,f,v,p,b,o} 中分配一个或多个标签,或者可以使用“标记”块来自集合 {S,U} 的单个标签。
CHiME-Home is a dataset for sound source recognition in home environments. It utilizes approximately 6.8 hours of home environment audio recordings. These recordings are sourced from the CHiME project (Computational Hearing in Multi-Source Environments), where the recording devices were deployed in a Victorian semi-detached house in the UK. The recordings are selected from 22 sessions totaling 19.5 hours, with each session conducted between 7:30 AM and 8:00 PM. In the selected recordings, the device was placed in the lounge (living room) near a door leading to a corridor that connects to a kitchen without a door. Since the lounge door is typically open, prominent sounds may originate from both the lounge and the kitchen. The selection of allowed labels is motivated by the sound sources present in the considered acoustic environment: human speakers (c, m, f); human activities (p); television (v); household appliances (b). Additional labels o, S, and U correspond to any other identifiable sound, silence, and unidentifiable sound, respectively. Labels S and U can only be assigned individually. Annotators are required to assign at least one label to each block: annotators may assign one or more labels from the set {c, m, f, v, p, b, o}, or use a single label from the set {S, U} to "mark" the block.

- CHiME-Home数据集首次发表,旨在为智能家居环境中的语音识别研究提供一个标准化的数据集。
- CHiME-Home数据集首次应用于语音识别挑战赛,推动了智能家居领域语音技术的研究与应用。
- CHiME-Home数据集的扩展版本发布,增加了更多的语音样本和环境噪声,提升了数据集的多样性和实用性。
- CHiME-Home数据集被广泛应用于学术研究和工业界,成为智能家居语音识别技术的重要基准数据集。
- CHiME-Home数据集的最新版本发布,引入了更多的多语言支持和跨文化语音样本,进一步丰富了数据集的内容。



