遇见数据集

Second DIHARD Challenge Development - SEEDLingS

收藏
DataCite Commons2021-11-15 更新2024-07-13 收录
官方服务:

资源简介:

<h3>Introduction</h3><br> <p>Second DIHARD Challenge Development - SEEDLinGS was developed by Duke University and LDC and contains approximately two hours of English child language recordings along with corresponding annotations used in support of the <a href="https://dihardchallenge.github.io/dihard2">Second DIHARD Challenge</a>.</p><br> <p>This release, when combined with Second DIHARD Challenge Development - Eleven Sources (LDC2021S10), contains the development set audio data and annotation, except for CHiME-5 audio files, which must be obtained from the <a href="https://licensing.sheffield.ac.uk/product/chime5">University of Sheffield</a>.</p><br> <p>The DIHARD Challenges are a set of shared tasks on diarization focusing on "hard" diarization; that is, speech diarization for challenging corpora where there was an expectation that existing state-of-the-art systems would fare poorly. As with the <a href="https://dihardchallenge.github.io/dihard1/">first challenge</a>, the second development and evaluation sets were drawn from a diverse sampling of sources including monologues, map task dialogues, broadcast interviews, sociolinguistic interviews, meeting speech, speech in restaurants, clinical recordings, extended child language acquisition recordings, and YouTube videos.</p><br> <h3>Data</h3><br> <p>Source data is from the <a href="https://homebank.talkbank.org/access/Password/Bergelson.html">SEEDLingS</a> (The Study of Environmental Effects on Developing Linguistic Skills) corpus, designed to investigate how infants' early linguistic and environmental input plays a role in their learning. Recordings were generated in the home environment of infants in the Rochester, New York area. A subset of that data was annotated by LDC for use in the First and Second DIHARD Challenges.</p><br> <p>The data in this release consists of files provided in the Second DIHARD Challenge as well as subsequently updated annotated files not provided to second challenge participants.</p><br> <p>All audio is provided in the form of 16 kHz, 16-bit, mono-channel FLAC files. The diarization for each recording is stored as a NIST Rich Transcription Time Marked (RTTM) file. RTTM files are space-separated text files containing one turn per line. Segmentation files are stored as HTK label files. Each of these files contains one speech segment per line. Scoring regions for each recording are specific by un-partitioned evaluation map (UEM) files. All annotation file types are encoded as UTF-8. More information about the file formats and data sources and domains are in the included documentation.</p><br> <h3>Updates</h3><br> <p>None at this time.</p></br> Portions © 2019, 2021 Duke University, © 2019, 2021 Trustees of the University of Pennsylvania

<h3>简介</h3><br><p>第二届DIHARD挑战赛开发数据集——SEEDLinGS 由杜克大学与LDC开发,包含约两小时的英语儿童语言录音及对应标注文件,用于支撑<a href="https://dihardchallenge.github.io/dihard2">第二届DIHARD挑战赛</a>。</p><br><p>本发布包与《第二届DIHARD挑战赛开发数据集——十一来源(LDC2021S10)》结合后,可提供开发集的音频数据与标注文件,但CHiME-5音频文件除外,此类文件需从<a href="https://licensing.sheffield.ac.uk/product/chime5">谢菲尔德大学</a>获取。</p><br><p>DIHARD挑战赛系列是一组聚焦于困难场景下说话人 diarization的共享任务,即针对现有最优系统预期表现不佳的挑战性语料库开展语音说话人 diarization研究。与<a href="https://dihardchallenge.github.io/dihard1/">首届挑战赛</a>一致,第二届开发集与评测集的数据源涵盖广泛,包括独白、地图任务对话、地图任务对话、广播访谈、社会语言学访谈、会议语音、餐厅场景语音、临床录音、长时儿童语言习得录音以及YouTube视频。</p><br><h3>数据</h3><br><p>本数据集的源数据来自<a href="https://homebank.talkbank.org/access/Password/Bergelson.html">SEEDLingS</a>语料库(全称:The Study of Environmental Effects on Developing Linguistic Skills,即儿童语言技能发展的环境影响研究),该语料库旨在探究婴儿早期语言输入与环境输入对其语言学习的作用。录音采集自纽约州罗切斯特地区的婴儿家庭环境。LDC对该数据集的一个子集进行了标注,用于首届与第二届DIHARD挑战赛。</p><br><p>本次发布的数据包含第二届DIHARD挑战赛中使用的文件,以及后续更新的、未向第二届挑战赛参与者提供的标注文件。</p><br><p>所有音频均采用16 kHz、16位单声道FLAC格式存储。每份录音的说话人 diarization结果存储为NIST Rich Transcription Time Marked(RTTM)文件。RTTM文件为空格分隔的文本文件,每行对应一个说话人轮次。分段文件存储为HTK标签文件,每份文件每行对应一个语音片段。每份录音的评分区域由未分区评测映射(UEM)文件指定。所有标注文件均采用UTF-8编码。关于文件格式、数据源与数据领域的更多信息,请参见随附的文档。</p><br><h3>更新说明</h3><br><p>暂无更新。</p><br>Portions © 2019, 2021 Duke University, © 2019, 2021 Trustees of the University of Pennsylvania

创建时间:
2021-11-09
搜集汇总
背景与挑战
背景概述
该数据集是第二次DIHARD挑战的开发集部分,由杜克大学和LDC开发,包含约两小时的英语儿童家庭环境录音及对应标注,用于支持'困难'语音分离任务研究。数据来源于SEEDLingS语料库,旨在分析婴儿语言学习与环境输入的关系,音频以16 kHz FLAC格式提供,标注包括RTTM和HTK文件。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务