遇见数据集

Multi-Language Conversational Telephone Speech 2011 -- Arabic Group

收藏
DataCite Commons2021-07-01 更新2025-04-16 收录
官方服务:

资源简介:

<h3>Introduction</h3><br> <p>Multi-Language Conversational Telephone Speech 2011 -- Arabic Group was developed by the Linguistic Data Consortium (LDC) and is comprised of approximately 117 hours of telephone speech in distinct dialects of colloquial Arabic: Iraqi, Levantine and Maghrebi.</p><br> <p>The data were collected primarily to support research and technology evaluation in automatic language identification, and portions of these telephone calls were used in the NIST 2011 Language Recognition Evaluation (<a href="https://www.nist.gov/itl/iad/mig/2011-language-recognition-evaluation">LRE</a>). LRE 2011 focused on language pair discrimination for 24 languages/dialects, some of which could be considered mutually intelligible or closely related.</p><br> <p>LDC has also released the following as part of the Multi-Language Conversational Telephone Speech 2011 series:</p><br> <ul><br> <li>Slavic Group (<a href="../../../LDC2016S11">LDC2016S11</a>)</li><br> <li>Turkish (<a href="../../../LDC2017S09">LDC2017S09</a>)</li><br> <li>South Asian (<a href="../../../LDC2017S14">LDC2017S14</a>)</li><br> <li>Central Asian (<a href="../../../LDC2018S03">LDC2018S03</a>)</li><br> <li>Central European (<a href="../../../LDC2018S08">LDC2018S08</a>)</li><br> <li>Spanish (<a href="../../../LDC2018S12">LDC2018S12</a>)</li><br> <li>English (<a href="../../../LDC2019S06">LDC2019S06</a>)</li><br> </ul><br> <h3>Data</h3><br> <p>Participants were recruited by native speakers who contacted acquaintances in their social network. Those native speakers made one call, up to 15 minutes, to each acquaintance. The data was collected using <a href="https://www.ldc.upenn.edu/about/facilities/human-subjects-collection">LDC's telephone collection infrastructure</a>, comprised of three computer telephony systems. Human auditors labeled calls for callee gender, dialect type and noise. Demographic information about the participants was not collected.</p><br> <p>All audio data are presented in FLAC-compressed MS-WAV (RIFF) file format (*.flac); when uncompressed, each file is 2 channels, recorded at 8000 samples/second with samples stored as 16-bit signed integers, representing a lossless conversion from the original mu-law sample data as captured digitally from the public telephone network. The following table summarizes the total number of calls, total number of hours of recorded audio, and the total size of compressed data:</p><br> <table border="1" cellpadding="2"><br> <tbody><br> <tr><br> <td>group</td><br> <td>lng</td><br> <td>#calls</td><br> <td>#hours</td><br> <td>#MB</td><br> </tr><br> <tr><br> <td>arabic</td><br> <td>iraqi</td><br> <td>210</td><br> <td>37.4</td><br> <td>1908</td><br> </tr><br> <tr><br> <td>arabic</td><br> <td>levantine</td><br> <td>225</td><br> <td>41.1</td><br> <td>2041</td><br> </tr><br> <tr><br> <td>arabic</td><br> <td>maghrebi</td><br> <td>207</td><br> <td>38.6</td><br> <td>2024</td><br> </tr><br> <tr><br> <td>arabic</td><br> <td>totals</td><br> <td>642</td><br> <td>117.1</td><br> <td>5973</td><br> </tr><br> </tbody><br> </table><br> <h3>Samples</h3><br> <p>Please view this <a href="desc/addenda/LDC2019S02.flac">audio sample</a>.</p><br> <h3>Updates</h3><br> <p>None at this time.</p></br> Portions © 2019 Trustees of the University of Pennsylvania

<h3>引言</h3><br><p>2011年多语言会话电话语音数据集——阿拉伯语组由语言数据联盟(Linguistic Data Consortium,LDC)开发,包含约117小时的口语阿拉伯语不同方言电话语音,涵盖伊拉克阿拉伯语、黎凡特阿拉伯语与马格里布阿拉伯语三种方言。</p><br><p>该数据集的采集主要用于支撑自动语言识别领域的研究与技术评测,其中部分通话数据曾应用于美国国家标准与技术研究院(National Institute of Standards and Technology,NIST)2011年语言识别评测(Language Recognition Evaluation,LRE)。2011年语言识别评测聚焦24种语言/方言的语言对判别任务,其中部分语言/方言可被视为相互可理解或亲缘关系密切。</p><br><p>语言数据联盟(LDC)还发布了2011年多语言会话电话语音数据集系列的以下子数据集:</p><br><ul><br><li><br>斯拉夫语组(<a href="../../../LDC2016S11">LDC2016S11</a>)</li><br><li><br>土耳其语组(<a href="../../../LDC2017S09">LDC2017S09</a>)</li><br><li><br>南亚语组(<a href="../../../LDC2017S14">LDC2017S14</a>)</li><br><li><br>中亚语组(<a href="../../../LDC2018S03">LDC2018S03</a>)</li><br><li><br>中欧语组(<a href="../../../LDC2018S08">LDC2018S08</a>)</li><br><li><br>西班牙语组(<a href="../../../LDC2018S12">LDC2018S12</a>)</li><br><li><br>英语组(<a href="../../../LDC2019S06">LDC2019S06</a>)</li><br></ul><br><h3>数据</h3><br><p>数据采集通过母语使用者招募其社交网络中的熟人完成,每位母语使用者会向每位熟人拨打一次通话,时长不超过15分钟。本次数据采集采用<a href="https://www.ldc.upenn.edu/about/facilities/human-subjects-collection">LDC电话采集基础设施</a>,该系统由三套计算机电话系统组成。人工审核员会对通话进行标注,标注内容包括被叫方性别、方言类型与背景噪音。本次采集未收集参与者的人口统计学信息。</p><br><p>所有音频数据均采用FLAC压缩的MS-WAV(RIFF)文件格式(文件扩展名为*.flac);解压后,每个音频文件为双声道,采样率为8000样本/秒,样本以16位有符号整数存储,该格式是对从公共电话网络数字化采集的原始μ律样本数据进行无损转换后的结果。下表汇总了总通话数、总录音时长以及压缩后数据的总大小:</p><br><table border="1" cellpadding="2"><br><tbody><br><tr><br><td>组别</td><br><td>语言/方言</td><br><td>通话数</td><br><td>时长(小时)</td><br><td>压缩数据大小(MB)</td><br></tr><br><tr><br><td>阿拉伯语</td><br><td>伊拉克阿拉伯语</td><br><td>210</td><br><td>37.4</td><br><td>1908</td><br></tr><br><tr><br><td>阿拉伯语</td><br><td>黎凡特阿拉伯语</td><br><td>225</td><br><td>41.1</td><br><td>2041</td><br></tr><br><tr><br><td>阿拉伯语</td><br><td>马格里布阿拉伯语</td><br><td>207</td><br><td>38.6</td><br><td>2024</td><br></tr><br><tr><br><td>阿拉伯语</td><br><td>总计</td><br><td>642</td><br><td>117.1</td><br><td>5973</td><br></tr><br></tbody><br></table><br><h3>样本示例</h3><br><p>请查看此<a href="desc/addenda/LDC2019S02.flac">音频样本</a>。</p><br><h3>更新记录</h3><br><p>暂无更新记录。</p><br><p>部分内容 © 2019 宾夕法尼亚大学托管委员会</p>

创建时间:
2020-11-30
二维码
社区交流群
二维码
科研交流群
商业服务