遇见数据集

IARPA Babel Lithuanian Language Pack IARPA-babel304b-v1.0b

收藏
DataCite Commons2024-02-12 更新2024-07-13 收录
官方服务:

资源简介:

<h3>Introduction</h3><br> <p>IARPA Babel Lithuanian Language Pack IARPA-babel304b-v1.0b was developed by <a href="http://www.appen.com/">Appen</a> for the IARPA (Intelligence Advanced Research Projects Activity) <a href="http://www.iarpa.gov/index.php/research-programs/babel">Babel</a> program. It contains approximately 210 hours of Lithuanian conversational and scripted telephone speech collected in 2013 and 2014 along with corresponding transcripts.</p><br> <p>The Babel program focuses on underserved languages and seeks to develop speech recognition technology that can be rapidly applied to any human language to support keyword search performance over large amounts of recorded speech.</p><br> <h3>Data</h3><br> <p>The Lithuanian speech in this release represents that spoken in the Auk&scaron;taitian and Samogitian dialect regions of Lithuania. The gender distribution among speakers is approximately equal; speakers' ages range from 16 years to 71 years. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments including the street, a home or office, a public place, and inside a vehicle.</p><br> <p>Audio data is presented as 8kHz 8-bit a-law encoded audio in sphere format and 48kHz 24-bit PCM encoded audio in wav format. Transcripts are encoded in UTF-8. Further information about transcription methodology is contained in the documentation accompanying this release.</p><br> <p>Evaluation data is available from <a href="http://www.nist.gov/">NIST</a> in support of <a href="http://www.nist.gov/itl/iad/mig/openkws.cfm">OpenKWS</a>.</p><br> <h3>Samples</h3><br> <p>Please view this <a href="desc/addenda/LDC2019S03.sph">audio sample</a> and <a href="desc/addenda/LDC2019S03.txt">transcript sample</a>.</p><br> <h3>Updates</h3><br> <p>None at this time.</p></br> Portions © 2015 U.S. Government</br> </br>The U.S. Government acquired this data from Appen Pty Ltd, which assigned the copyright to the data to the U.S. Government.

<h3>引言</h3><br><p>IARPA Babel立陶宛语数据包IARPA-babel304b-v1.0b由<a href="http://www.appen.com/">Appen</a>为美国情报高级研究计划局(Intelligence Advanced Research Projects Activity,IARPA)旗下的<a href="http://www.iarpa.gov/index.php/research-programs/babel">Babel</a>项目研发。该数据包包含2013年至2014年间采集的约210小时立陶宛语会话式与脚本式电话语音数据,以及对应的转写文本。</p><br><p>Babel项目聚焦于低资源语言,旨在研发可快速适配任意人类语言的语音识别技术,以支撑大规模录制语音中的关键词检索任务。</p><br><h3>数据概况</h3><br><p>本版本中的立陶宛语语音数据采集自立陶宛境内的奥克什泰蒂亚(Aukštaitian)与萨莫吉希亚(Samogitian)两大方言区。发言者的性别分布大致均衡,年龄跨度为16岁至71岁。通话使用了各类终端设备(如移动电话、固定电话),采集场景涵盖街道、家庭/办公室、公共场所及车内等多种环境。</p><br><p>语音数据提供两种格式:8kHz 8位A律编码的SPHERE格式音频,以及48kHz 24位脉冲编码调制(PCM)编码的WAV格式音频。转写文本采用UTF-8编码。本版本附带的文档中包含转写方法的详细说明。</p><br><p>评估数据可从美国国家标准与技术研究院(NIST)获取,用于支持开放关键词检索(OpenKWS)相关任务。</p><br><h3>样例</h3><br><p>请查看以下样例:<a href="desc/addenda/LDC2019S03.sph">语音样例</a>与<a href="desc/addenda/LDC2019S03.txt">转写样例</a>。</p><br><h3>更新情况</h3><br><p>暂无更新记录。</p><br><br>部分内容 © 2015 美国政府<br><br>本数据由美国政府从Appen Pty Ltd处获得,该公司已将本数据的版权转让予美国政府。

创建时间:
2020-11-30
搜集汇总
数据集介绍
IARPA Babel Lithuanian Language Pack IARPA-babel304b-v1.0b 数据集图片
背景与挑战
背景概述
该数据集是IARPA Babel项目的一部分,包含约210小时的立陶宛语电话语音数据(收集于2013-2014年),涵盖对话和脚本语音,并附带转录文本。数据来自立陶宛的Aukštaitian和Samogitian方言区域,说话者年龄范围为16-71岁,性别分布均衡,音频以多种格式提供,主要用于语音识别技术开发,尤其关注资源不足语言的支持。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务