2003 NIST Language Recognition Evaluation

Name: 2003 NIST Language Recognition Evaluation
Creator: Linguistic Data Consortium
Published: 2021-07-01 16:18:26
License: 暂无描述

DataCite Commons2021-07-01 更新2025-04-16 收录

下载链接：

https://catalog.ldc.upenn.edu/LDC2006S31

下载链接

链接失效反馈

官方服务：

资源简介：

<h3>Introduction</h3><br> <p>The goal of the <a href="https://www.nist.gov/itl/iad/mig/language-recognition">NIST Language Recognition Evaluation (LRE)</a> is to establish the baseline of current performance capability for language recognition of conversational telephone speech and to lay the groundwork for further research efforts in the field. The series had its first evaluation in 1996. 2003 NIST Language Recognition Evaluation (LRE-03) was part of this ongoing series of evaluations of language recognition technology.</p><br> <p>The task evaluated was the detection of a given target language. Languages included in this releasea are Arabic, English, Farsi, French, German, Hindi, Japanese, Korean, Mandarin, Russian, Spanish, Tamil and Vietnamese.<br />Given a test segment of speech, a target language was assigned as a test hypothesis, and the task was to determine whether this test hypothesis was true or false. This release contains both the 1996 and 2003 NIST Language Recognition Evaluations.</p><br> <p>LDC released other LREs as:</p><br> <ul><br> <li>2005 NIST Language Recognition Evaluation (<a href="../../../LDC2008S05">LDC2008S05</a>)</li><br> <li>2007 NIST Language Recognition Evaluation Test Set (<a href="../../../LDC2009S04">LDC2009S04</a>)</li><br> <li>2007 NIST Language Recognition Evaluation Supplemental Training Set (<a href="../../../LDC2009S05">LDC2009S05</a>)</li><br> <li>2009 NIST Language Recognition Evaluation Test Set (<a href="../../../LDC2014S06">LDC2014S06</a>)</li><br> <li>2011 NIST Language Recognition Evaluation Test Set (<a href="../../../LDC2018S06">LDC2018S06</a>)</li><br> </ul><br> <h3>Data</h3><br> <p>Each speech file is one side of a "four wire" telephone conversation represented as 8-bit, 8kHz mulaw data. There are 11,830 speech files in sphere(.sph) format for a total of around forty six hours of speech. The speech data was compiled from the LDC's CALLFRIEND, CALLHOME, and Switchboard-2 corpora. Each file contains one test segment. The test segments are divided into three-second, ten-second, and thirty-second tests, each in its own directory.</p><br> <h3>Samples</h3><br> <p>For an example of the data in this corpus, please listen to this audio <a href="desc/addenda/LDC2006S31.wav" rel="nofollow">sample</a>.</p><br> <h3>Updates</h3><br> <p>A typo was fixed in the index.html file. There are 11,830 sphere files, not 11,839. The updated index file is available in the <a href="docs/LDC2006S31/" rel="nofollow">online docs</a> folder.</p></br> Portions © 1996-2002, 2006 Trustees of the University of Pennsylvania

提供机构：

Linguistic Data Consortium

创建时间：

2020-11-30

5,000+

优质数据集

54 个

任务类型

进入经典数据集