遇见数据集

TIMIT Acoustic-Phonetic Continuous Speech (MS-WAV version)

收藏
DataCite Commons2021-07-01 更新2025-04-16 收录
官方服务:

资源简介:

<h3>Introduction</h3> <p>This version of the TIMIT Acoustic-Phonetic Continuous Speech Corpus (LDC93S1) has all the waveform files formatted with ms-wav / RIFF headers, to make the corpus more accessible to a wider audience.</p> <p>The TIMIT corpus of read speech is designed to provide speech data for acoustic-phonetic studies and for the development and evaluation of automatic speech recognition systems. TIMIT contains broadband recordings of 630 speakers of eight major dialects of American English, each reading ten phonetically rich sentences. The TIMIT corpus includes time-aligned orthographic, phonetic and word transcriptions as well as a 16-bit, 16kHz speech waveform file for each utterance. Corpus design was a joint effort among the Massachusetts Institute of Technology (MIT), SRI International (SRI) and Texas Instruments, Inc. (TI). The speech was recorded at TI, transcribed at MIT and verified and prepared for CD-ROM production by the National Institute of Standards and Technology (NIST). </p> <p>The TIMIT corpus transcriptions have been hand verified. Test and training subsets, balanced for phonetic and dialectal coverage, are specified. Tabular computer-searchable information is included as well as written documentation.</p><h3>Samples</h3> <ul> <li> <a href="./desc/addenda/LDC93S1.phn" rel="nofollow">phonemes</a> </li> <li><a href="./desc/addenda/LDC93S1.txt" rel="nofollow">transcripts</a></li> <li><a href="./desc/addenda/LDC93S1.wav" rel="nofollow">audio</a></li> <li><a href="./desc/addenda/LDC93S1.wrd" rel="nofollow">word list</a></li> </ul></br> Portions © 1993 Trustees of the University of Pennsylvania

<h3>引言</h3> <p>本版本的TIMIT声学-语音连续语音语料库(LDC93S1)将所有波形文件采用ms-wav/RIFF格式头部封装,以提升该语料库对更广泛用户群体的易用性。</p> <p>该朗读语音TIMIT语料库专为声学-语音研究以及自动语音识别(automatic speech recognition)系统的开发与评估提供语音数据而设计。TIMIT语料库包含630名说话者的宽带语音录音,涵盖美式英语八大主要方言分区,每名说话者朗读10个语音学特征丰富的句子。TIMIT语料库为每条语音片段均配有时间对齐的正字法转写、语音学转写与词级转写,同时附带16位、16kHz采样率的语音波形文件。本语料库的设计由麻省理工学院(Massachusetts Institute of Technology,简称MIT)、SRI国际公司(SRI International,简称SRI)与德州仪器公司(Texas Instruments, Inc.,简称TI)联合完成。语音数据由TI录制、MIT完成转写,后由美国国家标准与技术研究院(National Institute of Standards and Technology,简称NIST)完成校验并筹备CD-ROM出版。</p> <p>TIMIT语料库的转写文本均经过人工校验。语料库已划定针对语音学与方言覆盖度进行均衡优化的测试子集与训练子集。除配套书面文档外,同时提供可通过计算机检索的表格化数据。</p><h3>示例数据</h3> <ul> <li> <a href="./desc/addenda/LDC93S1.phn" rel="nofollow">音素(phonemes)</a> </li> <li><a href="./desc/addenda/LDC93S1.txt" rel="nofollow">转写文本(transcripts)</a></li> <li><a href="./desc/addenda/LDC93S1.wav" rel="nofollow">音频(audio)</a></li> <li><a href="./desc/addenda/LDC93S1.wrd" rel="nofollow">词表(word list)</a></li> </ul></br>部分内容 © 1993 宾夕法尼亚大学(University of Pennsylvania)校董会

创建时间:
2020-11-30
搜集汇总
数据集介绍
TIMIT Acoustic-Phonetic Continuous Speech (MS-WAV version) 数据集图片
背景与挑战
背景概述
TIMIT Acoustic-Phonetic Continuous Speech (MS-WAV version) 是一个用于声学语音学研究和自动语音识别系统开发与评估的英语语音数据集。它包含630名美国英语八种方言说话者的录音,每人朗读十个语音丰富的句子,并提供时间对齐的转录和16位、16kHz的波形文件,格式为MS-WAV以提高可访问性。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务