Mandarin Affective Speech
收藏资源简介:
<h3>Introduction</h3> <p>Mandarin Affective Speech is a database of emotional speech consisting of audio recordings and corresponding transcripts collected in 2005 at the Advance Computing and System Laboratory, College of Computer Science and Technology, Zhejiang University, Hangzhou, People's Republic of China. This corpus was designed with two goals: first, to serve as a tool for linguistic and prosodic feature investigation of emotional expression in Mandarin Chinese; and second, to provide a source of training and test data essential to support research in speaker recognition with affective speech. The speech database was recorded by eliciting speakers to express different emotional states in response to stimuli. The speakers read scenarios designed to elicite an emotional response such as a colleague's mistake for anger, a pleasant trip for elation, a hurry-up scene for panic and a puppy's death for sadness. The five emotional states recorded are characterized as follows: </p><ul> <li>Neutral - Simple statements without any emotion.</li> <li>Anger - A strong feeling of displeasure or hostility.</li> <li>Elation - Be glad or happy because of praise.</li> <li>Panic - A sudden, overpowering terror, often affecting many people at once.</li> <li>Sadness - Affected or characterized by sorrow or unhappiness</li> </ul><h3>Data</h3> <p>Over 100 speakers participated in the data collection. After screening, recordings from 68 speakers (23 females, 45 males) were used in this corpus. Most of the speakers were in their twenties at the time of collection. Information about the speakers is contained in "SpeakerInfo.doc." </p><p>Subjects were given a text to read that consisted of five phrases, fifteen sentences and two paragraphs designed to generate the emotional speech. The material included all the phonemes in Mandarin. Each subject read the phrases, paragraphs, and sentences portraying the five emotional states: neutral (unemotional), anger, elation, panic and sadness. Altogether this database contains 25,636 utterances. The read material was constructed as follows:</p> <ul> <li> 5 phrases - "yes", "no" and three nouns as "apple", "train", "tennis ball". In Chinese, these words contain many different basic vowels and consonants.</li> <li>20 sentences - These sentences include all the phonemes and most common consonant clusters in Mandarin. The types of sentences are: simple statements, a declarative sentence with an enumeration, general questions (yes/no question), alternative questions, imperative sentences, exclamatory sentences, special questions (whquestions).</li> <li>2 paragraphs - Two readings, one selected from a famous Chinese novel, and the other stating a normal fact.</li> </ul><p>All the data were recorded in a quiet office on an OLYMPUS DM-20 digital voice recorder with a sampling rate of 22050Hz. Afterwards, the recorded voice files were transferred to a personal computer by USB (Universal Serial Bus). The recordings were then converted into monophonic Windows PCM format at 8 kHz sampling frequency and 16 bits resolution. </p><p>Further information about the data and methodology in this corpus is contained in the authors' paper, "MASC: A Speech Corpus in Mandarin for Emotional Analysis and Affective Speaker Recognition," in "MASC.pdf." </p><h3>Samples </h3><p>For an example of the data in this corpus, please listen the following examples: </p><ul> <li> <a href="./desc/addenda/LDC2007S09_neutral.wav" rel="nofollow">Neutral</a> </li> <li><a href="./desc/addenda/LDC2007S09_anger.wav" rel="nofollow">Anger</a></li> <li><a href="./desc/addenda/LDC2007S09_elation.wav" rel="nofollow">Elation</a></li> </ul> </br> Portions © 2005-2007, Zhejiang University, Advance Computing and System Laboratory, © 2007 Trustees of the University of Pennsylvania
<h3>引言</h3> <p>普通话情感语音语料库(Mandarin Affective Speech)是2005年于中国杭州浙江大学计算机科学与技术学院先进计算与系统实验室采集的情感语音数据库,包含语音录音与对应转录文本。本语料库的构建目标有二:其一,作为研究普通话情感表达的语言学与韵律特征的工具;其二,为带情感语音的说话人识别研究提供不可或缺的训练与测试数据源。该语音语料库通过诱导说话人针对刺激内容表达不同情绪状态的方式录制完成。受试者需朗读预设情境以诱发情感反应,例如以同事犯错的情境诱发愤怒、以愉快旅行的情境诱发喜悦、以紧急赶工的情境诱发恐慌、以小狗离世的情境诱发悲伤。本次录制涵盖五种情绪状态,特征描述如下:</p> <ul> <li>中性(Neutral):无任何情绪的简单陈述。</li> <li>愤怒(Anger):强烈的不悦或敌意情绪。</li> <li>喜悦(Elation):因受赞扬而感到开心愉悦。</li> <li>恐慌(Panic):突发且难以抑制的恐惧,通常会同时影响多人。</li> <li>悲伤(Sadness):以悲伤或不快为特征的情绪状态。</li> </ul> <h3>数据</h3> <p>共有超过100名说话人参与数据采集,经筛选后,最终纳入本语料库的录音来自68名说话人(23名女性,45名男性),多数受试者在采集时年龄处于二十余岁区间。说话人相关信息详见"SpeakerInfo.doc"。</p> <p>受试者需阅读一段包含5个短语、15个句子与2段文本的材料,以生成带情感的语音。该材料覆盖普通话的所有音素。每名受试者需朗读对应五种情绪(中性、愤怒、喜悦、恐慌、悲伤)的短语、段落与句子。本语料库总计包含25636条语音片段。朗读材料的具体构成如下:</p> <ul> <li>5个短语:“是”“否”以及“苹果”“火车”“网球”——在中文语境中,这些词汇包含多种不同的基础元音与辅音。</li> <li>20个句子:覆盖普通话的所有音素与绝大多数常见辅音连缀,句子类型包括:简单陈述句、带枚举内容的陈述句、一般疑问句(是非问句)、选择疑问句、祈使句、感叹句、特殊疑问句(wh问句)。</li> <li>2段文本:一段选自中国经典小说,另一段为普通事实陈述。</li> </ul> <p>所有数据均在安静办公室内使用奥林巴斯(OLYMPUS)DM-20数字语音记录仪采集,采样率为22050Hz。后续通过USB(通用串行总线,Universal Serial Bus)将录音文件传输至个人计算机,并转换为单声道Windows PCM格式,采样频率为8kHz,采样精度为16比特。</p> <p>关于本语料库的数据与研究方法的更多信息,详见作者发表的论文《MASC:用于情感分析与情感说话人识别的普通话语音语料库》(MASC.pdf)。</p> <h3>示例</h3> <p>如需查看本语料库的数据示例,请聆听以下音频:</p> <ul> <li><a href="./desc/addenda/LDC2007S09_neutral.wav" rel="nofollow">中性语音</a></li> <li><a href="./desc/addenda/LDC2007S09_anger.wav" rel="nofollow">愤怒语音</a></li> <li><a href="./desc/addenda/LDC2007S09_elation.wav" rel="nofollow">喜悦语音</a></li> </ul> </br> 部分内容 © 2005-2007 浙江大学先进计算与系统实验室,© 2007 宾夕法尼亚大学董事会




