GALE Phase 1 Chinese Broadcast Conversation Parallel Text - Part 1
收藏资源简介:
<h3>Introduction:</h3> <p> GALE Phase 1 Chinese Broadcast Conversation Parallel Text - Part 1, Linguistic Data Consortium (LDC) catalog number LDC2009T02 and ISBN 1-58563-499-9, contains transcripts and English translations of 20.4 hours of Chinese broadcast conversation programming from China Central TV (CCTV) and Phoenix TV. It does not contain the audio files from which the transcripts and translations were generated. GALE Phase 1 Chinese Broadcast Conversation Parallel Text - Part 1, along with other corpora, was used as training data in year 1 (Phase 1) of the DARPA-funded GALE program. </p> <h3>Source Data:</h3> <p> A total of 20.4 hours of Chinese broadcast conversation programming were selected from two sources: CCTV (a broadcaster from Mainland China), and Phoenix TV (a Hong Kong -based satellite TV station). The transcripts and translations represent recordings of eight different programs. </p> <p>A manual selection procedure was used to choose data appropriate for the GALE program, namely conversation (talk) programs focusing on current events. Stories on topics such as sports, entertainment and business were excluded from the data set. The following table is a summary of the files included in this release. </p> <table><tbody> <tr> <td width="108"> <p> Source </p> </td> <td width="164"> <p> Program </p> </td> <td width="162"> <p> Epoch (YYYY.MM) </p> </td> <td width="96"> <p> #hours </p> </td> <td width="105"> <p> #characters </p> </td> </tr> <tr> <td rowspan="2" width="108"> <p> CCTV </p> </td> <td width="164"> <p> Across China </p> </td> <td width="162"> <p> 2005.08 </p> </td> <td width="96"> <p> 1.0 </p> </td> <td width="105"> <p> 9,924 </p> </td> </tr> <tr> <td width="164"> <p> Todays Focus </p> </td> <td width="162"> <p> 2005.11 </p> </td> <td width="96"> <p> 2.2 </p> </td> <td width="105"> <p> 33,805 </p> </td> </tr> <tr> <td colspan="1" rowspan="6" width="108"> <p> Phoenix TV </p> </td> <td width="164"> <p> Asian Journal </p> </td> <td width="162"> <p> 2005.09 </p> </td> <td width="96"> <p> 2.2 </p> </td> <td width="105"> <p> 26,656 </p> </td> </tr> <tr> <td width="164"> <p> Behind the Headlines </p> </td> <td width="162"> <p> 2005.03 - 2005.11 </p> </td> <td width="96"> <p> 1.5 </p> </td> <td width="105"> <p> 17,933 </p> </td> </tr> <tr> <td> <p> A Date With Lu Yu </p> </td> <td> <p> 2005.09 - 2005.10 </p> </td> <td> <p> 7.1 </p> </td> <td> <p> 89,987 </p> </td> </tr> <tr> <td> <p> News Hacker </p> </td> <td> <p> 2005.03 - 2005.10 </p> </td> <td> <p> 2.3 </p> </td> <td> <p> 39,388 </p> </td> </tr> <tr> <td> <p> Newsline </p> </td> <td> <p> 2005.10 - 2005.11 </p> </td> <td> <p> 1.6 </p> </td> <td> <p> 15,496 </p> </td> </tr> <tr> <td> <p> Social Watch </p> </td> <td> <p> 2005.09 - 2005.11 </p> </td> <td> <p> 2.5 </p> </td> <td> <p> 29,159 </p> </td> </tr> </tbody></table><h3>Transcription:</h3> <p> The selected audio snippets were carefully transcribed by LDC annotators and professional transcription agencies following LDCs Quick Rich Transcription specification. Manual sentence units/segments (SU) annotation was also performed as part of the transcription task. Three types of end of sentence SU are identified: </p> <ul> <li> <p>statement SU</p> </li> <li> <p>question SU</p> </li> <li> <p>incomplete SU</p> </li> </ul><h3>Translation:</h3> <p> After transcription and SU annotation, files were reformatted into a human-readable translation format and assigned to professional translators for careful translation. Translators followed LDCs GALE Translation guidelines which describe the makeup of the translation team, the source data format, the translation data format, best practices for translating certain linguistic features (such as names and speech disfluencies) and quality control procedures applied to completed translations. </p> <h3>TDF Format:</h3> <p> All final data are in Tab Delimited Format (TDF). TDF is compatible with other transcription formats, such as the Transcriber format and AG format, and it is easy to process. </p> <p> Each line of a TDF file corresponds to a speech segment and contains 13 tab delimited fields: </p> <table width="259"><tbody> <tr> <td> <p><strong> Field</strong></p> </td> <td> <p><strong>Data Type</strong></p> </td> </tr> <tr> <td><p> file </p></td> <td><p>unicode </p></td> </tr> <tr> <td><p>channel </p></td> <td><p>int </p></td> </tr> <tr> <td><p>start </p></td> <td><p>float </p></td> </tr> <tr> <td><p>end </p></td> <td><p>float </p></td> </tr> <tr> <td><p>speaker </p></td> <td><p>unicode </p></td> </tr> <tr> <td><p>speakerType </p></td> <td><p>unicode </p></td> </tr> <tr> <td><p> speakerDialect </p></td> <td><p>unicode </p></td> </tr> <tr> <td><p>transcript </p></td> <td><p>unicode </p></td> </tr> <tr> <td><p>section </p></td> <td><p>int </p></td> </tr> <tr> <td><p>turn </p></td> <td><p>int </p></td> </tr> <tr> <td><p>segment </p></td> <td><p>int </p></td> </tr> <tr> <td><p>sectionType </p></td> <td><p>unicode </p></td> </tr> <tr> <td><p>suType </p></td> <td><p>unicode </p></td> </tr> </tbody></table><p> A source TDF file and its translation are the same except that the transcript in the source TDF is replaced by its English translation. </p> <h3>Sponsorship</h3> <p> This work was supported in part by the Defense Advanced Research Projects Agency, GALE Program Grant No. HR0011-06-1-0003. The content of this publication does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred. </p> <h3>Samples</h3> For an example of the data in this corpus, please examine these images of the <a href="./desc/addenda/LDC2009T02_source.png" rel="nofollow">source</a> and<a href="./desc/addenda/LDC2009T02_trans.png" rel="nofollow"> translation</a>. </br> Portions © 2005 China Central TV, © 2005 Phoenix TV, © 2005 - 2007, 2009 Trustees of the University of Pennsylvania.
<h3>简介:</h3> <p>GALE第一阶段中文广播对话平行语料库(第一部分)由语言数据联盟(Linguistic Data Consortium, LDC)发布,目录号为LDC2009T02,ISBN为1-58563-499-9。该语料库包含来自中国中央电视台(CCTV)与凤凰卫视的20.4小时中文广播对话节目的转写文本及英文译文,未附带生成转写与译文的原始音频文件。本语料库与其他语料一同作为美国国防高级研究计划局(Defense Advanced Research Projects Agency, DARPA)资助的GALE项目第一年度(第一阶段)的训练数据使用。</p> <h3>源数据:</h3> <p>本次数据集共收录20.4小时的中文广播对话节目,素材来源于两大渠道:中国大陆的中国中央电视台(CCTV),以及总部位于香港的卫星电视台凤凰卫视。所有转写文本与译文对应8档不同节目。</p> <p>本次数据采用人工遴选流程,筛选适配GALE项目需求的内容,即聚焦时事的谈话类对话节目。体育、娱乐与商业类主题的内容均被排除在数据集之外。下表为本批次发布文件的汇总信息。</p> <table><tbody> <tr> <td width="108"> <p> 来源 </p> </td> <td width="164"> <p> 节目名称 </p> </td> <td width="162"> <p> 录制时段(YYYY.MM) </p> </td> <td width="96"> <p> 时长(小时) </p> </td> <td width="105"> <p> 字符数 </p> </td> </tr> <tr> <td rowspan="2" width="108"> <p> CCTV </p> </td> <td width="164"> <p> 走遍中国(Across China) </p> </td> <td width="162"> <p> 2005.08 </p> </td> <td width="96"> <p> 1.0 </p> </td> <td width="105"> <p> 9,924 </p> </td> </tr> <tr> <td width="164"> <p> 今日关注(Todays Focus) </p> </td> <td width="162"> <p> 2005.11 </p> </td> <td width="96"> <p> 2.2 </p> </td> <td width="105"> <p> 33,805 </p> </td> </tr> <tr> <td colspan="1" rowspan="6" width="108"> <p> 凤凰卫视(Phoenix TV) </p> </td> <td width="164"> <p> 亚洲新闻周刊(Asian Journal) </p> </td> <td width="162"> <p> 2005.09 </p> </td> <td width="96"> <p> 2.2 </p> </td> <td width="105"> <p> 26,656 </p> </td> </tr> <tr> <td width="164"> <p> 新闻背后(Behind the Headlines) </p> </td> <td width="162"> <p> 2005.03 - 2005.11 </p> </td> <td width="96"> <p> 1.5 </p> </td> <td width="105"> <p> 17,933 </p> </td> </tr> <tr> <td> <p> 鲁豫有约(A Date With Lu Yu) </p> </td> <td> <p> 2005.09 - 2005.10 </p> </td> <td> <p> 7.1 </p> </td> <td> <p> 89,987 </p> </td> </tr> <tr> <td> <p> 新闻黑客(News Hacker) </p> </td> <td> <p> 2005.03 - 2005.10 </p> </td> <td> <p> 2.3 </p> </td> <td> <p> 39,388 </p> </td> </tr> <tr> <td> <p> 新闻线(Newsline) </p> </td> <td> <p> 2005.10 - 2005.11 </p> </td> <td> <p> 1.6 </p> </td> <td> <p> 15,496 </p> </td> </tr> <tr> <td> <p> 社会观察(Social Watch) </p> </td> <td> <p> 2005.09 - 2005.11 </p> </td> <td> <p> 2.5 </p> </td> <td> <p> 29,159 </p> </td> </tr> </tbody></table><h3>转写流程:</h3> <p> 遴选后的音频片段由LDC标注人员及专业转写机构遵照LDC快速富转录(Quick Rich Transcription)规范完成转写。转写任务同时包含人工句子单元(Sentence Unit, SU)标注环节,共识别三类句子单元结尾类型:</p> <ul> <li> <p>陈述型句子单元</p> </li> <li> <p>疑问型句子单元</p> </li> <li> <p>不完整型句子单元</p> </li> </ul><h3>译文制作:</h3> <p> 完成转写与句子单元标注后,文件将被重新格式化为易读的译文格式,并交由专业译员进行精准翻译。译员需遵循LDC发布的GALE翻译指南,该指南涵盖翻译团队构成、源数据格式、译文数据格式、特定语言特征(如人名与言语不流畅现象)的翻译规范,以及译文完成后的质量管控流程。</p> <h3>制表符分隔格式(Tab Delimited Format, TDF):</h3> <p> 所有最终数据均采用制表符分隔格式(TDF)存储。该格式兼容Transcriber格式与AG格式等其他转录格式,且易于处理。</p> <p> 每个TDF文件的行对应一条语音片段,共包含13个制表符分隔的字段:</p> <table width="259"><tbody> <tr> <td> <p><strong> 字段 </strong></p> </td> <td> <p><strong> 数据类型 </strong></p> </td> </tr> <tr> <td><p> file </p></td> <td><p>unicode </p></td> </tr> <tr> <td><p> channel </p></td> <td><p>整数型 </p></td> </tr> <tr> <td><p> start </p></td> <td><p>浮点型(起始时间) </p></td> </tr> <tr> <td><p> end </p></td> <td><p>浮点型(结束时间) </p></td> </tr> <tr> <td><p> speaker </p></td> <td><p>unicode(发言者标识) </p></td> </tr> <tr> <td><p> speakerType </p></td> <td><p>unicode(发言者类型) </p></td> </tr> <tr> <td><p> speakerDialect </p></td> <td><p>unicode(发言者方言) </p></td> </tr> <tr> <td><p> transcript </p></td> <td><p>unicode(转写文本) </p></td> </tr> <tr> <td><p> section </p></td> <td><p>整数型(章节编号) </p></td> </tr> <tr> <td><p> turn </p></td> <td><p>整数型(轮次编号) </p></td> </tr> <tr> <td><p> segment </p></td> <td><p>整数型(片段编号) </p></td> </tr> <tr> <td><p> sectionType </p></td> <td><p>unicode(章节类型) </p></td> </tr> <tr> <td><p> suType </p></td> <td><p>unicode(句子单元类型) </p></td> </tr> </tbody></table><p> 源语言TDF文件与译文TDF文件的结构完全一致,仅将源语言TDF中的转写文本替换为对应的英文译文。</p> <h3>资助说明</h3> <p> 本工作部分由美国国防高级研究计划局(Defense Advanced Research Projects Agency, DARPA)的GALE项目资助,资助编号为HR0011-06-1-0003。本出版物内容未必代表美国政府的立场或政策,不应被视为获得官方背书。</p> <h3>样本</h3> <p> 如需查看本语料库的数据示例,请参阅以下链接中的<a href="./desc/addenda/LDC2009T02_source.png" rel="nofollow">源文件</a>与<a href="./desc/addenda/LDC2009T02_trans.png" rel="nofollow">译文文件</a>图像。</p> <br> <p> 部分内容 © 2005 中国中央电视台,© 2005 凤凰卫视,© 2005 - 2007, 2009 宾夕法尼亚大学理事会。</p>




