GALE Phase 4 Chinese Broadcast Conversation Speech
收藏资源简介:
<h3>Introduction</h3><br> <p>GALE Phase 4 Chinese Broadcast Conversation Speech was developed by the Linguistic Data Consortium (LDC) and is comprised of approximately 172 hours of Mandarin Chinese broadcast conversation speech collected in 2008 by LDC and Hong Kong University of Science and Technology (HKUST), Hong Kong, during Phase 4 of the DARPA GALE (Global Autonomous Language Exploitation) Program.</p><br> <p>Corresponding transcripts are released as GALE Phase 4 Chinese Broadcast Conversation Transcripts (<a href="http://catalog.ldc.upenn.edu/LDC2016T12">LDC2016T12</a>).</p><br> <p>Broadcast audio for the GALE program was collected at LDC’s Philadelphia, PA USA facilities and at three remote collection sites: HKUST (Chinese), Medianet (Tunis, Tunisia) (Arabic), and MTC (Rabat, Morocco) (Arabic). The combined local and outsourced broadcast collection supported GALE at a rate of approximately 300 hours per week of programming from more than 50 broadcast sources for a total of over 30,000 hours of collected broadcast audio over the life of the program.</p><br> <p>LDC’s local <a href="https://www.ldc.upenn.edu/about/facilities/broadcast-collection"> broadcast collection system</a> is highly automated, easily extensible and robust and capable of collecting, processing and evaluating hundreds of hours of content from several dozen sources per day. The broadcast material is served to the system by a set of free-to-air (FTA) satellite receivers, commercial direct satellite systems (DSS) such as DirecTV, direct broadcast satellite (DBS) receivers, and cable television (CATV) feeds. The mapping between receivers and recorders is dynamic and modular. All signal routing is performed under computer control, using a 256x64 A/V matrix switch. Programs are recorded in a high bandwidth A/V format and are then processed to extract audio, to generate keyframes and compressed audio/video, to produce time-synchronized closed captions (in the case of North American English) and to generate automatic speech recognition (ASR) output. An overview of the system, the sources recorded and the configuration of the recording laboratory are contained in the Guidelines for Broadcast Audio Collection Version 3.0 included in this release.</p><br> <p>LDC designed a portable platform for remote broadcast collection. This is a TiVO-style digital video recording (DVR) system that records two streams of A/V material simultaneously. It supports analog CATV (NTSC and PAL) and FTA DVB-S satellite programming and can operate outside of the United States. It has a small footprint, weighs less than 30 pounds and can be transported as carry-on luggage.</p><br> <p>HKUST collected Chinese broadcast programming using its internal recording system and a portable broadcast collection platform designed by LDC and installed at HKUST in 2006.</p><br> <h3>Data</h3><br> <p>The broadcast conversation recordings in this release feature interviews, call-in programs and roundtable discussions focusing principally on current events from the following sources: Beijing TV, a national television station in Mainland China; China Central TV (CCTV), a national and international broadcaster in Mainland China; Hubei TV, a regional television station in Mainland China, Hubei Province; Phoenix TV, a Hong Kong-based satellite television station ; and Voice of America (VOA), a U.S. government-funded broadcast programmer.</p><br> <p>This release contains 236 audio files presented in <a href="http://flac.sourceforge.net">FLAC</a>-compressed Waveform Audio File format (.flac), 16000 Hz single-channel 16-bit PCM. Each file was audited by a native Chinese speaker following Audit Procedure Specification Version 2.0 which is included in this release. The broadcast auditing process served three principal goals: as a check on the operation of the broadcast collection system equipment by identifying failed, incomplete or faulty recordings; as an indicator of broadcast schedule changes by identifying instances when the incorrect program was recorded; and as a guide for data selection by retaining information about a program’s genre, data type and topic.</p><br> <h3>Samples</h3><br> <p>Please listen to this <a href="desc/addenda/LDC2016S03.wav">sample</a>.</p><br> <h3>Updates</h3><br> <p>None at this time.</p><br> <h3>Acknowledgment</h3><br> <p>This work was supported in part by the Defense Advanced Research Projects Agency, GALE Program Grant No. HR0011-06-1-0003. The content of this publication does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred.</p></br> Portions © 2008 Beijing TV, China Central TV, Hubei TV, Phoenix TV, © 2008, 2011, 2016 Trustees of the University of Pennsylvania
<h3>引言</h3><br><p>GALE第四阶段汉语广播对话语音语料库由语言数据联盟(Linguistic Data Consortium, LDC)开发完成,包含约172小时汉语普通话广播对话语音数据。该语料采集工作由LDC与香港科技大学(Hong Kong University of Science and Technology, HKUST)于2008年开展,隶属于美国国防高级研究计划局(Defense Advanced Research Projects Agency, DARPA)主导的GALE(全球自主语言利用,Global Autonomous Language Exploitation)项目第四阶段。</p><br><p>配套的转写文本已以《GALE第四阶段汉语广播对话转写文本》(<a href="http://catalog.ldc.upenn.edu/LDC2016T12">LDC2016T12</a>)形式发布。</p><br><p>GALE项目的广播音频采集工作分别于LDC位于美国宾夕法尼亚州费城的本地机房,以及三处远程采集站点完成:香港科技大学(负责汉语语料)、突尼斯突尼斯市的Medianet站点(负责阿拉伯语语料)、摩洛哥拉巴特的MTC站点(负责阿拉伯语语料)。本项目采用本地与外包结合的广播采集方案,可从50余个广播源每周获取约300小时节目内容,整个项目周期内累计采集广播音频超过30000小时。</p><br><p>LDC本地的<a href="https://www.ldc.upenn.edu/about/facilities/broadcast-collection">广播采集系统</a>具备高度自动化、易扩展且鲁棒性强的特点,可每日从数十个信号源采集、处理并评估数百小时的内容。该系统的信号输入来源包括多套免费接收(Free-to-air, FTA)卫星接收机、DirecTV等商业直播卫星系统(Direct Satellite System, DSS)、直接广播卫星(Direct Broadcast Satellite, DBS)接收机以及有线电视(Cable Television, CATV)信号源。接收机与录像机的映射关系采用动态模块化设计,所有信号路由均通过计算机控制的256×64音视频矩阵切换器完成。节目首先以高带宽音视频格式录制,随后经过处理以提取音频、生成关键帧与压缩音视频文件、生成时间同步的隐藏式字幕(针对北美英语节目),并输出自动语音识别(Automatic Speech Recognition, ASR)结果。本发布包中附带的《广播音频采集指南V3.0》对本系统、采集信号源以及录制实验室的配置进行了详细说明。</p><br><p>LDC设计了一款用于远程广播采集的便携平台。该平台采用类TiVO的数字视频录像机(Digital Video Recorder, DVR)架构,可同时录制两路音视频流。其支持模拟有线电视(采用NTSC与PAL制式)以及免费接收DVB-S卫星节目,且可在美国境外部署。该平台体积小巧、重量不足30磅,可作为随身行李进行托运。</p><br><p>香港科技大学于2006年在本校部署了LDC设计的便携广播采集平台,并结合内部录制系统完成汉语广播节目的采集工作。</p><br><h3>数据</h3><br><p>本发布包中的广播对话录音内容涵盖访谈、热线节目与圆桌讨论,主题以时事为主,信号源来自以下机构:中国内地国家级电视台北京卫视、中国内地兼具国内与国际播出能力的中央广播电视总台(China Central Television, CCTV)、中国内地湖北省省级电视台湖北卫视、总部位于香港的卫星电视台凤凰卫视,以及由美国政府资助的广播机构美国之音(Voice of America, VOA)。</p><br><p>本发布包包含236个音频文件,采用<a href="http://flac.sourceforge.net">FLAC</a>压缩波形音频文件格式(.flac),采样率为16000Hz、单声道、16位脉冲编码调制(PCM)。本发布包附带的《审核流程规范V2.0》要求,所有音频文件均由汉语母语者完成人工审核。本次广播音频审核工作主要达成三大目标:一是通过识别失效、不完整或存在瑕疵的录音,对广播采集系统设备的运行状态进行校验;二是通过识别录错节目的案例,监测广播节目排期的变更情况;三是通过留存节目类型、数据类别与主题等信息,为数据筛选提供依据。</p><br><h3>示例</h3><br><p>请收听本<a href="desc/addenda/LDC2016S03.wav">示例音频</a>。</p><br><h3>更新说明</h3><br><p>暂无更新。</p><br><h3>致谢</h3><br><p>本研究部分得到美国国防高级研究计划局GALE项目资助(项目编号:HR0011-06-1-0003)。本出版物的内容不代表美国政府的立场或政策,不应被视为获得官方背书。</p><br><p>部分内容 © 2008 北京卫视、中央广播电视总台、湖北卫视、凤凰卫视;© 2008、2011、2016 宾夕法尼亚大学理事会</p>




