遇见数据集

TRAD Chinese-French Parallel Text -- Broadcast News

收藏
DataCite Commons2021-07-01 更新2025-04-16 收录
官方服务:

资源简介:

<h3>Introduction</h3><br> <p>TRAD Chinese-French Parallel Text -- Broadcast News was developed by <a href="http://elda.org/en/">ELDA</a> as part of the <a href="http://www.elra.info/en/projects/archived-projects/pea-trad/">PEA-TRAD project</a>. It contains French translations of a subset of approximately 30,000 Chinese characters from GALE Phase 1 Chinese Broadcast News Parallel Text - Part 3 (<a href="../../../LDC2008T18">LDC2008T18</a>).</p><br> <p>The PEA-TRAD project (Translation as a Support for Document Analysis) was supported by the French Ministry of Defense (DGA). Its purpose was to develop speech-to-speech translation technology for multiple languages (e.g., Arabic, Chinese, Pashto) from a variety of domains. ELDA developed several corpora for this effort.</p><br> <p>The Linguistic Data Consortium (LDC) has also released the following TRAD corpora:</p><br> <ul><br> <li>TRAD Chinese-French Parallel Text -- Blog (<a href="../../../LDC2018T02">LDC2018T02</a>)</li><br> <li>TRAD Arabic-French Parallel Text -- Newsgroup (<a href="../../../LDC2018T13">LDC2018T13</a>)</li><br> <li>TRAD Arabic-French Parallel Text -- Newswire (<a href="../../../LDC2018T21">LDC2018T21</a>)</li><br> </ul><br> <h3>Data</h3><br> <p>This release consists of 977 segments (translation units) from 139 documents. The source data is Chinese broadcast news collected and translated into English by LDC for the DARPA GALE (Global Autonomous Language Exploitation) program. Information about the ELDA translation team, translation guidelines and validation results is contained in the documentation accompanying this release.</p><br> <p>The Chinese source file contains 33,571 characters and the French reference translation contains 22,424 words. The data is presented in two unicode-encoded XML files along with an associated DTD.</p><br> <h3>Samples</h3><br> <p>Please view this <a href="desc/addenda/LDC2018T17.src.xml">source sample</a> and <a href="desc/addenda/LDC2018T17.ref.xml">reference sample</a>.</p><br> <h3>Updates</h3><br> <p>None at this time.</p></br> Portions © 2005, 2006 China Central TV, © 2005, 2006 Phoenix TV, © 2018 ELDA, © 2005-2006, 2008, 2018 Trustees of the University of Pennsylvania

<h3>引言</h3><br><p>TRAD汉法平行语料库——广播新闻由<a href="http://elda.org/en/">ELDA</a>开发,作为<a href="http://www.elra.info/en/projects/archived-projects/pea-trad/">PEA-TRAD项目(“翻译辅助文档分析”项目)</a>的一部分。该语料库包含源自GALE第一阶段中文广播新闻平行语料库第3部分(<a href="../../../LDC2008T18">LDC2008T18</a>)中约30000个汉字的子集的法文译稿。</p><br><p>PEA-TRAD项目(“翻译辅助文档分析”项目)由法国国防部军备总局(DGA)资助,其目标是研发面向多领域多语种(如阿拉伯语、汉语、普什图语)的语音到语音翻译技术,ELDA为该项目开发了多套语料库。</p><br><p>语言数据联盟(Linguistic Data Consortium,LDC)还发布了以下TRAD系列语料库:</p><br><ul><br><li>TRAD汉法平行语料库——博客文本(<a href="../../../LDC2018T02">LDC2018T02</a>)</li><br><li>TRAD阿法平行语料库——新闻组文本(<a href="../../../LDC2018T13">LDC2018T13</a>)</li><br><li>TRAD阿法平行语料库——新闻专线文本(<a href="../../../LDC2018T21">LDC2018T21</a>)</li><br></ul><br><h3>数据概况</h3><br><p>本次发布的语料包含来自139份文档的977个片段(翻译单元)。源数据为LDC为美国国防高级研究计划局(DARPA)GALE(全球自主语言利用,Global Autonomous Language Exploitation)计划采集并英译的中文广播新闻。本次发布附带的文档中包含ELDA翻译团队、翻译规范及验证结果的相关信息。</p><br><p>中文源文件包含33571个字符,法文参考译文字数为22424词。本语料以两份Unicode编码的XML文件及配套文档类型定义(DTD)文件形式呈现。</p><br><h3>示例</h3><br><p>请查看<a href="desc/addenda/LDC2018T17.src.xml">源文件示例</a>与<a href="desc/addenda/LDC2018T17.ref.xml">参考译文示例</a>。</p><br><h3>更新记录</h3><br><p>暂无更新。</p></br><p>部分内容 © 2005、2006 中国中央电视台,© 2005、2006 凤凰卫视,© 2018 ELDA,© 2005-2006、2008、2018 宾夕法尼亚大学托管方</p>

创建时间:
2020-11-30
搜集汇总
数据集介绍
TRAD Chinese-French Parallel Text -- Broadcast News 数据集图片
背景与挑战
背景概述
该数据集是TRAD项目下的中法平行文本,专门针对广播新闻领域,包含977个翻译单元(源自139个文档),中文源文本有33,571个字符,法语参考翻译有22,424个单词。它由ELDA开发,属于PEA-TRAD项目的一部分,旨在支持多语言语音到语音翻译技术,适用于语言建模和机器翻译等应用。数据以XML格式提供,聚焦于中文和法语之间的翻译任务。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务