遇见数据集

GALE Arabic-English Parallel Aligned Treebank -- Broadcast News Part 1

收藏
DataCite Commons2021-07-01 更新2025-04-16 收录
官方服务:

资源简介:

<p>GALE Arabic-English Parallel Aligned Treebank -- Broadcast News Part 1 was developed by the Linguistic Data Consortium (LDC) and contains 115,826 tokens of word aligned Arabic and English parallel text with treebank annotations. This material was used as training data in the DARPA GALE (Global Autonomous Language Exploitation) program.</p> <p>Parallel aligned treebanks are treebanks annotated with morphological and syntactic structures aligned at the sentence level and the sub-sentence level. Such data sets are useful for natural language processing and related fields, including automatic word alignment system training and evaluation, transfer-rule extraction, word sense disambiguation, translation lexicon extraction and cultural heritage and cross-linguistic studies. With respect to machine translation system development, parallel aligned treebanks may improve system performance with enhanced syntactic parsers, better rules and knowledge about language pairs and reduced word error rate.</p> <p> In this release, the source Arabic data was translated into English. Arabic and English treebank annotations were performed independently. The parallel texts were then word aligned. The material in this corpus corresponds to a portion of the Arabic treebanked data in Arabic Treebank - Broadcast News v1.0 (<a href="http://catalog.ldc.upenn.edu/LDC2012T07" rel="nofollow">LDC2012T07</a>).</p> <h3>Data</h3> <p>The source data consists of Arabic broadcast news programming collected by LDC in 2005 and 2006 from Alhurra, Aljazeera and Dubai TV. All data is encoded as UTF-8. A count of files, words, tokens and segments is below.</p> <table> <tr> <td>Language</td> <td>Files</td> <td>Words</td> <td>Tokens</td> <td>Segments</td> </tr> <tr> <td>Arabic</td> <td>28</td> <td>89,213</td> <td>115,826</td> <td>4,824</td> </tr> </table><p>Note: Word count is based on the untokenized Arabic source. Token count is based on the ATB-tokenized Arabic source.</p> <p>The purpose of the GALE word alignment task was to find correspondences between words, phrases or groups of words in a set of parallel texts. Arabic-English word alignment annotation consisted of the following tasks:</p> <ul> <li>Identifying different types of links: translated (correct or incorrect) and not translated (correct or incorrect)</li> <li>Identifying sentence segments not suitable for annotation, e.g., blank segments, incorrectly-segmented segments, segments with foreign languages</li> <li>Tagging unmatched words attached to other words or phrases</li> </ul><p>This release contains four types of files - raw, tokenized, treebank, and wa. The raw format contains the original Arabic and English sentences without any annotation. The tokenized format is the treebank tokenized version of the raw data which may contain <i>Empty Category</i> tokens (treebank leaves that have the POS label -NONE-). The treebank and wa files are treebank and word alignment annotations on the tokenized files.</p> <h3>Samples</h3> <p>Please view the below samples.</p> <ul> <li><a href="./desc/addenda/LDC2013T14.eng.raw.txt" rel="nofollow">English Raw</a></li> <li><a href="./desc/addenda/LDC2013T14.eng.tkn.txt" rel="nofollow">English Token</a></li> <li><a href="./desc/addenda/LDC2013T14.eng.tree.txt" rel="nofollow">English Tree</a></li> <li><a href="./desc/addenda/LDC2013T14.arb.raw.jpg" rel="nofollow">Arabic Raw</a></li> <li><a href="./desc/addenda/LDC2013T14.arb.tkn.jpg" rel="nofollow">Arabic Token</a></li> <li><a href="./desc/addenda/LDC2013T14.arb.tree.jpg" rel="nofollow">Arabic Tree</a></li> <li><a href="./desc/addenda/LDC2013T14.wa.txt" rel="nofollow">Word Alignment</a></li> </ul><h3>Sponsorship</h3> <p>This work was supported in part by the Defense Advanced Research Projects Agency, GALE Program Grant No. HR0011-06-1-0003. The content of this publication does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred.</p> <h3>Updates</h3> <p>None at this time. </p> </br> Portions © 2005-2006 Aljzaeera, © 2005 Dubai TV, © 2005-2006, 2012, 2013 Trustees of the University of Pennsylvania

<p>GALE阿英平行对齐树库(GALE Arabic-English Parallel Aligned Treebank)——广播新闻第一部分,由语言数据联盟(Linguistic Data Consortium, LDC)开发,包含115,826个经词对齐的阿英平行文本语料,并附带树库(treebank)标注。该语料曾作为国防高级研究计划局(Defense Advanced Research Projects Agency, DARPA)GALE(全球自主语言开发,Global Autonomous Language Exploitation)项目的训练数据使用。</p><p>平行对齐树库(parallel aligned treebank)是指在句子级与子句级完成对齐的、附带形态与句法结构标注的树库。此类语料集可应用于自然语言处理及相关领域,包括自动词对齐系统的训练与评估、迁移规则抽取、词义消歧、翻译词典抽取,以及文化遗产与跨语言研究。在机器翻译系统开发场景中,平行对齐树库可通过优化句法分析器、完善语言对相关规则与知识库,以及降低词错误率,从而提升系统性能。</p><p>本次发布的源语言为阿拉伯语文本,已被译为英语。阿英树库标注均独立完成,随后对平行文本进行词对齐处理。本语料库中的语料源自阿拉伯语树库——广播新闻v1.0(Arabic Treebank - Broadcast News v1.0)中的部分阿拉伯语树库语料(详见<a href="http://catalog.ldc.upenn.edu/LDC2012T07" rel="nofollow">LDC2012T07</a>)。</p><h3>数据概况</h3><p>源语料为语言数据联盟(LDC)于2005年至2006年间从Alhurra、半岛电视台(Aljazeera)以及迪拜电视台(Dubai TV)采集的阿拉伯语广播新闻节目。所有数据均采用UTF-8编码。以下为文件数、词数、Token数与分段数的统计:</p><table><tr><td>语言</td><td>文件数</td><td>词数</td><td>Token数</td><td>分段数</td></tr><tr><td>阿拉伯语</td><td>28</td><td>89,213</td><td>115,826</td><td>4,824</td></tr></table><p>注:词数统计基于未分词的阿拉伯语源文本,Token数统计基于ATB分词(ATB-tokenized)的阿拉伯语源文本。</p><p>GALE词对齐任务的目标是识别平行文本集合中词、短语或词组之间的对应关系。阿英词对齐标注包含以下任务:</p><ul><li>识别不同类型的对齐链接:已翻译(正确或错误)与未翻译(正确或错误)</li><li>识别不适用于标注的句子分段,例如空白分段、分词错误的分段以及包含外语的分段</li><li>标记依附于其他词或短语的未匹配词</li></ul><p>本次发布包含四种格式的文件:原始格式(raw)、分词格式(tokenized)、树库格式(treebank)以及词对齐格式(wa)。原始格式包含未经过任何标注的阿英原句;分词格式为原始数据的树库分词版本,可能包含空范畴(Empty Category)Token,即词性标注为-NONE-的树库叶节点;树库格式与wa格式文件分别为分词文件上的树库标注与词对齐标注。</p><h3>示例样本</h3><p>请查看以下示例样本:</p><ul><li><a href="./desc/addenda/LDC2013T14.eng.raw.txt" rel="nofollow">英语原始文本</a></li><li><a href="./desc/addenda/LDC2013T14.eng.tkn.txt" rel="nofollow">英语分词文本</a></li><li><a href="./desc/addenda/LDC2013T14.eng.tree.txt" rel="nofollow">英语树库文本</a></li><li><a href="./desc/addenda/LDC2013T14.arb.raw.jpg" rel="nofollow">阿拉伯语原始文本</a></li><li><a href="./desc/addenda/LDC2013T14.arb.tkn.jpg" rel="nofollow">阿拉伯语分词文本</a></li><li><a href="./desc/addenda/LDC2013T14.arb.tree.jpg" rel="nofollow">阿拉伯语树库文本</a></li><li><a href="./desc/addenda/LDC2013T14.wa.txt" rel="nofollow">词对齐文件</a></li></ul><h3>资助说明</h3><p>本研究部分由国防高级研究计划局(Defense Advanced Research Projects Agency, DARPA)通过GALE项目资助,资助编号HR0011-06-1-0003。本出版物内容不一定代表美国政府的立场或政策,不应视为获得官方背书。</p><h3>更新说明</h3><p>暂无更新。</p></br><p>部分内容 © 2005-2006 半岛电视台(Aljazeera)、© 2005 迪拜电视台、© 2005-2006、2012、2013 宾夕法尼亚大学托管会</p>

创建时间:
2020-11-30
搜集汇总
数据集介绍
GALE Arabic-English Parallel Aligned Treebank -- Broadcast News Part 1 数据集图片
背景与挑战
背景概述
该数据集是GALE项目中的阿拉伯语-英语平行对齐树库,专注于广播新闻领域,包含115,826个标记的平行文本,带有树库注释和词对齐信息。数据来源于2005-2006年的阿拉伯语广播新闻节目,用于机器翻译、信息检测等自然语言处理任务,支持系统训练和评估。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务