Chinese English News Magazine Parallel Text
收藏资源简介:
<h3>Introduction</h3> <p>This file contains documentation on the Chinese English News Magazine Parallel Text, Linguistic Data Consortium (LDC) catalog number LDC2005T10 and ISBN 1-58563-333-X. </p><p>This corpus contains Chinese news stories and their English translations LDC collected via Sinorama Magazine, Taiwan, from 1976 to 2004. It totals 6,366 story pairs, 365,568 sentence pairs, 20M Chinese characters and 9M English words. The corpus is aligned at sentence level.</p><h3>Data</h3> <p>Sinorama Magazine is published monthly in several languages, including Chinese, English, Japanese. LDC received its 1976 to 2000 publications on a single CD, and its 2001 to 2004 publications via Sinorama's website. </p> <p>The Sinorama Chinese text was encoded in Big5. The data came story aligned but were lack of sentence level alignment. The sentence alignment was done at the LDC using <a href="http://champollion.sourceforge.net" rel="nofollow">Champollion</a> v 1.1. </p> <p>The final data is put in the data directory, which contains subdirectories for Chinese documents, English documents, and the sentence level alignment, identified as "Chinese," "English," and "alignment." </p> <p>The English and Chinese files may contain one or more documents, with each document formatted in SGML as follows: </p><p> [English or Chinese text] </p><p> [English or Chinese text] [English or Chinese text] ... </p><p>Notes: * the </p><p> and tags are always assigned sequential numeric IDs, starting at one. * the tags are always placed on the same line with their contents, and are always separated from the contents by a space. </p> * if an English file and a Chinese file share the same file name, they contain the same documents. * all Chinese text is encoded in Big5. <p>Each alignment file contains the sentence level alignment of multiple documents, each being formatted in SGML as follows: ... </p><p>Notes: * the docid in an English file, its Chinese translation and the ALIGNMENT are the same. * EnglishSegId and ChineseSegId may have none, one, or more than one segment IDs. </p><h3>Samples</h3> <p>The following files provide an example of this corpus: </p><ul> <li> <a href="desc/addenda/LDC2005T10_Chinese.txt" rel="nofollow">Chinese</a> </li> <li> <a href="desc/addenda/LDC2005T10_English.txt" rel="nofollow">English</a> </li> <li> <a href="desc/addenda/LDC2005T10_align.txt" rel="nofollow">Alignment</a> </li> </ul> <p>Portions © 2005 Trustees of the University of Pennsylvania</p></br> Portions © 1976-2004 Sinorama Magazine <br><br>Portions © 2005 Trustees of the University of Pennsylvania
<h3>简介</h3> <p>本文件包含关于汉英新闻杂志平行语料库的说明文档,该语料库由语言数据联盟(Linguistic Data Consortium, LDC)发布,目录编号为LDC2005T10,ISBN为1-58563-333-X。</p><p>本语料库收录了1976年至2004年间,由中国台湾《光华杂志》(Sinorama Magazine)提供的中文新闻稿件及其英文译稿。语料库总计包含6,366对稿件、365,568对句子,涵盖2000万中文字符与900万英文单词,且已在句子级别完成对齐。</p><h3>数据</h3> <p>《光华杂志》(Sinorama Magazine)以多语言月刊形式发行,涵盖中文、英文、日文等版本。语言数据联盟获取了该刊1976年至2000年的出版物(存储于单张CD),以及2001年至2004年的网络版出版物。</p> <p>初始获取的《光华杂志》中文文本采用Big5编码,稿件层面已完成对齐,但未实现句子级对齐。语言数据联盟使用<a href="http://champollion.sourceforge.net" rel="nofollow">Champollion</a> v1.1工具完成了句子对齐工作。</p> <p>最终处理完成的语料存储于data目录,该目录下设三个子目录:分别存放中文文档(Chinese)、英文文档(English)以及句子级对齐文件(alignment)。</p> <p>英文与中文文件可包含一个或多个文档,每份文档均采用SGML格式编写,格式如下:</p><p> [英文或中文文本内容]</p><p> [英文或中文文本内容] [英文或中文文本内容] ...</p><p>注意事项:* <doc>与</doc>标签始终按从1开始的连续数字顺序分配ID。* <p>标签始终与其内容置于同一行,且与内容间以空格分隔。* 若英文文件与中文文件文件名一致,则二者包含的文档完全对应。* 所有中文文本均采用Big5编码。</p><p>每个对齐文件包含多篇文档的句子级对齐信息,格式同样采用SGML,示例如下:……</p><p>注意事项:* 英文文件、其中文译稿以及对齐文件的docid(文档ID)完全一致。* EnglishSegId与ChineseSegId可包含零个、一个或多个分段ID。</p><h3>示例</h3> <p>以下文件为本语料库的示例样本:</p><ul> <li> <a href="desc/addenda/LDC2005T10_Chinese.txt" rel="nofollow">中文样例</a> </li> <li> <a href="desc/addenda/LDC2005T10_English.txt" rel="nofollow">英文样例</a> </li> <li> <a href="desc/addenda/LDC2005T10_align.txt" rel="nofollow">对齐样例</a> </li> </ul> <p>本语料库部分内容 © 2005 宾夕法尼亚大学托管委员会</p></br> 本语料库部分内容 © 1976-2004 《光华杂志》 <br><br>本语料库部分内容 © 2005 宾夕法尼亚大学托管委员会




