Chinese Treebank 5.0
收藏资源简介:
<h3>Introduction</h3><br> <p>Chinese Treebank 5.0 was produced by Linguistic Data Consortium (LDC) catalog number LDC2005T01 and ISBN 1-58563-323-2.</p><br> <p>The Penn Chinese Treebank is an ongoing project that started in the summer of 1998. The goal of the project is to create a 500,000-word corpus of Chinese text with syntactic bracketing. Chinese Treebank 1.0 was first published in 2000, and it was later corrected and released in 2001 as <a href="http://catalog.ldc.upenn.edu/LDC2001T11" rel="nofollow">Chinese Treebank 2.0</a>. Another updated version was released in 2004 as <a href="http://catalog.ldc.upenn.edu/LDC2004T05" rel="nofollow"> Chinese Treebank 4.0</a>. More information about the project is available on the <a href="https://www.cs.brandeis.edu/~llc/page2/page2.html">Chinese Treebank</a> website.</p><br> <p>The content used in this corpus comes from the following newswire sources:</p><br> <table><br> <tbody><br> <tr><br> <td colspan="20%">698 articles</td><br> <td>Xinhua (1994-1998)</td><br> </tr><br> <tr><br> <td colspan="20%">55 articles</td><br> <td>Information Services Department of HKSAR (1997)</td><br> </tr><br> <tr><br> <td colspan="20%">132 articles</td><br> <td>Sinorama magazine, Taiwan (1996-1998 & 2000-2001)</td><br> </tr><br> </tbody><br> </table><br> <h3>Data</h3><br> <p>Chinese Treebank 5.0 contains 507,222 words, 824,983 Hanzi, 18,782 sentences, and 890 data files.</p><br> <p>All files are GB encoded. The format of Chinese Treebank 5.0 is the same as the Penn English Treebank. All files have been annotated at least twice. The first pass was done by one annotator, and the resulting files were checked by a second annotator (second pass). Some files were also double-blind annotated and then adjudicated to create gold standard files.</p><br> <p>The corpus provides four versions of files: bracketed, raw, segmented and postagged. The raw, segmented and postagged versions are generated from the bracketed version and so do not reflect the previous annotation stages. The bracketed files are sequentially named as follows: chtb_nnnn.fid, where nnnn is a sequential file number.</p><br> <h3>Samples</h3><br> <p>To see an example of Gold Standard file, please examine this <a href="../../../Catalog/desc/addenda/LDC2005T01_gold.fid" rel="nofollow">sample</a>.</p><br> <h3>Updates</h3><br> <p>The 5.1 update contains corrections to errors found in the earlier version. Specifically, sentences which had more than one top-level node have been modified. Additionally, some GB-encoded white spaces have been converted to ASCII. The 5.1 package is available as an additional download to all those who have licensed CTB5.0.</p></br> Portions © 1994-1998 Xinhua News Agency, © 1996-2001 Sinorama Magazine, © 1997 The Government of the Hong Kong Special Administrative Region, © 2001, 2004, 2005 Trustees of the University of Pennsylvania
<h3>简介</h3><br><p>中文树库5.0(Chinese Treebank 5.0)由语言数据联盟(Linguistic Data Consortium,LDC)制作,编号为LDC2005T01,ISBN为1-58563-323-2。</p><br><p>宾夕法尼亚大学中文树库(Penn Chinese Treebank)是一项始于1998年夏季的持续推进项目,其目标是构建一个包含50万字、带有句法括号标注的中文文本语料库。中文树库1.0于2000年首次发布,后经修订完善,于2001年以<a href="http://catalog.ldc.upenn.edu/LDC2001T11" rel="nofollow">中文树库2.0</a>的名称正式发布。2004年,另一项更新版本以<a href="http://catalog.ldc.upenn.edu/LDC2004T05" rel="nofollow">中文树库4.0</a>的形式发布。该项目的更多详情可访问<a href="https://www.cs.brandeis.edu/~llc/page2/page2.html">中文树库</a>官方网站。</p><br><p>本语料库的内容源自以下新闻类数据源:</p><br><table><br><tbody><br><tr><br><td colspan="20%">698篇文章</td><br><td>新华社(1994-1998年)</td><br></tr><br><tr><br><td colspan="20%">55篇文章</td><br><td>香港特别行政区政府新闻处(1997年)</td><br></tr><br><tr><br><td colspan="20%">132篇文章</td><br><td>台湾《Sinorama》杂志(1996-1998年及2000-2001年)</td><br></tr><br></tbody><br></table><br><h3>数据概况</h3><br><p>中文树库5.0共包含507222个词、824983个汉字、18782个句子以及890个数据文件。</p><br><p>所有文件均采用GB编码格式,其格式与宾夕法尼亚大学英文树库(Penn English Treebank)保持一致。所有文件均经过至少两轮标注:第一轮由一名标注人员完成,产出的文件经第二名标注人员复核(第二轮标注);部分文件还采用双盲标注方式,之后经审定以生成金标准文件。</p><br><p>本语料库提供四类文件版本:带句法括号标注版、原始文本版、分词版以及词性标注版。其中原始文本版、分词版和词性标注版均由带句法括号标注版生成,无法体现此前的标注阶段。带句法括号标注版的文件按顺序命名为`chtb_nnnn.fid`,其中`nnnn`为连续的文件编号。</p><br><h3>示例</h3><br><p>如需查看金标准文件示例,请访问<a href="../../../Catalog/desc/addenda/LDC2005T01_gold.fid" rel="nofollow">此示例</a>。</p><br><h3>更新说明</h3><br><p>5.1版本更新针对早期版本中的错误进行了修正:具体而言,修正了存在多个顶层节点的句子;此外,将部分GB编码的空白字符转换为ASCII字符。所有已授权获取中文树库5.0的用户,均可额外下载获取5.1版本更新包。</p><br><p>部分内容 © 1994-1998 新华通讯社,© 1996-2001 《Sinorama》杂志,© 1997 香港特别行政区政府,© 2001、2004、2005 宾夕法尼亚大学校董会</p>




