Arabic Treebank: Part 3 (full corpus) v 2.0 (MPG + Syntactic Analysis)
收藏资源简介:
<h3>Introduction</h3> <p> This file contains documentation on the Arabic Treebank: Part 3 (full corpus) v 2.0 (MPG + Syntactic Analysis), Linguistic Data Consortium (LDC) catalog number LDC2005T20 and ISBN 1-58563-341-0. </p><p> The goal of the Arabic Treebank project is to support the development of data-driven approaches to natural language processing (NLP), human language technologies, automatic content extraction (topic extraction and/or grammar extraction), cross-lingual information retrieval, information detection, and other forms of linguistic research on Modern Standard Arabic in general. The LDC was sponsored to develop an Arabic POS and Treebank of 1,000,000 words, and this corpus is part three of that project. In this release, we provide both syntactic (treebank) annotation and annotation on part of speech (POS), gloss, and word segmentation. </p><p> Treebanks are language resources that provide annotations of natural languages at various levels of structure: at the word level, the phrase level, and the sentence level. Treebanks have become crucially important for the development of data-driven approaches to natural language processing (NLP), human language technologies, automatic content extraction (topic extraction and/or grammar extraction), cross-lingual information retrieval, information detection, and other forms of linguistic research in general. </p><p> This corpus is designed for those who study and use languages either professionally or academically, and who need text corpora in their work. The Penn Arabic Treebank is particularly suitable for language developers, computational linguists and computer scientists who are interested in various aspects of natural language processing. </p><p> The Penn Arabic Treebank, which was part of the DARPA TIDES project, started in the Fall of 2001 with the objective of annotating via human intervention and automatically a large Arabic machine-readable text corpus. As in previous Penn Treebanks, two different kinds of information need to be produced by two different (human and computer) processes. The Arabic Treebank project consists therefore of two distinct phases: (a) Part-of-Speech (=POS) tagging, which divides the text into lexical tokens and gives relevant information about each token such as lexical category, inflectional features, and a gloss (referred to as POS for convenience, although it includes morphological and gloss information not traditionally included with part-of-speech annotation), and (b) Arabic Treebanking (=ArabicTB), which characterizes the constituent structures of word sequences, provides categories for each non-terminal node, and identifies null elements, co-reference, traces, etc. </p><p> Both tasks started in November 2001 with an initial pilot consisting of 734 files representing roughly 166K words of written Modern Standard Arabic newswire from the Agence France Presse corpus, which has since been released as "Arabic Treebank: Part 1 v 3.0," LDC Catalog No. LDC2005T02. The second part was released as the 168K-word corpus "Arabic Treebank: Part 2 v 2.0," LDC Catalog No. LDC2004T02. </p><p> The current Arabic Treebank: Part 3 corpus consists of 600 stories from the An Nahar News Agency. This corpus is also referred to as ANNAHAR. The new features include complete vocalization of all Imperfect Verb mood endings: Indicative, Subjunctive, and Jussive. </p><p> The POS only annotation of this ANNAHAR corpus was released in 2004 under the catalog number LDC2004T11 (Arabic Treebank: Part 3 v 1.0). In addition to the treebank annotation, this release (i.e., Arabic Treebank: Part 3 v 2.0) also includes the POS annotation in LDC2004T11. </p><h3>Samples</h3> The POS and treebank samples belowe provide an example the data contained in this corpus <ul> <li> <a href="desc/addenda/LDC2005T20_pos.xml" rel="nofollow">POS</a> </li> <li> <a href="desc/addenda/LDC2005T20_treebank.xml" rel="nofollow">Treebank</a> </li> </ul> </br> Portions © 2002 An Nahar, © 2003, 2004, 2005 Trustees of the University of Pennsylvania
<h3>引言</h3> <p>本文件包含阿拉伯语树库第三部分(完整语料库)v 2.0(MPG + 句法分析)的说明文档,该语料库由语言数据联盟(Linguistic Data Consortium, LDC)发行,编号为LDC2005T20,ISBN为1-58563-341-0。</p><p>阿拉伯语树库项目的目标是为面向现代标准阿拉伯语的、基于数据驱动的自然语言处理(Natural Language Processing, NLP)、人类语言技术、自动内容抽取(主题抽取和/或语法抽取)、跨语言信息检索、信息检测及其他各类语言学研究提供支持。语言数据联盟受资助开发包含100万词的阿拉伯语词性标注(Part-of-Speech, POS)树库,本次发布的语料库即该项目的第三部分。本版本同时提供句法(树库)标注以及词性标注、词义标注和分词标注。</p><p>树库是一类在多种结构层级上为自然语言提供标注的语言资源:涵盖词级、短语级与句级。树库对于发展基于数据驱动的自然语言处理、人类语言技术、自动内容抽取、跨语言信息检索、信息检测及其他各类语言学研究而言,已成为至关重要的资源。</p><p>本语料库面向以专业或学术方式研究、使用语言,且在工作中需要文本语料库的人员。宾夕法尼亚阿拉伯语树库尤其适合对自然语言处理各方面感兴趣的语言开发者、计算语言学家与计算机科学家。</p><p>作为DARPA TIDES项目的组成部分,宾夕法尼亚阿拉伯语树库于2001年秋季启动,目标是通过人工干预与自动化手段对大型阿拉伯语机器可读文本语料库进行标注。与此前的宾夕法尼亚树库一致,本次任务需要通过两种不同的(人工与计算机)流程生成两类不同的信息。因此,阿拉伯语树库项目包含两个独立阶段:(a) 词性标注(Part-of-Speech Tagging,即POS标注):将文本切分为词汇Token(Token),并为每个Token提供词汇类别、屈折特征以及词义标注(为方便起见,简称POS标注,尽管其包含传统词性标注未涵盖的形态与词义信息);(b) 阿拉伯语树库标注(Arabic Treebanking,即ArabicTB):刻画词序列的成分结构,为每个非终端节点提供类别,并识别空成分、共指关系、句法踪迹等。</p><p>两项任务均于2001年11月启动,初始试点包含734份文件,对应来自法新社语料库的约16.6万词书面现代标准阿拉伯语新闻文本,该试点语料库后续以《阿拉伯语树库:第一部分v3.0》之名发行,编号为LDC2005T02。第二部分则以16.8万词语料库《阿拉伯语树库:第二部分v2.0》之名发行,编号为LDC2004T02。</p><p>本次发布的阿拉伯语树库第三部分语料库包含来自《今日新闻》(An Nahar)通讯社的600篇新闻报道,该语料库亦被称为ANNAHAR。本次版本新增功能包括对所有未完成动词语气词尾(陈述式、虚拟式与命令式)的完整元音标注。</p><p>该ANNAHAR语料库的仅词性标注版本已于2004年以编号LDC2004T11发行(对应《阿拉伯语树库:第三部分v1.0》)。除树库标注外,本次发布版本(即《阿拉伯语树库:第三部分v2.0》)同时包含LDC2004T11中的词性标注内容。</p><h3>示例</h3> 以下词性标注与树库示例展示了本语料库包含的数据内容 <ul> <li> <a href="desc/addenda/LDC2005T20_pos.xml" rel="nofollow">词性标注</a> </li> <li> <a href="desc/addenda/LDC2005T20_treebank.xml" rel="nofollow">树库标注</a> </li> </ul> </br> 部分内容 © 2002 An Nahar,© 2003、2004、2005 宾夕法尼亚大学董事会




