Baidu Chinese Treebank (DuCTB)
收藏资源简介:
Baidu Chinese Treebank (DuCTB)是由百度创建的大规模中文依存句法分析数据集,包含约一百万个标注句子,来源于搜索日志、新闻电讯、论坛讨论及对话程序等多种数据源。数据集的创建旨在覆盖尽可能多的表达方式,包括常规和非常规句子,以满足工业应用的需求。DuCTB专注于分析句子中的语法结构而非语义,其标注指南旨在让普通用户易于理解。该数据集的应用领域广泛,特别是在自然语言处理任务中,如依存句法分析,旨在提高分析的准确性和效率。
Baidu Chinese Treebank (DuCTB) is a large-scale Chinese dependency parsing dataset created by Baidu. It contains approximately one million annotated sentences, sourced from multiple data sources including search logs, news articles, forum discussions, and dialogue systems. The dataset was developed to cover as many linguistic expressions as possible, including both conventional and unconventional sentences, to meet the demands of industrial applications. DuCTB focuses on analyzing the grammatical structure of sentences rather than their semantic meaning, and its annotation guidelines are designed to be easily understandable for ordinary users. This dataset has a wide range of application scenarios, especially in natural language processing tasks such as dependency parsing, aiming to improve the accuracy and efficiency of parsing analysis.




