CoAM: Corpus of All-Type Multiword Expressions
收藏资源简介:
CoAM数据集是由奈良先端科学技术大学院大学和Resolve Research共同创建的多词表达(MWE)识别数据集,包含1300个句子。该数据集旨在解决现有MWE识别数据集标注不一致、类型单一或规模有限的问题。数据集通过多步骤构建过程,包括人工标注、人工审查和自动化一致性检查,确保数据质量。数据集中的MWE被标记为不同类型(如名词、动词等),以便进行细粒度的错误分析。数据集的应用领域包括机器翻译和词汇复杂性评估等自然语言处理任务,旨在提高MWE识别的准确性和可靠性。
The CoAM dataset is a multi-word expression (MWE) recognition dataset jointly created by Nara Institute of Science and Technology (NAIST) and Resolve Research, consisting of 1,300 sentences. This dataset is designed to address the shortcomings of existing MWE recognition datasets, including inconsistent annotations, single category coverage, and limited scale. The dataset adopts a multi-step construction process that includes manual annotation, manual review, and automated consistency checks to ensure high data quality. MWEs in the dataset are labeled with various categories such as nouns, verbs, etc., to enable fine-grained error analysis. The dataset can be applied to natural language processing tasks like machine translation and lexical complexity assessment, aiming to improve the accuracy and reliability of MWE recognition.




