CantoNLU
收藏资源简介:
CantoNLU 是一个针对粤语自然语言理解(NLU)的基准数据集,由多伦多大学和安大略科技大学的研究团队创建。该数据集涵盖了七个任务,包括词义消歧、语言可接受性判断、语言检测、自然语言推理、情感分析、词性标注和依存句法分析。数据集由手动编译的词义消歧数据集、从错误跨度数据集改编的语言可接受性判断数据集以及从并行语料库中构建的语言检测数据集组成。CantoNLU旨在解决粤语语言处理领域缺乏评估框架的问题,并促进未来粤语自然语言处理研究的发展。
CantoNLU is a benchmark dataset for Cantonese Natural Language Understanding (NLU), developed by a research team from the University of Toronto and Ontario Tech University. This dataset covers seven tasks, including Word Sense Disambiguation (WSD), Linguistic Acceptability Judgment, Language Detection, Natural Language Inference (NLI), Sentiment Analysis, Part-of-Speech (POS) Tagging, and Dependency Parsing. It consists of three components: a manually compiled WSD dataset, a Linguistic Acceptability Judgment dataset adapted from the Error Span dataset, and a Language Detection dataset constructed from parallel corpora. CantoNLU aims to address the shortage of evaluation frameworks in the field of Cantonese language processing and facilitate the advancement of future Cantonese natural language processing research.




