CLSE
收藏资源简介:
CLSE数据集是由谷歌创建的,旨在支持自然语言生成(NLG)系统的开发和评估。该数据集包含34种语言和74种不同的语义类型,覆盖从航空订票到视频游戏等多种应用。数据集通过专家语言学家的标注,确保了语言属性的准确性和多样性。CLSE数据集特别适用于测试NLG系统的语言鲁棒性,帮助解决在处理命名实体时常见的语法错误问题。
The CLSE dataset was created by Google to support the development and evaluation of natural language generation (NLG) systems. It covers 34 languages and 74 distinct semantic types, spanning a wide range of applications from air travel booking to video games. The dataset is annotated by professional linguists to ensure the accuracy and diversity of its linguistic attributes. The CLSE dataset is particularly suitable for testing the linguistic robustness of NLG systems, and helps resolve common grammatical errors that occur when handling named entities.
CLSE: Corpus of Linguistically Significant Entities
描述
CLSE(语言学重要实体语料库)是一个由语言学专家注释的命名实体数据集。该数据集包含34种语言,涵盖74种不同的语义类型,支持从航空订票到视频游戏等多种应用。该语料库的目的是促进创建更多语言多样性的自然语言生成(NLG)数据集。
许可证
本仓库的内容根据CC-BY许可证进行许可。
论文
在使用此数据集时,请引用以下论文:
@inproceedings{clse2022, title={CLSE: Corpus of Linguistically Significant Entities}, author={Chuklin, Aleksandr and Zhao, Justin and Kale, Mihir}, booktitle={Proceedings of the 2nd Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2022) at EMNLP 2022}, year={2022} }

- 1CLSE: Corpus of Linguistically Significant Entities谷歌 · 2023年



