Potential Idiomatic Expression (PIE)-English
收藏资源简介:
本研究介绍了名为‘Potential Idiomatic Expression (PIE)-English’的数据集,由卢利奥理工大学创建,旨在为英语自然语言处理提供一个大规模的习语数据集。该数据集包含超过20,174条样本,涵盖近1,200个习语案例,分为10个类别,如隐喻、明喻等。数据主要来源于英国国家语料库和英国网页语料库。通过手动提取和标注,确保了数据的高质量和准确性。该数据集适用于机器翻译、词义消歧等NLP任务,有助于提升对话系统和信息检索的性能。
This study presents a dataset titled 'Potential Idiomatic Expression (PIE)-English', developed by Luleå University of Technology. This dataset aims to supply a large-scale idiom resource for English natural language processing (NLP). It encompasses more than 20,174 samples, covering nearly 1,200 idiom instances, and is classified into 10 categories including metaphor, simile, and others. The primary sources of the dataset are the British National Corpus and the UK Web Corpus. Manual extraction and annotation procedures are implemented to guarantee the high quality and accuracy of the collected data. This dataset is applicable to multiple NLP tasks such as machine translation, word sense disambiguation, and others, and can facilitate the performance enhancement of dialogue systems and information retrieval systems.
PIE-English: Corpus for Classes of Idioms
引用信息
- 标题: Potential Idiomatic Expression (PIE)-English: Corpus for Classes of Idioms
- 作者: Adewumi, Tosin and Vadoodi, Roshanak and Tripathy, Aparajia and Nikolaidou, Konstantina and Liwicki, Foteini and Liwicki, Marcus
- 会议: Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022)
- 时间: June 2022
- 地点: Marseille, France
- 出版商: European Language Resources Association (ELRA)
- URL: http://www.lrec-conf.org/proceedings/lrec2022/pdf/2022.lrec-1.72.pdf

- 1Potential Idiomatic Expression (PIE)-English: Corpus for Classes of Idioms卢利奥理工大学 · 2022年



