MACRONYM
收藏资源简介:
MACRONYM是一个大规模的多语言和多领域缩略语提取数据集,由俄勒冈大学计算机与信息科学系创建。该数据集包含27,200条来自6种不同语言(英语、丹麦语、西班牙语、法语、波斯语和越南语)和2个领域(法律和科学)的句子。数据集的创建过程涉及从联合国平行语料库和Europarl语料库中收集数据,并通过众包平台招募母语者进行标注。MACRONYM数据集旨在解决多语言和多领域文本处理中缩略语识别的问题,支持如问答和机器翻译等应用。
MACRONYM is a large-scale multilingual and multi-domain acronym extraction dataset created by the Department of Computer and Information Science at the University of Oregon. This dataset contains 27,200 sentences across 6 distinct languages (English, Danish, Spanish, French, Persian and Vietnamese) and 2 domains (law and science). The dataset was constructed by collecting data from the United Nations Parallel Corpus and the Europarl Corpus, with native speakers recruited via crowdsourcing platforms for annotation. The MACRONYM dataset aims to address the problem of acronym recognition in multilingual and multi-domain text processing, supporting applications such as question answering and machine translation.




