five

Dataset of "Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents"

收藏
NIAID Data Ecosystem2026-03-12 收录
下载链接:
https://zenodo.org/record/5167247
下载链接
链接失效反馈
官方服务:
资源简介:
This is the dataset used in the paper "Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents" at AIRE'21   In this paper, we explore the use of a multiword expression detection in combination with a knowledge-based word sense disambiguation to disambiguate expressions in requirements documents. The dataset comprises a gold standard for multiword expression detection and sense disambiguation for Wikipedia and WordNet 3.1. It covers 18 projects: CM1, EBT and GANTT as well as the 15 projects of the NFR dataset.   File format We use a tab-separated version of the DiMSUM file format and extended it with sense information. The nine original DiMSUM tab-separated columns: 1. token offset 2. word 3. lowercase lemma 4. POS 5. MWE tag 6. offset of parent token (i.e. previous token in the same MWE), if applicable 7. strength level encoded in the tag, if applicable. Currently not used 8. supersense label, Currently not used 9. sentence ID   and the two further columns for sense information: 10. Wikipedia article name 11. WordNet 3.1 synset   The last two columns might end with .1 or .0 indicating that the sense is a fully applicable or partial sense of a multiword expression. Attribution (of datasets used) The NFR Dataset can be attributed to Jane Cleland-Huang. Jane Cleland-Huang, Sepideh Mazrouee, Huang Liguo, & Dan Port. (2007). nfr [Data set]. Zenodo. Available: http://doi.org/10.5281/zenodo.268542   The CM1, EBT and GANTT datasets were retrieved from the Center of Excellence for Software & Systems Traceability (CoEST) http://coest.org/
创建时间:
2021-08-06
5,000+
优质数据集
54 个
任务类型
进入经典数据集
二维码
社区交流群

面向社区/商业的数据集话题

二维码
科研交流群

面向高校/科研机构的开源数据集话题

数据驱动未来

携手共赢发展

商业合作