遇见数据集

PURE: a Dataset of Public Requirements Documents

收藏
Zenodo2024-12-19 更新2026-05-25 收录
官方服务:

资源简介:

Please cite this dataset as Ferrari, A., Spagnolo, G. O., & Gnesi, S. (2017, September). PURE: A dataset of public requirements documents. In 2017 IEEE 25th International Requirements Engineering Conference (RE) (pp. 502-505). IEEE. https://ieeexplore.ieee.org/abstract/document/8049173 This dataset presents PURE (PUblic REquirements dataset), a dataset of 79 publicly available natural language requirements documents collected from the Web. The dataset includes 34,268 sentences and can be used for natural language processing tasks that are typical in requirements engineering, such as model synthesis, abstraction identification and document structure assessment. It can be further annotated to work as a benchmark for other tasks, such as ambiguity detection, requirements categorisation and identification of equivalent re-quirements. In the associated paper, we present the dataset and we compare its language with generic English texts, showing the peculiarities of the requirements jargon, made of a restricted vocabulary of domain-specific acronyms and words, and long sentences. We also present the common XML format to which we have manually ported a subset of the documents, with the goal of facilitating replication of NLP experiments. The XML documents are also available for download. The paper associated to the dataset can be found here: https://ieeexplore.ieee.org/document/8049173/ More info about the dataset is available here: http://nlreqdataset.isti.cnr.it Preprint of the paper available at ResearchGate: https://goo.gl/HxJD7X The dataset includes: - all the documents in PDF format - a subset of 19 documents in XML format - the .xsd schema of the XML files The dataset has been created by gathering data from web sources and we are not aware of license agreements or intellectual property rights on the requirements. The curator took utmost diligence in minimizing the risks of copyright infringement by using non-recent data that is less likely to be critical, by sampling a subset of the original requirements collection, and by qualitatively analyzing the requirements. In case of copyright infringement, please contact the dataset curator (Alessio Ferrari, alessio.ferrari@cnr.it, alessio.ferrari@ucd.ie) to discuss the possibility of removal of that dataset [see Zenodo's policies].

请按以下格式引用本数据集:Ferrari, A., Spagnolo, G. O. 与 Gnesi, S. (2017年9月)。PURE:公开需求文档数据集。收录于2017年第25届IEEE国际需求工程会议(RE)(第502-505页)。IEEE出版社。 https://ieeexplore.ieee.org/abstract/document/8049173 本数据集为PURE(Public REquirements dataset,公开需求数据集),包含从互联网采集的79份公开自然语言需求文档,共计34268条语句。该数据集可用于需求工程领域典型的自然语言处理(Natural Language Processing, NLP)任务,例如模型合成、抽象识别与文档结构评估。此外,其可通过进一步标注,作为歧义检测、需求分类以及等价需求识别等任务的基准数据集使用。在配套学术论文中,我们介绍了该数据集,并将其语言特征与通用英文文本进行对比,揭示了需求行话的独特性:这类行话由领域专属的缩略语与词汇构成,且语句普遍较长。我们还提供了统一XML格式的文档子集(由我们手动转换自原始文档),旨在降低NLP实验的复现难度,相关XML文档亦可下载获取。 本数据集的配套学术论文可通过以下链接获取:https://ieeexplore.ieee.org/document/8049173/ 更多数据集相关信息可访问以下链接:http://nlreqdataset.isti.cnr.it 论文预印本可在ResearchGate平台获取:https://goo.gl/HxJD7X 本数据集包含以下内容: - 全部PDF格式的原始文档 - 19份XML格式的文档子集 - XML文件的.xsd模式文件 本数据集通过采集网络公开资源构建,我们未核实相关需求文档的许可协议与知识产权归属情况。数据集管理者已尽最大努力降低版权侵权风险:选用时效性较弱、侵权可能性更低的非近期数据,对原始需求集进行采样,并对需求内容开展定性分析。若存在版权侵权问题,请联系数据集管理者(Alessio Ferrari,邮箱:alessio.ferrari@cnr.it、alessio.ferrari@ucd.ie),我们将依据Zenodo政策协商移除对应数据集的相关事宜。

提供机构:
Zenodo
创建时间:
2018-09-12
二维码
社区交流群
二维码
科研交流群
商业服务