遇见数据集

Dataset of Paper Titled: A Semi-Automated Approach for Detecting Ambiguities in Software Requirements Using SpanBERT and Named Entity Recognition

收藏
Zenodo2025-08-11 更新2026-05-26 收录
官方服务:

资源简介:

This dataset supports the study titled "A Semi-Automated Approach for Detecting Ambiguities in Functional Requirements Using SpanBERT." The dataset consists of 425 original functional requirements collected from 16 diverse software domains, such as finance, healthcare, education, and e-commerce. These requirements were curated to represent realistic and domain-relevant software specification statements written in natural language. Each requirement in the dataset has been manually annotated for three types of linguistic ambiguities: Anaphoric ambiguity (e.g., unclear references like "it" or "they"), Coordination ambiguity (e.g., ambiguous use of conjunctions like "and" or "or"), Missing condition ambiguity (e.g., implied conditions not explicitly stated). These annotations were used to train and evaluate a semi-automated ambiguity detection approach based on SpanBERT, a transformer-based natural language processing model. This dataset is intended to support further research on automated ambiguity detection, requirements engineering, and natural language understanding in software engineering. It is especially valuable for researchers and practitioners aiming to improve the clarity of software requirement specifications and reduce ambiguity-related issues during development. How to cite F. Talha, T. Tahir, and T. Nadeem, “ A Semiautomated Approach for Detecting Ambiguities in Software Requirements Using SpanBERT and Named Entity Recognition,” Journal of Software: Evolution and Process 37, no. 8 (2025): e70041, https://doi.org/10.1002/smr.70041.

提供机构:
Zenodo
创建时间:
2025-07-31
二维码
社区交流群
二维码
科研交流群
商业服务