Software Requirements Classification: From Bag-of-Words to Transformer [Dataset]
收藏资源简介:
We constructed a new dataset by merging data from other software requirement datasets. It consists of five different datasets that contain Functional Requirements (FRs) and Non-Functional Requirements (NFRs), with additional labels indicating the category in which the NFRs belong (i.e., their Specific Type). This makes the final dataset suitable both for binary and for multi-class requirement classification tasks. In case that labels where missing, manual labeling by the authors was conducted.The following datasets that consist it are: PROMISE: A software engineering repository publicly available for building predictive software models. PURE: A collection of public requirements retrieved from Web. IoTAC: A dataset that includes well-defined and well-structured security requirements. Kaggle: A publicly available dataset retrieved from Kaggle. SecReq: An open-source dataset that contains security and non-security requirements from three open-source projects. In this dataset, there was a lack of non-security requirements labels. Therefore, a manual labeling process was performed by the authors identifying whether they are FR or NFR, as well as their specific category in case of NFR



