RequiBERT-30M: Unlabeled Software Requirements Dataset for Pretraining
收藏资源简介:
{ "title": "RequiBERT-30M: Unlabeled Software Requirements Dataset for Pretraining", "description": "This dataset contains approximately 30 million unlabeled software requirement sentences, created by applying oversampling techniques on an original set of 23,558 functional and non-functional requirements. It is intended for unsupervised pretraining of domain-specific language models in software engineering, such as the RequiBERT model proposed in our research. The dataset is provided as multiple .txt files with one requirement per line.", "creators": [ { "name": "Kiramat Rahman", "affiliation": "University of Swat, Charbagh, Khyber Pakhtunkhwa, Pakistan" } ], "license": "CC-BY-4.0", "keywords": [ "RequiBERT", "requirements engineering", "functional requirements", "non-functional requirements", "software engineering", "unsupervised learning", "language model", "dataset" ], "upload_type": "dataset", "publication_date": "2025-06-03"}



