遇见数据集

RequiBERT-30M: Unlabeled Software Requirements Dataset for Pretraining

收藏
Zenodo2025-06-13 更新2026-05-26 收录
官方服务:

资源简介:

{ "title": "RequiBERT-30M: Unlabeled Software Requirements Dataset for Pretraining", "description": "This dataset contains approximately 30 million unlabeled software requirement sentences, created by applying oversampling techniques on an original set of 23,558 functional and non-functional requirements. It is intended for unsupervised pretraining of domain-specific language models in software engineering, such as the RequiBERT model proposed in our research. The dataset is provided as multiple .txt files with one requirement per line.", "creators": [ { "name": "Kiramat Rahman", "affiliation": "University of Swat, Charbagh, Khyber Pakhtunkhwa, Pakistan" } ], "license": "CC-BY-4.0", "keywords": [ "RequiBERT", "requirements engineering", "functional requirements", "non-functional requirements", "software engineering", "unsupervised learning", "language model", "dataset" ], "upload_type": "dataset", "publication_date": "2025-06-03"}

提供机构:
Zenodo
创建时间:
2025-06-13
二维码
社区交流群
二维码
科研交流群
商业服务