Defect Prediction Tool Validation Dataset 2
收藏资源简介:
<strong>This dataset is used to address the Research Questions in the study at Transactions on Software Engineering</strong>: <strong>Within-Project</strong> <strong>Defect Prediction of Infrastructure-as-Code using Product and Process Metrics. </strong> <strong>See also: https://github.com/stefanodallapalma/TSE-2020-05-0217.</strong> It provides * <strong>repositories.json</strong> - a list of repositories selected from open-source GitHub repositories based on the Ansible language. * <strong>fixing-commits.json</strong> - a list of defect-fixing commits extracted from those repositories. * <strong>fixed-files.json</strong> - a list of Ansible files fixed in those defect-fixing commits and respective bug-inducing commits. * <strong>failure-prone-files.json</strong> - a list of failure-prone files through the repository's commit history. * <strong>metrics.zip </strong>- csv files consisting of releases (set of files) and their IaC-oriented, delta and process metrics extracted from each analyzed repository * <strong>projects.zip </strong>- for each analyzed project, it contains the data (models, performance, and results of Recursive Feature Elimination) used to answer the Research Questions. <strong>Context</strong> <em>Infrastructure-as-code (IaC)</em> is the DevOps strategy that allows management and provisioning of infrastructure through the definition of machine-readable files and automation around them, rather than physical hardware configuration or interactive configuration tools. On the one hand, although IaC represents an ever-increasing widely adopted practice nowadays, still little is known concerning how to best maintain, speedily evolve, and continuously improve the code behind the IaC strategy in a measurable fashion. <br> On the other hand, source code measurements are often computed and analyzed to evaluate the different quality aspects of the software developed.<br> In particular, Infrastructure-as-Code is simply "code", as such it is prone to defects as any other programming languages. This dataset targets the YAML-based Ansible language to devise <strong>within-project defects prediction</strong> approaches for IaC based on Machine-learning. <strong>Content</strong> The dataset contains metrics extracted from 85 open-source GitHub repositories based on the Ansible language that satisfied the following criteria: * The repository has at least one push event to its master branch in the last six months;<br> * The repository has at least 2 releases;<br> * At least 10% of the files in the repository are IaC scripts;<br> * The repository has at least 2 core contributors;<br> * The repository has evidence of continuous integration practice, such as the presence of a .travis.yaml file;<br> * The repository has a comments ratio of at least 0.1%;<br> * The repository has commit frequency of at least 2 per month on average;<br> * The repository has an issue frequency of at least 0.01 events per month on average;<br> * The repository has evidence of a license, such as the presence of a LICENSE.md file<br> * The repository has at least 100 source lines of code. Metrics are grouped into three categories: * <strong>IaC-Oriented:</strong> metrics of structural properties derived from the source code of infrastructure scripts. Click [here](https://www.sciencedirect.com/science/article/pii/S0164121220301618) for more info. * <strong>Delta</strong>: metrics that capture the amount of change in a file between two successive releases, collected for each IaC-oriented metric. * <strong>Process</strong>: metrics that capture aspects of the development process rather than aspects about the code itself. Description of the process metrics in this dataset can be found [here](https://pydriller.readthedocs.io/en/latest/processmetrics.html). In addition to the metrics, the dataset contains the pre-trained models (*.joblib) in the folders rq1 and rq2 of projects.zip. You can load the model in Python as follows: ```<br> from joblib import load<br> model = load('projects/owner/repository/rq1/random_forest.joblib'), mmap_mode='r') best_estimator = model['estimator'] # The estimator that maximized the AUC-PR cv_results = model['cv_results'] # The results of each step of the validation procedure best_index = mode['best_index'] # The index to access the best cv_results<br> ``` <strong>Acknowledgements</strong> This work is supported by the European Commission grants no. 825040 (RADON H2020). <br> <strong>Inspiration</strong> What source code properties and properties about the development process are good predictors of defects in Infrastructure-as-Code scripts?



