Dataset for "How Software Engineering Operationalizes Open Science: An Ecosystem-Level Multiple-Case Study of Mature Free and Open Source Software Projects"
收藏资源简介:
# Empirical Data Package This repository contains the empirical datasets supporting the findingsreported in the article **"How Software Engineering Operationalizes OpenScience: An Ecosystem-Level Multiple-Case Study of Mature Free and OpenSource Software Projects."** The datasets document the projects investigated in the multiple-casestudy, the software engineering practices considered in the analysis,the evidence extracted from official project sources, and the knowledgebase connecting observed practices to artifacts, openness mechanisms,and Open Science dimensions. ## Files ### `01_Projects.xlsx` Contains the characterization of the ten mature Free and Open SourceSoftware (FOSS) projects included in the study. The dataset includes project identifiers, software domains, primarytechnologies, licenses, governance or sponsorship models, officialwebsites, repositories, primary development documentation, and selectionstatus. The ten projects constitute the cases investigated in the study. ### `02_Practices.xlsx` Contains the catalog of software engineering practices used in theempirical analysis. Each practice has a unique identifier and includes its operationaldefinition, purpose, accepted forms of evidence, primary artifacts,typical openness mechanisms, inclusion and exclusion criteria,borderline cases, examples, counterexamples, and reference basis. This dataset provides the coding framework used to identify and classifypractices in the investigated projects. ### `03_Evidence.xlsx` Contains the documentary evidence extracted from official sources of theten FOSS projects. The `Evidence` worksheet contains the final evidence retained aftersemantic auditing. Each evidence item is associated with a project and asoftware engineering practice and includes the source type, source URL,access date, evidence summary, observed artifact, confidence level, andpresence status. The `Semantic_Audit_Log` worksheet records the decisions made during thesemantic audit, including evidence that was kept, reclassified, orremoved and the justification for each decision. The `Semantic_Audit_Summary` worksheet summarizes the auditing processand the resulting evidence dataset. This file constitutes the primary empirical evidence supporting thefindings reported in the article. ### `04_Knowledge_Base.xlsx` Contains the empirically instantiated relationships derived from theobserved software engineering practices. Each relationship connects a practice to an artifact type, an opennessmechanism, and an Open Science dimension. The dataset also recordsassociation strength, justification, and empirical status. This knowledge base supports the analysis of how individual engineeringpractices contribute to broader Open Science capabilities and how theseelements form an interconnected ecosystem. ## Traceability Unique identifiers are used consistently across the datasets: - `PRxx` identifies projects.- Practice identifiers such as `GOV-01`, `CON-03`, and `IMP-01` identify software engineering practices.- Evidence identifiers such as `PR01-EV001` identify individual documentary evidence items.- `KBxxxx` identifies relationships represented in the knowledge base. These identifiers enable traceability across the empirical data. Inparticular, findings can be traced from Open Science dimensions andopenness mechanisms in the knowledge base to individual softwareengineering practices and from those practices to documentary evidenceobtained from the investigated FOSS projects. ## Data Sources The empirical evidence was collected from publicly accessible officialproject sources, including project websites, repositories, governancedocumentation, contribution guidelines, development documentation,issue-tracking infrastructure, release documentation, and other officialproject resources. Source URLs and access dates are provided in the evidence dataset tosupport inspection of the original sources. ## Intended Use These datasets are provided to support transparency, verification,replication, and further research. They can be used to inspect theempirical basis of the findings reported in the article, reproducedescriptive analyses, investigate individual projects or practices, andconduct additional analyses of the relationships between SoftwareEngineering and Open Science. ## Citation When using these datasets, please cite both the associated article andthe archived dataset record. The definitive bibliographic informationand DOI should be obtained from the corresponding Zenodo record.



