Replication Package for Architectural Degradation: How to Measure and to Remediate
收藏资源简介:
Replication Package The package contains the datasets, source code, validation results, inter-rater agreement results, and generated outputs used throughout the study "Architectural Degradation: How to Measure and to Remediate". It helps to reproduce the literature screening, full-text extraction, LLM-based analysis, validation, consolidation of the final categories, agreement analysis, and the analyses reported in the paper. Note: All hyperlinks referring to files in this replication package are intended to be used after downloading the package. Local file links may not work when browsing the package online. License All generated data is provided under DATA_LICENSE Creative Commons 4.0 Attribution License. All scripts are provided under SCRIPT_LICENSE MIT License. Package Structure Approaches for Architectural degradation/ │ ├── README.md ├── INSTALL.md ├── requirements.txt ├── DATA_LICENSE ├── SCRIPT_LICENSE ├── Final Full read and analysis (architecture approaches).xlsx ├── Validation Approaches paper.xlsx │ ├── Code and outputs/ │ ├── Title Abstract Round Code and output/ │ ├── T-A Validation code and Outputs/ │ ├── Full Read Code/ │ ├── Full Read Round Validation Outputs/ │ ├── Final round code and outputs/ │ └── Final Round Validation code and outputs/ │ ├── Inter Agreement(Gwet AC & Cohen K)/ └── Code for Figures/ Data Files 1. Final Full read and analysis (architecture approaches).xlsx This workbook contains the study process, screening results, extracted full-text data, final coded datasets, and analysis results used in the study. Workbook Contents Sheet Name Description Contents Provides an overview of the workbook contents. Process Documents the study workflow and the procedures followed for screening, extraction, coding, and analysis. White (titleabstract) Contains the title and abstract screening results for the white literature (peer-reviewed studies). Grey(titleabstract) Contains the title and abstract screening results for the grey literature. Full read (Accepted) Contains the extracted information from the studies accepted for full-text analysis. Tools Contains the final consolidated tools identified for measuring architectural degradation. Approaches to Measure Contains the final consolidated approaches identified for measuring architectural degradation. Metrics Contains the final consolidated metrics identified for measuring architectural degradation. Approches to remediate Contains the final consolidated approaches identified for remediating architectural degradation. 2. Validation Approaches paper.xlsx This workbook contains the validation datasets and summary statistics used to evaluate the reliability and consistency of the screening, extraction, and final coding stages. Workbook Contents Sheet Name Description Contents Provides an overview of the validation workbook. Stats Contains summary statistics and validation results used in the study. Title abstract (validation) Contains the validation data for the title and abstract screening stage. Approaches to measure (final round validation) Contains the final-round validation results for approaches to measure architectural degradation. tools (Final round validation) Contains the final-round validation results for tools. metrics (final round validation) Contains the final-round validation results for metrics. Approaches to remediate (final round validation) Contains the final-round validation results for remediation approaches. Approaches to measure (full read) Contains the full-read validation results for approaches to measure architectural degradation. Tools (full read) Contains the full-read validation results for tools. Metrics (full read) Contains the full-read validation results for metrics. Remediation Approches (full read) Contains the full-read validation results for remediation approaches. Code and Outputs The Code and outputs directory contains the source code, intermediate data, validation scripts, and generated outputs used throughout the study. Folder Description Title Abstract Round Code and output Contains the code and outputs used for title and abstract screening of the white and grey literature. T-A Validation code and Outputs Contains the code and outputs used to validate the title and abstract screening results. Full Read Code Contains the code and outputs used for full-text extraction and identification of measurement approaches, tools, metrics, and remediation approaches. Full Read Round Validation Outputs Contains the outputs produced by the independent LLM validators for the full-read extraction stage. Final round code and outputs Contains the multi-round consolidation code, intermediate outputs, and final outputs for approaches to measure, tools, metrics, and remediation approaches. Final Round Validation code and outputs Contains the code and outputs used to validate the final consolidated results. Inter Agreement (Gwet AC & Cohen K) Contains the code and outputs used to calculate Cohen's kappa and Gwet's AC1 agreement statistics. Code for Figures Contains the code/data and generated outputs used to create the figures reported in the paper. Replication of the Results This section describes the procedure for reproducing the title and abstract screening, full-text analysis, LLM-based extraction and validation, final-round consolidation, agreement analysis, and final results reported in the study. The workflow uses Marco-o1 as the primary model and Mistral, Qwen, and Llama as independent validation models. The computationally intensive LLM experiments were originally executed on a supercomputer. Setting Up the Environment Follow the instructions in INSTALL.md to configure the required software environment. The required Python dependencies are listed in requirements.txt. After configuring the environment, execute the replication stages below in the specified order. 1. Title and Abstract Screening Use the scripts and outputs provided in: Code and outputs/Title Abstract Round Code and output/ The scripts process the retrieved white and grey literature and generate the title and abstract screening decisions. The corresponding results are provided in the White (titleabstract) and Grey(titleabstract) sheets of: Final Full read and analysis (architecture approaches).xlsx Studies accepted during screening proceed to validation and full-text analysis. 2. Title and Abstract Validation Use the validation code and outputs provided in: Code and outputs/T-A Validation code and Outputs/ The title and abstract screening results are independently validated using Mistral, Qwen, and Llama. Human validation is performed on the selected validation sample according to the validation procedure described in the paper. The resulting validation data are provided in the Title abstract (validation) sheet of: Validation Approaches paper.xlsx The corresponding agreement analysis is available under: Inter Agreement(Gwet AC & Cohen K)/Title Abstract(GWET AC and Cohen K)/ 3. Full-Text Analysis with Marco-o1 Use the code and outputs provided in: Code and outputs/Full Read Code/ Marco-o1 is used as the primary model to analyze the accepted studies and extract information concerning approaches to measure architectural degradation, tools, metrics, and approaches to remediate architectural degradation. For the original supercomputer execution, the general command structure is: sbatch <SBATCH_SCRIPT> "MODEL_NAME" <PORT> <OUTPUT_FILE> The extracted full-read data are consolidated in the Full read (Accepted) sheet of: Final Full read and analysis (architecture approaches).xlsx Note: Where full-text publications are not distributed with the replication package, researchers must obtain the corresponding papers independently before reproducing PDF-based extraction. 4. Full-Read Validation After the primary-model extraction, the full-read results are independently validated using Mistral, Qwen, and Llama. The generated validation outputs are provided in: Code and outputs/Full Read Round Validation Outputs/ The resulting validation data are consolidated in the corresponding full-read sheets of: Validation Approaches paper.xlsx These include validation for: Approaches to measure architectural degradation Tools Metrics Remediation approaches The corresponding Cohen's kappa and Gwet's AC1 analyses are provided in the relevant folders under: Inter Agreement(Gwet AC & Cohen K)/ 5. Final-Round Analysis After completing the full-read extraction and validation, the extracted information is consolidated through additional analysis rounds. The code, intermediate outputs, and final outputs are provided in: Code and outputs/Final round code and outputs/ The directory contains separate analysis pipelines for: Approaches to Measure Tools Metrics Remediation The scripts progressively consolidate the extracted items into the final categories used in the study. The resulting final datasets are provided in the Tools, Approaches to Measure, Metrics, and Approches to remediate sheets of: Final Full read and analysis (architecture approaches).xlsx 6. Final-Round Validation The final consolidated results are independently validated using Mistral, Qwen, and Llama, together with human validation according to the procedure described in the paper. The validation code and outputs are provided in: Code and outputs/Final Round Validation code and outputs/ Separate validation outputs are provided for approaches to measure, tools, metrics, and remediation approaches. The resulting validation data are consolidated in the corresponding final round validation sheets of: Validation Approaches paper.xlsx These validated results constitute the final coded data used for the study analysis. 7. Inter-Rater Agreement Analysis The scripts and outputs used to calculate agreement between human and LLM validation decisions are provided in: Inter Agreement(Gwet AC & Cohen K)/ The agreement analysis includes Cohen's kappa and Gwet's AC1 for the title/abstract stage, full-read validation, and final-round validation where applicable. The directory contains the input validation datasets together with the generated pairwise agreement and agreement-summary outputs. 8. Data Analysis and Figures The final analysis uses the consolidated datasets from: Final Full read and analysis (architecture approaches).xlsx The corresponding analysis results are provided in the AnalysisResults sheet. The code/data and generated outputs used to create the figures reported in the paper are provided in: Code for Figures/ Local Replication Without a Supercomputer Researchers without access to a SLURM-based supercomputer can execute a smaller-scale version of the LLM pipeline locally using an OpenAI-compatible vLLM endpoint, provided that sufficient computational resources are available. Detailed environment requirements and configuration instructions are provided in INSTALL.md. For example, Marco-o1 can be served using: python -m vllm.entrypoints.openai.api_server \ --model AIDC-AI/Marco-o1 \ --port 8000 \ --tensor-parallel-size 1



