Replication Package for "Software Architecture Community Analysis. A 25-year retrospective."
收藏资源简介:
This is the Replication package for the paper “Software Architecture Community Analysis. A 25-year retrospective.” We analyzed 2,641 papers published between 2001 and 2025 in five major software architecture conference series — WICSA, ECSA, QoSA, ICSA and CBSE — to trace how the research themes, the leading institutions and countries, and the collaboration structures of the software architecture community have evolved over 25 years. Research topics were extracted from titles and abstracts by a Large Reasoning Model and classified into a taxonomy by a LangChain-based Agentic AI workflow, with three independent validation LLMs and three human experts assessing the results at every stage. Institutional affiliations were standardised through the Research Organization Registry (ROR), and collaboration networks were reconstructed from co-authorship at institution and country level. This package contains everything needed to verify or extend that analysis: the raw bibliographic exports, every intermediate result, the complete pipeline, the analysis notebooks, and the figures of the paper. Licence Scripts are released under the MIT License; data files are released under CC BY 4.0. What is in the package DataCollection/ — raw Scopus exports, one file per conference series. Script/01_LLMs/ — the LLM and Agentic AI pipeline for research topic extraction and two-round classification (Python scripts plus SLURM job scripts, served with vLLM). Script/02_ROR/ — affiliation standardisation against the ROR registry. Script/03_DataAnalysis (figure generation)/ — six Jupyter notebooks that reproduce every figure and table of the paper. Script/04_Citation_Extraction/ — year-by-year citation counts from the Scopus Search API. Results/ — all intermediate and final outputs, numbered in the order they are produced (00 to 09), including the LLM and human validation workbooks and the full taxonomy. Figures/ and ExtendedFigures/ — every figure of the paper, plus additional figures not included in the main text. Gephi/ — Gephi project files for the institutional collaboration networks. README.md and INSTALL.md — full documentation of the package structure, the replication steps, and the setup of each component. How to use it All intermediate outputs are included, so the figures and tables of the paper can be reproduced without a GPU, an HPC account or any API key. After unpacking the archive: pip install -r requirements.txt jupyter notebook "Script/03_DataAnalysis (figure generation)/" Then run the six notebooks in order. Re-running the LLM pipeline instead requires a Linux machine with NVIDIA GPUs and a Hugging Face access token; re-running the citation extraction requires an Elsevier/Scopus API key. Both are documented in INSTALL.md, together with the placeholders that must be filled in before the pipeline scripts can be submitted. Models used Large Reasoning Model (extraction and classification): Marco-o1 (7.6B). Validation model V1: Mistral-NeMo-Instruct-2407 (12.2B). Validation model V2: Qwen3-14B-Base (14.8B). Validation model V3: Llama-3.1-8B-Instruct (8B). All models were served locally with vLLM on the Mahti supercomputer hosted by CSC, Finnish IT Center for Science, with fixed inference parameters and a fixed random seed. LLM inference on GPU is not bit-reproducible, so small differences are expected when the pipeline is rerun; this is discussed in the threats to validity of the paper.



