遇见数据集

Replication package

收藏
Zenodo2026-09-26 更新2026-10-01 收录
官方服务:

资源简介:

How Context Shapes LLM-Based and Agentic Detection of Design and Architecture Smells Abstract Design and architecture smells degrade software maintainability, yet detecting them with Large Language Models (LLMs) remains challenging because relevant context spans multiple components. This replication package supports a study on context construction for LLM-based smell detection by comparing context representation, source code versus metrics and dependencies, and context acquisition, direct invocation versus agentic exploration. The study evaluates four LLMs on four design and architecture smells: Insufficient Modularization, Hub-like Modularization, God Component, and Unstable Dependency. The evaluation uses 160 manually validated instances from ten open-source Java projects. The package includes data, scripts, prompts, generated outputs, human evaluation material, and result files used to reproduce the quantitative and qualitative analyses reported in the paper. Repository Structure The replication package is organized as follows: detecting_smells_study/ ├── Architecture-Smell-Detection-using-Agents/ │ └── Mini-swe-agent-based smell detection workflow. │ ├── codex/ │ └── Codex-based agentic smell detection workflow. │ ├── human_evaluation/ │ └── Human evaluation material and reference labels. │ ├── metrics_deps/ │ └── LLM-based smell detection using metrics and dependencies. │ ├── raw_code/ │ └── LLM-based smell detection using raw source code. │ ├── results/ │ ├── rq1_rq2/ │ │ └── Quantitative results organized by smell, model, and strategy. │ │ │ └── rq3/ │ └── Qualitative analysis files, including RQ3_Analysis.xlsx. │ ├── clean_generated_data.sh ├── compute_metrics_deps_results.py ├── compute_minisweagent_results.py ├── compute_raw_code_results.py ├── requirements.txt └── README.md Requirements We recommend using a Python virtual environment before running the scripts: python3 -m venv venv source venv/bin/activate pip install -r requirements.txt Some workflows require external model-provider credentials. API keys are not included in the replication package. Configure the required environment variables according to the provider used in each experiment. Running the LLM-Based Strategies Raw-Code Strategy The raw-code strategy provides the model with the source code of the evaluated classes or packages. cd raw_code python main.py Metrics-and-Dependencies Strategy The metrics-and-dependencies strategy provides the model with DesigniteJava-generated metrics and dependency information. cd metrics_deps python main.py Running the Mini-swe-agent Strategy The Mini-swe-agent workflow is located in: Architecture-Smell-Detection-using-Agents/ This strategy gives the agent the smell detection task, the target class or package, and access to the repository. The agent then inspects the repository to acquire the context needed for detection. A typical execution follows the format below: cd Architecture-Smell-Detection-using-Agents python score.py --smell <smell> --model <model> --repetitions 1 Example: python score.py --smell unstable_dependency --model deepseek/deepseek-v3.2 --repetitions 1 The generated outputs can be consolidated using: cd .. python compute_minisweagent_results.py Running the Codex-Based Agentic Strategy The Codex-based workflow is located in: codex/ The main entry point receives a YAML configuration file through the --config argument. cd codex python scripts/run_experiment.py --config path/to/config.yaml Example configuration: model: kimik3 smell: Insufficient Modularization smell_definition: "when a class concentrates an **excessive** number of responsibilities, resulting in a large or complex implementation and an interface that is difficult to understand, use, or evolve." repo_path: data/repositories context: metrics After execution, the script prints a JSON summary containing the target element, smell, validity status, verdict path, Codex return code, and paths to the generated prompt, stdout, and stderr files. Computing Consolidated Results The repository includes scripts to compute consolidated result files for each strategy: python compute_raw_code_results.py python compute_metrics_deps_results.py python compute_minisweagent_results.py The consolidated outputs are stored under: results/rq1_rq2/ The qualitative analysis material used for RQ3 is stored under: results/rq3/ Results Organization The quantitative results are organized by smell, model, and strategy. For example: results/rq1_rq2/ ├── god_component/ │ ├── deepseek/ │ │ ├── metrics_deps/ │ │ │ └── results.json │ │ ├── mini-swe-agent/ │ │ └── raw_code/ │ ├── gpt/ │ ├── kimi-k3/ │ └── qwen/ │ ├── hublike_modularization/ ├── insufficient_modularization/ └── unstable_dependency/ The qualitative analysis files are organized under: results/rq3/ └── RQ3_Analysis.xlsx

提供机构:
Zenodo
创建时间:
2026-09-26
二维码
社区交流群
二维码
科研交流群
商业服务