Replication Package: Can Small LLMs Detect Defect-Prone Code Smells? An Empirical Evaluation of 8B--30B Models Against a Metric-Augmented SZZ Dataset
收藏官方服务:
资源简介:
Replication package for the paper "Can Small LLMs Detect Defect-Prone Code Smells? An Empirical Evaluation of 8B--30B Models Against a Metric-Augmented SZZ Dataset", submitted to EASE 2026. Contains: (1) a dataset of 1,200 stratified code smell samples labeled as harmful or not harmful via the SZZ algorithm, (2) 14,400 raw LLM evaluation outputs from 4 open-source models (Qwen3-Coder 30B, DeepSeek-Coder-V2 16B, Llama 3.1 8B, Granite 3.1 Dense 8B) across 3 prompting strategies (zero-shot, one-shot, chain-of-thought), (3) computed evaluation metrics (precision, recall, F1, MCC), and (4) all evaluation and analysis scripts.
提供机构:
Zenodo创建时间:
2026-02-28



