遇见数据集

Human-Aligned Instance-Level Technical Debt Severity Classification

收藏
Zenodo2026-09-26 更新2026-10-01 收录
官方服务:

资源简介:

# Human-Aligned Instance-Level Technical Debt Severity Classification Code and data for the paper *"Human-Aligned Instance-Level Technical Debt Severity Classification "* (IEEE SANER 2027, anonymized for double-blind review). ## OverviewTechnical-debt (TD) severity is commonly studied using severity labels assigned by static-analysis tools such as SonarQube. We show that these labels are fundamentally rule-level: across 67,540 SonarQube issues, rule identity alone reproduces the tool-assigned severity with 100.00% accuracy. In contrast, human judgments vary across instances of the same rule, indicating that TD severity is better treated as an instance-level property. We construct a human-annotated reference set of 2,000 issues and use a human-calibrated multi-LLM annotation process to extend the available supervision to a larger set of issues. We then study instance-level severity classification using three complementary views: code, structural metrics, and textual descriptions. Finally, we examine whether the learned textual severity signal generalizes across analyzer-generated descriptions and developer-authored SATD comments. This repository provides one representative script for each research question together with the three key datasets used in the study. It is intended to support inspection and reproduction of the main experiments rather than provide the complete data-collection and annotation pipeline. ## Contents ```├── data/ # 原始数据(新加)│ ├── sonarqube_issue.csv # 67,540 条 SonarQube 原始 issue (11MB)│ └── TD-Describe.csv # 20,167 条,code/codeSATD/SATD (26MB)├── experiment/│ ├── RQ1_model_compare.py # RQ1 — effectiveness vs. baselines (TF-IDF vs LLM embeddings; LR/SVM/XGBoost)│ ├── RQ2_ablation_sonarqube.py # RQ2 — which view (code / metrics / text) carries the signal│ ├── RQ3_pseudolabel_sonarqube.py # RQ3 — human-calibrated multi-LLM pseudo-labeling│ ├── RQ4_transfer_satd.py # RQ4 — cross-source transfer (description → SATD comment)│ ├── sonarqube_human_annotation_full.csv # human gold (2,000 issues)│ ├── pseudolabels_sonarqube.csv # LLM pseudo-labels (~30,000 issues)│ └── RQ4_panel_labels.csv # SATD same-problem set (538 items)├── README.md└── LICENSE``` ## Data | File | Description | Rows ||------|-------------|------|| `sonarqube_human_annotation_full.csv` | Instance-level human gold (Low/Medium/High/Not-a-real-issue), three developers + expert adjudication | 2,000 || `pseudolabels_sonarqube.csv` | LLM panel pseudo-labels over unlabeled issues | 30,019 || `rq5_panel_labels.csv` | SATD comments judged to describe the *same* problem as a paired SonarQube issue, with panel severity labels | 538 | The upstream raw SonarQube dump is not included. Scripts read raw data from `../data/` (a sibling `data/` directory); the LLM-based scripts call an OpenAI-compatible gateway and read the key from the environment variable `DEEPSEEK_API_KEY`. ## Setup ```bashpip install scikit-learn pandas xgboost openaiexport DEEPSEEK_API_KEY="your-key"``` ## License MIT — see [LICENSE](LICENSE).

提供机构:
Zenodo
创建时间:
2026-09-26
二维码
社区交流群
二维码
科研交流群
商业服务