遇见数据集

PRESTO: A Machine Learning Framework for Predicting Software System Performance Using SDLC Metrics — Research Compendium (Code, Synthetic Data, and Real-World Validation Adapters)

收藏
Zenodo2026-08-19 更新2026-08-20 收录
官方服务:

资源简介:

Source data, code, and paper for “PRESTO: A Machine Learning Framework for Predicting Software System Performance Using SDLC Metrics” (Rangari, Mishra, Senapati). Includes: (1) the Gaussian copula-based synthetic SDLC data generator and its three generated domains (ABC Cloud Provider, XYZ Sales Force, Card Payment Processor); (2) adapter/pipeline scripts and small derived outputs for four independent real-world validation datasets (TravisTorrent, Mozilla Perfherder, GHALogs, SQuaD) — bulk third-party raw data is not re-hosted, each source folder documents exactly how to fetch it from its original DOI/URL; (3) all ablation/robustness scripts (nested cross-validation tuning, Siegmund baseline, System-Availability ablation, bootstrap and target-correlation ablation, lag-shift leakage fix, Phase-Aware Recursive Feature Elimination reimplementation); (4) the figure-generation scripts; and (5) the manuscript LaTeX source. How well can you predict production performance from development process data alone? Modern DevOps pipelines record 164 metrics across eight SDLC phases, yet no public dataset links these signals to runtime outcomes. Across three enterprise-domain synthetic datasets, the best model achieved R² = 0.26–0.36 using only SDLC process features; a target-correlation ablation shows this largely reflects recovery of the generator's specified structure. Four independent real-world datasets provide bootstrap-CI-backed validation: TravisTorrent (R²=0.475 best of seven, median 0.005), Mozilla Perfherder (R²=0.17–0.91, autoregressive persistence), GHALogs (R²≈0.11, zero leakage risk, n=28,443), and SQuaD (R²=0.402 defect-fix, 0.483 enriched with real static-analysis features; R²=0.245 CVE-count). GHALogs and SQuaD's CVE-count result license a confirmatory claim: SDLC process signals carry genuine, transferable predictive value, though not uniformly across every real-world dataset examined. See README.md in this deposit for the full directory map and reproduction steps, and real-data/README.md for exactly what raw third-party data is excluded and how to fetch it (kept out to keep this compendium small and clonable, per source DOI/URL, not because it's unavailable).

提供机构:
Zenodo
创建时间:
2026-08-19
二维码
社区交流群
二维码
科研交流群
商业服务