遇见数据集

CloudGraph: evaluation dataset for graph-grounded verification of LLM-generated root cause analysis

收藏
Zenodo2026-08-15 更新2026-08-20 收录
官方服务:

资源简介:

Frozen evaluation results from CloudGraph: 36 chaos-injected Kubernetes failurescenarios drawn from RCAEval RE2, producing 3,685 extracted claims, each scoredindependently by two verifiers — Graph-Provenance Claim Scoring (GPCS) and aself-consistency baseline. The design is fully balanced: 3 systems x 6 fault types x 2 replicates, each rununder 3 context conditions and 3 retrieval methods. Confidence intervals arescenario-clustered paired bootstrap (10,000 resamples, seed 42) with Wilcoxonsigned-rank tests. Three of the four answered research questions went against the design'spredictions and are reported as measured: - GPCS flags 70.3% of claims unsupported against self-consistency's 57.9% (delta +0.1185, 95% CI [+0.0729, +0.1632], p < 0.0001)- On the 155 claims (4.2%) carrying correctness labels, neither verifier separates correct claims from incorrect ones (both gaps -0.8 pp)- Ranked retrieval did not beat an unranked context dump (p = 0.302)- Vector and hybrid retrieval are byte-identical across all 36 scenarios The "agreement" column records whether two verifiers reached the same verdict.It is not a measure of correctness, and nothing in this dataset supports a claimabout verifier accuracy. All scoring thresholds are hand-set; nothing iscalibrated. Includes the full request log (keys redacted), a SHA-256 manifest of every file,and the three published figures. Code: https://github.com/shivamshashank/CloudGraph

提供机构:
Zenodo
创建时间:
2026-08-15
二维码
社区交流群
二维码
科研交流群
商业服务