遇见数据集

47-Dataset Benchmark Validation Results for Pathway Subtyping Framework

收藏
Zenodo2026-07-30 更新2026-05-26 收录
官方服务:

资源简介:

Version 2.1 — description corrected (2026-07-29). The benchmark data are unchanged from v2.0; this version corrects the record's own description, which overstated what the corrected benchmark supports. (ERRATUM_2026-07-08.md is refreshed to name the concept DOI explicitly; corrected_benchmark_47datasets_v2.csv and threshold_model_real47_CORRECTED.json are byte-identical to v2.0, so anything computed against v2.0 remains valid without re-running.) The v2.0 description closed with the claim that the benchmark "shows that discrete-subtype reproducibility is dataset-dependent and low on most independent cohorts — a negative methodological result." That claim is withdrawn. It rests on the bootstrap_ari_5th_percentile column, and that column will not carry it, for two reasons: It is a 5th percentile, not a point estimate, so any fixed bar applied to it — 0.5, or the 0.80 used elsewhere — is near-zero by construction. A low pass rate is a property of the statistic, not a finding about subtypes. The column is internally inconsistent with the silhouette column. Well-separated clusters are trivially reproducible under resampling, so a high silhouette forces a high stability ARI. Yet the record contains rows such as GSE66351 (silhouette 0.967, bootstrap ARI −0.0095) and GSE16759 (silhouette 0.958, bootstrap ARI −0.1394), which are not physically consistent. The column does not measure bootstrap stability of the reported partition. No reproducibility-rate claim should be drawn from this record, in either direction. The withdrawal is not replaced with another rate. What the record does support is unchanged and remains the reason to cite it: the adaptive bootstrap-threshold model previously reported at R²=0.889 does not reproduce and is retracted. Refit on the released data it gives R²≈0.111 across all rows, ≈0.015 on valid rows, and ≈0.001 under a stricter ground-truth screen, with the slope reversing sign. Silhouette does not predict bootstrap-ARI reproducibility in these data, and no silhouette-calibrated threshold is supported by them. The corrections carried forward from v2.0 also stand: fourteen rows have degenerate ground truth (n_true_clusters=1) and are flagged invalid — the empty-input adjusted-Rand-index artifact gave 13 of them a spurious ARI=1.0, the 14th (GSE136196) returned 0.0; one row had impossible labels; the claim that the benchmark excludes the manuscript analysis dataset (TCGA-COAD) was incorrect and is withdrawn; and the counts are restated as 41 unique datasets, not 47, with 40,778 samples in passing rows, not 36,551. Known limitation, disclosed rather than fixed. The erratum's ground-truth rule (n_true_clusters > n_samples) is too weak. It catches only GSE92332 (1,957 labels / 533 samples) and misses rows where the label count approaches the sample count — GSE2109 (2,158/2,158), GSE5204 (79/79), GSE17537 (55/54), GSE44228 (72/69), GSE42127 (176/145), GSE66351 (190/96), GSE16759 (16/8). Seven rows still marked valid fail a ratio screen of n_true_clusters / n_samples >= 0.5. The data files are left as-is so that analyses run against v2.0 stay reproducible bit-for-bit; the stricter screen is applied in the analysis code instead, and the retraction conclusion holds under it (R²≈0.001). Users applying their own screen should apply the ratio test, not the erratum rule. Citing this record: use the concept DOI 10.5281/zenodo.19323753, which always resolves to the newest version. Version DOIs (v1.0 10.5281/zenodo.19323754, v1.1 10.5281/zenodo.19324360, v2.0 10.5281/zenodo.21262112) remain permanently resolvable per Zenodo DOI permanence — published records can be superseded but not deleted — and should not be cited in new work. See ERRATUM_2026-07-08.md in this record for the full correction notice.

提供机构:
Zenodo
创建时间:
2026-03-29
二维码
社区交流群
二维码
科研交流群
商业服务