遇见数据集

A measurement-validity framework for evaluating cross-cohort similarity between Hallmark correlation matrices: a hepatocellular carcinoma case study

收藏
Zenodo2026-08-13 更新2026-08-20 收录
官方服务:

资源简介:

Reproduction bundle for the manuscript "A measurement-validity framework for evaluating cross-cohort similarity between Hallmark correlation matrices: a hepatocellular carcinoma case study." It contains the source-bound objects underlying every value reported in the manuscript, each pinned by checksum, together with the analysis code, release configuration, and acceptance tests used to verify the deposit. WHAT THE STUDY ASKS. Hallmark program scores were computed separately in two hepatocellular carcinoma cohorts -- TCGA-LIHC (371 tumours) and ICGC LIRI-JP (232 donors) -- and their 50 x 50 correlation matrices were compared. Because Hallmark gene sets reuse member genes, the library's incidence and overlap architecture can induce covariance among program scores before biological interpretation. The study therefore asks what cross-cohort similarity between complete gene-set correlation matrices can support, in three analytical layers, and attaches to each answer the limit that its operator cannot exceed. WHAT IT FINDS. The two correlation matrices correlate at 0.7986. Against an ensemble of B = 10,000 objects that preserve each cohort's effective gene-set sizes and complete overlap structure while reassigning gene identity, the ensemble mean is 0.7169 and the observed value lies at the 87th percentile, with an upper-tail rank of 0.1295. A label-permutation null places the same named-axis correspondence at 1/10,001. Thus, the named correspondence is not reproduced by relabelling, but the magnitude of the raw similarity is compatible with the architecture-covariance scaffold generated by the overlap-preserving ensemble. After edgewise leave-one-object-out calibration relative to that scaffold, the residual field is broad and cross-cohort concordant: all 50 axes meet the prespecified 0.05 step-down criterion, 46 attain 6/10,001 or lower, and the two cohorts' residual patterns correlate at 0.8709. Mapping-based quadratic projections of the centered residual reach 1/10,001 in both the pathway-footprint and regulon families. These projections are sign-invariant and use study-derived Hallmark loadings; post-result diagnostics show that the loading fields are low-dimensional and overlapping rather than collections of independent directional findings. Five registered estimands were evaluated separately, each with its own population, score object, statistical operator, and null or directional alternative. None was promoted to the central claim, and one was a prespecified negative result. WHAT IS IN THE DEPOSIT. The deposit contains the frozen scientific-text source and public Supplement in Markdown; the typeset manuscript and Supplement in DOCX and PDF; Figures 1-5 in PDF and PNG; the analysis code and release configuration; per-pool score matrices, gene-support and membership objects, registered-estimand frames, and the immutable result objects cited by the manuscript; and the fifteen display derivatives named in the Supplement. Three catalogues document the release. MANIFEST.tsv records every bundled file, its byte count, SHA-256 checksum, class, and provenance chain. EXTERNAL.tsv records source objects that are named but not redistributed. CONTENT_DIGESTS.tsv records checksums computed over defined content rather than over an individual file. Where an object was written by one script and later rewritten by another, the manifest records the ordered production chain rather than attributing the final object to a single script that no longer regenerates it. HOW TO VERIFY IT. The release implements two verification paths with different tolerance contracts, together with a clean-install smoke test. The artifact-integrity path verifies every manifest row by byte count and checksum, recomputes the four content digests from their bundled sources, and reproduces deposit-to-deposit comparison values exactly. The regeneration path rebuilds seven score matrices from pinned inputs and compares manuscript-bearing quantities within tolerances fixed in writing before the tests were run; it also refits the two etiology estimands that operate directly on regenerated score matrices. The clean-install test verifies the release plumbing without making an additional scientific claim. The current acceptance record reports PASS = 22, FAIL = 0, SKIPPED = 0. The complete record is provided in tests/RESULTS.md. WHAT IS NOT INCLUDED. Three raw microarray series files from the Gene Expression Omnibus -- GSE14520, GSE89377, and GSE6764 -- are public and are not redistributed. They are identified in EXTERNAL.tsv by accession, byte count, and checksum. The GSE14520 gene-collapsed cache is bundled as the entry point for that analysis arm. TCGA-LIHC and ICGC LIRI-JP primary data remain available through the Genomic Data Commons and the ICGC open-access tier, respectively. LICENCE. Manuscript text, figures, and distributable derived data objects are released under CC BY 4.0. Code is released under the MIT License. Upstream resources and data from TCGA, ICGC, the Gene Expression Omnibus, MSigDB, PROGENy, and DoRothEA remain governed by their respective terms, which this deposit does not alter. See LICENCE for the complete release and upstream-resource statement.

提供机构:
Zenodo
创建时间:
2026-08-13
二维码
社区交流群
二维码
科研交流群
商业服务