RNA Portrait reproduction data: processed bulk RNA-seq matrices, validation profiles, model artifacts, and figure source data
收藏资源简介:
This dataset contains the non-code assets required to reproduce the analyses and figures reported in the study “Can genes be described in natural language like images?” (RNA Portrait). It accompanies the code repository at https://github.com/zhiluo20/rna-portrait and should be used with the matching repository revision. Contents - processed_data/training_matrix/ — prepared bulk RNA-seq training matrix containing 25,199 processed profiles from 117 public source identifiers and 4,096 selected genes, supplied as a monolithic matrix and project-level shards with aligned metadata.- processed_data/validation_profiles/ — External-180 (180 profiles) and MultiSource-450 (450 profiles), including per-profile expression and metadata files.- model_artifacts/bulk_multimodal_embedding/ — frozen model and runtime artifacts used in the reported analyses, including the RNA-language alignment backbone, prototype-attention portrait module, state-evidence scorer, age-like readout, benchmark outputs, and runtime configurations. See model_artifacts/README_model_artifacts.md for details.- source_data/ — 37 CSV files underlying the five main figures and seven Extended Data figures, with source_data_manifest.csv mapping files to figure panels.- figure_panel_sources/ — 14 SVG panel sources required for reconstruction of the Extended Data figures.- manuscript_analysis_tables/ — regenerated analysis tables supporting the reported comparisons.- software_environment/ — R 4.5.2 package-version records and installation instructions for the EPIC and MCP-counter analyses, together with the pinned third-party text-encoder revision.- quality_control/ — independent reproducibility checks and source-data comparisons.- data_availability/ — public-source accession and reproduction-asset manifests. Reproduction Clone the companion code repository and set the code and data-package paths as described in notebooks/External_Independent_Reproduction_Guide.ipynb. The release supports (1) retraining the RNA-language model from the deposited processed matrix, (2) rerunning the manuscript analyses from the frozen deposited artifacts, and (3) rebuilding the five main and seven Extended Data figures from source_data/ and figure_panel_sources/. Fresh training may produce numerically different weights across hardware and software versions. The deposited frozen artifacts define the outputs used in the reported analyses. Software environment The Python workflows require Python >= 3.11 and scikit-learn >= 1.8, together with the dependency versions specified in the companion code repository. The deconvolution analyses use R 4.5.2 and the package versions listed in software_environment/r_package_versions.csv. Data provenance The training and validation data were derived from public GEO, GDC/TCGA, and recount3 resources. Raw public RNA-seq data are not re-deposited. Original accessions and source identifiers are listed in data_availability/public_dataset_manifest.csv. Integrity The deposited archive is rna_portrait_data_package_20260817.tar.gz. Its SHA-256 checksum is provided in rna_portrait_data_package_20260817.tar.gz.sha256.



