遇见数据集

bioStream Standardized Reference Genomes and Multi-omics Indices

收藏
Zenodo2026-05-25 更新2026-05-26 收录
官方服务:

资源简介:

Overview A primary technical challenge in cross-study meta-analyses and multi-omics integration is the transcriptomic or epigenomic coordinate mismatching that arises from disparate reference assemblies. To eliminate this systemic variation and enforce strict analytical reproducibility, bioStream implements a standardized reference architecture anchored on the official 10x Genomics refdata-cellranger-arc ecosystem. This Zenodo repository hosts the pre-built standard downstream analytical indices (including STAR, RSEM, and Bowtie2) along with auxiliary quality control annotations. Programmatically extracted from a single root reference source, these indices guarantee that all subsequent alignment and quantification tasks across bulk (RNA, ATAC, ChIP) and single-cell modalities remain perfectly synchronized within the exact same genomic coordinate space. What is Included in this Repository This repository hosts the standalone standard bulk references for human and mouse. ⚠️ Note for Single-Cell / 10x Genomics Users: To minimize hosting redundancy, the official 10x Genomics GEX and ARC reference folders are not bundled inside these archives. They can be downloaded directly from the 10x Genomics official support site. The automated scripts to build, extend, and structure the entire reference directory layout can be found in our main repository: GitHub: JiekaiLab/bioStream. Supported Species & Contents 1. Homo sapiens (GRCh38 / hg38) Archive: hg38.tar.gz Coordinate Anchor: Strictly identical to 10x Genomics Cell Ranger ARC Reference 2024-A (refdata-cellranger-arc-GRCh38-2024-A). Components: Complete STAR index, RSEM transcript index, Bowtie2 genome index, ENCODE v2 blacklists, BED12 gene models, and RSeQC ribosomal RNA coordinates. 2. Mus musculus (GRCm38 / mm10) Archive: mm10.tar.gz Coordinate Anchor: Strictly identical to 10x Genomics Cell Ranger ARC Reference 2020-A (refdata-cellranger-arc-mm10-2020-A-2.0.0). Components: Complete STAR index, RSEM transcript index, Bowtie2 genome index, ENCODE v2 blacklists, BED12 gene models, and RSeQC ribosomal RNA coordinates. Data Components & Specifications Each zipped archive contains the following decoupled data layers: STAR Index: High-performance, splice-aware indices (Genome, SA, SAindex) optimized for RNA-seq alignment tasks. RSEM Index: Comprehensive transcript-level quantification references (.transcripts.fa, .ti, .seq, .grp) synced with the root GTF. Bowtie2 Index: Full .bt2 index set for DNA-level alignment tasks including bulk ATAC-seq and ChIP-seq. QC & Filtering Resources: bed12/ & RSeQC/: Full gene models in BED12 format and custom rRNA interval maps for post-alignment quality assessment. blacklists/: Curated ENCODE v2 Blacklist BED files to filter out high-signal anomalous regions (e.g., telomeres, centromeres). gene_anno/: A consolidated, flat metadata file (gene_anno_*.csv) mapping Gene IDs to Gene Symbols for simplified downstream Python/R analysis. Target Directory Tree Once downloaded and extracted alongside your 10x assets, the pipeline expects the following standardized tree layout under your reference/ path: reference ├── hg38 │ ├── bed12 # GRCh38_bed12.bed │ ├── blacklists # hg38-blacklist.v2.bed │ ├── bowtie # Bowtie2 index set (*.bt2) │ ├── gene_anno # gene_anno_GRCh38.csv │ ├── rsem # RSEM index set (*.ti, *.seq, etc.) │ ├── RSeQC # hg38_rRNA.bed │ └── star # STAR genome index folder ├── hg38_10x # [User-Downloaded] Official 10x GEX & ARC references │ ├── refdata-cellranger-arc-GRCh38-2024-A │ └── refdata-gex-GRCh38-2024-A ├── mm10 │ ├── bed12 # mm10_bed12.bed │ ├── blacklists # mm10-blacklist.v2.bed │ ├── bowtie # Bowtie2 index set (*.bt2) │ ├── gene_anno # gene_anno_mm10.csv │ ├── rsem # RSEM index set │ ├── RSeQC # mm10_rRNA.bed │ └── star # STAR genome index folder └── mm10_10x # [User-Downloaded] Official 10x GEX & ARC references ├── refdata-cellranger-arc-mm10-2020-A-2.0.0 └── refdata-gex-mm10-2020-A Technical Specifications & Consistency To guarantee absolute downstream comparability between bulk and single-cell cohorts processed through bioStream, the underlying sequence genome.fa and gene models genes.gtf are completely synchronized across every single sub-directory within a given species. For the complete step-by-step pipeline execution, deployment parameters, and multi-omics orchestration details, please refer to the main repository documentation at GitHub: JiekaiLab/bioStream.

提供机构:
Zenodo
创建时间:
2026-05-06
二维码
社区交流群
二维码
科研交流群
商业服务