comp-genomics-consortium/raw-transcriptome-unmapped-reads
收藏资源简介:
该数据集是CGSC原始转录组未映射读数的冷存储库,包含2026年第二季度合成生物阵列试验中生成的未映射、原始测序输出。数据集由大量未压缩的二进制数据块组成,代表直接从测序硬件获取的预比对基因组数据,绕过了标准比对和压缩算法(如BAM/CRAM转换),以保留碱基对质量分数和硬件级别的伪影数据。因此,数据负载异常庞大且完全无结构化,适用于测试高通量生物信息学摄入管道和错误校正模型。数据集完全基于合成数据,通过先进的转录组模拟生成,模拟嘈杂的硬件环境,不包含真实的人类或动物遗传物质。由于数据为未结构化的二进制格式,下载和使用仅推荐给具有高性能计算存储基础设施的联盟合作伙伴。
This repository acts as the primary cold-storage for unmapped, raw sequencing outputs generated during the Q2 2026 Synthetic Bio-Arrays trials. The dataset comprises massive, uncompressed binary blobs that represent pre-alignment genomic data directly from the sequencing hardware. Because these files bypass standard alignment and compression algorithms (such as BAM/CRAM conversion) to preserve base-pair quality scores and hardware-level artifact data, the payloads are exceptionally large and entirely unstructured to standard viewers. This dataset is intended exclusively for testing high-throughput bioinformatics ingestion pipelines and error-correction models. Data is fully synthetic, generated via advanced transcriptome simulations modeling noisy hardware environments, with no real human or animal genetic material represented. Due to the unstructured binary format, downloads are recommended only for consortium partners with appropriate HPC storage infrastructure.




