遇见数据集

lagosproject/ALPHAGenome-Embeddings

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

ALPHAGenome hg38嵌入数据集是一个预计算的DNA基础模型嵌入集合,针对整个人类基因组(参考版本hg38/GRCh38)。该数据集将人类基因组划分为约22,000个非重叠的131 KB区间(bins),每个区间的DNA序列通过ALPHAGenome基础模型嵌入到一个3,072维的潜在空间中,以捕捉基因组区域的潜在特征和关系。数据集按染色体组织,每个染色体对应一个嵌入文件(.npy格式,包含浮点数组)和一个元数据文件(.csv格式,包含染色体、起始和结束位置信息),总数据量约为312 MB。这些嵌入可用于基因组分析、可视化(如通过ALPHAGenome UMAP Explorer进行交互式浏览)和机器学习应用。数据集基于CC BY 4.0许可,允许研究和教育用途。

Pre-computed ALPHAGenome DNA foundation model embeddings for the entire human genome (hg38 / GRCh38). The human genome is divided into approximately 22,000 non-overlapping 131 KB bins, and each bins DNA sequence is embedded into a 3,072-dimensional latent space using the ALPHAGenome foundation model. The dataset is organized by chromosome, with each chromosome having an embeddings file (.npy format) and a metadata file (.csv format), totaling around 312 MB. These embeddings enable analysis, visualization (e.g., via the ALPHAGenome UMAP Explorer), and machine learning applications for genomic regions. The dataset is licensed under CC BY 4.0 for research and educational use.

提供机构:
lagosproject
二维码
社区交流群
二维码
科研交流群
商业服务