LiteFold/protenix-data
收藏资源简介:
Protenix Data是字节跳动开源的AlphaFold3复现模型Protenix所使用的完整预处理训练数据集,是目前公开的最大AF3风格训练语料库之一。该数据集基于wwPDB构建,专门用于AF3风格训练,支持蛋白质、DNA、RNA、配体、离子和修饰等多种生物分子的全原子扩散模型预测。数据集包含原始文件931,270个,总大小1.05 TiB,分为52个未压缩的tar分片。数据内容包含mmCIF文件、生物组装结构、MSA(多序列比对)模板、RNA MSA以及通过ColabFold流程生成的UniRef30和环境数据库比对结果。Protenix-v1版本将PDB数据截止时间扩展到2025-06-30,并增加了RNA MSA和基于HMMER的结构模板。数据集文件类型主要包括.a3m、.cif、.fasta、.pkl.gz、.csv等格式,适用于生物信息学、蛋白质结构预测和计算生物学研究。
Protenix Data is the full preprocessed training dataset released by ByteDance for training Protenix, an open-source PyTorch reproduction of AlphaFold3. It is one of the largest publicly available AF3-style training corpora, built from the wwPDB and designed for drop-in use in AF3-style training. The dataset supports a unified all-atom diffusion model for proteins, DNA, RNA, ligands, ions, and modifications. It contains 931,270 original files with a total payload of 1.05 TiB, organized into 52 uncompressed tar shards. The data includes mmCIF files, bioassembly structures, MSA templates, RNA MSAs, and alignments generated via the ColabFold pipeline against UniRef30 and environmental databases. The Protenix-v1 version extends the PDB cutoff to 2025-06-30 and adds RNA MSAs and HMMER-derived structural templates aligned with the AF3 template strategy. Common file formats include .a3m, .cif, .fasta, .pkl.gz, and .csv, making it suitable for bioinformatics, protein structure prediction, and computational biology research.




