该数据集卡片包含了用于评估蛋白质语言模型的预处理数据集,这些数据集用于论文《MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding》中的模型评估。数据集包括两部分:一是用于训练MULAN-small的预处理AF-0.5M数据集(AF05_pretraining.zip),二是论文中使用的所有下游数据集
# Dataset Card for Dayhoff Dayhoff is an Atlas of both protein sequence data and generative language models — a centralized resource that brings together 3.34 billion protein sequences across 1.7 bi
Data from The geometry of the hidden representation of large transformer models of the ESM2 model In the first release only the value of the intrinsic dimension (ID) and neighbor overlap with ground t
--- license: cc-by-nc-nd-4.0 tags: - bioinformatics - protein - drug-discovery --- # Protein Benchmark for ProtEnrich The paper is under review. \[[Github Repo](https://github.com/pcdslab/ProtEnric