遇见数据集

Curated Paired Antibody Dataset for Masked Language Model Training (AbCDR-MLM)

收藏
Zenodo2026-05-21 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains processed and curated paired heavy-light antibody sequences derived from the Observed Antibody Space (OAS) for masked language model training and evaluation, as as used in the paper "Preferential CDR masking in paired antibody language models improves binding affinity prediction". Contents: Clustered training set: 1,617,948 paired sequences (MMseqs2, 95% identity, 80% coverage) Validation set: 20,225 paired sequences Test set: 20,225 paired sequences IMGT-labeled CDR annotations for all sequences Citation If you use this dataset, please cite both the original OAS paper and our publication: Talaei, M., Walker, K. C., Hao, B., Jolley, E., Jin, Y., Kozakov, D., Misasi, J., Vajda, S., Paschalidis, I. Ch., & Joseph-McCarthy, D. (2025). Preferential CDR masking in paired antibody language models improves binding affinity prediction. bioRxiv. DOI: 10.1101/2025.10.31.685149

提供机构:
Zenodo
创建时间:
2026-02-24
二维码
社区交流群
二维码
科研交流群
商业服务