遇见数据集

SaeedLab/SeqScreen

收藏
Hugging Face2026-05-22 更新2026-05-31 收录
官方服务:

资源简介:

该README文件描述了SeqScreen基准测试中使用的两个数据集:ChEMBL和LIT-PCBA。ChEMBL数据集用于训练、验证和测试,包含1,492个蛋白质、394,190个分子和583,960个正对用于训练;334个蛋白质、101,586个分子和147,676个正对用于验证;326个蛋白质、105,048个分子和130,281个正对用于测试。该数据集基于ChEMBL 36生成,按蛋白质序列相似性分割。LIT-PCBA数据集包含1,452个蛋白质、375,120个分子和548,294个正对用于训练;326个蛋白质、97,213个分子和143,088个正对用于验证;15个蛋白质、404,586个分子和2,776,973个对(包括正负对)用于测试。训练和验证集使用过滤后的ChEMBL数据集生成,测试集为LIT-PCBA数据集。这些数据集旨在支持基于序列的虚拟筛选方法研究,用于计算药物发现中的蛋白质-分子结合预测。

The README file describes two datasets used in the SeqScreen benchmark: ChEMBL and LIT-PCBA. The ChEMBL dataset is used for training, validation, and testing, consisting of 1,492 proteins, 394,190 molecules, and 583,960 positive pairs for training; 334 proteins, 101,586 molecules, and 147,676 positive pairs for validation; 326 proteins, 105,048 molecules, and 130,281 positive pairs for testing. It was generated using ChEMBL 36 and split by protein sequence similarity. The LIT-PCBA dataset consists of 1,452 proteins, 375,120 molecules, and 548,294 positive pairs for training; 326 proteins, 97,213 molecules, and 143,088 positive pairs for validation; 15 proteins, 404,586 molecules, and 2,776,973 pairs (including positives and negatives) for testing. The training and validation sets were generated using a filtered version of the ChEMBL dataset, while the test set is the LIT-PCBA dataset. These datasets are designed to support sequence-based virtual screening methods for protein-molecule binding prediction in computational drug discovery.

提供机构:
SaeedLab
二维码
社区交流群
二维码
科研交流群
商业服务