遇见数据集

Proteome-wide binding affinity dataset produced by CASTER-DTA

收藏
Zenodo2025-10-15 更新2026-05-26 收录
官方服务:

资源简介:

This is the proteome-wide binding affinity dataset generated by CASTER-DTA, as described in the CASTER-DTA paper (preprint: https://www.biorxiv.org/content/10.1101/2024.11.25.625281v2; Briefings in Bioinformatics DOI 10.1093/bib/bbaf554). The files are provided in two forms: a Pandas pickle file and a Parquet file, as this maximizes compression of duplicate information. To load these, you can use the Pandas commands: affinity_df = pandas.read_pickle('proteome_affinity.pkl') or affinity_df = pandas.read_parquet('proteome_affinity.parquet') The two files should generate identical dataframes, with columns: protein_id: the ID of the protein in the pair (from the AlphaFold2 human proteome database) protein_sequence: the amino acid sequence of the protein molecule_id: the ZINC ID for the ligand/molecule in the pair (from the ZINC database) molecule_smiles: the SMILES string for the ligand/molecule affinity_score: the predicted binding affinity using our pretrained CASTER-DTA model (pK_d) protein_fragment: the protein fragment in the AlphaFold2 database (for proteins of long lengths, AlphaFold2 splits them in their human proteome download). Most proteins are "F1" representing a single fragment (the whole protein).

提供机构:
Zenodo
创建时间:
2025-10-15
二维码
社区交流群
二维码
科研交流群
商业服务