xiao-fei/PPI2Text-Dataset
收藏资源简介:
PPI2Text数据集是一个用于蛋白质-蛋白质相互作用(PPIs)的自由文本描述数据集。每个数据点包含两个UniProt登录号和一个描述相互作用的英文段落,总结相互作用的生物背景、条件、机制和功能后果。该数据集旨在训练和评估多模态模型,从蛋白质序列或结构输入生成PPI描述。在文本中,蛋白质名称被替换为Protein A和Protein B占位符,以确保模型必须从序列本身学习,而非通过名称记忆。数据集包含351,515个唯一蛋白质对,平均每个描述约1,770字符(约300词),以Parquet格式存储。构建过程基于IntAct、UniProt、STRING、PubMed、Reactome和3did等来源的数据,并使用Gemini-3-Pro-Preview生成证据基础的段落。数据集仅包含配对标识符和响应文本,不包含蛋白质序列或结构。
Free-text descriptions of protein–protein interactions (PPIs). Each example pairs two UniProt accessions with a single paragraph that summarizes the interactions biological context, conditions, mechanism, and functional consequences. The dataset was built to train and evaluate multimodal models that generate PPI descriptions from protein sequence/structure inputs. Protein names are replaced by the placeholders Protein A and Protein B in the text, so a model conditioned only on the two UniProt accessions cannot trivially recover the answer through name memorization and must learn from the sequences themselves. It contains 351,515 unique pairs with an average response length of approximately 1,770 characters (about 300 words), stored in Parquet format. The dataset was constructed by seeding pairs from IntAct and enriching with annotations from UniProt, STRING, PubMed, Reactome, 3did, etc., and using Gemini-3-Pro-Preview to generate evidence-grounded paragraphs. It only includes pair identifiers and response text, without protein sequences or structures.




