遇见数据集

AmelieSchreiber/binding_sites_random_split_by_family_550K

收藏
Hugging Face2023-09-13 更新2024-03-04 收录
官方服务:

资源简介:

该数据集通过UniProt搜索获得,包含带有家族和结合位点注释的蛋白质序列。数据集包括未审查(TrEMBL)和已审查的蛋白质序列,并且只包含注释得分为4的序列。数据集按家族排序和分割,随机选择家族作为测试数据集,直到大约20%的蛋白质序列被分离出来用于测试数据。排除了在结合位点注释中包含`<`、`>`或`?`的序列。此外,还包括了未列为结合位点的活性位点。对于长度超过1000个残基的序列,在训练测试分割后将其分割为不超过1000个氨基酸的非重叠部分。数据集还提供了包含训练/测试序列及其二进制标签的Pickle文件,可用于训练或验证训练/测试指标。

This dataset is obtained through UniProt searches, and comprises protein sequences annotated with protein family and binding site information. It includes both unreviewed (TrEMBL) and reviewed protein sequences, and only retains sequences with an annotation score of 4. The dataset is sorted and split by protein family: families are randomly selected to constitute the test dataset until approximately 20% of the total protein sequences are separated for testing. Sequences containing `<`, `>`, or `?` in their binding site annotations are excluded. Additionally, active sites that are not categorized as binding sites are also included in the dataset. For sequences longer than 1000 residues, they are split into non-overlapping segments of no more than 1000 amino acids each following the train-test split. The dataset also provides Pickle files containing training/test sequences and their binary labels, which can be used for model training or to validate training and testing metrics.

提供机构:
AmelieSchreiber
原始信息汇总

数据集概述

数据来源

  • 数据集来源于UniProt搜索,包含具有家族和结合位点注释的蛋白质序列。

数据内容

  • 包括未审核(TrEMBL)和已审核的蛋白质序列。
  • 仅包含注释分数为4的序列。
  • 排除了结合位点注释中包含<>?的序列。
  • 包含未列为结合位点的活性位点。

数据处理

  • 按家族分类,随机选择约20%的蛋白质序列作为测试数据。
  • 将长度超过1000个残基的序列分割成不超过1000个氨基酸的非重叠片段。
  • 提供仅包含训练/测试序列及其二进制标签的Pickle文件,可用于训练或验证。

数据规模

  • 数据集大小:100K<n<1M

标签

  • 包含“Binding-Active Sites”列,合并了结合位点和活性位点。

适用领域

  • 生物学
  • 蛋白质序列
  • 结合位点
  • 活性位点
二维码
社区交流群
二维码
科研交流群
商业服务