遇见数据集

Supplementary data for: "Rewriting protein alphabets with language models"

收藏
Zenodo2025-12-11 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the benchmarking data associated with the manuscript: "Rewriting protein alphabets with language models". The data is organized into three main archive files and one standalone result file, which together contain all necessary sequences, search results, and metadata required for reproducing the experiments and analyses presented in the paper. SCOPe40 Benchmarking Data (scope40.tar.gz) This file includes SCOPe40 (v2.08) sequences, sequence alignments, and the files necessary for reproducing all the benchmarked methods in the paper: Sequences: standard amino acid, TEA, 3Di and DSSP. Family/Superfamily/Fold sensitivities: MMseqs2-TEA/CV, MMseqs2-amino-acid sequences, MMseqs2-3Di, Foldseek, EBA, plus the ablations and the "exhaustive" searches combining different sequence representations. Multidomain Benchmarking Data (multidomain.tar.gz) This file includes the multidomain benchmark sequences and search results: Sequences: standard amino acid and TEA sequences for both the target and query sets. Search results: MMseqs2-TEA, MMseqs2-amino-acid sequences, Foldseek. 40k-pLDDT set (40kplddtset.tar.gz) Set described in the paper for entropy analysis and pLDDT mapping, includes: Sequences: standard amino acid, TEA, DSSP, and 3Di. Entropies: both average and per-residue TEA Shannon entropy values. pLDDT: from both AlphaFold2 and ESMFold predictions. Predicted Structures: ESMFold predicted structures. Singletons search results (singletons_vs_all_search_18_3.m8.gz) The search results from the singleton analysis discussed in Section 2.4 ("Improving functional annotation...") of the manuscript.

提供机构:
Zenodo
创建时间:
2025-12-11
二维码
社区交流群
二维码
科研交流群
商业服务