遇见数据集

SpecificityStudio

收藏
Zenodo2026-06-24 更新2026-06-28 收录
官方服务:

资源简介:

The raw data associated with the initial release of the benchmark dataset SpecificityStudio, associated with the manscript "Differences between protein fitness models can be used to design variants of altered specificity" by Samuel P. Berry, Rachelle Gaudet and Debora Marks. Just the dataset itself can be loaded as a pickled python dictionary, This dataset contains raw DMS data, multiple sequence alignments, trained model files, and predicted structures used for analyses. Please see the GitHub repository for the code used to analyze the data in this repository. A guide to the contents of this repository: SpecificityStudio_Jun2026.pkl: Each of the eight datasets with their specificity score classifications and scores from the 14 models included in the initial preprint. All main text figures of the SpecificityStudio manuscript can be reproduced from this .pkl file with the notebooks found in the GitHub repo. raw_data.tar.gz: The raw data for each of the eight datasets. Necessary if you want to re-process this data processed_data.tar.gz: The processed data in standardized, combined formats and with specificity classes assigned to each mutation msas.tar.gz: Multiple sequence alignments used in the initial study weights.tar.gz: Sequence weights calculated for sequences in each MSA, used by some models evcouplings_models.tar.gz: Trained EVcouplings models for each alignments in msas eve_models.tar.gz: Trained EVE models for each alignment in msas model_scores.tar.gz: The final model scores as .csv files (also in SpecificityStudio_Jun2026.pkl)

提供机构:
Zenodo
创建时间:
2026-06-24
二维码
社区交流群
二维码
科研交流群
商业服务