GROQ-seq Fitness and Function Measurements for TEV Protease Homolog Library
收藏资源简介:
(An updated version of this dataset is published at https://zenodo.org/records/20074407)This dataset contains GROQ-seq functional measurements of a sequence-diverse TEV protease library comprising 4,022 natural homologs identified by iterative JackHMMER searches of UniRef100, BFD, and MGnify, together with 422 AI-generated "shrunken" variants produced by the SCISOR discrete-diffusion model — yielding 11,722 unique amino-acid sequences (after synthesis-introduced variation) spanning Levenshtein distances up to 245 and a mean pairwise identity of 41.8% to TEV protease S219V. Fitness is measured via the same split-DHFR system as the TEV protease pilot, with the canonical TEV substrate (ENLYFQS) embedded as the linker; protease cleavage destroys DHFR and reduces trimethoprim resistance, translating activity into differences in cellular growth. Function is reported at the assay condition with the highest dynamic range (high TEV expression, low split-DHFR expression). This run of the assay was performed at the Living Measurements Systems Foundry (LMSF) at the National Institute of Standards and Technology. Part of a broader effort to generate sequence → function data across broad evolutionary sequence space, supporting machine-learning models of sequence-structure-function relationships and informing protein engineering across diverse protease families.
本数据集的更新版本已发布于https://zenodo.org/records/20074407。本数据集包含序列多样性TEV蛋白酶(TEV protease)文库的GROQ-seq功能检测数据,该文库包含通过对UniRef100、BFD及MGnify数据库进行迭代JackHMMER搜索鉴定得到的4022条天然同源序列,以及由SCISOR离散扩散模型(SCISOR discrete-diffusion model)生成的人工智能(AI)设计的“缩编”变体;经合成过程引入的变异校正后,共得到11722条独特的氨基酸序列,其莱文斯坦距离(Levenshtein distance)跨度可达245,与TEV蛋白酶S219V的平均两两序列一致性为41.8%。适应性通过与TEV蛋白酶先导实验相同的分裂DHFR(split-DHFR)系统进行测定,将经典TEV底物(ENLYFQS)作为接头嵌入体系;蛋白酶切割会破坏DHFR并降低甲氧苄啶抗性,从而将酶活转化为细胞生长的差异。功能数据以动态范围最高的检测条件(高TEV蛋白酶表达、低split-DHFR表达)进行报告。本次检测实验在美国国家标准与技术研究院(National Institute of Standards and Technology)的活体测量系统制造平台(LMSF)完成。本数据集属于一项更广泛研究计划的一部分,该计划旨在覆盖广阔的进化序列空间生成序列→功能数据,为序列-结构-功能关系的机器学习模型提供支撑,并为各类蛋白酶家族的蛋白质工程提供参考。



