GROQ-seq Fitness and Function Measurements for TEV Protease Homolog Library
收藏资源简介:
(An updated version of this dataset is published at https://zenodo.org/records/20074407)This dataset contains GROQ-seq functional measurements of a sequence-diverse TEV protease library comprising 4,022 natural homologs identified by iterative JackHMMER searches of UniRef100, BFD, and MGnify, together with 422 AI-generated "shrunken" variants produced by the SCISOR discrete-diffusion model — yielding 11,722 unique amino-acid sequences (after synthesis-introduced variation) spanning Levenshtein distances up to 245 and a mean pairwise identity of 41.8% to TEV protease S219V. Fitness is measured via the same split-DHFR system as the TEV protease pilot, with the canonical TEV substrate (ENLYFQS) embedded as the linker; protease cleavage destroys DHFR and reduces trimethoprim resistance, translating activity into differences in cellular growth. Function is reported at the assay condition with the highest dynamic range (high TEV expression, low split-DHFR expression). This run of the assay was performed at the Living Measurements Systems Foundry (LMSF) at the National Institute of Standards and Technology. Part of a broader effort to generate sequence → function data across broad evolutionary sequence space, supporting machine-learning models of sequence-structure-function relationships and informing protein engineering across diverse protease families.
本数据集的更新版本已发表于https://zenodo.org/records/20074407。本数据集包含序列多样性的TEV蛋白酶(TEV protease)文库的GROQ-seq功能检测数据,该文库包含通过对UniRef100、BFD和MGnify进行迭代JackHMMER搜索鉴定得到的4022个天然同源序列,以及由SCISOR离散扩散模型生成的422个人工智能(AI)设计的“缩窄型”变体;经合成引入的变异后,总计得到11722条独特的氨基酸序列,其莱文施泰因距离(Levenshtein distance)最高可达245,与TEV蛋白酶S219V的平均两两序列一致性为41.8%。适应性通过与TEV蛋白酶先导实验相同的split-DHFR系统进行测定,将经典TEV底物(ENLYFQS)作为连接肽嵌入;蛋白酶切割会破坏DHFR并降低甲氧苄啶(trimethoprim)抗性,从而将蛋白酶活性转化为细胞生长的差异。本研究在动态范围最高的实验条件(高TEV表达、低split-DHFR表达)下报告功能检测结果。 本次检测实验在美国国家标准与技术研究院(National Institute of Standards and Technology, NIST)的活体测量系统铸造厂(Living Measurements Systems Foundry, LMSF)完成。本数据集属于一项更广泛研究计划的一部分,该计划旨在跨广阔进化序列空间生成序列→功能数据,用于支撑序列-结构-功能关系的机器学习模型研发,并为多样化蛋白酶家族的蛋白质工程提供指导。



