Sequence-fitness dataset of O-methyltransferase activity for isovanillic acid production
收藏资源简介:
Corresponding dataset to the following publication: https://doi.org/10.1101/2025.10.24.684421This dataset is composed of two files data_seq_to_fitness_t_0_x_no_vif.csv: each row is a unique DNA sequence OMT.csv: each row is a unique amino acid sequence Thus, OMT.csv is a processed and aggregated version of data_seq_to_fitness_t_0_x_no_vif.csv. For more details on each, see below. 1. data_seq_to_fitness_t_0_x_no_vif.csv This dataset contains 572,013 unique datapoints obtained through MillionFull for sequence-fitness relationship for the enzyme activity protocatechuic acid + SAM -> isovanillic acid + SAH. For each variant we report precision‑weighted growth‑rate estimates, promoter, coding and translated protein sequences, exhaustive mutation annotations relative to a reference parent sequence, and quality‑control metrics derived from the underlying read clusters provided by DiMSum. For more details, see 1_readme.md 2. OMT.csv First, the following variants were removed from (1): variants with a stop codon variants with frameshift mutations (indels) in the DNA coding region variants with any mutations (substitution or indel) in the promoter region variants with indels in the protein sequence Then, for every unique protein sequence, growthrates from all degenerate nucleotide sequences were averaged to obtain the column growthrate_total. If multiple nucleotide sequences were present, growthrate_sigma_total was set to the standard deviation of each growthrate_total value per nucleotide sequnce. If only one nucleotide sequence was present for the protein sequence, then the nucleotide growthrate_sigma_total was used. Columns protein seq Protein sequence growthrate_total Growthrate total growthrate_sigma_total Growthrate sigma total mutations Mutations relative to parent, in list format. For example, ['S_28_T, E_50_K'] n mutations Number of mutations Parent Boolean, True/False normalized growthrate Growthrates normalized by parent growthrate. log normalized growthrate Log of normalized growthrate. Given a floor of -1.5
本数据集对应以下出版物:https://doi.org/10.1101/2025.10.24.684421。本数据集由两个文件组成: `data_seq_to_fitness_t_0_x_no_vif.csv`:每行对应一条唯一的DNA序列。 `OMT.csv`:每行对应一条唯一的氨基酸序列。 因此,`OMT.csv`是`data_seq_to_fitness_t_0_x_no_vif.csv`经处理与聚合后的版本。二者的详细说明如下。 1. `data_seq_to_fitness_t_0_x_no_vif.csv` 本数据集包含572,013条唯一数据点,通过MillionFull工具获取,用于研究原儿茶酸(protocatechuic acid)与S-腺苷甲硫氨酸(SAM)反应生成异香草酸(isovanillic acid)及S-腺苷同型半胱氨酸(SAH)的酶活序列-适合度关系。针对每个变体,本数据集提供以下信息: - 精度加权生长速率估计值 - 启动子(promoter)、编码区与翻译后蛋白质序列 - 相对于参考亲本序列的详尽突变注释 - 由DiMSum提供的原始读段簇衍生的质量控制指标。 更多细节请参阅`1_readme.md`。 2. `OMT.csv` 首先,从上述(1)的数据集中移除了以下变体: - 携带终止密码子(stop codon)的变体 - DNA编码区存在移码突变(插入缺失,indels)的变体 - 启动子区域存在任意突变(替换或插入缺失)的变体 - 蛋白质序列存在插入缺失的变体。 随后,针对每一条唯一的蛋白质序列,将所有简并核苷酸序列对应的生长速率取平均,得到`growthrate_total`列。若存在多条核苷酸序列,则将`growthrate_sigma_total`设为每条核苷酸序列对应的`growthrate_total`值的标准差;若某蛋白质序列仅对应一条核苷酸序列,则直接使用该核苷酸序列的`growthrate_sigma_total`。 各列说明如下: - `protein seq`:蛋白质序列 - `growthrate_total`:总生长速率 - `growthrate_sigma_total`:总生长速率标准差 - `mutations`:相对于亲本序列的突变,以列表格式呈现。例如:`['S_28_T, E_50_K']` - `n mutations`:突变数量 - `Parent`:布尔型字段,取值为True/False - `normalized growthrate`:以亲本生长速率归一化后的生长速率 - `log normalized growthrate`:归一化生长速率的对数值,设置下限为-1.5



