遇见数据集

CAZyme prediction in Ascomycetes yeast genomes guides discovery of novel xylanolytic species with diverse capacities of hemicellulose hydrolysis

收藏
Zenodo2021-02-18 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

<strong>Background</strong> This data is part of our publication "<em>title_placeholder</em>" (<em>link_placeholder</em>). The fasta files are originally from another publication (https://doi.org/10.1016/j.cell.2018.10.023), with data hosted on Figshare (https://doi.org/10.6084/m9.figshare.5854692). We have, however, further processed those fasta files by clustering them at 98% identity (and removed whitespace in the fasta headers). They are provided here to enable users to retrieve protein sequences for genes listed in the "332_yeast_genomes_enzyme_info_version_3.tsv" file. <strong>File description</strong> The main output file is the "332_yeast_genomes_enzyme_info_version_3.tsv" tab-separated output file. Each row in the data file indicated one gene with a single corresponding hmm hit at a specific position in the gene. A gene can (and often does) occur multiple times with different hmm model hits or hits with the same hmm model but at different positions inside the gene. Below follows a description of the data contained in each column of the output file. The name of each gene is specified and the corresponding protein sequence can be obtained from the organism fasta files obtained from the Figshare repository indicated above. <strong>The columns in the output file "332_yeast_genomes_enzyme_info_version_3.tsv" are as follows:</strong> column: organism<br> description: the organism name<br> value type: text <br> column: gene<br> description: the gene name as given inside fasta files in "protein_fasta.zip"<br> value type: text <br> column: hmm_model<br> description: the hmm model from signalp that gave the hit<br> value type: text <br> column: hmm_model_len<br> description: length of the hmm model, specified in the hmmer output file (there in the "qlen" column)<br> value type: integer <br> column: hmm_match_from<br> description: where in the hmm model the match with the gene starts, specified in the hmmer output file (there in the "hmm coord from" column)<br> value type: integer <br> column: hmm_match_to<br> description: where in the hmm model the match with the gene ends, specified in the hmmer output file (there in the "hmm coord to" column)<br> value type: integer <br> column: hmm_match_coverage<br> description: how much of hmm model actually matched to the gene from 0.35 to 1.0, computed as ("hmm_match_to" - "hmm_match_from")/"hmm_model_len"<br> value type: float <br> column: match_evalue<br> description: the e-value of the hmm model hit, specified in the hmmer output file (there in the "Evalue" column)<br> value type: float, scientific notation <br> column: gene_match_from<br> description: where in the gene the hmm model match starts, specified in the hmmer output file (there in the "ali coord from" column)<br> value type: integer <br> column: gene_match_to<br> description: where in the gene the hmm model match ends, specified in the hmmer output file (there in the "ali coord to" column)<br> value type: integer <br> column: enzyme<br> description: the full enzyme name, parsed from the hmm model name by excluding the ".hmm" file extension<br> value type: text <br> column: family<br> description: the enzyme name, excluding subfamily designations, parsed from the "enzyme" column<br> value type: text <br> column: enzyme_type<br> description: which main class of enzyme it is, GH, CBM, CE, etc., parsed from the "family" column<br> value type: text <br> column: signal_peptide<br> description: whether signal peptide is predicted (SP(Sec/SPI)) or not (OTHER), specified in the signalp output file (there in the "Prediction" column)<br> value type: text <br> column: signal_peptide_prob<br> description: probability that a signal peptide is present, specified in the signalp output file (there in the "SP(SEC/SPI)" column)<br> value type: float <br> column: sp_cut_pos<br> description: the position in the protein sequence where the signal peptide is predicted to be cleaved, specified in the signalp output file (there in the "CS Position" column)<br> value type: text <br> column: sp_cut_seq<br> description: the sequence at which the signal peptide is predicted to be cleaved, specified in the signalp output file (there in the "CS Position" column)<br> value type: text <br> column: sp_cut_prob<br> description: the probability of the cut-site prediction, specified in the signalp output file (there in the "CS Position" column)<br> value type: float <br> column: genes_in_fasta<br> description: the number of genes present in the organisms fasta file<br> value type: integer

创建时间:
2021-02-18
二维码
社区交流群
二维码
科研交流群
商业服务