Annotations of human skin-associated fungi
收藏资源简介:
Description This repository contains genome annotations for skin-associated fungi sequenced and analysed in the manuscript:Comparative Genomic Analysis of Clinically Relevant Human Skin-Associated Fungi, S. Agerbaek et al., https://doi.org/10.21203/rs.3.rs-6700810/v1 (under review) The data provided here represent the original genome assemblies and annotations generated prior to any downstream processing, and are deposited to ensure full reproducibility of the results reported in the study. Clean assemblies and raw reads are available under the Bioproject: PRJNA126044, which also contains metadata about the strains sequenced. Note that some genomes were excluded from the analysis following the initial upload, which is why there are fewer annotations here than assemblies in the Bioproject. Annotations available in this repo under Annotations/species_ID were annotated with Funannotate v1.8.13. Functional annotations were generated with Eggnog-mapper v2.1.7, InterProScan v5.56-89.0, antiSMASH v6.0.0 and SignalP v6.0. Content For each genome, the following files are provided: .scaffolds.fa Initial assembly .gff3 Structural gene annotation including gene, transcript, and CDS features .proteins.fa Predicted protein sequences .rn.proteins.fasta Predicted protein sequences with re-named fasta headers containing functional annotation in the format: >Species | Strain_ID | Gene accession | Gene name | Gene description | COG | Antismash | GO term | EC number(s) | Pfam term(s) | Interpro term(s) .cds-transcripts.fa Nucleotide coding sequences .annotations.txt Functional annotation summary table .stats.json Assembly and annotation statistics generated during the annotation process For questions regarding these data, please contact: Sofie Agerbaek, sofie.agerbaek@sund.ku.dk
### 数据集说明 本仓库包含本论文中完成测序与分析的皮肤关联真菌的基因组注释信息:《临床相关人类皮肤关联真菌的比较基因组学分析》(S. Agerbaek等,https://doi.org/10.21203/rs.3.rs-6700810/v1,处于审稿阶段)。 本文提供的数据为下游处理前生成的原始基因组组装结果与注释信息,旨在确保本研究报告结果的完全可复现性。 经过清理的组装序列与原始测序读段可在生物项目(Bioproject)PRJNA126044下获取,该项目同时包含测序菌株的元数据。请注意,部分基因组在初始上传后被排除出本次分析,因此本仓库中的注释条目数少于生物项目中的组装序列数。 本仓库`Annotations/species_ID`路径下的注释信息由Funannotate v1.8.13完成。功能注释通过Eggnog-mapper v2.1.7、InterProScan v5.56-89.0、antiSMASH v6.0.0与SignalP v6.0生成。 ### 内容说明 针对每个基因组,提供以下文件: 1. `.scaffolds.fa`:初始组装序列 2. `.gff3`:结构基因注释文件,包含基因、转录本与编码序列(CDS)特征 3. `.proteins.fa`:预测得到的蛋白质序列 4. `.rn.proteins.fasta`:经过重命名FASTA头部的预测蛋白质序列,头部格式为:>物种 | 菌株ID | 基因登录号 | 基因名称 | 基因描述 | COG | Antismash | GO术语 | EC编号 | Pfam术语 | InterPro术语 5. `.cds-transcripts.fa`:核苷酸编码序列 6. `.annotations.txt`:功能注释汇总表 7. `.stats.json`:注释过程中生成的组装与注释统计信息 若对这些数据有疑问,请联系: Sofie Agerbaek,邮箱:sofie.agerbaek@sund.ku.dk



