Single-nucleotide polymorphisms, genome assemblies, genome annotations, and gene predictions of Pyricularia oryzae isolates from rice
收藏资源简介:
We analyzed genetic diversity in isolates of the rice blast fungus (<em>Pyricularia oryzae</em>) covering a broad geographical range. We used genotyping data (Infinium beadchip) to infer population structure and whole genome resequencing data (Illumina sequencing) to investigate differences in repertoires of putative pathogenicity effectors. List of files: isolates_with_genotyping_data.xlsx: isolates with Infinium genotyping data isolates_with_resequencing_data.xlsx: isolates with whole genome sequencing data (including European Nucleotide Archive identifiers). Infinium-and-sequencing_SNPs_lineage1.txt: Allelic states of 204 P. oryzae isolates from lineage 1, genotyped at 3,686 markers using Infinium genotyping or whole genome sequencing. Genomic coordinates (#CHROM and POS) indicate position of the markers in the 70-15 reference genome (EnsemblFungi, assembly MG8). Line 'clonal_groups' shows the assignment of isolates to groups of multilocus genotypes repeated multiple times. Infinium-and-sequencing_SNPs.txt: Allelic states of 123 P. oryzae isolates genotyped at 3,686 markers using Infinium genotyping or whole genome sequencing. Genomic coordinates (#CHROM and POS) indicate position of the markers in the 70-15 reference genome (EnsemblFungi, assembly MG8). Assembly.zip: Fasta-format genome assemblies. Low-quality reads were removed using the software CUTADAPT. Reads were assembled using ABYSS 2.2.3 using different K-mer sizes and for each isolate we chose the assembled sequence with the highest N50 for further analyses. Repeated regions were masked using REPEATMASKER 4.1.0. Annotation.zip: GFF-format gene models, and corresponding protein and DNA sequences. Genes were predicted with BRAKER 2.1.5 using RNAseq data (Pordel et al. 2020 doi:10.1094/PHYTO-09-20-0423-R) as extrinsic evidence for model refinement. Genes were also predicted with AUGUSTUS 3.4.0 (training set=Magnaporthe grisea) and gene models that did not overlap with gene models identified with BRAKER were added to the GFF file generated with the latter. secretome_aln.zip: Alignment of all groups of orthologs corresponding to putative effector proteins. Homology relationships among predicted genes were established using ORTHOFINDER v2.4.0. Putative effector genes were identified as genes encoding proteins predicted to be secreted by at least two methods among three (SIGNALP 4.1, TARGETP and PHOBIUS), without predicted transmembrane domain based on TMHMM analysis, without predicted motif of retention in the endoplasmic reticulum based on PS-SCAN, and without CAZy annotation based on DBSCAN V7. Sequences for each group of orthologs were aligned and cleaned with TRANSLATORX using default parameters. non-secretome_aln.zip: Alignment of all groups of orthologs corresponding to genes that are not putative effector proteins (i.e. the portion of the gene space which is the complement of what is included in secretome_aln.zip) More details in https://doi.org/10.1101/2020.06.02.129296
本研究对覆盖广泛地理分布范围的稻瘟病菌(<em>Pyricularia oryzae</em>)分离株开展遗传多样性分析。本研究利用基因分型数据(Infinium芯片(Infinium beadchip))推断种群遗传结构,并通过全基因组重测序数据(Illumina测序(Illumina sequencing))探究潜在致病效应因子的组间差异。 文件清单如下: 1. isolates_with_genotyping_data.xlsx:携带Infinium基因分型数据的稻瘟病菌分离株数据集 2. isolates_with_resequencing_data.xlsx:携带全基因组测序数据(包含欧洲核苷酸档案库(European Nucleotide Archive)标识符)的稻瘟病菌分离株数据集 3. Infinium-and-sequencing_SNPs_lineage1.txt:来自谱系1的204株P. oryzae分离株的等位基因分型数据,通过Infinium基因分型或全基因组测序在3686个标记位点完成分型。其基因组坐标(#CHROM与POS)对应70-15参考基因组(EnsemblFungi数据库,组装版本MG8)中的标记位置。字段'clonal_groups'标注了分离株被划分为多位点基因型重复出现的克隆群的归属情况 4. Infinium-and-sequencing_SNPs.txt:123株P. oryzae分离株的等位基因分型数据,通过Infinium基因分型或全基因组测序在3686个标记位点完成分型。其基因组坐标(#CHROM与POS)对应70-15参考基因组(EnsemblFungi数据库,组装版本MG8)中的标记位置 5. Assembly.zip:Fasta格式的基因组组装序列。测序前使用CUTADAPT软件去除低质量reads,采用ABYSS 2.2.3软件结合不同K-mer长度对reads进行组装,并为每个分离株选取N50值最高的组装序列用于后续分析。使用REPEATMASKER 4.1.0软件屏蔽基因组中的重复序列区域 6. Annotation.zip:GFF格式的基因模型文件,以及对应的蛋白质与DNA序列。本研究使用BRAKER 2.1.5软件进行基因预测,以RNA-seq数据(Pordel等,2020,doi:10.1094/PHYTO-09-20-0423-R)作为外部证据优化基因预测模型;同时使用AUGUSTUS 3.4.0软件(训练集为Magnaporthe grisea)进行基因预测,并将未与BRAKER预测结果重叠的基因模型添加至BRAKER生成的GFF注释文件中 7. secretome_aln.zip:所有对应潜在致病效应因子蛋白的同源基因簇的比对序列。使用ORTHOFINDER v2.4.0软件确定预测基因间的同源关系。潜在致病效应因子基因的筛选标准为:编码的蛋白可被至少2种(共3种预测方法:SIGNALP 4.1、TARGETP与PHOBIUS)鉴定为分泌蛋白,经TMHMM分析无跨膜结构域,经PS-SCAN分析无内质网滞留基序,且无基于DBSCAN V7数据库的CAZy注释。每个同源基因簇的序列使用TRANSLATORX软件(默认参数)进行多序列比对与序列清洗 8. non-secretome_aln.zip:所有对应非潜在致病效应因子基因的同源基因簇的比对序列(即与secretome_aln.zip包含的基因空间互补的基因组区域) 更多详细信息请参见https://doi.org/10.1101/2020.06.02.129296



