遇见数据集

DNA methylation dynamics during stress-response in woodland strawberry (Fragaria vesca)

收藏
Zenodo2022-03-04 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Genome sequence and annotation of Fragaria vesca cv. Reine des Vallées</strong> In order to generate a reference genome for Fragaria vesca cv. Reine des Vallées, we used MinIon long-read sequencing data to substitute the <em>F. vesca</em> genome v.4.0.a2 genome. The detailed method used to obtain these results were the following: <em>Genome sequencing and assembly NIL Fb2</em> Genomic DNA from strawberry plants was extracted by a Hexadecyltrimethylammonium bromide (Cetrimonium bromide, CTAB) modified protocol (Healey, Furtado, Cooper, &amp; Henry, 2014) and purified with Agencourt AMPure XP beads (cat# A63880). Long-read sequencing was performed for the genome assembly; Genomic DNA by Ligation (Oxford Nanopore, cat# SQK-LSK109) library was prepared as described by the manufacturer and sequenced on a MinION for 72 h (Oxford Nanopore). <em>Reference genome polishing</em> Reads obtained from nanopore were filtered with Filtlong v0.2.1 (https://github.com/rrwick/Filtlong) using --min_mean_q 80 and --min_length 200. Cleaned reads were then aligned to the most recent version of the <em>F. vesca</em> genome v4.0.a2, downloaded from the Genome Database for Rosaceae (GDR) (https://www.rosaceae.org/species/fragaria_vesca/genome_v4.0.a2), using minimap2 v2.21 (H. Li, 2018) with parameters -aLx map-ont --MD -Y. The generated BAM file was then sorted and indexed with samtools v1.11 (H. Li et al., 2009). We used mosdepth v0.3.1 (Pedersen &amp; Quinlan, 2018) to verify that coverage on chromosomic scaffolds was over 50 X. Sniffles v1.0.12a (Sedlazeck et al., 2018) with parameters -s 10 -r 1000 -q 20 --genotype -l 30 -d 1000 was used to detect structural variations larger than 30 bp. The VCF files obtained from Sniffles was sorted and filtered with BCFtools v1.14 (Danecek et al., 2021) to keep only structural variants (SV) with smaller than 200,00 bp (we observed that larger SV were most of the time false positive caused by misalignments in regions with gaps or Ns), supported by 10 or more reads and with allelic frequencies above 0.8 (we were interested in homozygous changes). The complete filtering command used is “bcftools view -q 0.8 -Oz -i '(SVTYPE = "DUP" || SVTYPE = "INS" || SVTYPE = "DEL" || SVTYPE = "TRA" || SVTYPE = "INV" || SVTYPE = "INVDUP") &amp;&amp; %FILTER = "PASS" &amp;&amp; FMT/DV&gt;9 &amp;&amp; SVLEN&gt;29 &amp;&amp; SVLEN&lt;200000' “ From the VCF listing all the structural variants that we detected in our <em>F. vesca </em>accession, we generated a substituted genome version based on the reference <em>F. vesca</em> genome v.4.0.a2. The reference genome was first indexed with samtools faidx v1.11(Danecek et al., 2021) and a sequence dictionary was generated with Picard CreateSequenceDictionary v2.25.6 (https://broadinstitute.github.io/picard). The VCF containing the SV produced from our Nanopore sequencing was also indexed with gatk (Van der Auwera GA &amp; O'Connor BD, 2020) IndexFeatureFile v4.2.0.0 (https://gatk.broadinstitute.org/hc/en-us/articles/360037262651-IndexFeatureFile). FastaAlternateReferenceMaker v4.2.0.0 (https://gatk.broadinstitute.org/hc/en-us/articles/360037594571-FastaAlternateReferenceMaker) was then run with the reference genome and the VCF file to generate a substituted genome representative of our <em>Fragaria</em> accession. As substituting our genome with the detected structural variants changes genomic coordinates, we also corrected the public GFF genome annotation of <em>F. vesca</em> (Y, Pi, Gao, Liu, &amp; Kang, 2019) using liftoff v1.6.1 (Shumate &amp; Salzberg, 2021). Liftoff also detects and annotates duplications within the substituted genome. Transposable elements annotation was carried out using the EDTA transposable element annotation pipeline v. 1.9.6 (S. Ou et al., 2019) on the substituted genome using default parameters<em>.</em> <strong>Differentially methylated regions</strong> The file Stress_vs_control_DMRs.zip file contains the DMRs that were called using the reads submitted to ENA (ERP135585) and obtained as follows: First, bedGraph files from wgbs pipeline were pre-filtered for a minimum coverage of 5 reads using awk command. These output files were then used as input for the EpiDiverse/dmr bioinformatics analysis pipeline for non-model plant species to define DMRs (Nunn <em>et al</em>., 2021) with default parameters (minimum coverage threshold 5; maximum q-value 0.05; minimum differential methylation level 10%; 10 as minimum number of Cs; Minimum distance (bp) between Cs that are not to be considered as part of the same DMR is 146 bp). The pipeline uses metilene v.0.2.6.1 (https://www.bioinf.uni-leipzig.de/Software/metilene/) for pairwise comparison between groups and R-packages ggplot2 v.3.3.5 and gplots v.3.1.1, for visualization results (Fig. S1). Based on our <em>F. vesca</em> genome transcript annotation and methylation data (overlapped regions with DNA methylation cytosines and DMRs), we detected the methylated genes, promoters, 3’ UTRs, 5’UTR and transposable elements in strawberry. Global DNA methylation and DMR plots were performed with R-package ggplot2. Gene analyses by methylation patterns and analysis of per-family TE DNA methylation profiles were performed with deepTools v.3.5.0 (Ramírez <em>et al</em>., 2014). DMRs comparison between treatments were done by the Venn diagram v.1.7.0 R-package. We produced several genome browsers tracks with DMRs that we integrated in our local instance of JBrowse available at the following url: https://jbrowse.agroscope.info/jbrowse/?data=fragaria_sub

<strong>森林草莓(Fragaria vesca)‘Reine des Vallées’品种的基因组测序与注释</strong> 为构建森林草莓‘Reine des Vallées’的参考基因组,本研究利用MinION长读长测序数据替代了原版F. vesca基因组v4.0.a2版本。获取该结果的详细方法如下: <em>基因组测序与组装(近等基因系Fb2)</em> 草莓植株的基因组DNA采用十六烷基三甲基溴化铵(Hexadecyltrimethylammonium bromide,又称溴代十六烷基三甲胺,简称CTAB)改良方案提取(Healey, Furtado, Cooper, & Henry, 2014),并通过Agencourt AMPure XP磁珠(货号A63880)进行纯化。长读长测序用于基因组组装:基因组DNA连接建库试剂盒(Genomic DNA by Ligation,牛津纳米孔科技,货号SQK-LSK109)按照制造商说明书制备文库,并在MinION测序仪上测序72小时(牛津纳米孔科技)。 <em>参考基因组抛光</em> 通过纳米孔测序获得的reads使用Filtlong v0.2.1(https://github.com/rrwick/Filtlong)进行过滤,过滤参数为--min_mean_q 80与--min_length 200。将清理后的reads比对至最新版F. vesca基因组v4.0.a2,该基因组从蔷薇科基因组数据库(Genome Database for Rosaceae, GDR,https://www.rosaceae.org/species/fragaria_vesca/genome_v4.0.a2)下载,比对工具为minimap2 v2.21(H. Li, 2018),参数为-aLx map-ont --MD -Y。生成的BAM文件随后通过samtools v1.11(H. Li等, 2009)进行排序与索引。本研究使用mosdepth v0.3.1(Pedersen & Quinlan, 2018)验证染色体支架的覆盖度均超过50×。使用Sniffles v1.0.12a(Sedlazeck等, 2018),参数为-s 10 -r 1000 -q 20 --genotype -l 30 -d 1000,以检测长度大于30 bp的结构变异(SV)。通过BCFtools v1.14(Danecek等, 2021)对Sniffles生成的VCF文件进行排序与过滤,仅保留符合以下条件的结构变异:长度小于200000 bp(本研究观察到更大的结构变异多为间隙或N碱基区域比对错误导致的假阳性)、支持reads数≥10、等位基因频率≥0.8(本研究关注纯合变异)。完整的过滤命令为:`bcftools view -q 0.8 -Oz -i '(SVTYPE = "DUP" || SVTYPE = "INS" || SVTYPE = "DEL" || SVTYPE = "TRA" || SVTYPE = "INV" || SVTYPE = "INVDUP") && %FILTER = "PASS" && FMT/DV>9 && SVLEN>29 && SVLEN<200000'` 基于在本研究的F. vesca材料中检测到的所有结构变异VCF列表,本研究以原版F. vesca基因组v4.0.a2为参考,生成了替代版基因组。首先使用samtools faidx v1.11(Danecek等, 2021)对参考基因组建立索引,并通过Picard CreateSequenceDictionary v2.25.6(https://broadinstitute.github.io/picard)生成序列字典。将基于纳米孔测序获得的结构变异VCF文件通过GATK IndexFeatureFile v4.2.0.0(https://gatk.broadinstitute.org/hc/en-us/articles/360037262651-IndexFeatureFile)建立索引。随后使用FastaAlternateReferenceMaker v4.2.0.0(https://gatk.broadinstitute.org/hc/en-us/articles/360037594571-FastaAlternateReferenceMaker)结合参考基因组与VCF文件,生成代表本研究Fragaria属材料的替代版基因组。由于通过检测到的结构变异替换基因组序列会改变基因组坐标,本研究同时使用liftoff v1.6.1(Shumate & Salzberg, 2021)对公开的F. vesca GFF基因组注释(Y, Pi, Gao, Liu, & Kang, 2019)进行坐标校正。Liftoff还可在替代版基因组中检测并注释重复序列。转座元件注释通过EDTA转座元件注释流程v1.9.6(S. Ou等, 2019)完成,参数使用默认值,分析对象为替代版基因组。 <em>差异甲基化区域</em> Stress_vs_control_DMRs.zip文件包含通过提交至欧洲核苷酸档案馆(ERP135585)的reads所鉴定的差异甲基化区域(DMRs),获取流程如下:首先,使用awk命令对全基因组亚硫酸氢盐测序(WGBS)流程生成的bedGraph文件进行预过滤,保留覆盖度≥5的位点。将输出文件作为输入,使用EpiDiverse/dmr非模式植物生物信息学分析流程鉴定DMRs(Nunn等, 2021),参数使用默认值:最低覆盖度阈值5;最大q值0.05;最低差异甲基化水平10%;最少胞嘧啶(C)位点数量10;不属于同一DMR的胞嘧啶间最小距离为146 bp。该流程使用metilene v0.2.6.1(https://www.bioinf.uni-leipzig.de/Software/metilene/)进行组间两两比较,并使用R包ggplot2 v3.3.5与gplots v3.1.1进行结果可视化(图S1)。基于本研究的F. vesca基因组转录组注释与甲基化数据(与DNA甲基化胞嘧啶位点及DMRs重叠的区域),本研究鉴定了草莓中的甲基化基因、启动子、3’UTR、5’UTR与转座元件。全基因组DNA甲基化与DMR绘图通过R包ggplot2完成。基于甲基化模式的基因分析与转座元件家族特异性DNA甲基化谱分析通过deepTools v3.5.0(Ramírez等, 2014)完成。处理组间DMRs的比较通过R包Venn diagram v1.7.0完成。本研究生成了多个包含DMRs的基因组浏览器轨道,并整合至本地部署的JBrowse中,访问地址为:https://jbrowse.agroscope.info/jbrowse/?data=fragaria_sub

提供机构:
Zenodo
创建时间:
2022-03-04
二维码
社区交流群
二维码
科研交流群
商业服务