遇见数据集

Transposable element annotation Rhynchosporium commune isolate UK7

收藏
Zenodo2022-02-08 更新2026-06-04 收录
数据链接:
官方服务:

资源简介:

To obtain a consensus sequence for each TE family, RepeatModeler v. open-4.0.7 (http://www.repeatmasker.org/RepeatModeler/) was run on the <em>R. commune</em> UK7 reference genome. The classification was based on the GIRI Repbase (v. 2018) using RepeatMasker v. open-4.0.7. (Smit, Hubley, and P. 2015; Bao, Kojima, and Kohany 2015). We used WICKERsoft to finalize the classification of TE consensus sequences (Breen et al. 2010). Specifically, we used WICKERsoft to screen for copies of known consensus sequences from other fungal species with blastn filtering for sequence identity &gt; 80% and sequence length &gt; 80%. (Altschul et al. 1997). Then, using WICKERsoft, flanks of 10000 bp were added and visually inspected for sequence similarity and terminal repeats with dot plots. Subsequent multiple sequence alignments were performed with 10-15 sequences using ClustalW (Thompson, Higgins, and Gibson 1994). Alignment boundaries were visually inspected in WICKERsoft and trimmed if necessary. Using WICKERsoft, consensus sequences were classified according to the presence and type of terminal repeats, as well as homology of the encoded proteins based on blastx against the NCBI protein database. Consensus sequences were named according to the three-letter classification system (Wicker et al. 2007). The reference genome was annotated with the curated consensus sequences using RepeatMasker v. open-4.0.7 with a cut-off value of 250 (Smit, Hubley, and P. 2015). Simple repeats, low complexity regions and annotated elements shorter than 100 bp were filtered out and adjacent identical TEs overlapping by more than 100 bp were merged as belonging to the same TE family. Different TE families overlapping by more than 100 bp were considered as nested insertions and were renamed accordingly. Identical elements separated by less than 200 bp are indicative of interrupted elements and were grouped into a single element. TEs overlapping genes were recovered using the bedtools v. 2.27.1 suite and the “overlap” function (Quinlan and Hall 2010).

为获取每个转座子(Transposable Element,TE)家族的共有序列,我们以<em>R. commune</em> UK7参考基因组为研究对象,运行RepeatModeler v. open-4.0.7(http://www.repeatmasker.org/RepeatModeler/)进行分析。该分类基于GIRI Repbase数据库(v. 2018),并使用RepeatMasker v. open-4.0.7完成(Smit、Hubley与P. 2015;Bao、Kojima与Kohany 2015)。我们使用WICKERsoft对TE共有序列的分类进行最终确认(Breen et al. 2010)。具体而言,我们通过WICKERsoft,以BLASTN筛选序列一致性>80%、比对覆盖长度>80%的其他真菌物种已知共有序列拷贝(Altschul et al. 1997)。随后,借助WICKERsoft添加10000 bp的侧翼序列,并通过点图法目视检查序列相似性与末端重复序列。使用ClustalW对10~15条序列进行多序列比对(Thompson、Higgins与Gibson 1994),在WICKERsoft中目视检查比对边界,必要时进行修剪。通过WICKERsoft,依据末端重复序列的存在类型以及编码蛋白的同源性(基于BLASTX比对NCBI蛋白质数据库)对共有序列进行分类。共有序列按照Wicker等人2007年提出的三字母分类体系进行命名。使用阈值为250的RepeatMasker v. open-4.0.7,以整理得到的共有序列对参考基因组进行注释(Smit、Hubley与P. 2015)。过滤去除简单重复序列、低复杂度区域以及长度小于100 bp的注释元件;将相邻且重叠区域超过100 bp的相同TE家族元件合并为同一TE家族。不同TE家族间重叠区域超过100 bp的情况视为嵌套插入,并据此重新命名。间隔小于200 bp的相同元件视为中断型元件,将其合并为单个元件。使用bedtools v. 2.27.1套件的“重叠分析”功能,提取与基因存在重叠的TE序列(Quinlan与Hall 2010)。

提供机构:
Zenodo
创建时间:
2022-02-08
二维码
社区交流群
二维码
科研交流群
商业服务