遇见数据集

Simulated nucleotide sequences for testing alignment-free genome distance estimates

收藏
Zenodo2020-09-10 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This repository contains (12×500=)6,000 pairs of nucleotide sequences that have been simulated for testing alignment-free genome distance estimates, as described in Criscuolo (2019). Given an evolutionary distance <em>d</em> varying from 0.05 to 0.60 (step = 0.05), the program SeqGen was used to simulate the evolution of 500 nucleotide sequence pairs with <em>d</em> substitution events per character (GTR+Γ evolutionary model). For each of the 12 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 0.60, an XZ-compressed file containing 500 lines is available. Each line contains 18 fields separated by blank spaces:<br> [1] seed value used during simulation,<br> [2] true evolutionary distance <em>d</em> between the two simulated sequences,<br> [3] total number of simulated characters,<br> [4] number of non-indel characters with nucleotide mismatch,<br> [5] number of non-indel characters,<br> [6-9] A, C, G, T frequencies used during simulation,<br> [10-15] GTR parameters used during simulation,<br> [16] Γ distribution parameter used during simulation,<br> [17-18] two simulated sequences with indel events as gaps. Of note, each pair of aligned sequences without gaps can be regenerated using SeqGen v1.3.4 with parameters from fields [1,3,6-16] and the following two-leaf model tree: <pre>(t1:d,t2:0.000);</pre> where <em>d</em> is given in field [2]. ___ Criscuolo A (2019) <em>A fast alignment-free bioinformatics procedure to infer accurate distance-based phylogenetic trees from genome assemblies</em>. Research Ideas and Outcomes, 5:e36178. doi:10.3897/rio.5.e36178

本数据集仓库包含(12×500=)6000对核苷酸序列,均为测试无比对基因组距离估算(alignment-free genome distance estimates)方法而模拟生成,相关细节详见Criscuolo(2019)的研究。当进化距离*d*取值范围为0.05至0.60(步长0.05)时,使用SeqGen程序按照广义时间可逆+Γ(GTR+Γ)进化模型,为每个进化距离模拟生成500对核苷酸序列,每对序列的每个位点发生*d*次替换事件。针对12个进化距离值*d*=0.05、0.10、……、0.60,各提供一个经XZ压缩的数据文件,每个文件包含500行数据。每行包含18个以空格分隔的字段: [1] 模拟过程中使用的随机种子值 [2] 两条模拟序列间的真实进化距离*d* [3] 模拟产生的总位点数 [4] 无插入缺失(insertion-deletion, indel)且存在核苷酸错配的位点数 [5] 无插入缺失的位点数 [6-9] 模拟过程中使用的A、C、G、T碱基频率 [10-15] 模拟过程中使用的GTR参数 [16] 模拟过程中使用的Γ分布参数 [17-18] 两条带有插入缺失间隙的模拟序列。值得注意的是,若使用SeqGen v1.3.4,并结合字段[1,3,6-16]中的参数以及如下两叶系统发育树模型: (t1:d,t2:0.000); 其中*d*即为字段[2]中给出的数值,即可重新生成每一对不含间隙的比对序列。___ Criscuolo A (2019) 《一种快速无比对生物信息学方法,用于从基因组组装结果中推断基于距离的准确系统发育树》。Research Ideas and Outcomes, 5:e36178. doi:10.3897/rio.5.e36178

提供机构:
Zenodo
创建时间:
2020-09-10
二维码
社区交流群
二维码
科研交流群
商业服务