遇见数据集

A VCF file of all whole-genome sequenced heliothine indviduals, aligned to BACs (not those containing CYP337B1, 2 or 3) available on NCBI

收藏
DataONE2016-09-28 更新2024-06-26 收录
数据链接:
官方服务:

资源简介:

BAC descriptions are available in the supplementary document. Heliothine moths were collected between 2004 and 2014 from 16 different countries around the world across various climatic zones and altitudes (Tables S1 and S2), many of which are described in Behere et al. (2007); and Tay et al. (2013). Samples were collected as larvae from wild and crop host plants, as adult moths via light/pheromone traps, or as larvae after bioassay, and preserved in ethanol (>95%) or RNAlater, or stored at -20°C prior to DNA extraction. DNA was extracted from samples using DNeasy blood and tissue kits (Qiagen), before being quantified with a Qubit 2.0. Nextera libraries were produced following the manufacturer’s instructions and sequence was generated as 100 bp PE reads (Illumina HiSeq 2000, Biological Resources Facility, Australian National University, Canberra, Australia, as well as at Beijing Genomics Institute, Hong Kong). Sample and sequencing data are included in the supplementary material (Table S2). Raw reads were aligned to BAC sequences, originally derived from H. armigera and available on NCBI (accessions in supplementary document), using BBMap. Reads were trimmed when quality in at least 2 bases fell below Q10. Only uniquely aligning reads were included in the analysis, to prevent spuriously inferring evolutionary processes occurring independently on each BAC. Outputted BAM files were sorted before duplicate reads were removed and files were annotated with read groups using Picard v. 1.138 (http://picard.sourceforge.net). BAC reference sequences were indexed using Samtools v. 1.1.0 (Li et al. 2009). UnifiedGenotyper in GATK v. 3.3-0 (McKenna et al. 2010) was used to estimate genotypes across all individuals simultaneously, implementing a heterozygosity value of 0.01. Variant call format files containing SNP calls were reformatted into Plink format using VCFtools v. 0.1.12b (Danecek et al. 2011).

细菌人工染色体(Bacterial Artificial Chromosome,BAC)的描述详见补充文档。实夜蛾亚科蛾类样本于2004年至2014年间采集自全球16个不同国家,覆盖多样气候带与海拔梯度(详见附表S1、S2),其中多数样本的采集背景已在Behere等(2007)及Tay等(2013)的研究中详述。样本采集方式包括:从野生及作物寄主植物上采集幼虫、利用灯光/性诱捕器采集成虫,或经生物测定后采集幼虫;样本保存方式为置于体积分数≥95%的乙醇或RNAlater中,或在DNA提取前于-20℃冷冻保存。使用DNeasy血液与组织试剂盒(Qiagen)提取样本基因组DNA,随后通过Qubit 2.0荧光定量仪进行浓度定量。按照试剂盒说明书构建Nextera文库,并采用Illumina HiSeq 2000测序平台生成100 bp双端读段(PE reads),测序工作分别在澳大利亚堪培拉澳大利亚国立大学生物资源研究中心以及香港华大基因完成。样本信息与测序数据详见补充材料(附表S2)。原始测序读段使用BBMap软件比对至BAC序列,该序列最初来自棉铃虫(Helicoverpa armigera),可在NCBI数据库获取(序列登录号详见补充文档)。当至少2个连续碱基的质量值低于Q10时,将对测序读段进行修剪。仅保留唯一比对上的读段用于后续分析,以避免错误推断独立发生于各BAC的进化过程。输出的BAM文件先进行排序,随后移除重复读段,并使用Picard v.1.138版本(http://picard.sourceforge.net)为文件添加读段组注释信息。使用Samtools v.1.1.0版本(Li等,2009)对BAC参考序列进行索引构建。使用基因组分析工具包(Genome Analysis Toolkit,GATK)v.3.3-0版本(McKenna等,2010)中的UnifiedGenotyper工具,同时对所有个体的基因型进行估计,并设置杂合度参数为0.01。使用VCFtools v.0.1.12b版本(Danecek等,2011)将包含单核苷酸多态性(Single Nucleotide Polymorphism,SNP)检测结果的变异呼叫格式(VCF)文件重新格式化为Plink格式文件。

创建时间:
2016-09-28
二维码
社区交流群
二维码
科研交流群
商业服务