遇见数据集

Supporting data for "Chromosome-level reference genome for the medically important Arabian horned viper (<i>Cerastes gasperettii</i>)"

收藏
DataCite Commons2025-06-09 更新2025-04-15 收录
数据链接:
官方服务:

资源简介:

Venoms have traditionally been studied from a proteomic and/or transcriptomic perspective, often overlooking the true genetic complexity underlying venom production. The recent surge in genome-based venom research (sometimes called venomics) has proven to be instrumental in deepening our molecular understanding of venom evolution, particularly through the identification and mapping of toxin-coding loci across the broader chromosomal architecture. Although venomous snakes are a model system in venom research, the number of high-quality reference genomes in the group remains limited. In this study, we present a chromosome-resolution reference genome for the Arabian horned viper (<i>Cerastes gasperettii</i>), a venomous snake native to the Arabian Peninsula. Our highly-contiguous genome allowed us to explore macrochromosomal rearrangements within the Viperidae family, as well as across squamates. We identified the main highly-expressed toxin genes compousing the venoms core, in line with our proteomic results. We also compared microsyntenic changes in the main toxin gene clusters with those of other venomous snake species, highlighting the pivotal role of gene duplication and loss in the emergence and diversification of Snake Venom Metalloproteinases (SVMPs) and Snake Venom Serine Proteases (SVSPs) for <i>Cerastes gasperettii</i>. Using Illumina short-read sequencing data, we reconstructed the demographic history and genome-wide diversity of the species, revealing how historical aridity likely drove population expansions. Finally, this study highlights the importance of using long-read sequencing as well as chromosome-level reference genomes to disentangle the origin and diversification of toxin gene families in venomous species.

长期以来,毒液研究多从蛋白质组学(proteomics)或转录组学(transcriptomics)视角展开,却往往忽视了毒液生成背后的真实遗传复杂性。近年来,基于基因组的毒液研究(有时亦称毒液基因组学(venomics))蓬勃发展,现已被证实对深化我们对毒液演化的分子认知起到了关键作用,尤其是通过在更广泛的染色体架构中识别并定位毒素编码基因座(loci)。尽管有毒蛇类是毒液研究的经典模式系统,但该类群的高质量参考基因组数量仍然有限。本研究报道了阿拉伯角蝰(*Cerastes gasperettii*)的染色体级参考基因组,该物种为原产于阿拉伯半岛的有毒蛇类。我们的高连续性基因组使得我们能够探究蝰科(Viperidae)类群内部以及有鳞目(Squamata)中的大染色体重排事件。我们鉴定出了构成毒液核心的主要高表达毒素基因,该结果与我们的蛋白质组学分析相符。我们还将主要毒素基因簇中的微共线性变化与其他有毒蛇类进行了对比,阐明了基因重复与缺失在阿拉伯角蝰的蛇毒金属蛋白酶(Snake Venom Metalloproteinases,SVMPs)与蛇毒丝氨酸蛋白酶(Snake Venom Serine Proteases,SVSPs)的起源与多样化进程中所发挥的关键作用。我们利用Illumina短读长测序数据重构了该物种的种群历史与全基因组多样性特征,揭示了历史时期的干旱环境如何推动了其种群扩张。最后,本研究强调了结合长读长测序技术与染色体级参考基因组,以厘清有毒物种毒素基因家族的起源与多样化机制的重要性。

提供机构:
GigaScience Database
创建时间:
2025-02-21
二维码
社区交流群
二维码
科研交流群
商业服务