遇见数据集

<p>SNPs in spots for K = 6.</p>

收藏
NIAID Data Ecosystem2026-05-10 收录
官方服务:

资源简介:

Inferring the genetic structure at the subpopulation level is crucial for understanding the demographic histories that shape genetic diversity. Among the most widely used approaches are methods based on admixture and structure modeling—named after the respective software tools—which have become standard due to their intuitive, interpretable outputs. In this study, we address a key methodological question: how does the traditional admixture-based decomposition of genetic components in multilocus population data relate to clustering approaches that leverage machine learning, specifically Self-Organizing Maps (SOMs)? We implemented this approach through our custom SOM-based tool, SOMmelier, which enables the portrayal of genetic structure by identifying modules of co-mutated SNPs and arranging them in a topology-aware genetic landscape. Topology-awareness refers to the organization of genetic modules in a two-dimensional map, where their spatial proximity reflects mutual similarity. We applied Admixture and SOMmelier to investigate the population genetics of European grapevine. Based on prior literature, we considered up to six genetic components, which formed a genetic landscape that closely mirrors the geographic expanse of the classical Mediterranean world—from Western Asia through the Caucasus to Western Europe. The resulting topology reflects the dynamic spatial and temporal nature of grapevine domestication and diffusion. We demonstrate that SOMmelier can recover the genetic components identified by Admixture solely through statistical clustering. By integrating the topological structure of SNP co-variation, it offers perspectives on population structure, evolutionary history, and trait associations in grapevine—and has applicability to other species and systems in population genetics.

推断亚种群水平的遗传结构,对于解析塑造遗传多样性的种群历史至关重要。目前应用最为广泛的方法之一,是基于混合模型(ADMIXTURE)与群体结构分析(STRUCTURE)的建模方法——这类方法以对应的软件工具命名,凭借其直观且可解释的输出结果,已成为群体遗传学研究的主流标准方法。本研究旨在解决一个关键的方法学问题:基于传统混合模型的多位点种群数据遗传组分分解方法,与基于机器学习的聚类方法——尤其是自组织映射(Self-Organizing Maps, SOMs)——之间存在何种关联?我们通过自主开发的基于SOM的工具SOMmelier实现了该方法:该工具可通过识别共突变单核苷酸多态性(Single Nucleotide Polymorphism, SNPs)模块,并将其排布于拓扑感知的遗传景观中,实现遗传结构的可视化呈现。所谓拓扑感知,指的是将遗传模块排布于二维映射空间中,模块间的空间邻近性可反映其相互相似性。我们运用ADMIXTURE与SOMmelier两种工具,对欧洲葡萄的种群遗传学展开研究。基于已有文献基础,我们最多考虑了6个遗传组分,由此构建的遗传景观与古典地中海世界的地理范围高度吻合——从西亚经高加索地区延伸至西欧。该拓扑结构可反映葡萄驯化与传播过程中动态的时空特征。本研究证明,SOMmelier仅通过统计聚类即可复现ADMIXTURE所识别的遗传组分。通过整合SNP共变异的拓扑结构,该工具可为葡萄的种群结构、进化历史以及性状关联研究提供全新视角,同时也可推广应用于种群遗传学领域的其他物种与研究体系。

创建时间:
2026-02-20
二维码
社区交流群
二维码
科研交流群
商业服务