遇见数据集

Intensification

收藏
知名数据库2026-06-11 收录
官方服务:

资源简介:

Intensification is a database that contains the results for 12 repeat protein domains, from the amplification of population-genetic signal by constructing a motif-based multiple sequence alignment (motif-MSA). Because protein-coding regions are typically under high selective constraints, these variants occur at low frequencies, such that there is often insufficient statistics for downstream calculations. We make use of the modular structure of repeat motifs to amplify signals of selection from population genetics and traditional inter-species conservation. For each repeat protein repeat domain, we construct a motif-MSA and then accumulate single nucleotide variants (SNVs) across the human population based on the genomic coordinate system of the motif-MSA. This allows us to integrate all the corresponding SNV population-genetic profiles, including enrichment of rare variants, non-synonymous-to-synonymous ratio and delta DAFs, with the amino acid variation across the motif-MSA

Intensification数据库收录了12个重复蛋白质的重复结构域的相关研究结果,其数据源自通过构建基于基序的多序列比对(motif-MSA)以扩增群体遗传信号的研究。由于蛋白质编码区域通常处于较强的选择约束之下,此类变异的发生频率普遍偏低,致使下游计算往往缺乏足够的统计支持。本研究利用重复基序的模块化结构,从群体遗传学及传统的物种间保守性分析中扩增选择信号。针对每个重复蛋白质的重复结构域,我们首先构建基序-MSA(motif-MSA),随后基于该基序-MSA的基因组坐标系,汇总全球人群中的单核苷酸变异(Single Nucleotide Variant, SNV)数据。这使得我们能够将所有对应的单核苷酸变异群体遗传特征——包括罕见变异富集度、非同义突变与同义突变比值以及等位基因频率差值(delta DAFs)——与基序-MSA中的氨基酸变异进行整合。

提供机构:
耶鲁大学
搜集汇总
数据集介绍
Intensification 数据集图片
背景与挑战
背景概述
Intensification 是一个数据库,旨在通过构建基于基序的多序列比对来放大群体遗传信号,专门针对12个重复蛋白结构域。它整合了人类群体中的单核苷酸变异数据,以增强对选择信号的分析。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务