associations
收藏资源简介:
GWAS Catalog Associations 数据集包含来自 NHGRI-EBI GWAS Catalog 的经过整理的遗传关联结果。该数据集记录了同行评审研究中报告的 SNP-性状关联,每一行代表一个遗传变异(通常是单核苷酸多态性,SNP)与疾病或性状之间的关联。数据集以表格形式呈现,适用于下游分析、机器学习和基因组学研究工作流。数据集的主要领域是全基因组关联研究(GWAS),观察单位为 SNP-性状关联。典型用途包括基因组风险分析、变异注释流程、表型-基因型关系研究、遗传关联的机器学习以及 GWAS 发现的荟萃分析。数据集结构包括多个列,如日期、PubMed ID、作者、期刊、研究标题、疾病/性状、样本描述、染色体位置、基因信息、统计显著性等。数据集的整理过程包括文献识别、手动整理、标准化、注释和质量控制。该数据集仅包括 GWAS 显著关联,完整的汇总统计信息可从 GWAS Catalog 获取。数据集的使用需遵循 EMBL-EBI 服务的一般使用条款。
The GWAS Catalog Associations dataset contains curated genetic association results sourced from the NHGRI-EBI GWAS Catalog. This dataset documents SNP-trait associations reported in peer-reviewed research, with each row representing an association between a genetic variant (typically a single nucleotide polymorphism, SNP) and a disease or trait. Presented in tabular format, the dataset is suitable for downstream analysis, machine learning, and genomics research workflows. Its primary domain is genome-wide association studies (GWAS), with the observational unit being SNP-trait associations. Typical use cases include genomic risk analysis, variant annotation pipelines, phenotype-genotype relationship research, machine learning for genetic associations, and meta-analyses of GWAS findings. The dataset includes multiple columns such as date, PubMed ID, authors, journal, study title, disease/trait, sample description, chromosomal location, gene information, statistical significance, and others. The dataset curation process involves literature identification, manual curation, standardization, annotation, and quality control. This dataset only includes GWAS-significant associations; complete summary statistics can be obtained from the GWAS Catalog. Usage of the dataset is subject to the general terms of service of EMBL-EBI services.




