Bidirectional Two-Sample Mendelian Randomization Protocol: Causal Association Between Lipoprotein(a) and Chronic Obstructive Pulmonary Disease with Minor Allele Frequency Subgroup Stratification
收藏资源简介:
Description Study Background Elevated lipoprotein(a) [Lp(a)] is a genetically determined lipid biomarker with well-established causal effects on cardiovascular disease. Observational epidemiological studies have reported inconsistent correlations between circulating Lp(a) levels and chronic obstructive pulmonary disease (COPD), and it remains unclear whether the association reflects true causal effects, reverse causality, or confounding factors. Conventional observational research cannot eliminate residual confounding, which limits causal inference. Therefore, this study adopts two-sample Mendelian randomization (MR) to eliminate confounding and reverse causality bias and clarify the causal relationship between genetically predicted Lp(a) and COPD. Study Design This is a pre-specified bidirectional two-sample summary-data Mendelian randomization study. Two directions of causal inference will be performed: Forward MR: Genetically elevated Lp(a) as exposure, COPD disease status as the outcome; Reverse MR: Genetic liability to COPD as exposure, plasma Lp(a) concentration as the outcome. All GWAS summary datasets will be uniformly converted to the hg19 genome build, and only European ancestry samples will be included to reduce population stratification bias. A predefined subgroup analysis will stratify instrumental single-nucleotide polymorphisms (SNPs) into low-frequency variants (minor allele frequency, MAF < 0.05) and common variants (MAF ≥ 0.05), to explore heterogeneous causal effects between rare LPA variants and common variants on COPD risk. Instrumental Variable Selection Criteria SNPs reaching genome-wide significance threshold (P < 5 × 10⁻⁸) are extracted from corresponding publicly available GWAS summary statistics; Linkage disequilibrium clumping will be conducted via the online API function embedded in the TwoSampleMR package, using the European 1000 Genomes reference panel; clumping parameters: r² = 0.001, clumping window = 10000 kb; Palindromic SNPs, ambiguous strand variants, and SNPs with missing allele frequency information will be excluded; Remaining valid instrumental SNPs will be divided into two subgroups based on MAF for separate bidirectional MR analysis. Primary and Secondary Statistical Models Primary causal estimator: Random-effects inverse variance weighted (IVW) model; Secondary complementary models: MR-Egger regression and weighted median estimator, to verify the robustness of IVW results. Predefined Sensitivity Analyses Cochran’s Q statistic to evaluate heterogeneity across individual instrumental SNPs; MR-Egger intercept test to detect unbalanced directional horizontal pleiotropy; Leave-one-out analysis to identify whether the overall causal estimate is driven by a single outlier SNP; Visual diagnostic plots including scatter plot, funnel plot and forest plot will be generated for result validation. Software and Analytical Environment All analyses will be performed using R software (latest stable version). Core packages include TwoSampleMR, data.table, ggplot2. Linkage disequilibrium clumping will be completed through the built-in online API function of TwoSampleMR package, using European 1000 Genomes reference panel. All analytical procedures are fully pre-specified before formal data analysis to avoid p-hacking and post-hoc modification of analytical schemes.



