Dataset for: A Bayesian Confirmatory Factor Model for Multivariate Observations in the Form of Two-Way Tables of Data
收藏资源简介:
Researchers collected multiple measurements on schizophrenia (SZ) patients and their relatives, as well as control subjects and their relatives, to study vulnerability factors for schizophrenics and their near relatives. Observations across individuals from the same family are correlated, and also the multiple outcome measures on the same individuals are correlated. Traditional data analyses model outcomes separately and thus do not provide information about the interrelationships among outcomes. We propose a novel Bayesian Family Factor Model (BFFM), which extends the classical confirmatory factor analysis (CFA) model to explain the correlations among observed variables using a combination of family-member factors and outcome factors.Traditional methods for fitting CFA models, such as full information maximum likelihood (FIML) estimation using quasi-Newton optimization (QNO) can have convergence problems and Heywood cases (lack-of-convergence) caused by empirical under-identification. In contrast, modern Bayesian Markov chain Monte Carlo handles these inference problems easily. Simulations compare the BFFM to FIML-QNO in settings where the true covariance matrix is identified, close to not identified and not identified. For these settings, FIML-QNO fails to fit the data in $13\%$, $57\%$ and $85\%$ of the cases, respectively, while MCMC provides stable estimates. When both methods successfully fit the data, estimates from the BFFM have smaller variances and comparable mean squared errors. We illustrate the BFFM by analyzing data on data from schizophrenics and their family members.
研究人员针对精神分裂症(schizophrenia, SZ)患者及其亲属、对照个体及其亲属开展了多维度测量,以探究精神分裂症患者及其近亲属的易感风险因素。同一家族个体间的观测数据存在相关性,同一受试者的多项结局指标也存在相关性。传统数据分析方法会单独对结局指标进行建模,因此无法揭示不同结局指标间的内在关联。本研究提出一种新型贝叶斯家族因子模型(Bayesian Family Factor Model, BFFM),该模型拓展了经典验证性因子分析(confirmatory factor analysis, CFA)框架,通过结合家族成员因子与结局因子来解释观测变量间的相关性。传统的验证性因子分析模型拟合方法,如采用拟牛顿优化(quasi-Newton optimization, QNO)的全信息极大似然(full information maximum likelihood, FIML)估计,可能会出现收敛问题以及海伍德案例(Heywood cases,即经验识别不足导致的不收敛情况)。与之相比,现代贝叶斯马尔可夫链蒙特卡洛(Markov chain Monte Carlo, MCMC)方法可轻松解决这类推断难题。我们通过仿真实验,在真实协方差矩阵可识别、接近不可识别以及完全不可识别的三种场景下,对比了BFFM与FIML-QNO的性能。在这三类场景中,FIML-QNO分别在13%、57%和85%的案例中无法完成数据拟合,而MCMC则可生成稳定的参数估计。当两种方法均成功拟合数据时,BFFM得到的估计值具有更小的方差,且均方误差性能相当。最后,我们通过分析精神分裂症患者及其家庭成员的相关数据,展示了BFFM的实际应用效果。



