遇见数据集

Risky Business: Factor Analysis of Survey Data – Assessing the Probability of Incorrect Dimensionalisation

收藏
Figshare2016-10-31 更新2026-04-29 收录
官方服务:

资源简介:

This paper undertakes a systematic assessment of the extent to which factor analysis the correct number of latent dimensions (factors) when applied to ordered-categorical survey items (so-called Likert items). We simulate 2400 data sets of uni-dimensional Likert items that vary systematically over a range of conditions such as the underlying population distribution, the number of items, the level of random error, and characteristics of items and item-sets. Each of these datasets is factor analysed in a variety of ways that are frequently used in the extant literature, or that are recommended in current methodological texts. These include exploratory factor retention heuristics such as Kaiser’s criterion, Parallel Analysis and a non-graphical scree test, and (for exploratory and confirmatory analyses) evaluations of model fit. These analyses are conducted on the basis of Pearson and polychoric correlations. We find that, irrespective of the particular mode of analysis, factor analysis applied to ordered-categorical survey data very often leads to over-dimensionalisation. The magnitude of this risk depends on the specific way in which factor analysis is conducted, the number of items, the properties of the set of items, and the underlying population distribution. The paper concludes with a discussion of the consequences of over-dimensionalisation, and a brief mention of alternative modes of analysis that are much less prone to such problems.

本研究系统性评估了将因子分析应用于有序分类调查题项(即所谓李克特(Likert)题项)时,其准确识别潜在维度(因子)正确数量的能力边界。我们针对单维李克特题项生成了2400组仿真数据集,这些数据集的生成条件涵盖系统变化的多个维度,包括总体潜在分布、题项数量、随机误差水平,以及题项与题项集的特征。针对每一组数据集,我们采用现有文献中常用或当前方法论教材推荐的多种方式进行因子分析,具体包括探索性因子保留启发式方法,如凯泽准则(Kaiser’s criterion)、平行分析(Parallel Analysis)以及非图形碎石检验;同时针对探索性与验证性分析,还包括模型拟合度评估。所有分析均基于皮尔逊(Pearson)相关与多列相关(polychoric correlations)展开。研究结果表明:无论采用何种分析模式,将因子分析应用于有序分类调查数据时,往往会导致维度识别过度。该风险的严重程度取决于因子分析的具体实施方式、题项数量、题项集的属性以及总体潜在分布。本研究最后讨论了维度识别过度带来的后果,并简要提及了更不易出现此类问题的替代分析方法。

创建时间:
2016-10-31
二维码
社区交流群
二维码
科研交流群
商业服务