遇见数据集

How Robust Are Multirater Interrater Reliability Indices to Changes in Frequency Distribution?

收藏
Figshare2016-02-13 更新2026-04-29 收录
官方服务:

资源简介:

Interrater reliability studies are used in a diverse set of fields. Often, these investigations involve three or more raters, and thus, require the use of indices such as Fleiss’s kappa, Conger’s kappa, or Krippendorff’s alpha. Through two motivating examples—one theoretical and one from practice—this article exposes limitations of these indices when the units to be rated are not well-distributed across the rating categories. Then, using a Monte Carlo simulation and information visualizations, we argue for the use of two alternative indices, the Brennan–Prediger coefficient and Gwet’s AC2, because the agreement levels reported by these indices are more robust to variation in the distribution of units that raters encounter. The article concludes by exploring the complex, interwoven relationship between the number of levels in a rating instrument, the agreement level present among raters, and the distribution of units that are to be scored. Supplementary materials for this article are available online.

评分者信度(Interrater reliability)研究广泛应用于众多领域。此类研究通常涉及3名及以上评分者,因此需使用弗莱斯kappa值(Fleiss’s kappa)、康格kappa值(Conger’s kappa)或克里彭多夫α系数(Krippendorff’s alpha)等指标。本文通过两个引例——一个为理论示例、一个来自实际场景,揭示了当待评分单元在评分类别中分布不均时,上述指标存在的局限性。随后,本文借助蒙特卡洛模拟(Monte Carlo simulation)与信息可视化(information visualizations)手段,论证了采用布伦南-普里迪格系数(Brennan–Prediger coefficient)与格威特AC2(Gwet’s AC2)这两种替代指标的合理性:这两类指标所报告的一致性水平,对于评分者所接触的评分单元分布变化具备更强的稳健性。本文最后探讨了评分工具的等级数量、评分者间一致性水平以及待评分单元分布三者之间复杂且相互交织的关联。本文的补充材料可在线获取。

创建时间:
2016-02-13
二维码
社区交流群
二维码
科研交流群
商业服务