遇见数据集

How Robust Are Multi-Rater Inter-Rater Reliability Indices to Changes in Frequency Distribution?

收藏
DataCite Commons2020-09-04 更新2024-07-25 收录
官方服务:

资源简介:

Inter-rater reliability studies are used in a diverse set of fields. Often, these investigations involve three or more raters, and thus, require the use of indices such as Fleiss’s kappa, Conger’s kappa, or Krippendorff’s alpha. Through two motivating examples – one theoretical and one from practice – this paper exposes limitations of these indices when the units to be rated are not well-distributed across the rating categories. Then, using a Monte Carlo simulation and information visualizations, we argue for the use of two alternative indices, the Brennan-Prediger coefficient and Gwet’s AC2, because the agreement levels reported by these indices are more robust to variation in the distribution of units that raters encounter. The paper concludes by exploring the complex, interwoven relationship between the number of levels in a rating instrument, the agreement level present among raters, and the distribution of units that are to be scored.

提供机构:
Taylor & Francis
创建时间:
2016-02-12
二维码
社区交流群
二维码
科研交流群
商业服务