Application of Random Forests Methods to Diabetic Retinopathy Classification Analyses
收藏资源简介:
Background Diabetic retinopathy (DR) is one of the leading causes of blindness in the United States and world-wide. DR is a silent disease that may go unnoticed until it is too late for effective treatment. Therefore, early detection could improve the chances of therapeutic interventions that would alleviate its effects. Methodology Graded fundus photography and systemic data from 3443 ACCORD-Eye Study participants were used to estimate Random Forest (RF) and logistic regression classifiers. We studied the impact of sample size on classifier performance and the possibility of using RF generated class conditional probabilities as metrics describing DR risk. RF measures of variable importance are used to detect factors that affect classification performance. Principal Findings Both types of data were informative when discriminating participants with or without DR. RF based models produced much higher classification accuracy than those based on logistic regression. Combining both types of data did not increase accuracy but did increase statistical discrimination of healthy participants who subsequently did or did not have DR events during four years of follow-up. RF variable importance criteria revealed that microaneurysms counts in both eyes seemed to play the most important role in discrimination among the graded fundus variables, while the number of medicines and diabetes duration were the most relevant among the systemic variables. Conclusions and Significance We have introduced RF methods to DR classification analyses based on fundus photography data. In addition, we propose an approach to DR risk assessment based on metrics derived from graded fundus photography and systemic data. Our results suggest that RF methods could be a valuable tool to diagnose DR diagnosis and evaluate its progression.
## 背景 糖尿病视网膜病变(Diabetic Retinopathy, DR)是美国乃至全球范围内主要的致盲病因之一。该疾病属于沉默性病症,往往不易被早期察觉,待到确诊时已延误有效治疗时机。因此,早期筛查可提升治疗干预的成功率,从而减轻该疾病带来的危害。 ## 研究方法 本研究纳入了3443名ACCORD眼部研究(ACCORD-Eye Study)参与者的分级眼底照相数据与全身临床数据,用于构建随机森林(Random Forest)与逻辑回归(Logistic Regression)分类器。本研究探讨了样本量对分类器性能的影响,以及能否将随机森林生成的类别条件概率作为描述糖尿病视网膜病变风险的评估指标;同时利用随机森林的变量重要性指标,筛选影响分类性能的相关因素。 ## 主要研究结果 两类数据均能有效区分是否患有糖尿病视网膜病变的研究对象。基于随机森林的模型分类准确率远高于逻辑回归模型。联合两类数据虽未提升整体分类准确率,但可更精准地统计区分随访四年内发生或未发生糖尿病视网膜病变事件的健康受试者。随机森林变量重要性分析显示,在分级眼底照相相关变量中,双眼微血管瘤计数对分类的贡献度最高;而在全身变量中,用药种类与糖尿病病程是最具相关性的因素。 ## 结论与意义 本研究将随机森林方法应用于基于眼底照相数据的糖尿病视网膜病变分类分析中。此外,本研究提出了一种基于分级眼底照相与全身数据衍生指标的糖尿病视网膜病变风险评估方案。研究结果表明,随机森林方法可作为诊断糖尿病视网膜病变并评估其病情进展的有效工具。



