遇见数据集

Model performance.

收藏
Figshare2024-05-30 更新2026-04-28 收录
官方服务:

资源简介:

Detecting voice disorders from voice recordings could allow for frequent, remote, and low-cost screening before costly clinical visits and a more invasive laryngoscopy examination. Our goals were to detect unilateral vocal fold paralysis (UVFP) from voice recordings using machine learning, to identify which acoustic variables were important for prediction to increase trust, and to determine model performance relative to clinician performance. Patients with confirmed UVFP through endoscopic examination (N = 77) and controls with normal voices matched for age and sex (N = 77) were included. Voice samples were elicited by reading the Rainbow Passage and sustaining phonation of the vowel "a". Four machine learning models of differing complexity were used. SHapley Additive exPlanations (SHAP) was used to identify important features. The highest median bootstrapped ROC AUC score was 0.87 and beat clinician’s performance (range: 0.74–0.81) based on the recordings. Recording durations were different between UVFP recordings and controls due to how that data was originally processed when storing, which we can show can classify both groups. And counterintuitively, many UVFP recordings had higher intensity than controls, when UVFP patients tend to have weaker voices, revealing a dataset-specific bias which we mitigate in an additional analysis. We demonstrate that recording biases in audio duration and intensity created dataset-specific differences between patients and controls, which models used to improve classification. Furthermore, clinician’s ratings provide further evidence that patients were over-projecting their voices and being recorded at a higher amplitude signal than controls. Interestingly, after matching audio duration and removing variables associated with intensity in order to mitigate the biases, the models were able to achieve a similar high performance. We provide a set of recommendations to avoid bias when building and evaluating machine learning models for screening in laryngology.

通过语音录音识别语音障碍,可实现在昂贵的临床就诊与侵入性喉镜检查前,开展高频、远程且低成本的筛查。本研究旨在通过机器学习从语音录音中识别单侧声带麻痹(unilateral vocal fold paralysis, UVFP),明确对预测结果具有重要贡献的声学变量以提升模型可信度,并对比模型与临床医师的诊断性能。本研究纳入经内镜检查确诊的单侧声带麻痹患者77例,以及年龄、性别匹配的嗓音正常对照者77例。语音样本通过朗读彩虹短文(Rainbow Passage)与持续发元音/a/采集得到。本研究采用4种复杂度各异的机器学习模型,并使用Shapley可加解释(SHapley Additive exPlanations, SHAP)方法识别关键特征。4种模型中,经自助法(bootstrap)计算的最高中位受试者工作特征曲线下面积(Receiver Operating Characteristic Area Under the Curve, ROC AUC)得分为0.87,优于临床医师的诊断性能(0.74~0.81)。由于存储时的原始数据处理方式不同,单侧声带麻痹组与对照组的录音时长存在显著差异,而本研究证实,该时长差异本身即可实现两组的分类。且与直觉相悖的是,尽管单侧声带麻痹患者的嗓音通常较弱,但多数患者的录音强度高于对照组,这揭示了本数据集特有的偏倚,我们在后续补充分析中对该偏倚进行了校正。本研究证实,录音时长与强度的偏倚导致了患者组与对照组间的数据集特有差异,而模型正是利用了该差异以提升分类性能。此外,临床医师的评分进一步证实,患者存在过度发声的情况,其录音的信号振幅高于对照组。有趣的是,为校正偏倚,我们对录音时长进行匹配并移除与强度相关的变量后,模型仍可达到相近的高性能水平。本研究最后提出了一系列建议,旨在为喉科学领域的机器学习筛查模型构建与评估过程中规避偏倚提供参考。

创建时间:
2024-05-30
二维码
社区交流群
二维码
科研交流群
商业服务