FederatedChestCT
收藏DataCite Commons2024-03-14 更新2025-04-16 收录
下载链接:
https://ieee-dataport.org/documents/federatedchestct
下载链接
链接失效反馈官方服务:
资源简介:
While it is well known that population differences from genetics, sex, race, and environmental factors contribute to disease, AI studies in medicine have largely focused on locoregional patient cohorts with less diverse data sources. Such limitation stems from barriers to large-scale data share and ethical concerns over data privacy. Federated learning (FL) is one potential pathway for AI development that enables learning across hospitals without data share. In this study, we show the results of various FL strategies on one of the largest and most diverse chest CT datasets: 21 participating hospitals across five continents that comprise >10,000 patients (>1 million CT 2D frames). We propose an FL strategy that leverages synthetically-generated data to overcome class and size imbalances. We also describe the sources of data heterogeneity in the context of FL, and show how even among the correctly labeled populations, disparities can arise due to these biases. Our second high-level goal is to show our FL sensitivity to privacy and ways to defend against adversaries through existing algorithms. We illustrate that FL can provide some baseline level of privacy over centralized data sharing (CDS). However, even with FL and Differentially-Private training (DP) on FL, site and membership information are still vulnerable to inference attacks. Finally, we show that the use of good synthetic CT images can simultaneously provide an uplift in both FL performance and privacy. Please check https://github.com/edhlee/LargeScaleFLCT as this will serve as the official code and database.
提供机构:
IEEE DataPort
创建时间:
2024-03-14



