Journal of Artificial Intelligence Research & Advances Original Research
Federated Learning for Privacy-Preserving AI Model Training Across Distributed Healthcare Systems
Abstract
Building effective AI diagnostic tools in clinical environments presents a fundamental contradiction — the patient data most critical to model performance is precisely the data subject to the strictest legal and institutional restrictions. Regulations such as HIPAA and GDPR, while essential for protecting patient rights, render conventional centralized training pipelines largely impractical in real hospital settings where data cannot be transferred, pooled, or shared across institutional boundaries. This paper presents a federated learning framework specifically engineered to resolve that contradiction, enabling AI diagnostic models to be trained collaboratively across distributed hospital networks without any raw patient records leaving their origin systems. Unlike approaches that treat privacy as a compliance checkbox, the proposed design embeds differential privacy mechanisms and secure multi-party computation directly into the core training architecture, treating them as structural requirements rather than optional additions. An adaptive gradient aggregation strategy was further developed to manage the statistical heterogeneity that inevitably arises when training data originates from hospitals with meaningfully different patient demographics, equipment standards, and clinical documentation practices. The framework was evaluated across three clinically relevant tasks — chest X-ray pathology classification, diabetic retinopathy grading, and clinical note summarization — distributed across ten simulated hospital nodes of varying data volume and feature distribution. Experimental results demonstrate that the federated model achieved performance within 2.3% of a centrally trained baseline on all three tasks, while reducing measurable privacy leakage by over 94% as assessed through membership inference attack protocols.
Keywords
References (33)
- Regulation (EU) 2016/679 of the European Parliament and of the Council. General Data Protection Regulation. Off J Eur Union.
- 2016; L119: 1–88.
- Rieke N, Hancox J, Li W, et al. The future of digital health with federated learning. NPJ Digit Med. 2020; 3(1): 119.
- McMahan B, Moore E, Ramage D, et al. Communication-efficient learning of deep networks from decentralized data. In:
- Proceedings of AISTATS 2017; Fort Lauderdale, USA. PMLR; 2017. 1273–82p.
- Melis L, Song C, De Cristofaro E, et al. Exploiting unintended feature leakage in collaborative learning. In: Proceedings of IEEE
- Symposium on Security and Privacy 2019; San Francisco, USA. IEEE; 2019. 691–706p.
- Li T, Sahu AK, Zaheer M, et al. Federated optimization in heterogeneous networks. Proc Mach Learn Syst. 2020; 2: 429–50p.
- Sheller MJ, Reina GA, Edwards B, et al. Multi-institutional deep learning modeling without sharing patient data: a feasibility study
- on brain tumor segmentation. In: Crimi A, Bakas S, editors. Brainlesion: Glioma, Multiple Sclerosis, Stroke and TBI. Cham:
- Springer; 2019. 92–104p.
- Flores M, Raja A, Bhatt U, et al. Federated learning used for predicting outcomes in SARS-CoV-2 patients. Res Sq. 2021.
- doi:10.21203/rs.3.rs-126892/v1.
- Brisimi TS, Chen R, Mela T, et al. Federated learning of predictive models from federated electronic health records. Int J Med
- Inform. 2018; 112: 59–67p.
- Zhu L, Liu Z, Han S. Deep leakage from gradients. Adv Neural Inf Process Syst. 2019; 32: 14774–84p.
- Dwork C, McSherry F, Nissim K, et al. Calibrating noise to sensitivity in private data analysis. In: Proceedings of TCC 2006; New
- York, USA. Berlin: Springer; 2006. 265–84p.
- Abadi M, Chu A, Goodfellow I, et al. Deep learning with differential privacy. In: Proceedings of ACM CCS 2016; Vienna,
- Austria. New York: ACM; 2016. 308–18p.
- Kairouz P, McMahan HB, Avent B, et al. Advances and open problems in federated learning. Found Trends Mach Learn. 2021;
- 14(1–2): 1–210p.
- Mironov I. Rényi differential privacy of the Gaussian mechanism. In: Proceedings of IEEE CSF 2017; Santa Barbara, USA. IEEE;
- 2017. 263–75p.
- Bonawitz K, Ivanov V, Kreuter B, et al. Practical secure aggregation for privacy-preserving machine learning. In: Proceedings of
- ACM CCS 2017; Dallas, USA. New York: ACM; 2017. 1175–91p.
- Wang X, Peng Y, Lu L, et al. ChestX-ray8: Hospital-scale chest X-ray database and benchmarks. In: Proceedings of IEEE CVPR
- 2017; Honolulu, USA. IEEE; 2017. 2097–106p.
- Kaggle. Diabetic Retinopathy Detection [dataset online]. San Francisco: Kaggle; 2015 [cited 2024 Jan 10]. Available from:
- https://www.kaggle.com/c/diabetic-retinopathy-detection
- Johnson AE, Pollard TJ, Shen L, et al. MIMIC-III, a freely accessible critical care database. Sci Data. 2016; 3(1): 160035.
- Shokri R, Stronati M, Song C, et al. Membership inference attacks against machine learning models. In: Proceedings of IEEE
- Symposium on Security and Privacy 2017; San Jose, USA. IEEE; 2017. 3–18p.