<p>Due to privacy concerns and the non-Independent and Identically Distributed (non-IID) nature of patient data, traditional centralized learning approaches encounter significant challenges in medical diagnosis tasks. This paper presents a federated deep learning approach to address data heterogeneity in healthcare settings through the aggregation of cluster similarity-based client data. The central server performs data aggregation primarily using K-means clustering on the data received from contributing clients in a decentralized network. This similarity-based procedure addresses the non-IID nature of data by grouping clients with similar data distributions for local model training, thereby improving model convergence and reducing communication rounds. Various deep architectures are utilized to evaluate the model’s performance. Extensive classification experiments on a benchmarking dataset of Ulcerative Colitis endoscopic images have shown the model effectiveness, achieving an accuracy of 78%, F1-score of 78% with DenseNet-121, and a Quadratic Weighted Kappa of 77% with ResNet-50. Additionally, the clustering method significantly improved convergence speed, reducing both the training time and the number of training rounds by 55% on average. Finally, the proposed federated learning with K-means clustering consistently outperformed centralized methods in all evaluation metrics, demonstrating its efficacy in handling non-IID data distributions and providing a more robust privacy-preserving solution for healthcare systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Federated deep learning with clustering-based aggregation for medical diagnosis

  • Mohammed Al-Refai,
  • Shahed Alkhaza’leh,
  • Ahmad Alzu’bi

摘要

Due to privacy concerns and the non-Independent and Identically Distributed (non-IID) nature of patient data, traditional centralized learning approaches encounter significant challenges in medical diagnosis tasks. This paper presents a federated deep learning approach to address data heterogeneity in healthcare settings through the aggregation of cluster similarity-based client data. The central server performs data aggregation primarily using K-means clustering on the data received from contributing clients in a decentralized network. This similarity-based procedure addresses the non-IID nature of data by grouping clients with similar data distributions for local model training, thereby improving model convergence and reducing communication rounds. Various deep architectures are utilized to evaluate the model’s performance. Extensive classification experiments on a benchmarking dataset of Ulcerative Colitis endoscopic images have shown the model effectiveness, achieving an accuracy of 78%, F1-score of 78% with DenseNet-121, and a Quadratic Weighted Kappa of 77% with ResNet-50. Additionally, the clustering method significantly improved convergence speed, reducing both the training time and the number of training rounds by 55% on average. Finally, the proposed federated learning with K-means clustering consistently outperformed centralized methods in all evaluation metrics, demonstrating its efficacy in handling non-IID data distributions and providing a more robust privacy-preserving solution for healthcare systems.