Study and Analysis of the Heterogeneity of a Prostate Cancer Dataset: First Steps on the Release of a Multicenter Strongly-Annotated Dataset
摘要
Prostate cancer (PCa) is a prevalent and deadly disease, necessitating advancements in diagnostic approaches with the advent of clinical digitization. This shift involves the application of Deep Learning and Convolution Neural Networks for automated PCa diagnosis through image processing. The Gleason scale is pivotal in assessing PCa aggressiveness, but the notable inter-observer variability among pathologists in assigning Gleason scores poses a challenge. To address this, leveraging diverse and sufficiently heterogeneous datasets becomes crucial for training artificial intelligence systems to generalize effectively. In this study, a pivotal step is taken by introducing a dataset comprising labeled PCa samples from three distinct medical centers. The outcomes unveil insights into image variability and establish a ground truth using supervised learning, serving as a benchmark for researchers. This initiative aims to empower AI systems to provide precise diagnoses across diverse PCa samples, ultimately enhancing clinical decision-making and mitigating observer-dependent discrepancies.