Learning Algorithms for Breast Cancer Classification and Diagnosis
摘要
The comprehension about the key features of breast tumors that lead to their classification as benign or malignant is fundamental to improving the detection and diagnosis of breast cancer, contributing significantly to survival rates and treatment effectiveness. This study proposes a multidisciplinary approach that combines analytical methods and graphical visualizations to classify breast tumors as benign or malignant, using supervised and unsupervised learning algorithms. The adopted dataset in this study is from the repository of the University of Wisconsin (USA). It comprises 569 breast tumor biopsy samples, with 32 features measured from digitized images of biopsy slides. Initially, for unsupervised learning, Pearson correlation was used as a similarity metric for hierarchical grouping, resulting in the formation of six clusters through a dendrogram. In supervised learning, the Principal Component Analysis (PCA) technique was performed to reduce the number of features, m achieving the 10 most relevant. The Support Vector Machine (SVM) model was applied with and without the PCA results. The comparison between hierarchical grouping and SVM methods demonstrated a notable advantage of SVM in terms of accuracy in classifying breast tumors. The use of cross-validation showed the superiority of SVM over clustering for this specific purpose. The analysis of breast tumor features and the classification approaches offer important perspectives on improving breast cancer diagnosis and treatment practices. The use of classifiers such as SVM, together with dimensionality reduction techniques such as PCA, can result in significant improvements in diagnostic accuracy and effectiveness, directly benefiting patient care in this critical area of medicine.