Breast cancer remains a leading cause of mortality among women, emphasizing the need for faster and more accurate diagnostic tools. This study proposes a novel approach for the classification of breast tumors using Fine Needle Aspiration (FNA) cytology data from the Wisconsin Breast Cancer dataset. The predictive model utilizes a recurrent neural network (RNN) with long short-term memory (LSTM) layers, enhanced by feature selection via the SelectKBest algorithm and optimized through hyperparameter tuning using Optuna. Synthetic Minority Over-sampling Technique (SMOTE) was applied to address class imbalance, and early stopping was incorporated to mitigate overfitting during training. The proposed model achieved an impressive classification accuracy of 98.25% and an AUC-ROC score of 0.998, outperforming existing methods. This study not only demonstrates the model’s effectiveness in distinguishing between benign and malignant tumors but also highlights its potential for clinical applications in early detection and diagnosis. Despite its success, the model’s limitations include a need for validation on larger, diverse datasets and an analysis of its robustness across different patient populations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Breast Cancer Classification Using RNN with LSTM Layers and SelectKBest Feature Selection on Fine Needle Aspiration Data

  • Nidhi Umashankar,
  • K. Sai Geethanjali,
  • I. S. Rajesh,
  • M. A. Bharathi,
  • V. L. Sowmya

摘要

Breast cancer remains a leading cause of mortality among women, emphasizing the need for faster and more accurate diagnostic tools. This study proposes a novel approach for the classification of breast tumors using Fine Needle Aspiration (FNA) cytology data from the Wisconsin Breast Cancer dataset. The predictive model utilizes a recurrent neural network (RNN) with long short-term memory (LSTM) layers, enhanced by feature selection via the SelectKBest algorithm and optimized through hyperparameter tuning using Optuna. Synthetic Minority Over-sampling Technique (SMOTE) was applied to address class imbalance, and early stopping was incorporated to mitigate overfitting during training. The proposed model achieved an impressive classification accuracy of 98.25% and an AUC-ROC score of 0.998, outperforming existing methods. This study not only demonstrates the model’s effectiveness in distinguishing between benign and malignant tumors but also highlights its potential for clinical applications in early detection and diagnosis. Despite its success, the model’s limitations include a need for validation on larger, diverse datasets and an analysis of its robustness across different patient populations.