Improvements in the Imbalanced Hemogram Data Classification
摘要
The exponential growth of hospital information systems (HIS) has led to the accumulation of vast amounts of medical data, necessitating effective analysis methods to enhance the quality and efficiency of medical services. Machine learning has emerged as a valuable technology for the automated and accurate analysis of medical data, offering potential applications in disease diagnosis and treatment. This study aims to contribute to the advancement of classification methods and address data imbalance issues in the context of hematological data. Specifically, we propose an efficient algorithm for disease classification utilizing hemogram blood test samples, employing the random forest algorithm in conjunction with the synthetic minority oversampling technique. Experimental results using real hematological data from a local hospital demonstrate the superiority of the proposed method, achieving an impressive accuracy rate of up to 97.75% and an Area Under the Curve value of up to 98.65%. The findings underscore the value of leveraging machine learning techniques in diagnoses and treatment in clinical practice, especially when integrated into HIS systems.