Smoking Status Classification: A Comparative Analysis of Machine Learning Techniques with Clinical Real World Data
摘要
Electronic health records often lack consistent and organized documentation regarding lifestyle-related risk factors. This study addresses this by presenting methodologies aimed at standardizing the recording of patients’ smoking status. Different types of machine learning methods are applied to an anonymized set of German-language clinical narratives in order to categorize smoking status as a multi-class classification task utilizing SNOMED CT as a terminology standard. Our findings demonstrate the effectiveness of downstreaming medBERT.de, an openly available medical language model in German, achieving the best performance with an F1-measure of [0.969–0.976] 95% CI, in comparison to CNN, LSTM and an SVM baseline.