The aim of this study was to compare the performance of seven deep natural language processing (NLP) models in classifying brain tumor follow-up image reports, so as to quickly and standardised extract prognostic features from unstructured reports. We collected follow-up reports from patients diagnosed with brain tumors at two hospitals and manually classified them as “tumor/tumor free” and “tumor status (progressive or stable/improving)” as baseline data. Reports were randomly divided into training sets, verification sets, and test sets in a 7:2:1 ratio. Seven deep NLP models including one-dimensional convolutional neural network (CNN), recurrent neural network (RNN), gated cyclic unit (GRU), Long short-term memory network (LSTM), ClinicalBERT, BlueBERT and ELECTRA were used in the study. The verification set was used for frequent evaluation and parameter fine-tuning, and finally the model performance was evaluated using the test set data, and the consistency between Cohen’s kappa test and manual classification results was passed. In addition, the relationship between the extracted image features and overall survival was evaluated by multivariate Cox proportional hazard regression analysis. The results showed that in 10006 reports of 1580 patients, kappa values between manual annotators were 0.80 and 0.77, respectively. Except RNN, kappa values between other models and manual annotation results were between 0.78 and 0.80, showing a good consistency. The classification task AUC of the seven models exceeded 0.90, and the weighted F1 score, AUC, sensitivity and specificity of the ELECTRA model in the classification task of “with or without tumor” were 0.910, 0.96, 0.85 and 0.94, respectively. These measures in the “tumor status” classification task were 0.925, 0.96, 0.76, and 0.98, respectively. Survival analysis showed no significant difference in overall survival between the machine and manual groups. Patients classified as having tumors had 2.74 times the risk of death compared with those without tumors (2.84 times in the artificial group). Patients classified as having tumor progression had 2.25 times the risk of death compared to those in the stable/improved group (2.12 times in the artificial group). In summary, the ELECTRA model performs best among the seven deep NLP models, which can effectively classify tumor features in unstructured image reports and provide reliable risk stratification information for patients.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of Multi-model Fusion Deep NLP System in Classification of Brain Tumor Follow-Up Image Reports

  • Jinzhu Yang

摘要

The aim of this study was to compare the performance of seven deep natural language processing (NLP) models in classifying brain tumor follow-up image reports, so as to quickly and standardised extract prognostic features from unstructured reports. We collected follow-up reports from patients diagnosed with brain tumors at two hospitals and manually classified them as “tumor/tumor free” and “tumor status (progressive or stable/improving)” as baseline data. Reports were randomly divided into training sets, verification sets, and test sets in a 7:2:1 ratio. Seven deep NLP models including one-dimensional convolutional neural network (CNN), recurrent neural network (RNN), gated cyclic unit (GRU), Long short-term memory network (LSTM), ClinicalBERT, BlueBERT and ELECTRA were used in the study. The verification set was used for frequent evaluation and parameter fine-tuning, and finally the model performance was evaluated using the test set data, and the consistency between Cohen’s kappa test and manual classification results was passed. In addition, the relationship between the extracted image features and overall survival was evaluated by multivariate Cox proportional hazard regression analysis. The results showed that in 10006 reports of 1580 patients, kappa values between manual annotators were 0.80 and 0.77, respectively. Except RNN, kappa values between other models and manual annotation results were between 0.78 and 0.80, showing a good consistency. The classification task AUC of the seven models exceeded 0.90, and the weighted F1 score, AUC, sensitivity and specificity of the ELECTRA model in the classification task of “with or without tumor” were 0.910, 0.96, 0.85 and 0.94, respectively. These measures in the “tumor status” classification task were 0.925, 0.96, 0.76, and 0.98, respectively. Survival analysis showed no significant difference in overall survival between the machine and manual groups. Patients classified as having tumors had 2.74 times the risk of death compared with those without tumors (2.84 times in the artificial group). Patients classified as having tumor progression had 2.25 times the risk of death compared to those in the stable/improved group (2.12 times in the artificial group). In summary, the ELECTRA model performs best among the seven deep NLP models, which can effectively classify tumor features in unstructured image reports and provide reliable risk stratification information for patients.