Pancreatic cancer is notoriously difficult to detect, with diagnosis often relying on symptoms that only develop at advanced stages of the disease. Routine blood tests may signal a developing cancer before these symptoms appear. Limited research has investigated the use of time-varying information from laboratory tests before diagnosis. This study used UK primary care data to compare machine learning approaches for detecting pancreatic cancer at various time-intervals before diagnosis. The machine learning challenge is that such real-world data is irregular and sparse and therefore difficult to use for model creation. In this study, deep learning time-series models (LSTM and GRU-D) were compared to a feature engineering approach. We found that while predictive performance was strongest at diagnosis date (maximum AUROC of 0.85), cases could be detected 18 months before diagnosis, with GRU-D achieving an AUROC of 0.57. Closer to the diagnosis date, where diagnostic signals are stronger, feature engineering approaches outperformed the deep learning models. However, further from diagnosis, the deep learning models, particularly the GRU-D, maintained marginally better performance. Calibration of the models was good at the diagnosis date but was poor across all models at a lead time of greater than 6 months. This study demonstrates that routine blood tests show some predictive capacity for earlier detection of pancreatic cancer. However, this capacity quickly decreases further from diagnosis date, with poor discrimination beyond 6 months. These results should be of interest to researchers interested in using machine learning and electronic health records to support earlier diagnosis of cancer.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning Approaches to the Early Detection of Pancreatic Cancer from Time-Series Primary Care Data

  • Victoria Moglia,
  • Lesley Smith,
  • Gordon Cook,
  • Marc De Kamps,
  • Owen Johnson

摘要

Pancreatic cancer is notoriously difficult to detect, with diagnosis often relying on symptoms that only develop at advanced stages of the disease. Routine blood tests may signal a developing cancer before these symptoms appear. Limited research has investigated the use of time-varying information from laboratory tests before diagnosis. This study used UK primary care data to compare machine learning approaches for detecting pancreatic cancer at various time-intervals before diagnosis. The machine learning challenge is that such real-world data is irregular and sparse and therefore difficult to use for model creation. In this study, deep learning time-series models (LSTM and GRU-D) were compared to a feature engineering approach. We found that while predictive performance was strongest at diagnosis date (maximum AUROC of 0.85), cases could be detected 18 months before diagnosis, with GRU-D achieving an AUROC of 0.57. Closer to the diagnosis date, where diagnostic signals are stronger, feature engineering approaches outperformed the deep learning models. However, further from diagnosis, the deep learning models, particularly the GRU-D, maintained marginally better performance. Calibration of the models was good at the diagnosis date but was poor across all models at a lead time of greater than 6 months. This study demonstrates that routine blood tests show some predictive capacity for earlier detection of pancreatic cancer. However, this capacity quickly decreases further from diagnosis date, with poor discrimination beyond 6 months. These results should be of interest to researchers interested in using machine learning and electronic health records to support earlier diagnosis of cancer.