Purpose <p>Inflammatory bowel disease (IBD) cohort identification typically relies primarily on read/billing codes, which may miss some patients. However, a complete picture cannot typically be obtained due to database fragmentation/missingness. This study used novel cohort retrieval methods to identify the total IBD cohort from a large university teaching hospital with a specialist intestinal failure unit.</p> Methods <p>Between 2007 and 2023, 11 clinical databases (ICD10 codes, OPCS4 codes, clinician-entry IBD registry, IBD patient portal, prescriptions, biochemistry, flare line calls, clinic appointments, endoscopy, histopathology, and clinic letters) were identified as having the potential to help identify local patients with&#xa0;IBD. The 11 databases were statistically compared, and a penalized logistic regression (LR) classifier was robustly trained and validated.</p> Results <p>The gold-standard validation cohort comprised 2800 patients: 2092(75%) with IBD and 708(25%) without. All the databases contained unique patients that were not covered by the Casemix ICD-10 database. The penalizsed LR model (AUROC:0.85-Validation) confidently identified 8,159 patients with IBD (threshold: 0.496). By combining the likely true-positive predictions from the LR model with likely true-positive IBD clinic letters, a final estimate of <i>13,048</i> patients with IBD was obtained. ICD-10 codes combined with medication data identified only 8,048 patients, suggesting that present recapture methods missed <i>38.3%</i> of the local cohort.</p> Conclusion <p>Diagnostic billing codes and medication data alone cannot accurately identify complete cohorts of individuals with IBD in secondary care. A multimodal cross-database model can partially compensate for this deficit. However, to improve this situation in the future, more robust natural language processing (NLP)-based identification mechanisms will be required<i>.</i></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identification of Cohorts with Inflammatory Bowel Disease Amidst Fragmented Clinical Databases via Machine Learning

  • Matthew Stammers,
  • Stephanie Sartain,
  • J. R. Fraser Cummings,
  • Christopher Kipps,
  • Reza Nouraei,
  • Markus Gwiggner,
  • Cheryl Metcalf,
  • James Batchelor

摘要

Purpose

Inflammatory bowel disease (IBD) cohort identification typically relies primarily on read/billing codes, which may miss some patients. However, a complete picture cannot typically be obtained due to database fragmentation/missingness. This study used novel cohort retrieval methods to identify the total IBD cohort from a large university teaching hospital with a specialist intestinal failure unit.

Methods

Between 2007 and 2023, 11 clinical databases (ICD10 codes, OPCS4 codes, clinician-entry IBD registry, IBD patient portal, prescriptions, biochemistry, flare line calls, clinic appointments, endoscopy, histopathology, and clinic letters) were identified as having the potential to help identify local patients with IBD. The 11 databases were statistically compared, and a penalized logistic regression (LR) classifier was robustly trained and validated.

Results

The gold-standard validation cohort comprised 2800 patients: 2092(75%) with IBD and 708(25%) without. All the databases contained unique patients that were not covered by the Casemix ICD-10 database. The penalizsed LR model (AUROC:0.85-Validation) confidently identified 8,159 patients with IBD (threshold: 0.496). By combining the likely true-positive predictions from the LR model with likely true-positive IBD clinic letters, a final estimate of 13,048 patients with IBD was obtained. ICD-10 codes combined with medication data identified only 8,048 patients, suggesting that present recapture methods missed 38.3% of the local cohort.

Conclusion

Diagnostic billing codes and medication data alone cannot accurately identify complete cohorts of individuals with IBD in secondary care. A multimodal cross-database model can partially compensate for this deficit. However, to improve this situation in the future, more robust natural language processing (NLP)-based identification mechanisms will be required.