Background <p>Early introduction of peanut products to infants around 4- to 6- months of age may reduce peanut allergy incidence. However, clinician adherence to the National Institute of Allergy and Infectious Diseases’ 2017 Addendum Guidelines for the Prevention of Peanut Allergy (PPA), which recommend early peanut introduction based on risk levels, has been reportedly low. Documentation of clinician peanut introduction recommendations and peanut allergy risk in electronic health records (EHR) is variable and often in the form of unstructured data. Therefore, this study aims to develop and validate a Natural Language Processing (NLP) approach to identify and measure pediatric clinicians’ adherence to the PPA guidelines (peanut introduction recommendations and peanut allergy risk assessments, including eczema severity), as documented in EHR systems.</p> Methods <p>An NLP pipeline was developed to process clinical notes and patient instructions from EHRs in the Intervention to Reduce Early Peanut Allergy in Children (iREACH) trial. iREACH is a two-arm, cluster-randomized, controlled clinical trial that evaluates an intervention to enhance clinician adherence to the PPA guidelines. The database includes 4- and 6-month well-child care visits from 30 practices across three clinical networks in Illinois. The development of the NLP was organized into three main phases: exploratory (reviewed EHR notes to identify concepts for developing NLP algorithms), training (resulting in the first version of the NLP algorithm), and validation (based on gold standard datasets). Chart reviews were conducted to review the accuracy and reliability of the NLP model. The NLP pipeline assessed peanut introduction recommendations and severe eczema.</p> Results <p>NLP achieved high precision (0.98), recall (0.94) and overall performance F-measure (0.96) for identifying peanut introduction recommendations across the three networks. However, identifying severe eczema proved more challenging, with precision of 0.52, recall of 0.92, and overall performance F-measure of 0.67 across the three networks. Therefore, manual review was used to confirm the severe eczema results for the trial.</p> Conclusions <p>Considering the pragmatic study design, clinical notes were the only feasible source of data, yet documentation of severe eczema varied considerably across networks. Future refinement and validation of the severe eczema NLP pipeline are required. NLP is a valuable tool in assessing large EHR data sets in pragmatic, multi-site clinical trials.</p> Trial registration <p>This trial is registered on ClinicalTrials.gov (NCT04604431). Registered on October 14, 2020.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing pediatric clinician adherence to the guidelines for prevention of peanut allergy: a natural language processing study

  • Anthony F. Wong,
  • Lucy A. Bilaver,
  • Jialing Jiang,
  • Yuan Luo,
  • Ruchi S. Gupta,
  • Marc Rosenman,
  • Michael S. Carroll

摘要

Background

Early introduction of peanut products to infants around 4- to 6- months of age may reduce peanut allergy incidence. However, clinician adherence to the National Institute of Allergy and Infectious Diseases’ 2017 Addendum Guidelines for the Prevention of Peanut Allergy (PPA), which recommend early peanut introduction based on risk levels, has been reportedly low. Documentation of clinician peanut introduction recommendations and peanut allergy risk in electronic health records (EHR) is variable and often in the form of unstructured data. Therefore, this study aims to develop and validate a Natural Language Processing (NLP) approach to identify and measure pediatric clinicians’ adherence to the PPA guidelines (peanut introduction recommendations and peanut allergy risk assessments, including eczema severity), as documented in EHR systems.

Methods

An NLP pipeline was developed to process clinical notes and patient instructions from EHRs in the Intervention to Reduce Early Peanut Allergy in Children (iREACH) trial. iREACH is a two-arm, cluster-randomized, controlled clinical trial that evaluates an intervention to enhance clinician adherence to the PPA guidelines. The database includes 4- and 6-month well-child care visits from 30 practices across three clinical networks in Illinois. The development of the NLP was organized into three main phases: exploratory (reviewed EHR notes to identify concepts for developing NLP algorithms), training (resulting in the first version of the NLP algorithm), and validation (based on gold standard datasets). Chart reviews were conducted to review the accuracy and reliability of the NLP model. The NLP pipeline assessed peanut introduction recommendations and severe eczema.

Results

NLP achieved high precision (0.98), recall (0.94) and overall performance F-measure (0.96) for identifying peanut introduction recommendations across the three networks. However, identifying severe eczema proved more challenging, with precision of 0.52, recall of 0.92, and overall performance F-measure of 0.67 across the three networks. Therefore, manual review was used to confirm the severe eczema results for the trial.

Conclusions

Considering the pragmatic study design, clinical notes were the only feasible source of data, yet documentation of severe eczema varied considerably across networks. Future refinement and validation of the severe eczema NLP pipeline are required. NLP is a valuable tool in assessing large EHR data sets in pragmatic, multi-site clinical trials.

Trial registration

This trial is registered on ClinicalTrials.gov (NCT04604431). Registered on October 14, 2020.