Machine learning-based identification of natural history studies in rare diseases: a step toward understanding disease development and outcome
摘要
According to the Food and Drug Administration’s definition, natural history studies (NHS) are observational studies that collect data on the course of a disease from onset to resolution or death, in the absence of an intervention. NHS plays a critical role in rare disease research to support clinical development objectives, including estimating disease prevalence, identifying diagnostic biomarkers. In this study, we present a machine learning based approach for systematically identifying NHS from PubMed as a foundational step toward large-scale NHS analysis for drug development in rare diseases. We explored a manually curated NHS corpus to develop machine learning and deep learning predictive models for NHS classification. We evaluated both binary and multiclass classifications and found that models trained specifically for binary classification (distinguishing NHS-relevant from NHS-irrelevant studies) achieved substantially higher performance than those trained to distinguish among four NHS-related categories, namely, Unrelated, Irrelevant, Secondary, and Primary. Among all models tested, the pretrained PubMedBERT-base-uncased-abstract model achieved the best performance (precision = 0.8171, recall = 0.8079, F1 score = 0.8125, AUCPR = 0.8768). These results highlight the feasibility of automated NHS identification using deep learning and underscore the effectiveness of binary classification as an initial approach. This work demonstrates potential to accelerate NHS data collection and lays the foundation for NHS analysis to enhance our understanding of disease progression in rare diseases, and may extend to common diseases as well.