Transformer-Based Models to Assist Covid-19 Literature Screening
摘要
Purpose: During the Covid-19 pandemic, the surge in medical literature necessitated rapid extraction of actionable insights for researchers and decision-makers. Systematic reviews, though reliable, require significant time from expert curators. This work proposes leveraging transformer models pretrained on medical texts for automated abstract screening. Methods: Six transformer models with varied architectures and properties were fine-tuned and optimized using curated datasets comprising up to 36 labels and 50,000 examples. The data was integrated, augmented, and cleaned through web scraping, manual labeling, and semi-automatic labeling experiments. Key factors influencing performance were data quality, quantity, model size, architecture, and the pretraining corpus. Results: The top-performing models demonstrated consistently high recall and good F1-scores across all scenarios. This performance indicates a robust filtering solution for abstract screening, capable of significantly reducing reviewers’ effort while ensuring relevant information retrieval. Conclusion: The proposed automated abstract screening method using fine-tuned transformer models offers a viable solution to manage the overwhelming amount of medical literature. By maintaining high recall and good F1-scores, these models can effectively aid in systematic reviews, thereby enhancing efficiency for researchers and decision-makers.