LLM enabled classification of patient self-reported symptoms and needs in health systems across the USA
摘要
US health systems receive up to 200 M monthly website visitors. Connecting patient searches to the appropriate workflow requires accurate classification. A dataset of searches on ~15 US health system websites was annotated, characterized, and used to train and evaluate a multi-label, multi-class, deep neural network. This classifier was deployed to health systems touching patients in all 50 states and compared to an LLM. The training dataset contained 504 unique classes with performance of the model in classifying searches among those classes ranging from ~0.90 to ~0.70 across metrics depending on the number of classes included. GPT-4 performed similarly if given a master list and demonstrated value in providing added coverage to augment the supervised classifier’s performance. The collected data revealed characteristics of patient searches in the largest, multi-center, national study of US health systems to date.