Are LLMs Ready to Shoulder Biomedical Trust?
摘要
Large language models (LLMs), such as ChatGPT and Med-PaLM 2, have demonstrated transformative potential in biomedical applications such as diagnosis, medical prognosis, and patient care. This article explores the dominant LLMs developed for medical applications, and analyzes and compares their features, applications, advantages, and limitations. An analytical comparison of these models highlights their performance across various healthcare domains, while also addressing the ethical considerations surrounding their use. Furthermore, the article examines the risks of misleading medical prognoses due to issues such as overgeneralization, hallucinations, and biases inherent in LLMs. By emphasizing fine-tuning, explainability, and human oversight, this article provides a comprehensive overview of how these challenges can be mitigated, ultimately guiding the development of safer and more effective LLMs in biomedical contexts.