Textual similarity calculation techniques in the medical field: a retrospective review
摘要
The integration and application of digital technology in the medical field are accelerating the development of medical technology and generating a large amount of information resources. Such information resources are characterized by large-scale, heterogeneous, and scattered distribution, including medical guidelines, academic literature, electronic medical records, medical databases, and diagnostic reports. Text similarity computing is an intelligent technology that mines intrinsically related information, supports the integration of medical digital resources, explores the potential value of medical data, and directly or indirectly promotes the development of medical research, clinical diagnosis, and decision support technology. With the development of artificial intelligence technology and large models, medical text similarity computation has achieved remarkable results with this support. Similarity calculation, from the original string rule template matching to today’s highly accurate neural network models, has achieved significant performance improvements. Throughout the medical field, text similarity computing has achieved significant results in advancing the development of medicine and improving the doctor-patient relationships. Text similarity computing in the medical field also provides the basis and necessary foundation for the intelligent development of medical treatment in the new era. In this paper, we systematize the development of text similarity computation in the medical field and comprehensively review more than 200 related studies. We summarize and compare the traditional statistical methods, machine learning, and neural network models for medical text similarity computation. In particular, typical cases are analyzed for each text similarity computation method in the medical field to highlight their advantages and potential research value. In this paper, after an extensive surveys and summary, we discuss the research methods of text similarity computation in the medical field, commonly used datasets, and the analysis of the research results. Finally, it summarizes and looks forward to the future development of text similarity computing in the medical field.