A Review of Mathematical Information Retrieval: Bridging Symbolic Representation and Intelligent Retrieval
摘要
Mathematical information retrieval is a subfield of information retrieval that deals with retrieving mathematics-related content such as formulas, equations, symbols, and their texts. Its purpose is to help users access relevant mathematical expressions or associated documents, revisited through queries like keywords or mathematical expressions. Such a field has been changing over the years more progressively, overcoming the challenges of using text-based approaches to more modern AI techniques like vector embeddings, tree-based, and large language models. This review provides a comprehensive study of the approaches and techniques used and the issues faced in mathematical information retrieval. It discusses retrieval methods, text-based, tree-based, embedding-based, and large language model-based based focusing on their efficiency in formula search and cross-lingual retrieval. Moreover, the review covers benchmark datasets including arXiv, Wikipedia Corpus, and ARQMath, which are crucial for the advancement of mathematical information retrieval. The review also examines the major mathematical information retrieval systems, including MIaS, MathWebSearch, Tangent, and MCAT, assessing their retrieval efficiency and search capabilities. The key challenges, such as formula representation, lack of standard notation, and the need for multimodal retrieval, are addressed. The study identifies research gaps and suggests future directions, particularly enhancing cross-lingual mathematical information retrieval.