We present a comprehensive analysis of search engines from past MathIR competitions (ARQMath, NTCIR) and investigate the semantic representations required for effective formula retrieval. After evaluating these search engines’ methodologies and performance, we uncover a diverse range of equation-parsing and feature-extraction strategies. Inspired by these approaches, we present an ensemble-based retrieval system, integrating symbolic, tree-based, and textual representations of formulas. A key innovation of our work is the Leaf-to-Leaf (L2L) feature extraction method, which captures structural dependencies in formula trees by identifying the shortest paths of operators between pairs of leaf nodes. Our system combines L2L features with symbol-based and LaTeX-based rankings, surpassing the top performing ARQMath-3 search engines in nDCG \(^{\prime }\) and mAP \(^{\prime }\) metrics, while matching or exceeding the top NTCIR-12 performers in PR \(^{\prime }\) @{15, 20} and P \(^{\prime }\) @{5, 10, 15, 20} measures. These results demonstrate the effectiveness of leveraging multiple feature spaces to enhance mathematical formula retrieval.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Math Formula Search Using Diverse Structural and Symbolic Representations

  • Sumedh Vemuganti,
  • Ayu Seiya,
  • Nickvash Kani

摘要

We present a comprehensive analysis of search engines from past MathIR competitions (ARQMath, NTCIR) and investigate the semantic representations required for effective formula retrieval. After evaluating these search engines’ methodologies and performance, we uncover a diverse range of equation-parsing and feature-extraction strategies. Inspired by these approaches, we present an ensemble-based retrieval system, integrating symbolic, tree-based, and textual representations of formulas. A key innovation of our work is the Leaf-to-Leaf (L2L) feature extraction method, which captures structural dependencies in formula trees by identifying the shortest paths of operators between pairs of leaf nodes. Our system combines L2L features with symbol-based and LaTeX-based rankings, surpassing the top performing ARQMath-3 search engines in nDCG \(^{\prime }\) and mAP \(^{\prime }\) metrics, while matching or exceeding the top NTCIR-12 performers in PR \(^{\prime }\) @{15, 20} and P \(^{\prime }\) @{5, 10, 15, 20} measures. These results demonstrate the effectiveness of leveraging multiple feature spaces to enhance mathematical formula retrieval.