In the modern era of data science, it is increasingly common to work with high-dimensional datasets characterised by heterogeneous features, containing outliers, and distributed across multiple computational environments. Emerging research suggests that alternative distance measures, such as cosine similarity in text mining or embeddings, can outperform traditional metric distances in certain similarity search applications. However, most of the existing similarity search algorithms are tailored for metric distances, limiting their ability to exploit these alternative measures fully. To bridge this gap, we introduce GDASC (General Distributed Approximate Similarity Search with Clustering), a novel framework designed for distributed approximate similarity search. GDASC builds a multilevel index by employing a user-defined distance measure and following a rather unusual clustering approach. Thanks to its bottom-up building approach, which is particularly optimised for environments where information is distributed across multiple computational nodes, this index can efficiently perform approximate k-nearest neighbour searches. The flexibility offered by this method broadens the range of applicable distance measures beyond traditional metrics. Besides, it enables deeper exploration into the potential advantages of constructing solutions by combining regions derived from different dissimilarity spaces. This PhD thesis aims to thoroughly investigate the potential benefits of exploring diverse dissimilarity spaces. Thus, by delving into its combination, we strive to uncover performance improvements and generate new insights, potentially making a significant contribution to advancing the state of the art in similarity search methodologies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Combining Dissimilarity Spaces to Improve Approximate Similarity Search

  • Elena García-Morato,
  • Felipe Ortega,
  • Javier Gómez

摘要

In the modern era of data science, it is increasingly common to work with high-dimensional datasets characterised by heterogeneous features, containing outliers, and distributed across multiple computational environments. Emerging research suggests that alternative distance measures, such as cosine similarity in text mining or embeddings, can outperform traditional metric distances in certain similarity search applications. However, most of the existing similarity search algorithms are tailored for metric distances, limiting their ability to exploit these alternative measures fully. To bridge this gap, we introduce GDASC (General Distributed Approximate Similarity Search with Clustering), a novel framework designed for distributed approximate similarity search. GDASC builds a multilevel index by employing a user-defined distance measure and following a rather unusual clustering approach. Thanks to its bottom-up building approach, which is particularly optimised for environments where information is distributed across multiple computational nodes, this index can efficiently perform approximate k-nearest neighbour searches. The flexibility offered by this method broadens the range of applicable distance measures beyond traditional metrics. Besides, it enables deeper exploration into the potential advantages of constructing solutions by combining regions derived from different dissimilarity spaces. This PhD thesis aims to thoroughly investigate the potential benefits of exploring diverse dissimilarity spaces. Thus, by delving into its combination, we strive to uncover performance improvements and generate new insights, potentially making a significant contribution to advancing the state of the art in similarity search methodologies.