Measuring the similarity between objects is an important task in many different statistical and machine learning problems. In this work, we propose a new similarity measure. The proposal is based on the idea of rank transformation of distances and deploys the concept of Information Imbalance as a measure of variable importance. The specific characteristics captured by the proposed similarity are demonstrated in a toy example, where it is compared with the Euclidean distance and the cosine similarity on a two dimensional dataset. The proposed similarity can be embedded in clustering algorithms, k-Nearest Neighbour type models, or it can be tested in other applications such as image analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Non-parametric Similarity Measure Based on Information Imbalance

  • Elena Ballante,
  • Silvia Figini,
  • Antonietta Mira

摘要

Measuring the similarity between objects is an important task in many different statistical and machine learning problems. In this work, we propose a new similarity measure. The proposal is based on the idea of rank transformation of distances and deploys the concept of Information Imbalance as a measure of variable importance. The specific characteristics captured by the proposed similarity are demonstrated in a toy example, where it is compared with the Euclidean distance and the cosine similarity on a two dimensional dataset. The proposed similarity can be embedded in clustering algorithms, k-Nearest Neighbour type models, or it can be tested in other applications such as image analysis.