<p>The clustering of histogram data, which preserves important distributional information, remains a significant challenge due to the limitations of existing approaches. This paper presents a novel two-level clustering algorithm explicitly designed for histogram data. It integrates a self-organizing map (SOM) framework with the <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="521_2025_11530_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(L_2\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>L</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation> Wasserstein distance to capture distributional features such as location, scale, and shape. By enriching prototypes with local density and connectivity measures, the approach automatically infers cluster numbers without prior knowledge. Experiments on synthetic and real-world datasets underscore the quality of the obtained results in comparison with other distance-based or histogram-based approaches, the computational efficiency of the proposed approach, and its interpretability through topology-preserving visualizations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Histogram self-organizing map for cluster detection and visualization

  • Guénaël Cabanes,
  • Younès Bennani,
  • Parisa Rastin

摘要

The clustering of histogram data, which preserves important distributional information, remains a significant challenge due to the limitations of existing approaches. This paper presents a novel two-level clustering algorithm explicitly designed for histogram data. It integrates a self-organizing map (SOM) framework with the \(L_2\) L 2 Wasserstein distance to capture distributional features such as location, scale, and shape. By enriching prototypes with local density and connectivity measures, the approach automatically infers cluster numbers without prior knowledge. Experiments on synthetic and real-world datasets underscore the quality of the obtained results in comparison with other distance-based or histogram-based approaches, the computational efficiency of the proposed approach, and its interpretability through topology-preserving visualizations.