<p>Social media platforms generate large volumes of real-time, geo-referenced content during natural hazards, enabling multimodal topic models to uncover latent themes and their spatial organisation. However, the evaluation of such models typically relies on unimodal metrics, primarily semantic coherence, which overlook spatial structure, temporal dynamics, and the practical usefulness of discovered topics. This paper proposes a multidimensional evaluation framework for disaster-related social media analysis that integrates semantic, spatial, temporal, and operational indicators of topic quality. The framework distinguishes between Utility Information Value (UIV), capturing the informational usefulness of topic representations, and Diagnostic Information Value (DIV), capturing the structural and hazard-consistent validity of spatial and spatio-temporal topic patterns. We examine the framework through a comparative benchmarking analysis of multimodal topic models (MultiGraph and JSTTS) alongside a strong unimodal baseline, using eight geo-referenced datasets from X&#xa0;(Twitter) and Bluesky covering earthquakes, floods, hurricanes, and wildfires under a unified experimental setup. Results show that multimodal models achieve higher actionability (up to 0.75), while the unimodal baseline shows greater semantic diversity (up to 0.99). Spatio-temporal interaction varies across hazards (0.07-0.45), and spatial patterns exhibit hazard-dependent structure. UIV and DIV are strongly correlated (Pearson <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(r = 0.89\)</EquationSource> </InlineEquation>). However, mean performance remains similar across models (UIV <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\approx\)</EquationSource> </InlineEquation> 0.49-0.50, DIV <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\approx\)</EquationSource> </InlineEquation> 0.52-0.53), with no statistically significant differences (<InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(p&gt; 0.05\)</EquationSource> </InlineEquation>), indicating that performance differences are moderate and dataset-dependent rather than consistently model-driven.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Beyond unimodal metrics: a multimodal evaluation framework for geo-referenced content on social media

  • Ehsaneddin Jalilian,
  • David Hanny,
  • Bernd Resch

摘要

Social media platforms generate large volumes of real-time, geo-referenced content during natural hazards, enabling multimodal topic models to uncover latent themes and their spatial organisation. However, the evaluation of such models typically relies on unimodal metrics, primarily semantic coherence, which overlook spatial structure, temporal dynamics, and the practical usefulness of discovered topics. This paper proposes a multidimensional evaluation framework for disaster-related social media analysis that integrates semantic, spatial, temporal, and operational indicators of topic quality. The framework distinguishes between Utility Information Value (UIV), capturing the informational usefulness of topic representations, and Diagnostic Information Value (DIV), capturing the structural and hazard-consistent validity of spatial and spatio-temporal topic patterns. We examine the framework through a comparative benchmarking analysis of multimodal topic models (MultiGraph and JSTTS) alongside a strong unimodal baseline, using eight geo-referenced datasets from X (Twitter) and Bluesky covering earthquakes, floods, hurricanes, and wildfires under a unified experimental setup. Results show that multimodal models achieve higher actionability (up to 0.75), while the unimodal baseline shows greater semantic diversity (up to 0.99). Spatio-temporal interaction varies across hazards (0.07-0.45), and spatial patterns exhibit hazard-dependent structure. UIV and DIV are strongly correlated (Pearson \(r = 0.89\) ). However, mean performance remains similar across models (UIV \(\approx\) 0.49-0.50, DIV \(\approx\) 0.52-0.53), with no statistically significant differences ( \(p> 0.05\) ), indicating that performance differences are moderate and dataset-dependent rather than consistently model-driven.