<p>This paper introduces a domain-adaptive visual enhanced keyword spotting framework that overcomes the severe performance degradation of audio-only systems in noisy or distant conditions. The approach employs a visual-acoustic memory architecture, in which visual key memories from observed lip movements query a shared audio memory to generate candidate acoustic representations. These enriched visual embeddings are fused through a reflective attention mechanism that extracts bidirectional alignment signals from a unified attention map. Long-range temporal dynamics are modeled by a continuity-preserving attention network that enforces consistency via future feature prediction and input reconstruction objectives. Evaluations on the Multimodal Information Based Speech Processing challenge corpus show relative improvements of 24.2% on the development set and 14.8% on the test set compared with state-of-the-art baselines, demonstrating marked gains in both recognition accuracy and robustness under adverse conditions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Domain adaptative keyword spotting with multimodal enhancement

  • Longxi Chen,
  • Han Wang

摘要

This paper introduces a domain-adaptive visual enhanced keyword spotting framework that overcomes the severe performance degradation of audio-only systems in noisy or distant conditions. The approach employs a visual-acoustic memory architecture, in which visual key memories from observed lip movements query a shared audio memory to generate candidate acoustic representations. These enriched visual embeddings are fused through a reflective attention mechanism that extracts bidirectional alignment signals from a unified attention map. Long-range temporal dynamics are modeled by a continuity-preserving attention network that enforces consistency via future feature prediction and input reconstruction objectives. Evaluations on the Multimodal Information Based Speech Processing challenge corpus show relative improvements of 24.2% on the development set and 14.8% on the test set compared with state-of-the-art baselines, demonstrating marked gains in both recognition accuracy and robustness under adverse conditions.