<p>Cross-modal voice-face association learning is extensively researched, revealing inherent links between voice and facial data. Feature alignment is a crucial component in this task. However, traditional methods often rely on high-dimensional floating-point data for feature alignment, resulting in substantial memory usage and limited computational efficiency. To address these challenges, we propose a method called Hash-Deep Residual Alignment Network (H-DRAN). This model integrates a residual learning mechanism to effectively combine early and deep features, enhancing representational capability. Additionally, it employs a hash function in the final layer to convert input features into fixed-length hash codes, reducing dimensionality and memory requirements while improving computation speed. H-DRAN optimizes voice and facial hash code representations by minimizing a joint loss comprising inter-modal similarity loss and semi-triplet loss. The latter utilizes an online mining technique to filter out typical semi-hard triplets. Compared to state-of-the-art (SOTA) methods, H-DRAN improves performance in 1:2 matching by 1-9%, verification by 1-11%, and retrieval by 1-4%. Memory usage is reduced to 16 bytes per sample using hash codes, compared to traditional floating-point storage. Visualization results demonstrate that H-DRAN clusters the voice and face features of the same person while dispersing those of different individuals.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing cross-modal voice-face association with heterogeneous hashing network

  • Yanxia Liang,
  • Huanhuan Zhang,
  • Xin Liu,
  • Xiaopeng Yang,
  • Fuping Wang,
  • Jing Jiang

摘要

Cross-modal voice-face association learning is extensively researched, revealing inherent links between voice and facial data. Feature alignment is a crucial component in this task. However, traditional methods often rely on high-dimensional floating-point data for feature alignment, resulting in substantial memory usage and limited computational efficiency. To address these challenges, we propose a method called Hash-Deep Residual Alignment Network (H-DRAN). This model integrates a residual learning mechanism to effectively combine early and deep features, enhancing representational capability. Additionally, it employs a hash function in the final layer to convert input features into fixed-length hash codes, reducing dimensionality and memory requirements while improving computation speed. H-DRAN optimizes voice and facial hash code representations by minimizing a joint loss comprising inter-modal similarity loss and semi-triplet loss. The latter utilizes an online mining technique to filter out typical semi-hard triplets. Compared to state-of-the-art (SOTA) methods, H-DRAN improves performance in 1:2 matching by 1-9%, verification by 1-11%, and retrieval by 1-4%. Memory usage is reduced to 16 bytes per sample using hash codes, compared to traditional floating-point storage. Visualization results demonstrate that H-DRAN clusters the voice and face features of the same person while dispersing those of different individuals.