<p>Automatic modulation recognition (AMR) is a fundamental process between signal detection and demodulation. Despite recent advances in deep learning-based AMR, existing methods often fail to maintain robustness in severe noise scenarios. To address this, we propose MAFFNet, a noise-robust multi-modal architecture that synergistically processes raw in-phase/quadrature (IQ) signals and derived amplitude/phase (AP) information through a dual-branch vision transformer-LSTM framework. Specifically, the modified vision transformer (ViT) branch employs localized attention mechanisms to reduce computational complexity, while the LSTM branch incorporates phase-difference attention to model temporal dependencies in AP information. Additionally, a learnable feature fusion module with element-wise weights dynamically combines multi-domain features, complemented by an orthogonal constraint loss that reduces inter-branch redundancy. Extensive experiments on the RML2018.01 benchmark show that the proposed MAFFNet achieves 82.53<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7492_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> accuracy between 0–10 SNRs, outperforming other methods by 5.46–11.25%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MAFFNet: a multi-modal adaptive feature fusion net for signal modulation recognition

  • Yuncong Jiang,
  • Weifei Jia,
  • Quanlin Yu

摘要

Automatic modulation recognition (AMR) is a fundamental process between signal detection and demodulation. Despite recent advances in deep learning-based AMR, existing methods often fail to maintain robustness in severe noise scenarios. To address this, we propose MAFFNet, a noise-robust multi-modal architecture that synergistically processes raw in-phase/quadrature (IQ) signals and derived amplitude/phase (AP) information through a dual-branch vision transformer-LSTM framework. Specifically, the modified vision transformer (ViT) branch employs localized attention mechanisms to reduce computational complexity, while the LSTM branch incorporates phase-difference attention to model temporal dependencies in AP information. Additionally, a learnable feature fusion module with element-wise weights dynamically combines multi-domain features, complemented by an orthogonal constraint loss that reduces inter-branch redundancy. Extensive experiments on the RML2018.01 benchmark show that the proposed MAFFNet achieves 82.53 \(\%\) % accuracy between 0–10 SNRs, outperforming other methods by 5.46–11.25%.