A Multimodal Fusion Network for Automatic Modulation Classification
摘要
Automatic Modulation Classification (AMC) plays a critical role in intelligent wireless communication systems. In this paper, a Multimodal Fusion Network is proposed to improve AMC performance by jointly leveraging I/Q sequences and spectrograms. Spectral features are learned from spectrograms through a combined architecture of CNN and Transformer, while temporal features from I/Q sequences are captured by a complex-valued Transformer. To enhance the generalization capability of the model, tailored data augmentation techniques are designed for I/Q signals. Furthermore, to address the modality discrepancy and enable effective feature integration, contrastive learning is employed to align the latent representations of both modalities before fusion. Finally, a cross-attention mechanism is used to fuse complementary information from the aligned features. Experimental results demonstrate that the proposed method consistently outperforms existing state-of-the-art model.