Two-stream convolutional neural network for acoustic image recognition using time-frequency maps
摘要
Acoustic image recognition is crucial for underwater object tracking, ocean monitoring, and situational awareness. However, traditional methods relying on manual feature extraction are sensitive to environmental noise from complex underwater conditions, which degrades their accuracy and robustness. To address these issues, we propose a two-stream convolutional neural network for acoustic image recognition. The proposed method is characterized by processing raw time-frequency and differential feature maps in parallel to extract both global and local difference features, thereby enhancing recognition through weighted fusion. The method is validated on a simulated dataset generated by time-frequency analysis of underwater wake signals. An improved focal loss is introduced to handle class imbalance. In this context, the AdaBelief optimizer is employed to accelerate convergence and enhance stability. Experiments show our method achieves a Top-1 accuracy of 75.6%, a 7.7% improvement over backbone-only methods, while maintaining high computational efficiency. Ablation studies confirm the effectiveness of our method, boosting accuracy by 2.99%. The improved focal Loss and AdaBelief optimizer further increase accuracy by 2.49% and 4.29%, respectively, with low computational overhead. These results validate the effectiveness and practical applicability of our method.