Intelligent Transportation Target Recognition Based on Multi-Modal Data Fusion in Remote Sensing and UAV Images
摘要
Detecting small targets in remote sensing and UAV imagery remains challenging due to limited resolution and ambiguous features. To address these issues, we propose a multi-modal detection framework based on an enhanced YOLO architecture that fuses infrared and visible light data. We introduce a Multi-Modal Local-Global Fusion Block (MLGF) to extract globally normalized features and design a dedicated detection head to exploit cross-modal complementarity. Extensive experiments show that our method achieves 82.48% average accuracy on the VEDAI dataset and 81.6% on DroneVehicle, outperforming state-of-the-art methods. It also improves mAP@0.5 by 5.9% on the VisDrone dataset. These results demonstrate the effectiveness and practical potential of our approach for autonomous driving and intelligent transportation systems.