Advancing Multi-modal Visual Pattern Recognition: Object Detection
摘要
Recently, the integration of multimodal data systems has emerged as a critical area of research and development, significantly improving the robustness and accuracy of visual pattern recognition systems. In this paper, we propose a novel approach by creating a hybrid model architecture that leverages the detection capabilities of YOLOv8. Our data set consists of 5000 multimodal image pairs, encompassing 13 distinct classes: RGB, depth, and infrared, processed separately. The outputs of each are fused using late fusion techniques, and Non-Maximum Suppression (NMS) is applied to increase accuracy. We utilized stratified test-train split algorithms to validate the datasets. Our evaluations produced results demonstrating significant improvements in detection performance, with mean Average Precision (mAP) as the primary metric. An accuracy of 0.2048 was obtained by testing our model. Our findings suggest that improved multimodal detection can enhance the impact of real-world applications, creating new avenues for scrutiny in visual pattern recognition.