<p>In real-time, high-performance computing of computer vision, the classification of dangerous scenes in crowded environments, such as the rising popularity of tourist destinations, poses significant challenges. Traditional safety monitoring methods heavily rely on manual supervision, which is inadequate for real-time risk identification to ensure the safety of life and property. This is particularly challenging in crowded environments, where transitioning from manual monitoring to human–machine collaboration is imperative. Large language models (LLMs), inferred and trained on supercomputing platforms, offer new opportunities for enhancing safety systems. Their large parameterization improves the accuracy, interpretability, and generalization of vision models. Advances like DeepSeek, leveraging high-performance computing, have also lowered deployment costs. This study introduces Gemini-Scene, a novel multimodal framework leveraging LLMs and Mamba-CNN attention mechanisms to classify dangerous scenes accurately. We construct a comprehensive dataset focused on hazardous scenarios and develop the multimodal network Gemini-Scene. By integrating image embeddings, textual descriptions, and scene probabilities from pre-trained models such as CLIP and BERT, our framework aligns multimodal features through a Mamba-CNN attention mechanism. Experimental results demonstrate that Gemini-Scene achieves a classification accuracy of 97.5%, significantly outperforming existing methods with accuracies below 92.8%. This framework offers an advanced solution for next-generation safety monitoring in high-risk, densely populated areas. The code and dataset are available at <a href="https://github.com/2254886209/Gemini-Scene">https://github.com/2254886209/Gemini-Scene</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing dangerous scene classification with multimodal LLMs and attention mechanisms

  • Dong Zhang,
  • Tian Xie,
  • Shuai Wu,
  • Shukai Duan,
  • Lidan Wang

摘要

In real-time, high-performance computing of computer vision, the classification of dangerous scenes in crowded environments, such as the rising popularity of tourist destinations, poses significant challenges. Traditional safety monitoring methods heavily rely on manual supervision, which is inadequate for real-time risk identification to ensure the safety of life and property. This is particularly challenging in crowded environments, where transitioning from manual monitoring to human–machine collaboration is imperative. Large language models (LLMs), inferred and trained on supercomputing platforms, offer new opportunities for enhancing safety systems. Their large parameterization improves the accuracy, interpretability, and generalization of vision models. Advances like DeepSeek, leveraging high-performance computing, have also lowered deployment costs. This study introduces Gemini-Scene, a novel multimodal framework leveraging LLMs and Mamba-CNN attention mechanisms to classify dangerous scenes accurately. We construct a comprehensive dataset focused on hazardous scenarios and develop the multimodal network Gemini-Scene. By integrating image embeddings, textual descriptions, and scene probabilities from pre-trained models such as CLIP and BERT, our framework aligns multimodal features through a Mamba-CNN attention mechanism. Experimental results demonstrate that Gemini-Scene achieves a classification accuracy of 97.5%, significantly outperforming existing methods with accuracies below 92.8%. This framework offers an advanced solution for next-generation safety monitoring in high-risk, densely populated areas. The code and dataset are available at https://github.com/2254886209/Gemini-Scene.