EfficientViT: An Efficient Vision Transformer for Fire and Smoke Image Classification
摘要
Given the variety of fire and smoke, which are distinguished by variances in texture and color, it is extremely difficult to detect fire and smoke from visual imagery. A significant amount of economic and environmental harm has resulted from fire and smoke incidents around the world. In order to prevent property damage and save lives, fires must be detected as soon as possible, for example, forest fires, which frequently occur in the summer yet are easily avoidable. Using cutting-edge deep learning models including Convolutional Neural Network (CNN), VGG16, Vision Transformer (ViT), and MobileViT, we offer a thorough method for classifying fire and smoke in this study. We also provide an innovative model dubbed EfficientViT, which combines the advantages of the MobileViT and Inverted Residual layers. The correct classification of visual images into the categories of fire and smoke is our main goal. We show that EfficientViT beats the other developed models, obtaining an amazing accuracy of 0.92, through a thorough experimentation and evaluation. This demonstrates how well our suggested model works at correctly differentiating between incidents of fire and smoke.