FoodNet-CNN: A Deep Learning Model for Multi-Class Food Image Classification
摘要
The rapid proliferation of food imagery on social media and mobile devices, coupled with increasing demand for intelligent dietary monitoring, has driven significant advances in automated food recognition. However, most existing deep learning models focus on Western and East Asian cuisines, leaving Vietnamese dishes—characterized by high intra-class variability and subtle inter-class differences—largely unexplored. This study introduces FoodNet-CNN, a domain-specific deep learning framework designed for multi-class classification of ten iconic Vietnamese dishes. A novel, curated dataset of 5,030 images was constructed by integrating the public 30VNfoods dataset with manually collected samples from diverse real-world contexts to enhance variability and robustness. We systematically evaluate a custom-built CNN alongside state-of-the-art transfer learning architectures (VGG16, VGG19, and MobileNetV2), applying a two-stage fine-tuning strategy to optimize domain adaptation. Experimental results demonstrate that the fine-tuned VGG16 model achieves 98.2% test accuracy, significantly outperforming the baseline CNN (85.7%), and establishing a strong benchmark for Vietnamese food classification. Additionally, the model is deployed in a web-based real-time recognition system, illustrating its potential in applications such as digital nutrition tracking, smart restaurants, and health monitoring. By addressing a culturally significant yet underrepresented domain, this work expands the landscape of food recognition research and provides a scalable foundation for future developments in computational gastronomy.