错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Display and View—NSDN—Image Caption Generator and Headcount

  • Shreya Nahar,
  • Tanishka R. Jain,
  • Mihir Shah,
  • Golda Dilip

摘要

The automatic description of the content of an image is a challenge that combines computer vision with natural language processing in the field of artificial intelligence. Most existing models have limitations in accurately identifying subtle elements of an image and counting the number of people present. To address this challenge, we introduce the natural-language scene description network (NSDN), which incorporates recent developments in computer vision and machine translation to generate descriptive phrases that effectively depict images. This technology has the potential to improve the efficiency, capacity, reliability, and safety of crowd management tasks, particularly in diverse and adaptive crowd situations. Despite challenges such as clutter, occlusion, non-uniform object scale, and irregular object distribution, YOLO shows promise for intelligent crowd counting and analysis in images. This article reviews, categorizes, analyzes distinguishing features, and extensively assesses the effectiveness of crowd-counting methods that rely on convolutional neural networks and provides a detailed analysis of their performance.