This chapter explores how robots perceive their environment using various sensors and process visual data through CNNs and transformers for tasks like classification, segmentation, and object detection. It discusses the trade-offs between different CNN-based models and highlights the advantages of vision transformers (ViT) and detection transformers (DETR) in capturing global context. The chapter also covers scalability and emerging transformer-based methods beneficial for robotics applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robot Perception: Sensors and Image Processing

  • Alishba Imran,
  • Keerthana Gopalakrishnan

摘要

This chapter explores how robots perceive their environment using various sensors and process visual data through CNNs and transformers for tasks like classification, segmentation, and object detection. It discusses the trade-offs between different CNN-based models and highlights the advantages of vision transformers (ViT) and detection transformers (DETR) in capturing global context. The chapter also covers scalability and emerging transformer-based methods beneficial for robotics applications.