Understanding Visual SLAM
摘要
From the inception of computer vision, there has always been a dream that one day, intelligent hardware devices could, like humans, observe the world through their eyes, perceive objects around them, comprehend spatial environments, and explore the unknown. This enchanting and romantic vision has captivated countless researchers, driving them to tirelessly pursue it night and day. Although achieving this has proven challenging, in today’s era of increasing automation, many scenarios necessitate devices having the capability for spatial perception and positioning: indoor robotic vacuums need to model their working environment, autonomous vehicles on the roads must understand and locate themselves within those roads, drones in the sky require real-time tracking of their position and orientation, and virtual and augmented reality devices need to capture their viewpoint and comprehend real-world scenes. Without this capability, all these innovative entities could not manifest in our reality, which would be a significant loss. To enable computers to parse spatial environments and ascertain their position within them, SLAM technology has seen significant advances over the past 30 years. SLAM stands for Simultaneous Localization And Mapping, which involves an entity equipped with specific sensors, creating a model of its environment while in motion, without prior information about the environment, and simultaneously estimating its own movement. If the primary sensor is a camera, the technology is referred to as Visual SLAM. This chapter will primarily focus on Visual SLAM, the most widely applied form of SLAM in the industrial sector, to introduce this spatial intelligence computing technology.