<p>In environments where users must interact with distributed Internet of Things (IoT) devices, conventional gesture-based control systems independently relying on fixed cameras suffer from limited mobility, restricted sensing range, and infrastructure dependence. To overcome these limitations, this study proposes a human-following drone that provides gesture recognition for wireless control of IoT devices. This study utilized a Tello drone with MediaPipe’s body-landmarks detection to track and follow a human subject while recognizing gestures. The 3D coordinates of body parts were extracted from RGB images captured by the drone. Based on the 3D coordinates, a drone-control system employing proportional-integral (PI) controllers was designed to track a human in the 3D space. For gesture recognition, seven intuitive gestures were defined, and a feature vector was constructed using relative 3D displacements between selected body landmark pairs and joint angles formed by anatomically meaningful triplets. Classifier models including artificial neural network (ANN), support vector machine (SVM), random forest (RF), and k-nearest neighbor (kNN) were trained on 21,266 images. In real-world experiments, the drone reliably maintained a following distance of 1.5–2.3 m and an orientation within ±29.1°, while the RF classifier achieved the highest gesture recognition accuracy of 98.14%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Human-following Drone Providing Gesture Recognition to Control IoT Devices Based on 3D Body-landmark Detection

  • Kyeongmo Kang,
  • Won-jong Kim

摘要

In environments where users must interact with distributed Internet of Things (IoT) devices, conventional gesture-based control systems independently relying on fixed cameras suffer from limited mobility, restricted sensing range, and infrastructure dependence. To overcome these limitations, this study proposes a human-following drone that provides gesture recognition for wireless control of IoT devices. This study utilized a Tello drone with MediaPipe’s body-landmarks detection to track and follow a human subject while recognizing gestures. The 3D coordinates of body parts were extracted from RGB images captured by the drone. Based on the 3D coordinates, a drone-control system employing proportional-integral (PI) controllers was designed to track a human in the 3D space. For gesture recognition, seven intuitive gestures were defined, and a feature vector was constructed using relative 3D displacements between selected body landmark pairs and joint angles formed by anatomically meaningful triplets. Classifier models including artificial neural network (ANN), support vector machine (SVM), random forest (RF), and k-nearest neighbor (kNN) were trained on 21,266 images. In real-world experiments, the drone reliably maintained a following distance of 1.5–2.3 m and an orientation within ±29.1°, while the RF classifier achieved the highest gesture recognition accuracy of 98.14%.