Understanding human intent is crucial for human-computer interaction (HCI) and human-robot collaboration (HRC) in industrial environments. However, existing intention recognition models suffer from poor generalization, high computational costs, and latency issues, particularly when processing high-dimensional skeletal motion data. This study explores dimensionality reduction techniques, including Principal Component Analysis (PCA), Kernel PCA (kPCA), and Uniform Manifold Approximation and Projection (UMAP), to optimize feature selection while maintaining relevant motion features. A multimodal system using RGB, depth, and embedded sensors is developed to classify human intentions. The proposed system integrates neural network-based models to analyze spatial and temporal motion patterns, enabling adaptive robotic behavior through short-term pose prediction, collision sensing, and path correction. The OAK-D AI platform is employed for real-time depth perception and edge AI processing. While the applied dimensionality reduction techniques lead to reduced model accuracy, they significantly enhance computational efficiency, lower processing latency, and reduce power consumption. These trade-offs make the system more suitable for real-time, resource-constrained industrial applications, ensuring faster but slightly less accurate decision-making in AI-driven industrial automation and workplace safety monitoring.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Intention Recognition Efficiency: A Study on Skeletal Data Dimensionality Reduction and Neural Architectures

  • Luka Orsag,
  • Tomislav Stipancic,
  • Leon Koren,
  • Matija Zidaric

摘要

Understanding human intent is crucial for human-computer interaction (HCI) and human-robot collaboration (HRC) in industrial environments. However, existing intention recognition models suffer from poor generalization, high computational costs, and latency issues, particularly when processing high-dimensional skeletal motion data. This study explores dimensionality reduction techniques, including Principal Component Analysis (PCA), Kernel PCA (kPCA), and Uniform Manifold Approximation and Projection (UMAP), to optimize feature selection while maintaining relevant motion features. A multimodal system using RGB, depth, and embedded sensors is developed to classify human intentions. The proposed system integrates neural network-based models to analyze spatial and temporal motion patterns, enabling adaptive robotic behavior through short-term pose prediction, collision sensing, and path correction. The OAK-D AI platform is employed for real-time depth perception and edge AI processing. While the applied dimensionality reduction techniques lead to reduced model accuracy, they significantly enhance computational efficiency, lower processing latency, and reduce power consumption. These trade-offs make the system more suitable for real-time, resource-constrained industrial applications, ensuring faster but slightly less accurate decision-making in AI-driven industrial automation and workplace safety monitoring.