<p>This paper presents a dual-path deep learning architecture for geometric shape recognition, integrating temporal and spatial information from Inertial Measurement Unit (IMU) and grip pressure sensors. The model comprises a Bidirectional Long Short-Term Memory (BiLSTM) network for learning temporal dynamics and a 1D Convolutional Neural Network (1D CNN) for capturing spatial distributions. An attention mechanism dynamically fuses both feature streams, selectively emphasizing salient segments to enhance classification performance and interpretability. Framed as a supervised classification problem, the task is well-defined with ground truth labels, enabling effective training and evaluation. Experimental results demonstrate that the proposed model achieves 94.7% accuracy and a 93.3% F1-score, outperforming both traditional methods (e.g., DTW, HMM) and single-path deep learning models. The model also maintains high robustness under Gaussian noise, confirming the attention mechanism’s role in mitigating irrelevant input variations. This approach addresses existing limitations in multimodal shape recognition by offering an interpretable and scalable solution. The architecture shows promise for real-world applications, particularly in early-stage developmental screening tasks such as the Korean Developmental Screening Test (K-DST), where dynamic integration of temporal-spatial features is essential for diagnostic precision. The study provides a foundational framework for future work involving more complex shapes, multi-sensor data fusion, and child-focused cognitive assessments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prediction and Recognition of Geometric Shapes Using Dual-Path Architecture Based on Attention Mechanism

  • Chae-Young Lim,
  • Kyung-Ho Kim

摘要

This paper presents a dual-path deep learning architecture for geometric shape recognition, integrating temporal and spatial information from Inertial Measurement Unit (IMU) and grip pressure sensors. The model comprises a Bidirectional Long Short-Term Memory (BiLSTM) network for learning temporal dynamics and a 1D Convolutional Neural Network (1D CNN) for capturing spatial distributions. An attention mechanism dynamically fuses both feature streams, selectively emphasizing salient segments to enhance classification performance and interpretability. Framed as a supervised classification problem, the task is well-defined with ground truth labels, enabling effective training and evaluation. Experimental results demonstrate that the proposed model achieves 94.7% accuracy and a 93.3% F1-score, outperforming both traditional methods (e.g., DTW, HMM) and single-path deep learning models. The model also maintains high robustness under Gaussian noise, confirming the attention mechanism’s role in mitigating irrelevant input variations. This approach addresses existing limitations in multimodal shape recognition by offering an interpretable and scalable solution. The architecture shows promise for real-world applications, particularly in early-stage developmental screening tasks such as the Korean Developmental Screening Test (K-DST), where dynamic integration of temporal-spatial features is essential for diagnostic precision. The study provides a foundational framework for future work involving more complex shapes, multi-sensor data fusion, and child-focused cognitive assessments.