<p>This paper presents a DDPG-based solution to the inverse kinematics of a six-degree-of-freedom Stewart platform. The problem is posed directly in actuator-length space, which enables strict enforcement of hardware limits and avoids the ill conditioning that arises from coupling between position and orientation in Cartesian coordinates. The policy is further extended with residual correction and a predictive safety shield: a proportional baseline command is refined by a data-driven residual, and a shield reverts to the baseline whenever the mixed command is predicted to increase tracking error or violate constraints. A realistic simulator is used that injects sensor and actuator noise, models a first-order actuator lag, and applies small geometric perturbations. The controller is benchmarked against IK-Direct, Joint-space proportional control, and a TD3 policy configured with the same inference pipeline over a diverse set of translational trajectories including infinity, spiral, Lissajous in two and three dimensions, multisin, and stairs. Performance is reported using root-mean-square tracking error in millimeters, the longest streak of steps below a fixed error threshold, average control step time in milliseconds, and an explicit constraint-penalty term. The results show millimeter-level accuracy on smooth paths and competitive performance on challenging trajectories, with consistent gains over proportional baselines and very large improvements over direct reinforcement-learning variants without residuals or shielding, while maintaining sub-millisecond control latencies. A reward ablation together with a clear hyperparameter rationale is provided to substantiate learning stability and sensitivity. These findings indicate that reinforcement learning, when coupled with residual correction and safety shielding, is a practical pathway to precise, robust, and real-time control of parallel robotic mechanisms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Solving inverse kinematics for a 6-DOF Stewart platform using the DDPG algorithm

  • Soheil Sheikh Ahmadi,
  • Arash Rahmani

摘要

This paper presents a DDPG-based solution to the inverse kinematics of a six-degree-of-freedom Stewart platform. The problem is posed directly in actuator-length space, which enables strict enforcement of hardware limits and avoids the ill conditioning that arises from coupling between position and orientation in Cartesian coordinates. The policy is further extended with residual correction and a predictive safety shield: a proportional baseline command is refined by a data-driven residual, and a shield reverts to the baseline whenever the mixed command is predicted to increase tracking error or violate constraints. A realistic simulator is used that injects sensor and actuator noise, models a first-order actuator lag, and applies small geometric perturbations. The controller is benchmarked against IK-Direct, Joint-space proportional control, and a TD3 policy configured with the same inference pipeline over a diverse set of translational trajectories including infinity, spiral, Lissajous in two and three dimensions, multisin, and stairs. Performance is reported using root-mean-square tracking error in millimeters, the longest streak of steps below a fixed error threshold, average control step time in milliseconds, and an explicit constraint-penalty term. The results show millimeter-level accuracy on smooth paths and competitive performance on challenging trajectories, with consistent gains over proportional baselines and very large improvements over direct reinforcement-learning variants without residuals or shielding, while maintaining sub-millisecond control latencies. A reward ablation together with a clear hyperparameter rationale is provided to substantiate learning stability and sensitivity. These findings indicate that reinforcement learning, when coupled with residual correction and safety shielding, is a practical pathway to precise, robust, and real-time control of parallel robotic mechanisms.