Context-Aware Dynamic Actor-Critic for Multimodal Visual Reinforcement Learning
摘要
Integrating RGB frames with other visual modalities is gaining increasing traction in vision-based reinforcement learning (RL). Existing multimodal vision-based RL methods typically fuse all input modality features into a global modality feature, which is used as a unified environmental description for subsequent decision-making. However, such a straightforward “Fuse-and-Act” paradigm often overlooks the influence of modality heterogeneity on value estimation and policy learning, leading to modality value clash and unimodal policy superiority. To solve this problem, this paper introduces a Context-aware Dynamic Actor-Critic (ConDAC) framework for multimodal vision-based RL. Unlike “Fuse-and-Act”, ConDAC employs a “Discern-and-Decide” paradigm that dynamically coordinates the interaction of different modalities based on the environmental context before making decisions. It includes a Dual Consistent Value Estimation (DCVE) method to assess the contributions of individual modalities as well as their collective efficacy, ensuring more accurate value estimation. Additionally, a Targeted Progressive Policy Learning (TPPL) strategy systematically refines the final policy by progressively integrating the best insights from various modalities. Extensive experiments demonstrate the superiority of our approach in the challenging domains of autonomous driving and robotic control tasks, suggesting a promising direction for future enhancements in multimodal vision-based RL.