错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-supervised Meta Auxiliary Learning for Actor and Action Video Segmentation from Natural Language

  • Linwei Ye,
  • Zhenhua Wang

摘要

This paper addresses the problem of actor and action video segmentation from natural language. Given a video and a language query, the goal is to segment the actor and its action described by the query. Existing methods focus on exploring elaborated multimodal feature fusion networks to combine visual and linguistic features for an effective multimodal representation directly learnt from this labeled segmentation task. In this paper, we propose a novel self-supervised meta auxiliary learning method to improve the primary segmentation task by adding an auxiliary task for better generalization. The auxiliary task is established to reconstruct the input sentence representation so that the multimodal representation can be adapted to a specific query. In addition, the auxiliary task does not require additional labels. It can also be used in test time to update a multimodal representation according to a specific query in a self-supervised way.