错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised State Encoding in Video Sequences Using \(\beta \) -Variational Autoencoders

  • Stephan Mulder,
  • Mathys C. Du Plessis

摘要

Monitoring and providing feedback on the execution of sequential tasks is common in various domains, such as industrial processes for quality control, automated supervision for skill acquisition and even surgical procedures. This research explores the use of a Disentangled \(\beta \) -Variational Autoencoder ( \(\beta \) -VAE) to encode different states in video data depicting a sequence of actions being performed on a series of objects without explicit labels. We trained a Disentangled \(\beta \) -VAE on video data of a sequence being performed and evaluated its ability to distinguish between states using visualisations based on similarity metrics. The evaluation was performed using a set of sequences specifically designed to establish the validity and limits of \(\beta \) -VAE’s encoding of the states. These sequences included both unseen sequences which were similar to the training data, as well as out-of-distribution sequences which deviate from those seen in training. The results demonstrate that the \(\beta \) -VAE successfully learned to encode distinct states within the sequence, as evidenced by the visualisations. It is shown that \(\beta \) -VAE is capable of detecting states within a sequence. Furthermore, it is demonstrated that these learnt states inherently also have learnt dependencies and relationships regarding the sequence in which they are performed. These findings lay the foundation for the development of an overarching algorithm that monitors a sequence in progress and provides feedback when deviations from the expected sequence occur.