<p>Vision Transformers (ViTs) significantly improved the recognition accuracies of the image classification tasks by learning global cues on large datasets. Local cues like edges, textures, colors, and boundaries are often overlooked by them during training. One such complex application is Indian Classical Dance Pose Identification (ICDPI) in videos of stage performances where ViTs have learned to notice global features but were unable to master local essentials. Merging local cues with ViTs global features for robust and efficient decision making is the primary objective of this work. Wavelets have shown to identify local cues in ViTs but fail to handle similar object scale factors across data samples effectively. To overcome we propose wavelet convolutional ViTs (WCViT) with multi head attention module designed using multi scale spatial and frequency fusion. The proposed WCViT effectively identifies texture and edges during the training process and registered a higher overall accuracy of 6.1% than the traditional ViTs. Evaluation of WCViT is conducted on our challenging ’Bharatanatyam’ Indian classical dance video dataset constructed using stage performance online videos (BCDSPD 23) and a benchmark ’Let’s Dance’.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Wavelet convolutional vision transformer (WCViT) for Indian classical dance identification

  • P. V. V. Kishore,
  • D. Anil Kumar,
  • G. Hima Bindu,
  • B. Prasad,
  • P. Praveen Kumar,
  • R. Prasad,
  • E. Kiran Kumar

摘要

Vision Transformers (ViTs) significantly improved the recognition accuracies of the image classification tasks by learning global cues on large datasets. Local cues like edges, textures, colors, and boundaries are often overlooked by them during training. One such complex application is Indian Classical Dance Pose Identification (ICDPI) in videos of stage performances where ViTs have learned to notice global features but were unable to master local essentials. Merging local cues with ViTs global features for robust and efficient decision making is the primary objective of this work. Wavelets have shown to identify local cues in ViTs but fail to handle similar object scale factors across data samples effectively. To overcome we propose wavelet convolutional ViTs (WCViT) with multi head attention module designed using multi scale spatial and frequency fusion. The proposed WCViT effectively identifies texture and edges during the training process and registered a higher overall accuracy of 6.1% than the traditional ViTs. Evaluation of WCViT is conducted on our challenging ’Bharatanatyam’ Indian classical dance video dataset constructed using stage performance online videos (BCDSPD 23) and a benchmark ’Let’s Dance’.