<p>Pupil detection plays a crucial role in human-computer interaction and biomedical computing applications. In recent years, many pupil detection methods have been proposed to detect the pupil using one single frame of video. However, the performance of these methods may degrade due to blinking, eyelash occlusion, and eye-corner occlusion, and etc. To address this problem, we propose a deep neural network-based pupil detection method using multiple continuous frames of video. The method employs a U-shaped structure composed of multi-path encoders, spatiotemporal feature fusion architecture, and a single-path decoder. First, multiple continuous frames are fed into the multi-path encoders one on one, where each path of encoder extracts features of the input frame respectively. These features are then processed by the proposed spatiotemporal feature fusion architecture of the bidirectional SwinLSTM (BiSwinLSTM), which consists of two SwinLSTMs. The architecture combines the Swin Transformer and LSTM (SwinLSTM) to learn spatiotemporal features from multiple continuous frames. Finally, the single-path decoder receives the spatiotemporal features output by the hybrid deep architecture and generates the pupil detection results for the target frame. Utilizing multi-frame information improves the results of pupil detection. Experiments on the dataset demonstrate that the proposed algorithm outperforms other pupil detection methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A robust pupil detection method based on multiple continuous frames

  • Maosong Jiang,
  • Yanlu Cao,
  • Yeru Xia,
  • Yi Chang,
  • Yongzhong Lin,
  • Wenzhi Zhao,
  • Fei Teng,
  • Wenlong Liu

摘要

Pupil detection plays a crucial role in human-computer interaction and biomedical computing applications. In recent years, many pupil detection methods have been proposed to detect the pupil using one single frame of video. However, the performance of these methods may degrade due to blinking, eyelash occlusion, and eye-corner occlusion, and etc. To address this problem, we propose a deep neural network-based pupil detection method using multiple continuous frames of video. The method employs a U-shaped structure composed of multi-path encoders, spatiotemporal feature fusion architecture, and a single-path decoder. First, multiple continuous frames are fed into the multi-path encoders one on one, where each path of encoder extracts features of the input frame respectively. These features are then processed by the proposed spatiotemporal feature fusion architecture of the bidirectional SwinLSTM (BiSwinLSTM), which consists of two SwinLSTMs. The architecture combines the Swin Transformer and LSTM (SwinLSTM) to learn spatiotemporal features from multiple continuous frames. Finally, the single-path decoder receives the spatiotemporal features output by the hybrid deep architecture and generates the pupil detection results for the target frame. Utilizing multi-frame information improves the results of pupil detection. Experiments on the dataset demonstrate that the proposed algorithm outperforms other pupil detection methods.