A Vision Transformer Based Indoor Localization Using CSI Signals in IoT Networks
摘要
In recent years, the Channel State Information (CSI) based fingerprint localization method has shown promising growth in locating users indoors. However, deep-learning-based mapping of CSI signals into location remains a challenge due to the signal’s complex nature. The existing Convolutional Neural Network (CNN) algorithms are limited in capturing long-range CSI sub-carrier dependency, which represents location-specific information. In this paper, CNN-aided Vision Transformer is considered, which utilizes both local and global structures present in CSI for improved learning. CNN’s receptive field captures local structure among CSI sub-carriers aiding Transformer to learn global dependency by utilizing its self-attention mechanism. The proposed method outperforms baseline deep learning models such as CNN and Long-Short Term Memory (LSTM) on public CSI fingerprint testbeds.