Dynamic Temporal Residual Learning and Attention Rectification
摘要
This chapter further explores how to model long-term contextual dependencies in text images. A dynamic temporal residual learning mechanism is proposed to model the contextual information in feature sequences by introducing the residual learning method into the temporal dimension of an RNN encoder. An attention rectification method is also proposed to mitigate the attention drift problem in the decoder.