Detecting Areas of Interest for Blind People: Deep Learning Saliency Methods for Artworks
摘要
The purpose of this study is to explore human visual attention when observing artworks to create audio descriptions that will guide tactile exploration via a force feedback tablet F2T. To find the semantically important elements, we tested with an Eye-tracker people’s behaviour when observing scenes with and without audio description. The collected data and the small dataset constituted from images of the Bayeux Tapestry will be used to train a deep learning model. This model aims to predict saliency on other images. We use three models to predict images of Bayeux Tapestry: Resnet50, TransalNet, SAM-LSTM-Resnet. We can see that after the training phase on our dataset, the predictions of chosen model (SAM-LSTM-Resnet) are closer to the ground truth and have better correlation, which is a significant improvement over the same model learnt with only the original dataset.