Archiving Social Media Discussions in Time and Space: A Focus on Refugees from Middle East and Related War Conflicts During Jan 2015 – Apr 2016
摘要
Current research presents an approach for archiving social media discussions in time and space. The author generated Latent Dirichlet Allocation (LDA) models and extracted topics, from data, classified into weekly time periods. A thematic dataset of 1.4M tweets, in Greek, posted from 2015–01-01 00:00:01 to 2016–04-30 23:59:59 GMT was used. Through the use of a transformer, the author applied Named Entity Recognition (NER), extracted approximately 940700 geolocations and geocoded them. In total, 1780 topics were extracted. The next step was to reflect them into a human understandable classification schema. The latter was consisted of classes of three main categories: I. Refugee, II. War, III. Irrelevant. The topics were replicated according to the number of classes, and were weighted, estimating thus their percentage distribution. The topic distribution of the documents (tweets), was considered for reflecting geocoded tweets in the classification schema. Results include related figures and maps displaying the frequency of extracted geolocations in Syria and Iraq. The experimentation also revealed, insights regarding LDA model parametrization, while the use of transformers instead of list based geoparsing procedures provide impressive results especially when manipulating texts which refer to wide geographic areas or to the globe in general. Through current approach it can be revealed what were the most dominant discussions, when were those mostly discussed, regarding which geographic areas, queries vital in many disciplines.