Natural Disaster Twitter Data Classification Using CNN and Logistic Regression
摘要
Satellite-based natural disaster management and assessment are prevalent. A significant hurdle is faced in estimating the population affected as well as the internal damage of buildings which cannot be assessed from the top. Social media images and texts can estimate the affected population fairly well. This paper uses Twitter and Flickr data for sentiment analysis and classification using SVM, CNN, XGBoost, Logistic Regression, Gradient Boost, etc. The sentiment analysis gave us information about the panic situation among the people, i.e. panic, no panic or neutral. The best results for text classification were provided by Logistic Regression which gave an accuracy of 83.45% and 88.99% on test and train data, respectively. For image classification, CNN was used, which gave us an accuracy of 83.29%. Since social media reaction is immediate, our system can swiftly assist government agencies and organisations in providing required aid in affected regions based on priority.