Text Mining Data Summarization Approach for COVID-19 Using Hybrid Machine Learning Techniques
摘要
The absence of timely identification of COVID-19 cases is a significant obstacle for healthcare professionals, governmental bodies, institutions, and nations in their collective efforts to mitigate the rapid transmission of this lethal pathogen throughout many regions. In this particular context, the preceding body of information pertaining to epidemics has served as a catalyst for researchers to assume a substantial role in the identification and detection of COVID-19 tweets, using Machine Learning (ML) and Deep Learning (DL) approaches. This research presents a sentiment classification approach that utilizes a combination of feature extraction methods and hybrid machine learning approaches to analyze social media Twitter datasets gathered from Twitter API, where in the tweets were prepared for preprocessing and then categorized into three distinct groups: positive, neutral, and negative. In the third step, a range of techniques, such as lemmas, TF-IDF, Word2Vec etc. were used to extract different features from the tweets. These commonly used methods were leveraged to compile feature datasets. Various techniques were used to extract separate datasets for the individual characteristics. In the concluding stage, several machine learning classification algorithms are used to identify sentiment via the utilization of machine learning techniques. In the comprehensive experimental research, it was shown that the Bag-of-Words (BoW) approach yielded superior outcomes when combined with the Hybrid Machine Learning (HML) compared to other current machine learning techniques. The HML model exhibited greater performance compared to the other classifiers, with an accuracy rate of 98.15%. After the tweets have been accurately recognized The proposed multiclass HML obtains an accuracy rate of 95.5% for positive sentiment, which is superior than conventional machine learning classifiers.