The Effect of Text Normalization on Mining Portuguese Man-of-War Instagram Posts
摘要
Text extracted from social media have distinct characteristics, such as: use of hashtags, emojis, neologisms, mix of languages and informal writing. These features increase the challenge of determining which normalization techniques should be applied in the preprocessing step in order to maximize the performance of classifiers. In this work we evaluate the impact of applying these techniques on the result of models trained to classify Instagram posts as legitimate occurrences of the Portuguese man-of-war. We performed experiments with different techniques and individually evaluated their performance in comparison to models trained with raw data. The results showed that some normalization can be interesting to apply when training the language representation model BERT, while none of them significantly outperformed the classical models trained with raw text.