Text extracted from social media have distinct characteristics, such as: use of hashtags, emojis, neologisms, mix of languages and informal writing. These features increase the challenge of determining which normalization techniques should be applied in the preprocessing step in order to maximize the performance of classifiers. In this work we evaluate the impact of applying these techniques on the result of models trained to classify Instagram posts as legitimate occurrences of the Portuguese man-of-war. We performed experiments with different techniques and individually evaluated their performance in comparison to models trained with raw data. The results showed that some normalization can be interesting to apply when training the language representation model BERT, while none of them significantly outperformed the classical models trained with raw text.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Effect of Text Normalization on Mining Portuguese Man-of-War Instagram Posts

  • Heloisa F. Rocha,
  • Carlos A. Prolo,
  • Aurora R. Pozo,
  • Carmem S. Hara

摘要

Text extracted from social media have distinct characteristics, such as: use of hashtags, emojis, neologisms, mix of languages and informal writing. These features increase the challenge of determining which normalization techniques should be applied in the preprocessing step in order to maximize the performance of classifiers. In this work we evaluate the impact of applying these techniques on the result of models trained to classify Instagram posts as legitimate occurrences of the Portuguese man-of-war. We performed experiments with different techniques and individually evaluated their performance in comparison to models trained with raw data. The results showed that some normalization can be interesting to apply when training the language representation model BERT, while none of them significantly outperformed the classical models trained with raw text.