错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Content-Based System for Discrimination Detection of Ecuadorian Text

  • Diego Vallejo-Huanga,
  • Alexis Vallejo,
  • Gustavo Contreras

摘要

The increase in Internet users and the widespread use of social networks has led to higher rates of discrimination. The lack of technological tools and the inability of social networks to identify cyberbullying has generated psychological affection in minority social groups. This research compiled a glossary of segregative terms in the Ecuadorian context, according to four types of discrimination defined by the United Nations. The discriminatory dictionary was constructed using information from various blogs, research papers, and texts containing segregative expressions. This dictionary allowed the development of a content-based software prototype, using NLP-derived techniques and similarity metrics, to classify the type and degree of discrimination of texts in the Ecuadorian context. The prototype includes a web user interface, which allows for analyzing the corpus of a text and shows the degree of discrimination according to its taxonomy. The prototype’s performance was validated through functional tests by experimenting with a dataset of 40 sentences labeled as discriminatory and non-discriminatory. The experimental process used the pre-trained artificial intelligence ChatGPT model as an external validation method. The experiments showed that the cosine vector approach of the proposed prototype has an accuracy of 88%, in contrast to 50% of ChatGPT, for classifying a discriminative text in the Ecuadorian domain.