University News: A New Data Source for NLP Bias Research
摘要
This research explores the use of university news articles for Natural Language Processing (NLP) and gender bias detection. It emphasises the importance of ethical considerations in NLP, advocating for transparency and diversity in dataset selection to ensure fairness. Using techniques such as Sentiment Analysis (SA) and gender-specific language classification, the study reveals a bias towards male possessive terms, indicating gender imbalance in the content. While the Facebook BART-Large-Mnli model demonstrated strong accuracy, it struggled with neutral sentiment, suggesting areas for improvement. The study highlights university news as a valuable dataset for promoting equity and inclusivity in NLP tools, laying the foundation for fairer methodologies.