Improving Stress Detection with Synthetic Datasets: GPT-4o-Mini and Transformer Model Evaluation
摘要
The rise of mental health issues as a public health concern necessitates efficient detection methods. Social media is now an integral part of daily life, where users frequently share stress-related experiences. This study investigates synthetic data generation techniques to expand datasets and improve the accuracy of stress analysis in social media posts using machine learning models. Using the OpenAI GPT-4o-mini model through its API, we implement two text augmentation techniques, paraphrasing and inspirational augmentation, generating diverse datasets with original, paraphrased, and inspirationally augmented texts. Evaluations on BERT and DistilBERT show that the combined dataset led to the best result, with BERT achieving an accuracy of 84.74%, compared to 73.94% with the baseline. Our experimental evaluation, conducted on real-world user posts, demonstrates that synthetic datasets significantly improve stress analysis accuracy, highlighting its potential for mental health detection in social media.