Detecting Big-5 Personality Dimensions from Text Based on Large Language Models
摘要
Detecting personality from text has been a challenging problem to tackle for a variety of reasons. One is the over-reliance on small datasets. Another is the significant variation in the personality tests used. This work utilizes a large (thus far underutilized) dataset of comments from Reddit, labeled with the Big-5 dimensions - the most widely validated and well accepted measure of personality. This work combines large language models with additional prediction layers to produce personality predictions. For model evaluation, new metrics are adopted to assess the accuracy of the model at various levels of error tolerance. Additionally, a comparison of the Mean Squared Error (MSE) with the previous best results is provided.