Automated Classification of Android Games Using Word Embeddings
摘要
Application stores allow developers to organise their applications into general categories, such as social, music, games, etc. For the games category, they suggest a list of subcategories. The need to automate the games classification process to avoid misclassification is increasing as the mobile gaming industry continues to grow. However, there are few researchers focusing on game classification. As their results are poor, this research compares five methods to identify the category of mobile games: Decision Tree, Random Forest, Extreme Gradient Boosting, Gradient Boosting Machines, and Artificial Neural Networks. We use a large collection of 147,128 mobile games to ensure diversity in our data set. We extract features from game descriptions, manifest encodings and from application packages. To capture the semantic meaning of words, we use four embedding methods: Word2Vec, fastText, BERT, MiniLM. For each embedding method, we evaluate the performance of the classifiers. Our results show that Extreme Gradient Boosting can outperform existing work, achieving 81% precision in classification.