Exploring Cost-Effective Machine Learning Classifiers for Text Data: A Literature Review
摘要
The escalating size of Machine Learning (ML) models in Natural Language Processing (NLP) systems presents a challenge for researchers with limited computing resources. This assertion is substantiated by the current research landscape, where major industrial corporations dominate the State-of-The-Art (SoTA) NLP-ML model leaderboard. As substantial ML models entail high computing costs, an economical approach within this research domain is crucial. This paper aims to provide a literature review as a foundational analysis toward exploring the right direction in designing cost-effective or economical NLP-ML models, specifically for text data classification problem. Before delving into a domain-specific NLP-ML review, an introduction to ML algorithms is provided. This introductory section discusses an overview of the most common ML algorithms in recent years, along with the problem domains these algorithms attempt to address, with classification being one of those challenges. The review section comprises three interconnected yet independent subsections: ML for classification, ML for NLP, and the economics of NLP-ML models. Synthesizing information from these review subsections, the subsequent sections present a summary and a future direction. While the summary aims to provide a cohesive conclusion to the entire review, the future direction section includes an example case with a proposed methodology, illustrating where the economical ML model might be implemented. Collectively, these final sections provide a justified direction on how to design an economical NLP-ML classifier, with the hope that it becomes a basic guide for researchers interested in pursuing this field.