This paper investigates the application of Machine Learning techniques and use of Large Language Models (LLMs) to categorize research output in weakly classified academic fields, specifically focusing on classifying publications in Islamic Economics, Finance, Banking, and Business (IEFBB). Using a dataset of about 3,958 Scopus-indexed publications, we employed text classification techniques as a scalable enabler to analyze research trends and collaborations within this niche area. Initial experiments using traditional machine learning methods, among which Neural Network model provided the highest cross-validated accuracy of 85%, and 75% test set accuracy. This provided a benchmark that we then used to explore LLM-based approaches, leveraging GPT-4o-mini model with prompt engineering and fine-tuning. Our results demonstrate that zero-shot prompt approach achieved 82% and 71% accuracy on validation and test sets, respectively. While iterative prompt engineering attempts after error analysis improved classification accuracy to 86% and 73% accuracy on validation and test sets; fine-tuning GPT-4o-mini with a tailored prompt significantly improved classification accuracy, achieving up to 86% on validation and 83% on the test sets. These findings underscore the potential of LLMs in automating the categorization of research output in specialized fields that lack established categorization systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Using LLMs to Enhance Research Output Tracking in Weakly Categorized Research Areas: A Case Study to Classify Publications in Islamic Economics, Finance, Banking and Business

  • Asem Kasem,
  • Imran Alvi,
  • Elie Fares

摘要

This paper investigates the application of Machine Learning techniques and use of Large Language Models (LLMs) to categorize research output in weakly classified academic fields, specifically focusing on classifying publications in Islamic Economics, Finance, Banking, and Business (IEFBB). Using a dataset of about 3,958 Scopus-indexed publications, we employed text classification techniques as a scalable enabler to analyze research trends and collaborations within this niche area. Initial experiments using traditional machine learning methods, among which Neural Network model provided the highest cross-validated accuracy of 85%, and 75% test set accuracy. This provided a benchmark that we then used to explore LLM-based approaches, leveraging GPT-4o-mini model with prompt engineering and fine-tuning. Our results demonstrate that zero-shot prompt approach achieved 82% and 71% accuracy on validation and test sets, respectively. While iterative prompt engineering attempts after error analysis improved classification accuracy to 86% and 73% accuracy on validation and test sets; fine-tuning GPT-4o-mini with a tailored prompt significantly improved classification accuracy, achieving up to 86% on validation and 83% on the test sets. These findings underscore the potential of LLMs in automating the categorization of research output in specialized fields that lack established categorization systems.