Language models play a crucial role in day-to-day elementary tasks. They are broadly classified into Large Language Models (LLMs) and Small Language Models (SLMs). LLMs are advance AI models trained on massive datasets to understand and generate human-like text while SLMs are lightweight, efficient models with fewer parameters. The goal of this research is to tackle the challenges posed by LLMs, such as high computational costs, time consuming training of the models, large datasets, and more expensive solutions. This research leveraged Microsoft’s Phi-3 model, an SLM, that proved to be a cost effective alternative that is highly affordable and gives quick and accurate results. The proposed approach focuses on training the Phi-3 model for multi-tasking on three core tasks—reasoning, web search, and mathematical problem-solving. After fine-tuning the Phi-3.5-mini-instruct model, an accuracy of 91.51% has been achieved on Grade School Math 8 K (GSM8K) dataset for mathematical problem-solving, 75% accuracy on Reasoning over Paragraph Effects in Situations (ROPES) dataset for reasoning task. Also a comparative analysis of ROUGE scores for the web search summarization task has been done that showed significance improvement.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Multi-tasking with Small Language Models

  • Ambika Gupta,
  • Navya Gupta,
  • Aashi Gupta,
  • Ritika Kumari

摘要

Language models play a crucial role in day-to-day elementary tasks. They are broadly classified into Large Language Models (LLMs) and Small Language Models (SLMs). LLMs are advance AI models trained on massive datasets to understand and generate human-like text while SLMs are lightweight, efficient models with fewer parameters. The goal of this research is to tackle the challenges posed by LLMs, such as high computational costs, time consuming training of the models, large datasets, and more expensive solutions. This research leveraged Microsoft’s Phi-3 model, an SLM, that proved to be a cost effective alternative that is highly affordable and gives quick and accurate results. The proposed approach focuses on training the Phi-3 model for multi-tasking on three core tasks—reasoning, web search, and mathematical problem-solving. After fine-tuning the Phi-3.5-mini-instruct model, an accuracy of 91.51% has been achieved on Grade School Math 8 K (GSM8K) dataset for mathematical problem-solving, 75% accuracy on Reasoning over Paragraph Effects in Situations (ROPES) dataset for reasoning task. Also a comparative analysis of ROUGE scores for the web search summarization task has been done that showed significance improvement.