错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analysis of Parent with Fine Tuned Large Language Model

  • Vaishali Baviskar,
  • Shrinidhi Shedbalkar,
  • Varun More,
  • Sagar Waghmare,
  • Yash Wafekar,
  • Madhushi Verma

摘要

This paper offers a comparative examination of two cuttingedge large language models, Guanaco and Llama, within the realm of natural language comprehension and generation tasks. Guanaco is a model fine-tuned on the open-source LLM Llama itself using Qlora, while Llama is trained on a combination of proprietary and open-source datasets. The assessment encompasses their performance on benchmarks like Massively Multitask Language Understanding (MMLU), Vicuna, and ARC. MMLU benchmark is a comprehensive evaluation of large language models’ capabilities on a wide range of tasks, including summarization, question answering, and natural language inference. ELO rating is a dynamic rating system that calculates the relative skill levels of players in zero-sum games, taking into account the outcome of each game. The abstraction and reasoning corpus (ARC) LLM benchmark is a set of tasks that are designed to evaluate the ability of large language models (LLMs) to reason and solve problems using only their core knowledge. The tasks are based on simple abstract concepts, such as objects, goal states, counting, and basic geometry. It demonstrates that Guanaco achieves strong performance on the MMLU benchmark, even outperforming Llama on the ARC benchmark. On the other hand, Llama excels on the Vicuna benchmark, surpassing Guanaco fine-tuned on open-source data. In a qualitative analysis, both models exhibit strengths and weaknesses. Guanaco showcases the ability to demonstrate theory of mind capabilities, whereas Llama sometimes generates inaccurate or unreliable responses in specific scenarios. Overall, this study sheds light on the performance and attributes of Guanaco and Llama, emphasizing their potential in various language comprehension and generation tasks.