Fine-tuning Large Language Models (LLMs) currently requires significant computational resources and time, hindering their widespread adoption. Fine-tuning LLMs unlocks their potential for specialized tasks but often comes at a high cost in computational resources and time. This significantly limits their real-world applicability in research, especially in resource-constrained settings. This paper introduces a holistic strategy for optimizing LLM fine-tuning by leveraging QLoRA (Quantized Low-Rank Adaptation), which reduces the computational resources and time required for LLaMa2-7B fine-tuning. The effectiveness of this approach is demonstrated through comprehensive evaluations, comparing the Precision, Recall, and F1 scores by employing BERTScore against benchmarks such as the Gemini and GPT with the doctor’s answers as ground truth. The proposed method achieves a significant reduction in both computational resource consumption and fine-tuning time compared to baseline approaches. Notably, this efficiency gain is achieved while maintaining competitive performance with benchmark models like Gemini and GPT. This is evident from the F1-scores obtained using BERTScore, where the proposed method achieves a score of 0.841, while Gemini and GPT score 0.849 and 0.855, respectively. Despite achieving slightly lower F1 scores, the proposed method offers a smaller model size and a significantly faster fine-tuning process, making it a compelling alternative for resource-constrained environments. Our experiments explore optimizing fine-tuning within constrained resources, making them particularly relevant for seeking low-cost solutions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing LLaMa2 Fine-Tuning with LoRA and Quantization for Medical Question Answering

  • Trung-Tin Bui,
  • Dat Truong,
  • Lua Ngo

摘要

Fine-tuning Large Language Models (LLMs) currently requires significant computational resources and time, hindering their widespread adoption. Fine-tuning LLMs unlocks their potential for specialized tasks but often comes at a high cost in computational resources and time. This significantly limits their real-world applicability in research, especially in resource-constrained settings. This paper introduces a holistic strategy for optimizing LLM fine-tuning by leveraging QLoRA (Quantized Low-Rank Adaptation), which reduces the computational resources and time required for LLaMa2-7B fine-tuning. The effectiveness of this approach is demonstrated through comprehensive evaluations, comparing the Precision, Recall, and F1 scores by employing BERTScore against benchmarks such as the Gemini and GPT with the doctor’s answers as ground truth. The proposed method achieves a significant reduction in both computational resource consumption and fine-tuning time compared to baseline approaches. Notably, this efficiency gain is achieved while maintaining competitive performance with benchmark models like Gemini and GPT. This is evident from the F1-scores obtained using BERTScore, where the proposed method achieves a score of 0.841, while Gemini and GPT score 0.849 and 0.855, respectively. Despite achieving slightly lower F1 scores, the proposed method offers a smaller model size and a significantly faster fine-tuning process, making it a compelling alternative for resource-constrained environments. Our experiments explore optimizing fine-tuning within constrained resources, making them particularly relevant for seeking low-cost solutions.