<p>The rapid evolution of large language models (LLMs), GPT-4 Turbo, Google Gemini, Qwen, Meta’s LLaMA 3.1, and DeepSeek-R1 has redefined the landscape of artificial intelligence. In the study, we conduct a hybrid meta-analysis integrating publicly available benchmarks, model cards, technical reports, and open-source repositories to evaluate LLMs across both performance and operational dimensions. Quantitative data were aggregated from standardized tasks such as MMLU (reasoning), HumanEval (code generation), FLORES-200 (translation), and TyDiQA (multilingual Q&amp;A), complemented by efficiency metrics including FLOPs, GPU hours, inference latency, and subscription costs. A big data–driven KPI framework covering scalability index, data-throughput rate, energy per token, and training cost efficiency was applied to enable normalized, cross-model comparison. Results indicate that DeepSeek-R1 demonstrates strong coding and multilingual efficiency, ChatGPT-4 Turbo leads in reasoning accuracy, Gemini Ultra excels in multimodal inference, Qwen is competitive in Chinese-language tasks, and LLaMA 3.1 remains the most adaptable open-source option. Across datasets, DeepSeek-R1 achieved 80.2 ± 1.5% on HumanEval and 78.5 ± 1.8% on MMLU, compared with ChatGPT-4 Turbo’s 86.5 ± 1.9%; these gaps fall within observed heterogeneity (I<sup>2</sup> = 14.6%). The findings highlight trade-offs among accuracy, scalability, and cost efficiency, emphasizing the need for transparent, sustainable, and multimodal LLM development.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Meta-analysis of large language models: benchmarking DeepSeek-R1 against ChatGPT, Gemini, Qwen, and LLaMA

  • Shafique Ahmed Awan,
  • Muazzam Ali Khan Khattak,
  • Abdullah Ayub Khan,
  • Anwar Ali Sathio,
  • Jamil Abedalrahim Jamil Alsayaydeh,
  • Rex Bacarra,
  • Safarudin Gazali Herawan,
  • Rehman Aziz

摘要

The rapid evolution of large language models (LLMs), GPT-4 Turbo, Google Gemini, Qwen, Meta’s LLaMA 3.1, and DeepSeek-R1 has redefined the landscape of artificial intelligence. In the study, we conduct a hybrid meta-analysis integrating publicly available benchmarks, model cards, technical reports, and open-source repositories to evaluate LLMs across both performance and operational dimensions. Quantitative data were aggregated from standardized tasks such as MMLU (reasoning), HumanEval (code generation), FLORES-200 (translation), and TyDiQA (multilingual Q&A), complemented by efficiency metrics including FLOPs, GPU hours, inference latency, and subscription costs. A big data–driven KPI framework covering scalability index, data-throughput rate, energy per token, and training cost efficiency was applied to enable normalized, cross-model comparison. Results indicate that DeepSeek-R1 demonstrates strong coding and multilingual efficiency, ChatGPT-4 Turbo leads in reasoning accuracy, Gemini Ultra excels in multimodal inference, Qwen is competitive in Chinese-language tasks, and LLaMA 3.1 remains the most adaptable open-source option. Across datasets, DeepSeek-R1 achieved 80.2 ± 1.5% on HumanEval and 78.5 ± 1.8% on MMLU, compared with ChatGPT-4 Turbo’s 86.5 ± 1.9%; these gaps fall within observed heterogeneity (I2 = 14.6%). The findings highlight trade-offs among accuracy, scalability, and cost efficiency, emphasizing the need for transparent, sustainable, and multimodal LLM development.