错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An On-Device Evaluation Framework for LLMs with Budget-Constrained Subsets

  • Minghao Wang,
  • Enqi Liu,
  • Ping Zhang,
  • Jianhua Huang,
  • Dongdong Zhang,
  • Rui Zhang

摘要

Evaluating LLMs on mobile devices presents challenges, as standard benchmarks like MMLU comprise thousands of questions requiring substantial time and resources. This paper adapts item response theory (IRT) subset selection for on-device evaluation. We present an end-to-end framework for Android smartphones that integrates automated testing via ADB with comprehensive performance monitoring. Using an IRT-sampled subset of 300 questions, we validate that our approach achieves accuracy within 0.13% of full MMLU results while reducing evaluation time. Testing across 12 mobile devices reveals substantial performance variations—up to 2.78 \(\times \) in latency—demonstrating that thermal design, memory configuration, and system optimization are as critical as hardware specifications for practical LLM deployment.