An On-Device Evaluation Framework for LLMs with Budget-Constrained Subsets
摘要
Evaluating LLMs on mobile devices presents challenges, as standard benchmarks like MMLU comprise thousands of questions requiring substantial time and resources. This paper adapts item response theory (IRT) subset selection for on-device evaluation. We present an end-to-end framework for Android smartphones that integrates automated testing via ADB with comprehensive performance monitoring. Using an IRT-sampled subset of 300 questions, we validate that our approach achieves accuracy within 0.13% of full MMLU results while reducing evaluation time. Testing across 12 mobile devices reveals substantial performance variations—up to 2.78 \(\times \) in latency—demonstrating that thermal design, memory configuration, and system optimization are as critical as hardware specifications for practical LLM deployment.