We present an empirical evaluation of several compact large language models (LLMs) from the Llama 2 family. The models are small enough to be run on a typical consumer machine which allows them to be used offline, enabling increased privacy. Their efficiency at tasks such as giving instructions or explanations on a wide variety of subjects makes them a viable alternative to online language processing tools such as ChatGPT. We evaluate the impact of the GPU offloading, the number of threads and the size of the context on the token generation speed.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of Text Generation Speeds for a Set of Compact Large Language Models

  • Maciej Hojda,
  • Grzegorz Popek

摘要

We present an empirical evaluation of several compact large language models (LLMs) from the Llama 2 family. The models are small enough to be run on a typical consumer machine which allows them to be used offline, enabling increased privacy. Their efficiency at tasks such as giving instructions or explanations on a wide variety of subjects makes them a viable alternative to online language processing tools such as ChatGPT. We evaluate the impact of the GPU offloading, the number of threads and the size of the context on the token generation speed.