Multimodal Large Language Models (MLLMs) have demonstrated proficiency in processing diverse modalities, including text, images, and audio. These models leverage extensive pre-existing knowledge, enabling them to address complex problems with minimal to no specific training examples, as evidenced in few-shot and zero-shot in-context learning scenarios. This paper investigates the use of MLLMs’ visual capabilities to ‘eyeball’ solutions for the Traveling Salesman Problem (TSP) by analyzing images of point distributions on a two-dimensional plane. Our experiments aimed to validate the hypothesis that MLLMs can effectively ‘eyeball’ viable TSP routes. The results from zero-shot, few-shot, self-ensemble, and self-refine zero-shot evaluations show promising outcomes. We anticipate that these findings will inspire further exploration into MLLMs’ visual reasoning abilities to tackle other combinatorial problems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Eyeballing Combinatorial Problems: A Case Study of Using Multimodal Large Language Models to Solve Traveling Salesman Problems

  • Mohammed Elhenawy,
  • Ahmed Abdelhay,
  • Taqwa I. Alhadidi,
  • Huthaifa I. Ashqar,
  • Shadi Jaradat,
  • Ahmed Jaber,
  • Sebastien Glaser,
  • Andry Rakotonirainy

摘要

Multimodal Large Language Models (MLLMs) have demonstrated proficiency in processing diverse modalities, including text, images, and audio. These models leverage extensive pre-existing knowledge, enabling them to address complex problems with minimal to no specific training examples, as evidenced in few-shot and zero-shot in-context learning scenarios. This paper investigates the use of MLLMs’ visual capabilities to ‘eyeball’ solutions for the Traveling Salesman Problem (TSP) by analyzing images of point distributions on a two-dimensional plane. Our experiments aimed to validate the hypothesis that MLLMs can effectively ‘eyeball’ viable TSP routes. The results from zero-shot, few-shot, self-ensemble, and self-refine zero-shot evaluations show promising outcomes. We anticipate that these findings will inspire further exploration into MLLMs’ visual reasoning abilities to tackle other combinatorial problems.