Improving Tourism Image Classification with AI-Generated Descriptions: A Comparative Analysis of LLaVA and BLIP Models
摘要
This study examined the impact of AI-generated image descriptions on the classification of tourism-related images. It specifically compared two AI models, LLaVA and BLIP, and used BER Topic to cluster their generated descriptions. By analyzing two contrasting tourist destinations, Tokyo Disneyland and Kenrokuen, the research revealed that LLaVA is more effective for handling complex and diverse images, whereas BLIP excels with consistent imagery. These results show the significance of choosing an AI model that aligns with the dataset’s characteristics.