The rapid advancements in generative artificial intelligence (AI) have prompted researchers across various professional domains to explore the feasibility and implications of integrating generative AI models, particularly large language models (LLMs) like ChatGPT, into their fields. The emergence of multimodal LLMs such as ChatGPT-4o, capable of handling text, image, and audio, presents opportunities for applying generative AI technologies in structural engineering to address specialized tasks involving structural images. This study evaluates the performance of multimodal large language model, ChatGPT-4o, on visual based structural defect assessment including information extraction, cause analysis and safety assessment. Initially, text-based conceptual questions were used to test ChatGPT-4o's understanding and processing ability of professional knowledge in structural engineering. Subsequently, real structural defect images were employed to explore the potential applications of the model in structural defect assessment. The study analyzed the comprehensiveness and accuracy of the model's information extraction from structural images, focusing on identification, classification, and localization of structural defects. The results indicate that ChatGPT-4o performed well in understanding professional concepts and generating textual solutions, while accurately identifying and localizing common structural defects for analysis. However, the research also revealed ChatGPT-4o's limitations in accurately extracting information from images of complex environments or special types of defects. These findings highlight the significant potential of applying large language models in structural engineering, underscore the necessity of fine-tuning models for specific purposes and tasks, and open up prospects for future integration of detection technology and artificial intelligence.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the Performance of Large Language Model on Structural Defect Assessment

  • Yi Zhang,
  • Ray Kai Leung Su

摘要

The rapid advancements in generative artificial intelligence (AI) have prompted researchers across various professional domains to explore the feasibility and implications of integrating generative AI models, particularly large language models (LLMs) like ChatGPT, into their fields. The emergence of multimodal LLMs such as ChatGPT-4o, capable of handling text, image, and audio, presents opportunities for applying generative AI technologies in structural engineering to address specialized tasks involving structural images. This study evaluates the performance of multimodal large language model, ChatGPT-4o, on visual based structural defect assessment including information extraction, cause analysis and safety assessment. Initially, text-based conceptual questions were used to test ChatGPT-4o's understanding and processing ability of professional knowledge in structural engineering. Subsequently, real structural defect images were employed to explore the potential applications of the model in structural defect assessment. The study analyzed the comprehensiveness and accuracy of the model's information extraction from structural images, focusing on identification, classification, and localization of structural defects. The results indicate that ChatGPT-4o performed well in understanding professional concepts and generating textual solutions, while accurately identifying and localizing common structural defects for analysis. However, the research also revealed ChatGPT-4o's limitations in accurately extracting information from images of complex environments or special types of defects. These findings highlight the significant potential of applying large language models in structural engineering, underscore the necessity of fine-tuning models for specific purposes and tasks, and open up prospects for future integration of detection technology and artificial intelligence.