Building Information Modeling (BIM) boosts collaboration and efficiency in the construction industry by digitalizing data exchange throughout a building’s lifecycle. However, traditional manual methods for extracting information from unstructured documents for BIM validation are inefficient and error-prone. This study proposes integrating BIM with Natural Language Processing (NLP), specifically Named Entity Recognition (NER), utilizing pre-trained Large Language Models (LLMs) like GPT-4 and Mistral 7B, coupled with prompt engineering techniques to automate extracting requirements from Building Technical Specifications (BTS) documents. The capability of GPT-4 and Mistral 7B to extract BIM entities in zero-shot and few-shot scenarios was compared with Camembert_base, a BERT Transformer-based language model fine-tuned on a dataset of 72 BTS documents. The comparison aimed to assess LLMs’ viability as alternatives to traditional NER models and the impact of domain-specific fine-tuning. Findings indicate GPT4 and Mistral 7B show potential in recognizing BIM entities, yet their effectiveness varies, highlighting the limits of direct LLMs application without further adaptation. Camembert_base outperformed in few-shot scenarios, achieving over 90% F1-score across all entities. This underscores the importance of fine-tuning LLMs with domain-specific data for optimized NER tasks in BIM validation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Named Entity Recognition for Building Information Model Verification Using Large Language Models

  • Insaf Nahri,
  • Romain Pinquié,
  • Philippe Véron,
  • Nicolas Bus,
  • Mathieu Thorel

摘要

Building Information Modeling (BIM) boosts collaboration and efficiency in the construction industry by digitalizing data exchange throughout a building’s lifecycle. However, traditional manual methods for extracting information from unstructured documents for BIM validation are inefficient and error-prone. This study proposes integrating BIM with Natural Language Processing (NLP), specifically Named Entity Recognition (NER), utilizing pre-trained Large Language Models (LLMs) like GPT-4 and Mistral 7B, coupled with prompt engineering techniques to automate extracting requirements from Building Technical Specifications (BTS) documents. The capability of GPT-4 and Mistral 7B to extract BIM entities in zero-shot and few-shot scenarios was compared with Camembert_base, a BERT Transformer-based language model fine-tuned on a dataset of 72 BTS documents. The comparison aimed to assess LLMs’ viability as alternatives to traditional NER models and the impact of domain-specific fine-tuning. Findings indicate GPT4 and Mistral 7B show potential in recognizing BIM entities, yet their effectiveness varies, highlighting the limits of direct LLMs application without further adaptation. Camembert_base outperformed in few-shot scenarios, achieving over 90% F1-score across all entities. This underscores the importance of fine-tuning LLMs with domain-specific data for optimized NER tasks in BIM validation.