Within the construction and design engineering sphere, integrating advanced technologies has become indispensable for streamlining processes and enhancing productivity. This paper explores the development and implementation of a design assist tool, combining Natural Language Processing (NLP) methodologies to extract semantic information from digital sources. Moreover, different tools’ ability to extract data from email bodies and attachments has been studied. The extraction process, comprising steps such as file detachment with different formats, data labelling, pre-processing techniques such as tokenisation, and feature engineering, requires selecting appropriate techniques, which this paper examines. In this study, data labelling with six features has been done for 278 emails, each of which contained attachments in various formats such as JPEG, PDF, PNG, DWG, etc. Using various tools and patterns such as Regular Expression (Regex), Tokenisation, Stemming, etc., for pre-processing step, made the data ready for training the model. The paper delves into the utility of libraries and models like SpaCy, NLTK, and BERT for efficient data extraction and natural language processing tasks, offering insights into their comparative strengths and suitability for diverse textual analysis needs. The findings suggest that judicious selection of techniques in an NLP project can significantly streamline processes, resulting in time and resource efficiencies. Moreover, the results indicate automated data extraction substantially reduces internal design review time, translating to expedited design turnover and significant cost savings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Construction Design Efficiency: An Approach to Data Extraction with Natural Language Processing for Technical Drawing

  • Ghazal Salimi,
  • Farzad Rahimian,
  • Ebere Donatus Okonta,
  • Stephen Oliver,
  • Alessandro Di Stefano,
  • Edlira Vakaj,
  • Nick Lane

摘要

Within the construction and design engineering sphere, integrating advanced technologies has become indispensable for streamlining processes and enhancing productivity. This paper explores the development and implementation of a design assist tool, combining Natural Language Processing (NLP) methodologies to extract semantic information from digital sources. Moreover, different tools’ ability to extract data from email bodies and attachments has been studied. The extraction process, comprising steps such as file detachment with different formats, data labelling, pre-processing techniques such as tokenisation, and feature engineering, requires selecting appropriate techniques, which this paper examines. In this study, data labelling with six features has been done for 278 emails, each of which contained attachments in various formats such as JPEG, PDF, PNG, DWG, etc. Using various tools and patterns such as Regular Expression (Regex), Tokenisation, Stemming, etc., for pre-processing step, made the data ready for training the model. The paper delves into the utility of libraries and models like SpaCy, NLTK, and BERT for efficient data extraction and natural language processing tasks, offering insights into their comparative strengths and suitability for diverse textual analysis needs. The findings suggest that judicious selection of techniques in an NLP project can significantly streamline processes, resulting in time and resource efficiencies. Moreover, the results indicate automated data extraction substantially reduces internal design review time, translating to expedited design turnover and significant cost savings.