Building a specialised Hebrew textual corpus on construction, planning and architecture
摘要
Specialised Textual corpora are essential for advancing Natural Language Processing (NLP) applications, particularly in low-resourced languages. This article presents a detailed methodology for compiling a specialised textual corpus on construction, planning, and architecture (CPA) consisting of 22 million words extracted from 1218 Hebrew publications. The corpus was curated to represent contemporary language and its historical transformations since the 1950s and involved systematic source selection, digitisation, text cleaning, and morphological parsing applied to texts sourced from professional research reports, scholarly studies, guidelines, white papers, reviews, research articles, and legislative decrees. The article discusses specific challenges that resulted from the applied curation and processing methodology while providing a comparative analysis highlighting the significance and added value of constructing a profession-specific and diachronic corpus. The experience and insights gained from the process can benefit the development of a wide variety of specialised textual corpora representing professional discourse in every language.