PlanBERT: From Messy Zonal Plans to Informative Vector Embeddings
摘要
Text embedding models trained on vast web-scraped corpus generalize well to daily language. However, they often fall short when applied in specialized domains that require precise language and foreign terms, like law and medicine. This gap highlights the necessity for data-efficient methodologies to fine-tune these models for narrow-domain applications. This paper introduces PlanBERT, a new approach for enabling data-efficient domain adaptation and fine-tuning of embedding models. The approach builds on self-supervised contrastive pre-training, synthetic training data generated by large language models (LLMs), and decorrelation of embedding features. The paper also introduces the term “informative vector embeddings” to adjust the training objectives to incentivise more analytics-friendly embeddings and demonstrate that PlanBERT can learn the domain language of zonal plans and outperform larger and more complex state-of-the-art models in challenging real-world zonal plan tasks.