Text embedding models trained on vast web-scraped corpus generalize well to daily language. However, they often fall short when applied in specialized domains that require precise language and foreign terms, like law and medicine. This gap highlights the necessity for data-efficient methodologies to fine-tune these models for narrow-domain applications. This paper introduces PlanBERT, a new approach for enabling data-efficient domain adaptation and fine-tuning of embedding models. The approach builds on self-supervised contrastive pre-training, synthetic training data generated by large language models (LLMs), and decorrelation of embedding features. The paper also introduces the term “informative vector embeddings” to adjust the training objectives to incentivise more analytics-friendly embeddings and demonstrate that PlanBERT can learn the domain language of zonal plans and outperform larger and more complex state-of-the-art models in challenging real-world zonal plan tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PlanBERT: From Messy Zonal Plans to Informative Vector Embeddings

  • Henrik Brådland,
  • Morten Goodwin,
  • Per-Arne Andersen,
  • Alexander S. Nossum

摘要

Text embedding models trained on vast web-scraped corpus generalize well to daily language. However, they often fall short when applied in specialized domains that require precise language and foreign terms, like law and medicine. This gap highlights the necessity for data-efficient methodologies to fine-tune these models for narrow-domain applications. This paper introduces PlanBERT, a new approach for enabling data-efficient domain adaptation and fine-tuning of embedding models. The approach builds on self-supervised contrastive pre-training, synthetic training data generated by large language models (LLMs), and decorrelation of embedding features. The paper also introduces the term “informative vector embeddings” to adjust the training objectives to incentivise more analytics-friendly embeddings and demonstrate that PlanBERT can learn the domain language of zonal plans and outperform larger and more complex state-of-the-art models in challenging real-world zonal plan tasks.