Cce-SpERT: A Span-Based Pre-trained Transformer Model for Chinese Civil Engineering Information Joint Extraction
摘要
In this paper, we apply a word encoding averaging approach to Transformer incorporating embedded codes of domain lexicons, and propose a Span-based joint information extraction model Cce-SpERT for Chinese civil engineering information specifications. In the proposed model, we study the multi-domain specialized vocabulary in the domain of Chinese civil engineering, and build a specific domain vocabulary embedding approach based on the word encoding average mechanism. With the aid of domain adaptive and task adaptive pre-trainings, we combine BERT and Transformer to generate a Span-based Transformer model Cce-SpERT for extraction of the entity and inter-entity relation at one-time. The experimental results demonstrate the potential of our model on pre-training, negative sampling, Span-based joint extraction, domain word embedding, and relative position embedding, and show that our model achieves an excellent result of joint information extraction on the regulatory text of Chinese civil engineering.