错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pretraining an Encoder: The BERT Language Model

  • Pierre M. Nugues

摘要

Transformers have the capacity to store a huge quantity of information. This chapter describes how we can pretrain the encoder part of a transformer on very large raw corpora to fit their parameters from word associations. It then describes how to fine-tune the parameters for applications such as classification, sequence annotation, and question answering.