错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

One-Shot Transformer-Based Framework for Visually-Rich Document Understanding

  • Huynh Vu The,
  • Van Pham Hoai,
  • Jeff Yang

摘要

There is a growing need for efficient entity extraction (EE) from business documents. While recent EE models have shown good accuracy for a variety of document templates, fine-tuning these models and acquiring additional training data can be expensive. To address this problem, we propose a novel template-based system for the EE task which does not require model fine-tuning for new entities and templates. The system includes two one-shot transformer-based models: one for template recognition and the other for entity recognition. The document recognition model (OTDC) achieves high accuracy (over 93%) on more than 200 templates of public and private datasets. The entity recognition model (OTER) outperforms recent zero-shot models with regards to the full set of labeled entities in the public SROIE datasets. We have also gathered and annotated the public RVL-CDIP and invoice datasets to showcase the generalization of our OTER models for the EE task across a wide range of document templates, containing both single and multiple-region fields.