<p>In this study, we propose an exemplar sampling algorithm to enhance the performance of end-to-end instance incremental learning for imbalanced administrative and financial document classification. This method uses a Determinantal Point Process (DPP) to select diverse exemplars representative of the dataset. The proposed algorithm is evaluated on both private and public administrative imbalanced datasets and compared with four other sampling algorithms. On the private dataset, our method outperforms all other methods and addresses the forgetting issue in incremental learning entirely. The accuracy, mean recall, weighted precision and mean F<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(_1\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mn>1</mn> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> of DPP are 98.56%, 95.12%, 98.57% and 95.14%, respectively. It surpasses static learning, which uses all documents to train at once, in terms of mean recall by 0.3% while reducing accuracy by 0.07% and weighted precision &amp; mean F<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(_1\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mn>1</mn> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> by only 0.03%. Additionally, on the public RVL-CDIP imbalanced dataset, our sampling algorithm demonstrates superiority over the other sampling algorithms and static learning, especially in achieving the highest accuracy, mean recall and mean F<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(_1\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mn>1</mn> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> of 87.99%, 87.50% and 87.54%, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exemplar sampling algorithm for instance incremental learning on imbalanced document datasets

  • Tri-Cong Pham,
  • Mickaël Coustaty,
  • Aurélie Joseph,
  • Gaspar Deloin,
  • Vincent Poulain d’Andecy,
  • Antoine Doucet

摘要

In this study, we propose an exemplar sampling algorithm to enhance the performance of end-to-end instance incremental learning for imbalanced administrative and financial document classification. This method uses a Determinantal Point Process (DPP) to select diverse exemplars representative of the dataset. The proposed algorithm is evaluated on both private and public administrative imbalanced datasets and compared with four other sampling algorithms. On the private dataset, our method outperforms all other methods and addresses the forgetting issue in incremental learning entirely. The accuracy, mean recall, weighted precision and mean F \(_1\) 1 of DPP are 98.56%, 95.12%, 98.57% and 95.14%, respectively. It surpasses static learning, which uses all documents to train at once, in terms of mean recall by 0.3% while reducing accuracy by 0.07% and weighted precision & mean F \(_1\) 1 by only 0.03%. Additionally, on the public RVL-CDIP imbalanced dataset, our sampling algorithm demonstrates superiority over the other sampling algorithms and static learning, especially in achieving the highest accuracy, mean recall and mean F \(_1\) 1 of 87.99%, 87.50% and 87.54%, respectively.