Exemplar sampling algorithm for instance incremental learning on imbalanced document datasets
摘要
In this study, we propose an exemplar sampling algorithm to enhance the performance of end-to-end instance incremental learning for imbalanced administrative and financial document classification. This method uses a Determinantal Point Process (DPP) to select diverse exemplars representative of the dataset. The proposed algorithm is evaluated on both private and public administrative imbalanced datasets and compared with four other sampling algorithms. On the private dataset, our method outperforms all other methods and addresses the forgetting issue in incremental learning entirely. The accuracy, mean recall, weighted precision and mean F