错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Latent Dirichlet Allocation for Topic Modeling and Intelligent Document Classification

  • Rajdeep Chatterjee,
  • Chandan Mukherjee,
  • Siddhartha Chatterjee,
  • Biswaroop Nath

摘要

Document classification plays a pivotal role in facilitating faster and more intelligent information retrieval. This paper specifically focuses on document image classification, utilizing optical character recognition (OCR) to transform it into a natural language processing (NLP) problem. Our research aims to enhance the accuracy of popular models by incorporating Latent Dirichlet Allocation (LDA), a text mining and topic modeling technique based on probabilistic distributions. By integrating LDA, we aim to leverage its capabilities for filtering the OCR-generated textual data. This filtering process enables us to extract important features and improve the performance of our models. In our study, we employed several models, in which the Long Short-Term Memory (LSTM) model, a widely used deep learning architecture, to tackle the document image classification task achieved the most promising results, with the LSTM achieving a notable accuracy rate of 92%. To establish the robustness of our findings, we conduct extensive testing on a separate, larger test dataset, which further validates the accuracy and reliability of our proposed approach. Real-life test scenarios are also considered, wherein our LSTM model successfully identifies different types of documents. The successful application of our model in real-world scenarios highlights its potential for practical implementation.