The need for automating high-volume and unstructured data processing is becoming increasingly critical as organizations aim to improve operational efficiency. In a move to automate workflows, Optical Character Recognition (OCR) plays a key role in it, yet traditional OCR engines frequently struggle with accuracy and efficiency when dealing with complex document layouts and ambiguous text structures. This challenge becomes more pronounced in large-scale tasks, where both speed and precision are essential. In this paper, we address these challenges by introducing LMV-RPA, a new Novel Large Model Voting-based Robotic Process Automation (RPA) system designed to enhance the accuracy and efficiency of OCR workflows. The research problem centers around the limitations of single-engine OCR systems and the need for a more robust solution that would handle complex, high-volume data with greater precision and speed. Our key contribution is the implementation of a majority voting mechanism that integrates the outputs from multiple OCR engines—Paddle OCR, Tesseract OCR, Easy OCR, and DocTR—alongside Large Language Models (LLMs), such as LLaMA 3 and Gemini-1.5-pro. This mechanism improves the OCR output to convert it into structured JSON format, significantly enhancing accuracy, particularly for documents with complex and ambiguous layouts. The methodology of this research is the multi-phase pipeline, where each OCR engine’s text extraction is processed by LLMs. The results are then combined using a majority voting mechanism, to make sure that the most accurate text is selected for conversion. This approach not only increases accuracy but also optimizes processing speed, striking a balance between precision and efficiency. LMV-RPA achieves 99% accuracy in OCR tasks, compared to the baseline model with 94%, while reducing processing time by 80%. Benchmark evaluations validate the system’s scalability, demonstrating that LMV-RPA provides a faster, more reliable, and scalable solution for automating large-scale document processing tasks, outperforming existing OCR-RPA integrations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LMV-RPA: Large Model Voting-Based Robotic Process Automation

  • Osama Hosam Abdellaif,
  • Ahmed Ayman,
  • Ali Hamdi

摘要

The need for automating high-volume and unstructured data processing is becoming increasingly critical as organizations aim to improve operational efficiency. In a move to automate workflows, Optical Character Recognition (OCR) plays a key role in it, yet traditional OCR engines frequently struggle with accuracy and efficiency when dealing with complex document layouts and ambiguous text structures. This challenge becomes more pronounced in large-scale tasks, where both speed and precision are essential. In this paper, we address these challenges by introducing LMV-RPA, a new Novel Large Model Voting-based Robotic Process Automation (RPA) system designed to enhance the accuracy and efficiency of OCR workflows. The research problem centers around the limitations of single-engine OCR systems and the need for a more robust solution that would handle complex, high-volume data with greater precision and speed. Our key contribution is the implementation of a majority voting mechanism that integrates the outputs from multiple OCR engines—Paddle OCR, Tesseract OCR, Easy OCR, and DocTR—alongside Large Language Models (LLMs), such as LLaMA 3 and Gemini-1.5-pro. This mechanism improves the OCR output to convert it into structured JSON format, significantly enhancing accuracy, particularly for documents with complex and ambiguous layouts. The methodology of this research is the multi-phase pipeline, where each OCR engine’s text extraction is processed by LLMs. The results are then combined using a majority voting mechanism, to make sure that the most accurate text is selected for conversion. This approach not only increases accuracy but also optimizes processing speed, striking a balance between precision and efficiency. LMV-RPA achieves 99% accuracy in OCR tasks, compared to the baseline model with 94%, while reducing processing time by 80%. Benchmark evaluations validate the system’s scalability, demonstrating that LMV-RPA provides a faster, more reliable, and scalable solution for automating large-scale document processing tasks, outperforming existing OCR-RPA integrations.