Archaeology is a field of study where identifying and categorizing ancient artifacts is crucial, especially when it comes to inscriptions in languages like Pali. However, due to the unique shape and edges of Pali characters, handwriting recognition and retrieval techniques face several challenges. Manuscripts written in Pali hold significant value in India’s intangible cultural heritage. Unfortunately, the level of digitization and intelligence in preserving Pali manuscript culture is insufficient, as reading Pali text requires a specific skill set. This paper proposes a system to solve the problem of transforming characters into understandable forms with appropriate meaning. In addition, for information extraction and retrieval of data, we are combining machine learning and image processing principles with Natural Language Processing (NLP). NLP combines the field of linguistics and computer science to decipher Pali Prakrit language structure and grammatical phrases to create models that can comprehend, break down, and separate significant details from scripted literature. To validate the proposed method, we used the Government of India dataset, a popular dataset for machine learning experiments involving Pali Script. The outcomes demonstrate the benefits of the suggested strategy in terms of accuracy recognition of characters and memory footprint. The proposed model could help improve character retrieval accuracy as well as the preservation of ancient Pali manuscript characters.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of an AI-Based System for an Ancient Character Extraction and Transformation: The Tipitaka Image Dataset of Pali Manuscript Characters

  • S. R. Gudadhe,
  • A. A. Bardekar,
  • A. B. Ranit

摘要

Archaeology is a field of study where identifying and categorizing ancient artifacts is crucial, especially when it comes to inscriptions in languages like Pali. However, due to the unique shape and edges of Pali characters, handwriting recognition and retrieval techniques face several challenges. Manuscripts written in Pali hold significant value in India’s intangible cultural heritage. Unfortunately, the level of digitization and intelligence in preserving Pali manuscript culture is insufficient, as reading Pali text requires a specific skill set. This paper proposes a system to solve the problem of transforming characters into understandable forms with appropriate meaning. In addition, for information extraction and retrieval of data, we are combining machine learning and image processing principles with Natural Language Processing (NLP). NLP combines the field of linguistics and computer science to decipher Pali Prakrit language structure and grammatical phrases to create models that can comprehend, break down, and separate significant details from scripted literature. To validate the proposed method, we used the Government of India dataset, a popular dataset for machine learning experiments involving Pali Script. The outcomes demonstrate the benefits of the suggested strategy in terms of accuracy recognition of characters and memory footprint. The proposed model could help improve character retrieval accuracy as well as the preservation of ancient Pali manuscript characters.