Transcription of Ancient Indian Manuscripts Through Artificial Intelligence—Current Status of Technology and the Way Forward
摘要
India holds the largest known collection of ancient and medieval manuscripts, but many of these still remain unread and uncatalogued due to the immense amount of time and effort required to transcribe them manually. Automating this process would be highly beneficial, as it would reveal valuable information about our past that is currently hidden within these manuscripts. While these documents are written in various languages and scripts, Sanskrit is the most common. This paper takes a practical approach to explore the feasibility of using machine learning to transcribe these manuscript treasures using publicly available information on current technology. Some work has already been done on European platforms such as Transkribus and eScriptorium to transcribe texts in Sanskrit, Pali/Prakrit, and Tamil. It is possible to build on these efforts to develop a viable system for transcribing Indian manuscripts on a large scale, but this would require a dedicated effort to build and train models that can handle the complexity of these manuscripts with adequate accuracy and speed. While a deep learning model built by a private commercial enterprise in the USA called Nanonets has been found to be the most user-friendly and can provide transcriptions of thirteenth/fourteenth century Sanskrit and Tamil manuscripts without any training or visual enhancement of images, it is not yet adequately accurate and may be more expensive. Regardless of the method chosen, this project will require financial and technical support from government and/or private entities. The Government of India’s National Mission for Manuscripts could be the focal point for coordinating this activity, and the effort would likely be well worth it.