Discussions and Future Directions
摘要
This chapter introduces the applications and future directions of Open-set text recognition tasks. First, this chapter introduces some existing applications in OSTR tasks, from different perspectives: recognized language, text recognition granularity, the overall technical route, and open-set classifier design and model implementation. Then we discuss the influence of the Multi-modal Large Language Model on the Open-Set Text Recognition (OSTR) task. We introduce some typical Large Language Models and Multi-modal Large Language Models, and their application to Text Recognition. We also discuss some potentials to leverage the capabilities of MLLMs for the OSTR task: Auto/semi-auto data construction, Downstream task fine-tuning, and Semantic understanding enhancement. In the end, this chapter discusses four main trends in the development of open-set text recognition techniques: (a) Cross-language text recognition technology, (b) Character fine-grained analysis techniques, (c) New character induction discovery techniques, and (d) Incremental language model evolution techniques.