A DEEP LEARNING-BASED APPROACH FOR URDU AND ENGLISH HANDWRITTEN TEXT RECOGNITION

Authors

  • Nazia Rabnawaz
  • Aiesha Ahmad
  • Humera Batool Gill
  • Muhammad Ahsan Jamil
  • Naeem Aslam

Keywords:

Transformer-based models, multilingual, Optical character recognition (OCR), Urdu OCR, document analysis, and performance evaluation

Abstract

The difficulties presented by low-resource languages are significant when it comes to digitizing written material. Due to the simple linguistic resources that mostly abound within such languages, establishing efficient systems for accurate optical character recognition is an issue that calls for special consideration. Asides highlighting the importance and need to specialize on such languages, this paper introduces ViLanOCR, which is an innovative multilingual OCR system tailored specifically towards Urdu and English languages. The system takes advantage of sophisticated multilingual transformers within its language models that result in greater performances than other approaches, which pose considerable difficulties due to the nature of low resource languages. On the Urdu UHWR testing set, the proposed system achieves state-of-the-art performances with a CER measure of 1.1%. The testing findings show how effective the suggested method is, outperforming Cutting-edge beginnings in the digitization of Urdu handwriting

Downloads

Published

2026-03-31

How to Cite

Nazia Rabnawaz, Aiesha Ahmad, Humera Batool Gill, Muhammad Ahsan Jamil, & Naeem Aslam. (2026). A DEEP LEARNING-BASED APPROACH FOR URDU AND ENGLISH HANDWRITTEN TEXT RECOGNITION. Spectrum of Engineering Sciences, 4(3), 3157–3173. Retrieved from https://www.thesesjournal.com/index.php/1/article/view/3527