Offline Urdu OCR using Ligature based Segmentation for Nastaliq Script

Ankur Rana; Gurpreet Singh Lehal

doi:10.17485/ijst/2015/v8i35/124870

Offline Urdu OCR using Ligature based Segmentation for Nastaliq Script

Affiliations
1 Department of Computer Engineering, Punjabi University, Patiala - 147002, Punjab, India
2 Department of Computer Science, Punjabi University, Patiala - 147002, Punjab, India

Abstract
References
Article Metrics
Refbacks

There are two most popular writing styles of Urdu i.e. Naskh and Nastaliq. Considering Arabic OCR research, ample amount of work has been done on Naskh writing style; focusing on Urdu, which uses Arabic character set commonly used Nastaliq writing style. Due to Nastaliq writing style, Urdu OCR poses many distinct challenges like compactness, diagonal orientation and context character shape sensitivity etc., for OCR system to correctly recognize the Urdu text image. Due to compactness and slanting nature of Nastaliq writing style, existing methods for Naskh style would not give desirable results. Therefore, in this paper, we are presenting ligature based segmentation OCR system for Urdu Nastaliq script. We have discussed in detail various unique challenges for the Urdu OCR and different feature extraction techniques for Ligature recognition using SVM and kNN classifier. The system is trained to recognize 11,000 Urdu ligatures. We have achieved overall 90.29% accuracy tested on Urdu text images.