text
1 article · search the full text for this term
-
Optimized Text Extraction and E-Repository Development of Hindi and Punjabi Documents Using OCR and NLP Techniques
Abstract: Increase in digitalization of content in the form of text content necessitates powerful document and text extraction systems, particularly for Indian languages such as Hindi and Punjabi. The existing Optical Character Recognition (OCR) solutions support major scripts such as English, leaving a research opportunity for effective recognition of Devanagari and Gurmukhi scripts. This study recommends a modified text extraction algorithm based on Tesseract OCR, accompanied by preprocessing steps of conversion …
Published in Journal of Web Engineering & Technology · Vol. 12, Issue 3, 2025 · pp. 35–43 Read article