How the text is recognised in your browser
OCR turns a picture of text into real characters. Here each PDF page is drawn to a canvas at 2x scale and read by Tesseract.js on your own device, using the language you select.
You choose one input file (PDF, PNG, JPG or WebP, up to 100 MB) and one language: English, Urdu, Arabic or English plus Urdu. The result appears in a text box, with a page marker line between pages of a multi-page PDF, and can be downloaded as TXT. For English you can also download a searchable PDF. It contains the page images with invisible text placed over each recognised line, so you can search and select text in a PDF viewer.