OCR accuracy depends on four things: input DPI, the language pack you select, the contrast of your scan, and how straight the text is on the page. Tesseract, the engine behind GekiPDF's OCR tool, adds a searchable text layer to your PDF. For best results, feed it scans at 200-300 DPI – anything lower and small characters blur together; higher just makes larger files without improving recognition.
Your scan's contrast matters too: dark text on a light background works best, while grey-on-grey or faded ink will be misread. Skewed pages – even by a few degrees – break line detection, so straighten them before uploading. GekiPDF includes English and Chinese (simplified/traditional) among its language packs; pick the right one or you'll get gibberish for non-English text.
Frequently asked questions
What DPI should I use for OCR?
Use 200-300 DPI. Lower than 200 loses detail; higher than 300 just increases file size without improving accuracy.
Does GekiPDF support languages other than English?
Yes, it includes Chinese simplified and traditional among others. Select the correct language pack before running OCR.
Can I fix a skewed scan before OCR?
You can rotate or adjust the page using other tools in GekiPDF before running OCR. Straight text improves recognition significantly.
Do it in the full tool
Open OCR PDF for all options, or jump straight to another step below.
Related tools
- OCR a scanned PDF free (English + Chinese)
- OCR a Scanned PDF
- OCR PDF in English
- OCR PDF in Chinese
- Make a PDF Searchable
- OCR a Japanese PDF with the jpn pack plus English for mixed kanji, kana and Latin text
- PDF to Text
- PDF to Word
- Rotate PDF
- How to compress a PDF to 100 KB
- Compress a PDF to 200 KB (free, online)
- Compress a PDF to 500 KB
- How to compress a scanned PDF (without blurring it)