Getting Arabic text out of a PDF often produces اإلمارات instead of الإمارات, words in reverse order, or words split in the middle. The cause is how PDFs store text: as positioned glyphs, often in visual left-to-right order and sometimes as "presentation form" characters rather than normal letters. Here are the four common ways to extract the text, and when each works.
First: is your PDF digital or scanned?
Try to select a word in the PDF viewer. If you can, the PDF has a text layer (it is "digital", made by Word, a bank system or a browser). If you can't, it is a scan – a picture of text – and you need OCR.
| Method | Digital PDFs | Scanned PDFs | Typical problems |
|---|---|---|---|
| Copy and paste from a viewer | Sometimes | No | Reversed words, broken lam-alef, lost table structure |
| Open the PDF in Microsoft Word | Sometimes | Limited | Text boxes per line, mixed directions, layout drift |
| Google Drive → Open with Google Docs (OCR) | Yes (re-reads it as an image) | Yes | OCR errors on small fonts, tables become plain text, file is uploaded to Google |
| Bidi-aware converter (reads the text layer) | Yes, exact text | No | Fonts without a character map cannot be recovered |
Why a bidi-aware converter gives exact text
For digital PDFs, the exact characters are already in the file. A converter that knows Arabic can:
- reorder each line with the Unicode Bidirectional Algorithm, keeping English words, numbers and IBANs left-to-right inside Arabic text;
- map presentation-form glyphs (U+FE70–U+FEFF) back to standard letters, so search and sorting work;
- repair lam-alef ligatures and words split by spacing;
- rebuild tables from the glyph positions, right-to-left.
Our Arabic PDF → Excel / Word converter does this in your browser – the file is not uploaded. In tests on three Arabic guides published by the UAE Federal Tax Authority, the common pdftotext tool produced 48, 157 and 998 broken lam-alef words; our converter produced none. For more detail see why Arabic text comes out reversed.
Which to use
- Digital PDF, need exact text or tables (statements, invoices, government forms): a bidi-aware converter.
- Scanned PDF: OCR (Google Docs works for many Arabic documents; check names and numbers carefully).
- Confidential documents: prefer a tool that processes the file on your device.
Sources
- Unicode Standard Annex #9 – Unicode Bidirectional Algorithm
- Unicode Arabic Presentation Forms-B (U+FE70–U+FEFF)
Published 2026-10-02 by Karuna Labs. Our tools check file structure and checksums; always review outputs (and payment files in your bank's preview) before relying on them. This is general information, not financial, tax or legal advice.