If you copy Arabic from a PDF and paste it into Excel, Word or a chat, you may get one of these problems:
- Reversed words or lines – the PDF stores glyphs in visual (left-to-right) order and the copying program does not reverse them back;
- Broken lam-alef – السلام becomes السالم or الإمارات becomes اإلمارات, because the ligature glyph is mapped to two letters in the wrong order;
- Words split by spaces – تحويل becomes تحو يل, because letters that don't join are spaced like separate words;
- Strange characters – "presentation form" code points (U+FE70–U+FEFF) that look right but don't match when you search.
The fix
Our converter reads the PDF's text layer with the positions of every glyph and repairs these cases: it reorders each line using the Unicode bidirectional rules, fixes ligatures, merges split words, and normalises presentation forms to standard letters. In tests on three real Arabic guides published by the UAE Federal Tax Authority (created with Microsoft Word), the common pdftotext tool produced 48, 157 and 998 broken lam-alef words; this converter produced none.
It works on digital PDFs only. For scanned documents you need OCR software.