BlogPDF to Word6 min read

How to fix reversed Arabic text after converting PDF to Word

Reversed Arabic in Word comes from one of three causes: paragraphs set left-to-right, text extracted in visual order, or letters stored as presentation forms. Word's Right-to-left paragraph button fixes only the first. The other two need a reconversion or OCR that writes logical-order, standard Unicode text.

You converted an Arabic PDF to Word and the result looks wrong. The words run backwards, the letters sit apart instead of joining, or the line looks fine until you type into it and everything jumps. Before you start fixing it line by line, find out which problem you actually have. Some of them take one click to fix in Word. Others can't be fixed from inside the document at all, and every hour spent on them is wasted.

This guide is for anyone who has to deliver an editable Arabic file from a PDF: editors preparing a printed edition for re-typesetting, translators receiving source files, and office teams turning contracts or forms into documents they can edit. If you are still choosing how to convert, start with our Arabic PDF to Word page. This article is about diagnosing and repairing a file you already have.

The three causes, and how to tell them apart

Unicode stores right-to-left text in logical order, the order a person types it. The display is worked out afterwards by the bidirectional algorithm (Unicode Standard Annex #9), which also handles the left-to-right digits and Latin words inside an Arabic sentence. A converted Word file goes wrong when one of three things breaks this model.

SymptomLikely causeFixable in Word?
Words are in the right order but the line is left-aligned, and punctuation lands at the wrong endParagraph direction is set to left-to-rightYes
Every word, or every letter, appears in reverse orderThe converter wrote the text in visual order, the order the glyphs were painted in the PDFNot reliably
Letters look connected but search fails, or they show as isolated shapes that won't join when you editText stored as Arabic presentation forms instead of standard lettersPartly, outside Word

Three quick tests separate them:

  1. Copy one line into a plain-text editor such as Notepad or a browser's address bar. Plain-text editors apply only the standard bidirectional rules. If the words still read backwards there, the text itself is stored in visual order. If it reads correctly, the problem is Word's paragraph formatting.
  2. Search for a common word you can see on the page, such as "الله" or "على". If Find reports no matches while the word is plainly visible, suspect presentation forms or visual order.
  3. Type a single letter in the middle of a word. Correct text re-shapes around the new letter. Presentation-form text keeps its old shapes, and visual-order text often makes the whole line jump.

Fix 1: paragraph direction (one click)

If only the direction is wrong, Word can fix it. Microsoft documents that the Left-to-right and Right-to-left buttons appear in the Paragraph group on the Home tab, but only once a right-to-left language is enabled. If you don't see them:

  1. Go to File > Options > Language.
  2. Add an Arabic dialect as an editing (authoring) language and choose Set as Preferred.
  3. Restart Word, select the affected paragraphs (or the whole document with Ctrl+A) and click Right-to-left.

Then check the edges of each line. Full stops, closing brackets and quotation marks should sit on the left end of an Arabic line, and Latin words and digits should read left to right inside it. If they do, you are done with this cause.

Fix 2: visual-order text (reconvert instead)

When the converter has written the characters in the order they were drawn, the Word file contains Arabic that is literally backwards. Switching the paragraph direction only moves the problem around, because the bidirectional algorithm reorders the reversed characters again. Reversing strings by hand or with a macro looks tempting, but it breaks on every embedded number, Latin word and bracket, and those are exactly the places where errors end up in a contract or a footnote.

The realistic fix is to throw the file away and convert again with a tool that outputs logical-order Unicode. For a text-based PDF that can be a different converter. For a scanned PDF it has to be OCR built for Arabic, because there is no text layer to rescue. ScribeTools turns scanned and text-based Arabic PDFs into editable Word with right-to-left paragraphs, and you can try it on a few pages of your own file before deciding. Whatever you use, check the output with the tests above before you edit anything.

Fix 3: presentation forms (normalize the text)

Arabic letters change shape depending on their position: initial, medial, final or isolated. Unicode includes separate "presentation form" characters for these shapes, mainly for compatibility with older systems. Some PDFs store their text this way, and some converters copy it straight into Word. The text can look acceptable, but Word's search, spell check and any later processing all see different characters from the ones you would type.

Illustrative example. We ran a short Python check in this session. The four characters U+FEB3, U+FEE0, U+FE8E, U+FEE1 (the positional forms of س ل ا م) display as "سلام", but a plain search for سلام does not find them. After Unicode NFKC normalization (unicodedata.normalize("NFKC", text)) they become the standard letters U+0633, U+0644, U+0627, U+0645 and the search matches. If you have a technical colleague, they can normalize the extracted text this way. Be aware that NFKC also changes some other characters, such as certain ligatures and full-width forms, so run it on a copy and compare. Reconverting from the PDF is often simpler, especially if the file also has a direction problem.

Review checklist before you deliver the file

Use this on any converted Arabic Word file, whichever tool produced it:

  • A line copied into a plain-text editor reads in the right order
  • Find locates three common words you can see on the page
  • Every paragraph is set to right-to-left (Latin-only paragraphs excepted)
  • Full stops, brackets and quotation marks sit at the correct end of the line
  • Dates and numbers read the same as in the PDF, including the digit style (١٢٣ or 123)
  • Hijri and Gregorian dates keep their parts in the original order
  • Table columns run right to left and no cells have been merged or split
  • Diacritics (tashkeel) are still attached to their letters where the source has them
  • Footnote markers and numbered lists still point to the right items
  • A spot check of at least one full page against the PDF, read word by word

Limitations

  • These tests tell you what kind of problem you have. They don't prove the text is correct. OCR output can contain wrong letters or dropped dots that pass every test above, so a human read against the original is still needed before you publish or rely on the file.
  • Word's direction controls depend on your Office version and language settings. Microsoft notes that the buttons only appear once a right-to-left language is enabled, and menu names can differ between Windows, Mac and the web version.
  • Heavily vocalized text, marginal notes, mixed Arabic and Latin footnotes, and complex tables are the hardest cases for every converter. Budget extra review time for them.
  • If a PDF was produced from images with a poor-quality hidden text layer, extracting that layer will never give clean text. Running OCR on the page image is the only route.

For scanned books and long editions, see also our guide to digitizing Arabic books.

Sources

  1. Microsoft Support: Using right-to-left languages in Office
  2. Unicode Standard Annex #9: Unicode Bidirectional Algorithm
  3. Unicode FAQ: Bidirectional text