BlogGuides6 min read
Why Find can't locate Arabic words in a converted document
Arabic search fails when the stored characters differ from what you type, even though they look the same: hamza forms, tashkeel, tatweel, or Persian and Urdu letters such as keheh and Farsi yeh. Word's Arabic Find options cover the first three. Mixed letter codes need a Replace pass or a cleaner conversion.
You have an Arabic document that was converted from a scan or a PDF. The text looks right. You type a word you can see on the page into Find, and Word says there are no matches. Or a client searches the delivered file for a name and finds only half of its occurrences.
This matters more than it seems. An editor who can't search a source book can't check cross-references. A translator who can't find every instance of a term can't keep it consistent. An office team that can't search a contract archive by party name will end up reading files one by one. If a file will be searched after you hand it over, searchability is part of the job.
The cause is almost always the same: the characters stored in the file are not the characters you typed, even though both render identically. This guide covers the four mismatches that cause most failures, what Word can ignore for you, and what has to be fixed in the text itself. If the text is reversed or the letters won't join, that is a different problem. See how to fix reversed Arabic text after converting PDF to Word first.
The four mismatches
| What you typed | What the file may contain | Looks the same? |
|---|---|---|
| احمد (bare alef) | أحمد (alef with hamza above, U+0623), or the reverse | Close, easy to miss |
| كتاب (no vowel marks) | كِتَاب (with fatha, kasra and other tashkeel) | No, but searchers often forget |
| كتاب | كـتـاب (with tatweel, U+0640, stretching the joins) | Nearly, in justified text |
| كتاب (Arabic kaf, U+0643) | کتاب (keheh, U+06A9, the Persian and Urdu kaf) | Often identical in the font |
The last row is the one that wastes the most time. Arabic yeh (U+064A) and Farsi yeh (U+06CC) are another such pair. The W3C's Arabic and Persian layout requirements list keheh and Farsi yeh as Persian letters and Arabic kaf and yeh as the Arabic ones, so a keyboard set to Persian or Urdu produces different characters from an Arabic keyboard. A file built from mixed sources, or by a tool that guessed the wrong one, can contain both, and a search for one will skip the other. In some fonts the final and isolated Farsi yeh has no dots, so it can also be confused with alef maqsura (ى, U+0649).
The W3C document also describes tatweel as a character "inserted between two joined letters to widen their connection." When a converter reproduces justified print by inserting tatweel, every stretched word becomes a different string from its plain form.
Step 1: let Word ignore what it can
Word's Find has three Arabic-specific matching settings. Microsoft documents them as
MatchAlefHamza and MatchKashida, which apply "in an Arabic language document", and
MatchDiacritics, which applies "in a right-to-left language document". When each one is on,
Find requires an exact match on that feature. When it is off, Find can treat the variants as
the same.
In desktop Word on Windows with an Arabic editing language enabled (File > Options > Language), these appear as check boxes in the Find and Replace dialog after you click More >>. The labels and availability vary by version and language setup. If you don't see them, the same settings can be turned off from a macro:
With Selection.Find
.MatchAlefHamza = False
.MatchDiacritics = False
.MatchKashida = False
End With
Turn all three off and search again. If the word is now found, you know which kind of variation the file has. For a delivered file you may still want to clean it up (Step 3), because your client's PDF viewer or search system may not have these options.
Step 2: test for mixed letter codes
None of the three settings covers keheh against kaf or Farsi yeh against Arabic yeh. To test for them:
- Open Replace and search for ک (keheh). If you can't type it, use Insert > Symbol, choose the Arabic subset and pick U+06A9.
- Do the same for ی (Farsi yeh, U+06CC).
- If Find reports matches in a document that should be pure Arabic, the file mixes codes.
For an Urdu or Persian document, run the test the other way round: search for the Arabic kaf and yeh (U+0643 and U+064A) to find letters that came from an Arabic keyboard or an Arabic-trained tool.
Step 3: fix the text, not just the search
Decide which form the file should use, then make it consistent.
- Mixed letter codes: use Replace All to turn every ک into ك and every ی into ي, or the reverse for Urdu and Persian. Only do this on a copy, and only when the whole document is in one language. A bilingual Arabic–Urdu book needs a person to decide paragraph by paragraph.
- Tatweel: if it was inserted only for justification, replace U+0640 with nothing. Keep it where the source uses it on purpose, for example in headings or abbreviations.
- Hamza and tashkeel: don't strip these to make search easier. They carry meaning, and removing them changes the text your client paid to have transcribed. Rely on the search settings instead and tell the recipient how to use them.
Illustrative example. We ran a short Python check in this session. Unicode NFKC
normalization, which repairs presentation-form letters, does not solve any of these four
problems. After unicodedata.normalize("NFKC", text), كتاب written with Arabic kaf and
the same word written with keheh are still different strings. Tatweel and tashkeel survive
unchanged, and أ stays a separate character from ا. So a "normalize" button in a text tool is
not proof that the file is searchable. Consistency has to be checked letter by letter.
Checklist before delivery
Run this on any converted Arabic, Urdu or Persian file that someone will search:
- Three common words you can see on the page are found with all Arabic match options on
- The same words are found in the PDF viewer or system the recipient will use
- Searching for ک and ی finds nothing in an Arabic-only file (or ك and ي in an Urdu or Persian file)
- Tatweel appears only where the source uses it deliberately
- Hamza forms match the source edition, not a "corrected" spelling
- Tashkeel is present where the source has it and absent where it doesn't
- Proper names that recur (authors, places, parties to a contract) are spelled the same way every time
- One full page has been read word by word against the original
Where conversion fits in
Most of these problems are introduced when the text is created. A scan read by OCR that isn't built for Arabic often guesses between look-alike letters. A PDF typed on a Persian keyboard carries its letters into Word. Converting with an Arabic-first tool such as ScribeTools' Arabic OCR can reduce the cleanup. For vocalized texts, see also our answer on OCR and Arabic diacritics. You should still run the checklist above on its output, as on any other tool's.
Limitations
- Word's Arabic matching settings are documented in Microsoft's developer reference. Where they appear in the user interface depends on your Word version, platform and language settings, and Word for the web and for Mac may not offer them.
- Ignoring hamza or diacritics in Find also widens what it matches. A search can then hit different words that share the same base letters, so check each result before a Replace All.
- Replace All on letter codes is safe only for single-language text. In mixed Arabic, Urdu and Persian material it will corrupt the minority language.
- These checks confirm that the file is consistent. They don't prove that OCR read every letter correctly. A wrong letter that is stored consistently will pass every test here, so a human read against the original is still needed before you publish or rely on the file.