Arabic Word Counter

Paste Arabic text and read its words (كلمات), sentences (جمل) and length with and without vowel marks (تشكيل).

Sample text. Type or paste to replace it.

Words (كلمات)
18
Characters (حروف)
90
Characters with marks (حروف مع التشكيل)
97
Vowel marks (تشكيل)
7
Without spaces (بدون مسافات)
73
Sentences (جمل)
3
Paragraphs (فقرات)
1
Letters (حروف أبجدية)
68
Digits (أرقام)
1
Punctuation (علامات الترقيم)
4
Bytes in UTF-8 (الحجم بالبايت)
175

More options a limit for words (كلمات)

How it works

This Arabic word counter gives 11 counts for pasted text, labelled in English and Arabic: words (كلمات), characters, vowel marks (تشكيل) and eight more. Length comes with and without the marks, since كَتَبَ and كتب are the same 3 letters.

Read the full guide

What the sample shows

The sample gives 18 words, 90 characters, and 97 with its 7 marks. Paste your own text over it. The field Limit for words (كلمات) says كلمات where a brief says ٣٠٠ كلمة, because Arabic uses the plural only after 3 to 10.

One word, four parts

Arabic writes its one-letter words onto the next word. وسنعيده in the sample is and, will, we return and it, and counts once. A space typed after و adds one. Short vowels are usually left out, and a writer marks only a word that reads two ways: مَنْ is who, مِنْ is from.

Tips
  • Web forms count every mark. If one refuses your text, read Characters with marks (حروف مع التشكيل), not Characters (حروف).
  • Three full stops typed in mid-sentence (...) add a false sentence. The single ellipsis character … does not.
  • Each ـ stretching a word (جمـــيل) is a character and a letter. Remove them before checking a حرف limit.

Questions and answers

Do vowel marks count as characters?

Not in Characters (حروف): كَتَبَ is 3, the same as كتب. Counted with its marks it is 6, and each mark adds two bytes.

Is الكتاب one word or two?

One. The article ال is written onto its noun, and so are و, ف, ب, ل and ك, the future marker س and the pronoun that ends a verb. The page counts by spaces. A marker may count و separately, so ask how yours counts.

Does it count Persian or Urdu text?

Yes, under Arabic labels. The Urdu full stop ۔ (U+06D4) ends a sentence and the Persian digits ۰ to ۹ count. A zero width non-joiner (U+200C) inside a Persian word adds one to the length with marks and none to Characters (حروف). Other languages: word counters.

What the 11 counts mean

Length, the numbers a limit is set in 4

  • Characters (حروف) Everything in the box, spaces and line breaks included. A letter and the marks stacked on it are one character, so كَتَبَ is 3 and so is كتب. لا is two characters, lam and alef, although Arabic always draws the pair as one shape. Each ـ (U+0640 Arabic tatweel) that stretches a word counts as one. The everyday Arabic word for a character, حرف, is also the word for a letter, so ask what a limit set in it includes.
  • Characters with marks (حروف مع التشكيل) Every code point, so each vowel mark is counted as well as the letter under it: كَتَبَ is 6 and كتب is 3. A letter that carries a shadda and a vowel, دَّ, is 3 by itself. The length limit of a web form counts every mark too, which is why a fully marked sentence can be refused by a box it appears to fit. An emoji counts as two in a form and one here. With no marks, joiners or emoji in the text this number and the one above are equal.
  • Without spaces (بدون مسافات) The character count with every space, tab and line break taken out, no-break spaces included. Marks are not added back, so the sample gives 73 here with its marks and without them. The invisible right-to-left mark (U+200F), which keeps punctuation on the correct side where Arabic and Latin text meet, is not a space: it stays in this count and in both character counts, and nothing on screen shows it.
Show all 11 counts (8 more)

Length, the numbers a limit is set in, continued

  • Words (كلمات) Anything between two spaces that holds a letter or a digit. Arabic writes its one-letter words onto the word that follows, so وسنعيده, and we will return it, is four pieces of meaning and one word, and الكتاب, the book, is one. Marks change nothing: كَتَبَ and كتب are one word each. A و typed with a space after it counts as a word of its own, and a comma with no space after it glues two words into one.

Only in Arabic script 1

  • Vowel marks (تشكيل) The marks written above and below the letters: U+064E Arabic fatha, U+064F Arabic damma, U+0650 Arabic kasra, U+0652 Arabic sukun, U+0651 Arabic shadda, the three tanwin endings U+064B to U+064D, and any other combining mark. Most Arabic is written without them, so 0 is normal. مُدَرِّسَة, a woman teacher, carries 5 marks on 5 letters; bare, the same letters are مدرسة, which also spells school. A hamza stored apart from its alef (U+0654 Arabic hamza above) counts here as well.

Structure 2

  • Paragraphs (فقرات) Every row that has something on it; an empty row is skipped. A line of classical Arabic verse, a bayt, is one line made of two equal halves. Typed on one row it is one paragraph, and with a line break between the halves it is two, and two sentences as well, because a line break also ends a sentence. Text lifted from a PDF tends to break at the end of every printed row, which inflates this number the same way.
  • Sentences (جمل) A sentence ends at a full stop, at ؟ (U+061F Arabic question mark), at an exclamation mark and at a line break. The Arabic comma ، and semicolon ؛ never end one. This page reads د. before a name as the title doctor and not as an ending, so وصل د. سامي أمس. is one sentence, and a word that only ends in د still closes its own. Arabic has no capital letters, so the counter cannot tell a pause from an ending: three dots in the middle of a sentence add one that is not there. The single ellipsis character … (U+2026 horizontal ellipsis) does not.

What the text is made of 3

  • Digits (أرقام) Decimal digits in either set. The Eastern forms ٠ to ٩ (U+0660 to U+0669) and the Western 0 to 9 both count, and a number written in either is one word. The Persian digits (U+06F0 to U+06F9) are separate code points and count too. Inside right-to-left text a number still runs left to right, with its units on the right. An Eastern digit is two bytes and a Western one is one.
  • Letters (حروف أبجدية) Letters only: no digits, spaces, punctuation or marks, so كَتَبَ has 3. The alphabet has 28 letters, but more than 28 different ones can turn up here, because ة, ى, ء and the hamza carriers أ, إ, آ, ؤ and ئ each have a code point of their own. The stretching stroke ـ (U+0640) is a letter to Unicode, so جمـــيل has 7 letters where جميل has 4.
  • Punctuation (علامات الترقيم) Every punctuation character. Arabic has its own comma ، (U+060C Arabic comma), semicolon ؛ (U+061B) and question mark ؟ (U+061F), and shares the full stop, the colon and the exclamation mark with Latin text. The comma is turned upside down so that it cannot be taken for the vowel mark damma. The percent sign ٪ (U+066A) counts here, and the angle quotes « and » count one each.

How the text is stored 1

  • Bytes in UTF-8 (الحجم بالبايت) The size of the text once it is saved as UTF-8. Every letter and mark of ordinary Arabic, Persian and Urdu text and every Eastern digit takes two bytes, and a space, a full stop or a Western digit takes one. The ready-made joined forms kept for older systems, such as ﻻ (U+FEFB), take three. Arabic therefore runs close to two bytes a character, and marks add two each: كتب is 6 bytes and كَتَبَ is 12. The sample is 175 bytes for its 90 characters. The invisible right-to-left mark costs three.

How this list was compiled: Own work. The 11 counts are the ones this page shows, picked for what Arabic text does to a counter: marks that are optional, one-letter words written onto the next word, two sets of digits and its own comma, semicolon and question mark. Every example in a note was run through the counter this page uses, with the language set to Arabic and the short form this page adds (د), on 2026-09-20, and a note went in only when the result matched. Character names, code points and categories: https://www.unicode.org/Public/18.0.0/ucd/UnicodeData.txt, version 18.0.0. The marks, which texts carry them in full, and the pair school and teacher: https://en.wikipedia.org/wiki/Arabic_diacritics. The 28 letters, the forms outside them, the lam and alef pair and the direction of numbers: https://en.wikipedia.org/wiki/Arabic_alphabet. The attached article, prepositions, conjunctions and object pronouns: https://en.wikipedia.org/wiki/Arabic_grammar. The future marker: https://en.wikipedia.org/wiki/Arabic_verbs. One-letter conjunctions written as a prefix: https://en.wiktionary.org/wiki/%D9%88. The stretching stroke: https://en.wikipedia.org/wiki/Kashida. The two digit sets: https://en.wikipedia.org/wiki/Eastern_Arabic_numerals. The upside-down comma: https://en.wikipedia.org/wiki/Comma. The question mark: https://en.wikipedia.org/wiki/Question_mark. The right-to-left mark: https://en.wikipedia.org/wiki/Right-to-left_mark. The title abbreviations, listed there under Egyptian Arabic only: https://en.wiktionary.org/wiki/%D8%AF%D9%83%D8%AA%D9%88%D8%B1 (checked 2026-09-20). The zero width non-joiner in Persian, with می‌خواهم: https://en.wikipedia.org/wiki/Zero-width_non-joiner (checked 2026-09-20). The ready-made joined forms, U+FE70 to U+FEFF, kept only for compatibility with older standards: https://en.wikipedia.org/wiki/Arabic_Presentation_Forms-B (checked 2026-09-20). The bayt and its two halves: https://en.wikipedia.org/wiki/Bayt_(poetry). The Arabic terms: https://en.wiktionary.org/wiki/%D8%AD%D8%B1%D9%81 (letter, both plurals), https://ar.wikipedia.org/wiki/%D8%AE%D8%AF%D9%85%D8%A9_%D8%A7%D9%84%D8%B1%D8%B3%D8%A7%D8%A6%D9%84_%D8%A7%D9%84%D9%82%D8%B5%D9%8A%D8%B1%D8%A9 (the same word used for a character in a length limit) and the entries for word, sentence and paragraph with their plurals, https://en.wiktionary.org/wiki/%D9%83%D9%84%D9%85%D8%A9, https://en.wiktionary.org/wiki/%D8%AC%D9%85%D9%84%D8%A9 and https://en.wiktionary.org/wiki/%D9%81%D9%82%D8%B1%D8%A9 (checked 2026-09-20). How a web form measures a length limit: https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/Attributes/maxlength. Two-byte and three-byte ranges: https://en.wikipedia.org/wiki/UTF-8. All checked 2026-09-19 unless a later date is given. Licence: own work