Japanese Character Counter

Paste Japanese text and read its character count (文字数) and how many characters are kanji, hiragana and katakana.

Sample text. Type or paste to replace it.

Characters (文字数)
88
Without line breaks (改行を除く)
88
Without spaces (空白を除く)
87
Kanji (漢字)
17
Hiragana (ひらがな)
49
Katakana (カタカナ)
12
Words, estimated (単語数)
54
Sentences (文)
4
Paragraphs (段落)
1
Lines (行数)
1
Punctuation (約物)
8
Bytes in UTF-8 (バイト数)
262

Words in text written without spaces are an estimate: a dictionary splits them, and another browser may split them a little differently.

More options a limit for characters (文字数)

How it works

This Japanese character counter leads with characters (文字数), because Japanese length is set in characters, and splits them into kanji (漢字), hiragana and katakana. Words are marked as an estimate.

Read the full guide

The sample, counted

The sample is 88 characters: 17 kanji, 49 hiragana and 12 katakana. The other 10 are 8 marks, the digit 3 and one full-width space after !. Paste your own text over it. Your maximum goes into Limit for characters (文字数).

Squares, not words

The standard sheet of genkō yōshi, manuscript paper, has 400 squares, one character to a square, so Characters (文字数) is close to a square count. Paper differs: 。」 shares a square, a paragraph starts one square in, and a blank square follows ? or !. Type those blanks as full-width spaces and they count here too.

Tips
  • 3枚 in a brief means three sheets of 400 squares: 1200 at most, since the title and indents take squares.
  • Kanji (漢字) one short? Look for ヶ, which counts as katakana in 一ヶ月. The repeat mark 々 counts as kanji.
  • Bytes in UTF-8 (バイト数) counts three for each kana. Half-width katakana do not help: ガイド is 15 bytes, ガイド 9.

Questions and answers

Why is the word count only an estimate?

Japanese is written without spaces, so nothing marks where a word ends. The browser’s dictionary cuts the text and this page counts the pieces; another browser can cut the same sentence differently. Latin words and numbers are split at spaces and counted exactly.

Do 。 and 、 count as characters?

Yes. Each mark is one character, and each half of 「」 counts. They also show under Punctuation (約物), 8 in the sample. The long-vowel mark ー (U+30FC) is not punctuation: it counts as a kana, with the katakana in コーヒー and the hiragana in らーめん.

Why does 「行くよ。」と言った。 count as two sentences?

The counter ends a sentence at every 。, even one that closes speech inside a longer sentence. Without the inner mark, 「行くよ」と言った。 is one sentence. Subtract one from Sentences (文) for each 。」 followed by more text. Other languages have their own pages under word counters.

What the 12 counts mean

Length, the numbers a limit is set in 4

  • Characters (文字数) All of the text, with its spaces and its line breaks. A kanji, a kana, a small っ or ゃ and a mark such as 。 or 「 count one each, which is also how the squares of manuscript paper are filled. On paper the pair 。」 shares one square, and here it is two characters. An ellipsis written the usual way, ……, is two characters in both places. A rare kanji such as 𠮷 (U+20BB7, a variant of 吉) is one character, not two.
  • Without line breaks (改行を除く) The character count with every line break taken out and every space kept. An essay typed as four paragraphs holds three line breaks, so this number is three lower than the one above. It is the figure to give when a brief counts what you wrote and not the key presses between paragraphs. The full-width space that indents a paragraph still counts here.
  • Without spaces (空白を除く) Line breaks, spaces and tabs all taken out, the full-width space (U+3000 ideographic space) included. Japanese puts no spaces between words, so in most texts this differs from the count above only by the blank that indents each paragraph and the one that usually follows ? or !. The sample has one of those, after !, so it reads 87 here against 88.
  • Words, estimated (単語数) Japanese leaves no spaces between words, so a reader decides where each word ends, and so must a counter. The browser’s dictionary cuts the text into pieces and the page counts the pieces that hold a letter or a digit. Another browser may cut the same sentence differently, so the page marks this number as an estimate. Japanese briefs set their limits in characters, not words. A stretch in Latin letters or digits, Tokyo 2026, is counted by its spaces and is exact.
Show all 12 counts (8 more)

The three scripts 3

  • Hiragana (ひらがな) Every hiragana, small ones included: きょう is three. A long-vowel mark that follows a hiragana counts here, so らーめん on a shop sign is four, although standard spelling writes a long vowel in hiragana with another vowel kana. が stored as か plus a separate voicing mark (U+3099) is still one. Particles and verb endings are written in hiragana, which is how the sample reaches 49.
  • Kanji (漢字) Every Han character, whether or not it is one of the 2,136 jōyō kanji. The repeat mark 々 (U+3005) counts as a kanji, so 時々 is two, and so does 〇 (U+3007), the zero of a number written in kanji digits such as 五〇. A kanji beyond the basic plane, such as 𠮷, counts once. The small ヶ in 一ヶ月 is a katakana to Unicode and is counted there, not here.
  • Katakana (カタカナ) Every katakana, with the long-vowel mark ー (U+30FC) counted along with the kana before it: コーヒー is four. Half-width katakana count too. Their voicing marks are separate characters in the encoding, yet ガイド comes out as three katakana and three characters, the same as ガイド, because a mark is counted with its kana. The sample’s 12 are ドア, ラベル, コーヒー and ブーン.

Structure 3

  • Lines (行数) Blank lines count here, so this number is at least the paragraph count. A list typed on three lines shows 3, and a line break left after the last one makes it 4. These are the lines you typed, not the 20-square columns of manuscript paper: a long paragraph is one line here however many columns it would fill.
  • Paragraphs (段落) A paragraph here is a line with text on it. A blank line is not one, and neither is a line that holds nothing but the full-width space of an indent. A title, a name and three paragraphs of body text make five. A PDF tends to hand over its text with a break after each printed line, and every one of those lines is then a paragraph and the end of a sentence.
  • Sentences (文) A sentence ends at 。, at ? or ! in either width, and at a line break. No space is needed after the mark. Speech written 「行くよ。」と言った。 counts as two sentences, because the 。 inside the bracket ends one. Written 「行くよ」と言った。 it is one. A doubled mark such as !? ends one sentence and not two, and an ellipsis in mid-sentence ends none.

What the text is made of 1

  • Punctuation (約物) Every character Unicode classes as punctuation: 。 and 、, the brackets 「」 and 『』 (each half counts), the middle dot ・ and the ellipsis …. The long-vowel mark ー is a letter and is not counted. The wave dash in 5時〜6時 counts when it is U+301C wave dash. The look-alike U+FF5E fullwidth tilde is a maths symbol to Unicode, so the same range typed with it shows one mark fewer.

How the text is stored 1

  • Bytes in UTF-8 (バイト数) What the text weighs when it is stored as UTF-8. Every kana, kanji and full-width mark takes three bytes, where a, 1 and an ordinary space take one, so Japanese text weighs about three times its character count: the sample is 88 characters and 262 bytes. A kanji beyond the basic plane, such as 𠮷, takes four. Half-width katakana save nothing, since ガイド is 15 bytes and ガイド is 9.

How this list was compiled: Own work. The 12 counts are the ones this page shows, picked for how Japanese length is set (in characters, 文字数, and in sheets of 400-square manuscript paper) and for what a text in three scripts does to a counter. Every example in a note was run through the counter this page uses, with the language set to Japanese, on 2026-09-19, and a note went in only when the result matched. No Japanese word count is quoted, because that number comes from the dictionary of the browser doing the counting. Character names, code points, categories and the decomposition of が: https://www.unicode.org/Public/18.0.0/ucd/UnicodeData.txt, version 18.0.0. The three scripts and what each is used for, writing without spaces, and the 2,136 jōyō kanji: https://en.wikipedia.org/wiki/Japanese_writing_system. One square for each character, small kana and mark, the shared square of a full stop and a closing bracket, the indent of a paragraph, the two squares of an ellipsis and the columns of twenty squares: https://en.wikipedia.org/wiki/Genk%C5%8D_y%C5%8Dshi. Length quoted in sheets of 400 characters, and the word 文字数: https://ja.wikipedia.org/wiki/%E5%8E%9F%E7%A8%BF%E7%94%A8%E7%B4%99. The term 約物, the full-width space and the blank that usually follows ? and !, the six-dot ellipsis, the wave dash for ranges and the middle dot: https://en.wikipedia.org/wiki/Japanese_punctuation. The long-vowel mark in katakana and in らーめん and おーい: https://en.wikipedia.org/wiki/Ch%C5%8Donpu. The repeat mark in 時々 and 人々: https://en.wikipedia.org/wiki/Iteration_mark. Half-width katakana and their separate voicing marks: https://en.wikipedia.org/wiki/Half-width_kana. The zero 〇 in numbers written with kanji digits: https://en.wikipedia.org/wiki/Japanese_numerals. 𠮷 as a variant of 吉 in CJK Unified Ideographs Extension B: https://en.wiktionary.org/wiki/%F0%A0%AE%B7. One, three and four byte ranges: https://en.wikipedia.org/wiki/UTF-8. The terms 改行 (https://en.wiktionary.org/wiki/%E6%94%B9%E8%A1%8C), 空白 (https://en.wiktionary.org/wiki/%E7%A9%BA%E7%99%BD), 段落 (https://en.wiktionary.org/wiki/%E6%AE%B5%E8%90%BD) and 単語 (https://en.wiktionary.org/wiki/%E5%8D%98%E8%AA%9E), each opened 2026-09-20. How the browser finds words, sentences and characters: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Intl/Segmenter. That Japanese word boundaries come from a dictionary: https://unicode-org.github.io/icu/userguide/boundaryanalysis/ (checked 2026-09-20). All checked 2026-09-19 unless a later date is given. Licence: own work