Tools
Guides

Word Frequency Analyzer

Text

Count word and character n-gram frequencies in text — CJK via Intl.Segmenter, latin words via Unicode letters, or plain whitespace — with a Top-N table, share, and CSV export.

100% client-side No backend

Remote URLs are not fetched; paste your JSON directly.

Input
Output
Enter text to analyze word frequency.
On this page

What is a word frequency analyzer?#

A word frequency analyzer reads a body of text and counts how often each token appears, then ranks them. It is the tool behind keyword density checks for SEO, stopword discovery for search indexes, vocabulary-size estimates for language learning, and the “most common words” panel in any serious text editor. Count words by hand and you will be wrong; the moment case, punctuation and CJK segmentation enter the picture, intuition fails.

The hard part is deciding what counts as a token. This tool offers four tokenisation modes, because the right answer depends on the text:

  • Space — split on any whitespace. Language-agnostic; treats Hello, and hello as different tokens because of the comma and capital.
  • Regex — maximal runs of Unicode letters (\p{L}+). Drops digits and punctuation, so Hello, becomes Hello. The right default for Latin-script prose.
  • CJK — uses the browser’s Intl.Segmenter at word granularity for proper linguistic segmentation of Chinese/Japanese text (我爱北京 splits on actual word boundaries, not per character). When the browser lacks the segmenter, it falls back to per-Han-character counting and warns you.
  • N-gram — character n-grams of size 2–5 over the whitespace-collapsed stream, for repeat-pattern and typo-cluster analysis.

All counting runs locally; for inputs above a megabyte it moves to a background worker so the page stays responsive.

How to use it#

  1. Pick the mode in the toolbar: Space, Regex, CJK, or N-gram.
  2. Set Top N (default 20, max 1000). Only that many ranked rows are shown, but the totals below the table always reflect every token.
  3. If you chose N-gram, a size selector (2, 3, 4 or 5) appears — pick the gram length.
  4. Tick or untick Case-insensitive (on by default). When on, The and the merge into one count.
  5. Paste your text into the left pane, or hit Sample for a built-in passage.
  6. Read the right pane: a summary (total tokens, unique tokens) and a ranked table of token / count / share. Use Copy for the table or Download CSV for a token,count,share file you can open in any spreadsheet.

If the CJK mode had to fall back because your browser lacks Intl.Segmenter, a small note appears under the table so you know the segmentation is per-character rather than linguistic.

Key features#

  • Four tokenisation modes. Space for raw counts, Regex for clean word counts, CJK for proper Chinese/Japanese segmentation, N-gram for pattern discovery — instead of one mode that is wrong for three out of four texts.
  • Real CJK segmentation. Uses Intl.Segmenter with granularity: "word", the browser’s built-in linguistic segmenter, so 我爱北京天安门 is split on word boundaries, not hacked apart character by character.
  • Honest fallback. When the segmenter is unavailable, the tool degrades to per-Han-character counting and tells you it did, rather than silently producing different numbers.
  • Share, not just count. Every row shows its percentage of all tokens, computed over the full text — so even a Top-20 slice of a 10,000-word document reports meaningful percentages.
  • CSV export. One click downloads a token,count,share file, ready for a spreadsheet or a chart.
  • Deterministic ordering. Ties are broken alphabetically (token ascending), so re-running on identical input yields identical output across browsers and engines.
  • Heavy-input worker. Inputs above ~1 MB move to a background worker, so the main thread and the rest of the page stay responsive.
  • Local and safe. Tokens are rendered with textContent only — user-controlled text can never break out of a table cell.

Worked example#

Paste this short line into the input, with mode = Space and Case-insensitive on:

the quick brown fox the fox

The summary reports 6 total tokens and 4 unique. The ranked table is:

TokenCountShare
fox233.33%
the233.33%
brown116.67%
quick116.67%

Two things to notice. First, fox ranks above the despite the same count, because ties break alphabetically. Second, the shares add up to 100% across all four tokens — the percentages are computed over the full token total, not just the displayed rows, so they stay meaningful even when Top N hides the long tail.

Now switch mode to N-gram with size 2 over the same text (whitespace collapsed): the analyser produces bigrams like th, he, e , q … letting you see which letter pairs repeat, which is how you would hunt for duplicated prefixes or a stuck keyboard.

FAQ#

Which mode should I use for normal English text?#

Regex is the right default for Latin-script prose: it strips punctuation and digits, so word, word. and word all count as the same token, and you get a clean vocabulary count. Use Space only when punctuation itself is meaningful (for example, counting raw whitespace-delimited fields in a log).

How does the CJK mode actually split Chinese?#

It calls the browser’s Intl.Segmenter with granularity: "word", which is a real linguistic segmenter that knows where Chinese words begin and end. So 我爱北京天安门 splits into actual words rather than five separate characters. If your browser does not support Intl.Segmenter, the tool falls back to per-character counting and shows a note — the numbers are still usable for density comparison, just less linguistically precise.

What is the n-gram mode for?#

N-grams reveal repeating substrings. A 2-gram pass over a long text shows which letter or character pairs dominate; a 3-gram or 4-gram pass surfaces repeated prefixes, common morphemes, or typo clusters. It is the standard technique behind keyword-stuffing audits, cipher analysis, and autocomplete ranking.

Why are my percentages based on all tokens, not just the Top N shown?#

Because a percentage is only meaningful against the real total. If you ask for the Top 10 words in a 5,000-word essay, each word’s share should reflect its slice of the whole essay, not its slice of the 10 words displayed. Computing the ratio over the full token total keeps the numbers honest even when the table is truncated.

Is my text uploaded anywhere?#

No. Tokenisation, counting and CSV generation all run in your browser. Large inputs move to a background worker thread, but that thread is still part of this page — nothing is sent to a server.