Word Frequency Analyzer
TextCount word and character n-gram frequencies in text — CJK via Intl.Segmenter, latin words via Unicode letters, or plain whitespace — with a Top-N table, share, and CSV export.
Remote URLs are not fetched; paste your JSON directly.
On this page
What is a word frequency analyzer?#
A word frequency analyzer reads a body of text and counts how often each token appears, then ranks them. It is the tool behind keyword density checks for SEO, stopword discovery for search indexes, vocabulary-size estimates for language learning, and the “most common words” panel in any serious text editor. Count words by hand and you will be wrong; the moment case, punctuation and CJK segmentation enter the picture, intuition fails.
The hard part is deciding what counts as a token. This tool offers four tokenisation modes, because the right answer depends on the text:
- Space — split on any whitespace. Language-agnostic; treats
Hello,andhelloas different tokens because of the comma and capital. - Regex — maximal runs of Unicode letters (
\p{L}+). Drops digits and punctuation, soHello,becomesHello. The right default for Latin-script prose. - CJK — uses the browser’s
Intl.Segmenterat word granularity for proper linguistic segmentation of Chinese/Japanese text (我爱北京splits on actual word boundaries, not per character). When the browser lacks the segmenter, it falls back to per-Han-character counting and warns you. - N-gram — character n-grams of size 2–5 over the whitespace-collapsed stream, for repeat-pattern and typo-cluster analysis.
All counting runs locally; for inputs above a megabyte it moves to a background worker so the page stays responsive.
How to use it#
- Pick the mode in the toolbar: Space, Regex, CJK, or N-gram.
- Set Top N (default 20, max 1000). Only that many ranked rows are shown, but the totals below the table always reflect every token.
- If you chose N-gram, a size selector (2, 3, 4 or 5) appears — pick the gram length.
- Tick or untick Case-insensitive (on by default). When on,
Theandthemerge into one count. - Paste your text into the left pane, or hit Sample for a built-in passage.
- Read the right pane: a summary (total tokens, unique tokens) and a ranked table of token / count / share. Use Copy for the table or Download CSV for a
token,count,sharefile you can open in any spreadsheet.
If the CJK mode had to fall back because your browser lacks Intl.Segmenter, a small note appears under the table so you know the segmentation is per-character rather than linguistic.
Key features#
- Four tokenisation modes. Space for raw counts, Regex for clean word counts, CJK for proper Chinese/Japanese segmentation, N-gram for pattern discovery — instead of one mode that is wrong for three out of four texts.
- Real CJK segmentation. Uses
Intl.Segmenterwithgranularity: "word", the browser’s built-in linguistic segmenter, so我爱北京天安门is split on word boundaries, not hacked apart character by character. - Honest fallback. When the segmenter is unavailable, the tool degrades to per-Han-character counting and tells you it did, rather than silently producing different numbers.
- Share, not just count. Every row shows its percentage of all tokens, computed over the full text — so even a Top-20 slice of a 10,000-word document reports meaningful percentages.
- CSV export. One click downloads a
token,count,sharefile, ready for a spreadsheet or a chart. - Deterministic ordering. Ties are broken alphabetically (token ascending), so re-running on identical input yields identical output across browsers and engines.
- Heavy-input worker. Inputs above ~1 MB move to a background worker, so the main thread and the rest of the page stay responsive.
- Local and safe. Tokens are rendered with
textContentonly — user-controlled text can never break out of a table cell.
Worked example#
Paste this short line into the input, with mode = Space and Case-insensitive on:
the quick brown fox the fox
The summary reports 6 total tokens and 4 unique. The ranked table is:
| Token | Count | Share |
|---|---|---|
| fox | 2 | 33.33% |
| the | 2 | 33.33% |
| brown | 1 | 16.67% |
| quick | 1 | 16.67% |
Two things to notice. First, fox ranks above the despite the same count, because ties break alphabetically. Second, the shares add up to 100% across all four tokens — the percentages are computed over the full token total, not just the displayed rows, so they stay meaningful even when Top N hides the long tail.
Now switch mode to N-gram with size 2 over the same text (whitespace collapsed): the analyser produces bigrams like th, he, e , q … letting you see which letter pairs repeat, which is how you would hunt for duplicated prefixes or a stuck keyboard.
FAQ#
Which mode should I use for normal English text?#
Regex is the right default for Latin-script prose: it strips punctuation and digits, so word, word. and word all count as the same token, and you get a clean vocabulary count. Use Space only when punctuation itself is meaningful (for example, counting raw whitespace-delimited fields in a log).
How does the CJK mode actually split Chinese?#
It calls the browser’s Intl.Segmenter with granularity: "word", which is a real linguistic segmenter that knows where Chinese words begin and end. So 我爱北京天安门 splits into actual words rather than five separate characters. If your browser does not support Intl.Segmenter, the tool falls back to per-character counting and shows a note — the numbers are still usable for density comparison, just less linguistically precise.
What is the n-gram mode for?#
N-grams reveal repeating substrings. A 2-gram pass over a long text shows which letter or character pairs dominate; a 3-gram or 4-gram pass surfaces repeated prefixes, common morphemes, or typo clusters. It is the standard technique behind keyword-stuffing audits, cipher analysis, and autocomplete ranking.
Why are my percentages based on all tokens, not just the Top N shown?#
Because a percentage is only meaningful against the real total. If you ask for the Top 10 words in a 5,000-word essay, each word’s share should reflect its slice of the whole essay, not its slice of the 10 words displayed. Computing the ratio over the full token total keeps the numbers honest even when the table is truncated.
Is my text uploaded anywhere?#
No. Tokenisation, counting and CSV generation all run in your browser. Large inputs move to a background worker thread, but that thread is still part of this page — nothing is sent to a server.