Simplified ↔ Traditional Chinese Converter
ConvertConvert Chinese text between Simplified and Traditional with per-character diff highlighting. Choose general, Taiwan (regional vocabulary like 内存→記憶體), or Hong Kong variants. 100% client-side via opencc-js.
On this page
What is Simplified ↔ Traditional Chinese?#
Chinese is written in two parallel character sets. Simplified (简体) is used in mainland China, Singapore and Malaysia — the result of a 1950s reform that reduced stroke counts to boost literacy. Traditional (繁體) is used in Taiwan, Hong Kong, Macau, and by many overseas communities, and it is also the form you see in classical texts, calligraphy, and most pre-1950s sources.
Converting between the two is not a trivial 1:1 character swap, for two reasons:
- One-to-many characters. Several simplifications merged multiple traditional characters into one simplified form. The simplified 发 maps back to 發 (fā, “to send / prosper”) or 髮 (fà, “hair”) depending on meaning. The simplified 干 maps to 幹, 乾, or 干. A naive lookup table picks one and gets the others wrong.
- Regional vocabulary. The same concept is often a different word in different regions. “Memory” is 内存 in the mainland but 記憶體 in Taiwan; “software” is 软件 in the mainland but 軟體 in Taiwan; “information” is 信息 versus 資訊. Real localization is not done until these phrases are rewritten — swapping characters alone leaves the text sounding mainland-flavoured.
This page handles both. It does character-level conversion for the general case, and it applies region-aware phrase rewriting for Taiwan and Hong Kong, so the output reads naturally to readers in each region rather than sounding like machine-converted text. It also highlights exactly which characters changed, so you can proofread at a glance.
How to use it#
- Pick a direction: Simplified → Traditional or Traditional → Simplified.
- Pick a region:
- General — pure character-level conversion. Use this when you only need the script swapped and do not care about regional wording (classical text, names, or you will localize vocabulary yourself).
- Taiwan — character conversion plus Taiwan phrase variants (内存→記憶體, 软件→軟體, 信息→資訊).
- Hong Kong — character conversion plus Hong Kong phrase variants from its own regional dictionary.
- Paste your text into the Input pane on the left. The Output pane on the right shows the converted text, with every character that differs from the input highlighted.
- The status bar reports how many characters changed. Copy grabs the output; Sample loads a demo; Clear empties the input.
Key features#
- Region-aware phrase rewriting. Selecting Taiwan or Hong Kong does not just swap characters — it rewrites regional vocabulary using the same OpenCC dictionaries used by major localization pipelines, so 内存 genuinely becomes 記憶體 (not just the character swap 內存) for Taiwan readers.
- Diff highlighting. The output is rendered as a sequence of spans, with changed runs wrapped in a highlight. The diff is computed at Unicode code-point granularity (so supplementary-plane characters are handled correctly) and uses an LCS alignment for phrase rewrites where input and output lengths differ.
- Both directions, all regions. Simplified→Traditional and Traditional→Simplified are both first-class, and each works across general / Taiwan / Hong Kong.
- Dictionaries bundled locally. The full OpenCC dictionary ships in the page bundle — no network call, so the tool works offline and your text never leaves the browser.
- Safe rendering. Output is built with controlled DOM APIs (never
innerHTMLof user text), so even adversarial input cannot inject markup.
Worked example#
Start with 中国汉字转换 and direction Simplified → Traditional, region General:
Input: 中国汉字转换
Output: 中國漢字轉換 (4 characters changed: 国→國, 汉→漢, 转→轉, 换→換)
Now switch the region to Taiwan and try 内存 (“memory”). General mode only swaps characters; Taiwan mode rewrites the actual word:
General: 内存 → 內存 (character-level: 内→內)
Taiwan: 内存 → 記憶體 (phrase rewrite — the word Taiwan actually uses)
The Taiwan dictionary rewrites several such terms — 软件 becomes 軟體, 信息 becomes 資訊. Note that a literal term with no Taiwanese equivalent is left character-converted only: 计算机 stays 計算機 in every region, because the regional dictionaries convert only words that have a distinct local form.
To see the other direction, type 繁體中文 with Traditional → Simplified:
Input: 繁體中文
Output: 繁体中文 (1 character changed: 體→体)
Hong Kong uses its own variant dictionary on top of character conversion, distinct from Taiwan’s, so the same source text can come out slightly differently under the two regions — the quickest way to confirm which dictionary is active is to convert a regional term and compare.
FAQ#
What is the difference between General, Taiwan, and Hong Kong?#
General does a 1:1 character swap with no vocabulary changes — fine for names, classical text, or when you will handle wording yourself. Taiwan and Hong Kong layer region-specific phrase dictionaries on top, so 内存→記憶體, 软件→軟體, 信息→資訊 under Taiwan. If your audience is in Taiwan or Hong Kong, always pick the matching region — otherwise the script will be right but the wording will read as mainland-flavoured.
Will it pick the right traditional form for one-to-many characters (like 发 → 發/髮)?#
For phrase-level input it usually does, because the region dictionaries disambiguate by surrounding context. For an isolated ambiguous character with no context, no automatic converter can be certain — 发 on its own genuinely could be either. Always proofread the highlighted characters in legal, medical, or published text.
Why are some characters highlighted in the output?#
The highlight marks every character that changed during conversion, so you can focus your proofreading on the spots that actually moved. The diff is aligned at the code-point level, and for phrase rewrites (where input and output lengths differ) it uses a longest-common-subsequence alignment to keep the highlighted runs meaningful rather than smearing the change across the whole line.
Which direction should I pick if my text is mixed?#
Pick the direction of the majority. The converter does not auto-detect script — a fully correct auto-detect for mixed text needs more context than a character table can provide. If you have a Traditional document with a few Simplified quotes pasted in, run Traditional → Simplified; the Simplified quotes will simply pass through unchanged.