Tools
Guides

Chinese to Pinyin Converter

Convert

Convert Chinese characters (Hanzi) to pinyin in three tone forms at once — tone marks, tone numbers, and tone-less. Detects polyphones and supports surname mode.

100% client-side No backend
Input
Output
Tone marks
Tone numbers
No tones
Polyphone candidates
Enter Chinese characters to convert.
On this page

What is Pinyin?#

Pinyin is the official system for writing the sounds of Mandarin Chinese in the Latin alphabet. Each Chinese character (Hanzi) maps to one syllable, and that syllable is spelled out with an initial (a consonant), a final (a vowel group), and a tone — one of four pitch contours (plus the neutral tone) that change meaning. mā (mother), má (hemp), mǎ (horse) and mà (to scold) are four completely different words that differ only in tone.

A Hanzi-to-Pinyin converter takes Chinese text and produces its phonetic transcription. The hard part is not the spelling — it is disambiguation. Many characters are polyphones (多音字): the same character has different readings depending on the word it is part of. 重 is zhòng in 重要 (important) but chóng in 重庆 (Chongqing, “double celebration”). A naive dictionary lookup gets these wrong; this tool runs each character through a context-aware engine and additionally flags every polyphone it finds, so you can see all candidate readings at a glance.

This page renders three tone notations at once — tone marks (nǐ hǎo), tone numbers (ni3 hao3), and toneless (ni hao) — so whatever format your downstream system expects (flashcards, subtitles, search indexes, romanized filenames), you can copy it directly without converting twice.

How to use it#

  1. Paste Chinese text into the Input box on the left. Any Hanzi (simplified or traditional) is converted; Latin letters, digits, punctuation and whitespace pass through untouched, so you can mix Hello 你好 123 freely.
  2. The right pane fills in three rows: Symbol (tone marks over vowels), Num (trailing 1–4 digits), and None (no tones). Use the Copy button next to each row.
  3. Tick Surname mode when the leading characters are a family name. This matters because some surnames have a non-default reading: 单 is normally dān but Shàn as a surname; 区 is qū normally but Ōu as a surname.
  4. The Polyphone table lists every character in your input that has more than one possible reading, with all its candidates shown in tone-mark form. Use it to spot-check ambiguous characters.
  5. Copy all grabs all three tone modes at once; Sample loads a demo; Clear empties the input.

Key features#

  • Three tone modes from one pass. Symbol / num / none are computed together from the same engine call, so the three rows always agree — no risk of a num that does not match its symbol.
  • Polyphone (多音字) detection. Each Hanzi is re-queried for all its readings; any character with more than one is listed with every candidate, so you can confirm the engine picked the right one for your context.
  • Surname mode. Toggles the engine’s surname-aware path for the leading characters, fixing names like 单 (Shàn), 区 (Ōu), 解 (Xiè), 朴 (Piáo) that a generic lookup gets wrong.
  • Surrogate-safe counting. The character counter walks the string by Unicode code point, not UTF-16 code unit, so rare Extension-B ideographs (4-byte) count correctly.
  • Dictionary bundled locally. The CJK dictionary ships inside the page bundle — no network call, no API key, nothing leaves your browser.

Worked example#

Type 你好世界 (“hello world”):

Symbol: nǐ hǎo shì jiè
Num:    ni3 hao3 shi4 jie4
None:   ni hao shi jie
Polyphones: 好 → hǎo / hào

The polyphone table already flags 好, which has two readings (hǎo “good” and hào “to be fond of”). Here the engine correctly chose hǎo from context — the table is there so you can see the alternatives and override when the context is unusual.

Now try 重庆 (the city Chongqing). Here the engine reads 重 as chóng, not its more common zhòng:

Symbol: chóng qìng
Num:    chong2 qing4
None:   chong qing
Polyphones: 重 → chóng / zhòng

To see surname mode in action, enter 单先生 (“Mr. Shan”). With the box unchecked the engine reads 单 with its default dān; tick Surname mode and it switches to the surname reading shàn:

Surname off:  dān xiān shēng
Surname on:   shàn xiān shēng

FAQ#

Why does the same character sometimes have two readings?#

Because it is a polyphone (多音字) — one written form, multiple spoken forms, chosen by the word it appears in. 的 is read 4 ways: de (possessive 的), dí (的确 “indeed”), dì (目的 “purpose”), dī (的士 “taxi”). The engine picks the most likely reading from surrounding context, and the Polyphone table surfaces the alternatives so you can override it when the context is unusual.

Which tone format should I use?#

It depends where the text is going. Tone marks are right for human-facing display and for anything learners will read. Tone numbers are what older search systems, some flashcard import formats, and pinyin-sorting code expect. Toneless suits cases where tone is irrelevant — romanized filenames, URL slugs, a rough phonetic index. This page gives you all three so you do not have to convert twice.

Does it support Cantonese or other dialects?#

No — this tool transcribes Mandarin (Putonghua) pinyin only. Cantonese jyutping, Taiwanese bopomofo (zhuyin), and other romanizations are separate systems with their own dictionaries. If you need to switch a text between simplified and traditional characters rather than sound it out, use the simplified-traditional converter.

What does the character counter actually count?#

It counts Hanzi code points — Chinese ideographs — not bytes or UTF-16 units, and it handles 4-byte supplementary-plane characters correctly. 你好 counts as 2, and so does a pair of rare Extension-B characters that JavaScript’s string.length would report as 4. Punctuation and Latin text in your input are ignored by the counter.