Tools
Guides

Hex / ASCII / Unicode

Encoding

Convert between text, UTF-8 hex bytes and Unicode code points.

100% client-side No backend
Text
Hex (UTF-8)
Code points
On this page

What is a hex / ASCII / Unicode converter?#

The same piece of text can be written in three genuinely different ways, and this tool lets you look at all three at once:

  • Text — the literal characters as you would type and read them (Hello 世界 🌍).
  • Hex bytes — the raw UTF-8 byte sequence in hexadecimal, one byte per pair (48 65 6C 6C 6F 20 ...). This is what actually travels over a network or sits in a file.
  • Code points — the Unicode scalar values, one U+XXXX per character (U+0048 U+0065 ... U+1F30D). This is how Unicode itself numbers characters, independent of any byte encoding.

The reason these three views matter is that they are not the same length once you leave ASCII — and that is the single source of more text-encoding confusion than anything else. A character is one code point, but it can occupy one to four bytes in UTF-8. The earth-globe emoji 🌍 is a single character (U+1F30D), but in UTF-8 it is four bytes (F0 9F 8C 8D). If you have ever miscounted string length, seen a database reject a string as “too long”, or watched a substring slice land in the middle of a character, you have already met this gap.

Edit any one of the three representations and the other two are derived instantly. The parser is tolerant: hex input accepts spaces, colons, dashes, 0x prefixes; code-point input accepts U+, \u, 0x, or bare hex, separated by spaces, commas, or semicolons.

How to use it#

  1. Use the mode selector above the input box to declare what you are typing: Text, Hex, or Code points.
  2. Paste or edit in the input box. The other two views fill in automatically on the right.
  3. Use Copy next to any of the three result fields to grab that representation.
  4. Sample loads Hello 世界 🌍 so you can see how ASCII, CJK, and an emoji each map differently. Clear empties everything.

Key features#

  • Three views, one source of truth. Edit whichever representation you happen to have; the others are computed, never hand-maintained.
  • UTF-8 bytes, not Latin-1. Hex output is the real on-the-wire UTF-8 byte sequence, which is what almost every modern system uses — not the legacy one-byte-per-character assumption.
  • Code-point aware, including astral characters. for...of iteration means an emoji counts as a single code point, not two surrogate halves, so U+1F30D is what you see — never D83C DF0D.
  • Strict on the way back. Hex bytes that do not form valid UTF-8 (an odd-length string, or a dangling continuation byte) are reported with a clear error instead of decoded into garbled text. Code points outside the valid range, and lone surrogates (U+D800–U+DFFF), are rejected.

Worked example#

Load Sample to get Hello 世界 🌍, and the three derived views are:

Text:       Hello 世界 🌍
Hex bytes:  48 65 6C 6C 6F 20 E4 B8 96 E7 95 8C 20 F0 9F 8C 8D
Code points: U+0048 U+0065 U+006C U+006C U+006F U+0020 U+4E16 U+754C U+0020 U+1F30D

The ASCII letters H e l l o are the boring part: one byte each, code points U+0048 through U+006F. The interesting part is the rest:

  • 世 (U+4E16) is three UTF-8 bytes: E4 B8 96. 界 (U+754C) is also three bytes: E7 95 8C. One character, but not one byte.
  • 🌍 (U+1F30D) is four UTF-8 bytes: F0 9F 8C 8D. The code point is above U+FFFF, so it lives in Unicode’s astral plane and needs a 4-byte UTF-8 sequence — yet it is still a single character, one entry in the code-point list.

Run the experiment in reverse: switch the mode to Code points, paste U+1F600 U+1F604, and watch both the text (😄 😄-family emojis) and the UTF-8 hex appear. That round-trip is exactly how you confirm which bytes correspond to which character.

FAQ#

Why does my emoji take four hex bytes but only one code point?#

Because UTF-8 is a variable-width encoding. Code points below U+0080 fit in one byte; U+0080–U+07FF take two; U+0800–U+FFFF take three; and anything at or above U+10000 (most emojis, historic scripts, musical symbols) takes four. The code point is Unicode’s name for the character; the bytes are how UTF-8 happens to store it. Same character, two different numbers.

What is the difference between U+XXXX and \uXXXX?#

They refer to the same code point but come from different worlds. U+0041 is Unicode’s own notation. A is the escape used in JavaScript, Java, and C string literals. The catch: \uXXXX only holds four hex digits, so it cannot directly express an astral code point like U+1F30D — languages that use \u either need a surrogate pair (🌍) or an extended syntax like \u{1F30D}. This tool’s code-point input accepts either form.

I pasted hex bytes and got an “invalid UTF-8” error. What does that mean?#

The bytes you pasted do not form a legal UTF-8 sequence — most often because the hex string has an odd length (half a byte missing), a leading byte claims a multi-byte sequence that the following bytes do not complete, or the data was actually encoded in something other than UTF-8 (GBK, Shift-JIS, raw Latin-1). The decoder uses strict mode so the problem is reported instead of silently producing replacement characters.

Can I use this to count “real” characters in a string?#

Yes — the code-point count is the count a user would think of, and it is what most UIs should be limiting on. Notice that Hello 世界 🌍 is 10 characters by code points, even though it is 17 bytes in UTF-8 and 11 UTF-16 code units in JavaScript (string.length). If you need a character budget for a tweet or a form field, count code points, not bytes and not string.length.