Hex / ASCII / Unicode
EncodingConvert between text, UTF-8 hex bytes and Unicode code points.
- Text
- Hex (UTF-8)
- Code points
On this page
What is a hex / ASCII / Unicode converter?#
The same piece of text can be written in three genuinely different ways, and this tool lets you look at all three at once:
- Text — the literal characters as you would type and read them (
Hello 世界 🌍). - Hex bytes — the raw UTF-8 byte sequence in hexadecimal, one byte per pair (
48 65 6C 6C 6F 20 ...). This is what actually travels over a network or sits in a file. - Code points — the Unicode scalar values, one
U+XXXXper character (U+0048 U+0065 ... U+1F30D). This is how Unicode itself numbers characters, independent of any byte encoding.
The reason these three views matter is that they are not the same length once you leave ASCII — and that is the single source of more text-encoding confusion than anything else. A character is one code point, but it can occupy one to four bytes in UTF-8. The earth-globe emoji 🌍 is a single character (U+1F30D), but in UTF-8 it is four bytes (F0 9F 8C 8D). If you have ever miscounted string length, seen a database reject a string as “too long”, or watched a substring slice land in the middle of a character, you have already met this gap.
Edit any one of the three representations and the other two are derived instantly. The parser is tolerant: hex input accepts spaces, colons, dashes, 0x prefixes; code-point input accepts U+, \u, 0x, or bare hex, separated by spaces, commas, or semicolons.
How to use it#
- Use the mode selector above the input box to declare what you are typing: Text, Hex, or Code points.
- Paste or edit in the input box. The other two views fill in automatically on the right.
- Use Copy next to any of the three result fields to grab that representation.
- Sample loads
Hello 世界 🌍so you can see how ASCII, CJK, and an emoji each map differently. Clear empties everything.
Key features#
- Three views, one source of truth. Edit whichever representation you happen to have; the others are computed, never hand-maintained.
- UTF-8 bytes, not Latin-1. Hex output is the real on-the-wire UTF-8 byte sequence, which is what almost every modern system uses — not the legacy one-byte-per-character assumption.
- Code-point aware, including astral characters.
for...ofiteration means an emoji counts as a single code point, not two surrogate halves, soU+1F30Dis what you see — neverD83C DF0D. - Strict on the way back. Hex bytes that do not form valid UTF-8 (an odd-length string, or a dangling continuation byte) are reported with a clear error instead of decoded into garbled text. Code points outside the valid range, and lone surrogates (
U+D800–U+DFFF), are rejected.
Worked example#
Load Sample to get Hello 世界 🌍, and the three derived views are:
Text: Hello 世界 🌍
Hex bytes: 48 65 6C 6C 6F 20 E4 B8 96 E7 95 8C 20 F0 9F 8C 8D
Code points: U+0048 U+0065 U+006C U+006C U+006F U+0020 U+4E16 U+754C U+0020 U+1F30D
The ASCII letters H e l l o are the boring part: one byte each, code points U+0048 through U+006F. The interesting part is the rest:
世(U+4E16) is three UTF-8 bytes:E4 B8 96.界(U+754C) is also three bytes:E7 95 8C. One character, but not one byte.🌍(U+1F30D) is four UTF-8 bytes:F0 9F 8C 8D. The code point is aboveU+FFFF, so it lives in Unicode’s astral plane and needs a 4-byte UTF-8 sequence — yet it is still a single character, one entry in the code-point list.
Run the experiment in reverse: switch the mode to Code points, paste U+1F600 U+1F604, and watch both the text (😄 😄-family emojis) and the UTF-8 hex appear. That round-trip is exactly how you confirm which bytes correspond to which character.
FAQ#
Why does my emoji take four hex bytes but only one code point?#
Because UTF-8 is a variable-width encoding. Code points below U+0080 fit in one byte; U+0080–U+07FF take two; U+0800–U+FFFF take three; and anything at or above U+10000 (most emojis, historic scripts, musical symbols) takes four. The code point is Unicode’s name for the character; the bytes are how UTF-8 happens to store it. Same character, two different numbers.
What is the difference between U+XXXX and \uXXXX?#
They refer to the same code point but come from different worlds. U+0041 is Unicode’s own notation. A is the escape used in JavaScript, Java, and C string literals. The catch: \uXXXX only holds four hex digits, so it cannot directly express an astral code point like U+1F30D — languages that use \u either need a surrogate pair (🌍) or an extended syntax like \u{1F30D}. This tool’s code-point input accepts either form.
I pasted hex bytes and got an “invalid UTF-8” error. What does that mean?#
The bytes you pasted do not form a legal UTF-8 sequence — most often because the hex string has an odd length (half a byte missing), a leading byte claims a multi-byte sequence that the following bytes do not complete, or the data was actually encoded in something other than UTF-8 (GBK, Shift-JIS, raw Latin-1). The decoder uses strict mode so the problem is reported instead of silently producing replacement characters.
Can I use this to count “real” characters in a string?#
Yes — the code-point count is the count a user would think of, and it is what most UIs should be limiting on. Notice that Hello 世界 🌍 is 10 characters by code points, even though it is 17 bytes in UTF-8 and 11 UTF-16 code units in JavaScript (string.length). If you need a character budget for a tweet or a form field, count code points, not bytes and not string.length.