Unicode Converter — Code Points and UTF-8 Bytes
Convert text to and from Unicode code points, escapes and UTF-8 bytes
Runs in your browser · nothing is uploaded
8 visible characters · 8 code points · UTF-16 length 9
Results
Code points (U+XXXX)
U+0043 U+0061 U+0066 U+00E9 U+0020 U+1F600 U+0020 U+D55C
JavaScript / JSON (\uXXXX)
Caf\u00E9 \uD83D\uDE00 \uD55C
ES6 (\u{...})
Caf\u{E9} \u{1F600} \u{D55C}
HTML hex (&#x...;)
Café 😀 한
HTML decimal (&#...;)
Café 😀 한
UTF-8 bytes (hex)
43 61 66 C3 A9 20 F0 9F 98 80 20 ED 95 9C
UTF-16BE bytes (hex)
00 43 00 61 00 66 00 E9 00 20 D8 3D DE 00 00 20 D5 5C
Bytes per encoding
- UTF-8
- 14 bytes
- UTF-16
- 18 bytes
- EUC-KR (CP949)
- 7 bytes
- 2 characters cannot be written in EUC-KR (not counted)
EUC-KR (CP949) is the legacy Korean encoding: the 2,350 Hangul syllables of KS X 1001 and the other 8,822 modern syllables in the CP949 (UHC) extension all take 2 bytes. In UTF-8 a Hangul syllable takes 3 bytes and an emoji 4.
Character table
| Char | Code point | UTF-8 | UTF-16 | Decimal | Visible char |
|---|---|---|---|---|---|
| C | U+0043 | 43 | 0043 | 67 | 1 |
| a | U+0061 | 61 | 0061 | 97 | 2 |
| f | U+0066 | 66 | 0066 | 102 | 3 |
| é | U+00E9 | C3 A9 | 00E9 | 233 | 4 |
| (invisible) | U+0020 | 20 | 0020 | 32 | 5 |
| 😀 | U+1F600 | F0 9F 98 80 | D83D DE00 | 128512 | 6 |
| (invisible) | U+0020 | 20 | 0020 | 32 | 7 |
| 한 | U+D55C | ED 95 9C | D55C | 54620 | 8 |
What it is
Type some text and you get every common notation at once, each with a copy button: Unicode code points, JavaScript and JSON escapes, the ES6 form, HTML hex and decimal references, and the raw UTF-8 and UTF-16BE bytes. Use it to put non-ASCII text into source code safely, read the \uXXXX sequences in a log, or find out why text turned into mojibake.
It also compares byte counts in UTF-8, UTF-16 and the legacy Korean encoding EUC-KR (CP949), and lists each code point in a table. The reverse direction turns escapes or hex bytes back into text. Everything runs in your browser and nothing is stored.
How to use
- Pick a direction: Text → code or Code → text.
- For Text → code, type in the text box. The example “Café 😀 한” is already there, so results appear right away.
- Turn off “Keep ASCII as is” to escape every character, not just the non-ASCII ones.
- Copy the format you need. The byte comparison and the character table are below.
- For Code → text, keep Auto-detect for escapes and U+ values, even mixed together, or pick a hex byte format for byte lists.
- Copy the decoded text with its button.
How it works
- Code points are written as U+ followed by uppercase hex, padded to at least four digits (U+0041), separated by spaces.
- JavaScript / JSON escapes write one
\uXXXXper UTF-16 code unit, so anything from U+10000 up becomes a surrogate pair. The ES6 form writes one\u{...}per code point. - HTML references write one
&#x…;or&#…;per code point. - Keep ASCII as is leaves printable ASCII (U+0020 to U+007E) unchanged, except the backslash (JavaScript) and & < > “ ’ (HTML), so the output decodes back unambiguously. Line breaks and tabs are always escaped.
- UTF-8 bytes come from the browser’s TextEncoder. UTF-16 writes two bytes per code unit, most significant byte first (big-endian).
- EUC-KR (CP949): a reverse lookup is built by decoding every two-byte pair (lead 0x81 to 0xFE, trail 0x41 to 0xFE) with the browser’s EUC-KR decoder; ASCII is one byte. All 11,172 syllables from 가 to 힣 take two bytes: 2,350 from KS X 1001 and 8,822 from the CP949 (UHC) extension, which is filled in by CP949’s layout if a decoder only knows KS X 1001.
- Visible characters are grapheme clusters from the browser’s Intl.Segmenter (Unicode Standard Annex #29). The table has one row per code point, marks multi-code-point characters and shows the first 200.
- Decoding: U+ values take 4 to 8 hex digits, and spaces or commas between U+ values are treated as separators. Values above U+10FFFF and unpaired surrogates are reported with the character position. Byte input ignores spaces, commas, colons, hyphens and 0x, \x or % prefixes, and invalid UTF-8 is reported with the byte position.
- Input is limited to 20,000 characters.
Examples
| Text | Code point | JavaScript | UTF-8 bytes | UTF-8 / UTF-16 / EUC-KR |
|---|---|---|---|---|
| A | U+0041 | A (or \u0041 with Keep ASCII off) |
41 | 1 / 2 / 1 bytes |
| é | U+00E9 | \u00E9 |
C3 A9 | 2 / 2 / not encodable |
| 中 | U+4E2D | \u4E2D |
E4 B8 AD | 3 / 2 / 2 bytes |
| 한 | U+D55C | \uD55C |
ED 95 9C | 3 / 2 / 2 bytes |
| 😀 | U+1F600 | \uD83D\uDE00 |
F0 9F 98 80 | 4 / 4 / not encodable |
The default example “Café 😀 한” has 8 visible characters, 8 code points and a UTF-16 length of 9. It takes 14 bytes in UTF-8, 18 in UTF-16 and 7 in EUC-KR, where é and the emoji cannot be represented.
Going the other way, auto-detect turns Caf\u00E9 😀 U+D55C into “Café 😀 한”, and the UTF-8 bytes 43 61 66 C3 A9 decode to “Café”.
FAQ
What is the difference between a code point and an encoding?
A code point is the number Unicode gives a character, such as U+00E9 for é. An encoding is how that number is stored as bytes. The same é is C3 A9 in UTF-8 and 00 E9 in UTF-16, and the Korean syllable 한 (U+D55C) is ED 95 9C in UTF-8 but C7 D1 in the legacy Korean EUC-KR encoding.
Why does an emoji become two \uXXXX escapes?
JavaScript and JSON strings are sequences of UTF-16 code units. Characters above U+FFFF need two units, called a surrogate pair, so 😀 (U+1F600) is written \uD83D\uDE00. The ES6 form \u{1F600} names the code point directly, so it needs only one escape.
Why does one emoji count as several code points?
Many emoji are sequences. The family emoji is three person emoji joined by two zero width joiners (U+200D), so it is one visible character but five code points, eight UTF-16 units and 18 UTF-8 bytes. Skin tones, flags and decomposed Korean text work the same way, which is why the summary shows all three counts.
What does "cannot be written in EUC-KR" mean?
EUC-KR (CP949) is a legacy Korean encoding that covers Hangul, Hanja and a set of symbols, but not emoji or letters such as é. Those characters are left out of the EUC-KR byte total and counted separately. If your browser has no EUC-KR decoder, the figure is hidden.
Is my text sent anywhere?
No. Every conversion runs in your browser, and nothing you type is stored or uploaded.
Related tools
HTML Entity Encoder/Decoder
Encode and decode HTML entities for special and non-ASCII characters
URL Encoder/Decoder
A URL percent-encoder and decoder that handles Unicode and emoji
Invisible Character Copy
Copy blank and invisible Unicode characters, and detect or remove hidden ones
Morse Code Translator
Encode and decode International Morse code and hear it played
Fullwidth Halfwidth Converter
Switch text between fullwidth and halfwidth characters, option by option