EUC-KR to UTF-8 Converter (CP949, Shift_JIS)
A converter that re-encodes text and files between Korean, Japanese and Unicode encodings
Runs in your browser · nothing is uploaded
Input
Characters the target cannot hold
70 B in CP949 (extended EUC-KR, Korean)
First bytes (hex)
48 65 6C 6C 6F 21 20 BE C8 B3 E7 C7 CF BC BC BF E4 2E 20 AA B3 AA F3 AA CB AA C1 AA CF A1 A3 0A 6E 61 6D 65 2C 70 68 6F 6E 65 0A 48 6F 6E 67 20 …
Encoding tables follow the WHATWG Encoding Standard, the same rules browsers use for euc-kr and shift_jis. Browsers treat EUC-KR as CP949 (Microsoft’s extension with 8,822 more Hangul syllables); choose "EUC-KR (KS X 1001 only)" to refuse syllables outside the original 2,350. Everything runs in your browser.
What it is
This charset converter fixes files whose text turns into gibberish when they move between programs. Spreadsheets exported by Korean government sites or older Windows software are often saved as EUC-KR (CP949), and modern tools that expect UTF-8 show them as random symbols. Japanese files have the same problem with Shift_JIS. Open the file or paste the text, choose the target encoding and download the same content re-encoded.
How to use
- Choose “Paste text” or “Open a file”.
- For a file, keep the source encoding on “Detect automatically” and check the preview. If it looks broken, pick another source encoding.
- Choose the target encoding. For UTF-8 or UTF-16 you can add a byte order mark; for CP949, EUC-KR or Shift_JIS you choose how to replace characters the target cannot hold.
- Check the size and the first bytes in hex, then download. The file name gets a suffix such as -utf8 or -cp949.
How it works
- Encoding tables follow the WHATWG Encoding Standard, the same rules browsers use when they read euc-kr or shift_jis content.
- CP949 contains all 11,172 modern Hangul syllables: the 2,350 of KS X 1001 plus 8,822 more. With “EUC-KR (KS X 1001 only)” the extra syllables count as characters that cannot be encoded.
- Detection trusts a byte order mark first, then checks whether the bytes are valid UTF-8, and otherwise picks whichever of CP949, Shift_JIS and Windows-1252 reads most naturally.
- Files up to 20 MB can be opened.
Examples
| Text | Target | Bytes |
|---|---|---|
| 한글 | CP949 | C7 D1 B1 DB |
| 한글 | UTF-8 | ED 95 9C EA B8 80 |
| あ | Shift_JIS | 82 A0 |
| € | Windows-1252 | 80 |
| 가 | UTF-8 with BOM | EF BB BF EA B0 80 |
A Korean syllable takes two bytes in CP949 and three in UTF-8, so the UTF-8 version of the same text is larger.
FAQ
Are EUC-KR and CP949 the same?
Not quite. EUC-KR covers the 2,350 Hangul syllables of KS X 1001, while CP949 is Microsoft's extension that adds the other 8,822 modern syllables. Browsers and Windows often use the name EUC-KR for CP949. Pick "EUC-KR (KS X 1001 only)" when the receiving system only accepts the original set.
Why does Excel show broken Korean in my UTF-8 CSV?
Without a byte order mark, Excel on a Korean Windows system may read a CSV as CP949. Convert to UTF-8 with "Add a BOM" switched on, or convert to CP949 for older software.
What happens to characters the target cannot hold?
Emoji or rare characters missing from the target encoding are replaced with a question mark or with an HTML numeric reference such as 😀. The tool lists every character that was replaced.
What if automatic detection is wrong?
Choose the source encoding yourself. When the preview shows readable text with no � characters, you picked the right one. Very short files can look plausible in both CP949 and Shift_JIS.
Is my file uploaded?
No. Reading, detecting, converting and downloading all happen in your browser, and nothing is stored.
Related tools
Mojibake Fixer
A tool that repairs mojibake by reversing every likely pair of mixed-up encodings
Unicode Normalizer
A Unicode normalization tool for text and file names, including decomposed macOS names
CSV to JSON Converter
Convert CSV to JSON (delimiter auto-detect, headers, type conversion, nesting)
Base64 Encoder/Decoder
A Base64 encoder and decoder that gets Unicode right
Unicode Converter
Convert text to and from Unicode code points, escapes and UTF-8 bytes