Mojibake Fixer — Repair Garbled Text Encoding
A tool that repairs mojibake by reversing every likely pair of mixed-up encodings
Runs in your browser · nothing is uploaded
The example is UTF-8 text that was read as Windows-1252. Replace it with your own garbled text.
1 candidates. Most likely: saved as UTF-8, read as Windows-1252 (Latin-1)
Most likely
Café crème brûlée — naïve résumé
saved as UTF-8, read as Windows-1252 (Latin-1)
Every pair of UTF-8, CP949 (EUC-KR), Shift_JIS and Windows-1252 (12 pairs) is reversed, and text garbled twice is reversed once more. Candidates rank higher with more common letters (Latin, kana, the 2,350 common Hangul syllables of KS X 1001) and fewer �, control codes and stray Latin-1 symbols. Tables follow the WHATWG Encoding Standard; your text stays in your browser.
What it is
The mojibake fixer turns garbled text back into what it was meant to say. It helps when an email subject shows “é” instead of “é”, a database export is full of “’” where apostrophes should be, or Korean and Japanese text has turned into strings of odd symbols. You don’t need to know which encodings were mixed up: every likely pair is tried and the most natural results are listed first.
How to use
- Paste the garbled text. An example is filled in to start with.
- Look at the “Most likely” candidate first.
- Each candidate says how the text was garbled, for example “saved as UTF-8, read as Windows-1252”. If the problem keeps coming back, change that setting in the program that misread it.
- Copy the candidate you want.
- If nothing better turns up, or the text is full of �, open the original file with the charset converter.
How it works
- Encodings tried: UTF-8, CP949 (EUC-KR), Shift_JIS and Windows-1252 (Latin-1), all 12 pairs. Each reversal encodes the garbled text with the encoding it was wrongly read as, then decodes the bytes with the encoding they were really in.
- Candidates that lost nothing in the first step are reversed once more to catch double garbling.
- Ranking favours Latin letters and digits, Japanese kana and the 2,350 common Hangul syllables, and penalises �, control codes, stray Latin-1 symbols and Hangul glued in front of Latin letters. A candidate that creates new gaps of � is ranked down hard.
- When 20% or more of the input is �, the “cannot be recovered from the text” note and a link to the charset converter come first, and only candidates whose remaining letters read clearly (common Hangul or kana) are kept, so random-looking noise is never offered as a fix.
- Encoding tables follow the WHATWG Encoding Standard.
Examples
| Garbled | Cause | Fixed |
|---|---|---|
| Café crème | UTF-8 read as Windows-1252 | Café crème |
| it’s | UTF-8 read as Windows-1252 | it’s |
| 한글 | UTF-8 read as Windows-1252 | 한글 |
| À̸§ | CP949 read as Windows-1252 | 이름 |
| 譌・譛ャ隱槭� | UTF-8 Japanese read as Shift_JIS | 日本語… (partly) |
| �ȳ��ϼ��� | CP949 read as UTF-8 | cannot be recovered (note and converter link shown) |
FAQ
What is mojibake?
Text is stored as bytes. When it is saved in one encoding and read in another, the same bytes show up as the wrong characters. UTF-8 "café" read as Windows-1252 becomes "café", and Korean "한글" becomes "한글".
When can't it be fixed?
If the text is full of � (the replacement character), the original bytes were thrown away when it was misread, so they cannot be rebuilt from the text. A Korean EUC-KR file opened as UTF-8 is the classic case. Open the original file with the charset converter instead.
Which candidate should I trust?
The top one is the most natural-looking reading, but very short text can be ambiguous. Pick the one that reads as real sentences, and check the explanation of how it was garbled.
Does it handle text garbled twice?
Yes. Text that was misread, saved and misread again (é turning into Â…) is reversed in two steps. Texts over 5,000 characters only get one step to keep things fast.
Why do very short Japanese strings sometimes stay unfixed?
Candidates are ranked by how natural they look, and a very short garbled string can itself look plausible because it mixes kanji and kana. "ありがとう" garbled through Shift_JIS ("縺ゅj縺後→縺�") may get no candidate or a different order. Longer text ranks much more reliably, so paste the surrounding sentences too.
Is my text stored?
No. Everything runs in your browser.
Related tools
Charset Converter
A converter that re-encodes text and files between Korean, Japanese and Unicode encodings
Unicode Normalizer
A Unicode normalization tool for text and file names, including decomposed macOS names
URL Encoder/Decoder
A URL percent-encoder and decoder that handles Unicode and emoji
Unicode Converter
Convert text to and from Unicode code points, escapes and UTF-8 bytes
Archive Extractor
Extract archives without installing anything and repair garbled file names