Unicode Decoder
Convert between Unicode escapes and text. Supports \uXXXX, \UXXXXXXXX, &#XXXX; entities, and code point lookup. All processing done in your browser.
Result will appear here...
Character Information
| Character | Code Point | UTF-8 Bytes | HTML Entity | CSS Escape | JS Escape |
|---|
How to Use Unicode Decoder
Choose Mode
Select Decode (escape to text) or Encode (text to escape) mode.
Enter Input
Paste Unicode escapes or plain text depending on the mode.
Convert
Click Convert and review the result and character details.
What is Unicode Encoding?
Unicode encoding allows characters from any language or symbol system to be represented in different formats. Common encoding formats include \uXXXX (JavaScript/JSON), &#NNNNN; (HTML entities), and U+XXXX (Unicode code points). These encodings ensure text can be safely transmitted and displayed across different systems.
Encoding Formats
JavaScript (\uXXXX)
Used in JSON and JavaScript strings
HTML Entities (&#XXXX;)
Used in HTML to represent special characters
CSS Escapes (\XXXXXX)
Used in CSS content and identifiers
Fix Garbled Text (Mojibake Repair)
Garbled text (also called mojibake) appears when text encoded in one character set is decoded with another. The most common case: UTF-8 bytes are read as Windows-1252 or Latin-1, turning "café" into "café" and "张三" into "å¼ ä¸‰". Switch to the Fix Garbled Text mode above, paste the broken text, and the decoder will reconstruct the original characters by treating each character as a raw byte and re-decoding the result as UTF-8.
Common Garbled Text Examples
| Garbled | Repaired | Cause |
|---|---|---|
| café | café | UTF-8 → Latin-1 |
| å¼ ä¸‰ | 张三 | UTF-8 → Latin-1 |
| “hi†| “hi” | UTF-8 → double-encoded |
| ü ö ñ | ü ö ñ | UTF-8 → Windows-1252 |
| �� | � � | bytes already lost (U+FFFD) |
To prevent mojibake in your own apps, always declare the encoding explicitly: serve HTML with <meta charset="UTF-8">, set HTTP headers like Content-Type: text/html; charset=utf-8, and use UTF-8 consistently in databases and connections.
UTF-8 vs UTF-16 vs UTF-32 Comparison
| UTF-8 | UTF-16 | UTF-32 | |
|---|---|---|---|
| Encoding unit | 8 bit | 16 bit | 32 bit |
| A (U+0041) | 41 | 00 41 | 00 00 00 41 |
| é (U+00E9) | C3 A9 | 00 E9 | 00 00 00 E9 |
| 中 (U+4E2D) | E4 B8 AD | 4E 2D | 00 00 4E 2D |
| 🎉 (U+1F389) | F0 9F 8E 89 | D8 3C DF 89 | 00 01 F3 89 |
| Best suited for | Web, JSON, APIs — ASCII-compatible and compact | In-memory text in Java, JavaScript, C# | Direct code point access, rare in practice |
Common Unicode Code Points
Quick reference for frequently looked-up characters: punctuation, currency symbols, accented letters, and math signs, with their U+XXXX code points.
Punctuation & Symbols
| “U+201C | ”U+201D | ‘U+2018 | ’U+2019 |
| –U+2013 | —U+2014 | …U+2026 | U+A0 |
| ©U+A9 | ®U+AE | ™U+2122 | °U+B0 |
Currency Signs
| €U+20AC | £U+A3 | ¥U+A5 | ₹U+20B9 |
| ¢U+A2 | ₽U+20BD | ₩U+20A9 | ฿U+E3F |
Accented Letters
| éU+E9 | èU+E8 | êU+EA | üU+FC |
| öU+F6 | ñU+F1 | åU+E5 | øU+F8 |
| çU+E7 | ßU+DF | ąU+105 | žU+17E |
Math & Arrows
| ±U+B1 | ×U+D7 | ÷U+F7 | ≈U+2248 |
| ≠U+2260 | ≤U+2264 | ≥U+2265 | →U+2192 |
| ∞U+221E | ∑U+2211 | √U+221A | µU+B5 |
Frequently Asked Questions
This tool supports \uXXXX (BMP), \UXXXXXXXX (full Unicode), &#DDDD; (decimal HTML entity), &#xHHHH; (hex HTML entity), and U+XXXX (code point notation) for decoding.
UTF-8 encodes each Unicode code point using 1 to 4 bytes. ASCII characters (U+0000 to U+007F) use 1 byte, Latin and other common characters use 2 bytes, most CJK characters use 3 bytes, and emoji and rare characters use 4 bytes.
\uXXXX uses 4 hex digits and can represent characters in the Basic Multilingual Plane (U+0000 to U+FFFF). \UXXXXXXXX uses 8 hex digits and can represent any Unicode character including those beyond the BMP.
These are signs of mojibake: UTF-8 bytes were decoded as Windows-1252/Latin-1. "é" means the two UTF-8 bytes of "é" (0xC3 0xA9) were shown as two separate characters. Use the Fix Garbled Text mode to repair it. "�" is U+FFFD, the replacement character produced when bytes were already invalid — that information is lost and cannot be recovered.
UTF-8 encodes characters in 1–4 bytes and is fully ASCII-compatible, which makes it the default for the web, JSON, and most APIs. UTF-16 encodes characters in 2 or 4 bytes (using surrogate pairs for characters beyond U+FFFF) and is used internally by JavaScript, Java, and Windows. For most text on the web, UTF-8 is smaller and safer.
A code point is a number assigned to a character in the Unicode standard (U+00E9 for e-acute); a code unit is the storage unit an encoding uses. UTF-8 code units are 8-bit bytes - U+00E9 takes two - while UTF-16 code units are 16-bit, and characters beyond U+FFFF need two of them (a surrogate pair). JavaScript strings are UTF-16, so 'hello'.length is 5 but an emoji's .length is 2.
The same visible character can have multiple byte sequences: e-acute as one precomposed character (NFC) or e plus a combining accent (NFD). NFC composes, NFD decomposes, and NFKC/NFKD additionally fold compatibility characters (full-width forms, ligatures). Comparing or hashing user input without normalizing first is a classic source of duplicate accounts and failed lookups - fi ligature vs f+i.
The byte order mark is U+FEFF encoded at the start of a stream (EF BB BF in UTF-8) to signal the encoding. Editors usually hide it, but downstream it becomes an invisible first character: a CSV header gains a phantom key, JSON parsing fails on the stray token, and PHP sessions emit it before headers. UTF-8 needs no BOM since it has no byte-order ambiguity - save without it.
Related Tools
Hex Decoder
Convert between hexadecimal and text — supports configurable delimiter and prefix options
HTML Entity Decoder
Encode and decode HTML entities — convert between special characters and entity references
Base64 URL Decoder Parser
Decode and encode Base64, parse URL-encoded strings, preview Base64 images, and decode JWT header or payload segments
Authoritative References
Primary sources behind this tool - official standards and specifications, not secondhand summaries.