Unicode Decoder

Convert between Unicode escapes and text, look up code points, and fix garbled text (mojibake). Supports \uXXXX, \UXXXXXXXX, and &#XXXX; entities. All processing done in your browser.

Encode & Decode Character Info Client-Side Processing
Result will appear here...

How to Use Unicode Decoder

1

Choose Mode

Select Decode (escape to text) or Encode (text to escape) mode.

2

Enter Input

Paste Unicode escapes or plain text depending on the mode.

3

Convert

Click Convert and review the result and character details.

What is Unicode Encoding?

Unicode encoding allows characters from any language or symbol system to be represented in different formats. Common encoding formats include \uXXXX (JavaScript/JSON), &#NNNNN; (HTML entities), and U+XXXX (Unicode code points). These encodings ensure text can be safely transmitted and displayed across different systems.

Encoding Formats

JavaScript (\uXXXX)

Used in JSON and JavaScript strings

HTML Entities (&#XXXX;)

Used in HTML to represent special characters

CSS Escapes (\XXXXXX)

Used in CSS content and identifiers

Fix Garbled Text (Mojibake Repair)

Garbled text (also called mojibake) appears when text encoded in one character set is decoded with another. The most common case: UTF-8 bytes are read as Windows-1252 or Latin-1, turning "café" into "café" and "张三" into "å¼ ä¸‰". Switch to the Fix Garbled Text mode above, paste the broken text, and the decoder will reconstruct the original characters by treating each character as a raw byte and re-decoding the result as UTF-8.

Common Garbled Text Examples

Garbled Repaired Cause
café café UTF-8 → Latin-1
å¼ ä¸‰ 张三 UTF-8 → Latin-1
“hi†“hi” UTF-8 → double-encoded
ü ö ñ ü ö ñ UTF-8 → Windows-1252
�� � � bytes already lost (U+FFFD)

To prevent mojibake in your own apps, always declare the encoding explicitly: serve HTML with <meta charset="UTF-8">, set HTTP headers like Content-Type: text/html; charset=utf-8, and use UTF-8 consistently in databases and connections.

UTF-8 vs UTF-16 vs UTF-32 Comparison

UTF-8 UTF-16 UTF-32
Encoding unit 8 bit 16 bit 32 bit
A (U+0041) 41 00 41 00 00 00 41
é (U+00E9) C3 A9 00 E9 00 00 00 E9
中 (U+4E2D) E4 B8 AD 4E 2D 00 00 4E 2D
🎉 (U+1F389) F0 9F 8E 89 D8 3C DF 89 00 01 F3 89
Best suited for Web, JSON, APIs — ASCII-compatible and compact In-memory text in Java, JavaScript, C# Direct code point access, rare in practice

Common Unicode Code Points

Quick reference for frequently looked-up characters: punctuation, currency symbols, accented letters, and math signs, with their U+XXXX code points.

Punctuation & Symbols

“U+201C ”U+201D ‘U+2018 ’U+2019
–U+2013 —U+2014 …U+2026  U+A0
©U+A9 ®U+AE ™U+2122 °U+B0

Currency Signs

€U+20AC £U+A3 ¥U+A5 ₹U+20B9
¢U+A2 ₽U+20BD ₩U+20A9 ฿U+E3F

Accented Letters

éU+E9 èU+E8 êU+EA üU+FC
öU+F6 ñU+F1 åU+E5 øU+F8
çU+E7 ßU+DF ąU+105 žU+17E

Math & Arrows

±U+B1 ×U+D7 ÷U+F7 ≈U+2248
≠U+2260 ≤U+2264 ≥U+2265 →U+2192
∞U+221E ∑U+2211 √U+221A µU+B5

Frequently Asked Questions

What Unicode escape formats are supported?

This tool supports \uXXXX (BMP), \UXXXXXXXX (full Unicode), &#DDDD; (decimal HTML entity), &#xHHHH; (hex HTML entity), and U+XXXX (code point notation) for decoding.

How does UTF-8 encoding work?

UTF-8 encodes each Unicode code point using 1 to 4 bytes. ASCII characters (U+0000 to U+007F) use 1 byte, Latin and other common characters use 2 bytes, most CJK characters use 3 bytes, and emoji and rare characters use 4 bytes.

What is the difference between \u and \U?

\uXXXX uses 4 hex digits and can represent characters in the Basic Multilingual Plane (U+0000 to U+FFFF). \UXXXXXXXX uses 8 hex digits and can represent any Unicode character including those beyond the BMP.

Why does my text show characters like "é" or "�", and how do I fix it?

These are signs of mojibake: UTF-8 bytes were decoded as Windows-1252/Latin-1. "é" means the two UTF-8 bytes of "é" (0xC3 0xA9) were shown as two separate characters. Use the Fix Garbled Text mode to repair it. "�" is U+FFFD, the replacement character produced when bytes were already invalid — that information is lost and cannot be recovered.

What is the difference between UTF-8 and UTF-16?

UTF-8 encodes characters in 1–4 bytes and is fully ASCII-compatible, which makes it the default for the web, JSON, and most APIs. UTF-16 encodes characters in 2 or 4 bytes (using surrogate pairs for characters beyond U+FFFF) and is used internally by JavaScript, Java, and Windows. For most text on the web, UTF-8 is smaller and safer.

What is the difference between a code point and a code unit?

A code point is a number assigned to a character in the Unicode standard (U+00E9 for e-acute); a code unit is the storage unit an encoding uses. UTF-8 code units are 8-bit bytes - U+00E9 takes two - while UTF-16 code units are 16-bit, and characters beyond U+FFFF need two of them (a surrogate pair). JavaScript strings are UTF-16, so 'hello'.length is 5 but an emoji's .length is 2.

What are Unicode normalization forms and when do they matter?

The same visible character can have multiple byte sequences: e-acute as one precomposed character (NFC) or e plus a combining accent (NFD). NFC composes, NFD decomposes, and NFKC/NFKD additionally fold compatibility characters (full-width forms, ligatures). Comparing or hashing user input without normalizing first is a classic source of duplicate accounts and failed lookups - fi ligature vs f+i.

What is a BOM and why does it break my files?

The byte order mark is U+FEFF encoded at the start of a stream (EF BB BF in UTF-8) to signal the encoding. Editors usually hide it, but downstream it becomes an invisible first character: a CSV header gains a phantom key, JSON parsing fails on the stray token, and PHP sessions emit it before headers. UTF-8 needs no BOM since it has no byte-order ambiguity - save without it.

{-- * External Resources Component(#18 Phase 3b 内容佐证工程) * 工具页「权威引用」区块:RFC / W3C / WHATWG / ECMA / IANA / 官方规范站 / Wikipedia。 * * - 接受 :slug 属性 → 经 config/tool-sources.php 家族矩阵渲染该工具的权威引用 * - slug 未命中映射时不渲染(无权威来源的工具静默跳过) * - 链接 title 保持英文(引用源专名);description 经 * common.resources.descriptions.{key} 本地化,lang 未命中回退英文(线上不裸奔) * - 链接保持 dofollow(rel="noopener noreferrer") * * @param string|null $slug --}}