🔬 Unicode Character Inspector
Enter text to see each character broken down by Unicode codepoint, hex value and UTF-8 byte
sequence. This is a separate tool from
Convert to Unicode Escapes, which only
produces \u-escaped output — this one shows the full breakdown per character.
One row per code point, not per byte
The text is walked by code point, so an emoji made of a surrogate pair gets a single row here even though JavaScript's own length counts it as two. The last column shows how the character is stored in UTF-8:
| Char | Codepoint | Decimal | UTF-8 bytes |
|---|---|---|---|
| e | U+0065 | 101 | 65 |
| é | U+00E9 | 233 | C3 A9 |
| 日 | U+65E5 | 26085 | E6 97 A5 |
| 😀 | U+1F600 | 128512 | F0 9F 98 80 |
The pattern is worth remembering: ASCII takes one byte, accented Latin and Cyrillic two, CJK three, emoji four. That is why a 280 character limit and a 280 byte limit are different promises.
What it is actually good for
- Identifying a space that will not behave. U+0020 is an ordinary space, U+00A0 a no-break space out of Word, U+200B a zero-width space that no trim will remove.
- Spotting decomposed accents. If é comes back as two rows, U+0065 followed by U+0301, your text is in the NFD form macOS often produces - which breaks exact-match search and sorting.
- Taking an emoji apart. Family emoji appear as several code points joined by U+200D, skin tones as a modifier in the U+1F3FB range.
- Telling dashes apart before one breaks a CSV import: hyphen U+002D, en dash U+2013, em dash U+2014, minus sign U+2212.
Frequently asked questions
How do I find the Unicode codepoint for a character?
Paste the character into the box and read the Codepoint column. The U+ form is what you want for CSS content, regex ranges and documentation; the decimal is what an HTML numeric entity uses.
Why does one emoji show up as three or four rows?
Because it is a sequence, not a character. A profession or family emoji is several pictographs stitched together with zero-width joiners, and each of those joiners is a code point in its own right.
Why is my accented letter listed as two characters?
It is stored as a base letter plus a combining mark rather than one precomposed code point. Both look identical on screen, but only one matches when you search.