🔬 Unicode Character Inspector

Enter text to see each character broken down by Unicode codepoint, hex value and UTF-8 byte sequence. This is a separate tool from Convert to Unicode Escapes, which only produces \u-escaped output — this one shows the full breakdown per character.

One row per code point, not per byte

The text is walked by code point, so an emoji made of a surrogate pair gets a single row here even though JavaScript's own length counts it as two. The last column shows how the character is stored in UTF-8:

CharCodepointDecimalUTF-8 bytes
eU+006510165
éU+00E9233C3 A9
日U+65E526085E6 97 A5
😀U+1F600128512F0 9F 98 80

The pattern is worth remembering: ASCII takes one byte, accented Latin and Cyrillic two, CJK three, emoji four. That is why a 280 character limit and a 280 byte limit are different promises.

What it is actually good for

The field holds one line and every character produces a row, so inspect the fragment you are chasing rather than a whole document. There is no character-name column - the codepoint is the key to look the name up with.

Frequently asked questions

How do I find the Unicode codepoint for a character?

Paste the character into the box and read the Codepoint column. The U+ form is what you want for CSS content, regex ranges and documentation; the decimal is what an HTML numeric entity uses.

Why does one emoji show up as three or four rows?

Because it is a sequence, not a character. A profession or family emoji is several pictographs stitched together with zero-width joiners, and each of those joiners is a code point in its own right.

Why is my accented letter listed as two characters?

It is stored as a base letter plus a combining mark rather than one precomposed code point. Both look identical on screen, but only one matches when you search.