🗜️ Run-Length Encoder / Decoder

Run-length encoding (RLE) compresses repeated characters into a count + character form (e.g. aaabbc → 3a2b1c). Simple, and reversible.

Count first, character second - always

The encoder walks the text and closes a run whenever the character changes, writing the length in front of the character. Single characters are not exempt: they are written as a run of one.

aaabbbccccd  ->  3a3b4c1d     11 characters become 8
hello        ->  1h1e2l1o      5 characters become 8

That second line is the whole story of RLE. It only shrinks data that actually repeats; on ordinary prose, where runs of two are rare, it roughly doubles the size. Encoding is safe on any input, and the Input Length and Output Length rows tell you immediately whether you gained anything.

Decoding, and where a round trip breaks

This is a teaching format, not a compressor. To make a file smaller, use zip or gzip - they use LZ77 with Huffman coding, which finds repeated phrases rather than repeated single characters.

Frequently asked questions

What is run-length encoding used for?

Runs of identical bytes: fax transmission, BMP, PCX and TGA bitmaps, and the long white stretches of a scanned page. It is also a favourite first exercise in programming courses because the whole algorithm fits in ten lines.

Why is my encoded text longer than the original?

Because each character costs at least two characters once a count is added. Unless your text has long repeated stretches, RLE expands it - that is expected behaviour, not a fault.

How do I get my original text back?

Switch to Decode and paste the encoded string. As long as the original held no digits or line breaks, you get back exactly what you started with.