Why compress files?
- Compression makes files smaller, so they take up less storage space and are quicker to send or download, using less bandwidth.
Lossy compression
- Lossy compression permanently removes some data, usually detail that people are unlikely to notice. It greatly reduces the file size, but the original can't be recovered exactly.
- It is used for photos (JPEG), music (MP3) and video. It is not suitable for text or program files.
Lossless compression
- Lossless compression makes a file smaller without losing any data, so the original can be recovered exactly.
- It is used for text, program files and some images (PNG). It usually reduces the file size less than lossy compression.
Run-length encoding (RLE)
- RLE stores each run of repeated data as a count and a value: AAAAABBB becomes 5A 3B. In binary, 0000011100 becomes 5 0s, 3 1s and 2 0s.
- It works well when there are long runs of repeated values, but can make data bigger when there are few repeats.
Huffman coding (AQA)
- Huffman coding gives shorter binary codes to the characters that appear most often, using a Huffman tree. It is lossless.
- Bits needed = the total of (frequency × code length) for each character. A text with 4 As, 2 Bs and 1 C, coded A = 0, B = 10 and C = 11, needs 4 × 1 + 2 × 2 + 1 × 2 = 10 bits, compared with 7 × 7 = 49 bits in 7-bit ASCII.
Key terms
- Compression
- Making a file smaller.
- Lossy compression
- Compression that permanently removes some data.
- Lossless compression
- Compression that keeps all the data, so the original can be recovered exactly.
- Run-length encoding
- Storing runs of repeated data as a count and a value.
- Huffman coding
- Lossless compression that gives shorter codes to more frequent characters.
- Bandwidth
- The amount of data that can be sent over a network each second.