Data compression

GCSE Computer Science revision notes, key terms and practice questions.

Why compress files?

  • Compression makes files smaller, so they take up less storage space and are quicker to send or download, using less bandwidth.

Lossy compression

  • Lossy compression permanently removes some data, usually detail that people are unlikely to notice. It greatly reduces the file size, but the original can't be recovered exactly.
  • It is used for photos (JPEG), music (MP3) and video. It is not suitable for text or program files.

Lossless compression

  • Lossless compression makes a file smaller without losing any data, so the original can be recovered exactly.
  • It is used for text, program files and some images (PNG). It usually reduces the file size less than lossy compression.

Run-length encoding (RLE)

  • RLE stores each run of repeated data as a count and a value: AAAAABBB becomes 5A 3B. In binary, 0000011100 becomes 5 0s, 3 1s and 2 0s.
  • It works well when there are long runs of repeated values, but can make data bigger when there are few repeats.

Huffman coding (AQA)

  • Huffman coding gives shorter binary codes to the characters that appear most often, using a Huffman tree. It is lossless.
  • Bits needed = the total of (frequency × code length) for each character. A text with 4 As, 2 Bs and 1 C, coded A = 0, B = 10 and C = 11, needs 4 × 1 + 2 × 2 + 1 × 2 = 10 bits, compared with 7 × 7 = 49 bits in 7-bit ASCII.

Key terms

Compression
Making a file smaller.
Lossy compression
Compression that permanently removes some data.
Lossless compression
Compression that keeps all the data, so the original can be recovered exactly.
Run-length encoding
Storing runs of repeated data as a count and a value.
Huffman coding
Lossless compression that gives shorter codes to more frequent characters.
Bandwidth
The amount of data that can be sent over a network each second.

Practise Data compression: 12 questions