What is actually being counted?

This experiment uses a deliberately simple binary format. Store the first bit. Then store each run’s length minus one in six bits, which represents lengths from 1 to 64. Successive runs alternate between zero and one. The decoder knows the block contains exactly 64 bits.

For four runs, the payload is 1 + 6 × 4 = 25 bits. For an alternating sequence, every bit is a separate run: 1 + 6 × 64 = 385 bits. Both decode exactly to their original input.

The accounting boundary

These are bit counts for an agreed format, not actual browser storage or a downloaded file size. The decoder, the known block size, format identification, and file-container overhead are not counted. A general compressor would also need a way to identify the format or fall back to raw data.

Simple to us, difficult for this code

The alternating sequence is highly regular. A different representation could say “repeat 01.” This run-length encoder cannot exploit that kind of structure. Its failure is a limitation of the chosen code, not proof that the sequence is incompressible.

The experiment makes one modest point: compression is always relative to a representation and its costs. It does not measure semantic understanding, generalization, or intelligence.

Continue with “Compression is a lens on intelligence” →

Search the notebook

Search essays and pages. Everything stays in your browser.