Home·Error messages

python · Python

The message

UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte

What it means

You read a file or byte string as UTF-8 and hit a byte that UTF-8 does not allow; the message names both the position and the byte. Usually the file is cp1252 or another legacy encoding out of Windows, or it is not text at all — an image, a zip, or gzip data you forgot to decompress. Finding the real encoding and passing it as encoding= is the correct fix; errors="replace" lets the read finish but turns those bytes into question marks, so the data is silently damaged from then on.

The fix

There is no one-line command for this. The explanation says what to look at instead.

Printed by
python
Python
18

A Python traceback splits the answer in two — the last line says what went wrong, the frames above it say where — so reading only the last line gives you the name and loses the place, and for the value-is-missing errors such as NoneType and KeyError the cause almost always sits in a frame above the one that crashed.

Reading an error message

  • Read from the first line down. The lower you go the more it is about the tool’s internals; the cause is usually at the top.
  • If there is a file and a line number, start there — not the top stack frame, but the topmost line that names a file you wrote.
  • Search the message verbatim, but strip your own paths and variable names first; those are what stop the search from matching.
  • The same condition is worded differently across tool versions. If results look wrong, add the version number to the query.
  • Before pasting a fix, check what it throws away. Some of these cannot be undone.

Common questions

Q. What does “UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte” mean?

You read a file or byte string as UTF-8 and hit a byte that UTF-8 does not allow; the message names both the position and the byte. Usually the file is cp1252 or another legacy encoding out of Windows, or it is not text at all — an image, a zip, or gzip data you forgot to decompress. Finding the real encoding and passing it as encoding= is the correct fix; errors="replace" lets the read finish but turns those bytes into question marks, so the data is silently damaged from then on.

Q. How do I fix it?

There is no one-line command. The explanation above says what to look at instead.

Q. Which tool prints this?

python. It sits under Python, and the message runs to 13 words.

Errors nearby