python · Python
The message
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
What it means
You read a file or byte string as UTF-8 and hit a byte that UTF-8 does not allow; the message names both the position and the byte. Usually the file is cp1252 or another legacy encoding out of Windows, or it is not text at all — an image, a zip, or gzip data you forgot to decompress. Finding the real encoding and passing it as encoding= is the correct fix; errors="replace" lets the read finish but turns those bytes into question marks, so the data is silently damaged from then on.
The fix
There is no one-line command for this. The explanation says what to look at instead.
- Printed by
- python
- Python
- 18
A Python traceback splits the answer in two — the last line says what went wrong, the frames above it say where — so reading only the last line gives you the name and loses the place, and for the value-is-missing errors such as NoneType and KeyError the cause almost always sits in a frame above the one that crashed.
Reading an error message
- Read from the first line down. The lower you go the more it is about the tool’s internals; the cause is usually at the top.
- If there is a file and a line number, start there — not the top stack frame, but the topmost line that names a file you wrote.
- Search the message verbatim, but strip your own paths and variable names first; those are what stop the search from matching.
- The same condition is worded differently across tool versions. If results look wrong, add the version number to the query.
- Before pasting a fix, check what it throws away. Some of these cannot be undone.
Common questions
Q. What does “UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte” mean?
You read a file or byte string as UTF-8 and hit a byte that UTF-8 does not allow; the message names both the position and the byte. Usually the file is cp1252 or another legacy encoding out of Windows, or it is not text at all — an image, a zip, or gzip data you forgot to decompress. Finding the real encoding and passing it as encoding= is the correct fix; errors="replace" lets the read finish but turns those bytes into question marks, so the data is silently damaged from then on.
Q. How do I fix it?
There is no one-line command. The explanation above says what to look at instead.
Q. Which tool prints this?
python. It sits under Python, and the message runs to 13 words.