Docento.app
Close-up of a circuit board
All Posts

Fix Garbled Characters in CSV Files: Encoding Problems Explained

By The Docento.app TeamPublished 4 min read
Try Docento's free PDF editor — No sign-up, 100% private — sign, annotate, and stamp PDFs in your browser.Open the editor

You open a CSV and "José" has become "José", or a name with an accent shows a diamond with a question mark. The file is not corrupt. It is being read with the wrong character encoding. CSV has no way to say which encoding it uses, so the reader has to guess, and guesses go wrong.

What an encoding is

A text file is a sequence of bytes. An encoding is the table that maps bytes to characters. For plain English letters, nearly all encodings agree, which is why simple files always look fine. For accented letters, non-Latin scripts and symbols, they diverge.

Common encodings you will meet:

  • UTF-8: the dominant encoding on the web. It can represent every Unicode character, using one byte for basic Latin letters and more for others.
  • Windows-1252 (often called "ANSI" on Windows): a single-byte encoding for Western European languages.
  • ISO-8859-1 (Latin-1): similar, covering Western European languages.
  • UTF-16: uses two or more bytes per character, sometimes with a byte order mark.

The classic symptom

When UTF-8 text is read as Windows-1252, a character like é, which is two bytes in UTF-8, shows up as two characters: é. This pattern of an à or  followed by a symbol is a reliable sign of UTF-8 read as a single-byte encoding. The reverse mistake, a single-byte file read as UTF-8, tends to produce the replacement character, a diamond with a question mark, because the bytes are not valid UTF-8 sequences.

Why Excel is the usual culprit

Spreadsheet applications on Windows have long assumed that a CSV opened by double-clicking uses the system's legacy code page rather than UTF-8. A UTF-8 file without a byte order mark may therefore display incorrectly, even though the same file looks right in a text editor. Newer versions improve this, but behaviour varies by version and platform.

The byte order mark (BOM)

A BOM is a few bytes at the very start of a file, EF BB BF for UTF-8, which signal the encoding. Some programs use it to recognise UTF-8. For CSV opened in Excel, a UTF-8 file with a BOM is often displayed correctly where one without is not. But a BOM can be harmful elsewhere: some parsers treat it as part of the first column name, so a header id becomes id and lookups fail. See text file encodings explained.

How to fix a file that looks wrong

Option 1: Import with the correct encoding. Don't double-click. In Excel use the Data tab's text import (From Text/CSV), which has a File Origin or encoding selector; in LibreOffice the Text Import dialog has a Character set field; in Google Sheets, File then Import handles UTF-8 well. Choose UTF-8 and check the preview.

Option 2: Check in a neutral viewer. Open the file in a text editor or in Docento's Text & Markdown Editor, which reads text files as UTF-8. If the text looks right there, the file is UTF-8 and the problem is the program you used before.

Option 3: Re-save with the right encoding. In a text editor with an encoding option, re-save as UTF-8 (with BOM if the destination is old Excel on Windows, without BOM for most other systems).

Option 4: Repair mangled text. If the garbled version was saved over the original, the damage may be reversible only if the conversion was lossless. A file read as Windows-1252 and re-saved as UTF-8 can often be repaired by reversing the step, but a file that went through replacement characters has lost the original information. Always keep the original.

Avoiding the problem

  • Agree on UTF-8 with whoever produces or consumes the file, and say so explicitly.
  • Use UTF-8 with BOM only when the destination needs it.
  • Test with a sample containing accented characters, an emoji and non-Latin text before relying on a pipeline.
  • Prefer import dialogs to double-clicking.

Takeaway

Garbled characters in a CSV mean the reader guessed the wrong encoding. Look at the file in a neutral editor, import it with UTF-8 selected, and agree on one encoding with anyone who shares the file.

Try Docento's free PDF editor

No sign-up, 100% private — sign, annotate, and stamp PDFs in your browser.

Open the editor

Related Posts