Skip to content
Tutorial 11 min read

Garbled Scraped Text: Fix Character Encoding Before Parsing

Your scraper returns mojibake instead of product names. Trace the failure from missing charset headers to wrong decode steps and fix it before parsing.

FE
FineData Engineering · Editorial Policy
|
On this page

The Symptom: Your Product Names Look Like a Keyboard Accident

You ship a scraper for a storefront. Everything looks fine in the logs: HTTP 200, response received, HTML parsed, rows inserted. Then someone opens the database and finds this in the product_name column:

Cafetière Inox 1L â¬24,99
Schränke & Möbel – Neu

That’s mojibake. UTF-8 bytes decoded as Latin-1 (or cp1252), which turns every non-ASCII character into two or three wrong ones. The è became è. The euro sign became â¬. The en-dash became â€". Your data pipeline is technically running and completely unusable at the same time.

It gets worse. The damage is often silent at small scale. If the site is mostly English with occasional accented names, 97% of rows look fine and the broken 3% slip through validation until a customer or a downstream analyst notices. By then you’ve got thousands of poisoned rows, and fixing them after the fact means re-scraping everything or writing a repair pass that guesses at what the original bytes were.

The failure has a specific shape, and it happens at a specific point in your code. Let’s trace it.

Where Encoding Actually Breaks

Character encoding fails at exactly one place: the moment bytes become a string. Everything before that moment is bytes, everything after is a decoded assumption. If the assumption is wrong, the string is wrong, and no amount of downstream parsing fixes it.

There are four places this can go wrong, in order of how often I see them:

  1. Your HTTP client guesses the charset. Python’s requests does this. If the response has no Content-Type: text/html; charset=... header, requests falls back to ISO-8859-1 for text/* responses per an old HTTP spec quirk. The server sent UTF-8. Your client decoded it as Latin-1. Mojibake, instantly.

  2. The server lies. The header says charset=utf-8 but the HTML is actually cp1252. This happens on sites assembled from multiple templates or after a migration where one subsystem still emits legacy encoding. The header wins in every standard client, so you get replacement characters (“) on every byte sequence that isn’t valid UTF-8.

  3. You decode twice. Someone calls .encode('latin-1').decode('utf-8') as a “fix” on text that was already correct, or a middleware layer re-encodes a decoded string. Double-encoding produces the mirrored problem: fixing it requires knowing which layer did what.

  4. The bytes never were text. The response is gzip or brotli that didn’t get decompressed, or it’s actually JSON with an escaped-unicode payload, and you’re staring at binary garbage that no charset will fix.

The critical insight: the correct decode depends on information outside the bytes themselves — headers, HTML meta tags, BOMs. A decode function that only looks at the byte buffer is guessing. Guessing at scale is how you end up with a database full of é.

Naive Fix #1: Assume UTF-8 Everywhere

The internet is converging on UTF-8, so just decode everything as UTF-8, right?

text = response.content.decode('utf-8')

This fixes case 1 (the requests Latin-1 fallback) and it’s the right default. But it explodes on case 2 with a UnicodeDecodeError, and — worse — if you add errors='replace' to suppress the exception, you silently convert every legacy-encoded page into a field of “ characters. You’ve traded visible mojibake for invisible data loss. I’ve seen teams run that way for months because the rows “looked clean.” They looked clean because the evidence was destroyed.

Hard-coding UTF-8 also does nothing for the pages that are genuinely cp1252, Shift-JIS, or windows-1256, which still exist in large numbers on older e-commerce and regional sites.

Naive Fix #2: Run chardet on Everything

The next move everyone reaches for: statistical charset detection.

import chardet
guess = chardet.detect(response.content)
text = response.content.decode(guess['encoding'])

Two problems. First, detection is slow — chardet on a 500KB page can take hundreds of milliseconds, and if you’re scraping thousands of pages, that’s real money and latency. Second, and more important: statistical detection is a fallback, not a first resort. The page usually tells you its encoding, in two or three standard places, and those declarations are more reliable than a byte-frequency heuristic. Using chardet first means ignoring free, correct metadata in favor of a probabilistic guess that will occasionally return Windows-1254 with 72% confidence on a Turkish page that was actually UTF-8.

There’s a subtler failure too: chardet can’t distinguish cp1252 from ISO-8859-1 reliably because their byte layouts overlap almost entirely. It’ll pick one. Sometimes it picks wrong for characters in the 0x80–0x9F range, which is exactly where cp1252 has the smart quotes and em-dashes that break product descriptions.

If you want statistical detection as a last-ditch fallback, use charset-normalizer — it’s faster and maintained — but put it last in the chain, never first.

The Correct Decode Order

Here’s the resolution order that matches what browsers actually do, which is the behavior your scraper should mirror because sites are built and tested against browsers:

  1. Byte-order mark. If the payload starts with a UTF-8 BOM (EF BB BF) or UTF-16 BOM, that’s authoritative. Strip it after decoding.
  2. HTTP Content-Type header charset. The server’s declared intent.
  3. HTML <meta charset> or <meta http-equiv="Content-Type"> tag. Found in the first ~1KB of the document.
  4. Statistical detection on a sample.
  5. UTF-8 with replacement as the final default, plus a flag so you know it happened.

Browsers apply one more rule that matters: if the header charset produces invalid byte sequences but the meta tag charset decodes cleanly, the meta wins. The header lied; the document was right. You should implement that override.

Here’s the full function:

import re
from charset_normalizer import from_bytes

BOMS = [
    (b"\xef\xbb\xbf", "utf-8-sig"),
    (b"\xff\xfe", "utf-16-le"),
    (b"\xfe\xff", "utf-16-be"),
]

META_RE = re.compile(
    rb'<meta[^>]+charset\s*=\s*["\']?\s*([a-zA-Z0-9_\-]+)',
    re.IGNORECASE,
)

def decode_html(raw: bytes, header_charset: str | None) -> tuple[str, str]:
    """Return (text, encoding_used). Resolution order mirrors a browser."""
    # 1. BOM is authoritative.
    for bom, enc in BOMS:
        if raw.startswith(bom):
            return raw.decode(enc), enc

    # 2. Try the declared header charset, but validate.
    candidates = []
    if header_charset:
        candidates.append(header_charset)

    # 3. Look for a meta charset in the first 2KB.
    m = META_RE.search(raw[:2048])
    if m:
        candidates.append(m.group(1).decode("ascii", "ignore"))

    for enc in candidates:
        try:
            return raw.decode(enc), enc
        except (UnicodeDecodeError, LookupError):
            continue  # declared encoding lied — try the next candidate

    # 4. Statistical fallback on a sample.
    best = from_bytes(raw[:100_000]).best()
    if best and best.encoding:
        return str(best), f"{best.encoding} (detected)"

    # 5. Last resort. Flag it, don't hide it.
    return raw.decode("utf-8", errors="replace"), "utf-8 (fallback, replacement chars)"

Note the LookupError catch — servers occasionally send a header like charset=utf8 (no dash) or a made-up name, and .decode() raises LookupError for unknown codecs, not UnicodeDecodeError. Missing that exception is how decode functions crash on pages that would otherwise parse fine.

And note what the function returns: the encoding used, as data. If your pipeline logs encoding_used per page, you can query “which URLs fell back to detection or replacement” and fix the source of the problem instead of discovering it in a customer report.

Wiring It Into the Scrape Response

If you’re using a scraping API rather than raw requests, the same rules apply — you just need to know where the bytes and the header information land in the response. Here’s a synchronous scrape against a storefront page, requesting raw HTML so we control the decode ourselves:

import requests

resp = requests.post(
    "https://api.finedata.ai/api/v1/scrape",
    headers={"Authorization": "Bearer fd_your_api_key"},
    json={
        "url": "https://store.example.com/catalog/cafetieres",
        "formats": ["rawHtml"],
        "only_main_content": False,
        "use_antibot": True,
    },
    timeout=180,
)
data = resp.json()
raw_html = data["data"]["rawHtml"].encode("utf-8", errors="surrogateescape")
text, used = decode_html(raw_html, "utf-8")

One honest caveat about this shape: when a scraping service returns HTML inside a JSON envelope, the bytes have already been decoded once by the service. In the common case the service decoded correctly (it applies the header/meta resolution described above), and re-encoding to bytes so your own decode logic can run is redundant. But if you’re debugging an encoding problem, you want your decode function in the path so you can see what it decides and log it. Redundant decode that you control beats a black-box decode you can’t inspect. Once things are stable, drop the round-trip and trust the envelope.

The other thing worth doing: request markdown alongside rawHtml when your goal is text for downstream processing. Markdown output from the service has already gone through a sane text extraction, and if the mojibake is present in the markdown but absent in your local decode of the raw HTML, you’ve localized the bug to the service’s pipeline rather than yours. That two-format diff has saved me hours of pointing fingers in the wrong direction. For more on when to lean on pre-extracted formats, see Stop Cleaning HTML by Hand: Scrape Pages as Markdown.

The Repair Pass: Fixing Already-Poisoned Rows

Say you already have mojibake in the database. The good news: UTF-8-decoded-as-Latin-1 is a reversible transformation, because every byte of UTF-8 maps to a valid Latin-1 code point. This round-trips:

def fix_double_encoded(s: str) -> tuple[str, bool]:
    """Repair text that was UTF-8 bytes decoded as Latin-1/cp1252."""
    try:
        repaired = s.encode("latin-1").decode("utf-8")
    except (UnicodeEncodeError, UnicodeDecodeError):
        return s, False
    # Heuristic guard: only accept if the repair reduced the
    # count of suspicious sequences.
    suspicious = ("Ã", "Â", "â€", "€")
    before = sum(s.count(c) for c in suspicious)
    after = sum(repaired.count(c) for c in suspicious)
    return (repaired, True) if after < before else (s, False)

The guard matters. Run unconditionally, this function will “repair” legitimately French text like Café à Paris into nonsense, or mangle strings that were already correct. The heuristic isn’t perfect — a string with zero suspicious sequences before and after returns (s, False), which is the safe outcome.

The cp1252 variant is nastier because cp1252 has a few unmapped bytes (0x81, 0x8D, and friends), so encode('cp1252') can throw on strings that came from a Latin-1 decode. If your source sites are Windows-hosted, try cp1252 first and fall back to latin-1. And test the repair on a sample before running it against production — a bad repair pass is worse than the original damage because now you’ve corrupted the corruption.

My actual recommendation: don’t write a repair pass at all if you can re-scrape. Re-scraping with a correct decode is deterministic; repair is guesswork about which of four failure modes produced each broken string. The repair pass is only worth building when re-scraping is expensive or the pages have changed since the bad data was collected.

Normalization: The Bug That Isn’t Encoding

Once decoding is correct, you’ll hit a related problem that people constantly misdiagnose as encoding: Unicode normalization. The string Café can exist as four code points (C, a, f, e, combining-acute) or three (C, a, f, é). Both render identically. Both are “correct.” But they don’t compare equal, don’t match in SQL WHERE clauses, and break deduplication.

import unicodedata

name = unicodedata.normalize("NFC", scraped_name)  # compose to canonical form

Normalize to NFC on ingest, every time, before the string touches your database. Pick one form and enforce it at the pipeline boundary. This is also where you catch the other silent killer: invisible characters like zero-width spaces (U+200B) and non-breaking spaces (U+00A0) that sites inject into prices and names, which make "24,99" and "24,99" compare as different strings. Strip or map them explicitly:

CLEAN = str.maketrans({"\u00a0": " ", "\u200b": "", "\ufeff": ""})
name = name.translate(CLEAN).strip()

If your extracted fields are landing in the wrong columns on top of arriving garbled, that’s a separate extraction problem — Fix Scraped Data That Lands in the Wrong Fields covers the selector-side debugging.

Gotchas Worth Knowing Before You Ship

Gzip confusion looks like encoding failure but isn’t. If your decoded “text” starts with \x1f\x8b, you’re looking at uncompressed gzip payload. That’s a Content-Encoding handling bug in your client or middleware, and no charset will fix it. Check for the magic bytes before you blame encodings.

JSON endpoints dodge most of this — usually. JSON is UTF-8 by specification, so hidden API endpoints are typically a free pass on charset issues. But watch for \uXXXX escape sequences that decode into compatibility forms (the fi ligature, ㎡), which arrive as NFKC-only equivalents and won’t match their NFC counterparts without normalize("NFKC", ...). The trade-offs between scraping those endpoints versus the DOM are covered in Hidden JSON Endpoints vs DOM Parsing: What to Scrape.

Don’t trust Content-Length-based truncation checks on decoded text. If you validate payload completeness by comparing string length against the header, the units don’t match — the header counts bytes, your string counts code points. Compare against the raw bytes.

Log the encoding used, per page. This is the single highest-leverage thing in the whole post. A encoding_used column turns “some names look weird” into SELECT url FROM pages WHERE encoding_used LIKE '%fallback%' — a precise, answerable question.

Wrap-Up

Mojibake is a decode-time bug with a search-time symptom. The fix is boring and effective: resolve the charset in browser order (BOM, header, meta tag, detection, flagged fallback), validate each candidate by actually attempting the decode, normalize to NFC at the pipeline boundary, and log which encoding every page used. Statistical detection is your last resort, not your first. And if the damage is already in your database, prefer re-scraping over a repair pass — the repair only round-trips cleanly for the UTF-8-as-Latin-1 case, and guessing which case you’re in per row is a losing game.

The pages that break this are rarely the ones you tested. They’re the legacy category pages, the regional storefronts, the template that one contractor built a decade ago. Your decode function has to handle them anyway — and now it does, with evidence in the logs when it can’t.

#character-encoding #utf-8 #mojibake #data-quality #web-scraping #slot:problems-first

Related Articles