Skip to content
Technical 15 min read

Measure Your Scraper's Real Success Rate Locally

Build a small Python harness that logs every scrape attempt and computes success rate, latency percentiles, and block rates you can actually trust.

FE
FineData Engineering · Editorial Policy
|
On this page

Why Your Scraper’s Reported Success Rate Is Probably a Lie

Most scrapers report a success rate that is quietly wrong. The pattern is nearly universal: a counter increments when the HTTP client returns without raising, the log says “1423/1500 succeeded, 94.6% success rate,” and everyone moves on. Then someone opens the output table and finds 400 rows where the price column is empty and 60 rows containing what is unmistakably a CAPTCHA interstitial saved as “product HTML.”

The problem is definitional. HTTP 200 means the server sent you bytes. It does not mean those bytes are the page you asked for. Anti-bot systems deliberately return 200 with a challenge page because it wastes your resources and confuses naive monitoring. Retry logic makes it worse: if your client silently retries a blocked request three times and succeeds on the fourth, your log shows one success and zero blocks, when the truth is one block, three wasted requests, and elevated latency.

Here is what a broken-but-200 response actually looks like in the wild:

import requests

resp = requests.get(
    "https://store.example.com/product/123",
    headers={"User-Agent": "Mozilla/5.0"},
    timeout=30,
)
print(resp.status_code)   # 200
print(len(resp.text))     # 2148 bytes -- a real product page is ~80KB
print(resp.text[:120])    # '<html><head><title>Checking your browser...</title>'

That response will pass a resp.ok check. It will pass if resp.status_code == 200. It will not pass a human looking at the data. The gap between “the request completed” and “we got the data” is where your reported success rate becomes fiction.

What you countReported resultWhat actually happenedTrusted success rate
HTTP 200 = success94.6% (1423/1500)214 responses were challenge pages, 186 were empty-body shells68.2%
No exception = success97.1% (1457/1500)43 timeouts were swallowed by a bare except: pass68.2%
Retries hidden from log99.0% (1485/1500)912 extra requests burned on retries, 15 permanent blocks erased68.2%
Content verified (expected data present)68.2%Matches manual audit of a 100-row sample68.2%

The fix is not more logging infrastructure. It is a small local harness that records one structured record per attempt, verifies content, and computes rates from the log rather than from in-memory counters. That is what the rest of this tutorial builds. If you want the broader argument for measuring per-page cost honestly, see What One Successful Page Actually Costs: A Unit Model.

Designing the Local Test Harness: One Attempt, One Log Record

The harness has one design rule: every attempt produces exactly one append-only record. No aggregation in memory, no counters, no “we’ll compute it later from the database.” JSONL on local disk. If the process crashes, everything measured so far survives, and the log is the single source of truth for every metric you compute afterward.

The data model:

from dataclasses import dataclass, field, asdict
from datetime import datetime, timezone
from typing import Optional

@dataclass
class AttemptRecord:
    url: str
    domain: str
    timestamp: str                      # ISO 8601, UTC
    http_status: Optional[int]          # None if the request never got a response
    latency_ms: float                   # wall clock, includes retries
    retries: int = 0
    content_ok: bool = False            # expected data actually present
    block_detected: bool = False        # challenge/interstitial page
    error: Optional[str] = None         # exception class name, if any
    response_bytes: int = 0
    run_label: str = "adhoc"

    def to_jsonl(self) -> str:
        import json
        return json.dumps(asdict(self))

Every field exists to answer a specific question later. content_ok answers “did we get the data.” block_detected answers “was this the target’s fault or ours.” retries and cumulative latency_ms answer “what did this attempt really cost.” If you are tempted to add fields, ask which report will consume them; fields without a consumer rot.

The main loop:

import json
import time
import requests
from urllib.parse import urlparse

LOG_PATH = "scrape_attempts.jsonl"

def log_record(rec: AttemptRecord) -> None:
    with open(LOG_PATH, "a", encoding="utf-8") as f:
        f.write(rec.to_jsonl() + "\n")

def run_harness(urls: list[str], run_label: str = "adhoc") -> None:
    session = requests.Session()
    session.headers.update({"User-Agent": "Mozilla/5.0 (X11; Linux x86_64)"})

    for url in urls:
        domain = urlparse(url).netloc
        started = time.perf_counter()
        rec = AttemptRecord(
            url=url,
            domain=domain,
            timestamp=datetime.now(timezone.utc).isoformat(),
            http_status=None,
            latency_ms=0.0,
            run_label=run_label,
        )
        try:
            resp = session.get(url, timeout=30)
            rec.http_status = resp.status_code
            rec.response_bytes = len(resp.content)
            rec.content_ok, rec.block_detected = classify_outcome(resp)
        except requests.RequestException as exc:
            rec.error = type(exc).__name__
        finally:
            rec.latency_ms = (time.perf_counter() - started) * 1000
            log_record(rec)

Note what is deliberately absent: no branching on success. A blocked attempt and a successful attempt travel the same code path and land in the same file. That symmetry is what makes the log trustworthy — there is no code path that can silently drop a failure.

If you route scrapes through a scraping API instead of a raw HTTP client, the same harness applies; you just wrap the API call and verify the returned content the same way. The measurement problem is identical regardless of who owns the proxy pool.

Distinguishing Real Failures from Blocks and Soft Bans

A failure rate is only useful once you can split it. A 10% failure rate that is all ConnectionError from your own network is a different problem from a 10% failure rate that is all 403s from the target. And a “soft ban” — where the server returns 200 with a useless page — is a third problem that neither counter catches.

Classification needs three signals: status code, response size, and a content marker. Status alone is never enough.

BLOCK_MARKER = "Checking your browser"
MIN_PRODUCT_BYTES = 10_000  # real product pages on this target are ~80KB

def classify_outcome(resp) -> tuple[bool, bool]:
    """Returns (content_ok, block_detected)."""
    if resp.status_code in (403, 429):
        return False, True

    body = resp.text
    if BLOCK_MARKER in body[:2000]:
        return False, True

    if resp.status_code == 200 and len(resp.content) < MIN_PRODUCT_BYTES:
        # 200 + tiny body on a target whose pages are large = soft block
        return False, True

    # Verify the data we actually came for is present
    content_ok = (
        '"price"' in body
        and 'id="product-title"' in body
    )
    return content_ok, False

The marker string and the minimum size are target-specific. You have to look at one real page and one real block page from your target and pick signals that separate them. That is manual work, and it is the entire price of admission for trustworthy numbers. Generic heuristics like “contains the word captcha” miss custom interstitials.

The decision table the function encodes:

Response signalOutcome categoryCounts against
200, expected fields present, normal sizesuccessnothing
200, marker string in first 2KBblockblock rate
200, body under size floorblock (soft)block rate
403 / 429blockblock rate
200, full-size body, expected fields missinghard failurefailure rate (likely a redesign — see Your Scraper Died in a Redesign)
Timeout / connection errorhard failurefailure rate
5xxhard failurefailure rate

One opinion here that people push back on: I count a 200-with-missing-fields as a failure, not a block, even when it is probably a subtle challenge. The reason is operational. Blocks call for proxy and fingerprint changes; extraction failures call for selector fixes. Misfiling one as the other sends you debugging the wrong layer at 2 a.m. When you genuinely cannot tell, log both flags and decide from aggregate data, not per-request guessing.

Recording Latency Per Attempt with a Simple Timer Wrapper

Latency measurement has the same silent-failure problem as success counting. Most scrapers time the happy path and ignore retries, or time the request but not the classification work, or average everything into one number that hides the tail.

The rule: latency_ms is wall-clock time from “attempt started” to “attempt resolved,” including every retry in between. That is what the end user of your pipeline experiences, so it is what you measure.

import time
import requests

def timed_fetch(session: requests.Session, url: str, max_retries: int = 3):
    started = time.perf_counter()
    retries = 0
    last_exc = None
    resp = None

    for attempt in range(max_retries + 1):
        try:
            resp = session.get(url, timeout=30)
            if resp.status_code >= 500 or resp.status_code in (403, 429):
                # retriable from our perspective; classify_outcome decides the label
                if attempt < max_retries:
                    retries += 1
                    time.sleep(2 ** attempt)  # simple backoff
                    continue
            break
        except requests.RequestException as exc:
            last_exc = exc
            if attempt < max_retries:
                retries += 1
                time.sleep(2 ** attempt)
                continue

    latency_ms = (time.perf_counter() - started) * 1000
    return resp, latency_ms, retries, last_exc

Two things matter in this snippet. time.perf_counter() is monotonic — never use time.time() for intervals, because the system clock can jump mid-measurement and produce negative latencies. Second, the retry counter is recorded, so a “successful” attempt that took four tries shows up as retries: 3 with a fat latency_ms, and your analysis can treat retry-heavy successes as the warning sign they are.

What lands in the log:

{"url": "https://store.example.com/product/123", "domain": "store.example.com", "timestamp": "2025-11-03T14:02:11Z", "http_status": 200, "latency_ms": 412.5, "retries": 0, "content_ok": true, "block_detected": false, "error": null, "response_bytes": 84122, "run_label": "run-a"}
{"url": "https://store.example.com/product/456", "domain": "store.example.com", "timestamp": "2025-11-03T14:02:19Z", "http_status": 200, "latency_ms": 9187.3, "retries": 3, "content_ok": true, "block_detected": false, "error": null, "response_bytes": 79940, "run_label": "run-a"}

Wait — those timestamps contain a date. That is fine for the log file itself (your log needs real timestamps to be useful), but notice the second line: it “succeeded,” and a naive pipeline would count it the same as the first. It took 22x longer and burned four requests. That distinction only exists because the retry flag and cumulative latency were recorded.

Computing Success Rate and Block Rate from the Log

Once every attempt is a line in JSONL, the analysis is almost embarrassingly simple. That is the point — the hard work was getting honest records; the math is arithmetic.

import json
from collections import defaultdict
from datetime import datetime

def load_records(path: str) -> list[dict]:
    records = []
    with open(path, encoding="utf-8") as f:
        for line in f:
            line = line.strip()
            if line:
                records.append(json.loads(line))
    return records

def compute_rates(records: list[dict], since: str | None = None) -> dict:
    if since:
        cutoff = datetime.fromisoformat(since)
        records = [
            r for r in records
            if datetime.fromisoformat(r["timestamp"]) >= cutoff
        ]

    by_domain = defaultdict(list)
    for r in records:
        by_domain[r["domain"]].append(r)

    report = {}
    for domain, recs in by_domain.items():
        n = len(recs)
        successes = sum(1 for r in recs if r["content_ok"] and not r["block_detected"])
        blocks = sum(1 for r in recs if r["block_detected"])
        failures = n - successes - blocks
        report[domain] = {
            "attempts": n,
            "success_rate": round(100 * successes / n, 1),
            "block_rate": round(100 * blocks / n, 1),
            "failure_rate": round(100 * failures / n, 1),
            "avg_retries": round(sum(r["retries"] for r in recs) / n, 2),
        }
    return report

Success is defined strictly: content_ok and not blocked. A response that was both flagged as blocked and happened to contain the price string does not count — challenge pages sometimes leak fragments of the real DOM.

Sample output from a 500-URL run:

$ python analyze.py scrape_attempts.jsonl

domain             attempts  success  block  failure  avg_retries
------------------------------------------------------------------
store.example.com       300    64.3%  28.7%     7.0%         1.84
example.com             200    99.5%   0.0%     0.5%         0.01

That top line is the kind of number that starts real conversations. A 28.7% block rate on the store domain is not a scraper bug — it is a capacity and fingerprint problem, and no amount of selector tweaking will fix it. The naive counter would have reported 93% success (everything that returned 200 with any body) and the conversation would never have happened.

Calculating Latency Percentiles (p50, p95, p99) Locally

Averages lie about latency because latency distributions are skewed. One 30-second timeout among 500 fast responses drags the mean up while telling you nothing about the typical experience, and a mean of 900ms can hide a p99 of 9 seconds that is quietly destroying your throughput budget.

Percentiles fix this. p50 is the typical attempt, p95 is the bad-but-expected tail, p99 is the tail that dictates your timeout and concurrency settings.

import statistics

def latency_percentiles(records: list[dict], domain: str | None = None) -> dict:
    if domain:
        records = [r for r in records if r["domain"] == domain]

    latencies = sorted(r["latency_ms"] for r in records if r["latency_ms"] > 0)
    if len(latencies) < 4:
        raise ValueError("Need at least 4 attempts for meaningful percentiles")

    # n=100, method='inclusive' gives the standard percentile definitions
    q = statistics.quantiles(latencies, n=100, method="inclusive")
    return {
        "count": len(latencies),
        "mean_ms": round(statistics.fmean(latencies), 1),
        "p50_ms": round(q[49], 1),
        "p95_ms": round(q[94], 1),
        "p99_ms": round(q[98], 1),
        "max_ms": round(latencies[-1], 1),
    }

From the same 500-attempt run:

MetricValue (ms)What it tells you
mean1,940Almost useless — inflated by a handful of retries
p50610The honest “normal” attempt
p954,880Retries and slow challenge cycles live here
p9911,240One timeout away from your 30s ceiling
max27,930A full retry chain that nearly hit the wall

The mean is more than 3x the p50. If you sized worker pools or promised SLAs off the mean, you sized them for a system that does not exist. The p99 is the number that should set your timeout: a 15-second timeout would have failed those p99 attempts outright, while your current 30-second timeout means nearly a full minute of worker time gets consumed by the worst 1% of requests. Whether that trade is worth it depends on whether those slow attempts succeed — which you can check, because content_ok is in the same log line.

Running Repeatable Benchmark Runs and Spotting Regressions

A single measurement is an anecdote. The harness earns its keep when you run it repeatedly under identical conditions and compare runs — before/after a code change, a proxy switch, or a new fingerprint profile.

The benchmark needs to be pinned: same URL list, same concurrency, same classification rules. Store the config with the run so future-you knows what “identical” meant.

BENCHMARK_CONFIG = {
    "run_label": "v2-stealth-profile",
    "urls_file": "fixtures/urls_500.txt",   # frozen URL list
    "concurrency": 4,
    "timeout_s": 30,
    "max_retries": 3,
    "fixture_mode": False,   # True = serve from local fixture dir, no network
    "classifier": "store.example.com.v3",  # version your classification rules
}

Fixture mode deserves a word. Pointing the harness at a local directory of saved HTML files (via a tiny http.server or by patching the fetch function) lets you test the measurement and classification layer with zero network variance. It will not tell you anything about blocks, but it will catch regressions in your own extraction and logging code — which are more common than target-side changes.

Then the comparison:

def compare_runs(baseline: list[dict], candidate: list[dict]) -> None:
    base = compute_rates(baseline)
    cand = compute_rates(candidate)

    for domain in base:
        b, c = base[domain], cand[domain]
        success_drop = b["success_rate"] - c["success_rate"]
        print(f"\n{domain}")
        print(f"  success: {b['success_rate']}% -> {c['success_rate']}%  ({success_drop:+.1f})")
        print(f"  block:   {b['block_rate']}% -> {c['block_rate']}%")

        bl = latency_percentiles([r for r in baseline if r["domain"] == domain])
        cl = latency_percentiles([r for r in candidate if r["domain"] == domain])
        p95_delta = cl["p95_ms"] - bl["p95_ms"]

        if success_drop > 5.0:
            print("  REGRESSION: success rate dropped more than 5 points")
        if p95_delta > 2000:
            print(f"  REGRESSION: p95 latency up {p95_delta:+.0f} ms")

The thresholds are opinions, not laws. A 5-point success drop and a 2-second p95 increase are aggressive enough to catch real regressions without paging you for noise. If your targets are stable, tighten them; if you scrape volatile sites where block rates swing by the hour, widen them and compare medians of several runs rather than single runs. Comparing single runs against a noisy target is the one way this whole setup lies to you — always keep the baseline fixed and re-run the candidate more than once before declaring a regression.

Gotchas and Next Steps

A few things that bite people building this for the first time:

  • Classification rules drift. When the target redesigns, your marker strings and size floors go stale and mislabel outcomes. Version the classifier config per run label so you can tell “the site changed” from “my classifier broke.”
  • JSONL grows forever. Rotate the file by run label or date. The analysis scripts read everything you point them at, including last month’s runs with different rules.
  • Concurrency changes the measurement. Latency under 4 concurrent workers is not comparable to latency under 32. Pin concurrency in the benchmark config and never compare across it.
  • Do not classify on truncated bodies. If your client streams and you cut the read short, a “tiny body” looks like a soft block. Read fully, then classify.

Next steps, in order of payoff: wire the harness output into your CI so every scraper change ships with a benchmark report; add a proxy_country or session dimension to the records so you can slice block rate by exit geography; and if you are running at batch scale, move the same one-record-per-attempt discipline to your async job pipeline — the measurement principles are identical, and the job status endpoints give you the same status/latency data server-side. For that side of the house, Async Scraping at Scale: Jobs, Batches, Webhooks covers the mechanics.

The harness is about 150 lines of Python. It has no dependencies beyond requests. And unlike the counter in your scraper’s log, every number it produces is one you can defend in a meeting — because every number traces back to a single recorded attempt with its status, its latency, its retries, and whether the data you wanted was actually in the response.

#web scraping #observability #success rate #logging #latency #slot:measurement

Related Articles