Skip to content
Tutorial 14 min read

Re-Drive Only the Failures: Handling Partial Batch Results

A batch scrape rarely fails cleanly. Learn to reconcile partial results, classify why each URL failed, and re-drive only the misses without paying twice.

FE
FineData Engineering · Editorial Policy
|
On this page

When a Completed Batch Isn’t Complete

You submitted 10,000 URLs in a batch. The batch status says completed. You open the results and find 9,588 rows. Where are the other 412? Nobody flagged them. Nothing errored loudly. The batch just… finished, and 4% of your data silently didn’t arrive.

This is the actual failure mode of batch scraping at scale. Batches don’t fail atomically. They fail partially: a few jobs hit throttles, a few hit dead links, a worker dies mid-write, a response comes back with an empty body. If your pipeline treats “batch completed” as “all data present,” you ship a dataset with holes in it and don’t find out until a downstream report looks wrong.

The fix is boring and effective: keep a per-URL ledger, reconcile it against your input manifest, classify each miss, and re-drive only the misses. This tutorial builds that loop end to end, using the async batch endpoints as the substrate.

Modeling Partial Batch Results: The Completion Ledger Pattern

The core data structure is a ledger: one record per input URL, tracking what happened to it across every pass. Not per batch. Not per job. Per URL. The URL is the only stable identity you have across multiple batch submissions.

A ledger entry for store.example.com/product/123 looks like this:

{
  "url": "https://store.example.com/product/123",
  "status": "failed",
  "http_code": 429,
  "error_class": "backoff_required",
  "attempts": 2,
  "last_error_at": "2025-01-14T09:31:22Z",
  "job_id": "job_8f21a",
  "batch_id": "batch_c4d90",
  "formats": ["markdown"],
  "result": null
}

Every field earns its place. http_code and error_class drive retry decisions. attempts prevents infinite loops. job_id and batch_id let you trace a failure back to the originating API call when a customer asks why a specific product is missing from Tuesday’s dataset.

Here’s a minimal implementation. It’s a dict keyed by URL because that’s all you need; if you want durability, back it with SQLite or Postgres, but the in-memory shape is the same:

from dataclasses import dataclass, field, asdict
from typing import Optional
import json

@dataclass
class LedgerEntry:
    url: str
    status: str = "pending"          # pending | success | failed
    http_code: Optional[int] = None
    error_class: Optional[str] = None  # retryable | permanent | backoff_required
    attempts: int = 0
    last_error_at: Optional[str] = None
    job_id: Optional[str] = None
    batch_id: Optional[str] = None
    result: Optional[dict] = None

class BatchLedger:
    def __init__(self, urls):
        self.entries = {u: LedgerEntry(url=u) for u in urls}

    def mark_success(self, url, result, job_id=None):
        e = self.entries[url]
        e.status = "success"
        e.result = result
        e.job_id = job_id
        e.attempts += 1
        e.http_code = None
        e.error_class = None

    def mark_failed(self, url, http_code, error_class, job_id=None, ts=None):
        e = self.entries[url]
        e.status = "failed"
        e.http_code = http_code
        e.error_class = error_class
        e.job_id = job_id
        e.attempts += 1
        e.last_error_at = ts

    def to_manifest(self):
        return [asdict(e) for e in self.entries.values()]

One design rule I hold firmly: the ledger is keyed by the input URL, not by whatever the site redirected to. If store.example.com/product/123 301s to store.example.com/p/123, you still record the outcome under the original URL. Otherwise your reconciliation diff in the next section produces false positives, and you’ll re-scrape URLs that actually succeeded.

Reconciling Input URLs Against Scraped Output Before Anything Else

Before you classify anything, you need to know what’s actually missing. “The batch said 10,000 jobs, I got 9,588 rows” is three different problems wearing one coat:

  • 412 jobs genuinely failed and returned an error status.
  • Some jobs “succeeded” but the write to your results file got truncated or the worker crashed mid-flush.
  • Some rows exist but are unparseable garbage — a login page captured as markdown, a zero-byte body, JSON that didn’t survive the trip.

Reconciliation runs before classification because you can’t classify a row you never received. The batch itself goes in as a POST with a JSON body — one entry per URL under the top-level requests key, plus a callback_url that fires when every job in the batch completes:

curl -X POST "https://api.finedata.ai/api/v1/async/batch" \
  -H "Authorization: Bearer fd_your_api_key" \
  -H "Content-Type: application/json" \
  -d '{
    "callback_url": "https://example.com/webhooks/batch-complete",
    "requests": [
      {"url": "https://store.example.com/product/123", "formats": ["markdown"]},
      {"url": "https://store.example.com/product/456", "formats": ["markdown"]}
    ]
  }'

When the callback fires, pull the batch status with results included — GET /api/v1/async/batch/{batch_id}?include_results=true — and diff three ways:

def reconcile(input_manifest, results_file):
    input_urls = {row["url"] for row in input_manifest}
    results = load_jsonl(results_file)   # list of dicts with "url" and "content"

    seen_urls = set()
    missing_urls, empty_results, unparseable_rows = [], [], []

    for r in results:
        url = r.get("url")
        seen_urls.add(url)
        content = r.get("content") or ""
        if len(content.strip()) == 0:
            empty_results.append(url)
        elif not looks_like_product_page(content):  # your domain-specific sanity check
            unparseable_rows.append(url)

    missing_urls = [u for u in input_urls if u not in seen_urls]
    return missing_urls, empty_results, unparseable_rows

The gap types and how to catch each:

Gap typeWhat it looks likeDetection method
Missing outputURL in input, absent from results entirelySet difference: input URLs minus result URLs
Empty bodyRow exists, content is zero bytes or whitespaceLength check on the content field
Malformed recordRow exists, content is the wrong page (interstitial, login wall, error page)Domain-specific validator: expected field present, minimum length, expected domain in the final URL

The third one is the sneaky one. A job can return HTTP 200 with a fully rendered “access denied” interstitial and your pipeline will happily store it as a success. If you scrape store.example.com product pages, write a validator that checks the page actually contains a price. It’s ten lines of code and it catches the failure class that status codes never will.

This reconciliation is also where you catch the dumb infrastructure bugs. I’ve seen pipelines lose 200 rows because a log-rotation script truncated the output file mid-run. The batch API did its job perfectly. The gap was on our side, and re-driving those URLs through the API would have billed us for a bug in our own plumbing.

Classifying Failures: Retryable, Permanent, and Backoff-Required

Once you know which URLs missed, sort them by why. Retrying a 404 is wasted money. Retrying a 429 immediately is worse — it deepens the throttle. Retrying a DNS failure once is fine; retrying it five times in a minute is noise.

Error signatureClassReasoning
HTTP 429backoff_requiredThe host is throttling you; retry later with delay, possibly through a different exit IP
HTTP 503backoff_requiredOverloaded origin or soft block; delay and retry
HTTP 403retryable (once)Often an anti-bot interstitial; escalate stealth settings on retry
HTTP 404permanentThe page is gone; no retry will fix this
TimeoutretryableTransient by nature; retry with the same settings first
DNS errorretryableUsually transient infrastructure; retry once or twice
Empty body (200)retryableLikely an interstitial or render miss; retry with JS rendering

The 403 row is where people disagree with me. Many pipelines treat 403 as permanent because “the site blocked us.” In practice, a first-pass 403 frequently succeeds on a second pass with a residential exit IP or a stronger stealth profile — the block was fingerprint-based, not account-based. Budget one escalated retry for 403s, and only then call them permanent. That single policy change took one of my pipelines from a 4% permanent-failure rate to under 1%.

Here’s the classifier:

from enum import Enum

class ErrorClass(Enum):
    RETRYABLE = "retryable"
    PERMANENT = "permanent"
    BACKOFF_REQUIRED = "backoff_required"

def classify(record):
    """record: dict with http_code, error, content for a store.example.com URL."""
    code = record.get("http_code")
    err = (record.get("error") or "").lower()

    if code == 429:
        return ErrorClass.BACKOFF_REQUIRED, "throttled by origin"
    if code == 503:
        return ErrorClass.BACKOFF_REQUIRED, "origin overloaded or soft block"
    if code == 404:
        return ErrorClass.PERMANENT, "page removed"
    if code == 403:
        return ErrorClass.RETRYABLE, "possible anti-bot interstitial; escalate on retry"
    if "timeout" in err or code == 408:
        return ErrorClass.RETRYABLE, "request timed out"
    if "dns" in err:
        return ErrorClass.RETRYABLE, "transient DNS failure"
    if code == 200 and not (record.get("content") or "").strip():
        return ErrorClass.RETRYABLE, "empty body; likely interstitial"
    if code is None:
        return ErrorClass.RETRYABLE, "no response recorded; possible write loss"
    return ErrorClass.PERMANENT, f"unmapped status {code}"

Note the last fallback. Anything you can’t classify defaults to permanent, not retryable. An unmapped error that you blindly retry five times is how you burn through a token budget on a bug you haven’t diagnosed yet. Fail closed, log it, and promote it to retryable only after a human looks at one sample.

Building the Re-Drive Queue From the Ledger, Not the Original Input

The most expensive mistake in partial-batch handling is re-running the original input file. It feels safe. It’s also how you pay twice for 9,588 successes.

The re-drive queue comes exclusively from the ledger:

MAX_ATTEMPTS = 3

def build_redrive_manifest(ledger, max_attempts=MAX_ATTEMPTS):
    worklist = []
    for e in ledger.entries.values():
        if e.status == "success":
            continue
        if e.error_class == "permanent":
            continue
        if e.attempts >= max_attempts:
            continue
        worklist.append({
            "url": e.url,
            "prior_attempts": e.attempts,
            "prior_http_code": e.http_code,
            "escalate": e.error_class == "backoff_required" or e.http_code == 403,
        })
    return worklist

The escalate flag matters. A URL that failed with 429 through a datacenter exit IP will probably fail again the same way. The re-drive should change something: render the page, scroll to trigger lazy-loaded content, wait for the network to settle, allow more retries. Retrying with identical parameters and hoping for a different outcome is not a strategy.

import requests

def to_scrape_request(item):
    req = {"url": item["url"], "formats": ["markdown"]}
    if item["escalate"]:
        # Escalate using only documented request fields: render the page,
        # scroll to load lazy content, wait for the network to settle,
        # and allow more retries than the default of 5.
        req.update({
            "js_scroll": True,
            "js_wait_for": "networkidle",
            "max_retries": 8,
        })
    return req

def submit_redrive(worklist):
    body = {
        "callback_url": "https://example.com/webhooks/redrive-complete",
        "requests": [to_scrape_request(i) for i in worklist],
    }
    resp = requests.post(
        "https://api.finedata.ai/api/v1/async/batch",
        headers={"Authorization": "Bearer fd_your_api_key"},
        json=body,
    )
    return resp.json()["batch_id"]

Now the money argument, with the numbers from the opening scenario. Assume an average of 5 tokens per request once you count retries and rendering on the escalated subset:

StrategyURLs re-billedApprox. token cost
Re-run entire input10,000~50,000
Re-drive classified misses only412~2,400

That’s a 95% reduction in re-drive cost, and it compounds. A nightly pipeline that re-runs its full input every time any URL misses is paying roughly double its necessary cost every single night. The ledger pays for itself the first week. For more on when batching even makes sense versus single calls, see Batch Requests vs Single Calls: Scraping Pattern Trade-offs.

Respecting Rate Limits on the Retry Pass Without Slowing the Whole Batch

Here’s the trap: your first pass failed partly because you were going too fast. If the re-drive fires the same 412 URLs at the same concurrency into the same host, you’ll reproduce the exact conditions that caused the misses.

The re-drive pass gets its own throttle policy, tuned per host:

BACKOFF_CONFIG = {
    "store.example.com": {
        "base_delay_s": 30,       # first retry waits at least 30s
        "multiplier": 2.0,        # 30s, 60s, 120s...
        "jitter_range_s": (5, 15), # plus 5-15s random jitter
        "per_host_cap": 5,        # max 5 concurrent requests to this host
    }
}

Jitter is not optional decoration. If 200 URLs all got 429’d at the same moment, they all carry the same last_error_at. Retry them all at last_error_at + 30s and you’ve built a synchronized hammer. The jitter range breaks the cohort apart.

The retry wrapper reads attempt history from the ledger to decide eligibility:

import random, time
from datetime import datetime, timedelta, timezone

def parse_ts(s):
    """Z-safe ISO 8601 parse: datetime.fromisoformat() rejects a trailing 'Z',
    so normalize it to the equivalent +00:00 offset before parsing."""
    return datetime.fromisoformat(s.replace("Z", "+00:00"))

def next_eligible_time(entry, cfg):
    if entry.attempts == 0:
        return datetime.now(timezone.utc)
    delay = cfg["base_delay_s"] * (cfg["multiplier"] ** (entry.attempts - 1))
    jitter = random.uniform(*cfg["jitter_range_s"])
    last = parse_ts(entry.last_error_at)
    return last + timedelta(seconds=delay + jitter)

def is_eligible(entry, cfg):
    return datetime.now(timezone.utc) >= next_eligible_time(entry, cfg)

One nuance: backoff_required entries should also consider switching exit geography or proxy type on retry. A host that throttled your datacenter IP range may accept a residential request immediately — no delay needed at all. Backoff and IP rotation are substitutes; sometimes rotation is cheaper than waiting. There’s a longer treatment of this in Client-Side vs Service-Side Retries for Scraping Failures, but the short version: keep retry policy in your ledger loop, because only the ledger knows the attempt count.

Merging Re-Drive Results Back Into the Ledger Idempotently

The re-drive produces a second results file. Fold it into the same ledger with an upsert keyed by URL. First-pass successes are never touched; only retried entries update:

def merge_redrive(ledger, redrive_results, batch_id):
    for r in redrive_results:
        url = r["url"]
        e = ledger.entries.get(url)
        if e is None:
            # URL not in ledger: someone submitted outside the pipeline. Log it.
            continue
        if e.status == "success":
            continue  # never overwrite a confirmed success

        if r.get("status") == "completed" and (r.get("content") or "").strip():
            ledger.mark_success(url, result=r, job_id=r.get("job_id"))
        else:
            cls, reason = classify(r)
            ledger.mark_failed(
                url,
                http_code=r.get("http_code"),
                error_class=cls.value,
                job_id=r.get("job_id"),
                ts=r.get("finished_at"),
            )

The if e.status == "success": continue guard is what makes the merge idempotent. Run the merge twice, out of order, or on overlapping files — the ledger converges to the same state. That property is what lets you run a third pass if needed without any special-casing.

Here’s the before/after for three entries:

Before the re-drive:

[
  {"url": "https://store.example.com/product/123", "status": "failed",
   "http_code": 429, "error_class": "backoff_required", "attempts": 1},
  {"url": "https://store.example.com/product/456", "status": "failed",
   "http_code": 404, "error_class": "permanent", "attempts": 1},
  {"url": "https://store.example.com/product/789", "status": "failed",
   "http_code": 403, "error_class": "retryable", "attempts": 1}
]

After:

[
  {"url": "https://store.example.com/product/123", "status": "success",
   "http_code": null, "error_class": null, "attempts": 2},
  {"url": "https://store.example.com/product/456", "status": "failed",
   "http_code": 404, "error_class": "permanent", "attempts": 1},
  {"url": "https://store.example.com/product/789", "status": "failed",
   "http_code": 404, "error_class": "permanent", "attempts": 2}
]

Read that third entry carefully. The 403 was classified retryable, got its one escalated retry, and came back 404 — the product was delisted and the first pass just happened to catch an interstitial instead of the real answer. The classifier promoted it to permanent. That’s the system working: one retry bought you the truth, and the ledger records why.

Deciding When to Stop: Exit Criteria and Reporting the Residue

Re-drive loops need explicit stop conditions, or they run forever on a shrinking, stubborn set of URLs. Three criteria, any one of which terminates the loop:

loop:
    worklist = build_redrive_manifest(ledger)
    if worklist is empty:                      STOP (clean)
    if pass_count >= MAX_PASSES (3):          STOP (budget)
    if time_elapsed > TIME_BUDGET (45 min):   STOP (deadline)
    if len(worklist) < FLOOR (10 URLs):        STOP (diminishing returns)
    submit_redrive(worklist)
    wait for callback or poll batch status
    merge_redrive(ledger, results)

The floor criterion is the one people skip. If 10,000 URLs have shrunk to 7 unresolved entries, another full pass costs more in orchestration time than the data is worth. Below the floor, the residue goes to a report, not back into the queue.

The end-of-batch report is the deliverable that makes this whole pattern defensible to stakeholders. It says exactly what’s missing and why:

BATCH REPORT — store.example.com nightly product sync
Input: 10,000 URLs | Success: 9,995 | Unresolved: 5 | Passes: 2

Unresolved URLs:
1. https://store.example.com/product/456  — permanent (404, 2 attempts)
   Action: flag for delisting review; likely removed from catalog
2. https://store.example.com/product/789  — permanent (404 after escalation)
   Action: flag for delisting review
3. https://store.example.com/product/1011 — backoff_required (429, 3 attempts)
   Action: host still throttling after 3 passes; consider carrier exit IP
   or move this URL to the next scheduled window

Every unresolved entry gets a recommended action. “5 URLs failed” is an incident; “5 URLs failed, 2 are delisted products, 1 needs a different exit IP, here’s the plan” is just Tuesday.

Wrap-Up

Partial batch failure is the norm, not the exception, so build for it: a per-URL ledger as the single source of truth, a three-way reconciliation before classification, error classes that separate dead links from throttles, a re-drive queue built only from the misses, backoff with jitter on the retry pass, idempotent merging, and hard stop criteria with a residue report.

The economics alone justify the work — re-driving 412 classified misses instead of 10,000 URLs cuts retry cost by roughly 95%, and the ledger makes every failure explainable after the fact.

Gotchas worth repeating: key the ledger by input URL, not redirect targets; default unmapped errors to permanent, not retryable; treat empty 200 responses as failures, not successes; and always change something on an escalated retry — identical parameters reproduce identical failures.

Next steps if you’re wiring this into a real pipeline: read Async Scraping at Scale: Jobs, Batches, Webhooks for the callback mechanics that trigger the merge step, and Building ETL Pipelines with Scraping APIs for where the ledger sits relative to your transform and load stages.

#batch-scraping #data-pipelines #retry-strategy #error-classification #web-scraping-api #slot:pipeline

Related Articles