Re-Drive Only the Failures: Handling Partial Batch Results
A batch scrape rarely fails cleanly. Learn to reconcile partial results, classify why each URL failed, and re-drive only the misses without paying twice.
On this page
When a Completed Batch Isn’t Complete
You submitted 10,000 URLs in a batch. The batch status says completed. You open the results and find 9,588 rows. Where are the other 412? Nobody flagged them. Nothing errored loudly. The batch just… finished, and 4% of your data silently didn’t arrive.
This is the actual failure mode of batch scraping at scale. Batches don’t fail atomically. They fail partially: a few jobs hit throttles, a few hit dead links, a worker dies mid-write, a response comes back with an empty body. If your pipeline treats “batch completed” as “all data present,” you ship a dataset with holes in it and don’t find out until a downstream report looks wrong.
The fix is boring and effective: keep a per-URL ledger, reconcile it against your input manifest, classify each miss, and re-drive only the misses. This tutorial builds that loop end to end, using the async batch endpoints as the substrate.
Modeling Partial Batch Results: The Completion Ledger Pattern
The core data structure is a ledger: one record per input URL, tracking what happened to it across every pass. Not per batch. Not per job. Per URL. The URL is the only stable identity you have across multiple batch submissions.
A ledger entry for store.example.com/product/123 looks like this:
{
"url": "https://store.example.com/product/123",
"status": "failed",
"http_code": 429,
"error_class": "backoff_required",
"attempts": 2,
"last_error_at": "2025-01-14T09:31:22Z",
"job_id": "job_8f21a",
"batch_id": "batch_c4d90",
"formats": ["markdown"],
"result": null
}
Every field earns its place. http_code and error_class drive retry decisions. attempts prevents infinite loops. job_id and batch_id let you trace a failure back to the originating API call when a customer asks why a specific product is missing from Tuesday’s dataset.
Here’s a minimal implementation. It’s a dict keyed by URL because that’s all you need; if you want durability, back it with SQLite or Postgres, but the in-memory shape is the same:
from dataclasses import dataclass, field, asdict
from typing import Optional
import json
@dataclass
class LedgerEntry:
url: str
status: str = "pending" # pending | success | failed
http_code: Optional[int] = None
error_class: Optional[str] = None # retryable | permanent | backoff_required
attempts: int = 0
last_error_at: Optional[str] = None
job_id: Optional[str] = None
batch_id: Optional[str] = None
result: Optional[dict] = None
class BatchLedger:
def __init__(self, urls):
self.entries = {u: LedgerEntry(url=u) for u in urls}
def mark_success(self, url, result, job_id=None):
e = self.entries[url]
e.status = "success"
e.result = result
e.job_id = job_id
e.attempts += 1
e.http_code = None
e.error_class = None
def mark_failed(self, url, http_code, error_class, job_id=None, ts=None):
e = self.entries[url]
e.status = "failed"
e.http_code = http_code
e.error_class = error_class
e.job_id = job_id
e.attempts += 1
e.last_error_at = ts
def to_manifest(self):
return [asdict(e) for e in self.entries.values()]
One design rule I hold firmly: the ledger is keyed by the input URL, not by whatever the site redirected to. If store.example.com/product/123 301s to store.example.com/p/123, you still record the outcome under the original URL. Otherwise your reconciliation diff in the next section produces false positives, and you’ll re-scrape URLs that actually succeeded.
Reconciling Input URLs Against Scraped Output Before Anything Else
Before you classify anything, you need to know what’s actually missing. “The batch said 10,000 jobs, I got 9,588 rows” is three different problems wearing one coat:
- 412 jobs genuinely failed and returned an error status.
- Some jobs “succeeded” but the write to your results file got truncated or the worker crashed mid-flush.
- Some rows exist but are unparseable garbage — a login page captured as markdown, a zero-byte body, JSON that didn’t survive the trip.
Reconciliation runs before classification because you can’t classify a row you never received. The batch itself goes in as a POST with a JSON body — one entry per URL under the top-level requests key, plus a callback_url that fires when every job in the batch completes:
curl -X POST "https://api.finedata.ai/api/v1/async/batch" \
-H "Authorization: Bearer fd_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"callback_url": "https://example.com/webhooks/batch-complete",
"requests": [
{"url": "https://store.example.com/product/123", "formats": ["markdown"]},
{"url": "https://store.example.com/product/456", "formats": ["markdown"]}
]
}'
When the callback fires, pull the batch status with results included — GET /api/v1/async/batch/{batch_id}?include_results=true — and diff three ways:
def reconcile(input_manifest, results_file):
input_urls = {row["url"] for row in input_manifest}
results = load_jsonl(results_file) # list of dicts with "url" and "content"
seen_urls = set()
missing_urls, empty_results, unparseable_rows = [], [], []
for r in results:
url = r.get("url")
seen_urls.add(url)
content = r.get("content") or ""
if len(content.strip()) == 0:
empty_results.append(url)
elif not looks_like_product_page(content): # your domain-specific sanity check
unparseable_rows.append(url)
missing_urls = [u for u in input_urls if u not in seen_urls]
return missing_urls, empty_results, unparseable_rows
The gap types and how to catch each:
| Gap type | What it looks like | Detection method |
|---|---|---|
| Missing output | URL in input, absent from results entirely | Set difference: input URLs minus result URLs |
| Empty body | Row exists, content is zero bytes or whitespace | Length check on the content field |
| Malformed record | Row exists, content is the wrong page (interstitial, login wall, error page) | Domain-specific validator: expected field present, minimum length, expected domain in the final URL |
The third one is the sneaky one. A job can return HTTP 200 with a fully rendered “access denied” interstitial and your pipeline will happily store it as a success. If you scrape store.example.com product pages, write a validator that checks the page actually contains a price. It’s ten lines of code and it catches the failure class that status codes never will.
This reconciliation is also where you catch the dumb infrastructure bugs. I’ve seen pipelines lose 200 rows because a log-rotation script truncated the output file mid-run. The batch API did its job perfectly. The gap was on our side, and re-driving those URLs through the API would have billed us for a bug in our own plumbing.
Classifying Failures: Retryable, Permanent, and Backoff-Required
Once you know which URLs missed, sort them by why. Retrying a 404 is wasted money. Retrying a 429 immediately is worse — it deepens the throttle. Retrying a DNS failure once is fine; retrying it five times in a minute is noise.
| Error signature | Class | Reasoning |
|---|---|---|
| HTTP 429 | backoff_required | The host is throttling you; retry later with delay, possibly through a different exit IP |
| HTTP 503 | backoff_required | Overloaded origin or soft block; delay and retry |
| HTTP 403 | retryable (once) | Often an anti-bot interstitial; escalate stealth settings on retry |
| HTTP 404 | permanent | The page is gone; no retry will fix this |
| Timeout | retryable | Transient by nature; retry with the same settings first |
| DNS error | retryable | Usually transient infrastructure; retry once or twice |
| Empty body (200) | retryable | Likely an interstitial or render miss; retry with JS rendering |
The 403 row is where people disagree with me. Many pipelines treat 403 as permanent because “the site blocked us.” In practice, a first-pass 403 frequently succeeds on a second pass with a residential exit IP or a stronger stealth profile — the block was fingerprint-based, not account-based. Budget one escalated retry for 403s, and only then call them permanent. That single policy change took one of my pipelines from a 4% permanent-failure rate to under 1%.
Here’s the classifier:
from enum import Enum
class ErrorClass(Enum):
RETRYABLE = "retryable"
PERMANENT = "permanent"
BACKOFF_REQUIRED = "backoff_required"
def classify(record):
"""record: dict with http_code, error, content for a store.example.com URL."""
code = record.get("http_code")
err = (record.get("error") or "").lower()
if code == 429:
return ErrorClass.BACKOFF_REQUIRED, "throttled by origin"
if code == 503:
return ErrorClass.BACKOFF_REQUIRED, "origin overloaded or soft block"
if code == 404:
return ErrorClass.PERMANENT, "page removed"
if code == 403:
return ErrorClass.RETRYABLE, "possible anti-bot interstitial; escalate on retry"
if "timeout" in err or code == 408:
return ErrorClass.RETRYABLE, "request timed out"
if "dns" in err:
return ErrorClass.RETRYABLE, "transient DNS failure"
if code == 200 and not (record.get("content") or "").strip():
return ErrorClass.RETRYABLE, "empty body; likely interstitial"
if code is None:
return ErrorClass.RETRYABLE, "no response recorded; possible write loss"
return ErrorClass.PERMANENT, f"unmapped status {code}"
Note the last fallback. Anything you can’t classify defaults to permanent, not retryable. An unmapped error that you blindly retry five times is how you burn through a token budget on a bug you haven’t diagnosed yet. Fail closed, log it, and promote it to retryable only after a human looks at one sample.
Building the Re-Drive Queue From the Ledger, Not the Original Input
The most expensive mistake in partial-batch handling is re-running the original input file. It feels safe. It’s also how you pay twice for 9,588 successes.
The re-drive queue comes exclusively from the ledger:
MAX_ATTEMPTS = 3
def build_redrive_manifest(ledger, max_attempts=MAX_ATTEMPTS):
worklist = []
for e in ledger.entries.values():
if e.status == "success":
continue
if e.error_class == "permanent":
continue
if e.attempts >= max_attempts:
continue
worklist.append({
"url": e.url,
"prior_attempts": e.attempts,
"prior_http_code": e.http_code,
"escalate": e.error_class == "backoff_required" or e.http_code == 403,
})
return worklist
The escalate flag matters. A URL that failed with 429 through a datacenter exit IP will probably fail again the same way. The re-drive should change something: render the page, scroll to trigger lazy-loaded content, wait for the network to settle, allow more retries. Retrying with identical parameters and hoping for a different outcome is not a strategy.
import requests
def to_scrape_request(item):
req = {"url": item["url"], "formats": ["markdown"]}
if item["escalate"]:
# Escalate using only documented request fields: render the page,
# scroll to load lazy content, wait for the network to settle,
# and allow more retries than the default of 5.
req.update({
"js_scroll": True,
"js_wait_for": "networkidle",
"max_retries": 8,
})
return req
def submit_redrive(worklist):
body = {
"callback_url": "https://example.com/webhooks/redrive-complete",
"requests": [to_scrape_request(i) for i in worklist],
}
resp = requests.post(
"https://api.finedata.ai/api/v1/async/batch",
headers={"Authorization": "Bearer fd_your_api_key"},
json=body,
)
return resp.json()["batch_id"]
Now the money argument, with the numbers from the opening scenario. Assume an average of 5 tokens per request once you count retries and rendering on the escalated subset:
| Strategy | URLs re-billed | Approx. token cost |
|---|---|---|
| Re-run entire input | 10,000 | ~50,000 |
| Re-drive classified misses only | 412 | ~2,400 |
That’s a 95% reduction in re-drive cost, and it compounds. A nightly pipeline that re-runs its full input every time any URL misses is paying roughly double its necessary cost every single night. The ledger pays for itself the first week. For more on when batching even makes sense versus single calls, see Batch Requests vs Single Calls: Scraping Pattern Trade-offs.
Respecting Rate Limits on the Retry Pass Without Slowing the Whole Batch
Here’s the trap: your first pass failed partly because you were going too fast. If the re-drive fires the same 412 URLs at the same concurrency into the same host, you’ll reproduce the exact conditions that caused the misses.
The re-drive pass gets its own throttle policy, tuned per host:
BACKOFF_CONFIG = {
"store.example.com": {
"base_delay_s": 30, # first retry waits at least 30s
"multiplier": 2.0, # 30s, 60s, 120s...
"jitter_range_s": (5, 15), # plus 5-15s random jitter
"per_host_cap": 5, # max 5 concurrent requests to this host
}
}
Jitter is not optional decoration. If 200 URLs all got 429’d at the same moment, they all carry the same last_error_at. Retry them all at last_error_at + 30s and you’ve built a synchronized hammer. The jitter range breaks the cohort apart.
The retry wrapper reads attempt history from the ledger to decide eligibility:
import random, time
from datetime import datetime, timedelta, timezone
def parse_ts(s):
"""Z-safe ISO 8601 parse: datetime.fromisoformat() rejects a trailing 'Z',
so normalize it to the equivalent +00:00 offset before parsing."""
return datetime.fromisoformat(s.replace("Z", "+00:00"))
def next_eligible_time(entry, cfg):
if entry.attempts == 0:
return datetime.now(timezone.utc)
delay = cfg["base_delay_s"] * (cfg["multiplier"] ** (entry.attempts - 1))
jitter = random.uniform(*cfg["jitter_range_s"])
last = parse_ts(entry.last_error_at)
return last + timedelta(seconds=delay + jitter)
def is_eligible(entry, cfg):
return datetime.now(timezone.utc) >= next_eligible_time(entry, cfg)
One nuance: backoff_required entries should also consider switching exit geography or proxy type on retry. A host that throttled your datacenter IP range may accept a residential request immediately — no delay needed at all. Backoff and IP rotation are substitutes; sometimes rotation is cheaper than waiting. There’s a longer treatment of this in Client-Side vs Service-Side Retries for Scraping Failures, but the short version: keep retry policy in your ledger loop, because only the ledger knows the attempt count.
Merging Re-Drive Results Back Into the Ledger Idempotently
The re-drive produces a second results file. Fold it into the same ledger with an upsert keyed by URL. First-pass successes are never touched; only retried entries update:
def merge_redrive(ledger, redrive_results, batch_id):
for r in redrive_results:
url = r["url"]
e = ledger.entries.get(url)
if e is None:
# URL not in ledger: someone submitted outside the pipeline. Log it.
continue
if e.status == "success":
continue # never overwrite a confirmed success
if r.get("status") == "completed" and (r.get("content") or "").strip():
ledger.mark_success(url, result=r, job_id=r.get("job_id"))
else:
cls, reason = classify(r)
ledger.mark_failed(
url,
http_code=r.get("http_code"),
error_class=cls.value,
job_id=r.get("job_id"),
ts=r.get("finished_at"),
)
The if e.status == "success": continue guard is what makes the merge idempotent. Run the merge twice, out of order, or on overlapping files — the ledger converges to the same state. That property is what lets you run a third pass if needed without any special-casing.
Here’s the before/after for three entries:
Before the re-drive:
[
{"url": "https://store.example.com/product/123", "status": "failed",
"http_code": 429, "error_class": "backoff_required", "attempts": 1},
{"url": "https://store.example.com/product/456", "status": "failed",
"http_code": 404, "error_class": "permanent", "attempts": 1},
{"url": "https://store.example.com/product/789", "status": "failed",
"http_code": 403, "error_class": "retryable", "attempts": 1}
]
After:
[
{"url": "https://store.example.com/product/123", "status": "success",
"http_code": null, "error_class": null, "attempts": 2},
{"url": "https://store.example.com/product/456", "status": "failed",
"http_code": 404, "error_class": "permanent", "attempts": 1},
{"url": "https://store.example.com/product/789", "status": "failed",
"http_code": 404, "error_class": "permanent", "attempts": 2}
]
Read that third entry carefully. The 403 was classified retryable, got its one escalated retry, and came back 404 — the product was delisted and the first pass just happened to catch an interstitial instead of the real answer. The classifier promoted it to permanent. That’s the system working: one retry bought you the truth, and the ledger records why.
Deciding When to Stop: Exit Criteria and Reporting the Residue
Re-drive loops need explicit stop conditions, or they run forever on a shrinking, stubborn set of URLs. Three criteria, any one of which terminates the loop:
loop:
worklist = build_redrive_manifest(ledger)
if worklist is empty: STOP (clean)
if pass_count >= MAX_PASSES (3): STOP (budget)
if time_elapsed > TIME_BUDGET (45 min): STOP (deadline)
if len(worklist) < FLOOR (10 URLs): STOP (diminishing returns)
submit_redrive(worklist)
wait for callback or poll batch status
merge_redrive(ledger, results)
The floor criterion is the one people skip. If 10,000 URLs have shrunk to 7 unresolved entries, another full pass costs more in orchestration time than the data is worth. Below the floor, the residue goes to a report, not back into the queue.
The end-of-batch report is the deliverable that makes this whole pattern defensible to stakeholders. It says exactly what’s missing and why:
BATCH REPORT — store.example.com nightly product sync
Input: 10,000 URLs | Success: 9,995 | Unresolved: 5 | Passes: 2
Unresolved URLs:
1. https://store.example.com/product/456 — permanent (404, 2 attempts)
Action: flag for delisting review; likely removed from catalog
2. https://store.example.com/product/789 — permanent (404 after escalation)
Action: flag for delisting review
3. https://store.example.com/product/1011 — backoff_required (429, 3 attempts)
Action: host still throttling after 3 passes; consider carrier exit IP
or move this URL to the next scheduled window
Every unresolved entry gets a recommended action. “5 URLs failed” is an incident; “5 URLs failed, 2 are delisted products, 1 needs a different exit IP, here’s the plan” is just Tuesday.
Wrap-Up
Partial batch failure is the norm, not the exception, so build for it: a per-URL ledger as the single source of truth, a three-way reconciliation before classification, error classes that separate dead links from throttles, a re-drive queue built only from the misses, backoff with jitter on the retry pass, idempotent merging, and hard stop criteria with a residue report.
The economics alone justify the work — re-driving 412 classified misses instead of 10,000 URLs cuts retry cost by roughly 95%, and the ledger makes every failure explainable after the fact.
Gotchas worth repeating: key the ledger by input URL, not redirect targets; default unmapped errors to permanent, not retryable; treat empty 200 responses as failures, not successes; and always change something on an escalated retry — identical parameters reproduce identical failures.
Next steps if you’re wiring this into a real pipeline: read Async Scraping at Scale: Jobs, Batches, Webhooks for the callback mechanics that trigger the merge step, and Building ETL Pipelines with Scraping APIs for where the ledger sits relative to your transform and load stages.
Related Articles
Stop Cleaning HTML by Hand: Scrape Pages as Markdown
Chunking raw HTML for LLM pipelines wastes tokens and breaks parsers. See how one scrape request field returns clean, ready-to-chunk markdown instead.
TutorialGarbled Scraped Text: Fix Character Encoding Before Parsing
Your scraper returns mojibake instead of product names. Trace the failure from missing charset headers to wrong decode steps and fix it before parsing.
TutorialThroughput Decay Before the 429s: Fix Hidden Throttling
Your batch pace slows long before requests fail. Learn to tell target-site throttling from your own concurrency limits and fix it with adaptive pacing.