What One Successful Page Actually Costs: A Unit Model
Build a per-page cost model for scraping: retries, rendering, proxy traffic, and CAPTCHA solving, with worked numbers you can recompute for your own budget.
On this page
Why Average Cost Per Page Lies: The Case for a Unit Model
Ask an engineering team what a scraped page costs and you’ll usually get division, not analysis. Monthly invoice divided by pages delivered. One number, no structure, and completely useless for making decisions, because it averages across domains with wildly different difficulty, hides retry traffic you’re paying for, and tells you nothing about what happens when you scale up or swap targets.
Here’s a $500/month bill broken down three ways. Same spend, same 250,000 delivered pages:
| Lens | The math | What it tells you |
|---|---|---|
| Naive average | $500 ÷ 250,000 pages | $0.00200 per page. Feels cheap. Tells you nothing. |
| Per-domain average | news.example.com: $120 ÷ 200k = $0.00060; store.example.com: $310 ÷ 40k = $0.00775; forum.example.com: $70 ÷ 10k = $0.00700 | One domain costs 13x another per page. store.example.com eats 62% of the budget for 16% of the pages. |
| Per-successful-page unit | store.example.com: 71,000 attempts produced 40,000 delivered pages; $310 ÷ 40,000 = $0.00775, of which ~$0.00290 is retry and CAPTCHA overhead | The domain isn’t expensive because the pages are big. It’s expensive because most attempts don’t succeed on the first try. |
The third lens is the one that changes decisions. Kill store.example.com and the bill drops 62% — but you only see that if you track attempts, not just deliveries.
The unit model is one formula:
cost_per_success = (base + retries + render + proxy + captcha) / success_rate
Where:
- base — the price of a single fetch attempt, before any add-ons
- retries — the expected number of extra attempts per successful page
- render — the JavaScript-rendering surcharge, applied only to the rendered fraction of your mix
- proxy — transit cost, billed per-request or per-GB, multiplied by total attempts (not successes)
- captcha — probability-weighted solver spend per submitted page
- success_rate — the fraction of submitted pages that actually deliver
Every term is measurable. That’s the whole point. The rest of this post populates each term with numbers, then assembles them.
Counting the Real Fetch: Base Requests and Hidden Retries
Most retry setups are invisible. You mount a retrying adapter on your session, it does its job, and the attempt count exists only inside urllib3’s internals. When someone asks “how many requests did we actually send per delivered page?”, nobody knows.
For measurement, use an explicit ladder so every attempt is observable:
import time
import requests
def fetch_with_ladder(url, max_attempts=3):
attempts = 0
while attempts < max_attempts:
attempts += 1
r = requests.get(url, timeout=15)
print(f"attempt={attempts} status={r.status_code} "
f"bytes={len(r.content)}")
if r.status_code == 200:
return r, attempts
time.sleep(2 ** (attempts - 1)) # 1s, 2s, 4s
return None, attempts
page, n = fetch_with_ladder("https://example.com")
In production you’d use HTTPAdapter with urllib3.util.retry.Retry and a status_forcelist of [429, 500, 502, 503]. I prefer the explicit loop during measurement phases precisely because urllib3’s Retry won’t tell you the count without log forensics. Measure with the ugly version, then switch.
Now the math, and a result that surprises most people. If each attempt on a given endpoint succeeds with probability p, and retries face the same p, then total requests divided by successful pages equals exactly 1/p — regardless of how many retries you allow. Retries don’t make successful pages cheaper. They raise the completion rate.
| Policy | Max attempts | Completion @ p=0.70 | @ p=0.85 | @ p=0.95 | Requests per successful page |
|---|---|---|---|---|---|
| Single attempt | 1 | 70.0% | 85.0% | 95.0% | 1.43 / 1.18 / 1.05 |
| Fixed ladder | 3 | 97.2% | 99.7% | 99.99% | 1.43 / 1.18 / 1.05 |
| Exponential backoff | 5 | 99.76% | ~99.99% | ~100% | 1.43 / 1.18 / 1.05 |
Note that exponential backoff changes timing, not count. It’s politeness, not economics. The count column is identical across all three rows.
The 1/p invariance breaks in practice, and in the bad direction. Real failures correlate: once an exit IP gets flagged, the per-attempt success probability drops on every subsequent attempt from that IP. Your retry ladder then pays full base-plus-proxy cost for attempts that are less likely to succeed than the one that just failed. This is why client-side retry ladders beyond two or three attempts are usually waste — the deeper discussion of where retries belong is in client-side vs service-side retries, but the short version: measure your actual per-attempt p before assuming the table above describes your world.
When JavaScript Rendering Turns One Page into Twenty Requests
A static fetch is one HTTP request. A rendered page is a browser session: the document, then stylesheets, scripts, fonts, XHR calls, tracking beacons, lazy-loaded images. The cost multiplier isn’t the render flag itself — it’s everything the browser does after loading.
Count it directly:
from playwright.sync_api import sync_playwright
class Counter:
n = 0
def bump(self, *args):
self.n += 1
counter = Counter()
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.on("request", counter.bump)
page.goto("https://store.example.com/product/123",
wait_until="networkidle")
print(f"network requests: {counter.n}")
browser.close()
Run this against a real product page and you’ll typically see somewhere between 15 and 60 network requests for what you think of as “one page.” Fill in this worksheet from your own local run — the sample numbers below are from a mid-weight storefront page, yours will differ:
| Metric | Static HTTP GET | Full render |
|---|---|---|
| Network requests | 1 | 23 |
| Bytes transferred | ~150 KB | ~1.9 MB |
| CPU-seconds | ~0 | 3.5 |
| Wall time | 0.4 s | 6.8 s |
| Relative fetch cost | 1x | 12–20x |
The decision rule I use: render only when the data genuinely isn’t in the initial HTML. Check the static response first — many storefronts embed product JSON in a script tag or serve it from an XHR endpoint you can call directly, which collapses the render line item to zero. The trade-off analysis for when rendering is worth it is covered in JS rendering vs plain HTTP; for the unit model, the only thing that matters is the rendered fraction of your mix and the per-render surcharge.
One honest caveat: rendered pages also fail differently. Timeouts mid-render, selector-dependent waits that never resolve, memory pressure on the render fleet. Rendered pages usually have a lower per-attempt p than static ones, which quietly inflates the retry term too. The line items are not independent.
Pricing Proxy Traffic in Bytes, Not Requests
Proxy billing comes in two shapes: per-request and per-GB. They produce wildly different per-page costs depending on page weight, and most teams pick a tier without measuring the bytes.
Measure first:
import requests
def measure(url):
r = requests.get(url, timeout=15)
return {
"url": url,
"content_length_header": r.headers.get("Content-Length"),
"content_encoding": r.headers.get("Content-Encoding", "identity"),
"decompressed_bytes": len(r.content),
}
print(measure("https://example.com"))
print(measure("https://store.example.com/listing"))
If the response is chunked (no Content-Length) and gzipped, read the wire bytes with stream=True and r.raw.read() before requests decompresses anything. The gap between wire bytes and decompressed bytes is often 4–6x on HTML-heavy pages, and per-GB proxy plans bill on the wire side.
Now the tier comparison. Assume 1.43 requests per successful page (the p=0.70 column from earlier) and illustrative market rates — plug in your actual contract numbers:
| Pricing tier | Rate | Cost per 1,000 pages @ 150 KB avg | @ 1.2 MB avg |
|---|---|---|---|
| Per-request | $0.0008/request | 1,430 × $0.0008 = $1.14 | $1.14 (size-independent) |
| Per-GB | $4.00/GB | 214 MB → $0.86 | 1.72 GB → $6.86 |
| Unmetered/bundled | ~$0.75/1,000 pages flat | $0.75 | $0.75 |
The crossover is brutal. At 150 KB average pages, per-GB wins. At 1.2 MB — which is what you get once 30% of your mix is rendered storefront pages — per-GB costs 6x the per-request tier. If your page mix shifts toward rendering over the quarter, your proxy tier decision from last quarter silently inverts. This is exactly the kind of drift the naive average hides.
The CAPTCHA Line Item: Frequency, Solver Pricing, and Failure Loops
CAPTCHA cost is probability-weighted spend: how often pages challenge you, times what a solve costs, adjusted for solver failure.
Worked numbers. Assume a 12% challenge rate on your target domain, $0.002 per solve, 90% solver success, and one re-solve allowed before fallback:
solves per challenged page = 1 + (1 - 0.90) = 1.1
solver spend per challenge = 1.1 × $0.002 = $0.00220
per submitted page = 0.12 × $0.0022 = $0.000264
per successful page (÷0.95) = ≈ $0.000278
Roughly three hundredths of a cent. Sounds negligible. It isn’t, for two reasons.
First, the rate isn’t constant. It climbs as your exit IPs age, and it’s not evenly distributed across your page mix — the hardest 10% of your targets generate most of the challenges. Per-domain CAPTCHA accounting matters as much as per-domain spend accounting.
Second, the failure loop. If a failed solve triggers a full page retry, and the retried page — now from a flagged IP — gets challenged at 40% instead of 12%, the solver fees stay bounded but the retry traffic doesn’t. Each loop re-pays base, proxy, and possibly render. A policy object that caps this belongs in your scraping loop, not in your head:
{
"captcha_policy": {
"max_solve_attempts": 2,
"backoff_seconds": [0, 30],
"fallback_action": "blacklist_ip",
"escalate_to_residential": true
}
}
The fallback_action matters more than max_solve_attempts. Retrying the same page from the same flagged IP is the single most common way a modest CAPTCHA rate turns into a spending spiral. Blacklist the IP for that domain, rotate, and move on. The strategic question of when to avoid challenges instead of paying to solve them deserves its own treatment — see avoid vs solve CAPTCHA strategy — but the unit model only needs the expected-spend number.
Assembling the Full Unit Model: A Worked Example at Three Scales
Every term is now measurable, so put them together:
def cost_per_success(p, base_per_request, proxy_per_request,
render_fraction=0.0, render_cost=0.0,
captcha_rate=0.0, solve_price=0.0,
solve_success=1.0):
attempts = 1.0 / p # requests per successful page
fetch = (base_per_request + proxy_per_request) * attempts
render = render_fraction * render_cost * attempts
captcha = captcha_rate * solve_price \
* (1 + (1 - solve_success)) / p
return fetch + render + captcha
# 1K pages: static content, clean datacenter IPs
print(cost_per_success(p=0.95, base_per_request=0.001,
proxy_per_request=0.0008))
# -> 0.0019
# 100K pages: 30% rendered, light CAPTCHA exposure
print(cost_per_success(p=0.85, base_per_request=0.001,
proxy_per_request=0.0008,
render_fraction=0.3, render_cost=0.005,
captcha_rate=0.05, solve_price=0.002,
solve_success=0.9))
# -> 0.0040
# 10M pages: hard targets, residential proxies, 60% rendered
print(cost_per_success(p=0.75, base_per_request=0.001,
proxy_per_request=0.003,
render_fraction=0.6, render_cost=0.005,
captcha_rate=0.12, solve_price=0.002,
solve_success=0.9))
# -> 0.0097
If you’re using a scraping API, the line items map directly to request flags. A request like this carries an explicit, countable price in token adders — +5 for rendering, +3 for residential, +10 for CAPTCHA solving — which means the unit model isn’t an estimate, it’s arithmetic on your own payload:
import requests
resp = requests.post(
"https://api.finedata.ai/api/v1/scrape",
headers={"Authorization": "Bearer fd_your_api_key"},
json={
"url": "https://store.example.com/product/123",
"use_js_render": True, # +5 tokens
"use_residential": True, # +3 tokens
"solve_captcha": True, # +10 tokens
"max_retries": 3,
"formats": ["markdown"],
},
)
Now the scale table:
| Scale | Fetch (base+proxy) | Render | CAPTCHA | Success rate | Cost per successful page | Monthly variable |
|---|---|---|---|---|---|---|
| 1K pages | $0.0019 | — | — | 95% | $0.0019 | $1.90 (fixed infra dominates) |
| 100K pages | $0.0021 | $0.0018 | $0.0001 | 85% | $0.0040 | $401 |
| 10M pages | $0.0053 | $0.0040 | $0.0004 | 75% | $0.0097 | $96,900 |
Here’s the claim you might disagree with: in scraping, per-page cost usually rises with scale, it doesn’t fall. Normal software economics give you volume discounts. Scraping economics give you the opposite — at 10M pages/month you’ve exhausted clean datacenter IPs, your hardest domains dominate the mix, your CAPTCHA rate is up, and your success rate is down. The only genuinely sublinear component is fixed cost: infrastructure and engineering amortize to nothing. Everything variable gets worse per unit. Anyone forecasting 10M-page costs by multiplying their 100K-page costs by 100 is going to undershoot the budget by 2x or more. Plan for the per-page number to climb, and treat any forecast where it declines as a red flag demanding justification.
Recomputing the Model for Your Own Budget: A Measurement Checklist
The worked numbers above are placeholders. Yours come from five local measurements, each taking under an hour:
- Baseline fetch — hit
example.comand your lightest real target. Record status, bytes, duration. This calibrates your measurement harness itself. - Retry rate — run 100 fetches against your primary target domain with the explicit ladder from earlier. The fraction of first-attempt successes is your p. Don’t guess this; it drives every other term.
- Render multiplier — fetch one
store.example.comproduct page statically and via full render. Fill in the worksheet: request count, bytes, CPU-seconds, wall time. If the data exists in the static response, kill the render flag for that domain. - Byte weight — measure wire bytes on your heaviest page type, not your average one. Per-GB proxy tiers are decided by the tail.
- Challenge rate — over the same 100-fetch sample, record the fraction that returned a challenge page instead of content. That’s your CAPTCHA rate per domain.
Then make the measurements continuous. One JSON line per fetch attempt, piped into whatever log aggregation you already run, gives you a live feed of every unit-model input:
import json
import time
import requests
def fetch_logged(url, session, attempt=1):
t0 = time.monotonic()
r = session.get(url, timeout=20)
print(json.dumps({
"url": url,
"attempt": attempt,
"status": r.status_code,
"bytes": len(r.content),
"duration_ms": round((time.monotonic() - t0) * 1000),
"captcha_solved": False,
}))
return r
A week of these lines and you can compute per-domain p, byte distributions, and challenge rates with actual confidence intervals instead of vibes. Drop those into cost_per_success() and the output is your budget, not an analogy of your budget.
Wrap-Up
The naive average — invoice divided by delivered pages — is a number that can’t inform any decision. The unit model decomposes it into five measurable terms: base fetch, retry amplification at 1/p, render surcharge on the rendered fraction, proxy transit in bytes or requests, and probability-weighted CAPTCHA spend.
Three findings worth carrying forward. Retries buy completion rate, not cheaper pages — requests per success is fixed at 1/p while failures correlate in practice, so deep retry ladders are usually waste. Proxy tier choice inverts as your page mix gets heavier, so re-run the byte measurement whenever your rendered fraction changes. And per-page cost rises with scale in this business, because scale pushes you onto harder targets and dirtier IPs — forecast accordingly.
Measure the five inputs, recompute quarterly, and question any plan where the per-page number is assumed to shrink.
Related Articles
Forecast Scrape Spend: A Budget Model You Can Recompute
Turn success rates, retry multipliers, and rendering overhead into a scraping budget you can defend, with formulas and worked numbers you can recompute.
TechnicalOn-Demand Scraping vs Prefetched Data: Serving Trade-offs
Latency, freshness, and cost per served result: when to scrape in the request path versus harvest ahead into storage, and what failed fetches cost.
TechnicalAvoid CAPTCHAs vs Solve Them: Strategy Trade-offs
Two ways to deal with CAPTCHAs at scale: engineer them away with stealth or pay to solve each one. An honest comparison of cost, latency, and success rates.