Skip to content
Technical 12 min read

What One Successful Page Actually Costs: A Unit Model

Build a per-page cost model for scraping: retries, rendering, proxy traffic, and CAPTCHA solving, with worked numbers you can recompute for your own budget.

FE
FineData Engineering · Editorial Policy
|
On this page

Why Average Cost Per Page Lies: The Case for a Unit Model

Ask an engineering team what a scraped page costs and you’ll usually get division, not analysis. Monthly invoice divided by pages delivered. One number, no structure, and completely useless for making decisions, because it averages across domains with wildly different difficulty, hides retry traffic you’re paying for, and tells you nothing about what happens when you scale up or swap targets.

Here’s a $500/month bill broken down three ways. Same spend, same 250,000 delivered pages:

LensThe mathWhat it tells you
Naive average$500 ÷ 250,000 pages$0.00200 per page. Feels cheap. Tells you nothing.
Per-domain averagenews.example.com: $120 ÷ 200k = $0.00060; store.example.com: $310 ÷ 40k = $0.00775; forum.example.com: $70 ÷ 10k = $0.00700One domain costs 13x another per page. store.example.com eats 62% of the budget for 16% of the pages.
Per-successful-page unitstore.example.com: 71,000 attempts produced 40,000 delivered pages; $310 ÷ 40,000 = $0.00775, of which ~$0.00290 is retry and CAPTCHA overheadThe domain isn’t expensive because the pages are big. It’s expensive because most attempts don’t succeed on the first try.

The third lens is the one that changes decisions. Kill store.example.com and the bill drops 62% — but you only see that if you track attempts, not just deliveries.

The unit model is one formula:

cost_per_success = (base + retries + render + proxy + captcha) / success_rate

Where:

  • base — the price of a single fetch attempt, before any add-ons
  • retries — the expected number of extra attempts per successful page
  • render — the JavaScript-rendering surcharge, applied only to the rendered fraction of your mix
  • proxy — transit cost, billed per-request or per-GB, multiplied by total attempts (not successes)
  • captcha — probability-weighted solver spend per submitted page
  • success_rate — the fraction of submitted pages that actually deliver

Every term is measurable. That’s the whole point. The rest of this post populates each term with numbers, then assembles them.

Counting the Real Fetch: Base Requests and Hidden Retries

Most retry setups are invisible. You mount a retrying adapter on your session, it does its job, and the attempt count exists only inside urllib3’s internals. When someone asks “how many requests did we actually send per delivered page?”, nobody knows.

For measurement, use an explicit ladder so every attempt is observable:

import time
import requests

def fetch_with_ladder(url, max_attempts=3):
    attempts = 0
    while attempts < max_attempts:
        attempts += 1
        r = requests.get(url, timeout=15)
        print(f"attempt={attempts} status={r.status_code} "
              f"bytes={len(r.content)}")
        if r.status_code == 200:
            return r, attempts
        time.sleep(2 ** (attempts - 1))  # 1s, 2s, 4s
    return None, attempts

page, n = fetch_with_ladder("https://example.com")

In production you’d use HTTPAdapter with urllib3.util.retry.Retry and a status_forcelist of [429, 500, 502, 503]. I prefer the explicit loop during measurement phases precisely because urllib3’s Retry won’t tell you the count without log forensics. Measure with the ugly version, then switch.

Now the math, and a result that surprises most people. If each attempt on a given endpoint succeeds with probability p, and retries face the same p, then total requests divided by successful pages equals exactly 1/p — regardless of how many retries you allow. Retries don’t make successful pages cheaper. They raise the completion rate.

PolicyMax attemptsCompletion @ p=0.70@ p=0.85@ p=0.95Requests per successful page
Single attempt170.0%85.0%95.0%1.43 / 1.18 / 1.05
Fixed ladder397.2%99.7%99.99%1.43 / 1.18 / 1.05
Exponential backoff599.76%~99.99%~100%1.43 / 1.18 / 1.05

Note that exponential backoff changes timing, not count. It’s politeness, not economics. The count column is identical across all three rows.

The 1/p invariance breaks in practice, and in the bad direction. Real failures correlate: once an exit IP gets flagged, the per-attempt success probability drops on every subsequent attempt from that IP. Your retry ladder then pays full base-plus-proxy cost for attempts that are less likely to succeed than the one that just failed. This is why client-side retry ladders beyond two or three attempts are usually waste — the deeper discussion of where retries belong is in client-side vs service-side retries, but the short version: measure your actual per-attempt p before assuming the table above describes your world.

When JavaScript Rendering Turns One Page into Twenty Requests

A static fetch is one HTTP request. A rendered page is a browser session: the document, then stylesheets, scripts, fonts, XHR calls, tracking beacons, lazy-loaded images. The cost multiplier isn’t the render flag itself — it’s everything the browser does after loading.

Count it directly:

from playwright.sync_api import sync_playwright

class Counter:
    n = 0
    def bump(self, *args):
        self.n += 1

counter = Counter()

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.on("request", counter.bump)
    page.goto("https://store.example.com/product/123",
              wait_until="networkidle")
    print(f"network requests: {counter.n}")
    browser.close()

Run this against a real product page and you’ll typically see somewhere between 15 and 60 network requests for what you think of as “one page.” Fill in this worksheet from your own local run — the sample numbers below are from a mid-weight storefront page, yours will differ:

MetricStatic HTTP GETFull render
Network requests123
Bytes transferred~150 KB~1.9 MB
CPU-seconds~03.5
Wall time0.4 s6.8 s
Relative fetch cost1x12–20x

The decision rule I use: render only when the data genuinely isn’t in the initial HTML. Check the static response first — many storefronts embed product JSON in a script tag or serve it from an XHR endpoint you can call directly, which collapses the render line item to zero. The trade-off analysis for when rendering is worth it is covered in JS rendering vs plain HTTP; for the unit model, the only thing that matters is the rendered fraction of your mix and the per-render surcharge.

One honest caveat: rendered pages also fail differently. Timeouts mid-render, selector-dependent waits that never resolve, memory pressure on the render fleet. Rendered pages usually have a lower per-attempt p than static ones, which quietly inflates the retry term too. The line items are not independent.

Pricing Proxy Traffic in Bytes, Not Requests

Proxy billing comes in two shapes: per-request and per-GB. They produce wildly different per-page costs depending on page weight, and most teams pick a tier without measuring the bytes.

Measure first:

import requests

def measure(url):
    r = requests.get(url, timeout=15)
    return {
        "url": url,
        "content_length_header": r.headers.get("Content-Length"),
        "content_encoding": r.headers.get("Content-Encoding", "identity"),
        "decompressed_bytes": len(r.content),
    }

print(measure("https://example.com"))
print(measure("https://store.example.com/listing"))

If the response is chunked (no Content-Length) and gzipped, read the wire bytes with stream=True and r.raw.read() before requests decompresses anything. The gap between wire bytes and decompressed bytes is often 4–6x on HTML-heavy pages, and per-GB proxy plans bill on the wire side.

Now the tier comparison. Assume 1.43 requests per successful page (the p=0.70 column from earlier) and illustrative market rates — plug in your actual contract numbers:

Pricing tierRateCost per 1,000 pages @ 150 KB avg@ 1.2 MB avg
Per-request$0.0008/request1,430 × $0.0008 = $1.14$1.14 (size-independent)
Per-GB$4.00/GB214 MB → $0.861.72 GB → $6.86
Unmetered/bundled~$0.75/1,000 pages flat$0.75$0.75

The crossover is brutal. At 150 KB average pages, per-GB wins. At 1.2 MB — which is what you get once 30% of your mix is rendered storefront pages — per-GB costs 6x the per-request tier. If your page mix shifts toward rendering over the quarter, your proxy tier decision from last quarter silently inverts. This is exactly the kind of drift the naive average hides.

The CAPTCHA Line Item: Frequency, Solver Pricing, and Failure Loops

CAPTCHA cost is probability-weighted spend: how often pages challenge you, times what a solve costs, adjusted for solver failure.

Worked numbers. Assume a 12% challenge rate on your target domain, $0.002 per solve, 90% solver success, and one re-solve allowed before fallback:

solves per challenged page   = 1 + (1 - 0.90) = 1.1
solver spend per challenge   = 1.1 × $0.002   = $0.00220
per submitted page           = 0.12 × $0.0022 = $0.000264
per successful page (÷0.95)  =                 ≈ $0.000278

Roughly three hundredths of a cent. Sounds negligible. It isn’t, for two reasons.

First, the rate isn’t constant. It climbs as your exit IPs age, and it’s not evenly distributed across your page mix — the hardest 10% of your targets generate most of the challenges. Per-domain CAPTCHA accounting matters as much as per-domain spend accounting.

Second, the failure loop. If a failed solve triggers a full page retry, and the retried page — now from a flagged IP — gets challenged at 40% instead of 12%, the solver fees stay bounded but the retry traffic doesn’t. Each loop re-pays base, proxy, and possibly render. A policy object that caps this belongs in your scraping loop, not in your head:

{
  "captcha_policy": {
    "max_solve_attempts": 2,
    "backoff_seconds": [0, 30],
    "fallback_action": "blacklist_ip",
    "escalate_to_residential": true
  }
}

The fallback_action matters more than max_solve_attempts. Retrying the same page from the same flagged IP is the single most common way a modest CAPTCHA rate turns into a spending spiral. Blacklist the IP for that domain, rotate, and move on. The strategic question of when to avoid challenges instead of paying to solve them deserves its own treatment — see avoid vs solve CAPTCHA strategy — but the unit model only needs the expected-spend number.

Assembling the Full Unit Model: A Worked Example at Three Scales

Every term is now measurable, so put them together:

def cost_per_success(p, base_per_request, proxy_per_request,
                     render_fraction=0.0, render_cost=0.0,
                     captcha_rate=0.0, solve_price=0.0,
                     solve_success=1.0):
    attempts = 1.0 / p  # requests per successful page
    fetch = (base_per_request + proxy_per_request) * attempts
    render = render_fraction * render_cost * attempts
    captcha = captcha_rate * solve_price \
        * (1 + (1 - solve_success)) / p
    return fetch + render + captcha

# 1K pages: static content, clean datacenter IPs
print(cost_per_success(p=0.95, base_per_request=0.001,
                       proxy_per_request=0.0008))
# -> 0.0019

# 100K pages: 30% rendered, light CAPTCHA exposure
print(cost_per_success(p=0.85, base_per_request=0.001,
                       proxy_per_request=0.0008,
                       render_fraction=0.3, render_cost=0.005,
                       captcha_rate=0.05, solve_price=0.002,
                       solve_success=0.9))
# -> 0.0040

# 10M pages: hard targets, residential proxies, 60% rendered
print(cost_per_success(p=0.75, base_per_request=0.001,
                       proxy_per_request=0.003,
                       render_fraction=0.6, render_cost=0.005,
                       captcha_rate=0.12, solve_price=0.002,
                       solve_success=0.9))
# -> 0.0097

If you’re using a scraping API, the line items map directly to request flags. A request like this carries an explicit, countable price in token adders — +5 for rendering, +3 for residential, +10 for CAPTCHA solving — which means the unit model isn’t an estimate, it’s arithmetic on your own payload:

import requests

resp = requests.post(
    "https://api.finedata.ai/api/v1/scrape",
    headers={"Authorization": "Bearer fd_your_api_key"},
    json={
        "url": "https://store.example.com/product/123",
        "use_js_render": True,    # +5 tokens
        "use_residential": True,  # +3 tokens
        "solve_captcha": True,    # +10 tokens
        "max_retries": 3,
        "formats": ["markdown"],
    },
)

Now the scale table:

ScaleFetch (base+proxy)RenderCAPTCHASuccess rateCost per successful pageMonthly variable
1K pages$0.0019——95%$0.0019$1.90 (fixed infra dominates)
100K pages$0.0021$0.0018$0.000185%$0.0040$401
10M pages$0.0053$0.0040$0.000475%$0.0097$96,900

Here’s the claim you might disagree with: in scraping, per-page cost usually rises with scale, it doesn’t fall. Normal software economics give you volume discounts. Scraping economics give you the opposite — at 10M pages/month you’ve exhausted clean datacenter IPs, your hardest domains dominate the mix, your CAPTCHA rate is up, and your success rate is down. The only genuinely sublinear component is fixed cost: infrastructure and engineering amortize to nothing. Everything variable gets worse per unit. Anyone forecasting 10M-page costs by multiplying their 100K-page costs by 100 is going to undershoot the budget by 2x or more. Plan for the per-page number to climb, and treat any forecast where it declines as a red flag demanding justification.

Recomputing the Model for Your Own Budget: A Measurement Checklist

The worked numbers above are placeholders. Yours come from five local measurements, each taking under an hour:

  1. Baseline fetch — hit example.com and your lightest real target. Record status, bytes, duration. This calibrates your measurement harness itself.
  2. Retry rate — run 100 fetches against your primary target domain with the explicit ladder from earlier. The fraction of first-attempt successes is your p. Don’t guess this; it drives every other term.
  3. Render multiplier — fetch one store.example.com product page statically and via full render. Fill in the worksheet: request count, bytes, CPU-seconds, wall time. If the data exists in the static response, kill the render flag for that domain.
  4. Byte weight — measure wire bytes on your heaviest page type, not your average one. Per-GB proxy tiers are decided by the tail.
  5. Challenge rate — over the same 100-fetch sample, record the fraction that returned a challenge page instead of content. That’s your CAPTCHA rate per domain.

Then make the measurements continuous. One JSON line per fetch attempt, piped into whatever log aggregation you already run, gives you a live feed of every unit-model input:

import json
import time
import requests

def fetch_logged(url, session, attempt=1):
    t0 = time.monotonic()
    r = session.get(url, timeout=20)
    print(json.dumps({
        "url": url,
        "attempt": attempt,
        "status": r.status_code,
        "bytes": len(r.content),
        "duration_ms": round((time.monotonic() - t0) * 1000),
        "captcha_solved": False,
    }))
    return r

A week of these lines and you can compute per-domain p, byte distributions, and challenge rates with actual confidence intervals instead of vibes. Drop those into cost_per_success() and the output is your budget, not an analogy of your budget.

Wrap-Up

The naive average — invoice divided by delivered pages — is a number that can’t inform any decision. The unit model decomposes it into five measurable terms: base fetch, retry amplification at 1/p, render surcharge on the rendered fraction, proxy transit in bytes or requests, and probability-weighted CAPTCHA spend.

Three findings worth carrying forward. Retries buy completion rate, not cheaper pages — requests per success is fixed at 1/p while failures correlate in practice, so deep retry ladders are usually waste. Proxy tier choice inverts as your page mix gets heavier, so re-run the byte measurement whenever your rendered fraction changes. And per-page cost rises with scale in this business, because scale pushes you onto harder targets and dirtier IPs — forecast accordingly.

Measure the five inputs, recompute quarterly, and question any plan where the per-page number is assumed to shrink.

#cost modeling #unit economics #scraping budget #retry cost #web scraping #slot:cost-accounting

Related Articles