Skip to content
Technical 15 min read

Forecast Scrape Spend: A Budget Model You Can Recompute

Turn success rates, retry multipliers, and rendering overhead into a scraping budget you can defend, with formulas and worked numbers you can recompute.

FE
FineData Engineering · Editorial Policy
|
On this page

Why Raw Page Counts Understate Scrape Spend: The Full Cost Stack

Every scraping budget conversation I have ever sat in starts the same way: someone multiplies page count by per-request price and calls it the forecast. Ten thousand pages at eight-tenths of a cent per page, eighty dollars a month, done. That number survives exactly one contact with production traffic.

The problem is that “pages” is not the unit you pay for. You pay for attempts. Between the page you want and the invoice there are three wedges: some attempts fail and never produce a page, some pages need multiple attempts before one succeeds, and some pages cost five to twenty times more than others because they need a browser engine and a better proxy to load at all. A budget that ignores any of the three will be wrong, and it will be wrong in the expensive direction.

Start with the naive stack, line by line. Assume a target of 10,000 successful pages per month from store.example.com, a plain-fetch cost of $0.008 per request, a measured success rate of 85%, a hand-wavy retry multiplier of 1.4, and a 40% share of pages that need JavaScript rendering (which raises the average request cost to $0.016):

LineAdjustmentEffective requestsCost @ plain rateDelta vs. previous line
Nominalpages × price10,000$80.00—
Success rate 85%pages ÷ 0.8511,765$94.12+$14.12
Retry multiplier 1.4× 1.416,470$131.76+$37.64
40% render shareavg cost × 216,470$263.52+$131.76

The nominal estimate was off by 3.3x. And notice something uncomfortable: the biggest single line item is not retries or failures — it is rendering share. Most teams argue for hours about retry policy and never once ask what fraction of their corpus actually needs a browser.

The core formula for the first two wedges:

effective_requests = pages / success_rate * retry_multiplier
                    = 10,000 / 0.85 * 1.4
                    = 16,470

I will be honest about that 1.4: it is a lazy number, and if you apply it on top of a success-rate division the way the table does, you double-count. Dividing by 0.85 already assumes you re-drive every page until it succeeds; multiplying by 1.4 on top of that charges for retries twice. Later in this post I will derive the multiplier properly from the retry policy itself. For now, the point stands: even the sloppy stack lands closer to reality than pages-times-price, and the disciplined version lands closer still.

Measuring Your Real Success Rate from Local Logs

The 85% figure above cannot come from a vendor dashboard average. It has to come from your own traffic against your own target domains, because success rate is a property of the pair (your exit IPs, their anti-bot configuration), not of the scraping service in the abstract. A rate measured on easy blog pages tells you nothing about what you will pay on store.example.com.

The right source is a local response log. Every scrape you run should append one line: timestamp, URL, HTTP status, latency, and a content sanity check. Here is a parser that computes success rate per status bucket from a JSON-lines log:

import json
from collections import Counter

# log line format:
# {"ts": "...", "url": "https://store.example.com/p/123",
#  "status": 200, "latency_ms": 812, "body_bytes": 48213}

RETRYABLE = {429, 500, 502, 503, 504, "timeout"}
DEAD = {401, 403, 404}

def classify(rec):
    if rec["status"] == 200:
        # soft failure: 200 with a near-empty body is usually a
        # challenge shell or an unrendered SPA skeleton
        if rec["body_bytes"] < 2048:
            return "soft_fail"
        return "success"
    if rec["status"] in DEAD:
        return "dead"
    if rec["status"] in RETRYABLE or rec["status"] == "timeout":
        return "retryable"
    return "unknown"

buckets = Counter()
total = 0
with open("scrape_log.jsonl") as f:
    for line in f:
        rec = json.loads(line)
        if "store.example.com" not in rec["url"]:
            continue
        buckets[classify(rec)] += 1
        total += 1

success_rate = buckets["success"] / total
print(f"total={total}  success={success_rate:.3f}")
print(dict(buckets))

The classification rules matter more than the parser. Here is the bucket table I use:

BucketExamplesCounts asRationale
200, sane bodystatus 200, body > 2 KB, expected selector presentSuccessThe only bucket that produces a billable page
200, empty/challenge bodystatus 200, body < 2 KB, no expected contentSoft failure, retryableAnti-bot interstitials and SPA shells return 200
403 / 401status 403DeadRetrying with the same egress rarely helps; needs a tier change, not a retry
429status 429Retryable, with backoffRate limit — retrying immediately makes it worse
5xxstatus 500–504RetryableUsually transient origin-side
Timeoutno statusRetryableOften resolves on retry with a longer budget

Two opinions baked into that table, and you may disagree with both. First, I treat soft failures as retryable rather than dead, because a 200-with-empty-shell is the single most common misclassification I see — teams count those as successes, their parsers return None, and the failure only surfaces downstream. Second, I refuse to retry 403s. A 403 means the target recognized the client; burning four more attempts at the same tier is paying full price to get rejected five times. The fix for 403s is a different proxy tier or fingerprint, which is a cost-model decision, not a retry-policy one. If you want the measurement methodology in more depth, Measure Your Scraper’s Real Success Rate Locally covers the harness side.

Retry Multipliers: Pricing the Cost of Eventual Success

Once you have a per-attempt success probability p from the logs, the retry policy converts directly into a multiplier. Under a capped policy with at most N attempts per page, the expected number of attempts per started page is a truncated geometric series:

E[attempts] = 1 + q + q^2 + ... + q^(N-1)
             = (1 - q^N) / (1 - q)
where q = 1 - p

With p = 0.85, q = 0.15, and a cap of N = 3:

E[attempts] = (1 - 0.15^3) / 0.85
            = 0.996625 / 0.85
            = 1.1725

So each page you start consumes 1.1725 attempts on average, and the probability a started page eventually succeeds is 1 - q^N = 0.9966. To get 10,000 successful pages you therefore start 10,000 / 0.9966 = 10,034 pages and burn 10,034 × 1.1725 = 11,765 attempts. The attempts-per-successful-page figure is 1.1725 / 0.9966 = 1.1765.

Compare that to the hand-wavy 1.4 from the first section. The derived multiplier is 1.18, not 1.4 — the naive stack overstated retry cost by roughly 19%. That is the difference between a number you defend in a meeting and a number you hope nobody checks.

Here is the policy as config, with the multiplier-relevant fields annotated:

# retry policy for store.example.com
target: store.example.com
max_retries: 3          # N in the geometric series — caps the multiplier
backoff:
  base_seconds: 2      # does NOT affect cost, only wall-clock
  max_seconds: 60
retry_on:              # only these feed q; dead statuses excluded
  - 429
  - 5xx
  - timeout
  - soft_fail          # 200 with empty body
escalate_after: 2      # switch tier before attempt 3 — a cost event,
                       # not a retry; budget it separately

One subtlety: q in the formula should be the retryable failure rate, not the total failure rate. If 15% of attempts fail but a third of those are 403s you never retry, your effective q is closer to 0.10 and the multiplier drops to about 1.11. Excluding dead statuses from the retry budget is free money — it is the most common over-forecast I see.

The distinction between client-side and service-side retries also matters here. If the scraping API retries internally (the sync endpoint’s auto_retry and max_retries params do exactly this), you pay for the internal attempts but never see them in your logs, so your measured q will look better than your true per-attempt quality. Either disable internal retries for measurement periods or accept that your multiplier is already baked into the observed success rate. Client-Side vs Service-Side Retries for Scraping Failures dissects that split; for budgeting, the rule is: never let both layers retry, because you will pay for a Cartesian product of attempts.

Rendering Overhead: When a Page Costs 10x a Plain Fetch

The third wedge is tier mix, and it is the one that quietly dominates. On a token-metered service, a plain fetch with the default browser-like TLS fingerprint costs 2 tokens. Add residential proxy routing and it is 5. Add headless rendering and it is 7. But the expensive pages are not the average rendered page — they are the hard ones, where you stack rendering with a stealth engine: full fingerprint emulation at 20 tokens, headful premium mode at 30, plus the rendering and proxy tokens on top. A fully loaded headful premium render with residential egress runs around 40 tokens against 2 for a plain fetch. That is 20x per request, before you count wall-clock.

Wall-clock matters because it eats throughput. Measure it locally before you believe any tier decision. A quick harness: 50 plain fetches versus 50 headless renders of the same product page, https://example.com/product/123, recording medians and p95:

ModenMedian latencyp95 latency
Plain HTTP fetch500.9 s2.1 s
Headless render (use_js_render)506.8 s14.2 s

A render takes 7.5x as long as a fetch at the median. If you run a fixed pool of workers, your pages-per-hour ceiling drops by roughly the same factor, which means either more concurrency or missed SLAs — both of which have costs that never show up in a per-token invoice.

Now the per-1,000-requests view, assuming $0.004 per token:

TierTokens/requestCost per 1,000 requestsMultiplier vs. plain
Plain fetch (datacenter, antibot TLS)2$8.001.0x
Residential proxy fetch5$20.002.5x
Headless render (datacenter)7$28.003.5x
Headful premium render + residential~40$160.0020x

The uncomfortable takeaway: the tier ladder is steeper than the retry ladder. Moving your success rate from 85% to 95% saves you about 12% on attempts. Moving 40% of your corpus from plain fetch to rendering doubles your average request cost. If you want to trim spend, audit the render share before you touch anything else — I have seen pipelines render every page because some pages need it, which is like shipping every package overnight because some packages are urgent. The per-request economics are laid out in more detail in JS Rendering vs Plain HTTP: The Real Cost of Every Scrape.

Assembling the Budget Model: One Formula You Can Recompute Monthly

Everything collapses into one line:

monthly_spend = (pages / success_rate)
              * retry_multiplier
              * per_request_cost
              * tier_multiplier

Where each input has a precise definition:

  • pages — successful pages required per month (the business number)
  • success_rate — eventual success rate per started page, from the log-derived per-attempt rate and the retry cap: 1 - q^N
  • retry_multiplier — attempts per started page: (1 - q^N) / (1 - q)
  • per_request_cost — your blended cost per attempt, computed from the tier mix
  • tier_multiplier — 1.0 if you keep the blend inside per_request_cost; use it only when modeling a tier change as a scenario

Plugging in the numbers derived so far: 10,000 pages, eventual success 0.9966, multiplier 1.1725, and a tier mix of 60% plain fetch (2 tokens) and 40% render (7 tokens) — a blended 4 tokens, $0.016 per attempt:

monthly_spend = (10,000 / 0.9966) * 1.1725 * $0.016
              =  10,034 * 1.1725 * $0.016
              =  11,765 * $0.016
              =  $188.24

The naive pages-times-price estimate was $80. The defensible number is $188 — 2.35x, and every dollar of the gap is traceable to a named input.

As code, so finance can rerun it without you:

def forecast_monthly_spend(pages, success_rate, retry_multiplier,
                           per_request_cost):
    """All inputs measured or quoted; see inputs.json for provenance."""
    started = pages / success_rate
    attempts = started * retry_multiplier
    return attempts * per_request_cost

# tier blend: 60% plain fetch @ 2 tokens, 40% render @ 7 tokens,
# token price $0.004
blend = 0.6 * 2 * 0.004 + 0.4 * 7 * 0.004   # $0.016 per attempt

spend = forecast_monthly_spend(
    pages=10_000,
    success_rate=0.9966,     # 1 - q^3, from logs
    retry_multiplier=1.1725, # (1 - q^3) / (1 - q), from retry policy
    per_request_cost=blend,
)
print(f"attempts={int(10_000 / 0.9966 * 1.1725)}, spend=${spend:.2f}")
# attempts=11765, spend=$188.24

The function is deliberately four arguments. If you find yourself adding a fifth and sixth — “fudge factor,” “buffer” — stop. A buffer belongs in the presentation of the number, not in the model. Buffers hidden inside a formula are how forecasts drift 40% from reality with no explanation anyone can produce.

Sensitivity Analysis: Which Input Moves the Budget Most

A forecast is only as good as your knowledge of which assumption it leans on. Hold the retry multiplier at 1.18 per successful page and the token price at $0.004, vary success rate and render share, and read the monthly spend for the 10,000-page target:

Render share 0%Render share 50%Render share 100%
Success 95%$99$224$348
Success 85%$111$250$389
Success 70%$135$303$472

Read the table along both axes. Dropping success from 95% to 70% — a 26-point collapse — raises spend 27–36%. Moving render share from 0% to 100% raises it 250%. The model is three times more fragile to tier mix than to success rate. If you have one engineering hour to spend, spend it on classifying which URLs genuinely need rendering, not on squeezing another point of success rate. This is the non-obvious conclusion, and half the teams I present it to push back because failure rates feel like the thing that determines cost. The arithmetic says otherwise.

The break-even question is the sharper version: when does a more expensive tier pay for itself? Compare a residential proxy fetch (5 tokens) against a headless render (7 tokens) on store.example.com. Cost per successful page is tokens / success, so rendering breaks even when:

7 / s_render = 5 / s_residential
s_render     = 1.4 * s_residential

A 1.4x success improvement is a tall order when residential fetches already succeed at 85% — you would need 119% success, which is impossible. Rendering never pays off against a working cheaper tier. But flip the scenario: on a JavaScript-heavy catalog where plain fetches return 200-with-empty-shell two-thirds of the time, effective plain-fetch success is maybe 0.25. Now the break-even render success is 3.5 × 0.25 = 0.875, and a rendered success rate of 0.95 clears it comfortably. The rule: never upgrade tiers to fix a success rate you have not measured, and always upgrade when the cheaper tier’s effective success is below the token-ratio break-even. The measurement, not the instinct, decides.

Defending the Number: Presenting the Model to Finance and Engineering

A model nobody can audit is a guess with formatting. The fix is to ship the inputs as a versioned artifact where every parameter carries its provenance and its verification recency:

{
  "model_version": "3",
  "target": "store.example.com",
  "pages_per_month": 10000,
  "inputs": {
    "per_attempt_success": {
      "value": 0.85,
      "source": "log-derived",
      "evidence": "scrape_log.jsonl, 4,120 requests, bucket table v2",
      "verified": "12 days ago"
    },
    "retry_cap": {
      "value": 3,
      "source": "config",
      "evidence": "retry policy YAML, hash 9f2c1",
      "verified": "12 days ago"
    },
    "render_share": {
      "value": 0.4,
      "source": "measured",
      "evidence": "URL classification audit, 500-URL sample",
      "verified": "5 weeks ago"
    },
    "token_price": {
      "value": 0.004,
      "source": "vendor quote",
      "evidence": "plan pricing page, screenshot archived",
      "verified": "3 weeks ago"
    }
  }
}

Three source types, three trust levels. Log-derived numbers are the strongest — they are your traffic. Vendor quotes drift. Measured shares rot as the target site changes. The verified field exists so that when the forecast misses, the first question is not “whose fault” but “which input went stale.”

The monthly recompute cadence turns variance from an argument into a one-liner. Worked example: last month the model forecast $188. Actual came in at $208. Recompute with the fresh log-derived success rate — it fell from 0.85 to 0.78 per attempt, which drags eventual success to 0.9966… no, it drags it to 1 - 0.22^3 = 0.9894, and the multiplier to (1 - 0.0106) / 0.78 = 1.273. New attempts: 10,000 / 0.9894 × 1.273 = 12,873. At the $0.016 blend that is $206 — the variance is explained to within $2. The one-line justification for the deck: “429 share on the target rose from 6% to 11% after they tightened rate limits; success rate 0.85 → 0.78; recomputed delta +$18, residual $2 unexplained.”

That is what a defensible budget looks like. Not a smaller variance — an attributable one.

Wrap-Up

Pages-times-price is off by a factor of two to three on any non-trivial corpus, and the gap has three named causes: success rate divides your yield, retry policy multiplies your attempts, and tier mix multiplies your per-attempt cost. The full model is one formula — (pages / success_rate) × retry_multiplier × per_request_cost × tier_multiplier — where every input is measured, not guessed: success rate from your own response logs with soft failures classified honestly, the retry multiplier derived from the truncated geometric series of your actual retry cap, and the tier blend from a URL audit that separates pages that need a browser from pages that do not.

Two findings worth carrying out of this even if you skip the math. First, tier mix moves the budget roughly three times harder than success rate does, so audit your render share before touching anything else. Second, derive the retry multiplier from the policy — the hand-wavy 1.4 overstated cost by 19% against the computed 1.18, and overstated budgets get cut, and cuts get made against the wrong line item.

Recompute monthly, version the inputs, stamp every parameter with provenance and a verification date. When the number is challenged — and it will be — you want the answer to be a re-run, not a shrug.

#cost modeling #budget forecasting #unit economics #web scraping #capacity planning #slot:cost-accounting

Related Articles