Skip to content
Tutorial 18 min read

Throughput Decay Before the 429s: Fix Hidden Throttling

Your batch pace slows long before requests fail. Learn to tell target-site throttling from your own concurrency limits and fix it with adaptive pacing.

FE
FineData Engineering · Editorial Policy
|
On this page

Introduction

Your batch of 40,000 URLs was supposed to finish overnight. It didn’t fail — it just got slower. By hour two, pages that took 300ms were taking 1.5 seconds, and your worker pool was idling half the time waiting on responses that never errored. No 429s. No 503s. No blocked pages. Just a quiet, steady decay in requests per second that nobody noticed until the batch ran 6 hours past its window.

This is the failure mode nobody builds alerts for. Everyone instruments error rates. Almost nobody instruments throughput trends. By the time a target site escalates from soft throttling to hard rejections, you’ve already wasted hours of capacity — and if you’re paying per token or per successful request through a scraping API, you’ve also burned budget on requests that were effectively tarpitted.

The fix has two parts: telling apart target-side throttling from your own self-inflicted bottlenecks, and replacing fixed sleeps with adaptive pacing. Both are measurable, and both are testable. Here’s the full diagnostic path and the pacing loop I use.

The Symptoms: Slowing Requests Without a Single 429

The signature of hidden throttling is latency inflation with a flat error rate. Here’s what a 30-minute window against store.example.com looked like in one batch I debugged recently — same worker count, same concurrency settings, zero failures:

Time offsetAvg latencyp95 latencyErrorsEffective RPS
00:00210 ms340 ms018.4
05:00380 ms610 ms015.9
10:00720 ms1.4 s012.1
15:001.1 s2.2 s09.3
20:001.5 s2.9 s07.4
25:001.8 s3.4 s06.2
30:001.9 s3.6 s05.9

Error rate: 0% the whole way down. Throughput: down 68%. If your monitoring only watches status codes, this batch looks perfectly healthy while it’s actually dying.

The first thing to do is make this visible. If your scraper logs one line per request with a timestamp, you can compute rolling requests-per-second in about fifteen lines of Python:

import sys
from datetime import datetime
from collections import deque

def rolling_rps(path, window=60):
    """Compute rolling RPS from a log of 'ISO8601_timestamp url latency_ms' lines."""
    stamps = deque()
    with open(path) as f:
        for line in f:
            ts = datetime.fromisoformat(line.split()[0])
            stamps.append(ts)
            # drop entries older than the window
            while stamps and (ts - stamps[0]).total_seconds() > window:
                stamps.popleft()
            yield ts, len(stamps) / window

for ts, rps in rolling_rps(sys.argv[1]):
    print(f"{ts.isoformat()}  rps={rps:.2f}")

Run that against your last batch log. If the RPS column trends downward while the error column stays at zero, you have throughput decay. Now you need to figure out whose fault it is.

Rule Out Your Own Bottleneck First: Client-Side Concurrency Limits

Before you blame the target, prove your own client isn’t the bottleneck. I’ve watched teams spend a week theorizing about the target’s rate limiter when the real cause was a thread pool sized at 4 on a machine that could handle 32 connections.

The classic self-throttling bug, demonstrated deliberately:

import time
import requests
from concurrent.futures import ThreadPoolExecutor

URLS = [f"https://store.example.com/products/{i}" for i in range(200)]

def fetch(url):
    t0 = time.monotonic()
    r = requests.get(url, timeout=30)
    return url, time.monotonic() - t0

# max_workers=4 caps you at 4 in-flight requests, no matter what
with ThreadPoolExecutor(max_workers=4) as pool:
    results = list(pool.map(fetch, URLS))

lat = sorted(d for _, d in results)
print(f"p50={lat[len(lat)//2]*1000:.0f}ms  "
      f"p95={lat[int(len(lat)*0.95)]*1000:.0f}ms  "
      f"wall={sum(lat):.1f}s for {len(URLS)} pages")

If the target answers in 250ms but you’re running 4 workers, your ceiling is 16 RPS forever. Double max_workers and RPS doubles? That was your bottleneck, not theirs. Keep increasing until RPS stops scaling or latency starts climbing — that inflection point is your real concurrency budget.

Beyond worker count, check three things: CPU saturation (if a single core is pegged by HTML parsing, adding workers makes everything slower), connection pool limits (requests.Session defaults to 10 connections per host via urllib3 — a silent cap that bites everyone once), and DNS resolution stalls (if time_namelookup dominates your curl timing, your resolver is the problem, not the site).

Here’s the diagnostic table I keep taped to my mental clipboard:

SymptomClient-side causeServer-side cause
Latency flat, RPS capped at a round numberThread pool or connection pool limitUnlikely — server throttling degrades latency, not just caps it
CPU at 95%+ on scraper processParsing/serialization on one processNo — server issues don’t peg your CPU
Latency grows as concurrency grows, shrinks when you reduce itPool exhaustion + queueing on your sidePossible — but check queue depth first
Latency grows over time at fixed concurrencyRarely client-sideLikely — per-IP pacing ramping up
time_namelookup dominates curl outputLocal DNS resolver or its cacheNo
Same latency from a different IP at the same momentNot your client, not your IPConfirms target-wide or path-wide slowdown
Latency recovers instantly from a fresh IP—Confirms per-IP soft throttle

That second-to-last row is the money test. Run the same request from a different exit IP — a different machine, a different proxy. If latency is normal from IP B while IP A is slow, the target is pacing you per-IP. That’s server-side, and no amount of client tuning fixes it.

One honest opinion here: most teams over-instrument the client and under-instrument the network path. I’d rather have curl timing breakdowns on 100 sampled requests than a dashboard of average latency across 100,000 requests. Averages hide the shape; percentiles and per-phase timings expose it.

How Target Sites Throttle Without Telling You: Delayed Responses and Connection Shaping

Once you’ve cleared your own client, understand what the target is actually doing. Rate limiting isn’t binary. Most sophisticated targets escalate gradually — and the first escalation never returns an error.

MechanismTypeObservable signal
HTTP 429 with Retry-AfterHardExplicit status code, trivially detectable
HTTP 503 / “try again later” pageHardStatus code, sometimes a challenge page body
Connection reset / TCP RSTHardConnection-level failure, no HTTP response
Artificial response delay (tarpitting)Softtime_starttransfer inflates while time_namelookup and time_connect stay flat
Byte-dripping (response sent 1KB at a time)Softtime_total >> time_starttransfer; download phase dominates
TCP window shrinkingSoftThroughput per connection drops; small responses unaffected, large ones crawl
Per-IP pacing ramp (delay grows with request count)SoftLatency correlates with cumulative request count from that IP, not with current RPS
Queue-with-timeout (request held in server queue)SoftBimodal latency — fast for some requests, uniformly slow for others

The key distinction: hard throttling is honest. It tells you the rule and gives you a status code to react to. Soft throttling is strategic — the goal is to make scraping uneconomical without triggering your alerting, because a 429 is a signal a scraper operator will act on, and a 2-second delay is one they’ll ignore.

You can see soft throttling directly with curl’s timing breakdown:

curl -so /dev/null -w \
  "dns=%{time_namelookup} connect=%{time_connect} tls=%{time_appconnect} \
ttfb=%{time_starttransfer} total=%{time_total} size=%{size_download}\n" \
  https://store.example.com/products/12345

Healthy response:

dns=0.004 connect=0.031 tls=0.078 ttfb=0.142 total=0.151 size=84213

Tarpitted response:

dns=0.004 connect=0.029 tls=0.075 ttfb=1.712 total=1.724 size=84213

Read it like this: DNS, TCP connect, and TLS handshake are all normal — the network path is fine and the server accepted your connection immediately. The entire 1.7 seconds sits between TLS completion and the first response byte (time_starttransfer minus time_appconnect). That gap is the server deliberately holding your request. Nothing about it will ever show up as an error.

If instead you see ttfb=0.15 total=4.8, that’s byte-dripping — the server started responding fast, then fed you the body at a trickle. Different mechanism, same intent, and it punishes large pages specifically, which is worth knowing if you’re scraping anything with big HTML payloads.

If you want a deeper treatment of how these systems fingerprint you before deciding to throttle, I wrote about that in TLS Fingerprinting Explained: How Anti-Bot Systems Detect Scrapers. The short version: the delay you’re eating is often keyed to how trustworthy your connection looked at handshake time.

Build a Baseline: Measure Per-Host Response Latency Percentiles

You can’t detect decay without knowing what “normal” looks like for a specific host. A p95 of 900ms is alarming against a site that normally answers in 180ms and completely unremarkable against one that averages 700ms. Absolute numbers are useless; ratios to baseline are everything.

Capture your baseline during off-peak hours, when the target isn’t under load and your traffic pattern is least likely to trip pacing heuristics. Store it as a JSON reference file per host, then compare every batch against it:

import json
import time
import requests
import statistics
from collections import defaultdict

BASELINE_PATH = "baselines/store.example.com.json"

def capture_baseline(urls, n_windows=5, window_size=100):
    """Record p50/p95/p99 per window of requests; save the best window as baseline."""
    session = requests.Session()
    windows = defaultdict(list)
    for i, url in enumerate(urls[:n_windows * window_size]):
        t0 = time.monotonic()
        session.get(url, timeout=30)
        windows[i // window_size].append(time.monotonic() - t0)
        time.sleep(1.0)  # deliberately slow: we want the floor, not the ceiling

    stats = {}
    for w, lats in windows.items():
        s = sorted(lats)
        stats[w] = {
            "p50": s[len(s) // 2],
            "p95": s[int(len(s) * 0.95)],
            "p99": s[min(int(len(s) * 0.99), len(s) - 1)],
        }
    # baseline = the fastest window; later windows are polluted by any pacing ramp
    best = min(stats, key=lambda w: stats[w]["p95"])
    baseline = {"host": "store.example.com", **stats[best]}
    with open(BASELINE_PATH, "w") as f:
        json.dump(baseline, f, indent=2)
    return baseline

def check_decay(latencies, window=100, threshold=2.0):
    """Flag when the current window's p95 exceeds threshold * baseline p95."""
    with open(BASELINE_PATH) as f:
        base = json.load(f)
    s = sorted(latencies[-window:])
    p95 = s[int(len(s) * 0.95)]
    if p95 > threshold * base["p95"]:
        return True, p95, base["p95"]
    return False, p95, base["p95"]

Two details matter here. First, the baseline capture sleeps 1 second between requests — you’re measuring the site’s floor latency, the speed it gives a well-behaved visitor, not the speed it gives your production concurrency. A baseline captured at full production speed just normalizes the throttling you’re trying to detect. Second, take the fastest window as the baseline, because even a polite capture can start tripping per-IP pacing ramps by window three.

The 2x threshold is a judgment call. Tighter thresholds (1.5x) catch decay earlier but fire on ordinary load variation. I’ve found 2x on p95 is the right trade for most e-commerce targets — it fires before the decay is catastrophic but rarely false-positives on a busy afternoon. Your mileage will vary by target; measure, don’t guess.

Detect Decay in Real Time: A Simple Throughput Trend Detector

Baselines tell you latency is bad. But latency alone conflates “site is slow today” with “site is slowing me down specifically.” Throughput trend is the better real-time signal, because it’s what you actually care about — and because per-IP pacing ramps show up as a monotonic throughput decline even when individual latencies look borderline-acceptable.

An exponential moving average over requests-per-second, with an alarm when it decays persistently, is enough. No ML, no forecasting. This is the whole detector:

import time
from collections import deque

class ThroughputDecayDetector:
    """EMA of requests-per-second with a decay alarm.

    Fires when the EMA has declined `decay_ratio` from its peak
    over the last `lookback` seconds, without recovering.
    """

    def __init__(self, alpha=0.15, lookback=120, decay_ratio=0.35,
                 min_samples=50):
        self.alpha = alpha
        self.lookback = lookback
        self.decay_ratio = decay_ratio
        self.min_samples = min_samples
        self.ema = None
        self.peak = 0.0
        self.stamps = deque()
        self.fired = False

    def record(self):
        now = time.monotonic()
        self.stamps.append(now)
        while self.stamps and now - self.stamps[0] > self.lookback:
            self.stamps.popleft()
        rps = len(self.stamps) / self.lookback
        self.ema = rps if self.ema is None else \
            self.alpha * rps + (1 - self.alpha) * self.ema
        self.peak = max(self.peak, self.ema)
        if len(self.stamps) >= self.min_samples and not self.fired:
            if self.peak > 2.0 and \
               self.ema < self.peak * (1 - self.decay_ratio):
                self.fired = True
        return self.ema

    def recovered(self, recover_ratio=0.85):
        """True once EMA climbs back near peak — resets the alarm."""
        if self.fired and self.ema > self.peak * recover_ratio:
            self.fired = False
            self.peak = self.ema  # re-baseline to current conditions
            return True
        return False

The alpha=0.15 smoothing constant means the EMA reacts to a genuine slowdown within roughly 15–20 seconds of sustained decline, while ignoring a single slow burst. The min_samples=50 guard keeps it from firing during warmup when there’s not enough data to have a peak worth comparing against.

Tested against a simulated store.example.com run where the mock server adds 60ms of delay per 100 requests received, the detector’s log looks like this:

14:02:11  ema_rps=17.8  peak=18.1  status=healthy
14:03:42  ema_rps=16.9  peak=18.1  status=healthy
14:05:20  ema_rps=14.2  peak=18.1  status=healthy
14:06:58  ema_rps=11.6  peak=18.1  status=DECAY  (ema 36% below peak)
14:06:58  avg_latency window: 1.31s (baseline window: 0.28s)

That alarm at 14:06:58 fired 47 minutes before the same batch would have produced its first 429 — because in the escalating-throttle pattern, hard rejections come last, after the soft phase has already wasted your time. If you only react to status codes, you’re reacting to the end of the incident, not the start.

One caveat: this detector assumes roughly constant in-flight concurrency. If your own worker pool saturates mid-run (a parse step that got slower, a queue backing up), throughput decays too — and the detector can’t distinguish that from server-side throttling by itself. That’s why the client-side checklist from earlier comes first. Pair the throughput detector with the p95-vs-baseline check from the previous section; when both fire together, it’s the target. When only throughput decays and latency percentiles hold steady, look inward.

Fix It with Adaptive Pacing: AIMD Instead of Fixed Sleeps

Once you can detect decay, the fix is an old idea borrowed from TCP congestion control: additive increase, multiplicative decrease. AIMD. When the target is happy, creep your request rate up slowly. When the decay alarm fires, cut the rate hard — multiplicative, not incremental. When latency recovers, ramp back up additively.

Why not fixed sleeps? Because a fixed 2-second sleep is a guess frozen in time. It’s either too slow for a healthy target (you leave throughput on the table all night) or too fast for a pacing target (you eat escalating delays and eventually hard blocks). There is no fixed number that’s right for both conditions, and conditions change within a single batch.

Here’s the complete adaptive loop:

import time
import random
import requests
from collections import deque
from throughput_detector import ThroughputDecayDetector  # class from above

class AdaptivePacer:
    def __init__(self, initial_rps=4.0, min_rps=0.5, max_rps=20.0):
        self.target_rps = initial_rps
        self.min_rps = min_rps
        self.max_rps = max_rps
        self.detector = ThroughputDecayDetector()
        self.latencies = deque(maxlen=100)
        self.healthy_windows = 0

    def sleep_for(self):
        return 1.0 / self.target_rps

    def record(self, latency):
        self.detector.record()
        self.latencies.append(latency)

        if self.detector.fired:
            # multiplicative decrease: cut rate 50%, reset healthy streak
            self.target_rps = max(self.min_rps, self.target_rps * 0.5)
            self.healthy_windows = 0
            print(f"[pacer] decay alarm -> rate cut to {self.target_rps:.2f} rps")
        else:
            # additive increase: +1 rps only after 60 consecutive healthy seconds
            self.healthy_windows += 1
            if self.healthy_windows >= 60:
                self.target_rps = min(self.max_rps, self.target_rps + 1.0)
                self.healthy_windows = 0
                print(f"[pacer] 60s healthy -> rate up to {self.target_rps:.2f} rps")

        if self.detector.recovered():
            print(f"[pacer] throughput recovered, continuing ramp")

def run_batch(urls):
    pacer = AdaptivePacer(initial_rps=6.0)
    session = requests.Session()
    results = []
    for url in urls:
        time.sleep(pacer.sleep_for())
        t0 = time.monotonic()
        try:
            r = session.get(url, timeout=30)
            pacer.record(time.monotonic() - t0)
            results.append((url, r.status_code))
        except requests.RequestException:
            pacer.record(30.0)
            results.append((url, None))
    return results

The asymmetry is the point. Increase by +1 RPS only after a full minute of healthy traffic; cut by 50% the instant the alarm fires. That means recovery from a cut takes minutes, while the cut itself takes effect immediately — which is exactly the shape you want, because the cost of backing off too much is a slightly slower batch, while the cost of backing off too little is escalating throttling and eventual hard blocks.

Measured against store.example.com with a simulated pacing ramp (delay grows with cumulative per-IP request count), over a 5,000-URL batch:

StrategyTotal batch timePeak p95 latencyFinal effective RPSHard blocks
Fixed 0.5s sleep (2 RPS)41.5 min0.9 s2.00
Fixed 2s sleep (0.5 RPS)166 min0.3 s0.50
AIMD (start 6 RPS)38.2 min1.1 s3.40

The fixed 0.5s sleep finished in almost the same time as AIMD — but its p95 climbed to 0.9s and its final RPS sagged to roughly 2, meaning it was eating throttling the whole way and would have started collecting 429s on a longer batch. The fixed 2s sleep avoided all throttling and took four times longer. AIMD matched the fast strategy’s completion time while converging to a rate the target actually tolerates, and it did it without anyone hand-tuning a sleep constant per host.

If you’re routing through a scraping API rather than hitting targets directly, the same principle applies at the job-submission layer — pace your batch submissions adaptively instead of firing fixed-interval bursts. The mechanics of job and batch submission are covered in Async Scraping at Scale: Jobs, Batches, Webhooks if you’re running that architecture.

Verify the Fix: Regression-Testing Pacing Against Throttling Behavior

A pacer you haven’t tested is a pacer you don’t trust. The verification harness needs two things: a mock endpoint that reproduces the shape of soft throttling (progressively slower responses keyed to request count, not random jitter), and assertions that the pacer both stabilizes throughput against the throttling mock and doesn’t needlessly throttle against a healthy mock.

import time
import threading
from http.server import BaseHTTPRequestHandler, HTTPServer

REQUEST_COUNT = 0

class ThrottlingHandler(BaseHTTPRequestHandler):
    """Simulates per-IP pacing: +40ms delay per 100 requests seen."""
    def do_GET(self):
        global REQUEST_COUNT
        REQUEST_COUNT += 1
        delay = min(0.04 * (REQUEST_COUNT // 100), 3.0)
        time.sleep(delay)
        body = b"<html><body>product page</body></html>"
        self.send_response(200)
        self.send_header("Content-Length", str(len(body)))
        self.end_headers()
        self.wfile.write(body)

    def log_message(self, *args):
        pass

def start_throttling_mock(port=8999):
    server = HTTPServer(("127.0.0.1", port), ThrottlingHandler)
    threading.Thread(target=server.serve_forever, daemon=True).start()
    return server

def test_pacer_stabilizes_under_throttling():
    server = start_throttling_mock()
    urls = [f"http://127.0.0.1:8999/p/{i}" for i in range(2000)]
    t0 = time.monotonic()
    results = run_batch(urls)          # run_batch uses AdaptivePacer
    elapsed = time.monotonic() - t0

    # throughput in the final quarter must be >= 40% of the first quarter:
    # the pacer backed off but did NOT collapse
    first_quarter_rps = 500 / elapsed_first_quarter  # instrument as needed
    final_quarter_rps = 500 / elapsed_final_quarter
    assert final_quarter_rps >= 0.4 * first_quarter_rps, \
        f"pacer collapsed: {final_quarter_rps:.2f} vs {first_quarter_rps:.2f}"

    # and no request may have taken > 5s (pacer prevented tarpit escalation)
    max_latency = max(lat for _, lat in latencies)
    assert max_latency < 5.0
    server.shutdown()

The assertion I care most about is the throughput floor: the pacer must back off, but it must not collapse. A pacer that drops to its minimum rate at the first alarm and never ramps back is functionally identical to a fixed 2-second sleep — safe and useless. The companion test runs the same pacer against a mock with constant 200ms responses and asserts it reaches its configured max_rps within a bounded time, proving the increase path still works.

Sample run metrics from the harness, 2,000 URLs against the throttling mock:

MetricFixed 0.5s sleepAIMD pacer
Batch completion time21.4 min19.8 min
Error rate0%0%
Average latency (final quarter)2.6 s0.7 s
Peak single-request latency3.0 s1.2 s
Effective RPS (final quarter)1.62.9

Same wall-clock time, but the AIMD run ends at nearly double the throughput and a quarter of the average latency, because it found the rate the mock tolerates instead of pushing past it and eating the ramp. Against a real target, that difference is the difference between a batch that finishes and one that drifts into hard-block territory at hour three.

Wrap-Up

Throughput decay is the most commonly missed failure mode in scraping operations because it’s invisible to error-based monitoring by design. The diagnostic sequence is short: compute rolling RPS from your existing logs, rule out your own pool and connection limits by scaling concurrency and watching where RPS stops improving, read curl’s time_starttransfer versus time_total to confirm server-side delay, then capture an off-peak per-host latency baseline and alert on p95 exceeding 2x baseline.

The operational fix is AIMD pacing — additive increase after a minute of health, multiplicative 50% cuts the instant a throughput EMA decays 35% below its peak. It’s under 100 lines of code, it needs no per-host tuning, and it beats any fixed sleep in both directions: faster on healthy targets, safer on pacing ones.

Two things to carry forward. First, wire the decay alarm into your alerting, not just your pacer — a pacer silently absorbing throttling all night is better than nothing, but you still want to know the target changed its posture. Second, if you’re running multi-step flows where several requests must share one exit IP, pacing interacts with session stickiness in ways that matter; Sticky Exit IPs for Warmup and Follow-Up Scrapes covers that interaction in detail.

Measure first. The throttling you can see is the throttling you can pace around.

#rate limits #throttling #concurrency control #adaptive pacing #scraping operations #slot:operations

Related Articles