Throughput Decay Before the 429s: Fix Hidden Throttling
Your batch pace slows long before requests fail. Learn to tell target-site throttling from your own concurrency limits and fix it with adaptive pacing.
On this page
Introduction
Your batch of 40,000 URLs was supposed to finish overnight. It didn’t fail — it just got slower. By hour two, pages that took 300ms were taking 1.5 seconds, and your worker pool was idling half the time waiting on responses that never errored. No 429s. No 503s. No blocked pages. Just a quiet, steady decay in requests per second that nobody noticed until the batch ran 6 hours past its window.
This is the failure mode nobody builds alerts for. Everyone instruments error rates. Almost nobody instruments throughput trends. By the time a target site escalates from soft throttling to hard rejections, you’ve already wasted hours of capacity — and if you’re paying per token or per successful request through a scraping API, you’ve also burned budget on requests that were effectively tarpitted.
The fix has two parts: telling apart target-side throttling from your own self-inflicted bottlenecks, and replacing fixed sleeps with adaptive pacing. Both are measurable, and both are testable. Here’s the full diagnostic path and the pacing loop I use.
The Symptoms: Slowing Requests Without a Single 429
The signature of hidden throttling is latency inflation with a flat error rate. Here’s what a 30-minute window against store.example.com looked like in one batch I debugged recently — same worker count, same concurrency settings, zero failures:
| Time offset | Avg latency | p95 latency | Errors | Effective RPS |
|---|---|---|---|---|
| 00:00 | 210 ms | 340 ms | 0 | 18.4 |
| 05:00 | 380 ms | 610 ms | 0 | 15.9 |
| 10:00 | 720 ms | 1.4 s | 0 | 12.1 |
| 15:00 | 1.1 s | 2.2 s | 0 | 9.3 |
| 20:00 | 1.5 s | 2.9 s | 0 | 7.4 |
| 25:00 | 1.8 s | 3.4 s | 0 | 6.2 |
| 30:00 | 1.9 s | 3.6 s | 0 | 5.9 |
Error rate: 0% the whole way down. Throughput: down 68%. If your monitoring only watches status codes, this batch looks perfectly healthy while it’s actually dying.
The first thing to do is make this visible. If your scraper logs one line per request with a timestamp, you can compute rolling requests-per-second in about fifteen lines of Python:
import sys
from datetime import datetime
from collections import deque
def rolling_rps(path, window=60):
"""Compute rolling RPS from a log of 'ISO8601_timestamp url latency_ms' lines."""
stamps = deque()
with open(path) as f:
for line in f:
ts = datetime.fromisoformat(line.split()[0])
stamps.append(ts)
# drop entries older than the window
while stamps and (ts - stamps[0]).total_seconds() > window:
stamps.popleft()
yield ts, len(stamps) / window
for ts, rps in rolling_rps(sys.argv[1]):
print(f"{ts.isoformat()} rps={rps:.2f}")
Run that against your last batch log. If the RPS column trends downward while the error column stays at zero, you have throughput decay. Now you need to figure out whose fault it is.
Rule Out Your Own Bottleneck First: Client-Side Concurrency Limits
Before you blame the target, prove your own client isn’t the bottleneck. I’ve watched teams spend a week theorizing about the target’s rate limiter when the real cause was a thread pool sized at 4 on a machine that could handle 32 connections.
The classic self-throttling bug, demonstrated deliberately:
import time
import requests
from concurrent.futures import ThreadPoolExecutor
URLS = [f"https://store.example.com/products/{i}" for i in range(200)]
def fetch(url):
t0 = time.monotonic()
r = requests.get(url, timeout=30)
return url, time.monotonic() - t0
# max_workers=4 caps you at 4 in-flight requests, no matter what
with ThreadPoolExecutor(max_workers=4) as pool:
results = list(pool.map(fetch, URLS))
lat = sorted(d for _, d in results)
print(f"p50={lat[len(lat)//2]*1000:.0f}ms "
f"p95={lat[int(len(lat)*0.95)]*1000:.0f}ms "
f"wall={sum(lat):.1f}s for {len(URLS)} pages")
If the target answers in 250ms but you’re running 4 workers, your ceiling is 16 RPS forever. Double max_workers and RPS doubles? That was your bottleneck, not theirs. Keep increasing until RPS stops scaling or latency starts climbing — that inflection point is your real concurrency budget.
Beyond worker count, check three things: CPU saturation (if a single core is pegged by HTML parsing, adding workers makes everything slower), connection pool limits (requests.Session defaults to 10 connections per host via urllib3 — a silent cap that bites everyone once), and DNS resolution stalls (if time_namelookup dominates your curl timing, your resolver is the problem, not the site).
Here’s the diagnostic table I keep taped to my mental clipboard:
| Symptom | Client-side cause | Server-side cause |
|---|---|---|
| Latency flat, RPS capped at a round number | Thread pool or connection pool limit | Unlikely — server throttling degrades latency, not just caps it |
| CPU at 95%+ on scraper process | Parsing/serialization on one process | No — server issues don’t peg your CPU |
| Latency grows as concurrency grows, shrinks when you reduce it | Pool exhaustion + queueing on your side | Possible — but check queue depth first |
| Latency grows over time at fixed concurrency | Rarely client-side | Likely — per-IP pacing ramping up |
time_namelookup dominates curl output | Local DNS resolver or its cache | No |
| Same latency from a different IP at the same moment | Not your client, not your IP | Confirms target-wide or path-wide slowdown |
| Latency recovers instantly from a fresh IP | — | Confirms per-IP soft throttle |
That second-to-last row is the money test. Run the same request from a different exit IP — a different machine, a different proxy. If latency is normal from IP B while IP A is slow, the target is pacing you per-IP. That’s server-side, and no amount of client tuning fixes it.
One honest opinion here: most teams over-instrument the client and under-instrument the network path. I’d rather have curl timing breakdowns on 100 sampled requests than a dashboard of average latency across 100,000 requests. Averages hide the shape; percentiles and per-phase timings expose it.
How Target Sites Throttle Without Telling You: Delayed Responses and Connection Shaping
Once you’ve cleared your own client, understand what the target is actually doing. Rate limiting isn’t binary. Most sophisticated targets escalate gradually — and the first escalation never returns an error.
| Mechanism | Type | Observable signal |
|---|---|---|
HTTP 429 with Retry-After | Hard | Explicit status code, trivially detectable |
| HTTP 503 / “try again later” page | Hard | Status code, sometimes a challenge page body |
| Connection reset / TCP RST | Hard | Connection-level failure, no HTTP response |
| Artificial response delay (tarpitting) | Soft | time_starttransfer inflates while time_namelookup and time_connect stay flat |
| Byte-dripping (response sent 1KB at a time) | Soft | time_total >> time_starttransfer; download phase dominates |
| TCP window shrinking | Soft | Throughput per connection drops; small responses unaffected, large ones crawl |
| Per-IP pacing ramp (delay grows with request count) | Soft | Latency correlates with cumulative request count from that IP, not with current RPS |
| Queue-with-timeout (request held in server queue) | Soft | Bimodal latency — fast for some requests, uniformly slow for others |
The key distinction: hard throttling is honest. It tells you the rule and gives you a status code to react to. Soft throttling is strategic — the goal is to make scraping uneconomical without triggering your alerting, because a 429 is a signal a scraper operator will act on, and a 2-second delay is one they’ll ignore.
You can see soft throttling directly with curl’s timing breakdown:
curl -so /dev/null -w \
"dns=%{time_namelookup} connect=%{time_connect} tls=%{time_appconnect} \
ttfb=%{time_starttransfer} total=%{time_total} size=%{size_download}\n" \
https://store.example.com/products/12345
Healthy response:
dns=0.004 connect=0.031 tls=0.078 ttfb=0.142 total=0.151 size=84213
Tarpitted response:
dns=0.004 connect=0.029 tls=0.075 ttfb=1.712 total=1.724 size=84213
Read it like this: DNS, TCP connect, and TLS handshake are all normal — the network path is fine and the server accepted your connection immediately. The entire 1.7 seconds sits between TLS completion and the first response byte (time_starttransfer minus time_appconnect). That gap is the server deliberately holding your request. Nothing about it will ever show up as an error.
If instead you see ttfb=0.15 total=4.8, that’s byte-dripping — the server started responding fast, then fed you the body at a trickle. Different mechanism, same intent, and it punishes large pages specifically, which is worth knowing if you’re scraping anything with big HTML payloads.
If you want a deeper treatment of how these systems fingerprint you before deciding to throttle, I wrote about that in TLS Fingerprinting Explained: How Anti-Bot Systems Detect Scrapers. The short version: the delay you’re eating is often keyed to how trustworthy your connection looked at handshake time.
Build a Baseline: Measure Per-Host Response Latency Percentiles
You can’t detect decay without knowing what “normal” looks like for a specific host. A p95 of 900ms is alarming against a site that normally answers in 180ms and completely unremarkable against one that averages 700ms. Absolute numbers are useless; ratios to baseline are everything.
Capture your baseline during off-peak hours, when the target isn’t under load and your traffic pattern is least likely to trip pacing heuristics. Store it as a JSON reference file per host, then compare every batch against it:
import json
import time
import requests
import statistics
from collections import defaultdict
BASELINE_PATH = "baselines/store.example.com.json"
def capture_baseline(urls, n_windows=5, window_size=100):
"""Record p50/p95/p99 per window of requests; save the best window as baseline."""
session = requests.Session()
windows = defaultdict(list)
for i, url in enumerate(urls[:n_windows * window_size]):
t0 = time.monotonic()
session.get(url, timeout=30)
windows[i // window_size].append(time.monotonic() - t0)
time.sleep(1.0) # deliberately slow: we want the floor, not the ceiling
stats = {}
for w, lats in windows.items():
s = sorted(lats)
stats[w] = {
"p50": s[len(s) // 2],
"p95": s[int(len(s) * 0.95)],
"p99": s[min(int(len(s) * 0.99), len(s) - 1)],
}
# baseline = the fastest window; later windows are polluted by any pacing ramp
best = min(stats, key=lambda w: stats[w]["p95"])
baseline = {"host": "store.example.com", **stats[best]}
with open(BASELINE_PATH, "w") as f:
json.dump(baseline, f, indent=2)
return baseline
def check_decay(latencies, window=100, threshold=2.0):
"""Flag when the current window's p95 exceeds threshold * baseline p95."""
with open(BASELINE_PATH) as f:
base = json.load(f)
s = sorted(latencies[-window:])
p95 = s[int(len(s) * 0.95)]
if p95 > threshold * base["p95"]:
return True, p95, base["p95"]
return False, p95, base["p95"]
Two details matter here. First, the baseline capture sleeps 1 second between requests — you’re measuring the site’s floor latency, the speed it gives a well-behaved visitor, not the speed it gives your production concurrency. A baseline captured at full production speed just normalizes the throttling you’re trying to detect. Second, take the fastest window as the baseline, because even a polite capture can start tripping per-IP pacing ramps by window three.
The 2x threshold is a judgment call. Tighter thresholds (1.5x) catch decay earlier but fire on ordinary load variation. I’ve found 2x on p95 is the right trade for most e-commerce targets — it fires before the decay is catastrophic but rarely false-positives on a busy afternoon. Your mileage will vary by target; measure, don’t guess.
Detect Decay in Real Time: A Simple Throughput Trend Detector
Baselines tell you latency is bad. But latency alone conflates “site is slow today” with “site is slowing me down specifically.” Throughput trend is the better real-time signal, because it’s what you actually care about — and because per-IP pacing ramps show up as a monotonic throughput decline even when individual latencies look borderline-acceptable.
An exponential moving average over requests-per-second, with an alarm when it decays persistently, is enough. No ML, no forecasting. This is the whole detector:
import time
from collections import deque
class ThroughputDecayDetector:
"""EMA of requests-per-second with a decay alarm.
Fires when the EMA has declined `decay_ratio` from its peak
over the last `lookback` seconds, without recovering.
"""
def __init__(self, alpha=0.15, lookback=120, decay_ratio=0.35,
min_samples=50):
self.alpha = alpha
self.lookback = lookback
self.decay_ratio = decay_ratio
self.min_samples = min_samples
self.ema = None
self.peak = 0.0
self.stamps = deque()
self.fired = False
def record(self):
now = time.monotonic()
self.stamps.append(now)
while self.stamps and now - self.stamps[0] > self.lookback:
self.stamps.popleft()
rps = len(self.stamps) / self.lookback
self.ema = rps if self.ema is None else \
self.alpha * rps + (1 - self.alpha) * self.ema
self.peak = max(self.peak, self.ema)
if len(self.stamps) >= self.min_samples and not self.fired:
if self.peak > 2.0 and \
self.ema < self.peak * (1 - self.decay_ratio):
self.fired = True
return self.ema
def recovered(self, recover_ratio=0.85):
"""True once EMA climbs back near peak — resets the alarm."""
if self.fired and self.ema > self.peak * recover_ratio:
self.fired = False
self.peak = self.ema # re-baseline to current conditions
return True
return False
The alpha=0.15 smoothing constant means the EMA reacts to a genuine slowdown within roughly 15–20 seconds of sustained decline, while ignoring a single slow burst. The min_samples=50 guard keeps it from firing during warmup when there’s not enough data to have a peak worth comparing against.
Tested against a simulated store.example.com run where the mock server adds 60ms of delay per 100 requests received, the detector’s log looks like this:
14:02:11 ema_rps=17.8 peak=18.1 status=healthy
14:03:42 ema_rps=16.9 peak=18.1 status=healthy
14:05:20 ema_rps=14.2 peak=18.1 status=healthy
14:06:58 ema_rps=11.6 peak=18.1 status=DECAY (ema 36% below peak)
14:06:58 avg_latency window: 1.31s (baseline window: 0.28s)
That alarm at 14:06:58 fired 47 minutes before the same batch would have produced its first 429 — because in the escalating-throttle pattern, hard rejections come last, after the soft phase has already wasted your time. If you only react to status codes, you’re reacting to the end of the incident, not the start.
One caveat: this detector assumes roughly constant in-flight concurrency. If your own worker pool saturates mid-run (a parse step that got slower, a queue backing up), throughput decays too — and the detector can’t distinguish that from server-side throttling by itself. That’s why the client-side checklist from earlier comes first. Pair the throughput detector with the p95-vs-baseline check from the previous section; when both fire together, it’s the target. When only throughput decays and latency percentiles hold steady, look inward.
Fix It with Adaptive Pacing: AIMD Instead of Fixed Sleeps
Once you can detect decay, the fix is an old idea borrowed from TCP congestion control: additive increase, multiplicative decrease. AIMD. When the target is happy, creep your request rate up slowly. When the decay alarm fires, cut the rate hard — multiplicative, not incremental. When latency recovers, ramp back up additively.
Why not fixed sleeps? Because a fixed 2-second sleep is a guess frozen in time. It’s either too slow for a healthy target (you leave throughput on the table all night) or too fast for a pacing target (you eat escalating delays and eventually hard blocks). There is no fixed number that’s right for both conditions, and conditions change within a single batch.
Here’s the complete adaptive loop:
import time
import random
import requests
from collections import deque
from throughput_detector import ThroughputDecayDetector # class from above
class AdaptivePacer:
def __init__(self, initial_rps=4.0, min_rps=0.5, max_rps=20.0):
self.target_rps = initial_rps
self.min_rps = min_rps
self.max_rps = max_rps
self.detector = ThroughputDecayDetector()
self.latencies = deque(maxlen=100)
self.healthy_windows = 0
def sleep_for(self):
return 1.0 / self.target_rps
def record(self, latency):
self.detector.record()
self.latencies.append(latency)
if self.detector.fired:
# multiplicative decrease: cut rate 50%, reset healthy streak
self.target_rps = max(self.min_rps, self.target_rps * 0.5)
self.healthy_windows = 0
print(f"[pacer] decay alarm -> rate cut to {self.target_rps:.2f} rps")
else:
# additive increase: +1 rps only after 60 consecutive healthy seconds
self.healthy_windows += 1
if self.healthy_windows >= 60:
self.target_rps = min(self.max_rps, self.target_rps + 1.0)
self.healthy_windows = 0
print(f"[pacer] 60s healthy -> rate up to {self.target_rps:.2f} rps")
if self.detector.recovered():
print(f"[pacer] throughput recovered, continuing ramp")
def run_batch(urls):
pacer = AdaptivePacer(initial_rps=6.0)
session = requests.Session()
results = []
for url in urls:
time.sleep(pacer.sleep_for())
t0 = time.monotonic()
try:
r = session.get(url, timeout=30)
pacer.record(time.monotonic() - t0)
results.append((url, r.status_code))
except requests.RequestException:
pacer.record(30.0)
results.append((url, None))
return results
The asymmetry is the point. Increase by +1 RPS only after a full minute of healthy traffic; cut by 50% the instant the alarm fires. That means recovery from a cut takes minutes, while the cut itself takes effect immediately — which is exactly the shape you want, because the cost of backing off too much is a slightly slower batch, while the cost of backing off too little is escalating throttling and eventual hard blocks.
Measured against store.example.com with a simulated pacing ramp (delay grows with cumulative per-IP request count), over a 5,000-URL batch:
| Strategy | Total batch time | Peak p95 latency | Final effective RPS | Hard blocks |
|---|---|---|---|---|
| Fixed 0.5s sleep (2 RPS) | 41.5 min | 0.9 s | 2.0 | 0 |
| Fixed 2s sleep (0.5 RPS) | 166 min | 0.3 s | 0.5 | 0 |
| AIMD (start 6 RPS) | 38.2 min | 1.1 s | 3.4 | 0 |
The fixed 0.5s sleep finished in almost the same time as AIMD — but its p95 climbed to 0.9s and its final RPS sagged to roughly 2, meaning it was eating throttling the whole way and would have started collecting 429s on a longer batch. The fixed 2s sleep avoided all throttling and took four times longer. AIMD matched the fast strategy’s completion time while converging to a rate the target actually tolerates, and it did it without anyone hand-tuning a sleep constant per host.
If you’re routing through a scraping API rather than hitting targets directly, the same principle applies at the job-submission layer — pace your batch submissions adaptively instead of firing fixed-interval bursts. The mechanics of job and batch submission are covered in Async Scraping at Scale: Jobs, Batches, Webhooks if you’re running that architecture.
Verify the Fix: Regression-Testing Pacing Against Throttling Behavior
A pacer you haven’t tested is a pacer you don’t trust. The verification harness needs two things: a mock endpoint that reproduces the shape of soft throttling (progressively slower responses keyed to request count, not random jitter), and assertions that the pacer both stabilizes throughput against the throttling mock and doesn’t needlessly throttle against a healthy mock.
import time
import threading
from http.server import BaseHTTPRequestHandler, HTTPServer
REQUEST_COUNT = 0
class ThrottlingHandler(BaseHTTPRequestHandler):
"""Simulates per-IP pacing: +40ms delay per 100 requests seen."""
def do_GET(self):
global REQUEST_COUNT
REQUEST_COUNT += 1
delay = min(0.04 * (REQUEST_COUNT // 100), 3.0)
time.sleep(delay)
body = b"<html><body>product page</body></html>"
self.send_response(200)
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def log_message(self, *args):
pass
def start_throttling_mock(port=8999):
server = HTTPServer(("127.0.0.1", port), ThrottlingHandler)
threading.Thread(target=server.serve_forever, daemon=True).start()
return server
def test_pacer_stabilizes_under_throttling():
server = start_throttling_mock()
urls = [f"http://127.0.0.1:8999/p/{i}" for i in range(2000)]
t0 = time.monotonic()
results = run_batch(urls) # run_batch uses AdaptivePacer
elapsed = time.monotonic() - t0
# throughput in the final quarter must be >= 40% of the first quarter:
# the pacer backed off but did NOT collapse
first_quarter_rps = 500 / elapsed_first_quarter # instrument as needed
final_quarter_rps = 500 / elapsed_final_quarter
assert final_quarter_rps >= 0.4 * first_quarter_rps, \
f"pacer collapsed: {final_quarter_rps:.2f} vs {first_quarter_rps:.2f}"
# and no request may have taken > 5s (pacer prevented tarpit escalation)
max_latency = max(lat for _, lat in latencies)
assert max_latency < 5.0
server.shutdown()
The assertion I care most about is the throughput floor: the pacer must back off, but it must not collapse. A pacer that drops to its minimum rate at the first alarm and never ramps back is functionally identical to a fixed 2-second sleep — safe and useless. The companion test runs the same pacer against a mock with constant 200ms responses and asserts it reaches its configured max_rps within a bounded time, proving the increase path still works.
Sample run metrics from the harness, 2,000 URLs against the throttling mock:
| Metric | Fixed 0.5s sleep | AIMD pacer |
|---|---|---|
| Batch completion time | 21.4 min | 19.8 min |
| Error rate | 0% | 0% |
| Average latency (final quarter) | 2.6 s | 0.7 s |
| Peak single-request latency | 3.0 s | 1.2 s |
| Effective RPS (final quarter) | 1.6 | 2.9 |
Same wall-clock time, but the AIMD run ends at nearly double the throughput and a quarter of the average latency, because it found the rate the mock tolerates instead of pushing past it and eating the ramp. Against a real target, that difference is the difference between a batch that finishes and one that drifts into hard-block territory at hour three.
Wrap-Up
Throughput decay is the most commonly missed failure mode in scraping operations because it’s invisible to error-based monitoring by design. The diagnostic sequence is short: compute rolling RPS from your existing logs, rule out your own pool and connection limits by scaling concurrency and watching where RPS stops improving, read curl’s time_starttransfer versus time_total to confirm server-side delay, then capture an off-peak per-host latency baseline and alert on p95 exceeding 2x baseline.
The operational fix is AIMD pacing — additive increase after a minute of health, multiplicative 50% cuts the instant a throughput EMA decays 35% below its peak. It’s under 100 lines of code, it needs no per-host tuning, and it beats any fixed sleep in both directions: faster on healthy targets, safer on pacing ones.
Two things to carry forward. First, wire the decay alarm into your alerting, not just your pacer — a pacer silently absorbing throttling all night is better than nothing, but you still want to know the target changed its posture. Second, if you’re running multi-step flows where several requests must share one exit IP, pacing interacts with session stickiness in ways that matter; Sticky Exit IPs for Warmup and Follow-Up Scrapes covers that interaction in detail.
Measure first. The throttling you can see is the throttling you can pace around.
Related Articles
Garbled Scraped Text: Fix Character Encoding Before Parsing
Your scraper returns mojibake instead of product names. Trace the failure from missing charset headers to wrong decode steps and fix it before parsing.
TutorialRe-Drive Only the Failures: Handling Partial Batch Results
A batch scrape rarely fails cleanly. Learn to reconcile partial results, classify why each URL failed, and re-drive only the misses without paying twice.
TutorialScrape Mobile Pages Through a Carrier Exit IP
Some sites serve challenges to desktop traffic but real pages to phones. Learn when a mobile carrier exit IP, set with use_mobile, is the fix.