Scrape Mobile Pages Through a Carrier Exit IP
Some sites serve challenges to desktop traffic but real pages to phones. Learn when a mobile carrier exit IP, set with use_mobile, is the fix.
On this page
Why Desktop Requests Get Challenges While Phones Get Real Pages
The symptom is maddeningly specific. Your scraper requests https://store.example.com/product/123 with a standard desktop user-agent, and the response is a 403 with 14 KB of JavaScript challenge boilerplate instead of the product page. Then you open the same URL on your phone, on mobile data, and the full page loads instantly. No CAPTCHA. No interstitial. Just content.
Same URL. Same moment. Completely different outcomes.
This happens because anti-bot systems score requests, and desktop traffic from datacenter ranges scores badly on two axes at once: the TLS/HTTP fingerprint of your HTTP client and the ASN of your exit IP. Mobile visitors on carrier networks score well on both, so many sites apply their strictest rules only to the desktop path. The mobile path gets a lighter policy — sometimes no policy at all.
Here is what the divergence actually looks like when you log both attempts against the same product page:
| Request profile | HTTP status | Page title | Response size | Content |
|---|---|---|---|---|
| Desktop UA, datacenter IP | 403 | ”Just a moment…” | ~14 KB | Challenge interstitial |
| Mobile UA, carrier IP | 200 | ”Anodized Widget — Store Example” | ~212 KB | Full product HTML |
And the raw log excerpt from a curl run that reproduces it:
$ curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
-A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" \
https://store.example.com/product/123
403 14203
$ curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
-A "Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) AppleWebKit/605.1.15" \
https://store.example.com/product/123
200 212447
The naive fix people try first is swapping the user-agent string. It works for about ten minutes, until the site correlates your datacenter ASN with a mobile UA and starts challenging that combination too — a mobile UA from a cloud provider’s IP range is a contradiction no risk engine misses. If you want the mobile treatment, you have to actually arrive from a mobile network. That is the whole problem, and it has a specific solution.
How Mobile Carrier Exit IPs Differ from Datacenter and Residential IPs
Not all proxy pools are equal in the eyes of a risk engine. The ASN your exit IP belongs to is one of the cheapest, most reliable signals a site can check, and the three pool types land in very different places on it.
| Signal | Datacenter IP | Residential proxy | Mobile carrier exit IP |
|---|---|---|---|
| ASN type | Hosting/cloud provider | Consumer ISP | Mobile network operator |
| Typical blocklists | On most public blocklists | On some, inconsistently | Rarely listed |
| Users per IP | Hundreds of servers | A handful of households | Thousands of phones |
| Observed challenge rate on strict sites | High | Medium | Low |
| Cost per request | Cheapest | Mid | Highest |
The reason carrier IPs get lenient treatment is structural. A mobile carrier runs carrier-grade NAT, meaning thousands of legitimate phones share a single public IP at any moment. Blocking that IP breaks real customers, so sites are heavily disincentivized from challenging it aggressively. A datacenter IP serves no humans at all — blocking it costs the site nothing. Residential sits in between.
You can verify which pool you are actually on with an ASN lookup against the exit IP your request used:
# Parse the exit IP from your scrape response, then:
$ whois -h whois.cymru.com " -v 2001:db8:1234:abcd::1"
Bulk mode; whois.cymru.com [2025-01-05 18:22:31 +0000]
2001:db8:1234:abcd::1 | 64512 | EXAMPLE-HOSTING | example.com | - | - | - | - | US
That EXAMPLE-HOSTING line is the diagnosis. If the ASN resolves to a hosting provider, no amount of header tuning will buy you the mobile treatment. You want that lookup to return something like ATT-MOBILITY-AS20057 or VODAFONE — an operator ASN. If you want the background on how these signals feed into detection, TLS Fingerprinting Explained covers the other half of the scoring equation.
Setting use_mobile to Route Requests Through a Carrier Exit IP
The core change is one boolean. The scraping API exposes use_mobile on the sync scrape endpoint, which routes your request through a mobile carrier proxy instead of the default datacenter pool. It costs +4 tokens per request versus the default route, so treat it as a targeted tool, not a default.
Here is the configuration as a JSON payload, which is how I keep scrape configs in a pipeline repo:
{
"url": "https://store.example.com/catalog",
"use_mobile": true,
"formats": ["rawHtml"],
"use_antibot": true,
"timeout": 120
}
And the Python version, which submits the same request and prints the status plus the exit IP the request actually used:
import requests
API = "https://api.finedata.ai/scrape" # POST /api/v1/scrape
HEADERS = {"Authorization": "Bearer fd_your_api_key"}
payload = {
"url": "https://store.example.com/catalog",
"use_mobile": True,
"formats": ["rawHtml"],
"use_antibot": True,
"timeout": 120,
}
resp = requests.post(API, json=payload, headers=HEADERS, timeout=180)
data = resp.json()
print("status:", data.get("status"))
# The response body includes the exit IP used for the request.
# Parse it -- never assume it.
exit_ip = data.get("exit_ip") or data.get("data", {}).get("exit_ip")
print("exit_ip:", exit_ip)
One detail matters here: use_antibot stays on. The carrier IP fixes your network-level reputation; the antibot layer keeps the browser-like TLS fingerprint so the request doesn’t look like a raw requests call that happens to come from a phone network. You need both halves consistent. A mobile IP carrying a Python-requests TLS handshake is as contradictory as a desktop UA on a mobile IP.
Matching User-Agent and Viewport to the Mobile Exit IP
The carrier exit IP is necessary but not sufficient. Every layer of the request has to agree that this is a phone, because risk engines cross-check exactly these consistencies. If your exit IP is carrier-grade but your user-agent says Windows Chrome, you have built a fingerprint that no real device could produce. That contradiction is itself a detection signal — often a stronger one than either fact alone.
So pair use_mobile with an explicit mobile user-agent and mobile-appropriate headers:
payload = {
"url": "https://store.example.com/catalog",
"use_mobile": True,
"use_antibot": True,
"formats": ["rawHtml"],
"headers": {
"User-Agent": (
"Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) "
"AppleWebKit/605.1.15 (KHTML, like Gecko) "
"Version/17.0 Mobile/15E148 Safari/604.1"
),
"Accept": "text/html,application/xhtml+xml",
"Accept-Language": "en-US,en;q=0.9",
"Viewport": "width=device-width, initial-scale=1.0",
"Sec-CH-UA-Mobile": "?1",
},
}
If the site serves a responsive layout, the Viewport and client-hint headers matter when JS rendering is involved — the rendered DOM differs between a mobile and desktop viewport, and some sites only hydrate their mobile layout when the viewport says so. Add use_js_render: True when the catalog is client-side rendered.
The before/after on a strict product page makes the consistency argument concrete:
| Configuration | Exit IP type | User-Agent | Result |
|---|---|---|---|
| Default | Datacenter | Desktop Chrome | 403 challenge |
| Mobile IP, mismatched profile | Carrier | Desktop Chrome | 403 challenge (inconsistency flagged) |
| Mobile IP, partial match | Carrier | Mobile UA, desktop viewport | 200, but desktop DOM served |
| Fully consistent mobile profile | Carrier | Mobile UA + mobile viewport | 200, mobile DOM served |
Row two is the one people skip and then blame the proxy. The mismatch is detectable and gets detected. Row three is sneakier: you get a 200, but the site serves you the desktop layout because your viewport hints said desktop, and your extraction selectors — written against the mobile DOM you saw on your phone — silently return nothing. Fix Scraped Data That Lands in the Wrong Fields covers that failure mode in detail.
Verifying the Fix: Measuring Response Quality Before and After
Never trust a fix you haven’t measured. “It seemed to work” is how a pipeline ends up storing 14 KB of challenge HTML for three weeks before anyone notices. Run both routes against the same URL and compare content length, the title tag, and a CAPTCHA marker before you commit the change.
import re
import requests
API = "https://api.finedata.ai/scrape"
HEADERS = {"Authorization": "Bearer fd_your_api_key"}
MOBILE_UA = (
"Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) "
"AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Mobile/15E148 Safari/604.1"
)
DESKTOP_UA = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36"
)
CAPTCHA_MARKERS = ("cf-challenge", "g-recaptcha", "h-captcha", "px-captcha")
def scrape(url, use_mobile, user_agent):
payload = {
"url": url,
"use_mobile": use_mobile,
"use_antibot": True,
"formats": ["rawHtml"],
"headers": {"User-Agent": user_agent},
}
r = requests.post(API, json=payload, headers=HEADERS, timeout=180)
r.raise_for_status()
return r.json()
def profile(result):
html = str(result.get("data", {}).get("rawHtml", ""))
title = re.search(r"<title[^>]*>(.*?)</title>", html, re.S)
return {
"http_status": result.get("status"),
"title": (title.group(1).strip() if title else "")[:60],
"content_bytes": len(html),
"captcha_present": any(m in html for m in CAPTCHA_MARKERS),
}
url = "https://store.example.com/product/123"
desktop = profile(scrape(url, use_mobile=False, user_agent=DESKTOP_UA))
mobile = profile(scrape(url, use_mobile=True, user_agent=MOBILE_UA))
for name, p in (("desktop", desktop), ("mobile", mobile)):
print(f"{name:8} status={p['http_status']} bytes={p['content_bytes']:>7} "
f"captcha={p['captcha_present']} title={p['title']!r}")
The output tells you unambiguously whether the route change did anything:
| Route | HTTP status | Title | Content bytes | CAPTCHA marker |
|---|---|---|---|---|
| Desktop default | 403 | ”Just a moment…“ | 14,203 | true |
use_mobile: true | 200 | ”Anodized Widget — Store Example” | 212,447 | false |
A 15x difference in content size plus a real product title is the signature of a genuine page. If the mobile route returns similar byte counts to the desktop challenge response, the site challenges mobile traffic too, and you need a different escalation path — heavier stealth options or, better, avoiding the CAPTCHA entirely rather than paying to solve it.
Put this comparison script in your CI or your pipeline’s pre-flight checks. When the site changes its policy six months from now, you want the alarm to fire before your warehouse fills with interstitials.
Handling Pagination and Session Stickiness on Mobile Routes
Once the mobile route works for a single page, the next problem is crawling a category without every page request landing on a different carrier IP. Some sites bind your session to the first IP they see; a new IP mid-session with the same cookies looks like session hijacking and gets you challenged again. The fix is session_id, which pins all requests carrying the same ID to one exit IP, with session_ttl controlling how long that binding lives (default 1800 seconds, max 86400).
import re
import uuid
import requests
API = "https://api.finedata.ai/scrape"
HEADERS = {"Authorization": "Bearer fd_your_api_key"}
MOBILE_UA = (
"Mozilla/5.0 (Linux; Android 14; Pixel 8) "
"AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Mobile Safari/537.36"
)
session_id = str(uuid.uuid4()) # one session, one exit IP, whole crawl
base = "https://store.example.com/c"
page = 1
seen_urls = set()
while True:
payload = {
"url": f"{base}?page={page}",
"use_mobile": True,
"use_antibot": True,
"formats": ["rawHtml"],
"session_id": session_id,
"session_ttl": 3600,
"headers": {"User-Agent": MOBILE_UA},
"only_main_content": True,
}
r = requests.post(API, json=payload, headers=HEADERS, timeout=180)
r.raise_for_status()
html = str(r.json().get("data", {}).get("rawHtml", ""))
if not html or "Just a moment" in html:
print(f"page {page}: challenge or empty, stopping")
break
# Follow rel=next in the mobile-rendered markup
next_link = re.search(r'<link[^>]+rel=["\']next["\'][^>]+href=["\']([^"\']+)["\']', html)
if not next_link:
# Fallback: some mobile templates use an anchor instead
next_link = re.search(r'<a[^>]+rel=["\']next["\'][^>]+href=["\']([^"\']+)["\']', html)
print(f"page {page}: {len(html)} bytes")
if not next_link:
break
next_url = next_link.group(1)
if next_url in seen_urls:
break # pagination looped -- stop
seen_urls.add(next_url)
page += 1
Two things in that loop deserve comment. First, the rel=next detection tries both <link> and <a> forms because mobile templates are often a different codebase from the desktop site — a separate mobile theme, not a responsive layout — and the markup is not guaranteed to match what you saw in a desktop browser. Second, the loop guard on seen_urls matters because some mobile templates paginate with JavaScript state instead of URLs and will happily serve the same page forever if you trust the page counter alone.
If you are running this at volume, move it to the async batch endpoint (POST /api/v1/async/batch) with a callback_url, so you are not holding 120-second timeouts open per page. The session mechanics work the same way across batch jobs. Keep One Exit IP Across Multi-Step Scrape Requests goes deeper on session warmup patterns.
When Not to Use a Mobile Exit IP: Costs and Trade-offs
Here is where I will argue against the tool this post just taught you. use_mobile costs +4 tokens per request on top of the base, and for many workloads it buys nothing. Carrier IPs are the right answer to a specific problem — a site that challenges desktop/datacenter traffic but serves mobile traffic — and an expensive habit everywhere else.
| Crawl scenario | Recommended exit route | Why |
|---|---|---|
| Low-volume detail pages on a strict site | Mobile carrier | The +4 token cost is trivial at dozens of pages/day |
| High-volume sitemap crawl, lenient site | Datacenter (default) | No reputation problem to solve; +4 tokens x 500K pages is pure waste |
| Geo-restricted content | proxy_country + residential | You need a specific country, not a carrier ASN |
| Multi-step flow needing one warm IP | session_id + residential, sticky | Carrier NAT makes stickiness less predictable at high volume |
| Desktop-only web app | Datacenter + use_js_render | Mobile route may serve a stripped-down template |
That last row is the trap nobody warns you about. Some sites serve less data in their mobile HTML than the desktop version — deliberately. A common case: the target site hides the full specification table and customer reviews behind an accordion that only renders in the desktop DOM, and the mobile template simply omits the markup. Your scraper gets a clean 200, extracts the title and price, and silently never collects the specs — because the element does not exist in the response.
Check before committing:
# Compare the same URL across routes before choosing one
desktop_html = scrape("https://store.example.com/product/123", use_mobile=False)
mobile_html = scrape("https://store.example.com/product/123", use_mobile=True)
for marker in ('id="specifications"', 'class="reviews"', 'data-price-history'):
print(f"{marker}: desktop={marker in desktop_html} mobile={marker in mobile_html}")
If the marker table shows the mobile route missing elements you need, the mobile route is the wrong tool even though it gets you past the challenge. In that case the better escalation is use_residential or the stealth options, which keep the desktop DOM while improving IP reputation.
My rule of thumb: default to the cheapest route that returns real content, measured by the verification script above. Escalate one step at a time — datacenter, then residential, then mobile — and only when the measurement shows the cheaper route failing. Pinning your whole pipeline to carrier IPs because one strict site needed it is how token budgets double overnight.
Wrap-up
The pattern is consistent: strict sites score requests across multiple signals at once, and the mobile path often carries a lighter policy because carrier NAT makes blocking it expensive for the site. To exploit that, you need the full profile to agree — use_mobile: true for the carrier exit IP, a mobile user-agent, mobile viewport hints, and use_antibot for the TLS layer. Any single mismatched layer can undo the rest.
The parts worth remembering:
- Measure before and after. Content length, title tag, and a CAPTCHA marker will tell you the truth about any route change.
- Use
session_idfor paginated crawls so the whole session shares one exit IP. - Check whether the mobile template actually contains the elements you need — a 200 with missing data is worse than a 403 you can see.
- Save carrier IPs for the sites that force the issue. At +4 tokens per request, the default route should stay your default.
If the challenge layer is your real bottleneck rather than IP reputation, Handling CAPTCHAs When Web Scraping covers the solve-versus-avoid calculus. And if you’re weighing the sync approach here against batch jobs for a full catalog crawl, Async Scraping at Scale is the natural next read.
Related Articles
Stop Cleaning HTML by Hand: Scrape Pages as Markdown
Chunking raw HTML for LLM pipelines wastes tokens and breaks parsers. See how one scrape request field returns clean, ready-to-chunk markdown instead.
TutorialScrape Localized Storefronts With a Country Exit Code
Storefronts change price, language, and stock by visitor country. Pin the scrape exit with an ISO-2 code and handle 422 when the country is unsupported.
TutorialRoute Scraping Traffic Through Proxies You Already Own
When targets allowlist your IPs or you already pay for residential proxies, attach a proxy profile so scrape requests exit through your pool.