Skip to content
Tutorial 13 min read

Scrape Mobile Pages Through a Carrier Exit IP

Some sites serve challenges to desktop traffic but real pages to phones. Learn when a mobile carrier exit IP, set with use_mobile, is the fix.

FE
FineData Engineering · Editorial Policy
|
On this page

Why Desktop Requests Get Challenges While Phones Get Real Pages

The symptom is maddeningly specific. Your scraper requests https://store.example.com/product/123 with a standard desktop user-agent, and the response is a 403 with 14 KB of JavaScript challenge boilerplate instead of the product page. Then you open the same URL on your phone, on mobile data, and the full page loads instantly. No CAPTCHA. No interstitial. Just content.

Same URL. Same moment. Completely different outcomes.

This happens because anti-bot systems score requests, and desktop traffic from datacenter ranges scores badly on two axes at once: the TLS/HTTP fingerprint of your HTTP client and the ASN of your exit IP. Mobile visitors on carrier networks score well on both, so many sites apply their strictest rules only to the desktop path. The mobile path gets a lighter policy — sometimes no policy at all.

Here is what the divergence actually looks like when you log both attempts against the same product page:

Request profileHTTP statusPage titleResponse sizeContent
Desktop UA, datacenter IP403”Just a moment…”~14 KBChallenge interstitial
Mobile UA, carrier IP200”Anodized Widget — Store Example”~212 KBFull product HTML

And the raw log excerpt from a curl run that reproduces it:

$ curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
    -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" \
    https://store.example.com/product/123
403 14203

$ curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
    -A "Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) AppleWebKit/605.1.15" \
    https://store.example.com/product/123
200 212447

The naive fix people try first is swapping the user-agent string. It works for about ten minutes, until the site correlates your datacenter ASN with a mobile UA and starts challenging that combination too — a mobile UA from a cloud provider’s IP range is a contradiction no risk engine misses. If you want the mobile treatment, you have to actually arrive from a mobile network. That is the whole problem, and it has a specific solution.

How Mobile Carrier Exit IPs Differ from Datacenter and Residential IPs

Not all proxy pools are equal in the eyes of a risk engine. The ASN your exit IP belongs to is one of the cheapest, most reliable signals a site can check, and the three pool types land in very different places on it.

SignalDatacenter IPResidential proxyMobile carrier exit IP
ASN typeHosting/cloud providerConsumer ISPMobile network operator
Typical blocklistsOn most public blocklistsOn some, inconsistentlyRarely listed
Users per IPHundreds of serversA handful of householdsThousands of phones
Observed challenge rate on strict sitesHighMediumLow
Cost per requestCheapestMidHighest

The reason carrier IPs get lenient treatment is structural. A mobile carrier runs carrier-grade NAT, meaning thousands of legitimate phones share a single public IP at any moment. Blocking that IP breaks real customers, so sites are heavily disincentivized from challenging it aggressively. A datacenter IP serves no humans at all — blocking it costs the site nothing. Residential sits in between.

You can verify which pool you are actually on with an ASN lookup against the exit IP your request used:

# Parse the exit IP from your scrape response, then:
$ whois -h whois.cymru.com " -v 2001:db8:1234:abcd::1"
Bulk mode; whois.cymru.com [2025-01-05 18:22:31 +0000]
2001:db8:1234:abcd::1 | 64512 | EXAMPLE-HOSTING | example.com | - | - | - | - | US

That EXAMPLE-HOSTING line is the diagnosis. If the ASN resolves to a hosting provider, no amount of header tuning will buy you the mobile treatment. You want that lookup to return something like ATT-MOBILITY-AS20057 or VODAFONE — an operator ASN. If you want the background on how these signals feed into detection, TLS Fingerprinting Explained covers the other half of the scoring equation.

Setting use_mobile to Route Requests Through a Carrier Exit IP

The core change is one boolean. The scraping API exposes use_mobile on the sync scrape endpoint, which routes your request through a mobile carrier proxy instead of the default datacenter pool. It costs +4 tokens per request versus the default route, so treat it as a targeted tool, not a default.

Here is the configuration as a JSON payload, which is how I keep scrape configs in a pipeline repo:

{
  "url": "https://store.example.com/catalog",
  "use_mobile": true,
  "formats": ["rawHtml"],
  "use_antibot": true,
  "timeout": 120
}

And the Python version, which submits the same request and prints the status plus the exit IP the request actually used:

import requests

API = "https://api.finedata.ai/scrape"  # POST /api/v1/scrape
HEADERS = {"Authorization": "Bearer fd_your_api_key"}

payload = {
    "url": "https://store.example.com/catalog",
    "use_mobile": True,
    "formats": ["rawHtml"],
    "use_antibot": True,
    "timeout": 120,
}

resp = requests.post(API, json=payload, headers=HEADERS, timeout=180)
data = resp.json()

print("status:", data.get("status"))
# The response body includes the exit IP used for the request.
# Parse it -- never assume it.
exit_ip = data.get("exit_ip") or data.get("data", {}).get("exit_ip")
print("exit_ip:", exit_ip)

One detail matters here: use_antibot stays on. The carrier IP fixes your network-level reputation; the antibot layer keeps the browser-like TLS fingerprint so the request doesn’t look like a raw requests call that happens to come from a phone network. You need both halves consistent. A mobile IP carrying a Python-requests TLS handshake is as contradictory as a desktop UA on a mobile IP.

Matching User-Agent and Viewport to the Mobile Exit IP

The carrier exit IP is necessary but not sufficient. Every layer of the request has to agree that this is a phone, because risk engines cross-check exactly these consistencies. If your exit IP is carrier-grade but your user-agent says Windows Chrome, you have built a fingerprint that no real device could produce. That contradiction is itself a detection signal — often a stronger one than either fact alone.

So pair use_mobile with an explicit mobile user-agent and mobile-appropriate headers:

payload = {
    "url": "https://store.example.com/catalog",
    "use_mobile": True,
    "use_antibot": True,
    "formats": ["rawHtml"],
    "headers": {
        "User-Agent": (
            "Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) "
            "AppleWebKit/605.1.15 (KHTML, like Gecko) "
            "Version/17.0 Mobile/15E148 Safari/604.1"
        ),
        "Accept": "text/html,application/xhtml+xml",
        "Accept-Language": "en-US,en;q=0.9",
        "Viewport": "width=device-width, initial-scale=1.0",
        "Sec-CH-UA-Mobile": "?1",
    },
}

If the site serves a responsive layout, the Viewport and client-hint headers matter when JS rendering is involved — the rendered DOM differs between a mobile and desktop viewport, and some sites only hydrate their mobile layout when the viewport says so. Add use_js_render: True when the catalog is client-side rendered.

The before/after on a strict product page makes the consistency argument concrete:

ConfigurationExit IP typeUser-AgentResult
DefaultDatacenterDesktop Chrome403 challenge
Mobile IP, mismatched profileCarrierDesktop Chrome403 challenge (inconsistency flagged)
Mobile IP, partial matchCarrierMobile UA, desktop viewport200, but desktop DOM served
Fully consistent mobile profileCarrierMobile UA + mobile viewport200, mobile DOM served

Row two is the one people skip and then blame the proxy. The mismatch is detectable and gets detected. Row three is sneakier: you get a 200, but the site serves you the desktop layout because your viewport hints said desktop, and your extraction selectors — written against the mobile DOM you saw on your phone — silently return nothing. Fix Scraped Data That Lands in the Wrong Fields covers that failure mode in detail.

Verifying the Fix: Measuring Response Quality Before and After

Never trust a fix you haven’t measured. “It seemed to work” is how a pipeline ends up storing 14 KB of challenge HTML for three weeks before anyone notices. Run both routes against the same URL and compare content length, the title tag, and a CAPTCHA marker before you commit the change.

import re
import requests

API = "https://api.finedata.ai/scrape"
HEADERS = {"Authorization": "Bearer fd_your_api_key"}

MOBILE_UA = (
    "Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) "
    "AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Mobile/15E148 Safari/604.1"
)
DESKTOP_UA = (
    "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
    "AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36"
)

CAPTCHA_MARKERS = ("cf-challenge", "g-recaptcha", "h-captcha", "px-captcha")

def scrape(url, use_mobile, user_agent):
    payload = {
        "url": url,
        "use_mobile": use_mobile,
        "use_antibot": True,
        "formats": ["rawHtml"],
        "headers": {"User-Agent": user_agent},
    }
    r = requests.post(API, json=payload, headers=HEADERS, timeout=180)
    r.raise_for_status()
    return r.json()

def profile(result):
    html = str(result.get("data", {}).get("rawHtml", ""))
    title = re.search(r"<title[^>]*>(.*?)</title>", html, re.S)
    return {
        "http_status": result.get("status"),
        "title": (title.group(1).strip() if title else "")[:60],
        "content_bytes": len(html),
        "captcha_present": any(m in html for m in CAPTCHA_MARKERS),
    }

url = "https://store.example.com/product/123"
desktop = profile(scrape(url, use_mobile=False, user_agent=DESKTOP_UA))
mobile = profile(scrape(url, use_mobile=True, user_agent=MOBILE_UA))

for name, p in (("desktop", desktop), ("mobile", mobile)):
    print(f"{name:8} status={p['http_status']} bytes={p['content_bytes']:>7} "
          f"captcha={p['captcha_present']} title={p['title']!r}")

The output tells you unambiguously whether the route change did anything:

RouteHTTP statusTitleContent bytesCAPTCHA marker
Desktop default403”Just a moment…“14,203true
use_mobile: true200”Anodized Widget — Store Example”212,447false

A 15x difference in content size plus a real product title is the signature of a genuine page. If the mobile route returns similar byte counts to the desktop challenge response, the site challenges mobile traffic too, and you need a different escalation path — heavier stealth options or, better, avoiding the CAPTCHA entirely rather than paying to solve it.

Put this comparison script in your CI or your pipeline’s pre-flight checks. When the site changes its policy six months from now, you want the alarm to fire before your warehouse fills with interstitials.

Handling Pagination and Session Stickiness on Mobile Routes

Once the mobile route works for a single page, the next problem is crawling a category without every page request landing on a different carrier IP. Some sites bind your session to the first IP they see; a new IP mid-session with the same cookies looks like session hijacking and gets you challenged again. The fix is session_id, which pins all requests carrying the same ID to one exit IP, with session_ttl controlling how long that binding lives (default 1800 seconds, max 86400).

import re
import uuid
import requests

API = "https://api.finedata.ai/scrape"
HEADERS = {"Authorization": "Bearer fd_your_api_key"}

MOBILE_UA = (
    "Mozilla/5.0 (Linux; Android 14; Pixel 8) "
    "AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Mobile Safari/537.36"
)

session_id = str(uuid.uuid4())  # one session, one exit IP, whole crawl
base = "https://store.example.com/c"

page = 1
seen_urls = set()

while True:
    payload = {
        "url": f"{base}?page={page}",
        "use_mobile": True,
        "use_antibot": True,
        "formats": ["rawHtml"],
        "session_id": session_id,
        "session_ttl": 3600,
        "headers": {"User-Agent": MOBILE_UA},
        "only_main_content": True,
    }
    r = requests.post(API, json=payload, headers=HEADERS, timeout=180)
    r.raise_for_status()
    html = str(r.json().get("data", {}).get("rawHtml", ""))

    if not html or "Just a moment" in html:
        print(f"page {page}: challenge or empty, stopping")
        break

    # Follow rel=next in the mobile-rendered markup
    next_link = re.search(r'<link[^>]+rel=["\']next["\'][^>]+href=["\']([^"\']+)["\']', html)
    if not next_link:
        # Fallback: some mobile templates use an anchor instead
        next_link = re.search(r'<a[^>]+rel=["\']next["\'][^>]+href=["\']([^"\']+)["\']', html)

    print(f"page {page}: {len(html)} bytes")
    if not next_link:
        break

    next_url = next_link.group(1)
    if next_url in seen_urls:
        break  # pagination looped -- stop
    seen_urls.add(next_url)
    page += 1

Two things in that loop deserve comment. First, the rel=next detection tries both <link> and <a> forms because mobile templates are often a different codebase from the desktop site — a separate mobile theme, not a responsive layout — and the markup is not guaranteed to match what you saw in a desktop browser. Second, the loop guard on seen_urls matters because some mobile templates paginate with JavaScript state instead of URLs and will happily serve the same page forever if you trust the page counter alone.

If you are running this at volume, move it to the async batch endpoint (POST /api/v1/async/batch) with a callback_url, so you are not holding 120-second timeouts open per page. The session mechanics work the same way across batch jobs. Keep One Exit IP Across Multi-Step Scrape Requests goes deeper on session warmup patterns.

When Not to Use a Mobile Exit IP: Costs and Trade-offs

Here is where I will argue against the tool this post just taught you. use_mobile costs +4 tokens per request on top of the base, and for many workloads it buys nothing. Carrier IPs are the right answer to a specific problem — a site that challenges desktop/datacenter traffic but serves mobile traffic — and an expensive habit everywhere else.

Crawl scenarioRecommended exit routeWhy
Low-volume detail pages on a strict siteMobile carrierThe +4 token cost is trivial at dozens of pages/day
High-volume sitemap crawl, lenient siteDatacenter (default)No reputation problem to solve; +4 tokens x 500K pages is pure waste
Geo-restricted contentproxy_country + residentialYou need a specific country, not a carrier ASN
Multi-step flow needing one warm IPsession_id + residential, stickyCarrier NAT makes stickiness less predictable at high volume
Desktop-only web appDatacenter + use_js_renderMobile route may serve a stripped-down template

That last row is the trap nobody warns you about. Some sites serve less data in their mobile HTML than the desktop version — deliberately. A common case: the target site hides the full specification table and customer reviews behind an accordion that only renders in the desktop DOM, and the mobile template simply omits the markup. Your scraper gets a clean 200, extracts the title and price, and silently never collects the specs — because the element does not exist in the response.

Check before committing:

# Compare the same URL across routes before choosing one
desktop_html = scrape("https://store.example.com/product/123", use_mobile=False)
mobile_html = scrape("https://store.example.com/product/123", use_mobile=True)

for marker in ('id="specifications"', 'class="reviews"', 'data-price-history'):
    print(f"{marker}: desktop={marker in desktop_html} mobile={marker in mobile_html}")

If the marker table shows the mobile route missing elements you need, the mobile route is the wrong tool even though it gets you past the challenge. In that case the better escalation is use_residential or the stealth options, which keep the desktop DOM while improving IP reputation.

My rule of thumb: default to the cheapest route that returns real content, measured by the verification script above. Escalate one step at a time — datacenter, then residential, then mobile — and only when the measurement shows the cheaper route failing. Pinning your whole pipeline to carrier IPs because one strict site needed it is how token budgets double overnight.

Wrap-up

The pattern is consistent: strict sites score requests across multiple signals at once, and the mobile path often carries a lighter policy because carrier NAT makes blocking it expensive for the site. To exploit that, you need the full profile to agree — use_mobile: true for the carrier exit IP, a mobile user-agent, mobile viewport hints, and use_antibot for the TLS layer. Any single mismatched layer can undo the rest.

The parts worth remembering:

  • Measure before and after. Content length, title tag, and a CAPTCHA marker will tell you the truth about any route change.
  • Use session_id for paginated crawls so the whole session shares one exit IP.
  • Check whether the mobile template actually contains the elements you need — a 200 with missing data is worse than a 403 you can see.
  • Save carrier IPs for the sites that force the issue. At +4 tokens per request, the default route should stay your default.

If the challenge layer is your real bottleneck rather than IP reputation, Handling CAPTCHAs When Web Scraping covers the solve-versus-avoid calculus. And if you’re weighing the sync approach here against batch jobs for a full catalog crawl, Async Scraping at Scale is the natural next read.

#mobile scraping #carrier ip #use_mobile #anti-bot #user agent #slot:api-capability

Related Articles