Amazon Scraper API vs Proxy: Choose a Product Data Source
Distinguish authorized Amazon APIs, licensed product data and proxy-based page checks by record quality and access rights.
Read guideBy the end of this guide, you’ll have a concrete, repeatable playbook for:

By the end of this guide, you’ll have a concrete, repeatable playbook for:
This is a practical tutorial, not theory. Each step stands on its own, so you can copy and adapt the patterns into your stack.
Before you start, you should have:
curlpython 3.9+playwright (Python) with one browser installedIf you’re still deciding when to use residential proxies vs a VPN, see our in-depth guide: Residential proxies vs VPN: when data teams need more than a VPN tunnel.
To avoid bans with residential proxies, you first need to understand what modern anti-bot systems actually score and block.
Cloudflare’s 2026 bot guidance makes this very clear:
Other providers behave similarly. That means bans often come from behavior, not just IP lists.
Residential proxies help you:
But they do not override:
Common failure at this step: assuming “residential = safe” and cranking up concurrency without rethinking sessions and fingerprints. Proxies are one layer; behavior still matters.
The single biggest lever for ban avoidance is session design, not raw IP volume. With ProxyLane, you can choose between:
Use rotating IPs for:
Rotation helps you:
Use sticky IPs for:
If you rotate mid-flow, many sites will:
ProxyLane’s positioning aligns with this: session control as a first-class concept, not an afterthought. Map sticky vs rotating to your workflow explicitly.
Common failure at this step: using pure "rotating everything" for logins and then blaming IPs for bans. If identity changes mid-session, the application will reject you by design.
This step gives you a concrete Python baseline that you can reuse in scripts and AI agents.
requestsAssume you have ProxyLane credentials:
residential.proxylane.dev8000PL_USERPL_PASSimport requests
PROXY_HOST = "residential.proxylane.dev"
PROXY_PORT = 8000
PROXY_USER = "PL_USER"
PROXY_PASS = "PL_PASS"
proxies = {
"http": f"http://{PROXY_USER}:{PROXY_PASS}@{PROXY_HOST}:{PROXY_PORT}",
"https": f"http://{PROXY_USER}:{PROXY_PASS}@{PROXY_HOST}:{PROXY_PORT}",
}
headers = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) \
AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36",
}
resp = requests.get("https://httpbin.org/ip", proxies=proxies, headers=headers, timeout=30)
print(resp.status_code, resp.text)
time.sleep(random.uniform(1.5, 4.0))Common failure at this step: sharing one requests.Session across many threads and targets, which reuses cookies and headers in a way that creates suspicious patterns and cross-contamination.
Playwright’s architecture supports exactly the control points that proxy workflows need: global proxy, per-context proxy, and isolated browser contexts.
from playwright.sync_api import sync_playwright
PROXY_HOST = "residential.proxylane.dev"
PROXY_PORT = 8000
PROXY_USER = "PL_USER"
PROXY_PASS = "PL_PASS"
proxy_server = f"http://{PROXY_HOST}:{PROXY_PORT}"
proxy_credentials = {
"username": PROXY_USER,
"password": PROXY_PASS,
}
with sync_playwright() as p:
browser = p.chromium.launch(
headless=True,
proxy={
"server": proxy_server,
"username": PROXY_USER,
"password": PROXY_PASS,
},
)
# Each context = isolated session + cookies + cache
context = browser.new_context(
user_agent=(
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/122.0.0.0 Safari/537.36"
)
)
page = context.new_page()
page.goto("https://httpbin.org/ip", wait_until="networkidle")
print(page.text_content("pre"))
context.close()
browser.close()
route and wait_until carefully:
wait_until="networkidle") to avoid replays.ProxyLane’s own guidance recommends: route the browser through an authenticated proxy, align the session with its egress, and validate the target result before extracting records.
Common failure at this step: reusing a single context to visit many accounts or users, which creates a weird “super session” with mixed cookies and history that looks nothing like a normal browser.
Rotation strategy is where most teams either stay safe or get banned quickly.
Providers like SOAX emphasize rotation that preserves context; ProxyLane aligns with this: rotation should keep the behavioral fingerprint expected by the workflow while changing egress when needed.
Pseudocode for a ban-aware loop in Python:
import time
import random
import requests
# Assume get_proxy() returns a new ProxyLane rotating endpoint each call
def fetch_url(url):
proxies = get_proxy()
headers = {"User-Agent": random_ua()}
for attempt in range(3):
resp = requests.get(url, proxies=proxies, headers=headers, timeout=30)
if resp.status_code in (200, 201):
return resp
if resp.status_code in (403, 429):
# change identity and back off
proxies = get_proxy()
time.sleep(random.uniform(15, 45))
continue
# For 5xx, shorter backoff and retry
if 500 <= resp.status_code < 600:
time.sleep(random.uniform(5, 15))
continue
break
return None
Common failure at this step: rotating too aggressively (every single request) on flows that expect stable identity (e.g., long scroll, multi-step form), resulting in behavior that gets flagged even if IPs are clean.
Cloudflare’s docs highlight volumetric scraping detection based on ASN and fingerprint, plus patterns like repeated low-score requests. The goal is to avoid looking like bursty, uniform automation.
Avoid burst traffic:
Randomize intervals:
Vary fingerprints:
Respect target semantics:
Geo-align traffic:
Common failure at this step: running all jobs from one country or ISP because “it’s cheaper or easier,” creating a strange geographic concentration that stands out.
Residential proxies are increasingly used behind AI agents, no-code tools, and workflow engines (like n8n) that orchestrate many small jobs.
ProxyLane maintains open-source “skills” that teach agents to set up and verify proxies in a provider-agnostic way. You can reuse the same principles:
For a single HTTP request node:
HTTP Method: GETURL: target URLOptions → Proxy:
HTTP Proxy: http://PL_USER:PL_PASS@residential.proxylane.dev:8000For multi-step flows:
Common failure at this step: pooling all proxy credentials into a single global config and letting many agents share them, which leads to shared cookies and cross-contamination of bans.
ProxyLane’s core viewpoint: measure proxies by the response your target returns and the cost per successful run, not just headline bandwidth or pool size.
Check IP and geo:
https://ipinfo.io or similar via your proxy.Check response quality:
Track key metrics per target:
Relate cost to outcomes:

For ban avoidance, pool size alone is not decisive; session behavior, geo precision, and rotation design matter more.
Common failure at this step: blaming proxies for bans without tracking any metrics. If you don’t know whether 10% or 60% of requests are blocked, you can’t tune rotation or concurrency.
Not necessarily. 403/429 often mean:
Try:
If blocks persist across many IPs and sessions, investigate whether the target disallows automation entirely.
There is no universal number, but safe starting ranges are:
Use your metrics: if success drops sharply after 3 requests per IP, reduce that cap.
No. Rotating everything is a common mistake.
Use rotating:
Use sticky:
Modern anti-bot systems expect continuity of identity; if you change IP mid-session, you often get flagged.
Signals include:
Mitigations:
Yes. A key differentiator of ProxyLane is that unused residential traffic never expires. You can:
This is especially valuable for teams doing periodic experiments or irregular enrichment jobs.
Avoiding bans with residential proxies is less about “more IPs” and more about session design, rotation strategy, and realistic behavior. With ProxyLane’s non-expiring traffic, granular geo/ISP targeting, and developer-first integrations (Python, Playwright, AI agents), you can:
Start small, measure carefully, and let the data guide how you tune rotation, concurrency, and session policies.
Scale your workflows with the tools you already use. Rotating or sticky sessions up to 24 hours.
Distinguish authorized Amazon APIs, licensed product data and proxy-based page checks by record quality and access rights.
Read guideChoose the right anti-detect browser proxy guide, map the shared verification steps, and keep account access separate from the network route.
Read guideConnect a buyer-owned proxy to an Apify Actor, keep the session boundary clear, and validate records instead of counting requests.
Read guide