Amazon Scraper API vs Proxy: Choose a Product Data Source
Distinguish authorized Amazon APIs, licensed product data and proxy-based page checks by record quality and access rights.
Read guideFetch HTML in an existing browser session to cut proxy traffic. Understand the Botasaurus 97% claim, missing-data risks, and cost per valid record.

Browser-based fetch requests can reduce proxy costs by retrieving HTML inside an existing browser session instead of navigating and loading each page's assets. The saving applies when the response contains the required data. Confirm it with matching validated records and provider-billed traffic; Botasaurus's reported 97% cost reduction is a case-specific claim.
Audience: scraping developers who understand browser navigation, HTML extraction, and proxy traffic billing. The decision is whether a collection job needs a rendered page for every record or can reuse a browser session for document requests.
A normal navigation loads the document and may start requests for stylesheets, scripts, fonts, images, and application data. Blocking images removes one resource category, while the browser still processes the page and loads whatever remains allowed. Repeating that process across a large catalog can transfer far more data than the records require.
A browser fetch request uses the Fetch application programming interface (API) to retrieve a response without navigating the tab. Reading its HTML as text does not render that document or load its referenced assets. A collector can parse the response separately while keeping the browser open on the target site.
The useful boundary is data already in HTML. Product identifiers, titles, and prices present in the response can be extracted without rendering. Fields added later by page JavaScript require another data source or a rendered-page path. A smaller response with a missing price is a failed collection result.
| Collection path | What it retrieves | Suitable when |
|---|---|---|
| Full navigation | Document plus allowed page resources | Data requires JavaScript execution or interaction |
| Navigation with resource blocking | Document plus the resources left enabled | Rendering is required, but some assets are unnecessary |
| Fetch from an existing browser session | Requested response body | Required fields are present in HTML or an authorized data response |
| Standalone HTTP client | Requested response outside the browser | The target works without browser-established state |
Browser fetch and a standalone client have different execution contexts. A request made inside the browser follows browser security and credential rules. Moving the same URL into a separate Python client does not automatically preserve those properties.
The Botasaurus example describes roughly 100,000 pages. Its author reports 250 GB and about USD 1,050 before the change, despite image blocking, followed by 5 GB and USD 30 with browser fetch. These are upstream reported figures, checked on 2026-10-07, rather than a ProxyLane benchmark.
| Quantity | Reported before | Reported after | Reduction calculated from those figures |
|---|---|---|---|
| Proxy traffic (GB) | 250 | 5 | 98.00% |
| Proxy expense (USD) | 1,050 | 30 | 97.14% |
The README gives USD 4.20/GB for the original run. At that same rate, 5 GB would cost USD 21; the reported USD 30 implies USD 6.00/GB. The source does not reconcile the difference or provide a reproducible dataset and billing ledger. Its headline is useful motivation, but 97% is not a forecast for another workload.
The implementation opens the site for a new browser, enables browser reuse, and retrieves subsequent responses with driver.requests.get. It then parses response HTML with soupify. The upstream code also fetches the first URL after its initial navigation, so setup traffic belongs in the measurement.
The browser can retain relevant cookies between navigation and fetch. Under the Fetch credential rules, same-origin requests include credentials by default, subject to cookie rules. Cross-origin requests depend on credentials configuration and cross-origin resource sharing (CORS). Browser fetch does not grant access to an arbitrary authenticated domain.
In the inspected Botasaurus Driver implementation, driver.requests.get runs window.fetch with explicit credentials: "include" and mode: "cors" options. The credentials option differs from bare fetch's same-origin default; cookie restrictions and CORS still apply. Check the installed driver version before relying on the wrapper's behavior.
Preserve the collection context during a comparison: the browser profile, proxy route, account, locale, and target origin. Check an exit-IP endpoint through the same browser before the sample. That observation confirms the route for that check; it does not prove that every request uses the same IP or that the target accepts it. The sticky-session guide explains route continuity separately from browser cookies.
Reading response text does not execute the fetched page's JavaScript. Nor does browser fetch remove server rate limits or guarantee that a target accepts document requests after navigation. A site can return a login page, challenge, or partial record while the transport succeeds.
A useful pilot compares a fixed, permitted URL sample under both collection paths. Include an ordinary product, a variant, and a page known to populate data after loading. Define required fields and acceptable freshness before collecting, so the cheaper path cannot pass by dropping difficult records.
Warning: An external pilot can consume paid proxy quota and use authenticated sessions. Start with a sandbox or permitted low-volume sample, cap total attempts and concurrency, and keep credentials out of code and logs. Stop the job if the budget or validation rule fails; spent traffic cannot be restored, and the original collection path is the fallback.
Record these values for each mode:
| Measurement | What belongs in the record |
|---|---|
| Validated output | Distinct record IDs and required fields; missing or changed values |
| Attempts | Initial navigation, fetches, redirects, retries, and fallback navigations |
| Browser traffic | Transferred bytes with the same cache policy in both modes |
| Billed traffic | Provider usage before and after the run, with metering delay allowed |
| Other expense | Browser time, compute, and any request or platform fee |
Use cost per accepted record as the decision metric. For each mode, divide the run's attributable cost by its distinct validated records. Calculate the percentage reduction as 100 × (1 - fetch_cost_per_record / navigation_cost_per_record); the baseline and accepted-record count must both be greater than zero.
Browser-side bytes help diagnose the saving. Chrome DevTools distinguishes transferred and uncompressed resource sizes; HTML string length is not either a complete wire measurement or a provider invoice. Reconcile dashboard usage with the proxy bandwidth billing method, including failed attempts and setup traffic.
| Symptom | Check | Response |
|---|---|---|
| Required price or stock is missing | Compare response HTML with the rendered record | Keep rendering or use an authorized endpoint supplying that field |
| Login or challenge content appears | Inspect final URL, content, and session expiry | Re-establish the session within the attempt budget; reject the response as data |
| Browser blocks response access | Inspect origin, credentials, and CORS errors | Use a permitted same-origin path or a supported authorized API |
| Target rate limit is reached | Inspect status and the server's retry instruction | Pause or stop within the budget; reduce concurrency |
| Provider usage barely falls | Include setup, retries, fallback traffic, and metering delay | Reconcile the whole run before increasing volume |
For HTTP 429, respect the server's Retry-After instruction when supplied. The upstream README suggests a fixed delay, but a target's limit must be established for that target. A universal requests-per-second rule would make the pilot unreliable.
Cookie resets also change the test conditions: deleting cookies can remove authentication and established session state. Diagnose a cookie-related failure before resetting the profile, and count any new navigation needed to recover.
If required fields are present and validated records match, browser fetch is a candidate for the repetitive document-fetching part of the job. Keep full navigation for actions that require rendering or interaction. Recheck after target markup, authentication, or data-loading behavior changes.
If rendering remains necessary, the Playwright resource-blocking guide covers reducing asset traffic while preserving required data. For the next bounded decision, compare the fixed sample's cost per successful result and provider usage before expanding the run.
Scale your workflows with the tools you already use. Rotating or sticky sessions up to 72 hours.
Distinguish authorized Amazon APIs, licensed product data and proxy-based page checks by record quality and access rights.
Read guideChoose the right anti-detect browser proxy guide, map the shared verification steps, and keep account access separate from the network route.
Read guideSeparate Australian egress from en-AU content, AUD pricing, GST display, postcode validation and the state delivery context.
Read guide