# How many GB of proxies do you need? We measured pages per GB

Source: https://ipvolt.com/blog/how-many-gb-of-proxies-do-i-need
Markdown: https://ipvolt.com/blog/how-many-gb-of-proxies-do-i-need.md

[Home](https://ipvolt.com/index.md) / [Blog](https://ipvolt.com/blog.md) / How Many GB of Proxies Do You Need? Pages per GB, Measured

Category: Benchmark
Published: 2026-10-05
Updated: 2026-10-05
Author: ipvolt
Reading time: 8 minutes
Tags: proxies, choosing

We loaded 34 public pages three ways through a metering proxy. 1 GB fits about 44,000 HTML-only fetches, or about 900 full Playwright page loads.

1 GB holds about 44,000 pages if you fetch only the HTML with an HTTP client, and about 900 if you load the same pages in a browser. How you fetch changed the answer about 49 times over; the kind of site changed it far less.

We loaded 34 public pages on 5 October 2026, three times in each of three modes, through a local metering proxy. The bytes are counted on the wire in both directions, so they include TLS handshakes, HTTP headers, everything sent upstream and an estimate of the proxy's CONNECT exchange, not only page weight.

| How the page is fetched | Median per page | 90th percentile | Pages per GB, median page | Pages per GB, 90th percentile page |
| --- | --- | --- | --- | --- |
| HTTP client, HTML only (Requests 2.34.2) | 22.7 KB | 59.2 KB | 44,083 | 16,880 |
| Playwright 1.63.0, full load | 1.10 MB | 2.45 MB | 906 | 408 |
| Playwright, images, media and fonts blocked | 454.6 KB | 2.28 MB | 2,199 | 437 |

1 GB is 10^9 bytes. Each page's value is the median of its three loads, and the 90th percentile is the fourth-heaviest of the 34 pages. If your targets are heavy, plan with that column: blocking images barely moved it (2.45 MB to 2.28 MB), because the heaviest pages were heavy with scripts.

## Estimate your own GB

```text
GB per day = pages per day × bytes per page × attempts per page ÷ 1,000,000,000
```

A worked example: 50,000 pages a day in a full browser at the 1.10 MB median, with 1.2 attempts per page, is 50,000 × 1,103,256 × 1.2 ÷ 10^9 = 66.2 GB a day. The same job with an HTTP client at the 22.7 KB median is 1.4 GB a day. The 1.2 is an assumption for the example (one retry for every five pages), not something measured here.

Count attempts, not pages, because a retry downloads the page again; [Proxy retries: one lost response, two created jobs](/blog/proxy-retries-duplicate-jobs) covers when a retry is safe at all. And these medians describe our 34 pages, not yours. [Scrapescope](/guides/scraper-bandwidth-report), the free meter used for this post, measures your own pages without a proxy account.

## Pages per GB by page type

Median bytes per page, with pages per GB in brackets:

| Page type | HTML only | Full load | Images, media and fonts blocked |
| --- | --- | --- | --- |
| Docs and reference (7 pages) | 37.2 KB (26,912) | 589.6 KB (1,696) | 433.5 KB (2,306) |
| News articles (5) | 51.7 KB (19,355) | 1.32 MB (758) | 2.24 MB (445) |
| Blog posts (5) | 24.9 KB (40,202) | 601.1 KB (1,663) | 444.3 KB (2,250) |
| Forum threads (3) | 20.8 KB (48,167) | 2.45 MB (408) | 2.26 MB (441) |
| Shop category pages (5) | 20.0 KB (50,042) | 1.33 MB (751) | 388.6 KB (2,573) |
| Shop product pages (5) | 21.3 KB (46,983) | 1.16 MB (858) | 347.9 KB (2,874) |
| JavaScript apps (4) | 11.1 KB (90,009) | 1.39 MB (717) | 751.0 KB (1,331) |

The groups are small, so read them as examples. The shop pages are scraping sandboxes and a demo store, because we left out retailers whose terms forbid automated access; real retail pages may be heavier. For the JavaScript apps, the 11 KB is the HTML document alone. An HTTP client fetches that one document and nothing a script would load, and we measured bytes, not whether the data you need is in that HTML.

## What changed the bytes most

In order of effect in this data:

1. **Browser or HTTP client.** A full browser load moved a median of 31 times the bytes of the HTML fetch for the same page, from 2.1 times on a Python documentation page to 195 times on a JavaScript-rendered demo page.
2. **Scripts, more than images.** Scripts were 49% of all full-load bytes; images, media and fonts together were 35%. The three forum threads stayed between 2.2 MB and 2.6 MB in both browser modes.
3. **Other sites' hosts.** Half of the full-load bytes (49.9%) went to hosts outside the page's own site. The median page sent 38% there, and 7 pages sent nothing. The largest single host was `www.googletagmanager.com`: it appeared on 17 of the 34 pages at about 308 KB per load, 12% of all full-load bytes. One news article contacted 32 hosts.
4. **Blocking images, media and fonts.** The median saving was 28%. It was over 50% on 10 pages, up to 84%, and under 10% on 6. Three of those 6 moved more bytes with blocking on. One news article pulled 1.14 MB from `www.gstatic.com` in two of its three blocked loads, a host that was not among its three largest in any full load. On another, a YouTube embed moved 1.9 MB to 2.0 MB with blocking and 1.2 MB without. Blocking changes what a page does, so measure a second run before relying on it; [Playwright's network guide](https://playwright.dev/python/docs/network#abort-requests) documents the route call used here, and [Playwright proxy setup](/guides/playwright-proxy-setup) covers the proxy side.

Compression is not ranked because it was on in every run: all 34 HTML responses came back compressed (17 Brotli, 14 gzip, 3 Zstandard). The median HTML document was 83.5 KB after decoding, against 22.7 KB on the wire for the whole fetch. A client that does not ask for compression downloads closer to the decoded size; see [which clients skip compression](/guides/curl-compressed-accept-encoding).

## Upload share and tunnel overhead

| Mode | Upstream share of all bytes | Median upstream bytes per page |
| --- | --- | --- |
| HTTP client, HTML only | 6.3% | 1.9 KB |
| Playwright, full load | 2.7% | 28.9 KB |
| Playwright, blocked | 2.9% | 17.3 KB |

Upload is a small share overall and a large one on small pages: up to 21% of a single HTML fetch, because the TLS handshake and the request cost the same however small the reply is. Check whether your provider counts both directions.

The estimated CONNECT exchange added a median of 109 bytes to an HTML fetch and 2.6 KB to a full browser page, one exchange per tunnel. The browser fetched nothing in the background by itself: scrapescope's background, idle-preconnect and unattributed buckets were 0 bytes in all 204 browser runs. That is Playwright's bundled Chromium with default launch settings; other browser builds were not tested.

## Cost per 1,000 pages

Median page, with the 90th percentile page in brackets:

| Price per GB | HTML only | Full load | Images, media and fonts blocked |
| --- | --- | --- | --- |
| $1.00 | $0.02 ($0.06) | $1.10 ($2.45) | $0.45 ($2.28) |
| $3.50 | $0.08 ($0.21) | $3.86 ($8.57) | $1.59 ($7.99) |
| $5.00 | $0.11 ($0.30) | $5.52 ($12.24) | $2.27 ($11.42) |
| $10.00 | $0.23 ($0.59) | $11.03 ($24.48) | $4.55 ($22.83) |

$3.50 is ipvolt's listed price: the [pricing section](https://ipvolt.com/#pricing) shows "From $3.50/GB", and the FAQ says "Traffic starts from $3.50/GB at launch." The other three are round reference points, not anyone's price list; for market prices see [what unlimited residential plans cost per GB](/blog/unlimited-residential-proxies-cost-per-gb). Each figure is bytes ÷ 10^9 × price. It estimates transfer and is not a bill.

## Compared with the 2025 Web Almanac

The [Page Weight chapter](https://almanac.httparchive.org/en/2025/page-weight) of the 2025 Web Almanac says "The median home page in 2025 was 2.86 MB on desktop and 2.56 MB on mobile." Its inner-page chart ends at 1,963 KB on desktop and 1,769 KB on mobile in July 2025, and its median HTML response is 35 KB on desktop and 33 KB on mobile.

The chapter defines page weight as "the total volume of bytes transferred to a user’s device" and notes that its JavaScript figures are "bytes for compressed JavaScript files", so these are transfer sizes, not uncompressed sizes. It does not repeat that statement next to the HTML figure.

Our full-load median of 1.10 MB is below the Almanac's inner-page medians. The set leans on documentation and sandbox pages and is mostly inner pages. The HTML figures are close: 22.7 KB on the wire here, 33 KB to 35 KB there. The two count different things. The Almanac counts what a page's responses weigh; this test counts every byte through the tunnel in both directions.

## Method and limits

- **When and where.** 5 October 2026, 08:37 to 09:04 UTC, from a server in Helsinki, Finland, on a direct connection. No paid proxy was used.
- **Tools.** scrapescope 0.1.0 in direct sizing mode as the local metering proxy, Python 3.14.4, Requests 2.34.2 with urllib3 2.8.0, and Playwright 1.63.0 with its bundled Chromium headless shell 153.0.8010.12.
- **Each load.** A new process, a new browser and an empty cache, with no cookies and no consent given. The browser used a 1280×900 viewport, waited for the load event and then up to 10 seconds for the network to go idle. It did not scroll or click, so content that loads on scroll is not counted. The HTTP client sent one GET and followed redirects.
- **Pages.** 34 URLs that robots.txt allowed, with no logins and no paywalls. One forum candidate was dropped before measuring because its site terms forbid automated access. All 306 loads returned HTTP 200.
- **Variation.** The median gap between a page's heaviest and lightest load was 0.05% for HTML, 0.2% for a full load and 0.3% with blocking. The largest was 54%, on the news article described above.
- **Limits.** A direct connection reproduces the transport, not a proxy's exit location, blocks or retries; the scrapescope README calls sizing mode "a lower bound" for protected sites. The CONNECT bytes are estimated. Shares by resource type are allocated by the tool from host totals, not measured per request. One location, one day, and pages change.

Downloads: the [URL list](https://ipvolt.com/downloads/how-many-gb-of-proxies-do-i-need/urls.csv), the [raw CSV](https://ipvolt.com/downloads/how-many-gb-of-proxies-do-i-need/results.csv) with all 306 runs, the [computed summary](https://ipvolt.com/downloads/how-many-gb-of-proxies-do-i-need/summary.json), and the [measurement scripts](https://ipvolt.com/downloads/how-many-gb-of-proxies-do-i-need/pages-per-gb-lab.zip) with a README.

ipvolt is building proxy infrastructure for developers and agents. [Join the waitlist](https://ipvolt.com/#waitlist-hero) for one email when access opens.

## Sources

- [Web Almanac 2025: Page Weight (HTTP Archive)](https://almanac.httparchive.org/en/2025/page-weight)
- [scrapescope README at the commit used: sizing mode, what is counted, limits](https://github.com/ipvolt/scrapescope/blob/e33e8fb90b28ea32fba92cd26e70aef474e2eccd/README.md)
- [Playwright for Python: network, aborting requests with route()](https://playwright.dev/python/docs/network#abort-requests)
- [Playwright for Python: Request.resource_type](https://playwright.dev/python/docs/api/class-request#request-resource-type)
- [Requests: response content and automatic decoding of compressed bodies](https://requests.readthedocs.io/en/latest/user/quickstart/#response-content)

## Know when ipvolt access opens.

ipvolt is in development. Leave your email and we’ll notify you once when access opens.

Consent: One email when access opens. Nothing else.

[Notify me](https://ipvolt.com/blog/how-many-gb-of-proxies-do-i-need#waitlist-blog-end). Use the email form on this page to join the interest list.

[Privacy](https://ipvolt.com/privacy)

## Related posts

- [Unlimited residential proxies: what they cost per GB](https://ipvolt.com/blog/unlimited-residential-proxies-cost-per-gb.md) (Comparison, Sep 27, 2026, 18 min read): Turn Mbps- and thread-based unlimited residential proxy plans into cost per GB, find the break-even with your per-GB rate, and check terms before you pay.
- [Multiple Meta Ad Accounts: An Agency Setup Checklist](https://ipvolt.com/blog/manage-multiple-meta-ad-accounts.md) (Analysis, Sep 19, 2026, 6 min read): Manage multiple Meta ad accounts with a client access record, clear team responsibilities and a checklist for access errors, security incidents and routing.
- [AI agent proxies and compute: choose by responsibility](https://ipvolt.com/blog/ai-agent-proxies-and-compute.md) (Analysis, Sep 17, 2026, 6 min read): Map an AI agent's retrieval, runtime and model requirements before choosing services, with a release-note monitoring example and editable worksheet.

## Related guides

- [Configure a proxy in Playwright](https://ipvolt.com/guides/playwright-proxy-setup.md): Set an authenticated HTTP proxy for a Playwright browser context, keep test state isolated, and diagnose navigation separately from subresources.
- [curl --compressed: which clients skip compression](https://ipvolt.com/guides/curl-compressed-accept-encoding.md): curl, Wget 1.x, urllib, node:http, PHP curl, Guzzle and java.net.http.HttpClient send no Accept-Encoding (or identity) by default. Tested fixes and byte checks.
- [Python Requests proxy: proxies dict, auth, SOCKS5](https://ipvolt.com/guides/python-requests-proxy.md): Configure a proxy in Python Requests: the proxies dictionary, Session defaults, credentials, SOCKS5 via requests[socks], environment variables and ProxyError.

## About ipvolt

Technical analysis from the ipvolt team.

ipvolt access is not open yet.
