Referral Program →
HomeGlossaryHeadless Browser
Glossary

Headless Browser

A real web browser that runs without a graphical interface, controlled by code to load, render, and interact with pages.

A headless browser is a full browser - usually Chrome or Firefox - running without a visible window, driven programmatically. It loads pages, executes JavaScript, and renders the DOM exactly as a normal browser would, but under the control of a script. Playwright, Puppeteer, and Selenium are the common tools for driving them.

Headless browsers exist to scrape modern sites that build their content with JavaScript. A plain HTTP request returns the initial HTML, which for many single-page apps is nearly empty - the real data loads afterward via JavaScript. A headless browser runs that JavaScript, so you can scrape the fully rendered page instead of an empty shell.

The trade-off is cost and detectability. Rendering a full page uses far more CPU, memory, and bandwidth than a simple HTTP request, so headless scraping is slower and more expensive at scale. Headless browsers also carry automation fingerprints (navigator.webdriver flags, unusual WebGL renderers) that anti-bot systems specifically look for.

Use a headless browser when the target genuinely requires JavaScript rendering, and pair it with residential proxies and fingerprint hardening so it looks like a real user's browser. For static HTML or available APIs, plain HTTP requests are dramatically faster and should be preferred.

A real browser with no window

A headless browser is a normal browser engine running without a visible interface, driven by code through Playwright, Puppeteer or Selenium. It executes JavaScript, renders pages and holds cookies exactly as the visible version does.

That makes it the tool for sites that build their content in the browser rather than on the server. An HTTP client fetching such a page receives a shell with no data in it.

The cost is traffic and time. A headless browser downloads scripts, styles, fonts and images unless told otherwise, so a page that would be 80 KB as HTML becomes several hundred kilobytes, which matters directly on metered proxies.

Headless mode also used to be detectable through obvious signals. Modern versions are much closer to the visible browser, but automation frameworks still leave traces that dedicated defences look for.

Running one through our proxies

Per-context proxies are the feature worth having, because they let one browser hold several identities:

Playwright with a proxy per context

ctx = browser.new_context(
    proxy={
        "server": "http://proxy.sotaproxy.com:10000",
        "username": "login_c_US_s_7_ttl_1h",
        "password": "password",
    },
    locale="en-US", timezone_id="America/New_York",
)
  • Block images, media and fonts through request routing. On residential this cuts the bill several times over.
  • Use a sticky session for the life of the context. A rotation mid-page loads half the assets from another address.
  • Match locale and timezone to the proxy country, or the address and the browser will contradict each other.
  • For account work prefer an antidetect browser over raw headless automation: it adds fingerprint and storage isolation you would otherwise build yourself.

Headless misconceptions

Headless is not required for most scraping

If the data is in the HTML or an API, an HTTP client is faster and much cheaper.

It does not hide automation by itself

Framework traces remain, and defences look for them specifically.

Traffic is not comparable to curl

A rendered page can be five times the bytes of its HTML.

A proxy does not fix detection

The address and the browser are separate signals and both are checked.

Ready to use headless browser?

SotaProxy gives you access to rotating residential, mobile, datacenter, and ISP proxies. No minimum commitment.

Get started