Cabeceras de solicitud
Metadatos en formato clave-valor enviados con cada solicitud HTTP que describen al cliente, el contenido aceptado y el contexto de la solicitud.
Las cabeceras de solicitud son los metadatos que un cliente envía con cada solicitud HTTP, como pares clave-valor antes del cuerpo. Informan al servidor sobre el cliente y la solicitud: User-Agent (el navegador), Accept y Accept-Language (qué contenido e idioma se desean), Referer (la página anterior), Cookie (estado de sesión) y muchas más. El servidor las usa para decidir cómo responder.
Para el scraping, las cabeceras son una superficie de detección importante. Los navegadores reales envían un conjunto específico y coherente de cabeceras en un orden concreto. Las librerías HTTP por defecto envían un conjunto escaso e inusual - a menudo solo un user-agent de bot y poco más. Los sistemas antibot comparan tus cabeceras con las que enviaría un navegador real y marcan las discrepancias.
Acertar con las cabeceras es más que poner un User-Agent de navegador. El conjunto completo debe ser coherente: las cabeceras Accept, Accept-Language, Accept-Encoding y Sec-* deben coincidir con el navegador que dices ser, en el orden en que ese navegador las envía. Un user-agent de Chrome con cabeceras que Chrome nunca enviaría es una señal evidente.
El anonimato también vive en las cabeceras. Un proxy transparente añade X-Forwarded-For revelando tu IP real; un proxy elite no envía ninguna de ellas. Cuando compruebas el anonimato de un proxy, el verificador inspecciona exactamente qué cabeceras reenvía el proxy al destino - la diferencia entre elite y transparente está escrita en las cabeceras de solicitud.
The metadata that arrives with every request
Headers accompany each HTTP request and describe the client and what it wants: which content types it accepts, which languages, which encodings, whether it is continuing a session, where it came from.
Real browsers send a specific set in a specific order, and that order is remarkably stable per browser and version. HTTP libraries send fewer headers in a different order, which is a cheap and reliable signal for anyone looking.
Headers also have to agree with everything else. A German exit address sending Accept-Language: ru-RU while the TLS handshake says Python is not a visitor from Germany, and no single one of those signals had to be wrong for the combination to be.
Getting headers right alongside the address
The proxy chooses where you appear to be. Headers have to tell the same story:
A coherent set
Accept-Language: en-US,en;q=0.9 matches login_c_US
Accept-Encoding: gzip saves money on metered products
User-Agent: generated by the browser, not typed by hand
Referer: present when a person would have arrived from somewhere- Send Accept-Encoding: gzip on residential. Compression is the largest single saving available on a metered product.
- Match Accept-Language to the proxy country. Sites use it as much as the address to decide what to serve.
- Do not hand-assemble browser headers. Use a real browser or an impersonation library that also gets the order right.
- Verify what actually arrives: fetch a header echo service through the proxy and read the result.
Header misconceptions
Editing headers does not disguise a library
Order and TLS fingerprint still identify it, and now they disagree with your headers.
More headers is not more human
Real browsers send a specific set. Extra ones stand out.
A proxy does not add or remove headers here
Ours add none. What arrives is what your client sent.
Referer is not always harmless
Sending one that could not exist is a contradiction like any other.
Términos relacionados
Ver esto en práctica
¿Listo para usar cabeceras de solicitud?
SotaProxy te da acceso a proxies residenciales rotativos, móviles, de centro de datos e ISP. Sin compromiso mínimo.
Empezar