d48cd5b855
The `ignoreSslErrors` option stopped working during the v4 HTTP client interface rework: it was still folded into got-style request options, but those never reach `httpClient.sendRequest()`. In v3 the option works, and actors commonly expose it in their input schemas and pass it into crawler options (e.g. actor-scraper), so this keeps it working instead of removing it. The option is renamed to `ignoreTlsErrors`, matching `session.proxyInfo.ignoreTlsErrors`, the browser pool, and the impit client (the old name is dropped, documented in the upgrading guide; actors migrating to v4 rename it in their own code, the SDK does not touch this option). The crawler forwards it (still defaulting to `true`, same as v3) as a new `SendRequestOptions.ignoreTlsErrors` flag, which `BaseHttpClient` also enables for MITM proxy sessions (previously equally dead). The impit client honors the flag; for custom clients it is best effort, and `FetchHttpClient` cannot disable TLS verification at all. The dead plumbing in `getRequestOptions()` is removed and unit tests cover the forwarding chain.
@crawlee/impit-client
This package provides a Crawlee-compliant HttpClient interface for the impit package.
To use the impit package directly without Crawlee, check out impit on NPM.
Example usage
Simply pass the ImpitHttpClient instance to the httpClient option of the crawler constructor:
import { CheerioCrawler, Dictionary } from '@crawlee/cheerio';
import { ImpitHttpClient, Browser } from '@crawlee/impit-client';
const crawler = new CheerioCrawler({
httpClient: new ImpitHttpClient({
browser: Browser.Firefox,
http3: true,
ignoreTlsErrors: true,
}),
async requestHandler({ $, request }) {
// Extract the title of the page.
const title = $('title').text();
console.log(`Title of the page ${request.url}: ${title}`);
},
});
crawler.run([
'http://www.example.com/page-1',
'http://www.example.com/page-2',
]);