NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3519 most downloaded on PyPI
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & scraper
Last release 11 days ago
23 Sep 2026
Ships fairly regularly
a new release about every 3 weeks
Most releases are documented
notes for 37 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
72 releases · first in 2024
One column per month.
docker pull unclecode/crawl4ai:0.9.4 docker pull unclecode/crawl4ai:latest
PyPI:
pip install crawl4ai==0.9.4Docker:
docker pull unclecode/crawl4ai:0.9.4
docker pull unclecode/crawl4ai:latestNote: Docker images are being built and will be available shortly.
Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
0.9.4 is a security release. It closes three coordinated-disclosure advisories: two SSRF paths that bypassed the Docker server's egress controls, and a trust-boundary bypass that let a non-admin API client read server environment variables. It also makes content pruning about 10x faster with the new lxml-native PruningContentFilterLXML, now the default, and ships the bug fixes that accumulated on develop since 0.9.3. There are no breaking changes. Users who self-host the Docker server should upgrade.
RobotsParser.can_fetch() fetched /robots.txt on a bare aiohttp client that followed redirects and re-resolved the host, so check_robots_txt in an untrusted request body could make the Docker server reach internal, loopback, and cloud-metadata addresses. The fetch now goes through the server's pinning egress proxy, which checks every hop and dials the pinned IP. Credit: arpe1618. (GHSA-f77g-77vp-r96v)link_preview_config (CWE-918, high): the URL seeder fetched every link on a crawled page with its own httpx client, outside the egress controls, and returned each page's parsed <head> to the caller. The seeder's fetches now go through the same pinning egress proxy, and LinkPreviewConfig gets caps on max_links, concurrency, and timeout for untrusted bodies. Credit: Ibrahim AlJaafreh (LinkedIn), Cystack RedTeam (cystack.ps). (GHSA-wh5w-hmj3-vgg7)LLMConfig in {"type": "dict", "value": {...}} slipped it past UNTRUSTED_ALLOWED_TYPES, and from_kwargs then rebuilt it as trusted. A non-admin client could read any server environment variable, including LLM keys and SECRET_KEY. The unwrapped value is now re-checked under the untrusted gate, and from_kwargs carries the caller's provenance instead of defaulting to trusted. Credit: Adam Jordan (adamyordan). (GHSA-5w5p-vcv6-mm3f)The two SSRF fixes share one mechanism: the new crawl4ai/egress_policy.py holds a process-wide egress proxy URL for the library's own HTTP clients. The Docker server registers its existing PinningProxy there at boot. A plain library caller sets nothing and sees no change.
All reporters are credited in SECURITY-CREDITS.md. GitHub Security Advisories accompany this release.
PruningContentFilterLXML: an lxml-native pruning filter. It computes every per-node metric in one bottom-up pass instead of re-walking each subtree, so pruning is O(N) instead of super-linear. Output is byte-identical to PruningContentFilter. Measured pruning time: medium page 134 to 13 ms, 6000-card page 2200 to 260 ms. It is now the default for the Docker server's fit filter and the CLI pruning filter.CRAWL4AI_MAX_TIMEOUT_MS sets the ceiling for page_timeout, wait_for_timeout, and body_visibility_timeout on untrusted configs. The default stays 60000 ms. (#2212, thanks @damusix; #2266)crawler.pool.max_pages_before_recycle (default 200) recycles a pooled browser context after it serves that many pages. A context gets slower with sustained use, and the idle janitor never fires on a busy server. Set it to 0 to disable. (#2232, issue #2231)PruningContentFilter emits a DeprecationWarning on direct use. Switch to PruningContentFilterLXML, which takes the same arguments and gives the same output. Existing import paths keep working.Crawler and core
rowspan and colspan are expanded into a grid, and <th> row headers are kept instead of shifting the row left. Spans are clamped, so one cell cannot hang the parse. (#2261, issue #2258)Disallow: /*? no longer blocks the whole site. (#2229, thanks @Nalhin)Allow: precedence. (#2278)--headless=new on macOS arm64. OptimizationHints is no longer disabled. (#2241, issue #2239, thanks @Zsanz3)Docker server
code field that the untrusted boundary rejects. (#2262, issue #2260)md and llm runs skip the /config/dump pre-flight, which failed on the legacy code field. (#2224, issue #2222)Documentation and CI
SECURITY.md lists 0.9.x as supported. (#2269, thanks @nightcityblade)tests/unit/test_egress_policy.py and deploy/docker/tests/test_security_ssrf_seeder.py: seeder and robots.txt fetches through the egress proxy.tests/unit/test_config_provenance.py: direct, wrapped, and nested forbidden types refused under the untrusted gate.tests/async/test_content_filter_prune_lxml.py: PruningContentFilterLXML output matches PruningContentFilter.None.
docker pull unclecode/crawl4ai:0.9.3 docker pull unclecode/crawl4ai:latest
PyPI:
pip install crawl4ai==0.9.3Docker:
docker pull unclecode/crawl4ai:0.9.3
docker pull unclecode/crawl4ai:latestNote: Docker images are being built and will be available shortly.
Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
0.9.3 is a security release. It closes five coordinated-disclosure advisories in the PDF processing path and the Docker Playground UI, and ships the 33 bug fixes that accumulated on develop since 0.9.2, most of them in the Docker server. There are no new features and no breaking changes. Users who accept untrusted URLs on the Docker server, or who open PDFs from sources they do not control, should upgrade.
The PDF path was the common thread. PDFContentScrapingStrategy is selectable from an untrusted Docker API request body, and it fetches with requests outside the browser, so none of the Chromium-side egress or resource controls applied to it.
PDFContentScrapingStrategy had no field allowlist, so an untrusted request body could set save_images_locally and image_save_dir and make the server write extracted images to a path of the caller's choosing. Those fields are now filtered at the trust boundary and extract_images is forced off for untrusted bodies. Credit: Zhixi "Jace" Sun (manus-use). (GHSA-xpp7-j28w-2gvx)max_pdf_bytes (100 MiB default), enforced on the running total rather than the caller-supplied content-length, and parsing stops at max_pdf_pages (2000 default). Untrusted bodies cannot raise their own caps. The Docker config now ships a non-zero limits.wall_clock_s of 300 seconds. Credit: Nguyen Tran Thanh Lam (c240030). (GHSA-v2rm-hvrj-2x9q)cleaned_html (CWE-79, medium): paragraph text taken verbatim from a PDF was written into cleaned_html without escaping, so markup embedded in a PDF survived into the result and executed when rendered. Paragraph text is now escaped like every other sink in that function. Credit: Nguyen Tran Thanh Lam (c240030). (GHSA-7g3g-vhm6-79f3)element.innerHTML = element.textContent, which re-parsed attacker-controlled crawled content as live HTML in the operator's session. The round trip is removed; highlight.js renders safely from textContent. Credit: e1codes. (GHSA-m446-hp3q-qfxp)All reporters are credited in SECURITY-CREDITS.md. GitHub Security Advisories accompany this release.
This release also carries the bug fixes that accumulated on develop since 0.9.2.
Docker server
PDFContentScrapingStrategy are routed to PDFCrawlerStrategy so the pairing works without extra configuration. (#2094, #2150)HTTP_PROXY / HTTPS_PROXY instead of ignoring it. (#2142)CRAWL4AI_API_TOKEN is forwarded through compose, and .llm.env is optional rather than required. (#2094)hooks.code field from other rejections. (#2094)output_path is declared a deprecated no-op rather than silently ignored. (#2094)GET /monitor redirects to the dashboard UI. (#2157, issue #2091)mcp is capped below 2 so the v1 low-level API used by mcp_bridge keeps working. (#2148, thanks @weike-zhang)Crawler and core
ManagedBrowser no longer leaks a Playwright driver process when the browser fails to launch inside __aenter__. (#2160)PDFCrawlerStrategy placeholder responses are no longer vetoed as anti-bot blocks, which previously failed every PDF crawl and burned the retry budget. (#2138, issue #2135)setTimeout waits are removed from the overlay and consent removal scripts, which could hang a crawl under a restrictive CSP. (#2139)remove_overlay_elements no longer removes <body> when the body carries a global popup class. (#2163, thanks @Nalhin)verbose=False. (#2117, #2131, #2145, issues #2116, #2129, #2144)Documentation
PDFCrawlerStrategy plus PDFContentScrapingStrategy pairing requirement is documented.tests/unit/test_pdf_download_limits.py: 22 tests covering per-hop destination validation, DNS rebinding, redirect bounds, byte and page caps, and untrusted-body clamping.tests/unit/test_pdf_html_escaping.py: escaping of PDF paragraph text in cleaned_html.deploy/docker/tests/test_security_pdf_image_write.py: rejection of image-write fields from untrusted bodies.crawler_configs PDF guard.None.
docker pull unclecode/crawl4ai:0.9.2 docker pull unclecode/crawl4ai:latest
PyPI:
pip install crawl4ai==0.9.2Docker:
docker pull unclecode/crawl4ai:0.9.2
docker pull unclecode/crawl4ai:latestNote: Docker images are being built and will be available shortly.
Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
July 2026 - 2 min read
I'm releasing Crawl4AI v0.9.2, a small maintenance patch that cleans up a dispatcher resource leak and fixes a handful of Docker and GPU build issues. No new features, no breaking changes.
If you're on v0.9.1, upgrade freely.
MemoryAdaptiveDispatcher no longer leaves crawl tasks and browser pages running after a streaming crawl is closedENABLE_GPU=true Docker builds no longer fail on the CUDA toolkitMemoryAdaptiveDispatcher.run_urls_stream() was closed or cancelled mid-stream, its per-URL crawl tasks kept running in the background. A later arun_many() on the same crawler could then overlap the old tasks while they still held the same browser, contexts, and pages — leaking pages and emitting stray Playwright TargetClosedErrors. Closing the stream now cancels and awaits every in-flight task and drains any queued-but-unstarted URLs before returning. (#2071, thanks @reallav0)/config/dump requires a type, but the playground's pyConfigToJson only sent code. The request now includes the config type, aligns the stream fallback with the dump shape, and teaches shouldUseStream both nestings. (#2059, thanks @Pitchfork-and-Torch)/monitor/ws endpoint 500'd because a router-level token_dep is HTTP-request-only and can't inject into WebSocket scopes. Auth is now enforced solely by AuthGateMiddleware; admin routes keep require_admin. (#2060, thanks @Pitchfork-and-Torch)ENABLE_GPU=true builds failed with "Package has no installation candidate" because nvidia-cuda-toolkit lives in Debian Bookworm's non-free component, which the python:3.12-slim-bookworm base image omits. The Dockerfile now adds the non-free (and contrib) apt sources before the GPU install block. (#2020, thanks @harshmathurx)None.
pip install -U crawl4ai
crawl4ai-doctor # verify installation
Docker users: pull the latest image once the Docker release workflow finishes.
python docs/releases_review/demo_v0.9.2.py
Thanks to the community contributors who made this release possible: @reallav0 (#2071, #2067), @Pitchfork-and-Torch (#2059, #2060), @harshmathurx (#2020).
docker pull unclecode/crawl4ai:0.9.1 docker pull unclecode/crawl4ai:latest
PyPI:
pip install crawl4ai==0.9.1Docker:
docker pull unclecode/crawl4ai:0.9.1
docker pull unclecode/crawl4ai:latestNote: Docker images are being built and will be available shortly.
Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
July 2026 - 3 min read
I'm releasing Crawl4AI v0.9.1, a patch release that ships 12 bug fixes across Docker, browser, core, and extraction, plus one new feature for PruningContentFilter.
No breaking changes. If you're on v0.9.0, upgrade freely.
preserve_classes / preserve_tags parameters to protect specific elements from density-based pruningchannel='chromium' no longer crashes Playwright on Windowspage_timeout was passed in milliseconds to aiohttp (which expects seconds), effectively disabling timeouts in HTTP modePruningContentFilter's density-based scoring is great at stripping boilerplate, but it sometimes takes short metadata elements — author names, timestamps, attribution lines — along with it. The new preserve_classes and preserve_tags parameters let you whitelist specific CSS classes or HTML tags that should never be pruned, regardless of their density score.
from crawl4ai.content_filter_strategy import PruningContentFilter
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator
filter = PruningContentFilter(
threshold=0.48,
preserve_classes=["author", "byline", "dateline"],
preserve_tags=["time", "address"],
)
generator = DefaultMarkdownGenerator(content_filter=filter)
config = CrawlerRunConfig(markdown_generator=generator)
Whitelisted nodes skip scoring entirely. Default is empty sets — no behavior change for existing users. (#1900, thanks @hafezparast)
authFetch() with Bearer token. API routes remain fail-closed. (#2037)channel='chromium' (the default) caused Playwright to look for a system Chrome install instead of the bundled binary, crashing on Windows with TargetClosedError. The default channel is no longer passed to Playwright. (#2051, thanks @fstark96)page_timeout (60000ms) was passed directly to aiohttp.ClientTimeout which expects seconds, making the effective timeout 16.7 hours. Now correctly divided by 1000. (#1894, thanks @hafezparast)BestFirstCrawlingStrategy stabilized for deterministic crawl order. (#1998, thanks @nightcityblade)html2text now preserves all attributes on table tags when bypass_tables is enabled. (#2007)<6 to <7 so crawl4ai can co-install with packages requiring lxml 6.x (e.g. scrapling). (#2019)normalize_url duplicates and accidental adaptive_crawler copy. (thanks @RajanChavada)pip install -U crawl4ai
crawl4ai-doctor # verify installation
Docker users: pull the latest image once the Docker release workflow finishes.
Thanks to the community contributors who made this release possible: @hafezparast (#1894, #1900), @nightcityblade (#1998, #1999, #2025, #2027), @fstark96 (#2051), @TobiasWallura-xitaso (#2047), @harshmathurx (#2040), @RajanChavada (#2042).
docker pull unclecode/crawl4ai:0.9.0 docker pull unclecode/crawl4ai:latest
PyPI:
pip install crawl4ai==0.9.0Docker:
docker pull unclecode/crawl4ai:0.9.0
docker pull unclecode/crawl4ai:latestNote: Docker images are being built and will be available shortly.
Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
0.9.0 is a major, secure-by-default release of the Crawl4AI Docker API server. The out-of-the-box deployment is now hardened with defense in depth: authentication is on by default, the server binds loopback unless you give it a token, and the network request body is treated as an untrusted trust boundary. This release contains breaking changes for the self-hosted HTTP server only. The core pip library (SDK / in-process use) is unchanged.
What changed: the Docker server moved from an open, trust-the-caller posture to a closed, secure-by-default one. Defaults that used to be permissive (open bind, no auth, request-supplied browser internals, TLS verification off, Redis with no password) are now safe by default and gated behind explicit configuration.
What you must do: set CRAWL4AI_API_TOKEN and re-issue any tokens, then review whether you relied on any of the request fields or features that are now configured server-side. Most plain "crawl these URLs" users only need the two steps in the "Everyone" section of the migration guide. The full guide is at deploy/docker/MIGRATION.md.
This release completes the secure-by-default hardening of the Docker API server begun in 0.8.7 and 0.8.8. It moves the worst remaining issues from mitigation to architecture: unauthenticated access and request-supplied code/config are eliminated by design rather than patched in place. Every change is hardening; users self-hosting the Docker server should upgrade and follow the migration guide.
0.0.0.0. With no token it binds 127.0.0.1 and prints a one-off local token; exposing it requires CRAWL4AI_API_TOKEN and Authorization: Bearer <token> on every request except GET /health.O_NOFOLLOW, closing a path-traversal-to-file-write class. Credit: Y4tacker./crawl/stream and /crawl with stream=true now validate the destination and return HTTP 400 for disallowed targets, matching the non-streaming handlers. Credit: KOH Jun Sheng.browser_config.extra_args rejected (CWE-94): launch arguments can no longer be supplied over the network, closing a Chromium launch-arg injection class. Credit: Y4tacker, UDU_RisePho (hoanggxyuuki).All reporters are credited in SECURITY-CREDITS.md. GitHub Security Advisories accompany this release.
These apply to the self-hosted Docker API server only. The pip library is unaffected. See deploy/docker/MIGRATION.md for the step-by-step migration and deploy/docker/SECURITY-VERIFY.md for the deployment checklist.
CRAWL4AI_API_TOKEN and send Authorization: Bearer <token>. With no token the server binds loopback only.0.0.0.0 without a token; put a TLS-terminating reverse proxy in front when you expose it.POST /token.js_code, js_code_before_wait, c4a_script, proxy / proxy_config, extra_args, user_data_dir, cdp_url, cookies, headers, init_scripts, base_url, deep_crawl_strategy, simulate_user, magic, process_in_browser, and nested LLM config objects are rejected with HTTP 400 when sent over the network. Configure them server-side or use the in-process SDK. Unknown fields are dropped; timeouts, viewport, and scroll counts are clamped.hooks.code is replaced by a fixed action set (block_resources, add_cookies, set_headers, scroll_to_bottom, wait_for_timeout). See GET /hooks/info.output_path removed, replaced by an artifact id: /screenshot and /pdf store the result and return artifact_id + URL; fetch via authenticated GET /artifacts/{artifact_id} (TTL and quota apply).base_url removed: /md, /llm, and /llm/job select a provider by name only; endpoint and key are configured server-side and constrained by config.llm.allowed_providers.POST /monitor/actions/* and /monitor/stats/reset need an admin-scope principal.security.cors_allow_origins.CRAWL4AI_ALLOW_INSECURE_TLS=true, CRAWL4AI_ALLOW_INTERNAL_URLS=true.REDIS_PASSWORD.0 = unbounded).{"error": "Internal server error", "correlation_id": "…"}; match the id in the logs for detail.Y4tacker, KOH Jun Sheng, and UDU_RisePho (hoanggxyuuki). See SECURITY-CREDITS.md.
docker pull unclecode/crawl4ai:0.8.9 docker pull unclecode/crawl4ai:latest
PyPI:
pip install crawl4ai==0.8.9Docker:
docker pull unclecode/crawl4ai:0.8.9
docker pull unclecode/crawl4ai:latestNote: Docker images are being built and will be available shortly.
Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
0.8.9 is a follow-up, backward-compatible security patch for the self-hosted Docker API server, closing an SSRF path that 0.8.8 did not cover. Upgrade in place; no configuration changes required.
A security advisory accompanies this release.
/crawl, /crawl/stream, or /crawl/job request could set browser_config.proxy_config.server (or the deprecated browser_config.proxy, or crawler_config.proxy_config, or a --proxy-server / --host-resolver-rules flag in extra_args) to an internal address and route the browser through it, reaching internal services and cloud-metadata endpoints. All proxy destinations are now validated with the same global-routability check before the browser is built, and proxy/DNS-redirecting flags are stripped from extra_args. A legitimate public proxy still works. Credit: Geo (geo-chen).Backward compatible. Note: raw --proxy-server / --host-resolver-rules / --proxy-bypass-list / --proxy-pac-url flags passed via extra_args are now ignored; configure proxies through proxy_config (which is validated).
docker pull unclecode/crawl4ai:0.8.8 docker pull unclecode/crawl4ai:latest
PyPI:
pip install crawl4ai==0.8.8Docker:
docker pull unclecode/crawl4ai:0.8.8
docker pull unclecode/crawl4ai:latestNote: Docker images are being built and will be available shortly.
Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
0.8.8 is a focused, backward-compatible security patch for the self-hosted Docker API server. Upgrade in place; no configuration changes are required. If you run the Docker server, upgrade. If it is exposed to a network, also set CRAWL4AI_API_TOKEN.
Security advisories accompany this release.
64:ff9b::/96, 6to4 2002::/16, IPv4-mapped, and the unspecified ::), which previously bypassed the explicit blocklist and could reach internal services and cloud-metadata endpoints. SSRF errors no longer echo the resolved address. Credit: internal security audit.output_path hardened (CWE-59/22): /screenshot and /pdf now resolve symlinks and re-check containment before writing, and write with O_NOFOLLOW, closing a symlink/TOCTOU bypass of the directory restriction. output_path behavior is unchanged for normal use. Credit: internal security audit./md, /llm, /llm/job) ignore a request-supplied base_url, so the configured provider key can no longer be redirected to an attacker endpoint. LLMConfig additionally refuses to resolve protected environment variables via the env: token form. The base_url field is still accepted but no longer honored. Credit: Geo (geo-chen); the env: hardening from internal security audit.All changes are backward compatible.
The next release is a larger, secure-by-default update for the self-hosted Docker API server, with intentional breaking changes. We are giving advance notice so you can prepare. If you run the Docker server, start planning now and test in staging before upgrading:
CRAWL4AI_API_TOKEN) is configured./screenshot and /pdf return an artifact id instead of a file path, and the LLM endpoint is selected by provider name.A full migration guide will accompany the pre-announcement on Discord and X.
PyPI: `bash pip install crawl4ai==0.8.7 `
PyPI:
pip install crawl4ai==0.8.7
Docker:
docker pull unclecode/crawl4ai:0.8.7
docker pull unclecode/crawl4ai:latest
Note: Docker images are being built and will be available shortly. Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
0.8.7 is a security-hardening release. It bundles every responsibly-disclosed vulnerability patched since 0.8.6, plus the new DomainMapper feature and a batch of scraping, deep-crawl, and LLM fixes.
This release fixes multiple critical vulnerabilities in the Docker API server. If you self-host the Docker API, upgrade immediately. Two GitHub Security Advisories accompany this release.
gi_frame.f_back frame-chain escape in the computed-field eval() path. Removed eval() from computed fields entirely and deleted _safe_eval_expression. Credit: Song Binglin (q1uf3ng).asyncio, json, re) carried a full __builtins__, bypassing the __import__ block. Stripped injected builtins and removed dangerous allowlist entries. Credit: by111 (August829)."mysecret" allowed token forgery. Removed the default, reject weak/short secrets, and auto-generate an ephemeral key when JWT is enabled with no key set. Credit: by111 (August829).output_path (CVSS 9.1, CWE-22): /screenshot and /pdf wrote to any path. Restricted writes to CRAWL4AI_OUTPUT_DIR and reject .. traversal. Credit: Jeongbean Jeon, wulonchia./crawl/job and /llm/job could reach internal and cloud-metadata IPs. Added a blocklist and follow_redirects=False. Credit: Jeongbean Jeon./crawl, /md, and /llm fetched arbitrary URLs, and IPv6-mapped IPv4 addresses ([::ffff:169.254.169.254]) bypassed naive checks. Added destination validation on all entry points and normalize IPv6-mapped IPv4 before the blocklist check. Credit: secsys_codex, Velayutham Selvaraj, IcySun./execute_js (CVSS 8.1, CWE-94): disabled by default via CRAWL4AI_EXECUTE_JS_ENABLED, removed --disable-web-security from default browser args, and added an SSRF blocklist on the destination. Credit: by111 (August829)./monitor/* routes, including destructive actions, were unauthenticated. Added token_dep to the router and an explicit token check on the WebSocket endpoint. Credit: Jeongbean Jeon.innerHTML without escaping. Added server-side html.escape() and a client-side escapeHtml() wrapper. Credit: Jeongbean Jeon./config/dump: replaced with JSON input validated by Pydantic.markdown_generator type in CrawlerRunConfig to reject malformed JSON (#1880).include_subdomains flag and a per-source timeout.rowspan/colspan in cleaned HTML (#1920).tail text when removing empty elements (#1938)NlpSentenceChunking (#1909)set(False) instead of reset(token) (#1917)semaphore_count into the auto-created MemoryAdaptiveDispatcher and default it to 10 (#1927)LLMExtractionStrategy.extraction_type to schemaLLMTableExtraction to the Docker deserialization allowlistsuccess=True for binary downloads and skip the block check when downloaded_files is set<base href> in prefetch quick_extract_links (#752)AsyncLogger output to stderr by default (#1968) and use Console(width=200) for non-TTY contextsensure_ascii=False in the MCP bridge to preserve CJK characters (#1967)browser_adapter now uses the Stealth import, fixing a stealth import mismatch (#1960)arun() return type to CrawlResultContainer (#1898)Song Binglin (q1uf3ng), by111 (August829), Jeongbean Jeon, wulonchia, secsys_codex, Velayutham Selvaraj, and IcySun. See SECURITY-CREDITS.md.
Nothing published for this version
PyPI: `bash pip install crawl4ai==0.8.5 `
PyPI:
pip install crawl4ai==0.8.5
Docker:
docker pull unclecode/crawl4ai:0.8.5
docker pull unclecode/crawl4ai:latest
Note: Docker images are being built and will be available shortly. Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
March 2026 • 10 min read
I'm releasing Crawl4AI v0.8.5—our biggest release since v0.8.0. This update brings automatic anti-bot detection with proxy escalation, Shadow DOM flattening, deep crawl cancellation, and over 60 bug fixes from both our team and the community. If you're running crawls at scale or dealing with protected sites, this one's for you.
cancel() or should_cancel callbackset_defaults() / get_defaults() / reset_defaults()avoid_ads / avoid_css| col1 | col2 | pipe delimiters in markdown outputThis is the headline feature. Crawl4AI now automatically detects when a page is blocked by anti-bot protection and takes action—retrying with different proxies or falling back to an alternative fetch method.
The detection uses three tiers:
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig
from crawl4ai.async_configs import ProxyConfig
config = CrawlerRunConfig(
# Try direct first, then proxy on bot detection
proxy_config=[
ProxyConfig.DIRECT,
ProxyConfig(server="http://my-proxy:8080"),
],
max_retries=2,
# Optional: fallback when all proxies fail
fallback_fetch_function=my_web_unlocker_function,
)
async with AsyncWebCrawler() as crawler:
result = await crawler.arun("https://protected-site.com", config=config)
# Check what happened
stats = result.crawl_stats
print(f"Resolved by: {stats['resolved_by']}") # "direct", "proxy", or "fallback_fetch"
print(f"Proxies tried: {len(stats['proxies_used'])}")
The system errs on the side of caution—false positives are cheap (the fallback rescues them), but false negatives mean garbage results. After 5 iterations of real-world testing, it handles everything from Cloudflare challenges to Reddit's 180KB SPA block pages.
Web components with shadow DOM hide their content from regular DOM traversal. The new flatten_shadow_dom option serializes shadow DOM content into the light DOM before extraction.
config = CrawlerRunConfig(flatten_shadow_dom=True)
async with AsyncWebCrawler() as crawler:
result = await crawler.arun("https://some-web-component-site.com", config=config)
# Shadow DOM content is now visible in result.html, cleaned_html, and markdown
The implementation patches attachShadow to force-open closed shadow roots, recursively resolves <slot> projections, and strips only shadow-scoped <style> tags. It also reorders the JS execution pipeline—js_code now runs after wait_for + delay_before_return_html so your scripts operate on the fully-hydrated page. If you need JS to run before waiting, use the new js_code_before_wait parameter.
All deep crawl strategies (BFS, DFS, BestFirst) now support graceful cancellation:
from crawl4ai.deep_crawling import DFSDeepCrawlStrategy
pages_found = 0
def should_stop():
return pages_found >= 50 # Stop after finding enough pages
async def on_state(state):
nonlocal pages_found
pages_found = state["pages_crawled"]
strategy = DFSDeepCrawlStrategy(
max_depth=3,
max_pages=1000,
should_cancel=should_stop, # Sync or async callback
on_state_change=on_state,
)
config = CrawlerRunConfig(deep_crawl_strategy=strategy)
async with AsyncWebCrawler() as crawler:
results = await crawler.arun("https://example.com", config=config)
print(f"Cancelled: {strategy.cancelled}")
You can also call strategy.cancel() directly from another thread or coroutine.
Tired of repeating the same parameters? Set defaults once and they apply to every new instance:
from crawl4ai import BrowserConfig, CrawlerRunConfig
# Set organization-wide defaults
BrowserConfig.set_defaults(headless=True, text_mode=True)
CrawlerRunConfig.set_defaults(verbose=False, remove_consent_popups=True)
# All new instances inherit defaults
bc = BrowserConfig() # headless=True, text_mode=True
rc = CrawlerRunConfig() # verbose=False, remove_consent_popups=True
# Explicit params always override
bc2 = BrowserConfig(text_mode=False) # text_mode=False, headless still True
# Inspect and reset
print(BrowserConfig.get_defaults()) # {"headless": True, "text_mode": True}
BrowserConfig.reset_defaults() # Back to normal
Many sites split a single item's data across sibling elements (think Hacker News, where title and score are in separate <tr> rows). The new "source" field navigates to a sibling before extracting:
from crawl4ai.extraction_strategy import JsonCssExtractionStrategy
schema = {
"name": "HackerNewsItems",
"baseSelector": "tr.athing",
"fields": [
{"name": "title", "selector": ".titleline > a", "type": "text"},
{"name": "link", "selector": ".titleline > a", "type": "attribute", "attribute": "href"},
# Navigate to the NEXT sibling <tr> to get the score
{"name": "score", "selector": ".score", "type": "text", "source": "+ tr"},
{"name": "author", "selector": ".hnuser", "type": "text", "source": "+ tr"},
]
}
strategy = JsonCssExtractionStrategy(schema=schema)
Works in both JsonCssExtractionStrategy and JsonXPathExtractionStrategy. Falls back gracefully when siblings don't exist.
A single flag auto-dismisses cookie consent banners from 40+ CMP platforms:
config = CrawlerRunConfig(remove_consent_popups=True)
Covers OneTrust, Cookiebot, Didomi, Quantcast, Sourcepoint, Google FundingChoices, TrustArc, ConsentManager, Osano, Iubenda, Complianz, LiveRamp, CookieYes, Klaro, Termly, and many more.
Block ad trackers and CSS resources at the network level for faster, leaner crawls:
config = BrowserConfig(
avoid_ads=True, # Blocks doubleclick, google-analytics, etc.
avoid_css=True, # Blocks .css, .less, .scss resources
)
For long-running crawl sessions:
config = BrowserConfig(
memory_saving_mode=True, # Aggressive cache/V8 heap flags
max_pages_before_recycle=100, # Auto-restart browser after N pages
)
This prevents memory leaks during sustained crawling. The recycling uses a version-based approach that's safe under concurrent load—we fixed three separate deadlock bugs to get this right.
Tables in markdown output now have proper GitHub-Flavored Markdown pipe delimiters:
Before (v0.8.0):
Name | Age | City
---|---|---
Alice | 30 | NYC
After (v0.8.5):
| Name | Age | City |
| --- | --- | --- |
| Alice | 30 | NYC |
query_llm_config: Separate LLM config for adaptive crawler query expansion (#1682)force_viewport_screenshot: Screenshot only the viewport, not the full pagedevice_scale_factor: Configurable screenshot DPI via BrowserConfig (#1463)redirected_status_code: Now available on CrawlResult (#1435)wait_for_images: Wait for images to load before taking screenshots (#1792)score_threshold: Filter low-quality URLs in BestFirstCrawlingStrategy (#1804)link_preview_timeout: Configurable timeout in AdaptiveConfig (#1793)--json-ensure-ascii: CLI flag for Unicode preservation in JSON output (#1668)type-list pipeline: Chained extraction like ["attribute", "regex"] in JsonCssExtractionStrategy (#1290)Severity: CRITICAL Affected: Docker API deployment (v0.8.0 and earlier)
The /crawl endpoint's deserialization logic used eval() for certain object types. I removed this entirely and added an allowlist (ALLOWED_DESERIALIZE_TYPES) so only known config classes can be instantiated.
Affected: Docker deployments using Redis
Upgraded Redis to 7.2.7 which patches the Lua use-after-free vulnerability.
/token endpoint now requires api_token when configured (#1795)sec-ch-ua synced with User-Agent, WebGL kept alive in stealth modecreate_isolated_context=Falseadd_init_script (#1768)simulate_user destroying page content via ArrowDown keypressERR_INVALID_AUTH_CREDENTIALS (#1281)can_process_url() to receive normalized URLtotal_score not calculated for links that fail head extractionFilterChain.add_filter AttributeError on tuple immutabilityis_external_url port comparison (#1783)<base> tag ignored in html2text relative link resolution (#1721)cleaned_html (#1364)class and id attributes in cleaned_html (#1782)force_json_response path for LLM extractionfinish_reason (#1788)agenerate_schema() JSON parsing for Anthropic modelsfrom_serializable_dict ignoring plain data dicts with "type" keycss_selector ignored in LXML scraping for raw:// URLs (#1484)CRAWL4_AI_BASE_DIRECTORY env var (#1296)UnicodeEncodeError in URL seeder, strip zero-width chars (#1784)scroll_delay ignored in full-page screenshot scroller/llm per-request provider override, Redis config from host/port/password (#1611, #1817)scan_full_page=False (#1750)arun_many dispatcher bypass (#1818, #1509)tf-playwright-stealth with playwright-stealth (#1553)script.js in package distribution (#1711)text → string) (#1077)chardet.detect in thread executor (#1751)mean_delay/max_range from CrawlerRunConfig into dispatcher rate limiter (#1786)Added a comprehensive 291-test regression suite covering all major subsystems: core crawl, content processing, extraction strategies, deep crawling, browser management, config serialization, utilities, and edge cases.
cleaned_html Now Preserves class and id AttributesIf you have downstream code that parses cleaned_html and assumes no class/id attributes are present, this may need updating. This change enables users to do CSS-based analysis on cleaned HTML.
If you pin Redis versions in your deployment, update to 7.2.7 or later.
pip install --upgrade crawl4ai
# or
pip install crawl4ai==0.8.5
docker pull unclecode/crawl4ai:0.8.5
docker run -d -p 11235:11235 --shm-size=1g unclecode/crawl4ai:0.8.5
Run the verification tests to confirm all features are working:
python docs/releases_review/demo_v0.8.5.py
This runs 13 actual tests that crawl real URLs and verify each feature end-to-end.
This release includes contributions from a large number of community members. Thank you to everyone who submitted PRs, reported issues, and provided reproduction steps. Special thanks to all contributors listed in CONTRIBUTORS.md.
Issues fixed: #462, #880, #943, #1031, #1077, #1183, #1213, #1251, #1281, #1290, #1296, #1308, #1354, #1364, #1370, #1374, #1424, #1435, #1463, #1484, #1487, #1489, #1494, #1503, #1509, #1512, #1520, #1553, #1594, #1601, #1606, #1611, #1622, #1635, #1640, #1658, #1666, #1667, #1668, #1671, #1682, #1686, #1711, #1715, #1716, #1721, #1730, #1731, #1746, #1750, #1751, #1754, #1758, #1762, #1768, #1770, #1776, #1782, #1783, #1784, #1786, #1788, #1789, #1790, #1792, #1793, #1794, #1795, #1796, #1797, #1801, #1803, #1804, #1805, #1815, #1817, #1818, #1824
This is a massive release—10 new features, critical security patches, and 60+ bug fixes. Whether you're dealing with anti-bot protection, shadow DOM sites, or just want more reliable crawls at scale, v0.8.5 has you covered. Thank you for your continued support!
Happy crawling!
- unclecode
PyPI: `bash pip install crawl4ai==0.8.0 `
PyPI:
pip install crawl4ai==0.8.0
Docker:
docker pull unclecode/crawl4ai:0.8.0
docker pull unclecode/crawl4ai:latest
Note: Docker images are being built and will be available shortly. Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
__import__ from hook allowed builtins
CRAWL4AI_HOOKS_ENABLED environment variablefile://, javascript:, data: URLs on /execute_js, /screenshot, /pdf, /htmlhttp://, https://, and raw: URLsCRAWL4AI_HOOKS_ENABLED=true to enableresume_state and on_state_change for BFS/DFS/Best-First strategiesprocess_in_browser parametercache_ttl_hours and validate_sitemap_lastmod parameters# character (CSS color codes like #eee)Release Date: January 2026 Previous Version: v0.7.6 Status: Release Candidate
What changed: Hooks are now disabled by default on the Docker API.
Why: Security fix for Remote Code Execution (RCE) vulnerability.
Who is affected: Users of the Docker API who use the hooks parameter in /crawl requests.
Migration:
# To re-enable hooks (only if you trust all API users):
export CRAWL4AI_HOOKS_ENABLED=true
What changed: The endpoints /execute_js, /screenshot, /pdf, and /html now reject file:// URLs.
Why: Security fix for Local File Inclusion (LFI) vulnerability.
Who is affected: Users who were reading local files via the Docker API.
Migration: Use the Python library directly for local file processing:
# Instead of API call with file:// URL, use library:
from crawl4ai import AsyncWebCrawler
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(url="file:///path/to/file.html")
Severity: CRITICAL (CVSS 10.0)
Affected: Docker API deployment (all versions before v0.8.0)
Vector: POST /crawl with malicious hooks parameter
Details: The __import__ builtin was available in hook code, allowing attackers to import os, subprocess, etc. and execute arbitrary commands.
Fix:
__import__ from allowed builtinsCRAWL4AI_HOOKS_ENABLED=false)Severity: HIGH (CVSS 8.6)
Affected: Docker API deployment (all versions before v0.8.0)
Vector: POST /execute_js (and other endpoints) with file:///etc/passwd
Details: API endpoints accepted file:// URLs, allowing attackers to read arbitrary files from the server.
Fix: URL scheme validation now only allows http://, https://, and raw: URLs.
Discovered by Neo by ProjectDiscovery (projectdiscovery.io) - December 2025
Pre-page-load JavaScript injection for stealth evasions.
config = BrowserConfig(
init_scripts=[
"Object.defineProperty(navigator, 'webdriver', {get: () => false})"
]
)
ws://, wss://)cdp_cleanup_on_close=TrueAll deep crawl strategies (BFS, DFS, Best-First) now support crash recovery:
from crawl4ai.deep_crawling import BFSDeepCrawlStrategy
strategy = BFSDeepCrawlStrategy(
max_depth=3,
resume_state=saved_state, # Resume from checkpoint
on_state_change=save_callback # Persist state in real-time
)
Generate PDFs and MHTML from cached HTML content.
Render cached HTML and capture screenshots.
Proper URL resolution for raw: HTML processing:
config = CrawlerRunConfig(base_url='https://example.com')
result = await crawler.arun(url='raw:{html}', config=config)
Fast link extraction without full page processing:
config = CrawlerRunConfig(prefetch=True)
Enhanced proxy rotation with sticky sessions support.
Non-browser crawler now supports proxies.
New process_in_browser parameter for browser operations on local content:
config = CrawlerRunConfig(
process_in_browser=True, # Force browser processing
screenshot=True
)
result = await crawler.arun(url='raw:<html>...</html>', config=config)
Intelligent cache invalidation for sitemaps:
config = SeedingConfig(
cache_ttl_hours=24,
validate_sitemap_lastmod=True
)
Problem: CSS color codes like #eee were being truncated.
Before: raw:body{background:#eee} → body{background:
After: raw:body{background:#eee} → body{background:#eee}
Various fixes to cache validation and persistence.
Update the package:
pip install --upgrade crawl4ai
Docker API users:
export CRAWL4AI_HOOKS_ENABLED=truefile:// URLs no longer work on API (use library directly)Review security settings:
# config.yml - recommended for production
security:
enabled: true
jwt_enabled: true
Test your integration before deploying to production
hooks parameter in API callsfile:// URLs via the APISee CHANGELOG.md for complete version history.
Thanks to all contributors who made this release possible.
Special thanks to Neo by ProjectDiscovery for responsible security disclosure.
For questions or issues, please open a GitHub Issue.
PyPI: `bash pip install crawl4ai==0.7.8 `
PyPI:
pip install crawl4ai==0.7.8
Docker:
docker pull unclecode/crawl4ai:0.7.8
docker pull unclecode/crawl4ai:latest
Note: Docker images are being built and will be available shortly. Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
December 2025
I'm releasing Crawl4AI v0.7.8—a focused stability release that addresses 11 bugs reported by the community. While there are no new features in this release, these fixes resolve important issues affecting Docker deployments, LLM extraction, URL handling, and dependency compatibility.
The Problem: When sending deep crawl requests to the Docker API with ContentRelevanceFilter, the server failed to deserialize the filter, causing requests to fail.
The Fix: I added ContentRelevanceFilter to the public exports and enhanced the deserialization logic with dynamic imports.
# This now works correctly in Docker API
import httpx
request = {
"urls": ["https://docs.example.com"],
"crawler_config": {
"deep_crawl_strategy": {
"type": "BFSDeepCrawlStrategy",
"max_depth": 2,
"filter_chain": [
{
"type": "ContentRelevanceFilter",
"query": "API documentation",
"threshold": 0.3
}
]
}
}
}
async with httpx.AsyncClient() as client:
response = await client.post("http://localhost:11235/crawl", json=request)
# Previously failed, now works!
The Problem: BrowserConfig.to_dict() failed when proxy_config was set because ProxyConfig wasn't being serialized to a dictionary.
The Fix: ProxyConfig.to_dict() is now called during serialization.
from crawl4ai import BrowserConfig
from crawl4ai.async_configs import ProxyConfig
proxy = ProxyConfig(
server="http://proxy.example.com:8080",
username="user",
password="pass"
)
config = BrowserConfig(headless=True, proxy_config=proxy)
# Previously raised TypeError, now works
config_dict = config.to_dict()
json.dumps(config_dict) # Valid JSON
The Problem: The .cache folder in the Docker image had incorrect permissions, causing crawling to fail when caching was enabled.
The Fix: Corrected ownership and permissions during image build.
# Cache now works correctly in Docker
docker run -d -p 11235:11235 \
--shm-size=1g \
-v ./my-cache:/app/.cache \
unclecode/crawl4ai:0.7.8
The Problem: The LLM rate limiting backoff parameters were hardcoded, making it impossible to adjust retry behavior for different API rate limits.
The Fix: LLMConfig now accepts three new parameters for complete control over retry behavior.
from crawl4ai import LLMConfig
# Default behavior (unchanged)
default_config = LLMConfig(provider="openai/gpt-4o-mini")
# backoff_base_delay=2, backoff_max_attempts=3, backoff_exponential_factor=2
# Custom configuration for APIs with strict rate limits
custom_config = LLMConfig(
provider="openai/gpt-4o-mini",
backoff_base_delay=5, # Wait 5 seconds on first retry
backoff_max_attempts=5, # Try up to 5 times
backoff_exponential_factor=3 # Multiply delay by 3 each attempt
)
# Retry sequence: 5s -> 15s -> 45s -> 135s -> 405s
The Problem: LLMExtractionStrategy always sent markdown to the LLM, but some extraction tasks work better with HTML structure preserved.
The Fix: Added input_format parameter supporting "markdown", "html", "fit_markdown", "cleaned_html", and "fit_html".
from crawl4ai import LLMExtractionStrategy, LLMConfig
# Default: markdown input (unchanged)
markdown_strategy = LLMExtractionStrategy(
llm_config=LLMConfig(provider="openai/gpt-4o-mini"),
instruction="Extract product information"
)
# NEW: HTML input - preserves table/list structure
html_strategy = LLMExtractionStrategy(
llm_config=LLMConfig(provider="openai/gpt-4o-mini"),
instruction="Extract the data table preserving structure",
input_format="html"
)
# NEW: Filtered markdown - only relevant content
fit_strategy = LLMExtractionStrategy(
llm_config=LLMConfig(provider="openai/gpt-4o-mini"),
instruction="Summarize the main content",
input_format="fit_markdown"
)
The Problem: When using url="raw:<html>...", the entire HTML content was being passed to extraction strategies as the URL parameter, polluting LLM prompts.
The Fix: The URL is now correctly set to "Raw HTML" for raw HTML inputs.
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig
html = "<html><body><h1>Test</h1></body></html>"
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(
url=f"raw:{html}",
config=CrawlerRunConfig(extraction_strategy=my_strategy)
)
# extraction_strategy receives url="Raw HTML" instead of the HTML blob
The Problem: When JavaScript caused a page redirect, relative links were resolved against the original URL instead of the final URL.
The Fix: redirected_url now captures the actual page URL after all JavaScript execution completes.
from crawl4ai import AsyncWebCrawler
async with AsyncWebCrawler() as crawler:
# Page at /old-page redirects via JS to /new-page
result = await crawler.arun(url="https://example.com/old-page")
# BEFORE: redirected_url = "https://example.com/old-page"
# AFTER: redirected_url = "https://example.com/new-page"
# Links are now correctly resolved against the final URL
for link in result.links['internal']:
print(link['href']) # Relative links resolved correctly
The Problem: PyPDF2 was deprecated in 2022 and is no longer maintained.
The Fix: Replaced with the actively maintained pypdf library.
# Installation (unchanged)
pip install crawl4ai[pdf]
# The PDF processor now uses pypdf internally
# No code changes required - API remains the same
The Problem: Using the deprecated class Config syntax caused deprecation warnings with Pydantic v2.
The Fix: Migrated to model_config = ConfigDict(...) syntax.
# No more deprecation warnings when importing crawl4ai models
from crawl4ai.models import CrawlResult
from crawl4ai import CrawlerRunConfig, BrowserConfig
# All models are now Pydantic v2 compatible
The Problem: The EmbeddingStrategy in AdaptiveCrawler had commented-out LLM code and was using hardcoded mock query variations instead.
The Fix: Uncommented and activated the LLM call for actual query expansion.
# AdaptiveCrawler query expansion now actually uses the LLM
# Instead of hardcoded variations like:
# variations = {'queries': ['what are the best vegetables...']}
# The LLM generates relevant query variations based on your actual query
The Problem: When extracting code from web pages, import statements were sometimes concatenated without proper line separation.
The Fix: Import statements now maintain proper newline separation.
# BEFORE: "import osimport sysfrom pathlib import Path"
# AFTER:
# import os
# import sys
# from pathlib import Path
None! This release is fully backward compatible.
pip install --upgrade crawl4ai
# or
pip install crawl4ai==0.7.8
# Pull the latest version
docker pull unclecode/crawl4ai:0.7.8
# Run
docker run -d -p 11235:11235 --shm-size=1g unclecode/crawl4ai:0.7.8
Run the verification tests to confirm all fixes are working:
python docs/releases_review/demo_v0.7.8.py
This runs actual tests that verify each bug fix is properly implemented.
Thank you to everyone who reported these issues and provided detailed reproduction steps. Your bug reports make Crawl4AI better for everyone.
Issues fixed: #1642, #1638, #1629, #1621, #1412, #1269, #1268, #1181, #1178, #1116, #678
This stability release ensures Crawl4AI works reliably across Docker deployments, LLM extraction workflows, and various edge cases. Thank you for your continued support and feedback!
Happy crawling!
- unclecode
Updated pyOpenSSL from >=24.3.0 to >=25.3.0 (security vulnerability fix)
This release introduces a complete self-hosting platform with enterprise-grade real-time monitoring. This release transforms Crawl4AI Docker from a simple containerized crawler into a production-ready platform with full operational transparency and control.
Major Feature: Real-time Monitoring & Self-Hosting Platform
Docker deployment now includes:
🐛 Critical Bug Fixes
Configuration & Features
Docker & Infrastructure
Security
PyPI:
pip install crawl4ai==0.7.7
Docker:
docker pull unclecode/crawl4ai:0.7.7
docker pull unclecode/crawl4ai:latest
Note: Docker images are being built and will be available shortly. Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
November 14, 2025 • 10 min read
Today I'm releasing Crawl4AI v0.7.7—the Self-Hosting & Monitoring Update. This release transforms Crawl4AI Docker from a simple containerized crawler into a complete self-hosting platform with enterprise-grade real-time monitoring, full operational transparency, and production-ready observability.
The Problem: Running Crawl4AI in Docker was like flying blind. Users had no visibility into what was happening inside the container—memory usage, active requests, browser pools, or errors. Troubleshooting required checking logs, and there was no way to monitor performance or manually intervene when issues occurred.
My Solution: I built a complete real-time monitoring system with an interactive dashboard, comprehensive REST API, WebSocket streaming, and manual control actions. Now you have full transparency and control over your crawling infrastructure.
Before v0.7.7, Docker was just a containerized crawler. After v0.7.7, it's a complete self-hosting platform that gives you:
Access the dashboard at http://localhost:11235/dashboard to see:
The dashboard updates every 2 seconds via WebSocket, giving you live visibility into your crawling operations.
The Problem: Monitoring dashboards are great for humans, but automation and integration require programmatic access.
My Solution: A comprehensive REST API that exposes all monitoring data for integration with your existing infrastructure.
import httpx
import asyncio
async def monitor_system_health():
async with httpx.AsyncClient() as client:
response = await client.get("http://localhost:11235/monitor/health")
health = response.json()
print(f"Container Metrics:")
print(f" CPU: {health['container']['cpu_percent']:.1f}%")
print(f" Memory: {health['container']['memory_percent']:.1f}%")
print(f" Uptime: {health['container']['uptime_seconds']}s")
print(f"\nBrowser Pool:")
print(f" Permanent: {health['pool']['permanent']['active']} active")
print(f" Hot Pool: {health['pool']['hot']['count']} browsers")
print(f" Cold Pool: {health['pool']['cold']['count']} browsers")
print(f"\nStatistics:")
print(f" Total Requests: {health['stats']['total_requests']}")
print(f" Success Rate: {health['stats']['success_rate_percent']:.1f}%")
print(f" Avg Latency: {health['stats']['avg_latency_ms']:.0f}ms")
asyncio.run(monitor_system_health())
async def track_requests():
async with httpx.AsyncClient() as client:
response = await client.get("http://localhost:11235/monitor/requests")
requests_data = response.json()
print(f"Active Requests: {len(requests_data['active'])}")
print(f"Completed Requests: {len(requests_data['completed'])}")
# See details of recent requests
for req in requests_data['completed'][:5]:
status_icon = "✅" if req['success'] else "❌"
print(f"{status_icon} {req['endpoint']} - {req['latency_ms']:.0f}ms")
async def monitor_browser_pool():
async with httpx.AsyncClient() as client:
response = await client.get("http://localhost:11235/monitor/browsers")
browsers = response.json()
print(f"Pool Summary:")
print(f" Total Browsers: {browsers['summary']['total_count']}")
print(f" Total Memory: {browsers['summary']['total_memory_mb']} MB")
print(f" Reuse Rate: {browsers['summary']['reuse_rate_percent']:.1f}%")
# List all browsers
for browser in browsers['permanent']:
print(f"🔥 Permanent: {browser['browser_id'][:8]}... | "
f"Requests: {browser['request_count']} | "
f"Memory: {browser['memory_mb']:.0f} MB")
async def get_endpoint_stats():
async with httpx.AsyncClient() as client:
response = await client.get("http://localhost:11235/monitor/endpoints/stats")
stats = response.json()
print("Endpoint Analytics:")
for endpoint, data in stats.items():
print(f" {endpoint}:")
print(f" Requests: {data['count']}")
print(f" Avg Latency: {data['avg_latency_ms']:.0f}ms")
print(f" Success Rate: {data['success_rate_percent']:.1f}%")
The Monitor API includes these endpoints:
GET /monitor/health - System health with pool statisticsGET /monitor/requests - Active and completed request trackingGET /monitor/browsers - Browser pool details and efficiencyGET /monitor/endpoints/stats - Per-endpoint performance analyticsGET /monitor/timeline?minutes=5 - Time-series data for chartsGET /monitor/logs/janitor?limit=10 - Cleanup activity logsGET /monitor/logs/errors?limit=10 - Error logs with contextPOST /monitor/actions/cleanup - Force immediate cleanupPOST /monitor/actions/kill_browser - Kill specific browserPOST /monitor/actions/restart_browser - Restart browserPOST /monitor/stats/reset - Reset accumulated statisticsThe Problem: Polling the API every few seconds wastes resources and adds latency. Real-time dashboards need instant updates.
My Solution: WebSocket streaming with 2-second update intervals for building custom real-time dashboards.
import websockets
import json
import asyncio
async def monitor_realtime():
uri = "ws://localhost:11235/monitor/ws"
async with websockets.connect(uri) as websocket:
print("Connected to real-time monitoring stream")
while True:
# Receive update every 2 seconds
data = await websocket.recv()
update = json.loads(data)
# Access all monitoring data
print(f"\n--- Update at {update['timestamp']} ---")
print(f"Memory: {update['health']['container']['memory_percent']:.1f}%")
print(f"Active Requests: {len(update['requests']['active'])}")
print(f"Total Browsers: {update['browsers']['summary']['total_count']}")
if update['errors']:
print(f"⚠️ Recent Errors: {len(update['errors'])}")
asyncio.run(monitor_realtime())
Expected Real-World Impact:
The Problem: Creating a new browser for every request is slow and memory-intensive. Traditional browser pools are static and inefficient.
My Solution: A smart 3-tier browser pool that automatically adapts to usage patterns.
import httpx
async def demonstrate_browser_pool():
async with httpx.AsyncClient() as client:
# Request 1-3: Default config → Uses permanent browser
print("Phase 1: Using permanent browser")
for i in range(3):
await client.post(
"http://localhost:11235/crawl",
json={"urls": [f"https://httpbin.org/html?req={i}"]}
)
print(f" Request {i+1}: Reused permanent browser")
# Request 4-6: Custom viewport → Cold pool (first use)
print("\nPhase 2: Custom config creates cold pool browser")
viewport_config = {"viewport": {"width": 1280, "height": 720}}
for i in range(4):
await client.post(
"http://localhost:11235/crawl",
json={
"urls": [f"https://httpbin.org/json?v={i}"],
"browser_config": viewport_config
}
)
if i < 2:
print(f" Request {i+1}: Cold pool browser")
else:
print(f" Request {i+1}: Promoted to hot pool! (after 3 uses)")
# Check pool status
response = await client.get("http://localhost:11235/monitor/browsers")
browsers = response.json()
print(f"\nPool Status:")
print(f" Permanent: {len(browsers['permanent'])} (always active)")
print(f" Hot: {len(browsers['hot'])} (frequently used configs)")
print(f" Cold: {len(browsers['cold'])} (on-demand)")
print(f" Reuse Rate: {browsers['summary']['reuse_rate_percent']:.1f}%")
asyncio.run(demonstrate_browser_pool())
Pool Tiers:
Expected Real-World Impact:
The Problem: Long-running crawlers accumulate idle browsers and consume memory over time.
My Solution: An automatic janitor system that monitors and cleans up idle resources.
async def monitor_janitor_activity():
async with httpx.AsyncClient() as client:
response = await client.get("http://localhost:11235/monitor/logs/janitor?limit=5")
logs = response.json()
print("Recent Cleanup Activities:")
for log in logs:
print(f" {log['timestamp']}: {log['message']}")
# Example output:
# 2025-11-14 10:30:00: Cleaned up 2 cold pool browsers (idle > 5min)
# 2025-11-14 10:25:00: Browser reuse rate: 85.3%
# 2025-11-14 10:20:00: Hot pool browser promoted (10 requests)
The Problem: Sometimes you need to manually intervene—kill a stuck browser, force cleanup, or restart resources.
My Solution: Manual control actions via the API for operational troubleshooting.
async def force_cleanup():
async with httpx.AsyncClient() as client:
response = await client.post("http://localhost:11235/monitor/actions/cleanup")
result = response.json()
print(f"Cleanup completed:")
print(f" Browsers cleaned: {result.get('cleaned_count', 0)}")
print(f" Memory freed: {result.get('memory_freed_mb', 0):.1f} MB")
async def kill_stuck_browser(browser_id: str):
async with httpx.AsyncClient() as client:
response = await client.post(
"http://localhost:11235/monitor/actions/kill_browser",
json={"browser_id": browser_id}
)
if response.status_code == 200:
print(f"✅ Browser {browser_id} killed successfully")
async def reset_stats():
async with httpx.AsyncClient() as client:
response = await client.post("http://localhost:11235/monitor/stats/reset")
print("📊 Statistics reset for fresh monitoring")
# Export metrics for Prometheus scraping
async def export_prometheus_metrics():
async with httpx.AsyncClient() as client:
health = await client.get("http://localhost:11235/monitor/health")
data = health.json()
# Export in Prometheus format
metrics = f"""
# HELP crawl4ai_memory_usage_percent Memory usage percentage
# TYPE crawl4ai_memory_usage_percent gauge
crawl4ai_memory_usage_percent {data['container']['memory_percent']}
# HELP crawl4ai_request_success_rate Request success rate
# TYPE crawl4ai_request_success_rate gauge
crawl4ai_request_success_rate {data['stats']['success_rate_percent']}
# HELP crawl4ai_browser_pool_count Total browsers in pool
# TYPE crawl4ai_browser_pool_count gauge
crawl4ai_browser_pool_count {data['pool']['permanent']['active'] + data['pool']['hot']['count'] + data['pool']['cold']['count']}
"""
return metrics
async def check_alerts():
async with httpx.AsyncClient() as client:
health = await client.get("http://localhost:11235/monitor/health")
data = health.json()
# Memory alert
if data['container']['memory_percent'] > 80:
print("🚨 ALERT: Memory usage above 80%")
# Trigger cleanup
await client.post("http://localhost:11235/monitor/actions/cleanup")
# Success rate alert
if data['stats']['success_rate_percent'] < 90:
print("🚨 ALERT: Success rate below 90%")
# Check error logs
errors = await client.get("http://localhost:11235/monitor/logs/errors")
print(f"Recent errors: {len(errors.json())}")
# Latency alert
if data['stats']['avg_latency_ms'] > 5000:
print("🚨 ALERT: Average latency above 5s")
CRITICAL_METRICS = {
"memory_usage": {
"current": "container.memory_percent",
"target": "<80%",
"alert_threshold": ">80%",
"action": "Force cleanup or scale"
},
"success_rate": {
"current": "stats.success_rate_percent",
"target": ">95%",
"alert_threshold": "<90%",
"action": "Check error logs"
},
"avg_latency": {
"current": "stats.avg_latency_ms",
"target": "<2000ms",
"alert_threshold": ">5000ms",
"action": "Investigate slow requests"
},
"browser_reuse_rate": {
"current": "browsers.summary.reuse_rate_percent",
"target": ">80%",
"alert_threshold": "<60%",
"action": "Check pool configuration"
},
"total_browsers": {
"current": "browsers.summary.total_count",
"target": "<15",
"alert_threshold": ">20",
"action": "Check for browser leaks"
},
"error_frequency": {
"current": "len(errors)",
"target": "<5/hour",
"alert_threshold": ">10/hour",
"action": "Review error patterns"
}
}
This release includes significant bug fixes that improve stability and performance:
The Problem: LLM extraction was blocking async execution, causing URLs to be processed sequentially instead of in parallel (issue #1055).
The Fix: Resolved the blocking issue to enable true parallel processing for LLM extraction.
# Before v0.7.7: Sequential processing
# After v0.7.7: True parallel processing
async with AsyncWebCrawler() as crawler:
urls = ["url1", "url2", "url3", "url4"]
# Now processes truly in parallel with LLM extraction
results = await crawler.arun_many(
urls,
config=CrawlerRunConfig(
extraction_strategy=LLMExtractionStrategy(...)
)
)
# 4x faster for parallel LLM extraction!
Expected Impact: Major performance improvement for batch LLM extraction workflows.
The Problem: DFS (Depth-First Search) deep crawl strategy had implementation issues.
The Fix: Enhanced DFSDeepCrawlStrategy with proper seen URL tracking and improved documentation.
The Problem: Documentation didn't match the actual async_configs.py implementation.
The Fix: Updated all configuration documentation to accurately reflect the current implementation.
The Problem: Sitemap parsing and URL normalization issues in AsyncUrlSeeder (issue #1559).
The Fix: Added comprehensive tests and fixes for sitemap namespace parsing and URL normalization.
The Problem: The remove_overlay_elements functionality wasn't working (issue #1396).
The Fix: Fixed by properly calling the injected JavaScript function.
The Problem: Viewport configuration wasn't working in managed browsers (issue #1490).
The Fix: Added proper viewport size configuration support for browser launch.
The Problem: CDP (Chrome DevTools Protocol) endpoint verification had timing issues causing connection failures (issue #1445).
The Fix: Added exponential backoff for CDP endpoint verification to handle timing variations.
fit_html property serialization in /crawl and /crawl/stream endpointsNone! This release is fully backward compatible.
# Pull the latest version
docker pull unclecode/crawl4ai:0.7.7
# Or use the latest tag
docker pull unclecode/crawl4ai:latest
# Run with monitoring enabled (default)
docker run -d \
-p 11235:11235 \
--shm-size=1g \
--name crawl4ai \
unclecode/crawl4ai:0.7.7
# Access the monitoring dashboard
open http://localhost:11235/dashboard
# Upgrade to latest version
pip install --upgrade crawl4ai
# Or install specific version
pip install crawl4ai==0.7.7
Run the comprehensive demo that showcases all monitoring features:
python docs/releases_review/demo_v0.7.7.py
The demo includes:
docs/releases_review/demo_v0.7.7.py - Working examples/dashboard to get familiar with the monitoring systemThank you to our community for the feedback, bug reports, and feature requests that shaped this release. Special thanks to everyone who contributed to the issues that were fixed in this version.
The monitoring system was built based on real user needs for production deployments, and your input made it comprehensive and practical.
http://localhost:11235/dashboard (when running)Crawl4AI v0.7.7 delivers complete self-hosting with enterprise-grade monitoring. You now have full visibility and control over your web crawling infrastructure. The monitoring dashboard, comprehensive API, and WebSocket streaming give you everything needed for production deployments. Try the self-hosting platform—it's a game changer for operational excellence!
Happy crawling with full visibility! 🕷️📊
- unclecode
Crawl4AI v0.7.6 - Webhook Support for Docker Job Queue API
Crawl4AI v0.7.6 - Webhook Support for Docker Job Queue API
Users can now:
PyPI:
pip install crawl4ai==0.7.6
Docker:
docker pull unclecode/crawl4ai:0.7.6
docker pull unclecode/crawl4ai:latest
Note: Docker images are being built and will be available shortly. Check the Docker Release workflow for build status.
See CHANGELOG.md for details.
Release Date: October 22, 2025
I'm excited to announce Crawl4AI v0.7.6, featuring a complete webhook infrastructure for the Docker job queue API! This release eliminates polling and brings real-time notifications to both crawling and LLM extraction workflows.
The headline feature of v0.7.6 is comprehensive webhook support for asynchronous job processing. No more constant polling to check if your jobs are done - get instant notifications when they complete!
Key Capabilities:
/crawl/job and /llm/job endpoints now support webhooksconfig.yml for all jobscrawl and llm_extraction tasksInstead of constantly checking job status:
OLD WAY (Polling):
# Submit job
response = requests.post("http://localhost:11235/crawl/job", json=payload)
task_id = response.json()['task_id']
# Poll until complete
while True:
status = requests.get(f"http://localhost:11235/crawl/job/{task_id}")
if status.json()['status'] == 'completed':
break
time.sleep(5) # Wait and try again
NEW WAY (Webhooks):
# Submit job with webhook
payload = {
"urls": ["https://example.com"],
"webhook_config": {
"webhook_url": "https://myapp.com/webhook",
"webhook_data_in_payload": True
}
}
response = requests.post("http://localhost:11235/crawl/job", json=payload)
# Done! Webhook will notify you when complete
# Your webhook handler receives the results automatically
curl -X POST http://localhost:11235/crawl/job \
-H "Content-Type: application/json" \
-d '{
"urls": ["https://example.com"],
"browser_config": {"headless": true},
"crawler_config": {"cache_mode": "bypass"},
"webhook_config": {
"webhook_url": "https://myapp.com/webhooks/crawl-complete",
"webhook_data_in_payload": false,
"webhook_headers": {
"X-Webhook-Secret": "your-secret-token"
}
}
}'
curl -X POST http://localhost:11235/llm/job \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/article",
"q": "Extract the article title, author, and publication date",
"schema": "{\"type\":\"object\",\"properties\":{\"title\":{\"type\":\"string\"}}}",
"provider": "openai/gpt-4o-mini",
"webhook_config": {
"webhook_url": "https://myapp.com/webhooks/llm-complete",
"webhook_data_in_payload": true
}
}'
Success (with data):
{
"task_id": "llm_1698765432",
"task_type": "llm_extraction",
"status": "completed",
"timestamp": "2025-10-22T10:30:00.000000+00:00",
"urls": ["https://example.com/article"],
"data": {
"extracted_content": {
"title": "Understanding Web Scraping",
"author": "John Doe",
"date": "2025-10-22"
}
}
}
Failure:
{
"task_id": "crawl_abc123",
"task_type": "crawl",
"status": "failed",
"timestamp": "2025-10-22T10:30:00.000000+00:00",
"urls": ["https://example.com"],
"error": "Connection timeout after 30s"
}
from flask import Flask, request, jsonify
app = Flask(__name__)
@app.route('/webhook', methods=['POST'])
def handle_webhook():
payload = request.json
task_id = payload['task_id']
task_type = payload['task_type']
status = payload['status']
if status == 'completed':
if 'data' in payload:
# Process data directly
data = payload['data']
else:
# Fetch from API
endpoint = 'crawl' if task_type == 'crawl' else 'llm'
response = requests.get(f'http://localhost:11235/{endpoint}/job/{task_id}')
data = response.json()
# Your business logic here
print(f"Job {task_id} completed!")
elif status == 'failed':
error = payload.get('error', 'Unknown error')
print(f"Job {task_id} failed: {error}")
return jsonify({"status": "received"}), 200
app.run(port=8080)
None! This release is fully backward compatible.
No migration needed! Webhooks are opt-in:
webhook_config to your job payload# Just add webhook_config to your existing payload
payload = {
# Your existing configuration
"urls": ["https://example.com"],
"browser_config": {...},
"crawler_config": {...},
# NEW: Add webhook configuration
"webhook_config": {
"webhook_url": "https://myapp.com/webhook",
"webhook_data_in_payload": True
}
}
webhooks:
enabled: true
default_url: "https://myapp.com/webhooks/default" # Optional
data_in_payload: false
retry:
max_attempts: 5
initial_delay_ms: 1000
max_delay_ms: 32000
timeout_ms: 30000
headers:
User-Agent: "Crawl4AI-Webhook/1.0"
# Pull the latest image
docker pull unclecode/crawl4ai:0.7.6
# Or use latest tag
docker pull unclecode/crawl4ai:latest
# Run with webhook support
docker run -d \
-p 11235:11235 \
--env-file .llm.env \
--name crawl4ai \
unclecode/crawl4ai:0.7.6
pip install --upgrade crawl4ai
Try the release demo:
python docs/releases_review/demo_v0.7.6.py
This comprehensive demo showcases:
Thank you to the community for the feedback that shaped this feature! Special thanks to everyone who requested webhook support for asynchronous job processing.
Happy crawling with webhooks! 🕷️🪝
- unclecode
Proxy Parameter Deprecated - Use new proxy_config structure
Inject custom Python functions at 8 key pipeline points for authentication, performance optimization, and content processing.
Function-Based API with IDE support:
from crawl4ai import hooks_to_string
async def on_page_context_created(page, context, **kwargs):
"""Block images to speed up crawling"""
await context.route("**/*.{png,jpg,jpeg,gif,webp}", lambda route: route.abort())
return page
hooks_code = hooks_to_string({"on_page_context_created": on_page_context_created})
8 Available Hook Points:
on_browser_created, on_page_context_created, before_goto, after_goto, on_user_agent_updated, on_execution_started, before_retrieve_html, before_return_html
🤖 Enhanced LLM Integration
🔒 HTTPS Preservation
New preserve_https_for_internal_links option maintains secure protocols throughout crawling — critical for authenticated sessions and security-conscious applications.
🛠️ Major Bug Fixes
📦 Installation
PyPI: pip install crawl4ai==0.7.5
Docker: docker pull unclecode/crawl4ai:0.7.5 docker pull unclecode/crawl4ai:latest
Platforms Supported: Linux/AMD64, Linux/ARM64 (Apple Silicon, AWS Graviton)
⚠️ Breaking Changes
📚 Resources
🙏 Contributors
Thank you to everyone who reported issues, provided feedback, and contributed to this release!
Full Changelog: https://github.com/unclecode/crawl4ai/compare/v0.7.4...v0.7.5
September 29, 2025 • 8 min read
Today I'm releasing Crawl4AI v0.7.5—focused on extensibility and security. This update introduces the Docker Hooks System for pipeline customization, enhanced LLM integration, and important security improvements.
hooks_to_string() utility with Docker client auto-conversionEvery scraping project needs custom logic—authentication, performance optimization, content processing. Traditional solutions require forking or complex workarounds. Docker Hooks let you inject custom Python functions at 8 key points in the crawling pipeline.
import requests
# Real working hooks for httpbin.org
hooks_config = {
"on_page_context_created": """
async def hook(page, context, **kwargs):
print("Hook: Setting up page context")
# Block images to speed up crawling
await context.route("**/*.{png,jpg,jpeg,gif,webp}", lambda route: route.abort())
print("Hook: Images blocked")
return page
""",
"before_retrieve_html": """
async def hook(page, context, **kwargs):
print("Hook: Before retrieving HTML")
# Scroll to bottom to load lazy content
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
await page.wait_for_timeout(1000)
print("Hook: Scrolled to bottom")
return page
""",
"before_goto": """
async def hook(page, context, url, **kwargs):
print(f"Hook: About to navigate to {url}")
# Add custom headers
await page.set_extra_http_headers({
'X-Test-Header': 'crawl4ai-hooks-test'
})
return page
"""
}
# Test with Docker API
payload = {
"urls": ["https://httpbin.org/html"],
"hooks": {
"code": hooks_config,
"timeout": 30
}
}
response = requests.post("http://localhost:11235/crawl", json=payload)
result = response.json()
if result.get('success'):
print("✅ Hooks executed successfully!")
print(f"Content length: {len(result.get('markdown', ''))} characters")
Available Hook Points:
on_browser_created: Browser setupon_page_context_created: Page context configurationbefore_goto: Pre-navigation setupafter_goto: Post-navigation processingon_user_agent_updated: User agent changeson_execution_started: Crawl initializationbefore_retrieve_html: Pre-extraction processingbefore_return_html: Final HTML processingWriting hooks as strings works, but lacks IDE support and type checking. v0.7.5 introduces a function-based approach with automatic conversion!
Option 1: Using the hooks_to_string() Utility
from crawl4ai import hooks_to_string
import requests
# Define hooks as regular Python functions (with full IDE support!)
async def on_page_context_created(page, context, **kwargs):
"""Block images to speed up crawling"""
await context.route("**/*.{png,jpg,jpeg,gif,webp}", lambda route: route.abort())
await page.set_viewport_size({"width": 1920, "height": 1080})
return page
async def before_goto(page, context, url, **kwargs):
"""Add custom headers"""
await page.set_extra_http_headers({
'X-Crawl4AI': 'v0.7.5',
'X-Custom-Header': 'my-value'
})
return page
# Convert functions to strings
hooks_code = hooks_to_string({
"on_page_context_created": on_page_context_created,
"before_goto": before_goto
})
# Use with REST API
payload = {
"urls": ["https://httpbin.org/html"],
"hooks": {"code": hooks_code, "timeout": 30}
}
response = requests.post("http://localhost:11235/crawl", json=payload)
Option 2: Docker Client with Automatic Conversion (Recommended!)
from crawl4ai.docker_client import Crawl4aiDockerClient
# Define hooks as functions (same as above)
async def on_page_context_created(page, context, **kwargs):
await context.route("**/*.{png,jpg,jpeg,gif,webp}", lambda route: route.abort())
return page
async def before_retrieve_html(page, context, **kwargs):
# Scroll to load lazy content
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
await page.wait_for_timeout(1000)
return page
# Use Docker client - conversion happens automatically!
client = Crawl4aiDockerClient(base_url="http://localhost:11235")
results = await client.crawl(
urls=["https://httpbin.org/html"],
hooks={
"on_page_context_created": on_page_context_created,
"before_retrieve_html": before_retrieve_html
},
hooks_timeout=30
)
if results and results.success:
print(f"✅ Hooks executed! HTML length: {len(results.html)}")
Benefits of Function-Based Hooks:
Enhanced LLM integration with custom providers, temperature control, and base URL configuration.
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig
from crawl4ai.extraction_strategy import LLMExtractionStrategy
# Test with different providers
async def test_llm_providers():
# OpenAI with custom temperature
openai_strategy = LLMExtractionStrategy(
provider="gemini/gemini-2.5-flash-lite",
api_token="your-api-token",
temperature=0.7, # New in v0.7.5
instruction="Summarize this page in one sentence"
)
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(
"https://example.com",
config=CrawlerRunConfig(extraction_strategy=openai_strategy)
)
if result.success:
print("✅ LLM extraction completed")
print(result.extracted_content)
# Docker API with enhanced LLM config
llm_payload = {
"url": "https://example.com",
"f": "llm",
"q": "Summarize this page in one sentence.",
"provider": "gemini/gemini-2.5-flash-lite",
"temperature": 0.7
}
response = requests.post("http://localhost:11235/md", json=llm_payload)
New Features:
temperature parameter for creativity controlbase_url for custom API endpointsThe Problem: Modern web apps require HTTPS everywhere. When crawlers downgrade internal links from HTTPS to HTTP, authentication breaks and security warnings appear.
Solution: HTTPS preservation maintains secure protocols throughout crawling.
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, FilterChain, URLPatternFilter, BFSDeepCrawlStrategy
async def test_https_preservation():
# Enable HTTPS preservation
url_filter = URLPatternFilter(
patterns=["^(https:\/\/)?quotes\.toscrape\.com(\/.*)?$"]
)
config = CrawlerRunConfig(
exclude_external_links=True,
preserve_https_for_internal_links=True, # New in v0.7.5
deep_crawl_strategy=BFSDeepCrawlStrategy(
max_depth=2,
max_pages=5,
filter_chain=FilterChain([url_filter])
)
)
async with AsyncWebCrawler() as crawler:
async for result in await crawler.arun(
url="https://quotes.toscrape.com",
config=config
):
# All internal links maintain HTTPS
internal_links = [link['href'] for link in result.links['internal']]
https_links = [link for link in internal_links if link.startswith('https://')]
print(f"HTTPS links preserved: {len(https_links)}/{len(internal_links)}")
for link in https_links[:3]:
print(f" → {link}")
proxy parameter deprecated)This release addresses multiple issues reported by the community through GitHub issues and Discord discussions:
# Old proxy config (deprecated)
# browser_config = BrowserConfig(proxy="http://proxy:8080")
# New enhanced proxy config
browser_config = BrowserConfig(
proxy_config={
"server": "http://proxy:8080",
"username": "optional-user",
"password": "optional-pass"
}
)
proxy_config structurecssselect for better CSS handling# Install latest version
pip install crawl4ai==0.7.5
# Docker deployment
docker pull unclecode/crawl4ai:latest
docker run -p 11235:11235 unclecode/crawl4ai:latest
Try the Demo:
# Run working examples
python docs/releases_review/demo_v0.7.5.py
Resources:
Happy crawling! 🕷️
PyPI: `bash pip install crawl4ai==0.7.4 `
PyPI:
pip install crawl4ai==0.7.4
Docker:
docker pull unclecode/crawl4ai:0.7.4
docker pull unclecode/crawl4ai:latest
See CHANGELOG.md for details.
Welcome to Crawl4AI v0.7.3! This release brings powerful new capabilities for stealth crawling, intelligent URL configuration, memory optimization, an
Welcome to Crawl4AI v0.7.3! This release brings powerful new capabilities for stealth crawling, intelligent URL configuration, memory optimization, and enhanced data extraction. Whether you're dealing with bot-protected sites, mixed content types, or large-scale crawling operations, this update has you covered.
After powering 51,000+ developers and becoming the #1 trending web crawler, we're launching GitHub Sponsors to ensure Crawl4AI stays independent and innovative forever.
Why sponsor? Own your data pipeline. No API limits. Direct access to the creator.
Become a Sponsor → | See Benefits
Break through sophisticated bot detection systems with our new stealth capabilities:
from crawl4ai import AsyncWebCrawler, BrowserConfig
# Enable stealth mode for undetectable crawling
browser_config = BrowserConfig(
browser_type="undetected", # Use undetected Chrome
headless=True, # Can run headless with stealth
extra_args=[
"--disable-blink-features=AutomationControlled",
"--disable-web-security"
]
)
async with AsyncWebCrawler(config=browser_config) as crawler:
# Successfully bypass Cloudflare, Akamai, and custom bot detection
result = await crawler.arun("https://protected-site.com")
print(f"✅ Bypassed protection! Content: {len(result.markdown)} chars")
What it enables:
Apply different crawling strategies to different URL patterns automatically:
from crawl4ai import CrawlerRunConfig
# Define specialized configs for different content types
configs = [
# Documentation sites - aggressive caching, include links
CrawlerRunConfig(
url_matcher=["*docs*", "*documentation*"],
cache_mode="write",
markdown_generator_options={"include_links": True}
),
# News/blog sites - fresh content, scroll for lazy loading
CrawlerRunConfig(
url_matcher=lambda url: 'blog' in url or 'news' in url,
cache_mode="bypass",
js_code="window.scrollTo(0, document.body.scrollHeight/2);"
),
# API endpoints - structured extraction
CrawlerRunConfig(
url_matcher=["*.json", "*api*"],
extraction_strategy=LLMExtractionStrategy(
provider="openai/gpt-4o-mini",
extraction_type="structured"
)
),
# Default fallback for everything else
CrawlerRunConfig()
]
# Crawl multiple URLs with perfect configurations
results = await crawler.arun_many([
"https://docs.python.org/3/", # → Uses documentation config
"https://blog.python.org/", # → Uses blog config
"https://api.github.com/users", # → Uses API config
"https://example.com/" # → Uses default config
], config=configs)
Perfect for:
Track and optimize memory usage during large-scale operations:
from crawl4ai.memory_utils import MemoryMonitor
# Monitor memory during crawling
monitor = MemoryMonitor()
monitor.start_monitoring()
# Perform memory-intensive operations
results = await crawler.arun_many([
"https://heavy-js-site.com",
"https://large-images-site.com",
"https://dynamic-content-site.com"
] * 100) # Large batch
# Get detailed memory report
report = monitor.get_report()
print(f"Peak memory usage: {report['peak_mb']:.1f} MB")
print(f"Memory efficiency: {report['efficiency']:.1f}%")
# Automatic optimization suggestions
if report['peak_mb'] > 1000: # > 1GB
print("💡 Consider batch size optimization")
print("💡 Enable aggressive garbage collection")
Benefits:
Direct pandas DataFrame conversion from web tables:
result = await crawler.arun("https://site-with-tables.com")
# New streamlined approach
if result.tables:
print(f"Found {len(result.tables)} tables")
import pandas as pd
for i, table in enumerate(result.tables):
# Instant DataFrame conversion
df = pd.DataFrame(table['data'])
print(f"Table {i}: {df.shape[0]} rows × {df.shape[1]} columns")
print(df.head())
# Rich metadata available
print(f"Source: {table.get('source_xpath', 'Unknown')}")
print(f"Headers: {table.get('headers', [])}")
# Old way (now deprecated)
# tables_data = result.media.get('tables', []) # ❌ Don't use this
Improvements:
Switch between LLM providers without rebuilding images:
# Option 1: Direct environment variables
docker run -d \
-e LLM_PROVIDER="groq/llama-3.2-3b-preview" \
-e GROQ_API_KEY="your-key" \
-p 11235:11235 \
unclecode/crawl4ai:0.7.3
# Option 2: Using .llm.env file (recommended for production)
docker run -d \
--env-file .llm.env \
-p 11235:11235 \
unclecode/crawl4ai:0.7.3
Create .llm.env file:
LLM_PROVIDER=openai/gpt-4o-mini
OPENAI_API_KEY=your-openai-key
GROQ_API_KEY=your-groq-key
Override per request when needed:
# Use cheaper models for simple tasks, premium for complex ones
response = requests.post("http://localhost:11235/crawl", json={
"url": "https://complex-page.com",
"extraction_strategy": {
"type": "llm",
"provider": "openai/gpt-4" # Override default
}
})
result.media to result.tables# Fresh install
pip install crawl4ai==0.7.3
# Upgrade from previous version
pip install --upgrade crawl4ai==0.7.3
# Specific version
docker pull unclecode/crawl4ai:0.7.3
# Latest (points to 0.7.3)
docker pull unclecode/crawl4ai:latest
# Version aliases
docker pull unclecode/crawl4ai:0.7 # Minor version
docker pull unclecode/crawl4ai:0 # Major version
result.tables replaces result.media.get('tables')browser_type="undetected"url_matcher parameter in CrawlerRunConfigThis release sets the foundation for even more advanced features coming in v0.8:
Live Long and import crawl4ai
Crawl4AI continues to evolve with your needs. This release makes it stealthier, smarter, and more scalable. Try the new undetected browser and multi-config features—they're game changers!
- The Crawl4AI Team
📝 This release draft was composed and edited by human but rewritten and finalized by AI. If you notice any mistakes, please raise an issue.
🕵️ Undetected Browser Support: New browser adapter pattern with stealth capabilities
browser_adapter.py with undetected Chrome integration🎨 Multi-URL Configuration System: URL-specific crawler configurations for batch processing
"*.pdf", "*/blog/*")🧠 Memory Monitoring & Optimization: Comprehensive memory usage tracking
memory_utils.py module for memory monitoring and optimization📊 Enhanced Table Extraction: Improved table access and DataFrame conversion
result.tables interface replacing generic result.media approachpd.DataFrame(table['data'])💰 GitHub Sponsors Integration: 4-tier sponsorship system
🐳 Docker LLM Provider Flexibility: Environment-based LLM configuration
LLM_PROVIDER environment variable support for dynamic provider switching.llm.env file support for secure configuration managementasync_crawler_strategy.py to backupresult.media to result.tablesAugust 6, 2025 • 5 min read
Today I'm releasing Crawl4AI v0.7.3—the Multi-Config Intelligence Update. This release brings smarter URL-specific configurations, flexible Docker deployments, important bug fixes, and documentation improvements that make Crawl4AI more robust and production-ready.
The Problem: You're crawling a mix of documentation sites, blogs, and API endpoints. Each needs different handling—caching for docs, fresh content for news, structured extraction for APIs. Previously, you'd run separate crawls or write complex conditional logic.
My Solution: I implemented URL-specific configurations that let you define different strategies for different URL patterns in a single crawl batch. First match wins, with optional fallback support.
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, MatchMode
# Define specialized configs for different content types
configs = [
# Documentation sites - aggressive caching, include links
CrawlerRunConfig(
url_matcher=["*docs*", "*documentation*"],
cache_mode="write",
markdown_generator_options={"include_links": True}
),
# News/blog sites - fresh content, scroll for lazy loading
CrawlerRunConfig(
url_matcher=lambda url: 'blog' in url or 'news' in url,
cache_mode="bypass",
js_code="window.scrollTo(0, document.body.scrollHeight/2);"
),
# API endpoints - structured extraction
CrawlerRunConfig(
url_matcher=["*.json", "*api*"],
extraction_strategy=LLMExtractionStrategy(
provider="openai/gpt-4o-mini",
extraction_type="structured"
)
),
# Default fallback for everything else
CrawlerRunConfig() # No url_matcher = matches everything
]
# Crawl multiple URLs with appropriate configs
async with AsyncWebCrawler() as crawler:
results = await crawler.arun_many(
urls=[
"https://docs.python.org/3/", # → Uses documentation config
"https://blog.python.org/", # → Uses blog config
"https://api.github.com/users", # → Uses API config
"https://example.com/" # → Uses default config
],
config=configs
)
Matching Capabilities:
"*.pdf", "*/blog/*"Expected Real-World Impact:
The Problem: Hardcoded LLM providers in Docker deployments. Want to switch from OpenAI to Groq? Rebuild and redeploy. Testing different models? Multiple Docker images.
My Solution: Configure LLM providers via environment variables. Switch providers without touching code or rebuilding images.
# Option 1: Direct environment variables
docker run -d \
-e LLM_PROVIDER="groq/llama-3.2-3b-preview" \
-e GROQ_API_KEY="your-key" \
-p 11235:11235 \
unclecode/crawl4ai:latest
# Option 2: Using .llm.env file (recommended for production)
# Create .llm.env file:
# LLM_PROVIDER=openai/gpt-4o-mini
# OPENAI_API_KEY=your-openai-key
# GROQ_API_KEY=your-groq-key
docker run -d \
--env-file .llm.env \
-p 11235:11235 \
unclecode/crawl4ai:latest
Override per request when needed:
# Use default provider from .llm.env
response = requests.post("http://localhost:11235/crawl", json={
"url": "https://example.com",
"extraction_strategy": {"type": "llm"}
})
# Override to use different provider for this specific request
response = requests.post("http://localhost:11235/crawl", json={
"url": "https://complex-page.com",
"extraction_strategy": {
"type": "llm",
"provider": "openai/gpt-4" # Override default
}
})
Expected Real-World Impact:
.llm.env file, not in commandsThis release includes several important bug fixes that improve stability and reliability:
Based on community feedback, we've updated:
Thanks to our contributors and the entire community for feedback and bug reports.
Crawl4AI continues to evolve with your needs. This release makes it smarter, more flexible, and more stable. Try the new multi-config feature and flexible Docker deployment—they're game changers!
Happy Crawling! 🕷️
- The Crawl4AI Team
No breaking changes - direct upgrade from v0.7.0 or v0.7.1.
July 25, 2025 • 3 min read
This release introduces automated CI/CD pipelines for seamless releases and optimizes dependencies for a lighter, more efficient package.
The new automated release process ensures consistent, reliable releases:
# Trigger releases with a simple tag
git tag v0.7.2
git push origin v0.7.2
# Automatically:
# ✅ Validates version consistency
# ✅ Builds and publishes to PyPI
# ✅ Builds multi-platform Docker images
# ✅ Pushes to Docker Hub with proper tags
# ✅ Creates GitHub release
Default installation is now significantly smaller:
# Core installation (smaller, faster)
pip install crawl4ai==0.7.2
# With ML features (includes sentence-transformers)
pip install crawl4ai[transformer]==0.7.2
# Full installation
pip install crawl4ai[all]==0.7.2
Enhanced Docker support with multi-platform images:
# Pull the latest version
docker pull unclecode/crawl4ai:0.7.2
docker pull unclecode/crawl4ai:latest
# Available tags:
# - unclecode/crawl4ai:0.7.2 (specific version)
# - unclecode/crawl4ai:0.7 (minor version)
# - unclecode/crawl4ai:0 (major version)
# - unclecode/crawl4ai:latest
sentence-transformers moved from required to optional dependenciespip install crawl4ai==0.7.2
No breaking changes - direct upgrade from v0.7.0 or v0.7.1.
Questions? Issues?
P.S. The new CI/CD pipeline will make future releases faster and more reliable. Thanks for your patience as we improve our release process!
No breaking changes - upgrade directly from v0.7.0.
July 17, 2025 • 2 min read
A small maintenance release that removes unused code and improves documentation.
crawl4ai/browser_manager.pyRemoved unused StealthConfig import and configuration that wasn't being used anywhere in the codebase. The project uses its own custom stealth implementation through JavaScript injection instead.
# Removed unused code:
from playwright_stealth import StealthConfig
stealth_config = StealthConfig(...) # This was never used
pip install crawl4ai==0.7.1
No breaking changes - upgrade directly from v0.7.0.
Questions? Issues?
*January 28, 2025 • 10 min read*
January 28, 2025 • 10 min read
Today I'm releasing Crawl4AI v0.7.0—the Adaptive Intelligence Update. This release introduces fundamental improvements in how Crawl4AI handles modern web complexity through adaptive learning, intelligent content discovery, and advanced extraction capabilities.
The Problem: Websites change. Class names shift. IDs disappear. Your carefully crafted selectors break at 3 AM, and you wake up to empty datasets and angry stakeholders.
My Solution: I implemented an adaptive learning system that observes patterns, builds confidence scores, and adjusts extraction strategies on the fly. It's like having a junior developer who gets better at their job with every page they scrape.
The Adaptive Crawler maintains a persistent state for each domain, tracking:
from crawl4ai import AdaptiveCrawler, AdaptiveConfig, CrawlState
# Initialize with custom learning parameters
config = AdaptiveConfig(
confidence_threshold=0.7, # Min confidence to use learned patterns
max_history=100, # Remember last 100 crawls per domain
learning_rate=0.2, # How quickly to adapt to changes
patterns_per_page=3, # Patterns to learn per page type
extraction_strategy='css' # 'css' or 'xpath'
)
adaptive_crawler = AdaptiveCrawler(config)
# First crawl - crawler learns the structure
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(
"https://news.example.com/article/12345",
config=CrawlerRunConfig(
adaptive_config=config,
extraction_hints={ # Optional hints to speed up learning
"title": "article h1",
"content": "article .body-content"
}
)
)
# Crawler identifies and stores patterns
if result.success:
state = adaptive_crawler.get_state("news.example.com")
print(f"Learned {len(state.patterns)} patterns")
print(f"Confidence: {state.avg_confidence:.2%}")
# Subsequent crawls - uses learned patterns
result2 = await crawler.arun(
"https://news.example.com/article/67890",
config=CrawlerRunConfig(adaptive_config=config)
)
# Automatically extracts using learned patterns!
Expected Real-World Impact:
The Problem: Modern web apps only render what's visible. Scroll down, new content appears, old content vanishes into the void. Traditional crawlers capture that first viewport and miss 90% of the content. It's like reading only the first page of every book.
My Solution: I built Virtual Scroll support that mimics human browsing behavior, capturing content as it loads and preserving it before the browser's garbage collector strikes.
from crawl4ai import VirtualScrollConfig
# For social media feeds (Twitter/X style)
twitter_config = VirtualScrollConfig(
container_selector="[data-testid='primaryColumn']",
scroll_count=20, # Number of scrolls
scroll_by="container_height", # Smart scrolling by container size
wait_after_scroll=1.0, # Let content load
capture_method="incremental", # Capture new content on each scroll
deduplicate=True # Remove duplicate elements
)
# For e-commerce product grids (Instagram style)
grid_config = VirtualScrollConfig(
container_selector="main .product-grid",
scroll_count=30,
scroll_by=800, # Fixed pixel scrolling
wait_after_scroll=1.5, # Images need time
stop_on_no_change=True # Smart stopping
)
# For news feeds with lazy loading
news_config = VirtualScrollConfig(
container_selector=".article-feed",
scroll_count=50,
scroll_by="page_height", # Viewport-based scrolling
wait_after_scroll=0.5,
wait_for_selector=".article-card", # Wait for specific elements
timeout=30000 # Max 30 seconds total
)
# Use it in your crawl
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(
"https://twitter.com/trending",
config=CrawlerRunConfig(
virtual_scroll_config=twitter_config,
# Combine with other features
extraction_strategy=JsonCssExtractionStrategy({
"tweets": {
"selector": "[data-testid='tweet']",
"fields": {
"text": {"selector": "[data-testid='tweetText']", "type": "text"},
"likes": {"selector": "[data-testid='like']", "type": "text"}
}
}
})
)
)
print(f"Captured {len(result.extracted_content['tweets'])} tweets")
Key Capabilities:
Expected Real-World Impact:
The Problem: You crawl a page and get 200 links. Which ones matter? Which lead to the content you actually want? Traditional crawlers force you to follow everything or build complex filters.
My Solution: I implemented a three-layer scoring system that analyzes links like a human would—considering their position, context, and relevance to your goals.
from crawl4ai import LinkPreviewConfig
# Configure intelligent link analysis
link_config = LinkPreviewConfig(
# What to analyze
include_internal=True,
include_external=True,
max_links=100, # Analyze top 100 links
# Relevance scoring
query="machine learning tutorials", # Your interest
score_threshold=0.3, # Minimum relevance score
# Performance
concurrent_requests=10, # Parallel processing
timeout_per_link=5000, # 5s per link
# Advanced scoring weights
scoring_weights={
"intrinsic": 0.3, # Link quality indicators
"contextual": 0.5, # Relevance to query
"popularity": 0.2 # Link prominence
}
)
# Use in your crawl
result = await crawler.arun(
"https://tech-blog.example.com",
config=CrawlerRunConfig(
link_preview_config=link_config,
score_links=True
)
)
# Access scored and sorted links
for link in result.links["internal"][:10]: # Top 10 internal links
print(f"Score: {link['total_score']:.3f}")
print(f" Intrinsic: {link['intrinsic_score']:.1f}/10") # Position, attributes
print(f" Contextual: {link['contextual_score']:.1f}/1") # Relevance to query
print(f" URL: {link['href']}")
print(f" Title: {link['head_data']['title']}")
print(f" Description: {link['head_data']['meta']['description'][:100]}...")
Scoring Components:
Intrinsic Score (0-10): Based on link quality indicators
Contextual Score (0-1): Relevance to your query
Total Score: Weighted combination for final ranking
Expected Real-World Impact:
The Problem: You want to crawl an entire domain but only have the homepage. Or worse, you want specific content types across thousands of pages. Manual URL discovery? That's a job for machines, not humans.
My Solution: I built Async URL Seeder—a turbocharged URL discovery engine that combines multiple sources with intelligent filtering and relevance scoring.
from crawl4ai import AsyncUrlSeeder, SeedingConfig
# Basic discovery - find all product pages
seeder_config = SeedingConfig(
# Discovery sources
source="sitemap+cc", # Sitemap + Common Crawl
# Filtering
pattern="*/product/*", # URL pattern matching
ignore_patterns=["*/reviews/*", "*/questions/*"],
# Validation
live_check=True, # Verify URLs are alive
max_urls=5000, # Stop at 5000 URLs
# Performance
concurrency=100, # Parallel requests
hits_per_sec=10 # Rate limiting
)
seeder = AsyncUrlSeeder(seeder_config)
urls = await seeder.discover("https://shop.example.com")
# Advanced: Relevance-based discovery
research_config = SeedingConfig(
source="crawl+sitemap", # Deep crawl + sitemap
pattern="*/blog/*", # Blog posts only
# Content relevance
extract_head=True, # Get meta tags
query="quantum computing tutorials",
scoring_method="bm25", # Or "semantic" (coming soon)
score_threshold=0.4, # High relevance only
# Smart filtering
filter_nonsense_urls=True, # Remove .xml, .txt, etc.
min_content_length=500, # Skip thin content
force=True # Bypass cache
)
# Discover with progress tracking
discovered = []
async for batch in seeder.discover_iter("https://physics-blog.com", research_config):
discovered.extend(batch)
print(f"Found {len(discovered)} relevant URLs so far...")
# Results include scores and metadata
for url_data in discovered[:5]:
print(f"URL: {url_data['url']}")
print(f"Score: {url_data['score']:.3f}")
print(f"Title: {url_data['title']}")
Discovery Methods:
Expected Real-World Impact:
This release includes significant performance improvements through optimized resource handling, better concurrency management, and reduced memory footprint.
# Before v0.7.0 (slow)
results = []
for url in urls:
result = await crawler.arun(url)
results.append(result)
# After v0.7.0 (fast)
# Automatic batching and connection pooling
results = await crawler.arun_batch(
urls,
config=CrawlerRunConfig(
# New performance options
batch_size=10, # Process 10 URLs concurrently
reuse_browser=True, # Keep browser warm
eager_loading=False, # Load only what's needed
streaming_extraction=True, # Stream large extractions
# Optimized defaults
wait_until="domcontentloaded", # Faster than networkidle
exclude_external_resources=True, # Skip third-party assets
block_ads=True # Ad blocking built-in
)
)
# Memory-efficient streaming for large crawls
async for result in crawler.arun_stream(large_url_list):
# Process results as they complete
await process_result(result)
# Memory is freed after each iteration
Performance Gains:
PDF extraction is now natively supported in Crawl4AI.
# Extract data from PDF documents
result = await crawler.arun(
"https://example.com/report.pdf",
config=CrawlerRunConfig(
pdf_extraction=True,
extraction_strategy=JsonCssExtractionStrategy({
# Works on converted PDF structure
"title": {"selector": "h1", "type": "text"},
"sections": {"selector": "h2", "type": "list"}
})
)
)
link_extractor renamed to link_preview (better reflects functionality)CrawlerConfig split into CrawlerRunConfig and BrowserConfig# Old (v0.6.x)
from crawl4ai import CrawlerConfig
config = CrawlerConfig(timeout=30000)
# New (v0.7.0)
from crawl4ai import CrawlerRunConfig, BrowserConfig
browser_config = BrowserConfig(timeout=30000)
run_config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)
I'm currently working on bringing advanced automation capabilities to Crawl4AI. This includes:
These features are under active development and will revolutionize how we approach web automation. Stay tuned!
pip install crawl4ai==0.7.0
Check out the updated documentation.
Questions? Issues? I'm always listening:
Happy crawling! 🕷️
P.S. If you're using Crawl4AI in production, I'd love to hear about it. Your use cases inspire the next features.
January 28, 2025 • 10 min read
Today I'm releasing Crawl4AI v0.7.0—the Adaptive Intelligence Update. This release introduces fundamental improvements in how Crawl4AI handles modern web complexity through adaptive learning, intelligent content discovery, and advanced extraction capabilities.
The Problem: Websites change. Class names shift. IDs disappear. Your carefully crafted selectors break at 3 AM, and you wake up to empty datasets and angry stakeholders.
My Solution: I implemented an adaptive learning system that observes patterns, builds confidence scores, and adjusts extraction strategies on the fly. It's like having a junior developer who gets better at their job with every page they scrape.
The Adaptive Crawler maintains a persistent state for each domain, tracking:
from crawl4ai import AsyncWebCrawler, AdaptiveCrawler, AdaptiveConfig
import asyncio
async def main():
# Configure adaptive crawler
config = AdaptiveConfig(
strategy="statistical", # or "embedding" for semantic understanding
max_pages=10,
confidence_threshold=0.7, # Stop at 70% confidence
top_k_links=3, # Follow top 3 links per page
min_gain_threshold=0.05 # Need 5% information gain to continue
)
async with AsyncWebCrawler(verbose=False) as crawler:
adaptive = AdaptiveCrawler(crawler, config)
print("Starting adaptive crawl about Python decorators...")
result = await adaptive.digest(
start_url="https://docs.python.org/3/glossary.html",
query="python decorators functions wrapping"
)
print(f"\n✅ Crawling Complete!")
print(f"• Confidence Level: {adaptive.confidence:.0%}")
print(f"• Pages Crawled: {len(result.crawled_urls)}")
print(f"• Knowledge Base: {len(adaptive.state.knowledge_base)} documents")
# Get most relevant content
relevant = adaptive.get_relevant_content(top_k=3)
print(f"\nMost Relevant Pages:")
for i, page in enumerate(relevant, 1):
print(f"{i}. {page['url']} (relevance: {page['score']:.2%})")
asyncio.run(main())
Expected Real-World Impact:
The Problem: Modern web apps only render what's visible. Scroll down, new content appears, old content vanishes into the void. Traditional crawlers capture that first viewport and miss 90% of the content. It's like reading only the first page of every book.
My Solution: I built Virtual Scroll support that mimics human browsing behavior, capturing content as it loads and preserving it before the browser's garbage collector strikes.
from crawl4ai import VirtualScrollConfig
# For social media feeds (Twitter/X style)
twitter_config = VirtualScrollConfig(
container_selector="[data-testid='primaryColumn']",
scroll_count=20, # Number of scrolls
scroll_by="container_height", # Smart scrolling by container size
wait_after_scroll=1.0 # Let content load
)
# For e-commerce product grids (Instagram style)
grid_config = VirtualScrollConfig(
container_selector="main .product-grid",
scroll_count=30,
scroll_by=800, # Fixed pixel scrolling
wait_after_scroll=1.5 # Images need time
)
# For news feeds with lazy loading
news_config = VirtualScrollConfig(
container_selector=".article-feed",
scroll_count=50,
scroll_by="page_height", # Viewport-based scrolling
wait_after_scroll=0.5 # Wait for content to load
)
# Use it in your crawl
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(
"https://twitter.com/trending",
config=CrawlerRunConfig(
virtual_scroll_config=twitter_config,
# Combine with other features
extraction_strategy=JsonCssExtractionStrategy({
"tweets": {
"selector": "[data-testid='tweet']",
"fields": {
"text": {"selector": "[data-testid='tweetText']", "type": "text"},
"likes": {"selector": "[data-testid='like']", "type": "text"}
}
}
})
)
)
print(f"Captured {len(result.extracted_content['tweets'])} tweets")
Key Capabilities:
Expected Real-World Impact:
The Problem: You crawl a page and get 200 links. Which ones matter? Which lead to the content you actually want? Traditional crawlers force you to follow everything or build complex filters.
My Solution: I implemented a three-layer scoring system that analyzes links like a human would—considering their position, context, and relevance to your goals.
import asyncio
from crawl4ai import CrawlerRunConfig, CacheMode, AsyncWebCrawler
from crawl4ai.adaptive_crawler import LinkPreviewConfig
async def main():
# Configure intelligent link analysis
link_config = LinkPreviewConfig(
include_internal=True,
include_external=False,
max_links=10,
concurrency=5,
query="python tutorial", # For contextual scoring
score_threshold=0.3,
verbose=True
)
# Use in your crawl
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(
"https://www.geeksforgeeks.org/",
config=CrawlerRunConfig(
link_preview_config=link_config,
score_links=True, # Enable intrinsic scoring
cache_mode=CacheMode.BYPASS
)
)
# Access scored and sorted links
if result.success and result.links:
for link in result.links.get("internal", []):
text = link.get('text', 'No text')[:40]
print(
text,
f"{link.get('intrinsic_score', 0):.1f}/10" if link.get('intrinsic_score') is not None else "0.0/10",
f"{link.get('contextual_score', 0):.2f}/1" if link.get('contextual_score') is not None else "0.00/1",
f"{link.get('total_score', 0):.3f}" if link.get('total_score') is not None else "0.000"
)
asyncio.run(main())
Scoring Components:
Intrinsic Score: Based on link quality indicators
Contextual Score: Relevance to your query using BM25 algorithm
Total Score: Combined score for final ranking
Expected Real-World Impact:
The Problem: You want to crawl an entire domain but only have the homepage. Or worse, you want specific content types across thousands of pages. Manual URL discovery? That's a job for machines, not humans.
My Solution: I built Async URL Seeder—a turbocharged URL discovery engine that combines multiple sources with intelligent filtering and relevance scoring.
import asyncio
from crawl4ai import AsyncUrlSeeder, SeedingConfig
async def main():
async with AsyncUrlSeeder() as seeder:
# Discover Python tutorial URLs
config = SeedingConfig(
source="sitemap", # Use sitemap
pattern="*python*", # URL pattern filter
extract_head=True, # Get metadata
query="python tutorial", # For relevance scoring
scoring_method="bm25",
score_threshold=0.2,
max_urls=10
)
print("Discovering Python async tutorial URLs...")
urls = await seeder.urls("https://www.geeksforgeeks.org/", config)
print(f"\n✅ Found {len(urls)} relevant URLs:")
for i, url_info in enumerate(urls[:5], 1):
print(f"\n{i}. {url_info['url']}")
if url_info.get('relevance_score'):
print(f" Relevance: {url_info['relevance_score']:.3f}")
if url_info.get('head_data', {}).get('title'):
print(f" Title: {url_info['head_data']['title'][:60]}...")
asyncio.run(main())
Discovery Methods:
Expected Real-World Impact:
This release includes significant performance improvements through optimized resource handling, better concurrency management, and reduced memory footprint.
# Optimized crawling with v0.7.0 improvements
results = []
for url in urls:
result = await crawler.arun(
url,
config=CrawlerRunConfig(
# Performance optimizations
wait_until="domcontentloaded", # Faster than networkidle
cache_mode=CacheMode.ENABLED # Enable caching
)
)
results.append(result)
Performance Gains:
link_extractor renamed to link_preview (better reflects functionality)CrawlerConfig split into CrawlerRunConfig and BrowserConfig# Old (v0.6.x)
from crawl4ai import CrawlerConfig
config = CrawlerConfig(timeout=30000)
# New (v0.7.0)
from crawl4ai import CrawlerRunConfig, BrowserConfig
browser_config = BrowserConfig(timeout=30000)
run_config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS)
I'm currently working on bringing advanced automation capabilities to Crawl4AI. This includes:
These features are under active development and will revolutionize how we approach web automation. Stay tuned!
pip install crawl4ai==0.7.0
Check out the updated documentation.
Questions? Issues? I'm always listening:
Happy crawling! 🕷️
P.S. If you're using Crawl4AI in production, I'd love to hear about it. Your use cases inspire the next features.
…callers now must explicitly close pages (BREAKING CHANGE)
Release 0.6.3 (unreleased)
Features
RegexExtractionStrategy for pattern-based extraction, including built-in patterns for emails, URLs, phones, dates, support for custom regexes, an LLM-assisted pattern generator, optimized HTML preprocessing via fit_html, and enhanced network response body capture (9b5ccac)POST /crawl/job & GET /crawl/job/{task_id} for crawls, POST /llm/job & GET /llm/job/{task_id} for LLM tasks—backed by Redis task management with configurable TTL, moved schemas to schemas.py, and added demo_docker_polling.py example (94e9959)Fixes
take_screenshot and take_screenshot_naive, preventing premature teardown; callers now must explicitly close pages (BREAKING CHANGE) (a3e9ef9)Documentation
docs/apps/linkdin/README.md so examples copy & paste cleanly (87d4b0f)litellm argument details for correct script usage (bd5a9ac)Refactoring
Enum in async_logger, browser_profiler, content_filter_strategy and related modules for cleaner, type-safe formatting (cd2b490)Experimental
rich (WIP, work ongoing) (b2f3cb0)New RegexExtractionStrategy for fast pattern-based extraction without requiring LLM
RegexExtractionStrategy for fast pattern-based extraction without requiring LLM
generate_pattern utility for LLM-assisted pattern creation (one-time use)fit_html as a top-level field in CrawlResult for optimized HTML extractionNew dedicated tables field in CrawlResult model for better table extraction handling
tables field in CrawlResult model for better table extraction handlingSwitch to DefaultMarkdownGenerator to silence deprecation warnings.
crun_cfg = CrawlerRunConfig(
url="https://browserleaks.com/geo", # test page that shows your location
locale="en-US", # Accept-Language & UI locale
timezone_id="America/Los_Angeles", # JS Date()/Intl timezone
geolocation=GeolocationConfig( # override GPS coords
latitude=34.0522,
longitude=-118.2437,
accuracy=10.0,
)
)
df = pd.DataFrame(result.media["tables"][0]["rows"], columns=result.media["tables"][0]["headers"]) and get CSV or pandas without extra parsing.tests/memory) for 1 k+ URL runs.ProxyConfig moved to async_configs.crawl4ai/browser/*.DefaultMarkdownGenerator and warn.crawl4ai/browser/* to the new pooled browser modules.AsyncPlaywrightCrawlerStrategy.get_page adopt the new signature.DefaultMarkdownGenerator to silence deprecation warnings.121 files changed, ≈36 223 insertions, ≈4 975 deletions
<table>s into DataFrames or CSV with one flagtests/memory and API load scriptsProxyConfig moved to async_configscrawl4ai/browser/* superseded by the new pooled browser layerDefaultMarkdownGenerator and emit warningscrawl4ai/browser/* to the new pooled browser modulesAsyncPlaywrightCrawlerStrategy.get_page, adopt the new signatureDefaultMarkdownGenerator (or silence the deprecation warning)121 files changed, ≈36 223 insertions, ≈4 975 deletions :contentReference[oaicite:0]{index=0}:contentReference[oaicite:1]{index=1}
content_source parameter allows choosing between cleaned_html, raw_html, and fit_htmlcleaned_html behaviorWe're excited to announce the release of Crawl4AI v0.6.0, our biggest and most feature-rich update yet. This version introduces major architectural upgrades, brand-new capabilities for geo-aware crawling, high-efficiency scraping, and real-time streaming support for scalable deployments.
Crawl as if you’re anywhere in the world. With v0.6.0, each crawl can simulate:
Example:
CrawlerRunConfig(
url="https://browserleaks.com/geo",
locale="en-US",
timezone_id="America/Los_Angeles",
geolocation=GeolocationConfig(
latitude=34.0522,
longitude=-118.2437,
accuracy=10.0
)
)
Great for accessing region-specific content or testing global behavior.
Extract HTML tables directly into usable formats like Pandas DataFrames or CSV with zero parsing hassle. All table data is available under result.media["tables"].
Example:
raw_df = pd.DataFrame(
result.media["tables"][0]["rows"],
columns=result.media["tables"][0]["headers"]
)
This makes it ideal for scraping financial data, pricing pages, or anything tabular.
We've overhauled browser management. Now, multiple browser instances can be pooled and pages pre-warmed for ultra-fast launches:
This powers the new Docker Playground experience and streamlines heavy-load crawling.
Need full visibility? You can now capture:
No more guesswork on what happened during your crawl.
We’re exposing MCP socket and SSE endpoints, allowing:
This is a major step towards making Crawl4AI real-time ready.
Want to test performance under heavy load? v0.6.0 includes a new memory stress-test suite that supports 1,000+ URL workloads. Ideal for:
crawl4ai/browser/* modules are removed. Update imports accordingly.AsyncPlaywrightCrawlerStrategy.get_page now uses a new function signature.DefaultMarkdownGenerator with warning.Want a visual walkthrough of all these updates? Watch the video: 🔗 https://youtu.be/9x7nVcjOZks
If you're new to Crawl4AI, start here: 🔗 https://www.youtube.com/watch?v=xo3qK6Hg9AA&t=15s
We’ve just opened up our Discord for the public. Join us to:
💬 https://discord.gg/wpYFACrHR4
pip install -U crawl4ai
Live long and import crawl4ai. 🖖
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
*(crawler)* Add experimental parameters dictionary to CrawlerRunConfig to support beta features
Nothing published for this version
Nothing published for this version
Nothing published for this version
This release contains several breaking changes. Please review the full release notes for migration guidance.
Release Theme: Power, Flexibility, and Scalability
Crawl4AI v0.5.0 is a major release focused on significantly enhancing the library's power, flexibility, and scalability.
crwl CLI provides convenient access to all features with intuitive commandslxml library for 10-20x speedup with complex pagesThis release contains several breaking changes. Please review the full release notes for migration guidance.
For complete details, visit: https://docs.crawl4ai.com/blog/releases/0.5.0/
*(profiles)* Add BrowserProfiler class for dedicated browser profile management
Release Theme: Power, Flexibility, and Scalability
Crawl4AI v0.5.0 is a major release focused on significantly enhancing the library's power, flexibility, and scalability. Key improvements include a new deep crawling system, a memory-adaptive dispatcher for handling large-scale crawls, multiple crawling strategies (including a fast HTTP-only crawler), Docker deployment options, and a powerful command-line interface (CLI). This release also includes numerous bug fixes, performance optimizations, and documentation updates.
Important Note: This release contains several breaking changes. Please review the "Breaking Changes" section carefully and update your code accordingly.
Crawl4AI now supports deep crawling, allowing you to explore websites beyond the
initial URLs. This is controlled by the deep_crawl_strategy parameter in
CrawlerRunConfig. Several strategies are available:
BFSDeepCrawlStrategy (Breadth-First Search): Explores the website level
by level. (Default)DFSDeepCrawlStrategy (Depth-First Search): Explores each branch as
deeply as possible before backtracking.BestFirstCrawlingStrategy: Uses a scoring function to prioritize which
URLs to crawl next.import time
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, BFSDeepCrawlStrategy
from crawl4ai.content_scraping_strategy import LXMLWebScrapingStrategy
from crawl4ai.deep_crawling import DomainFilter, ContentTypeFilter, FilterChain, URLPatternFilter, KeywordRelevanceScorer, BestFirstCrawlingStrategy
import asyncio
# Create a filter chain to filter urls based on patterns, domains and content type
filter_chain = FilterChain(
[
DomainFilter(
allowed_domains=["docs.crawl4ai.com"],
blocked_domains=["old.docs.crawl4ai.com"],
),
URLPatternFilter(patterns=["*core*", "*advanced*"],),
ContentTypeFilter(allowed_types=["text/html"]),
]
)
# Create a keyword scorer that prioritises the pages with certain keywords first
keyword_scorer = KeywordRelevanceScorer(
keywords=["crawl", "example", "async", "configuration"], weight=0.7
)
# Set up the configuration
deep_crawl_config = CrawlerRunConfig(
deep_crawl_strategy=BestFirstCrawlingStrategy(
max_depth=2,
include_external=False,
filter_chain=filter_chain,
url_scorer=keyword_scorer,
),
scraping_strategy=LXMLWebScrapingStrategy(),
stream=True,
verbose=True,
)
async def main():
async with AsyncWebCrawler() as crawler:
start_time = time.perf_counter()
results = []
async for result in await crawler.arun(url="https://docs.crawl4ai.com", config=deep_crawl_config):
print(f"Crawled: {result.url} (Depth: {result.metadata['depth']}), score: {result.metadata['score']:.2f}")
results.append(result)
duration = time.perf_counter() - start_time
print(f"\n✅ Crawled {len(results)} high-value pages in {duration:.2f} seconds")
asyncio.run(main())
Breaking Change: The max_depth parameter is now part of CrawlerRunConfig
and controls the depth of the crawl, not the number of concurrent crawls. The
arun() and arun_many() methods are now decorated to handle deep crawling
strategies. Imports for deep crawling strategies have changed. See the
Deep Crawling documentation for more details.
The new MemoryAdaptiveDispatcher dynamically adjusts concurrency based on
available system memory and includes built-in rate limiting. This prevents
out-of-memory errors and avoids overwhelming target websites.
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, MemoryAdaptiveDispatcher
import asyncio
# Configure the dispatcher (optional, defaults are used if not provided)
dispatcher = MemoryAdaptiveDispatcher(
memory_threshold_percent=80.0, # Pause if memory usage exceeds 80%
check_interval=0.5, # Check memory every 0.5 seconds
)
async def batch_mode():
async with AsyncWebCrawler() as crawler:
results = await crawler.arun_many(
urls=["https://docs.crawl4ai.com", "https://github.com/unclecode/crawl4ai"],
config=CrawlerRunConfig(stream=False), # Batch mode
dispatcher=dispatcher,
)
for result in results:
print(f"Crawled: {result.url} with status code: {result.status_code}")
async def stream_mode():
async with AsyncWebCrawler() as crawler:
# OR, for streaming:
async for result in await crawler.arun_many(
urls=["https://docs.crawl4ai.com", "https://github.com/unclecode/crawl4ai"],
config=CrawlerRunConfig(stream=True),
dispatcher=dispatcher,
):
print(f"Crawled: {result.url} with status code: {result.status_code}")
print("Dispatcher in batch mode:")
asyncio.run(batch_mode())
print("-" * 50)
print("Dispatcher in stream mode:")
asyncio.run(stream_mode())
Breaking Change: AsyncWebCrawler.arun_many() now uses
MemoryAdaptiveDispatcher by default. Existing code that relied on unbounded
concurrency may require adjustments.
Crawl4AI now offers two crawling strategies:
AsyncPlaywrightCrawlerStrategy (Default): Uses Playwright for
browser-based crawling, supporting JavaScript rendering and complex
interactions.AsyncHTTPCrawlerStrategy: A lightweight, fast, and memory-efficient
HTTP-only crawler. Ideal for simple scraping tasks where browser rendering is
unnecessary.from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, HTTPCrawlerConfig
from crawl4ai.async_crawler_strategy import AsyncHTTPCrawlerStrategy
import asyncio
# Use the HTTP crawler strategy
http_crawler_config = HTTPCrawlerConfig(
method="GET",
headers={"User-Agent": "MyCustomBot/1.0"},
follow_redirects=True,
verify_ssl=True
)
async def main():
async with AsyncWebCrawler(crawler_strategy=AsyncHTTPCrawlerStrategy(browser_config =http_crawler_config)) as crawler:
result = await crawler.arun("https://example.com")
print(f"Status code: {result.status_code}")
print(f"Content length: {len(result.html)}")
asyncio.run(main())
Crawl4AI can now be easily deployed as a Docker container, providing a consistent and isolated environment. The Docker image includes a FastAPI server with both streaming and non-streaming endpoints.
# Build the image (from the project root)
docker build -t crawl4ai .
# Run the container
docker run -d -p 8000:8000 --name crawl4ai crawl4ai
API Endpoints:
/crawl (POST): Non-streaming crawl./crawl/stream (POST): Streaming crawl (NDJSON)./health (GET): Health check./schema (GET): Returns configuration schemas./md/{url} (GET): Returns markdown content of the URL./llm/{url} (GET): Returns LLM extracted content./token (POST): Get JWT tokenBreaking Changes:
.llm.env file for API keys.config.yml structure.supervisord instead of direct process management.See the Docker deployment documentation for detailed instructions.
A new CLI (crwl) provides convenient access to Crawl4AI's functionality from
the terminal.
# Basic crawl
crwl https://example.com
# Get markdown output
crwl https://example.com -o markdown
# Use a configuration file
crwl https://example.com -B browser.yml -C crawler.yml
# Use LLM-based extraction
crwl https://example.com -e extract.yml -s schema.json
# Ask a question about the crawled content
crwl https://example.com -q "What is the main topic?"
# See usage examples
crwl --example
See the CLI documentation for more details.
Added LXMLWebScrapingStrategy for faster HTML parsing using the lxml
library. This can significantly improve scraping performance, especially for
large or complex pages. Set scraping_strategy=LXMLWebScrapingStrategy() in
your CrawlerRunConfig.
Breaking Change: The ScrapingMode enum has been replaced with a strategy
pattern. Use WebScrapingStrategy (default) or LXMLWebScrapingStrategy.
Added ProxyRotationStrategy abstract base class with RoundRobinProxyStrategy
concrete implementation.
import re
from crawl4ai import (
AsyncWebCrawler,
BrowserConfig,
CrawlerRunConfig,
CacheMode,
RoundRobinProxyStrategy,
)
import asyncio
from crawl4ai import ProxyConfig
async def main():
# Load proxies and create rotation strategy
proxies = ProxyConfig.from_env()
#eg: export PROXIES="ip1:port1:username1:password1,ip2:port2:username2:password2"
if not proxies:
print("No proxies found in environment. Set PROXIES env variable!")
return
proxy_strategy = RoundRobinProxyStrategy(proxies)
# Create configs
browser_config = BrowserConfig(headless=True, verbose=False)
run_config = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS,
proxy_rotation_strategy=proxy_strategy
)
async with AsyncWebCrawler(config=browser_config) as crawler:
urls = ["https://httpbin.org/ip"] * (len(proxies) * 2) # Test each proxy twice
print("\n📈 Initializing crawler with proxy rotation...")
async with AsyncWebCrawler(config=browser_config) as crawler:
print("\n🚀 Starting batch crawl with proxy rotation...")
results = await crawler.arun_many(
urls=urls,
config=run_config
)
for result in results:
if result.success:
ip_match = re.search(r'(?:[0-9]{1,3}\.){3}[0-9]{1,3}', result.html)
current_proxy = run_config.proxy_config if run_config.proxy_config else None
if current_proxy and ip_match:
print(f"URL {result.url}")
print(f"Proxy {current_proxy.server} -> Response IP: {ip_match.group(0)}")
verified = ip_match.group(0) == current_proxy.ip
if verified:
print(f"✅ Proxy working! IP matches: {current_proxy.ip}")
else:
print("❌ Proxy failed or IP mismatch!")
print("---")
asyncio.run(main())
LLMContentFilter for intelligent markdown generation. This new
filter uses an LLM to create more focused and relevant markdown output.from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, DefaultMarkdownGenerator
from crawl4ai.content_filter_strategy import LLMContentFilter
from crawl4ai import LLMConfig
import asyncio
llm_config = LLMConfig(provider="gemini/gemini-1.5-pro", api_token="env:GEMINI_API_KEY")
markdown_generator = DefaultMarkdownGenerator(
content_filter=LLMContentFilter(llm_config=llm_config, instruction="Extract key concepts and summaries")
)
config = CrawlerRunConfig(markdown_generator=markdown_generator)
async def main():
async with AsyncWebCrawler() as crawler:
result = await crawler.arun("https://docs.crawl4ai.com", config=config)
print(result.markdown.fit_markdown)
asyncio.run(main())
Added: URL redirection tracking. The crawler now automatically follows
HTTP redirects (301, 302, 307, 308) and records the final URL in the
redirected_url field of the CrawlResult object. No code changes are
required to enable this; it's automatic.
Added: LLM-powered schema generation utility. A new generate_schema
method has been added to JsonCssExtractionStrategy and
JsonXPathExtractionStrategy. This greatly simplifies creating extraction
schemas.
from crawl4ai import JsonCssExtractionStrategy
from crawl4ai import LLMConfig
llm_config = LLMConfig(provider="gemini/gemini-1.5-pro", api_token="env:GEMINI_API_KEY")
schema = JsonCssExtractionStrategy.generate_schema(
html="<div class='product'><h2>Product Name</h2><span class='price'>$99</span></div>",
llm_config = llm_config,
query="Extract product name and price"
)
print(schema)
Expected Output (may vary slightly due to LLM)
{
"name": "ProductExtractor",
"baseSelector": "div.product",
"fields": [
{"name": "name", "selector": "h2", "type": "text"},
{"name": "price", "selector": ".price", "type": "text"}
]
}
Added: robots.txt compliance support. The crawler can now respect
robots.txt rules. Enable this by setting check_robots_txt=True in
CrawlerRunConfig.
config = CrawlerRunConfig(check_robots_txt=True)
Added: PDF processing capabilities. Crawl4AI can now extract text, images,
and metadata from PDF files (both local and remote). This uses a new
PDFCrawlerStrategy and PDFContentScrapingStrategy.
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig
from crawl4ai.processors.pdf import PDFCrawlerStrategy, PDFContentScrapingStrategy
import asyncio
async def main():
async with AsyncWebCrawler(crawler_strategy=PDFCrawlerStrategy()) as crawler:
result = await crawler.arun(
"https://arxiv.org/pdf/2310.06825.pdf",
config=CrawlerRunConfig(
scraping_strategy=PDFContentScrapingStrategy()
)
)
print(result.markdown) # Access extracted text
print(result.metadata) # Access PDF metadata (title, author, etc.)
asyncio.run(main())
Added: Support for frozenset serialization. Improves configuration serialization, especially for sets of allowed/blocked domains. No code changes required.
Added: New LLMConfig parameter. This new parameter can be passed for
extraction, filtering, and schema generation tasks. It simplifies passing
provider strings, API tokens, and base URLs across all sections where LLM
configuration is necessary. It also enables reuse and allows for quick
experimentation between different LLM configurations.
from crawl4ai import LLMConfig
from crawl4ai import LLMExtractionStrategy
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig
# Example of using LLMConfig with LLMExtractionStrategy
llm_config = LLMConfig(provider="openai/gpt-4o", api_token="YOUR_API_KEY")
strategy = LLMExtractionStrategy(llm_config=llm_config, schema=...)
# Example usage within a crawler
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(
url="https://example.com",
config=CrawlerRunConfig(extraction_strategy=strategy)
)
Breaking Change: Removed old parameters like provider, api_token,
base_url, and api_base from LLMExtractionStrategy and
LLMContentFilter. Users should migrate to using the LLMConfig object.
Changed: Improved browser context management and added shared data support.
(Breaking Change: BrowserContext API updated). Browser contexts are now
managed more efficiently, reducing resource usage. A new shared_data
dictionary is available in the BrowserContext to allow passing data between
different stages of the crawling process. Breaking Change: The
BrowserContext API has changed, and the old get_context method is
deprecated.
Changed: Renamed final_url to redirected_url in CrawledURL. This
improves consistency and clarity. Update any code referencing the old field
name.
Changed: Improved type hints and removed unused files. This is an internal improvement and should not require code changes.
Changed: Reorganized deep crawling functionality into dedicated module.
(Breaking Change: Import paths for DeepCrawlStrategy and related classes
have changed). This improves code organization. Update imports to use the new
crawl4ai.deep_crawling module.
Changed: Improved HTML handling and cleanup codebase. (Breaking
Change: Removed ssl_certificate.json file). This removes an unused file.
If you were relying on this file for custom certificate validation, you'll
need to implement an alternative approach.
Changed: Enhanced serialization and config handling. (Breaking Change:
FastFilterChain has been replaced with FilterChain). This change
simplifies config and improves the serialization.
Added: Modified the license to Apache 2.0 with a required attribution
clause. See the LICENSE file for details. All users must now clearly
attribute the Crawl4AI project when using, distributing, or creating
derivative works.
Fixed: Prevent memory leaks by ensuring proper closure of Playwright pages. No code changes required.
Fixed: Make model fields optional with default values (Breaking
Change: Code relying on all fields being present may need adjustment).
Fields in data models (like CrawledURL) are now optional, with default
values (usually None). Update code to handle potential None values.
Fixed: Adjust memory threshold and fix dispatcher initialization. This is an internal bug fix; no code changes are required.
Fixed: Ensure proper exit after running doctor command. No code changes are required.
Fixed: JsonCss selector and crawler improvements.
Fixed: Not working long page screenshot (#403)
Documentation: Updated documentation URLs to the new domain.
Documentation: Added SERP API project example.
Documentation: Added clarifying comments for CSS selector behavior.
Documentation: Add Code of Conduct for the project (#410)
MemoryAdaptiveDispatcher is now the default for
arun_many(), changing concurrency behavior. The return type of arun_many
depends on the stream parameter.max_depth is now part of CrawlerRunConfig and controls
crawl depth. Import paths for deep crawling strategies have changed.BrowserContext API has been updated.ScrapingMode enum replaced by strategy pattern
(WebScrapingStrategy, LXMLWebScrapingStrategy).content_filter parameter from
CrawlerRunConfig. Use extraction strategies or markdown generators with
filters instead.WebCrawler, CLI, and docs management functionality.DeepCrawlStrategy,
BreadthFirstSearchStrategy, and related classes due to the new
deep_crawling module structure.CrawlerRunConfig: Move max_depth to CrawlerRunConfig. If using
content_filter, migrate to an extraction strategy or a markdown generator
with a filter.arun_many(): Adapt code to the new MemoryAdaptiveDispatcher behavior
and the return type.BrowserContext: Update code using the BrowserContext API.None values for optional fields in data
models.ScrapingMode enum with WebScrapingStrategy or
LXMLWebScrapingStrategy.crwl command and update any scripts using the
old CLI.Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
SSL certificate validation options in extraction strategies
Browser and SSL Handling
Content Processing
JSON Extraction
Field Types
computed, conditional, aggregate, templatePerformance
Error Handling
evalNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Deleted several deprecated files and examples that are no longer relevant.
This release introduces several powerful new features, including robots.txt compliance, dynamic proxy support, LLM-powered schema generation, and improved documentation.
Robots.txt Compliance:
check_robots_txt parameter in CrawlerRunConfig to enable robots.txt checking before crawling a URL.AsyncWebCrawler with 403 status codes for blocked URLs.Proxy Configuration:
CrawlerRunConfig, allowing dynamic proxy settings per crawl request.LLM-Powered Schema Generation:
URL Redirection Tracking:
redirected_url field of the AsyncCrawlResponse object.Enhanced Streamlined Documentation:
Improved Browser Context Management:
shared_data parameter in CrawlerRunConfig to pass data between hooks.Memory Dispatcher System:
MemoryAdaptiveDispatcher and SemaphoreDispatcher for improved resource management.RateLimiter for rate limiting support.CrawlerMonitor for real-time monitoring of crawler operations.Streaming Support:
stream parameter in CrawlerRunConfig.Content Scraping Strategy:
LXMLWebScrapingStrategy for faster content scraping.scraping_strategy parameter in CrawlerRunConfig.Browser Path Management:
Memory Threshold:
Pydantic Model Fields:
Documentation Structure:
Scraping Mode:
ScrapingMode enum with a strategy pattern for more flexible content scraping.Version Update:
0.4.248.Code Cleanup:
Updated dependencies:
Ignored certain patterns and directories:
.gitignore and .codeiumignore to ignore additional patterns and directories, streamlining the development environment.Simplified Personal Story in README:
README.md for clarity.Removed Deprecated Files:
Previous Releases:
computed, conditional, aggregate, and template field types.configure_windows_event_loop to resolve NotImplementedError for asyncio subprocesses on Windows. (#utils.py, #tutorials/async-webcrawler-basics.md)page_need_scroll Method: Added a method to determine if a page requires scrolling before taking actions in AsyncPlaywrightCrawlerStrategy. (#async_crawler_strategy.py)0.4.246 to 0.4.247. (#version.py)AsyncPlaywrightCrawlerStrategy by adding a scroll_delay parameter for better control. (#async_crawler_strategy.py)hello_world.py example to reflect the latest API changes and better illustrate features. (#examples/hello_world.py)content_scraping_strategy.py. (#content_scraping_strategy.py)finally block to ensure pages are closed when no session_id is provided.async_crawler_strategy.pyfinally:
# If no session_id is given we should close the page
if not config.session_id:
await page.close()
_get_elements in JsonCssExtractionStrategy to return all matching elements instead of just the first one, ensuring comprehensive extraction. (#extraction_strategy.py)We're excited to announce Crawl4AI 0.4.3, focusing on three key areas: Speed & Efficiency, LLM Integration, and Core Platform Improvements. This relea
We're excited to announce Crawl4AI 0.4.3, focusing on three key areas: Speed & Efficiency, LLM Integration, and Core Platform Improvements. This release significantly improves crawling performance while adding powerful new LLM-powered features.
The new dispatcher system provides intelligent resource management and real-time monitoring:
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, DisplayMode
from crawl4ai.async_dispatcher import MemoryAdaptiveDispatcher, CrawlerMonitor
async def main():
urls = ["https://example1.com", "https://example2.com"] * 50
# Configure memory-aware dispatch
dispatcher = MemoryAdaptiveDispatcher(
memory_threshold_percent=80.0, # Auto-throttle at 80% memory
check_interval=0.5, # Check every 0.5 seconds
max_session_permit=20, # Max concurrent sessions
monitor=CrawlerMonitor( # Real-time monitoring
display_mode=DisplayMode.DETAILED
)
)
async with AsyncWebCrawler() as crawler:
results = await dispatcher.run_urls(
urls=urls,
crawler=crawler,
config=CrawlerRunConfig()
)
Process crawled URLs in real-time instead of waiting for all results:
config = CrawlerRunConfig(stream=True)
async with AsyncWebCrawler() as crawler:
async for result in await crawler.arun_many(urls, config=config):
print(f"Got result for {result.url}")
# Process each result immediately
New LXML scraping strategy offering up to 20x faster parsing:
config = CrawlerRunConfig(
scraping_strategy=LXMLWebScrapingStrategy(),
cache_mode=CacheMode.ENABLED
)
Smart content filtering and organization using LLMs:
config = CrawlerRunConfig(
markdown_generator=DefaultMarkdownGenerator(
content_filter=LLMContentFilter(
provider="openai/gpt-4o",
instruction="Extract technical documentation and code examples"
)
)
)
Generate extraction schemas instantly using LLMs instead of manual CSS/XPath writing:
schema = JsonCssExtractionStrategy.generate_schema(
html_content,
schema_type="CSS",
query="Extract product name, price, and description"
)
Integrated proxy support with automatic rotation and verification:
config = CrawlerRunConfig(
proxy_config={
"server": "http://proxy:8080",
"username": "user",
"password": "pass"
}
)
Built-in robots.txt support with SQLite caching:
config = CrawlerRunConfig(check_robots_txt=True)
result = await crawler.arun(url, config=config)
if result.status_code == 403:
print("Access blocked by robots.txt")
Track final URLs after redirects:
result = await crawler.arun(url)
print(f"Initial URL: {url}")
print(f"Final URL: {result.redirected_url}")
pip install -U crawl4ai
For complete examples, check our demo repository.
Happy crawling! 🕷️
`text_mode` (boolean): Enables text-only mode, disables images, JavaScript, and GPU-related features for faster, minimal rendering.
crawl4ai/async_crawler_strategy.pytext_mode (boolean): Enables text-only mode, disables images, JavaScript, and GPU-related features for faster, minimal rendering.light_mode (boolean): Optimizes the browser by disabling unnecessary background processes and features for efficiency.viewport_width and viewport_height: Dynamically adjusts based on text_mode mode (default values: 800x600 for text_mode, 1920x1080 otherwise).extra_args: Adds browser-specific flags for text_mode mode.adjust_viewport_to_content: Dynamically adjusts the viewport to the content size for accurate rendering.viewport adjustments: Dynamically computed based on text_mode or custom configuration.light_mode and text_mode by adding specific browser arguments to reduce resource consumption.create_session method:
adjust_viewport_to_content:
adjust_viewport_to_content).scan_full_page).self.viewport_width, self.viewport_height).delay_before_return_html parameter.light_mode by disabling unnecessary browser features such as extensions, background timers, and sync.chrome, firefox, webkit).docs/examples/quickstart_async.pyLLMExtractionStrategy:
OpenAIModelFee.schema()OpenAIModelFee.model_json_schema()OpenAIModelFee class and its JSON schema.self.viewport_width and self.viewport_height throughout the code.This post was generated with the help of ChatGPT, take everything with a grain of salt. 🧂
Hi everyone,
I just finished putting together version 0.4.1 of Crawl4AI, and there are a few changes in here that I think you’ll find really helpful. I’ll explain what’s new, why it matters, and exactly how you can use these features (with the code to back it up). Let’s get into it.
One thing that always bugged me with crawlers is how often they miss lazy-loaded content, especially images. In this version, I made sure Crawl4AI waits for all images to load before moving forward. This is useful because many modern websites only load images when they’re in the viewport or after some JavaScript executes.
Here’s how to enable it:
await crawler.crawl(
url="https://example.com",
wait_for_images=True # Add this argument to ensure images are fully loaded
)
What this does is:
This single change handles the majority of lazy-loading cases you’re likely to encounter.
Sometimes, you don’t need to download images or process JavaScript at all. For example, if you’re crawling to extract text data, you can enable text-only mode to speed things up. By disabling images, JavaScript, and other heavy resources, this mode makes crawling 3-4 times faster in most cases.
Here’s how to turn it on:
crawler = AsyncPlaywrightCrawlerStrategy(
text_mode=True # Set this to True to enable text-only crawling
)
When text_mode=True, the crawler automatically:
viewport_width and viewport_height).If you need to crawl thousands of pages where you only care about text, this mode will save you a ton of time and resources.
Another useful addition is the ability to dynamically adjust the viewport size to match the content on the page. This is particularly helpful when you’re working with responsive layouts or want to ensure all parts of the page load properly.
Here’s how it works:
To enable this, use:
await crawler.crawl(
url="https://example.com",
adjust_viewport_to_content=True # Dynamically adjusts the viewport
)
This approach makes sure the entire page gets loaded into the viewport, especially for layouts that load content based on visibility.
Some websites load data dynamically as you scroll down the page. To handle these cases, I added support for full-page scanning. It simulates scrolling to the bottom of the page, checking for new content, and capturing it all.
Here’s an example:
await crawler.crawl(
url="https://example.com",
scan_full_page=True, # Enables scrolling
scroll_delay=0.2 # Waits 200ms between scrolls (optional)
)
What happens here:
If you’ve ever had to deal with infinite scroll pages, this is going to save you a lot of headaches.
By default, every time you crawl a page, a new browser context (or tab) is created. That’s fine for small crawls, but if you’re working on a large dataset, it’s more efficient to reuse the same session.
I added a method called create_session for this:
session_id = await crawler.create_session()
# Use the same session for multiple crawls
await crawler.crawl(
url="https://example.com/page1",
session_id=session_id # Reuse the session
)
await crawler.crawl(
url="https://example.com/page2",
session_id=session_id
)
This avoids creating a new tab for every page, speeding up the crawl and reducing memory usage.
Here are a few smaller updates I’ve made:
light_mode=True to disable background processes, extensions, and other unnecessary features, making the browser more efficient.delay_before_return_html (now set to 0.1 seconds).You can install or upgrade to version 0.4.1 like this:
pip install crawl4ai --upgrade
As always, I’d love to hear your thoughts. If there’s something you think could be improved or if you have suggestions for future versions, let me know!
Enjoy the new features, and happy crawling! 🕷️
The 0.4.0 release introduces significant improvements to content filtering, multi-threaded environment handling, user-agent generation, and test cover
The 0.4.0 release introduces significant improvements to content filtering, multi-threaded environment handling, user-agent generation, and test coverage. Key highlights include the introduction of the PruningContentFilter, designed to automatically identify and extract the most valuable parts of an HTML document, as well as enhancements to the BM25ContentFilter to extend its versatility and effectiveness.
This release significantly enhances the content extraction capabilities of Crawl4ai with the introduction of the PruningContentFilter, improved supervised filtering with BM25ContentFilter, and robust multi-threaded handling. Additionally, the user-agent generator provides much-needed versatility, resolving compatibility issues faced by many users.
Users are encouraged to experiment with the new content filtering methods to determine which best suits their needs.
Enhanced Docker Support (Nov 29, 2024)
basic-amd64, all-amd64, gpu-amd64 for AMD64.basic-arm64, all-arm64, gpu-arm64 for ARM64.docs/examples/quickstart_async.py to be more useful and user-friendly.requirements.txt with a new pydantic dependency.crawl4ai/__version__.py to 0.3.746.main.py which might affect existing deployments relying on static content.post_install method in crawl4ai/install.py to streamline post-installation setup tasks.crawl4ai/migrations.py with enhanced logging for better error visibility.docker-compose.yml to support local and hub services for different architectures, enhancing build and deploy capabilities.docs/examples/docker_example.py to facilitate comprehensive testing.Updated README with new docker commands and setup instructions. Enhanced installation instructions and guidance.
Added post-install script functionality.
Introduced post_install method for automation of post-installation tasks.
Improved migration logging. Refined migration processes and added better logging.
Refactored docker-compose for better service management. Updated to define services for different platforms and versions.
Updated dependencies.
Added pydantic to requirements file.
Updated version number. Bumped version number to 0.3.746.
Enhanced example scripts. Uncommented example usage in async guide for user functionality.
Refactored code to improve maintainability. Streamlined app structure by removing static pages code.
Nothing published for this version
Nothing published for this version
Enhance features and documentation
Enhance features and documentation
Added new contributors and pull request details. Updated community contributions and acknowledged pull requests.
Version update. Bumped version to 0.3.743.
Improved ManagedBrowser configuration. Enhanced browser initialization with configurable host and debugging port; improved hook execution.
Optimized HTML processing. Implemented 'fast_format_html' for optimized HTML formatting; applied it when 'prettiify' is enabled.
Enhanced markdown generation strategy. Updated to use DefaultMarkdownGenerator and improved markdown generation with filters option.
Refactored markdown generation class. Renamed DefaultMarkdownGenerationStrategy to DefaultMarkdownGenerator; added content filter handling.
Enhanced utility functions. Improved input sanitization and enhanced HTML formatting method.
Improved documentation for hooks. Updated code examples to include cookies in crawler strategy initialization.
Refactored tests to match class renaming. Updated tests to use renamed DefaultMarkdownGenerator class.
Nothing published for this version
Nothing published for this version
Support for raw HTML and local file crawling via URL prefixes ('raw:', 'file://')
fit_markdown flag for optional markdown generation__del__ method from AsyncPlaywrightCrawlerStrategy to prevent async cleanup issuesRemoved deprecated: crawl4ai/content_cleaning_strategy.py.
This changelog details the updates and changes introduced in Crawl4AI version 0.3.74. It's designed to inform developers about new features, modifications to existing components, removals, and other important information.
downloads_path parameter in the AsyncWebCrawler constructor or the arun method. If not specified, downloads are saved to a "downloads" folder within the .crawl4ai directory.CrawlResult object. Successfully downloaded files are listed in the downloaded_files attribute, providing their paths.accept_downloads parameter to the crawler strategies (defaults to False). If set to True you can add JS code and wait_for parameter for file download.Example:
import asyncio
import os
from pathlib import Path
from crawl4ai import AsyncWebCrawler
async def download_example():
downloads_path = os.path.join(Path.home(), ".crawl4ai", "downloads")
os.makedirs(downloads_path, exist_ok=True)
async with AsyncWebCrawler(
accept_downloads=True,
downloads_path=downloads_path,
verbose=True
) as crawler:
result = await crawler.arun(
url="https://www.python.org/downloads/",
js_code="""
const downloadLink = document.querySelector('a[href$=".exe"]');
if (downloadLink) { downloadLink.click(); }
""",
wait_for=5 # To ensure download has started
)
if result.downloaded_files:
print("Downloaded files:")
for file in result.downloaded_files:
print(f"- {file}")
asyncio.run(download_example())
RelevanceContentFilter strategy (and its implementation BM25ContentFilter) for extracting relevant content from web pages, replacing Fit Markdown and other content cleaning strategy. This new strategy leverages the BM25 algorithm to identify chunks of text relevant to the page's title, description, keywords, or a user-provided query.fit_markdown flag in the content scraper is used to filter content based on title, meta description, and keywords.Example:
from crawl4ai import AsyncWebCrawler
from crawl4ai.content_filter_strategy import BM25ContentFilter
async def filter_content(url, query):
async with AsyncWebCrawler() as crawler:
content_filter = BM25ContentFilter(user_query=query)
result = await crawler.arun(url=url, extraction_strategy=content_filter, fit_markdown=True)
print(result.extracted_content) # Or result.fit_markdown for the markdown version
print(result.fit_html) # Or result.fit_html to show HTML with only the filtered content
asyncio.run(filter_content("https://en.wikipedia.org/wiki/Apple", "fruit nutrition health"))
file:// prefix for local file paths.raw: prefix for raw HTML strings.Example:
async def crawl_local_or_raw(crawler, content, content_type):
prefix = "file://" if content_type == "local" else "raw:"
url = f"{prefix}{content}"
result = await crawler.arun(url=url)
if result.success:
print(f"Markdown Content from {content_type.title()} Source:")
print(result.markdown)
# Example usage with local file and raw HTML
async def main():
async with AsyncWebCrawler() as crawler:
# Local File
await crawl_local_or_raw(
crawler, os.path.abspath('tests/async/sample_wikipedia.html'), "local"
)
# Raw HTML
await crawl_raw_html(crawler, "<h1>Raw Test</h1><p>This is raw HTML.</p>")
asyncio.run(main())
ManagedBrowser class introduced for improved browser session handling, offering features like persistent browser sessions between requests (using session_id parameter) and browser process monitoring.use_managed_browser, use_persistent_context, and chrome_channel parameters to AsyncPlaywrightCrawlerStrategy.Example:
async def browser_management_demo():
user_data_dir = os.path.join(Path.home(), ".crawl4ai", "user-data-dir")
os.makedirs(user_data_dir, exist_ok=True) # Ensure directory exists
async with AsyncWebCrawler(
use_managed_browser=True,
user_data_dir=user_data_dir,
use_persistent_context=True,
verbose=True
) as crawler:
result1 = await crawler.arun(
url="https://example.com", session_id="my_session"
)
result2 = await crawler.arun(
url="https://example.com/anotherpage", session_id="my_session"
)
asyncio.run(browser_management_demo())
CacheMode enum (ENABLED, DISABLED, READ_ONLY, WRITE_ONLY, BYPASS) and always_bypass_cache parameter in AsyncWebCrawler for fine-grained cache control. This replaces bypass_cache, no_cache_read, no_cache_write, and always_by_pass_cache.crawl4ai/content_cleaning_strategy.py.bypass_cache, disable_cache, no_cache_read, no_cache_write, and always_by_pass_cache. These have been superseded by cache_mode.crawl4ai/__version__.py.crawl4ai/cache_context.py.crawl4ai/version_manager.py.crawl4ai/migrations.py.crawl4ai-migrate entry point.NEED_MIGRATION and SHOW_DEPRECATION_WARNINGS.CRAWL4AI_API_TOKEN environment variable. This enhances API security./crawl_sync for immediate result retrieval, and direct crawl endpoint /crawl_direct bypassing the task queue.WebCrawler is being phased out. While still available via crawl4ai[sync], it will eventually be removed. Transition to AsyncWebCrawler is strongly recommended. Boolean cache control flags in arun are also deprecated, migrate to using the cache_mode parameter. See examples in the "New Features" section above for correct usage.__del__ method and ensuring the browser context is closed explicitly using context managers.WebScrapingStrategy. More detailed error messages and suggestions for debugging will minimize frustration when running into unexpected issues.Old way:
crawler = AsyncWebCrawler(always_by_pass_cache=True)
result = await crawler.arun(url="https://example.com", bypass_cache=True)
New way:
from crawl4ai import CacheMode
crawler = AsyncWebCrawler(always_bypass_cache=True)
result = await crawler.arun(url="https://example.com", cache_mode=CacheMode.BYPASS)
Your coding agent can read these notes before it upgrades. Set up the MCP server →