NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Packagist · #718 most downloaded on Packagist
Crawl all internal links found on a website
Last release 2 months ago
07 Aug 2026
Release timing varies
gaps range from 8 days to 5 months
Most releases are documented
notes for 52 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
11 years old
129 releases · first in 2015
Fix SitemapUrlParser extracting no URLs from namespaced sitemaps by @witheez in #512
Full Changelog: https://github.com/spatie/crawler/compare/9.4.1...9.4.2
One column per quarter.
Fix Guzzle 8 exception handling for responses and retries by @GrahamCampbell in #510
Full Changelog: 9.4.0...9.4.1
Full Changelog: https://github.com/spatie/crawler/compare/9.4.0...9.4.1
Added support for Guzzle 8 alongside Guzzle 7.
Added support for Guzzle 8 alongside Guzzle 7.
Skip signal handler registration when pcntl functions are disabled by @poldixd in #508
Full Changelog: https://github.com/spatie/crawler/compare/9.3.1...9.3.2
Skip fnmatch for URLs longer than FILENAME_MAX by @mattiasgeniar in #505
Full Changelog: 9.3.0...9.3.1
Full Changelog: https://github.com/spatie/crawler/compare/9.3.0...9.3.1
Add a shouldStopCallback hook for graceful external stops by @kissifrot in #504
Full Changelog: 9.2.1...9.3.0
Full Changelog: https://github.com/spatie/crawler/compare/9.2.1...9.3.0
Close response body streams after processing by @freekmurze in #503
Full Changelog: 9.0.1...9.2.1
Full Changelog: https://github.com/spatie/crawler/compare/9.0.1...9.2.1
Full Changelog : 9.1.0...9.2.0
Full Changelog: 9.1.0...9.2.0
Full Changelog: https://github.com/spatie/crawler/compare/9.1.0...9.2.0
Add urlParser() method to allow setting a custom URL parser
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: freekmurze freekmurze@users.noreply.github.com
When the crawler reached its time limit, crawl limit, or was interrupted via a signal, any HTTP requests that were still in-flight in the Guzzle pool
When the crawler reached its time limit, crawl limit, or was interrupted via a signal, any HTTP requests that were still in-flight in the Guzzle pool would eventually fail (e.g. due to their per-request timeout) and get reported to crawlFailed. These URLs were dispatched but never truly crawled, so reporting them as failures was misleading.
The CrawlRequestFailed handler now checks whether the crawler has reached its limits before reporting a failure. If it has, the failure is silently ignored.
Regular request failures (timeouts, connection errors, server errors) that happen during normal crawling are still reported as before.
allow_redirects default: changed from false to ['track_redirects' => true] so redirects are followed and the redirect history header is populated correctlyallowedMimeTypes) now notify observers via crawled() with an empty body instead of being silently skipped/../, /./) in extracted URLs are now normalized per RFC 3986Crawler::create() now merge with defaults instead of replacing them (pass null to remove a default)CrawlRequestFailed now wraps non-RequestException errors so observers always receive a RequestExceptionUrl, ResponseWithCachedBody, InvalidUrlstream() method to opt-in to streaming HTTP responses for reduced memory usagematchWww() method to treat www.example.com and example.com as equivalent when using internalOnly()includeSubdomains() now works as a flag on internalOnly() and composes with matchWww()CrawlResponse::redirectHistory() and CrawlResponse::wasRedirected() for inspecting redirect chainsCrawlObserver::crawlFailed() now receives a ?TransferStatistics parameter for detecting timeoutsaddObserver() now accepts variadic arguments: addObserver($obs1, $obs2)CrawlRequestFailed now preserves the original request from ConnectException (retaining custom headers like X-Started-At)Major rewrite. See UPGRADING.md for a full list of breaking changes.
Major rewrite. See UPGRADING.md for a full list of breaking changes.
UriInterface with plain string URLs throughout the APIResponseInterface with CrawlResponse in observer callbacksCrawlProfile is now an interface instead of an abstract classCrawlObserverCollection no longer implements ArrayAccess or Iteratorhttp to httpssuggest)UrlParser interface redesigned to return ExtractedUrl[] instead of adding to queue directlyCrawlQueue::has() now accepts string instead of CrawlUrl|UriInterfacestart() now returns a FinishReason enumCrawler::create()CrawlResponse object with status(), body(), dom(), header(), transferStats(), and moreCrawlProgress tracking with urlsCrawled, urlsFailed, urlsFound, urlsPendingFinishReason enum: Completed, CrawlLimitReached, TimeLimitReached, InterruptedonCrawled(), onFailed(), onFinished(), onWillCrawl()foundUrls() to collect all URLs as CrawledUrl objectsfake() for testing without HTTP requestsinternalOnly(), includeSubdomains(), shouldCrawl()depth(), concurrency(), delay(), limit(), userAgent()FixedDelayThrottle and AdaptiveThrottlealsoExtract(), extractAll(), ResourceType enumArrayCrawlQueuealwaysCrawl() and neverCrawl() pattern overridesretry() for automatic retries on connection errors and 5xx responsesTransferStatistics with typed timing accessorsCloudflareRenderer for JavaScript renderingJavaScriptRenderer interface for custom renderersbasicAuth(), token(), withoutVerifying(), proxy(), cookies(), queryParameters(), middleware()CrawlUrl::create() static factory (use new CrawlUrl(...) instead)Spatie\Crawler\Url classResponseWithCachedBody (replaced by CrawlResponse)nicmart/tree dependencyspatie/browsershot as a required dependency (moved to suggest)setBrowsershot() and getBrowsershot() methodsstartCrawling() method (use start())setUrlParserClass() (use parseSitemaps() or pass a UrlParser directly)Add Laravel 13 support
Add Laravel 13 support
Update nicmart/tree dependency version to ^0.10 by @robinmiau in https://github.com/spatie/crawler/pull/497
Full Changelog: https://github.com/spatie/crawler/compare/8.4.6...8.4.7
Nothing published for this version
When fetching robots.txt, use the same User-Agent as defined by the user by @mattiasgeniar in https://github.com/spatie/crawler/pull/491
Full Changelog: https://github.com/spatie/crawler/compare/8.4.4...8.4.5
Update issue template by @AlexVanderbist in https://github.com/spatie/crawler/pull/488
Full Changelog: https://github.com/spatie/crawler/compare/8.4.3...8.4.4
Do not try robots.txt when ignored by @kissifrot in https://github.com/spatie/crawler/pull/485
Full Changelog: https://github.com/spatie/crawler/compare/8.4.2...8.4.3
set spatie/browsershot minimal version to 5.0.5 by @grafst in https://github.com/spatie/crawler/pull/484
Full Changelog: https://github.com/spatie/crawler/compare/8.4.1...8.4.2
Full Changelog: https://github.com/spatie/crawler/compare/8.4.0...8.4.1
Full Changelog: https://github.com/spatie/crawler/compare/8.4.0...8.4.1
Add execution time limit by @VincentLanglet in https://github.com/spatie/crawler/pull/480
Full Changelog: https://github.com/spatie/crawler/compare/8.3.1...8.4.0
Upgrade spatie/browsershot to 5.0 by @hasansoyalan in https://github.com/spatie/crawler/pull/478
Full Changelog: https://github.com/spatie/crawler/compare/8.3.0...8.3.1
Add support for PHP 8.4 by @pascalbaljet in https://github.com/spatie/crawler/pull/477
Full Changelog: https://github.com/spatie/crawler/compare/8.2.3...8.3.0
Fix setParsableMimeTypes() by @superpenguin612 in https://github.com/spatie/crawler/pull/470
Full Changelog: https://github.com/spatie/crawler/compare/8.2.2...8.2.3
Nothing published for this version
Check original URL against depth tree when visited link is a redirect by @superpenguin612 in https://github.com/spatie/crawler/pull/467
Full Changelog: https://github.com/spatie/crawler/compare/8.2.0...8.2.1
Fix wording in documentation by @adamtomat in https://github.com/spatie/crawler/pull/460
Full Changelog: https://github.com/spatie/crawler/compare/8.1.0...8.2.0
feat: custom link parser by @Velka-DEV in https://github.com/spatie/crawler/pull/458
Full Changelog: https://github.com/spatie/crawler/compare/8.0.4...8.1.0
- allow Browsershot v4
Fix return type by @riesjart in https://github.com/spatie/crawler/pull/452
Full Changelog: https://github.com/spatie/crawler/compare/8.0.2...8.0.3
Define only needed methods in observer implementation by @buismaarten in https://github.com/spatie/crawler/pull/449
Full Changelog: https://github.com/spatie/crawler/compare/8.0.1...8.0.2
Check if rel attribute contains nofollow by @robbinbenard in https://github.com/spatie/crawler/pull/445
Full Changelog: https://github.com/spatie/crawler/compare/8.0.0...8.0.1
add linkText to crawl observer methods
- support Laravel 10
Feat/convert phpunit tests to pest by @mansoorkhan96 in https://github.com/spatie/crawler/pull/401
Full Changelog: https://github.com/spatie/crawler/compare/7.1.1...7.1.2
Nothing published for this version
- allow Laravel 9 collections
Nothing published for this version
Nothing published for this version
Nothing published for this version
- allow psr7 v2
- change response type hint
- require PHP 8+ - drop support for PHP 7.x - convert syntax to PHP 8 - no API changes have been made
Nothing published for this version
bugfix: infinite loops when a CrawlProfile prevents crawling
add setCurrentCrawlLimit and setTotalCrawlLimit
setCurrentCrawlLimit and setTotalCrawlLimit- add support for PHP 8.0
tweak variable naming in ArrayCrawlQueue
ArrayCrawlQueue (#326)remove all deprecated functions and classes
Nothing published for this version
treat connection exceptions as request exceptions
fix: method and property name error
add crawler option to allow crawl links with rel="nofollow"
only crawl links that are completely parsed
- fix curl streaming responses
- add setParseableMimeTypes()
setParseableMimeTypes() (#293)fix LinkAdder not receiving the updated DOM
- allow tightenco/collect 7
respect maximum response size when checking Robots Meta tags
- allow Guzzle 7
- allow symfony 5 components
Your coding agent can read these notes before it upgrades. Set up the MCP server →