NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1905 most downloaded on PyPI
A high-level Web Crawling and Web Scraping framework
Last release 24 days ago
10 Sep 2026
Release timing varies
gaps range from 2 weeks to 5 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
17 years old
115 releases · first in 2009
New RemoteControl extension which allows inspecting and controlling a running crawl over HTTP, used by the Scrapy MCP server
RemoteControl extension which allows inspecting and controlling a running crawl over HTTP, used by the Scrapy MCP serveraiohttp-based download handler (now the default when running without a reactor)Highlights:
New RemoteControl extension which allows inspecting and controlling a running crawl over HTTP, used by the Scrapy MCP server
Experimental aiohttp -based download handler (now the default when running without a reactor)
Added support for Python 3.15. ( #7511 )
New dependencies:
aiohttp >= 3.13.3
charset-normalizer >= 3.4.0
platformdirs >= 2.0.0
( #7866 , #8054 )
When running without a Twisted reactor , i.e. with TWISTED_REACTOR_ENABLED set to False , the default download handler for http and https is now AiohttpDownloadHandler instead of HttpxDownloadHandler . You can configure HttpxDownloadHandler in the DOWNLOAD_HANDLERS setting if you want. ( #8118 )
scrapy.http.cookies.WrappedRequest.is_unverifiable() now always returns False , and the undocumented is_unverifiable request meta key that it used to read is ignored. ( #7933 )
The RANDOMIZE_DOWNLOAD_DELAY setting is deprecated. Use the new DOWNLOAD_DELAY_JITTER setting instead. Similarly, the randomize_delay key of the DOWNLOAD_SLOTS setting is deprecated in favor of a new jitter key, the randomize_delay attribute of scrapy.core.downloader.Slot is deprecated in favor of its new jitter attribute, and the randomize_delay attribute of scrapy.core.downloader.Downloader is deprecated in favor of the DOWNLOAD_DELAY_JITTER setting. ( #7881 )
The install_root_handler parameter of configure_logging() , CrawlerProcess and AsyncCrawlerProcess is deprecated. Use the new LOG_INSTALL_ROOT_HANDLER setting instead. ( #4793 , #8041 )
The scrapy.dupefilters.RFPDupeFilter.fingerprints attribute is deprecated. Overriding scrapy.dupefilters.RFPDupeFilter.request_fingerprint() is deprecated as well; set the REQUEST_FINGERPRINTER_CLASS setting instead. ( #5517 , #7943 )
Overriding the add_pre_hook() or add_post_hook() methods of Contract is deprecated. Define pre_process() or post_process() instead. ( #6681 , #7886 )
Added a RemoteControl extension, enabled by default, which allows connecting to crawl processes via HTTP and running code inside them. ( #7866 )
Added an experimental HTTP download handler based on aiohttp , AiohttpDownloadHandler . ( #8118 )
Added SQLite-backed scheduler queues: PickleFifoSQLiteQueue , PickleLifoSQLiteQueue , MarshalFifoSQLiteQueue and MarshalLifoSQLiteQueue . They write each request within its own transaction, so that an unclean shutdown cannot corrupt the on-disk queue, at the cost of slower scheduling. ( #845 , #7877 )
Added a DOWNLOAD_DELAY_JITTER setting, and a jitter key for DOWNLOAD_SLOTS , which set the magnitude of the random variation applied to DOWNLOAD_DELAY . They replace the RANDOMIZE_DOWNLOAD_DELAY setting and the randomize_delay key, which could only toggle a fixed ±50%. ( #7881 )
Added a LOG_COLOR setting, which colorizes log output by log level when logging to a terminal. It needs the new color extra. ( #8091 )
The encoding of a response that doesn’t declare one is now detected with charset-normalizer when the response body is neither ASCII nor UTF-8, instead of always falling back to cp1252 . ( #3135 , #8054 )
Added a LOG_INSTALL_ROOT_HANDLER setting, which replaces the deprecated install_root_handler parameter and, unlike it, can also be set from a settings.py module or from the command line. ( #4793 , #8041 )
Contracts now support callbacks defined with async def , including asynchronous generators. ( #6681 , #7886 )
Added the MethodContract ( @method ), BodyContract ( @body ), HeaderContract ( @header ) and CookieContract ( @cookie ) contracts, which set the corresponding attributes of the sample request, and an -a option for the check command, to set spider arguments as in the crawl command. ( #1918 , #8053 )
Added Response.to_dict() , Response.from_dict() and response_from_dict() , and used them in the built-in HTTP cache storages , which now restore cached responses of any response class, including those of third-party plugins, with all their attributes. ( #1450 , #7908 )
Cached responses now indicate when they were stored, through the new cache_timestamp request meta key. ( #2221 , #8034 )
The format key of FEEDS is now inferred from the file extension of the feed URI when not set, e.g. json for a URI ending in .json . ( #1158 , #8031 )
Added find_projects() , which yields the root directory of every Scrapy project in a directory tree. ( #8024 )
Added a public scheduler attribute to ExecutionEngine . ( #8099 )
StatsCollector objects now implement str() , which returns the pretty-printed stats. ( #2746 )
open_in_browser() now also supports responses that are neither HTML nor plain text, picking a file extension based on the Content-Type header. ( #3902 , #8067 )
The -t / --template option of the genspider command now also takes a path to a .tmpl file, so that a custom template can be used without setting TEMPLATES_DIR . ( #8071 )
The --pdb command-line option now uses ipdb instead of pdb when ipdb is installed. ( #4284 , #8052 )
Log records about item processing, i.e. about items scraped, dropped or raising an exception, now carry the item in their extra dict, under the item key, so that custom logging handlers can read it. ( #8085 )
RFPDupeFilter now keeps request fingerprints as bytes , to reduce memory usage. The requests.seen file that it writes in the job directory is now a binary file. ( #5517 , #7943 )
The --pdb command-line option now starts a post-mortem debugging session on every logged error that comes with a traceback, instead of on every Failure object built anywhere in the process. ( #3552 , #8037 )
The Scrapy shell no longer prints the DEBUG messages of parso , an indirect dependency of IPython, when LOG_LEVEL is DEBUG . ( #8060 )
Other code refactoring and improvements. ( #8123 )
HTTP11DownloadHandler now sets certificate and ip_address attributes even for responses without a body. ( #4466 , #8048 )
HTTP connection pool keys in HTTP11DownloadHandler and H2DownloadHandler now include the request bind address (see DOWNLOAD_BIND_ADDRESS ), so that requests bound to different addresses no longer reuse each other’s pooled connections. ( #3565 , #8081 )
HttpCacheMiddleware now updates the headers of a cached response, and stores it again, when a revalidation request gets a 304 response. ( #3778 , #8068 )
An exception raised in a spider callback that no process_spider_exception() method handles is now offered to each of those methods only once. ( #4729 , #7996 )
Defining allowed_domains as a property no longer breaks the commands that match a URL to a spider, i.e. shell , fetch and parse . A property cannot be evaluated on a spider class, so it is now ignored, with a warning, instead of raising TypeError . ( #3119 , #8094 )
get_project_settings() now adds the project directory to sys.path also when the SCRAPY_SETTINGS_MODULE environment variable is set, so that the settings module that the variable points to can be imported. ( #4780 , #8042 )
scrapy.utils.datatypes.LocalCache no longer evicts the oldest item when updating an existing one. ( #8113 )
Added a page about using Scrapy with coding agents , covering the official agent plugin and the Scrapy MCP server. ( #8120 )
Documented the request metadata keys that were missing from the list of special keys , and documented that keys whose name starts with an underscore are internal. ( #3585 , #5564 , #7933 )
Documented how to write custom Scrapy commands , how to write custom spider templates , how to access the response of a failed media download and how to set request headers for media pipeline requests . ( #2046 , #2504 , #3056 , #6844 , #6904 , #8049 , #8056 , #8063 )
Documented how to subclass RedirectMiddleware to allow or deny redirects based on the target URL, and how to set the priority of the requests that CrawlSpider generates from its rules. ( #3613 , #4009 , #8061 , #8070 )
Documented that get_project_settings() returns project settings only, that spiders running concurrently in the same process get independent crawlers, middlewares and settings, and that Scrapy sets the request attribute of the Failure objects that it passes to errbacks. ( #2378 , #4253 , #6408 , #7901 , #8051 , #8055 )
Documented ScrapyCommand and StatsCollector . ( #6844 , #6904 , #8072 )
Improved the contents of llms-full.txt and Markdown versions of documentation pages. ( #8083 , #8087 , #8088 , #8089 , #8117 , #8119 )
Reached 100% test coverage. ( #8021 )
Improved and fixed type hints. ( #8017 , #8077 , #8102 )
Improved CodSpeed benchmarks. ( #8030 , #8080 )
CI and test improvements and fixes. ( #8027 , #8039 , #8059 , #8064 , #8073 , #8075 , #8076 , #8079 , #8092 , #8103 , #8115 )
One column per quarter.
Highlights:
New RemoteControl extension which allows inspecting and controlling a running crawl over HTTP, used by the Scrapy MCP server
Experimental aiohttp-based download handler (now the default when running without a reactor)
Added support for Python 3.15. (7511)
New dependencies:
aiohttp_ >= 3.13.3
charset-normalizer_ >= 3.4.0
platformdirs_ >= 2.0.0
(7866, 8054)
When running without a Twisted reactor, i.e. with TWISTED_REACTOR_ENABLED set to False, the default download handler for http and https is now ~scrapy.core.downloader.handlers._aiohttp.AiohttpDownloadHandler instead of ~scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler. You can configure ~scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler in the DOWNLOAD_HANDLERS setting if you want. (8118)
scrapy.http.cookies.WrappedRequest.is_unverifiable() now always returns False, and the undocumented is_unverifiable request meta key that it used to read is ignored. (7933)
The RANDOMIZE_DOWNLOAD_DELAY setting is deprecated. Use the new DOWNLOAD_DELAY_JITTER setting instead. Similarly, the randomize_delay key of the DOWNLOAD_SLOTS setting is deprecated in favor of a new jitter key, the randomize_delay attribute of scrapy.core.downloader.Slot is deprecated in favor of its new jitter attribute, and the randomize_delay attribute of scrapy.core.downloader.Downloader is deprecated in favor of the DOWNLOAD_DELAY_JITTER setting. (7881)
The install_root_handler parameter of ~scrapy.utils.log.configure_logging, ~scrapy.crawler.CrawlerProcess and ~scrapy.crawler.AsyncCrawlerProcess is deprecated. Use the new LOG_INSTALL_ROOT_HANDLER setting instead. (4793, 8041)
The scrapy.dupefilters.RFPDupeFilter.fingerprints attribute is deprecated. Overriding scrapy.dupefilters.RFPDupeFilter.request_fingerprint is deprecated as well; set the REQUEST_FINGERPRINTER_CLASS setting instead. (5517, 7943)
Overriding the add_pre_hook() or add_post_hook() methods of ~scrapy.contracts.Contract is deprecated. Define pre_process() or post_process() instead. (6681, 7886)
Added a ~scrapy.extensions.remote_control.RemoteControl extension, enabled by default, which allows connecting to crawl processes via HTTP and running code inside them. (7866)
Added an experimental HTTP download handler based on aiohttp_, ~scrapy.core.downloader.handlers._aiohttp.AiohttpDownloadHandler. (8118)
Added SQLite-backed scheduler queues: ~scrapy.squeues.PickleFifoSQLiteQueue, ~scrapy.squeues.PickleLifoSQLiteQueue, ~scrapy.squeues.MarshalFifoSQLiteQueue and ~scrapy.squeues.MarshalLifoSQLiteQueue. They write each request within its own transaction, so that an unclean shutdown cannot corrupt the on-disk queue, at the cost of slower scheduling. (845, 7877)
Added a DOWNLOAD_DELAY_JITTER setting, and a jitter key for DOWNLOAD_SLOTS, which set the magnitude of the random variation applied to DOWNLOAD_DELAY. They replace the RANDOMIZE_DOWNLOAD_DELAY setting and the randomize_delay key, which could only toggle a fixed ±50%. (7881)
Added a LOG_COLOR setting, which colorizes log output by log level when logging to a terminal. It needs the new color extra. (8091)
The encoding of a response that doesn't declare one is now detected with charset-normalizer_ when the response body is neither ASCII nor UTF-8, instead of always falling back to cp1252. (3135, 8054)
Added a LOG_INSTALL_ROOT_HANDLER setting, which replaces the deprecated install_root_handler parameter and, unlike it, can also be set from a settings.py module or from the command line. (4793, 8041)
Contracts now support callbacks defined with async def, including asynchronous generators. (6681, 7886)
Added the ~scrapy.contracts.default.MethodContract (@method), ~scrapy.contracts.default.BodyContract (@body), ~scrapy.contracts.default.HeaderContract (@header) and ~scrapy.contracts.default.CookieContract (@cookie) contracts, which set the corresponding attributes of the sample request, and an -a option for the check command, to set spider arguments as in the crawl command. (1918, 8053)
Added Response.to_dict(), Response.from_dict() and ~scrapy.utils.response.response_from_dict, and used them in the built-in HTTP cache storages, which now restore cached responses of any response class, including those of third-party plugins, with all their attributes. (1450, 7908)
Cached responses now indicate when they were stored, through the new cache_timestamp request meta key. (2221, 8034)
The format key of FEEDS is now inferred from the file extension of the feed URI when not set, e.g. json for a URI ending in .json. (1158, 8031)
Added ~scrapy.utils.project.find_projects, which yields the root directory of every Scrapy project in a directory tree. (8024)
Added a public ~scrapy.core.engine.ExecutionEngine.scheduler attribute to ~scrapy.core.engine.ExecutionEngine. (8099)
~scrapy.statscollectors.StatsCollector objects now implement __str__(), which returns the pretty-printed stats. (2746)
~scrapy.utils.response.open_in_browser now also supports responses that are neither HTML nor plain text, picking a file extension based on the Content-Type header. (3902, 8067)
The -t/--template option of the genspider command now also takes a path to a .tmpl file, so that a custom template can be used without setting TEMPLATES_DIR. (8071)
The --pdb command-line option now uses ipdb_ instead of pdb when ipdb_ is installed. (4284, 8052)
Log records about item processing, i.e. about items scraped, dropped or raising an exception, now carry the item in their extra dict, under the item key, so that custom logging handlers can read it. (8085)
~scrapy.dupefilters.RFPDupeFilter now keeps request fingerprints as bytes, to reduce memory usage. The requests.seen file that it writes in the job directory is now a binary file. (5517, 7943)
The --pdb command-line option now starts a post-mortem debugging session on every logged error that comes with a traceback, instead of on every ~twisted.python.failure.Failure object built anywhere in the process. (3552, 8037)
The Scrapy shell no longer prints the DEBUG messages of parso_, an indirect dependency of IPython, when LOG_LEVEL is DEBUG. (8060)
Other code refactoring and improvements. (8123)
~scrapy.core.downloader.handlers.http11.HTTP11DownloadHandler now sets ~scrapy.http.Response.certificate and ~scrapy.http.Response.ip_address attributes even for responses without a body. (4466, 8048)
HTTP connection pool keys in ~scrapy.core.downloader.handlers.http11.HTTP11DownloadHandler and ~scrapy.core.downloader.handlers.http2.H2DownloadHandler now include the request bind address (see DOWNLOAD_BIND_ADDRESS), so that requests bound to different addresses no longer reuse each other's pooled connections. (3565, 8081)
~scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware now updates the headers of a cached response, and stores it again, when a revalidation request gets a 304 response. (3778, 8068)
An exception raised in a spider callback that no ~scrapy.spidermiddlewares.SpiderMiddleware.process_spider_exception method handles is now offered to each of those methods only once. (4729, 7996)
Defining ~scrapy.Spider.allowed_domains as a property no longer breaks the commands that match a URL to a spider, i.e. shell, fetch and parse. A property cannot be evaluated on a spider class, so it is now ignored, with a warning, instead of raising TypeError. (3119, 8094)
~scrapy.utils.project.get_project_settings now adds the project directory to sys.path also when the SCRAPY_SETTINGS_MODULE environment variable is set, so that the settings module that the variable points to can be imported. (4780, 8042)
scrapy.utils.datatypes.LocalCache no longer evicts the oldest item when updating an existing one. (8113)
Added a page about using Scrapy with coding agents, covering the official agent plugin and the Scrapy MCP server. (8120)
Documented the request metadata keys that were missing from the list of special keys, and documented that keys whose name starts with an underscore are internal. (3585, 5564, 7933)
Documented how to write custom Scrapy commands, how to write custom spider templates, how to access the response of a failed media download and how to set request headers for media pipeline requests. (2046, 2504, 3056, 6844, 6904, 8049, 8056, 8063)
Documented how to subclass ~scrapy.downloadermiddlewares.redirect.RedirectMiddleware to allow or deny redirects based on the target URL, and how to set the priority of the requests that ~scrapy.spiders.CrawlSpider generates from its rules. (3613, 4009, 8061, 8070)
Documented that ~scrapy.utils.project.get_project_settings returns project settings only, that spiders running concurrently in the same process get independent crawlers, middlewares and settings, and that Scrapy sets the request attribute of the ~twisted.python.failure.Failure objects that it passes to errbacks. (2378, 4253, 6408, 7901, 8051, 8055)
Documented ~scrapy.commands.ScrapyCommand and ~scrapy.statscollectors.StatsCollector. (6844, 6904, 8072)
Improved the contents of llms-full.txt and Markdown versions of documentation pages. (8083, 8087, 8088, 8089, 8117, 8119)
Reached 100% test coverage. (8021)
Improved and fixed type hints. (8017, 8077, 8102)
Improved CodSpeed benchmarks. (8030, 8080)
CI and test improvements and fixes. (8027, 8039, 8059, 8064, 8073, 8075, 8076, 8079, 8092, 8103, 8115)
HttpxDownloadHandler now uses httpx2
HttpxDownloadHandler now uses httpx2brotli and Zstandard support are now always available, and optional extras cover the rest of the optional featuresCrawler attributes, such as stats, now raise RuntimeError instead of being None before the crawl startsMetaCopyDetectionMiddlewareHighlights:
HttpxDownloadHandler now uses httpx2
The Twisted-based HTTP/2 download handler is no longer experimental
brotli and Zstandard support are now always available, and optional extras cover the rest of the optional features
Late Crawler attributes, such as stats , now raise RuntimeError instead of being None before the crawl starts
Item exporters now export fields in declaration order
New MetaCopyDetectionMiddleware
New optimization page and built-in stats reference
brotli ( brotlicffi on PyPy) and Zstandard support (the standard library compression.zstd module on Python 3.14 and higher, the backports.zstd package on earlier versions) are now required, so br and zstd are always included in the Accept-Encoding header of requests, and Brotli- and Zstandard-compressed responses are always decoded. Websites may now serve such responses to crawls that previously did not advertise support for them.
The minimum required versions are brotli 1.2.0, brotlicffi 1.2.0.0 and backports.zstd 1.3.0.
( #4698 , #6978 , #7083 , #7929 , #8009 )
Increased the minimum versions of the following dependencies:
cryptography : 37.0.0 → 41.0.5
pyOpenSSL : 22.0.0 → 24.3.0
queuelib : 1.4.2 → 1.6.1
service_identity : 23.1.0 → 24.2.0
w3lib : 1.17.0 → 2.1.1
( #7841 , #7874 , #7879 , #8001 )
The IPython shell requires IPython 8.15.0 or higher. Install the ipython extra to get a compatible version. ( #5447 , #7596 , #7816 )
The following runtime usage of zope.interface interfaces is removed:
SpiderLoader and DummySpiderLoader are no longer marked as implementing the ISpiderLoader interface.
get_spider_loader() no longer checks that the configured spider loader implements the ISpiderLoader interface.
BlockingFeedStorage , FileFeedStorage and StdoutFeedStorage are no longer marked as implementing the IFeedStorage interface.
H2DownloadHandler no longer checks that the DOWNLOADER_CLIENTCONTEXTFACTORY class implements the IPolicyForHTTPS interface.
( #6585 , #7731 )
The engine , extensions , logformatter , request_fingerprinter and stats attributes of Crawler raise RuntimeError when read before the crawl starts, instead of being None until then.
Code that reads them from the spider_opened signal handler onwards is unaffected, and no longer needs to narrow their type. Code that checked whether they were set, e.g. if crawler.stats: , must be updated, since reading them now raises instead of returning None .
( #6136 , #7882 )
Item exporters now export the fields of an item in declaration order, i.e. the order in which they are defined in the item class , instead of the order in which they were populated, as CsvItemExporter already did. dict items, which have no declared fields, keep using the key order of each item. ( #6662 , #6854 , #7824 )
scrapy.utils.serialize.ScrapyJSONEncoder , used by JSON feed exports , the telnet console and the PeriodicLog extension, now serializes datetime , date and time objects in ISO 8601 format, e.g. 2023-08-03T23:24:57.148903+00:00 instead of 2023-08-03 23:24:57 , keeping microseconds and time zone information.
Its DATE_FORMAT and TIME_FORMAT attributes are removed.
( #2087 , #7918 )
scrapy.utils.trackref.live_refs is now a WeakKeyDictionary instead of a collections.defaultdict , so that classes defined at run time are released once they are no longer used. Reading the entry of a class with no tracked instances now raises KeyError instead of creating and returning an empty mapping. ( #5995 , #7922 )
The MEMDEBUG_NOTIFY setting is removed. It had no effect, but code reading it now gets None instead of its default value, which was an empty list. ( #7737 )
scrapy.utils.log.logformatter_adapter() no longer passes the whole dict returned by a log formatter method as logging arguments when that dict has no args key, or its args are empty, and its msg has no %(name)s placeholders. Such messages are now logged verbatim, so a literal % in them no longer breaks logging.
An args tuple is now expanded into one logging argument per item, so that % -style placeholders work with it as they do with a dict .
( #5570 , #5572 , #7936 )
FEEDS keys and FEED_URI values that are pathlib.Path objects are now used as paths, instead of being converted into file:// URIs. This makes them keep working when they contain URI parameters or characters that URI conversion would percent-encode. ( #5794 , #6425 , #6611 , #7674 )
Selector and TextResponse.selector no longer force the html selector type for responses that are neither HtmlResponse nor XmlResponse objects. A JsonResponse gets the json type, and for any other response parsel determines the type from the body.
The response class, and hence the selector type, comes from the content type that the website reports. When a website reports the wrong content type, recast the response, e.g. response.replace(cls=HtmlResponse) .
( #4627 , #5291 , #6025 , #7924 , #7972 )
AutoThrottle no longer sets the download_delay attribute of the running spider to define the starting delay of download slots. The starting delay is still applied, but code that reads that attribute at run time no longer sees it. ( #7167 , #7175 , #7833 )
The check command now ignores ITEM_PIPELINES and FEEDS , since contracts check the output of callbacks instead of sending it to item processing, so a check run no longer triggers their side effects, e.g. writing an empty output file. Use the -s command-line option to set them back for a check run. ( #3385 , #7957 )
XMLFeedSpider and CSVFeedSpider no longer raise NotConfigured when parse_node() or parse_row() is not defined; the resulting AttributeError is reported instead. ( #7768 )
scrapy.utils.misc.md5sum() , deprecated since Scrapy 2.12.0, is removed. ( #6264 , #8023 )
scrapy.utils.iterators.xmliter() , deprecated since Scrapy 2.11.1 because it is vulnerable to ReDoS attacks, is removed. Use xmliter_lxml() instead. ( #7765 )
scrapy.utils.datatypes.CaselessDict , deprecated since Scrapy 2.10.0, is removed. Use CaseInsensitiveDict instead. ( #5146 , #8023 )
The download_delay spider attribute is deprecated. Use the DOWNLOAD_DELAY setting, or DOWNLOAD_SLOTS to set a delay for specific domains, instead.
The max_concurrent_requests spider attribute, deprecated since Scrapy 2.13.0, now sets the CONCURRENT_REQUESTS_PER_DOMAIN setting, which is what it always mapped to, and warns accordingly.
Both attributes are ignored, with a different warning, when the corresponding setting is already set at the spider priority or higher.
( #7167 , #7175 , #7833 )
The Spider.log() method is deprecated. Use the methods of Spider.logger instead. ( #7739 )
The scrapy.interfaces module and its ISpiderLoader interface are deprecated. Custom spider loaders only need to follow SpiderLoaderProtocol . ( #6585 , #7731 )
scrapy.extensions.feedexport.IFeedStorage is deprecated. Custom feed storages only need to follow scrapy.extensions.feedexport.FeedStorageProtocol . ( #6585 , #7731 )
scrapy.utils.python.re_rsearch() is deprecated. ( #7765 )
Importing FileException from scrapy.pipelines.files is deprecated. Import it from scrapy.pipelines.media instead. ( #7544 , #7673 , #7973 )
Setting request.meta["is_secure"] to False to send an s3:// request over plaintext HTTP is deprecated. The flag will be ignored in a future Scrapy version. ( #7738 )
The unused multiplier attribute of PeriodicLog is deprecated. ( #7809 , #7982 )
Returning, from a log formatter method, a msg with %(name)s placeholders and no args is deprecated. Those placeholders are still interpolated with the returned dict , but in a future Scrapy version the message will be logged verbatim. Return those values under args instead. ( #5570 , #7971 )
Added optional extras for every optional dependency of Scrapy: bpython , gcs , httpx , images , ipython , ptpython , robotparser , s3 , twisted-http2 and uvloop . For example, pip install scrapy[s3,images] . ( #7596 )
HttpxDownloadHandler now uses httpx2 , the successor of httpx , which the new httpx extra installs together with its HTTP/2 and SOCKS proxy support. httpx is still used when httpx2 is not installed, but it is no longer tested. ( #7762 )
Added a robots_parsed signal, sent by RobotsTxtMiddleware after it parses a robots.txt file. It supports asynchronous handlers .
Added a crawl_delay() method to RobotParser , implemented by all built-in robots.txt parsers .
( #7830 )
Added a Request.to_curl() method, the inverse of from_curl() . ( #7743 , #7746 , #7802 )
Added a depth_reset request meta key that gives a request depth 0 instead of the depth of its source response plus 1. ( #891 , #7913 )
Added MetaCopyDetectionMiddleware , enabled by default, which warns once per crawl when a spider yields a request carrying internal meta keys that were likely copied from response.meta , and a META_COPY_WARN_SKIP_KEYS setting to exclude keys from that check. ( #7588 )
Added an AWS_MAX_POOL_CONNECTIONS setting, which defines the connection pool size of the AWS clients of the S3 feed storage backend and the S3 media pipeline storage backend , and defaults to REACTOR_THREADPOOL_MAXSIZE . It is also exposed as a max_pool_connections parameter of S3FeedStorage and as an AWS_MAX_POOL_CONNECTIONS attribute of S3FilesStore . ( #4985 , #7794 )
Added a scrapy.utils.asyncio.sleep() function, which works both with and without a Twisted reactor. ( #7843 )
CONCURRENT_REQUESTS can now be set to 0 for no limit. ( #7840 )
H2DownloadHandler is no longer experimental, and it now sends the bytes_received and headers_received signals and supports StopDownload . ( #5046 , #5047 , #5055 , #7896 , #7986 )
An exception raised by Spider.start() is now reported through the spider_error signal and the spider_exceptions/count and spider_exceptions/{exception} stats, and closes the spider with the new start_error finish_reason instead of finished . See Handling start errors .
CloseSpider raised from Spider.start() now closes the spider with the given reason, instead of being reported as a start error.
( #3463 , #4058 , #4182 , #6148 , #7884 )
CloseSpider can now also be raised while the spider is starting, e.g. from a spider_opened signal handler or from the open_spider() method of an item pipeline , to close the spider before it starts crawling. Every component still gets started, and stopped, before the spider is closed with the given reason. ( #3435 , #7905 )
Added an FTPS feed storage backend , i.e. support for the ftps URI scheme in FEEDS , which uploads the feed over a TLS connection, verifying the certificate of the server. ( #4180 , #7953 )
Changes to Spider.allowed_domains during a crawl are now taken into account by OffsiteMiddleware , whose should_follow() method is now documented as the way to implement a different offsite policy. ( #3257 , #3412 , #7903 , #7912 )
BaseSettings methods that take settings, such as update() and the settings parameter of crawler classes, now also accept an iterable of (name, value) tuples. ( #7759 , #7763 )
The cookies parameter of Request now also accepts bool , float and int values, and the formdata parameter of FormRequest now accepts any mapping or iterable of key-value pairs. ( #7858 , #7864 )
Added a scrapy.utils.reactorless.uninstall_reactor_import_hook() function, which AsyncCrawlerProcess.start() now uses to uninstall the twisted.internet.reactor import hook when it exits. ( #7747 )
Added the depth/request_ignored_count and httpcache/retrieve_error stats. ( #1308 , #2222 , #7805 , #7916 )
The get_serialized_fields() method of item exporters , previously named _get_serialized_fields() , is now public and documented, for custom item exporters to use. ( #5706 , #7931 )
Scrapy now writes the session keys of its HTTPS connections to the file that the SSLKEYLOGFILE environment variable points to, so that traffic analysis tools such as Wireshark can decrypt them. See Decrypting TLS traffic . ( #4368 , #7948 )
Added an HTTP2_MAX_FRAME_SIZE setting, which allows raising the maximum HTTP/2 frame size that servers may send, previously fixed at 16384, above which connections failed. ( #5050 , #7988 )
The crawl , parse and runspider commands now warn when FEEDS is set, e.g. through -o or -O , but the FeedExporter extension is disabled, so that no item is exported. ( #5970 , #6082 , #6373 , #7902 )
Log formatters ( LOG_FORMATTER ), item processors ( ITEM_PROCESSOR ) and robots.txt parsers ( ROBOTSTXT_PARSER ) are now built as components , so they no longer need a from_crawler() method. ( #7808 )
HttpCacheMiddleware now logs a warning and handles the request as a cache miss when reading a cache entry raises an exception, e.g. because the entry is corrupted, instead of letting the exception propagate. It also counts those entries in the new httpcache/retrieve_error stat. ( #2222 , #7805 )
Feed URIs now only expand %(...)s parameters, keeping any other percent character as is, so that percent-encoded URIs, e.g. one with %20 in a path or with percent-encoded FTP credentials, are no longer misinterpreted as printf-style formatting directives. ( #5794 , #6425 , #7674 )
Feed exports now start storing a FEED_EXPORT_BATCH_ITEM_COUNT batch as soon as it is complete, instead of waiting until the spider closes. ( #7730 , #7733 )
CsvItemExporter now warns when the fields that it took from the first item do not cover the fields of a later item, i.e. when it silently drops data. ( #4002 , #4053 , #7613 , #7651 )
GCSFeedStorage no longer requires the storage.buckets.get permission. ( #5475 , #7945 )
Media pipelines now log media requests that were filtered out, e.g. as offsite requests, at the DEBUG level and without a traceback, instead of reporting them as download errors. ( #7544 , #7673 )
OffsiteMiddleware now raises IgnoreRequest with a message, e.g. Filtered offsite request to 'offsite.example' , which errbacks and log messages that report that exception now include. ( #7544 , #7673 )
HTTP11DownloadHandler now skips response header lines that have no colon, logging them at the DEBUG level, as web browsers do, instead of being unable to download such a response at all. ( #210 , #7806 )
CookiesMiddleware now sends domain cookies to hosts without a dot in their name and to hosts given as an IP address. ( #6410 , #7900 )
TextResponse.json() now decodes bodies that are not valid UTF-8, UTF-16 or UTF-32 using TextResponse.encoding , instead of raising UnicodeDecodeError . ( #6456 , #7897 )
scrapy.resolver.CachingHostnameResolver now caches addresses without a port, and sets the requested port on cache hits, so that a cached address no longer carries the port of the request that populated the cache. ( #6442 , #7772 )
DownloaderAwarePriorityQueue now removes the directory of a download slot from the JOBDIR directory once that slot is drained. ( #5275 , #7955 )
TelnetConsole no longer raises an exception on shutdown when it could not listen on any of the TELNETCONSOLE_PORT ports. ( #2702 , #7910 )
The DOWNLOAD_WARNSIZE warning is no longer logged twice for a response whose Content-Length header already exceeded the limit. ( #2476 , #7963 )
HttpCompressionMiddleware now logs a warning when it drops a response for exceeding DOWNLOAD_MAXSIZE during decompression. ( #6616 , #7742 )
DepthMiddleware now logs only the first request ignored for exceeding DEPTH_LIMIT , and counts them all in the new depth/request_ignored_count stat. ( #1308 , #7916 )
parse now sets the callback it uses on the request of the response it passes to that callback. ( #3095 , #3124 , #7803 )
The IPython shell now works when an asyncio event loop is already running in the same thread, e.g. when calling scrapy.shell.inspect_response() from a callback while using the asyncio reactor. ( #5447 , #7816 )
Request.from_curl() now merges repeated -d , --data and --data-raw options into a single body joined with & , as curl does, instead of keeping only the last one. ( #7728 )
The copy() method and the |= operator of scrapy.utils.datatypes.CaseInsensitiveDict no longer leave the internal mapping of original key spellings shared or out of date. ( #7783 )
ExecutionEngine.download_async() no longer recurses once per returned request, e.g. once per redirect. ( #7544 , #7673 )
LinkExtractor now canonicalizes each extracted URL once instead of twice when canonicalize is True . ( #7961 )
Items yielded from Spider.start() now keep the spider busy until the item pipelines are done with them, so that a spider that only yields items from start() no longer closes before processing them. ( #7029 , #7891 )
open_in_browser() now also adds its base tag to HTML responses that have no head element, and it now overrides a base tag already present in the response, so that relative URLs resolve against the response URL in every case. ( #6550 , #7879 )
Spider.start() implementations that are not asynchronous generators now raise TypeError with a message that says so, instead of failing in a way that does not point at the cause. ( #5426 , #7946 )
The check , fetch and parse commands now return the exit code 1 when a component fails to initialize, as crawl and runspider already did. ( #4292 , #7920 )
HttpCompressionMiddleware no longer hangs on a deflate response body followed by extra bytes. ( #7841 )
scrapy.utils.python.get_func_args() now reports the parameters that a functools.partial object binds by position, instead of an empty list. ( #7841 )
Fixed NameError exceptions on Python 3.14, where PEP 649 made annotation evaluation lazy, when inspecting the signature of a callable with annotations imported only for type checking. ( #7796 , #7818 )
scrapy.utils.decorators.deprecated can now be used both as @deprecated and as @deprecated(...) without confusing type checkers. ( #7797 )
The default download handlers can now download from domains with emoji characters or underscores, which were previously rejected. ( #3321 , #4330 , #7846 )
Callbacks and media pipeline results no longer wait 100 ms before proceeding. ( #8019 )
Shutting down a crawl no longer risks raising an unhandled RuntimeError if the code interrupted by the shutdown signal was itself writing to the log. ( #8022 )
Nested selectors, e.g. the result of calling jmespath() on a selector, now let parsel determine their type instead of forcing the html type, so that they no longer return the wrong type or value. ( #8038 , #8040 )
Added a built-in stats reference , covering every stat that Scrapy sets. ( #6351 , #7814 )
Replaced the broad crawls page with a new optimization page, about finding the bottleneck of a crawl before changing any setting, which covers broad crawls as one of its sections. ( #4737 , #7938 )
Added a concepts page, a quick map of Scrapy’s main building blocks, what problem each one solves, and when to reach for it. ( #1569 , #8025 )
Added a cookies page, which gathers what used to be spread across the request and downloader middleware pages. ( #7947 )
Added callbacks and errbacks sections to the request and response page, covering callback assignment , how to write a callback and supported callback output . ( #5054 , #6437 , #7821 , #7898 )
Documented the ITEM_PROCESSOR setting and the ItemProcessorProtocol protocol that its value must implement. ( #7983 )
Documented how to write an item exporter , how to test an item pipeline , how to download a request from a downloader middleware , how to name media files after the response , how to add objects to the shell and how to run spiders inside an existing application or in a Jupyter notebook . ( #915 , #1199 , #2594 , #5706 , #6554 , #6594 , #7751 , #7872 , #7876 , #7889 , #7909 , #7931 )
Added an inspecting live traffic section to the debugging page, covering Wireshark and mitmproxy. ( #5222 , #8007 )
Documented how TextResponse resolves the response encoding, and how to resolve it differently, e.g. to give the encoding declared in the response body precedence over the Content-Type header. ( #4933 , #7977 )
Documented the memory use of response parsing and the parser limits that Scrapy lifts, in the security page. ( #5700 , #7930 )
Documented that signal handlers run in an undefined order , that scheduler_empty must only be awaited from start() , that concurrency and politeness settings apply per crawler when running multiple spiders in the same process , and that a JOBDIR directory cannot be shared across Scrapy versions. ( #3191 , #5330 , #5522 , #7861 , #7883 , #7907 , #7941 )
Documented that the html iterator of XMLFeedSpider can silently mangle tags that HTML treats as void elements, e.g. <link> , dropping their content and closing tag. ( #4675 , #8045 )
Documented that the keep_fragments parameter of fingerprint() is not a substitute for rendering JavaScript to reach content that a headless browser loads based on the URL fragment. ( #4789 , #8033 )
Documented how to derive the job directory from the spider name . ( #4748 , #8035 )
Documented that the project name from scrapy.cfg also appears elsewhere by default, and which of those uses actually require it to match. ( #2484 , #8044 )
Many other corrections and improvements. ( #4589 , #4796 , #5532 , #5548 , #6053 , #6184 , #6627 , #6787 , #6943 , #6989 , #7710 , #7725 , #7737 , #7767 , #7769 , #7771 , #7774 , #7775 , #7777 , #7779 , #7780 , #7817 , #7832 , #7835 , #7862 , #7871 , #7875 , #7880 , #7890 , #7903 , #7913 , #7917 , #7939 , #7940 , #7962 , #7965 )
Improved and fixed type hints. ( #7712 , #7785 , #7858 , #7864 , #7865 , #7867 )
Added CPU benchmarks, tracked on CodSpeed, so that performance regressions are caught before they are merged and performance work can be measured. ( #7831 , #7839 , #7870 , #7887 , #7914 , #7954 )
Added a nightly job that runs the test suite against the development branches of dependencies, so that incompatibilities are found before those dependencies are released. ( #5291 , #6025 , #7924 , #7960 )
CI and test improvements and fixes. ( #5049 , #5620 , #5837 , #6478 , #6794 , #7262 , #7437 , #7702 , #7720 , #7724 , #7727 , #7736 , #7741 , #7749 , #7753 , #7755 , #7768 , #7778 , #7782 , #7792 , #7793 , #7795 , #7797 , #7798 , #7809 , #7829 , #7834 , #7836 , #7838 , #7841 , #7844 , #7848 , #7853 , #7854 , #7857 , #7859 , #7863 , #7895 , #7906 , #7928 , #7935 , #7966 , #7968 , #7974 , #7979 , #7985 , #7990 , #7993 , #7995 , #8000 , #8001 , #8002 )
HTTP/2 and SOCKS proxy support for HttpxDownloadHandler
Highlights:
Security bug fixes
HTTP/2 and SOCKS proxy support for HttpxDownloadHandler
Improved settings for changing allowed TLS versions
s3:// requests now use HTTPS by default, instead of plaintext HTTP.
Previously, S3DownloadHandler sent signed S3 requests over plaintext HTTP unless request.meta["is_secure"] was set to a true value, exposing the request path, the AWS Authorization header, the X-Amz-Security-Token header (when using temporary credentials), and the response contents to network attackers, who could also tamper with responses. See the 76g3-c3x4-crvx security advisory for details.
To restore the previous behavior for a given request, set request.meta["is_secure"] to False .
The DOWNLOADER_CLIENT_TLS_METHOD setting is deprecated. You should use the DOWNLOAD_TLS_MIN_VERSION and/or DOWNLOAD_TLS_MAX_VERSION settings instead if you want to change the TLS method selection. ( #3288 , #6546 )
The following spider attributes are deprecated in favor of settings:
http_user (use HTTPAUTH_USER )
http_pass (use HTTPAUTH_PASS )
http_auth_domain (use HTTPAUTH_DOMAIN )
( #7590 )
The scrapy.commands.ScrapyCommand.help() method is deprecated. It was never called by Scrapy. ( #7626 , #7633 )
The following TLS-related functions and constants, intended for internal use, are deprecated:
scrapy.core.downloader.tls.METHOD_TLS
scrapy.core.downloader.tls.METHOD_TLSv10
scrapy.core.downloader.tls.METHOD_TLSv11
scrapy.core.downloader.tls.METHOD_TLSv12
scrapy.core.downloader.tls.openssl_methods
scrapy.core.downloader.tls.DEFAULT_CIPHERS
scrapy.utils.ssl.ffi_buf_to_string()
scrapy.utils.ssl.get_temp_key_info()
scrapy.utils.ssl.x509name_to_string()
( #6546 , #7619 , #7665 )
The CRAWLSPIDER_FOLLOW_LINKS setting is deprecated. You can set follow=False in your rules to achieve the same effect. ( #7592 )
Instantiating HttpCompressionMiddleware without a crawler argument is deprecated. ( #7655 )
Instantiating RefererMiddleware without a settings argument is deprecated. ( #7664 )
Added support for HTTP/2 requests to HttpxDownloadHandler . It requires setting the new HTTPX_HTTP2_ENABLED setting to True . ( #7575 )
Added support for SOCKS proxies to HttpxDownloadHandler . ( #747 , #7575 )
Added DOWNLOAD_TLS_MIN_VERSION and DOWNLOAD_TLS_MAX_VERSION settings as replacements for the DOWNLOADER_CLIENT_TLS_METHOD setting (which is now deprecated). Compared to the old setting, they support specifying a range of allowed versions and support newer TLS versions. ( #4821 , #6546 )
Added HTTPAUTH_USER , HTTPAUTH_PASS and HTTPAUTH_DOMAIN settings and http_user , http_pass and http_auth_domain meta keys as more flexible ways to set HTTP authentication data. ( #7590 )
Added a verbatim_url meta key that can be set to True to skip request URL canonicalization. ( #7473 )
Added deny_tags and deny_attrs arguments to LinkExtractor . ( #6321 , #7679 )
scrapy.Item.fields now returns the fields in the definition order instead of the alphabetical one. ( #7015 , #7694 )
Added a RETRY_GIVE_UP_LOG_LEVEL setting, a give_up_log_level meta key and a give_up_log_level argument of the get_retry_request() function that allow changing the log level of the message logged when the retry limit has been reached. ( #4622 , #5297 , #7567 )
It’s now possible to set DOWNLOADER_CLIENT_TLS_CIPHERS to None to use the default ciphers of the underlying TLS implementation. ( #7499 , #7665 )
FormRequest is no longer deprecated, only its from_response() method is still deprecated. ( #7561 , #7671 )
Switched the item definition in the default project template from a scrapy.item.Item to a dataclass. ( #7493 , #7513 )
Fixed deprecation warnings with pyOpenSSL 26.3.0. ( #7619 )
Removed the runtime warnings for Spider.allowed_domains containing URLs or domains with ports instead of just domains and for spider classes having a start_url attribute instead of start_urls . Please use scrapy-lint to find mistakes in your spider code instead. ( #4421 , #7627 )
scrapy.utils.test.get_crawler() now disables TELNETCONSOLE_ENABLED by default. ( #7644 )
Other code refactoring and improvements. ( #7409 , #7593 , #7594 , #7611 , #7649 )
HttpxDownloadHandler no longer ignores proxy credentials for redirected or retried requests. ( #7601 , #7630 )
GCSFeedStorage now closes the temporary file after the upload. ( #7546 )
Fixed scrapy shell <URL> running a full spider crawl when there is a spider for the requested URL. This bug was introduced in Scrapy 2.13.0. ( #7552 , #7557 )
The IMAGES_STORE_S3_ACL and IMAGES_STORE_GCS_ACL settings are no longer ignored. This bug was introduced in Scrapy 2.12.0. ( #7597 , #7614 )
FTPDownloadHandler now closes the connection after making the request. ( #7602 , #7667 )
Removed the deprecated spider argument from the pipeline defined in the default project template. ( #7676 )
Fixed scrapy genspider --edit not working. ( #7260 , #7683 )
When a Crawler instance is passed to AsyncCrawlerRunner.create_crawler() or CrawlerRunner.create_crawler() , settings from both classes are now merged, previously only the settings from the Crawler instance were used. ( #1280 , #7647 )
Fixed several issues with cookie handling in scrapy.utils.request.request_to_curl() . ( #7603 , #7675 , #7684 )
Fixed scrapy.resolver.CachingThreadedResolver not disabling the cache when DNSCACHE_ENABLED is set to False . ( #7663 )
Fixed scrapy.utils.response.open_in_browser() not removing comments when looking for the <base> tag. ( #7506 )
Fixed checking for deprecated methods in custom ITEM_PROCESSOR implementations. ( #7589 )
Fixed scrapy.utils.url.strip_url() corrupting some URLs with credentials. ( #7604 , #7605 )
scrapy.utils.misc.rel_has_nofollow() now ignores the case when looking for “nofollow” strings. ( #7632 )
Fixed an exception in scrapy.utils.sitemap.Sitemap when parsing some malformed sitemaps. ( #7686 , #7687 )
Mentioned scrapy-lint in the docs. ( #4421 , #7627 )
Added the docs about security considerations . ( #7389 , #7678 )
Improved the item pipeline docs . ( #2350 , #7676 )
Documented which stats are collected by CoreStats . ( #7421 )
Switched documentation examples from using scrapy.item.Item to using dataclasses. ( #7493 , #7513 )
Added feature comparison tables to the download handler docs. ( #7575 )
Improved the docs for logging settings . ( #6909 , #7668 )
Documented a way to improve startup time and memory usage by using SPIDER_MODULES . ( #7576 , #7600 )
Clarified handling of the type argument of Selector . ( #7704 )
Other documentation improvements and fixes. ( #4954 , #6120 , #7286 , #7564 , #7573 , #7598 , #7599 , #7698 )
Fixed deprecation warnings with pytest 9.1.0. ( #7621 )
Type hints improvements and fixes. ( #6958 , #7586 )
CI and test improvements and fixes. ( #5954 , #7002 , #7017 , #7247 , #7508 , #7545 , #7566 , #7574 , #7585 , #7595 , #7608 , #7610 , #7612 , #7616 , #7625 , #7637 , #7639 , #7640 , #7641 , #7642 , #7643 , #7644 , #7645 , #7646 , #7654 , #7655 , #7664 , #7672 , #7677 , #7680 , #7682 , #7692 )
Official support for Python 3.14
Highlights:
Official support for Python 3.14
Support for Twisted 26.4.0+
Increased the minimum versions of the following dependencies:
service_identity : 18.1.0 → 23.1.0
( #7347 )
Added support for Twisted 26.4.0+. ( #7347 , #7505 , #7520 )
Added support for Python 3.14. ( #6604 , #7460 )
The following classes and functions, intended for internal use by HTTP11DownloadHandler and H2DownloadHandler , have been made private:
scrapy.core.downloader.handlers.http11.ScrapyAgent
scrapy.core.downloader.handlers.http11.ScrapyProxyAgent
scrapy.core.downloader.handlers.http11.TunnelingAgent
scrapy.core.downloader.handlers.http11.TunnelingTCP4ClientEndpoint
scrapy.core.downloader.handlers.http11.tunnel_request_data()
scrapy.core.downloader.handlers.http2.ScrapyH2Agent
( #7496 , #7510 )
scrapy.FormRequest is deprecated. You can use the form2request library instead, see Creating requests that submit HTML forms . ( #6438 )
scrapy.utils.python.MutableChain is deprecated. ( #7504 )
The start_requests() method of Spider , deprecated in 2.13.0, is removed and no longer called. Use start() instead, or both to maintain support for lower Scrapy versions. ( #7490 )
Support for process_start_requests() methods of spider middlewares , deprecated in 2.13.0, is removed. Use process_start() instead, or both to maintain support for lower Scrapy versions. ( #7490 )
Support for synchronous process_spider_output() methods of spider middlewares, deprecated in Scrapy 2.13.0, is removed. You should upgrade the affected middlewares to have asynchronous process_spider_output() methods. ( #7504 )
The spider arguments of the following methods of Scraper , deprecated in Scrapy 2.13.0, are removed:
close_spider()
enqueue_scrape()
handle_spider_error()
handle_spider_output()
( #7487 )
HTTP/1.0 support code, deprecated in Scrapy 2.13.0, is removed. This includes:
scrapy.core.downloader.handlers.http10.HTTP10DownloadHandler
The scrapy.core.downloader.webclient module.
The DOWNLOADER_HTTPCLIENTFACTORY setting.
( #7486 )
The following functions, deprecated in Scrapy 2.13.0, are removed, you should import them from w3lib.url directly instead:
scrapy.utils.url.add_or_replace_parameter()
scrapy.utils.url.add_or_replace_parameters()
scrapy.utils.url.any_to_uri()
scrapy.utils.url.canonicalize_url()
scrapy.utils.url.file_uri_to_path()
scrapy.utils.url.is_url()
scrapy.utils.url.parse_data_uri()
scrapy.utils.url.parse_url()
scrapy.utils.url.path_to_file_uri()
scrapy.utils.url.safe_download_url()
scrapy.utils.url.safe_url_string()
scrapy.utils.url.url_query_cleaner()
scrapy.utils.url.url_query_parameter()
( #7487 )
The following test-related code, deprecated in Scrapy 2.13.0, is removed:
the scrapy.utils.testproc module
the scrapy.utils.testsite module
scrapy.utils.test.assert_gcs_environ()
scrapy.utils.test.get_ftp_content_and_delete()
scrapy.utils.test.get_gcs_content_and_delete()
scrapy.utils.test.mock_google_cloud_storage()
scrapy.utils.test.skip_if_no_boto()
scrapy.utils.test.TestSpider
( #7487 )
scrapy.utils.versions.scrapy_components_versions() , deprecated in Scrapy 2.13.0, is removed, you can use scrapy.utils.versions.get_versions() instead. ( #7487 )
scrapy.downloadermiddlewares.ajaxcrawl.AjaxCrawlMiddleware and scrapy.utils.url.escape_ajax() , deprecated in Scrapy 2.13.0, are removed. ( #7487 )
The init() method of priority queue classes (see SCHEDULER_PRIORITY_QUEUE ) now needs to support a keyword-only start_queue_cls parameter, not supporting it was deprecated in Scrapy 2.13.0. ( #7487 )
scrapy.spiders.init.InitSpider , deprecated in Scrapy 2.13.0, is removed. ( #7487 )
New features and improvements for HttpxDownloadHandler :
Support for proxies.
Support for the download_latency meta key.
Support for Response.certificate .
Default headers set by the httpx library are no longer added to requests.
( #7441 , #7524 )
HTTP11DownloadHandler now skips HTTPS proxy certificate verification when the DOWNLOAD_VERIFY_CERTIFICATES setting is set to False . ( #7496 )
time.monotonic() is used instead of time.time() to calculate elapsed time in various places. ( #7377 )
Improved extraction of the file extension from the URL in FilesPipeline . ( #4225 , #7414 )
Other code refactoring and improvements. ( #7401 )
HTTP11DownloadHandler now raises an exception when a request has an https:// destination and an https:// proxy, which is not supported by this handler. Previously it tried to connect to the proxy via HTTP in this case. ( #7496 )
H2DownloadHandler now raises an exception for requests with http:// URLs instead of trying to connect, which is not supported by this handler. ( #7496 )
H2DownloadHandler no longer adds the :status pseudo-header to Response.headers . ( #7441 )
Fixed scrapy.utils.response.open_in_browser() removing the <head> tag when adding the <base> tag. ( #7459 )
Documented that HTTP11DownloadHandler doesn’t support HTTPS proxies for HTTPS destinations and that H2DownloadHandler doesn’t support proxies at all. ( #7496 )
Added an example of using logging.handlers.TimedRotatingFileHandler to rotate Scrapy logs. ( #3628 , #7501 )
Added a CITATION.cff file. ( #7502 , #7519 )
Mentioned DOWNLOADER_CLIENT_TLS_METHOD in Avoiding getting banned . ( #5232 , #7518 )
Other documentation improvements and fixes. ( #7417 , #7463 , #7472 , #7480 , #7489 , #7503 , #7507 )
Added tests that connect to https://books.toscrape.com/ to test the behavior with a real website. These tests are marked with the requires_internet pytest mark and can be skipped with e.g. -m 'not requires_internet' if you cannot or don’t want to run them. ( #7520 )
Type hints improvements and fixes. ( #7492 , #7532 )
CI and test improvements and fixes. ( #7441 , #7466 , #7491 , #7496 )
Fixed links in https://docs.scrapy.org/llms.txt
Full Changelog: 2.15.1...2.15.2
Bug fixes Full Changelog
Sharing of the SSL context between multiple connections, introduced in Scrapy 2.15.0, is reverted as it caused problems and wasn’t actually needed. ( #7445 , #7450 )
Fixed scrapy.settings.BaseSettings.getwithbase() failing on keys with dots that aren’t import names. It now works the way it worked before Scrapy 2.15.0, without trying to match class objects and import path. A separate method, get_component_priority_dict_with_base() , was added that does that, and it is now used for component priority dictionaries . ( #7426 , #7449 )
Documentation rendering improvements. ( #7452 , #7454 )
Experimental support for running without a Twisted reactor
httpx-based download handlerHighlights:
Experimental support for running without a Twisted reactor
Experimental httpx -based download handler
The built-in HTTP download handlers now raise Scrapy-specific exceptions instead of implementation-specific ones, see Exceptions raised by download handlers . This can affect user code that handles downloader exceptions, such as process_exception() methods of custom downloader middlewares . ( #7208 )
In order to fix a long-standing bug with handling of asynchronous storages, the following changes were made to media pipeline classes, which can impact some of the user code that subclasses them or calls their methods directly:
overrides of scrapy.pipelines.media.MediaPipeline.media_downloaded() and file_downloaded() can now return coroutines
media_downloaded() , file_downloaded() and image_downloaded() now return coroutines
( #2183 , #6369 , #7182 )
Request and Response objects: slots and setter changes:
scrapy.http.Request and scrapy.http.Response now define slots . Assigning arbitrary attributes to instances (for example, response.foo = 1 ) will raise AttributeError . Store per-request/response data in the request/response meta mapping instead of attaching new attributes to the objects.
If you maintain custom Request or Response subclasses that relied on dynamic instance attributes, either add 'dict' to your subclass slots to allow dynamic attributes, or migrate per-instance state to meta or explicit documented attributes.
The setters for headers , flags and cookies no longer coerce falsy values into None . For example, request.headers = {} now stores an empty scrapy.http.headers.Headers instance (not None ), and request.flags = [] remains an empty list instead of being set to None . Update code that relied on is None checks or the previous coercion behaviour.
( #7036 , #7367 , #7374 )
The context factory class set as the value of the DOWNLOADER_CLIENTCONTEXTFACTORY setting is now required to support the method argument of init() , recommended since Scrapy 1.2.0. ( #7353 )
scrapy.mail.MailSender is deprecated. Please use smtplib , twisted.mail.smtp or other 3rd party email libraries. ( #7249 , #7263 )
The scrapy.extensions.statsmailer.StatsMailer extension is deprecated. You can instead implement your own notifications by handling the spider_closed signal. ( #7249 , #7263 )
The MEMUSAGE_NOTIFY_MAIL setting is deprecated. You can instead implement your own notifications by handling the memusage_warning_reached and spider_closed signals. ( #7249 , #7263 )
The DNS_RESOLVER setting was renamed to TWISTED_DNS_RESOLVER and the old name is deprecated. ( #7350 , #7361 )
The DOWNLOADER_CLIENTCONTEXTFACTORY setting is deprecated. If you were using it to switch to scrapy.core.downloader.contextfactory.BrowserLikeContextFactory , please use the new DOWNLOAD_VERIFY_CERTIFICATES setting instead. If you cannot use the default context factory for some other reason, please subclass the download handler instead. ( #7352 , #7379 )
scrapy.core.downloader.contextfactory.BrowserLikeContextFactory is deprecated. You can set the new DOWNLOAD_VERIFY_CERTIFICATES setting to True instead. ( #7379 )
The following implementation details of the context factory handling code are deprecated:
scrapy.core.downloader.contextfactory.AcceptableProtocolsContextFactory
scrapy.core.downloader.contextfactory.load_context_factory_from_settings()
scrapy.core.downloader.contextfactory.ScrapyClientContextFactory
scrapy.core.downloader.tls.ScrapyClientTLSOptions
( #7353 , #7391 )
Passing str instead of bytes to scrapy.utils.sitemap.Sitemap and scrapy.utils.sitemap.sitemap_urls_from_robots() is deprecated. ( #7007 )
scrapy.utils.misc.walk_modules() is deprecated. You can use scrapy.utils.misc.walk_modules_iter() instead. ( #7388 )
scrapy.shell.Shell.inthread is deprecated. You can use scrapy.shell.Shell.fetch_available instead to check if fetch() can be used. ( #7395 )
scrapy.commands.ScrapyCommand.set_crawler() is deprecated. ( #7276 )
Added an experimental mode for running Scrapy without installing a Twisted reactor: set TWISTED_REACTOR_ENABLED to False to enable it. This mode has limitations, refer to its documentation for details. As long as it’s experimental, its behavior and related features and APIs may change in future Scrapy releases in a breaking way. ( #6219 , #7185 , #7186 , #7187 , #7188 , #7190 , #7197 , #7199 , #7209 , #7228 , #7355 , #7366 , #7385 , #7395 )
Added the scrapy.utils.reactorless.is_reactorless() function that checks if there is a running asyncio event loop but no Twisted reactor. ( #7185 , #7199 )
Changed scrapy.utils.asyncio.is_asyncio_available() to return True if there is a running asyncio loop, even if no Twisted reactor is installed. ( #7185 , #7199 )
Added an experimental download handler that uses the httpx library and doesn’t require a Twisted reactor: HttpxDownloadHandler . As long as it’s experimental, its behavior may change in future Scrapy releases in a breaking way. ( #6805 , #7239 , #7368 , #7384 )
Added the DOWNLOAD_BIND_ADDRESS setting as a global counterpart to the per-request bindaddress meta key. ( #7266 , #7283 )
Added the DOWNLOAD_VERIFY_CERTIFICATES setting that can be set to True to make Scrapy abort HTTPS requests when the server certificate is invalid or doesn’t match the domain. ( #7379 )
The built-in HTTP download handlers now raise Scrapy-specific exceptions instead of implementation-specific ones, to allow unified handling of similar problems caused by different implementations. The default value of the RETRY_EXCEPTIONS setting was updated replacing Twisted-specific exceptions with these new ones. The exceptions:
CannotResolveHostError
DownloadCancelledError
DownloadConnectionRefusedError
DownloadFailedError
DownloadTimeoutError
ResponseDataLossError
UnsupportedURLSchemeError
( #7208 )
Added the memusage_warning_reached signal emitted by the MemoryUsage extension when the memory usage reaches MEMUSAGE_WARNING_MB . ( #7249 , #7263 )
Added Headers.to_tuple_list() that returns headers as a list of (key, value) tuples. ( #7239 )
S3DownloadHandler now uses the download handler configured for the "https" scheme to make requests instead of always using HTTP11DownloadHandler . ( #7369 , #7370 )
Added scrapy.utils.misc.walk_modules_iter() as a replacement for scrapy.utils.misc.walk_modules() that returns an iterable instead of a list. ( #7388 )
asyncio.to_thread() is now used instead of twisted.internet.threads.deferToThread() in the built-in feed storages, media pipeline storages and the scrapy.utils.decorators.inthread() decorator when available. ( #7183 , #7184 , #7349 )
Improved memory footprint of Request and Response objects by adding slots and omitting empty lists and dicts in some internal attributes. ( #7036 , #7367 , #7374 )
_ScrapyClientContextFactory no longer mutates the SSL context, to avoid the behavior that was deprecated in pyOpenSSL 25.1.0. ( #6859 , #7353 )
Improved memory usage of SitemapSpider and scrapy.utils.sitemap.Sitemap . ( #3529 , #7007 )
Improved the scheduling behavior of DownloaderAwarePriorityQueue when crawling multiple domains. ( #7293 , #7351 )
HTTP11DownloadHandler and H2DownloadHandler now handle TLS verbose logging (see DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING ) directly instead of relying on _ScrapyClientContextFactory . ( #7387 )
The server certificate verification code now correctly handles certificates with IP addresses in subjectAltName . ( #7353 )
Improved reliability of scrapy.utils.trackref.get_oldest() . ( #1758 , #7375 )
Other code refactoring and improvements. ( #7210 , #7238 , #7376 , #7386 , #7395 , #7405 , #7410 )
Media pipelines should now wait for uploads to asynchronous storages (e.g. S3FilesStore ) to complete. ( #2183 , #6369 , #7182 )
Fixed merging *_BASE settings (e.g. merging DOWNLOADER_MIDDLEWARES with DOWNLOADER_MIDDLEWARES_BASE ) when a component is referred to by a class object in one setting and by a string import path in the other one. ( #6912 , #6993 )
scrapy runspider and scrapy crawl now set the exit code to 1 if an exception happened early (this was broken since Scrapy 2.13.0). ( #6820 , #7255 )
Fixed repeated warnings about data loss (see DOWNLOAD_FAIL_ON_DATALOSS ) not being suppressed in HTTP11DownloadHandler . ( #7222 )
Improved FTP connection management in scrapy.pipelines.files.FTPFilesStore . ( #7256 )
Fixed the spider variable in the shell , which wasn’t available since Scrapy 2.13.0. ( #7395 )
The llms.txt and llms-full.txt files and Markdown versions of pages are now generated when the HTML documentation is built. ( #7380 )
Added a “Copy as Markdown” button to the HTML documentation. ( #7380 )
Added docs for using Pydantic models as items . ( #6955 , #6966 )
Documented job directory contents . ( #4842 , #5260 )
Improved docs for dont_filter . ( #6398 , #7245 )
Clarified that settings related to TWISTED_DNS_RESOLVER are only taken into account if the selected resolver supports them. ( #7385 )
Other documentation improvements and fixes. ( #7248 , #7274 , #7406 , #7408 )
Added the no-reactor test environment that doesn’t install a Twisted reactor and uses pytest-asyncio instead of pytest-twisted to run asynchronous test functions. ( #6952 , #7189 , #7233 , #7234 , #7254 , #7259 )
Fixed running tests with pytest-xdist . ( #7216 , #7257 )
Type hints improvements and fixes. ( #7300 , #7331 )
CI and test improvements and fixes. ( #7060 , #7223 , #7232 , #7241 , #7250 , #7256 , #7276 , #7277 , #7279 , #7329 , #7363 , #7381 , #7402 )
Values from the Referrer-Policy header of HTTP responses are no longer executed as Python callables. See the cwxj-rr6w-m6w7 security advisory for deta
Referrer-Policy header of HTTP responses are no longer executed as Python callables. See the cwxj-rr6w-m6w7 security advisory for details.Values from the Referrer-Policy header of HTTP responses are no longer executed as Python callables. See the cwxj-rr6w-m6w7 security advisory for details.
In line with the standard , 301 redirects of POST requests are converted into GET requests.
Converting to a GET request implies not only a method change, but also omitting the body and Content-* headers in the redirect request. On cross-origin redirects (for example, cross-domain redirects), this is effectively a security bug fix for scenarios where the body contains secrets.
Passing a response URL string as the first positional argument to scrapy.spidermiddlewares.referer.RefererMiddleware.policy() is deprecated. Pass a Response instead.
The parameter has also been renamed to response to reflect this change. The old parameter name ( resp_or_url ) is deprecated.
Added a new setting, REFERRER_POLICIES , to allow customizing supported referrer policies.
Made additional redirect scenarios convert to GET in line with the standard :
Only POST 302 redirects are converted into GET requests; other methods are preserved.
HEAD 303 redirects are not converted into GET requests.
GET 303 redirects do not have their body or standard Content-* headers removed.
Redirects where the original request body is dropped now also have their Content-Encoding , Content-Language and Content-Location headers removed, in addition to the Content-Type and Content-Length headers that were already being removed.
Redirects now preserve the source URL fragment if the redirect URL does not include one. This is useful when using browser-based download handlers, such as scrapy-playwright or scrapy-zyte-api , while letting Scrapy handle redirects.
The Referer header is now removed on redirect if RefererMiddleware is disabled.
The handling of the Referer header on redirects now takes into account the Referer-Policy header of the response that triggers the redirect.
Replace deprecated Codecov CI action
maybeDeferred_coro(){open,close}_spider()scrapy.utils.defer.maybeDeferred_coro() is deprecated. ( #7212 )
Fixed custom stats collectors that require a spider argument in their open_spider() and close_spider() methods not receiving the argument when called by the engine.
Note, however, that the spider argument is now deprecated and will stop being passed in a future version of Scrapy.
( #7213 )
Replaced deprecated codecov/test-results-action@v1 GitHub Action with codecov/codecov-action@v5 . ( #7180 , #7215 )
More coroutine-based replacements for Deferred-based APIs
DownloaderAwarePriorityQueueHighlights:
More coroutine-based replacements for Deferred-based APIs
The default priority queue is now DownloaderAwarePriorityQueue
Dropped support for Python 3.9 and PyPy 3.10
Improved and documented the API for custom download handlers
Dropped support for Python 3.9. ( #7121 )
Dropped support for PyPy 3.10. ( #7050 )
Increased the minimum versions of the following dependencies:
lxml : 4.6.0 → 4.6.4
Pillow (optional dependency): 8.0.0 → 8.3.2
botocore (optional dependency): 1.4.87 → 1.13.45
Restored support for brotlicffi dropped in Scrapy 2.13.4. Its minimum supported version is now 1.2.0.0 . ( #7160 )
If you set the TWISTED_REACTOR setting to a non-asyncio value at the spider level , you may now need to set the FORCE_CRAWLER_PROCESS setting to True when running Scrapy via its command-line tool to avoid a reactor mismatch exception. ( #6845 )
The log_count/* stats no longer count some of the early messages that they counted before. While the earliest log messages, emitted before the counter is initialized, were never counted, the counter initialization now happens later than in previous Scrapy versions. You may need to adjust expected values if you retrieve and compare values of these stats in your code. ( #7046 )
The classes listed below are now abstract base classes . They cannot be instantiated directly and their subclasses need to override the abstract methods listed below to be able to be instantiated. If you previously instantiated these classes directly, you will now need to subclass them and provide trivial (e.g. empty) implementations for the abstract methods.
scrapy.commands.ScrapyCommand
run()
short_desc()
scrapy.exporters.BaseItemExporter
export_item()
scrapy.extensions.feedexport.BlockingFeedStorage
_store_in_thread()
scrapy.middleware.MiddlewareManager
_get_mwlist_from_settings()
scrapy.spidermiddlewares.referer.ReferrerPolicy
referrer()
( #6930 )
Scrapy no longer passes a spider argument to any methods of the stats collector . It wasn’t passed in many of the calls even in older Scrapy versions, so we don’t expect existing custom stats collector implementations to require a spider argument. If your implementation needs a Spider instance, you can get it from the Crawler instance passed to the constructor. ( #7011 )
scrapy.middleware.MiddlewareManager no longer includes code for handling open_spider() and close_spider() component methods. As this code was only used for pipelines it was moved into scrapy.pipelines.ItemPipelineManager . This change should only affect custom subclasses of MiddlewareManager . The following code was moved:
scrapy.middleware.MiddlewareManager.open_spider()
scrapy.middleware.MiddlewareManager.close_spider()
Code in scrapy.middleware.MiddlewareManager._add_middleware() that processes open_spider() and close_spider() component methods.
( #7006 )
scrapy.downloadermiddlewares.robotstxt.RobotsTxtMiddleware.process_request() now returns a coroutine, previously it returned a Deferred object or None . The robot_parser() method was also changed to return a coroutine. This change only impacts code that subclasses RobotsTxtMiddleware or calls its methods directly. ( #6802 )
The built-in download handlers have been refactored, changing the signatures of their methods. This change should only affect user code that subclasses any of these handlers or calls their methods directly. ( #6778 , #7164 )
scrapy.pipelines.media.MediaPipeline.process_item() now returns a coroutine, previously it returned a Deferred object. This change only impacts code that calls this method directly. ( #7177 )
The from_settings() method of the following components, deprecated in Scrapy 2.12.0, is removed. You should use from_crawler() instead.
scrapy.dupefilters.RFPDupeFilter
scrapy.mail.MailSender
scrapy.middleware.MiddlewareManager
scrapy.core.downloader.contextfactory.ScrapyClientContextFactory
scrapy.pipelines.files.FilesPipeline
scrapy.pipelines.images.ImagesPipeline
( #7126 )
Scrapy no longer calls from_settings() methods of 3rd-party components , deprecated in Scrapy 2.12.0. You should define a from_crawler() method instead. ( #7126 )
The initialization flow of scrapy.pipelines.media.MediaPipeline and its subclasses was simplified, it now mandates from_crawler() methods and crawler arguments of init() methods. Not using these was deprecated in Scrapy 2.12.0. ( #7126 )
The REQUEST_FINGERPRINTER_IMPLEMENTATION setting, deprecated in Scrapy 2.12.0, is removed. ( #7126 )
The scrapy.utils.misc.create_instance() function, deprecated in Scrapy 2.12.0, is removed. Use scrapy.utils.misc.build_from_crawler() instead. ( #7126 )
The scrapy.core.downloader.Downloader._get_slot_key() function, deprecated in Scrapy 2.12.0, is removed. Use scrapy.core.downloader.Downloader.get_slot_key() instead. ( #7126 )
The scrapy.twisted_version attribute, deprecated in Scrapy 2.12.0, is removed. You should instead use the twisted.version attribute directly. ( #7126 )
The following utility functions, deprecated in Scrapy 2.12.0, are removed:
scrapy.utils.defer.process_chain_both()
scrapy.utils.python.equal_attributes()
scrapy.utils.python.flatten()
scrapy.utils.python.iflatten()
scrapy.utils.request.request_authenticate()
scrapy.utils.test.assert_samelines()
( #7126 )
scrapy.utils.serialize.ScrapyJSONDecoder , deprecated in Scrapy 2.12.0, is removed. ( #7126 )
The scrapy.extensions.feedexport.build_storage() function, deprecated in Scrapy 2.12.0, is removed, you can instead call the builder callable directly. ( #7126 )
scrapy.spidermiddlewares.offsite.OffsiteMiddleware , deprecated in Scrapy 2.11.2, is removed. scrapy.downloadermiddlewares.offsite.OffsiteMiddleware should be used instead. ( #6926 )
The following methods that return a Deferred are deprecated in favor of their coroutine-based replacements:
scrapy.core.downloader.handlers.DownloadHandlers
download_request() (use download_request_async() )
scrapy.core.downloader.middleware.DownloaderMiddlewareManager
download() (use download_async() )
scrapy.core.engine.ExecutionEngine
start() (use start_async() )
stop() (use stop_async() )
close() (use close_async() )
open_spider() (use open_spider_async() )
close_spider() (use close_spider_async() )
download() (use download_async() )
scrapy.core.scraper.Scraper
open_spider() (use open_spider_async() )
call_spider() (use call_spider_async() )
close_spider() (use close_spider_async() )
handle_spider_output() (use handle_spider_output_async() )
start_itemproc() (use start_itemproc_async() )
scrapy.core.spidermw.SpiderMiddlewareManager
scrape_response() (use scrape_response_async() )
scrapy.crawler.Crawler
stop() (use stop_async() )
scrapy.pipelines.ItemPipelineManager
process_item() (use process_item_async() )
open_spider() (use open_spider_async() )
close_spider() (use close_spider_async() )
scrapy.signalmanager.SignalManager
send_catch_log_deferred() (use send_catch_log_async() )
scrapy.utils.signal.send_catch_log_deferred() (use scrapy.utils.signal.send_catch_log_async() )
( #6791 , #6842 , #6979 , #6997 , #6999 , #7005 , #7043 , #7069 , #7161 , #7164 )
The following spider attributes are deprecated in favor of settings:
download_maxsize (use DOWNLOAD_MAXSIZE )
download_timeout (use DOWNLOAD_TIMEOUT )
download_warnsize (use DOWNLOAD_WARNSIZE )
max_concurrent_requests (use CONCURRENT_REQUESTS_PER_DOMAIN )
user_agent (use USER_AGENT )
( #6988 , #6994 , #7038 , #7039 , #7117 , #7176 )
Returning a Deferred from the following user-defined functions is deprecated in favor of defining them as coroutine functions:
spider callbacks and errbacks (which was never officially supported and may work incorrectly)
the process_request() , process_response() and process_exception() methods of custom downloader middlewares
the process_item() , open_spider() and close_spider() methods of custom pipelines
signal handlers
the download_request() and close() methods of custom download handlers
( #6718 , #6778 , #7069 , #7147 , #7148 , #7149 , #7150 , #7151 , #7161 , #7164 , #7179 )
Passing a spider argument to the following methods is deprecated:
scrapy.core.spidermw.SpiderMiddlewareManager.process_start()
scrapy.core.downloader.Downloader.fetch()
scrapy.core.downloader.Downloader._get_slot()
scrapy.core.downloader.handlers.DownloadHandlers.download_request()
all public methods of scrapy.statscollectors.StatsCollector
scrapy.spidermiddlewares.base.BaseSpiderMiddleware.process_spider_output()
scrapy.spidermiddlewares.base.BaseSpiderMiddleware.process_spider_output_async()
all process_*() methods of built-in downloader middlewares
all process_*() methods of built-in spider middlewares
scrapy.pipelines.media.MediaPipeline.open_spider()
scrapy.pipelines.media.MediaPipeline.process_item()
( #6750 , #6927 , #6984 , #7006 , #7011 , #7033 , #7037 , #7045 , #7178 )
Instantiating subclasses of scrapy.middleware.MiddlewareManager without a Crawler instance is deprecated. ( #6984 )
For the following user-defined functions and methods requiring a spider argument is deprecated, if you need a Spider instance inside them you should get it from the Crawler instance (you may need to refactor your code to save that instance in e.g. the from_crawler() method):
the process_request() , process_response() and process_exception() methods of custom downloader middlewares
the process_spider_input() , process_spider_output() , process_spider_output_async() and process_spider_exception() methods of custom spider middlewares
the process_item() method of custom pipelines
the fetch() method of a custom DOWNLOADER
( #6927 , #6984 , #7006 , #7037 )
The following things in custom download handlers are deprecated:
not having a lazy attribute (you should define it as True if you want to keep the current behavior)
returning a Deferred from the download_request() method (you should refactor it to return a coroutine; you also need to remove the spider argument when doing this)
not having a close() method, having a synchronous one or one that returns a Deferred (you should refactor it to return a coroutine or add an empty one if you don’t have it)
( #6778 , #7164 )
Custom implementations of ITEM_PROCESSOR should now define process_item_async() , open_spider_async() and close_spider_async() methods instead of, or in addition to, process_item() , open_spider() and close_spider() . ( #7005 , #7043 )
The CONCURRENT_REQUESTS_PER_IP setting is deprecated, use CONCURRENT_REQUESTS_PER_DOMAIN instead. ( #6917 , #6921 )
The scrapy.core.downloader.handlers.http module is deprecated. You should import scrapy.core.downloader.handlers.http11.HTTP11DownloadHandler directly instead of importing the scrapy.core.downloader.handlers.http.HTTPDownloadHandler alias. ( #7079 )
The scrapy.utils.decorators.defers() decorator is deprecated, you can use twisted.internet.defer.maybeDeferred() directly or reimplement this decorator in your code. ( #7164 )
scrapy.spiders.CrawlSpider._parse_response() is deprecated, use scrapy.spiders.CrawlSpider.parse_with_rules() instead. ( #4463 , #6804 )
The functions that add a delay to a Deferred are deprecated, their underlying Twisted functions can be used instead, either directly if a delay isn’t needed, or with some explicit way to add a delay if it’s needed:
scrapy.utils.defer.mustbe_deferred() (you can use twisted.internet.defer.maybeDeferred() )
scrapy.utils.defer.defer_succeed() (you can use twisted.internet.defer.succeed() )
scrapy.utils.defer.defer_fail() (you can use twisted.internet.defer.fail() )
scrapy.utils.defer.defer_result() (you can use twisted.internet.defer.succeed() and twisted.internet.defer.fail() )
( #6937 )
Added scrapy.crawler.AsyncCrawlerProcess and scrapy.crawler.AsyncCrawlerRunner as counterparts to CrawlerProcess and CrawlerRunner that offer coroutine-based APIs. ( #6789 , #6790 , #6796 , #6817 , #6845 , #7034 )
Added coroutine counterparts to some of the Deferred-based APIs:
scrapy.core.downloader.handlers.DownloadHandlers
download_request_async() (to download_request() )
scrapy.core.downloader.middleware.DownloaderMiddlewareManager
download_async() (to download() )
scrapy.core.engine.ExecutionEngine
start_async() (to start() )
stop_async() (to stop() )
close_async() (to close() )
open_spider_async() (to open_spider() )
close_spider_async() (to close_spider() )
download_async() (to download() )
scrapy.core.scraper.Scraper
open_spider_async() (to open_spider() )
close_spider_async() (to close_spider() )
start_itemproc_async() (to start_itemproc() )
scrapy.crawler.Crawler
crawl_async() (to crawl() )
stop_async() (to stop() )
scrapy.pipelines.ItemPipelineManager
process_item_async() (to process_item() )
open_spider_async() (to open_spider() )
close_spider_async() (to close_spider() )
scrapy.signalmanager.SignalManager
send_catch_log_async() (to send_catch_log_deferred() )
( #6781 , #6791 , #6792 , #6795 , #6801 , #6817 , #6842 , #6997 , #7005 , #7043 , #7069 ,:gh: 7164 , #7202 )
The default value of the SCHEDULER_PRIORITY_QUEUE setting is now 'scrapy.pqueues.DownloaderAwarePriorityQueue' . ( #6924 , #6940 )
Added scrapy.extensions.logcount.LogCount , an enabled-by-default extension that is responsible for the log_count/* stats. Previously, this code was in scrapy.crawler.Crawler and couldn’t be disabled. ( #7046 )
Added scrapy.spiders.CrawlSpider.parse_with_rules() as a public replacement for _parse_response() . ( #4463 , #6804 )
Added scrapy.utils.asyncio.is_asyncio_available() as an alternative to scrapy.utils.reactor.is_asyncio_reactor_installed() with a future-proof name and semantics. ( #6827 )
The API for download handlers , previously undocumented, has been modernized and documented. An optional base class, scrapy.core.downloader.handlers.base.BaseDownloadHandler , has been added to simplify writing custom download handlers that conform to the current API. ( #4944 , #6778 , #7164 )
Added scrapy.utils.defer.ensure_awaitable() , which can be helpful to call user-defined functions that can return coroutines, Deferreds or values directly. ( #7005 )
The requests.seen file, written by RFPDupeFilter when job persistence is enabled, now uses line buffering to reduce data loss in spider crashes. ( #6019 , #7094 )
Images downloaded by ImagesPipeline are now automatically transposed based on EXIF data. ( #6525 , #6975 )
Refactored internal functions to use coroutines instead of Deferreds. ( #6795 , #6852 , #6855 , #6858 , #7159 )
Commands that don’t need a CrawlerProcess instance no longer create it. ( #6824 )
Improved shell help formatting when using IPython 9+. ( #6915 , #6980 )
Setting FILES_STORE or IMAGES_STORE to None now correctly disables the respective pipeline. ( #6964 , #6969 )
MetaRefreshMiddleware now uses the URL set in the <base> tag as the base URL when redirecting to a relative URL. ( #7042 , #7047 )
Passing None as a value of the download_slot request meta key is now handled in the same way as not setting this meta key at all. ( #7172 )
Fixed parsing of the first line of robots.txt files that have a BOM. ( #6195 , #7095 )
Added documentation about download handlers, their API and built-in handlers. ( #4944 , #7164 )
Added a section about the scrapy-spider-metadata library to the spider argument docs . ( #6676 , #6957 , #7116 )
Improved the docs about coroutine-based and Deferred-based APIs. ( #6800 , #7146 )
Other documentation improvements and fixes. ( #7058 , #7076 , #7109 , #7195 , #7198 )
Switched from twisted.trial to pytest-twisted and replaced remaining unittest and twisted.trial features with pytest ones. ( #6658 , #6873 , #6884 , #6938 )
Enabled fancy pytest asserts. ( #6888 )
Added Sphinx Lint to the pre-commit configuration. ( #6920 )
CI and test improvements and fixes. ( #6649 , #6769 , #6821 , #6835 , #6836 , #6846 , #6883 , #6885 , #6889 , #6905 , #6928 , #6933 , #6941 , #6942 , #6945 , #6947 , #6960 , #6968 , #6972 , #6974 , #6996 , #7003 , #7012 , #7013 , #7050 , #7059 , #7070 , #7073 , #7118 , #7127 , #7141 , #7143 , #7145 , #7173 )
Code cleanups. ( #6803 , #6838 , #6849 , #6875 , #6876 , #6892 , #6930 , #6949 , #6970 , #6977 , #6986 , #7008 , #7177 )
Fix for the CVE-2025-6176 security issue: improved protection against decompression bombs in HttpCompressionMiddleware for responses compressed using…
Fix for the CVE-2025-6176 security issue: improved protection against decompression bombs in HttpCompressionMiddleware for responses compressed using the br and deflate methods. Requires brotli >= 1.2.0.
Improved protection against decompression bombs in HttpCompressionMiddleware for responses compressed using the br and deflate methods: if a single compressed chunk would be larger than the response size limit (see DOWNLOAD_MAXSIZE ) when decompressed, decompression is no longer carried out. This is especially important for the br (Brotli) method that can provide a very high compression ratio. Please, see the CVE-2025-6176 and GHSA-2qfp-q593-8484 security advisories for more information. ( #7134 )
The minimum supported version of the optional brotli package is now 1.2.0 . ( #7134 )
The brotlicffi and brotlipy packages can no longer be used to decompress Brotli-compressed responses. Please install the brotli package instead. ( #7134 )
Restricted the maximum supported Twisted version to 25.5.0 , as Scrapy currently uses some private APIs changed in later Twisted versions. ( #7142 )
Stopped setting the COVERAGE_CORE environment variable in tests, it didn’t have an effect but caused the coverage module to produce a warning or an error. ( #7137 )
Removed the documentation build dependency on the deprecated sphinx-hoverxref module. ( #6786 , #6922 )
Changed the values for DOWNLOAD_DELAY (from 0 to 1) and CONCURRENT_REQUESTS_PER_DOMAIN (from 8 to 1) in the default project template.
DOWNLOAD_DELAY (from 0 to 1) and CONCURRENT_REQUESTS_PER_DOMAIN (from 8 to 1) in the default project template.Changed the values for DOWNLOAD_DELAY (from 0 to 1 ) and CONCURRENT_REQUESTS_PER_DOMAIN (from 8 to 1 ) in the default project template. ( #6597 , #6918 , #6923 )
Improved scrapy.core.engine.ExecutionEngine logic related to initialization and exception handling, fixing several cases where the spider would crash, hang or log an unhandled exception. ( #6783 , #6784 , #6900 , #6908 , #6910 , #6911 )
Fixed a Windows issue with feed exports using scrapy.extensions.feedexport.FileFeedStorage that caused the file to be created on the wrong drive. ( #6894 , #6897 )
Allowed running tests with Twisted 25.5.0+ again. Pytest 8.4.1+ is now required for running tests in non-pinned envs as support for the new Twisted version was added in that version. ( #6893 )
Fixed running tests with lxml 6.0.0+. ( #6919 )
Added a deprecation notice for scrapy.spidermiddlewares.offsite.OffsiteMiddleware to the Scrapy 2.11.2 release notes . ( #6926 )
Updated contribution docs to refer to ruff instead of black . ( #6903 )
Added .venv/ and .vscode/ to .gitignore . ( #6901 , #6907 )
Fixed a bug introduced in Scrapy 2.13.0 that caused results of request errbacks to be ignored when the errback was called because of a downloader erro
Fixed a bug introduced in Scrapy 2.13.0 that caused results of request errbacks to be ignored when the errback was called because of a downloader error. ( #6861 , #6863 )
Added a note about the behavior change of scrapy.utils.reactor.is_asyncio_reactor_installed() to its docs and to the “Backward-incompatible changes” section of the Scrapy 2.13.0 release notes . ( #6866 )
Improved the message in the exception raised by scrapy.utils.test.get_reactor_settings() when there is no reactor installed. ( #6866 )
Updated the scrapy.crawler.CrawlerRunner examples in Common Practices to install the reactor explicitly, to fix reactor-related errors with Scrapy 2.13.0 and later. ( #6865 )
Fixed scrapy fetch not working with scrapy-poet . ( #6872 )
Fixed an exception produced by scrapy.core.engine.ExecutionEngine when it’s closed before being fully initialized. ( #6857 , #6867 )
Improved the README, updated the Scrapy logo in it. ( #6831 , #6833 , #6839 )
Restricted the Twisted version used in tests to below 25.5.0, as some tests fail with 25.5.0. ( #6878 , #6882 )
Updated type hints for Twisted 25.5.0 changes. ( #6882 )
Removed the old artwork. ( #6874 )
Give callback requests precedence over start requests when priority values are the same.
Give callback requests precedence over start requests when priority values are the same.
This makes changes from 2.13.0 to start request handling more intuitive and backward compatible. For scenarios where all requests have the same priorities, in 2.13.0 all start requests were sent before the first callback request. In 2.13.1, same as in 2.12 and lower, start requests are only sent when there are not enough pending callback requests to reach concurrency limits.
( #6828 )
Added a deepwiki badge to the README. ( #6793 )
Fixed a typo in the code example of Delaying start request iteration . ( #6812 , #6815 )
Fixed a typo in the Supported callables section of the documentation. ( #6822 )
Made this page more prominently listed in PyPI project links. ( #6826 )
Give callback requests precedence over start requests when priority values are the same.
This makes changes from 2.13.0 to start request handling more intuitive and backward compatible. For scenarios where all requests have the same priorities, in 2.13.0 all start requests were sent before the first callback request. In 2.13.1, same as in 2.12 and lower, start requests are only sent when there are not enough pending callback requests to reach concurrency limits.
Added a deepwiki badge to the README. (6793)
Fixed a typo in the code example of start-requests-lazy. (6812, 6815)
Fixed a typo in the coroutine-support section of the documentation. (6822)
Made this page more prominently listed in PyPI project links. (6826)
Spider middlewares that don't support asynchronous spider output are deprecated
start_requests() (sync) with start() (async) and changed how it is iterated.allow_offsite request meta keyHighlights:
The asyncio reactor is now enabled by default
Replaced start_requests() (sync) with start() (async) and changed how it is iterated
Added the allow_offsite request meta key
Spider middlewares that don’t support asynchronous spider output are deprecated
Added a base class for universal spider middlewares
Dropped support for PyPy 3.9. ( #6613 )
Added support for PyPy 3.11. ( #6697 )
The default value of the TWISTED_REACTOR setting was changed from None to "twisted.internet.asyncioreactor.AsyncioSelectorReactor" . This value was used in newly generated projects since Scrapy 2.7.0 but now existing projects that don’t explicitly set this setting will also use the asyncio reactor. You can change this setting in your project to use a different reactor. ( #6659 , #6713 )
The iteration of start requests and items no longer stops once there are requests in the scheduler, and instead runs continuously until all start requests have been scheduled.
To reproduce the previous behavior, see Delaying start request iteration . ( #6729 )
An unhandled exception from the open_spider() method of a spider middleware no longer stops the crawl. ( #6729 )
In scrapy.core.engine.ExecutionEngine :
The second parameter of open_spider() , start_requests , has been removed. The start requests are determined by the spider parameter instead (see start() ).
The slot attribute has been renamed to _slot and should not be used.
( #6729 )
In scrapy.core.engine , the Slot class has been renamed to _Slot and should not be used. ( #6729 )
The slot telnet variable has been removed. ( #6729 )
In scrapy.core.spidermw.SpiderMiddlewareManager , process_start_requests() has been replaced by process_start() . ( #6729 )
The scrape_func callable passed to scrapy.core.spidermw.SpiderMiddlewareManager.scrape_response() is now called with 2 parameters, response and request , instead of 3, and must return a Deferred instead of an iterable. ( #6787 )
The now-deprecated start_requests() method, when it returns an iterable instead of being defined as a generator, is now executed after the scheduler instance has been created. ( #6729 )
When using JOBDIR , start requests are now serialized into their own, s -suffixed priority folders. You can set SCHEDULER_START_DISK_QUEUE to None or "" to change that, but the side effects may be undesirable. See SCHEDULER_START_DISK_QUEUE for details. ( #6729 )
The URL length limit, set by the URLLENGTH_LIMIT setting, is now also enforced for start requests. ( #6777 )
Calling scrapy.utils.reactor.is_asyncio_reactor_installed() without an installed reactor now raises an exception instead of installing a reactor. This shouldn’t affect normal Scrapy use cases, but it may affect 3rd-party test suites that use Scrapy internals such as Crawler and don’t install a reactor explicitly. If you are affected by this change, you most likely need to install the reactor before running Scrapy code that expects it to be installed. ( #6732 , #6735 )
The from_settings() method of UrlLengthMiddleware , deprecated in Scrapy 2.12.0, is removed earlier than the usual deprecation period (this was needed because after the introduction of the BaseSpiderMiddleware base class and switching built-in spider middlewares to it those middlewares need the Crawler instance at run time). Please use from_crawler() instead. ( #6693 )
scrapy.utils.url.escape_ajax() is no longer called when a Request instance is created. It was only useful for websites supporting the escaped_fragment feature which most modern websites don’t support. If you still need this you can modify the URLs before passing them to Request . ( #6523 , #6651 )
Removed old deprecated name aliases for some signals:
stats_spider_opened (use spider_opened instead)
stats_spider_closing and stats_spider_closed (use spider_closed instead)
item_passed (use item_scraped instead)
request_received (use request_scheduled instead)
( #6654 , #6655 )
The start_requests() method of Spider is deprecated, use start() instead, or both to maintain support for lower Scrapy versions. ( #456 , #3477 , #4467 , #5627 , #6729 )
The process_start_requests() method of spider middlewares is deprecated, use process_start() instead, or both to maintain support for lower Scrapy versions. ( #456 , #3477 , #4467 , #5627 , #6729 )
The init method of priority queue classes (see SCHEDULER_PRIORITY_QUEUE ) should now support a keyword-only start_queue_cls parameter. ( #6752 )
Spider middlewares that don’t support asynchronous spider output are deprecated. The async iterable downgrading feature, needed for using such middlewares with asynchronous callbacks and with other spider middlewares that produce asynchronous iterables, is also deprecated. Please update all such middlewares to support asynchronous spider output. ( #6664 )
Functions that were imported from w3lib.url and re-exported in scrapy.utils.url are now deprecated, you should import them from w3lib.url directly. They are:
scrapy.utils.url.add_or_replace_parameter()
scrapy.utils.url.add_or_replace_parameters()
scrapy.utils.url.any_to_uri()
scrapy.utils.url.canonicalize_url()
scrapy.utils.url.file_uri_to_path()
scrapy.utils.url.is_url()
scrapy.utils.url.parse_data_uri()
scrapy.utils.url.parse_url()
scrapy.utils.url.path_to_file_uri()
scrapy.utils.url.safe_download_url()
scrapy.utils.url.safe_url_string()
scrapy.utils.url.url_query_cleaner()
scrapy.utils.url.url_query_parameter()
( #4577 , #6583 , #6586 )
HTTP/1.0 support code is deprecated. It was disabled by default and couldn’t be used together with HTTP/1.1. If you still need it, you should write your own download handler or copy the code from Scrapy. The deprecations include:
scrapy.core.downloader.handlers.http10.HTTP10DownloadHandler
scrapy.core.downloader.webclient.ScrapyHTTPClientFactory
scrapy.core.downloader.webclient.ScrapyHTTPPageGetter
Overriding scrapy.core.downloader.contextfactory.ScrapyClientContextFactory.getContext()
( #6634 )
The following modules and functions used only in tests are deprecated:
the scrapy.utils.testproc module
the scrapy.utils.testsite module
scrapy.utils.test.assert_gcs_environ()
scrapy.utils.test.get_ftp_content_and_delete()
scrapy.utils.test.get_gcs_content_and_delete()
scrapy.utils.test.mock_google_cloud_storage()
scrapy.utils.test.skip_if_no_boto()
If you need to use them in your tests or code, you can copy the code from Scrapy. ( #6696 )
scrapy.utils.test.TestSpider is deprecated. If you need an empty spider class you can use scrapy.utils.spider.DefaultSpider or create your own subclass of scrapy.Spider . ( #6678 )
scrapy.downloadermiddlewares.ajaxcrawl.AjaxCrawlMiddleware is deprecated. It was disabled by default and isn’t useful for most of the existing websites. ( #6523 , #6651 , #6656 )
scrapy.utils.url.escape_ajax() is deprecated. ( #6523 , #6651 )
scrapy.spiders.init.InitSpider is deprecated. If you find it useful, you can copy its code from Scrapy. ( #6708 , #6714 )
scrapy.utils.versions.scrapy_components_versions() is deprecated, use scrapy.utils.versions.get_versions() instead. ( #6582 )
BaseDupeFilter.log() is deprecated. It does nothing and shouldn’t be called. ( #4151 )
Passing the spider argument to the following methods of Scraper is deprecated:
close_spider()
enqueue_scrape()
handle_spider_error()
handle_spider_output()
( #6764 )
You can now yield the start requests and items of a spider from the start() spider method and from the process_start() spider middleware method, both asynchronous generators .
This makes it possible to use asynchronous code to generate those start requests and items, e.g. reading them from a queue service or database using an asynchronous client, without workarounds. ( #456 , #3477 , #4467 , #5627 , #6729 )
Start requests are now scheduled as soon as possible.
As a result, their priority is now taken into account as soon as CONCURRENT_REQUESTS is reached. ( #456 , #3477 , #4467 , #5627 , #6729 )
Crawler.signals has a new wait_for() method. ( #6729 )
Added a new scheduler_empty signal. ( #6729 )
Added new settings: SCHEDULER_START_DISK_QUEUE and SCHEDULER_START_MEMORY_QUEUE . ( #6729 )
Added StartSpiderMiddleware , which sets is_start_request to True on start requests . ( #6729 )
Exposed a new method of Crawler.engine : needs_backout() . ( #6729 )
Added the allow_offsite request meta key that can be used instead of the more general dont_filter request attribute to skip processing of the request by OffsiteMiddleware (but not by other code that checks dont_filter ). ( #3690 , #6151 , #6366 )
Added an optional base class for spider middlewares, BaseSpiderMiddleware , which can be helpful for writing universal spider middlewares without boilerplate and code duplication. The built-in spider middlewares now inherit from this class. ( #6693 , #6777 )
Scrapy add-ons can now define a class method called update_pre_crawler_settings() to update pre-crawler settings . ( #6544 , #6568 )
Added helpers for modifying component priority dictionary settings. ( #6614 )
Responses that use an unknown/unsupported encoding now produce a warning. If Scrapy knows that installing an additional package (such as brotli ) will allow decoding the response, that will be mentioned in the warning. ( #4697 , #6618 )
Added the spider_exceptions/count stat which tracks the total count of exceptions (tracked also by per-type spider_exceptions/* stats). ( #6739 , #6740 )
Added the DEFAULT_DROPITEM_LOG_LEVEL setting and the scrapy.exceptions.DropItem.log_level attribute that allow customizing the log level of the message that is logged when an item is dropped. ( #6603 , #6608 )
Added support for the -b, --cookie curl argument to scrapy.Request.from_curl() . ( #6684 )
Added the LOG_VERSIONS setting that allows customizing the list of software whose versions are logged when the spider starts. ( #6582 )
Added the WARN_ON_GENERATOR_RETURN_VALUE setting that allows disabling run time analysis of callback code used to warn about incorrect return statements in generator-based callbacks. You may need to disable this setting if this analysis breaks on your callback code. ( #6731 , #6738 )
Removed or postponed some calls of itemadapter.is_item() to increase performance. ( #6719 )
Improved the error message when running a scrapy command that requires a project (such as scrapy crawl ) outside of a project directory. ( #2349 , #3426 )
Added an empty ADDONS setting to the settings.py template for new projects. ( #6587 )
Yielding an item from Spider.start or from SpiderMiddleware.process_start no longer delays the next iteration of starting requests and items by up to 5 seconds. ( #6729 )
Fixed calculation of items_per_minute and responses_per_minute stats. ( #6599 )
Fixed an error initializing scrapy.extensions.feedexport.GCSFeedStorage . ( #6617 , #6628 )
Fixed an error running scrapy bench . ( #6632 , #6633 )
Fixed duplicated log messages about the reactor and the event loop. ( #6636 , #6657 )
Fixed resolving type annotations of SitemapSpider._parse_sitemap() at run time, required by tools such as scrapy-poet . ( #6665 , #6671 )
Calling scrapy.utils.reactor.is_asyncio_reactor_installed() without an installed reactor now raises an exception instead of installing a reactor. ( #6732 , #6735 )
Restored support for the x-gzip content encoding. ( #6618 )
Documented the setting values set in the default project template. ( #6762 , #6775 )
Improved the docs about asynchronous iterable support in spider middlewares. ( #6688 )
Improved the docs about using Deferred -based APIs in coroutine-based code and included a list of such APIs. ( #6677 , #6734 , #6776 )
Improved the contribution docs . ( #6561 , #6575 )
Removed the Splash recommendation from the headless browser suggestion. We no longer recommend using Splash and recommend using other headless browser solutions instead. ( #6642 , #6701 )
Added the dark mode to the HTML documentation. ( #6653 )
Other documentation improvements and fixes. ( #4151 , #6526 , #6620 , #6621 , #6622 , #6623 , #6624 , #6721 , #6723 , #6780 )
Switched from setup.py to pyproject.toml . ( #6514 , #6547 )
Switched the build backend from setuptools to hatchling . ( #6771 )
Replaced most linters with ruff . ( #6565 , #6576 , #6577 , #6581 , #6584 , #6595 , #6601 , #6631 )
Improved accuracy and performance of collecting test coverage. ( #6255 , #6610 )
Fixed an error that prevented running tests from directories other than the top level source directory. ( #6567 )
Reduced the amount of mockserver calls in tests to improve the overall test run time. ( #6637 , #6648 )
Fixed tests that were running the same test code more than once. ( #6646 , #6647 , #6650 )
Refactored tests to use more pytest features instead of unittest ones where possible. ( #6678 , #6680 , #6695 , #6699 , #6700 , #6702 , #6709 , #6710 , #6711 , #6712 , #6725 )
Type hints improvements and fixes. ( #6578 , #6579 , #6593 , #6605 , #6694 )
CI and test improvements and fixes. ( #5360 , #6271 , #6547 , #6560 , #6602 , #6607 , #6609 , #6613 , #6619 , #6626 , #6679 , #6703 , #6704 , #6716 , #6720 , #6722 , #6724 , #6741 , #6743 , #6766 , #6770 , #6772 , #6773 )
Code cleanups. ( #6600 , #6606 , #6635 , #6764 )
Dropped support for Python 3.8, added support for Python 3.13
start_requests can now yield itemsscrapy.http.JsonResponseCLOSESPIDER_PAGECOUNT_NO_ITEM settingHighlights:
Dropped support for Python 3.8, added support for Python 3.13
scrapy.Spider.start_requests() can now yield items
Added JsonResponse
Added CLOSESPIDER_PAGECOUNT_NO_ITEM
Dropped support for Python 3.8. ( #6466 , #6472 )
Added support for Python 3.13. ( #6166 )
Minimum versions increased for these dependencies:
Twisted : 18.9.0 → 21.7.0
cryptography : 36.0.0 → 37.0.0
pyOpenSSL : 21.0.0 → 22.0.0
lxml : 4.4.1 → 4.6.0
Removed setuptools from the dependency list. ( #6487 )
User-defined cookies for HTTPS requests will have the secure flag set to True unless it’s set to False explicitly. This is important when these cookies are reused in HTTP requests, e.g. after a redirect to an HTTP URL. ( #6357 )
The Reppy-based robots.txt parser, scrapy.robotstxt.ReppyRobotParser , was removed, as it doesn’t support Python 3.9+. ( #5230 , #6099 , #6499 )
The initialization API of scrapy.pipelines.media.MediaPipeline and its subclasses was improved and it’s possible that some previously working usage scenarios will no longer work. It can only affect you if you define custom subclasses of MediaPipeline or create instances of these pipelines via from_settings() or init() calls instead of from_crawler() calls.
Previously, MediaPipeline.from_crawler() called the from_settings() method if it existed or the init() method otherwise, and then did some additional initialization using the crawler instance. If the from_settings() method existed (like in FilesPipeline ) it called init() to create the instance. It wasn’t possible to override from_crawler() without calling MediaPipeline.from_crawler() from it which, in turn, couldn’t be called in some cases (including subclasses of FilesPipeline ).
Now, in line with the general usage of from_crawler() and from_settings() and the deprecation of the latter the recommended initialization order is the following one:
All init() methods should take a crawler argument. If they also take a settings argument they should ignore it, using crawler.settings instead. When they call init() of the base class they should pass the crawler argument to it too.
A from_settings() method shouldn’t be defined. Class-specific initialization code should go into either an overridden from_crawler() method or into init() .
It’s now possible to override from_crawler() and it’s not necessary to call MediaPipeline.from_crawler() in it if other recommendations were followed.
If pipeline instances were created with from_settings() or init() calls (which wasn’t supported even before, as it missed important initialization code), they should now be created with from_crawler() calls.
( #6540 )
The response_body argument of ImagesPipeline.convert_image is now positional-only, as it was changed from optional to required. ( #6500 )
The convert argument of scrapy.utils.conf.build_component_list() is now positional-only, as the preceding argument ( custom ) was removed. ( #6500 )
The overwrite_output argument of scrapy.utils.conf.feed_process_params_from_cli() is now positional-only, as the preceding argument ( output_format ) was removed. ( #6500 )
Removed the scrapy.utils.request.request_fingerprint() function, deprecated in Scrapy 2.7.0. ( #6212 , #6213 )
Removed support for value "2.6" of setting REQUEST_FINGERPRINTER_IMPLEMENTATION , deprecated in Scrapy 2.7.0. ( #6212 , #6213 )
RFPDupeFilter subclasses now require supporting the fingerprinter parameter in their init method, introduced in Scrapy 2.7.0. ( #6102 , #6113 )
Removed the scrapy.downloadermiddlewares.decompression module, deprecated in Scrapy 2.7.0. ( #6100 , #6113 )
Removed the scrapy.utils.response.response_httprepr() function, deprecated in Scrapy 2.6.0. ( #6111 , #6116 )
Spiders with spider-level HTTP authentication, i.e. with the http_user or http_pass attributes, must now define http_auth_domain as well, which was introduced in Scrapy 2.5.1. ( #6103 , #6113 )
Media pipelines methods file_path() , file_downloaded() , get_images() , image_downloaded() , media_downloaded() , media_to_download() , and thumb_path() must now support an item parameter, added in Scrapy 2.4.0. ( #6107 , #6113 )
The init() and from_crawler() methods of feed storage backend classes must now support the keyword-only feed_options parameter, introduced in Scrapy 2.4.0. ( #6105 , #6113 )
Removed the scrapy.loader.common and scrapy.loader.processors modules, deprecated in Scrapy 2.3.0. ( #6106 , #6113 )
Removed the scrapy.utils.misc.extract_regex() function, deprecated in Scrapy 2.3.0. ( #6106 , #6113 )
Removed the scrapy.http.JSONRequest class, replaced with JsonRequest in Scrapy 1.8.0. ( #6110 , #6113 )
scrapy.utils.log.logformatter_adapter no longer supports missing args , level , or msg parameters, and no longer supports a format parameter, all scenarios that were deprecated in Scrapy 1.0.0. ( #6109 , #6116 )
A custom class assigned to the SPIDER_LOADER_CLASS setting that does not implement the ISpiderLoader interface will now raise a zope.interface.verify.DoesNotImplement exception at run time. Non-compliant classes have been triggering a deprecation warning since Scrapy 1.0.0. ( #6101 , #6113 )
Removed the --output-format / -t command line option, deprecated in Scrapy 2.1.0. -O <URI>:<FORMAT> should be used instead. ( #6500 )
Running crawl() more than once on the same Crawler instance, deprecated in Scrapy 2.11.0, now raises an exception. ( #6500 )
Subclassing HttpCompressionMiddleware without support for the crawler argument in init() and without a custom from_crawler() method, deprecated in Scrapy 2.5.0, is no longer allowed. ( #6500 )
Removed the EXCEPTIONS_TO_RETRY attribute of RetryMiddleware , deprecated in Scrapy 2.10.0. ( #6500 )
Removed support for S3 feed exports without the boto3 package installed, deprecated in Scrapy 2.10.0. ( #6500 )
Removed the scrapy.extensions.feedexport._FeedSlot class, deprecated in Scrapy 2.10.0. ( #6500 )
Removed the scrapy.pipelines.images.NoimagesDrop exception, deprecated in Scrapy 2.8.0. ( #6500 )
The response_body argument of ImagesPipeline.convert_image is now required, not passing it was deprecated in Scrapy 2.8.0. ( #6500 )
Removed the custom argument of scrapy.utils.conf.build_component_list() , deprecated in Scrapy 2.10.0. ( #6500 )
Removed the scrapy.utils.reactor.get_asyncio_event_loop_policy() function, deprecated in Scrapy 2.9.0. Use asyncio.get_event_loop() and related standard library functions instead. ( #6500 )
The from_settings() methods of the Scrapy components that have them are now deprecated. from_crawler() should now be used instead. Affected components:
scrapy.dupefilters.RFPDupeFilter
scrapy.mail.MailSender
scrapy.middleware.MiddlewareManager
scrapy.core.downloader.contextfactory.ScrapyClientContextFactory
scrapy.pipelines.files.FilesPipeline
scrapy.pipelines.images.ImagesPipeline
scrapy.spidermiddlewares.urllength.UrlLengthMiddleware
( #6540 )
It’s now deprecated to have a from_settings() method but no from_crawler() method in 3rd-party Scrapy components . You can define a simple from_crawler() method that calls cls.from_settings(crawler.settings) to fix this if you don’t want to refactor the code. Note that if you have a from_crawler() method Scrapy will not call the from_settings() method so the latter can be removed. ( #6540 )
The initialization API of scrapy.pipelines.media.MediaPipeline and its subclasses was improved and some old usage scenarios are now deprecated (see also the “Backward-incompatible changes” section). Specifically:
It’s deprecated to define an init() method that doesn’t take a crawler argument.
It’s deprecated to call an init() method without passing a crawler argument. If it’s passed, it’s also deprecated to pass a settings argument, which will be ignored anyway.
Calling from_settings() is deprecated, use from_crawler() instead.
Overriding from_settings() is deprecated, override from_crawler() instead.
( #6540 )
The REQUEST_FINGERPRINTER_IMPLEMENTATION setting is now deprecated. ( #6212 , #6213 )
The scrapy.utils.misc.create_instance() function is now deprecated, use scrapy.utils.misc.build_from_crawler() instead. ( #5523 , #5884 , #6162 , #6169 , #6540 )
scrapy.core.downloader.Downloader._get_slot_key() is deprecated, use scrapy.core.downloader.Downloader.get_slot_key() instead. ( #6340 , #6352 )
scrapy.utils.defer.process_chain_both() is now deprecated. ( #6397 )
scrapy.twisted_version is now deprecated, you should instead use twisted.version directly (but note that it’s an incremental.Version object, not a tuple). ( #6509 , #6512 )
scrapy.utils.python.flatten() and scrapy.utils.python.iflatten() are now deprecated. ( #6517 , #6519 )
scrapy.utils.python.equal_attributes() is now deprecated. ( #6517 , #6519 )
scrapy.utils.request.request_authenticate() is now deprecated, you should instead just set the Authorization header directly. ( #6517 , #6519 )
scrapy.utils.serialize.ScrapyJSONDecoder is now deprecated, it didn’t contain any code since Scrapy 1.0.0. ( #6517 , #6519 )
scrapy.utils.test.assert_samelines() is now deprecated. ( #6517 , #6519 )
scrapy.extensions.feedexport.build_storage() is now deprecated. You can instead call the builder callable directly. ( #6540 )
scrapy.utils.misc.md5sum() is now deprecated. ( #6264 )
scrapy.Spider.start_requests() can now yield items. ( #5289 , #6417 )
Note
Some spider middlewares may need to be updated for Scrapy 2.12 support before you can use them in combination with the ability to yield items from start_requests() .
Added a new Response subclass, JsonResponse , for responses with a JSON MIME type . ( #6069 , #6171 , #6174 )
The LogStats extension now adds items_per_minute and responses_per_minute to the stats when the spider closes. ( #4110 , #4111 )
Added CLOSESPIDER_PAGECOUNT_NO_ITEM which allows closing the spider if no items were scraped in a set amount of time. ( #6434 )
User-defined cookies can now include the secure field. ( #6357 )
Added component getters to Crawler : get_addon() , get_downloader_middleware() , get_extension() , get_item_pipeline() , get_spider_middleware() . ( #6181 )
Slot delay updates by the AutoThrottle extension based on response latencies can now be disabled for specific requests via the autothrottle_dont_adjust_delay meta key. ( #6246 , #6527 )
If SPIDER_LOADER_WARN_ONLY is set to True , SpiderLoader does not raise SyntaxError but emits a warning instead. ( #6483 , #6484 )
Added support for multiple-compressed responses (ones with several encodings in the Content-Encoding header). ( #5143 , #5964 , #6063 )
Added support for multiple standard values in REFERRER_POLICY . ( #6381 )
Added support for brotlicffi (previously named brotlipy ). brotli is still recommended but only brotlicffi works on PyPy. ( #6263 , #6269 )
Added MetadataContract that sets the request meta. ( #6468 , #6469 )
Extended the list of file extensions that LinkExtractor ignores by default. ( #6074 , #6125 )
scrapy.utils.httpobj.urlparse_cached() is now used in more places instead of urllib.parse.urlparse() . ( #6228 , #6229 )
MediaPipeline is now an abstract class and its methods that were expected to be overridden in subclasses are now abstract methods. ( #6365 , #6368 )
Fixed handling of invalid @ -prefixed lines in contract extraction. ( #6383 , #6388 )
Importing scrapy.extensions.telnet no longer installs the default reactor. ( #6432 )
Reduced log verbosity for dropped requests that was increased in 2.11.2. ( #6433 , #6475 )
Added SECURITY.md that documents the security policy. ( #5364 , #6051 )
Example code for running Scrapy from a script no longer imports twisted.internet.reactor at the top level, which caused problems with non-default reactors when this code was used unmodified. ( #6361 , #6374 )
Documented the SpiderState extension. ( #6278 , #6522 )
Other documentation improvements and fixes. ( #5920 , #6094 , #6177 , #6200 , #6207 , #6216 , #6223 , #6317 , #6328 , #6389 , #6394 , #6402 , #6411 , #6427 , #6429 , #6440 , #6448 , #6449 , #6462 , #6497 , #6506 , #6507 , #6524 )
Added py.typed , in line with PEP 561 . ( #6058 , #6059 )
Fully covered the code with type hints (except for the most complicated parts, mostly related to twisted.web.http and other Twisted parts without type hints). ( #5989 , #6097 , #6127 , #6129 , #6130 , #6133 , #6143 , #6191 , #6268 , #6274 , #6275 , #6276 , #6279 , #6325 , #6326 , #6333 , #6335 , #6336 , #6337 , #6341 , #6353 , #6356 , #6370 , #6371 , #6384 , #6385 , #6387 , #6391 , #6395 , #6414 , #6422 , #6460 , #6466 , #6472 , #6494 , #6498 , #6516 )
Improved Bandit checks. ( #6260 , #6264 , #6265 )
Added pyupgrade to the pre-commit configuration. ( #6392 )
Added flake8-bugbear , flake8-comprehensions , flake8-debugger , flake8-docstrings , flake8-string-format and flake8-type-checking to the pre-commit configuration. ( #6406 , #6413 )
CI and test improvements and fixes. ( #5285 , #5454 , #5997 , #6078 , #6084 , #6087 , #6132 , #6153 , #6154 , #6201 , #6231 , #6232 , #6235 , #6236 , #6242 , #6245 , #6253 , #6258 , #6259 , #6270 , #6272 , #6286 , #6290 , #6296 #6367 , #6372 , #6403 , #6416 , #6435 , #6489 , #6501 , #6504 , #6511 , #6543 , #6545 )
Code cleanups. ( #6196 , #6197 , #6198 , #6199 , #6254 , #6257 , #6285 , #6305 , #6343 , #6349 , #6386 , #6415 , #6463 , #6470 , #6499 , #6505 , #6510 , #6531 , #6542 )
Issue tracker improvements. ( #6066 )
Mostly bug fixes, including security bug fixes.
Mostly bug fixes, including security bug fixes.
Redirects to non-HTTP protocols are no longer followed. Please, see the 23j4-mw76-5v7h security advisory for more information. ( #457 )
The Authorization header is now dropped on redirects to a different scheme ( http:// or https:// ) or port, even if the domain is the same. Please, see the 4qqq-9vqf-3h3f security advisory for more information.
When using system proxy settings that are different for http:// and https:// , redirects to a different URL scheme will now also trigger the corresponding change in proxy settings for the redirected request. Please, see the jm3v-qxmh-hxwv security advisory for more information. ( #767 )
Spider.allowed_domains is now enforced for all requests, and not only requests from spider callbacks. ( #1042 , #2241 , #6358 )
xmliter_lxml() no longer resolves XML entities. ( #6265 )
defusedxml is now used to make scrapy.http.request.rpc.XmlRpcRequest more secure. ( #6250 , #6251 )
scrapy.spidermiddlewares.offsite.OffsiteMiddleware (a spider middleware) is now deprecated and not enabled by default. The new downloader middleware with the same functionality, scrapy.downloadermiddlewares.offsite.OffsiteMiddleware , is enabled instead. ( #2241 , #6358 )
Restored support for brotlipy , which had been dropped in Scrapy 2.11.1 in favor of brotli . ( #6261 )
Note
brotlipy is deprecated, both in Scrapy and upstream. Use brotli instead if you can.
Make METAREFRESH_IGNORE_TAGS ["noscript"] by default. This prevents MetaRefreshMiddleware from following redirects that would not be followed by web browsers with JavaScript enabled. ( #6342 , #6347 )
During feed export , do not close the underlying file from built-in post-processing plugins . ( #5932 , #6178 , #6239 )
LinkExtractor now properly applies the unique and canonicalize parameters. ( #3273 , #6221 )
Do not initialize the scheduler disk queue if JOBDIR is an empty string. ( #6121 , #6124 )
Fix Spider.logger not logging custom extra information. ( #6323 , #6324 )
robots.txt files with a non-UTF-8 encoding no longer prevent parsing the UTF-8-compatible (e.g. ASCII) parts of the document. ( #6292 , #6298 )
scrapy.http.cookies.WrappedRequest.get_header() no longer raises an exception if default is None . ( #6308 , #6310 )
Selector now uses scrapy.utils.response.get_base_url() to determine the base URL of a given Response . ( #6265 )
The media_to_download() method of media pipelines now logs exceptions before stripping them. ( #5067 , #5068 )
When passing a callback to the parse command, build the callback callable with the right signature. ( #6182 )
Add a FAQ entry about creating blank requests . ( #6203 , #6208 )
Document that scrapy.Selector.type can be "json" . ( #6328 , #6334 )
Make builds reproducible. ( #5019 , #6322 )
Packaging and test fixes. ( #6286 , #6290 , #6312 , #6316 , #6344 )
- Security bug fixes. - Support for Twisted >= 23.8.0. - Documentation improvements. See the full changelog.
Highlights:
Security bug fixes.
Support for Twisted >= 23.8.0.
Documentation improvements.
Addressed ReDoS vulnerabilities :
scrapy.utils.iterators.xmliter is now deprecated in favor of xmliter_lxml() , which XMLFeedSpider now uses.
To minimize the impact of this change on existing code, xmliter_lxml() now supports indicating the node namespace with a prefix in the node name, and big files with highly nested trees when using libxml2 2.7+.
Fixed regular expressions in the implementation of the open_in_browser() function.
Please, see the cc65-xxvf-f7r9 security advisory for more information.
DOWNLOAD_MAXSIZE and DOWNLOAD_WARNSIZE now also apply to the decompressed response body. Please, see the 7j7m-v7m3-jqm7 security advisory for more information.
Also in relation with the 7j7m-v7m3-jqm7 security advisory , the deprecated scrapy.downloadermiddlewares.decompression module has been removed.
The Authorization header is now dropped on redirects to a different domain. Please, see the cw9j-q3vf-hrrv security advisory for more information.
The Twisted dependency is no longer restricted to < 23.8.0. ( #6024 , #6064 , #6142 )
The OS signal handling code was refactored to no longer use private Twisted functions. ( #6024 , #6064 , #6112 )
Improved documentation for Crawler initialization changes made in the 2.11.0 release. ( #6057 , #6147 )
Extended documentation for Request.meta . ( #5565 )
Fixed the dont_merge_cookies documentation. ( #5936 , #6077 )
Added a link to Zyte’s export guides to the feed exports documentation. ( #6183 )
Added a missing note about backward-incompatible changes in PythonItemExporter to the 2.11.0 release notes. ( #6060 , #6081 )
Added a missing note about removing the deprecated scrapy.utils.boto.is_botocore() function to the 2.8.0 release notes. ( #6056 , #6061 )
Other documentation improvements. ( #6128 , #6144 , #6163 , #6190 , #6192 )
Added Python 3.12 to the CI configuration, re-enabled tests that were disabled when the pre-release support was added. ( #5985 , #6083 , #6098 )
Fixed a test issue on PyPy 7.3.14. ( #6204 , #6205 )
Highlights:
Security bug fixes.
Support for Twisted >= 23.8.0.
Documentation improvements.
Addressed ReDoS vulnerabilities:
scrapy.utils.iterators.xmliter is now deprecated in favor of ~scrapy.utils.iterators.xmliter_lxml, which ~scrapy.spiders.XMLFeedSpider now uses.
To minimize the impact of this change on existing code, ~scrapy.utils.iterators.xmliter_lxml now supports indicating the node namespace with a prefix in the node name, and big files with highly nested trees when using libxml2 2.7+.
Fixed regular expressions in the implementation of the ~scrapy.utils.response.open_in_browser function.
Please, see the cc65-xxvf-f7r9 security advisory for more information.
DOWNLOAD_MAXSIZE and DOWNLOAD_WARNSIZE now also apply to the decompressed response body. Please, see the 7j7m-v7m3-jqm7 security advisory for more information.
Also in relation with the 7j7m-v7m3-jqm7 security advisory, the deprecated scrapy.downloadermiddlewares.decompression module has been removed.
The Authorization header is now dropped on redirects to a different domain. Please, see the cw9j-q3vf-hrrv security advisory for more information.
The Twisted dependency is no longer restricted to < 23.8.0. (6024, 6064, 6142)
The OS signal handling code was refactored to no longer use private Twisted functions. (6024, 6064, 6112)
Improved documentation for ~scrapy.crawler.Crawler initialization changes made in the 2.11.0 release. (6057, 6147)
Extended documentation for .Request.meta. (5565)
Fixed the dont_merge_cookies documentation. (5936, 6077)
Added a link to Zyte's export guides to the feed exports documentation. (6183)
Added a missing note about backward-incompatible changes in ~scrapy.exporters.PythonItemExporter to the 2.11.0 release notes. (6060, 6081)
Added a missing note about removing the deprecated scrapy.utils.boto.is_botocore() function to the 2.8.0 release notes. (6056, 6061)
Other documentation improvements. (6128, 6144, 6163, 6190, 6192)
Added Python 3.12 to the CI configuration, re-enabled tests that were disabled when the pre-release support was added. (5985, 6083, 6098)
Fixed a test issue on PyPy 7.3.14. (6204, 6205)
Spiders can now modify settings in their from_crawler methods, e.g. based on spider arguments.
from_crawler methods, e.g. based on spider arguments.Highlights:
Spiders can now modify settings in their from_crawler() methods, e.g. based on spider arguments .
Periodic logging of stats.
Most of the initialization of scrapy.crawler.Crawler instances is now done in crawl() , so the state of instances before that method is called is now different compared to older Scrapy versions. We do not recommend using the Crawler instances before crawl() is called. ( #6038 )
scrapy.Spider.from_crawler() is now called before the initialization of various components previously initialized in scrapy.crawler.Crawler.init() and before the settings are finalized and frozen. This change was needed to allow changing the settings in scrapy.Spider.from_crawler() . If you want to access the final setting values and the initialized Crawler attributes in the spider code as early as possible you can do this in scrapy.Spider.start_requests() or in a handler of the engine_started signal. ( #6038 )
The TextResponse.json method now requires the response to be in a valid JSON encoding (UTF-8, UTF-16, or UTF-32). If you need to deal with JSON documents in an invalid encoding, use json.loads(response.text) instead. ( #6016 )
PythonItemExporter used the binary output by default but it no longer does. ( #6006 , #6007 )
Removed the binary export mode of PythonItemExporter , deprecated in Scrapy 1.1.0. ( #6006 , #6007 )
Note
If you are using this Scrapy version on Scrapy Cloud with a stack that includes an older Scrapy version and get a “TypeError: Unexpected options: binary” error, you may need to add scrapinghub-entrypoint-scrapy >= 0.14.1 to your project requirements or switch to a stack that includes Scrapy 2.11.
Removed the CrawlerRunner.spiders attribute, deprecated in Scrapy 1.0.0, use CrawlerRunner.spider_loader instead. ( #6010 )
The scrapy.utils.response.response_httprepr() function, deprecated in Scrapy 2.6.0, has now been removed. ( #6111 )
Running crawl() more than once on the same scrapy.crawler.Crawler instance is now deprecated. ( #1587 , #6040 )
Spiders can now modify settings in their from_crawler() method, e.g. based on spider arguments . ( #1305 , #1580 , #2392 , #3663 , #6038 )
Added the PeriodicLog extension which can be enabled to log stats and/or their differences periodically. ( #5926 )
Optimized the memory usage in TextResponse.json by removing unnecessary body decoding. ( #5968 , #6016 )
Links to .webp files are now ignored by link extractors . ( #6021 )
Fixed logging enabled add-ons. ( #6036 )
Fixed MailSender producing invalid message bodies when the charset argument is passed to send() . ( #5096 , #5118 )
Fixed an exception when accessing self.EXCEPTIONS_TO_RETRY from a subclass of RetryMiddleware . ( #6049 , #6050 )
scrapy.settings.BaseSettings.getdictorlist() , used to parse FEED_EXPORT_FIELDS , now handles tuple values. ( #6011 , #6013 )
Calls to datetime.utcnow() , no longer recommended to be used, have been replaced with calls to datetime.now() with a timezone. ( #6014 )
Updated a deprecated function call in a pipeline example. ( #6008 , #6009 )
Extended typing hints. ( #6003 , #6005 , #6031 , #6034 )
Pinned brotli to 1.0.9 for the PyPy tests as 1.1.0 breaks them. ( #6044 , #6045 )
Other CI and pre-commit improvements. ( #6002 , #6013 , #6046 )
Marked Twisted >= 23.8.0 as unsupported.
Marked Twisted >= 23.8.0 as unsupported.
Added Python 3.12 support, dropped Python 3.7 support.
Highlights:
Added Python 3.12 support, dropped Python 3.7 support.
The new add-ons framework simplifies configuring 3rd-party components that support it.
Exceptions to retry can now be configured.
Many fixes and improvements for feed exports.
Dropped support for Python 3.7. ( #5953 )
Added support for the upcoming Python 3.12. ( #5984 )
Minimum versions increased for these dependencies:
lxml : 4.3.0 → 4.4.1
cryptography : 3.4.6 → 36.0.0
pkg_resources is no longer used. ( #5956 , #5958 )
boto3 is now recommended instead of botocore for exporting to S3. ( #5833 ).
The value of the FEED_STORE_EMPTY setting is now True instead of False . In earlier Scrapy versions empty files were created even when this setting was False (which was a bug that is now fixed), so the new default should keep the old behavior. ( #872 , #5847 )
When a function is assigned to the FEED_URI_PARAMS setting, returning None or modifying the params input parameter, deprecated in Scrapy 2.6, is no longer supported. ( #5994 , #5996 )
The scrapy.utils.reqser module, deprecated in Scrapy 2.6, is removed. ( #5994 , #5996 )
The scrapy.squeues classes PickleFifoDiskQueueNonRequest , PickleLifoDiskQueueNonRequest , MarshalFifoDiskQueueNonRequest , and MarshalLifoDiskQueueNonRequest , deprecated in Scrapy 2.6, are removed. ( #5994 , #5996 )
The property open_spiders and the methods has_capacity and schedule of scrapy.core.engine.ExecutionEngine , deprecated in Scrapy 2.6, are removed. ( #5994 , #5998 )
Passing a spider argument to the spider_is_idle() , crawl() and download() methods of scrapy.core.engine.ExecutionEngine , deprecated in Scrapy 2.6, is no longer supported. ( #5994 , #5998 )
scrapy.utils.datatypes.CaselessDict is deprecated, use scrapy.utils.datatypes.CaseInsensitiveDict instead. ( #5146 )
Passing the custom argument to scrapy.utils.conf.build_component_list() is deprecated, it was used in the past to merge FOO and FOO_BASE setting values but now Scrapy uses scrapy.settings.BaseSettings.getwithbase() to do the same. Code that uses this argument and cannot be switched to getwithbase() can be switched to merging the values explicitly. ( #5726 , #5923 )
Added support for Scrapy add-ons . ( #5950 )
Added the RETRY_EXCEPTIONS setting that configures which exceptions will be retried by RetryMiddleware . ( #2701 , #5929 )
Added the possiiblity to close the spider if no items were produced in the specified time, configured by CLOSESPIDER_TIMEOUT_NO_ITEM . ( #5979 )
Added support for the AWS_REGION_NAME setting to feed exports. ( #5980 )
Added support for using pathlib.Path objects that refer to absolute Windows paths in the FEEDS setting. ( #5939 )
Fixed creating empty feeds even with FEED_STORE_EMPTY=False . ( #872 , #5847 )
Fixed using absolute Windows paths when specifying output files. ( #5969 , #5971 )
Fixed problems with uploading large files to S3 by switching to multipart uploads (requires boto3 ). ( #960 , #5735 , #5833 )
Fixed the JSON exporter writing extra commas when some exceptions occur. ( #3090 , #5952 )
Fixed the “read of closed file” error in the CSV exporter. ( #5043 , #5705 )
Fixed an error when a component added by the class object throws NotConfigured with a message. ( #5950 , #5992 )
Added the missing scrapy.settings.BaseSettings.pop() method. ( #5959 , #5960 , #5963 )
Added CaseInsensitiveDict as a replacement for CaselessDict that fixes some API inconsistencies. ( #5146 )
Documented scrapy.Spider.update_settings() . ( #5745 , #5846 )
Documented possible problems with early Twisted reactor installation and their solutions. ( #5981 , #6000 )
Added examples of making additional requests in callbacks. ( #5927 )
Improved the feed export docs. ( #5579 , #5931 )
Clarified the docs about request objects on redirection. ( #5707 , #5937 )
Added support for running tests against the installed Scrapy version. ( #4914 , #5949 )
Extended typing hints. ( #5925 , #5977 )
Fixed the test_utils_asyncio.AsyncioTest.test_set_asyncio_event_loop test. ( #5951 )
Fixed the test_feedexport.BatchDeliveriesTest.test_batch_path_differ test on Windows. ( #5847 )
Enabled CI runs for Python 3.11 on Windows. ( #5999 )
Simplified skipping tests that depend on uvloop . ( #5984 )
Fixed the extra-deps-pinned tox env. ( #5948 )
Implemented cleanups. ( #5965 , #5986 )
Compatibility with new cryptography and new parsel.
Highlights:
Per-domain download settings.
Compatibility with new cryptography and new parsel .
JMESPath selectors from the new parsel .
Bug fixes.
scrapy.extensions.feedexport._FeedSlot is renamed to scrapy.extensions.feedexport.FeedSlot and the old name is deprecated. ( #5876 )
Settings corresponding to DOWNLOAD_DELAY , CONCURRENT_REQUESTS_PER_DOMAIN and RANDOMIZE_DOWNLOAD_DELAY can now be set on a per-domain basis via the new DOWNLOAD_SLOTS setting. ( #5328 )
Added TextResponse.jmespath() , a shortcut for JMESPath selectors available since parsel 1.8.1. ( #5894 , #5915 )
Added feed_slot_closed and feed_exporter_closed signals. ( #5876 )
Added scrapy.utils.request.request_to_curl() , a function to produce a curl command from a Request object. ( #5892 )
Values of FILES_STORE and IMAGES_STORE can now be pathlib.Path instances. ( #5801 )
Fixed a warning with Parsel 1.8.1+. ( #5903 , #5918 )
Fixed an error when using feed postprocessing with S3 storage. ( #5500 , #5581 )
Added the missing scrapy.settings.BaseSettings.setdefault() method. ( #5811 , #5821 )
Fixed an error when using cryptography 40.0.0+ and DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING is enabled. ( #5857 , #5858 )
The checksums returned by FilesPipeline for files on Google Cloud Storage are no longer Base64-encoded. ( #5874 , #5891 )
scrapy.utils.request.request_from_curl() now supports $-prefixed string values for the curl --data-raw argument, which are produced by browsers for data that includes certain symbols. ( #5899 , #5901 )
The parse command now also works with async generator callbacks. ( #5819 , #5824 )
The genspider command now properly works with HTTPS URLs. ( #3553 , #5808 )
Improved handling of asyncio loops. ( #5831 , #5832 )
LinkExtractor now skips certain malformed URLs instead of raising an exception. ( #5881 )
scrapy.utils.python.get_func_args() now supports more types of callables. ( #5872 , #5885 )
Fixed an error when processing non-UTF8 values of Content-Type headers. ( #5914 , #5917 )
Fixed an error breaking user handling of send failures in scrapy.mail.MailSender.send() . ( #1611 , #5880 )
Expanded contributing docs. ( #5109 , #5851 )
Added blacken-docs to pre-commit and reformatted the docs with it. ( #5813 , #5816 )
Fixed a JS issue. ( #5875 , #5877 )
Fixed make htmlview . ( #5878 , #5879 )
Fixed typos and other small errors. ( #5827 , #5839 , #5883 , #5890 , #5895 , #5904 )
Extended typing hints. ( #5805 , #5889 , #5896 )
Tests for most of the examples in the docs are now run as a part of CI, found problems were fixed. ( #5816 , #5826 , #5919 )
Removed usage of deprecated Python classes. ( #5849 )
Silenced include-ignored warnings from coverage. ( #5820 )
Fixed a random failure of the test_feedexport.test_batch_path_differ test. ( #5855 , #5898 )
Updated docstrings to match output produced by parsel 1.8.1 so that they don’t cause test failures. ( #5902 , #5919 )
Other CI and pre-commit improvements. ( #5802 , #5823 , #5908 )
This is a maintenance release, with minor features, bug fixes, and cleanups.
This is a maintenance release, with minor features, bug fixes, and cleanups.
This is a maintenance release, with minor features, bug fixes, and cleanups.
The scrapy.utils.gz.read1 function, deprecated in Scrapy 2.0, has now been removed. Use the read1() method of GzipFile instead. ( #5719 )
The scrapy.utils.python.to_native_str function, deprecated in Scrapy 2.0, has now been removed. Use scrapy.utils.python.to_unicode() instead. ( #5719 )
The scrapy.utils.python.MutableChain.next method, deprecated in Scrapy 2.0, has now been removed. Use next() instead. ( #5719 )
The scrapy.linkextractors.FilteringLinkExtractor class, deprecated in Scrapy 2.0, has now been removed. Use LinkExtractor instead. ( #5720 )
Support for using environment variables prefixed with SCRAPY_ to override settings, deprecated in Scrapy 2.0, has now been removed. ( #5724 )
Support for the noconnect query string argument in proxy URLs, deprecated in Scrapy 2.0, has now been removed. We expect proxies that used to need it to work fine without it. ( #5731 )
The scrapy.utils.python.retry_on_eintr function, deprecated in Scrapy 2.3, has now been removed. ( #5719 )
The scrapy.utils.python.WeakKeyCache class, deprecated in Scrapy 2.4, has now been removed. ( #5719 )
The scrapy.utils.boto.is_botocore() function, deprecated in Scrapy 2.4, has now been removed. ( #5719 )
scrapy.pipelines.images.NoimagesDrop is now deprecated. ( #5368 , #5489 )
ImagesPipeline.convert_image must now accept a response_body parameter. ( #3055 , #3689 , #4753 )
Applied black coding style to files generated with the genspider and startproject commands. ( #5809 , #5814 )
FEED_EXPORT_ENCODING is now set to "utf-8" in the settings.py file that the startproject command generates. With this value, JSON exports won’t force the use of escape sequences for non-ASCII characters. ( #5797 , #5800 )
The MemoryUsage extension now logs the peak memory usage during checks, and the binary unit MiB is now used to avoid confusion. ( #5717 , #5722 , #5727 )
The callback parameter of Request can now be set to scrapy.http.request.NO_CALLBACK() , to distinguish it from None , as the latter indicates that the default spider callback ( parse() ) is to be used. ( #5798 )
Enabled unsafe legacy SSL renegotiation to fix access to some outdated websites. ( #5491 , #5790 )
Fixed STARTTLS-based email delivery not working with Twisted 21.2.0 and better. ( #5386 , #5406 )
Fixed the finish_exporting() method of item exporters not being called for empty files. ( #5537 , #5758 )
Fixed HTTP/2 responses getting only the last value for a header when multiple headers with the same name are received. ( #5777 )
Fixed an exception raised by the shell command on some cases when using asyncio . ( #5740 , #5742 , #5748 , #5759 , #5760 , #5771 )
When using CrawlSpider , callback keyword arguments ( cb_kwargs ) added to a request in the process_request callback of a Rule will no longer be ignored. ( #5699 )
The images pipeline no longer re-encodes JPEG files. ( #3055 , #3689 , #4753 )
Fixed the handling of transparent WebP images by the images pipeline . ( #3072 , #5766 , #5767 )
scrapy.shell.inspect_response() no longer inhibits SIGINT (Ctrl+C). ( #2918 )
LinkExtractor with unique=False no longer filters out links that have identical URL and text. ( #3798 , #3799 , #4695 , #5458 )
RobotsTxtMiddleware now ignores URL protocols that do not support robots.txt ( data:// , file:// ). ( #5807 )
Silenced the filelock debug log messages introduced in Scrapy 2.6. ( #5753 , #5754 )
Fixed the output of scrapy -h showing an unintended commands line. ( #5709 , #5711 , #5712 )
Made the active project indication in the output of commands more clear. ( #5715 )
Documented how to debug spiders from Visual Studio Code . ( #5721 )
Documented how DOWNLOAD_DELAY affects per-domain concurrency. ( #5083 , #5540 )
Improved consistency. ( #5761 )
Fixed typos. ( #5714 , #5744 , #5764 )
Applied black coding style , sorted import statements, and introduced pre-commit . ( #4654 , #4658 , #5734 , #5737 , #5806 , #5810 )
Switched from os.path to pathlib . ( #4916 , #4497 , #5682 )
Addressed many issues reported by Pylint. ( #5677 )
Improved code readability. ( #5736 )
Improved package metadata. ( #5768 )
Removed direct invocations of setup.py . ( #5774 , #5776 )
Removed unnecessary OrderedDict usages. ( #5795 )
Removed unnecessary str definitions. ( #5150 )
Removed obsolete code and comments. ( #5725 , #5729 , #5730 , #5732 )
Fixed test and CI issues. ( #5749 , #5750 , #5756 , #5762 , #5765 , #5780 , #5781 , #5782 , #5783 , #5785 , #5786 )
Relaxed the restriction introduced in 2.6.2 so that the Proxy-Authentication header can again be set explicitly in certain cases, restoring compatibil
Proxy-Authentication header can again be set explicitly in certain cases, restoring compatibility with scrapy-zyte-smartproxy 2.1.0 and olderRelaxed the restriction introduced in 2.6.2 so that the Proxy-Authorization header can again be set explicitly, as long as the proxy URL in the proxy metadata has no other credentials, and for as long as that proxy URL remains the same; this restores compatibility with scrapy-zyte-smartproxy 2.1.0 and older ( #5626 ).
Using -O / --overwrite-output and -t / --output-format options together now produces an error instead of ignoring the former option ( #5516 , #5605 ).
Replaced deprecated asyncio APIs that implicitly use the current event loop with code that explicitly requests a loop from the event loop policy ( #5685 , #5689 ).
Fixed uses of deprecated Scrapy APIs in Scrapy itself ( #5588 , #5589 ).
Fixed uses of a deprecated Pillow API ( #5684 , #5692 ).
Improved code that checks if generators return values, so that it no longer fails on decorated methods and partial methods ( #5323 , #5592 , #5599 , #5691 ).
Upgraded the Code of Conduct to Contributor Covenant v2.1 ( #5698 ).
Fixed typos ( #5681 , #5694 ).
Re-enabled some erroneously disabled flake8 checks ( #5688 ).
Ignored harmless deprecation warnings from typing in tests ( #5686 , #5697 ).
Modernized our CI configuration ( #5695 , #5696 ).
Added Python 3.11 support, dropped Python 3.6 support
Highlights:
Added Python 3.11 support, dropped Python 3.6 support
Improved support for asynchronous callbacks
Asyncio support is enabled by default on new projects
Output names of item fields can now be arbitrary strings
Centralized request fingerprinting configuration is now possible
Python 3.7 or greater is now required; support for Python 3.6 has been dropped. Support for the upcoming Python 3.11 has been added.
The minimum required version of some dependencies has changed as well:
lxml : 3.5.0 → 4.3.0
Pillow ( images pipeline ): 4.0.0 → 7.1.0
zope.interface : 5.0.0 → 5.1.0
( #5512 , #5514 , #5524 , #5563 , #5664 , #5670 , #5678 )
ImagesPipeline.thumb_path must now accept an item parameter ( #5504 , #5508 ).
The scrapy.downloadermiddlewares.decompression module is now deprecated ( #5546 , #5547 ).
The process_spider_output() method of spider middlewares can now be defined as an asynchronous generator ( #4978 ).
The output of Request callbacks defined as coroutines is now processed asynchronously ( #4978 ).
CrawlSpider now supports asynchronous callbacks ( #5657 ).
New projects created with the startproject command have asyncio support enabled by default ( #5590 , #5679 ).
The FEED_EXPORT_FIELDS setting can now be defined as a dictionary to customize the output name of item fields, lifting the restriction that required output names to be valid Python identifiers, e.g. preventing them to have whitespace ( #1008 , #3266 , #3696 ).
You can now customize request fingerprinting through the new REQUEST_FINGERPRINTER_CLASS setting, instead of having to change it on every Scrapy component that relies on request fingerprinting ( #900 , #3420 , #4113 , #4762 , #4524 ).
jsonl is now supported and encouraged as a file extension for JSON Lines files ( #4848 ).
ImagesPipeline.thumb_path now receives the source item ( #5504 , #5508 ).
When using Google Cloud Storage with a media pipeline , FILES_EXPIRES now also works when FILES_STORE does not point at the root of your Google Cloud Storage bucket ( #5317 , #5318 ).
The parse command now supports asynchronous callbacks ( #5424 , #5577 ).
When using the parse command with a URL for which there is no available spider, an exception is no longer raised ( #3264 , #3265 , #5375 , #5376 , #5497 ).
TextResponse now gives higher priority to the byte order mark when determining the text encoding of the response body, following the HTML living standard ( #5601 , #5611 ).
MIME sniffing takes the response body into account in FTP and HTTP/1.0 requests, as well as in cached requests ( #4873 ).
MIME sniffing now detects valid HTML 5 documents even if the html tag is missing ( #4873 ).
An exception is now raised if ASYNCIO_EVENT_LOOP has a value that does not match the asyncio event loop actually installed ( #5529 ).
Fixed Headers.getlist() returning only the last header ( #5515 , #5526 ).
Fixed LinkExtractor not ignoring the tar.gz file extension by default ( #1837 , #2067 , #4066 )
Clarified the return type of Spider.parse ( #5602 , #5608 ).
To enable HttpCompressionMiddleware to do brotli compression , installing brotli is now recommended instead of installing brotlipy , as the former provides a more recent version of brotli.
Signal documentation now mentions coroutine support and uses it in code examples ( #4852 , #5358 ).
Avoiding getting banned now recommends Common Crawl instead of Google cache ( #3582 , #5432 ).
The new Components topic covers enforcing requirements on Scrapy components, like downloader middlewares , extensions , item pipelines , spider middlewares , and more; Enforcing asyncio as a requirement has also been added ( #4978 ).
Settings now indicates that setting values must be picklable ( #5607 , #5629 ).
Removed outdated documentation ( #5446 , #5373 , #5369 , #5370 , #5554 ).
Fixed typos ( #5442 , #5455 , #5457 , #5461 , #5538 , #5553 , #5558 , #5624 , #5631 ).
Fixed other issues ( #5283 , #5284 , #5559 , #5567 , #5648 , #5659 , #5665 ).
Added a continuous integration job to run twine check ( #5655 , #5656 ).
Addressed test issues and warnings ( #5560 , #5561 , #5612 , #5617 , #5639 , #5645 , #5662 , #5671 , #5675 ).
Cleaned up code ( #4991 , #4995 , #5451 , #5487 , #5542 , #5667 , #5668 , #5672 ).
Applied minor code improvements ( #5661 ).
Makes pip install Scrapy work again.
Makes pip install Scrapy work again.
It required making changes to support pyOpenSSL 22.1.0. We had to drop support for SSLv3 as a result.
We also upgraded the minimum versions of some dependencies.
See the changelog.
Added support for pyOpenSSL 22.1.0, removing support for SSLv3 ( #5634 , #5635 , #5636 ).
Upgraded the minimum versions of the following dependencies:
cryptography : 2.0 → 3.3
pyOpenSSL : 16.2.0 → 21.0.0
service_identity : 16.0.0 → 18.1.0
Twisted : 17.9.0 → 18.9.0
zope.interface : 4.1.3 → 5.0.0
( #5621 , #5632 )
Fixes test and documentation issues ( #5612 , #5617 , #5631 ).
Fixes a security issue around HTTP proxy usage, and addresses a few regressions introduced in Scrapy 2.6.0.
Fixes a security issue around HTTP proxy usage, and addresses a few regressions introduced in Scrapy 2.6.0.
See the changelog.
Security bug fix:
When HttpProxyMiddleware processes a request with proxy metadata, and that proxy metadata includes proxy credentials, HttpProxyMiddleware sets the Proxy-Authorization header, but only if that header is not already set.
There are third-party proxy-rotation downloader middlewares that set different proxy metadata every time they process a request.
Because of request retries and redirects, the same request can be processed by downloader middlewares more than once, including both HttpProxyMiddleware and any third-party proxy-rotation downloader middleware.
These third-party proxy-rotation downloader middlewares could change the proxy metadata of a request to a new value, but fail to remove the Proxy-Authorization header from the previous value of the proxy metadata, causing the credentials of one proxy to be sent to a different proxy.
To prevent the unintended leaking of proxy credentials, the behavior of HttpProxyMiddleware is now as follows when processing a request:
If the request being processed defines proxy metadata that includes credentials, the Proxy-Authorization header is always updated to feature those credentials.
If the request being processed defines proxy metadata without credentials, the Proxy-Authorization header is removed unless it was originally defined for the same proxy URL.
To remove proxy credentials while keeping the same proxy URL, remove the Proxy-Authorization header.
If the request has no proxy metadata, or that metadata is a falsy value (e.g. None ), the Proxy-Authorization header is removed.
It is no longer possible to set a proxy URL through the proxy metadata but set the credentials through the Proxy-Authorization header. Set proxy credentials through the proxy metadata instead.
Also fixes the following regressions introduced in 2.6.0:
CrawlerProcess supports again crawling multiple spiders ( #5435 , #5436 )
Installing a Twisted reactor before Scrapy does (e.g. importing twisted.internet.reactor somewhere at the module level) no longer prevents Scrapy from starting, as long as a different reactor is not specified in TWISTED_REACTOR ( #5525 , #5528 )
Fixed an exception that was being logged after the spider finished under certain conditions ( #5437 , #5440 )
The --output / -o command-line parameter supports again a value starting with a hyphen ( #5444 , #5445 )
The scrapy parse -h command no longer throws an error ( #5481 , #5482 )
Fixes a regression introduced in 2.6.0 that would unset the request method when following redirects.
Fixes a regression introduced in 2.6.0 that would unset the request method when following redirects.
Fixes a regression introduced in 2.6.0 that would unset the request method when following redirects.
Security fixes for cookie handling (see details below)
pathlib.Path output paths and per-feed item filtering and post-processingWhen a Request object with cookies defined gets a redirect response causing a new Request object to be scheduled, the cookies defined in the original Request object are no longer copied into the new Request object.
If you manually set the Cookie header on a Request object and the domain name of the redirect URL is not an exact match for the domain of the URL of the original Request object, your Cookie header is now dropped from the new Request object.
The old behavior could be exploited by an attacker to gain access to your cookies. Please, see the cjvr-mfj7-j4j8 security advisory for more information.
Note: It is still possible to enable the sharing of cookies between different domains with a shared domain suffix (e.g. example.com and any subdomain) by defining the shared domain suffix (e.g. example.com) as the cookie domain when defining your cookies. See the documentation of the Request class for more information.
When the domain of a cookie, either received in the Set-Cookie header of a response or defined in a Request object, is set to a public suffix <https://publicsuffix.org/>_, the cookie is now ignored unless the cookie domain is the same as the request domain.
The old behavior could be exploited by an attacker to inject cookies from a controlled domain into your cookiejar that could be sent to other domains not controlled by the attacker. Please, see the mfjm-vh54-3f96 security advisory for more information.
Highlights:
Security fixes for cookie handling
Python 3.10 support
asyncio support is no longer considered experimental, and works out-of-the-box on Windows regardless of your Python version
Feed exports now support pathlib.Path output paths and per-feed item filtering and post-processing
When a Request object with cookies defined gets a redirect response causing a new Request object to be scheduled, the cookies defined in the original Request object are no longer copied into the new Request object.
If you manually set the Cookie header on a Request object and the domain name of the redirect URL is not an exact match for the domain of the URL of the original Request object, your Cookie header is now dropped from the new Request object.
The old behavior could be exploited by an attacker to gain access to your cookies. Please, see the cjvr-mfj7-j4j8 security advisory for more information.
Note
It is still possible to enable the sharing of cookies between different domains with a shared domain suffix (e.g. example.com and any subdomain) by defining the shared domain suffix (e.g. example.com ) as the cookie domain when defining your cookies. See the documentation of the Request class for more information.
When the domain of a cookie, either received in the Set-Cookie header of a response or defined in a Request object, is set to a public suffix , the cookie is now ignored unless the cookie domain is the same as the request domain.
The old behavior could be exploited by an attacker to inject cookies from a controlled domain into your cookiejar that could be sent to other domains not controlled by the attacker. Please, see the mfjm-vh54-3f96 security advisory for more information.
The h2 dependency is now optional, only needed to enable HTTP/2 support . ( #5113 )
The formdata parameter of FormRequest , if specified for a non-POST request, now overrides the URL query string, instead of being appended to it. ( #2919 , #3579 )
When a function is assigned to the FEED_URI_PARAMS setting, now the return value of that function, and not the params input parameter, will determine the feed URI parameters, unless that return value is None . ( #4962 , #4966 )
In scrapy.core.engine.ExecutionEngine , methods crawl() , download() , schedule() , and spider_is_idle() now raise RuntimeError if called before open_spider() . ( #5090 )
These methods used to assume that ExecutionEngine.slot had been defined by a prior call to open_spider() , so they were raising AttributeError instead.
If the API of the configured scheduler does not meet expectations, TypeError is now raised at startup time. Before, other exceptions would be raised at run time. ( #3559 )
The _encoding field of serialized Request objects is now named encoding , in line with all other fields ( #5130 )
scrapy.http.TextResponse.body_as_unicode , deprecated in Scrapy 2.2, has now been removed. ( #5393 )
scrapy.item.BaseItem , deprecated in Scrapy 2.2, has now been removed. ( #5398 )
scrapy.item.DictItem , deprecated in Scrapy 1.8, has now been removed. ( #5398 )
scrapy.Spider.make_requests_from_url , deprecated in Scrapy 1.4, has now been removed. ( #4178 , #4356 )
When a function is assigned to the FEED_URI_PARAMS setting, returning None or modifying the params input parameter is now deprecated. Return a new dictionary instead. ( #4962 , #4966 )
scrapy.utils.reqser is deprecated. ( #5130 )
Instead of request_to_dict() , use the new Request.to_dict() method.
Instead of request_from_dict() , use the new scrapy.utils.request.request_from_dict() function.
In scrapy.squeues , the following queue classes are deprecated: PickleFifoDiskQueueNonRequest , PickleLifoDiskQueueNonRequest , MarshalFifoDiskQueueNonRequest , and MarshalLifoDiskQueueNonRequest . You should instead use: PickleFifoDiskQueue , PickleLifoDiskQueue , MarshalFifoDiskQueue , and MarshalLifoDiskQueue . ( #5117 )
Many aspects of scrapy.core.engine.ExecutionEngine that come from a time when this class could handle multiple Spider objects at a time have been deprecated. ( #5090 )
The has_capacity() method is deprecated.
The schedule() method is deprecated, use crawl() or download() instead.
The open_spiders attribute is deprecated, use spider instead.
The spider parameter is deprecated for the following methods:
spider_is_idle()
crawl()
download()
Instead, call open_spider() first to set the Spider object.
scrapy.utils.response.response_httprepr() is now deprecated. ( #4972 )
You can now use item filtering to control which items are exported to each output feed. ( #4575 , #5178 , #5161 , #5203 )
You can now apply post-processing to feeds, and built-in post-processing plugins are provided for output file compression. ( #2174 , #5168 , #5190 )
The FEEDS setting now supports pathlib.Path objects as keys. ( #5383 , #5384 )
Enabling asyncio while using Windows and Python 3.8 or later will automatically switch the asyncio event loop to one that allows Scrapy to work. See Windows-specific notes . ( #4976 , #5315 )
The genspider command now supports a start URL instead of a domain name. ( #4439 )
scrapy.utils.defer gained 2 new functions, deferred_to_future() and maybe_deferred_to_future() , to help await on Deferreds when using the asyncio reactor . ( #5288 )
Amazon S3 feed export storage gained support for temporary security credentials ( AWS_SESSION_TOKEN ) and endpoint customization ( AWS_ENDPOINT_URL ). ( #4998 , #5210 )
New LOG_FILE_APPEND setting to allow truncating the log file. ( #5279 )
Request.cookies values that are bool , float or int are cast to str . ( #5252 , #5253 )
You may now raise CloseSpider from a handler of the spider_idle signal to customize the reason why the spider is stopping. ( #5191 )
When using HttpProxyMiddleware , the proxy URL for non-HTTPS HTTP/1.1 requests no longer needs to include a URL scheme. ( #4505 , #4649 )
All built-in queues now expose a peek method that returns the next queue object (like pop ) but does not remove the returned object from the queue. ( #5112 )
If the underlying queue does not support peeking (e.g. because you are not using queuelib 1.6.1 or later), the peek method raises NotImplementedError .
Request and Response now have an attributes attribute that makes subclassing easier. For Request , it also allows subclasses to work with scrapy.utils.request.request_from_dict() . ( #1877 , #5130 , #5218 )
The open() and close() methods of the scheduler are now optional. ( #3559 )
HTTP/1.1 TunnelError exceptions now only truncate response bodies longer than 1000 characters, instead of those longer than 32 characters, making it easier to debug such errors. ( #4881 , #5007 )
ItemLoader now supports non-text responses. ( #5145 , #5269 )
The TWISTED_REACTOR and ASYNCIO_EVENT_LOOP settings are no longer ignored if defined in custom_settings . ( #4485 , #5352 )
Removed a module-level Twisted reactor import that could prevent using the asyncio reactor . ( #5357 )
The startproject command works with existing folders again. ( #4665 , #4676 )
The FEED_URI_PARAMS setting now behaves as documented. ( #4962 , #4966 )
Request.cb_kwargs once again allows the callback keyword. ( #5237 , #5251 , #5264 )
Made scrapy.utils.response.open_in_browser() support more complex HTML. ( #5319 , #5320 )
Fixed CSVFeedSpider.quotechar being interpreted as the CSV file encoding. ( #5391 , #5394 )
Added missing setuptools to the list of dependencies. ( #5122 )
LinkExtractor now also works as expected with links that have comma-separated rel attribute values including nofollow . ( #5225 )
Fixed a TypeError that could be raised during feed export parameter parsing. ( #5359 )
asyncio support is no longer considered experimental. ( #5332 )
Included Windows-specific help for asyncio usage . ( #4976 , #5315 )
Rewrote Using a headless browser with up-to-date best practices. ( #4484 , #4613 )
Documented local file naming in media pipelines . ( #5069 , #5152 )
Frequently Asked Questions now covers spider file name collision issues. ( #2680 , #3669 )
Provided better context and instructions to disable the URLLENGTH_LIMIT setting. ( #5135 , #5250 )
Documented that Reppy parser does not support Python 3.9+. ( #5226 , #5231 )
Documented the scheduler component . ( #3537 , #3559 )
Documented the method used by media pipelines to determine if a file has expired . ( #5120 , #5254 )
Running multiple spiders in the same process now features scrapy.utils.project.get_project_settings() usage. ( #5070 )
Running multiple spiders in the same process now covers what happens when you define different per-spider values for some settings that cannot differ at run time. ( #4485 , #5352 )
Extended the documentation of the StatsMailer extension. ( #5199 , #5217 )
Added JOBDIR to Settings . ( #5173 , #5224 )
Documented Spider.attribute . ( #5174 , #5244 )
Documented TextResponse.urljoin . ( #1582 )
Added the body_length parameter to the documented signature of the headers_received signal. ( #5270 )
Clarified SelectorList.get usage in the tutorial . ( #5256 )
The documentation now features the shortest import path of classes with multiple import paths. ( #2733 , #5099 )
quotes.toscrape.com references now use HTTPS instead of HTTP. ( #5395 , #5396 )
Added a link to our Discord server to Getting help . ( #5421 , #5422 )
The pronunciation of the project name is now officially /ˈskreɪpaɪ/. ( #5280 , #5281 )
Added the Scrapy logo to the README. ( #5255 , #5258 )
Fixed issues and implemented minor improvements. ( #3155 , #4335 , #5074 , #5098 , #5134 , #5180 , #5194 , #5239 , #5266 , #5271 , #5273 , #5274 , #5276 , #5347 , #5356 , #5414 , #5415 , #5416 , #5419 , #5420 )
Added support for Python 3.10. ( #5212 , #5221 , #5265 )
Significantly reduced memory usage by scrapy.utils.response.response_httprepr() , used by the DownloaderStats downloader middleware, which is enabled by default. ( #4964 , #4972 )
Removed uses of the deprecated optparse module. ( #5366 , #5374 )
Extended typing hints. ( #5077 , #5090 , #5100 , #5108 , #5171 , #5215 , #5334 )
Improved tests, fixed CI issues, removed unused code. ( #5094 , #5157 , #5162 , #5198 , #5207 , #5208 , #5229 , #5298 , #5299 , #5310 , #5316 , #5333 , #5388 , #5389 , #5400 , #5401 , #5404 , #5405 , #5407 , #5410 , #5412 , #5425 , #5427 )
Implemented improvements for contributors. ( #5080 , #5082 , #5177 , #5200 )
Implemented cleanups. ( #5095 , #5106 , #5209 , #5228 , #5235 , #5245 , #5246 , #5292 , #5314 , #5322 )
Highlights:
Security fixes for cookie handling
Python 3.10 support
asyncio support is no longer considered experimental, and works out-of-the-box on Windows regardless of your Python version
Feed exports now support pathlib.Path output paths and per-feed item filtering and post-processing
When a ~scrapy.Request object with cookies defined gets a redirect response causing a new ~scrapy.Request object to be scheduled, the cookies defined in the original ~scrapy.Request object are no longer copied into the new ~scrapy.Request object.
If you manually set the Cookie header on a ~scrapy.Request object and the domain name of the redirect URL is not an exact match for the domain of the URL of the original ~scrapy.Request object, your Cookie header is now dropped from the new ~scrapy.Request object.
The old behavior could be exploited by an attacker to gain access to your cookies. Please, see the cjvr-mfj7-j4j8 security advisory for more information.
Note
It is still possible to enable the sharing of cookies between different domains with a shared domain suffix (e.g. example.com and any subdomain) by defining the shared domain suffix (e.g. example.com) as the cookie domain when defining your cookies. See the documentation of the ~scrapy.Request class for more information.
When the domain of a cookie, either received in the Set-Cookie header of a response or defined in a ~scrapy.Request object, is set to a public suffix, the cookie is now ignored unless the cookie domain is the same as the request domain.
The old behavior could be exploited by an attacker to inject cookies from a controlled domain into your cookiejar that could be sent to other domains not controlled by the attacker. Please, see the mfjm-vh54-3f96 security advisory for more information.
The h2 dependency is now optional, only needed to enable HTTP/2 support. (5113)
The formdata parameter of ~scrapy.FormRequest, if specified for a non-POST request, now overrides the URL query string, instead of being appended to it. (2919, 3579)
When a function is assigned to the FEED_URI_PARAMS setting, now the return value of that function, and not the params input parameter, will determine the feed URI parameters, unless that return value is None. (4962, 4966)
In scrapy.core.engine.ExecutionEngine, methods ~scrapy.core.engine.ExecutionEngine.crawl, ~scrapy.core.engine.ExecutionEngine.download, ~scrapy.core.engine.ExecutionEngine.schedule, and ~scrapy.core.engine.ExecutionEngine.spider_is_idle now raise RuntimeError if called before ~scrapy.core.engine.ExecutionEngine.open_spider. (5090)
These methods used to assume that ExecutionEngine.slot had been defined by a prior call to ~scrapy.core.engine.ExecutionEngine.open_spider, so they were raising AttributeError instead.
If the API of the configured scheduler does not meet expectations, TypeError is now raised at startup time. Before, other exceptions would be raised at run time. (3559)
The _encoding field of serialized ~scrapy.Request objects is now named encoding, in line with all other fields (5130)
scrapy.http.TextResponse.body_as_unicode, deprecated in Scrapy 2.2, has now been removed. (5393)
scrapy.item.BaseItem, deprecated in Scrapy 2.2, has now been removed. (5398)
scrapy.item.DictItem, deprecated in Scrapy 1.8, has now been removed. (5398)
scrapy.Spider.make_requests_from_url, deprecated in Scrapy 1.4, has now been removed. (4178, 4356)
When a function is assigned to the FEED_URI_PARAMS setting, returning None or modifying the params input parameter is now deprecated. Return a new dictionary instead. (4962, 4966)
scrapy.utils.reqser is deprecated. (5130)
Instead of ~scrapy.utils.reqser.request_to_dict, use the new .Request.to_dict method.
Instead of ~scrapy.utils.reqser.request_from_dict, use the new scrapy.utils.request.request_from_dict function.
In scrapy.squeues, the following queue classes are deprecated: ~scrapy.squeues.PickleFifoDiskQueueNonRequest, ~scrapy.squeues.PickleLifoDiskQueueNonRequest, ~scrapy.squeues.MarshalFifoDiskQueueNonRequest, and ~scrapy.squeues.MarshalLifoDiskQueueNonRequest. You should instead use: ~scrapy.squeues.PickleFifoDiskQueue, ~scrapy.squeues.PickleLifoDiskQueue, ~scrapy.squeues.MarshalFifoDiskQueue, and ~scrapy.squeues.MarshalLifoDiskQueue. (5117)
Many aspects of scrapy.core.engine.ExecutionEngine that come from a time when this class could handle multiple ~scrapy.Spider objects at a time have been deprecated. (5090)
The ~scrapy.core.engine.ExecutionEngine.has_capacity method is deprecated.
The ~scrapy.core.engine.ExecutionEngine.schedule method is deprecated, use ~scrapy.core.engine.ExecutionEngine.crawl or ~scrapy.core.engine.ExecutionEngine.download instead.
The ~scrapy.core.engine.ExecutionEngine.open_spiders attribute is deprecated, use ~scrapy.core.engine.ExecutionEngine.spider instead.
The spider parameter is deprecated for the following methods:
~scrapy.core.engine.ExecutionEngine.spider_is_idle
~scrapy.core.engine.ExecutionEngine.crawl
~scrapy.core.engine.ExecutionEngine.download
Instead, call ~scrapy.core.engine.ExecutionEngine.open_spider first to set the ~scrapy.Spider object.
scrapy.utils.response.response_httprepr is now deprecated. (4972)
You can now use item filtering to control which items are exported to each output feed. (4575, 5178, 5161, 5203)
You can now apply post-processing to feeds, and built-in post-processing plugins are provided for output file compression. (2174, 5168, 5190)
The FEEDS setting now supports pathlib.Path objects as keys. (5383, 5384)
Enabling asyncio while using Windows and Python 3.8 or later will automatically switch the asyncio event loop to one that allows Scrapy to work. See asyncio-windows. (4976, 5315)
The genspider command now supports a start URL instead of a domain name. (4439)
scrapy.utils.defer gained 2 new functions, ~scrapy.utils.defer.deferred_to_future and ~scrapy.utils.defer.maybe_deferred_to_future, to help await on Deferreds when using the asyncio reactor. (5288)
Amazon S3 feed export storage gained support for temporary security credentials (AWS_SESSION_TOKEN) and endpoint customization (AWS_ENDPOINT_URL). (4998, 5210)
New LOG_FILE_APPEND setting to allow truncating the log file. (5279)
Request.cookies values that are bool, float or int are cast to str. (5252, 5253)
You may now raise ~scrapy.exceptions.CloseSpider from a handler of the spider_idle signal to customize the reason why the spider is stopping. (5191)
When using ~scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware, the proxy URL for non-HTTPS HTTP/1.1 requests no longer needs to include a URL scheme. (4505, 4649)
All built-in queues now expose a peek method that returns the next queue object (like pop) but does not remove the returned object from the queue. (5112)
If the underlying queue does not support peeking (e.g. because you are not using queuelib 1.6.1 or later), the peek method raises NotImplementedError.
~scrapy.Request and ~scrapy.http.Response now have an attributes attribute that makes subclassing easier. For ~scrapy.Request, it also allows subclasses to work with scrapy.utils.request.request_from_dict. (1877, 5130, 5218)
The ~scrapy.core.scheduler.BaseScheduler.open and ~scrapy.core.scheduler.BaseScheduler.close methods of the scheduler are now optional. (3559)
HTTP/1.1 ~scrapy.core.downloader.handlers.http11.TunnelError exceptions now only truncate response bodies longer than 1000 characters, instead of those longer than 32 characters, making it easier to debug such errors. (4881, 5007)
~scrapy.loader.ItemLoader now supports non-text responses. (5145, 5269)
The TWISTED_REACTOR and ASYNCIO_EVENT_LOOP settings are no longer ignored if defined in ~scrapy.Spider.custom_settings. (4485, 5352)
Removed a module-level Twisted reactor import that could prevent using the asyncio reactor. (5357)
The startproject command works with existing folders again. (4665, 4676)
The FEED_URI_PARAMS setting now behaves as documented. (4962, 4966)
Request.cb_kwargs once again allows the callback keyword. (5237, 5251, 5264)
Made scrapy.utils.response.open_in_browser support more complex HTML. (5319, 5320)
Fixed CSVFeedSpider.quotechar being interpreted as the CSV file encoding. (5391, 5394)
Added missing setuptools to the list of dependencies. (5122)
LinkExtractor now also works as expected with links that have comma-separated rel attribute values including nofollow. (5225)
Fixed a TypeError that could be raised during feed export parameter parsing. (5359)
asyncio support is no longer considered experimental. (5332)
Included Windows-specific help for asyncio usage. (4976, 5315)
Rewrote topics-headless-browsing with up-to-date best practices. (4484, 4613)
Documented local file naming in media pipelines. (5069, 5152)
faq now covers spider file name collision issues. (2680, 3669)
Provided better context and instructions to disable the URLLENGTH_LIMIT setting. (5135, 5250)
Documented that Reppy parser does not support Python 3.9+. (5226, 5231)
Documented the scheduler component. (3537, 3559)
Documented the method used by media pipelines to determine if a file has expired. (5120, 5254)
run-multiple-spiders now features scrapy.utils.project.get_project_settings usage. (5070)
run-multiple-spiders now covers what happens when you define different per-spider values for some settings that cannot differ at run time. (4485, 5352)
Extended the documentation of the ~scrapy.extensions.statsmailer.StatsMailer extension. (5199, 5217)
Added JOBDIR to topics-settings. (5173, 5224)
Documented Spider.attribute. (5174, 5244)
Documented TextResponse.urljoin. (1582)
Added the body_length parameter to the documented signature of the headers_received signal. (5270)
Clarified SelectorList.get usage in the tutorial. (5256)
The documentation now features the shortest import path of classes with multiple import paths. (2733, 5099)
quotes.toscrape.com references now use HTTPS instead of HTTP. (5395, 5396)
Added a link to our Discord server to getting-help. (5421, 5422)
The pronunciation of the project name is now officially /ˈskreɪpaɪ/. (5280, 5281)
Added the Scrapy logo to the README. (5255, 5258)
Fixed issues and implemented minor improvements. (3155, 4335, 5074, 5098, 5134, 5180, 5194, 5239, 5266, 5271, 5273, 5274, 5276, 5347, 5356, 5414, 5415, 5416, 5419, 5420)
Added support for Python 3.10. (5212, 5221, 5265)
Significantly reduced memory usage by scrapy.utils.response.response_httprepr, used by the ~scrapy.downloadermiddlewares.stats.DownloaderStats downloader middleware, which is enabled by default. (4964, 4972)
Removed uses of the deprecated optparse module. (5366, 5374)
Extended typing hints. (5077, 5090, 5100, 5108, 5171, 5215, 5334)
Improved tests, fixed CI issues, removed unused code. (5094, 5157, 5162, 5198, 5207, 5208, 5229, 5298, 5299, 5310, 5316, 5333, 5388, 5389, 5400, 5401, 5404, 5405, 5407, 5410, 5412, 5425, 5427)
Implemented improvements for contributors. (5080, 5082, 5177, 5200)
Implemented cleanups. (5095, 5106, 5209, 5228, 5235, 5245, 5246, 5292, 5314, 5322)
If you use `HttpAuthMiddleware` (i.e. the http_user and http_pass spider attributes) for HTTP authentication, any request exposes your credentials to
Security bug fix:
If you use HttpAuthMiddleware (i.e. the http_user and http_pass spider attributes) for HTTP authentication, any request exposes your credentials to the request target.
To prevent unintended exposure of authentication credentials to unintended domains, you must now additionally set a new, additional spider attribute, http_auth_domain, and point it to the specific domain to which the authentication credentials must be sent.
If the http_auth_domain spider attribute is not set, the domain of the first request will be considered the HTTP authentication target, and authentication credentials will only be sent in requests targeting that domain.
If you need to send the same HTTP authentication credentials to multiple domains, you can use w3lib.http.basic_auth_header instead to set the value of the Authorization header of your requests.
If you really want your spider to send the same HTTP authentication credentials to any domain, set the http_auth_domain spider attribute to None.
Finally, if you are a user of scrapy-splash, know that this version of Scrapy breaks compatibility with scrapy-splash 0.7.2 and earlier. You will need to upgrade scrapy-splash to a greater version for it to continue to work.
New get_retry_request() function to retry requests from spider callbacks
Highlights:
Official Python 3.9 support
Experimental HTTP/2 support
New get_retry_request() function to retry requests from spider callbacks
New headers_received signal that allows stopping downloads early
New Response.protocol attribute
Removed all code that was deprecated in 1.7.0 and had not already been removed in 2.4.0 . ( #4901 )
Removed support for the SCRAPY_PICKLED_SETTINGS_TO_OVERRIDE environment variable, deprecated in 1.8.0 . ( #4912 )
The scrapy.utils.py36 module is now deprecated in favor of scrapy.utils.asyncgen . ( #4900 )
Experimental HTTP/2 support through a new download handler that can be assigned to the https protocol in the DOWNLOAD_HANDLERS setting. ( #1854 , #4769 , #5058 , #5059 , #5066 )
The new scrapy.downloadermiddlewares.retry.get_retry_request() function may be used from spider callbacks or middlewares to handle the retrying of a request beyond the scenarios that RetryMiddleware supports. ( #3590 , #3685 , #4902 )
The new headers_received signal gives early access to response headers and allows stopping downloads . ( #1772 , #4897 )
The new Response.protocol attribute gives access to the string that identifies the protocol used to download a response. ( #4878 )
Stats now include the following entries that indicate the number of successes and failures in storing feeds :
feedexport / success_count /< storage type > feedexport / failed_count /< storage type >
Where <storage type> is the feed storage backend class name, such as FileFeedStorage or FTPFeedStorage .
( #3947 , #4850 )
The UrlLengthMiddleware spider middleware now logs ignored URLs with INFO logging level instead of DEBUG , and it now includes the following entry into stats to keep track of the number of ignored URLs:
urllength / request_ignored_count
( #5036 )
The HttpCompressionMiddleware downloader middleware now logs the number of decompressed responses and the total count of resulting bytes:
httpcompression / response_bytes httpcompression / response_count
( #4797 , #4799 )
Fixed installation on PyPy installing PyDispatcher in addition to PyPyDispatcher, which could prevent Scrapy from working depending on which package got imported. ( #4710 , #4814 )
When inspecting a callback to check if it is a generator that also returns a value, an exception is no longer raised if the callback has a docstring with lower indentation than the following code. ( #4477 , #4935 )
The Content-Length header is no longer omitted from responses when using the default, HTTP/1.1 download handler (see DOWNLOAD_HANDLERS ). ( #5009 , #5034 , #5045 , #5057 , #5062 )
Setting the handle_httpstatus_all request meta key to False now has the same effect as not setting it at all, instead of having the same effect as setting it to True . ( #3851 , #4694 )
Added instructions to install Scrapy in Windows using pip . ( #4715 , #4736 )
Logging documentation now includes additional ways to filter logs . ( #4216 , #4257 , #4965 )
Covered how to deal with long lists of allowed domains in the FAQ . ( #2263 , #3667 )
Covered scrapy-bench in Benchmarking . ( #4996 , #5016 )
Clarified that one extension instance is created per crawler. ( #5014 )
Fixed some errors in examples. ( #4829 , #4830 , #4907 , #4909 , #5008 )
Fixed some external links, typos, and so on. ( #4892 , #4899 , #4936 , #4942 , #5005 , #5063 )
The list of Request.meta keys is now sorted alphabetically. ( #5061 , #5065 )
Updated references to Scrapinghub, which is now called Zyte. ( #4973 , #5072 )
Added a mention to contributors in the README. ( #4956 )
Reduced the top margin of lists. ( #4974 )
Made Python 3.9 support official ( #4757 , #4759 )
Extended typing hints ( #4895 )
Fixed deprecated uses of the Twisted API. ( #4940 , #4950 , #5073 )
Made our tests run with the new pip resolver. ( #4710 , #4814 )
Added tests to ensure that coroutine support is tested. ( #4987 )
Migrated from Travis CI to GitHub Actions. ( #4924 )
Fixed CI issues. ( #4986 , #5020 , #5022 , #5027 , #5052 , #5053 )
Implemented code refactorings, style fixes and cleanups. ( #4911 , #4982 , #5001 , #5002 , #5076 )
Fixed feed exports overwrite support
Fixed feed exports overwrite support
Fixed the asyncio event loop handling, which could make code hang
Fixed the IPv6-capable DNS resolver CachingHostnameResolver for download handlers that call reactor.resolve
Fixed the output of the genspider command showing placeholders instead of the import part of the generated spider module (issue 4874)
Fixed feed exports overwrite support ( #4845 , #4857 , #4859 )
Fixed the AsyncIO event loop handling, which could make code hang ( #4855 , #4872 )
Fixed the IPv6-capable DNS resolver CachingHostnameResolver for download handlers that call reactor.resolve ( #4802 , #4803 )
Fixed the output of the genspider command showing placeholders instead of the import path of the generated spider module ( #4874 )
Migrated Windows CI from Azure Pipelines to GitHub Actions ( #4869 , #4876 )
Python 3.5 support has been dropped.
Hihglights:
Python 3.5 support has been dropped.
The file_path method of media pipelines can now access the source item.
This allows you to set a download file path based on item data.
The new item_export_kwargs key of the FEEDS setting allows to define keyword parameters to pass to item exporter classes.
You can now choose whether feed exports overwrite or append to the output file.
For example, when using the crawl or runspider commands, you can use the -O option instead of -o to overwrite the output file.
Zstd-compressed responses are now supported if zstandard is installed.
In settings, where the import path of a class is required, it is now possible to pass a class object instead.
Highlights:
Python 3.5 support has been dropped.
The file_path method of media pipelines can now access the source item .
This allows you to set a download file path based on item data.
The new item_export_kwargs key of the FEEDS setting allows to define keyword parameters to pass to item exporter classes
You can now choose whether feed exports overwrite or append to the output file.
For example, when using the crawl or runspider commands, you can use the -O option instead of -o to overwrite the output file.
Zstd-compressed responses are now supported if zstandard is installed.
In settings, where the import path of a class is required, it is now possible to pass a class object instead.
Python 3.6 or greater is now required; support for Python 3.5 has been dropped
As a result:
When using PyPy, PyPy 7.2.0 or greater is now required
For Amazon S3 storage support in feed exports or media pipelines , botocore 1.4.87 or greater is now required
To use the images pipeline , Pillow 4.0.0 or greater is now required
( #4718 , #4732 , #4733 , #4742 , #4743 , #4764 )
CookiesMiddleware once again discards cookies defined in Request.headers .
We decided to revert this bug fix, introduced in Scrapy 2.2.0, because it was reported that the current implementation could break existing code.
If you need to set cookies for a request, use the Request.cookies parameter.
A future version of Scrapy will include a new, better implementation of the reverted bug fix.
( #4717 , #4823 )
scrapy.extensions.feedexport.S3FeedStorage no longer reads the values of access_key and secret_key from the running project settings when they are not passed to its init method; you must either pass those parameters to its init method or use S3FeedStorage.from_crawler ( #4356 , #4411 , #4688 )
Rule.process_request no longer admits callables which expect a single request parameter, rather than both request and response ( #4818 )
In custom media pipelines , signatures that do not accept a keyword-only item parameter in any of the methods that now support this parameter are now deprecated ( #4628 , #4686 )
In custom feed storage backend classes , init method signatures that do not accept a keyword-only feed_options parameter are now deprecated ( #547 , #716 , #4512 )
The scrapy.utils.python.WeakKeyCache class is now deprecated ( #4684 , #4701 )
The scrapy.utils.boto.is_botocore() function is now deprecated, use scrapy.utils.boto.is_botocore_available() instead ( #4734 , #4776 )
The following methods of media pipelines now accept an item keyword-only parameter containing the source item :
In scrapy.pipelines.files.FilesPipeline :
file_downloaded()
file_path()
media_downloaded()
media_to_download()
In scrapy.pipelines.images.ImagesPipeline :
file_downloaded()
file_path()
get_images()
image_downloaded()
media_downloaded()
media_to_download()
( #4628 , #4686 )
The new item_export_kwargs key of the FEEDS setting allows to define keyword parameters to pass to item exporter classes ( #4606 , #4768 )
Feed exports gained overwrite support:
When using the crawl or runspider commands, you can use the -O option instead of -o to overwrite the output file
You can use the overwrite key in the FEEDS setting to configure whether to overwrite the output file ( True ) or append to its content ( False )
The init and from_crawler methods of feed storage backend classes now receive a new keyword-only parameter, feed_options , which is a dictionary of feed options
( #547 , #716 , #4512 )
Zstd-compressed responses are now supported if zstandard is installed ( #4831 )
In settings, where the import path of a class is required, it is now possible to pass a class object instead ( #3870 , #3873 ).
This includes also settings where only part of its value is made of an import path, such as DOWNLOADER_MIDDLEWARES or DOWNLOAD_HANDLERS .
Downloader middlewares can now override response.request .
If a downloader middleware returns a Response object from process_response() or process_exception() with a custom Request object assigned to response.request :
The response is handled by the callback of that custom Request object, instead of being handled by the callback of the original Request object
That custom Request object is now sent as the request argument to the response_received signal, instead of the original Request object
( #4529 , #4632 )
When using the FTP feed storage backend :
It is now possible to set the new overwrite feed option to False to append to an existing file instead of overwriting it
The FTP password can now be omitted if it is not necessary
( #547 , #716 , #4512 )
The init method of CsvItemExporter now supports an errors parameter to indicate how to handle encoding errors ( #4755 )
When using asyncio , it is now possible to set a custom asyncio loop ( #4306 , #4414 )
Serialized requests (see Jobs: pausing and resuming crawls ) now support callbacks that are spider methods that delegate on other callable ( #4756 )
When a response is larger than DOWNLOAD_MAXSIZE , the logged message is now a warning, instead of an error ( #3874 , #3886 , #4752 )
The genspider command no longer overwrites existing files unless the --force option is used ( #4561 , #4616 , #4623 )
Cookies with an empty value are no longer considered invalid cookies ( #4772 )
The runspider command now supports files with the .pyw file extension ( #4643 , #4646 )
The HttpProxyMiddleware middleware now simply ignores unsupported proxy values ( #3331 , #4778 )
Checks for generator callbacks with a return statement no longer warn about return statements in nested functions ( #4720 , #4721 )
The system file mode creation mask no longer affects the permissions of files generated using the startproject command ( #4722 )
scrapy.utils.iterators.xmliter now supports namespaced node names ( #861 , #4746 )
Request objects can now have about: URLs, which can work when using a headless browser ( #4835 )
The FEED_URI_PARAMS setting is now documented ( #4671 , #4724 )
Improved the documentation of link extractors with an usage example from a spider callback and reference documentation for the Link class ( #4751 , #4775 )
Clarified the impact of CONCURRENT_REQUESTS when using the CloseSpider extension ( #4836 )
Removed references to Python 2’s unicode type ( #4547 , #4703 )
We now have an official deprecation policy ( #4705 )
Our documentation policies now cover usage of Sphinx’s versionadded and versionchanged directives, and we have removed usages referencing Scrapy 1.4.0 and earlier versions ( #3971 , #4310 )
Other documentation cleanups ( #4090 , #4782 , #4800 , #4801 , #4809 , #4816 , #4825 )
Extended typing hints ( #4243 , #4691 )
Added tests for the check command ( #4663 )
Fixed test failures on Debian ( #4726 , #4727 , #4735 )
Improved Windows test coverage ( #4723 )
Switched to formatted string literals where possible ( #4307 , #4324 , #4672 )
Modernized super() usage ( #4707 )
Other code and test cleanups ( #1790 , #3288 , #4165 , #4564 , #4651 , #4714 , #4738 , #4745 , #4747 , #4761 , #4765 , #4804 , #4817 , #4820 , #4822 , #4839 )
Feed exports now support Google Cloud Storage as a storage backend
Hihglights:
Feed exports now support Google Cloud Storage as a storage backend
The new FEED_EXPORT_BATCH_ITEM_COUNT setting allows to deliver output items in batches of up to the specified number of items.
It also serves as a workaround for delayed file delivery, which causes Scrapy to only start item delivery after the crawl has finished when using certain storage backends (S3, FTP, and now GCS).
The base implementation of item loaders has been moved into a separate library, itemloaders, allowing usage from outside Scrapy and a separate release schedule
Highlights:
Feed exports now support Google Cloud Storage as a storage backend
The new FEED_EXPORT_BATCH_ITEM_COUNT setting allows to deliver output items in batches of up to the specified number of items.
It also serves as a workaround for delayed file delivery , which causes Scrapy to only start item delivery after the crawl has finished when using certain storage backends ( S3 , FTP , and now GCS ).
The base implementation of item loaders has been moved into a separate library, itemloaders , allowing usage from outside Scrapy and a separate release schedule
Removed the following classes and their parent modules from scrapy.linkextractors :
htmlparser.HtmlParserLinkExtractor
regex.RegexLinkExtractor
sgml.BaseSgmlLinkExtractor
sgml.SgmlLinkExtractor
Use LinkExtractor instead ( #4356 , #4679 )
The scrapy.utils.python.retry_on_eintr function is now deprecated ( #4683 )
Feed exports support Google Cloud Storage ( #685 , #3608 )
New FEED_EXPORT_BATCH_ITEM_COUNT setting for batch deliveries ( #4250 , #4434 )
The parse command now allows specifying an output file ( #4317 , #4377 )
Request.from_curl() and curl_to_request_kwargs() now also support --data-raw ( #4612 )
A parse callback may now be used in built-in spider subclasses, such as CrawlSpider ( #712 , #732 , #781 , #4254 )
Fixed the CSV exporting of dataclass items and attr.s items ( #4667 , #4668 )
Request.from_curl() and curl_to_request_kwargs() now set the request method to POST when a request body is specified and no request method is specified ( #4612 )
The processing of ANSI escape sequences in enabled in Windows 10.0.14393 and later, where it is required for colored output ( #4393 , #4403 )
Updated the OpenSSL cipher list format link in the documentation about the DOWNLOADER_CLIENT_TLS_CIPHERS setting ( #4653 )
Simplified the code example in Working with dataclass items ( #4652 )
The base implementation of item loaders has been moved into itemloaders ( #4005 , #4516 )
Fixed a silenced error in some scheduler tests ( #4644 , #4645 )
Renewed the localhost certificate used for SSL tests ( #4650 )
Removed cookie-handling code specific to Python 2 ( #4682 )
Stopped using Python 2 unicode literal syntax ( #4704 )
Stopped using a backlash for line continuation ( #4673 )
Removed unneeded entries from the MyPy exception list ( #4690 )
Automated tests now pass on Windows as part of our continuous integration system ( #4458 )
Automated tests now pass on the latest PyPy version for supported Python versions in our continuous integration system ( #4504 )
The `startproject` command no longer makes unintended changes to the permissions of files in the destination folder, such as removing execution permis
The startproject command no longer makes unintended changes to the permissions of files in the destination folder, such as removing execution permissions.
The startproject command no longer makes unintended changes to the permissions of files in the destination folder, such as removing execution permissions (4662, 4666)
dataclass objects and attrs objects are now valid item types
Highlights:
TextResponse.json methodbytes_received signal that allows canceling response downloadCookiesMiddleware fixesHighlights:
Python 3.5.2+ is required now
dataclass objects and attrs objects are now valid item types
New TextResponse.json method
New bytes_received signal that allows canceling response download
CookiesMiddleware fixes
Support for Python 3.5.0 and 3.5.1 has been dropped; Scrapy now refuses to run with a Python version lower than 3.5.2, which introduced typing.Type ( #4615 )
TextResponse.body_as_unicode() is now deprecated, use TextResponse.text instead ( #4546 , #4555 , #4579 )
scrapy.item.BaseItem is now deprecated, use scrapy.item.Item instead ( #4534 )
dataclass objects and attrs objects are now valid item types , and a new itemadapter library makes it easy to write code that supports any item type ( #2749 , #2807 , #3761 , #3881 , #4642 )
A new TextResponse.json method allows to deserialize JSON responses ( #2444 , #4460 , #4574 )
A new bytes_received signal allows monitoring response download progress and stopping downloads ( #4205 , #4559 )
The dictionaries in the result list of a media pipeline now include a new key, status , which indicates if the file was downloaded or, if the file was not downloaded, why it was not downloaded; see FilesPipeline.get_media_requests for more information ( #2893 , #4486 )
When using Google Cloud Storage for a media pipeline , a warning is now logged if the configured credentials do not grant the required permissions ( #4346 , #4508 )
Link extractors are now serializable, as long as you do not use lambdas for parameters; for example, you can now pass link extractors in Request.cb_kwargs or Request.meta when persisting scheduled requests ( #4554 )
Upgraded the pickle protocol that Scrapy uses from protocol 2 to protocol 4, improving serialization capabilities and performance ( #4135 , #4541 )
scrapy.utils.misc.create_instance() now raises a TypeError exception if the resulting instance is None ( #4528 , #4532 )
CookiesMiddleware no longer discards cookies defined in Request.headers ( #1992 , #2400 )
CookiesMiddleware no longer re-encodes cookies defined as bytes in the cookies parameter of the init method of Request ( #2400 , #3575 )
When FEEDS defines multiple URIs, FEED_STORE_EMPTY is False and the crawl yields no items, Scrapy no longer stops feed exports after the first URI ( #4621 , #4626 )
Spider callbacks defined using coroutine syntax no longer need to return an iterable, and may instead return a Request object, an item , or None ( #4609 )
The startproject command now ensures that the generated project folders and files have the right permissions ( #4604 )
Fix a KeyError exception being sometimes raised from scrapy.utils.datatypes.LocalWeakReferencedCache ( #4597 , #4599 )
When FEEDS defines multiple URIs, log messages about items being stored now contain information from the corresponding feed, instead of always containing information about only one of the feeds ( #4619 , #4629 )
Added a new section about accessing cb_kwargs from errbacks ( #4598 , #4634 )
Covered chompjs in Parsing JavaScript code ( #4556 , #4562 )
Removed from Coroutines the warning about the API being experimental ( #4511 , #4513 )
Removed references to unsupported versions of Twisted ( #4533 )
Updated the description of the screenshot pipeline example , which now uses coroutine syntax instead of returning a Deferred ( #4514 , #4593 )
Removed a misleading import line from the scrapy.utils.log.configure_logging() code example ( #4510 , #4587 )
The display-on-hover behavior of internal documentation references now also covers links to commands , Request.meta keys, settings and signals ( #4495 , #4563 )
It is again possible to download the documentation for offline reading ( #4578 , #4585 )
Removed backslashes preceding *args and **kwargs in some function and method signatures ( #4592 , #4596 )
Adjusted the code base further to our style guidelines ( #4237 , #4525 , #4538 , #4539 , #4540 , #4542 , #4543 , #4544 , #4545 , #4557 , #4558 , #4566 , #4568 , #4572 )
Removed remnants of Python 2 support ( #4550 , #4553 , #4568 )
Improved code sharing between the crawl and runspider commands ( #4548 , #4552 )
Replaced chain(*iterable) with chain.from_iterable(iterable) ( #4635 )
You may now run the asyncio tests with Tox on any Python version ( #4521 )
Updated test requirements to reflect an incompatibility with pytest 5.4 and 5.4.1 ( #4588 )
Improved SpiderLoader test coverage for scenarios involving duplicate spider names ( #4549 , #4560 )
Configured Travis CI to also run the tests with Python 3.5.2 ( #4518 , #4615 )
Added a Pylint job to Travis CI ( #3727 )
Added a Mypy job to Travis CI ( #4637 )
Made use of set literals in tests ( #4573 )
Cleaned up the Travis CI configuration ( #4517 , #4519 , #4522 , #4537 )
New `FEEDS` setting to export to multiple feeds
Highlights:
FEEDS setting to export to multiple feedsResponse.ip_address attributeHighlights:
New FEEDS setting to export to multiple feeds
New Response.ip_address attribute
AssertionError exceptions triggered by assert statements have been replaced by new exception types, to support running Python in optimized mode (see -O ) without changing Scrapy’s behavior in any unexpected ways.
If you catch an AssertionError exception from Scrapy, update your code to catch the corresponding new exception.
( #4440 )
The LOG_UNSERIALIZABLE_REQUESTS setting is no longer supported, use SCHEDULER_DEBUG instead ( #4385 )
The REDIRECT_MAX_METAREFRESH_DELAY setting is no longer supported, use METAREFRESH_MAXDELAY instead ( #4385 )
The ChunkedTransferMiddleware middleware has been removed, including the entire scrapy.downloadermiddlewares.chunked module; chunked transfers work out of the box ( #4431 )
The spiders property has been removed from Crawler , use CrawlerRunner.spider_loader or instantiate SPIDER_LOADER_CLASS with your settings instead ( #4398 )
The MultiValueDict , MultiValueDictKeyError , and SiteNode classes have been removed from scrapy.utils.datatypes ( #4400 )
The FEED_FORMAT and FEED_URI settings have been deprecated in favor of the new FEEDS setting ( #1336 , #3858 , #4507 )
A new setting, FEEDS , allows configuring multiple output feeds with different settings each ( #1336 , #3858 , #4507 )
The crawl and runspider commands now support multiple -o parameters ( #1336 , #3858 , #4507 )
The crawl and runspider commands now support specifying an output format by appending :<format> to the output file ( #1336 , #3858 , #4507 )
The new Response.ip_address attribute gives access to the IP address that originated a response ( #3903 , #3940 )
A warning is now issued when a value in allowed_domains includes a port ( #50 , #3198 , #4413 )
Zsh completion now excludes used option aliases from the completion list ( #4438 )
Request serialization no longer breaks for callbacks that are spider attributes which are assigned a function with a different name ( #4500 )
None values in allowed_domains no longer cause a TypeError exception ( #4410 )
Zsh completion no longer allows options after arguments ( #4438 )
zope.interface 5.0.0 and later versions are now supported ( #4447 , #4448 )
Spider.make_requests_from_url , deprecated in Scrapy 1.4.0, now issues a warning when used ( #4412 )
Improved the documentation about signals that allow their handlers to return a Deferred ( #4295 , #4390 )
Our PyPI entry now includes links for our documentation, our source code repository and our issue tracker ( #4456 )
Covered the curl2scrapy service in the documentation ( #4206 , #4455 )
Removed references to the Guppy library, which only works in Python 2 ( #4285 , #4343 )
Extended use of InterSphinx to link to Python 3 documentation ( #4444 , #4445 )
Added support for Sphinx 3.0 and later ( #4475 , #4480 , #4496 , #4503 )
Removed warnings about using old, removed settings ( #4404 )
Removed a warning about importing StringTransport from twisted.test.proto_helpers in Twisted 19.7.0 or newer ( #4409 )
Removed outdated Debian package build files ( #4384 )
Removed object usage as a base class ( #4430 )
Removed code that added support for old versions of Twisted that we no longer support ( #4472 )
Fixed code style issues ( #4468 , #4469 , #4471 , #4481 )
Removed twisted.internet.defer.returnValue() calls ( #4443 , #4446 , #4489 )
Response.follow_all now supports an empty URL iterable as input (#4408, #4420)
Response.follow_all now supports an empty URL iterable as input (#4408, #4420)TWISTED_REACTOR (#4401, #4406)Response.follow_all now supports an empty URL iterable as input (4408, 4420)
Removed top-level ~twisted.internet.reactor imports to prevent errors about the wrong Twisted reactor being installed when setting a different Twisted reactor using TWISTED_REACTOR (4401, 4406)
Fixed tests (4422)
Python 2 support has been removed
Highlights:
Highlights:
Python 2 support has been removed
Partial coroutine syntax support and experimental asyncio support
New Response.follow_all method
FTP support for media pipelines
New Response.certificate attribute
IPv6 support through DNS_RESOLVER
Python 2 support has been removed, following Python 2 end-of-life on January 1, 2020 ( #4091 , #4114 , #4115 , #4121 , #4138 , #4231 , #4242 , #4304 , #4309 , #4373 )
Retry gaveups (see RETRY_TIMES ) are now logged as errors instead of as debug information ( #3171 , #3566 )
File extensions that LinkExtractor ignores by default now also include 7z , 7zip , apk , bz2 , cdr , dmg , ico , iso , tar , tar.gz , webm , and xz ( #1837 , #2067 , #4066 )
The METAREFRESH_IGNORE_TAGS setting is now an empty list by default, following web browser behavior ( #3844 , #4311 )
The HttpCompressionMiddleware now includes spaces after commas in the value of the Accept-Encoding header that it sets, following web browser behavior ( #4293 )
The init method of custom download handlers (see DOWNLOAD_HANDLERS ) or subclasses of the following downloader handlers no longer receives a settings parameter:
scrapy.core.downloader.handlers.datauri.DataURIDownloadHandler
scrapy.core.downloader.handlers.file.FileDownloadHandler
Use the from_settings or from_crawler class methods to expose such a parameter to your custom download handlers.
( #4126 )
We have refactored the scrapy.core.scheduler.Scheduler class and related queue classes (see SCHEDULER_PRIORITY_QUEUE , SCHEDULER_DISK_QUEUE and SCHEDULER_MEMORY_QUEUE ) to make it easier to implement custom scheduler queue classes. See Changes to scheduler queue classes below for details.
Overridden settings are now logged in a different format. This is more in line with similar information logged at startup ( #4199 )
The Scrapy shell no longer provides a sel proxy object, use response.selector instead ( #4347 )
LevelDB support has been removed ( #4112 )
The following functions have been removed from scrapy.utils.python : isbinarytext , is_writable , setattr_default , stringify_dict ( #4362 )
Using environment variables prefixed with SCRAPY_ to override settings is deprecated ( #4300 , #4374 , #4375 )
scrapy.linkextractors.FilteringLinkExtractor is deprecated, use scrapy.linkextractors.LinkExtractor instead ( #4045 )
The noconnect query string argument of proxy URLs is deprecated and should be removed from proxy URLs ( #4198 )
The next method of scrapy.utils.python.MutableChain is deprecated, use the global next() function or MutableChain.next instead ( #4153 )
Added partial support for Python’s coroutine syntax and experimental support for asyncio and asyncio -powered libraries ( #4010 , #4259 , #4269 , #4270 , #4271 , #4316 , #4318 )
The new Response.follow_all method offers the same functionality as Response.follow but supports an iterable of URLs as input and returns an iterable of requests ( #2582 , #4057 , #4286 )
Media pipelines now support FTP storage ( #3928 , #3961 )
The new Response.certificate attribute exposes the SSL certificate of the server as a twisted.internet.ssl.Certificate object for HTTPS responses ( #2726 , #4054 )
A new DNS_RESOLVER setting allows enabling IPv6 support ( #1031 , #4227 )
A new SCRAPER_SLOT_MAX_ACTIVE_SIZE setting allows configuring the existing soft limit that pauses request downloads when the total response data being processed is too high ( #1410 , #3551 )
A new TWISTED_REACTOR setting allows customizing the reactor that Scrapy uses, allowing to enable asyncio support or deal with a common macOS issue ( #2905 , #4294 )
Scheduler disk and memory queues may now use the class methods from_crawler or from_settings ( #3884 )
The new Response.cb_kwargs attribute serves as a shortcut for Response.request.cb_kwargs ( #4331 )
Response.follow now supports a flags parameter, for consistency with Request ( #4277 , #4279 )
Item loader processors can now be regular functions, they no longer need to be methods ( #3899 )
Rule now accepts an errback parameter ( #4000 )
Request no longer requires a callback parameter when an errback parameter is specified ( #3586 , #4008 )
LogFormatter now supports some additional methods:
download_error for download errors
item_error for exceptions raised during item processing by item pipelines
spider_error for exceptions raised from spider callbacks
( #374 , #3986 , #3989 , #4176 , #4188 )
The FEED_URI setting now supports pathlib.Path values ( #3731 , #4074 )
A new request_left_downloader signal is sent when a request leaves the downloader ( #4303 )
Scrapy logs a warning when it detects a request callback or errback that uses yield but also returns a value, since the returned value would be lost ( #3484 , #3869 )
Spider objects now raise an AttributeError exception if they do not have a start_urls attribute nor reimplement scrapy.spiders.Spider.start_requests() , but have a start_url attribute ( #4133 , #4170 )
BaseItemExporter subclasses may now use super().init(**kwargs) instead of self._configure(kwargs) in their init method, passing dont_fail=True to the parent init method if needed, and accessing kwargs at self._kwargs after calling their parent init method ( #4193 , #4370 )
A new keep_fragments parameter of scrapy.utils.request.request_fingerprint allows to generate different fingerprints for requests with different fragments in their URL ( #4104 )
Download handlers (see DOWNLOAD_HANDLERS ) may now use the from_settings and from_crawler class methods that other Scrapy components already supported ( #4126 )
scrapy.utils.python.MutableChain.iter now returns self , allowing it to be used as a sequence. ( #4153 )
The crawl command now also exits with exit code 1 when an exception happens before the crawling starts ( #4175 , #4207 )
LinkExtractor.extract_links no longer re-encodes the query string or URLs from non-UTF-8 responses in UTF-8 ( #998 , #1403 , #1949 , #4321 )
The first spider middleware (see SPIDER_MIDDLEWARES ) now also processes exceptions raised from callbacks that are generators ( #4260 , #4272 )
Redirects to URLs starting with 3 slashes ( /// ) are now supported ( #4032 , #4042 )
Request no longer accepts strings as url simply because they have a colon ( #2552 , #4094 )
The correct encoding is now used for attach names in MailSender ( #4229 , #4239 )
RFPDupeFilter , the default DUPEFILTER_CLASS , no longer writes an extra \r character on each line in Windows, which made the size of the requests.seen file unnecessarily large on that platform ( #4283 )
Z shell auto-completion now looks for .html files, not .http files, and covers the -h command-line switch ( #4122 , #4291 )
Adding items to a scrapy.utils.datatypes.LocalCache object without a limit defined no longer raises a TypeError exception ( #4123 )
Fixed a typo in the message of the ValueError exception raised when scrapy.utils.misc.create_instance() gets both settings and crawler set to None ( #4128 )
API documentation now links to an online, syntax-highlighted view of the corresponding source code ( #4148 )
Links to unexisting documentation pages now allow access to the sidebar ( #4152 , #4169 )
Cross-references within our documentation now display a tooltip when hovered ( #4173 , #4183 )
Improved the documentation about LinkExtractor.extract_links and simplified Link Extractors ( #4045 )
Clarified how ItemLoader.item works ( #3574 , #4099 )
Clarified that logging.basicConfig() should not be used when also using CrawlerProcess ( #2149 , #2352 , #3146 , #3960 )
Clarified the requirements for Request objects when using persistence ( #4124 , #4139 )
Clarified how to install a custom image pipeline ( #4034 , #4252 )
Fixed the signatures of the file_path method in media pipeline examples ( #4290 )
Covered a backward-incompatible change in Scrapy 1.7.0 affecting custom scrapy.core.scheduler.Scheduler subclasses ( #4274 )
Improved the README.rst and CODE_OF_CONDUCT.md files ( #4059 )
Documentation examples are now checked as part of our test suite and we have fixed some of the issues detected ( #4142 , #4146 , #4171 , #4184 , #4190 )
Fixed logic issues, broken links and typos ( #4247 , #4258 , #4282 , #4288 , #4305 , #4308 , #4323 , #4338 , #4359 , #4361 )
Improved consistency when referring to the init method of an object ( #4086 , #4088 )
Fixed an inconsistency between code and output in Scrapy at a glance ( #4213 )
Extended intersphinx usage ( #4147 , #4172 , #4185 , #4194 , #4197 )
We now use a recent version of Python to build the documentation ( #4140 , #4249 )
Cleaned up documentation ( #4143 , #4275 )
Re-enabled proxy CONNECT tests ( #2545 , #4114 )
Added Bandit security checks to our test suite ( #4162 , #4181 )
Added Flake8 style checks to our test suite and applied many of the corresponding changes ( #3944 , #3945 , #4137 , #4157 , #4167 , #4174 , #4186 , #4195 , #4238 , #4246 , #4355 , #4360 , #4365 )
Improved test coverage ( #4097 , #4218 , #4236 )
Started reporting slowest tests, and improved the performance of some of them ( #4163 , #4164 )
Fixed broken tests and refactored some tests ( #4014 , #4095 , #4244 , #4268 , #4372 )
Modified the tox configuration to allow running tests with any Python version, run Bandit and Flake8 tests by default, and enforce a minimum tox version programmatically ( #4179 )
Cleaned up code ( #3937 , #4208 , #4209 , #4210 , #4212 , #4369 , #4376 , #4378 )
The following changes may impact any custom queue classes of all types:
The push method no longer receives a second positional parameter containing request.priority * -1 . If you need that value, get it from the first positional parameter, request , instead, or use the new priority() method in scrapy.core.scheduler.ScrapyPriorityQueue subclasses.
The following changes may impact custom priority queue classes:
In the init method or the from_crawler or from_settings class methods:
The parameter that used to contain a factory function, qfactory , is now passed as a keyword parameter named downstream_queue_cls .
A new keyword parameter has been added: key . It is a string that is always an empty string for memory queues and indicates the JOBDIR value for disk queues.
The parameter for disk queues that contains data from the previous crawl, startprios or slot_startprios , is now passed as a keyword parameter named startprios .
The serialize parameter is no longer passed. The disk queue class must take care of request serialization on its own before writing to disk, using the request_to_dict() and request_from_dict() functions from the scrapy.utils.reqser module.
The following changes may impact custom disk and memory queue classes:
The signature of the init method is now init(self, crawler, key) .
The following changes affect specifically the ScrapyPriorityQueue and DownloaderAwarePriorityQueue classes from scrapy.core.scheduler and may affect subclasses:
In the init method, most of the changes described above apply.
init may still receive all parameters as positional parameters, however:
downstream_queue_cls , which replaced qfactory , must be instantiated differently.
qfactory was instantiated with a priority value (integer).
Instances of downstream_queue_cls should be created using the new ScrapyPriorityQueue.qfactory or DownloaderAwarePriorityQueue.pqfactory methods.
The new key parameter displaced the startprios parameter 1 position to the right.
The following class attributes have been added:
crawler
downstream_queue_cls (details above)
key (details above)
The serialize attribute has been removed (details above)
The following changes affect specifically the ScrapyPriorityQueue class and may affect subclasses:
A new priority() method has been added which, given a request, returns request.priority * -1 .
It is used in push() to make up for the removal of its priority parameter.
The spider attribute has been removed. Use crawler.spider instead.
The following changes affect specifically the DownloaderAwarePriorityQueue class and may affect subclasses:
A new pqueues attribute offers a mapping of downloader slot names to the corresponding instances of downstream_queue_cls .
( #3884 )
Security bug fixes. See the full changelog.
Security bug fixes.
Security bug fixes:
Due to its ReDoS vulnerabilities , scrapy.utils.iterators.xmliter is now deprecated in favor of xmliter_lxml() , which XMLFeedSpider now uses.
To minimize the impact of this change on existing code, xmliter_lxml() now supports indicating the node namespace as a prefix in the node name, and big files with highly nested trees when using libxml2 2.7+.
Please, see the cc65-xxvf-f7r9 security advisory for more information.
DOWNLOAD_MAXSIZE and DOWNLOAD_WARNSIZE now also apply to the decompressed response body. Please, see the 7j7m-v7m3-jqm7 security advisory for more information.
Also in relation with the 7j7m-v7m3-jqm7 security advisory , use of the scrapy.downloadermiddlewares.decompression module is discouraged and will trigger a warning.
The Authorization header is now dropped on redirects to a different domain. Please, see the cw9j-q3vf-hrrv security advisory for more information.
Security bug fixes:
Due to its `ReDoS vulnerabilities`_, scrapy.utils.iterators.xmliter is now deprecated in favor of ~scrapy.utils.iterators.xmliter_lxml, which ~scrapy.spiders.XMLFeedSpider now uses.
To minimize the impact of this change on existing code, ~scrapy.utils.iterators.xmliter_lxml now supports indicating the node namespace as a prefix in the node name, and big files with highly nested trees when using libxml2 2.7+.
Please, see the `cc65-xxvf-f7r9 security advisory`_ for more information.
DOWNLOAD_MAXSIZE and DOWNLOAD_WARNSIZE now also apply to the decompressed response body. Please, see the `7j7m-v7m3-jqm7 security advisory`_ for more information.
Also in relation with the `7j7m-v7m3-jqm7 security advisory`_, use of the scrapy.downloadermiddlewares.decompression module is discouraged and will trigger a warning.
The Authorization header is now dropped on redirects to a different domain. Please, see the cw9j-q3vf-hrrv security advisory for more information.
Fixes a security issue around HTTP proxy usage. See the [changelog](https://docs.scrapy.org/en/1.8/news.html#scrapy-1-8-3-2022-07-25) for details.
Fixes a security issue around HTTP proxy usage. See the changelog for details.
Security bug fix:
When HttpProxyMiddleware processes a request with proxy metadata, and that proxy metadata includes proxy credentials, HttpProxyMiddleware sets the Proxy-Authorization header, but only if that header is not already set.
There are third-party proxy-rotation downloader middlewares that set different proxy metadata every time they process a request.
Because of request retries and redirects, the same request can be processed by downloader middlewares more than once, including both HttpProxyMiddleware and any third-party proxy-rotation downloader middleware.
These third-party proxy-rotation downloader middlewares could change the proxy metadata of a request to a new value, but fail to remove the Proxy-Authorization header from the previous value of the proxy metadata, causing the credentials of one proxy to be sent to a different proxy.
To prevent the unintended leaking of proxy credentials, the behavior of HttpProxyMiddleware is now as follows when processing a request:
If the request being processed defines proxy metadata that includes credentials, the Proxy-Authorization header is always updated to feature those credentials.
If the request being processed defines proxy metadata without credentials, the Proxy-Authorization header is removed unless it was originally defined for the same proxy URL.
To remove proxy credentials while keeping the same proxy URL, remove the Proxy-Authorization header.
If the request has no proxy metadata, or that metadata is a falsy value (e.g. None ), the Proxy-Authorization header is removed.
It is no longer possible to set a proxy URL through the proxy metadata but set the credentials through the Proxy-Authorization header. Set proxy credentials through the proxy metadata instead.
When a `Request` object with cookies defined gets a redirect response causing a new `Request` object to be scheduled, the cookies defined in the origi
When a Request object with cookies defined gets a redirect response causing a new Request object to be scheduled, the cookies defined in the original Request object are no longer copied into the new Request object.
If you manually set the Cookie header on a Request object and the domain name of the redirect URL is not an exact match for the domain of the URL of the original Request object, your Cookie header is now dropped from the new Request object.
The old behavior could be exploited by an attacker to gain access to your cookies. Please, see the cjvr-mfj7-j4j8 security advisory for more information.
Note: It is still possible to enable the sharing of cookies between different domains with a shared domain suffix (e.g. example.com and any subdomain) by defining the shared domain suffix (e.g. example.com) as the cookie domain when defining your cookies. See the documentation of the Request class for more information.
When the domain of a cookie, either received in the Set-Cookie header of a response or defined in a Request object, is set to a public suffix <https://publicsuffix.org/>_, the cookie is now ignored unless the cookie domain is the same as the request domain.
The old behavior could be exploited by an attacker to inject cookies from a controlled domain into your cookiejar that could be sent to other domains not controlled by the attacker. Please, see the mfjm-vh54-3f96 security advisory for more information.
Security bug fixes:
When a ~scrapy.Request object with cookies defined gets a redirect response causing a new ~scrapy.Request object to be scheduled, the cookies defined in the original ~scrapy.Request object are no longer copied into the new ~scrapy.Request object.
If you manually set the Cookie header on a ~scrapy.Request object and the domain name of the redirect URL is not an exact match for the domain of the URL of the original ~scrapy.Request object, your Cookie header is now dropped from the new ~scrapy.Request object.
The old behavior could be exploited by an attacker to gain access to your cookies. Please, see the cjvr-mfj7-j4j8 security advisory for more information.
Note
It is still possible to enable the sharing of cookies between different domains with a shared domain suffix (e.g. example.com and any subdomain) by defining the shared domain suffix (e.g. example.com) as the cookie domain when defining your cookies. See the documentation of the ~scrapy.Request class for more information.
When the domain of a cookie, either received in the Set-Cookie header of a response or defined in a ~scrapy.Request object, is set to a public suffix, the cookie is now ignored unless the cookie domain is the same as the request domain.
The old behavior could be exploited by an attacker to inject cookies into your requests to some other domains. Please, see the mfjm-vh54-3f96 security advisory for more information.
If you use `HttpAuthMiddleware` (i.e. the http_user and http_pass spider attributes) for HTTP authentication, any request exposes your credentials to
Security bug fix:
If you use HttpAuthMiddleware (i.e. the http_user and http_pass spider attributes) for HTTP authentication, any request exposes your credentials to the request target.
To prevent unintended exposure of authentication credentials to unintended domains, you must now additionally set a new, additional spider attribute, http_auth_domain, and point it to the specific domain to which the authentication credentials must be sent.
If the http_auth_domain spider attribute is not set, the domain of the first request will be considered the HTTP authentication target, and authentication credentials will only be sent in requests targeting that domain.
If you need to send the same HTTP authentication credentials to multiple domains, you can use w3lib.http.basic_auth_header instead to set the value of the Authorization header of your requests.
If you really want your spider to send the same HTTP authentication credentials to any domain, set the http_auth_domain spider attribute to None.
Finally, if you are a user of scrapy-splash, know that this version of Scrapy breaks compatibility with scrapy-splash 0.7.2 and earlier. You will need to upgrade scrapy-splash to a greater version for it to continue to work.
See also Deprecation removals below.
Highlights:
Dropped Python 3.4 support and updated minimum requirements; made Python 3.8 support official
New Request.from_curl() class method
New ROBOTSTXT_PARSER and ROBOTSTXT_USER_AGENT settings
New DOWNLOADER_CLIENT_TLS_CIPHERS and DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING settings
Python 3.4 is no longer supported, and some of the minimum requirements of Scrapy have also changed:
cssselect 0.9.1
cryptography 2.0
lxml 3.5.0
pyOpenSSL 16.2.0
queuelib 1.4.2
service_identity 16.0.0
six 1.10.0
Twisted 17.9.0 (16.0.0 with Python 2)
zope.interface 4.1.3
( #3892 )
JSONRequest is now called JsonRequest for consistency with similar classes ( #3929 , #3982 )
If you are using a custom context factory ( DOWNLOADER_CLIENTCONTEXTFACTORY ), its init method must accept two new parameters: tls_verbose_logging and tls_ciphers ( #2111 , #3392 , #3442 , #3450 )
ItemLoader now turns the values of its input item into lists:
item = MyItem () >>> item [ "field" ] = "value1" >>> loader = ItemLoader ( item = item ) >>> item [ "field" ] ['value1']
This is needed to allow adding values to existing fields ( loader.add_value('field', 'value2') ).
( #3804 , #3819 , #3897 , #3976 , #3998 , #4036 )
See also Deprecation removals below.
A new Request.from_curl class method allows creating a request from a cURL command ( #2985 , #3862 )
A new ROBOTSTXT_PARSER setting allows choosing which robots.txt parser to use. It includes built-in support for RobotFileParser , Protego (default), Reppy, and Robotexclusionrulesparser , and allows you to implement support for additional parsers ( #754 , #2669 , #3796 , #3935 , #3969 , #4006 )
A new ROBOTSTXT_USER_AGENT setting allows defining a separate user agent string to use for robots.txt parsing ( #3931 , #3966 )
Rule no longer requires a LinkExtractor parameter ( #781 , #4016 )
Use the new DOWNLOADER_CLIENT_TLS_CIPHERS setting to customize the TLS/SSL ciphers used by the default HTTP/1.1 downloader ( #3392 , #3442 )
Set the new DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING setting to True to enable debug-level messages about TLS connection parameters after establishing HTTPS connections ( #2111 , #3450 )
Callbacks that receive keyword arguments (see Request.cb_kwargs ) can now be tested using the new @cb_kwargs spider contract ( #3985 , #3988 )
When a @scrapes spider contract fails, all missing fields are now reported ( #766 , #3939 )
Custom log formats can now drop messages by having the corresponding methods of the configured LOG_FORMATTER return None ( #3984 , #3987 )
A much improved completion definition is now available for Zsh ( #4069 )
ItemLoader.load_item() no longer makes later calls to ItemLoader.get_output_value() or ItemLoader.load_item() return empty data ( #3804 , #3819 , #3897 , #3976 , #3998 , #4036 )
Fixed DummyStatsCollector raising a TypeError exception ( #4007 , #4052 )
FilesPipeline.file_path and ImagesPipeline.file_path no longer choose file extensions that are not registered with IANA ( #1287 , #3953 , #3954 )
When using botocore to persist files in S3, all botocore-supported headers are properly mapped now ( #3904 , #3905 )
FTP passwords in FEED_URI containing percent-escaped characters are now properly decoded ( #3941 )
A memory-handling and error-handling issue in scrapy.utils.ssl.get_temp_key_info() has been fixed ( #3920 )
The documentation now covers how to define and configure a custom log format ( #3616 , #3660 )
API documentation added for MarshalItemExporter and PythonItemExporter ( #3973 )
API documentation added for BaseItem and ItemMeta ( #3999 )
Minor documentation fixes ( #2998 , #3398 , #3597 , #3894 , #3934 , #3978 , #3993 , #4022 , #4028 , #4033 , #4046 , #4050 , #4055 , #4056 , #4061 , #4072 , #4071 , #4079 , #4081 , #4089 , #4093 )
scrapy.xlib has been removed ( #4015 )
The LevelDB storage backend ( scrapy.extensions.httpcache.LeveldbCacheStorage ) of HttpCacheMiddleware is deprecated ( #4085 , #4092 )
Use of the undocumented SCRAPY_PICKLED_SETTINGS_TO_OVERRIDE environment variable is deprecated ( #3910 )
scrapy.item.DictItem is deprecated, use Item instead ( #3999 )
Minimum versions of optional Scrapy requirements that are covered by continuous integration tests have been updated:
botocore 1.3.23
Pillow 3.4.2
Lower versions of these optional requirements may work, but it is not guaranteed ( #3892 )
GitHub templates for bug reports and feature requests ( #3126 , #3471 , #3749 , #3754 )
Continuous integration fixes ( #3923 )
Code cleanup ( #3391 , #3907 , #3946 , #3950 , #4023 , #4031 )
Highlights:
Dropped Python 3.4 support and updated minimum requirements; made Python 3.8 support official
New .Request.from_curl class method
New ROBOTSTXT_PARSER and ROBOTSTXT_USER_AGENT settings
New DOWNLOADER_CLIENT_TLS_CIPHERS and DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING settings
Python 3.4 is no longer supported, and some of the minimum requirements of Scrapy have also changed:
cssselect 0.9.1
cryptography_ 2.0
lxml_ 3.5.0
pyOpenSSL_ 16.2.0
queuelib_ 1.4.2
service_identity_ 16.0.0
six_ 1.10.0
Twisted_ 17.9.0 (16.0.0 with Python 2)
zope.interface_ 4.1.3
JSONRequest is now called ~scrapy.http.JsonRequest for consistency with similar classes (3929, 3982)
If you are using a custom context factory (DOWNLOADER_CLIENTCONTEXTFACTORY), its __init__ method must accept two new parameters: tls_verbose_logging and tls_ciphers (2111, 3392, 3442, 3450)
~scrapy.loader.ItemLoader now turns the values of its input item into lists:
>>> item = MyItem()
>>> item["field"] = "value1"
>>> loader = ItemLoader(item=item)
>>> item["field"]
['value1']
This is needed to allow adding values to existing fields (loader.add_value('field', 'value2')).
(3804, 3819, 3897, 3976, 3998, 4036)
See also 1.8-deprecation-removals below.
A new Request.from_curl class method allows creating a request from a cURL command (2985, 3862)
A new ROBOTSTXT_PARSER setting allows choosing which robots.txt_ parser to use. It includes built-in support for RobotFileParser, Protego (default), Reppy, and Robotexclusionrulesparser, and allows you to implement support for additional parsers (754, 2669, 3796, 3935, 3969, 4006)
A new ROBOTSTXT_USER_AGENT setting allows defining a separate user agent string to use for robots.txt_ parsing (3931, 3966)
~scrapy.spiders.Rule no longer requires a LinkExtractor parameter (781, 4016)
Use the new DOWNLOADER_CLIENT_TLS_CIPHERS setting to customize the TLS/SSL ciphers used by the default HTTP/1.1 downloader (3392, 3442)
Set the new DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING setting to True to enable debug-level messages about TLS connection parameters after establishing HTTPS connections (2111, 3450)
Callbacks that receive keyword arguments (see .Request.cb_kwargs) can now be tested using the new @cb_kwargs spider contract (3985, 3988)
When a @scrapes spider contract fails, all missing fields are now reported (766, 3939)
Custom log formats can now drop messages by having the corresponding methods of the configured LOG_FORMATTER return None (3984, 3987)
A much improved completion definition is now available for Zsh_ (4069)
ItemLoader.load_item() no longer makes later calls to ItemLoader.get_output_value() or ItemLoader.load_item() return empty data (3804, 3819, 3897, 3976, 3998, 4036)
Fixed ~scrapy.statscollectors.DummyStatsCollector raising a TypeError exception (4007, 4052)
FilesPipeline.file_path and ImagesPipeline.file_path no longer choose file extensions that are not `registered with IANA`_ (1287, 3953, 3954)
When using botocore_ to persist files in S3, all botocore-supported headers are properly mapped now (3904, 3905)
FTP passwords in FEED_URI containing percent-escaped characters are now properly decoded (3941)
A memory-handling and error-handling issue in scrapy.utils.ssl.get_temp_key_info has been fixed (3920)
The documentation now covers how to define and configure a custom log format (3616, 3660)
API documentation added for ~scrapy.exporters.MarshalItemExporter and ~scrapy.exporters.PythonItemExporter (3973)
API documentation added for ~scrapy.item.BaseItem and ~scrapy.item.ItemMeta (3999)
Minor documentation fixes (2998, 3398, 3597, 3894, 3934, 3978, 3993, 4022, 4028, 4033, 4046, 4050, 4055, 4056, 4061, 4072, 4071, 4079, 4081, 4089, 4093)
scrapy.xlib has been removed (4015)
The LevelDB_ storage backend (scrapy.extensions.httpcache.LeveldbCacheStorage) of ~scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware is deprecated (4085, 4092)
Use of the undocumented SCRAPY_PICKLED_SETTINGS_TO_OVERRIDE environment variable is deprecated (3910)
scrapy.item.DictItem is deprecated, use ~scrapy.item.Item instead (3999)
Minimum versions of optional Scrapy requirements that are covered by continuous integration tests have been updated:
botocore_ 1.3.23
Pillow_ 3.4.2
Lower versions of these optional requirements may work, but it is not guaranteed (3892)
GitHub templates for bug reports and feature requests (3126, 3471, 3749, 3754)
Continuous integration fixes (3923)
Code cleanup (3391, 3907, 3946, 3950, 4023, 4031)
Revert the fix for #3804 (#3819), which has a few undesired side effects (#3897, #3976).
Revert the fix for #3804 (#3819), which has a few undesired side effects (#3897, #3976).
Revert the fix for #3804 ( #3819 ), which has a few undesired side effects ( #3897 , #3976 ).
As a result, when an item loader is initialized with an item, ItemLoader.load_item() once again makes later calls to ItemLoader.get_output_value() or ItemLoader.load_item() return empty data.
Enforce lxml 4.3.5 or lower for Python 3.4 (#3912, #3918)
Enforce lxml 4.3.5 or lower for Python 3.4 (#3912, #3918)
Fix Python 2 support (#3889, #3893, #3896)
Fix Python 2 support (#3889, #3893, #3896)
Re-packaging of Scrapy 1.7.0, which was missing some changes in PyPI.
Re-packaging of Scrapy 1.7.0, which was missing some changes in PyPI.
Improvements for crawls targeting multiple domains
Highlights:
Note
Make sure you install Scrapy 1.7.1. The Scrapy 1.7.0 package in PyPI is the result of an erroneous commit tagging and does not include all the changes described below.
Highlights:
Improvements for crawls targeting multiple domains
A cleaner way to pass arguments to callbacks
A new class for JSON requests
Improvements for rule-based spiders
New features for feed exports
429 is now part of the RETRY_HTTP_CODES setting by default
This change is backward incompatible . If you don’t want to retry 429 , you must override RETRY_HTTP_CODES accordingly.
Crawler , CrawlerRunner.crawl and CrawlerRunner.create_crawler no longer accept a Spider subclass instance, they only accept a Spider subclass now.
Spider subclass instances were never meant to work, and they were not working as one would expect: instead of using the passed Spider subclass instance, their from_crawler method was called to generate a new instance.
Non-default values for the SCHEDULER_PRIORITY_QUEUE setting may stop working. Scheduler priority queue classes now need to handle Request objects instead of arbitrary Python data structures.
An additional crawler parameter has been added to the init method of the Scheduler class. Custom scheduler subclasses which don’t accept arbitrary parameters in their init method might break because of this change.
For more information, see SCHEDULER .
See also Deprecation removals below.
A new scheduler priority queue, scrapy.pqueues.DownloaderAwarePriorityQueue , may be enabled for a significant scheduling improvement on crawls targeting multiple web domains, at the cost of no CONCURRENT_REQUESTS_PER_IP support ( #3520 )
A new Request.cb_kwargs attribute provides a cleaner way to pass keyword arguments to callback methods ( #1138 , #3563 )
A new JSONRequest class offers a more convenient way to build JSON requests ( #3504 , #3505 )
A process_request callback passed to the Rule init method now receives the Response object that originated the request as its second argument ( #3682 )
A new restrict_text parameter for the LinkExtractor init method allows filtering links by linking text ( #3622 , #3635 )
A new FEED_STORAGE_S3_ACL setting allows defining a custom ACL for feeds exported to Amazon S3 ( #3607 )
A new FEED_STORAGE_FTP_ACTIVE setting allows using FTP’s active connection mode for feeds exported to FTP servers ( #3829 )
A new METAREFRESH_IGNORE_TAGS setting allows overriding which HTML tags are ignored when searching a response for HTML meta tags that trigger a redirect ( #1422 , #3768 )
A new redirect_reasons request meta key exposes the reason (status code, meta refresh) behind every followed redirect ( #3581 , #3687 )
The SCRAPY_CHECK variable is now set to the true string during runs of the check command, which allows detecting contract check runs from code ( #3704 , #3739 )
A new Item.deepcopy() method makes it easier to deep-copy items ( #1493 , #3671 )
CoreStats also logs elapsed_time_seconds now ( #3638 )
Exceptions from ItemLoader input and output processors are now more verbose ( #3836 , #3840 )
Crawler , CrawlerRunner.crawl and CrawlerRunner.create_crawler now fail gracefully if they receive a Spider subclass instance instead of the subclass itself ( #2283 , #3610 , #3872 )
process_spider_exception() is now also invoked for generators ( #220 , #2061 )
System exceptions like KeyboardInterrupt are no longer caught ( #3726 )
ItemLoader.load_item() no longer makes later calls to ItemLoader.get_output_value() or ItemLoader.load_item() return empty data ( #3804 , #3819 )
The images pipeline ( ImagesPipeline ) no longer ignores these Amazon S3 settings: AWS_ENDPOINT_URL , AWS_REGION_NAME , AWS_USE_SSL , AWS_VERIFY ( #3625 )
Fixed a memory leak in scrapy.pipelines.media.MediaPipeline affecting, for example, non-200 responses and exceptions from custom middlewares ( #3813 )
Requests with private callbacks are now correctly unserialized from disk ( #3790 )
FormRequest.from_response() now handles invalid methods like major web browsers ( #3777 , #3794 )
A new topic, Selecting dynamically-loaded content , covers recommended approaches to read dynamically-loaded data ( #3703 )
Speeding up broad crawls now features information about memory usage ( #1264 , #3866 )
The documentation of Rule now covers how to access the text of a link when using CrawlSpider ( #3711 , #3712 )
A new section, Writing your own storage backend , covers writing a custom cache storage backend for HttpCacheMiddleware ( #3683 , #3692 )
A new FAQ entry, How to split an item into multiple items in an item pipeline? , explains what to do when you want to split an item into multiple items from an item pipeline ( #2240 , #3672 )
Updated the FAQ entry about crawl order to explain why the first few requests rarely follow the desired order ( #1739 , #3621 )
The LOGSTATS_INTERVAL setting ( #3730 ), the FilesPipeline.file_path and ImagesPipeline.file_path methods ( #2253 , #3609 ) and the Crawler.stop() method ( #3842 ) are now documented
Some parts of the documentation that were confusing or misleading are now clearer ( #1347 , #1789 , #2289 , #3069 , #3615 , #3626 , #3668 , #3670 , #3673 , #3728 , #3762 , #3861 , #3882 )
Minor documentation fixes ( #3648 , #3649 , #3662 , #3674 , #3676 , #3694 , #3724 , #3764 , #3767 , #3791 , #3797 , #3806 , #3812 )
The following deprecated APIs have been removed ( #3578 ):
scrapy.conf (use Crawler.settings )
From scrapy.core.downloader.handlers :
http.HttpDownloadHandler (use http10.HTTP10DownloadHandler )
scrapy.loader.ItemLoader._get_values (use _get_xpathvalues )
scrapy.loader.XPathItemLoader (use ItemLoader )
scrapy.log (see Logging )
From scrapy.pipelines :
files.FilesPipeline.file_key (use file_path )
images.ImagesPipeline.file_key (use file_path )
images.ImagesPipeline.image_key (use file_path )
images.ImagesPipeline.thumb_key (use thumb_path )
From both scrapy.selector and scrapy.selector.lxmlsel :
HtmlXPathSelector (use Selector )
XmlXPathSelector (use Selector )
XPathSelector (use Selector )
XPathSelectorList (use Selector )
From scrapy.selector.csstranslator :
ScrapyGenericTranslator (use parsel.csstranslator.GenericTranslator )
ScrapyHTMLTranslator (use parsel.csstranslator.HTMLTranslator )
ScrapyXPathExpr (use parsel.csstranslator.XPathExpr )
From Selector :
_root (both the init method argument and the object property, use root )
extract_unquoted (use getall )
select (use xpath )
From SelectorList :
extract_unquoted (use getall )
select (use xpath )
x (use xpath )
scrapy.spiders.BaseSpider (use Spider )
From Spider (and subclasses):
DOWNLOAD_DELAY (use download_delay )
set_crawler (use from_crawler() )
scrapy.spiders.spiders (use SpiderLoader )
scrapy.telnet (use scrapy.extensions.telnet )
From scrapy.utils.python :
str_to_unicode (use to_unicode )
unicode_to_str (use to_bytes )
scrapy.utils.response.body_or_str
The following deprecated settings have also been removed ( #3578 ):
SPIDER_MANAGER_CLASS (use SPIDER_LOADER_CLASS )
The queuelib.PriorityQueue value for the SCHEDULER_PRIORITY_QUEUE setting is deprecated. Use scrapy.pqueues.ScrapyPriorityQueue instead.
process_request callbacks passed to Rule that do not accept two arguments are deprecated.
The following modules are deprecated:
scrapy.utils.http (use w3lib.http )
scrapy.utils.markup (use w3lib.html )
scrapy.utils.multipart (use urllib3 )
The scrapy.utils.datatypes.MergeDict class is deprecated for Python 3 code bases. Use ChainMap instead. ( #3878 )
The scrapy.utils.gz.is_gzipped function is deprecated. Use scrapy.utils.gz.gzip_magic_number instead.
It is now possible to run all tests from the same tox environment in parallel; the documentation now covers this and other ways to run tests ( #3707 )
It is now possible to generate an API documentation coverage report ( #3806 , #3810 , #3860 )
The documentation policies now require docstrings ( #3701 ) that follow PEP 257 ( #3748 )
Internal fixes and cleanup ( #3629 , #3643 , #3684 , #3698 , #3734 , #3735 , #3736 , #3737 , #3809 , #3821 , #3825 , #3827 , #3833 , #3857 , #3877 )
Note
Make sure you install Scrapy 1.7.1. The Scrapy 1.7.0 package in PyPI is the result of an erroneous commit tagging and does not include all the changes described below.
Highlights:
Improvements for crawls targeting multiple domains
A cleaner way to pass arguments to callbacks
A new class for JSON requests
Improvements for rule-based spiders
New features for feed exports
429 is now part of the RETRY_HTTP_CODES setting by default
This change is backward incompatible. If you don’t want to retry 429, you must override RETRY_HTTP_CODES accordingly.
~scrapy.crawler.Crawler, CrawlerRunner.crawl and CrawlerRunner.create_crawler no longer accept a ~scrapy.spiders.Spider subclass instance, they only accept a ~scrapy.spiders.Spider subclass now.
~scrapy.spiders.Spider subclass instances were never meant to work, and they were not working as one would expect: instead of using the passed ~scrapy.spiders.Spider subclass instance, their ~scrapy.spiders.Spider.from_crawler method was called to generate a new instance.
Non-default values for the SCHEDULER_PRIORITY_QUEUE setting may stop working. Scheduler priority queue classes now need to handle ~scrapy.Request objects instead of arbitrary Python data structures.
An additional crawler parameter has been added to the __init__ method of the ~scrapy.core.scheduler.Scheduler class. Custom scheduler subclasses which don't accept arbitrary parameters in their __init__ method might break because of this change.
For more information, see SCHEDULER.
See also 1.7-deprecation-removals below.
A new scheduler priority queue, scrapy.pqueues.DownloaderAwarePriorityQueue, may be enabled for a significant scheduling improvement on crawls targeting multiple web domains, at the cost of no CONCURRENT_REQUESTS_PER_IP support (3520)
A new .Request.cb_kwargs attribute provides a cleaner way to pass keyword arguments to callback methods (1138, 3563)
A new JSONRequest class offers a more convenient way to build JSON requests (3504, 3505)
A process_request callback passed to the ~scrapy.spiders.Rule __init__ method now receives the ~scrapy.http.Response object that originated the request as its second argument (3682)
A new restrict_text parameter for the LinkExtractor __init__ method allows filtering links by linking text (3622, 3635)
A new FEED_STORAGE_S3_ACL setting allows defining a custom ACL for feeds exported to Amazon S3 (3607)
A new FEED_STORAGE_FTP_ACTIVE setting allows using FTP’s active connection mode for feeds exported to FTP servers (3829)
A new METAREFRESH_IGNORE_TAGS setting allows overriding which HTML tags are ignored when searching a response for HTML meta tags that trigger a redirect (1422, 3768)
A new redirect_reasons request meta key exposes the reason (status code, meta refresh) behind every followed redirect (3581, 3687)
The SCRAPY_CHECK variable is now set to the true string during runs of the check command, which allows detecting contract check runs from code (3704, 3739)
A new Item.deepcopy() method makes it easier to deep-copy items (1493, 3671)
~scrapy.extensions.corestats.CoreStats also logs elapsed_time_seconds now (3638)
Exceptions from ~scrapy.loader.ItemLoader input and output processors are now more verbose (3836, 3840)
~scrapy.crawler.Crawler, CrawlerRunner.crawl and CrawlerRunner.create_crawler now fail gracefully if they receive a ~scrapy.spiders.Spider subclass instance instead of the subclass itself (2283, 3610, 3872)
~scrapy.spidermiddlewares.SpiderMiddleware.process_spider_exception is now also invoked for generators (220, 2061)
System exceptions like KeyboardInterrupt_ are no longer caught (3726)
ItemLoader.load_item() no longer makes later calls to ItemLoader.get_output_value() or ItemLoader.load_item() return empty data (3804, 3819)
The images pipeline (~scrapy.pipelines.images.ImagesPipeline) no longer ignores these Amazon S3 settings: AWS_ENDPOINT_URL, AWS_REGION_NAME, AWS_USE_SSL, AWS_VERIFY (3625)
Fixed a memory leak in scrapy.pipelines.media.MediaPipeline affecting, for example, non-200 responses and exceptions from custom middlewares (3813)
Requests with private callbacks are now correctly unserialized from disk (3790)
.FormRequest.from_response now handles invalid methods like major web browsers (3777, 3794)
A new topic, topics-dynamic-content, covers recommended approaches to read dynamically-loaded data (3703)
topics-broad-crawls now features information about memory usage (1264, 3866)
The documentation of ~scrapy.spiders.Rule now covers how to access the text of a link when using ~scrapy.spiders.CrawlSpider (3711, 3712)
A new section, httpcache-storage-custom, covers writing a custom cache storage backend for ~scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware (3683, 3692)
A new FAQ entry, faq-split-item, explains what to do when you want to split an item into multiple items from an item pipeline (2240, 3672)
Updated the FAQ entry about crawl order to explain why the first few requests rarely follow the desired order (1739, 3621)
The LOGSTATS_INTERVAL setting (3730), the FilesPipeline.file_path and ImagesPipeline.file_path methods (2253, 3609) and the Crawler.stop() method (3842) are now documented
Some parts of the documentation that were confusing or misleading are now clearer (1347, 1789, 2289, 3069, 3615, 3626, 3668, 3670, 3673, 3728, 3762, 3861, 3882)
Minor documentation fixes (3648, 3649, 3662, 3674, 3676, 3694, 3724, 3764, 3767, 3791, 3797, 3806, 3812)
The following deprecated APIs have been removed (3578):
scrapy.conf (use Crawler.settings)
From scrapy.core.downloader.handlers:
http.HttpDownloadHandler (use http10.HTTP10DownloadHandler)
scrapy.loader.ItemLoader._get_values (use _get_xpathvalues)
scrapy.loader.XPathItemLoader (use ~scrapy.loader.ItemLoader)
scrapy.log (see topics-logging)
From scrapy.pipelines:
files.FilesPipeline.file_key (use file_path)
images.ImagesPipeline.file_key (use file_path)
images.ImagesPipeline.image_key (use file_path)
images.ImagesPipeline.thumb_key (use thumb_path)
From both scrapy.selector and scrapy.selector.lxmlsel:
HtmlXPathSelector (use ~scrapy.Selector)
XmlXPathSelector (use ~scrapy.Selector)
XPathSelector (use ~scrapy.Selector)
XPathSelectorList (use ~scrapy.Selector)
From scrapy.selector.csstranslator:
ScrapyGenericTranslator (use parsel.csstranslator.GenericTranslator_)
ScrapyHTMLTranslator (use parsel.csstranslator.HTMLTranslator_)
ScrapyXPathExpr (use parsel.csstranslator.XPathExpr_)
From ~scrapy.Selector:
_root (both the __init__ method argument and the object property, use root)
extract_unquoted (use getall)
select (use xpath)
From ~scrapy.selector.SelectorList:
extract_unquoted (use getall)
select (use xpath)
x (use xpath)
scrapy.spiders.BaseSpider (use ~scrapy.spiders.Spider)
From ~scrapy.spiders.Spider (and subclasses):
DOWNLOAD_DELAY (use download_delay)
set_crawler (use ~scrapy.spiders.Spider.from_crawler)
scrapy.spiders.spiders (use ~scrapy.spiderloader.SpiderLoader)
scrapy.telnet (use scrapy.extensions.telnet)
From scrapy.utils.python:
str_to_unicode (use to_unicode)
unicode_to_str (use to_bytes)
scrapy.utils.response.body_or_str
The following deprecated settings have also been removed (3578):
SPIDER_MANAGER_CLASS (use SPIDER_LOADER_CLASS)
The queuelib.PriorityQueue value for the SCHEDULER_PRIORITY_QUEUE setting is deprecated. Use scrapy.pqueues.ScrapyPriorityQueue instead.
process_request callbacks passed to ~scrapy.spiders.Rule that do not accept two arguments are deprecated.
The following modules are deprecated:
scrapy.utils.http (use w3lib.http)
scrapy.utils.markup (use w3lib.html)
scrapy.utils.multipart (use urllib3)
The scrapy.utils.datatypes.MergeDict class is deprecated for Python 3 code bases. Use ~collections.ChainMap instead. (3878)
The scrapy.utils.gz.is_gzipped function is deprecated. Use scrapy.utils.gz.gzip_magic_number instead.
It is now possible to run all tests from the same tox_ environment in parallel; the documentation now covers this and other ways to run tests (3707)
It is now possible to generate an API documentation coverage report (3806, 3810, 3860)
The documentation policies now require docstrings_ (3701) that follow `PEP 257`_ (3748)
Internal fixes and cleanup (3629, 3643, 3684, 3698, 3734, 3735, 3736, 3737, 3809, 3821, 3825, 3827, 3833, 3857, 3877)
Clean-up of the deprecated code
Highlights:
Highlights:
better Windows support;
Python 3.7 compatibility;
big documentation improvements, including a switch from .extract_first() + .extract() API to .get() + .getall() API;
feed exports, FilePipeline and MediaPipeline improvements;
better extensibility: item_error and request_reached_downloader signals; from_crawler support for feed exporters, feed storages and dupefilters.
scrapy.contracts fixes and new features;
telnet console security improvements, first released as a backport in Scrapy 1.5.2 (2019-01-22) ;
clean-up of the deprecated code;
various bug fixes, small new features and usability improvements across the codebase.
While these are not changes in Scrapy itself, but rather in the parsel library which Scrapy uses for xpath/css selectors, these changes are worth mentioning here. Scrapy now depends on parsel >= 1.5, and Scrapy documentation is updated to follow recent parsel API conventions.
Most visible change is that .get() and .getall() selector methods are now preferred over .extract_first() and .extract() . We feel that these new methods result in a more concise and readable code. See extract() and extract_first() for more details.
Note
There are currently no plans to deprecate .extract() and .extract_first() methods.
Another useful new feature is the introduction of Selector.attrib and SelectorList.attrib properties, which make it easier to get attributes of HTML elements. See Selecting element attributes .
CSS selectors are cached in parsel >= 1.5, which makes them faster when the same CSS path is used many times. This is very common in case of Scrapy spiders: callbacks are usually called several times, on different pages.
If you’re using custom Selector or SelectorList subclasses, a backward incompatible change in parsel may affect your code. See parsel changelog for a detailed description, as well as for the full list of improvements.
Backward incompatible : Scrapy’s telnet console now requires username and password. See Telnet Console for more details. This change fixes a security issue ; see Scrapy 1.5.2 (2019-01-22) release notes for details.
from_crawler support is added to feed exporters and feed storages. This, among other things, allows to access Scrapy settings from custom feed storages and exporters ( #1605 , #3348 ).
from_crawler support is added to dupefilters ( #2956 ); this allows to access e.g. settings or a spider from a dupefilter.
item_error is fired when an error happens in a pipeline ( #3256 );
request_reached_downloader is fired when Downloader gets a new Request; this signal can be useful e.g. for custom Schedulers ( #3393 ).
new SitemapSpider sitemap_filter() method which allows to select sitemap entries based on their attributes in SitemapSpider subclasses ( #3512 ).
Lazy loading of Downloader Handlers is now optional; this enables better initialization error handling in custom Downloader Handlers ( #3394 ).
Expose more options for S3FilesStore: AWS_ENDPOINT_URL , AWS_USE_SSL , AWS_VERIFY , AWS_REGION_NAME . For example, this allows to use alternative or self-hosted AWS-compatible providers ( #2609 , #3548 ).
ACL support for Google Cloud Storage: FILES_STORE_GCS_ACL and IMAGES_STORE_GCS_ACL ( #3199 ).
Exceptions in contracts code are handled better ( #3377 );
dont_filter=True is used for contract requests, which allows to test different callbacks with the same URL ( #3381 );
request_cls attribute in Contract subclasses allow to use different Request classes in contracts, for example FormRequest ( #3383 ).
Fixed errback handling in contracts, e.g. for cases where a contract is executed for URL which returns non-200 response ( #3371 ).
more stats for RobotsTxtMiddleware ( #3100 )
INFO log level is used to show telnet host/port ( #3115 )
a message is added to IgnoreRequest in RobotsTxtMiddleware ( #3113 )
better validation of url argument in Response.follow ( #3131 )
non-zero exit code is returned from Scrapy commands when error happens on spider initialization ( #3226 )
Link extraction improvements: “ftp” is added to scheme list ( #3152 ); “flv” is added to common video extensions ( #3165 )
better error message when an exporter is disabled ( #3358 );
scrapy shell --help mentions syntax required for local files ( ./file.html ) - #3496 .
Referer header value is added to RFPDupeFilter log messages ( #3588 )
fixed issue with extra blank lines in .csv exports under Windows ( #3039 );
proper handling of pickling errors in Python 3 when serializing objects for disk queues ( #3082 )
flags are now preserved when copying Requests ( #3342 );
FormRequest.from_response clickdata shouldn’t ignore elements with input[type=image] ( #3153 ).
FormRequest.from_response should preserve duplicate keys ( #3247 )
Docs are re-written to suggest .get/.getall API instead of .extract/.extract_first. Also, Selectors docs are updated and re-structured to match latest parsel docs; they now contain more topics, such as Selecting element attributes or Extensions to CSS Selectors ( #3390 ).
Using your browser’s Developer Tools for scraping is a new tutorial which replaces old Firefox and Firebug tutorials ( #3400 ).
SCRAPY_PROJECT environment variable is documented ( #3518 );
troubleshooting section is added to install instructions ( #3517 );
improved links to beginner resources in the tutorial ( #3367 , #3468 );
fixed RETRY_HTTP_CODES default values in docs ( #3335 );
remove unused DEPTH_STATS option from docs ( #3245 );
other cleanups ( #3347 , #3350 , #3445 , #3544 , #3605 ).
Compatibility shims for pre-1.0 Scrapy module names are removed ( #3318 ):
scrapy.command
scrapy.contrib (with all submodules)
scrapy.contrib_exp (with all submodules)
scrapy.dupefilter
scrapy.linkextractor
scrapy.project
scrapy.spider
scrapy.spidermanager
scrapy.squeue
scrapy.stats
scrapy.statscol
scrapy.utils.decorator
See Module Relocations for more information, or use suggestions from Scrapy 1.5.x deprecation warnings to update your code.
Other deprecation removals:
Deprecated scrapy.interfaces.ISpiderManager is removed; please use scrapy.interfaces.ISpiderLoader.
Deprecated CrawlerSettings class is removed ( #3327 ).
Deprecated Settings.overrides and Settings.defaults attributes are removed ( #3327 , #3359 ).
All Scrapy tests now pass on Windows; Scrapy testing suite is executed in a Windows environment on CI ( #3315 ).
Python 3.7 support ( #3326 , #3150 , #3547 ).
Testing and CI fixes ( #3526 , #3538 , #3308 , #3311 , #3309 , #3305 , #3210 , #3299 )
scrapy.http.cookies.CookieJar.clear accepts “domain”, “path” and “name” optional arguments ( #3231 ).
additional files are included to sdist ( #3495 );
code style fixes ( #3405 , #3304 );
unneeded .strip() call is removed ( #3519 );
collections.deque is used to store MiddlewareManager methods instead of a list ( #3476 )
The fix is backward incompatible , it enables telnet user-password authentication by default with a random generated password. If you can’t upgrade ri…
Security bugfix : Telnet console extension can be easily exploited by rogue websites POSTing content to http://localhost:6023 , we haven’t found a way to exploit it from Scrapy, but it is very easy to trick a browser to do so and elevates the risk for local development environment.
The fix is backward incompatible , it enables telnet user-password authentication by default with a random generated password. If you can’t upgrade right away, please consider setting TELNETCONSOLE_PORT out of its default value.
See telnet console documentation for more info
Backport CI build failure under GCE environment due to boto import error.
This is a maintenance release with important bug fixes, but no new features:
This is a maintenance release with important bug fixes, but no new features:
O(N^2) gzip decompression issue which affected Python 3 and PyPy is fixed ( #3281 );
skipping of TLS validation errors is improved ( #3166 );
Ctrl-C handling is fixed in Python 3.5+ ( #3096 );
testing fixes ( #3092 , #3263 );
documentation improvements ( #3058 , #3059 , #3089 , #3123 , #3127 , #3189 , #3224 , #3280 , #3279 , #3201 , #3260 , #3284 , #3298 , #3294 ).
This release brings small new features and improvements across the codebase. Some highlights:
This release brings small new features and improvements across the codebase. Some highlights:
This release brings small new features and improvements across the codebase. Some highlights:
Google Cloud Storage is supported in FilesPipeline and ImagesPipeline.
Crawling with proxy servers becomes more efficient, as connections to proxies can be reused now.
Warnings, exception and logging messages are improved to make debugging easier.
scrapy parse command now allows to set custom request meta via --meta argument.
Compatibility with Python 3.6, PyPy and PyPy3 is improved; PyPy and PyPy3 are now supported officially, by running tests on CI.
Better default handling of HTTP 308, 522 and 524 status codes.
Documentation is improved, as usual.
Scrapy 1.5 drops support for Python 3.3.
Default Scrapy User-Agent now uses https link to scrapy.org ( #2983 ). This is technically backward-incompatible ; override USER_AGENT if you relied on old value.
Logging of settings overridden by custom_settings is fixed; this is technically backward-incompatible because the logger changes from [scrapy.utils.log] to [scrapy.crawler] . If you’re parsing Scrapy logs, please update your log parsers ( #1343 ).
LinkExtractor now ignores m4v extension by default, this is change in behavior.
522 and 524 status codes are added to RETRY_HTTP_CODES ( #2851 )
Support <link> tags in Response.follow ( #2785 )
Support for ptpython REPL ( #2654 )
Google Cloud Storage support for FilesPipeline and ImagesPipeline ( #2923 ).
New --meta option of the “scrapy parse” command allows to pass additional request.meta ( #2883 )
Populate spider variable when using shell.inspect_response ( #2812 )
Handle HTTP 308 Permanent Redirect ( #2844 )
Add 522 and 524 to RETRY_HTTP_CODES ( #2851 )
Log versions information at startup ( #2857 )
scrapy.mail.MailSender now works in Python 3 (it requires Twisted 17.9.0)
Connections to proxy servers are reused ( #2743 )
Add template for a downloader middleware ( #2755 )
Explicit message for NotImplementedError when parse callback not defined ( #2831 )
CrawlerProcess got an option to disable installation of root log handler ( #2921 )
LinkExtractor now ignores m4v extension by default
Better log messages for responses over DOWNLOAD_WARNSIZE and DOWNLOAD_MAXSIZE limits ( #2927 )
Show warning when a URL is put to Spider.allowed_domains instead of a domain ( #2250 ).
Fix logging of settings overridden by custom_settings ; this is technically backward-incompatible because the logger changes from [scrapy.utils.log] to [scrapy.crawler] , so please update your log parsers if needed ( #1343 )
Default Scrapy User-Agent now uses https link to scrapy.org ( #2983 ). This is technically backward-incompatible ; override USER_AGENT if you relied on old value.
Fix PyPy and PyPy3 test failures, support them officially ( #2793 , #2935 , #2990 , #3050 , #2213 , #3048 )
Fix DNS resolver when DNSCACHE_ENABLED=False ( #2811 )
Add cryptography for Debian Jessie tox test env ( #2848 )
Add verification to check if Request callback is callable ( #2766 )
Port extras/qpsclient.py to Python 3 ( #2849 )
Use getfullargspec under the scenes for Python 3 to stop DeprecationWarning ( #2862 )
Update deprecated test aliases ( #2876 )
Fix SitemapSpider support for alternate links ( #2853 )
Added missing bullet point for the AUTOTHROTTLE_TARGET_CONCURRENCY setting. ( #2756 )
Update Contributing docs, document new support channels ( #2762 , #3038 )
Include references to Scrapy subreddit in the docs
Fix broken links; use https:// for external links ( #2978 , #2982 , #2958 )
Document CloseSpider extension better ( #2759 )
Use pymongo.collection.Collection.insert_one() in MongoDB example ( #2781 )
Spelling mistake and typos ( #2828 , #2837 , #2884 , #2924 )
Clarify CSVFeedSpider.headers documentation ( #2826 )
Document DontCloseSpider exception and clarify spider_idle ( #2791 )
Update “Releases” section in README ( #2764 )
Fix rst syntax in DOWNLOAD_FAIL_ON_DATALOSS docs ( #2763 )
Small fix in description of startproject arguments ( #2866 )
Clarify data types in Response.body docs ( #2922 )
Add a note about request.meta['depth'] to DepthMiddleware docs ( #2374 )
Add a note about request.meta['dont_merge_cookies'] to CookiesMiddleware docs ( #2999 )
Up-to-date example of project structure ( #2964 , #2976 )
A better example of ItemExporters usage ( #2989 )
Document from_crawler methods for spider and downloader middlewares ( #3019 )
Release notes at https://doc.scrapy.org/en/latest/news.html#scrapy-1-4-0-2017-05-18
Release notes at https://doc.scrapy.org/en/latest/news.html#scrapy-1-4-0-2017-05-18
Scrapy 1.4 does not bring that many breathtaking new features but quite a few handy improvements nonetheless.
Scrapy now supports anonymous FTP sessions with customizable user and password via the new FTP_USER and FTP_PASSWORD settings. And if you’re using Twisted version 17.1.0 or above, FTP is now available with Python 3.
There’s a new response.follow method for creating requests; it is now a recommended way to create Requests in Scrapy spiders . This method makes it easier to write correct spiders; response.follow has several advantages over creating scrapy.Request objects directly:
it handles relative URLs;
it works properly with non-ascii URLs on non-UTF8 pages;
in addition to absolute and relative URLs it supports Selectors; for <a> elements it can also extract their href values.
For example, instead of this:
for href in response . css ( 'li.page a::attr(href)' ) . extract (): url = response . urljoin ( href ) yield scrapy . Request ( url , self . parse , encoding = response . encoding )
One can now write this:
for a in response . css ( 'li.page a' ): yield response . follow ( a , self . parse )
Link extractors are also improved. They work similarly to what a regular modern browser would do: leading and trailing whitespace are removed from attributes (think href=" http://example.com" ) when building Link objects. This whitespace-stripping also happens for action attributes with FormRequest .
Please also note that link extractors do not canonicalize URLs by default anymore. This was puzzling users every now and then, and it’s not what browsers do in fact, so we removed that extra transformation on extracted links.
For those of you wanting more control on the Referer: header that Scrapy sends when following links, you can set your own Referrer Policy . Prior to Scrapy 1.4, the default RefererMiddleware would simply and blindly set it to the URL of the response that generated the HTTP request (which could leak information on your URL seeds). By default, Scrapy now behaves much like your regular browser does. And this policy is fully customizable with W3C standard values (or with something really custom of your own if you wish). See REFERRER_POLICY for details.
To make Scrapy spiders easier to debug, Scrapy logs more stats by default in 1.4: memory usage stats, detailed retry stats, detailed HTTP error code stats. A similar change is that HTTP cache path is also visible in logs now.
Last but not least, Scrapy now has the option to make JSON and XML items more human-readable, with newlines between items and even custom indenting offset, using the new FEED_EXPORT_INDENT setting.
Enjoy! (Or read on for the rest of changes in this release.)
Default to canonicalize=False in scrapy.linkextractors.LinkExtractor ( #2537 , fixes #1941 and #1982 ): warning, this is technically backward-incompatible
Enable memusage extension by default ( #2539 , fixes #2187 ); this is technically backward-incompatible so please check if you have any non-default MEMUSAGE_*** options set.
EDITOR environment variable now takes precedence over EDITOR option defined in settings.py ( #1829 ); Scrapy default settings no longer depend on environment variables. This is technically a backward incompatible change .
Spider.make_requests_from_url is deprecated ( #1728 , fixes #1495 ).
Accept proxy credentials in proxy request meta key ( #2526 )
Support brotli-compressed content; requires optional brotlipy ( #2535 )
New response.follow shortcut for creating requests ( #1940 )
Added flags argument and attribute to Request objects ( #2047 )
Support Anonymous FTP ( #2342 )
Added retry/count , retry/max_reached and retry/reason_count/<reason> stats to RetryMiddleware ( #2543 )
Added httperror/response_ignored_count and httperror/response_ignored_status_count/<status> stats to HttpErrorMiddleware ( #2566 )
Customizable Referrer policy in RefererMiddleware ( #2306 )
New data: URI download handler ( #2334 , fixes #2156 )
Log cache directory when HTTP Cache is used ( #2611 , fixes #2604 )
Warn users when project contains duplicate spider names (fixes #2181 )
scrapy.utils.datatypes.CaselessDict now accepts Mapping instances and not only dicts ( #2646 )
Media downloads , with FilesPipeline or ImagesPipeline , can now optionally handle HTTP redirects using the new MEDIA_ALLOW_REDIRECTS setting ( #2616 , fixes #2004 )
Accept non-complete responses from websites using a new DOWNLOAD_FAIL_ON_DATALOSS setting ( #2590 , fixes #2586 )
Optional pretty-printing of JSON and XML items via FEED_EXPORT_INDENT setting ( #2456 , fixes #1327 )
Allow dropping fields in FormRequest.from_response formdata when None value is passed ( #667 )
Per-request retry times with the new max_retry_times meta key ( #2642 )
python -m scrapy as a more explicit alternative to scrapy command ( #2740 )
LinkExtractor now strips leading and trailing whitespaces from attributes ( #2547 , fixes #1614 )
Properly handle whitespaces in action attribute in FormRequest ( #2548 )
Buffer CONNECT response bytes from proxy until all HTTP headers are received ( #2495 , fixes #2491 )
FTP downloader now works on Python 3, provided you use Twisted>=17.1 ( #2599 )
Use body to choose response type after decompressing content ( #2393 , fixes #2145 )
Always decompress Content-Encoding: gzip at HttpCompressionMiddleware stage ( #2391 )
Respect custom log level in Spider.custom_settings ( #2581 , fixes #1612 )
‘make htmlview’ fix for macOS ( #2661 )
Remove “commands” from the command list ( #2695 )
Fix duplicate Content-Length header for POST requests with empty body ( #2677 )
Properly cancel large downloads, i.e. above DOWNLOAD_MAXSIZE ( #1616 )
ImagesPipeline: fixed processing of transparent PNG images with palette ( #2675 )
Tests: remove temp files and folders ( #2570 ), fixed ProjectUtilsTest on macOS ( #2569 ), use portable pypy for Linux on Travis CI ( #2710 )
Separate building request from _requests_to_follow in CrawlSpider ( #2562 )
Remove “Python 3 progress” badge ( #2567 )
Add a couple more lines to .gitignore ( #2557 )
Remove bumpversion prerelease configuration ( #2159 )
Add codecov.yml file ( #2750 )
Set context factory implementation based on Twisted version ( #2577 , fixes #2560 )
Add omitted self arguments in default project middleware template ( #2595 )
Remove redundant slot.add_request() call in ExecutionEngine ( #2617 )
Catch more specific os.error exception in scrapy.pipelines.files.FSFilesStore ( #2644 )
Change “localhost” test server certificate ( #2720 )
Remove unused MEMUSAGE_REPORT setting ( #2576 )
Binary mode is required for exporters ( #2564 , fixes #2553 )
Mention issue with FormRequest.from_response() due to bug in lxml ( #2572 )
Use single quotes uniformly in templates ( #2596 )
Document ftp_user and ftp_password meta keys ( #2587 )
Removed section on deprecated contrib/ ( #2636 )
Recommend Anaconda when installing Scrapy on Windows ( #2477 , fixes #2475 )
FAQ: rewrite note on Python 3 support on Windows ( #2690 )
Rearrange selector sections ( #2705 )
Remove nonzero from SelectorList docs ( #2683 )
Mention how to disable request filtering in documentation of DUPEFILTER_CLASS setting ( #2714 )
Add sphinx_rtd_theme to docs setup readme ( #2668 )
Open file in text mode in JSON item writer example ( #2729 )
Clarify allowed_domains example ( #2670 )
Release notes at https://doc.scrapy.org/en/latest/news.html#scrapy-1-3-3-2017-03-10
Release notes at https://doc.scrapy.org/en/latest/news.html#scrapy-1-3-3-2017-03-10
Make SpiderLoader raise ImportError again by default for missing dependencies and wrong SPIDER_MODULES . These exceptions were silenced as warnings since 1.3.0. A new setting is introduced to toggle between warning or exception if needed ; see SPIDER_LOADER_WARN_ONLY for details.
- Preserve request class when converting to/from dicts (utils.reqser) ( #2510 ).
Preserve request class when converting to/from dicts (utils.reqser) ( #2510 ).
Use consistent selectors for author field in tutorial ( #2551 ).
Fix TLS compatibility in Twisted 17+ ( #2558 )
Preserve request class when converting to/from dicts (utils.reqser) (2510).
Use consistent selectors for author field in tutorial (2551).
Fix TLS compatibility in Twisted 17+ (2558)
- Support 'True' and 'False' string values for boolean settings ( #2519 ); you can now do something like scrapy crawl myspider -s REDIRECT_ENABLED=Fal
Support 'True' and 'False' string values for boolean settings ( #2519 ); you can now do something like scrapy crawl myspider -s REDIRECT_ENABLED=False .
Support kwargs with response.xpath() to use XPath variables and ad-hoc namespaces declarations ; this requires at least Parsel v1.1 ( #2457 ).
Add support for Python 3.6 ( #2485 ).
Run tests on PyPy (warning: some tests still fail, so PyPy is not supported yet).
Enforce DNS_TIMEOUT setting ( #2496 ).
Fix view command ; it was a regression in v1.3.0 ( #2503 ).
Fix tests regarding *_EXPIRES settings with Files/Images pipelines ( #2460 ).
Fix name of generated pipeline class when using basic project template ( #2466 ).
Fix compatibility with Twisted 17+ ( #2496 , #2528 ).
Fix scrapy.Item inheritance on Python 3.6 ( #2511 ).
Enforce numeric values for components order in SPIDER_MIDDLEWARES , DOWNLOADER_MIDDLEWARES , EXTENSIONS and SPIDER_CONTRACTS ( #2420 ).
Reword Code of Conduct section and upgrade to Contributor Covenant v1.4 ( #2469 ).
Clarify that passing spider arguments converts them to spider attributes ( #2483 ).
Document formid argument on FormRequest.from_response() ( #2497 ).
Add .rst extension to README files ( #2507 ).
Mention LevelDB cache storage backend ( #2525 ).
Use yield in sample callback code ( #2533 ).
Add note about HTML entities decoding with .re()/.re_first() ( #1704 ).
Typos ( #2512 , #2534 , #2531 ).
Remove redundant check in MetaRefreshMiddleware ( #2542 ).
Faster checks in LinkExtractor for allow/deny patterns ( #2538 ).
Remove dead code supporting old Twisted versions ( #2544 ).
- HttpErrorMiddleware now logs errors with INFO level instead of DEBUG ; this is technically backward incompatible so please check your log parsers.
This release comes rather soon after 1.2.2 for one main reason: it was found out that releases since 0.18 up to 1.2.2 (included) use some backported code from Twisted ( scrapy.xlib.tx.* ), even if newer Twisted modules are available. Scrapy now uses twisted.web.client and twisted.internet.endpoints directly. (See also cleanups below.)
As it is a major change, we wanted to get the bug fix out quickly while not breaking any projects using the 1.2 series.
MailSender now accepts single strings as values for to and cc arguments ( #2272 )
scrapy fetch url , scrapy shell url and fetch(url) inside Scrapy shell now follow HTTP redirections by default ( #2290 ); See fetch and shell for details.
HttpErrorMiddleware now logs errors with INFO level instead of DEBUG ; this is technically backward incompatible so please check your log parsers.
By default, logger names now use a long-form path, e.g. [scrapy.extensions.logstats] , instead of the shorter “top-level” variant of prior releases (e.g. [scrapy] ); this is backward incompatible if you have log parsers expecting the short logger name part. You can switch back to short logger names using LOG_SHORT_NAMES set to True .
Scrapy now requires Twisted >= 13.1 which is the case for many Linux distributions already.
As a consequence, we got rid of scrapy.xlib.tx.* modules, which copied some of Twisted code for users stuck with an “old” Twisted version
ChunkedTransferMiddleware is deprecated and removed from the default downloader middlewares.
- Packaging fix: disallow unsupported Twisted versions in setup.py
Packaging fix: disallow unsupported Twisted versions in setup.py
Remove page on (deprecated & unsupported) Ubuntu packages from ToC
open_spider() (#2011)Your coding agent can read these notes before it upgrades. Set up the MCP server →