NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1442 most downloaded on PyPI
Library of web-related functions
Last release 4 days ago
30 Sep 2026
Release timing varies
gaps range from 1 weeks to 2.2 years
Nearly every release is documented
notes for 39 of 40 stable releases
Nothing withdrawn
no release was ever pulled
15 years old
40 releases · first in 2011
Bump version: 2.4.1 → 2.5.0
Bump version: 2.4.1 → 2.5.0
New features:
Added support for Python 3.15 (#295).
Added ~w3lib.url.add_http_if_no_scheme, ported from Scrapy, which adds http as the default scheme to a URL that has none (#309).
w3lib.url.parse_qsl_to_bytes, ~w3lib.url.url_query_parameter, ~w3lib.url.add_or_replace_parameter, ~w3lib.url.add_or_replace_parameters and ~w3lib.url.canonicalize_url now accept a separator (or query_separator for ~w3lib.url.canonicalize_url) keyword argument, to support query strings that use a separator other than & (#167).
~w3lib.html.get_base_url and ~w3lib.html.get_meta_refresh now accept a max_scan keyword argument, an upper bound on how much of the document they look at (#336).
Improved the performance of most w3lib.url functions, ~w3lib.url.safe_url_string in particular (#257), and of some w3lib.html functions (#256).
~w3lib.html.get_base_url and ~w3lib.html.get_meta_refresh are now several times faster on typical pages, which have no <base> tag and no <meta> refresh tag (#331, #332, #334, #338).
Deprecations and removals:
The w3lib.util module is deprecated, and now emits a DeprecationWarning on import (#322).
The undocumented w3lib_replace codec error handler is no longer registered (#318).
Security and correctness fixes:
~w3lib.url.safe_url_string, ~w3lib.url.canonicalize_url and w3lib.url.parse_url no longer disagree with how browsers and urllib.parse read a URL in the following cases, which could let a URL resolve to a different host or path than the one these functions reported:
\ is now treated like / in the authority and path of special-scheme URLs (#285).
An NFKC-normalized backslash in the host is now rejected, like other normalized authority delimiters already were (#280).
ASCII tab, carriage return and line feed characters are now stripped from the whole URL, including the host (#301).
Additionally, ~w3lib.url.safe_url_string now adds a / path to a URL that has a query but no path, since some HTTP clients otherwise send the query alone as the request target (#297), and no longer raises UnicodeDecodeError for a host that cannot be IDNA-encoded, such as one with an empty label, when encoding is not UTF-8 (#319).
~w3lib.url.canonicalize_url now resolves dot segments (. and ..) in the path and normalizes IPv6 addresses in the host, so that equivalent URLs canonicalize the same (#308), and keeps percent-encoded characters whose decoding would change their meaning, such as a semicolon (%3B, a path parameter delimiter, #323) or a percent sign (%25, which would otherwise decode to a bare %, #306), in the path encoded.
~w3lib.url.canonicalize_url no longer IDNA-encodes the userinfo of a URL together with its host, which changed both, e.g. http://user@éxample.com became http://xn--user@xample-fbb.com (#279).
~w3lib.url.safe_download_url now resolves percent-encoded dot segments, such as %2e or %2e%2e, like . and .. (#320), and decides whether the path keeps its trailing / based on the path alone, rather than on whether the whole URL ends with / (#311).
~w3lib.url.url_query_parameter now accepts a bytes URL, instead of parsing its b'…' representation (#325), and ~w3lib.url.url_query_cleaner accepts one as well, instead of raising TypeError, while the type hint of its parameterlist argument no longer claims to accept bytes items, which never matched anything (#257, #326).
~w3lib.url.path_to_file_uri and ~w3lib.url.file_uri_to_path now round-trip Windows UNC paths, \\server\share\file becoming file://server/share/file (#324), and ~w3lib.url.file_uri_to_path no longer raises for a path that starts with an extra pair of slashes, such as ////foo/bar (#327).
Fixed excessive backtracking on malformed input, which could be used for denial of service, in ~w3lib.html.get_base_url (#264), ~w3lib.html.get_meta_refresh (#266, #288), ~w3lib.html.remove_tags_with_content (#271), ~w3lib.html.replace_tags and ~w3lib.html.remove_tags (#268, #286), and ~w3lib.html.unquote_markup on unclosed CDATA sections (#289).
~w3lib.html.get_base_url no longer picks up a <base> tag written as text inside <script> or <noscript>, where a browser would ignore it, or one broken by a comment, such as <base h<!--c-->ref="…">, which a browser does not read as a tag at all (#303).
~w3lib.html.get_base_url now reads only the first href attribute of the <base> tag, as browsers do, instead of the last attribute whose name ends in href, such as data-href, and now also reads an unquoted href value (#349).
~w3lib.html.get_meta_refresh and ~w3lib.html.remove_tags_with_content no longer treat a <script> or <noscript> element as closed only by a bare </script>, now also recognizing a closing tag followed by whitespace, / or another token (#302).
~w3lib.url.parse_data_uri no longer raises on an empty quoted parameter value, such as charset="" (#310).
~w3lib.html.replace_entities and ~w3lib.html.get_meta_refresh now only treat ASCII digits as numeric character references and refresh intervals, matching browsers (#292).
~w3lib.html.replace_entities now resolves a null or surrogate numeric character reference, such as � or �, to U+FFFD, as browsers do, instead of producing a NUL or a lone surrogate that cannot be encoded (#321).
~w3lib.encoding.resolve_encoding, and with it every function that detects the encoding of a response, no longer resolves UTF-7, which browsers do not support and which lets ASCII-only bytes smuggle markup past filters (#307).
~w3lib.encoding.http_content_type_encoding now recognizes a quoted charset parameter value, and no longer stops matching a valid charset because of what follows it or of a quoted-string parameter value that itself contains charset (#291, #293).
~w3lib.encoding.html_body_declared_encoding now skips HTML comments, like ~w3lib.html.get_base_url and ~w3lib.html.get_meta_refresh already did (#305), and finds a charset declaration regardless of the order of the attributes in the tag, of how the http-equiv pragma is spelled, e.g. httpequiv="ContentType", and of what surrounds charset= within a content attribute value (#299, #337).
~w3lib.encoding.html_to_unicode now decodes a lead 0x80 byte as the euro sign when the encoding is GB18030, matching the WHATWG Encoding Standard (#300), and decodes a BOM-less utf-16 or utf-32 encoding declared in the body or auto-detected as big-endian, as it already did for one declared in the Content-Type header (#317).
~w3lib.html.replace_tags, ~w3lib.html.remove_tags and ~w3lib.html.remove_tags_with_content no longer end a tag at an angle bracket inside a quoted attribute value, such as the > of <img alt="a>b" src=x> (#343).
~w3lib.html.replace_entities and ~w3lib.html.unquote_markup no longer consume keep when it is an iterator, such as a generator, which left some entities unkept (#342, #345), and ~w3lib.html.remove_tags no longer raises ValueError when which_ones or keep is an empty iterator and the other one is not (#345).
~w3lib.html.get_meta_refresh no longer reports a redirect for a refresh content value that browsers do not follow because it uses non-ASCII whitespace, such as U+00A0 (#353).
~w3lib.html.get_meta_refresh no longer prints the HTML content it was given to stdout when a UnicodeDecodeError is raised (#256).
Tests, benchmarking and CI improvements (#247, #259, #262, #281, #296, #304, #313, #314, #329, #335, #340, #344, #350).
One column per quarter.
Bump version: 2.4.0 → 2.4.1
Bump version: 2.4.0 → 2.4.1
~w3lib.url.safe_url_string now preserves IPv6 brackets in the URL netloc (#253).
Dropped support for Python 3.9 and PyPy 3.10 ( #250 ).
headers_raw_to_dict and headers_dict_to_raw (#246).hatchling (#243).sphinx-hoverxref extension is no longer used to build the docs (#244).Dropped support for Python 3.9 and PyPy 3.10 (#250).
Added support for Python 3.14 and PyPy 3.11 (#241, #245).
Improved performance of ~w3lib.http.headers_raw_to_dict and ~w3lib.http.headers_dict_to_raw (#246).
Switched the build system to hatchling (#243).
The obsolete sphinx-hoverxref extension is no longer used to build the docs (#244).
Tests and CI improvements (#237, #238, #240, #242, #248, #251).
No code changes from v2.3.0.
Removed the following functions, deprecated in 2.0.0:
canonicalize_url() no longer applies lowercase to the userinfo URL component. ( #229 , #230 )
~w3lib.url.canonicalize_url no longer applies lowercase to the userinfo URL component. (#229, #230)
Dropped Python 3.7 support ( #214 ).
.readthedocs.yml (#219).pre-commit configuration, code reformatted with black (#220).Fix test failures on Python 3.11.4+ ( #212 , #213 ).
safe_url_string, canonicalize_url: apply stripping from the URL living standard by @Gallaecio in #207
Full Changelog: v2.1.0...v2.1.1
~w3lib.url.safe_url_string, ~w3lib.url.safe_download_url and ~w3lib.url.canonicalize_url now strip whitespace and control characters urls according to the URL living standard.
safe_url_string() , safe_download_url() and canonicalize_url() now strip whitespace and control characters urls according to the URL living standard.
update type annotation of auto_detect_fun param in html_to_unicode() by @BurnzZ in https://github.com/scrapy/w3lib/pull/190
OverflowError exception on convert_entity by @Laerte in https://github.com/scrapy/w3lib/pull/202Full Changelog: https://github.com/scrapy/w3lib/compare/v2.0.1...v2.1.0
OverflowError exception on convert_entity by @Laerte in #202Full Changelog: v2.0.1...v2.1.0
Dropped Python 3.6 support, and made Python 3.11 support official. (#195, #200)
~w3lib.url.safe_url_string now generates safer URLs.
To make URLs safer for the URL living standard:
;= are percent-encoded in the URL username.
;:= are percent-encoded in the URL password.
' is percent-encoded in the URL query if the URL scheme is special.
To make URLs safer for RFC 2396 and RFC 3986, |[] are percent-encoded in URL paths, queries, and fragments.
(#80, #203)
~w3lib.encoding.html_to_unicode now checks for the byte order mark before inspecting the Content-Type header when determining the content encoding, in line with the URL living standard. (#189, #191)
~w3lib.url.canonicalize_url now strips spaces from the input URL, to be more in line with the URL living standard. (#132, #136)
~w3lib.html.get_base_url now ignores HTML comments. (#70, #77)
Fixed ~w3lib.url.safe_url_string re-encoding percent signs on the URL username and password even when they were being used as part of an escape sequence. (#187, #196)
Fixed ~w3lib.http.basic_auth_header using the wrong flavor of base64 encoding, which could prevent authentication in rare cases. (#181, #192)
Fixed ~w3lib.html.replace_entities raising OverflowError in some cases due to a bug in CPython. (#199, #202)
Improved typing and fixed typing issues. (#190, #206)
Made CI and test improvements. (#197, #198)
Adopted a Code of Conduct. (#194)
Backwards incompatible changes:
Backwards incompatible changes:
w3lib.url.safe_url_string and w3lib.url.canonicalize_urlw3lib.url.canonicalize_url is going to change, and so, ifDeprecation removals (#169):
w3lib.form module is removed.w3lib.html.remove_entities function is removed.w3lib.url.urljoin_rfc function is removed.The following functions are deprecated, and will be removed in future releases
(#170):
w3lib.util.str_to_unicodew3lib.util.unicode_to_strw3lib.util.to_native_strOther improvements and bug fixes:
w3lib.html.get_meta_refresh for <meta> tags wherehttp-equiv is written after content (#179).w3lib.url.safe_url_string for IDNA domains with ports (#174).w3lib.url.url_query_cleaner no longer adds an unneeded # whenkeep_fragments=True is passed, and the URL doesn't have a fragmentMinor documentation fix (release date is set in the changelog).
Backwards incompatible changes:
Backwards incompatible changes:
Python 2 is no longer supported; Python 3.6+ is required now (#168, #175).
w3lib.url.safe_url_string and w3lib.url.canonicalize_url no longer convert "%23" to "#" when it appears in the URL path. This is a bug fix. It's listed as a backward-incomatible change because in some cases the output of w3lib.url.canonicalize_url is going to change, and so, if this output is used to generate URL fingerprints, new fingerprints might be incompatible with those created with the previous w3lib versions (#141).
Deprecation removals (#169):
The w3lib.form module is removed.
The w3lib.html.remove_entities function is removed.
The w3lib.url.urljoin_rfc function is removed.
The following functions are deprecated, and will be removed in future releases (#170):
w3lib.util.str_to_unicode
w3lib.util.unicode_to_str
w3lib.util.to_native_str
Other improvements and bug fixes:
Type annotations are added (#172, #184).
Added support for Python 3.9 and 3.10 (#168, #176).
Fixed w3lib.html.get_meta_refresh for <meta> tags where http-equiv is written after content (#179).
Fixed w3lib.url.safe_url_string for IDNA domains with ports (#174).
w3lib.url.url_query_cleaner no longer adds an unneeded # when keep_fragments=True is passed, and the URL doesn't have a fragment (#159).
Removed a workaround for an ancient pathname2url bug (#142)
CI is migrated to GitHub Actions (#166, #177); other CI improvements (#160, #182).
The code is formatted using black (#173).
Python 3.4 is no longer supported (issue #156)
w3lib.url.safe_url_string now supports an optional quote_path
parameter to disable the percent-encoding of the URL path (issue #119)w3lib.url.add_or_replace_parameter and
w3lib.url.add_or_replace_parameters no longer remove duplicate
parameters from the original query string that are not being added or
replaced (issue #126)w3lib.html.remove_tags now raises a ValueError exception
instead of AssertionError when using both the which_ones and the
keep parameters (issue #154)Add the encoding and path_encoding parameters to w3lib.url.safe_download_url (issue #118)
Add the encoding and path_encoding parameters to w3lib.url.safe_download_url (issue #118)
w3lib.url.safe_url_string now also removes tabs and new lines (issue #133)
w3lib.html.remove_comments now also removes truncated comments (issue #129)
w3lib.html.remove_tags_with_content no longer removes tags which start with the same text as one of the specified tags (issue #114)
Recommend pytest instead of nose to run tests (issue #124)
Fix url_query_cleaner to do not append "?" to urls without a query string (issue #109)
w3lib.url.add_or_replace_parameters helper (issue #117)…w3lib.encoding functions. This is technically backwards incompatible because it changes the way non-decodable bytes are replaced (in some cases instea…
\ufffd chars you can get one).
As a side effect, the fix speeds up decoding in Python 3.4+.Include additional assets used for distribution packages in the source tarball
[ and ] as safe characters in path and query components
of URLs, i.e. they are not escaped anymoreAdd w3lib.url.parse_data_uri helper for parsing "data:" URIs
w3lib.url.parse_data_uri helper for parsing "data:" URIsw3lib.html.strip_html5_whitespace function to strip leading and
trailing whitespace as per W3C recommendations, e.g. for cleaning
"href" attribute valuesw3lib.http.headers_raw_to_dict for multiple headers with same namecanonicalize_url() and safe_url_string(): strip ":" when no port is specified (as per RFC 3986_; see also https://github.com/scrapy/scrapy/issues/2377
canonicalize_url() and safe_url_string():
strip ":" when no port is specified (as per RFC 3986_;
see also https://github.com/scrapy/scrapy/issues/2377)url_query_cleaner(): support new keep_fragments argument
(defaulting to False)Add `canonicalize_url()` to w3lib.url
canonicalize_url() to w3lib.urlHandle IDNA encoding failures in safe_url_string() (issue #62)
Bugfix release:
safe_url_string() (issue #62)fix function import for (deprecated) urljoin_rfc (issue #51)
Bugfix release:
urljoin_rfc (issue #51)w3lib.url, via __all__
(see issue #54, https://github.com/scrapy/scrapy/issues/1917)For bytes URLs, when supplied encoding (or default UTF8) is wrong, safe_url_string falls back to percent-encoding offending bytes.
Bugfix release:
safe_url_string falls back to percent-encoding offending bytes.proper handling of non-ASCII characters in Python2 and Python3
Changes to safe_url_string:
path_encoding to override default UTF-8 when serializing non-ASCII
characters before percent-encodinghtml_body_declared_encoding also detects encoding when not sole attribute in <meta>.
Package is now properly marked as zip_safe.
remove_tags removes uppercase tags as well;
meta_refresh regex now handles leading newlines and whitespaces in the url;
url_query_cleaner now supports str or list parameters;
- reverted all 1.9.0 changes.
all url-related functions accept bytes and unicode and now return bytes.
w3lib.http.basic_auth_header now returns bytes
add support for big5-hkscs encoding.
PY3 fixed headers_raw_to_dict and headers_dict_to_raw;
Nothing published for this version
w3lib.form.encode_multipart is deprecated;
- Python 2.6 support is dropped.
get_meta_refresh encoding handling is fixed;
support non-standard gb_2312_80 encoding;
Detect encoding for content attr before http-equiv in meta tag.
w3lib.url.urljoin_rfc is deprecated.
First release of w3lib.
First release of w3lib.
Your coding agent can read these notes before it upgrades. Set up the MCP server →