NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #91 most downloaded on PyPI
Powerful and Pythonic XML processing library combining libxml2/libxslt with the ElementTree API.
Last release 1 months ago
02 Sep 2026
Release timing varies
gaps range from 8 days to 7 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
19 years old
130 releases · first in 2007
LP#2143321: A NameError in doctestcompare when the doctests are processed by a non-standard module. Original patch by Rahul Kumar.
====================
LP#2143321: A NameError in doctestcompare when the doctests are processed
by a non-standard module.
Original patch by Rahul Kumar.
GH#492: Failures in htmldiff() to string-join the diff result are now propagated
instead of silently returning an empty result.
Patch by Juan Carlos Carvajal B.
* .find() and .findall() now take advantage of their read-locked tree to skip tree modification checks, making them visibly faster.
7.0.0a3 (2026-06-16)
====================
Features added
--------------
* ``.find()`` and ``.findall()`` now take advantage of their read-locked tree
to skip tree modification checks, making them visibly faster.
* Python Element proxy instantiation was tuned another bit to shave off more of the
performance disadvantages added by tree locking.
* The path compilation cache used by ElementPath uses an LRU eviction scheme.
It previously just cleared the whole cache when overflowing.
Bugs fixed
----------
* Attribute iteration now remembers a safe integer index instead of a tree attribute pointer
that is potentially subject of user side modifications. This may make attribute iteration
slightly slower for long sequences of attributes but allows for attribute manipulation
between iteration steps, giving similar iteration+manipulation semantics as Python lists.
* Several issues (race conditions, memory leaks on error handling, crashes)
were found via an automated review and subsequently resolved.
Review contributed by devdanzin.
One column per quarter.
* GH#511: The ElementPath implementation (.find() etc) was rewritten based on native traversal of the libxml2 tree instead of passing through Element
7.0.0a2 (2026-06-09)
====================
Features added
--------------
* GH#511: The ElementPath implementation (``.find()`` etc) was rewritten based on
native traversal of the libxml2 tree instead of passing through Element proxy
objects. This makes some path operations like child steps or predicate evaluation
several times faster. Indexing can be more than 100x faster.
Note that some operations are also slower due to the new lock usages in lxml 7.0.
Usually, the more selective the path, the faster it runs.
Original idea and implementation by mahomaho.
* GH#477,LP#2111289: Experimental support for freethreading Python is enabled in Python 3.14t and later.
7.0.0a1 (2026-05-29)
====================
Features added
--------------
* GH#477,LP#2111289: Experimental support for freethreading Python is enabled
in Python 3.14t and later.
* GH#466: The shared parser name dict is now local to a parser (as opposed to global),
which allows to control its lifetime and cross-document usage more easily.
It is now also unbounded in size if the ``huge_tree=True`` option is provided.
* The default chunk size for reading from file-likes in ``iterparse()`` was increased
from 32 KiB to 64 KiB and is now configurable with a new ``chunk_size`` argument.
* Writing to Python file objects minimises data copying by passing buffers around.
* GH#502: Predicate evaluation in ``ElementPath`` (``.find*()``) is faster.
* HTML diffing is faster, following optimisations in CPython 3.15.
* If multiple Python exceptions get raised from within an lxml operation, they no longer
overwrite old ones but get stacked in the exception context of the latest exception.
Bugs fixed
----------
* Reading from file-like objects that start with a UTF-32 BOM failed to detect the BOM.
Other changes
-------------
* Support for Python 3.8 was removed. Python 3.9 will continue to be supported for
several years due to its usage in LTS Linux distributions.
* Usage on PyPy can currently crash due to the Freethreading Python changes.
This is expected to get fixed in the further releases towards 7.0 final.
* Some internal adaptations were made for libxml2 2.14.x and 2.15.x.
* The ``xmlDict...`` C functions of libxml2 are now declared in ``tree.pxd``.
They remain in ``xmlparser.pxd`` for legacy reasons but Cython code that uses them
should migrate the import to ``tree.pxd``.
* The list of software licenses in LICENSES.txt was clarified and updated
to match the actual set of software being shipped.
* Built using Cython 3.2.5.
LP#2165901: External parameter entity parsing was allowed by default (with resolve_entities="internal" ). Issue found by Tomer Fichman.
==================
resolve_entities="internal").GH#526: Some build files were missing in the sdist. Patch by Nicola Soranzo.
==================
GH#526: Some build files were missing in the sdist.
Patch by Nicola Soranzo.
Some minor corrections for error handling cases.
* The Linux wheels use a patched libxslt 1.1.43, fixing CVE-2025-7424 and CVE-2025-11731 .
6.1.1 (2026-05-18)
==================
Bugs fixed
----------
* The known link attributes in ``lxml.html.defs.link_attrs`` were missing ``xlink:href``,
which can be used for URL bypass attacks in embedded SVG/MathML/etc. content.
GHSA-4jhm-jv67-739f
* The Linux wheels use a patched libxslt 1.1.43, fixing CVE-2025-7424 and CVE-2025-11731.
* The Windows wheels use libxslt 1.1.45, fixing CVE-2025-7424 and CVE-2025-11731.
This release fixes a possible external entity injection (XXE) vulnerability in iterparse() and the ETCompatXMLParser.
6.1.0 (2026-04-17)
==================
This release fixes a possible external entity injection (XXE) vulnerability in
``iterparse()`` and the ``ETCompatXMLParser``.
Features added
--------------
* GH#486: The HTML ARIA accessibility attributes were added to the set of safe attributes
in ``lxml.html.defs``. This allows ``lxml_html_clean`` to pass them through.
Patch by oomsveta.
* The default chunk size for reading from file-likes in ``iterparse()`` is now configurable
with a new ``chunk_size`` argument.
Bugs fixed
----------
* LP#2146291: The ``resolve_entities`` option was still set to ``True`` for
``iterparse`` and ``ETCompatXMLParser``, allowing for external entity injection (XXE)
when using these parsers without setting this option explicitly.
The default was now changed to ``'internal'`` only (as for the normal XML and HTML parsers
since lxml 5.0).
Issue found by Sihao Qiu as CVE-2026-41066.
* LP#2148019: Spurious MemoryError during namespace cleanup.
6.0.4 (2026-04-12)
==================
Bugs fixed
----------
* LP#2148019: Spurious MemoryError during namespace cleanup.
* Several out of memory error cases now raise MemoryError that were not handled before.
6.0.3 (2026-04-09)
==================
Bugs fixed
----------
* Several out of memory error cases now raise ``MemoryError`` that were not handled before.
* Slicing with large step values (outside of ``+/- sys.maxsize``) could trigger undefined C behaviour.
* LP#2125399: Some failing tests were fixed or disabled in PyPy.
* LP#2138421: Memory leak in error cases when setting the ``public_id`` or ``system_url`` of a document.
* Memory leak in case of a memory allocation failure when copying document subtrees.
* When mapping an XPath result to Python failed, the result memory could leak.
* When preparing an XSLT transform failed, the XSLT parameter memory could leak.
Other changes
-------------
* Built using Cython 3.2.4.
* Binary wheels use zlib 1.3.2.
* LP#2125278: Compilation with libxml2 2.15.0 failed. Original patch by Xi Ruoyao.
6.0.2 (2025-09-21)
==================
Bugs fixed
----------
* LP#2125278: Compilation with libxml2 2.15.0 failed.
Original patch by Xi Ruoyao.
* Setting ``decompress=True`` in the parser had no effect in libxml2 2.15.
* Binary wheels on Linux and macOS use the library version libxml2 2.14.6.
See https://gitlab.gnome.org/GNOME/libxml2/-/releases/v2.14.6
* Test failures in libxml2 2.15.0 were fixed.
Other changes
-------------
* Binary wheels for Py3.9-3.11 on the ``riscv64`` architecture were added.
* Error constants were updated to match libxml2 2.15.0.
* Built using Cython 3.1.4.
LP#2116333: lxml.sax._getNsTag() could fail with an exception on malformed input.
LP#2116333: lxml.sax._getNsTag() could fail with an exception on malformed input.
GH#467: Some test adaptations were made for libxml2 2.15. Patch by Nick Wellnhofer.
LP2119510, GH#473: A Python compatibility test was fixed for Python 3.14+. Patch by Lumír Balhar.
GH#471: Wheels for "riscv64" on recent Python versions were added. Patch by ffgan.
GH#469: The wheel build no longer requires the wheel package unconditionally.
Patch by Miro Hrončok.
Binary wheels use the library version libxml2 2.14.5. See https://gitlab.gnome.org/GNOME/libxml2/-/releases/v2.14.5
Windows binary wheels continue to use a security patched library version libxml2 2.11.9.
The Schematron class is deprecated and will become non-functional in a future lxml version. The feature will soon be removed from libxml2 and stop bei…
GH#463: lxml.html.diff is faster and provides structurally better diffs.
Original patch by Steven Fernandez.
GH#405: The factories Element and ElementTree can now be used in type hints.
GH#448: Parsing from memoryview and other buffers is supported to allow zero-copy parsing.
GH#437: lxml.html.builder was missing several HTML5 tag names.
Patch by Nick Tarleton.
GH#458: CDATA can now be written into the incremental xmlfile() writer.
Original patch by Lane Shaw.
A new parser option decompress=False was added that controls the automatic
input decompression when using libxml2 2.15.0 or later. Disabling this option
by default will effectively prevent decompression bombs when handling untrusted
input. Code that depends on automatic decompression must enable this option.
Note that libxml2 2.15.0 was not released yet, so this option currently has no
effect but can already be used.
The set of compile time / runtime supported libxml2 feature names is available as
etree.LIBXML_COMPILED_FEATURES and etree.LIBXML_FEATURES.
This currently includes
catalog, ftp, html, http, iconv, icu,
lzma, regexp, schematron, xmlschema, xpath, zlib.
GH#353: Predicates in .find*() could mishandle tag indices if a default namespace is provided.
Original patch by Luise K.
GH#272: The head and body properties of lxml.html elements failed if no such element
was found. They now return None instead.
Original patch by FVolral.
Tag names provided by code (API, not data) that are longer than INT_MAX
could be truncated or mishandled in other ways.
.text_content() on lxml.html elements accidentally returned a "smart string"
without additional information. It now returns a plain string.
LP#2109931: When building lxml with coverage reporting, it now disables the sys.monitoring
support due to the lack of support in https://github.com/nedbat/coveragepy/issues/1790
Support for Python < 3.8 was removed.
Parsing directly from zlib (or lzma) compressed data is now considered an optional feature in lxml. It may get removed from libxml2 at some point for security reasons (compression bombs) and is therefore no longer guaranteed to be available in lxml.
As of this release, zlib support is still normally available in the binary wheels
but may get disabled or removed in later (x.y.0) releases. To test the availability,
use "zlib" in etree.LIBXML_FEATURES.
The Schematron class is deprecated and will become non-functional in a future lxml version.
The feature will soon be removed from libxml2 and stop being available.
GH#438: Wheels include the arm7l target.
GH#465: Windows wheels include the arm64 target.
Patch by Finn Womack.
Binary wheels use the library versions libxml2 2.14.4 and libxslt 1.1.43.
Note that this disables direct HTTP and FTP support for parsing from URLs.
Use Python URL request tools instead (which usually also support HTTPS).
To test the availability, use "http" in etree.LIBXML_FEATURES.
Windows binary wheels use the library versions libxml2 2.11.9, libxslt 1.1.39 and libiconv 1.17. They are now based on VS-2022.
Built using Cython 3.1.2.
The debug methods MemDebug.dump() and MemDebug.show() were removed completely.
libxml2 2.13.0 discarded this feature.
LP#2107279: Binary wheels use libxml2 2.13.8 and libxslt 1.1.43 to resolve several CVEs. (Binary wheels for Windows continue to use a patched libxml2
This release resolves CVE-2025-24928 as described in https://gitlab.gnome.org/GNOME/libxml2/-/issues/847
This release resolves CVE-2025-24928 as described in https://gitlab.gnome.org/GNOME/libxml2/-/issues/847
Binary wheels use libxml2 2.12.10 and libxslt 1.1.42.
Binary wheels for Windows use a patched libxml2 2.11.9 and libxslt 1.1.39.
GH#450: iterparse() internally triggered the DeprecationWarning` added in lxml 5.3.0 when parsing HTML.
GH#440: Some tests were adapted for libxml2 2.14.0. Patch by Nick Wellnhofer.
LP#2097175: DTD(external_id="…") erroneously required a byte string as ID value.
GH#450: iterparse() internally triggered the `DeprecationWarning`` added in lxml 5.3.0 when parsing HTML.
-flat_namespace.LP#2067707: The strip_cdata option in HTMLParser() turned out to be useless and is now deprecated.
CDATA sections are no longer rejected but split on output
to represent ]]> correctly.
Patch by Gertjan Klein.LP#2060160: Attribute values serialised differently in xmlfile.element() and xmlfile.write().
LP#2058177: The ISO-Schematron implementation could fail on unknown prefixes. Patch by David Lakin.
LP#2067707: The strip_cdata option in HTMLParser() turned out to be useless and is now deprecated.
Binary wheels use the library versions libxml2 2.12.9 and libxslt 1.1.42.
Windows binary wheels use the library versions libxml2 2.11.8 and libxslt 1.1.39.
Built with Cython 3.0.11.
GH#417: The test_feed_parser test could fail if lxml_html_clean was not installed. It is now skipped in that case.
GH#417: The test_feed_parser test could fail if lxml_html_clean was not installed.
It is now skipped in that case.
LP#2059910: The minimum CPU architecture for the Linux x86 binary wheels was set back to "core2", without SSE 4.2.
If libxml2 uses iconv, the compile time version is available as etree.ICONV_COMPILED_VERSION.
LP#2059910: The minimum CPU architecture for the Linux x86 binary wheels was set back to "core2", but with SSE 4.2 enabled.
LP#2059910: The minimum CPU architecture for the Linux x86 binary wheels was set back to "core2", but with SSE 4.2 enabled.
LP#2059977: Element.iterfind("//absolute_path") failed with a SyntaxError
where it should have issued a warning.
GH#416: The documentation build was using the non-standard which command.
Patch by Michał Górny.
Projects that use lxml without "lxml.html.clean" will not notice any difference, except that they won't have potentially vulnerable code installed. Th…
LP#1958539: The lxml.html.clean implementation suffered from several (only if used)
security issues in the past and was now extracted into a separate library:
https://github.com/fedora-python/lxml_html_clean
Projects that use lxml without "lxml.html.clean" will not notice any difference, except that they won't have potentially vulnerable code installed. The module is available as an "extra" setuptools dependency "lxml[html_clean]", so that Projects that need "lxml.html.clean" will need to switch their requirements from "lxml" to "lxml[html_clean]", or install the new library themselves.
The minimum CPU architecture for the Linux x86 binary wheels was upgraded to "sandybridge" (launched 2011), and glibc 2.28 / gcc 12 (manylinux_2_28) wheels were added.
Built with Cython 3.0.10.
LP#2048920: iterlinks() in lxml.html rejected bytes input in 5.1.0.
LP#2048920: iterlinks() in lxml.html rejected bytes input in 5.1.0.
High source line numbers from the parser are no longer truncated
(up to a C long) when using libxml2 2.11 or later.
GH#407: A compatibility test was adapted to recent expat versions. Patch by Miro Hrončok.
Binary wheels use the library versions libxml2 2.12.6 and libxslt 1.1.39.
Windows binary wheels use the library versions libxml2 2.11.7 and libxslt 1.1.39.
Built with Cython 3.0.9.
Parsing ASCII strings is slightly faster.
Cleaner() interpreted an accidentally provided string parameter
for the host_whitelist as list of characters and silently failed to reject any hosts.
Passing a non-collection is now rejected.Support for Python 2.7 and Python versions < 3.6 was removed.
The wheel build was migrated to use cibuildwheel.
Patch by Primož Godec.
GH#407: A compatibility test was adapted to recent expat versions. Patch by Miro Hrončok.
GH#407: A compatibility test was adapted to recent expat versions. Patch by Miro Hrončok.
Binary wheels use the library versions libxml2 2.12.6 and libxslt 1.1.39.
Built with Cython 3.0.9.
LP#2046208: Parsing non-BMP Python Unicode strings could fail on macOS.
LP#2046208: Parsing non-BMP Python Unicode strings could fail on macOS.
LP#2044225: When incrementally parsing broken HTML, reporting start events on missing structural tags failed and could lead to subsequent exceptions.
LP#2045435: Some (not all) issues with stricter C compilers were resolved.
The binary wheels in the 5.0.0 release did not validate cleanly (but installed ok).
.. _latest_release:
GH#385: The long deprecated unittest.m̀akeSuite() function is no longer used. Patch by Miro Hrončok.
Character escaping in C14N2 serialisation now uses a single pass over the text
instead of searching for each unescaped character separately.
Early support for Python 3.13a2 was added.
LP#1976304: The Element.addnext() method previously inserted the new element
before existing tail text. The tail text of both sibling elements now stays on
the respective elements.
LP#1980767, GH#379: TreeBuilder.close() could fail with a TypeError after
parsing incorrect input. Original patch by Enrico Minack.
Element.itertext(with_tail=False) returned the tail text of comments and
processing instructions, despite the explicit option.
GH#370: A crash with recent libxml2 2.11.x versions was resolved. Patch by Michael Schlenker.
A compile problem with recent libxml2 2.12.x versions was resolved.
The internal exception handling in C callbacks was improved for Cython 3.0.
The exception declarations of xmlInputReadCallback, xmlInputCloseCallback,
xmlOutputWriteCallback and xmlOutputCloseCallback in tree.pxd were
corrected to prevent running Python code or calling into the C-API with a live
exception set.
GH#385: The long deprecated unittest.m̀akeSuite() function is no longer used.
Patch by Miro Hrončok.
LP#1522052: A file-system specific test is now optional and should no longer fail on systems that don't support it.
GH#392: Some tests were adapted for libxml2 2.13. Patch by Nick Wellnhofer.
Contains all fixes from lxml 4.9.4.
LP#1742885: lxml no longer expands external entities (XXE) by default to prevent
the security risk of loading arbitrary files and URLs. If this feature is needed,
it can be enabled in a backwards compatible way by using a parser with the option
resolve_entities=True. The new default is resolve_entities='internal'.
With libxml2 2.10.4 and later (as provided by the lxml 5.0 binary wheels),
parsing HTML tags with "prefixes" no longer builds a namespace dictionary
in nsmap but considers the prefix:name string the actual tag name.
With older libxml2 versions, since 2.9.11, the prefix was removed. Before
that, the prefix was parsed as XML prefix.
lxml 5.0 does not try to hide this difference but now changes the ElementPath
implementation to let element.find("part1:part2") search for the tag
part1:part2 in documents parsed as HTML, instead of looking only for part2.
LP#2024343: The validation of the schema file itself is now optional in the
ISO-Schematron implementation. This was done because some lxml distributions
discard the RNG validation schema file due to licensing issues. The validation
can now always be disabled with Schematron(..., validate_schema=False).
It is enabled by default if available and disabled otherwise. The module
constant lxml.isoschematron.schematron_schema_valid_supported can be used
to detect whether schema file validation is available.
Some redundant and long deprecated methods were removed:
parser.setElementClassLookup(),
xslt_transform.apply(),
xpath.evaluate().
Some incorrect declarations were removed from python.pxd. In general, this file
should not be used by external Cython code. Use the C-API declarations provided by
Cython itself instead.
Binary wheels use the library versions libxml2 2.12.3 and libxslt 1.1.39.
Built with Cython 3.0.7, updated to follow recent changes in Cython 3.1-dev.
LP#2046398: Inserting/replacing an ancestor into a node's children could loop indefinitely.
LP#2046398: Inserting/replacing an ancestor into a node's children could loop indefinitely.
LP#1980767, GH#379: TreeBuilder.close() could fail with a TypeError after
parsing incorrect input. Original patch by Enrico Minack.
LP#1522052: A file-system specific test is now optional and should no longer fail on systems that don't support it.
Wheels include zlib 1.3, libxml2 2.10.3 and libxslt 1.1.39 (zlib 1.2.12, libxml2 2.10.3 and libxslt 1.1.37 on Windows).
Built with Cython 0.29.37.
LP#2008911: lxml.objectify accepted non-decimal numbers like ²²² as integers.
LP#2008911: lxml.objectify accepted non-decimal numbers like ²²² as integers.
A memory leak in lxml.html.clean was resolved by switching to Cython 0.29.34+.
GH#348: URL checking in the HTML cleaner was improved. Patch by Tim McCormack.
GH#371, GH#373: Some regex strings were changed to raw strings to fix Python warnings. Patches by Jakub Wilk and Anthony Sottile.
Wheels include zlib 1.2.13, libxml2 2.10.3 and libxslt 1.1.38 (zlib 1.2.12, libxml2 2.10.3 and libxslt 1.1.37 on Windows).
Built with Cython 0.29.36 to adapt to changes in Python 3.12.
CVE-2022-2309: A Bug in libxml2 2.9.1[0-4] could let namespace declarations from a failed parser run leak into later parser runs. This bug was worked…
LP#1981760: Element.attrib now registers as collections.abc.MutableMapping.
lxml now has a static build setup for macOS on ARM64 machines (not used for building wheels). Patch by Quentin Leffray.
A crash was resolved when using iterwalk() (or canonicalize()) after parsing certain incorrect input. Note that iterwalk() can crash on *valid* input
iterwalk() (or canonicalize())
after parsing certain incorrect input. Note that iterwalk() can crash
on valid input parsed with the same parser after failing to parse the
incorrect input.GH#341: The mixin inheritance order in lxml.html was corrected. Patch by xmo-odoo.
lxml.html was corrected.
Patch by xmo-odoo.Built with Cython 0.29.30 to adapt to changes in Python 3.11 and 3.12.
Wheels include zlib 1.2.12, libxml2 2.9.14 and libxslt 1.1.35 (libxml2 2.9.12+ and libxslt 1.1.34 on Windows).
GH#343: Windows-AArch64 build support in Visual Studio. Patch by Steve Dower.
GH#337: Path-like objects are now supported throughout the API instead of just strings. Patch by Henning Janssen.
GH#337: Path-like objects are now supported throughout the API instead of just strings. Patch by Henning Janssen.
The ElementMaker now supports QName values as tags, which always override
the default namespace of the factory.
Chunked Unicode string parsing via parser.feed() now encodes the input data to the native UTF-8 encoding directly, instead of going through Py_UNICODE
parser.feed() now encodes the input data
to the native UTF-8 encoding directly, instead of going through Py_UNICODE /
wchar_t encoding first, which previously required duplicate recoding in most cases.The standard namespace prefixes were mishandled during "C14N2" serialisation on Python 3. See https://mail.python.org/archives/list/lxml@python.org/thread/6ZFBHFOVHOS5GFDOAMPCT6HM5HZPWQ4Q/
lxml.objectify previously accepted non-XML numbers with underscores (like "1_000")
as integers or float values in Python 3.6 and later. It now adheres to the number
format of the XML spec again.
LP#1939031: Static wheels of lxml now contain the header files of zlib and libiconv (in addition to the already provided headers of libxml2/libxslt/libexslt).
A vulnerability (GHSL-2021-1038) in the HTML cleaner allowed sneaking script content through SVG images (CVE-2021-43818).
A vulnerability (GHSL-2021-1038) in the HTML cleaner allowed sneaking script content through SVG images (CVE-2021-43818).
A vulnerability (GHSL-2021-1037) in the HTML cleaner allowed sneaking script content through CSS imports and other crafted constructs (CVE-2021-43818).
GH#317: A new property system_url was added to DTD entities. Patch by Thirdegree.
GH#317: A new property system_url was added to DTD entities.
Patch by Thirdegree.
GH#314: The STATIC_* variables in setup.py can now be passed via env vars.
Patch by Isaac Jurado.
A vulnerability (CVE-2021-28957) was discovered in the HTML Cleaner by Kevin Chung, which allowed JavaScript to pass through. The cleaner now removes…
formaction attribute.A vulnerability (CVE-2020-27783) was discovered in the HTML Cleaner by Yaniv Nizry, which allowed JavaScript to pass through. The cleaner now removes…
A vulnerability was discovered in the HTML Cleaner by Yaniv Nizry, which allowed JavaScript to pass through. The cleaner now removes more sneaky "styl…
GH#310: lxml.html.InputGetter supports __len__() to count the number of input fields. Patch by Aidan Woolley.
GH#310: lxml.html.InputGetter supports __len__() to count the number of input fields.
Patch by Aidan Woolley.
lxml.html.InputGetter has a new .items() method to ease processing all input fields.
lxml.html.InputGetter.keys() now returns the field names in document order.
GH-309: The API documentation is now generated using sphinx-apidoc.
Patch by Chris Mayo.
LP#1869455: C14N 2.0 serialisation failed for unprefixed attributes when a default namespace was defined.
TreeBuilder.close() raised AssertionError in some error cases where it
should have raised XMLSyntaxError. It now raises a combined exception to
keep up backwards compatibility, while switching to XMLSyntaxError as an
interface.
Cleaner() now validates that only known configuration options can be set.
Cleaner() now validates that only known configuration options can be set.
LP#1882606: Cleaner.clean_html() discarded comments and PIs regardless of the
corresponding configuration option, if remove_unknown_tags was set.
LP#1880251: Instead of globally overwriting the document loader in libxml2, lxml now sets it per parser run, which improves the interoperability with other users of libxml2 such as libxmlsec.
LP#1881960: Fix build in CPython 3.10 by using Cython 0.29.21.
The setup options "--with-xml2-config" and "--with-xslt-config" were accidentally renamed to "--xml2-config" and "--xslt-config" in 4.5.1 and are now available again.
LP#1570388: Fix failures when serialising documents larger than 2GB in some cases.
LP#1570388: Fix failures when serialising documents larger than 2GB in some cases.
LP#1865141, GH#298: QName values were not accepted by the el.iter() method.
Patch by xmo-odoo.
LP#1863413, GH#297: The build failed to detect libraries on Linux that are only configured via pkg-config. Patch by Hugh McMaster.
A new function indent() was added to insert tail whitespace for pretty-printing an XML tree.
indent() was added to insert tail whitespace for pretty-printing
an XML tree.MacOS builds are 64-bit-only by default. Set CFLAGS and LDFLAGS explicitly to override it.
Linux/MacOS Binary wheels now use libxml2 2.9.10 and libxslt 1.1.34.
LP#1840234: The package version number is now available as lxml.__version__.
LP#1844674: itertext() was missing tail text of comments and PIs since 4.4.0.
itertext() was missing tail text of comments and PIs since 4.4.0.LP#1835708: ElementInclude incorrectly rejected repeated non-recursive includes as recursive. Patch by Rainer Hausdorf.
ElementInclude incorrectly rejected repeated non-recursive
includes as recursive.
Patch by Rainer Hausdorf.LP#1838252: The order of an OrderedDict was lost in 4.4.0 when passing it as attrib mapping during element creation.
LP#1838252: The order of an OrderedDict was lost in 4.4.0 when passing it as attrib mapping during element creation.
LP#1838521: The package metadata now lists the supported Python versions.
The ElementTree.write_c14n() method has been deprecated in favour of the long preferred ElementTree.write(f, method="c14n"). It will be removed in a f…
Element.clear() accepts a new keyword argument keep_tail=True to clear
everything but the tail text. This is helpful in some document-style use cases
and for clearing the current element in iterparse() and pull parsing.
When creating attributes or namespaces from a dict in Python 3.6+, lxml now preserves the original insertion order of that dict, instead of always sorting the items by name. A similar change was made for ElementTree in CPython 3.8. See https://bugs.python.org/issue34160
Integer elements in lxml.objectify implement the __index__() special method.
GH#269: Read-only elements in XSLT were missing the nsmap property.
Original patch by Jan Pazdziora.
ElementInclude can now restrict the maximum inclusion depth via a max_depth
argument to prevent content explosion. It is limited to 6 by default.
The target object of the XMLParser can have start_ns() and end_ns()
callback methods to listen to namespace declarations.
The TreeBuilder has new arguments comment_factory and pi_factory to
pass factories for creating comments and processing instructions, as well as
flag arguments insert_comments and insert_pis to discard them from the
tree when set to false.
A C14N 2.0 <https://www.w3.org/TR/xml-c14n2/>_ implementation was added as
etree.canonicalize(), a corresponding C14NWriterTarget class, and
a c14n2 serialisation method.
When writing to file paths that contain the URL escape character '%', the file path could wrongly be mangled by URL unescaping and thus write to a different file or directory. Code that writes to file paths that are provided by untrusted sources, but that must work with previous versions of lxml, should best either reject paths that contain '%' characters, or otherwise make sure that the path does not contain maliciously injected '%XX' URL hex escapes for paths like '../'.
Assigning to Element child slices with negative step could insert the slice at the wrong position, starting too far on the left.
Assigning to Element child slices with overly large step size could take very long, regardless of the length of the actual slice.
Assigning to Element child slices of the wrong size could sometimes fail to raise a ValueError (like a list assignment would) and instead assign outside of the original slice bounds or leave parts of it unreplaced.
The comment and pi events in iterwalk() were never triggered, and
instead, comments and processing instructions in the tree were reported as
start elements. Also, when walking an ElementTree (as opposed to its root
element), comments and PIs outside of the root element are now reported.
LP#1827833: The RelaxNG compact syntax support was broken with recent versions
of rnc2rng.
LP#1758553: The HTML elements source and track were added to the list
of empty tags in lxml.html.defs.
Registering a prefix other than "xml" for the XML namespace is now rejected.
Failing to write XSLT output to a file could raise a misleading exception.
It now raises IOError.
Support for Python 3.4 was removed.
When using Element.find*() with prefix-namespace mappings, the empty string
is now accepted to define a default namespace, in addition to the previously
supported None prefix. Empty strings are more convenient since they keep
all prefix keys in a namespace dict strings, which simplifies sorting etc.
The ElementTree.write_c14n() method has been deprecated in favour of the
long preferred ElementTree.write(f, method="c14n"). It will be removed
in a future release.
Rebuilt with Cython 0.29.13 to support Python 3.8.
Rebuilt with Cython 0.29.10 to support Python 3.8.
Fix leak of output buffer and unclosed files in _XSLTResultTree.write_output().
_XSLTResultTree.write_output().Crash in 4.3.1 when appending a child subtree with certain text nodes.
The module lxml.sax is compiled using Cython in order to speed it up.
The module lxml.sax is compiled using Cython in order to speed it up.
GH#267: lxml.sax.ElementTreeProducer now preserves the namespace prefixes.
If two prefixes point to the same URI, the first prefix in alphabetical order
is used. Patch by Lennart Regebro.
Updated ISO-Schematron implementation to 2013 version (now MIT licensed) and the corresponding schema to the 2016 version (with optional "properties").
GH#270, GH#271: Support for Python 2.6 and 3.3 was removed. Patch by hugovk.
The minimum dependency versions were raised to libxml2 2.9.2 and libxslt 1.1.27, which were released in 2014 and 2012 respectively.
Built with Cython 0.29.2.
LP#1799755: Fix a DeprecationWarning in Py3.7+.
LP#1799755: Fix a DeprecationWarning in Py3.7+.
Import warnings in Python 3.6+ were resolved.
Javascript URLs that used URL escaping were not removed by the HTML cleaner. Security problem found by Omar Eissa. (CVE-2018-19787)
GH#259: Allow using pkg-config for build configuration. Patch by Patrick Griffis.
pkg-config for build configuration.
Patch by Patrick Griffis.Element.insert().
Patch by Alexander Weggerle.Reverted GH#265: lxml links against zlib as a shared library again.
GH#266: Fix sporadic crash during GC when parse-time schema validation is used and the parser participates in a reference cycle. Original patch by Jul
GH#266: Fix sporadic crash during GC when parse-time schema validation is used and the parser participates in a reference cycle. Original patch by Julien Greard.
GH#265: lxml no longer links against zlib as a shared library, only on static builds. Patch by Nehal J Wani.
LP#1755825: iterwalk() failed to return the 'start' event for the initial element if a tag selector is used.
LP#1755825: iterwalk() failed to return the 'start' event for the initial
element if a tag selector is used.
LP#1756314: Failure to import 4.2.0 into PyPy due to a missing library symbol.
LP#1727864, GH#258: Add "-isysroot" linker option on MacOS as needed by XCode 9.
GH#255: SelectElement.value returns more standard-compliant and browser-like defaults for non-multi-selects. If no option is selected, the value of th
GH#255: SelectElement.value returns more standard-compliant and
browser-like defaults for non-multi-selects. If no option is selected, the
value of the first option is returned (instead of None). If multiple options
are selected, the value of the last one is returned (instead of that of the
first one). If no options are present (not standard-compliant)
SelectElement.value still returns None.
GH#261: The HTMLParser() now supports the huge_tree option.
Patch by stranac.
LP#1551797: Some XSLT messages were not captured by the transform error log.
LP#1737825: Crash at shutdown after an interrupted iterparse run with XMLSchema validation.
Rebuild with Cython 0.27.3 to improve support for Py3.7.
ElementPath supports text predicates for current node, like "[.='text']".
ElementPath supports text predicates for current node, like "[.='text']".
ElementPath allows spaces in predicates.
Custom Element classes and XPath functions can now be registered with a decorator rather than explicit dict assignments.
Static Linux wheels are now built with link time optimisation (LTO) enabled. This should have a beneficial impact on the overall performance by providing a tighter compiler integration between lxml and libxml2/libxslt.
PythonElementClassLookup could fail with a TypeError.Note: This is a backwards incompatible change of the default configuration. If your code parses byte strings/streams and depends on character detectio…
The ElementPath implementation is now compiled using Cython,
which speeds up the .find*() methods quite significantly.
The modules lxml.builder, lxml.html.diff and lxml.html.clean
are also compiled using Cython in order to speed them up.
xmlfile() supports async coroutines using async with and await.
iterwalk() has a new method skip_subtree() that prevents walking into
the descendants of the current element.
RelaxNG.from_rnc_string() accepts a base_url argument to
allow relative resource lookups.
The XSLT result object has a new method .write_output(file) that serialises
output data into a file according to the <xsl:output> configuration.
GH#251: HTML comments were handled incorrectly by the soupparser. Patch by mozbugbox.
LP#1654544: The html5parser no longer passes the useChardet option
if the input is a Unicode string, unless explicitly requested. When parsing
files, the default is to enable it when a URL or file path is passed (because
the file is then opened in binary mode), and to disable it when reading from
a file(-like) object.
Note: This is a backwards incompatible change of the default configuration.
If your code parses byte strings/streams and depends on character detection,
please pass the option guess_charset=True explicitly, which already worked
in older lxml versions.
LP#1703810: etree.fromstring() failed to parse UTF-32 data with BOM.
LP#1526522: Some RelaxNG errors were not reported in the error log.
LP#1567526: Empty and plain text input raised a TypeError in soupparser.
LP#1710429: Uninitialised variable usage in HTML diff.
LP#1415643: The closing tags context manager in xmlfile() could continue
to output end tags even after writing failed with an exception.
LP#1465357: xmlfile.write() now accepts and ignores None as input argument.
Compilation under Py3.7-pre failed due to a modified function signature.
lxml.*.pyx to plain
*.pyx (e.g. etree.pyx) to simplify their handling in the build
process. Care was taken to keep the old header files as fallbacks for
code that compiles against the public C-API of lxml, but it might still
be worth validating that third-party code does not notice this change.Your coding agent can read these notes before it upgrades. Set up the MCP server →