NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #109 most downloaded on PyPI
Powerful and Pythonic XML processing library combining libxml2/libxslt with the ElementTree API.
Last release 27 days ago
02 Sep 2026
Release timing varies
gaps range from 8 days to 7 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
19 years old
130 releases · first in 2007
The previously undocumented docstring option in ElementTree.write() produces a deprecation warning and will eventually be removed.
ElementTree.write() has a new option doctype that writes out a
doctype string before the serialisation, in the same way as tostring().
GH#220: xmlfile allows switching output methods at an element level.
Patch by Burak Arslan.
LP#1595781, GH#240: added a PyCapsule Python API and C-level API for passing externally generated libxml2 documents into lxml.
GH#244: error log entries have a new property path with an XPath
expression (if known, None otherwise) that points to the tree element
responsible for the error. Patch by Bob Kline.
The namespace prefix mapping that can be used in ElementPath now injects a default namespace when passing a None prefix.
GH#238: Character escapes were not hex-encoded in the xmlfile serialiser.
Patch by matejcik.
GH#229: fix for externally created XML documents. Patch by Theodore Dubois.
LP#1665241, GH#228: Form data handling in lxml.html no longer strips the option values specified in form attributes but only the text values. Patch by Ashish Kulkarni.
LP#1551797: revert previous fix for XSLT error logging as it breaks multi-threaded XSLT processing.
LP#1673355, GH#233: fromstring() html5parser failed to parse byte strings.
docstring option in ElementTree.write()
produces a deprecation warning and will eventually be removed.One column per quarter.
GH#218 was ineffective in Python 3.
GH#218 was ineffective in Python 3.
GH#222: lxml.html.submit_form() failed in Python 3.
Patch by Jakub Wilk.
GH#220: xmlfile allows switching output methods at an element level. Patch by Burak Arslan.
xmlfile allows switching output methods at an element level.
Patch by Burak Arslan.Work around installation problems in recent Python 2.7 versions due to FTP download failures.
GH#219: xmlfile.element() was not properly quoting attribute values.
Patch by Burak Arslan.
GH#218: xmlfile.element() was not properly escaping text content of
script/style tags. Patch by Burak Arslan.
No source changes, issued only to solve problems with the binary packages released for 3.7.0.
GH#217: XMLSyntaxError now behaves more like its SyntaxError baseclass. Patch by Philipp A.
GH#217: XMLSyntaxError now behaves more like its SyntaxError
baseclass. Patch by Philipp A.
GH#216: HTMLParser() now supports the same collect_ids parameter
as XMLParser(). Patch by Burak Arslan.
GH#210: Allow specifying a serialisation method in xmlfile.write().
Patch by Burak Arslan.
GH#203: New option default_doctype in HTMLParser that allows
disabling the automatic doctype creation. Patch by Shadab Zafar.
GH#201: Calling the method .set('attrname') without value argument
(or None) on HTML elements creates an attribute without value that
serialises like <div attrname></div>. Patch by Daniel Holth.
GH#197: Ignore form input fields in form_values() when they are
marked as disabled in HTML. Patch by Kristian Klemon.
Log entries no longer allow anything but plain string objects as message text and file name.
zlib is included in the list of statically built libraries.
GH#204, LP#1614693: build fix for MacOS-X.
LP#1614603: change linker flags to build multi-linux wheels
LP#1614603: release without source changes to provide cleanly built Linux wheels
GH#180: Separate option inline_style for Cleaner that only removes style attributes instead of all styles. Patch by Christian Pedersen.
GH#180: Separate option inline_style for Cleaner that only removes style
attributes instead of all styles. Patch by Christian Pedersen.
GH#196: Windows build support for Python 3.5. Contribution by Maximilian Hils.
GH#199: Exclude file fields from FormElement.form_values (as browsers do).
Patch by Tomas Divis.
GH#198, LP#1568167: Try to provide base URL from Resolver.resolve_string().
Patch by Michael van Tellingen.
GH#191: More accurate float serialisation in objectify.FloatElement.
Patch by Holger Joukl.
LP#1551797: Repair XSLT error logging. Patch by Marcus Brinkmann.
GH#187: Now supports (only) version 5.x and later of PyPy. Patch by Armin Rigo.
GH#187: Now supports (only) version 5.x and later of PyPy. Patch by Armin Rigo.
GH#181: Direct support for .rnc files in RelaxNG() if rnc2rng
is installed. Patch by Dirkjan Ochtman.
GH#189: Static builds honour FTP proxy configurations when downloading the external libs. Patch by Youhei Sakurai.
GH#186: Soupparser failed to process entities in Python 3.x. Patch by Duncan Morris.
GH#185: Rare encoding related TypeError on import was fixed.
Patch by Petr Demin.
Deprecated API usage in doctestcompare.
Unicode string results failed XPath queries in PyPy.
LP#1497051: HTML target parser failed to terminate on exceptions and continued parsing instead.
Deprecated API usage in doctestcompare.
An ElementTree compatibility test added in lxml 3.4.3 that failed in Python 3.4+ was removed again.
Expression cache in ElementPath was ignored. Fix by Changaco.
Expression cache in ElementPath was ignored. Fix by Changaco.
LP#1426868: Passing a default namespace and a prefixed namespace mapping
as nsmap into xmlfile.element() raised a TypeError.
LP#1421927: DOCTYPE system URLs were incorrectly quoted when containing double quotes. Patch by Olli Pottonen.
LP#1419354: meta-redirect URLs were incorrectly processed by
iterlinks() if preceded by whitespace.
LP#1415907: Crash when creating an XMLSchema from a non-root element of an XML document.
LP#1415907: Crash when creating an XMLSchema from a non-root element of an XML document.
LP#1369362: HTML cleaning failed when hitting processing instructions with pseudo-attributes.
CDATA() wrapped content was rejected for tail text.
CDATA sections were not serialised as tail text of the top-level element.
New htmlfile HTML generator to accompany the incremental xmlfile serialisation API. Patch by Burak Arslan.
htmlfile HTML generator to accompany the incremental xmlfile
serialisation API. Patch by Burak Arslan.lxml.sax.ElementTreeContentHandler did not initialise its superclass.xmlfile(buffered=False) disables output buffering and flushes the content after each API operation (starting/ending element blocks or writes). A new m
xmlfile(buffered=False) disables output buffering and flushes the
content after each API operation (starting/ending element blocks or writes).
A new method xf.flush() can alternatively be used to explicitly flush
the output.
lxml.html.document_fromstring has a new option ensure_head_body=True
which will add an empty head and/or body element to the result document if
missing.
lxml.html.iterlinks now returns links inside meta refresh tags.
New XMLParser option collect_ids=False to disable ID hash table
creation. This can substantially speed up parsing of documents with many
different IDs that are not used.
The parser uses per-document hash tables for XML IDs. This reduces the load of the global parser dict and speeds up parsing for documents with many different IDs.
ElementTree.getelementpath(element) returns a structural ElementPath
expression for the given element, which can be used for lookups later.
xmlfile() accepts a new argument close=True to close file(-like)
objects after writing to them. Before, xmlfile() only closed the file
if it had opened it internally.
Allow "bytearray" type for ASCII text input.
LP#400588: decoding errors have become hard errors even in recovery mode. Previously, they could lead to an internal tree representation in a mixed encoding state, which lead to very late errors or even silently incorrect behaviour during tree traversal or serialisation.
Requires Python 2.6, 2.7, 3.2 or later. No longer supports Python 2.4, 2.5 and 3.1, use lxml 3.3.x for those.
Requires libxml2 2.7.0 or later and libxslt 1.1.23 or later, use lxml 3.3.x with older versions.
Prevent tree cycle creation when adding Elements as siblings.
Prevent tree cycle creation when adding Elements as siblings.
LP#1361948: crash when deallocating Element siblings without parent.
LP#1354652: crash when traversing internally loaded documents in XSLT extension functions.
HTML cleaning could fail to strip javascript links that mix control characters into the link scheme.
Source line numbers above 65535 are available on Elements when using libxml2 2.9 or later.
lxml.html.fragment_fromstring() failed for bytes input in Py3.LP#1287118: Crash when using Element subtypes with __slots__.
__slots__._LogEntry and _Attrib can no longer be
subclassed from Python code.The properties resolvers and version, as well as the methods set_element_class_lookup() and makeelement(), were lost from iterparse objects in 3.3.0.
The properties resolvers and version, as well as the methods
set_element_class_lookup() and makeelement(), were lost from
iterparse objects in 3.3.0.
LP#1222132: instances of XMLSchema, Schematron and RelaxNG
did not clear their local error_log before running a validation.
LP#1238500: lxml.doctestcompare mixed up "expected" and "actual" in attribute values.
Some file I/O tests were failing in MS-Windows due to non-portable temp file usage. Initial patch by Gabi Davar.
LP#910014: duplicate IDs in a document were not reported by DTD validation.
LP#1185332: tostring(method="html") did not use HTML serialisation
semantics for trailing tail text. Initial patch by Sylvain Viollon.
LP#1281139: .attrib value of Comments lost its mutation methods
in 3.3.0. Even though it is empty and immutable, it should still
provide the same interface as that returned for Elements.
LP#1014290: HTML documents parsed with parser.feed() failed to find elements during tag iteration.
LP#1014290: HTML documents parsed with parser.feed() failed to find
elements during tag iteration.
LP#1273709: Building in PyPy failed due to missing support for
PyUnicode_Compare() and PyByteArray_*() in PyPy's C-API.
LP#1274413: Compilation in MSVC failed due to missing "stdint.h" standard header file.
LP#1274118: iterparse() failed to parse BOM prefixed files.
The heuristic that distinguishes file paths from URLs was tightened to produce less false negatives.
Crash in xmlfile() when closing open elements out of order in an error case.
Crash in xmlfile() when closing open elements out of order in an error case.
Crash in target parsing with attributes.
LP#1255132: crash when trying to run validation over non-Element (e.g. comment or PI).
Memory leak when creating an XPath evaluator in a thread.
Memory leak when creating an XPath evaluator in a thread.
LP#1228881: repr(XSLTAccessControl) failed in Python 3.
Raise ValueError when trying to append an Element to itself or
to one of its own descendants.
LP#1206077: htmldiff discarded whitespace from the output.
Compressed plain-text serialisation to file-like objects was broken.
Fix support for Python 2.4 which was lost in 3.2.2.
LP#1185701: spurious XMLSyntaxError after finishing iterparse().
LP#1185701: spurious XMLSyntaxError after finishing iterparse().
Crash in lxml.objectify during xsi annotation.
The methods apply_templates() and process_children() of XSLT extension elements have gained two new boolean options elements_only and remove_blank_tex
apply_templates() and process_children() of XSLT
extension elements have gained two new boolean options elements_only
and remove_blank_text that discard either all strings or whitespace-only
strings from the result list.When moving Elements to another tree, the namespace cleanup mechanism no longer drops namespace prefixes from attributes for which it finds a default namespace declaration, to prevent them from appearing as unnamespaced attributes after serialisation.
Returning non-type objects from a custom class lookup method could lead to a crash.
Instantiating and using subtypes of Comments and ProcessingInstructions crashed.
LP#690319: Leading whitespace could change the behaviour of the string parsing functions in lxml.html.
LP#690319: Leading whitespace could change the behaviour of the string
parsing functions in lxml.html.
LP#599318: The string parsing functions in lxml.html are more robust
in the face of uncommon HTML content like framesets or missing body tags.
Patch by Stefan Seelmann.
LP#712941: I/O errors while trying to access files with paths that contain
non-ASCII characters could raise UnicodeDecodeError instead of properly
reporting the IOError.
LP#673205: Parsing from in-memory strings disabled network access in the default parser and made subsequent attempts to parse from a URL fail.
LP#971754: lxml.html.clean appends 'nofollow' to 'rel' attributes instead of overwriting the current value.
LP#715687: lxml.html.clean no longer discards scripts that are explicitly allowed by the user provided whitelist. Patch by Christine Koppelt.
LP#1136509: Passing attributes through the namespace-unaware API of the sax bridge (i.e. the handler.startElement() method) failed with a TypeError. P
LP#1136509: Passing attributes through the namespace-unaware API of
the sax bridge (i.e. the handler.startElement() method) failed
with a TypeError. Patch by Mike Bayer.
LP#1123074: Fix serialisation error in XSLT output when converting the result tree to a Unicode string.
GH#105: Replace illegal usage of xmlBufLength() in libxml2 2.9.0
by properly exported API function xmlBufUse().
LP#1160386: Write access to lxml.html.FormElement.fields raised an AttributeError in Py3.
LP#1160386: Write access to lxml.html.FormElement.fields raised
an AttributeError in Py3.
Illegal memory access during cleanup in incremental xmlfile writer.
lxml.etree._BaseParser was removed
from the module dict.GH#89: lxml.html.clean allows overriding the set of attributes that it considers 'safe'. Patch by Francis Devereux.
LP#1104370: copy.copy(el.attrib) raised an exception. It now returns
a copy of the attributes as a plain Python dict.
GH#95: When used with namespace prefixes, the el.find*() methods
always used the first namespace mapping that was provided for each
path expression instead of using the one that was actually passed
in for the current run.
LP#1092521, GH#91: Fix undefined C symbol in Python runtimes compiled without threading support. Patch by Ulrich Seidl.
Fix crash during interpreter shutdown by switching to Cython 0.17.3 for building.
End-of-file handling was incorrect in iterparse() when reading from a low-level C file stream and failed in libxml2 2.9.0 due to its improved consiste
Passing long Unicode strings into the feed() parser interface failed to read the entire string.
feed() parser interface
failed to read the entire string.Crash when merging text nodes in element.remove().
Crash when merging text nodes in element.remove().
Crash in sax/target parser when reporting empty doctype.
Crash when building an nsmap (Element property) with empty namespace URIs.
Crash when building an nsmap (Element property) with empty namespace URIs.
Crash due to race condition when errors (or user messages) occur during threaded XSLT processing.
XSLT stylesheet compilation could ignore compilation errors.
lxml.html.tostring() gained new serialisation options with_tail and doctype.
lxml.html.tostring() gained new serialisation options
with_tail and doctype.Fixed a crash when using iterparse() for HTML parsing and
requesting start events.
Fixed parsing of more selectors in cssselect. Whitespace before pseudo-elements and pseudo-classes is significant as it is a descendant combinator. "E :pseudo" should parse the same as "E *:pseudo", not "E:pseudo". Patch by Simon Sapin.
lxml.html.diff no longer raises an exception when hitting 'img' tags without 'src' attribute.
lxml.objectify.deannotate() has a new boolean option cleanup_namespaces to remove the objectify namespace declarations (and generally clean up the nam
lxml.objectify.deannotate() has a new boolean option
cleanup_namespaces to remove the objectify namespace
declarations (and generally clean up the namespace declarations)
after removing the type annotations.
lxml.objectify gained its own SubElement() function as a
copy of etree.SubElement to avoid an otherwise redundant import
of lxml.etree on the user side.
Fixed the "descendant" bug in cssselect a second time (after a first fix in lxml 2.3.1). The previous change resulted in a serious performance regression for the XPath based evaluation of the translated expression. Note that this breaks the usage of some of the generated XPath expressions as XSLT location paths that previously worked in 2.3.1.
Fixed parsing of some selectors in cssselect. Whitespace after combinators ">", "+" and "~" is now correctly ignored. Previously it was parsed as a descendant combinator. For example, "div> .foo" was parsed the same as "div>* .foo" instead of "div>.foo". Patch by Simon Sapin.
New option kill_tags in lxml.html.clean to remove specific tags and their content (i.e. their whole subtree).
New option kill_tags in lxml.html.clean to remove specific
tags and their content (i.e. their whole subtree).
pi.get() and pi.attrib on processing instructions to parse
pseudo-attributes from the text content of processing instructions.
lxml.get_include() returns a list of include paths that can be
used to compile external C code against lxml.etree. This is
specifically required for statically linked lxml builds when code
needs to compile against the exact same header file versions as lxml
itself.
Resolver.resolve_file() takes an additional option
close_file that configures if the file(-like) object will be
closed after reading or not. By default, the file will be closed,
as the user is not expected to keep a reference to it.
HTML cleaning didn't remove 'data:' links.
The html5lib parser integration now uses the 'official' implementation in html5lib itself, which makes it work with newer releases of the library.
In lxml.sax, endElementNS() could incorrectly reject a plain
tag name when the corresponding start event inferred the same plain
tag name to be in the default namespace.
When an open file-like object is passed into parse() or
iterparse(), the parser will no longer close it after use. This
reverts a change in lxml 2.3 where all files would be closed. It is
the user's responsibility to properly close the file(-like) object,
also in error cases.
Assertion error in lxml.html.cleaner when discarding top-level elements.
In lxml.cssselect, use the xpath 'A//B' (short for
'A/descendant-or-self::node()/B') instead of 'A/descendant::B' for
the css descendant selector ('A B'). This makes a few edge cases
like "div *:last-child" consistent with the selector behavior in
WebKit and Firefox, and makes more css expressions valid location
paths (for use in xsl:template match).
In lxml.html, non-selected <option> tags no longer show up in the
collected form values.
Adding/removing <option> values to/from a multiple select form
field properly selects them and unselects them.
--download-dir option.When looking for children, lxml.objectify takes '{}tag' as meaning an empty namespace, as opposed to the parent namespace.
lxml.objectify takes '{}tag' as
meaning an empty namespace, as opposed to the parent namespace.When finished reading from a file-like object, the parser
immediately calls its .close() method.
When finished parsing, iterparse() immediately closes the input
file.
Work-around for libxml2 bug that can leave the HTML parser in a non-functional state after parsing a severely broken document (fixed in libxml2 2.7.8).
marque tag in HTML cleanup code is correctly named marquee.
Crash in newer libxml2 versions when moving elements between documents that had attributes on replaced XInclude nodes.
Crash in newer libxml2 versions when moving elements between documents that had attributes on replaced XInclude nodes.
Import fix for urljoin in Python 3.1+.
Crash in XSLT when generating text-only result documents with a stylesheet created in a different thread.
Fixed several Python 3 regressions by building with Cython 0.11.3.
Support for running XSLT extension elements on the input root node (e.g. in a template matching on "/").
Crash in XPath evaluation when reading smart strings from a document other than the original context document.
Support recent versions of html5lib by not requiring its
XHTMLParser in htmlparser.py anymore.
Manually instantiating the custom element classes in
lxml.objectify could crash.
Invalid XML text characters were not rejected by the API when they appeared in unicode strings directly after non-ASCII characters.
lxml.html.open_http_urllib() did not work in Python 3.
The functions strip_tags() and strip_elements() in
lxml.etree did not remove all occurrences of a tag in all cases.
Crash in XSLT extension elements when the XSLT context node is not an element.
Static build of libxml2/libxslt was broken.
The resolve_entities option did not work in the incremental feed parser.
The resolve_entities option did not work in the incremental feed
parser.
Looking up and deleting attributes without a namespace could hit a namespaced attribute of the same name instead.
Late errors during calls to SubElement() (e.g. attribute related
ones) could leave a partially initialised element in the tree.
Modifying trees that contain parsed entity references could result in an infinite loop.
ObjectifiedElement.setattr created an empty-string child element when the attribute value was rejected as a non-unicode/non-ascii string
Syntax errors in lxml.cssselect could result in misleading error
messages.
Invalid syntax in CSS expressions could lead to an infinite loop in
the parser of lxml.cssselect.
CSS special character escapes were not properly handled in
lxml.cssselect.
CSS Unicode escapes were not properly decoded in lxml.cssselect.
Select options in HTML forms that had no explicit value
attribute were not handled correctly. The HTML standard dictates
that their value is defined by their text content. This is now
supported by lxml.html.
XPath raised a TypeError when finding CDATA sections. This is now fully supported.
Calling help(lxml.objectify) didn't work at the prompt.
The ElementMaker in lxml.objectify no longer defines the default
namespaces when annotation is disabled.
Feed parser failed to honour the 'recover' option on parse errors.
Diverting the error logging to Python's logging system was broken.
New helper functions strip_attributes(), strip_elements(), strip_tags() in lxml.etree to remove attributes/subtrees/tags from a subtree.
strip_attributes(), strip_elements(),
strip_tags() in lxml.etree to remove attributes/subtrees/tags
from a subtree.Namespace cleanup on subtree insertions could result in missing namespace declarations (and potentially crashes) if the element defining a namespace was deleted and the namespace was not used by the top element of the inserted subtree but only in deeper subtrees.
Raising an exception from a parser target callback didn't always terminate the parser.
Only {true, false, 1, 0} are accepted as the lexical representation for BoolElement ({True, False, T, F, t, f} not any more), restoring lxml <= 2.0 behaviour.
Injecting default attributes into a document during XML Schema validation (also at parse time).
Injecting default attributes into a document during XML Schema validation (also at parse time).
Pass huge_tree parser option to disable parser security
restrictions imposed by libxml2 2.7.
The script for statically building libxml2 and libxslt didn't work in Py3.
XMLSchema() also passes invalid schema documents on to libxml2
for parsing (which could lead to a crash before release 2.6.24).
Support for standalone flag in XML declaration through tree.docinfo.standalone and by passing standalone=True/False on serialisation.
standalone flag in XML declaration through
tree.docinfo.standalone and by passing standalone=True/False
on serialisation.Potential memory leak on exception handling. This was due to a problem in Cython, not lxml itself.
Potential memory leak on exception handling. This was due to a problem in Cython, not lxml itself.
Failing import on systems that have an io module.
Crash when using an XPath evaluator in multiple threads.
Ref-count leaks when lxml enters a try-except statement while an outside exception lives in sys.exc_*(). This was due to a problem in Cython, not lxml
Ref-count leaks when lxml enters a try-except statement while an outside exception lives in sys.exc_*(). This was due to a problem in Cython, not lxml itself.
Parser Unicode decoding errors could get swallowed by other exceptions.
Name/import errors in some Python modules.
Internal DTD subsets that did not specify a system or public ID were not serialised and did not appear in the docinfo property of ElementTrees.
Fix a pre-Py3k warning when parsing from a gzip file in Py2.6.
Test suite fixes for libxml2 2.7.
Resolver.resolve_string() did not work for non-ASCII byte strings.
Resolver.resolve_file() was broken.
Overriding the parser encoding didn't work for many encodings.
lxml.etree now tries to find the absolute path name of files when parsing from a file-like object. This helps custom resolvers when resolving relative
Memory problem when passing documents between threads.
Target parser did not honour the recover option and raised an
exception instead of calling .close() on the target.
Crash when parsing XSLT stylesheets in a thread and using them in another.
Crash when parsing XSLT stylesheets in a thread and using them in another.
Encoding problem when including text with ElementInclude under Python 3.
Smart strings can be switched off in XPath (smart_strings keyword option).
Smart strings can be switched off in XPath (smart_strings
keyword option).
lxml.html.rewrite_links() strips links to work around documents
with whitespace in URL attributes.
Custom resolvers were not used for XMLSchema includes/imports and XInclude processing.
CSS selector parser dropped remaining expression after a function with parameters.
objectify.enableRecursiveStr() was removed, use
objectify.enable_recursive_str() instead
Speed-up when running XSLTs on documents from other threads
Crash when using an XPath evaluator in multiple threads.
Ref-count leaks when lxml enters a try-except statement while an outside exception lives in sys.exc_*(). This was due to a problem in Cython, not lxml
Memory problem when passing documents between threads.
Memory problem when passing documents between threads.
Target parser did not honour the recover option and raised an
exception instead of calling .close() on the target.
lxml.html.rewrite_links() strips links to work around documents with whitespace in URL attributes.
lxml.html.rewrite_links() strips links to work around documents
with whitespace in URL attributes.Crash when parsing XSLT stylesheets in a thread and using them in another.
CSS selector parser dropped remaining expression after a function with parameters.
Your coding agent can read these notes before it upgrades. Set up the MCP server →