NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #858 most downloaded on PyPI
HTML cleaner from lxml project
Last release 4 months ago
20 May 2026
Ships fairly regularly
a new release about every 3 months
Nearly every release is documented
notes for 13 of 13 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
13 releases · first in 2024
Fixed a security vulnerability where javascript: URLs in xlink:href attributes were not sanitized whensafe_attrs_only=False, allowing cross-site scrip…
Fixed a security vulnerability where javascript: URLs in xlink:href attributes were not sanitized when``safe_attrs_only=False``, allowing cross-site scripting (XSS) attacks. The fix requires lxml>=6.1.1, which adds xlink:href to the set of link attributes iterated by rewrite_links(). Reported by Guillem Lefait (@glefait).
Add more tests for different combinations of backslashes and unicode
Add more tests for different combinations of backslashes and unicode
One column per month.
Fixed a bug where Unicode escapes in CSS were not properly decoded before security checks. This prevents attackers from bypassing filters using escape sequences. (CVE-2026-28348)
Fixed a security issue where <base> tags could be used for URL hijacking attacks. The <base> tag is now automatically removed whenever the <head> tag is removed (via page_structure=True or manual configuration), as <base> must be inside <head> according to HTML specifications. (CVE-2026-28350)
Tests updated to work correctly with new lxml and libxml2 releases.
Tests updated to work correctly with new lxml and libxml2 releases.
Python 3.6 and 3.7 are no longer tested.
Improved documentation about CSS removal behavior.
lxml_html_clean now correctly handles HTML input as bytes as it did before the 0.2.0 release.
lxml_html_clean now correctly handles HTML input as bytes as it did before the 0.2.0 release.
Removed superfluous debug prints.
Removed superfluous debug prints.
Remove only the CSS comment if a suspicious content is detected
Remove only the CSS comment if a suspicious content is detected
The Cleaner() now scans for hidden JavaScript code embedded within CSS comments. In certain contexts, such as within <svg> or <math> tags, <style> tags may lose their intended function, allowing comments like /* foo */ to potentially be executed by the browser. If a suspicious content is detected, only the comment is removed. (CVE-2024-52595)
Do not parse URL addresses when it is not necessary.
Do not parse URL addresses when it is not necessary.
Parsing of URL addresses has been enhanced and Cleaner removes ambiguous URLs.
Parsing of URL addresses has been enhanced and Cleaner removes ambiguous URLs.
Fixes: #15
Fixes: #15
sdist now includes all test files and changelog.
Memory efficiency is now much better for HTML pages where cleaner removes a lot of elements.
Memory efficiency is now much better for HTML pages where cleaner removes a lot of elements. (#14)
ASCII control characters (except HT, VT, CR and LF) are now removed from string inputs before they're parsed by lxml/libxml2.
ASCII control characters (except HT, VT, CR and LF) are now removed from string inputs before they're parsed by lxml/libxml2.
Regular expresion for image data URLs now supports multiple data URLs on a single line.
Regular expresion for image data URLs now supports multiple data URLs on a single line.
First official release of the split project.
First official release of the split project.
Your coding agent can read these notes before it upgrades. Set up the MCP server →