NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1742 most downloaded on PyPI
A fast HTML5 parser with CSS selectors, written in Cython, using the Lexbor engine.
Last release today
03 Oct 2026
Release timing varies
gaps range from 1 weeks to 3 months
Nearly every release is documented
notes for 50 of 52 stable releases
1 version withdrawn
withdrawn after publishing
8 years old
53 releases · first in 2018
This release contains breaking changes . It focuses on speed improvements and corner cases.
This release contains breaking changes. It focuses on speed improvements and corner cases.
The main change is that the Modest backend is no longer available.
It is outdated and unmaintained since 2021, contains bugs, and does not follow modern HTML5 standards.
The second is that parse_fragment() is gone. It re-implemented fragment parsing by guessing
whether the input was a document or a fragment; LexborHTMLParser(html, is_fragment=True) already
does this properly, so the guessing layer has been removed.
The rest of the changes fix corner cases where the bugs were happening rarely,
usually when heavily modifying the tree.
selectolax.parser is now a stub that raises ImportErrorfrom selectolax.lexbor import LexborHTMLParser) instead.parse_fragment().css_matches and any_css_matches missing matches outside the first top-level node of an HTMLcss.tags and strip_tags doing nothing on an HTML fragment. tags returned an empty list andstrip_tags removed nothing, because the fragment was not reachable from the document node.<html> wrapper of an HTML fragment being reported as a match.scripts_contain and script_srcs_contain missing matches outside the first top-levelcss.scripts_contain and script_srcs_contain answering from another scope's cachedselect() missing matches outside the first top-level node of an HTML fragment.css.traverse() covering only the first top-level node of an HTML fragment, andattribute_longer_than and any_attribute_longer_than returning inconsistent resultsattributes and attrs now report an attribute's qualified name instead of its local name.href and xlink:href used to collapse them into a single href key.iter() skipping the remaining children when a node is removed during iterationLexborHTMLParser('a<span>s</span>', is_fragment=True).text() returned 'a' instead of 'as'.text() silently returning truncated text instead of raising when a fragment fails to be collected.scripts_contain and script_srcs_contain sometimes returning wrong results due to HTML mutations.html, inner_html, html_pretty,attributes, attrs, id, tag, text_lexbor and text_content raise UnicodeDecodeError.text() already did.text_lexbor sometimes holding temporary memory longer than neededmerge_text_nodes(). It is now up to 10 times faster.__eq__attrs.items() and attrs.values().inner_html setter attaching element children to non-element nodeshead and body dangling after setting inner_html on the <html> elementinner_html setter freeing the replaced childrenroot going stale on an HTML fragment.unwrap().text method. Text fragments are now concatenated as raw bytes, up to 5x faster.skip_empty being ignored by text(deep=True)text() raising UnicodeDecodeError on undecodable bytes when deep=False. It now substitutesid method on text nodesattrs[key] = value raising AttributeError instead of TypeError when value is not a stringLexborNode or LexborAttributes directly.unwrap() corrupting the tree in some cases.head and body going stale once <head>/<body> is removed from the document.attrs reading freed memory when it outlives the node it was obtained from.attrs[key] = None; it leaked the value buffer header on every call.clone()encoding=True to LexborHTMLParser, which detects the encoding of bytes input andOne column per quarter.
Support lexbor-only builds with --disable-modest
--disable-modestselectolax no longer automatically imports selectolax.parser and selectolax.lexborAdd options parameter to LexborHTMLParser to pass Lexbor document parsing options (for example, LexborDocumentOptions.WO_EVENTS to disable mutation ev
Add NULL checks for CSS selectors module to prevent crashes
Do not destroy nodes when stripping tags
Add an ability to specify tags and namespace for fragmented parser
html5testtext(deep=True) on a text nodeAdd Add html_pretty , inner_html_pretty methods
html_pretty, inner_html_pretty methodsmerge_text_nodeslexborBreaking changes: Empty tags are now serialized to <div value=""> instead of <div value> ( Commit 4530fed ).
.text() and iter() for HTML fragments when there are multiple nodes at the root level. Resolves #209.<div value=""> instead of <div value>unwrap_tags and merge_text_nodes.Fix HTML parsing in fragment parser for LexborHTMLParser
LexborHTMLParserskip_empty parameter for text methods @pygarapcomment_content method @pygarapcreate_tag method to LexborHTMLParser.select()) when attributes are empty.Add is_fragment parameter to LexborHTMLParser @pygarap
Fix missing description on PyPi.
lexborBump version: 0.4.1 → 0.4.2
Bump version: 0.4.1 → 0.4.2
Fix parsing of CSS selectors that contain Unicode characters.
Fix incorrect default value in docstrings for strict argument
any_css_matchescss_first methodmerge_text_nodes for lexbor backend.inner_html property. Allows to get and set inner HTML of a node.css_first in lexbor backend.clone method to lexbor backend. Resolve #117..decompose. Resolves #179..unwrap. Resolves #169.Lexbor backend now supports :lexbor-contains("abc" i) CSS pseudo-class to match text nodes.
:lexbor-contains("abc" i) CSS pseudo-class to match text nodes.Add merge_text_nodes to lexbor backend. Fixes #170. @amirshukayev
merge_text_nodes to lexbor backend. Fixes #170. @amirshukayevUpdate lexbor. The new version of Lexbor fixes bugs related to CSS selectors. https://github.com/lexbor/lexbor/issues/279
Released
Improve type hints, add docstrings to type hints. Fixes #167
Released
Expose SelectolaxError exception in lexbor.pyi
SelectolaxError exception in lexbor.pyiReleased
SelectolaxError exception in lexbor.pyiFeat: Add unwrap empty tags functionality. Fixes #159.
Released
Fix: Update lexbor and improve HTML serialization speed. Fixes #153.
LexborHTMLParser.__init__. Fixes #144.- Fix: Header detected as head
Released - Improve type hints
Released
Feat: Add parse_fragment() and create_tag()
parse_fragment() and create_tag()Node.insert_child()Node.parser to access the HTMLParser to which the node belongsAdd Node.insert_child method to lexbor and modest backends. Thanks to @JuroOravec
Node.insert_child method to lexbor and modest backends. Thanks to @JuroOravecReleased
Node.insert_child method to lexbor and modest backends- Add Python 3.13 wheels - Update lexbor
*Breaking change*: lexbor backend now includes the root node when querying CSS selectors. Same as Modest backend.
lexbor backend now includes the root node when querying CSS selectors. Same as Modest backend.css_matches and any_css_matches methods for Modest backend on some compilersFixup for max HTML size increase (from 0.3.19 release, #110 )
lexbor backend (#104). Thanks to @lexborisovReleased
lexbor backendIncrease maximum HTML size to 2.4GB
Fix memory leak when using CSS selectors, lexbor backend
lexbor backendReleased
lexbor backend- Update lexbor - Add Python 3.12 wheels
lexbor- Make HTML nodes hashable - Pin Cython version
Improve typing. Thanks to @nesb1
Fix memory leak for lexbor backend
lexbor backendChanges: - Updated lexbor
Changes:
lexbor- Update lexbor - Add Python 3.11 wheels
lexborFix out-of-bounds bug for merge_text_nodes method.
merge_text_nodes method.Released
merge_text_nodes method.Due to a typo in the previous release version, we had to rerelease v0.3.9 as v0.3.10. If your package manager installed v3.8.9 -- please ignore it. Fo
Warning
Due to a typo in the previous release version, we had to rerelease v0.3.9 as v0.3.10.
If your package manager installed v3.8.9 -- please ignore it.
For more information, please see #70.
New changes
text(deep=True, separator='x').merge_text_nodes method for Modest backend.Released
This release does not contain any changes. Due to a typo in the version number (#70), we need to make a new release.
Fix incorrect text handling when using text(deep=True) on a text node.
text(deep=True) on a text node.Fix return type of HTMLParser.tags
Add binary builds for Python 3.10 and ARM builds for MacOS and Linux
Released
Added type annotations. Thanks to @stranger-danger-zamu
Released - Fix HTMLParser.html
Released
HTMLParser.htmlImprove text extraction for lexbor
select method for lexborReleased
selector method for lexbor- Fix setup.py for Windows
setup.py for WindowsAdded advanced Selector (the select method)
select method)strip_tagsclone method for the HtmlParser objectdetect_encoding, decode_errors, use_meta_tags, raw_html attributes for HtmlParsersget method to the attrs propertyhead propertyDon't throw exception when encoding text as UTF-8 bytes fails (#40).
- Build wheels for Apple Silicon
Fix strip argument is ignored for the root node (#35).
Fix root node property. The root property now points to the html tag.
root property now points to the html tag.Released
root property now points to the html tag.Add raw_value attribute for Node objects
raw_value attribute for Node objects (#22 )Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →