NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
A fast, fully typed HTML toolkit for Python, powered by a C-accelerated core.
Last release today
10 Oct 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 26 of 26 stable releases
Nothing withdrawn
no release was ever pulled
4 months old
26 releases · first in 2026
One column per month.
feat(python): default to Python 3.15 by @gaborbernat in #1280
Full Changelog: 1.14.1...1.15.0
Add turbohtml.etree views for extraction code that uses ElementTree text, tail, mutation and XPath operations. The C adapter shares native DOM nodes and supports copying across lxml API boundaries. ( #1284 )
👷 ci(type): pin the type env to Python 3.14 by @gaborbernat in #1279
Full Changelog: 1.14.0...1.14.1
Retain Markdown fence languages across class ordering and whitespace around <code> ; recognize language-* classes on <pre> . ( #1281 )
Speed up XPath re:test() and matches() filters with many alternatives or repeated patterns. ( #1282 )
♻️ refactor: bring the security fixes to house style by @gaborbernat in #978
Note truncated.
The standalone IDNA fuzz target checks URL and Unicode normalization invariants. ( #1010 )
Add bounded HTML grammar generation and export materialized production sweeps for corpus seeding. ( #1018 )
Reduce saved HTML, CSS and JavaScript findings with language-aware deletion. ( #1019 )
minify_js() shortens empty static blocks, new X() , async ()=> and nested if / else . ( #1111 )
CSS minification shortens escaped names and preserves their token boundaries. ( #1112 )
minify_js() drops non-directive constant statements and prints an empty loop body as ; . ( #1113 )
minify_css() drops the space next to a string in an at-rule prelude. ( #1134 )
Detect URL output changes and rejection during invariant fuzzing. ( #1145 )
Add encoding stream and decoded-text invariants to the fuzz oracles. ( #1152 )
Add a bounded public URL-host oracle for Unicode mappings, NFC and canonical Punycode. ( #1155 )
Add fixed-target and repeated link-resolution checks to the fuzz oracles. ( #1159 )
Add fixed text and encoding checks for streaming byte parsing. ( #1160 )
Add bounded DOM sequence and observer callback invariant checks. ( #1162 )
Add bounded iterator edit-sequence invariant checks. ( #1164 )
Add bounded custom-property grammar generation with independent payload checks. ( #1167 )
Add bounded HTML grammar generation with literal sibling-tree checks. ( #1168 )
Add bounded XML generation with literal names and decoded whitespace expectations. ( #1170 )
Add bounded XML generation with literal comment and CDATA expectations. ( #1172 )
Fuzz canonical-equivalent IDNA hosts against independent normalized URL targets. ( #1176 )
Generate SVG integration trees with independent namespace and parent expectations. ( #1177 )
Share IDNA conversion between URL normalization and the standalone harness. ( #1179 )
Extend encoding fuzz fixtures with BOM and conflicting declarations. ( #1180 )
Extend encoding fuzz fixtures with both content-type declaration attribute orders. ( #1181 )
Add encoding fuzz fixtures for UTF-16 declarations that HTML maps to UTF-8. ( #1183 )
Generate finite nested lists with independent expectations for each item’s parent. ( #1184 )
Generate finite tables with independent row group and cell parent expectations. ( #1185 )
Add encoding fuzz fixtures for labels that decode as windows-1252. ( #1186 )
Expand HTML corpus generation with pinned Domato tag, attribute and value alternatives, document-only tag profiles, and formatting and SVG integration shapes. ( #1192 )
Check URL split/recomposition and reparse invariants through public normalization. ( #1193 )
Add bounded XML structure generation with materialized corpus bytes and separate node and expansion budgets. ( #1196 )
Compare inline XSLT and schema outcomes with libxml2 and libxslt, including repaired seed-tree mutations. ( #1198 )
Add bounded CSS structure generation from pinned Domato and Blink tables, with distinct variable/keyframe bindings and independent AST budget checks. ( #1201 )
Add bounded DOM byte programs with independent state and register-order checks. ( #1203 )
Add bounded Markdown source and HTML grammars with independent first-conversion checks and materialized corpus bytes. ( #1204 )
Add bounded encoded corpus generation with explicit codec expectations. ( #1205 )
Admit generated XML, CSS, Markdown and encoding bytes through public fuzz consumers, and seed coverage-guided CI from their generation contexts. ( #1207 )
Add a fuzz-only C build mode that poisons arena gaps and parked wrappers and checks its sanitizer at startup. ( #1217 )
RelaxNG reads include and externalRef files resolved against the new base_url argument, confined to include_root when set. ( #1218 )
Fuzz retained wrappers, nested callbacks, parser reentry, hostile argument conversions and garbage collection with byte programs that a native tree verifier checks after every step. ( #1223 )
Run the IDNA, phone, JavaScript and CSS standalone fuzz harnesses under MemorySanitizer, and the DOM lifecycle programs in parallel threads under ThreadSanitizer. ( #1244 )
Add an allocation-failure sweep that fails each PyMem allocation of the fuzz seeds in turn and expects MemoryError. ( #1277 )
📝 docs(changelog): number the XSD pattern fix by @gaborbernat in #948
Full Changelog: 1.13.0...1.13.1
Raise MemoryError for an EXSLT str:padding length above sys.maxsize // 4 characters instead of writing past a small buffer. ( #933 )
Split sanitized CSS declarations where a browser does, reading a backslash outside strings as an escape and an unquoted url(...) as one token, and drop a declaration whose string a newline ends (GHSA-9x48-74fj-x342). ( #934 )
Raise ValueError on a circular xsl:attribute-set and RecursionError past 400 levels of use-attribute-sets and template nesting instead of crashing, and apply a shared attribute set once per element. ( #935 )
Keep XSLT comment and processing-instruction text inside its node, and raise ValueError when the html method would write > inside a processing instruction. ( #936 )
Leave out an XSLT element, attribute or processing instruction whose computed name is invalid, keeping an element’s content, instead of writing the name as markup. ( #937 )
Raise ValueError in rewrite() when set_attribute gets a name, or a comment’s set_text gets text, that would inject markup. ( #938 )
Keep < and / apart when minifying CSS, and return a stylesheet unchanged when its minified form would spell a </style its source did not (GHSA-wh2c-9vrv-w597). ( #939 )
Cap the CSS minifier at 100 nesting levels, dropping a deeper block and copying a deeper function argument through unchanged, instead of overflowing the C stack. ( #940 )
Escape " as " in attribute values under Formatter.MINIMAL . ( #941 )
Escape < and > in attribute values under every formatter and in the start tags rewrite() writes for edited elements. ( #942 )
Stop minify_js() from creating a </script or <!-- the source did not spell, and emit an inline script verbatim when its minified form still holds either sequence. ( #943 )
Raise RuntimeError from a serialize_iter() stream after a tree edit such as extract or clear , instead of crashing or streaming a cut document. ( #944 )
Space apart the sequences that close DOM-built leaf data: a leading > or -> in a comment outside canonical XML, -- before > or !> in an HTML comment, a - before another - or at the end in an XML or canonical comment, and ?> in XML instruction data, and split XML CDATA at ]]> (GHSA-7vrw-9mvp-2x7j). ( #945 )
Match unsafe names and the per-name policy options in any ASCII case when sanitizing parse_xml() trees, and replace an attribute spelled in another case when writing one (GHSA-pjrh-5m66-3wwf). ( #946 )
Stop re:test , re:replace , matches and replace from hanging on patterns such as (a+)+$ ; oversized patterns and runaway back-references raise ValueError . ( #947 )
Raise ValueError for a schema pattern nested deeper than 250 groups, needing more than 2,000,000 automaton states, or holding a {n,m} quantifier with n above m , and match long alternations without exhausting the stack. ( #948 )
🐛 fix(sanitize): check meta refresh redirect URLs by @gaborbernat in #920
Full Changelog: 1.12.0...1.13.0
Add turbohtml.transform.strparam() to pass text as an XSLT parameter without hand-quoting it. ( #928 )
fix(markdown): preserve selected list markers by @gaborbernat in #914
Full Changelog: 1.11.0...1.12.0
Use iter_elements() for lazy, tag-filtered traversal of large trees. The iterator keeps its next match across edits and follows the tree that then contains it. ( #915 )
🐛 fix(markdown): keep non-item list content by @gaborbernat in #869
Full Changelog: 1.10.0...1.11.0
Add Markdown.Escaping(mode="none") , which leaves prose as written in to_markdown() apart from the asterisks and underscores choices and a | inside a table cell. ( #874 )
Assigning Element.tag renames the element in place, keeping its namespace, attributes, children and position. ( #883 )
Add turbohtml.migration.bleach.attribute_policy() for retaining bleach attribute rules with native CSS settings. Preserve callable precedence, source values before safety checks, and changes to rule mappings. Treat * as an attribute name. ( #895 )
Add Policy.attribute_predicate to inspect attributes before allowlist and safety checks. ( #901 )
👷 ci(codspeed): make the benchmark gate deterministic by @gaborbernat in #803
Full Changelog: 1.9.0...1.10.0
Add DocumentFragment : Range.extract_contents and Range.clone_contents return one, and every insertion method moves a fragment’s children into place and leaves it empty, as the DOM insert algorithm does. ( #857 )
✨ feat(clean): transform trees before serialization by @gaborbernat in #797
Full Changelog: 1.8.0...1.9.0
Add native DOM whitespace collapse, comment removal, and ordered transformation composition. Serialize subtree children with inner=True on serialize , encode , and serialize_iter . ( #791 )
Speed up JavaScript minification of long guard-return sequences. ( #795 )
Speed up JavaScript minification of declarations with single-use initializers. ( #795 )
Speed up descendant :has() selectors on nested elements. ( #795 )
Speed up JavaScript minification of libraries containing boolean literals. ( #795 )
Speed up canonical serialization with explicit options. ( #795 )
Speed up importing turbohtml from an installed wheel. ( #795 )
Speed up XPath translate() with large character maps. ( #795 )
Speed up sanitization of elements with many rejected attributes. ( #795 )
Speed up XSLT numbering with explicit element, wildcard or document-root patterns. ( #795 )
Speed up XML parsing of long text without character references. ( #795 )
Speed up HTML parsing with deeply nested formatting elements. ( #795 )
Speed up repeated computed-style lookups on an unchanged tree. ( #795 )
Speed up HTML parsing with repeated scope checks in nested structures. ( #795 )
Speed up XSD validation with inherited simple-type facets. ( #795 )
Speed up repeated named-slot assignment queries on wide shadow roots. ( #795 )
Speed up CSS minification with long conflicting declaration values. ( #795 )
Speed up parent queries over selections with shared parents. ( #795 )
Speed up HTML serialization with explicit options. ( #795 )
Speed up XML parsing of attributes with many namespace prefixes. ( #795 )
Speed up generation of the Unicode normalization tables. ( #795 )
Speed up form-data collection from deeply nested disabled fieldsets. ( #795 )
Speed up language detection on text with many distinct trigrams. ( #795 )
Speed up queries combining many roots supplied out of document order. ( #795 )
Speed up article extraction from deeply nested content. ( #795 )
Speed up XPath equality comparisons between node sets. ( #795 )
Speed up positional CSS queries over wide sibling lists. ( #795 )
Speed up repeated default XSLT level="any" numbering. ( #795 )
Speed up computed-style matching of class-qualified selectors. ( #795 )
Speed up XPath inequality comparisons between node sets. ( #795 )
Speed up XPath str:replace() with sparse matches of long strings. ( #795 )
Speed up HTML parsing of long text with carriage returns. ( #795 )
Speed up HTML parsing that merges text around tables. ( #795 )
Speed up Markdown wrapping of long lines. ( #795 )
Speed up JavaScript minification of mixed retained and removable declarations. ( #795 )
Speed up XSD validation of elements with many instance attributes. ( #795 )
Speed up RELAX NG validation of optional groups and interleaves. ( #795 )
Speed up microdata extraction when referenced properties are already ordered. ( #795 )
Speed up closest-ancestor queries over overlapping selections. ( #795 )
Speed up XPath set:intersection() on large node sets. ( #795 )
Speed up JavaScript minification that removes declarations from long statement lists. ( #795 )
Speed up XPath id() when its argument contains many short nodes. ( #795 )
Speed up streaming rewrites that add many attributes to a start tag. ( #795 )
Speed up XPath str:concat() over many short text nodes. ( #795 )
Speed up compilation of schema patterns with large character classes. ( #795 )
Speed up external-link extraction from URLs with many subdomains. ( #795 )
Speed up flattening nested shadow-slot assignments. ( #795 )
Speed up XSLT applications with many template rules. ( #795 )
Speed up canonical serialization of deeply nested trees with xlink attributes. ( #795 )
Speed up JavaScript literal propagation within large multi-binding declarations. ( #795 )
Speed up plain-text serialization with explicit options. ( #795 )
Speed up repeated default XSLT sibling numbering. ( #795 )
Speed up adding many attributes to an element. ( #795 )
Speed up CSS minification with many rules sharing declaration bodies. ( #795 )
Speed up mutations observed only for unrelated event kinds. ( #795 )
Speed up repeated element path generation on an unchanged tree. ( #795 )
Speed up table extraction with large row and column spans. ( #795 )
Speed up Unicode normalization of long combining-mark sequences. ( #795 )
Speed up cloning partially selected Range ancestors. ( #795 )
Speed up serialization of elements with large attribute sets. ( #795 )
Speed up publication-date extraction from text containing many distinct dates. ( #795 )
Speed up pruning selections that share matching ancestors. ( #795 )
Speed up shadow-slot assignment queries on hosts with many children. ( #795 )
Speed up DOM normalization of empty text nodes. ( #795 )
Speed up reading attributes from tokens with many attributes. ( #795 )
Speed up Range operations on fully contained sibling intervals. ( #795 )
Speed up JavaScript minification of long expression sequences. ( #795 )
Speed up XSLT numbering with static predicates and union patterns. ( #795 )
Speed up XSD validation of values with named-type facets. ( #795 )
Speed up XSLT stylesheets with repeated numbering patterns. ( #795 )
Speed up JavaScript minification that merges long declaration lists. ( #795 )
Speed up computed styles for complex selector alternatives. ( #795 )
Speed up repeated validation with schema patterns. ( #795 )
Speed up queries combining many detached roots. ( #795 )
Speed up Markdown conversion of long runs of punctuation. ( #795 )
Speed up Node.equals() for elements with many attributes. ( #795 )
Speed up article extraction from pages with many content candidates. ( #795 )
Speed up XML parsing of elements with many distinct attributes. ( #795 )
Speed up queries spanning many documents. ( #795 )
Speed up JavaScript initialization-order checks across many variable pairs. ( #795 )
Speed up explicit XSLT numbering across repeated source-node visits. ( #795 )
Speed up XPath set:difference() on large node sets. ( #795 )
Speed up JavaScript literal propagation across many interleaved declarations. ( #795 )
Speed up XPath ordered numeric comparisons between node sets. ( #795 )
Speed up JavaScript minification of integer arrays. ( #795 )
Speed up XSD validation of elements with many declared attributes. ( #795 )
Speed up Range operations whose boundary is near the start of a sibling list. ( #795 )
Speed up iteration over SAX element records . ( #795 )
Speed up encoding detection of short inputs. ( #795 )
Speed up repeated document-wide radio-group updates. ( #795 )
Speed up XML parsing of long attribute values without character references. ( #795 )
Speed up XSD validation of decimals without numeric bounds. ( #795 )
Speed up XPath set:distinct() on large node sets. ( #795 )
Speed up XPath equality comparisons of long strings. ( #795 )
Speed up DOM normalization of adjacent text nodes. ( #795 )
Speed up XPath set:has-same-node() on large node sets. ( #795 )
Speed up HTML parsing with many active formatting elements. ( #795 )
Speed up sibling queries over selections with shared parents. ( #795 )
Speed up microdata extraction with repeated item references. ( #795 )
Speed up microdata extraction when properties are local to their item. ( #795 )
Speed up HTML parsing of long text containing isolated NUL characters. ( #795 )
Speed up feed extraction from entries with extension fields. ( #795 )
Speed up XPath unions of results already in document order. ( #795 )
Speed up CSS minification with many mergeable media blocks. ( #795 )
Speed up JavaScript minification with interleaved unused bindings. ( #795 )
Speed up DOM linkification of long non-ASCII text. ( #795 )
Speed up XSLT applications with many named declarations. ( #795 )
Speed up XPath translate() on text with repeated characters. ( #795 )
Speed up streaming encoding detection of short inputs. ( #795 )
Speed up structured-data extraction from documents without metadata. ( #795 )
Skip repeated NUL scans in IncrementalParser . ( #799 )
Speed up parse_into() by reusing namespace strings within each parse. ( #799 )
Avoid repeated duplicate scans when Transform indexes XSLT keys. ( #799 )
Avoid shifting existing variable bindings when Transform enters and leaves scopes. ( #799 )
Speed up escaping disallowed tags with many attributes in Sanitizer . ( #799 )
Skip disqualified encoding candidates in EncodingDetector . ( #799 )
Speed up text by determining string width while collecting text. ( #799 )
Speed up text emission in Transform . ( #799 )
Reduce rule matching work in to_annotated_text() with many annotation rules. ( #799 )
Speed up check() on nested sections without headings. ( #799 )
Speed up exact text filters in find() and find_all() . ( #799 )
Speed up Element construction. ( #799 )
Speed up template-safe sanitize() . ( #799 )
Speed up Unicode range sorting in minify_css() . ( #799 )
Skip impossible composition lookups in normalize() . ( #799 )
Avoid redundant string copies and integer conversions in Transform numeric sorting. ( #799 )
Select only the highest-ranked trigrams when detect_language() scores text. ( #799 )
Avoid temporary text copies when xpath() concatenates node strings. ( #799 )
Speed up zero terms with large exponents in minify_css() . ( #799 )
Speed up boolean attribute filters in find() and find_all() . ( #799 )
Avoid temporary function-rendering buffers in minify_css() . ( #799 )
Avoid rescanning ASCII hosts in normalize_url() . ( #799 )
Speed up unescaped query keys in normalize_url() . ( #799 )
Speed up dot-segment handling in normalize_url() . ( #799 )
Speed up child and sibling :has() selectors in select() . ( #799 )
Speed up sanitize() when checking allowed attribute prefixes. ( #799 )
Use indexed country-code lookup when formatting and validating PhoneNumber . ( #799 )
Speed up phone table generation. ( #799 )
Speed up Unicode normalization table generation. ( #799 )
Skip unused attribute-name conversions in Sanitizer prefix checks. ( #799 )
Reuse namespace declarations for consecutive copies of a literal element in Transform . ( #799 )
Speed up literal conversion in css_to_xpath() . ( #799 )
Skip unchanged prefixes in normalize() . ( #799 )
Index Unicode digit ranges when parsing PhoneNumber . ( #799 )
Reuse each form’s first submit control when matching :default in select() . ( #799 )
Stop temporal text scanning after the first valid date in dates() . ( #799 )
Reuse adjacent sibling positions in select() nth selectors. ( #799 )
Use direct UTF decoders when parse() reads UTF-8 or UTF-16 bytes. ( #799 )
Speed up single-character renaming in minify_js() . ( #799 )
Speed up named-entity serialization with Html . ( #799 )
Avoid intermediate metadata copies in opengraph() . ( #799 )
Speed up parse_into() by avoiding temporary text copies. ( #799 )
Speed up attribute sorting with Html for elements with many attributes. ( #799 )
Speed up is_valid() and is_valid() for invalid documents. ( #799 )
Use direct combining-class lookup for common marks in normalize() . ( #799 )
🐛 fix(markdown): escape a table cell pipe where it is written by @gaborbernat in #765
Full Changelog: 1.7.0...1.8.0
PhoneNumbers(collapse_whitespace=True) reads a run of HTML whitespace between the parts of a number as the one space it renders as, so a number a source formatter or a template broke across a line still links. ( #758 )
Detect a URL whose authority is a single-label host or an IP literal when its scheme is written, so http://localhost:8000/path , http://intranet/ and http://[::1]:8080/ link. A bare domain still needs a dot and a known top-level domain to be told apart from an ordinary word.
Recognize an internationalized top-level domain written as its Unicode label, so президент.рф links the way президент.xn--p1ai already did. ( #768 )
Add unique=True to LinkDetector.find , which keeps the first span of each distinct URL. ( #770 )
Move the work behind the shipped link callbacks nofollow() and target_blank() into the C core. They keep their names, their signatures and their behavior, with one correction: the web-scheme test now matches the whole scheme, so ht: and httpx: no longer count as web links. ( #774 )
Move the Query facade’s set algebra into the C core: deduplicating by node identity, collecting siblings, the combined text , the joined attribute read, and the four class operations. Behavior is unchanged. ( #775 )
Move the encoding detector’s candidate ranking into the C core: deduplicating the scored candidates, normalizing each score to its share, promoting the detector’s own winner, and applying the Detection allowlist, exclusions, language preference and confidence floor. Results are unchanged. ( #776 )
Move the boilerplate classifier into the C core: segmenting a page into paragraph units, collapsing each unit’s text, and deciding which units are article content against the Extraction thresholds. Results are unchanged. ( #777 )
Move the structured-data shaping into the C core: which JSON-LD blocks carry data, and rendering a MicrodataItem as nested plain dicts. Decoding each block stays with the standard library’s JSON parser. Results are unchanged. ( #778 )
Move the publication-date stages into the C core: the canonical-URL, <meta> , JSON-LD, <time> and visible-text signals dates() reads now run in one walk each. Results are unchanged. ( #779 )
Move the URL cleaning pipeline into the C core: normalize_url() , clean_url() and extract_links() now run their normalization, gates and anchor walk in one C call each. Results are unchanged. ( #780 )
Move the conformance verdict and the severity views into the C core: check() reads valid from the walk, and errors / warnings / infos filter there. Results are unchanged. ( #781 )
Move the query facade’s traversal into the C core: parent() , children() and closest() walk there, and a Matcher applies its limit and filters a node’s children inside the walk. Results are unchanged. ( #782 )
Move StyleDeclaration ’s accessors into the C core: which declaration wins a repeated property and the text serialization are computed there. Results are unchanged. ( #783 )
Move the encoding and language detectors’ last decisions into the C core: which label a whatwg-* codec name resolves to, when a byte-order mark settles an EncodingDetector , the no-match answer, and the language confidence floor. Results are unchanged. ( #784 )
Move the sanitizer’s policy compilation into the C core: the rel value, the value allowlists, the style patterns and the transform rules a Sanitizer indexes are built there. Results are unchanged. ( #785 )
Move the bleach shim’s attributes translation into the C core: the flat, per-tag and callable shapes of turbohtml.migration.bleach.clean() compile there, and the per-tag predicates run through a C-bound filter. Results are unchanged. ( #786 )
Move the link and phone detectors’ configuration folding into the C core: tag, top-level-domain, scheme and word lists, region codes, the phone type mask and the E.164 check of PhoneNumber are computed there. Results are unchanged. ( #787 )
Move the whitespace fold of turbohtml.migration.markupsafe.Markup.striptags() into the C core. Results are unchanged. ( #789 )
Move turbohtml.build ’s argument sorting into the C core: a leading mapping becomes the attributes and a string becomes a text node there, for E and document() alike. Results are unchanged. ( #790 )
Add sanitize_node() and sanitize_report_node() , the tree-to-tree sanitizer for a pipeline that parses once and serializes once; sanitize() and sanitize_report() also accept a parsed node. ( #793 )
Add linkify_node() , which links URLs, email addresses and phone numbers in an already parsed tree in place; linkify() also accepts a parsed node. ( #794 )
PhoneNumbers(collapse_whitespace=True) reads a run of HTML whitespace between the parts of a number as the one space it renders as, so a number a source formatter or a template broke across a line still links. (758)
Detect a URL whose authority is a single-label host or an IP literal when its scheme is written, so http://localhost:8000/path, http://intranet/ and http://[::1]:8080/ link. A bare domain still needs a dot and a known top-level domain to be told apart from an ordinary word.
Recognize an internationalized top-level domain written as its Unicode label, so президент.рф links the way президент.xn--p1ai already did. (768)
Add unique=True to LinkDetector.find, which keeps the first span of each distinct URL. (770)
Move the work behind the shipped link callbacks ~turbohtml.clean.nofollow and ~turbohtml.clean.target_blank into the C core. They keep their names, their signatures and their behavior, with one correction: the web-scheme test now matches the whole scheme, so ht: and httpx: no longer count as web links. (774)
Move the ~turbohtml.query.Query facade's set algebra into the C core: deduplicating by node identity, collecting siblings, the combined text, the joined attribute read, and the four class operations. Behavior is unchanged. (775)
Move the encoding detector's candidate ranking into the C core: deduplicating the scored candidates, normalizing each score to its share, promoting the detector's own winner, and applying the ~turbohtml.detect.Detection allowlist, exclusions, language preference and confidence floor. Results are unchanged. (776)
Move the boilerplate classifier into the C core: segmenting a page into paragraph units, collapsing each unit's text, and deciding which units are article content against the ~turbohtml.extract.Extraction thresholds. Results are unchanged. (777)
Move the structured-data shaping into the C core: which JSON-LD blocks carry data, and rendering a ~turbohtml.MicrodataItem as nested plain dicts. Decoding each block stays with the standard library's JSON parser. Results are unchanged. (778)
Move the publication-date stages into the C core: the canonical-URL, <meta>, JSON-LD, <time> and visible-text signals ~turbohtml.extract.dates reads now run in one walk each. Results are unchanged. (779)
Move the URL cleaning pipeline into the C core: ~turbohtml.extract.normalize_url, ~turbohtml.extract.clean_url and ~turbohtml.extract.extract_links now run their normalization, gates and anchor walk in one C call each. Results are unchanged. (780)
Move the conformance verdict and the severity views into the C core: ~turbohtml.conformance.check reads valid from the walk, and errors/warnings/infos filter there. Results are unchanged. (781)
Move the query facade's traversal into the C core: ~turbohtml.query.Query.parent, ~turbohtml.query.Query.children and ~turbohtml.query.Query.closest walk there, and a ~turbohtml.query.Matcher applies its limit and filters a node's children inside the walk. Results are unchanged. (782)
Move ~turbohtml.cssom.StyleDeclaration's accessors into the C core: which declaration wins a repeated property and the text serialization are computed there. Results are unchanged. (783)
Move the encoding and language detectors' last decisions into the C core: which label a whatwg-* codec name resolves to, when a byte-order mark settles an ~turbohtml.detect.EncodingDetector, the no-match answer, and the language confidence floor. Results are unchanged. (784)
Move the sanitizer's policy compilation into the C core: the rel value, the value allowlists, the style patterns and the transform rules a ~turbohtml.clean.Sanitizer indexes are built there. Results are unchanged. (785)
Move the bleach shim's attributes translation into the C core: the flat, per-tag and callable shapes of turbohtml.migration.bleach.clean compile there, and the per-tag predicates run through a C-bound filter. Results are unchanged. (786)
Move the link and phone detectors' configuration folding into the C core: tag, top-level-domain, scheme and word lists, region codes, the phone type mask and the E.164 check of ~turbohtml.clean.PhoneNumber are computed there. Results are unchanged. (787)
Move the whitespace fold of turbohtml.migration.markupsafe.Markup.striptags into the C core. Results are unchanged. (789)
Move turbohtml.build's argument sorting into the C core: a leading mapping becomes the attributes and a string becomes a text node there, for ~turbohtml.build.E and ~turbohtml.build.document alike. Results are unchanged. (790)
Add ~turbohtml.clean.sanitize_node and ~turbohtml.clean.sanitize_report_node, the tree-to-tree sanitizer for a pipeline that parses once and serializes once; ~turbohtml.clean.sanitize and ~turbohtml.clean.sanitize_report also accept a parsed node. (793)
Add ~turbohtml.clean.linkify_node, which links URLs, email addresses and phone numbers in an already parsed tree in place; ~turbohtml.clean.linkify also accepts a parsed node. (794)
🔧 chore: batch dependency updates weekly on Tuesday by @gaborbernat in #756
Full Changelog: 1.6.1...1.7.0
Link phone numbers as tel: anchors with a PhoneNumbers setting; each match carries a PhoneNumber , the class that parses and formats numbers. ( #758 )
Link phone numbers as tel: anchors with a ~turbohtml.clean.PhoneNumbers setting; each match carries a ~turbohtml.clean.PhoneNumber, the class that parses and formats numbers. (758)
📄 docs: publish llms.txt from the docs build by @gaborbernat in #749
Full Changelog: 1.6.0...1.6.1
Parse a document, or feed a first stream chunk, that opens with a carriage return without touching the not yet allocated input buffer; UndefinedBehaviorSanitizer reported the empty append as a zero offset applied to a null pointer. ( #752 )
👷 ci(coverage): pin ctrace core so test contexts record by @gaborbernat in #703
Full Changelog: 1.5.1...1.6.0
Compile reusable XSLT stylesheet state during Transform construction. ( #713 )
Add allow_imports and import_root controls for XSLT import filesystem access. ( #714 )
Compile reusable XSLT stylesheet state during ~turbohtml.transform.Transform construction. (713)
Add allow_imports and import_root controls for XSLT import filesystem access. (714)
fix(tests): cast attr id for sorted to satisfy ty by @gaborbernat in #694
Full Changelog: 1.5.0...1.5.1
Match tag and attribute names case-sensitively on a turbohtml.parse_xml() document: an uppercase name like Y reads through attrs instead of raising KeyError , CSS selectors tell [Attr] from [attr] and Child from child , and pickle.loads() / copy.deepcopy() round-trip the tree without folding its names to the HTML parser’s lowercase. ( #695 )
Normalize line endings in a turbohtml.parse_xml() document per XML 1.0 §2.11 across attribute values, CDATA sections, comments, and processing instructions: a literal CRLF folds to a single LF (one space in an attribute value) rather than two, matching the text-content path and every other conformant XML parser. ( #696 )
⚡ perf(normalize): skip default table ranges by @gaborbernat in #649
Full Changelog: 1.4.0...1.5.0
Speed the C core across the read, query, validation, and extraction paths. Measured against the 1.4.0 release on a fixed workload, an id-selective XPath lookup runs about 260x faster, overlapping text search over 100x faster, Unicode normalization about 6x, microdata itemref resolution about 5x, Relax NG compilation about 4x, and CSS minification, conformance checking, attribute rewriting, date extraction, and computed-style resolution roughly 2x to 3x. Linear table scans give way to hashed or indexed lookups for tag atoms, XPath id tokens, microdata itemrefs, the Relax NG and XSD global tables, the encoding labels, and the entity names ( #659 , #672 , #671 , #657 , #656 , #661 , #660 , #651 , #654 ). The substring and literal-regex searches, the transform sort, and the CSS rule merge now bound their worst case instead of scanning to the end ( #665 , #664 , #666 , #652 , #662 ). The tokenizer name buffers, the table grid rows, and the DOM arena grow by doubling so a large document amortizes its allocation ( #669 , #668 , #667 ). The GB18030 decoder, the NFC normalizer, the attribute rewriter, and the IDNA mapper drop the branchy table walks that a common input never needed ( #674 , #673 , #649 , #670 , #650 ). Stylesheets cache their parse, list filtering batches, the language detector merges its trigram profiles, and date extraction trims its scan ( #653 , #663 , #648 , #658 ).
Serialize documents faster. The clean-run scan that hunts the next character needing an escape steps four UCS-4 code points per SIMD probe rather than two, and since the tree stores text as UCS-4 this scan runs on every serialization. The serialize benchmark improves by about 24%, and the gain carries into the strip and sanitize operations that serialize their result ( #676 ). IDNA host normalization rejects a non-composing mark pair with one comparison ahead of the composition-table search, saving the probes a run of unaccented text would otherwise spend ( #680 , #679 ). ( #676 )
Number a run of siblings in one pass. xsl:number computes a level’s number as the count of preceding siblings that match the count criteria, which rescanned the whole run for each sibling and cost time quadratic in the run length; number_counts was 10.13% of all instructions executed by the instruction-dense transform workload, the largest single entry ahead of every CPython symbol. The engine now carries the previous answer forward, since a node’s number is its previous sibling’s number plus whether that sibling counted, and holds the criteria alongside it so a run numbered under different criteria does not reuse it. Callgrind measures 5.9% fewer instructions over that workload ( #687 ). ( #687 )
Dispatch the html.parser adapter’s tokens from C. turbohtml.migration.stdlib.HTMLParser used to loop in Python over the token stream, building a token object and reading six of its fields for every token; the tokenizer now calls the handle_* methods itself, binding them once per feed and passing the strings straight through. Under Callgrind that removes 30.8% of the instructions the whole workload executes and 46.7% of its indirect branches, which is what the interpreter spends on dispatch. The adapter now leads lxml’s native target parser by 1.7 to 2.7 times where it previously trailed it, and html.parser by 7.5 to 11.2 times. ( #689 )
Move the read path’s per-item loops into C. turbohtml.saxparse.sax_parse() built a tuple, a typed record, and an isinstance chain for every event; the tokenizer now binds the six handler methods once and drives them off the tree walk, for 43.7% fewer instructions with instruction-cache misses down 71% on the WHATWG specification. URL cleaning ( turbohtml.extract.clean_url() , normalize_url() , extract_links() ), the dates() <meta> stage, turbohtml.query.escape_identifier() , and turbohtml.clean.linkify() ’s scheme classification each moved their loops to C as well, for 3.5% to 56.8% fewer instructions on those operations. Differential runs against the retired Python across 1,484,640, 809,481, 360,282, 302,762, 38,590, and 4,828 inputs found one divergence: the URL resolver matches %2e in either case as the WHATWG URL path state specifies, which no public call reaches because percent-encoding uppercases escape hex first. iter_events() keeps its typed records, since those are its public product.
Other hot paths keep their Python entry and shed work. XSD validation resolves each schema element’s qname once at compile time into a binary-searched table for 17.2% fewer instructions, built for XSD only after the eager cache slowed RELAX NG compilation by 8.15%. The link walk mints element wrappers on first use (14.8% fewer) and skips the URL join that already returns an absolute reference (67.8% fewer); the <meta> encoding prescan jumps inert bytes with memchr (8.67% fewer on ASCII); the gb18030 decoder maps astral pointers by arithmetic (13.4% fewer); and CSS property dispatch gates on the first byte (9.3% fewer on bootstrap.css ). The migration and benchmark tables also stopped over-claiming: a caveat marks only the rows of the operation it describes, a shared measurement shows in both columns, XPath labels render as inline code, and the CSS-stripper note gives the measured 1-3% size difference. ( #692 )
🐛 fix(dom, xslt): lock subtree copies by @gaborbernat in #635
Full Changelog: 1.3.1...1.4.0
Add capture_attributes=False for tag-and-text token streams, and reduce DOM allocation for large documents. ( #640 )
🔧 chore(meta): enrich PyPI links and issue templates by @gaborbernat in #633
Full Changelog: 1.3.0...1.3.1
✨ feat(pypy): run the C core on PyPy, publish 3.10 + 3.11 wheels by @gaborbernat in #623
Full Changelog: 1.2.0...1.3.0
Run the C core on PyPy 3.10 and 3.11 through cpyext , with the same conformance and byte-identical output as CPython; see Interpreters for what cpyext costs and for the three behaviors that do not carry over. ( #623 )
Decode the legacy encodings 1.2x to 1.5x faster by inlining each WHATWG decoder into its specialized loop, and stop reallocating turbohtml.IncrementalParser ’s decode scratch on every feed() . ( #624 )
🐛 fix(dom): lock threaded tree reads by @gaborbernat in #617
Full Changelog: 1.1.1...1.2.0
turbohtml.detect.detect() reports windows-1252 , not ascii , for pure-ASCII input; EncodingMatch gained a trailing codec field; and IncrementalParser takes WHATWG labels, so latin-1 raises where iso-8859-1 works. ( #622 )
turbohtml.detect.detect reports windows-1252, not ascii, for pure-ASCII input; ~turbohtml.detect.EncodingMatch gained a trailing codec field; and ~turbohtml.IncrementalParser takes WHATWG labels, so latin-1 raises where iso-8859-1 works. (622)
🔧 chore: correct package metadata and widen discoverability by @gaborbernat in #616
Full Changelog: 1.1.0...1.1.1
✨ feat(bench): name why each empty benchmark cell is blank by @gaborbernat in #526
<template shadowrootmode>) by @gaborbernat in #577Full Changelog: 1.0.0...1.1.0
Policy gained strip_template_markers : with it on, sanitizing collapses template-engine expressions ( {{ }} , ${ } , <% %> ) in kept text and attribute values to a single space, so the output cannot re-inject when a template engine renders it. This matches DOMPurify’s SAFE_FOR_TEMPLATES . ( #527 )
turbohtml.clean.sanitize_report() (and turbohtml.clean.Sanitizer.sanitize_report() ) sanitize a fragment and return what the policy dropped alongside the cleaned HTML: one turbohtml.clean.Removed record per removed element or stripped attribute, in walk order. This matches DOMPurify’s DOMPurify.removed . ( #528 )
turbohtml.convert.css_specificity() returns the (a, b, c) specificity of each selector in a comma-separated list, per CSS Selectors Level 4 §17, the value cssselect exposes as Selector.specificity() . It weighs the parsed selector in one C pass, with :is() / :not() / :has() taking their most specific argument and :where() contributing zero. ( #529 )
turbohtml.extract.feed() normalizes an RSS 2.0, Atom 1.0, or RDF/RSS-1.0 document into one frozen, typed Feed of Entry records, the feedparser.parse entry point over turbohtml.Document.feed() . It detects the format from the root element and maps each dialect’s spelling of a field – the entry title , link , id , updated / published , summary / content , and author – onto one shape in a single C walk of the parsed tree, over 12x faster than feedparser on a 30-item feed. ( #530 )
turbohtml.clean.Policy gains transform_tags : a map that renames elements while sanitizing, sanitize-html’s transformTags . Map a source tag to a string to rename it, or to a turbohtml.clean.Transform to rename it and add attributes (sanitize-html’s simpleTransform ). The rename runs before the allowlist in the same C walk, so the renamed element is re-checked from scratch – a transform decides an element’s name but never its safety: mapping a tag to script still drops it, and an added attribute is scrubbed like the element’s own. ( #531 )
A new turbohtml.saxparse module adds a DOM-less, event-driven parse. turbohtml.saxparse.sax_parse() drives a document through the WHATWG tree builder and fires a callback on a turbohtml.saxparse.SaxHandler subclass for each construct it builds – a start tag, an end tag, a run of text, a comment, the doctype, and a <?...> processing instruction – while turbohtml.saxparse.iter_events() yields the same stream as typed records. The events reflect the fully spec-correct tree (implied html / head / body , foster parenting, the adoption agency), so unlike html.parser.HTMLParser you see the tree the parser built; no per-node Python object is created and nothing is retained after the parse, so a one-pass extraction never builds a document-sized object graph. The tokenization, tree construction, and walk all run in C. ( #532 )
turbohtml.clean.Policy gains allowed_styles , a per-element, per-property value allowlist for the style attribute keyed {tag: {property: [pattern, ...]}} with "*" matching every tag. A declaration survives only when its value matches one of the property’s patterns, porting sanitize-html’s allowedStyles . It narrows css_properties by value and never weakens the baseline that drops expression() and disallowed-scheme url() . ( #533 )
Policy gained isolate_named_props : with it on, sanitizing prefixes every kept id and name value with user-content- , moving it out of the property namespace so it cannot shadow a built-in document or form property through named access (DOM clobbering, where <input name="attributes"> makes form.attributes resolve to the input and <img name="body"> hides document.body ). An already-prefixed value is left alone, so re-sanitizing is a fixpoint. This matches DOMPurify’s SANITIZE_NAMED_PROPS . ( #534 )
Html gained xml : with xml=True , serialize() , encode() , and serialize_iter() emit XML/XHTML instead of HTML – the equivalent of lxml’s tostring(method="xml") . Every empty element self-closes ( <br/> ), foreign SVG and MathML subtrees carry their namespace declarations, and text and attribute values follow the XML escaping rules, with no HTML void-element or raw-text special casing. It composes with sort_attributes and an Indent layout. ( #535 )
Policy gained a predicate-based custom-element allowance and split content profiles, porting DOMPurify’s CUSTOM_ELEMENT_HANDLING and USE_PROFILES . custom_element_check keeps an unlisted hyphenated custom element ( my-widget , x-card ) when a caller-supplied matcher admits its name, custom_attribute_check extends the same idea to that element’s attributes, and allow_customized_builtins keeps an is attribute whose value names a custom element. allow_html , allow_svg , and allow_mathml gate each namespace independently, so a policy can keep SVG but drop MathML, or the reverse. All of it runs in the one C sanitize walk, and the non-configurable safety baseline – on* handlers, javascript: URLs, unsafe tags – still applies to whatever a matcher keeps. ( #536 )
turbohtml.transform adds a full XSLT 1.0 processor, the job lxml ’s etree.XSLT does. turbohtml.transform.Transform compiles a stylesheet (parsed with turbohtml.parse_xml() ) and applies it to source documents, and turbohtml.transform.transform() does both in one call. The whole transform runs in the C extension, reusing turbohtml’s XPath 1.0 engine for every match pattern and select expression. It covers the entire XSLT 1.0 instruction set: templates with match / name / mode / priority , apply-templates with sort and with-param , call-template , for-each , if , choose , value-of , copy / copy-of , element / attribute / text , variable / param , multi-level number , key with the key() function, strip-space / preserve-space , attribute-set with use-attribute-sets , namespace-alias , fallback , simplified literal-result-element stylesheets, xsl:import with import precedence (resolved against a base_url ), cdata-section-elements , and the xml / html / text output methods (html auto-selected for a null-namespace html root, with <meta> injection). Validated against libxslt’s XSLT 1.0 Recommendation corpus at 76 of 79 cases byte-for-byte; the three remaining need a locale-collation, DTD, or XPath-namespace-axis layer turbohtml does not carry. ( #537 )
turbohtml.Node.canonicalize() serializes a subtree to Canonical XML (c14n), the byte-exact form an XML signature signs. A turbohtml.Canonical config selects the algorithm: Canonical XML 1.0 or 1.1, the exclusive variant that renders only the namespaces a subtree visibly uses, the with-comments variant, and an inclusive_ns_prefixes prefix list for exclusive mode. Attributes are reordered (namespace declarations first, then by namespace URI and local name), redundant namespace declarations are dropped, empty elements are written as start-end pairs, and character references are normalized, matching lxml ’s tostring(method="c14n") byte-for-byte over the same infoset. ( #538 )
turbohtml.validate.XMLSchema and turbohtml.validate.RelaxNG validate a document parsed with turbohtml.parse_xml() against an XSD 1.0 or RELAX NG schema, mirroring lxml’s etree.XMLSchema / etree.RelaxNG . A schema compiles once in the C core and each validate() returns a ValidationResult – a valid flag plus one ValidationError per violation, each with the document-order path that located it. XSD covers the element/attribute declarations, the sequence/choice/all content models with minOccurs / maxOccurs , references, complex/simple types with extension, the built-in datatypes, and the constraining facets; RELAX NG covers the full XML-syntax pattern set (including interleave ) through the derivative algorithm. ( #539 )
turbohtml.parse_xml() parses a document under XML 1.0 well-formedness instead of the WHATWG HTML tree builder, returning the same navigable Document . Names stay case-sensitive, <x/> self-closes any element, CDATA sections and processing instructions become CData and ProcessingInstruction nodes, only the five predefined entities and numeric references resolve, and a namespace prefix must be declared with xmlns . Names follow the exact XML 1.0 NameStartChar / NameChar productions, and the Namespaces in XML 1.0 well-formedness constraints hold in full: the reserved xml and xmlns prefixes and their namespace names cannot be rebound, a prefix declaration cannot be empty, a processing-instruction target carries no colon, and no two attributes share an expanded name. There is no HTML recovery: the first well-formedness violation – a mismatched or unclosed tag, an undeclared prefix, an undefined entity, a duplicate attribute – raises HTMLParseError . This is the equivalent of lxml.etree.fromstring / etree.XMLParser over turbohtml’s dependency-free, fully typed node API. ( #540 )
Added an HTML5 authoring-conformance checker with a severity model. turbohtml.conformance.check() walks a parsed document and returns a ConformanceReport – a valid verdict plus every ConformanceMessage , each carrying a stable code , a severity ( "error" , "warning" , or "info" ), a human-readable message, and a source line and column. It flags the document-conformance requirements the parser does not raise as a ParseError : a missing img alt, obsolete presentational elements and attributes, duplicate ids, invalid or redundant ARIA roles, empty headings, a section without a heading, and a document with no title or lang . The document is valid exactly when nothing is an error, so warnings and info notes never change the verdict. The whole walk runs in the C core against the WHATWG authoring rules and WAI-ARIA 1.2, the model the Nu Html Checker (validator.nu) uses; check_html() parses a markup string first. ( #541 )
xpath() gained the string subset of XPath 2.0: ends-with , string-join(seq, sep) , lower-case and upper-case (Unicode case mapping), and the regex matches(input, pattern[, flags]) and replace(input, pattern, repl[, flags]) spellings, where replace reads $1 -style group references and rewrites every match. They dispatch in the compiled-C engine alongside the XPath 1.0 core and the EXSLT namespaces, so an expression ported from elementpath , lxml , or htmlquery that leans on them runs without registration. ( #542 )
turbohtml.detect.normalize() returns text in a Unicode normalization form (UAX #15) – NFC , NFD , NFKC , or NFKD – the C successor to unicodedata.normalize() , and turbohtml.detect.is_normalized() tests membership. Both run over tables generated from the interpreter’s own unicodedata , so they agree with it exactly, and a quick check returns already-normalized text without allocating. ( #543 )
turbohtml.rewrite.rewrite() transforms HTML in a single streaming pass without building a tree, the model Cloudflare’s lol-html popularized. It runs the WHATWG tokenizer over the input while tracking only the open-element stack, hands each element a CSS selector matches – and, on request, each run of text, each comment, and the doctype – to a Python handler that edits it in place (set or remove an attribute, insert markup before, after, or around it, replace its inner content, unwrap it, or drop it), and emits the result incrementally. Working memory stays proportional to the open-element depth, not the document size, so a multi-megabyte page rewrites in a fixed footprint, and an untouched construct is reproduced verbatim. Because the pass never looks ahead, the matchable selector subset is the one decidable from an element and its ancestors – type, universal, id, class, and attribute selectors, the descendant and child combinators, :root , and :is() / :where() / :not() over that subset; a sibling combinator, a positional or structural pseudo-class, or :has() raises SelectorSyntaxError . ( #544 )
A new turbohtml.treebuild module retargets the parser at a tree of your own. turbohtml.treebuild.parse_into() runs the WHATWG tree builder and drives a builder object – a create_* method per node kind plus an append that links a child under its parent – to construct the tree directly, returning whatever the builder made its document root. No navigable turbohtml.Node is materialized and the tree is walked only once, so an index, a diff tree, or another library’s nodes is populated straight from the parse rather than by a second descent. Each element carries its namespace URI and its attributes as (name, value) pairs, a <template> ’s content is appended under the template handle, and a bogus <?...> construct reaches a distinct create_pi . This is Rust html5ever’s TreeSink and Node parse5’s TreeAdapter in turbohtml shape; the tree construction, the walk, and the string extraction all run in C. ( #545 )
turbohtml.cssom runs the CSS Object Model cascade: computed_style() resolves the getComputedStyle of an element by collecting every <style> sheet plus the inline style along its ancestor chain, matching the native selector engine, ordering the declarations by origin importance, the style attribute, specificity, and source order, then applying inheritance, shorthand expansion, and each property’s initial value – all in the C core under the per-tree critical section. Alongside it, StyleSheet , RuleList , StyleRule , and StyleDeclaration are the read-only, turbohtml-native spelling of the CSSOM CSSStyleSheet / CSSRuleList / CSSStyleRule / CSSStyleDeclaration interfaces. The returned value is the computed value, not the used value: turbohtml runs no layout, so lengths and percentages come back as written, the same boundary jsdom and cssstyle draw. Shorthand expansion covers the distributive families ( margin , padding , border-width / style / color , overflow ) and the <line-width> || <line-style> || <color> shorthands border , each border-<side> , and outline , whose components resolve in any order and reset every longhand they cover. ( #546 )
turbohtml.Node.to_source() losslessly serializes a tree back to HTML, re-emitting the verbatim source bytes of every element and text run a parse left untouched and reserializing only the parts a mutation changed. Parse with source_locations=True and an unedited round trip reproduces the input byte for byte – author quoting, tag-name case, character-reference spelling, and insignificant whitespace intact – for markup that parsed without implied elements or content reordering; after an edit only the changed node’s markup is rewritten while every untouched sibling and subtree keeps its original span. It is the tree-based counterpart to the streaming turbohtml.rewrite.rewrite() , the model Cloudflare’s lol-html popularized. ( #547 )
turbohtml.parse() , turbohtml.parse_fragment() , and IncrementalParser gained a source_locations flag (default False ) that records the granular source spans parse5 exposes as sourceCodeLocationInfo . With it on, each element’s source_location returns a SourceLocation giving the SourceSpan of its start tag, its end tag ( None when the source never closed it), and each attribute’s whole name="value" , every span carrying start/end line, column, and code-point offset so source[start_offset:end_offset] slices the construct out. The tokenizer stamps the spans in C as it runs and the tree builder hangs the record off each element, so the feature is zero-overhead when off; it implies positions when on, keeping source_line and position available beside the spans. ( #548 )
Added declarative Shadow DOM to the parser. When the tree builder meets a <template> carrying a shadowrootmode of open or closed on a valid shadow host, it attaches a shadow root to the template’s parent and parses the template’s content into it, reusing the Shadow DOM tree model – the template element never joins the light tree. shadowrootdelegatesfocus and shadowrootclonable set the matching flags, readable as the new delegates_focus and clonable properties. Following the WHATWG per-document flag, turbohtml.parse() honors declarative shadow roots by default (a browser navigation) while turbohtml.parse_fragment() does not (an innerHTML assignment); the new allow_declarative_shadow_roots argument flips either default, matching setHTMLUnsafe when turned on for a fragment. ( #549 )
Added the DOM Living Standard traversal objects turbohtml.TreeWalker and turbohtml.NodeIterator , with the turbohtml.NodeFilter constants for the what_to_show bitmask and the filter verdicts. A TreeWalker is a movable cursor over a subtree – parent_node , first_child , last_child , next_sibling , previous_sibling , next_node , previous_node – while a NodeIterator is the flat, filtered forward/backward view and iterates directly in a for loop. Both take a what_to_show node-type mask and an optional filter callback returning FILTER_ACCEPT , FILTER_REJECT , or FILTER_SKIP ; reject drops a node and its whole subtree while skip drops only the node, so a TreeWalker prunes where a NodeIterator (having no subtree) treats the two alike. The state machine and the what_to_show test run in the C core; the filter is the one callback into Python. This ports traversal code written against the browser DOM or jsdom. ( #550 )
turbohtml.parse() and turbohtml.parse_fragment() gained a scripting flag (default False ). With it on, turbohtml sets the WHATWG scripting flag: <noscript> becomes a raw-text element, so its content is one text run rather than parsed markup and serializes back unescaped, reproducing the tree a scripting browser builds. The flag is a property of the parsed tree, so the serializer and inner_html stay consistent with how it was parsed. parse5 and html5ever default this on for browser fidelity; turbohtml keeps it off so <noscript> fallback content stays navigable. ( #551 )
Added the DOM Living Standard Range and StaticRange types. A Range holds two boundary points – each a (container, offset) pair – and carries the full boundary API ( set_start() / set_end() and their _before / _after variants, select_node() , select_node_contents() , collapse() ), the derived collapsed and common_ancestor_container properties, the comparisons ( compare_boundary_points() , compare_point() , is_point_in_range() , intersects_node() ), and the content operations ( clone_contents() , extract_contents() , delete_contents() , insert_node() , surround_contents() , clone_range() ), each following the WHATWG boundary-point ordering and extract/clone/delete algorithms in C under the per-tree critical section. StaticRange is the immutable four-value snapshot. Offsets index code points in character data and children elsewhere, so a Python string’s own indexing lines up with a text-node offset. ( #552 )
Added the DOM Living Standard Shadow DOM tree model. attach_shadow() attaches an open or closed shadow tree and returns a ShadowRoot – a document-fragment-like root held off the light tree, so it never appears among the host’s children or in its serialization – reachable through shadow_root ( None for a closed root) and carrying mode , host , set_inner_html() , and append() . <slot> elements assign the host’s children by name (the unnamed default slot takes the rest): assigned_nodes() and assigned_elements() read what a slot received, with a flatten option that falls back to a slot’s own children and expands nested shadow slots, and assigned_slot gives the slot a child landed in. flattened_children returns the composed tree with every slot replaced by its assigned nodes. The assignment and flattening algorithms run in C under the per-tree critical section and are computed on demand, so they always reflect the current tree. ( #553 )
Added MutationObserver , a synchronous take on the DOM MutationObserver for recording tree edits. Register a node with observe() and the DOM options ( child_list , attributes , character_data , subtree , attribute_old_value , character_data_old_value , attribute_filter ); every change made through the mutation API queues a MutationRecord carrying the added and removed nodes, the surrounding siblings, and the attribute name and old value when asked, following the WHATWG “queue a mutation record” algorithm in C under the per-tree critical section. Because turbohtml has no event loop, delivery is synchronous rather than microtask-scheduled: take_records() returns and clears the queued batch, and deliver() drains it and calls the observer’s callback. disconnect() stops observing and discards pending records. ( #554 )
Policy gained xml : with it on, the sanitizer serializes the cleaned tree as well-formed XML/XHTML instead of HTML. Every kept empty element self-closes ( <br/> ), foreign SVG and MathML subtrees declare their namespace, text and attribute values follow the XML escaping rules, and a kept comment, a control character outside XML’s Char production, or an attribute name XML cannot hold is neutralized, so the output always reparses through turbohtml.parse_xml() . The walk and the safety baseline are unchanged, so an XML-mode policy is exactly as safe as its HTML-mode twin. This clones DOMPurify’s PARSER_MEDIA_TYPE: 'application/xhtml+xml' and replaces the brittle .replace("<br>", "<br/>") a bleach-based cleaner needs to feed a strict XHTML consumer such as Reportlab’s RML. turbohtml.Node.inner_xml exposes the same children-only XML serialization for any node. ( #565 )
The rewrite benchmark now runs against a fair in-process peer. lxml and BeautifulSoup do the same edits – rel=nofollow on every link, loading=lazy on every image, every comment dropped – through the parse, mutate, and serialize round trip that turbohtml.rewrite.rewrite() skips, and the table reports each party’s peak resident memory beside throughput, so the tree the streaming rewriter never builds shows up as memory it never holds. The lol-html migration guide carries the numbers. ( #612 )
The 1.0 release finishes the native-C port, settles one canonical public API, and closes the feature gap against the libraries turbohtml replaces. The notes below fold the whole 0.4.0 to 1.0.0 span into one overview; the anchor issues point at the epics behind each theme.
- Give the public surface one name per concept. CSS matching folds from turbohtml.match into turbohtml.query ; the sanitizer, linkifier, and every min
Give the public surface one name per concept. CSS matching folds from turbohtml.match into turbohtml.query ; the sanitizer, linkifier, and every minifier gather under turbohtml.clean ; a malformed selector raises one turbohtml.SelectorSyntaxError from every parse path; the two Detector classes split into turbohtml.detect.EncodingDetector and turbohtml.clean.LinkDetector ; and each surface with more than six arguments takes one frozen options config. ( #478 )
Select a serialization mode with a single layout argument in place of indent , so serialize(indent=2) becomes serialize(layout=Indent(2)) and Minify selects minified output. ( #171 )
Report a valueless attribute ( <x a> ) as the empty string rather than None in turbohtml.Element.attrs , matching the WHATWG tokenizer and the DOM. ( #87 )
- Build and edit the tree, not just read it: construct Element , Text , and Comment nodes and rearrange them with the full set of insert, wrap, extrac
Build and edit the tree, not just read it: construct Element , Text , and Comment nodes and rearrange them with the full set of insert, wrap, extract, and normalize methods, with attrs and .text / .data as live setters. copy , deepcopy , and pickle duplicate a subtree - by @gaborbernat . ( #19 )
Round out the node model: ProcessingInstruction and CData join the hierarchy, Doctype exposes its public_id and system_id , and every node type supports structural pattern matching - by @gaborbernat . ( #22 )
Build and edit the tree, not just read it: construct ~turbohtml.Element, ~turbohtml.Text, and ~turbohtml.Comment nodes and rearrange them with the full set of insert, wrap, extract, and normalize methods, with ~turbohtml.Element.attrs and .text/.data as live setters. copy, deepcopy, and pickle duplicate a subtree - by gaborbernat. (19)
Round out the node model: ~turbohtml.ProcessingInstruction and ~turbohtml.CData join the hierarchy, ~turbohtml.Doctype exposes its ~turbohtml.Doctype.public_id and ~turbohtml.Doctype.system_id, and every node type supports structural pattern matching - by gaborbernat. (22)
- Query any node with CSS through select() and select_one() , a native matcher covering type, universal, #id , .class , and attribute selectors (all o
Query any node with CSS through select() and select_one() , a native matcher covering type, universal, #id , .class , and attribute selectors (all operators plus the case-sensitivity flag) across the descendant, child, adjacent, and sibling combinators, returning comma groups in document order. An invalid selector raises ValueError - by @gaborbernat . ( #14 )
Search with a richer find() and find_all() filter grammar: match the tag and attributes by string, regex, bool, callable, or list (including class_ and the attrs mapping), and choose the search direction with the axis keyword. find_all takes a limit and returns a list - by @gaborbernat . ( #15 )
Test a node against a selector with matches() and closest() : matches() reports whether the node satisfies a CSS selector in context, and closest() returns the nearest matching ancestor (or the node itself), or None - by @gaborbernat . ( #16 )
Walk the tree by axis with new iterators: next_siblings , previous_siblings , document-order following and preceding , plus the strings and stripped_strings text iterators - by @gaborbernat . ( #17 )
Read HTML token-list attributes ( class , rel , headers , sizes , sandbox , and the rest) as a list[str] in turbohtml.Element.attrs , split on ASCII whitespace; other attributes stay strings and valueless ones stay None - by @gaborbernat . ( #18 )
Control serialization on any node: inner_html returns the children, while serialize() and encode() take a formatter (the Formatter enum picks the escape policy) and an indent for pretty output. The default stays WHATWG-conformant HTML - by @gaborbernat . ( #20 )
Parse bytes directly: turbohtml.parse() sniffs the encoding with the WHATWG algorithm (BOM, encoding argument, <meta> charset, then windows-1252), decodes with U+FFFD replacement, and reports the result in encoding - by @gaborbernat . ( #21 )
Query any node with CSS through ~turbohtml.Node.select and ~turbohtml.Node.select_one, a native matcher covering type, universal, #id, .class, and attribute selectors (all operators plus the case-sensitivity flag) across the descendant, child, adjacent, and sibling combinators, returning comma groups in document order. An invalid selector raises ValueError - by gaborbernat. (14)
Search with a richer ~turbohtml.Node.find and ~turbohtml.Node.find_all filter grammar: match the tag and attributes by string, regex, bool, callable, or list (including class_ and the attrs mapping), and choose the search direction with the axis keyword. find_all takes a limit and returns a list - by gaborbernat. (15)
Test a node against a selector with ~turbohtml.Node.matches and ~turbohtml.Node.closest: matches() reports whether the node satisfies a CSS selector in context, and closest() returns the nearest matching ancestor (or the node itself), or None - by gaborbernat. (16)
Walk the tree by axis with new iterators: ~turbohtml.Node.next_siblings, ~turbohtml.Node.previous_siblings, document-order ~turbohtml.Node.following and ~turbohtml.Node.preceding, plus the ~turbohtml.Node.strings and ~turbohtml.Node.stripped_strings text iterators - by gaborbernat. (17)
Read HTML token-list attributes (class, rel, headers, sizes, sandbox, and the rest) as a list[str] in turbohtml.Element.attrs, split on ASCII whitespace; other attributes stay strings and valueless ones stay None - by gaborbernat. (18)
Control serialization on any node: ~turbohtml.Node.inner_html returns the children, while ~turbohtml.Node.serialize and ~turbohtml.Node.encode take a formatter (the ~turbohtml.Formatter enum picks the escape policy) and an indent for pretty output. The default stays WHATWG-conformant HTML - by gaborbernat. (20)
Parse bytes directly: turbohtml.parse sniffs the encoding with the WHATWG algorithm (BOM, encoding argument, <meta> charset, then windows-1252), decodes with U+FFFD replacement, and reports the result in ~turbohtml.Document.encoding - by gaborbernat. (21)
- Tokenize HTML directly with a WHATWG-conformant tokenizer: turbohtml.tokenize() for whole strings, the streaming turbohtml.Tokenizer , and the turbo
Tokenize HTML directly with a WHATWG-conformant tokenizer: turbohtml.tokenize() for whole strings, the streaming turbohtml.Tokenizer , and the turbohtml.Token / turbohtml.TokenType types, validated against the html5lib-tests tokenizer conformance suite. ( #6 )
Run turbohtml.escape() and turbohtml.unescape() faster: vectorized scanning and bulk copying speed up both calls, with unescaping of real escaped HTML about three times faster than the general lookup path. The benchmark now uses pyperf over multi-MiB real documents - by @gaborbernat . ( #7 )
- Install reliably from PyPI again: publishing each wheel in its own job keeps PEP 740 attestations within the Sigstore identity’s lifetime, fixing th
Install reliably from PyPI again: publishing each wheel in its own job keeps PEP 740 attestations within the Sigstore identity’s lifetime, fixing the sigstore.oidc.ExpiredIdentity failure that blocked the first upload - by @gaborbernat . ( #4 )
Your coding agent can read these notes before it upgrades. Set up the MCP server →