NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3886 most downloaded on PyPI
Extract text from HTML
Last release 12 months ago
06 Oct 2025
Release timing varies
gaps range from 3 weeks to 3.7 years
Nearly every release is documented
notes for 15 of 15 stable releases
Nothing withdrawn
no release was ever pulled
10 years old
15 releases · first in 2016
Bump version: 0.7.0 → 0.7.1
Bump version: 0.7.0 → 0.7.1
Added support for Python 3.14.
Explicitly re-export public names.
Migrated the build system to hatchling.
CI improvements.
Bump version: 0.6.2 → 0.7.0
Bump version: 0.6.2 → 0.7.0
Removed support for Python 3.8.
Added support for Python 3.13.
Added type hints and py.typed.
CI improvements.
One column per quarter.
Bump version: 0.6.1 → 0.6.2
Bump version: 0.6.1 → 0.6.2
Bump version: 0.6.0 → 0.6.1
Bump version: 0.6.0 → 0.6.1
Fixed HTML comment and processing instruction handling.
Use lxml-html-clean instead of lxml[html_clean] in setup.py, to avoid https://github.com/jazzband/pip-tools/issues/2004
Bump version: 0.5.2 → 0.6.0
Bump version: 0.5.2 → 0.6.0
Moved the Git repository to https://github.com/zytedata/html-text.
Added official support for Python 3.9-3.12.
Removed support for Python 2.7 and 3.5-3.7.
Switched the lxml dependency to lxml[html_clean] to support lxml >= 5.2.0.
Switch from Travis CI to GitHub Actions.
CI improvements.
Bump version: 0.5.1 → 0.5.2
Bump version: 0.5.1 → 0.5.2
Handle lxml Cleaner exceptions (a workaround for https://bugs.launchpad.net/lxml/+bug/1838497 );
Python 3.8 support;
testing improvements.
Bump version: 0.5.0 → 0.5.1
Bump version: 0.5.0 → 0.5.1
Fixed whitespace handling when guess_punct_space is False: html-text was producing unnecessary spaces after newlines.
Bump version: 0.4.1 → 0.5.0
Bump version: 0.4.1 → 0.5.0
Parsel dependency is removed in this release, though parsel is still supported.
parsel package is no longer required to install and use html-text;
html_text.etree_to_text function allows to extract text from lxml Elements;
html_text.cleaner is an lxml.html.clean.Cleaner instance with options tuned for text extraction speed and quality;
test and documentation improvements;
Python 3.7 support.
Fixed a regression in 0.4.0 release: text was empty when html_text.extract_text is called with a node with text, but without children.
Fixed a regression in 0.4.0 release: text was empty when html_text.extract_text is called with a node with text, but without children.
Bump version: 0.3.0 → 0.4.0
Bump version: 0.3.0 → 0.4.0
This is a backwards-incompatible release: by default html_text functions now add newlines after elements, if appropriate, to make the extracted text to look more like how it is rendered in a browser.
To turn it off, pass guess_layout=False option to html_text functions.
guess_layout option to to make extracted text look more like how it is rendered in browser.
Add tests of layout extraction for real webpages.
Expose functions that operate on selectors, use .//text() to extract text from selector.
Expose functions that operate on selectors, use .//text() to extract text from selector.
Packaging fix (include CHANGES.rst)
Packaging fix (include CHANGES.rst)
Fix unwanted joins of words with inline tags: spaces are added for inline tags too, but a heuristic is used to preserve punctuation without extra spac
Fix unwanted joins of words with inline tags: spaces are added for inline tags too, but a heuristic is used to preserve punctuation without extra spaces.
Accept parsed html trees.
Travis-CI and codecov.io integrations added
Travis-CI and codecov.io integrations added
* First release on PyPI.
First release on PyPI.
Your coding agent can read these notes before it upgrades. Set up the MCP server →