NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1029 most downloaded on PyPI
Fast and robust extraction of original and updated publication dates from URLs and web pages.
Last release 2 days ago
02 Oct 2026
Release timing varies
gaps range from 2 weeks to 10 months
Nearly every release is documented
notes for 57 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
9 years old
61 releases · first in 2017
performance: 2x faster extraction
One column per quarter.
maintenance: modernize code and packaging
maintenance: remove LXML version constraint and update setup with @zhemaituk ( #184 , #186 , #187 )
maintenance: explicitly support Python 3.13
maintenance: explicit re-export and code quality
fix: more robust copyright parsing
fix: more restrictive YYYYMM pattern to prevent ValueError with @b3n4kh
update docs with @EkaterineSheshelidze
fix meta property updated vs. original behavior
fix for MacOS: pin LXML dependency with @adamh-oai
focus on precision, stricter extraction patterns (#103, #105, #106, #112)
Full Changelog: https://github.com/adbar/htmldate/compare/v1.5.2...v1.6.0
fix for missing months keys in custom extractor
try_date_expr() (#101)fix regression for fast extraction introduced in e8b3538
slightly higher accuracy with revised heuristics
maintenance release: upgrade urllib3 dependency
urllib3 dependencysupport min_date/max_date as datetimes or datetime strings with @kernc
Full Changelog: https://github.com/adbar/htmldate/compare/v1.4.1...v1.4.2
better coverage of relevant HTML attributes
%Y-%m-%dFull Changelog: https://github.com/adbar/htmldate/compare/v1.4.0...v1.4.1
additional search of free text in whole document
technical release: explicit support for Python 3.11 and logo
fix for use of min_date & max_date
min_date & max_date (#62)Entirely type-checked code base
clear_caches() (#57)Full Changelog: https://github.com/adbar/htmldate/compare/v1.2.3...v1.3.0
- fix for memory leak (#56) - docs updated Full Changelog: https://github.com/adbar/htmldate/compare/v1.2.2...v1.2.3
Full Changelog: https://github.com/adbar/htmldate/compare/v1.2.2...v1.2.3
slightly higher accuracy & faster extensive extraction
Full Changelog: https://github.com/adbar/htmldate/compare/v1.2.1...v1.2.2
better extraction coverage, simpler code
Full Changelog: https://github.com/adbar/htmldate/compare/v1.2.0...v1.2.1
remove unnecessary ciso8601 dependency
Full Changelog: https://github.com/adbar/htmldate/compare/v1.1.1...v1.2.0
improved extraction coverage (#47) by @liulinlin90
Full Changelog: https://github.com/adbar/htmldate/compare/v1.1.0...v1.1.1
better handling of file encodings
Full Changelog: https://github.com/adbar/htmldate/compare/v1.0.1...v1.1.0
maintenance release, code base cleaned
--version addedFull Changelog: https://github.com/adbar/htmldate/compare/v1.0.0...v1.0.1
faster and more accurate encoding detection
improved generic date parsing (thanks @RadhiFadlillah)
- improved exhaustive search - simplified code - bug fixes - removed support for Python 3.4
- bugfixes
dateparser and regex modules fully integrated
dateparser and regex modules fully integrateddependencies updated and reduced: switch from requests to bare urllib3, make chardet standard and cchardet optional
requests to bare urllib3, make chardet standard and cchardet optionalOverflowError in extraction- compatibility with Python 3.9 - better speed and accuracy
technical release: package requirements and docs wording
code base and performance improved
- more efficient code - additional evaluation data
performance and documentation improved
htmldate finds original and updated publication dates of any web page. All the steps needed from web page download to HTML parsing, scraping and text
htmldate finds original and updated publication dates of any web page. All the steps needed from web page download to HTML parsing, scraping and text analysis are included.
In a nutshell, with Python:
from htmldate import find_date find_date('http://blog.python.org/2016/12/python-360-is-now-available.html') '2016-12-23' find_date('https://netzpolitik.org/2016/die-cider-connection-abmahnungen-gegen-nutzer-von-creative-commons-bildern/', original_date=True) '2016-06-23'
On the command-line:
$ htmldate -u http://blog.python.org/2016/12/python-360-is-now-available.html '2016-12-23'
Releases used in production and meant to be archived on Zenodo for reproducibility and citability.
For more information see htmldate.readthedocs.io
reduced number of packages dependencies
First release used in production and meant to be archived on Zenodo for reproducibility and citability.
First release used in production and meant to be archived on Zenodo for reproducibility and citability.
- tests on Linux & MacOS - bugs removed
- coverage extension
small bugs and coverage issues removed
bugs corrected and cleaner code
significant speed-up after code profiling
- fixed lxml dependency - reordered XPath-expressions
refined and combined XPath-expressions
Nothing published for this version
Nothing published for this version
Nothing published for this version
improved consistency and further tests
improvements in markup analysis along with more tests
tested for Python2 and 3 with tox and coverage stats
extensive search can be disabled
refined targeting of HTML structure
- better extraction - logging - further tests - settings
tests functions (tox and pytest)
Your coding agent can read these notes before it upgrades. Set up the MCP server →