NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1039 most downloaded on PyPI
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML.
Last release 2 days ago
02 Oct 2026
Ships unpredictably
gaps range from 2 weeks to 1.5 years
Nearly every release is documented
notes for 53 of 53 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
53 releases · first in 2019
Downloads and decompression overhaul by @adbar
Major changes:
Fixes:
Maintenance:
One column per quarter.
Revamped text recovery and extraction sequence: better recall overall and better extraction on forums by @adbar
Major changes:
Fixes:
Maintenance:
Full Changelog: https://github.com/adbar/trafilatura/compare/v2.1.0...v2.2.0
Dependencies updated, lxml in particular (with minimal changes in the code)
Major changes:
Fixes:
Nones in code blocks by @crackcomm in (#797)Maintenance:
as_dict deprecation warning → use .as_dict() method on return value
Breaking changes:
bare_extraction():
Document class by defaultas_dict deprecation warning → use .as_dict() method on return value (#730)bare_extraction() and extract(): no_fallback deprecation warning → use fast instead (#730)decode argument in fetch_url() → use fetch_response instead (#724)max_tree_size parameter to settings.cfg (#742)Fixes:
options.source before raising error on empty doc tree by @dmoklaf (#707)options.source (#717)Metadata:
Command-line interface:
--list with @gremid (#744)Maintenance:
pyproject.toml file (#715)main_extractor (#714)__all__ (#740)Documentation:
docs/index.html by @nzw0301 (#711)downloads: add support for SOCKS proxies with @gremid
prune_xpath parameter added by @felipehertzer (#684)max_sitemaps parameter added by @felipehertzer (#690)crawler: add parameter class and types, breaking change for undocumented functions
Navigation:
Bugfixes:
AttributeError in element deletion (#668)MemoryError in table header columns (#665)Docs:
enforce fixed list of output formats, deprecate -out on the CLI
Breaking change:
-out on the CLI (#647)Faster, more accurate extraction:
Bugfixes and maintenance:
include_formatting (#649)MemoryError & ValueError during conversion to text (#658)Documentation:
crawls.rst: known is an unexpected argument, by @tommytyc in #638metadata now skipped by default (#613), to trigger inclusion in all output formats:
Breaking change:
with_metadata=True (Python)--with-metadata (CLI)Extraction:
Evaluation:
Maintenance:
raise errors on deprecated CLI and function arguments
Breaking changes:
trafilatura.hashing → trafilatura.deduplicationExtraction:
Downloads:
Maintenance:
Docs:
add markdown as explicit output
Extraction:
Metadata:
Maintenance:
process_record() (#549)Pin LXML to prevent broken dependency
Better precision by @felipehertzer (#509, #520)
Extraction:
Downloads and Navigation:
is_live_page() (#501)Maintenance:
Response class: convenience functions added (#497)lxml.html.Cleaner removed (#491)add advanced fetch_response() function → pending deprecation for fetch_url(decode=False)
Extraction:
html2txt() function (#483)Downloads:
fetch_response() function
→ pending deprecation for fetch_url(decode=False)Maintenance:
MacOS: fix setup, update htmldate and add tests
Maintenance:
Navigation:
MAX_REDIRECTS config setting and fix urllib3 redirect handling by @vbarbaresi in #461Documentation:
preserve space in certain elements with @idoshamun
Extraction:
Metadata:
htmldate extensive search parameter in config (#434)Navigation:
Documentation:
improved code block support with @idoshamun (#372, #401)
Extraction:
Metadata:
Command-line interface:
--probe option to CLI to check for extractable content (#378, #392)Maintenance:
htmldate and courlan)minor fixes: tables in figures (#301), headings (#354) and lists
Extraction:
Metadata:
additionalName by @awwitecki in #363Navigation:
Full Changelog: https://github.com/adbar/trafilatura/compare/v1.6.0...v1.6.1
fix deprecation warning with @sdondley in #321
Extraction:
Command-line interface:
Navigation
is_live test() using HTTP HEAD request (#327)Maintenance
urllib3 version 2.0+Full Changelog: https://github.com/adbar/trafilatura/compare/v1.5.0...v1.6.0
fixes for metadata extraction with @felipehertzer (#295, #296), @andremacola (#282, #310), and @edkrueger
Extraction:
Navigation:
Maintenance:
Full Changelog: https://github.com/adbar/trafilatura/compare/v1.4.1...v1.5.0
review argument consistency and add deprecation warnings
Extraction:
Metadata:
Command-line interface:
Setup:
Full Changelog: https://github.com/adbar/trafilatura/compare/v1.4.0...v1.4.1
Extraction:
Metadata:
Command-line interface:
Setup:
Impact on extraction and output format:
Impact on extraction and output format:
Smaller changes in convenience functions:
Updates:
Full Changelog: https://github.com/adbar/trafilatura/compare/v1.3.0...v1.4.0
prepared deprecation of old process_record() function
html2txt() function added (#221)process_record() functionFull Changelog: https://github.com/adbar/trafilatura/compare/v1.2.2...v1.3.0
more efficient rules for extraction
Full Changelog: https://github.com/adbar/trafilatura/compare/v1.2.1...v1.2.2
--precision and --recall arguments added to the CLI
--precision and --recall arguments added to the CLIFull Changelog: https://github.com/adbar/trafilatura/compare/v1.2.0...v1.2.1
efficiency: replaced module readability-lxml by trimmed fork
Full Changelog: https://github.com/adbar/trafilatura/compare/v1.1.0...v1.2.0
encodings: better detection, output NFC-normalized Unicode
Full Changelog: https://github.com/adbar/trafilatura/compare/v1.0.0...v1.1.0
compress HTML backup files & seamlessly open .gz files
pycurl, language identification with py3langidFull Changelog: https://github.com/adbar/trafilatura/compare/v0.9.3...v1.0.0
better, faster encoding detection: replaced chardet with charset_normalizer
Full Changelog: https://github.com/adbar/trafilatura/compare/v0.9.2...v0.9.3
first precision- and recall-oriented presets defined
CLI: option names normalized (heed deprecation warnings), new option explore
explorefocused crawling functions including politeness rules
better handling of formatting, links and images, title type as attribute in XML formats
extraction trade-off: slightly better recall
breaking change: the extract function now reads target format from output_format argument only
extract function now reads target format from output_format argument onlycustomizable configuration file to parametrize extraction and downloads
requests replaced with bare urllib3 and custom decodingadded bare_extraction function returning Python variables
bare_extraction function returning Python variables- link discovery in sitemaps - compatibility with Python 3.9 - extraction coverage improved - deduplication now optional - bug fixes
optional language detector changed: langid → pycld3
langid → pycld3bare_extraction()courlan), more complete metadataextended and more convenient command-line options
faster and more robust text and metadata extraction
better metadata extraction and integration (XML & XML-TEI)
improved "fast" mode (accuracy and speed)
support for Python 3.4 reactivated
code base re-structured for clarity and readability
added metadata to the XML output
better handling of nested elements, quotes and tables
- handling of line breaks - element trimming simplified
First release used in production and meant to be archived on Zenodo for reproducibility and citability.
First release used in production and meant to be archived on Zenodo for reproducibility and citability.
- optional dependencies - bugs in parsing removed
- code profiling and speed-up
better handling of non-p elements
improvements in extraction recall
first release, minimum viable package
Your coding agent can read these notes before it upgrades. Set up the MCP server →