NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1038 most downloaded on PyPI
Clean, filter and sample URLs to optimize data collection – includes spam, content type and language filters.
Last release 4 months ago
01 Jun 2026
Ships unpredictably
gaps range from 8 days to 1.6 years
Nearly every release is documented
notes for 32 of 32 stable releases
Nothing withdrawn
no release was ever pulled
6 years old
32 releases · first in 2020
One column per quarter.
More robust FIND_LINKS_REGEX expression by @danishashko
FIND_LINKS_REGEX expression by @danishashko (#130)UrlStore() parameter trailing renamed to trailing_slashfilter_links() and load_store() are now exported at the top levelextract_links() : deprecate base_url parameter
extract_links() : review and document, add deprecation warning for base_url argument
maintenance: deprecate Python 3.6 & 3.7, add pyproject.toml setup file ( #59 , #105 )
pyproject.toml setup file (#59, #105)UrlStore maintenance: deprecate timelimit argument
timelimit argument (#101)replace langcodes by babel and use its information on locales (#89, #92)
langcodes by babel and use its information on locales (#89, #92)timelimit parameter by time_limit (#91)license change from GPLv3+ to Apache 2.0
write() method and load_store() function added (#83)trailing_slash to keep of discard slashes at the end of URLs (#52)clean_url() (#77), simplify code (#79)IRI to URI normalization: encode path, query and fragments (#58, #60)
is_valid_url() (#63)Full Changelog: https://github.com/adbar/courlan/compare/v0.9.4...v0.9.5
is_valid_url() (#63)Full Changelog: v0.9.4...v0.9.5
new UrlStore functions: add_from_html() ( #42 ), discard() ( #44 ), get_unvisited_domains
add_from_html() (#42), discard() (#44), get_unvisited_domains--samplesize, use --sample with an integer instead (#54)Full Changelog: v0.9.3...v0.9.4
refined link extraction and link filters ( #30 , #36 )
get_unvisited_domains() method to UrlStore (#40)Full Changelog: v0.9.2...v0.9.3
add blogspot archives to type filter
network tests: larger throughput
reset() (#22) and get_all_counts() methodssignal in #18, total_url_numberFull Changelog: https://github.com/adbar/courlan/compare/v0.9.0...v0.9.1
hardening of filters and URL parses
UrlStore: get_crawl_delay(), print_unvisited_urls()UrlStore now triggers exit code 1 when interruptedextract_links(): no_filterFull Changelog: https://github.com/adbar/courlan/compare/v0.8.3...v0.9.0
fixed bug in domain name extraction
Full Changelog: https://github.com/adbar/courlan/compare/v0.8.2...v0.8.3
- full type hinting - maintenance: code linted Full Changelog: https://github.com/adbar/courlan/compare/v0.8.1...v0.8.2
Full Changelog: https://github.com/adbar/courlan/compare/v0.8.1...v0.8.2
add type annotations and check with mypy
mypyurl_filter() function moved from Trafilaturablackfast track for domain extraction (extract_domain(url, fast=True)), now taking subdomains into account
extract_domain(url, fast=True)), now taking subdomains into accountFull Changelog: https://github.com/adbar/courlan/compare/v0.7.2...v0.8.0
UrlStore: threading lock and convenience functions added
UrlStore: threading lock and convenience functions addedUrlStore: validation by default
UrlStore: validation by defaultFull Changelog: https://github.com/adbar/courlan/compare/v0.7.0...v0.7.1
UrlStore class added: data store containing URLs with relevant information
UrlStore class added: data store containing URLs with relevant informationFull Changelog: https://github.com/adbar/courlan/compare/v0.6.0...v0.7.0
reviewed code base: simplicity and execution speed
more complex language heuristics, use langcodes
- enhanced cleaning - fixed language filter
keep trailing slashes to avoid redirection
URL manipulation tools added: extract parts, fix relative URLs
- improve filter precision
reduced dependencies: replace requests with bare urllib3, and tldextract with tld for Python 3.6 upwards
requests with bare urllib3, and tldextract with tld for Python 3.6 upwards- Python 3.9 compatibility - Simplified imports - Bug fixes
English and German language filters
- Less aggressive strict filters - CLI bug fixed
Cleaner and more efficient filtering
urllib.parseCleaning and filtering targeting non-spam HTML pages with primarily text
Your coding agent can read these notes before it upgrades. Set up the MCP server →