NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1441 most downloaded on PyPI
Pure-Python robots.txt parser with support for modern conventions
Last release 13 days ago
21 Sep 2026
Release timing varies
gaps range from 9 days to 2.2 years
Most releases are documented
notes for 11 of 15 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
16 releases · first in 2019
Bump version: 0.6.2 → 0.7.0
Bump version: 0.6.2 → 0.7.0
Backward-incompatible: Protego.parse() now raises a more suitable TypeError instead of a ValueError when content is not a string.
Added official support for Python 3.15.
can_fetch() now allows /robots.txt whatever the rules say, as required by RFC 9309.
A group now applies to a user agent only if its product token appears in that user agent at a token boundary, so that User-agent: bot no longer applies to mybot. A product token within a whole User-Agent header value, e.g. Mozilla/5.0 (compatible; mybot/1.0), still matches.
Percent-encoding is now normalized in the query string and the parameters of a URL, and not only in its path, so that a rule matches whichever spelling either of them uses, e.g. Disallow: /a?b=ツ now matches /a?b=%E3%83%84. The fragment is now left out of matching altogether.
Visit-time now also accepts a hyphen between the two times, e.g. 0400-0845, and times with a single-digit hour, e.g. 400.
Malformed Crawl-delay values, such as a negative or infinite number, and Request-rate values with zero requests or zero seconds, are now ignored, as other malformed values already were.
Improved the handling of invalid and misspelled lines.
Sitemap and Host lines are now read anywhere in the file, including before any User-agent line, and no longer merge the groups they sit between. A User-agent line without a value still ends the preceding group. Lines written without a colon are now salvaged for every directive, Visit-time included, and with any whitespace as the separator. A line whose field is not a known directive is no longer read as one.
Fixed matching of URLs whose path contains =, which no rule could match because it was percent-encoded in the URL but not in the rule.
Fixed matching of URLs whose path starts with //, which got extra slashes before being matched.
Fixed Allow: …/index.html, which, besides allowing the parent directory as intended, also allowed URLs with a $ right after that directory.
Improved matching and parsing performance.
Documentation improvements, including a rewritten parser comparison table generated from benchmarks.
One column per quarter.
Wildcard matching is now performed without regular expressions. Please, see the CVE-2026-55520 and GHSA-wjmf-p669-5m5p security advisories for more in…
Fixed a ReDoS (regular expression denial of service) vulnerability: URL
patterns from robots.txt Allow and Disallow directives were
compiled into regular expressions, where multiple * wildcards could
cause exponential backtracking. A server could exploit this to cause denial
of service by serving a crafted robots.txt file. Wildcard matching is
now performed without regular expressions. Please, see the
CVE-2026-55520 and GHSA-wjmf-p669-5m5p security advisories for more
information.
Fixed parsing of Request-rate values where the seconds field has no time-unit suffix (e.g. 1/60 instead of 1/60s ). Previously the last digit of the n
Request-rate values where the seconds field has no time-unit suffix (e.g. 1/60 instead of 1/60s). Previously the last digit of the number was silently dropped.Added official support for Python 3.14.
Restructured the code, splitting the single protego.py file into multiple modules. The public API remains the same but some internal names may now be
protego.py file intopy.typed.setuptools to hatchling.setup.py to pyproject.toml.Dropped Python 3.8 support, added official Python 3.13 support.
Added official support for Python 3.12.
Added official support for Python 3.12.
= is no longer percent-encoded in patterns, fixing many scenarios where
patterns included query strings.
Dropped support for Python 2.7, 3.5, 3.6, and 3.7, and added support for 3.11 and for the upcoming 3.12.
Changed requirements:
Dropped support for Python 2.7, 3.5, 3.6, and 3.7, and added support
for 3.11 and for the upcoming 3.12.
six is no longer a dependency.
Added support for the Visit-Time directive.
Fixed leading asterisks in allow and disallow values not being properly
interpreted.
Protego.parse() now raises value error when content is not a string.
Fixes incorrect readme content-type specified in setup.py
content-type specified in setup.py (#21)Fixes disallow not working when no path is provided
Protego 0.1.16 fixes ( #8 ) an issue with robots.txt files containing absolute URLs as values for allow and disallow directives, where their path woul
Protego 0.1.16 fixes (#8) an issue with robots.txt files containing absolute URLs as values for allow and disallow directives, where their path would be interpreted ignoring their protocol and netloc, leading to unexpected results (#4). Now, values of allow and disallow directives are always interpreted as if they started with a /, as Google’s specification dictates that they must always start with /.
We’ve also improved (#3) our test coverage for malformed disallow directives.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →