NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #5126 most downloaded on PyPI
TatSu takes a grammar in a variation of EBNF as input, and outputs a memoizing PEG/Packrat parser in Python.
Last release 2 months ago
09 Jul 2026
Release timing varies
gaps range from 9 days to 3 months
Nearly every release is documented
notes for 34 of 36 stable releases
19 versions withdrawn
withdrawn after publishing
9 years old
55 releases · first in 2017
See the [CHANGELOG] for details.
See the CHANGELOG for details.
See the CHANGELOG for details Full Changelog: https://github.com/neogeny/TatSu/compare/v5.22.0...v5.23.0
See the CHANGELOG for details
Full Changelog: https://github.com/neogeny/TatSu/compare/v5.22.0...v5.23.0
One column per quarter.
See the CHANGELOG for details.
See the CHANGELOG for details.
See the CHANGELOG for details.
See the CHANGELOG for details.
See the CHANGELOG for details.
See the CHANGELOG for details.
<!-- Copyright (c) 2017-2026 Juancarlo Añez (apalala@gmail.com) SPDX-License-Identifier: BSD-4-Clause -->
<!-- Copyright (c) 2017-2026 Juancarlo Añez (apalala@gmail.com) SPDX-License-Identifier: BSD-4-Clause -->
There are ports of TatSu to Go and Rust. They are functionaly complete with except for features (like synthetic classes) that rely on the dynamic nature of Python.
铁修 TieXiu is the port of TatSu to Rust. It features a PyO3 interface os it's also a Python library, but the benchmarks show that the pure-Python parsers generated by TatSu are still more performant when hosting from Python. See the TieXiu README for a discussion of the performance limits of PEG parsers.
⻰OGoPEGo is the port of TatSu to Go. The implementation, being the most mature, is beutifully concise, and using the generated parsers has a simplicity closer to the Python style allowed by TatSu.
The algorithm for left-recursion analysis went over another round of simplification and optimization. Then the analysis done in pegen, a more efficient and theoretically-sound approach, was evaluated. All tests pass with the pegen's SCC (Strongly Connected Componets) algorithm, so the old-and-tried algorithm in TatSu was replaced.
Although left-recursion analysis is performed once per Grammar, before any parsing, a simpler implementation makes this core part of TatSu easier to maintain.
The g2e (ANTLR grammar to TatSu) translator has been revived and significantly simplified. A working example (python3.tatsu, 551 lines) is generated from Python 3's full ANTLR grammar and passes tatsu.compile().
Removed the regex conversion approach for ANTLR token rules. ANTLR lexer patterns (notably \uXXXX escapes) are not viable as Python regex patterns. Non-trivial token rules now emit Fail() instead of Pattern. The classes and methods TokenPattern, _token_expr_to_regex, _token_expr_to_regex_verbose, _decode_antlr_string, and _char_to_regex have been removed.
g2e substitutes simple token definitions (like OPEN_PAREN : '(' {opened++;};) for their right hand side (just '(') for better looking grammars. For complex token definitions ANTLR uses a special syntax which is not that of Python-compatible (PCRE2) regular expressions, so g2e omits them, leaving it to the user to decide how to handle those tokens. In many cases a single pattern match is enough for the grammar of interest, and a semantic rule may be added to validate additional conditions that the parsed token should meet.
Streamlined generated grammar output — removed unnecessary parenthesization:
(NEWLINE) → NEWLINE.[...], {...}, {...}+ unwrapped: [('as' NAME)] → ['as' NAME], {('.' NAME)} → {'.' NAME}.tokens {} declarations that collide with defined rules (e.g. INDENT/DEDENT).Token name resolution now uses uppercase names consistently.
The g2e example (examples/g2e) uses the old, LL(1) Python grammar. Now, since Python's PEG parser the actual grammar is a much simpler one. The example is kept as it was to demonstrate g2e's behavior over a complex grammar.
A new --recursion-limit (-R1) option was added to the tatsu CLI tool so it can handle large and deeply recursive input grammars. When used as a library, the host program should call sys.setrecursionlimit() when required by the grammar complexity.
Added better rendering to FailedParse.__str__(). Now a code fragment and line numbers are shown, as in many modern tools.
error: expecting 'world'
--> example:1:7
|
1 | hello missing
| ^ expecting 'world'
-> start
tatsu.ebnf define rules for JSON literals, so true, false, and null,
may be used where previously only True, False, and None were recognized.
The Python literals are still honored as before, as well as the boolean rule resolving to True for non-falsy values. These literals are only used in grammar directives, as parsing is only interested in the strings that match a Token or Pattern.
Now a Grammar can be imported from the JSON produced by model.asjson(). Roundtrip has been tested and it works. New methods Grammar.load(value: Any) -> Grammar and Grammar.loads(json: str) -> Grammar make the functionality available.
class Grammar:
@staticmethod
def load(value: Any) -> Grammar:
from .json import load_grammar
return load_grammar(value)
@staticmethod
def loads(value: str) -> Grammar:
from .json import loads_grammar
return loads_grammar(value)
The definition of the DEDENT rule in the TatSu grammar is used to support EBNF notations with no rule-terminatiors and grammars with no blank lines * rules. The pattern used in the rule was incorrectly consuming the first non-space character starting the next rule. Fixed.
Now this is a valid EBNF definition:
grammar = r"""
@@grammar :: MiniJSON
@@nameguard :: False
@@whitespace :: /\s+/
start: value $
value: object | array | string | number | 'true' | 'false' | 'null'
object: '{' members? '}'
array: '[' elements? ']'
members: pair (',' pair)*
elements: value (',' value)*
pair: string ':' value
string: '"' CONTENT '"'
CONTENT: /[^"]*/
number: /-?\d+(\.\d+)?/
"""
Lookaheads are always memoized. Configuration settings for disabling it have been deprecated and disabled.
The $-> (EOL) expression was introduced in the grammar language to match and consume the whitespace up to and including the next line break, using the Python semantics of os.linesep. The match interprets whitespace using the Python definition as implemented by str.isspace(), so beware when
a particular definition of whitespace is part of the language to parse.
The @nostak decorator for rules was added to the grammar. The setting hints the tracer and error handler that the rule should not be part of the call stack. The setting is useful to avoid noise in traces when low-level rules (like those for qualified or attributed identifiers) form their own small hierarchy.
The file extension for TatSu grammars is now .ebnf. The grammar language is, after all, an extension of the most known forms of EBNF syntax. Syntax highlighters may recognize the extension
The benchmark in tatsu.tool.bench was used over several large grammars and large input sets to evaluate parser strategies. The result is that there is a 1.3x performance advantage in generating a Python program versus using the in-memory model of the parsed TatSu grammar for parsing. In tests with complex projects (Java) the performance difference is not perceivable. The codspeed benchmark that runs with unit tests on GitHub doesn't see the performance difference either.
Now TatSu uses for bootstrap a module that loads its own grammar model as the main parser (the one used by tatsu.compile()). The previous kind of parser can still be generated with tatsu.to_python_sourcecode(), which remains well tested in several unit tests. The new model-based kind of parser can be generated with tatsu.to_parsermodel_sourcecode().
Note that you don't need to generate any source code for a parser in your own projects. TatSu does generate a module to make it faster to bootstrap a parser from its own grammar. In your projects you can run the usual steps to have a performant parser:
import tatsu
grammartext = ...
model = tatsu.compile(grammartext, asmodel=True)
output = model.parse(input)
Generating a module with classes for the type definitions in the grammar is still useful.
from pathlib import Path
import tatsu
grammartext = ...
sourcecode = tatsu.to_python_model(grammartext)
Path('./modelclases.py').write_text(sourcecode)
Optimizations in the parser logic produce parsing speeds comparable to those of TatSu v5.16 with any parsing strategy (model or generated code).
The old parser and model generator modules in tatsu.codegen have been deleted. Using pyrefly revealed that they are both incorrect and non-working. Their defunctness was caused by the lack of unit tests and their lack of use since tatsu.ngcodegen was introduced several years ago. The helper modules codegen.cgbase and codegen.rendering remain in case any old projects use them for their own code generation.
The g2e example in ./examples/g2e was removed. The example had become irrelevant now that the new PEG parser in Python uses a pegen-style grammar for the language that is less than a 1000 lines long. The TatSu grammar for ANTLR in ./examples/g2e/antlr.tatsu can still parse ANTLR grammars, but there's no test case for it. The semantics in g2e.semanrics.ANTLRSemantics try to do everything on a single pass (like substituting simple TOKEN rules by their value), when transformation of the parsed input grammar model should be more stable and easier to understand with a simplerr approach.
There's no longer a separate stack for the state of cut. The state of cut is kept in the general state stack.
A new @statescope context manager takes care of handling the state stack in most cases.
Lookaheads are always memoized. Configuration settings for disabling it have been deprecated and disabled.
A new PaserConfig.perlinememos: float configuration sets a (perlinememos * linecount)
bound on the total number of memoization entries that are allowed on each parse.
Incorporated zuban to the set of type linters.
Introduced objectmodel.ctx.CanParse(Protocol) defining the parse() method
for entry point to parsing.
An important refactoring was done to get rid of the legacy names "tokenizing" and "tokenizer" which didn't abide to theory and practice of parsing. Now the names are tatsu.input, tatsu.input.text, and tatsu.input.text.Text. The old names are still available as legacy for backwards compatibility.
Rule invludes (RuleInclude) kept an atcutal copy of the included rule in the model. To preserve consistent semantics, the only mentions of Rule in a model are at the top-level, in Grammar.rules and Grammar.rulemap.
Grammar models that haven't been compiled from a grammar but instead loaded from the JSON or Python representations don't need to be analyzed for left recursion, because the markers of the analysis are already in the loaded models. A new Grammar.analyzed: bool attribute was added to quickly check if a grammar model from any source has already been analyzed.
Support for #include in grammars has been dropped. It was always a bad idea. Text-to-text preprocessing doesn't belong in the grammar in part because it doesn't apply to input sources that are not text, like that of tokenizers or streams. The class tatsu.input.buffer.Buffer still has all the infrastrucure for supporting C-style or COBOL-style textual includes, and its definition of BufferCursor honors it. Buffer keeps track of which file was the source of each line of input, something essential for good error reporting. During compilation of grammar text to a Grammar object, the grammar text is the parser's input, so the Cursor semantics regarding the parsing still apply.
The CLI tool now has a --json option to produce the JSON version of the model for a grammar. Re-importing of a JSON model is not yet implemented in TatSu, but TieXiu uses them successfully as the fast way to import a TatSu grammar model.
Now 竜 TatSu’s own grammar is written in EBNF notation. Examples and documentation are converging to the syntax used by pegen, the base for the Python
Now 竜 TatSu’s own grammar is written in EBNF notation. Examples and documentation are converging to the syntax used by pegen, the base for the Python PEG parser.
Multi-line string literals in the grammar are now supported. Use triple
quotes like """…""" or '''…''' for multi-line string literals
in the grammar.
The (?:…) expression was added to grammars. It works like a (…) group but the expression parsed is not captured.
There’s now a copy of the 竜 TatSu grammar under the main package at ./tatsu/_grammar.tatsu. The grammar text is available as tatsu.grammar. The grammar remains available at ./grammar/tatsu.tatsu by a symbolic link.
The semantics of grammar representations and parsing where spread all over the codebase, in a way that made them hard to understand and maintain. Now each aspect of the semantics is defined in a single, small module. Some modules of interest, worth studying, are:
The module hierarchy has been refactored to avoid internal dependency cycles, to simplify imports, and to not require partial module parses by Python. Large modules have been factored at semantic boundaries into smaller modules within a package.
Left recursion is detected at grammar model creation time, and a GrammarError is raised if left recursion was not enabled in the grammar or configuration parameters.
Generated parsers and models now use @tatsu.rule and @tatsu.dataclass as appropriate.
Generated parsers now use anonymous functions to implement choices and closures:
with ctx.choice() as ch:
@ch.option
def _():
self.word(ctx)
@ch.option
def _():
self.string(ctx)
ch.expecting('<string>', '<word>')
Now in addition to the existing rules and rulemap attributes of Grammar, there is a rule attribute that allows access to rules as attributes:
class Grammar(Model):
rules: tuple[Rule, ...]
rulemap: dict[str, Rule]
rule: SimpleNamespace
model = tatsu.compile(grammar)
rule = mode.rule.start
print(m.rule.longone.exp.token)
Rules and the different forms of closures are back to returning list instead of tuple. The tuples were introduced as a not-thought-up shortcut and it has taken until now to fix that. Square brackets are much easier to discern in programs and output in a context in which parenthesis are already overloaded.
str(model) returns the standard __str__() output and no longer returns model.pretty(). To obtain a grammar representation of a grammar model use model.pretty() directly.repr(model) no longer returns asjsons(model) but instead returns a representation that can be used to reconstruct the model: model = tatsu.parse(grammar, asmodel=True)
evalmodel = eval(repr(model), globals=vars(grammars))
assert repr(model) == repr(evalmodel)
m.pretty() -> str: # pretty-printed grammar
m.asjson() -> Any: # object compatible with json.dumps()
m.asjsons() -> str: # json.dumps(m.asjson(), indent=2)
m.railroads() -> str: # a railroads diagram in Text/ASCII art
repr(m) -> str: # as an expression can reconstruct the model
Tokenizer classes provide implementations of the Cursor protocol which holds the state (e.g. the position) of a parse, while the tokenizer acts as the source of the input stream.ParseContext was moved to StateStack that only takes a Cursor as initialization parameter. It’s possible to have more than one parse on the same input because the state of the parsing is separate from the ParseContext and the Tokenizer.Parser, but can be declared in any class. The convention of naming the methods with a leading and a trailing underscore was removed so methods are now named like the grammar rules they represent.ctx: Ctx parameter to access and pass the invocation ParsContext. Ctx is the protocol that defines only the interface that methods for rules require to perform a parse according to the input grammar.regex + regex, adding regular expressions.?/…/? regexes in the grammar.ParseContext.substate was removed.TatSu can now produce railroad diagrams in *ASCII/Text Art*. There's an example for TatSu's own grammar in the current README.
TatSu can now produce railroad diagrams in ASCII/Text Art. There's an example for TatSu's own grammar in the current README.
The functionality is available from the tatsu.railroads module, and also on the command line:
--railroad, -r output a railroad diagram of the grammar in ASCII/Text Art
This release fixes the incorrect and forceful import of the rich and multiprocessing modules (#397/#399).
Now the Python source code in the project is formatted with black with skip-string-normalization = true.
*CAVEAT:* Several functions, methods, and argument names were deprecated. They can still be used, but *warnings* will be issued at runtime.
Maintenance and contributions to TatSu have been more difficult than necessary because of the way the code evolved through its lifetime.
This release is a major refactoring of the code in TatSu.
For the details about the many changes please take a look at the commit log.
Every effort has been made to preserve backwards compatibility by keeping most unit tests intact and testing with projects with large grammars and complex processing. If something escaped those tests, there will be a bugfix release with the fixes soon enough.
The TatSu documentation has been improved and expanded, and it has a better look&feel with improved navigation.
TatSu doesn't care about file names, but the default extension used in unit tests, examples, and documentation for grammars is now .tatsu
EBNF, both ISO and the classic variations, is fully supported as grammar input format
Now tatsu.parse(...., asmodel=True) produces a model that matches the ::Type declarations in their grammar (see the models documentation for a thorough review of the features).
walkers.NodeWalker now handles all known types of input.
Also:
DepthFirstWalker was reimplemented to ensure DFS semanticsPostOrderDepthFirstWalker walks children before parentsPreOrderWalker was broken and crazy. It was rewritten as a BreadthFirstWalker with the correct semanticsConstant expressions in a grammar are now evaluated deeply with multiple passes of eval() as to produce results that are intuitively correct:
def test_constant_math():
grammar = r"""
start = a:`7` b:`2` @:```{a} / {b}``` $ ;
"""
result = parse(grammar, '', trace=True)
assert result == 3.5
Evaluation of Python expressions by the parsing engine now use safe_eval(), a hardened firewall around most security attacks targeting eval() (see the safeeval module for details)
Because None is a valid initial value for attributes and a frequent return value for callables, the required logic for undefined values was moved to the notnone module, which declares Undefined as an alias for notnone.NotNone
In [1]: from tatsu.util.undefined import Undefined
In [2]: u = Undefined
In [3]: u is None
Out[3]: False
In [4]: u is Undefined
Out[4]: True
In [5]: Undefined is None
Out[5]: False
In [6]: d = u or 'OK'
In [7]: d
Out[7]: 'OK'
objectmodel.Node was rewritten to give it clear semantics and efficiency
Node after initialization generate a warning if the name of a method is being shadowed. This change avoids confusing @dataclass, which is used in generated object models.Node equality is explicitly defined as object identity. No attempts are made at comparing Node structurally.Node.children() has the expected semantics, and is much more efficient.Node.parseinfo is now honored by the parsing engine (previously, only results of type AST could have a parseinfo). Generation of parseinfo is disabled by default, and is enabled by passing pareseinfo=True to the API entry points.
def test_node_parseinfo(self):
grammar = """
@@grammar :: Test
start::Test = true | false ;
true = "test" @:`True` $;
false = "test" @:`False` $;
"""
text = 'test'
node = tatsu.parse(grammar, text, asmodel=True, parseinfo=True, )
assert type(node).__name__ == 'Test'
assert node.ast is True
assert node.parseinfo is not None
assert node.parseinfo.pos == 0
assert node.parseinfo.endpos == len(text)
Synthetic classes created by synth.synthetize() during parsing with ModelBuilderSemantics behave more consistently, and now have a base class of class SynthNode(BaseNode)
Now ast.AST has consistent semantics of a dict that allows access to contents using the attribute interface
asjson() and friends now cover all known cases with improved consistency and efficiency, so there are less demands over clients of the API
Entry points no longer list a large subset of the configuration options defined in ParserConfig, but still accept them through **settings keyword arguments. Now ParserConfig verifies that the settings passed to are valid, eliminating the frustration of passing an incorrect setting name (a typo) and hoping it has the intended effect.
TatSu still has no library dependencies for its core functionality, but several libraries
are used during its development and testing. The TatSu development configuration uses uv and hatch. Several requirements-xyz.txt files are generated in favor of those using pip with pyenv, virtualenvwrapper, or virtualenv
All attempts at recovering comments from parsed input were removed. It never worked, so it had no use. Comment recovery may be attempted in the future.
All pre-existing grammars are compatible with this version of TatSu.
Previously generated Python parsers and models, work with this version of TatSu, yet you should consider generating them anew to take advantage of the improved speed, layout, and features.
CAVEAT: Several functions, methods, and argument names were deprecated. They can still be used, but warnings will be issued at runtime.
CAVEAT: If there are invalid strings or regex patterns in your grammars YOU MUST fix them because now the grammar parser validates strings and patterns.
Many of the functions that TatSu defines for its own use are useful in other contexts. Some examples are:
from tatsu.safeeval import is_eval_safe
from tatsu.safeeval import hasshable
from tatsu.safeeval import make_hashable
from tatsu.util import safe_name
from tatsu.util.misc import find_from_rematch
from tatsu.util.misc import topsort
from tatsu.util.undefined import Undefined
# ...
Several refactorings, optimizations, and deprecations
ParserConfig away from infosParserConfig is now documentedtokenizer.Tokenizer is now a Protocol./test to ./tests[refactor] fix imports of asjson() and related
[contexts] reinstate cache pruning on cut
ownerlimit cuts to the closest choice
Bug fixes and suggestions by contributors
promote TatSu-LTS to support of previous versions of Python
remove comments_re and eol_comments_re from parser configuration (ParserConfig). Use comments and/or eol_comments instead
comments_re and eol_comments_re from parser configuration (ParserConfig). Use comments and/or eol_comments instead (#351)re.MULTILINE to compiled regexes. Users must add (?m) to the expressions for multiline (#351)Test agains Py313 and latest libraries
Test agains Py313 and latest libraries
Honor @@whitespace::None` in generated Python parsers
make sure that the names in @@keywords are always of type str
@@keywords are always of type strasjson()In #333 it was reported that pip install tatsu would also install a test package. This is fixed now.
In #333 it was reported that pip install tatsu would also install a test package. This is fixed now.
Nothing published for this version
Do not to resolve a model name when the ::Annotation in the grammar is a basic type like int or bool.
Do not to resolve a model name when the ::Annotation in the grammar is a basic type like int or bool.
[docs] deprecate declarative translation abd refactor
This release uses the new procedural (nor declarative) code and model generation throughout.
The previous codegen remains available and unchanged for backwards compatibility.
ngcodegenngcodegenThe undocumented parproc module helps to easily run parsing and translation batches in parallel.
The undocumented parproc module helps to easily run parsing and translation batches in parallel.
print() statements stranded in buffering.py
print() statements stranded in buffering.py
[buffering] do not re.escape regex for whitespace
include ./examples in source distributions
pyproject.toml./examples in source distributionsmake sure generated parsers and models pass ruff
ruff rulesruff[exceptions] generators in exceptions break multiprocessing
…on the __json__() protocol and thus create a backwards incompatibility.
This release adds fixes, enhancements, optimizations, and documentation to the asjson() protocol that is used to view ASTs and models.
The bump in the minor version number because the changes to asjson() will likely break code that relies on the __json__() protocol and thus create a backwards incompatibility.
[codegen] pass whitespace setting to parser
This release makes TatSu compatible with Python >= 3.8, but still states that the compatible version is Python >= 3.10.
This release makes TatSu compatible with Python >= 3.8, but still states that the compatible version is Python >= 3.10.
Fix walker cache logic in NodeWalker (@by-Exist)
NodeWalker (@by-Exist)Fix pickling of AST. The change also affects parsing within multiprocessing (@fizbin)
repr() for types not handled by json(@dnicolodi)Make sure that the generated parser is the same as the bootstrap parser.
re.findall(pattern, text)[0]. Now groups that should not be
returned when parsing should use the (?:) syntax. Now patterns
align with Python re independently of the
@@ignorecase setting for the grammar.{} interpolation in `constant[ expressions with the
semantics of ]{.title-ref}[str.format()]{.title-ref}`[constant]{.title-ref} as a multiline version of
`constant`.[constant]{.title-ref} and ^`constant[ as syntax
for an `alert]{.title-ref} expression. Alerts produce no tokens bug
get registed in [parseinfo]{.title-ref} records.-> skip expression always stop at EOF.-> skip expression go over comments, and not log while
skipping.FailedCut while running parsing in
parallel.parproc~ cut expressions to parse tracesNothing published for this version
Fix that settings passed to Context.parse() were ignored. Add Context.active_config for the configuration active during a parse
Context.parse() were ignored. Add Context.active_config for the configuration active during a parseNode._parent as part of the @dataclassMake AST and Node hashable. Necessary for caching Node.children()
AST and Node hashable. Necessary for caching Node.children()Node.__eq__() in terms of identity or Node.ast.__eq__()__Node.ast (was removed because of problems with __eq__())Node.children() from Node.ast when there are no attributes defined for the Node. This restores the desired behavior while developing a parse model.Now config: ParserConfig is used in __init__() and parse() methods of contexts.ParseContext, grammars.Grammar, and elsewhere to avoid the very long pa
config: ParserConfig is used in __init__() and parse() methods of contexts.ParseContext, grammars.Grammar, and elsewhere to avoid the very long parameter lists that abounded. ParseContext also provides clean and clear ways of overridinga group of settings with anotherAST_. Names within optionals that did not match will have their values set to None, and closures that did not match will be set to ``[]`setup.py in favor of setup.cfg and pyproject.toml (@KOLANICH_)Node.children() is now computed only when required, and cached@dataclassNothing published for this version
Fix bug in which rule fields were forced on empty AST (@Victorious3)
AST (@Victorious3)Several important refactorings in contexts.ParseContext
contexts.ParseContextignorecase settings apply to defined @@keywordsParseContextASTtatsu/bootstrap.py) to the generated parsermain() only outputs the JSON for the parse ASTTest with Python 3.9 (@apalala)
The default regexp for whitespace was changed to \` (?s)s+
//) like Python does@namechars@@whitespace :: None and @@whitespace :: False@keyword throughout the grammarAST or Node even if the associated expression was not parsed@@directive* #66 Fix multiline ( (?x) ) patterns not properly supported in grammar (@pdw-mb) * #70 Important upgrade to ModelBuilder and grammar specification of
Fix typos in documentation (@mjdominus
* #42 Rename vim files from grako.vim to tatsu.vim (@fcoelho) * #51 Fix inconsistent code generation for whitespace (@fpom) * #54 Only care about case
grako.vim to tatsu.vim (@fcoelho)whitespace (@fpom)* #37 Regression: The #include pragma works by using the EBNFBuffer from grammars.py. Somehow the default EBNFBootstrapBuffer from bootstrap.py has be
#37 Regression: The #include pragma works by using the EBNFBuffer from grammars.py. Somehow the default EBNFBootstrapBuffer from bootstrap.py has been used instead (@gegenschall).
#38 Documentation: Use of json.dumps() requires ast.asjson() (@davidchen).
* #27 Undo the fixes to dropped input on left recursion because they broke previously expected behavior. * #33 Fixes to the calc example and mini tuto
#27 Undo the fixes to dropped input on left recursion because they broke previously expected behavior.
#33 Fixes to the calc example and mini tutorial (@heronils)
#34 More left-recursion test cases (@manueljacob).
Documentation fixes (@manueljacob, @paulhoule)
#27 Left-recursive parsers would drop or skip input on many combinations of grammars and correct/incorrect inputs(@manueljacob)
Documentation fixes (@manueljacob, @paulhoule)
Parse speeds on large files reduced by 5-20% by optimizing parse contexts and closures, and unifying the AST_ and CST_ stacks.
Parse speeds on large files reduced by 5-20% by optimizing parse contexts and closures, and unifying the AST_ and CST_ stacks.
Added the "skip to" expression ( ->), useful for writing recovery rules. The parser will advance over input, one character at time, until the expression matches. Whitespace and comments will be skipped at each step.
Added the any expression ( /./) for matching the next character in the input.
The ANTLR_ grammar for Python3_ to the g2e example, and udate g2e to handle more ANTLR_ syntax.
Check typing with Mypy_.
Removed the very old regex example.
Make parse traces more compact. Add a sample to the docs.
Explain Grako_ compatibility in docs.
tatus.objectmodel.Node was not setting attributes from AST
tatus.objectmodel.Node was not setting attributes from ASTNew support for *left recursion* with correct associativity. All test cases pass.
@@left_recursion :: False directive to diasable it.@tatsumasu.tatsu.contexts.ParseContext for clarity.@@ignorecase directive and the ignorecase= parameter no
longer appy to regular expressions (patterns) in grammars. Use
(?i) in the pattern to ignore the case in a particular pattern.tatsu.g2e is a library and executable module for translating
ANTLR grammars to TatSu.calc example and made it part of the documentation
as Mini Tutorial.This is the initial release of TATSU.
This is the initial release of TATSU.
Your coding agent can read these notes before it upgrades. Set up the MCP server →