NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2112 most downloaded on PyPI
a modern parsing library
Last release 5 years ago
no release in 18 months
Ships fairly regularly
a new release about every 6 weeks
Some releases are documented
notes for 34 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
10 years old
66 releases · first in 2017
This is likely to be the last major release that supports Python 2 !
We are now working on a Python3.6+ only v1.0 branch, which will soon become the default. See the work in progress: https://github.com/lark-parser/lark/pull/925
We also have a new online IDE! Check it out here: https://lark-parser.github.io/ide
Lark can now generate standalone Javascript parsers! Check it out here: https://github.com/lark-parser/Lark.js (still in beta)
Using rule repeat (~ syntax) is now much much faster for large numbers, thanks to @MegaIng
Bugfix for the propagate_positions option. Added option value propagate_positions='ignore_ws'.
Fixed reconstructor for when keep_all_tokens=True
Added merge_transformers (Thanks Robin!)
Many minor bugfixes, and improvements to code and docs
Lark now tracks changes in imported grammars (%import), and updates the cache if necessary
%import), and updates the cache if necessaryLark.parse_interactive() for starting the parser in interactive modeAdded ast_utils, to assist in tranforming lark.Tree into a customized AST.
Better docs
Bugfixes
In the near future, Lark will drop support for Python 2. We will continue to develop for Python 3.6+ only, which will simplify the code and ease development.
Old releases (including this one) will still work, of course, and should be stable enough to accompany the remaining Python 2 users into the sunset.
If you have any objections, feel free to voice them here: https://github.com/lark-parser/lark/discussions/874
Thanks for everyone who helped make Lark better!
One column per quarter.
Better grammar re-use with the %override and %extend statements, which allow to rewrite and extend imported rules and tokens, similarly to class inher
%override and %extend statements, which allow to rewrite and extend imported rules and tokens, similarly to class inheritance. (See this example: https://github.com/lark-parser/lark/blob/master/examples/advanced/extend_python.py)Indenter now throws DedentError instead of AssertionError
Improved the Python3 grammar, now works with reconstructor. (See this example: https://github.com/lark-parser/lark/blob/master/examples/advanced/reconstruct_python.py)
Lots of refactoring for a better tomorrow.
rule/terminals names can now be in unicode. (thanks @julienmalard)
Better errors.
Better type hints.
lark.lark is now part of the standard library.
Earley:
Nothing published for this version
The LALR parser now supports priority in rules, as a way to resolve collision errors
LALR parser
The LALR parser now supports priority in rules, as a way to resolve collision errors
Improvements to the standalone tool, including more command-line options, like optional compression for the json data.
Improvements to the puppet error handling interface
Better error reporting on LALR collisions
Bugfixes in Earley
Misc
Added support for syntax highlighting in Atom
Fixes and improvements for the cache option. cache=True now uses a temporary directory instead of working directory.
Lark can now be imported directly from a zip (See: ed5c8ec51c4c6e8bd0ac80caff6afcb90a97d218)
Added more terminals to the grammar library (available for %import).
Nearley tools now supports case insensitive strings
Deprecated some interfaces
Improvements to docs, stubs, and various bugfixes
Thanks to @MegaIng for helping with Lark's maintenance, and to @ldbo, @chanicpanic, @michael-k, @ThatXliner and everyone else for their help and contributions.
Nothing published for this version
Nothing published for this version
Complete overhaul of documentation. Now using sphinx to generate API docs from docstrings. (commit 0664cbd3d3c19e321cae8df044839e7baf7135af. Thank you
Complete overhaul of documentation. Now using sphinx to generate API docs from docstrings. (commit 0664cbd3d3c19e321cae8df044839e7baf7135af. Thank you @chsasank !)
New and friendlier Earley SPPF interface! (commit 555b268eb26bcbfce64991ea7517338dee85a840. Thank you @chanicpanic !)
Added the ambiguity='forest' option. Added ForestTransformer and TreeForestTranformer.
Various Bugfixes to improve the handling of ambiguous results.
Read the docs here: https://lark-parser.readthedocs.io/en/latest/forest.html
New Vim syntax highlighting for Lark (https://github.com/lark-parser/vim-lark-syntax Thank you @omega16 !)
Lark now loads faster from cache (commit 7dc00179e63efa6e98d688bfba3265d382db79c4)
Terminals can now be composed of regexps and strings with different flags, if using Python 3.6+ (commit e6fc3c9b00306e3a8661210fcc93bf50479ee229)
Added support for parsing byte-strings, with the use_bytes flag (commit 9ee8428f3f6ad285ad93e2b62ec47d33fff54768).
UnexpectedToken exception now has the accepts attribute, which contains a list of terminals that would be accepted by the parser instead (in addition to the expects attribute, which is guided by the lexer and may include terminals that won't be accepted by the parser) (commit a7bcd0bc2d3cb96030d9e77523c0007e8034ce49)
Allow multiline regexes with the x flag (commit 9923987e94547ded8a17d7a03840c4cebce39188)
Lark no longer uses the default logger. Instead uses lark.LOGGER. (commit 7010f96825b5fbac79522d1b30689065df53dc8c)
Lark now notifies on unused terminals/rules through logging.debug.
Standalone generator now creates smaller files (without comments and docstrings). Also undergone various fixes. (commit bf2d9bf7b16cddb39f2e0ea3cefecc8de5269e2c)
Wheel distribution due to (somewhat) popular demand.
Lots of small bugfixes and improvements!
Many thanks to @MegaIng for his continued work on many of these new features and fixes, and to everyone else who contributed to Lark and helped make it even better.
on_error option to Lark.parse(). Read here: https://lark-parser.readthedocs.io/en/latest/classes/#larkparse
Added error handling to LALR!
on_error option to Lark.parse().
Read here: https://lark-parser.readthedocs.io/en/latest/classes/#larkparseSupport for better regexps with the regex module, when using Lark(..., regex=True)
Read here: https://lark-parser.readthedocs.io/en/latest/classes/#using-unicode-character-classes-with-regex
The last two releases were wrong. I apologize.
The last two releases were wrong. I apologize.
Hopefully that's the last of it, and we'll be back on track with periodic and accurate releases.
Nothing published for this version
Nothing published for this version
The main features for this release:
The main features for this release:
Grammar caching: It's now possible to cache the results of the LALR grammar analysis, for x2 to x3 faster loading. Use Lark(..., cache=True) or specify a file name. See here: https://lark-parser.readthedocs.io/en/latest/classes/
Grammar templates: Added support for grammar "functions" that expand in preprocessing. No docs yet, but see here for examples: https://github.com/lark-parser/lark/blob/master/tests/test_parser.py#L845
Lark online IDE: Technically not a feature, but it's possible to run Lark in the browser. Now we also have a simple IDE on github pages: https://lark-parser.github.io/lark/ide/app.html
Other changes:
Improved performance for large grammars
More debug prints when in debug mode
Better support for PyInstaller
Lots of bugfixes: mypy stubs, v_args, docs, and more.
Nothing published for this version
Nothing published for this version
Added the g_regex_flags option, to allow applying flags to all terminals.
g_regex_flags option, to allow applying flags to all terminals.end_pos for Earley, when using propagate_positionsAdded type stubs for all public APIs, in order to support type checking and completion using MyPy (or others)
Changes in this version are:
Added type stubs for all public APIs, in order to support type checking and completion using MyPy (or others)
Added two new methods to the Lark class: Lark.save() and Lark.load(). Both methods pickle and unpickle (respectively) the class instance into/from file objects. These can be used to allow faster loading times. (future versions will implement an automatic caching feature)
The standalone parser is now MPL2, instead of GPL. The Mozilla Public License is much less restrictive, so this shouldn't affect anyone who's already using the standalone parser. But it should make it easier for other users to adopt it.
Reverted maybe_placeholders to False by default. It didn't obey the semantic versioning standard.
Reverted maybe_placeholders to False by default. It didn't obey the semantic versioning standard.
Bugfix in standalone parser
The biggest change to this release is a new LALR engine, that is capable of dealing with a few edge cases that the previous parser couldn't.
- Better LALR
The biggest change to this release is a new LALR engine, that is capable of dealing with a few edge cases that the previous parser couldn't.
This parser is supposed to be fully backwards-compatible with the previous one, but that is hard to verify!
Thank you, @Raekye, for this great contribution to Lark!
For more details, see issue #418
- Transformers now visit tokens, as well as rules (an alternative to lexer_callbacks)
Transformer now visit tokens, in addition to rules.
Simply define a method with the correct name (uppercase, of course), and the transformer will visit your tokens before the rules that contain them.
It's possible to disable this, for backwards compatibility, or for the slight performance gain.
- Other Changes
Added visit_topdown methods to Visitor classes
Lark now allows line comments in its rule definitions
Better error messages
Improvements to documentation
Bugfixes
maybe_placeholders is now the default (backwards-incompatible)** (REVERTED in 0.8.1)
Improved error messages for EOF in Earley, recursive terminals, UnexpectedToken
Improved error messages for EOF in Earley, recursive terminals, UnexpectedToken
Bugfix for declared terminals, UnexpectedToken, unicode support in Python2,
Fixed a bug in Earley where running it from different threads produced bad results
Fixed a bug in Earley where running it from different threads produced bad results
Improved error reporting when using LALR
Added 'edit_terminals' option, to allow programmatical manipulation of terminals, for example to support keywords in different languages.
Note: This release skips 0.7.6, due to simple oversight on my part. Hopefully that shouldn't be a problem.
Nothing published for this version
Lark transformers can now visit tokens as well. Use like this:
Lark transformers can now visit tokens as well. Use like this:
class MyTransformer(Transformer):
def TOKEN1(self, tok):
return tok.upper()
def rule_as_usual(self, children):
return children
MyTransformer(visit_tokens=True).transform(tree)
Fixed a few regressions that I accidentally added to 0.7.4
Fixed long-standing non-determinism and prioritization bugs in Earley.
Fixed long-standing non-determinism and prioritization bugs in Earley.
Serialize tool now supports multiple start symbols
iter_subtrees, find_data and find_pred methods are now included in standalone parser
Bugfixes for the transformer interface, for the custom lexer, for grammar imports, and many more
Added a new tool called Serialize, that stores Lark's internal state as JSON. That will allow for integration with other languages. I have already sta
Added a new tool called Serialize, that stores Lark's internal state as JSON. That will allow for integration with other languages. I have already started such a project for Julia: https://github.com/erezsh/Lark_Julia (It's working, but still in early stages)
Minor bugfix regarding line-counting and the \s regex
Lark now allows you to specify the start symbol when calling Lark.parse() (requires pre-declaration of all possible start states, see the start option
New features:
Lark now allows you to specify the start symbol when calling Lark.parse() (requires pre-declaration of all possible start states, see the start option)
Negative priority now allows in rules and terminals (default value is still 1, may change in 0.8)
Also includes many minor bugfixes, optimizations, and improvements to documentation
Lark can now serialize its parsers, resulting in simplified stand-alone code.
Lark can now serialize its parsers, resulting in simplified stand-alone code.
Bugfix for v_args (Issue #350)
Improvements and bugfixes for importing rules from grammar files
Performance improvement for the reconstructor feature
This new version includes a brand new Earley implementation, with support for a Shared Packed Parse Forest (or SPPF), providing much better performanc
This new version includes a brand new Earley implementation, with support for a Shared Packed Parse Forest (or SPPF), providing much better performance for ambiguous grammars, both in terms of run time and memory consumption.
Users might notice a small degradation in Earley's run-time performance for deterministic grammars. Hopefully future releases will take care of that as well.
Big thanks to @night199uk for his great contribution.
Other features worth mentioning (though they exist in the previous release):
Added support for importing rules between grammars. The import mechanism is namespace-aware.
Added the maybe_placeholders option, which causes optionals of the form [expr] to return None when not matched, instead of just not appearing. (optionals of the form expr? maintain the previous behavior of not appearing unless matched)
Plenty of bugfixes, better errors, and better docs
Fixes regarding VisitError, and Discard
Fixes regarding VisitError, and Discard
Fixes to the lexer callbacks feature
Fix for indenter, allowing re-use of same instance
(Points to the wrong commit, due to technical issues. Actually refers to commit 13ddc43782e9f3f34fc5331081682c3678df0598)
(Points to the wrong commit, due to technical issues. Actually refers to commit 13ddc43782e9f3f34fc5331081682c3678df0598)
This release includes:
Better error reporting
Several bugfixes
A new experimental feature: "maybe_placeholders", which replaces missing "maybe"s with a None, instead of removing them.
For example:
>>> p = Lark("""!start: "a"? "b"? "c"? """, maybe_placeholders=True)
>>> p.parse('b')
Tree(start, [None, Token(B, 'b'), None])
>>> p.parse('ac')
Tree(start, [Token(A, 'a'), None, Token(C, 'c')])
Standalone parser now uses the contextual lexer
New features:
Standalone parser now uses the contextual lexer
Experimental support for importing rules in the grammar
Bugfixes:
Also improved error messages and documentation
Added MkDocs documentation (will replace github's wiki as official docs)
Added support for relative imports: %import .local_grammar.TERM will look for ./local_grammar.lark
Added support for relative imports: %import .local_grammar.TERM will look for ./local_grammar.lark
Added syntax for importing several terminals on one line: %import common (NUMBER LETTER FLOAT)
Several bug-fixes
Fixed a bug in PropagatePositions, that could occur when the user supplied Lark with a reducing transformer
Fixed a bug in PropagatePositions, that could occur when the user supplied Lark with a reducing transformer
Added support for v_args(tree=True), which supplies the decorated transformer method with the entire tree (instead of just the list of children)
Lark grammars are now utf8 by default
Lark grammars are now utf8 by default
Added option to provide a custom lexer (with example)
Fixed issue where Lark would throw RecursionError for huge grammars
Improved error messages
Revised transformers and visitors. Added v_args. inline_args and InlineTransformer are now deprecated. Instead, use v_args(inline=True) as a decorator…
Breaking changes:
Changes to Tree: Added meta attribute, in addition to data and children. Line & column attributes, when using propagate_positions, moved to meta (i.e. tree.meta.line)
Revised transformers and visitors. Added v_args. inline_args and InlineTransformer are now deprecated. Instead, use v_args(inline=True) as a decorator on transformer methods and class definition.
Changed default Earley lexing behavior - now returns the maximum match only. The original behavior, that attempts to match all appearances of a terminal, has been moved to the "dynamic_complete" lexer. (commit 6ea4588)
Restructured exceptions - UnexpectedInput is now superclass of UnexpectedToken and UnexpectedCharacters, all of which support the get_context() and match_examples() methods. (commit 5c6df8e)
Columns now start at 1 (instead of 0)
Default LALR lexer in now contextual
Removed "scanless parsing" mode (wasn't useful)
Other changes:
Added %declare directive for plugin support.
Default extension for Lark grammars is now .lark (with syntax highlighting for popular editors)
Improved error reporting
Lots of bugfixes, better performance, and cleaner code.
Bugfix in Earley prioritization
Also better error messages (there's still much to do in that regard)
Fixed an inefficiency in the tree builder, significant for big inputs
Fixed propagate positions feature.
Fixed propagate positions feature.
Added support for ranged-repeat (see commit for details)
Reconstruct now working again (experimental module)
A few bugfixes and refactors
Also contains several usability fixes
Also contains several usability fixes
Added a tool to generate standalone parsers for LALR(1)
Added a tool to generate standalone parsers for LALR(1)
Significant improvements to Earley and the way it handles ambiguity and priority.
Significant improvements to Earley and the way it handles ambiguity and priority.
Users can now return Discard in transformer callback to drop a child branch from its parent.
Terminals now accept every regexps flag.
Cleaned up the LALR(1) parser implementation
Many bugfixes.
Started making releases :)
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →