NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #772 most downloaded on PyPI
Lightweight static analysis for many languages. Find bug variants with patterns that look like source code.
Last release today
02 Oct 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 54 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
6 years old
360 releases · first in 2020
One column per quarter.
Scala: ellipsis are now allowed in for loop headers, so you can write patterns like for (...; $X <- $Y if $COND; ...) { ... } to match nested for loop
for (...; $X <- $Y if $COND; ...) { ... } to match nested for loops. (#5650)--verbose no longer toggles the display of timing information, use
--verbose --time to display this information.semgrep ci: CI runs in GitHub Actions failed to checkout the commit assoociated with the head branch, and is fixed here.
semgrep ci: CI runs in GitHub Actions failed to checkout the commit assoociated with the head branch, and is fixed here.Bash: Support for subshell syntax i.e. commands in parentheses
semgrep ci: CI runs were failing to checkout the PR head in GitHub Actions, which is
corrected here.pattern-propagators now works correclty when the
from or to metavariables match a function call. For example, given
sqlBuilder.append(page.getOrderBy()), we can now propagate taint from
page.getOrderBy() to sqlBuilder.Removed the following deprecated semgrep scan options: --json-stats, --json-time, --debugging-json, --save-test-output-tar, --synthesize-patterns, --g…
pattern-propagators feature that allows to specify
arbitrary patterns for the propagation of taint by side-effect. In particular,
this allows to specify how taint propagates through side-effectful function calls.
For example, you can specify that when tainted data is added to an array then the
array itself becomes tainted. (#4509)--config auto no longer sends the name of the repository being scanned to the Semgrep Registry.
As of June 21st, this data is not recorded by the Semgrep Registry backend, even if an old Semgrep version sends it.
Also as of June 21st, none of the previously collected repository names are retained by the Semgrep team;
any historical data has been wiped.semgrep scan options:
--json-stats, --json-time, --debugging-json, --save-test-output-tar, --synthesize-patterns,
--generate-config/-g, --dangerously-allow-arbitrary-code-execution-from-rules,
and --apply (which was an easter egg for job applications, not the same as --autofix)with context expressions where the value is not
bound (#5513)New language R with experimental support (#2360) Thanks to Zythosec for some contributions.
SEMGREP_ENABLE_VERSION_CHECK=0{...foo}) are now translated into the Dataflow ILsemgrep lsp --config auto!semgrep --config auto run on the semgrep Python package in 14s instead of 16s.--disable-version-check would still send a request
when a scan resulted in zero findings./src without notice.
This also could cause permission issues when running the image.$X() no longer matches new Foo(), for consistency with other languages (#5510)($X: C) matches new C(). (#5540)Dataflow: XML elements (e.g. JSX elements) have now a basic translation to the Dataflow IL, meaning that dataflow analysis (constant propagation, tain
package $X, which is useful to bind the package
name and use it in the error message.semgrep ci should be clear it is exiting with error code 0
when there are findings but none of them being blockersyarn.lock files with no depenencies, and with dependencies that lack URLs, now parseGeneric mode: new option generic_ellipsis_max_span for controlling how many lines an ellipsis can match
generic_ellipsis_max_span for controlling
how many lines an ellipsis can match (#5211)generic_comment_style for ignoring
comments that follow the specified syntax (C style, C++ style, or
Shell style) (#3428):include instruction in a .semgrepignore file.
These strings will NOT include user data or specific settings. As an example, with semgrep scan --output=secret.txt we might send "option/output" but will NOT send "option/output=secret.txt".Sarif output format now includes fixes section
fixes sectionr2c-internal-project-depends-on: support for poetry and gradle lockfilesSEMGREP_BASELINE_REF as alias for SEMGREP_BASELINE_COMMITr2c-internal-project-depends-on:
ci CLI command will now include ignored matches in output formats
that dictate they should always be included$X in a message to interpolate the variable captured
by a metavariable named $X, but there was no way to access the underlying value.
However, sometimes that value is more important than the captured variable.
Now you can use the syntax value($X) to interpolate the underlying
propagated value if it exists (if not, it will just use the variable name).
Example:
Take a target file that looks likex = 42
log(x)
Now take a rule to find that log command:- id: example_log
message: Logged $SECRET: value($SECRET)
pattern: log(42)
languages: [python]
Before, this would have given you the message Logged x: value(x). Now, it
will give the message Logged x: 42.return for taint analysis (#4975)metavariable-regex now supports an optional constant-propagation key. When this is set to true, information learned from constant propagation will be
metavariable-regex now supports an optional constant-propagation key.
When this is set to true, information learned from constant propagation
will be used when matching the metavariable against the regex. By default
it is set to falseENVshouldafound - False Negative reporting via the CLItaint(x) makes x tainted by side-effect.
Previously, we had to rely on a trick that declared that any occurrence of
x inside taint(x); ... was as taint source. If x was overwritten with
safe data, this was not recognized by the taint engine. Also, if taint(x)
occurred inside e.g. an if block, any occurrence of x outside that block
was not considered tainted. Now, if you specify that the code variable itself
is a taint source (using focus-metavariable), the taint engine will handle
this as expected, and it will not suffer from the aforementioned limitations.
We believe that this change should not break existing taint rules, but please
report any regressions that you may find.sanitize(x) sanitizes x by side-effect.
Previously, we had to rely on a trick that declared that any occurrence of
x inside sanitize(x); ... was sanitized. If x later overwritten with
tainted data, the taint engine would still regard x as safe. Now, if you
specify that the code variable itself is sanitized (using focus-metavariable),
the taint engine will handle this as expected and it will not suffer from such
limitation. We believe that this change should not break existing taint rules,
but please report any regressions that you may find.semgrep scan --config auto on the semgrep repo itself
went from 50-54 seconds to 28-30 seconds.
:include .gitignore and .git/
from the default .semgrepignore patterns.
This should not cause any difference in which files are targeted
as other parts of Semgrep ignore these files already.override keyword (#4220, #4798)(null)(foo) (#4468)func foo() (..., error, ...) {}) (#4896)with context expressions
(e.g., with (open(x) as a, open(y) as b): pass) (#5092)Files where only some part of the code had to be skipped due to a parse failure will now be listed as "partially scanned" in the end-of-scan skip repo
semgrep ci used to incorrectly report the base branch as a CI job's branch
when running on a pull_request_target event in GitHub Actions.
By fixing this, Semgrep App can now track issue status history with on: pull_request_target jobs.PRIVACY.md had already documented a timestamp field.Datafow: The dataflow engine now handles if-then-else expressions as in OCaml, Ruby, etc. Previously it only handled if-then-else statements.
class Foo(...) {} (#5180)fixed_lines is once again included in JSON output when running with --autofix --dryrunThe JSON output of semgrep scan is now fully specified using ATD (https://atd.readthedocs.io/) and jsonschema (https://json-schema.org/). See the semg
semgrep scan is now fully specified using
ATD (https://atd.readthedocs.io/) and jsonschema (https://json-schema.org/).
See the semgrep-interfaces submodule under interfaces/
(e.g., interfaces/semgrep-interfaces/Semgrep_output_v0.atd for the ATD spec)semgrep scan now contains a "version": field with the
version of Semgrep used to generate the match results.focus-metavariable can be used to
precisely specify that a function parameter is a source of taint, and the taint
engine will handle this as expected.let {x} = E, Semgrep will now infer that x
is tainted if E is tainted.--core-opts flag to send options to semgrep-core. For internal use: no guarantees made for semgrep-core options
--core-opts flag to send options to semgrep-core. For internal use: no guarantees made for semgrep-core options (#5111)Join mode now supports inline rules via the rules: key underneath the join: key.
rules: key underneath the join: key.Bash/Dockerfile: Add support for named ellipses such as in echo $...ARGS
echo $...ARGS (#4887)({ params }: Request) => { } with ({$VAR} : $REQ) => {...}. (#5004)Scala support is now officially GA
-> (P) {Q} where P and Q are sub-patterns. (#4950)semgrep install-deep-semgrep command for DeepSemgrep beta (#4993)lang.json file not found error while building the docker imageEXPOSE 12345 will now parse 12345 as an int instead of a string,
allowing metavariable-comparison with integers (#4875)def f[@an A, @an B](x : A, y : B) = ...)r2c-internal-project-depends-on:
A _reachable_ finding is one with both a dependency match and a pattern match: a vulnerable dependency was found and the vulnerable part of the depend…
focus-metavariable operator that lets you focus (or "zoom in") the match
on the code region delimited by a metavariable. This operator is useful for
narrowing down the code matched by a rule, to focus on what really matters. (#4453)semgrep ci uses "GITHUB_SERVER_URL" to generate urls if it is availableNO_COLOR=1 to force-disable colored outputpattern-sinks, plus the subset of metavariables bound by pattern-sources
that do not collide with the ones bound by pattern-sinks. We do not expect
this change to break many taint rules because source-sink metavariable
unification had a bug (see #4464) that prevented metavariables bound by a
pattern-inside to be unified, thus limiting the usefulness of the feature.
Nonetheless, it is still possible to force metavariable unification by setting
taint_unify_mvars: true in the rule's options.r2c-internal-project-depends-on: this is now a rule key, and not part of the pattern language.
The depends-on-either key can be used analgously to pattern-eitherr2c-internal-project-depends-on: each rule with this key will now distinguish between
reachable and unreachable findings. A reachable finding is one with both a dependency match
and a pattern match: a vulnerable dependency was found and the vulnerable part of the dependency
(according to the patterns in the rule) is used somewhere in code. An unreachable finding
is one with only a dependency match. Reachable findings are reported as coming from the
code that was pattern matched. Unreachable findings are reported as coming from the lockfile
that was dependency matched. Both kinds of findings specify their kind, along with all matched
dependencies, in the extra field of semgrep's JSON output, using the dependency_match_only
and dependency_matches fields, respectively.r2c-internal-project-depends-on: a finding will only be considered reachable if the file
containing the pattern match actually depends on the dependencies in the lockfile containing the
dependency match. A file depends on a lockfile if it is the nearest lockfile going up the
directory tree.semgrep as the entrypoint.
This means that semgrep is no longer prepended automatically to any command you run in the image.
This makes it possible to use the image in CI executors that run provisioning commands within the image.- is now parsed as a valid identifier in Scalanew $OBJECT(...) will now work properly as a taint sink (#4858)...{$X}... will no longer match strpattern-inside are now available to the
rule message. (#4464)SEMGREP_URL or SEMGREP_APP_URL
now updates the URL used both for Semgrep App communication,
and for fetching Semgrep Registry rules.# Changed - pin urllib3 to ~=1.26
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
C#: use latest tree-sitter-c-sharp with support for most C# 10.0 features
CMD ... to match both CMD ls and CMD ["ls"] (#4770).Fixed Deep expression matching and metavariables interaction. Semgrep will
not stop anymore at the first match and will enumarate all possible matchings
if a metavariable is used in a deep expression pattern
(e.g., <... $X ...>). This can introduce some performance regressions.
JSX: ellipsis in JSX body (e.g., <div>...</div>) now matches any
children (#4678 and #4717)
ℹ️ During a
--baseline-commitscan, Semgrep temporarily deletes files that were created since the baseline commit, and restores them at the end of the scan.
Previously, when scanning a subdirectory of a git repo with --baseline-commit,
Semgrep would delete all newly created files under the repo root,
but restore only the ones in the subdirectory.
Now, Semgrep only ever deletes files in the scanned subdirectory.
Previous releases allowed incompatible versions (21.1.0 & 21.2.0)
of the attrs dependency to be installed.
semgrep now correctly requires attrs 21.3.0 at the minimum.
package-lock.json parsing defaults to packages instead of dependencies as the source of dependencies
package-lock.json parsing will ignore dependencies with non-standard versions, and will succesfully parse
dependencies with no integrity field
File targeting logic has been mostly rewritten. (#4776) These inconsistencies were fixed in the process:
ℹ️ "Explicitly targeted file" refers to a file that's directly passed on the command line.
Previously, explicitly targeted files would be unaffected by most global filtering:
global include/exclude patterns and the file size limit.
Now .semgrepignore patterns don't affect them either,
so they are unaffected by all global filtering,
ℹ️ With
--skip-unknown-extensions, Semgrep scans only the explicitly targeted files that are applicable to the language you're scanning.
Previously, --skip-unknown-extensions would skip based only on file extension,
even though extensionless shell scripts expose their language via the shebang of the first line.
As a result, explicitly targeted shell files were always skipped when --skip-unknown-extensions was set.
Now, this flag decides if a file is the correct language with the same logic as other parts of Semgrep:
taking into account both extensions and shebangs.
Semgrep scans with --baseline-commit are now much faster.
These optimizations were added:
ℹ️ When
--baseline-commitis set, Semgrep first runs the current scan, then switches to the baseline commit, and runs the baseline scan.
The current scan now excludes files
that are unchanged between the baseline and the current commit
according to git status output.
The baseline scan now excludes rules and files that had no matches in the current scan.
When git ls-files is unavailable or --disable-git-ignore is set,
Semgrep walks the file system to find all target files.
Semgrep now walks the file system 30% faster compared to previous versions.
The output format has been updated to visually separate lines with headings and indentation.
new --show-supported-languages CLI flag to display the list of languages supported by semgrep. Thanks to John Wu for his contribution!
--validate will check that metavariable-x doesn't use an invalid
metavariable<foo /> used to
match also more complex JSX elements, e.g., <foo >some child</foo>.
This can now be disabled via rule options:
with xml_singleton_loose_matching: false (#4730)xml_attrs_implicit_ellipsis that allows
disabling the implicit ... that was added to JSX attributes patterns.--strict--config auto (#4674)--dry-runs where one fix changes the line numbers in a file that also has a second autofix.yarn.lock dependencies that do not specify a hashproject-depends-on rules with only pattern-inside at their leaves…of regular expression denial-of-service vulnerabilities (ReDoS, redos analyzer, #4700) and high-entropy string detection (entropy analyzer, #4672).
~/.semgrep/last.log-->, for join mode rules for recursively
chaining together Semgrep rules based on metavariable contents.paths.scanned key.--verbose, the skipped paths are also listed under the
paths.skipped key.metavariable-analysis feature
supporting two kinds of analyses: prediction of regular expression
denial-of-service vulnerabilities (ReDoS, redos analyzer, #4700)
and high-entropy string detection (entropy analyzer, #4672).semgrep publish allows users to upload private,
unlisted, or public rules to the Semgrep Registrymetavariable-regex or
metavariable-pattern. Previously, Semgrep had problems analyzing e.g. embedded
YAML content. (#4582)escapeshellarg and
htmlspecialchars_decode, if these functions are given constant arguments,
then Semgrep assumes that their output is also constantSEMGREP_LOGIN_TOKEN to SEMGREP_APP_TOKENExperimental baseline scanning. Run with --baseline-commit GIT_COMMIT to only show findings that currently exist but did not exist in GIT_COMMIT
--baseline-commit GIT_COMMIT to only
show findings that currently exist but did not exist in GIT_COMMIT--verbose mode will list all skipped paths along with the reason they were skippedimport module file names, thus
speeding up matching of patterns like import { $X } from 'foo'E as T will be matched correctly. E.g. previously
a pattern like v as $T would match v but not v as any, now it
correctly matches v as any but not v. (#4515)Dockerfile language: metavariables and ellipses are now supported in most places where it makes sense (#4556, #4577)
Dockerfile: add support for metavariables where argument expansion is already supported
Add an experimental key for internal team use: r2c-internal-project-depends-on that allows rules to filter based on the presence of 3rd-party dependen
r2c-internal-project-depends-on that
allows rules to filter based on the presence of 3rd-party dependencies at specific
version ranges.semgrep login and semgrep logout to store API token from semgrep.devsemgrep --config policy that uses stored API token to
retrieve configured rule policy on semgrep.dev--verbose) appear once per file,
not once per rule/filefor(...) patterns (#4530)Pre-alpha support for Dockerfile as a new target language
x = foo.bar() followed by a call x.baz(), Semgrep will keep
track of x's definition, and it will successfully match x.baz() with a
pattern like foo.bar().baz(). This feature should help writing simple yet
powerful rules, by letting the dataflow engine take care of any intermediate
assignments. Symbolic propagation is still experimental and it is disabled by
default, it must be enabled in a per-rule basis using options: and setting
symbolic_propagation: true. (#2783, #2859, #3207)--verbose outputs a timing and file breakdown summary at the endmetavariable-comparison now handles metavariables that bind to arbitrary
constant expressions (instead of just code variables)New language Solidity with experimental support.
semgrep login and semgrep logout commands to save api tokenval List(x,y,z) = List(1,2,3) to the generic ASTNew language Solidity with experimental support.
Fixed bug where the presence of .semgrepignore would cause runs to fail on files that were not subpaths of the directory where semgrep was being run
Improved filtering of rules based on file content (important speedup for nodejsscan rules notably)
class Foo<...> (#4335)class $X { ...} will now match class Foo<T> { }$SINK->method), and the LHS operand is a tainted
variable (#4320)-filter_irrelevant_rules on rules with very large pattern-eithers (#4305)--debug dumps--output Semgrep will no longer print search results to stdout,
but it will only save/post them to the specified file/URLsemgrep-ci relies on --disable-nosem still tagging findings with is_ignored correctly. Reverting optimization in 0.74.0 that left this field None when
--disable-nosem still tagging findings with is_ignored
correctly. Reverting optimization in 0.74.0 that left this field None when said
flag was usedSupport for method chaining patterns in Python, Golang, Ruby, and C# (#4300), so all GA languages now have method chaining
$X.map($F) matches xs map fprofiling_times object in --time --json output for more fine
grained visibility into slow parts of semgrep"..." (#3881)f(...) and f($X) correctly match f(x) in f(x) { |n| puts n } (#3880)switch had no other statement following it, and the last
statement of the switch's default case was a statement, such as throw,
that can exit the execution of the current function, this caused break
statements within the switch to not be resolved during the construction of
the CFG. This could led to e.g. constant propagation incorrectly flagging
variables as constants. (#4265)Constant propagation: Avoid "Impossible" errors due to unhandled cases
Java: Add partial support for synchronized blocks in the dataflow IL
synchronized blocks in the dataflow IL (#4150)await, yield, &, and other expressionsoptions: with flddef_assign: true (#4187)options: with
arrow_is_function: false (#4187)options: with let_is_var: falsex.f(y), if x is a constant then
it will be recognized as suchcase object within blockscase _x : Int => ...sh as an alias for bashcase class within blocksmetavariable-comparison: if a metavariable binds to a code variable that
is known to be constant, then we use that constant value in the comparison (#3727)~ when resolving config pathsMetavariable equality is enforced across sources/sanitizers/sinks in taint mode, and these metavariables correctly appear in match messages
semgrep --validate runs metachecks on the rule(=> Int) => IntGo: support ... in import list (#4067), for example import (... "error" ...)
import (... "error" ...)o. ... .foo() will now
also match just o.foo().The --enable-metrics flag is now always a flag, does not optionally take an argument
--enable-metrics flag is now always a flag, does not optionally take an argumentC: support ... in parameters and sizeof arguments
@interface pattern (#4030)not_conflicting: true. This affects the change made in 0.68.0
that allowed a sanitizer like - pattern: $F(...) to work, but turned out to
affect our ability to specify sanitization by side-effect. Now the default
semantics of sanitizers is reverted back to the same as before 0.68.0, and
- pattern: $F(...) is supported via the new not-conflicting sanitizers.Respect --skip-unknown-extensions even for files with no extension (treat no extension as an unknown extension)
Added support for raise/throw expressions in the dataflow engine and improved existing support for try-catch-finally
raise/throw expressions in the dataflow engine and improved
existing support for try-catch-finallyInput can be derived from subshells: semgrep --config ... <(...)
semgrep --config ... <(...)- pattern: $F(...) for declaring that any other
function is a sanitizersource(...) and built-in sanitizer
sanitize(...) used for convenience during early development, this was causing
some unexpected behavior in real code that e.g. had a function called source!p/ci) use new rule cdn and do client-side hydrationAdded support for break and continue in the dataflow engine
options: with
attr_expr: false (#3489)<... x ...> now matches sub-expressions of statements-filter_irrelevant_rules causing Semgrep to
incorrectly skip a file (#3755)HCL (a.k.a Terraform) experimental support
This project adheres to Semantic Versioning.
This project adheres to Semantic Versioning.
--vim and --emacs at the same time)pattern: $X optimization ("empty And; no positive terms in And")Enable associative matching for string concatenation
C#: support ellipsis in declarations
... in pattern-insides to simply match anything leftOCaml: support module aliasing, so looking for List.map will also find code that renamed List as L via module L = List.
List.map will also
find code that renamed List as L via module L = List.pattern-regex with completely empty files (#3705)--sarif exit code with suppressed findings (#3680)pattern: $X will not be evaluated on its own, but will look at the context and find $X within the metavariables bound, which should be significantly fasterDeprecated the following experimental features:
if $X = $Y)$X_, $F_OO).
Instead, $FOO will match everything else (lowercase identifiers,
full expressions, types, patterns, etc.).C/C++: Fixed stack overflows (segmentation faults) when processing very large files
1 and 0x1 as equal (#3579)foo.x is now detected as tainted if foo is a source of taintA new experimental 'join' mode. This mode runs multiple Semgrep rules on a codebase and "joins" the results based on metavariable contents. This lets
pattern-not-regex now works (#3503)pattern: $X in the presence of interpolated strings now works (#3560)Significant speed improvements, but the binary is now 95MB (from 47MB in 0.58.1, but it was 170MB in 0.58.0)
The --debug option now displays which files are currently processed incrementally; it will not wait until semgrep-core completely finishes.
New iteration of taint-mode that allows to specify sources/sanitizers/sinks using arbitrary pattern formulas. This provides plenty of flexibility. Not
- source(...) must now be written as - pattern: source(...).implicit_ellipsis that allows disabling the implicit
... that are added to record patterns, plus allow matching "spread fields"
(JS ...x) at any position (#3120)**) syntax in path include/exclude (#3173)try { ... }) for Java (#3417)with statements (#3402)pattern: $X optimization (#3476)pattern or
pattern-regexnew options: field in a YAML rule to enable/disable certain features (e.g., constant propagation). See https://github.com/returntocorp/semgrep/blob/de
options: field in a YAML rule to enable/disable certain features
(e.g., constant propagation). See https://github.com/returntocorp/semgrep/blob/develop/semgrep-core/src/core/Config_semgrep.atd
for the list of available features one can enable/disable.foo(:$ATOM))foo(/.../))__makeref, __reftype, __refvalue (#3364)pattern: $X will not be evaluated on its own, but will look at the context and find $X within the metavariables bound, which should be significantly fasterAssociative-commutative matching for Boolean AND and OR operations
foo("$VAR"))foo(:$ATOM))Add helpUri to sarif output if rule source metadata is defined
Your coding agent can read these notes before it upgrades. Set up the MCP server →