NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #805 most downloaded on PyPI
Always know what to expect from your data.
Last release 9 days ago
25 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
11 versions withdrawn
withdrawn after publishing
9 years old
369 releases · first in 2017
The changelog entries for releases 1.17.1 through 1.23.0 are rewritten in the structured format, and the deprecation timeline gains rows for the gx-re…
Compatibility: sqlalchemy now <2.1 (extras snowflake, databricks)
Fixes GX on SQLAlchemy 2.1 — SQLAlchemy 2.1.0, released 2026-09-24, broke GX 1.23.1 and earlier on Python 3.11+, where every SQL extra resolves it by default. Depending on the backend, import great_expectations failed whenever snowflake-sqlalchemy was installed, every Databricks query failed, driverless postgresql:// URLs could not load a driver, BigQuery queries comparing against a float failed, SQL Server reported mixed-case and upper-case tables as missing, and expect_column_values_to_be_of_type(type_="Numeric") failed on float columns. 1.23.2 fixes all of these: the snowflake and databricks extras stay below SQLAlchemy 2.1 until their dialects support it, and every other SQL extra runs on 2.1. Python 3.10 is unaffected, since SQLAlchemy 2.1 requires Python 3.11. If you can't upgrade yet, pin sqlalchemy<2.1; do the same if you install snowflake-sqlalchemy or databricks-sqlalchemy outside GX's extras. (#12269)
pip install --upgrade 'great_expectations[snowflake]' # include your extras so the SQLAlchemy cap appliesRegex Expectations work on ClickHouse — The four regex Expectations now run on ClickHouse, which does not support regexp_like(). ClickHouse is covered by integration tests for these Expectations. (#12222)
gx.expectations.ExpectColumnValuesToMatchRegex(column="name", regex="^A")Correct substring matching for regex Expectations on Snowflake — Snowflake's REGEXP operator anchors patterns to the whole value, so unanchored patterns behaved differently there than on other backends. Regex Expectations on Snowflake now use substring semantics that match every other GX backend, while preserving the user's pattern. (#12221)
gx.expectations.ExpectColumnValuesToMatchRegex(column="name", regex="ell")Nested struct columns supported in Spark value-counts Expectations — On the Spark engine, Expectations that rely on the column.value_counts metric — including ExpectColumnMostCommonValueToBeInSet and ExpectColumnKLDivergenceToBeLessThan — now work for dotted nested struct paths such as address.city, instead of returning an empty result with an unresolved-column error. (#12231)
gx.expectations.ExpectColumnMostCommonValueToBeInSet(column="address.city", value_set=["Springfield"])snowflake and databricks extras are capped below 2.1, a broken snowflake-sqlalchemy install no longer prevents import great_expectations, driverless postgresql:// URLs fall back to psycopg2 when psycopg is unavailable, BigQuery renders Double as FLOAT64, SQL Server reflects mixed-case tables, the Numeric type name matches float columns again, and database URL masking keeps the database and query string verbatim. (#12269)Decimal value — for example the maximum, mean or sum of a PostgreSQL numeric column or a pandas column of Decimal values — now serialize as float infinity instead of raising decimal.InvalidOperation. (#12254)UnexpectedRowsExpectation no longer misreads a query as containing a JOIN when the letters appear inside a string literal, a column name such as join_date, a quoted identifier or a comment; such queries are aliased correctly again and no longer fail with a syntax error on MySQL and SQL Server. (#12249)regexp_like(); other SQL dialects are unchanged. (#12222)Nullable(T) wrapper, so ExpectColumnValuesToBeOfType and ExpectColumnValuesToBeInTypeList behave as expected. (#12219)column.value_counts metric now resolves nested struct columns such as address.city, so Expectations built on it return results instead of an unresolved-column error. (#12231)gx-redshift extra alias and for CloudDataContext / cloud mode of get_context. (#12225)Thanks to @adimalkar, @Rayan-and-beyond (first contribution), @nanjeshramesh, @feiiiiii5, @alibro005, @pentaoa (first contribution).
One column per quarter.
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation .
Source distribution for great-expectations 1.23.2 File Size Uploaded
great_expectations-1.23.2.tar.gz 37.0 MB Sep 25, 2026 Details
Table of built distributions (wheels) for great-expectations 1.23.2 File Interpreter ABI Platform Reset
great_expectations-1.23.2-py3-none-any.whl 5.7 MB Sep 25, 2026 Python 3 none any Details
Total release size: 42.7 MB
The changelog now carries a deprecation timeline table listing every deprecated item, the version that deprecated it, and the version that removes it.
Spark now evaluates each regex independently with match_on="all" — On Spark, ExpectColumnValuesToMatchRegexList with match_on="all" now checks every regex separately against each column value, so patterns anchored at different positions (such as ^A and [0-9]{3}$) both match a value that satisfies them. This matches the behavior already seen on Pandas and SQL. (#12198)
gxe.ExpectColumnValuesToMatchRegexList(
column="id",
regex_list=["^A", "[0-9]{3}$"],
match_on="all",
)Each Validator reports results for its own Batch when a datasource is reused — Validators built on the same datasource no longer borrow one another's Batch. Running two validation definitions on threads, or creating two validators from one datasource on a single thread, now evaluates and reports each validator's own data, with the correct batch_id, batch_spec and batch_definition on the result. This fixes a long-standing latent bug made reachable by #12148 in the 1.23.0 release. (#12211)
validator_a = context.get_validator(batch_request=request_a)
validator_b = context.get_validator(batch_request=request_b)
# validator_a still validates request_a's batch
result = validator_a.expect_table_row_count_to_equal(value=3)A Spark schema saved in great_expectations.yml reloads correctly — A persisted spark_schema is now read back through StructType.fromJson, so reopening a File Data Context round-trips the schema instead of failing inside PySpark. Values that are not an accepted schema form now raise a validation error naming the field and the accepted types. (#12200)
context = gx.get_context(mode="file")
asset = context.data_sources.get("spark_ds").get_asset("my_asset")
assert asset.spark_schema is not NoneGX config files are read and written as UTF-8 regardless of locale — config_variables.yml and great_expectations.yml, and the .gitignore read while scaffolding a project, are now opened with an explicit UTF-8 encoding. Projects containing non-ASCII values or comments can be created and reloaded on hosts with a non-UTF-8 locale, and a project YAML file that is not valid UTF-8 now raises an error naming the file instead of a bare decode error. (#12182, #12204)
A gallery-wide test tier for data sources — A new gallery support tier asserts a measured test result across the entire shipped expectation library — one case per registered expectation, each pairing a passing and a failing configuration — and nine data sources (pandas in-memory and filesystem CSV, SQLite, MySQL, PostgreSQL, Trino, BigQuery, Databricks and Redshift) now declare it after being measured against the full gallery. (#12150)
Thanks to @lakshayxi (first contribution), @feiiiiii5 (first contribution), @ptimizeroracle (first contribution), @alibro005 (first contribution), @yigitcan-ozturk, @toyeshhm (first contribution), @nanjeshramesh, @p-mandale (first contribution).
Known issue: on Python 3.11+, this release resolves SQLAlchemy 2.1 (released 2026-09-24), which it does not support. Upgrade to 1.23.2, or pin sqlalchemy<2.1.
Spark now evaluates each regex independently with match_on="all" — On Spark, ExpectColumnValuesToMatchRegexList with match_on="all" now checks every regex separately against each column value, so patterns anchored at different positions (such as ^A and [0-9]{3}$) both match a value that satisfies them. This matches the behavior already seen on Pandas and SQL. (#12198)
gxe.ExpectColumnValuesToMatchRegexList(
column="id",
regex_list=["^A", "[0-9]{3}$"],
match_on="all",
)
Each Validator reports results for its own Batch when a datasource is reused — Validators built on the same datasource no longer borrow one another's Batch. Running two validation definitions on threads, or creating two validators from one datasource on a single thread, now evaluates and reports each validator's own data, with the correct batch_id, batch_spec and batch_definition on the result. This fixes a long-standing latent bug made reachable by #12148 in the 1.23.0 release. (#12211)
validator_a = context.get_validator(batch_request=request_a)
validator_b = context.get_validator(batch_request=request_b)
# validator_a still validates request_a's batch
result = validator_a.expect_table_row_count_to_equal(value=3)
A Spark schema saved in great_expectations.yml reloads correctly — A persisted spark_schema is now read back through StructType.fromJson, so reopening a File Data Context round-trips the schema instead of failing inside PySpark. Values that are not an accepted schema form now raise a validation error naming the field and the accepted types. (#12200)
context = gx.get_context(mode="file")
asset = context.data_sources.get("spark_ds").get_asset("my_asset")
assert asset.spark_schema is not None
GX config files are read and written as UTF-8 regardless of locale — config_variables.yml and great_expectations.yml, and the .gitignore read while scaffolding a project, are now opened with an explicit UTF-8 encoding. Projects containing non-ASCII values or comments can be created and reloaded on hosts with a non-UTF-8 locale, and a project YAML file that is not valid UTF-8 now raises an error naming the file instead of a bare decode error. (#12182, #12204)
A gallery-wide test tier for data sources — A new gallery support tier asserts a measured test result across the entire shipped expectation library — one case per registered expectation, each pairing a passing and a failing configuration — and nine data sources (pandas in-memory and filesystem CSV, SQLite, MySQL, PostgreSQL, Trino, BigQuery, Databricks and Redshift) now declare it after being measured against the full gallery. (#12150)
<details> <summary>Maintenance</summary>
</details>
Thanks to @lakshayxi (first contribution), @feiiiiii5 (first contribution), @ptimizeroracle (first contribution), @alibro005 (first contribution), @yigitcan-ozturk, @toyeshhm (first contribution), @nanjeshramesh, @p-mandale (first contribution).
[BUGFIX] Cache the SQL execution engine across calls instead of rebuilding it every time
Always know what to expect from your data.
pip install great-expectations Copy PIP instructions
Always know what to expect from your data. (See https://github.com/great-expectations/great_expectations for full description).
Homepage
Download
PyPI data Data sourced directly from PyPI's database.
PyPI data Data sourced directly from PyPI's database.
danklynn josh.stauffer
Author: The Great Expectations Team
Apache Software License (Apache-2.0)
Python <3.14, >=3.10
athena teradata singlestore mysql vertica dremio spark oracle redshift cloud gcs snowflake postgresql sql-server bigquery arrow hive spark-connect databricks azure excel gx-redshift pagerduty trino clickhouse test aws-secrets azure-secrets fabric gcp s3
data science testing pipeline data quality dataquality validation datavalidation
Development Status
5 - Production/Stable
Intended Audience
Developers
Other Audience
Science/Research
License
OSI Approved :: Apache Software License
Programming Language
Python :: 3
Python :: 3.10
Python :: 3.11
Python :: 3.12
Python :: 3.13
Topic
Scientific/Engineering
Software Development
Software Development :: Testing
Report project as malware
Download the file for your platform. If you're not sure which to choose, learn more about installing packages .
great_expectations-1.23.0.tar.gz (36.9 MB view details ) Uploaded Sep 10, 2026 Source
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names .
Copy a direct link to the current filters Copy
File name
Interpreter Interpreter py3
ABI ABI none
Platform Platform any
great_expectations-1.23.0-py3-none-any.whl (5.7 MB view details ) Uploaded Sep 10, 2026 Python 3
Details for the file great_expectations-1.23.0.tar.gz .
Download URL: great_expectations-1.23.0.tar.gz
Upload date: Sep 10, 2026
Size: 36.9 MB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/7.0.0 CPython/3.13.14
Hashes for great_expectations-1.23.0.tar.gz Algorithm Hash digest
SHA256 009335ebf49394cfbba646f1f0fe7b46ed6cf74a335723c6a529a1222657886e Copy
MD5 b5c947e2bbcf6a326dc73c88d1ad4b70 Copy
BLAKE2b-256 53195c6d136adc9c3895801a6a5e03db415dc02137fb7c0b88fd179bfdc801f2 Copy
See more details on using hashes here.
The following attestation bundles were made for great_expectations-1.23.0.tar.gz :
Publisher: ci.yml on fivetran/great_expectations Attestations: Values shown here reflect the state when the release was signed and may no longer be current.
Statement:
Statement type: https://in-toto.io/Statement/v1
Predicate type: https://docs.pypi.org/attestations/publish/v1
Subject name: great_expectations-1.23.0.tar.gz
Subject digest: 009335ebf49394cfbba646f1f0fe7b46ed6cf74a335723c6a529a1222657886e
Sigstore transparency entry: 2784397633
Sigstore integration time: Sep 10, 2026 Source repository:
Permalink: fivetran/great_expectations@8210c2e280ebcd7f097e9d8850d81255de3026f6
Branch / Tag: refs/tags/1.23.0
Owner: https://github.com/fivetran
Access: public Publication detail:
Token Issuer: https://token.actions.githubusercontent.com
Runner Environment: github-hosted
Publication workflow: ci.yml@8210c2e280ebcd7f097e9d8850d81255de3026f6
Trigger Event: push
Details for the file great_expectations-1.23.0-py3-none-any.whl .
Download URL: great_expectations-1.23.0-py3-none-any.whl
Upload date: Sep 10, 2026
Size: 5.7 MB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/7.0.0 CPython/3.13.14
Hashes for great_expectations-1.23.0-py3-none-any.whl Algorithm Hash digest
SHA256 ed42eb95ac5c4e7a5b1edb7cb811d635acc0b7b456ef4c8543731f57641ccddb Copy
MD5 f9e520c396f62c60bf0232c744bbd921 Copy
BLAKE2b-256 f9447d2a85d84191764130e078d4fc90024a9a554136095a2cb97304dada8fe3 Copy
See more details on using hashes here.
The following attestation bundles were made for great_expectations-1.23.0-py3-none-any.whl :
Publisher: ci.yml on fivetran/great_expectations Attestations: Values shown here reflect the state when the release was signed and may no longer be current.
Statement:
Statement type: https://in-toto.io/Statement/v1
Predicate type: https://docs.pypi.org/attestations/publish/v1
Subject name: great_expectations-1.23.0-py3-none-any.whl
Subject digest: ed42eb95ac5c4e7a5b1edb7cb811d635acc0b7b456ef4c8543731f57641ccddb
Sigstore transparency entry: 2784397684
Sigstore integration time: Sep 10, 2026 Source repository:
Permalink: fivetran/great_expectations@8210c2e280ebcd7f097e9d8850d81255de3026f6
Branch / Tag: refs/tags/1.23.0
Owner: https://github.com/fivetran
Access: public Publication detail:
Token Issuer: https://token.actions.githubusercontent.com
Runner Environment: github-hosted
Publication workflow: ci.yml@8210c2e280ebcd7f097e9d8850d81255de3026f6
Trigger Event: push
This release
1.23.0 This release
Sep 10, 2026 2 files
1.22.0
Aug 31, 2026 2 files
1.21.0
Aug 19, 2026 2 files
1.20.0
Aug 7, 2026 2 files
1.19.1
Jul 24, 2026 2 files
1.19.0
Jul 13, 2026 2 files
1.18.2
Jun 26, 2026 2 files
1.18.1
Jun 11, 2026 2 files
1.18.0
Jun 2, 2026 2 files
1.17.2
May 14, 2026 2 files
1.17.1
May 5, 2026 2 files
1.17.0
Apr 22, 2026 2 files
1.16.1
Apr 15, 2026 2 files
1.16.0
Apr 9, 2026 2 files
1.15.2
Apr 1, 2026 2 files
1.15.1
Mar 13, 2026 2 files
1.15.0
Mar 11, 2026 2 files
1.14.0
Mar 4, 2026 2 files
1.13.1
Mar 4, 2026 2 files
1.13.0
Feb 26, 2026 2 files
1.12.3
Feb 13, 2026 2 files
1.11.3
Jan 29, 2026 2 files
1.11.2
Jan 23, 2026 2 files
1.11.1
Jan 20, 2026 2 files
1.11.0
Jan 12, 2026 2 files
1.10.0
Dec 18, 2025 2 files
1.9.3
Dec 11, 2025 2 files
1.9.2
Dec 3, 2025 2 files
1.9.1
Nov 20, 2025 2 files
1.9.0
Nov 7, 2025 2 files
1.8.1
Oct 30, 2025 2 files
1.8.0
Oct 23, 2025 2 files
1.7.1
Oct 15, 2025 2 files
1.7.0
Oct 9, 2025 2 files
1.6.4
Oct 1, 2025 2 files
1.6.3
Sep 24, 2025 2 files
1.6.2
Sep 20, 2025 2 files
1.6.1
Sep 16, 2025 2 files
1.6.0
Sep 12, 2025 2 files
1.5.11
Sep 4, 2025 2 files
1.5.10
Aug 27, 2025 2 files
1.5.9
Aug 20, 2025 2 files
1.5.8
Aug 7, 2025 2 files
1.5.7
Jul 31, 2025 2 files
1.5.6
Jul 24, 2025 2 files
1.5.5
Jul 11, 2025 2 files
1.5.4
Jul 2, 2025 2 files
1.5.3
Jun 25, 2025 2 files
1.5.2
Jun 18, 2025 2 files
1.5.1
Jun 11, 2025 2 files
1.5.0
Jun 5, 2025 2 files
Yanked
1.4.7
Jun 5, 2025 2 files
Yanked reason: The version accidently was set to 1.4.7 instead of 1.5.0
1.4.6
May 28, 2025 2 files
1.4.5
May 22, 2025 2 files
1.4.4
May 14, 2025 2 files
1.4.3
May 7, 2025 2 files
1.4.2
Apr 24, 2025 2 files
1.4.1
Apr 22, 2025 2 files
1.4.0
Apr 15, 2025 2 files
1.3.14
Apr 9, 2025 2 files
1.3.13
Apr 3, 2025 2 files
1.3.12
Mar 26, 2025 2 files
1.3.11
Mar 19, 2025 2 files
1.3.10
Mar 12, 2025 2 files
1.3.9
Mar 5, 2025 2 files
1.3.8
Feb 26, 2025 2 files
1.3.7
Feb 19, 2025 2 files
1.3.6
Feb 14, 2025 2 files
1.3.5
Feb 3, 2025 2 files
1.3.4
Jan 29, 2025 2 files
1.3.3
Jan 22, 2025 2 files
1.3.2
Jan 17, 2025 2 files
1.3.1
Jan 8, 2025 2 files
1.3.0
Dec 19, 2024 2 files
1.2.6
Dec 11, 2024 2 files
1.2.5
Dec 4, 2024 2 files
1.2.4
Nov 20, 2024 2 files
1.2.3
Nov 14, 2024 2 files
1.2.2
Nov 7, 2024 2 files
1.2.1
Oct 31, 2024 2 files
1.2.0
Oct 24, 2024 2 files
1.1.3
Oct 15, 2024 2 files
1.1.2
Oct 10, 2024 2 files
1.1.1
Oct 8, 2024 2 files
1.1.0
Oct 3, 2024 2 files
1.0.6
Oct 1, 2024 2 files
1.0.5
Sep 19, 2024 2 files
1.0.4
Sep 16, 2024 2 files
Yanked
1.0.3
Sep 12, 2024 2 files
1.0.2
Sep 5, 2024 2 files
1.0.1
Aug 29, 2024 2 files
1.0.0
Aug 22, 2024 2 files
Pre-release
1.0.0a6
Aug 21, 2024 2 files
Pre-release
1.0.0a5
Aug 6, 2024 2 files
Pre-release
1.0.0a4
May 16, 2024 2 files
Pre-release
1.0.0a3
Apr 29, 2024 2 files
Pre-release
1.0.0a2
Apr 15, 2024 2 files
Pre-release
1.0.0a1
Feb 15, 2024 2 files
0.18.22
Oct 25, 2024 2 files
0.18.21
Sep 18, 2024 2 files
0.18.20
Sep 10, 2024 2 files
0.18.19
Jul 15, 2024 2 files
0.18.18
Jul 3, 2024 2 files
0.18.17
Jun 28, 2024 2 files
0.18.16
Jun 18, 2024 2 files
0.18.15
May 28, 2024 2 files
0.18.14
May 22, 2024 2 files
0.18.13
Apr 29, 2024 2 files
0.18.12
Mar 20, 2024 2 files
0.18.11
Mar 14, 2024 2 files
0.18.10
Feb 26, 2024 2 files
0.18.9
Feb 16, 2024 2 files
0.18.8
Jan 11, 2024 2 files
0.18.7
Dec 22, 2023 2 files
Yanked
0.18.6
Dec 20, 2023 2 files
0.18.5
Dec 14, 2023 2 files
0.18.4
Dec 8, 2023 2 files
0.18.3
Nov 16, 2023 2 files
0.18.2
Nov 9, 2023 2 files
0.18.1
Nov 2, 2023 2 files
0.18.0
Oct 30, 2023 2 files
0.17.23
Oct 20, 2023 2 files
0.17.22
Oct 12, 2023 2 files
0.17.21
Oct 6, 2023 2 files
Yanked
0.17.20
Sep 28, 2023 2 files
Yanked reason: Breaks GX Agent
0.17.19
Sep 21, 2023 2 files
0.17.18
Sep 20, 2023 2 files
0.17.17
Sep 18, 2023 2 files
0.17.16
Sep 15, 2023 2 files
0.17.15
Sep 7, 2023 2 files
0.17.14
Sep 1, 2023 2 files
Yanked
0.17.13
Aug 31, 2023 2 files
0.17.12
Aug 24, 2023 2 files
0.17.11
Aug 17, 2023 2 files
0.17.9
Aug 10, 2023 2 files
0.17.8
Aug 4, 2023 2 files
0.17.7
Jul 27, 2023 2 files
0.17.6
Jul 21, 2023 2 files
0.17.5
Jul 13, 2023 2 files
0.17.4
Jul 10, 2023 2 files
0.17.3
Jul 7, 2023 2 files
0.17.2
Jun 29, 2023 2 files
0.17.1
Jun 22, 2023 2 files
0.17.0
Jun 15, 2023 2 files
0.16.16
Jun 8, 2023 2 files
0.16.15
Jun 1, 2023 2 files
0.16.14
May 26, 2023 2 files
0.16.13
May 18, 2023 2 files
0.16.12
May 11, 2023 2 files
0.16.11
May 4, 2023 2 files
0.16.10
Apr 28, 2023 2 files
Yanked
0.16.9
Apr 28, 2023 2 files
0.16.8
Apr 20, 2023 2 files
0.16.7
Apr 13, 2023 2 files
0.16.6
Apr 6, 2023 2 files
0.16.5
Apr 2, 2023 2 files
Yanked
0.16.4
Mar 31, 2023 2 files
0.16.3
Mar 24, 2023 2 files
Yanked
0.16.2
Mar 23, 2023 2 files
0.16.1
Mar 16, 2023 2 files
0.16.0
Mar 10, 2023 2 files
0.15.50
Feb 23, 2023 2 files
0.15.49
Feb 17, 2023 2 files
0.15.48
Feb 9, 2023 2 files
0.15.47
Feb 2, 2023 2 files
0.15.46
Jan 26, 2023 2 files
0.15.45
Jan 26, 2023 2 files
0.15.44
Jan 19, 2023 2 files
0.15.43
Jan 12, 2023 2 files
0.15.42
Jan 5, 2023 2 files
0.15.41
Dec 15, 2022 2 files
0.15.40
Dec 13, 2022 2 files
Yanked
0.15.39
Dec 10, 2022 2 files
Yanked
0.15.38
Dec 9, 2022 2 files
Yanked
0.15.37
Dec 8, 2022 2 files
0.15.36
Dec 1, 2022 2 files
0.15.35
Dec 1, 2022 2 files
0.15.34
Nov 18, 2022 2 files
0.15.33
Nov 17, 2022 2 files
0.15.32
Nov 10, 2022 2 files
0.15.31
Nov 4, 2022 2 files
0.15.30
Nov 3, 2022 2 files
0.15.29
Oct 28, 2022 2 files
0.15.28
Oct 20, 2022 2 files
0.15.27
Oct 13, 2022 2 files
0.15.26
Sep 29, 2022 2 files
0.15.25
Sep 23, 2022 2 files
0.15.24
Sep 19, 2022 2 files
0.15.23
Sep 16, 2022 2 files
0.15.22
Sep 8, 2022 2 files
0.15.21
Sep 1, 2022 2 files
0.15.20
Aug 25, 2022 2 files
0.15.19
Aug 18, 2022 2 files
0.15.18
Aug 11, 2022 2 files
0.15.17
Aug 4, 2022 2 files
0.15.16
Jul 29, 2022 2 files
0.15.15
Jul 21, 2022 2 files
0.15.14
Jul 14, 2022 2 files
0.15.13
Jul 7, 2022 2 files
0.15.12
Jun 30, 2022 2 files
0.15.11
Jun 22, 2022 2 files
0.15.10
Jun 15, 2022 2 files
0.15.9
Jun 9, 2022 2 files
0.15.8
Jun 2, 2022 2 files
0.15.7
May 26, 2022 2 files
0.15.6
May 19, 2022 2 files
0.15.5
May 12, 2022 2 files
0.15.4
May 5, 2022 2 files
0.15.3
Apr 28, 2022 2 files
0.15.2
Apr 21, 2022 2 files
0.15.1
Apr 14, 2022 2 files
0.15.0
Apr 8, 2022 2 files
0.14.13
Mar 31, 2022 2 files
0.14.12
Mar 24, 2022 2 files
0.14.11
Mar 17, 2022 2 files
0.14.10
Mar 10, 2022 2 files
0.14.9
Mar 4, 2022 2 files
0.14.8
Feb 24, 2022 2 files
0.14.7
Feb 17, 2022 2 files
0.14.6
Feb 10, 2022 2 files
0.14.5
Feb 3, 2022 2 files
0.14.4
Jan 28, 2022 2 files
0.14.3
Jan 27, 2022 2 files
0.14.2
Jan 20, 2022 2 files
0.14.1
Jan 13, 2022 2 files
0.14.0
Jan 6, 2022 2 files
0.13.49
Dec 24, 2021 2 files
0.13.48
Dec 23, 2021 2 files
0.13.47
Dec 18, 2021 2 files
0.13.46
Dec 9, 2021 2 files
0.13.45
Dec 2, 2021 2 files
0.13.44
Nov 24, 2021 2 files
0.13.43
Nov 18, 2021 2 files
0.13.42
Nov 12, 2021 2 files
0.13.41
Nov 4, 2021 2 files
0.13.40
Oct 27, 2021 2 files
0.13.39
Oct 21, 2021 2 files
0.13.38
Oct 14, 2021 2 files
0.13.37
Oct 7, 2021 2 files
0.13.36
Sep 30, 2021 2 files
0.13.35
Sep 23, 2021 2 files
0.13.34
Sep 16, 2021 2 files
0.13.33
Sep 9, 2021 2 files
0.13.32
Sep 2, 2021 2 files
0.13.31
Aug 26, 2021 2 files
0.13.30
Aug 23, 2021 2 files
0.13.29
Aug 19, 2021 2 files
0.13.28
Aug 13, 2021 2 files
0.13.27
Aug 12, 2021 2 files
0.13.26
Aug 5, 2021 2 files
0.13.25
Jul 30, 2021 2 files
0.13.24
Jul 22, 2021 2 files
0.13.23
Jul 15, 2021 2 files
0.13.22
Jul 9, 2021 2 files
0.13.21
Jun 30, 2021 2 files
0.13.20
Jun 23, 2021 2 files
0.13.19
Apr 23, 2021 2 files
0.13.18
Apr 22, 2021 2 files
0.13.17
Apr 2, 2021 2 files
0.13.16
Apr 1, 2021 2 files
0.13.15
Mar 26, 2021 2 files
0.13.14
Mar 17, 2021 2 files
0.13.13
Mar 12, 2021 2 files
0.13.12
Mar 5, 2021 2 files
0.13.11
Feb 25, 2021 2 files
0.13.10
Feb 13, 2021 2 files
0.13.9
Feb 8, 2021 2 files
0.13.8
Jan 28, 2021 2 files
0.13.7
Jan 23, 2021 2 files
0.13.6
Jan 21, 2021 2 files
0.13.5
Jan 19, 2021 2 files
0.13.4
Dec 23, 2020 2 files
0.13.3
Dec 15, 2020 2 files
0.13.2
Dec 8, 2020 2 files
0.13.1
Dec 3, 2020 2 files
0.13.0
Dec 1, 2020 2 files
0.12.10
Nov 25, 2020 1 file
0.12.9
Nov 17, 2020 2 files
0.12.8
Nov 16, 2020 2 files
0.12.7
Oct 29, 2020 2 files
0.12.6
Oct 20, 2020 2 files
0.12.5
Oct 19, 2020 2 files
0.12.4
Oct 7, 2020 2 files
0.12.3
Sep 28, 2020 2 files
0.12.2
Sep 22, 2020 2 files
0.12.1
Sep 2, 2020 2 files
0.12.0
Aug 13, 2020 2 files
0.11.9
Jul 30, 2020 2 files
0.11.8
Jul 16, 2020 2 files
0.11.7
Jul 2, 2020 2 files
0.11.6
Jun 27, 2020 2 files
0.11.5
Jun 19, 2020 2 files
0.11.4
Jun 13, 2020 2 files
0.11.3
Jun 13, 2020 2 files
0.11.2
Jun 5, 2020 2 files
0.11.1
May 29, 2020 2 files
0.11.0
May 23, 2020 2 files
Pre-release
0.11.0b0
May 10, 2020 2 files
0.10.12
May 20, 2020 2 files
0.10.11
May 15, 2020 2 files
0.10.10
May 14, 2020 2 files
0.10.9
May 8, 2020 2 files
0.10.8
May 4, 2020 2 files
0.10.7
May 1, 2020 2 files
0.10.6
May 1, 2020 2 files
0.10.5
Apr 29, 2020 2 files
0.10.4
Apr 24, 2020 2 files
0.10.3
Apr 22, 2020 2 files
0.10.2
Apr 21, 2020 2 files
0.10.1
Apr 16, 2020 2 files
0.10.0
Apr 15, 2020 2 files
0.9.11
Apr 10, 2020 2 files
0.9.10
Apr 8, 2020 2 files
0.9.9
Apr 7, 2020 2 files
0.9.8
Apr 3, 2020 2 files
0.9.7
Mar 19, 2020 2 files
0.9.6
Mar 18, 2020 2 files
0.9.5
Mar 13, 2020 2 files
0.9.4
Mar 11, 2020 2 files
0.9.3
Mar 6, 2020 2 files
0.9.2
Feb 22, 2020 2 files
0.9.1
Feb 21, 2020 2 files
0.9.0
Feb 19, 2020 2 files
Pre-release
0.9.0b2
Feb 11, 2020 2 files
Pre-release
0.9.0b1
Dec 12, 2019 2 files
Pre-release
0.9.0b0
Nov 26, 2019 2 files
0.8.8
Feb 7, 2020 2 files
0.8.7
Jan 15, 2020 2 files
0.8.6
Dec 4, 2019 2 files
0.8.5
Nov 19, 2019 2 files
0.8.4.post0
Nov 7, 2019 2 files
0.8.4
Nov 6, 2019 2 files
0.8.3
Oct 29, 2019 2 files
0.8.2.post0
Oct 24, 2019 2 files
0.8.2
Oct 23, 2019 2 files
0.8.1
Oct 16, 2019 2 files
0.8.0
Oct 16, 2019 2 files
Pre-release
0.8.0a4
Oct 11, 2019 2 files
Pre-release
0.8.0a3
Oct 7, 2019 2 files
Pre-release
0.8.0a2
Oct 3, 2019 2 files
Pre-release
0.8.0a1
Sep 30, 2019 2 files
0.7.11
Oct 4, 2019 2 files
0.7.10
Sep 19, 2019 2 files
0.7.9
Sep 18, 2019 2 files
0.7.8
Sep 4, 2019 2 files
0.7.7
Aug 19, 2019 2 files
0.7.6
Aug 12, 2019 2 files
0.7.5
Aug 3, 2019 2 files
0.7.4
Aug 3, 2019 2 files
0.7.3
Jul 29, 2019 2 files
0.7.2
Jul 22, 2019 2 files
0.7.1
Jul 13, 2019 2 files
0.7.0
Jul 4, 2019 2 files
0.6.1
Jun 3, 2019 1 file
0.6.0
May 24, 2019 2 files
0.5.1
Apr 30, 2019 2 files
0.5.0
Apr 25, 2019 2 files
0.4.5
Dec 19, 2018 2 files
0.4.4
Aug 29, 2018 2 files
0.4.3
Jul 12, 2018 2 files
0.4.2
May 17, 2018 2 files
0.4.1
Mar 24, 2018 2 files
0.4.0
Mar 23, 2018 2 files
0.3.2
Feb 8, 2018 2 files
0.3.1
Feb 8, 2018 2 files
0.3.0
Dec 22, 2017 2 files
Pre-release
0.0.1111.post0.dev17
Aug 4, 2023 2 files
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page
"PyPI", "Python Package Index", and the blocks logos are registered trademarks of the Python Software Foundation .
© 2026 Python Software Foundation
Site map
Deployed from 490a846
Compatibility: new extra oracle
great_expectations[oracle] is a supported install — Oracle is now a published install path: installing the oracle extra brings in the Oracle driver and floors SQLAlchemy at 2.0, so the oracle+oracledb dialect the connection string needs is always available. The SQL dialect installation-commands table documents the new row. (#12091)
pip install 'great_expectations[oracle]'
Daily and monthly Batch Definitions work on Oracle query assets — Adding a daily or monthly Batch Definition to a query asset on Oracle previously failed with ORA-00907: missing right parenthesis, reported misleadingly as the partition column not being verifiable as a date or datetime. A query asset's SQL is now wrapped whole, so the Batch Definition can be created. As a side effect, a query whose SQL ends in a trailing line comment no longer breaks Batch Definition creation on any backend. (#12162)
asset = datasource.add_query_asset(name="orders", query="SELECT id, created_at FROM my_table")
asset.add_batch_definition_daily(name="daily", column="created_at")
Consistent verdict for z-score checks on zero or undefined variance — ExpectColumnValueZScoresToBeLessThan used to disagree by backend on a constant column: pandas flagged every row as an outlier, PostgreSQL and SQL Server surfaced a division-by-zero error, and SQLite and MySQL quietly succeeded. All engines now agree that a column with zero or undefined standard deviation succeeds with no unexpected values. Anyone who relied on this Expectation to catch a stuck or constant column should use ExpectColumnStdevToBeBetween with a non-zero min_value instead. (#12145)
gxe.ExpectColumnValueZScoresToBeLessThan(column="constant", threshold=1.96, double_sided=True)
# success=True, unexpected_count=0 on every backend
Mixed-case column names no longer break uniqueness checks on SQL backends — ExpectColumnValuesToBeUnique raised KeyError: '<column>' on case-insensitive SQL dialects (Databricks, PostgreSQL, Snowflake, SQL Server, Trino) whenever the column name was not already lower case and the result format asked for rows or unexpected indices. It now evaluates normally, so users who pinned to 1.19.1 for this reason can unpin. (#12180)
gxe.ExpectColumnValuesToBeUnique(column="CustomerID") # result_format="COMPLETE"
SQL execution engines are reused instead of rebuilt on every validation — Every validation against a SQL datasource used to build a new execution engine, with its own SQLAlchemy engine and connection pool, leaking an idle pooled connection per validation and re-running dialect setup each time. The engine is now cached as intended and rebuilt only when the datasource's connection configuration changes; a validation that follows a schema change still reflects the table afresh. (#12148)
datasource.get_execution_engine() is datasource.get_execution_engine() # now True
ExpectColumnStdevToBeBetween on SQLite now reports an undefined standard deviation as an observed_value of None — matching every other backend — instead of returning a result with no observed_value and an opaque "user-defined function raised exception" error, for columns with fewer than two non-null values and for empty tables. (#12168)ExpectColumnValuesToBeUnique no longer raises KeyError on SQL backends when a column name is not lower case and the result format requests rows or unexpected indices. (#12180)ExpectColumnValueZScoresToBeLessThan now succeeds with no unexpected values on columns whose standard deviation is zero or undefined, on every backend, instead of failing on pandas or raising a division-by-zero error on PostgreSQL and SQL Server. (#12145)UnicodeDecodeError under a non-UTF-8 system locale: filesystem store reads and project-configuration reads and writes are now pinned to UTF-8. (#12125)ORA-00907: missing right parenthesis; a query asset whose SQL ends in a line comment also works on every backend now. (#12162)-- comment no longer fails validation: the raw SQL is normalized before it is wrapped, so appended text cannot land inside a trailing comment. (#12124)fivetran/great_expectations repository. (#12126)<details> <summary>Maintenance</summary>
pip install 'great_expectations[oracle]' is now a documented, supported install path, with SQLAlchemy floored at 2.0 so the Oracle dialect is available, and an install row added to the SQL dialect installation-commands table. (#12091)conftest.py. (#12174)tests/integration/ into the type check, correcting their annotations and removing the covering exclude patterns. (#12173)tests/data_context/ into the type check and removed their exclude patterns. (#12171)match_on value key from the not-match-like-pattern-list metric, so the metric declares only the options it actually reads; ExpectColumnValuesToNotMatchLikePatternList never accepted match_on and still rejects it. (#12157)tests/core/ into the type check and removed their exclude patterns. (#12140)tests/render/ into the type check with annotation-only changes, leaving rendering behavior unchanged. (#12152)tests/test_utils.py, tests/actions/ and tests/checkpoint/test_checkpoint.py into the type check, including fixes so a failed database connection surfaces its original error instead of an AttributeError from the cleanup path. (#12146)</details>
Thanks to @siddharthgaur1 (first contribution), @Star-cloud626 (first contribution), @nanjeshramesh, @adimalkar (first contribution), @AnandkumarMall (first contribution), @Ryota-Di (first contribution), @yigitcan-ozturk (first contribution), @MannXo, @iamfeldman (first contribution).
[BUGFIX] Give the date-part string cast a length Oracle accepts
Compatibility: marshmallow minimum 3.7.1 → 3.18.0
Marshmallow 4 is now supported — Great Expectations now installs and runs against both Marshmallow 3 and Marshmallow 4, so it can be installed alongside deployments that pin Marshmallow 4 (such as Apache Airflow 3.3). The supported range is now marshmallow>=3.18.0 with no upper bound; the declared 3.7.1 floor was unreachable in practice, so no currently-working environment is excluded. (#12092, #12118)
pip install great_expectations marshmallow==4.3.1
Oracle is now a live-tested backend, with three Oracle defects fixed — Oracle joins the SQL test harness with live curated coverage, and the gaps that coverage exposed are fixed: regex Expectations now execute on Oracle instead of raising, query-based Expectations such as UnexpectedRowsExpectation now run because the derived-table alias is rendered in the form Oracle's grammar accepts, and daily and monthly batch definitions now work because the date-part string cast carries a length Oracle accepts. No other backend's rendered SQL changes. (#12085, #12103, #12104, #12102)
batch_definition = asset.add_batch_definition_daily(
name="daily", column="event_date"
)
add_store no longer crashes and empties great_expectations.yml — Calling context.add_store() with an existing store's name and a config containing a store_backend key crashed and left great_expectations.yml at 0 bytes, making the project unloadable. The config is now serialized before the file is opened, so a serialization failure leaves the existing file byte-for-byte intact, and the context id is written as a string that YAML can represent. (#12081)
current = context.config.stores[context.expectations_store_name]
context.add_store(
name=context.expectations_store_name,
config={
"class_name": current["class_name"],
"store_backend": dict(current["store_backend"]),
},
)
Clearer errors for unsupported regex dialects and masked Azure account keys — Regex Expectations run against a SQL dialect with no regex support now report Regex is not supported for dialect <name> in exception_info instead of an empty message, and Azure connection strings are masked regardless of field order so an account key can no longer appear unmasked in a StoreConfigurationError. (#12109, #12094)
Data Docs pages render for validation results with no run_id — context.build_data_docs() silently dropped a validation result's page when the result's meta had no "run_id" key. Such results now render, with run name and run time shown as __none__. (#12098)
context.build_data_docs() # renders a page for every persisted result
Documentation for the bundled agent skills — A new environment-setup page teaches how to install, verify, use, upgrade, and remove the agent skills that ship inside the great_expectations package, including the overwrite contract and the three skills as one path. (#12074, #12106)
python -m great_expectations skills install
python -m great_expectations skills list
marshmallow>=3.18.0 and the <4.0.0 cap removed; config_version bounds checking behaves identically on both majors. (#12092)context.add_store() no longer crashes and truncate great_expectations.yml to 0 bytes when re-supplying a store's own config containing a store_backend key; the project config is now serialized before the file is opened, the context id is persisted as a string, and an absent context id stays empty rather than becoming the string "None". (#12081)exception_info instead of an empty exception message, with the dialect name rendered cleanly. (#12109)ExpectColumnPairValuesToBeInSet on pandas now returns a verdict instead of a MetricResolutionError when the evaluated rows do not use a zero-based consecutive index, such as after null filtering or with a custom DataFrame index. (#12097)usedforsecurity=False, so batch identification, dataframe fingerprinting, and partitioning/sampling work on FIPS-enabled hosts. No computed digests change. (#12099)EndpointSuffix no longer raises a StoreConfigurationError containing the raw URL and account key. (#12094)context.build_data_docs() now renders a page for a validation result whose meta has no "run_id" key (or whose run_id is None), defaulting run name and run time to __none__ instead of silently dropping the page. (#12098)UnexpectedRowsExpectation now execute on Oracle: the derived-table alias is rendered through one shared helper that omits AS only for the grammar that rejects it, and the literal-boolean predicate rewrite now also applies to Oracle. No other backend's rendered SQL changes. (#12104)NotImplementedError, via a new Oracle branch in the dialect-regex helper that renders REGEXP_LIKE in both positive and negated forms. No other dialect's rendered SQL changes. (#12103)add_batch_definition_daily and add_batch_definition_monthly now work on Oracle: the multi-date-part partition query's string cast supplies an explicit length for the dialect that requires one, so batch retrieval no longer fails with ORA-00906. Curated coverage for daily and monthly batch definitions was added for every curated backend. (#12102)--symlink failure wording matches the installer's actual behavior. (#12106)python -m great_expectations skills install and skills list, the overwrite contract, the three skills as one path, and how to upgrade and remove them. (#12074)<details> <summary>Maintenance</summary>
__init__.py under a hyphenated, unimportable directory and gave tests/integration/test_script_runner.py its own shell helper instead of importing one from assets/. (#12113)oracledb driver requirement for the test lane, a pinned Oracle 21c container, an oracle pytest marker and CI lane, and curated-tier coverage; no great_expectations[oracle] extra is published yet. (#12085)</details>
Thanks to @Dev-iL (first contribution), @ArjunPakhan (first contribution), @joebasrawi (first contribution), @nanjeshramesh, @dkling-it (first contribution), @hemalrajput18 (first contribution), @MannXo (first contribution).
[MAINTENANCE] Sync SparkDBFSDatasource schema with its deprecation notice
Compatibility: new extra gcs; gx-sqlalchemy-redshift removed (extra gx-redshift); sqlalchemy-redshift added (extra gx-redshift); sqlalchemy minimum → 1.4.0 (extra redshift)
Integer batch parameters work on every datasource family — Numeric batch parameters such as year and month now accept integers on file, directory, and SQL assets alike, so a single batch_parameters dict drives one checkpoint spanning files and a warehouse. Digit strings still work but now emit a deprecation warning. A SQL request that matches nothing also explains why, distinguishing an empty table or column from candidates that exist but did not match, and naming the offending parameter and value. (#12065)
checkpoint.run(batch_parameters={"year": 2020, "month": 4})
Agent-skill guidance ships with the package — Great Expectations now bundles version-matched guidance for coding agents covering data source configuration, expectation authoring, and checkpoint orchestration, installable into a project's agent discovery directories. The guidance names the right optional dependency group for a missing driver, offers a batching cadence instead of assuming one, carries reuse-safe worked examples, and will not install packages, edit configuration files, create a project directory, or save files unless the user asked for it. (#12061, #12062, #12063, #12068, #12073)
python -m great_expectations skills install --target all
python -m great_expectations skills list
Two experimental expectations promoted into the core library — ExpectColumnValuesToNotBeOutliers (IQR and standard-deviation methods) and the multicolumn values-equal expectation are now supported core expectations on Pandas, SQL, and Spark, with null-safe evaluation, Gallery metadata, prescriptive rendering, and public exports. (#12011, #12018)
gx.expectations.ExpectColumnValuesToNotBeOutliers(
column="fare_amount", method="iqr", multiplier=1.5
)
Redshift installs the upstream SQLAlchemy dialect — pip install 'great_expectations[redshift]' now resolves the upstream sqlalchemy-redshift 1.0.0 dialect with SQLAlchemy 2, replacing the Great Expectations fork and lifting a sqlalchemy<2.0.0 pin that had been holding the dialect back at a three-year-old release. The gx-redshift extra keeps working as a deprecated alias that resolves identically. (#12044)
pip install 'great_expectations[redshift]'
S3 requests are attributable to Great Expectations — S3 clients built by Great Expectations now send a great-expectations/<version> user-agent suffix, appended to any user-supplied agent string rather than replacing it, so operators and S3-compatible providers can see which requests originate from Great Expectations. The S3 Data Source docs also clarify that endpoint_url is how you connect to a non-AWS S3-compatible store. (#11937)
Data Docs no longer advertises a removed suite-editing workflow — The "How to Edit This Suite" button and its popup, which pointed at a CLI command and notebook workflow that no longer exist, are gone from expectation suite and validation results pages, and expectation suite, profiling, and site index pages no longer render an empty "Actions" card. Validation results pages keep the Actions card and its Show All / Failed Only filter. (#12078)
\{"year": "2024", "month": "02"}) are deprecated; pass integers instead (\{"year": 2024, "month": 2}). Removal in 2.0.0. (#12065)gx-redshift install extra is deprecated and is now an alias that resolves identically to redshift; use great_expectations[redshift]. Removal in 2.0.0. (#12044)great-expectations/<version> user-agent suffix, appended to any user-supplied agent string, and the S3 Data Source docs clarify that endpoint_url connects to an S3-compatible object store. (#11937)ExpectColumnValuesToNotBeOutliers is now a supported core expectation on Pandas, SQL, and Spark, with IQR and standard-deviation detection, consistent null handling, inclusive threshold boundaries, and a clear error for unsupported methods. (#12011)batch_parameters dict drives a checkpoint spanning files and SQL; digit strings still work but warn, and a SQL request matching nothing now explains whether the data is absent or the parameter did not match. (#12065)python -m great_expectations skills install and listable with skills list; the installer leaves already-correct destinations alone, refuses directories it did not write, and requires --force to replace user-edited copies. (#12061)<details> <summary>Maintenance</summary>
show_how_to_buttons site config option still loads but gates nothing. (#12078)SparkDBFSDatasource JSON schema description now includes the deprecation notice the Python API has carried since 1.16.0, so schema-driven consumers see it too. (#12077)add_or_update_* factory method and each expectation to its schema, description, data quality issues, and supported data sources. (#12055)gx-redshift CI launch key now that the Redshift lane selects the canonical redshift marker; the deprecated gx-redshift install extra is unaffected. (#12060)redshift extra now installs upstream sqlalchemy-redshift 1.0.0 with sqlalchemy>=1.4.0 instead of the Great Expectations fork, fixing a pin that had been silently installing a three-year-old dialect; gx-redshift remains as a deprecated alias resolving identically. (#12044)--gcs flag, and the Spark-on-GCS guide is migrated to the current asset API so it no longer documents a call that raises. (#12059)gcs_deps marker and requirements file so they actually run in CI. (#12058)GX_GCS_TEST_BUCKET environment variable and raise clearly when it is unset; the published GCS guides now show a my_bucket placeholder instead of the real CI bucket. (#12056):latest. (#12054)</details>
Thanks to @goanpeca (first contribution), @chavalasantosh (first contribution), @AtomicGlance (first contribution).
[BUGFIX] Support standing up a FileDataContext on a read-only filesystem
Quantile expectations are correct on SQLite and no longer error on all-null columns — ExpectColumnQuantileValuesToBeBetween now selects the right rank on SQLite and ignores null values when computing quantiles, so observed quantiles match the other backends. A column with no non-null values now reports an unmet expectation — success: false with null observed values and per-quantile success details — on every backend instead of raising a TypeError on SQL backends or an IndexError on Spark. (#12008, #12026)
import great_expectations.expectations as gxe
suite.add_expectation(
gxe.ExpectColumnQuantileValuesToBeBetween(
column="passenger_count",
quantile_ranges={"quantiles": [0.25, 0.5], "value_ranges": [[1, 2], [1, 3]]},
)
)
Faster expect_column_values_to_be_unique on wide SQL tables — The SQLAlchemy implementation of column_values.unique now scans the source table once through a narrow window over only the target column, and only retrieves full rows (via a narrow duplicate-key join) when SUMMARY or COMPLETE result formats are requested. Wide column-store tables — where the previous query was cancelled by Redshift's workload-management timeouts — now validate reliably. (#11863)
import great_expectations.expectations as gxe
gxe.ExpectColumnValuesToBeUnique(column="id")
Validating multiple expectations on the same metric no longer fails on strict SQL backends — When several expectations in a suite depend on the same underlying metric, the generated SQL now gives each metric a unique alias, so backends such as Postgres no longer reject the query with Duplicated field name in view schema. (#11905)
File-backed Data Contexts work on read-only, version-controlled projects — gx.get_context(mode="file") now recognizes a project as already set up based on a committed great_expectations.yml alone, instead of requiring the gitignored uncommitted/ runtime directories. A clean checkout on a read-only filesystem is no longer mistaken for an unscaffolded project and no longer crashes during initialization. (#12000)
import great_expectations as gx
context = gx.get_context(mode="file", project_root_dir="/path/to/checkout")
ExpectColumnValuesToMatchStrftimeFormat is now a supported core Expectation — The Expectation now carries full support metadata and a generated schema, appears in the Expectation Gallery with a properly rendered docstring and examples, and declares a backend matrix of Pandas and Spark (SQL is out of scope). (#12009)
import great_expectations.expectations as gxe
gxe.ExpectColumnValuesToMatchStrftimeFormat(
column="event_date",
strftime_format="%Y-%m-%d",
mostly=0.95,
)
Adding a table asset is much faster on projects with many schemas — TableAsset.test_connection() now probes the table first and only lists server schemas if that probe fails, purely to refine the error message. On backends where schema listing is a server-wide metadata operation — for example a BigQuery project with thousands of datasets — adding a table asset no longer pays that cost. Table configurations whose schema name did not match the normalized schema listing but were otherwise accessible now succeed. (#12020)
asset = datasource.add_table_asset(name="my_asset", table_name="my_table", schema_name="my_schema")
ExpectColumnValuesToMatchStrftimeFormat is promoted to a supported core Expectation, with a corrected and Gallery-formatted docstring, support metadata, a generated JSON schema, a declared Pandas and Spark backend matrix, and expanded test coverage including mostly thresholds. (#12009)ExpectColumnQuantileValuesToBeBetween no longer reports a quantile one rank too low on SQLite and no longer raises on columns containing null values; quantile ranks are computed from non-null counts with exact fractional arithmetic, and the MySQL query applies the same null filter. (#12008)ExpectColumnQuantileValuesToBeBetween now reports an unmet expectation, with null observed values and per-quantile success details, for a column that has no non-null values, instead of raising on SQL backends and Spark; the Spark metric returns one null per requested quantile so column.quantile_values has the same shape on every backend. (#12026)expect_column_values_to_be_unique on SQLAlchemy backends now runs a single narrow pass over the target column, and only joins back to the source for full-row details under SUMMARY/COMPLETE result formats, eliminating the Redshift workload-management timeouts seen on very wide tables. (#11863)great_expectations.yml rather than from gitignored uncommitted/ directories, so a clean checkout is no longer destructively re-scaffolded. (#12000)<details> <summary>Maintenance</summary>
fast-uri dependency from 3.1.4 to 3.1.5, which includes a security fix. (#12019)gx_ci_test_ prefix, and the stale-schema cleanup patterns were corrected to match hex suffixes so stale schemas are actually swept. No library behavior changes. (#12015)brace-expansion dependency from 1.1.16 to 1.1.18. (#12014)postcss dependency from 8.5.12 to 8.5.25. (#12013)/assign-me on an issue that is not yet labeled ready for work now gets a posted explanation of why the claim was declined and where to find issues open for claiming, instead of silently doing nothing. (#11999)</details>
Thanks to @SreeramaYeshwanthGowd (first contribution), @TemidayoA (first contribution), @leodrivera, @nanjeshramesh (first contribution).
[MAINTENANCE] Ignore pyOpenSSL X509.get_subject deprecation warning for snowflake
Data Docs no longer errors when unexpected indices contain only id/pk columns — Validation results that report unexpected indices made up solely of the configured id/pk columns — including Spark and SQL runs and any run with unexpected values excluded — now render a count and index table in Data Docs instead of failing the result page with "No group keys passed!". (#11935)
Contributor License Agreement checks are now run by the project itself — The verification/cla-signed check is posted by the repository's own workflows rather than a third-party hosted app: it is reported on pull request heads and on merge-queue commits, fails closed when contributor status cannot be confirmed, leaves a single guiding comment naming any unsigned or unidentified committer, and keeps the cla-signed / cla-not-signed labels in sync with the check result. CLA signing links now point at the current forms. (#11985, #11983, #11980, #11992, #11982, #11974)
Distinct-values set Expectations document their observed_value contract — Documentation and JSON schemas for the distinct values set Expectations now state that observed_value is always None, and their code examples show unexpected_count and partial_unexpected_list (plus the missing-value variants) instead. (#11934)
<details> <summary>Maintenance</summary>
</details>
Thanks to @anxkhn, @EshwarCVS.
[BUGFIX] Preserve date-like strings in SQL distinct value sets ( #11947 ) (thanks @yuricavalcanti06 )
Compatibility: zstandard added (extra spark-connect)
Spark 4 and ANSI mode support — Great Expectations now works with Spark 4, including ANSI mode, so you can validate Spark DataFrames on the latest Spark release without pinning to Spark 3. The spark-connect extra now also installs zstandard. (#11969)
import great_expectations as gx
context = gx.get_context()
data_source = context.data_sources.add_spark(name="my_spark")
asset = data_source.add_dataframe_asset(name="my_df")
batch = asset.add_batch_definition_whole_dataframe("batch").get_batch(
batch_parameters={"dataframe": spark_df}
)
Date-like strings stay strings in SQL value sets — Distinct-value expectations against SQL data sources no longer turn non-ISO, date-like text such as "10-20" into a date, so text bins in a value_set compare correctly against text columns. Only strict YYYY-MM-DD strings are converted to dates. (#11947)
import great_expectations.expectations as gxe
gxe.ExpectColumnDistinctValuesToBeInSet(
column="bin",
value_set=["10-20", "20-30"],
)
Clearer failure for an empty regex_list — ExpectColumnValuesToMatchRegexList now rejects an empty regex_list at construction time with the message "regex_list must not be empty", instead of failing later during validation with an opaque "No objects to concatenate" error. This matches the behavior of ExpectColumnValuesToNotMatchRegexList. (#11958)
import great_expectations.expectations as gxe
gxe.ExpectColumnValuesToMatchRegexList(column="my_col", regex_list=[])
# pydantic.ValidationError: ... regex_list must not be empty
ExpectColumnValuesToMatchRegexList now fails at construction with "regex_list must not be empty" when given an empty list, instead of raising an opaque error at validation time. (#11958)value_set are no longer parsed into dates for SQL distinct-value expectations; only strict YYYY-MM-DD strings are converted. (#11947)<details> <summary>Maintenance</summary>
/assign-me, /unassign-me) with automatic release of idle claims, and narrowed issue staleness to issues labeled as needing more information. (#11957)</details>
Thanks to @anxkhn (first contribution), @yuricavalcanti06 (first contribution).
[MAINTENANCE] Fix pytest parametrize non-Collection iterable deprecation breaking scheduled CI
Spark Connect compatibility for distinct-values expectations — Expectations that rely on a column's distinct values — including expect_column_distinct_values_to_equal_set, expect_column_distinct_values_to_contain_set, and expect_column_distinct_values_to_be_subset_of — now run successfully against a Spark Connect session (for example Databricks serverless via an sc:// URL) instead of failing with a MetricResolutionError. Classic Spark sessions behave exactly as before. (#11922)
import great_expectations as gx
batch = ... # a Spark Connect-backed batch
batch.validate(
gx.expectations.ExpectColumnDistinctValuesToEqualSet(
column="color", value_set=["red", "green", "yellow"]
)
)
<details> <summary>Maintenance</summary>
</details>
Thanks to @zozo123 (first contribution).
[BUGFIX] Regex angle brackets not HTML-escaped in Data Docs
Data Docs now renders regex and other parameter values containing <, >, or & correctly — Expectation parameter values are HTML-escaped before being substituted into Data Docs render templates. Previously, a regex containing angle brackets — for example the negative lookbehind (?<!\s) — was emitted raw into the HTML, where the browser treated <! as the start of a comment and silently truncated the rendered pattern. Such values now display literally in Data Docs. The public API and serialized Expectation format are unchanged; only the HTML rendering layer is affected. (#11909)
gx.expectations.ExpectColumnValuesToMatchRegex(
column="my_column",
regex=r"(?<!\s)foo",
)
# The regex now appears in full in the generated Data Docs page.
<, >, or & — such as regexes using a negative lookbehind — are now HTML-escaped and render correctly in Data Docs instead of being truncated or hidden. (#11909)<details> <summary>Maintenance</summary>
</details>
[MINORBUMP] GX Cloud shutdown: raise on CloudDataContext construction and remove cloud test suites
GX Cloud paths now fail immediately with a clear explanation — GX Cloud has been shut down. Constructing a CloudDataContext directly, or asking get_context(...) for a cloud context (via mode="cloud", cloud_mode=True, a complete set of cloud_* arguments, or GX_CLOUD_* environment configuration), now raises a GreatExpectationsError right away instead of failing later with an opaque connection error. The message states that GX Cloud has been shut down and that these entry points will be removed in great_expectations 2.0. Non-cloud usage is unchanged, and the cloud classes and parameters remain importable with unchanged signatures through the 1.x line. (#11894)
import great_expectations as gx
# Raises GreatExpectationsError:
# "GX Cloud has been shut down, so this no longer functions and will be
# removed in great_expectations 2.0."
context = gx.get_context(mode="cloud")
# Non-cloud contexts still work as before
context = gx.get_context(mode="file")
CloudDataContext and the GX Cloud branch of get_context(...) (the cloud_* parameters, mode="cloud", cloud_mode=True, and GX_CLOUD_* environment configuration) are deprecated and now raise an error; the cloud-only exception, store, config, and identifier symbols remain importable only as shells. Use a non-cloud context such as gx.get_context(mode="file") or gx.get_context(mode="ephemeral"). Removal in 2.0.0. (#11894)CloudDataContext or requesting a cloud context from get_context(...) now raises a GreatExpectationsError explaining the shutdown instead of failing with an opaque connection error. Cloud classes and parameters stay importable with unchanged signatures until they are removed in great_expectations 2.0, and non-cloud usage is unaffected. (#11894)<details> <summary>Maintenance</summary>
BINARY(8388608) observed type that the Snowflake connector reports for VARBINARY columns. (#11892)</details>
[BUGFIX] Preserve boolean values passed to add_csv_asset (fixes #11206 ) ( #11867 ) (thanks @EshwarCVS )
SQLAlchemy 1.4 users can run uniqueness expectations again — Expectations that resolve the column_values.unique.condition metric no longer fail on SQLAlchemy 1.4 with AttributeError: module 'sqlalchemy' has no attribute 'Select', restoring compatibility for dialects still pinned to SQLAlchemy 1.x (such as ClickHouse, Redshift, and Teradata). (#11876)
Boolean options passed to pandas assets are preserved — Boolean values such as index_col=False handed to add_csv_asset are kept as booleans instead of being silently converted to strings, so pandas interprets them as flags rather than column names. This applies to boolean options across the pandas asset types. (#11867)
data_source.add_csv_asset(name="my_asset", path="data.csv", index_col=False)
AttributeError on SQLAlchemy 1.4 when an expectation resolved the column_values.unique.condition metric, restoring SQLAlchemy 1.4 compatibility for unique-value expectations. (#11876)add_csv_asset and other pandas assets, such as index_col=False, are no longer coerced to strings and are now applied as the boolean flags pandas expects. (#11867)<details> <summary>Maintenance</summary>
</details>
Thanks to @ranophoenix (first contribution), @EshwarCVS (first contribution).
[BUGFIX] Data Docs uses vulnerable jQuery 3.4.1
Compatibility: pytest-split added (extra test)
Data Docs now load a patched jQuery — Data Docs pages generated by Great Expectations now reference jQuery 3.7.1 instead of the vulnerable 3.4.1 (CVE-2020-11022, CVE-2020-11023). The Data Docs UI is unchanged and security scanners no longer flag GX-generated pages for this issue. (#11856)
Spark column names containing dots now work — Spark-backed data assets whose column names contain dots (for example Data.Entrega) can now be used in expectations without the spurious "The column X in BatchData does not exist" error. (#11851)
batch.validate(gxe.ExpectColumnValuesToNotBeNull(column="Data.Entrega"))
Nested Spark struct paths supported in unexpected_index_column_names — Referencing a nested Spark struct path such as Data.evt.id in unexpected_index_column_names no longer raises InvalidMetricAccessorDomainKwargsKeyError; unexpected rows are surfaced keyed by the full dotted path. (#11835)
batch.validate(
gxe.ExpectColumnValuesToBeInSet(column="value", value_set=[1, 2]),
result_format={
"result_format": "COMPLETE",
"unexpected_index_column_names": ["Data.evt.id"],
},
)
Compound uniqueness expectations work on Spark timestamps with Pandas 2.x — expect_compound_columns_to_be_unique and expect_select_column_values_to_be_unique_within_record no longer fail with ValueError: Passing in 'datetime64' dtype with no precision is not allowed. on Spark DataFrames that contain timestamp columns. Entries in partial_unexpected_list are now native Python values (for example datetime.datetime and None) rather than pandas/numpy equivalents. (#11861)
Expectation subclasses can use aliased Pydantic fields — Subclassing a built-in Expectation and declaring a field with Field(alias=...) no longer causes ValidationError: extra fields not permitted when validating a batch. (#11854)
class MyExpectation(gxe.ExpectColumnValuesToStartWith):
regex: str = pydantic.Field(alias="pattern")
batch.validate(MyExpectation(column="name", pattern="^a"))
Data.evt.id can now be used for expectation columns and unexpected index columns without raising an invalid-domain-kwargs error, and results are keyed by the full dotted path. (#11835)<details> <summary>Maintenance</summary>
pytest-split is now part of the development test requirements. (#11850)Expectation._atomic_prescriptive_template method and its add_values_with_json_schema_from_list_in_params helper, both deprecated in v0.15.43; use _prescriptive_template instead. (#11847)</details>
[MAINTENANCE] bump jest-environment-jsdom 30.2.0 → 30.3.0 ( CVE-2026-33671 )
Compatibility: new extra singlestore
SingleStore support — Great Expectations now works against SingleStore databases: SingleStoreDB is recognized as its own SQL dialect, regex and uniqueness expectations produce correct results, quoted identifiers are handled, and setup is covered in the documentation. Install with the new singlestore extra. (#11828, #11839, #11837, #11842)
pip install 'great_expectations[singlestore]'
import great_expectations as gx
context = gx.get_context()
data_source = context.data_sources.add_sql(
name="my_singlestore",
connection_string="singlestoredb://user:password@host:3306/my_db",
)
strict_min and strict_max now respected in expect_column_value_lengths_to_be_between on Spark and SQL — Passing strict_min=True or strict_max=True to expect_column_value_lengths_to_be_between previously produced inclusive-bound results on the Spark and SQL backends. Both backends now apply strictly exclusive bounds, matching the documented semantics and the Pandas backend. Non-strict usage is unchanged. (#11834, #11836)
import great_expectations.expectations as gxe
suite.add_expectation(
gxe.ExpectColumnValueLengthsToBeBetween(
column="name", min_value=2, max_value=4, strict_min=True, strict_max=True
)
)
Timezone-aware strftime formats accepted — ExpectColumnValuesToMatchStrftimeFormat no longer raises a validation error when the format contains %z, and the Spark implementation of the underlying metric now validates timezone-aware formats correctly. (#11812, #11817)
import great_expectations.expectations as gxe
gxe.ExpectColumnValuesToMatchStrftimeFormat(
column="ts", strftime_format="%Y-%m-%d %H:%M:%S%z"
)
Forecast store bounds used for windowed expectations — Windowed expectations now send each expectation's batch definition to the expectation-parameters endpoint, so users on the asynchronous forecast store path receive stored forecast bounds instead of falling back to inline training. Checkpoints spanning several batch definitions fetch and merge parameters for each one. (#11831)
expect_column_value_lengths_to_be_between on Spark now honors strict_min and strict_max, excluding boundary lengths as documented. (#11834)expect_column_value_lengths_to_be_between on SQL data sources now honors strict_min and strict_max, applying strictly exclusive bounds as documented. (#11836)%z validate correctly. (#11817)ExpectColumnValuesToMatchStrftimeFormat no longer raises a validation error when the format string includes the %z timezone directive. (#11812)expect_column_proportion_of_non_null_values_to_be_between supports a forecasted range. (#11821)<details> <summary>Maintenance</summary>
data_context, datasource_name, batch_parameters, and batch_kwargs arguments (and their read-only properties) from the internal Batch class. (#11843)ColumnMetricProvider alias and its deprecation-warning metaclass, which were deprecated in favor of ColumnAggregateMetricProvider. (#11832)run_id to Validator.validate(); a run identifier or dict is required. (#11826)</details>
[FEATURE] Add Pact contract tests for datasource API coverage gaps
<details> <summary>Maintenance</summary>
</details>
Thanks to @Adeyinka1 (first contribution).
[BUGFIX] Pin invoke==3.0.0 to avoid breaking change in 3.0.2
column.unique_proportion in metric list runs (#11786)Compatibility: pact-python added (extra cloud); invoke minimum 2.0.0 removed (extra test); pact-python minimum 2.0.1 → 3.1.0 (extra test)
column.unique_proportion available in metric list runs — Metric list runs can now compute column.unique_proportion alongside the existing column metrics, so you can retrieve the proportion of unique values per column without a separate run. (#11786)
from great_expectations.experimental.metric_repository.metrics import MetricTypes
metrics = [MetricTypes.COLUMN_UNIQUE_PROPORTION]
Suites added to an ephemeral context now pass freshness checks — context.suites.add(suite) now returns a suite that matches what was stored, so passing that suite straight into context.validation_definitions.add() in an ephemeral context no longer raises a freshness error. (#11758)
suite = context.suites.add(suite)
context.validation_definitions.add(
gx.ValidationDefinition(name="vd", data=batch_definition, suite=suite)
)
Correct exact_match default in ExpectTableColumnsToMatchSet output — Rendered descriptions for ExpectTableColumnsToMatchSet now reflect the expectation's real default for exact_match, so the rendered text no longer contradicts how the expectation actually validates. (#11785)
Microsoft Teams and Jira integration documentation — The documentation now covers setting up the Microsoft Teams and Jira integrations for GX notifications and issue tracking. (#11761, #11741)
No more unclosed-SQLite resource warnings on Python 3.13 — SQLAlchemy execution engines now dispose their connection pool when they are garbage collected, eliminating ResourceWarning: unclosed database noise and the spurious failures it caused on Python 3.13. (#11766)
PandasDBFSDatasource and SparkDBFSDatasource is deprecated; use datasources backed by Unity Catalog volumes, external locations, or workspace files. Removal in 2.0.0. (#11759)column.unique_proportion metric, computed alongside other column-level metrics. (#11786)PandasDBFSDatasource and SparkDBFSDatasource are now marked as deprecated, following Databricks' deprecation of DBFS; the classes still work and will be removed in a future major release. (#11759)ExpectTableColumnsToMatchSet renderers now use the correct default value for exact_match in their rendered output. (#11785)invoke dependency to 3.0.0 to avoid a breaking change introduced in 3.0.2. (#11781)<details> <summary>Maintenance</summary>
Gx-Version request header with a pattern instead of a literal value, so generated contracts are stable across commits. (#11791)brace-expansion from 1.1.12 to 1.1.13 in the documentation site dependencies. (#11777)lodash from 4.17.23 to 4.18.1 in the documentation site dependencies, picking up prototype-pollution and template code-injection fixes. (#11776)great_expectations.profile module and its tests as dead code. (#11763)</details>
[FEATURE] Refactor rest_contracts/conftest.py for client-driven Pact testing (GX-2725)
BigQuery datasource methods now surface in IDE autocomplete and type checking — The typed stub for context.data_sources now declares add_bigquery, update_bigquery, add_or_update_bigquery, and delete_bigquery, so BigQuery-specific datasource methods are discoverable in editor autocomplete and recognized by type checkers instead of pushing you toward the generic add_sql method. (#11736)
datasource = context.data_sources.add_bigquery(
name="my_bigquery_ds",
connection_string="bigquery://my-project/my_dataset",
)
New how-to guide: retrieve all unexpected rows — The documentation now includes a "Retrieve all unexpected rows" guide under Run Validations, with a runnable example showing how to get the full set of unexpected rows from a validation definition, plus cross-references from the custom SQL Expectation guide and the result format reference table. (#11712)
unexpected_rows = validation_definition.get_unexpected_rows(batch_parameters={})
run_rest_api_pact_test REST contract test helper is deprecated; use the client-driven Pact test approach built on the pact_cloud_context fixture. Removal in 2.0.0. (#11753)pact_cloud_context fixture builds a CloudDataContext against the Pact mock server without real cloud credentials, a shared data-context configuration response and interaction helper are available for reuse, and the older run_rest_api_pact_test helper is marked deprecated. (#11753)add_bigquery, update_bigquery, add_or_update_bigquery, delete_bigquery) and the BigQueryDatasource type are now declared in the datasources type stub, so they appear in IDE autocomplete and type checking instead of requiring the generic add_sql method. (#11736)<details> <summary>Maintenance</summary>
yaml dependency from 1.10.2 to 1.10.3. (#11746)flatted dependency from 3.3.3 to 3.4.2. (#11737)GenericSQLDatasourceTestConfig) to make it easier to try out new SQL datasources. (#11718)</details>
Thanks to @Julian901 (first contribution).
[MAINTENANCE] remove nested actions format
<details> <summary>Maintenance</summary>
</details>
[MINORBUMP] SQL Server and Fabric Data Sources
Fetch all unexpected rows from an UnexpectedRowsExpectation — ValidationDefinition.get_unexpected_rows() returns every failing row for an UnexpectedRowsExpectation, without the 200-row cap applied to validation results. Validation results also gained an ExpectationValidationResult.expectation property and an ExpectationSuiteValidationResult.batch_parameters property, so you can feed a failed result straight back in to retrieve its rows. (#11711)
result = validation_definition.run(batch_parameters={"year": 2026, "month": 3})
for evr in result.results:
if not evr.success:
rows = validation_definition.get_unexpected_rows(
evr.expectation,
batch_parameters=result.batch_parameters,
)
if rows:
write_to_quarantine(rows)
Documentation for SQL Server and Fabric data sources — The docs now cover creating and using SQL Server and Microsoft Fabric data sources. (#11686)
A single failing metric no longer fails the whole batch of metrics — When bulk metric resolution hits an error, metrics are now retried individually, so one problematic metric no longer causes every metric in the run to error. (#11708)
ValidationDefinition.get_unexpected_rows() to fetch all failing rows for an UnexpectedRowsExpectation without the 200-row cap, plus an ExpectationValidationResult.expectation property and an ExpectationSuiteValidationResult.batch_parameters property for post-run workflows. (#11711)partial_unexpected_count controls the number of values shown in partial_missing_list. (#11705)<details> <summary>Maintenance</summary>
trino package began shipping type information. (#11707)</details>
[MAINTENANCE] Deprecate schema_name on all TableAssets
trust_server_certificate to SQLServerDatasource (#11694)docs-creds-needed (#11698)schema_name on all TableAssets (#11689)Identical to 1.13.1, re-published the same day as a minor version: the release above carried a deprecation, which the minor number signals. No changes beyond 1.13.1.
[MAINTENANCE] Deprecate schema_name on all TableAssets
trust_server_certificate to SQLServerDatasource (#11694)docs-creds-needed (#11698)schema_name on all TableAssets (#11689)Trust a SQL Server certificate without turning off encryption — SQL Server and Fabric data sources accept a new trust_server_certificate option, so you can connect to a server presenting a self-signed or otherwise untrusted certificate while keeping encryption enabled instead of weakening encrypt to "Optional". The option works with both SQL Server authentication and Entra ID. (#11694)
import great_expectations as gx
context = gx.get_context()
datasource = context.data_sources.add_sql_server(
name="my_sql_server",
host="my-host",
database="my_database",
username="my_user",
password="my_password",
trust_server_certificate=True,
)
Config variable substitution errors no longer echo secret text — When a password or secret contains a literal $, Great Expectations no longer includes the text following the $ in the resulting missing-config-variable error message, so part of the secret is not leaked in logs. The error guidance also no longer points at the retired $MY_CONFIG_VAR substitution syntax. (#11693)
schema_name parameter on TableAsset and add_table_asset is deprecated; use the schema configured on the SQL data source's connection string. Removal in 2.0.0. (#11689)trust_server_certificate option, letting you trust a self-signed or untrusted server certificate while keeping the connection encrypted, with both SQL Server authentication and Entra ID. (#11694)*.service-now.com in the default allowed email domains. (#11669)<details> <summary>Maintenance</summary>
schema_name parameter on TableAsset and add_table_asset is deprecated; table assets now resolve their schema from the SQL data source they belong to, so specify the schema in the data source's connection configuration instead. (#11689)SENTRY_DSN environment variable and no performance data is collected. (#11695)$ in a password or secret, avoiding partial secret leakage, and their guidance no longer references the removed $MY_CONFIG_VAR syntax. (#11693)qs dependency from 6.14.1 to 6.14.2. (#11660)</details>
[MAINTENANCE] remove deprecated store backends
ExpectColumnDistinctValuesToBeInSet with database-pushed comparison (#11614)ExpectColumnDistinctValuesToContainSet with database-pushed comparison (#11615)ExpectColumnDistinctValuesToEqualSet with database-pushed comparison (#11616)FabricDatasource (#11685)SQLServerDatasource schema (#11662)Compatibility: altair minimum 4.2.1 → 5.0.0; new extra fabric; removed extra mssql; new extra sql-server
Microsoft Fabric datasource — You can now connect to Microsoft Fabric with the new Fabric datasource, which authenticates with an Entra ID service principal. Install it with the new fabric extra. (#11685, #11662)
import great_expectations as gx
context = gx.get_context()
datasource = context.data_sources.add_fabric(
name="my_fabric",
host="my-workspace.datawarehouse.fabric.microsoft.com",
database="my_warehouse",
client_id="<client-id>",
client_secret="<client-secret>",
)
Distinct-value set expectations now compare in the database — ExpectColumnDistinctValuesToBeInSet, ExpectColumnDistinctValuesToContainSet, and ExpectColumnDistinctValuesToEqualSet now push set comparison into the database instead of pulling every distinct value into memory, so they stay fast and produce small results on high-cardinality columns. Results no longer include the full list of distinct values as observed_value; instead they report unexpected_count/partial_unexpected_list and/or missing_count/partial_missing_list, each capped at 20 values. (#11614, #11615, #11616)
import great_expectations.expectations as gxe
suite.add_expectation(
gxe.ExpectColumnDistinctValuesToBeInSet(
column="my_col",
value_set=["a", "b", "c"],
)
)
"SQL Server" naming throughout, including the pip extra — User-facing references to MSSQL are now written as SQL Server. Install SQL Server support with the renamed extra. (#11674)
pip install 'great_expectations[sql-server]'
pandas 3 support — The upper pin on pandas has been removed, so Great Expectations can be installed alongside pandas 3. BigQuery reads fall back to pandas_gbq.read_gbq, and chart rendering works with pandas 3's new string dtype default (requires altair 5). (#11677)
ExpectAI documentation for the agent — The documentation now covers ExpectAI for the agent, including its prerequisites. (#11644, #11678)
ExpectColumnDistinctValuesToEqualSet now compares the value set inside the database rather than loading all distinct values into memory. Results return observed_value: None along with unexpected_count, partial_unexpected_list, missing_count, and partial_missing_list (each capped at 20 values), and the rendered output marks unexpected and missing values accordingly. (#11616)ExpectColumnDistinctValuesToBeInSet now compares the value set inside the database rather than loading all distinct values into memory. Results return observed_value: None along with unexpected_count and partial_unexpected_list (capped at 20 values), and the descriptive value-counts bar chart is no longer produced. (#11614)ExpectColumnDistinctValuesToContainSet now compares the value set inside the database rather than loading all distinct values into memory. Results return observed_value: None along with missing_count and partial_missing_list (capped at 20 values), and the rendered output marks missing values. (#11615)add_fabric(), update_fabric(), and delete_fabric() APIs, authenticated with an Entra ID service principal. (#11685){batch} placeholder on Databricks no longer fail with a cast error, because batch queries are now compiled with the datasource's own dialect so identifiers are quoted correctly. (#11671)<details> <summary>Maintenance</summary>
ExpectColumnValuesToBeOfType/ExpectColumnValuesToBeInTypeList, compared case-insensitively and without COLLATE clauses in the type string. (#11684)pandas_gbq.read_gbq and chart rendering updated for the new string dtype default. (#11677)gx.get_context(mode="ephemeral") instead of the removed build_in_memory_runtime_context helper, eliminating a source of flaky tests. (#11683)_get_default_value called with key ... but it is not a known field INFO log messages emitted during checkpoint validation. (#11626)sql-server instead of mssql, and related enum members, helper names, pytest markers, and the test CLI flag were renamed to match. The SQLAlchemy dialect value mssql is unchanged. (#11674)</details>
[DOCS] remove info about deprecated DBFS
SQLServerDatasource with SQL Server authentication (#11640)start_period to mercury healthcheck to prevent flaky CI failures (#11655)SQL Server datasources with Azure AD password authentication — You can now connect to SQL Server with a flat set of connection keyword arguments, including Azure Active Directory password authentication, without hand-building a connection string. (#11645, #11640, #11643)
context.data_sources.add_sql_server(
name="my_sql_server",
host="my-server.database.windows.net",
database="my_db",
username="user@example.com",
password="${MY_PASSWORD}",
)
Broader Microsoft SQL Server support — SQL Server now works with schemas, with bracket-quoted identifiers such as [my column], and with UnexpectedRowsExpectation queries. (#11649, #11652, #11646)
unexpected_index_query is returned for ExpectCompoundColumnsToBeUnique on SQL — ExpectCompoundColumnsToBeUnique run against SQL data sources now returns unexpected_index_query when you request it with return_unexpected_index_query=True or use the COMPLETE result format, so you can retrieve every failing row beyond the 200-row unexpected_list limit. COMPLETE also now honors return_unexpected_index_query=False when you set it explicitly. (#11639)
result = batch.validate(
ExpectCompoundColumnsToBeUnique(column_list=["a", "b"]),
result_format={"result_format": "COMPLETE"},
)
print(result.result["unexpected_index_query"])
pandas Timestamp values accepted in datetime comparisons — Datetime comparison expectations now handle pandas.Timestamp values correctly, checking the most specific type first so Timestamps are no longer mis-handled as plain dates. (#11637)
<details> <summary>Maintenance</summary>
</details>
Thanks to @subediparas5, @teixeirazeus (first contribution).
[BUGFIX] Fix Redshift fallback column detection for schema-qualified tables (#11606) (thanks @jni-bot)
Row conditions work again on SQLAlchemy 1.x data sources — Validating an expectation with a row_condition against a SQLAlchemy 1.x data source no longer fails with AttributeError: module 'sqlalchemy' has no attribute 'ColumnElement'. Row conditions now work on both SQLAlchemy 1.x and 2.x. (#11612)
Redshift column detection works for tables in non-default schemas — Redshift assets backed by a table in a non-default schema (for example bi_db.my_table) no longer fail column detection with relation "my_table" does not exist; the fallback lookup is now schema-qualified. (#11606)
row_condition with a SQLAlchemy 1.x data source no longer raises AttributeError: module 'sqlalchemy' has no attribute 'ColumnElement'. (#11612)airflow-provider-great-expectations documentation and removed an unused CI script. (#11621)<details> <summary>Maintenance</summary>
</details>
Thanks to @subediparas5 (first contribution).
[MAINTENANCE] Upper bound pandas to be below 3.0.0
Compatibility: pandas minimum set to 1.3.0 (python_version >= "3.12")
pandas 3.0 is excluded from supported versions — Installations now resolve a pandas version below 3.0.0, so environments no longer pick up an incompatible pandas 3.x release. On Python 3.12 and newer, the minimum supported pandas version is 1.3.0. (#11607)
Refreshed Result format documentation — The documentation covering result format has been reworked so it is easier to find the right result format setting and understand what each one returns. (#11596)
<details> <summary>Maintenance</summary>
pandas version to below 3.0.0. (#11607)</details>
[MAINTENANCE] Fix npm security vulnerabilities in docusaurus
Result-format levels are now respected for Custom SQL and Multi-Source Expectations — Validation results for Custom SQL and Multi-Source Expectations no longer include row-level data at result-format levels below COMPLETE. BOOLEAN_ONLY returns only success; BASIC and SUMMARY add the observed value (Custom SQL) or unexpected count and percent (Multi-Source); unexpected and missing rows appear only with COMPLETE. Multi-Source Expectations also render correctly when the result is empty. (#11601)
get_context is recognized as a public export by type checkers — The top-level great_expectations module now declares its public symbols explicitly, so static type checkers such as Pyright no longer report get_context and other promoted symbols as not exported. (#11578)
import great_expectations as gx
context = gx.get_context()
<details> <summary>Maintenance</summary>
</details>
Thanks to @ipriyankalimbad (first contribution).
[MAINTENANCE] Ignore DeprecationWarning emitted by deps
unexpected_rows as dicts in map expectation validation results (#11591)--pty and --no-pty flags to invoke deps (#11586)unexpected_rows for all Map expectations when opt-in flag is provided (#11583)DeprecationWarning emitted by deps (#11587)Unexpected rows are returned as dictionaries for Map expectations — Map expectation validation results now report unexpected_rows as dictionaries keyed by column name instead of database-specific row objects rendered as tuples, so results are easier to parse and no longer depend on an opt-in flag. (#11591, #11583)
result = batch.validate(expectation)
# result["result"]["unexpected_rows"]
# [{"col_a": 1.0, "col_b": 1.0, "col_c": 2.0}]
Column-based validations work on Redshift batches — batch.columns() no longer returns an empty list for Redshift batches on clusters with restricted information_schema access, and table names given as "schema.table" are resolved correctly, so column-based expectations run instead of failing with a metric domain error. (#11534)
batch = batch_definition.get_batch()
print(batch.columns())
Unexpected index query available with SUMMARY result format — return_unexpected_index_query is now supported with the SUMMARY result format, matching what BASIC already offered. (#11594)
result = batch.validate(
expectation,
result_format={
"result_format": "SUMMARY",
"unexpected_index_column_names": ["pk"],
"return_unexpected_index_query": True,
},
)
unexpected_rows as dictionaries by default, and the map_expectation_unexpected_rows_as_dict opt-in flag is no longer needed. (#11591)batch.columns() returning an empty list for Redshift batches, which caused column-based expectations to fail; column names are now retrieved via a fallback query when information_schema is inaccessible, and table_name values of the form "schema.table" are parsed correctly. (#11534)<details> <summary>Maintenance</summary>
return_unexpected_index_query is now supported with the SUMMARY result format, so SUMMARY is no longer more limited than BASIC. (#11594)DeprecationWarnings emitted by dependencies so local test runs are not failed by them. (#11587)map_expectation_unexpected_rows_as_dict Checkpoint setting that serializes unexpected_rows as dictionaries for all Map expectations on SQLAlchemy and Spark, with the default output unchanged. (#11583)invoke deps accepts --pty and --no-pty flags so automated environments can control pseudo-terminal usage. (#11586)pyparsing deprecation warnings via the compatibility layer, and documented the type-checking workflow so local runs match CI. (#11574)</details>
Thanks to @leodrivera (first contribution).
[FEATURE] Support query and PK columns on BOOLEAN_ONLY and BASIC
Unexpected-index columns and unexpected queries on BOOLEAN_ONLY and BASIC result formats — Result formats BOOLEAN_ONLY and BASIC now support returning primary-key/unexpected-index columns and the unexpected-rows query, so you can identify failing rows without switching to a more verbose result format. (#11563)
result = batch.validate(
expectation,
result_format={
"result_format": "BASIC",
"unexpected_index_column_names": ["pk_1"],
},
)
<details> <summary>Maintenance</summary>
</details>
[FEATURE] infer primary keys during column_types metric fetch
Primary key information in column type metrics — The table.column_types metric now reports whether each column is part of the table's primary key when using a SQL (SQLAlchemy) execution engine. Single-column, composite, and quoted primary keys are all detected, and columns in tables without a primary key are reported as not primary keys. (#11554)
# Each entry in the metric value now includes a `primary_key` flag:
# [
# {"name": "id", "type": "UUID", "primary_key": True},
# {"name": "created_at", "type": "TIMESTAMP WITH TIME ZONE", "primary_key": False},
# ]
Oracle query assets no longer get an unwanted FROM DUAL clause — Querying Oracle data sources through SQLAlchemy no longer appends a spurious FROM DUAL clause to an already well-formed SQL query, so query assets against Oracle run as written. Verified against Oracle 19c and PostgreSQL 10.16. (#11538)
table.column_types metric now includes a primary_key flag for each column when read through a SQL execution engine, covering single-column, composite, and quoted primary keys. (#11554)FROM DUAL clause when the query is already properly formatted. (#11538)<details> <summary>Maintenance</summary>
expect_query_results_to_match_comparison again include columns whose values are null, so column headers line up with the data. (#11548)invoke command is found during the docs build step. (#11545)</details>
Thanks to @konnor-b (first contribution).
[DOCS] validations with the Cloud API
Fluent Snowflake datasource update methods — Snowflake datasources can now be updated or upserted through the fluent API with update_snowflake and add_or_update_snowflake, matching the methods already available for other datasource types. (#11520)
context.data_sources.add_or_update_snowflake(
name="my_snowflake_ds",
connection_string="snowflake://<user>@<account>/<database>/<schema>?warehouse=<wh>&role=<role>",
)
Documentation for running validations with the GX Cloud API — The documentation now explains how to run validations using the GX Cloud API, including what is supported and how results are handled. (#11400)
Documentation for Custom Actions in GX Cloud — New documentation describes how to configure and use Custom Actions in GX Cloud. (#11521)
Dependency compatibility reference expanded — The compatibility reference in the docs now lists supported dependencies, so you can check which versions work with your installation before upgrading. (#11530)
private_key inside a Snowflake datasource's kwargs is deprecated; use the datasource's dedicated private_key connection argument. Removal in 2.0.0. (#11520)<details> <summary>Maintenance</summary>
private_key through kwargs now emits a deprecation warning, and the fluent API gained update_snowflake and add_or_update_snowflake. (#11520)</details>
[MAINTENANCE] Deprecate string-style row_conditions
row_conditions (#11515)row_condition parameter on expectations is deprecated; use Condition objects, such as Column("age") > 18. Removal in 2.0.0. (#11515)condition_parser parameter on expectations is deprecated; use Condition objects, such as Column("age") > 18. Removal in 2.0.0. (#11515)<details> <summary>Maintenance</summary>
row_condition, or supplying condition_parser, now raises a DeprecationWarning pointing to Condition objects (for example Column("age") > 18) instead. (#11515)</details>
[MAINTENANCE] Ignore boto warning about deprecating python 3.9 support
Column positional (#11497)ComparisonCondition with in/not in operators (#11494)bool values in Column.is_in( ) / Column.is_not_in() (#11500)Column import (#11506)row_conditions (#11515)unexpected_index_column_names in ExpectColumnValuesToNotBeNull results (#11513) (thanks @chay0112)Compatibility: Python <3.14,>=3.9 → <3.14,>=3.10; numpy removed (python_version == "3.9"); pandas removed (python_version == "3.9")
Row conditions: documented, importable, and rendered in Data Docs — Row conditions are now a supported way to scope an Expectation to a subset of rows. The condition classes (including Column and the comparison, nullity, and boolean conditions) are part of the public API and are imported from great_expectations.expectations.row_conditions; the previous great_expectations.expectations.conditions import path still works. Conditions are validated more strictly (an in/not-in parameter must be an iterable whose members share a single type, and boolean members are rejected), a bare string condition is always turned into a condition object even when no condition parser is given, and Data Docs now renders every condition when an Expectation carries more than one. New documentation pages, screenshots, and notes on the minimum GX Cloud API and agent versions required for certain row-condition features round this out. (#11478, #11494, #11497, #11500, #11504, #11506, #11507, #11509, #11511, #11512)
from great_expectations.expectations.row_conditions import Column
condition = Column("age") > 21
Python 3.10 is now the minimum supported version — Great Expectations no longer supports Python 3.9. Install on Python 3.10 or newer; the documentation now states 3.10 as the minimum. (#11501, #11485)
unexpected_index_column_names returned by ExpectColumnValuesToNotBeNull — When you request unexpected index columns in result_format, ExpectColumnValuesToNotBeNull now includes unexpected_index_column_names in its validation result, matching the other column-value Expectations. (#11513)
result_format={"result_format": "COMPLETE", "unexpected_index_column_names": ["customer_id"]}
unexpected_index_column_names in its validation results when they are requested through result_format. (#11513)<details> <summary>Maintenance</summary>
Column class and its siblings are now imported from great_expectations.expectations.row_conditions; the previous great_expectations.expectations.conditions import path continues to work. (#11506)Column.is_in() and Column.is_not_in(). (#11500)Column class now takes its column name as a positional argument, so it can be constructed as Column("age"). (#11497)</details>
Thanks to @chay0112 (first contribution).
[DOCS] snowflake password deprecation
condition_parser field intact for backwards compatibility (#11484)row_condition groups (#11488)None parameter in Conditions (#11491)Null checks in row conditions — Row conditions now express null comparisons explicitly with is_null() and is_not_null() on a column, and passing None as the value of a comparison operator is rejected instead of silently producing an invalid condition. (#11491)
import great_expectations.expectations as gxe
from great_expectations.core.expectation_condition import Column
gxe.ExpectColumnValuesToBeBetween(
column="amount",
min_value=0,
row_condition=Column("cancelled_at").is_null(),
)
Legacy row condition strings keep working alongside condition objects — Existing string-based row_condition values are accepted and converted into the new condition objects, the condition_parser field is preserved, pandas and Spark conditions have a passthrough path, and rendered expectation content displays the new condition types correctly. (#11474, #11484, #11480, #11481)
Clearer limits on combining row conditions — Nested AndConditions are flattened automatically, while OrConditions nested inside AndConditions or other OrConditions now raise an explicit error, as does supplying more than 100 conditions. (#11488)
Updated Snowflake connection documentation — The Snowflake documentation now covers the deprecation of password authentication and gives corrected guidance for configuring private key authentication. (#11416, #11490)
google.api_core Python 3.10 end-of-life warning. (#11493)<details> <summary>Maintenance</summary>
None as the value of a column comparison in a row condition is now rejected; use is_null() or is_not_null() for null checks, and the comparison parameter is required. (#11491)AndConditions are flattened, OrConditions nested inside AndConditions or OrConditions raise an error, and more than 100 conditions raises an error. (#11488)condition_parser field is retained for backwards compatibility so single-condition row conditions still convert to string syntax on older API versions. (#11484)row_condition values are transformed into the new condition objects. (#11474)</details>
[FEATURE] Snowflake Key Pair Auth API
Snowflake key pair authentication — Snowflake data sources now accept key pair authentication as a first-class part of the connection API, so you can configure a Snowflake data source with a private key instead of a password. (#11395)
context.data_sources.add_snowflake(
name="my_snowflake",
connection_details={
"account": "myOrg-my_account",
"user": "my_user",
"database": "my_db",
"schema": "my_schema",
"warehouse": "my_wh",
"role": "my_role",
"private_key": "<PEM-encoded private key>",
},
)
Row conditions are honored by Volume Expectations — Volume Expectations now apply the configured row condition, so expected row counts are evaluated against the filtered rows rather than the whole batch. (#11467)
Documented schema handling in Redshift and PostgreSQL connection strings — The connection documentation now explains how to specify a schema in Redshift and PostgreSQL connection strings. (#11433)
GX Cloud Data Health documentation for failed Expectations — New GX Cloud documentation covers the Data Health view for failed Expectations, with screenshots refreshed to match the current interface, alongside new GX Cloud architecture supporting content. (#11419, #11458, #11439)
<details> <summary>Maintenance</summary>
</details>
[BUGFIX] Fix ExpectColumnValuesToBeOfType for trino
unexpected_index_query (#11437)RedshiftConnectionDetails to type stub (#11434)Databricks SQL parameters are now compiled in unexpected_index_query — Validation results for Databricks now return an unexpected_index_query with its parameters fully rendered, so the query can be copied and run as-is. The query compilation is also no longer sensitive to unfamiliar bind-parameter patterns or to the ordering of parameter values. (#11437)
ExpectColumnValuesToBeOfType works against Trino — ExpectColumnValuesToBeOfType now evaluates correctly when validating data in Trino. (#11438)
import great_expectations as gx
gx.expectations.ExpectColumnValuesToBeOfType(column="id", type_="INTEGER")
unexpected_index_query, so the returned query is complete and runnable, and no longer depends on bind-parameter naming patterns or dictionary ordering. (#11437)ExpectColumnValuesToBeOfType so it evaluates correctly against Trino. (#11438)<details> <summary>Maintenance</summary>
</details>
[MINORBUMP] Remove Pandas Upper Bound Constraint
PandasS3Datasource using boto3_options (#11412)Compatibility: Python <3.13,>=3.9 → <3.14,>=3.9; numpy added (python_version >= "3.13"); pandas added (python_version >= "3.13"); posthog removed; pandas removed (extra snowflake) (python_version >= "3.9")
Python 3.13 support — Great Expectations now installs and runs on Python 3.13, in addition to the previously supported 3.9 through 3.12. (#11426)
Works with pandas 2.2 and newer — The <2.2 upper bound on pandas has been removed, so you can install Great Expectations alongside pandas 2.2.0 and later and pick up the newest pandas features and fixes. (#11423)
pip install great_expectations "pandas>=2.2"
Usage analytics removed — Great Expectations no longer collects or sends usage analytics, and the posthog dependency is no longer installed with the library. (#11420)
Reassigning a Snowflake connection string now works as expected — Setting a new connection string on an existing SQL data source — including Snowflake — is now converted to the proper connection type, so the data source stays usable after the reassignment. (#11410)
datasource.connection_string = "snowflake://user:password@account/db/schema?warehouse=wh&role=role"
Renderer class is no longer part of the public API. (#10866)<2.2 upper bound on pandas so Great Expectations can be used with pandas 2.2.0 and above. (#11423)PandasS3Datasource through boto3_options rather than environment variables. (#11412)<details> <summary>Maintenance</summary>
schema field, so a schema can be supplied separately when configuring a Redshift connection. (#11431)posthog dependency. (#11420)</details>
[MAINTENANCE] Run Athena tests as a separate step
<details> <summary>Maintenance</summary>
</details>
[DOCS] integration point diagrams
Compatibility: new extra test
Tutorial for validating unstructured data in GX Cloud — A new tutorial walks through validating unstructured data in GX Cloud end to end. (#11380)
Documentation for severity tagging — The GX Cloud documentation now covers severity tagging, including refreshed screenshots that match the current UI, plus new diagrams illustrating GX integration points. (#11354, #11394, #11391)
<details> <summary>Maintenance</summary>
column.non_null_count to the recognized metric types. (#11397)</details>
[BUGFIX] Fix ExpectColumnValuesToBeInTypeList for Trino
Compatibility: pyarrow removed (extra arrow); new extra arrow; new extra snowflake; removed extra snowflake; pyarrow removed (extra test); new extra test
Expectation reference documentation now describes the severity parameter — Every Expectation type's reference documentation now lists severity under "Other Parameters", with a link to the severity documentation, so you can see how to set failure severity directly from the Expectation reference. (#11387)
Type-list validation works against Trino — ExpectColumnValuesToBeInTypeList now compares column types correctly when validating data through the Trino dialect, instead of misreporting matching types. (#11386)
Documentation for workspaces — The documentation site now covers workspaces. (#11366)
ExpectColumnValuesToBeInTypeList now handles type comparisons correctly for the Trino dialect. (#11386)severity description with a documentation link to the "Other Parameters" section of every Expectation type, and capitalized "Expectation" in the FailureSeverity description. (#11387)<details> <summary>Maintenance</summary>
pyarrow>=14 for Python 3.12 in the development arrow requirements so Snowflake marker test jobs install a prebuilt wheel instead of failing to build from source; no runtime behavior changes. (#11388)</details>
[BUGFIX] Make workspaces optional for cloud_user_info
[FEATURE] Make GX Context workspace aware
unexpected_rows (#11368)workspace_id to store_backend dict (#11371)GX Cloud workspace awareness — Data Contexts are now workspace aware, laying the groundwork for GX Cloud's multi-workspace support. A workspace id supplied to a Cloud context is carried through to Cloud requests and to the credentials used by its stores. (#11369, #11371, #11373)
# GX_CLOUD_WORKSPACE_ID is read alongside your Cloud access token and organization id
import great_expectations as gx
context = gx.get_context(mode="cloud")
S3 data assets read past the first page of results — Listing files in an S3 directory or bucket with more results than fit in a single response now returns all of them instead of raising an error part-way through. (#11361)
More robust handling of quoted and mixed-case SQL identifiers — Schema and table names that are quoted, use data-source-specific quote characters, or use mixed case are now handled correctly when building queries and when collecting column metadata. (#11367, #11365)
<details> <summary>Maintenance</summary>
</details>
Thanks to @pawel99k (first contribution).
[FEATURE] Checkpoint actions notify on severity
Severity-aware Checkpoint notifications — Expectations that carry a severity value can now be validated and acted on end to end: validation results expose the highest-severity failure they contain, and built-in Checkpoint actions use it to decide whether to notify. (#11341, #11343, #11347)
result = checkpoint.run()
validation_result = result.run_results[next(iter(result.run_results))]
max_severity = validation_result.get_max_severity_failure()
Quoted table names stay quoted in GX Cloud — A table asset whose table name is quoted keeps its quoting when it is sent to and fetched back from GX Cloud, so the name continues to be treated as quoted. (#11357)
<details> <summary>Maintenance</summary>
</details>
[DOCS] Completeness anomaly detection now uses forecasted range
<details> <summary>Maintenance</summary>
</details>
[MAINTENANCE] Fix webpack-dev-server and form-data vulnerabilities
Expectation JSON schemas now carry failure severity — Expectations can express a failure severity, and the published JSON schemas now include the new severity field backed by a FailureSeverity enum. (#11337)
Documentation for Cloud API version 0.18 sunset — The compatibility reference and related documentation now reflect the sunset of Cloud API version 0.18. (#11334)
severity field, with a new FailureSeverity enum describing Expectation failure severity. (#11337)<details> <summary>Maintenance</summary>
</details>
[BUGFIX] Handle SQL parameter limit for Databricks
Validations against Databricks no longer fail on large bundled metric queries — Metric queries that exceed Databricks' 256 query-parameter limit are now split into smaller batches automatically, so validating batches with many parameters against Databricks completes instead of erroring. (#11317)
Documentation for retrying Expectation generation with your own input — The documentation now describes the retry workflows for supplying user input when generating Expectations, so you can guide generation when the first attempt isn't what you wanted. (#11325)
<details> <summary>Maintenance</summary>
</details>
Thanks to @Abdelkrim (first contribution).
[BUGFIX] change pyspark column reference from DataFrame.__getitem__ to F.col() (#11286) (thanks @alansk97)
<details> <summary>Maintenance</summary>
</details>
Thanks to @alansk97 (first contribution).
[BUGFIX] Remove incompatible min/max types for some range expectations
New documentation for Data Health, SQL generation, and pipeline architecture — The docs now cover the Data Health dashboard (including a screenshot of it), generating SQL, the newly supported data sources, and a pipeline architecture diagram. (#11294, #11307, #11289, #11293, #11298)
Range expectations reject incompatible date and datetime bounds — ExpectColumnUniqueValueCountToBeBetween, ExpectColumnStdevToBeBetween, and ExpectColumnValueLengthsToBeBetween no longer accept date or datetime values for their min and max inputs, so these expectations now only allow bounds that make sense for the value they measure. (#11305)
<details> <summary>Maintenance</summary>
</details>
[FEATURE] Add new Postgres "flavor" Data Source classes
BigQuery data source — A dedicated BigQuery data source class is now available, so BigQuery connections can be declared as their own data source type rather than as a generic SQL connection. (#11296)
Postgres-compatible data source flavors — New data source classes cover Postgres-compatible services — Google Cloud AlloyDB, Amazon Aurora, Citus, and Neon — so each of these backends can be selected directly when connecting to data. (#11290)
Disabling analytics is now fully respected — When analytics is disabled in the Data Context configuration, analytics initialization is no longer performed at all. This resolves permission-denied errors raised while looking for a user-level configuration file in restricted environments such as Databricks streaming jobs. (#11276)
<details> <summary>Maintenance</summary>
</details>
Thanks to @jmcorreia.
[BUGFIX] B/gx 1174/generalize schema expectation
Corrected result summary for ExpectTableColumnsToMatchSet — Validation results for ExpectTableColumnsToMatchSet now render correctly, so the expectation's summary reads accurately wherever results are displayed. (#11281)
New documentation for Anomaly Detection expectations — The docs now cover the Anomaly Detection expectation drawer and its underlying model, so you can understand how anomaly detection expectations are configured and how they behave. (#11234)
<details> <summary>Maintenance</summary>
</details>
[BUGFIX] Make ExpectTableColumnsToMatchSet case insensitive
ExpectTableColumnsToMatchSet now matches column names case-insensitively — On SQL dialects where column names are compared case-insensitively (PostgreSQL, Databricks SQL, and Snowflake), ExpectTableColumnsToMatchSet no longer fails when the expected column set differs only by letter casing from the table's actual columns. It now behaves consistently with the other column-name expectations such as expect_table_columns_to_match_ordered_list and expect_column_to_exist. (#11266)
import great_expectations.expectations as gxe
# Passes against a table whose columns are PASSENGER_COUNT and TRIP_DISTANCE
suite.add_expectation(
gxe.ExpectTableColumnsToMatchSet(column_set=["passenger_count", "trip_distance"])
)
<details> <summary>Maintenance</summary>
</details>
[FEATURE] Add SuiteParameterDict to all Expectation Kwargs (#11222) (thanks @Pascal06S)
ValidationError in ExpectColumnPairValuesToHaveDifferenceOfCustomPercentage (#11209) (thanks @sariaslaso)add_batch_definition_whole_directory only loading 1 file (#11254)min_value and max_value parameters (#11259)Suite parameters accepted in every expectation argument — All expectation arguments now accept suite parameters, so any keyword argument of an expectation can be supplied at validation time instead of being hard-coded when the expectation is defined. (#11222)
import great_expectations as gx
expectation = gx.expectations.ExpectColumnValuesToBeBetween(
column="passenger_count",
min_value={"$PARAMETER": "min_passengers"},
max_value={"$PARAMETER": "max_passengers"},
)
results = batch.validate(
expectation,
expectation_parameters={"min_passengers": 1, "max_passengers": 6},
)
Whole-directory batch definitions read every file again — Batch definitions created with add_batch_definition_whole_directory on S3, Azure Blob Storage, and Google Cloud Storage data assets now read all files in the directory instead of only one. (#11254)
asset = data_source.add_directory_csv_asset(name="my_asset", s3_prefix="data/")
batch_definition = asset.add_batch_definition_whole_directory("all_files")
batch = batch_definition.get_batch()
min_value and max_value parameters of ExpectColumnProportionOfNonNullValuesToBeBetween and ExpectColumnProportionOfUniqueValuesToBeBetween no longer accept date or datetime values; they now accept only numbers (or a suite parameter), and their published schemas reflect this. (#11259)add_batch_definition_whole_directory reading only a single file for S3, Azure Blob Storage, and Google Cloud Storage data assets; the whole directory is now read as one batch. (#11254)ValidationError raised when using ExpectColumnPairValuesToHaveDifferenceOfCustomPercentage; the expectation now declares its required percentage argument. (#11209)<details> <summary>Maintenance</summary>
brace-expansion documentation-site dependency from 1.1.11 to 1.1.12. (#11251)</details>
Thanks to @Pascal06S (first contribution), @sariaslaso (first contribution).
[FEATURE] Add ColumnAggregateNonNullCount metric
ExpectColumnProportionOfUniqueValuesToBeBetween (#11235)New expectation: ExpectColumnProportionOfUniqueValuesToBeBetween — You can now assert that the proportion of unique values in a column falls within an expected range, letting you catch columns that become unexpectedly duplicated or unexpectedly high-cardinality. (#11235)
import great_expectations.expectations as gxe
expectation = gxe.ExpectColumnProportionOfUniqueValuesToBeBetween(
column="passenger_count",
min_value=0.1,
max_value=0.9,
)
Non-null count available as a column metric — A column aggregate metric for the number of non-null values in a column is now available, so expectations and custom checks can reason about how much data a column actually contains. (#11229)
ExpectColumnProportionOfUniqueValuesToBeBetween, which validates that the proportion of unique values in a column falls between a minimum and maximum value. (#11235)ColumnAggregateNonNullCount metric that reports the number of non-null values in a column. (#11229)<details> <summary>Maintenance</summary>
--snowflake flag, so those tests run again in CI via the snowflake marker. (#11230)</details>
[MINORBUMP] docs for Multi-source Expectations
ExpectQueryResultsToMatchComparison docstring (#11221)pkg_resources dependency (#11213)Multi-source Expectations documentation — The documentation now covers Multi-source Expectations, explaining how to compare data across two different data sources. (#11165)
Redshift geometry and super column types supported — Redshift data sources now recognize the GEOMETRY and SUPER column types, so assets containing these columns can be introspected and validated. (#11194)
No more pkg_resources dependency — GX Core no longer depends on the deprecated pkg_resources package, removing its import-time deprecation warnings on modern Python installs. (#11213)
GEOMETRY and SUPER column types. (#11194)<details> <summary>Maintenance</summary>
min_value and max_value in ExpectColumnMaxToBeBetween. (#11225)pkg_resources dependency, replacing requirements parsing with a pip compatibility module and a self-contained parser in setup.py. (#11213)ExpectQueryResultsToMatchComparison docstring so parameter names and descriptions read consistently. (#11221)ExpectQueryResultsToMatchComparison. (#11216)</details>
Thanks to @VolkovGeoPhy.
Yanked reason: The version accidently was set to 1.4.7 instead of 1.5.0
Jun 5, 2025 2 files
Yanked reason: The version accidently was set to 1.4.7 instead of 1.5.0
[MAINTENANCE] limit pyspark to <4.0 due to breaking changes in types
ExpectQueryResultsToMatchComparison (#11203)Clearer errors for unhashable column types in ExpectQueryResultsToMatchComparison — When a query returns unhashable data types such as JSONB, ExpectQueryResultsToMatchComparison now raises a helpful error that names the first column containing unhashable data instead of failing with an unclear message. (#11193)
Case-insensitive column type checks on Databricks, Snowflake, and Postgres — expect_column_values_to_be_of_type now treats unquoted identifiers in column_name and column_type as case-insensitive on Databricks, Postgres, and Snowflake, so type expectations pass regardless of the casing you write. (#11192)
suite.add_expectation(
gxe.ExpectColumnValuesToBeOfType(column="my_column", type_="varchar")
)
Documentation for anomaly detection — The documentation now covers anomaly detection, alongside refreshed wording for the terms "Core" and "platform" and updated guidance noting that both tables and views are supported as data assets. (#11172, #11187, #11198, #11205)
<details> <summary>Maintenance</summary>
</details>
[DOCS] ExpectAI for all Data Sources
Redshift GEOMETRY and SUPER column types supported — Redshift data sources now recognize the GEOMETRY and SUPER column types, so assets using those columns can be used without an unsupported-type error. (#11183)
ExpectAI documentation now covers all Data Sources — The ExpectAI documentation has been reorganized so it applies to every supported Data Source rather than a subset. (#11178)
Fewer secret-store lookups when resolving config secrets — Secret substitution now reuses a cached secrets store client instead of rebuilding it on every lookup, avoiding repeated calls to the secrets backend when loading configuration. (#11184)
<details> <summary>Maintenance</summary>
</details>
Thanks to @VolkovGeoPhy.
[MAINTENANCE] Resolve datetime deprecation warnings (#11134) (thanks @emmanuel-ferdman)
ExpectQueryResultsToMatchSource docstring (#11158)ExpectQueryResultsToMatchSource (#11160)ExpectQueryResultsToMatchSource DQI (#11164)ExpectQueryResultsToMatchSource (#11174)New Expectation: ExpectQueryResultsToMatchSource — You can now compare the results of a SQL query run against your Data Source with the results of a query run against another Data Source, and require that at least a mostly fraction of records match. Supported on PostgreSQL, Snowflake, Databricks (SQL), Redshift, and SQLite. (#11144)
import great_expectations as gx
expectation = gx.expectations.ExpectQueryResultsToMatchSource(
target_query="SELECT id, amount FROM orders",
source_data_source_name="my_source_data_source",
source_query="SELECT id, amount FROM orders",
mostly=0.95,
)
Richer results and rendering for ExpectQueryResultsToMatchSource — Validation results for ExpectQueryResultsToMatchSource now report the specific rows missing from or unexpected in the target query results, and those differences are presented as a diagnostic table — with a simplified presentation when the source and target queries each return a single column. The Expectation also renders a readable summary showing the target query and the source Data Source it is compared against. (#11161, #11168, #11173, #11160)
mostly fraction of records match. (#11144)<details> <summary>Maintenance</summary>
</details>
Thanks to @esadek (first contribution), @emmanuel-ferdman (first contribution).
[FEATURE] Test infra to support source to target expectations
SupportedDataSources enum (#11143)Redshift data source support in the public API — Redshift data sources are now exposed through the public API decorator, and new documentation walks through connecting Great Expectations Cloud to Redshift. (#11097, #11095)
QueryDataSourceTable metric and provider — A new QueryDataSourceTable metric and its provider are available, enabling queries against a data source table as part of metric computation. (#11149)
ExpectAI approval workflow documentation — The Cloud documentation now describes the ExpectAI approval workflow for generating and approving Expectations. (#11072)
<details> <summary>Maintenance</summary>
mostly parameter description shown in Expectation docstrings and schemas. (#11147)</details>
[FEATURE] Changes to Redshift connection_string validator to support a dict type
connection_string validator to support a dict type (#11119)MetricErrorResult and Batch.compute_metrics() API typing (#11127)Redshift connection strings can be supplied as a dictionary — When adding a Redshift datasource, connection_string may now be given as a dictionary of connection components in addition to a string URL. (#11119)
import great_expectations as gx
context = gx.get_context()
datasource = context.data_sources.add_redshift(
name="my_redshift",
connection_string={
"drivername": "redshift+psycopg2",
"username": "my_user",
"password": "my_password",
"host": "my-cluster.redshift.amazonaws.com",
"port": 5439,
"database": "my_database",
},
)
Expectation coverage for Redshift assets — Expectations running against Redshift assets are now verified to the same level as Postgres, so Redshift users can rely on the same set of expectations behaving as documented. (#11128)
More type information shipped with the package — The published distribution now exposes more of the library's type information, so type checkers resolve Great Expectations types in your own code more completely. (#11115)
connection_string provided as a dictionary in addition to a string. (#11119)<details> <summary>Maintenance</summary>
Batch.compute_metrics() is now typed to include MetricErrorResult, and metric error types were simplified and consolidated. (#11127)</details>
[FEATURE] Allow user to provide connection details to connect to Redshift
ColumnDescriptiveStats metric (#11108)New ColumnDescriptiveStats metric — You can now compute a column's minimum, maximum, mean, and standard deviation in a single metric with ColumnDescriptiveStats, available on the pandas, SQL, and Spark backends. (#11108, #11109)
from great_expectations.metrics import ColumnDescriptiveStats
result = batch.compute_metrics(ColumnDescriptiveStats(column="passenger_count"))
print(result.value.min, result.value.max, result.value.mean, result.value.standard_deviation)
New ColumnValuesNotMatchRegexCount metric — You can now count the values in a column that do not match a regular expression with ColumnValuesNotMatchRegexCount, available on the pandas, SQL, and Spark backends. (#11103)
from great_expectations.metrics import ColumnValuesNotMatchRegexCount
result = batch.compute_metrics(
ColumnValuesNotMatchRegexCount(column="vendor_id", regex="^(a|d).+")
)
print(result.value)
Connect to Redshift with connection details — A Redshift data source can now be configured by supplying individual connection details instead of a full connection string. (#11105)
Redshift schema introspection no longer raises a TypeError — Using the gx-redshift extra to introspect schema information, such as computing column descriptive metrics, no longer fails with a runtime TypeError. (#11112)
ColumnDescriptiveStats metric, which returns a column's minimum, maximum, mean, and standard deviation on pandas, SQL, and Spark backends. (#11108)connection_string. (#11105)ColumnValuesNotMatchRegexCount metric, which counts column values that do not match a given regular expression on pandas, SQL, and Spark backends. (#11103)TypeError when using the gx-redshift extra to perform schema introspection, such as computing column descriptive metrics. (#11112)ExpectColumnValuesToBeBetween now correctly rejects configurations where both min_value and max_value are omitted, None, or empty strings. (#11102)MicrosoftTeamsNotificationAction failing with a 400 Bad Request when sending notifications. (#11106)<details> <summary>Maintenance</summary>
great_expectations.metrics under consistent module names. (#11109)</details>
Thanks to @jwalant-dattani (first contribution).
Your coding agent can read these notes before it upgrades. Set up the MCP server →