NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1519 most downloaded on PyPI
Data Quality eXtended (DQX) is a Python library for data quality checks and data quality monitoring
Last release 1 months ago
13 Aug 2026
Ships fairly regularly
a new release about every 4 weeks
Nearly every release is documented
notes for 32 of 32 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
32 releases · first in 2025
One column per month.
…), matching the File/Volume backends. See Breaking Changes for the user_metadata at-rest encoding change on the Delta backend.
DQMetricsObserver. The built-in DQAlert action can send notifications to Slack, Microsoft Teams, a generic HTTPS webhook, or the log, so pipelines can react to data quality regressions without custom plumbing. You can create your own custom actions as well, and custom alerting is possible via the callback destination, which invokes an in-process Python callable for each alert.DQEngine.compute_summary_metrics(...) produces the same row counts, per-check breakdown, and custom observer metrics as a lazy aggregation over the results DataFrame, so metrics can be computed inside Spark Declarative Pipelines where the observer- and streaming-listener-based paths cannot be used.aggr_matches_dataset dataset-level check (#1309). The check compares an aggregate metric (row count by default, or any curated/built-in aggregate) computed on the checked DataFrame against the same aggregate computed on a reference (upstream) table, enabling reconciliation-style validations against a source of truth.has_no_gaps_per_time_window dataset-level check (#1370). It detects gaps in a time series — windows of a configurable size that contain no rows between windows that do — with optional grouping and trailing-gap handling.ChecksSemanticValidator inspects declarative check metadata and reports duplicate rules (same function, arguments, criticality, and filter) and conflicting rules, surfacing authoring mistakes before checks run.has_valid_string_case row-level check (#1347). It validates consistent string casing with upper, lower, title, and sentence modes, casting non-string columns to strings before comparison.is_valid_national_id row-level check (#1346). It validates national identification numbers per country (default US), covering format, ranges, and obvious structural errors; it does not verify that a number was actually issued.is_valid_currency_code row-level check (#1368). It validates values against ISO 4217 currency codes, supporting both the three-letter alphabetic (e.g. USD) and three-digit numeric (e.g. 840) representations via code_format.is_valid_country_code row-level check (#1369). It validates values against ISO 3166-1 country codes in alpha-2 (default), alpha-3, or numeric form via code_format.is_valid_language_code row-level check (#1403). It validates values against ISO 639 language codes in alpha-2 (ISO 639-1) or alpha-3 (ISO 639-3) form.is_valid_subdivision_code row-level check (#1404). It validates values against ISO 3166-2 country subdivision codes (e.g. US-CA, GB-ENG), with optional cross-column country consistency via country_column.is_valid_uuid row-level check (#1436). It validates values against the canonical RFC 9562 UUID string form (case-insensitive), mirroring the other pure pattern-match checks.has_no_outliers check (#1317). The profiler can now generate a has_no_outliers check, and the MAD-based calculations and profiler defaults were refactored into shared constants. Disabled by default to retain existing performance.pydantic ValidationError — every entry point still raises DQX error types with the pre-migration message format.message_expr and typed user_metadata), matching the File/Volume backends. See Breaking Changes for the user_metadata at-rest encoding change on the Delta backend.sql_query rules against unsafe SQL (#1275). Both LLM-assisted rule-generation paths now drop any generated sql_query rule whose query contains unsafe (DML/DDL) SQL before returning it to the caller.filter and row_filter against unsafe SQL (#1303). All filter compile sites now route through a shared safe_filter_expr helper that rejects destructive SQL keywords, and a check with an unsafe filter is treated as skipped rather than failing the run.check_funcs.py, making any reordering that would break positional callers visible during review.fr) locale (#1330). French joins the existing English, Brazilian Portuguese, Italian, and Spanish translations.make mcp-deploy is a single end-to-end command.log_telemetry no longer makes a blocking per-check control-plane call that could raise a TimeoutError and terminate Structured Streaming jobs; telemetry is now non-throwing, deduplicated per process, and bounded by a short timeout.install_logger(), which previously removed existing root handlers and overwrote the logging configuration of applications using DQX as a library.is_unique violations (#1442). is_unique now requires the current row to match its filter before reporting a duplicate, so unfiltered rows sharing a key with filtered rows are no longer falsely flagged.* combined with a row filter (#1453). Dataset-level aggregate checks that aggregate over * with a row_filter no longer raise INVALID_USAGE_OF_STAR_OR_REGEX when constructed with F.col("*"); unfiltered count/count_distinct over * continue to work, and unsupported star/aggregate combinations now raise a clear InvalidParameterError.is_valid_email, is_valid_ipv4_address, is_valid_uuid, is_valid_national_id, and the is_ipv4_address_in_cidr value path now reject values with a trailing newline, which Java/Spark rlike previously accepted because $ also matches before a final line terminator.has_no_gaps_per_time_window now preserves gap violations for groups with null key components.Customer Name) now validate via a two-pass fallback.TimestampType and TimestampNTZType.build_app.py now appends the .cmd suffix to Node binaries on Windows, and CI sweeps orphaned jobs.set_utc_timezone test fixture to actually apply UTC (#1402).BREAKING CHANGES!
is_in_list, is_not_in_list, and is_not_null_and_is_in_list now resolve their allowed / forbidden string values as column expressions (consistent with the comparison checks), not string literals. A bare string is interpreted as a column reference, a numeric string (e.g. "3") is parsed as a number, and an ISO-date string (e.g. "2024-01-01") as a date. To match a string literal, single-quote the value (e.g. 'value') or wrap it in F.lit("value"). Existing checks that relied on bare strings being treated as literals must quote them. (#1419)user_metadata saved through the Delta table storage backend is now JSON-encoded at rest to preserve non-string types through the MAP<STRING, STRING> column. Save→load via DQX is transparent (you get the original typed value back), but the stored representation changes: direct SQL/dashboard consumers now read JSON-encoded values (decode with from_json), existing tables are not migrated, and legacy string values that look like JSON atoms ("true", "1", "null") read back as typed values (True / 1 / None) — re-save affected rule sets after upgrading to normalize. The File/Volume (YAML/JSON) and Lakebase (JSONB) backends are unaffected. (#1319)Full Changelog: v0.15.0...v0.16.0
@mwojtyczka, @ghanse, @vb-dbrks, @SreeramaYeshwanthGowd, @mattfaltyn, @fedeflowers, @aarushisingh04, @IvannKurchenko, @abhyuday1203, @arnoN7, @AtomicGlance, @berrybluecode, @laurencewells, @neeraj-bhadani-08, @SaptarshiAcharyya99, @souravg-db2, @STEFANOVIVAS, @SyedIshmumAhnaf, @Vsatyam013
Added LLM-generated AI explanations for row-level anomaly detection ( #1129 ). The has_no_row_anomalies check now attaches a plain-language ai_explana
has_no_row_anomalies check now attaches a plain-language ai_explanation to each flagged row under _dq_info[].anomaly, describing the likely cause, business impact, suggested action, the top contributing features, and the group's size and average severity. Explanations are generated vis Spark ai_query function against a Databricks Model Serving endpoint — no extra Python dependencies and no driver-side LLM calls — and anomalous rows are grouped by segment and top contributing features so the model is called once per group, keeping cost predictable on large datasets (bounded by max_groups). AI explanations are enabled by default and does not require additional settings. New parameters with good set of defaults include enable_ai_explanation, ai_explanation_llm_model_config, redact_columns (to keep sensitive columns out of the prompt and grouping), and max_groups. If the serving endpoint is unreachable, explanations are left null with a warning and scoring still succeeds. LLMModelConfig also gains max_tokens, temperature, timeout, and max_retries to bound LLM cost and latency and expose tuning parameters for the users if required.sample_by option to perform stratified sampling based on column values. Users control the sampling fraction with either a single sample_fraction applied equally across all strata, or a dictionary mapping each stratum to its own fraction. When sample_by is omitted, the profiler continues to use uniform sampling across all rows.is_valid_email (#1158). A new is_valid_email check validates email addresses against a pragmatic, ReDoS-safe subset of RFC 5321/5322. Like the IP-address checks, it ignores null values (no violation reported).is_geo_contains, is_geo_covers, is_geo_intersects, is_geo_touches, and is_geo_within. By default they use exact, meter-level precision built on the ST_* family of functions; is_geo_covers and is_geo_intersects additionally support an approximate mode built on H3_* cell indexing with a configurable resolution for faster checks on large datasets. The reference geometry can be a literal WKT/WKB/EWKT/EWKB value or another column, with optional try_to_geometry conversion of either side. Running these checks requires Databricks serverless compute or runtime 17.1 or above.save_results_in_table and the corresponding workflow path can now persist summary metrics without requiring an output or quarantine table, supporting observability-focused pipelines that only need the metrics table. Batch observations are triggered before metrics are saved so the metrics table is populated correctly, and streaming and no-observer cases now raise explicit errors. Existing configurations with an output or quarantine table are unaffected.DQRule now accepts an optional message_expr parameter that lets users define custom failure messages as a Spark Column or a SQL expression string. The same option is supported for checks defined declaratively in metadata (YAML/JSON), specified as a top-level message_expr key on the check definition alongside criticality and check. When omitted, the default message behavior is preserved; when provided, the custom message replaces the default message for failed rows.name now store the same autogenerated name and name-inclusive rule_fingerprint that apply_checks writes to _errors/_warnings (named checks and for_each_column rules are byte-identical to before). Requesting summary metrics via metrics_config without a configured observer now fails fast with an InvalidParameterError instead of silently skipping the metrics table.localStorage with no server-side or table changes, and the change is frontend-only. Non-English translations are AI-assisted and not yet reviewed by native speakers.apx package. scripts/build_app.py generates the FastAPI OpenAPI schema, runs orval, builds the frontend with Vite, and produces the application wheel (with a build-tagged local-version segment so successive deploys at the same commit always reinstall fresh code). scripts/dev.py runs uvicorn with reload alongside the Vite dev server, forwarding signals and tearing down both processes together. The bundle and warehouse-grant scripts were updated to support both bundle-managed and external (reuse) SQL warehouse modes. There is no runtime behavior change in the app itself.prevent_destroy lifecycle protection, and make app-bind adopts pre-existing resources. OLTP tables (rules, settings, RBAC, comments, schedules) move to Postgres via a migration runner, while analytical tables (validation runs, profiling, quarantine, metrics) stay on Delta. Error, warning, and input row counts from the DQX observer are now persisted and surfaced in the UI, label badges and label filtering were added to rule selection and scheduling, and a Spark Connect Observation.get mutability bug that overwrote total row counts was fixed.apply_checks_and_save_in_table and apply_checks_by_metadata_and_save_in_table previously raised AttributeError when called with output_config=None and a quarantine_config. output_config is now optional and skipped when unset, so quarantine-only runs write just the invalid records; passing neither configuration raises a clear InvalidParameterError.DQLLMEngine was imported unconditionally in contract_rules_generator.py purely for a type annotation, causing an ImportError when the [llm] extra was not installed and producing a misleading "install datacontract-cli" error. The import is now guarded behind TYPE_CHECKING, so generate_rules_from_contract(..., process_text_rules=False) works without the [llm] extra.user to client_id in the LakebaseChecksStorageConfig documentation to match the actual configuration field (#1201).BREAKING CHANGES!
enable_contributions defaults to True (was False), adding scoring cost (requires the shap library, already included in the [anomaly] extra). Set enable_contributions=False to restore the previous behaviour. (#1129)enable_ai_explanation defaults to True, so existing anomaly checks will make LLM calls against a Databricks Model Serving endpoint (default databricks-claude-sonnet-4-5) and incur cost. This requires Foundation Model APIs to be available in the workspace; if the endpoint is unreachable, explanations are skipped (null) with a warning rather than failing. Set enable_ai_explanation=False to opt out entirely. (#1129)_dq_info[].anomaly output now contains an additional ai_explanation struct. Downstream consumers that assert on the exact anomaly struct schema should account for the new field. (#1129)metrics_config without a configured observer now raises InvalidParameterError instead of silently skipping the metrics table. (#1193)Full Changelog: v0.14.0...v0.15.0
Contributors:
@mwojtyczka, @fedeflowers, @ghanse, @mvanhorn, @IvannKurchenko, @lfbraz, @aarushisingh04, @berrybluecode, @GewoonMaarten, @SAY-5, @cornzyblack, @acreese11, @Vsatyam013, @ruslan-basyrov,
ML-based row-level anomaly detection ( #990 , #1055 , #1062 ). DQX now offers ML-based row anomaly detection that automatically identifies unusual row
skills/ that teach AI assistants (Databricks Genie Code, Claude Code, Cursor, Copilot, and other tools following the open standard) how to use DQX correctly. The skills cover the public-API capabilities and are accompanied by an AGENTS.md canonical onboarding guide for AI coding agents, with a thin CLAUDE.md redirect for tools that look for it. A new docs guide documents installation and usage for each supported tool.has_no_aggr_outliers, has been introduced that detects outliers in time-series aggregates using a stateless rolling-window sigma method. The check is suitable for monitoring metrics such as daily transaction counts, hourly throughput, or any aggregate where deviations from a rolling baseline indicate quality issues, and complements the existing has_no_outliers MAD-based row-level check.are_polygons_mutually_disjoint, validates whether polygons in a column are mutually disjoint using ST_Intersects. The check supports row_filter, handles nulls and invalid geometries gracefully, and uses native Spark spatial intersections (rather than H3 indexing) for compatibility with Photon's spatial optimizations.foreign_key check now accepts a null_safe parameter. By default, NULL values in the foreign key columns are ignored (SQL ANSI behavior). When null_safe=True, NULL foreign-key values are matched against NULL reference values. Note: enabling null_safe=True on a previously non-null-safe single-column FK changes the auto-generated rule name (a_not_exists_in_ref_b → struct_a_as_a_not_exists_in_ref_struct_b_as_a) and the violation message format.{{ placeholder }} syntax for reusable templates, resolved at load time via a new variables parameter on load_checks() and load_checks_from_local_file(), or via default variables passed through ExtraParams at engine construction. The new resolve_variables() utility recursively replaces placeholders in all string fields of check definitions in a single pass and supports scalar types (str, int, float, bool, Decimal, datetime.date, datetime.datetime, datetime.time). Unresolved placeholders are logged as warnings.suppress_skipped: bool = False option in ExtraParams allows checks skipped due to missing columns or invalid filters to produce no entry in _errors/_warnings and not cause rows to appear in the invalid DataFrame. Additionally, a new skipped boolean field has been added to dq_result_item_schema so skipped checks can be identified structurally without string-parsing the violation message.DQMetricsObserver now emits a new check_metrics row alongside the existing aggregates (input_row_count, error_row_count, warning_row_count, valid_row_count). The value is a JSON array of structs — one per check — with check_name, error_count, and warning_count, fitting the existing metric_name/metric_value schema without widening it. The change is backward compatible: existing metrics are unchanged and the new row is additive.rule_fingerprint, rule_set_fingerprint, and created_at fields when saved to Delta or Lakebase storage, and rule_set_fingerprint is also stamped on summary metrics so every metric row can be traced back to the exact rule version that produced it. Each save creates a new versioned entry rather than overwriting prior history.OutputConfig now accepts partition_by and cluster_by fields, allowing users to save DataFrames as partitioned or clustered tables. Liquid clustering is automatically applied the first time checks are saved to a liquid-clustered table, and the integration tests verify both partitioning and clustering behaviour end to end.error or warn) for generated rules, allowing users to control rule severity at generation time rather than relying on a hardcoded default.generate_schema_validation, defaulting to True), ensuring dataset schemas match contract definitions. A new InvalidPhysicalTypeError provides clearer error handling when physical types are missing or invalid in schema properties.apply_checks_and_save_in_table and apply_checks_by_metadata_and_save_in_table now optionally load checks directly from a storage location (table or file), in addition to the existing option of using preloaded checks. Best-practice documentation has been updated with the recommended end-to-end patterns.demos/dqx_demo_industry/: a Banking demo (dqx_banking_demo.py) focused on fraud detection and transaction monitoring, and a rebuilt Fashion demo (dqx_fashion_demo.py) with industry-specific custom check functions and 11 quality rules. The Manufacturing demo has been moved into the same subdirectory for consistency, and the demo documentation has been updated with a new "Industry Accelerators" section.llms.txt format via the @signalwire/docusaurus-plugin-llms-txt plugin, with hierarchical organization so AI assistants and LLM-powered tools can consume DQX documentation more efficiently.min and max (lexicographic min/max is not meaningful for text data), and a count_distinct metric is now included for all column types in the profiler's summary stats output.py.typed marker file so downstream tools (mypy, pyright, etc.) recognise its existing type annotations instead of treating all databricks.labs.dqx imports as untyped.databricks labs uninstall dqx command now prompts for a custom workspace folder path (mirroring the install flow) and uses the new install_folder parameter on InstallationService.current() to locate installations outside the default /Users/<user>/.dqx location.sql_expression checks in Serverless v5 when the check name is auto-derived, made has_valid_schema compatible with Spark < 4, improved validation of required check function arguments, added agent guidelines, and added documentation on configuring DQX with Lakeflow Declarative Pipelines (LDP/DLT) for Materialized View incrementalization.has_valid_schema silently skipped validation for columns missing from the checked DataFrame has been fixed; missing columns are now reported as schema violations.save_results_in_table now correctly handles the case where the calling DQEngine has an associated observer but no observation or metrics configuration is passed. The bundle has also been updated to use the direct deployment engine.TableChecksStorageHandler now uses WorkspaceClient to check for table existence when saving checks, replacing previous spark.catalog calls and improving compatibility across compute environments.hatch to uv for dependency and build management, GHA workflows have been refactored to increase infrastructure isolation and remove the Azure-login dependency, and performance benchmarks have been moved from per-PR runs to nightly. Dependency versions have been tightened, GitHub Actions are now pinned by SHA, and lock files have been cleaned up to remove registry-specific URLs and unused entries.pyspark.testing.utils.assertDataFrameEqual instead of chispa.assert_df_equality. The chispa test dependency has been removed, the centralized assert_df_equality_ignore_fingerprints wrapper has been updated to translate chispa-style kwargs (ignore_nullable, ignore_column_order, ignore_row_order) to PySpark equivalents, and chispa-specific transforms handling in the e2e PII notebook has been migrated to apply transforms before assertion.BREAKING CHANGES!
overwrite to append. Rules are now versioned going forward — every save produces a new entry stamped with created_at, rule_set_fingerprint, and rule_fingerprint. To preserve the previous overwrite behaviour, explicitly pass mode="overwrite" when saving checks. (#1044)_errors / _warnings array columns) gained two new nullable fields: rule_fingerprint and rule_set_fingerprint (#1044) and skipped (#1063). Pipelines that write to pre-existing Delta output tables created against the older schema will fail with a schema-mismatch error on the next write. Mitigation: pass {"mergeSchema": "true"} in OutputConfig.options (and similarly for the quarantine and metrics outputs) so Delta evolves the table on the first run after upgrade. Code that constructs dq_result_item_schema manually or asserts against an explicit StructType for _errors / _warnings must be updated to match the new shape.apply_checks_and_save_in_table and apply_checks_by_metadata_and_save_in_table. Update callers accordingly — see the methods' updated docstrings for the new signature. (#1064)@ghanse, @mwojtyczka, @vb-dbrks, @fedeflowers, @sundarshankar89, @berrybluecode, @laurencewells, @Roshan1299, @sheeluvikas, @cait-c, @balgaly, @moomindani, @Swayam-arora-2004, @STEFANOVIVAS, @IvannKurchenko, @vpottam-nvidia, @alexott
Full Changelog: v0.13.0...v0.14.0
New DQX Data Quality Dashboard ( #1019 ). The data quality dashboard has been significantly enhanced to provide a centralized view of data quality met
is_null, is_empty, and is_null_or_empty, which enable verification of column values as null, empty strings, or both, complementing existing checks like is_not_null, is_not_empty, and is_not_null_and_not_empty. The functions also support optional arguments, like trim_strings to trim spaces from strings.is_equal_to, is_not_equal_to, is_aggr_equal and is_aggr_not_equal checks, allowing for more flexible and precise control over data validation. The introduction of tolerance logic, which checks for absolute and relative differences within specified thresholds via abs_tolerance and rel_tolerance parameters, provides more nuanced comparisons for numeric data.sql_expression) has been updated to support new lines in its expression argument, allowing for more complex and formatted SQL expressions.filter field in checks definition now correctly pushes down the filter condition defined at the check-level as row_filter to the check function, allowing checks to operate on the relevant subset of rows before aggregation. The documentation has been updated to advice users to use op-level filter condition for consistency instead of row_filter parameter. Overall, these changes aim to enhance the overall user experience.lakebase_user parameter has been replaced with lakebase_client_id, an optional service principal client ID used to connect to Lakebase, defaulting to the caller's identity if not provided. This change enhances the security and reliability of the authentication process, making it easier to work with Lakebase as a checks storage.has_valid_schema check has been enhanced to provide more flexibility in schema validation by introducing an optional exclude_columns parameter, allowing users to specify columns to ignore during validation. This parameter can be used to exclude metadata columns or other columns not relevant to schema validation, and it takes precedence over the columns list.dqx with the current version when it is missing, ensuring that product information is always set for telemetry purposes.@mwojtyczka @ghanse @alexott @nehamilak-db @cornzyblack @laurencewells @renardeinside @tlgnr @pierre-monnet @sheeluvikas @ashwin-911 @dwanneruchi @bpm1993 @Jgprog117
Full Changelog: v0.12.0...v0.13.0
AI-Assisted rules generation from data profiles ( #963 ). AI-assisted data quality rule generation was added, leveraging summary statistics from a pro
DQGenerator class includes a generate_dq_rules_ai_assisted method that can generate rules with or without user-provided input, using summary statistics to inform the rule creation process. This method offers flexibility in rule generation, allowing for both automated and user-guided creation of data quality rules.is_valid_json, has_json_keys, and has_valid_json_schema. The is_valid_json check verifies whether values in a specified column are valid JSON strings, while the has_json_keys check confirms the presence of specific keys in the outermost JSON object, allowing for optional parameters to require all keys to be present. The has_valid_json_schema check ensures that JSON strings conform to an expected schema, ignoring extra fields not defined in the schema.is_area_not_less_than, is_area_not_greater_than, is_area_equal_to, is_area_not_equal_to, is_num_points_not_less_than, is_num_points_not_greater_than, is_num_points_equal_to, and is_num_points_not_equal_to. These checks allow users to validate geometric data based on specific criteria, with options to specify the spatial reference system (SRID) and use geodesic area calculations. These changes enable more effective validation and quality control of geometric data, and are supported in Databricks serverless compute or runtime versions 17.1 and later.save_results_in_table method now accepts output configurations with volume paths, and the OutputConfig object has been updated to support table names with 2 or 3-level namespace, storage paths including Volume paths, S3, ADLS, or GCS, and optional trigger settings for streaming output. Furthermore, the code now supports saving DataFrames to both Delta tables and storage paths, with the save_dataframe_as_table function taking an output_config object that determines whether to save the DataFrame to a table or a path. The functionality includes support for batch and streaming writes, input validation, and error handling, with the existing functionality of saving to Delta tables preserved and new functionality added for saving to storage paths.aggr_params parameter to pass parameters to aggregate functions, such as percentile calculations, and supports two-stage aggregation for window-incompatible aggregates like count_distinct. Additionally, the function includes improved error handling, human-readable violation messages, and performance benchmarks for various aggregation scenarios, enabling advanced data quality monitoring and validation capabilities for data engineers and analysts.is_not_in_list, has been added to verify that values in a specified column are not present in a given list of forbidden values, allowing for null values and optional case-insensitive comparisons. This function is suitable for columns that are not of type MapType or StructType, and for optimal performance with large lists of forbidden values, it is recommended to use the foreign_key dataset-level check with the negate argument set to Trueumn to check, the list of forbidden values, and optionally the case sensitivity of the comparison, and its implementation includes input validation and custom error messages, with additional benchmark tests to measure its performance.is_not_greater_than functions based on the provided minimum and maximum limits, ensuring correct comparison by verifying that both limit values are of the same type. This update preserves the existing numeric behavior and introduces support for timestamp and date checks, while maintaining the ability to handle Python numeric types without stringification.sql_query check has been enhanced to support both row-level and dataset-level validation, allowing for more flexible data validation scenarios. In row-level validation, the check joins query results back to the input data to mark specific rows, whereas in dataset-level validation, the check result applies to all rows, making it suitable for aggregate validations with custom metrics. The merge_columns parameter is now optional, and when not provided, the check performs a dataset-level validation, providing a convenient way to validate entire datasets without requiring specific column mappings. Additionally, the check has been made more robust with input validation and error handling, ensuring that users can perform checks at both the row and dataset levels while preventing incorrect usage with informative error messages.has_no_outliers function has been introduced to detect outliers in numeric columns using the Median Absolute Deviation (MAD) method, which calculates the lower and upper limits as median - 3.5 * MAD and median + 3.5 * MAD, respectively, and considers values outside these limits as outliers. The function is designed to work with numeric columns of type int, float, long, and decimal, and it raises an error if the specified column is not of numeric type. The addition of this function enables the detection of outlier numeric values, enhancing the overall data validation capabilities.has_json_keys function has been updated to treat NULL values as valid, ensuring consistent behavior across ANSI and non-ANSI modes. Additionally, the functionality of saving DataFrames as tables has been improved, with updated regular expression patterns for table names and enhanced handling of streaming and non-streaming DataFrames.has_valid_schema check to accept a reference dataframe or table (#960). The has_valid_schema check has been enhanced to support validation against a reference dataframe or table, in addition to the existing expected schema. This allows users to verify the schema of their input dataframe against a reference dataframe or table by specifying either the ref_df_name or ref_table parameter, with exactly one of expected_schema, ref_df_name, or ref_table required. The check can be performed in strict mode for exact schema matching or in non-strict mode, which permits extra columns, and users can also specify particular columns to validate using the columns parameter. The function's update includes improved parameter validation, ensuring that only one valid schema source is specified, and new test cases have been added to cover various scenarios, including the use of reference tables and dataframes for schema validation, as well as parameter validation logic.is_not_null_island has been introduced to verify whether values in a specified column are NULL island geometries, such as POINT(0 0), POINTZ(0 0 0), or POINTZM(0 0 0 0). The is_not_null_island function requires Databricks serverless compute or runtime version 17.1 or higher.@mwojtyczka @ghanse @souravg-db2 @vb-dbrks @alfredzimmer @AdityaMandiwal @Escanor1996 @larsmoan @STEFANOVIVAS @tdikland @cornzyblack @alfredzimmer @bsr-the-mngrm
Full Changelog: v0.11.1...v0.12.0
Hotfix to update log level for spark connect to suppress dlt telemetry warnings in non-dlt serverless clusters.
Contributors: @mwojtyczka
Generationg of DQX rules from ODCS Data Contracts ( #932 ). The Data Contract Quality Rules Generation feature has been introduced, enabling users to
DQProfiler class now includes a detect_primary_keys_with_llm method, which returns a dictionary containing the primary key detection result, including the table name, success status, detected primary key columns, confidence level, reasoning, and error message if any. The DQGenerator class has been extended to utilize uniqueness profiles from the profiler for AI-assisted uniqueness rules generation. Various updates have been made to the configuration options, including the addition of an llm_primary_key_detection option, which allows users to control whether AI-assisted primary key detection is enabled or disabled.generate_dq_rules_ai_assisted method now accepts an InputConfig object, which allows users to specify the location and format of the input data, enabling more flexible input handling and filtering capabilities. The feature includes test cases to verify its functionality, including manual tests, unit tests, and integration tests, and the documentation has been updated with minor changes to reflect the new functionality. Additionally, the code has been modified to capitalize keywords to stabilize integration tests, and the DQGenerator class has been updated to accommodate the changes, allowing users to generate data quality rules from a variety of input sources. The InputConfig class provides a flexible way to configure the input data, including its location and format, and the get_column_metadata function has been introduced to retrieve column metadata from a given location. Overall, these updates aim to enhance the functionality and usability of the AI-assisted rules generation feature, providing more flexibility and accuracy in generating data quality rules.is_in_list and is_not_null_and_is_in_list check functions have been enhanced to support case-insensitive comparison, allowing users to choose between case-sensitive and case-insensitive comparisons via an optional case_sensitive boolean flag that defaults to True. These checks verify if values in a specified column are present in a list of allowed values, with the is_not_null_and_is_in_list check also requiring the values to be non-null. The updated checks provide more flexibility in data validation, enabling users to configure parameters such as the column to check, the list of allowed values, and the case sensitivity flag. However, it is recommended to use the foreign_key dataset-level check for large lists of allowed values or for columns of type MapType or StructType, as these checks are not suitable for such scenarios.--install-folder argument has been introduced, allowing users to specify a custom installation folder when running various CLI commands, such as opening dashboards, workflows, logs, and profiles. This argument override the default installation location to support scenarios where the user installs DQX in a custom location. The library's dependency on sqlalchemy has also been updated to require a version greater than or equal to 2.0 and less than 3.0 to avoid dependency issues in older DBRs.BREAKING CHANGES!
level parameter to criticality in generate_dq_rules method of DQGenerator for consistency.table: str parameter with input_config: InputConfig in profile_table method of DQProfiler for greater flexibility.table_name: str parameter with input_config: InputConfig in generate_dq_rules_ai_assisted method of DQGenerator for greater flexibility.Contributors: @dinbab1984, @mwojtyczka, @ghanse, @vb-dbrks, @jominjohny, @AdityaMandiwal
Added new field run_id to the detailed per-row quality results. This may or may not be a breaking change for you depending on how you leverage the res…
DQMetricsObserver class has been introduced to manage Spark observations and track summary metrics on datasets checked with the engine. The DQEngine class has been updated to optionally return the Spark observation associated with a given run, allowing users to access and save summary metrics. The engine now supports also writing summary metrics to a table using the metrics_config parameter, and a new save_summary_metrics method has been added to save data quality summary metrics to a table. Additionally, the engine has been updated to include a unique run_id field in the detailed per-row quality results, enabling cross-referencing with summary metrics. The changes also include updates to the configuration file to support the storage of summary metrics. Overall, these enhancements provide a more comprehensive and flexible data quality checking capability, allowing users to track and analyze data quality issues more effectively.DQGenerator class now includes a generate_dq_rules_ai_assisted method, which takes user input in natural language and optionally a schema from an input table to generate data quality rules. These rules are then validated for correctness. The AI-assisted rules generation feature supports both programmatic and no-code approaches. Additionally, the feature enables the use of different LLM models and gives the possibility to use custom check functions. The release also includes various updates to the documentation, configuration files, and testing framework to support the new AI-assisted rules generation feature, ensuring a more streamlined and efficient process for defining and applying data quality rules.checks_location resolution has been updated to accommodate Lakebase, supporting both table and file storage, with flexible formatting options, including "catalog.schema.table" and "database.schema.table". The Lakebase checks storage backend is configurable through the LakebaseChecksStorageConfig class, which includes fields for instance name, user, location, port, run configuration name, and write mode. This update provides users with more flexibility in storing and loading quality checks, ensuring that checks are saved correctly regardless of the specified location format.ConfigSerializer class, which handles the serialization and deserialization of workspace and run configurations.hatch-fancy-pypi-readme to fix images in PyPi (#601). The image source path for the logo in the README has been modified to correctly display the logo image when rendered, particularly on PyPi.databricks-labs-pytester version from 0.7.2 to 0.7.4, and code refactoring has been done to use a single Lakebase instance for all integration tests, with retry logic added to handle cases where the workspace quota limit for the number of Lakebase instances is exceeded, enhancing the testing infrastructure and improving test reliability. Furthermore, documentation updates have been made to clarify the application of quality checks to data using DQX. These changes aim to improve the efficiency, reliability, and clarity of the project's testing and documentation infrastructure.BREAKING CHANGES!
LIMITATIONS
Contributors: @mwojtyczka, @ghanse, @souravg-db2, @vb-dbrks, @alexott, @tlgnr
Added support for running checks on multiple tables ( #566 ). Added more flexibility and functionality in running data quality checks, allowing users
profiler_max_parallelism and quality_checker_max_parallelism. A new demo has been added to showcases how to use the profiler and apply checks across multiple tables. The changes aim to improve scalability of DQX.is_valid_ipv6_address check function), and validation if IPv6 address is within provided CIDR block (is_ipv6_address_in_cidr check function).has_valid_schema check function has been introduced to validate whether a DataFrame conforms to a specified schema, with results reported at the row level for consistency with other checks. This function can operate in non-strict mode, where it verifies the existence of expected columns with compatible types, or in strict mode, where it enforces an exact schema match, including column order and types. It accepts parameters such as the expected schema, which can be defined as a DDL string or a StructType object, and optional arguments to specify columns to validate and strict mode.is_latitude, is_longitude, is_geometry, is_geography, is_point, is_linestring, is_polygon, is_multipoint, is_multilinestring, is_multipolygon, is_ogc_valid, is_non_empty_geometry, has_dimension, has_x_coordinate_between, and has_y_coordinate_between. The addition of these geospatial data validation checks enhances the overall data quality capabilities, allowing for more accurate and reliable geospatial data processing and analysis. Running these checks requires Databricks serverless or cluster with runtime 17.1 or above.compare_datasets check has been enhanced with the introduction of absolute and relative tolerance parameters, enabling more flexible comparisons of decimal values. These tolerances can be applied to numeric columns.BREAKING CHANGES!
DQEngine: load_checks_from_local_file, load_checks_from_workspace_file, load_checks_from_table, load_checks_from_installation, save_checks_in_local_file, save_checks_in_workspace_file, save_checks_in_table,, save_checks_in_installation, load_run_config. For loading and saving checks, users are advised to use load_checks and save_checks of the DQEngine described here, which support various storage types.Contributors: @mwojtyczka, @ghanse, @tdikland, @Divya-Kovvuru-0802, @cornzyblack, @STEFANOVIVAS
Added performance benchmarks (#548). Performance tests are run to ensure performance does not degrade by more than 25% by any change. Benchmark result
deserialize_checks_to_dataframe function has been enhanced to correctly handle columns for sql_expression by removing the unnecessary check for DQDatasetRule instance and directly verifying if dq_rule_check.columns is not None.Contributors: @mwojtyczka @ghanse @cornzyblack @gchandra10
Added quality checker and end to end workflows (#519). This release introduces no-code solution for applying checks. The following workflows were adde
does_not_contain_pii check function and can be customized to suit specific use cases. The check requires pii extras to be installed: pip install databricks-labs-dqx[pii]. Furthermore, a new enum class NLPEngineConfig has been introduced to define various NLP engine configurations for PII detection. Overall, these updates aim to provide more robust and customizable quality checking capabilities for detecting PII data.is_equal_to and is_not_equal_to, have been introduced to enable equality checks on column values, allowing users to verify whether the values in a specified column are equal to or not equal to a given value, which can be a numeric literal, column expression, string literal, date literal, or timestamp literal.BREAKING CHANGES!
ExtraParams was moved from databricks.labs.dqx.rule module to databricks.labs.dqx.configContributors: @mwojtyczka @ghanse @renardeinside @cornzyblack @bsr-the-mngrm @dinbab1984 @AdityaMandiwal
If you are loading or saving checks from a storage (file, workspace file, table, installation), you are affected. We are deprecating the below methods…
is_data_fresh, has been introduced to identify stale data resulting from delayed pipelines, enabling early detection of upstream issues. This function assesses whether the values in a specified timestamp column are within a specified number of minutes from a base timestamp column. The function takes three parameters: the column to check, the maximum age in minutes before data is considered stale, and an optional base timestamp column, defaulting to the current timestamp if not provided.is_data_fresh_per_time_window, has been added to validate whether at least a specified minimum number of records arrive within every specified time window, ensuring data freshness. This function is customizable, allowing users to define the time window, minimum records per window, and lookback period.null_safe_row_matching and null_safe_column_value_matching, have been introduced to control how null values are handled, both defaulting to True. These parameters allow for null-safe primary key matching and column value matching, ensuring accurate comparison results even when null values are present in the data. The check now excludes specific columns from value comparison using the exclude_columns parameter while still considering them for row matching.round=False option, which was previously ignored. The code now handles the OverflowError that occurs when rounding up the maximum datetime value by capping the result and logging a warning.checks_location, replacing the previous checks_file and checks_table fields, to simplify the configuration and remove ambiguity by ensuring only one storage location can be defined per run configuration. The checks_location field can point to a file in the local path, workspace, installation folder, or Unity Catalog Volume, providing users with more flexibility and clarity when managing their quality checks.DQEngine class has undergone significant changes to improve modularity and maintainability, including the unification of methods for loading and saving checks under the load_checks and save_checks methods, which take a config parameter to determine the storage type, such as FileChecksStorageConfig, WorkspaceFileChecksStorageConfig, TableChecksStorageConfig, or InstallationChecksStorageConfig.DQRule objects and Python dictionaries, allowing for flexibility in check definition and usage. The serialize_checks method converts a list of DQRule instances into a dictionary representation, while the deserialize_checks method performs the reverse operation, converting a dictionary representation back into a list of DQRule instances. Additionally, the DQRule class now includes a to_dict method to convert a DQRule instance into a structured dictionary, providing a standardized representation of the rule's metadata. These changes enable users to work with checks in both formats, store and retrieve checks easily, and improve the overall management and storage of data quality checks. The conversion process supports local execution and handles non-complex column expressions, although complex PySpark expressions or Python functions may not be fully reconstructable when converting from class to metadata format.BREAKING CHANGES!
checks_file and checks_table fields have been removed from the installation run configuration. They are now consolidated into the single checks_location field. This change simplifies the configuration and clearly defines where checks are stored.load_run_config method has been moved to config_loader.RunConfigLoader, as it is not intended for direct use and falls outside the DQEngine core responsibilities.DEPRECIATION CHANGES!
If you are loading or saving checks from a storage (file, workspace file, table, installation), you are affected. We are deprecating the below methods. We are keeping the methods in the DQEngine but you should update your code as these methods will be removed in future versions.
load_checks method. The following methods have been removed from the DQEngine:
load_checks_from_local_file, load_checks_from_workspace_file, load_checks_from_installation, load_checks_from_table.load_checks method. The following methods have been removed from the DQEngine:
save_checks_in_local_file, save_checks_in_workspace_file, save_checks_in_installation, save_checks_in_table.The save_checks and load_checks take config as a parameter, which determines the storage types used. The following storage configs are currently supported:
FileChecksStorageConfig: file in the local filesystem (YAML or JSON)WorkspaceFileChecksStorageConfig: file in the workspace (YAML or JSON)TableChecksStorageConfig: a tableInstallationChecksStorageConfig: storage defined in the installation context, using either the checks_table or checks_file field from the run configuration.Contributors: @mwojtyczka, @karthik-ballullaya-db, @bsr-the-mngrm, @ajinkya441, @cornzyblack, @ghanse, @jominjohny, @dinbab1984
Added type validation for apply checks method (#465). The library now enforces stricter type validation for data quality rules, ensuring all elements
DQRule. If invalid types are encountered, a TypeError is raised with a descriptive error message, suggesting alternative methods for passing checks as dictionaries. Additionally, input attribute validation has been enhanced to verify the criticality value, which must be either warn or "error", and raises a ValueError for invalid values.compare_datasets, has been introduced to compare two DataFrames at both row and column levels, providing detailed information about differences, including new or missing rows and column-level changes. This check compares only columns present in both DataFrames, excludes map type columns, and can be customized to exclude specific columns or perform a FULL OUTER JOIN to identify missing records. The compare_datasets check can be used with a reference DataFrame or table name, and its results include information about missing and extra rows, as well as a map of changed columns and their differences.is_valid_ipv4_address and is_ipv4_address_in_cidr, have been introduced to verify whether values in a specified column are valid IPv4 addresses and whether they fall within a given CIDR block, respectively.Contributors: @mwojtyczka, @cornzyblack, @ghanse, @grusin-db
Only users using config.yml are affected. The breaking change is for the required field: output_location which is now stored inside input_config.locat…
DQEngine class has been updated to utilize InputConfig and OutputConfig objects to handle input and output configurations, providing more flexibility in the quality checking flow. The apply_checks_and_write_to_table and apply_checks_by_metadata_and_write_to_table methods have been introduced to support this functionality, applying checks using DQX classes and configuration, respectively. Additionally, the profiler configuration options have been reorganized into input_config and profiler_config sections, making it easier to understand and customize the profiling process. The changes aim to provide a more streamlined and efficient way to perform end-to-end quality checking and data validation, with improved configuration flexibility and readability.is_aggr_equal and is_aggr_not_equal, which enable users to perform equality checks on aggregate values, such as count, sum, average, minimum, and maximum, allowing verification that an aggregation on a column or group of columns is equal to or not equal to a specified limit. These checks can be configured with a criticality level of either error or warn and can be applied to specific columns or groups of columns. Additionally, the foreign_key check has been updated with a negate option, allowing the condition to be negated so that the check fails when the foreign key values exist in the reference dataframe or table, rather than when they do not exist. This expanded functionality enhances the library's data quality checking capabilities, providing more flexibility and power in validating data integrity.DLT has been renamed to Lakeflow Pipeline in documentation and docstrings, to maintain consistency in terminology. The quick demo has been enhanced to showcase defining checks using DQX classes, providing a more comprehensive approach to data quality validation. Additionally, performance information related to dataset-level checks has been added to the documentation, and instructions on how to use the Environment to install DQX in Lakeflow Pipelines have been provided.sql_expression now supports optional columns argument that is propagated to the results.BREAKING CHANGES!
config.yml are affected. The breaking change is for the required field: output_location which is now stored inside input_config.location.Full example:
log_level: INFO
version: 1
profiler_override_clusters: # <- optional dictionary mapping job cluster names to existing cluster IDs
main: your-existing-cluster-id # <- existing cluster Id to use
profiler_spark_conf: # <- optional spark configuration to use for the profiler job
spark.sql.ansi.enabled: true
run_configs:
- name: default # <- unique name of the run config (default used during installation)
input_config: # <- optional input data configuration
location: s3://iot-ingest/raw # <- input location of the data (table or cloud path)
format: delta # <- format, required if cloud path provided
is_streaming: false # <- whether the input data should be read using streaming (default is false)
schema: col1 int, col2 string # <- schema of the input data (optional), applicable if reading csv and json files
options: # <- additional options for reading from the input location (optional)
versionAsOf: '0'
output_config: # <- output data configuration
location: main.iot.silver # <- output location (table), used as input for quality dashboard ir quarantine locaiton is not provided
format: delta # <- format of the output table
mode: append # <- write mode for the output table (append or overwrite)
options: # <- additional options for writing to the output table (optional)
mergeSchema: 'true'
#checkpointLocation: /Volumes/catalog1/schema1/checkpoint # <- only applicable if input_config.is_streaming is enabled
trigger: # <- streaming trigger, only applicable if input_config.is_streaming is enabled
availableNow: true
quarantine_config: # <- quarantine data configuration, if specified, bad data is written to quarantine table
location: main.iot.silver_quarantine # <- quarantine location (table), used as input for quality dashboard
format: delta # <- format of the quarantine table
mode: append # <- write mode for the quarantine table (append or overwrite)
options: # <- additional options for writing to the quarantine table (optional)
mergeSchema: 'true'
#checkpointLocation: /Volumes/catalog1/schema1/checkpoint # <- only applicable if input_config.is_streaming is enabled
trigger: # <- streaming trigger, only applicable if input_config.is_streaming is enabled
availableNow: true
checks_file: iot_checks.yml # <- relative location of the quality rules (checks) defined in json or yaml file
checks_table: main.iot.checks # <- table storing the quality rules (checks)
profiler_config: # <- profiler configuration
summary_stats_file: iot_summary_stats.yml # <- relative location of profiling summary stats
sample_fraction: 0.3 # <- fraction of data to sample in the profiler (30%)
sample_seed: 30 # <- optional seed for reproducible sampling
limit: 1000 # <- limit the number of records to profile
warehouse_id: your-warehouse-id # <- warehouse id for refreshing dashboard
- name: another_run_config # <- unique name of the run config
...
Contributors: @mwojtyczka, @ghanse
Input parameters remain the same as before. This is a breaking change for checks defined using DQX classes. Yaml/Json definitions are not affected.
Release v0.6.0 (#395)
DQDatasetRule class has been added to define dataset-level checks, and several new check functions have been added including the foreign_key and sql_query dataset-level checks. The library now also supports custom dataset-level checks using arbitrary SQL queries and provides the ability to define checks on multiple DataFrames or Tables. The DQEngine class has been modified to optionally accept Spark session as a parameter in its constructor, allowing users to pass their own Spark session. Major internal refactorization has been carried out to improve code maintenance and structure.git commit -S command, and if any unsigned commits are found, it posts a comment with instructions on how to properly sign
commits.profile_table and profile_tables, which enable direct profiling of Delta tables, allowing users to generate summary statistics and candidate data quality rules. These methods provide a convenient way to profile data stored in Delta tables, with profile_table generating a profile from a single Delta table and profile_tables generating profiles from multiple Delta tables using explicit table lists or regex patterns for inclusion and exclusion. The profiling process is highly customizable, supporting extensive configuration options such as sampling, outlier detection, null value handling, and string handling. The generated profiles can be used to create Delta Live Tables expectations for enforcing data quality rules, and the profiling results can be stored in a table or file as YAML or JSON for easy management and reuse.BREAKING CHANGES!
is_unique, is_aggr_not_greater_than and is_aggr_not_less_than checks under dataset-level checks umbrella. These checks must be defined using DQDatasetRule class and not DQRowRule anymore. Input parameters remain the same as before. This is a breaking change for checks defined using DQX classes. Yaml/Json definitions are not affected.DQRowRuleForEachCol has been renamed to DQForEachColRule to make it generic and handle both row and dataset level rules.column_names to result_column_names in the ExtraParams for clarity as they may be confused with column(s) specified for the rules itself. This is a breaking change!Contributors: @mwojtyczka @alexott @nehamilak-db @ghanse @Escoto
Fix spark remote version detection in CI by @mwojtyczka in https://github.com/databrickslabs/dqx/pull/342
Corrected is_unique check to consistently handle nulls. To conform with the SQL ANSI Standard), null values are treated now as unknown, thus not duplicates, e.g. "(NULL, NULL) not equals (NULL, NULL); (1, NULL) not equals (1, NULL). A new parameter called nulls_distinct has been added to the check which is set to True by default. If set to False, null values are treated as duplicates, e.g. eg. (1, NULL) equals (1, NULL) and (NULL, NULL) equals (NULL, NULL) Added columns parameter to the is_unique check to be able to pass list of columns to handle composite keys.
This release introduces changes to the syntax of quality rules to ensure we can scale DQX functionality in the future. Please update your quality rules accordingly.
col_name field in the checks to column to better reflect the fact that it can be a column name or column expression.From:
- check:
function: is_not_null
arguments:
col_name: col1
To:
- check:
function: is_not_null
arguments:
column: col1
column instead of col_name as first argument.col_name to columns in the reporting columns. DQX stores list of columns instead of a single column now to handle checks that take multiple columns as input.col_names to for_each_column and moved it from check function arguments the check level.From:
- check:
function: is_not_null
arguments:
col_names:
- col1
- col2
To:
- check:
function: is_not_null
for_each_column:
- col1
- col2
DQColRule to DQRowRule for clarity. This only affects checks defined using classes.DQColSetRule to DQRowRuleForEachCol for clarity. This only affects checks defined using classes.row_checks module containing check functions to check_funcs. This only affects checks defined using classes.The documentation has been updated accordingly: https://databrickslabs.github.io/dqx/docs/reference/quality_rules/
Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.4.0...v0.5.0
Added input spark options and schema for reading from the storage (https://github.com/databrickslabs/dqx/issues/312). This commit enhances the data qu
Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.3.1...v0.4.0
Removed usage of lambda in quality checking (#310). We have replaced the usage of lambda functions n the quality checking with a more efficient implem
Contributors: @mwojtyczka
In addition, the col_functions module has been renamed to col_check_functions. This introduces a breaking change!. It is recommended to to update any…
DQRule and DQRuleColSet have been renamed to DQRuleCol and DQColSetRule, respectively, to support the addition of more rule types in the future, such as DQDatasetRule. The renaming includes corresponding changes in imports and method calls throughout the codebase. A deprecation warning has been added to the old classes. In addition, the col_functions module has been renamed to col_check_functions. This introduces a breaking change!. It is recommended to to update any references to the old class names in your code to ensure a smooth transition.sql_expression check to fail if the condition is not met, introducing a potential breaking change.Contributors: @mwojtyczka, @ghanse, @pierre-monnet
…have been clarified. This change includes a breaking change for some of the checks. Users are advised to review and test the changes before implementa…
is_not_less_than and is_not_greater_than functions now accept column names or expressions as limits. The input parameters for range checks have been unified, and the logic of is_not_in_range has been updated to be inclusive of the boundaries. The project's documentation has been improved, with the addition of comprehensive examples, and the contribution guidelines have been clarified. This change includes a breaking change for some of the checks. Users are advised to review and test the changes before implementation to ensure compatibility and avoid any disruptions. Resolves issues: 131, 197, 175, 205validate_checks method has been updated to accept a dictionary of custom check functions instead of global variables. However, globals() can still be specified for backward compatibility. This improvement resolves the issue #48.Contributors: @mwojtyczka
Fixed cli installation and demo (#177). In this release, changes have been made to adjust the dashboard name, ensuring compliance with new API naming
is_in_range and is_not_in_range quality rule functions have been updated to support a column as the minimum or maximum limit, in addition to a literal value. This change is accomplished through the introduction of optional min_limit_col_expr and max_limit_col_expr arguments, allowing users to specify a column expression as the minimum or maximum limit. Extensive testing, including unit tests and integration tests, has been conducted to ensure the correct behavior of the new functionality. These enhancements offer increased flexibility when defining quality rules, catering to a broader range of use cases and scenarios.Contributors: @karthik-ballullaya-db, @mwojtyczka
Fixed installation process for Serverless (#150). This commit removes the pyspark dependency from the librar to avoid spark version conflicts in Serve
Contributors: @mwojtyczka
Provided option to customize reporting column names (#127). In this release, the DQEngine library has been enhanced to allow for customizable reportin
filepath instead of "path". Additionally, new unit and integration tests have been added and manually tested to ensure the correct functionality of the updated code. Contributors @mwojtyczkatry_cast Spark function and replace it with cast and isNull checks to improve code compatibility, particularly for runtimes where try_cast is not available. The affected functionality includes null and empty column checks, checking if a column value is in a list, and checking if a column value is a valid date or timestamp. We have added unit and integration tests to ensure functionality is working as intended. Contributors @mwojtyczkaFull Changelog: https://github.com/databrickslabs/dqx/compare/v0.1.8...v0.1.11
Fixed docs-build by @mwojtyczka in https://github.com/databrickslabs/dqx/pull/129
Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.1.8...v0.1.10
Fixed docs-build by @mwojtyczka in https://github.com/databrickslabs/dqx/pull/129
Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.1.8...v0.1.9
Updated docs by @mwojtyczka in https://github.com/databrickslabs/dqx/pull/117
Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.1.7...v0.1.8
Set cache invalidation for pypi badge by @mwojtyczka in https://github.com/databrickslabs/dqx/pull/102
Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.1.6...v0.1.7
Fix for image links in README on PyPi by @alexott in https://github.com/databrickslabs/dqx/pull/95
Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.1.5...v0.1.6
Fix README on PyPi by using hatch-fancy-pypi-readme in the build by @alexott in https://github.com/databrickslabs/dqx/pull/81
hatch-fancy-pypi-readme in the build by @alexott in https://github.com/databrickslabs/dqx/pull/81Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.1.4...v0.1.5
* Updated release process
## What's Changed * Updated release process * Fixed installation via cli Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.1.0...v0.1.1
Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.1.0...v0.1.1
Fix warning about deprecated ruff usage by @alexott in https://github.com/databrickslabs/dqx/pull/25
make fmt by @nfx in https://github.com/databrickslabs/dqx/pull/7ruff usage by @alexott in https://github.com/databrickslabs/dqx/pull/25isinstance check to fix incompatibility with Spark Connect by @alexott in https://github.com/databrickslabs/dqx/pull/24Full Changelog: https://github.com/databrickslabs/dqx/compare/v0.0.0...v0.1.0
Your coding agent can read these notes before it upgrades. Set up the MCP server →