NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3935 most downloaded on PyPI
The datacontract CLI is an open source command-line tool for working with Data Contracts. It uses data contract YAML files to lint the data contract, connect to data sources and execute schema and quality tests, detect breaking changes, and export to different formats. The tool is written in Python. It can be used as a standalone CLI tool, in a CI/CD pipeline, or directly as a Python library.
Last release 9 days ago
25 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
92 releases · first in 2024
datacontract api : --contract-variables and --allow-local-files options, as alternatives to their environment variables
datacontract api: --contract-variables and --allow-local-files options, as alternatives to their environment variablesdatacontract import odata creates a datacontract from OData 4 metadata at a URL or from a local file (#1649 @jahlen)datacontract export writes SQL Server VECTOR(n) and VECTOR(n, float16) vector typesdatacontract test checks constraints and quality rules of nested properties on servers read through DuckDB (#1278)datacontract lint and datacontract test: a quality rule the CLI cannot run is reported as a warning instead of being silently droppeddatacontract edit) is updated to 0.1.14${VAR} references are accepted in fields with a fixed set of values (quality.type, quality.metric, quality.dimension, logicalType, servers[].type); datacontract test fails when one resolves into anything elsedatacontract test reports nested checks it cannot run and unsupported type checks on parquet files as warnings (#1278)datacontract api:
${VAR} only from the variables allow-listed with --contract-variables${X:-local} server type no longer slips past the checks for local files and environment-held credentialsdatacontract test: a SQL quality rule on a mysql server can no longer reach the MySQL server or its credentials; rules on mysql and iceberg servers are read as DuckDB SQLarguments, so invalidValues and missingValues rules lost their configurationdatacontract test --dry-run: a check that could not be planned no longer reports the run as skippeddatacontract test: SQL quality rule placeholders are quoted when the name needs it (e.g. a column with a space), and a query that does not parse reports the parse error (#1653)datacontract import spark preserves varchar(...) and char(...) types nested inside struct, array and map columns instead of widening them to string (#1634 @IchEssBlumen)datacontract import jsonschema keeps the type of nullable anyOf/oneOf properties and warns about union types (#1278)datacontract export sodacl warns about the nested properties it leaves out (#1278)One column per month.
Support for SAP HANA and Exasol
datacontract test:
hana extra (#1332 @ToniLippmann)exasol extra (#1335)datacontract export excel and datacontract import excel support all versions of the Excel template (ODCS v3.0.2, v3.1.0, v3.2.0)datacontract import sql detects table-level PRIMARY KEY and FOREIGN KEY constraints, derives unique: true for single-column keys, and emits property-level relationships in the ODCS shorthand format (#1618 @dmaresma)datacontract export great-expectations --checks restricts the suite to quality and/or properties expectations (#1617 @julienguilhempartner-spec)datacontract import warns once per import about column types that have no mapping to a logicalType (#1629 @regdat)datacontract lint validates against the ODCS schema for the apiVersion the contract declares, instead of always the newest onedatacontract-cli[csv] installs DuckDB, which datacontract import csv needs, instead of pandasdatacontract test and datacontract import sqlserver escape the host, database and driver in the SQL Server connection string, so a ; in the contract's servers block can no longer inject connection keywordsdatacontract api:
authoritativeDefinitions only against the Entropy Data host configured on the server, without following redirects, and answers a failed lookup on every endpoint with 422 and the URL, never with what the host answeredx-api-key with 403 instead of 500 and warns on failed API key checks/test no longer logs the submitted contractdatacontract changelog and datacontract breaking no longer interpret contract values as terminal markupdatacontract test:
--filter reports a duplicate key whose other occurrence lies outside the filtered rows (#1593,#1619 @gkhnelbstn)NULL values as duplicates of each other in uniqueness checkscli auth works again on macOS/Linux, and ActiveDirectoryInteractive fails fast off Windows instead of timing out (#1603 @johannzv)freshness service level it cannot interpret as a single failed check instead of aborting the whole run; freshness now also accepts an ISO-8601 duration as its value, like retentiondatacontract import:
string or date to the right logicalType, including in avro and spark imports (#1629 @regdat)s3, gcs and adls accept --format again, so Delta tables can be imported (#1628 @regdat)sql warns about statements it cannot parse and therefore skips, such as a CREATE TABLE with unquoted hyphens in its name (#686 @dwelden)datacontract dbt sync writes the generated-column marker under config.meta instead of a top-level meta so the model YAML parses under dbt Fusion; files written by an earlier version are migrated on the next sync (#1633 @FredrikBakken)datacontract dbt sync removes a column it generated once the property leaves the contract, without --prune; a column the user added their own tests or settings to is keptdatacontract changelog and datacontract breaking:
quality rulesbreaking recognizes changes to quality rulesContracts declaring apiVersion: v3.2.0 lint and round-trip, including enum , map , vector , semanticType , synonyms , deprecated , context , variables…
This release adds support for the Open Data Contract Standard v3.2.0 (#1557).
datacontract test checks vector element types when the data source reports them (#1585)datacontract test resolves variables in property names, enum values, nested type options, library quality arguments, and service levels (#1583)apiVersion: v3.2.0 lint and round-trip, including enum, map, vector, semanticType, synonyms, deprecated, context, variables in string values, and the new server types (#1558)datacontract export avro-idl writes map<...> and array<float> for map and vector properties; datacontract export sqlalchemy writes JSON and ARRAY(Float) (#1558)datacontract test resolves ${VAR} and ${VAR:-default} references in server fields and SQL quality queries from the environment; an unset variable without a default fails the run with its name, and export keeps the references (#1559)port may be a string such as ${DB_PORT}, including in Excel imports; config files accept ${VAR:-default} (#1559)enum on properties is read by datacontract test, the jsonschema, avro, avro-idl, protobuf, pydantic-model, dcs, great-expectations, sodacl and data-caterer exporters, and dbt test mapping, ahead of logicalTypeOptions.enum, the enum custom property and the invalidValues rule; HTML export lists the values with labels and descriptions (#1560)datacontract import from JSON Schema, Avro, Protobuf and DCS writes allowed values as enum entries instead of an invalidValues quality rule or custom properties (#1560)datacontract export pydantic-model types an enumerated string or integer property as typing.Literal (#1560)logicalType: map with a map block (key and value as full property definitions) is a first-class type: datacontract test checks the key and value types, including nested objects and maps, on DuckDB, Databricks, Snowflake, Trino, Kafka and the file sources (#1562)datacontract export writes native map types for snowflake, databricks, dataframe, duckdb (local, s3), clickhouse, trino, spark, iceberg, avro, avro-idl, protobuf, pydantic-model, go and dcs, JSON for postgres, mysql, sqlserver, oracle and bigquery, and additionalProperties for jsonschema; HTML shows the key and value types (#1562)datacontract import from sql, databricks/unity, spark, glue, iceberg, avro, parquet and dcs writes logicalType: map with the key and value instead of physicalType: map with custom properties; the mapKeyType, mapValueType, mapKeys and mapValues custom properties are still read but deprecated (#1562)logicalType: vector with logicalTypeOptions.dimensions is a first-class type: datacontract test accepts a native vector or an array of numbers and compares the dimensions when the column states them; export writes vector(n) for postgres, VECTOR(FLOAT, n) for snowflake, FLOAT[n] for duckdb, vector(n) for mysql, arrays of floats for databricks, dataframe, trino, clickhouse, bigquery, spark, iceberg, avro, avro-idl, protobuf, pydantic-model, go and dcs, and a fixed-length array of numbers for jsonschema (#1563)datacontract import from sql, postgres and snowflake reads vector(n), halfvec(n) and VECTOR(FLOAT, n) columns as logicalType: vector with their dimensions, and parquet reads a fixed-size list of floats the same way (#1563)datacontract export html renders semanticType, synonyms, deprecated and context on schema objects and properties, the contract-level context, vendor on custom properties, and customProperties and authoritativeDefinitions on SLA entries; export rdf writes enum entries, synonyms, map definitions, customProperties and other nested definitions as nodes of their own instead of dropping them or printing a model repr (#1561)datacontract changelog matches synonyms, enum entries, context statements and constraints, and the nested lists of SLA entries by their natural key, and relationships by id when present, instead of by position (#1561)datacontract test reads CSV files from local, s3, gcs and azure servers and Kafka JSON messages in the server's declared encoding, and uses the Athena workgroup (also DATACONTRACT_ATHENA_WORKGROUP), which makes stagingDir optional (#1564)hana, iceberg, exasol, teradata, ingres, vectorwise, versant and poet lint and export; test explains that it cannot connect to them yet, and the synonyms fastobjects and btrieve resolve to poet and zen (#1564)datacontract export sql --sql-server-type auto warns before falling back to the snowflake dialect for a server type without one (#1564)datacontract test supports the iceberg server type: tables are read from the REST catalog named by catalogUrl, namespace and warehouse with pyiceberg (credentials via DATACONTRACT_ICEBERG_CREDENTIAL or DATACONTRACT_ICEBERG_TOKEN, data files via the S3 options; DATACONTRACT_ICEBERG_CATALOG_TYPE selects a sql, glue or hive catalog instead of rest, DATACONTRACT_ICEBERG_S3_ENDPOINT points at an S3-compatible store, Amazon S3 Tables and the Glue REST endpoint are signed with SigV4 from the AWS credentials, DATACONTRACT_ICEBERG_PROPERTIES passes further catalog properties through), and datacontract import iceberg --catalog-url creates a contract from a catalog table with a ready-to-test server (#1565)datacontract breaking detects enum restrictions and vector shape changes (#1584)datacontract init and all importers write apiVersion: v3.2.0 (#1558)datacontract edit) is updated to 0.1.13 (#1566)open-data-contract-standard dependency bumped to 3.2.x (#1558).#/\@!%&^, as in the released ODCS v3.2.0 schemadatacontract breaking command and POST /breaking endpoint for breaking change detection ( #1016 @pierre-monnet )
authoritativeDefinitions link can reference a file next to the contract, either a property in another contract (url: business.odcs.yaml#schema/orders/properties/order_id) or a file that is the definition itself (url: definitions/order_id.odcs.yaml), resolved relative to the referencing contract (#1453)datacontract breaking command and POST /breaking endpoint for breaking change detection (#1016 @pierre-monnet)DATACONTRACT_KAFKA_GROUP_PREFIX environment variable to customise the consumer group ID prefix used during Kafka testing (#1553 @philipp-lutz)datacontract test checks the ODCS array options minItems, maxItems and uniqueItems (#1514 @OGsiji)datacontract export odcs defaults status to draft when the source DCS contract has no info.status (#1542 @michal-swiatowy)datacontract export great-expectations covers the logicalTypeOptions constraints, attaches contract metadata to every expectation, checks the column set instead of the column order and takes a --suite-name (#1544 @julienguilhempartner-spec)datacontract test --dry-run reports the checks a run would execute without connecting to the server or reading any data (#1510 @OGsiji)datacontract export jsonschema, datacontract export avro and datacontract test on local files use a property's physicalName as the field name when set, instead of the logical name (#1494 @philipp-lutz)datacontract import sql takes the server's database and schema from a qualified CREATE TABLE, instead of always writing placeholders (#651 @ReguiguiMohamed)datacontract import sql no longer fails on a DDL file that contains CREATE SCHEMA (#1529 @ReguiguiMohamed)datacontract test sums every component of an ISO 8601 retention period, instead of reading only the first one (#1538)datacontract test reports each freshness and retention check result on its own check, instead of writing every result to the first one (#1515 @erikgrip2)datacontract test --publish with an empty value runs without publishing, instead of failing (#1491)open-data-contract-standard dependency is pinned to 3.1.x; ODCS v3.2.0 support will come with a dedicated releaseNew server type duckdb to test the tables inside a DuckDB database file (the file is opened read-only)
duckdb to test the tables inside a DuckDB database file (the file is opened read-only)--dry-run flag for datacontract dbt sync that reports the same plan as a real sync, but writes nothing to disk (#1513 @q-maze)quality rule a check comes from (#1526)postgresql is accepted as the ODCS synonym of the postgres server typedatacontract import pydantic-model reads the contract description from the module docstring (#1507 @OGsiji)engine field of test results is now always one of datacontract-cli, dbt or jsonschema; the values datacontract, ibis and dbt-sync no longer occur (#1505)datacontract test and datacontract ci report a failed publish of the test results on the console and exit with code 1, instead of silently succeedingdatacontract api:
ENTROPY_DATA_HOST; other publish_url targets are refused (#1541)Access-Control-Allow-Origin: * headers; it serves no CORS headers at all, since the only browser client is the same-origin Swagger UI--reload to enable it (development only)skipped, not infodatacontract export html and datacontract catalog now HTML-escape data contract field values, closing a stored cross-site scripting hole--publish URL points at; set ENTROPY_DATA_HOST for a self-hosted deploymentquality.type: sql rule must be a read-only query; DDL, DML, COPY, ATTACH and the like are reported as a failed check instead of being executed, for every data sourcedatacontract api:
servers[].type: local, so a posted data contract cannot read the files of the server running it; set DATACONTRACT_CLI_API_ALLOW_LOCAL_FILES=true to allow itservers sectionx-api-key header in constant time to avoid a timing side channelauthoritativeDefinitions lookup instead of echoing the server-side error reason, which is now logged insteadquality.type: sql rule is read in the SQL dialect of its server type, so dialect-specific syntax (BigQuery backticks, Snowflake SAMPLE, SQL Server TOP) is no longer mistaken for an invalid queryDATACONTRACT_BIGQUERY_PROJECT, DATACONTRACT_POSTGRES_SCHEMA, and the like) now apply to the whole test run, not just to the connectiondatacontract test and datacontract export sodacl freshness and retention checks now honor the schema object's and property's physicalName (#1488 @erikgrip2)datacontract test --publish no longer fails for runs with skipped checks (skipped checks are omitted from the published test results)DATACONTRACT_BIGQUERY_BILLING_PROJECT no longer overrides the server's project, so tables are still read from the data project (#1358)datacontract export pydantic-model exports the ODCS timestamp and time logical types as datetime.datetime and datetime.time instead of guessing from the physical type, which turned a TIMESTAMPTZ column into a str (#1507 @OGsiji)datacontract export pydantic-model exports the ODCS date logical type as datetime.date instead of datetime.datetime (#1507 @OGsiji)datacontract import pydantic-model maps datetime.datetime to the timestamp logical type and datetime.time to time (#1507 @OGsiji)datacontract export dcs exports a pii custom property as a boolean instead of the string 'True'Test results name the check fields qualityId and failedSamples instead of quality_id and failed_samples ; the old names are deprecated, but still acce…
${dataset}, ${project}, ${catalog}, and ${database} placeholders for the server's valuesdatacontract test --metadata-only runs only checks that read the schema (field presence and types) and reports checks that read row values as skippeddatacontract test --checks accepts the ODCS terms properties and slaProperties, keeping schema and servicelevel as legacy aliasesdatacontract import pydantic-model creates a data contract from Pydantic modelsdatacontract export excel uses the ODCS Excel template bundled with the CLI instead of downloading it, so the export works offline; use --template for a custom templatedatacontract api reports the CLI version, describes every endpoint, parameter, and response model, and names its operations testDataContract, lintDataContract, exportDataContract, and changelogBetweenDataContractsDATACONTRACT_CLI_API_KEY now also protects POST /lint and POST /export, which previously answered without an API keyqualityId and failedSamples instead of quality_id and failed_samples; the old names are deprecated, but still accepted as input and still written next to the new onesdatacontract import unity no longer writes the databricksType custom property, which duplicated physicalTypedatacontract export sql --server databricks keeps the declared length of varchar(n) and char(n) instead of exporting STRINGdatacontract test for API servers with delimiter: array no longer fails JSON schema validation with data must be object (#1495)POST /export answers 422 instead of 500 when the posted data contract cannot be parseddatacontract test for Databricks no longer fails all checks of a model with a GEOGRAPHY or GEOMETRY column (#1483)datacontract export pydantic-model no longer emits an unparseable empty class for an object property without propertiesThis release drops the pyspark compile-time dependency. The server types dataframe and databricks still work with a provided Spark session. This remov
This release drops the pyspark compile-time dependency. The server types dataframe and databricks still work with a provided Spark session.
This removes the JVM dependency, makes the images much, much smaller (Docker image from 777 MB to 277 MB), and many CVEs are resolved.
DATACONTRACT_KAFKA_MAX_MESSAGES limits how many messages datacontract test reads from a topic, and DATACONTRACT_KAFKA_TIMEOUT how long it waits for onedatacontract test for Kafka no longer needs PySpark or a Java runtime: datacontract-cli[kafka] now installs confluent-kafka, fastavro, and DuckDB insteaddatacontract-cli[databricks] and datacontract-cli[dataframe] no longer install PySpark, so they can no longer shadow the build a Databricks Runtime or EMR cluster provides; supply your own PySpark for the Spark session these server types take, and databricks-runtime is now an alias of databricksdatacontract-cli[all] installs no PySpark at all, and so needs no Java runtimeRUN pip install, so datacontract dbt test, which needs a dbt adapter, is not available in the containerdatacontract import unity resolves struct and array columns into nested properties without PySpark installedspark_exporter.to_spark_schema(), to_struct_type(), to_struct_field(), and to_spark_data_type() return the exporter's own SparkDataType instead of pyspark.sql.types objects; use to_spark_dict() or the new to_pyspark_schema() for real PySpark schemasdatacontract export spark and datacontract export great-expectations --engine spark no longer require PySpark to be installedspark_exporter.to_spark_dict() reports which PySpark version lacks a type instead of raising a bare AttributeError, e.g. VariantType before PySpark 4.0datacontract test for Databricks no longer creates ibis's memtable staging volume, so read-only service principals without CREATE VOLUME permission can run testsDATAMESH_MANAGER_API_KEY , DATAMESH_MANAGER_HOST , DATACONTRACT_MANAGER_API_KEY , and DATACONTRACT_MANAGER_HOST (and the matching Config fields): use
DATAMESH_MANAGER_API_KEY, DATAMESH_MANAGER_HOST, DATACONTRACT_MANAGER_API_KEY, and DATACONTRACT_MANAGER_HOST (and the matching Config fields): use ENTROPY_DATA_API_KEY and ENTROPY_DATA_HOST insteaddatacontract export html fits to datacontract-editor visualizationdatacontract test --filter and --filters test only the rows matching a SQL predicate, e.g., the latest partition; also available as filter/filters query parameters on the API server's POST /test (#1463)DataContract(config=...) and DataContract.import_from_source(..., config=...), using the typed datacontract.Config class or a dict keyed by the environment variable namesPOST /test via datacontract-* headers (e.g. datacontract-snowflake-password)--config-file (defaults to ./datacontract-config.yaml or ~/.datacontract/config.yaml) or Config.from_yaml(), with ${VAR} references resolved from the environmentDATACONTRACT_SNOWFLAKE_TOKEN, DATACONTRACT_SNOWFLAKE_PASSCODE, DATACONTRACT_SNOWFLAKE_PRIVATE_KEY, DATACONTRACT_SNOWFLAKE_NETWORK_TIMEOUT, DATACONTRACT_SNOWFLAKE_SOCKET_TIMEOUT, DATACONTRACT_SNOWFLAKE_HOST, DATACONTRACT_SNOWFLAKE_PORTDATACONTRACT_POSTGRES_HOST (#1076)s3:// URLs for commands that read contracts, including lint, test, export, publish, changelog, and cidatacontract test against Snowflake only forwards the documented DATACONTRACT_SNOWFLAKE_* options to the connector; unknown variables are ignored with a warningDATACONTRACT_SNOWFLAKE_ACCOUNT overrides the contract's account instead of being ignored with a warningDATACONTRACT_DATABRICKS_SERVER_HOSTNAME overrides the contract's host; previously the contract value won when both were setdatacontract dbt sync writes a description on all dbt tests, not only newly-generated onesdatacontract test reads the Avro schema of a Kafka topic from the Confluent Schema Registry via DATACONTRACT_KAFKA_SCHEMA_REGISTRY_URL , DATACONTRACT_
datacontract test reads the Avro schema of a Kafka topic from the Confluent Schema Registry via DATACONTRACT_KAFKA_SCHEMA_REGISTRY_URL, DATACONTRACT_KAFKA_SCHEMA_REGISTRY_USERNAME, and DATACONTRACT_KAFKA_SCHEMA_REGISTRY_PASSWORD (#1347)DATACONTRACT_SNOWFLAKE_PRIVATE_KEY_PATH, DATACONTRACT_SNOWFLAKE_PRIVATE_KEY_PASSPHRASE, and DATACONTRACT_SNOWFLAKE_CONNECTION_TIMEOUT: use DATACONTRACT_SNOWFLAKE_PRIVATE_KEY_FILE, DATACONTRACT_SNOWFLAKE_PRIVATE_KEY_FILE_PWD, and DATACONTRACT_SNOWFLAKE_LOGIN_TIMEOUT insteadDATACONTRACT_SQLSERVER_TRUSTED_CONNECTION: use DATACONTRACT_SQLSERVER_AUTHENTICATION=windows insteaddatacontract test against a Kafka topic reports Avro messages it cannot decode as such, instead of reading every field as null (#1347)datacontract test connects to Databricks with an OAuth service principal again, instead of failing with Error during request to server on databricks-sql-connector 4.3.0 and later (#1389)DATACONTRACT_SQLSERVER_TRUSTED_CONNECTION no longer overrides an explicitly set DATACONTRACT_SQLSERVER_AUTHENTICATION, so a leftover flag cannot silently downgrade an Entra ID login to Windows authenticationdatacontract test passes DATACONTRACT_IMPALA_AUTH_MECHANISM, DATACONTRACT_IMPALA_USE_HTTP_TRANSPORT, and DATACONTRACT_IMPALA_HTTP_PATH to Impala again, so a Cloudera Virtual Warehouse can be reached instead of failing with TSocket read 0 bytesdatacontract test passes the servers block catalog to Athena again, instead of always querying awsdatacatalogdatacontract test supports DATACONTRACT_BIGQUERY_IMPERSONATION_ACCOUNT again to impersonate a service accountdatacontract test applies the documented Snowflake key-pair and timeout variables instead of silently ignoring them: the names they were documented under are not accepted by the Snowflake driver, and now map to the ones that aredatacontract dbt sync does not assume severity: warn as default anymore: tests now fail with dbt's default severity unless the contract declares a non-blocking quality.severitydatacontract dbt sync no longer drops or misplaces YAML comments that introduce the next column, test, or keydatacontract test treats a server typed mssql as SQL Server, so contracts carrying the ODBC/dbt spelling are testable (ODCS itself only defines sqlser
datacontract test treats a server typed mssql as SQL Server, so contracts carrying the ODBC/dbt spelling are testable (ODCS itself only defines sqlserver)datacontract test --quality-id runs a single quality rule by its ODCS quality.id, and --tag runs every quality rule declaring one of the given quality.tags (#1080)quality_id and tags of the quality rule a check comes fromdatabricks-runtime extra for installing inside a Databricks Runtime, where the cluster already provides PySpark: pip install datacontract-cli[databricks-runtime] (#1211 @chifu1234)--version and --system-truststore, and every import and export guide links to its command page and backdatacontract test --dimension runs only the checks measuring one data quality dimension, e.g. --dimension uniqueness; it matches the ODCS quality.dimension of a rule and the schema and service level checks that measure the same aspectdataframe extra installs just what testing Spark DataFrames needs: pip install datacontract-cli[dataframe]datacontract import trino creates a data contract from a Trino catalog, including a ready-to-test servers blockdatacontract import oracle creates a data contract from a live Oracle database, including a ready-to-test servers blockdatacontract import gcs and datacontract import adls create a data contract from files in Google Cloud Storage or Azure Blob Storage, including a ready-to-test servers blockdatacontract import sqlserver creates a data contract from a live SQL Server database, including a ready-to-test servers blockdatacontract import mysql creates a data contract from a live MySQL database, including a ready-to-test servers blockdatacontract import s3 creates a data contract from files in an S3 bucket, including a ready-to-test servers blockdatacontract import athena creates a data contract from an Amazon Athena database, including a ready-to-test servers blockdatacontract import unity is now datacontract import databricks; the unity format name keeps workingDATACONTRACT_REDSHIFT_AUTHENTICATION is no longer required and remains as an overridedatacontract import postgres creates a data contract from a live Postgres schema, including a ready-to-test servers blockdatacontract import redshift creates a data contract from an Amazon Redshift schema, including a ready-to-test servers blockdatacontract import bigquery now generates a servers block, so datacontract test works right after the importdatacontract test verifies declared primary keys: each key column must have no missing values, and the key must have no duplicates (a composite key is checked as a tuple) (#1220 @DMZ22)datacontract test checked physicalType against a same-named table in another schema when one existed, because the native column types were read from the catalog without the contract's schemadatacontract test against SQL Server no longer fails every check with "Could not read model" when server.schema differs from the login's default schemadatacontract-cli[s3] could not run datacontract test, and datacontract-cli[gcs] was missing the AWS duckdb extension the GCS connection loads; each data source extra now installs everything its guide needsdatacontract-cli[duckdb] nowdatacontract test told users to install datacontract-cli[local], an extra that does not exist, and datacontract-cli[api], which installs the web server rather than a test backend; both now point at duckdbdatacontract import gcs wrote type: gcs, which is not an ODCS server type, so the imported contract failed datacontract lint and datacontract test; GCS is now written as an s3 server on the Google interoperability endpointendpointUrl, which is interpolated into the statement that stores the S3, GCS and Azure credentials; every value is escaped nowdatacontract api server accepted a local file path as the schema query parameter, so a caller could have it read files from the server's filesystem; only http(s) URLs are accepted nowinformation_schema has no length or precision columns, so the catalog query failed and a wrong physicalType still passeddatacontract import athena and datacontract import glue now honour DATACONTRACT_S3_ACCESS_KEY_ID and DATACONTRACT_S3_SECRET_ACCESS_KEY; the Glue catalog was read with ambient AWS credentials onlyaws sso login, AWS_PROFILE, instance roles) when no access key is configured; previously such a setup failed with 403 Forbiddenaws sso login, AWS_PROFILE, instance roles); static access keys were presented as the only optionregionName in an Athena servers block was ignored, so the region could only be set via DATACONTRACT_S3_REGIONdatacontract import glue mapped timestamp columns to logicalType: date instead of timestampcodec not available in Python: 'UNICODE'pip install "botocore[crt]"logicalType: timenumber fields without a declared precision and scaleemail, uuid, date-time) to logicalTypeOptions.format instead of a custom property, so they are validated by datacontract testTIME types with precision or time zone (e.g. TIME(9)) to logicalType: time; previously the logical type was left unsetdatacontract test --checks quality now runs rowCount quality rules, which were wrongly categorized as schema checksdatacontract import dbt derives the contract id from the dbt manifest's project_name instead of always emitting the placeholder my-data-contract (#1221 @DMZ22)physicalType declaring a zero scale (NUMBER(38,0), decimal(18,0)) failed against its own column on Snowflake, Oracle, SQL Server and Databricks (#1377 @DMZ22)physicalType checks failed for structured OBJECT, ARRAY and MAP columns, whose fields are now compared field by field instead of as rendered strings (#1377 @DMZ22)datacontract test on Athena failed a physicalType written in the Hive spelling datacontract import athena produces (array<string> against the reported array(varchar)) (#1377 @DMZ22)physicalType declaring fractional seconds (TIMESTAMP_NTZ(9), datetime2(7), timestamp(3)) failed against its own column, so every timestamp column imported from Snowflake failed the first datacontract testdatacontract test and datacontract import oracle read Oracle character lengths in bytes, so an NVARCHAR2(50) column was reported and checked as NVARCHAR2(100)datacontract test on Databricks could not check the element types of ARRAY, MAP and STRUCT columns, which the catalog reports as a bare type namedatacontract export protobuf supports a customizable package name via the protoPackageName custom property ( #1381 @Schokuroff )
datacontract export protobuf supports a customizable package name via the protoPackageName custom property (#1381 @Schokuroff)datacontract export protobuf emits optional for non-required message/object fields (#1390 @Schokuroff)datacontract dbt sync resolves {object}/{property} placeholders in custom sql quality checks to the dbt ref() and column name (#1397)datacontract/lint/resolve.py:72: SyntaxWarning: 'return' in a 'finally' block return except_message is handled properly (#1384 @Cupprum)datacontract test no longer reports "backend is not installed" for Athena and other ibis SQL backends when packaging is missing from the environmentdatacontract test reports the nested types of a property that declares properties: or items: as a separate check
datacontract test reports the nested types of a property that declares properties: or items: as a separate checkdatacontract test on Snowflake verifies the physicalType of nested properties against the real column typedatacontract test on Snowflake matches a physicalType against the alias the catalog reports, such as BIGINT on a NUMBER(38,0) columndatacontract test no longer reports a mismatch for a physicalType without precision, such as NUMBER on a NUMBER(12,2) columndatacontract test now recursively verifies nested logicalType for Snowflake structured OBJECT / ARRAY columns
datacontract import can now import BigQuery type INTERVAL (#1367,#1372 @fantastisch)
datacontract import can now import BigQuery type INTERVAL (#1367,#1372 @fantastisch)ENTROPY_DATA_HOST value to set when the IRI host does not match the configured entropy-data host.datacontract test no longer reports a physical type mismatch for BigQuery type aliases, such as a physicalType of INTEGER on an INT64 column (#1371 @fantastisch)datacontract test no longer fails with CANNOT_CONVERT_COLUMN_INTO_BOOL on Databricks when a Spark session is usedextended datacontract dbt sync:
datacontract dbt sync:
--prune flag to remove everything that's not specified in the contract (models, tags, checks) - per default, only generated content gets removedversions: block)meta: blockdatacontract dbt test: Use local dbt to run all tests that have been generated using datacontract dbt sync earlier (scoped to a single data contract, or all data contracts in the opened dbt project)
datacontract dbt sync:
--run-tests or run datacontract dbt test afterwards)datacontract test now verifies a field's physicalType against the column's real native type from the platform catalog (length and precision included), taking precedence over logicalType (#1354)datacontract test JSON output now includes datacontractCliVersion (#1353 @hk8suva)datacontract test type-check errors now report the first failing field and a count of the remaining errors (#1334 @jorgengranseth)datacontract test no longer fails the type check for SQL Server uniqueidentifier (UUID) columns with "the column type could not be determined" (#1354)datacontract test against BigQuery no longer fails SQL quality checks with 'RowIterator' object has no attribute 'fetchone'datacontract import against BigQuery applies correct logicalType for BigQuery types TIMESTAMP, DATETIME, and TIME (#1366 @fantastisch)datacontract test against Kafka no longer reports every field as null for plain (non-Confluent Schema Registry) Avro messages
datacontract test against Kafka no longer reports every field as null for plain (non-Confluent Schema Registry) Avro messages (#1344)pattern argument in invalidValues quality checks (#1346)datacontract test against Redshift no longer fails with relation "pg_catalog.pg_enum" does not exist. Redshift rides the Postgres ibis backend, whose
datacontract test against Redshift no longer fails with relation "pg_catalog.pg_enum" does not exist. Redshift rides the Postgres ibis backend, whose schema introspection joins pg_catalog.pg_enum to detect enum columns — a relation Redshift does not expose. Introspection now omits that join (Redshift has no enum types).New global option --system-truststore (env DATACONTRACT_SYSTEM_TRUSTSTORE) to verify TLS using the operating system's certificate trust store, for use
--system-truststore (env DATACONTRACT_SYSTEM_TRUSTSTORE) to verify TLS using the operating system's certificate trust store, for use behind corporate proxies or with internal CAs.datacontract test against Redshift no longer fails with column "current_schema" does not exist. Redshift rides the Postgres ibis backend, whose intros
datacontract test against Redshift no longer fails with column "current_schema" does not exist. Redshift rides the Postgres ibis backend, whose introspection resolved the active schema with SELECT current_schema (no parentheses) — valid on PostgreSQL but rejected by Redshift, which only supports current_schema(). The configured server schema is now passed explicitly during introspection, skipping that query.datacontract test now only supports logicalTypes. Previously physicalType was preferrerd and used even if logicalType did not exist.
datacontract test now only supports logicalTypes. Previously physicalType was preferrerd and used even if logicalType did not exist.datacontract test field type check now compares the full structured type tree for object and array logical types.map type is not supported until ODCS version v3.2.0 and is also ignored.datacontract --help no longer fails with ModuleNotFoundError: No module named 'ibis' when the optional ibis extra is not installed.datacontract test against Oracle now qualifies tables with the configured server schema (owner), fixing Could not read model '<table>': <table> when the login user differs from the table owner.datacontract test validates Azure Blob Storage / ADLS Gen2 file metadata against a data contract used as a storage policy
datacontract test validates Azure Blob Storage / ADLS Gen2 file metadata against a data contract used as a storage policy (#1227)datacontract test against Trino now supports DATACONTRACT_TRINO_AUTHENTICATION=jwt with DATACONTRACT_TRINO_JWT_TOKEN, and DATACONTRACT_TRINO_AUTHENTICATION=oauth2 for the interactive browser flow.datacontract export sql --dialect clickhouse: export data contracts to ClickHouse SQL DDL. (#1293)examples/, used as the worked examples on the docs export and import pages.datacontract import unity now imports struct and array columns as structured ODCS types instead of only a flat type string: structs get nested properties, arrays get items (parsed recursively from Unity's type_json, including descriptions on nested fields). Arrays also get the correct logicalType: array (previously object); this logical type fix applies to the sql and snowflake importers as well. Map columns and other unmappable SQL types now leave logicalType unset (instead of the invalid object) until ODCS v3.2 adds logicalType: map (RFC 0030). (#1280)datacontract export dcs no longer crashes on data contracts with a structured description or a standard server.datacontract test against Oracle releases before 23ai no longer fails every check with ORA-00923: FROM keyword not found where expected during schema introspection.datacontract test now recognizes Oracle VARCHAR2/NVARCHAR2 (with or without a length, e.g. VARCHAR2(4000)) as string types in the field type check.datacontract import powerbi imports a data contract from a Power BI semantic model (.pbit, .bim, or .json) (#1232,#1233 @dmaresma)
datacontract import powerbi imports a data contract from a Power BI semantic model (.pbit, .bim, or .json) (#1232,#1233 @dmaresma)datacontract edit shows the edited filename in the editor header, no longer offers New/Load Example/Open, and Cancel reverts to the file on diskdatacontract test no longer aborts remaining quality checks after a SQL quality check falls back to native execution on DuckDB-backed sources (#1302)DATACONTRACT_S3_SESSION_TOKEN to the connection again, fixing UnrecognizedClientException with temporary AWS credentials (#1309)datacontract edit now accepts a URL: it asks to download a local copy and opens that in the editor, using the configured API key (e.g., ENTROPY_DATA_A
datacontract edit now accepts a URL: it asks to download a local copy and opens that in the editor, using the configured API key (e.g., ENTROPY_DATA_API_KEY) to fetch the file.DATACONTRACT_BIGQUERY_BILLING_PROJECT environment variable. When set, query jobs are submitted to (and billed against) the billing project while data is read from the project specified in the server config. This enables cross-project and cross-organisation validation without requiring bigquery.jobUser on the data project.datacontract test no longer fails on Spark/Databricks tables containing columns of types ibis cannot represent, such as VariantType (#1296)postgres and redshift extras now install psycopg[binary], whose wheels bundle the libpq client library. Previously, datacontract test against PostgreSQL/Redshift failed on machines without PostgreSQL client libraries installed (couldn't import psycopg: no pq wrapper available).Environment variables (e.g., credentials for datacontract test) are now also loaded from a .env file in the current working directory, walking up pare
datacontract test) are now also loaded from a .env file in the current working directory, walking up parent directories until one is found. Already-set environment variables take precedence over .env values, so exported variables and CI secrets keep working unchanged. (#1295)datacontract edit <file> opens a local data contract file in the Data Contract Editor (web UI). Starts a local web server (default port 4243, requires the api extra) that serves the editor, writes saves directly back to the file, and doubles as the editor's test runner: "Run test" in the editor executes the data contract tests locally via the server's own /test endpoint. The editor is bundled with the CLI and works offline; --editor-version loads a specific editor version from the CDN instead, --editor-assets-url a self-hosted build. If the file does not exist, the command offers to initialize a new data contract (same template as datacontract init). Saving gives feedback in the editor and on the console.datacontract test no longer fails with Compilation rule for RegexSearch operation is not defined when a contract uses logicalTypeOptions.pattern (or pattern) against SQL Server. SQL Server has no native regex operator, so the ibis mssql backend cannot compile re_search; pattern checks now fall back to a PATINDEX(...) > 0 LIKE match, restoring the behaviour of the former soda-core engine. Patterns that use real regex syntax (anchors, quantifiers, groups, .) cannot be expressed as a T-SQL LIKE pattern and now raise a clear error instead of failing cryptically. (#1284)Replaced the Soda Core quality/test engine with ibis. datacontract test now compiles schema and quality checks into ibis expressions (dialect-correct
datacontract test now compiles schema and quality checks into ibis expressions (dialect-correct SQL per backend via sqlglot, local/remote files via DuckDB) instead of generating SodaCL. Install extras now pull ibis-framework[<backend>] instead of soda-core-*. Check semantics and pass/fail results are preserved for the supported sources (postgres, redshift, mysql, snowflake, bigquery, databricks, sqlserver, oracle, trino, athena, impala, kafka/dataframe via the ibis Spark backend, and local/S3/GCS/Azure files).quality.type: custom with engine: soda) are no longer executed and now report a warning. Migrate them to quality.type: sql or a library metric (e.g. metric: rowCount).requires-python now allows 3.10–3.14). On 3.13/3.14 the Spark-backed extras resolve to PySpark 4.0 (Spark 3.5 has no 3.13+ build); the Kafka/Avro connector jars already adapt to the runtime Spark/Scala version. create_spark_session now pins PYSPARK_PYTHON/PYSPARK_DRIVER_PYTHON to the running interpreter so Spark's Python workers match the driver.datacontract test against Databricks now supports more authentication methods beyond the personal access token (DATACONTRACT_DATABRICKS_TOKEN): an OAuth service principal for machine-to-machine auth (DATACONTRACT_DATABRICKS_CLIENT_ID + DATACONTRACT_DATABRICKS_CLIENT_SECRET), a local config profile via the Databricks SDK unified auth (DATACONTRACT_DATABRICKS_PROFILE, also covers Azure CLI/MSI), and an explicit connector auth type (DATACONTRACT_DATABRICKS_AUTH_TYPE, e.g. databricks-oauth for the interactive browser flow).datacontract test now records structured diagnostics on each check explaining why it passed or failed: the metric, measured value, threshold, and (for "bad row" metrics) the total row count and failed fraction. invalid_count checks also report the validity rule they enforced (e.g. {"max_length": 20}, {"pattern": "^.+@.+$"}). The diagnostics surface in the JSON output and the JUnit failure text. This replaces the Soda-specific diagnostics payload that the ibis migration had left unpopulated.datacontract test now honors ODCS quality.unit: percent on the count-of-bad-rows library metrics (nullValues, missingValues, invalidValues). The threshold is then compared against the failed fraction (0–100) of the model row count instead of an absolute count, so e.g. metric: nullValues, unit: percent, mustBeLessThan: 5 passes when fewer than 5% of rows are null. Percent on metrics where a row fraction has no meaning (rowCount, duplicateValues) logs a warning and falls back to the absolute count. The measured percent is added to the check diagnostics.datacontract test now honors ODCS quality.severity: a non-blocking severity (info, warning, low, minor, trivial) downgrades a failing quality check to a warning instead of a failed, so it no longer fails the run. Any other severity (or none) still fails. The severity is recorded in the check diagnostics.datacontract test --include-failed-samples collects a small sample of the rows that failed each missing/invalid/duplicate check (off by default). Each sample is restricted to the contract's identifier columns (unique / primaryKey fields) plus the offending column; duplicate checks report the duplicated key values and their counts. Columns whose ODCS classification marks them sensitive (pii, personal, confidential, restricted, sensitive, secret) are omitted. Samples are capped at 5 rows per check and surface on Check.failed_samples in the JSON output and in the JUnit failure text. This is local-only and needs no Soda Cloud (soda-core itself collects failed-row samples only via Soda Cloud).mysql extension instead of ibis's native MySQL backend, so the mysql extra stays pure-pip (no mysqlclient C build / system MySQL client libraries required).duckdb-extension-* wheels (httpfs/aws/azure) pinned in lockstep. The 1.5.x extension wheels publish arm64 Linux builds, so air-gapped installs on arm64 Linux are now supported (the previous platform skip markers were removed).kafka and databricks extras allow pyspark<5.0.import protobuf now uses the pure-Python proto-schema-parser instead of the protoc system binary. The protobuf extra no longer requires protoc (or the protobuf runtime), so .proto import works out of the box, including transitive imports across subdirectories.pip / uv installs at build time are routed through Socket Firewall Free, which blocks malicious dependencies. (#1275, #1277)SparkSession startup because the base image had no JVM. End users pulling datacontract/cli are unaffected by the build-side changes. (#1277)Secret Validation Failure) on DuckDB ≥1.5: explicit KEY_ID/SECRET now use the default config provider instead of CREDENTIAL_CHAIN.import csv no longer fails with a DuckDB binder error on DuckDB ≥1.5 (the uniqueness probe now uses count(DISTINCT ...) via SQL).soda-core runtime dependency and all soda-core-* install extras, plus the setuptools runtime shim they required. The sodacl export format (datacontract export sodacl) is unchanged and is now generated independently of any Soda runtime.details field on test-result checks (Run.checks[].details). It was a Soda-era placeholder that was never populated; the new structured diagnostics field replaces it. The JUnit failure text no longer prints a Details: line.Resolve authoritativeDefinitions[type=definition] on schema properties, filling fields the contract author left unset
authoritativeDefinitions[type=definition] on schema properties, filling fields the contract author left unset (#1261)
authoritativeDefinitions[type in {definition, semantics}] now rejects the contract on lint, test, ci, export, and changelog. (#1261)authoritativeDefinitions[type=semantics] (and the legacy type=semantic) the same way. A url: that points at the configured entropy-data host is fetched directly; a url: that's an IRI (host doesn't match) is routed through GET /api/semantics?iri=... on the configured host, which uses the API key's organization to resolve. (#1262)--no-inline-references flag to skip the HTTP fetch (#1261)--json-schema points at a custom JSON Schema, the ODCS Pydantic step now accepts extra top-level fields the schema allows (#1266)test, lint, and ci now infer --output-format from the --output file extension when not given (.json → json, .xml → junit) (#1156 @dallylee)varchar(n) columns on Databricks with PySpark 4.0+, and for map and varchar types nested inside struct columns; affected columns emit a warning and skip the type check instead (#1219,#1245 @IchEssBlumen)WARNING/ERROR log messages are no longer hidden by default for import, export, changelog, catalog, dbt, and publish (#1264)datacontract test against S3, GCS, and Azure no longer fails with Failed to download extension in air-gapped containers. The required DuckDB extensions are now bundled via the s3/gcs/azure install extras (#1191 @ParenParikh)field_is_present check now correctly detects missing columns (#1065,#1163 @hieusats)import dbt --model <name>.vN now correctly imports the specified version of a versioned dbt model. Previously the filter compared the full name.vN string against node["name"] (which is always the bare base name), producing a silent empty contract (#1249 @willbowditch)datacontract test now logs the Data Contract CLI version and whether it ran as a local CLI or through the FastAPI server (including the request URL) a
datacontract test now logs the Data Contract CLI version and whether it ran as a local CLI or through the FastAPI server (including the request URL) as part of the test result logsphysicalName when set (falling back to name), matching the existing table-level resolution and the SQL/BigQuery exporters. Previously a property whose logical name differed from its physical column (e.g. name: brand with physicalName: BRAND) failed the presence and type checks even though the physical column existed (#1246)new datacontract dbt sync command: generate dbt tests from an ODCS contract, then run dbt test for them, and optionally publish the results to Entropy
datacontract dbt sync command: generate dbt tests from an ODCS contract, then run dbt test for them, and optionally publish the results to Entropy Data (#1222, #1235)redshift server type for datacontract test (requires pip install datacontract-cli[redshift]). (#1236)decimal/numeric per dialect (Postgres → numeric, MySQL → decimal) so test's column-type check matches information_schema (#1237)impala extra (pip install datacontract-cli[impala]) — pulls in soda-core-impala. Impala engine support landed in #965 but the install extra was never
impala extra (pip install datacontract-cli[impala]) — pulls in soda-core-impala. Impala engine support landed in #965 but the install extra was never added; users had to install soda-core-impala manually. Also included in [all].dbt extra and the dbt-core dependency. import dbt now reads manifest.json directly with no third-party dependency, and works without installing any extra. Minimum supported manifest schema version is v9 (dbt 1.5+). Users who installed datacontract-cli[dbt] should switch to plain datacontract-cli.protobuf extra now requires the protoc compiler installed on the system. Replaces the bundled grpcio-tools (~50 MB of platform-specific protoc binaries) with the lighter protobuf runtime (>=3.20,<7.0). import protobuf raises a clear error with platform-specific install hints if protoc is not on PATH. Install with brew install protobuf (macOS), sudo apt install protobuf-compiler (Debian/Ubuntu), etc. — see README.csv, excel, and oracle extras. The matching [project.optional-dependencies] entries already existed but were undocumented.{object} and ${object} placeholder in SQL quality queries as the ODCS-spec name for the current schema object (alias for {model}/{table}) (#676)changelog command help text now advertises (url or path) for V1/V2 arguments, clarifying that HTTP/HTTPS URLs are accepted (#1162)test command now exits non-zero when a server is specified, but soda-core fails to connect or authenticate (#1181)check_type labels model_qualty_sql and field_quality_sql (#1187)import spark now emits a native Spark SQL physicalType (e.g. string) instead of Python repr (e.g. StringType()). Contracts imported using Spark in v0.11.0–v0.12.1 did not perform type checks and must be re-imported. (#1048)setuptools as a base dependency. soda-core's env_helper.py imports from distutils.util import strtobool; distutils was removed from stdlib in Python 3.12 and stripped from python-build-standalone 3.11 builds. setuptools provides the distutils shim. Previously pulled in transitively via grpcio-tools; now required explicitly. Reverts #1199 — see soda-core#2091.This release introduces several changes to improve the usability of datacontract-cli for AI Agents.
This release introduces several changes to improve the usability of datacontract-cli for AI Agents.
Fix in v0.12.1: re-added
--schemaas alias for the new--json-schema(will be removed in v0.13.0)
| Command | Old option | New option |
|---|---|---|
lint, test, ci, publish, catalog |
--schema <PATH> (still works until v0.13.0) |
--json-schema <PATH> |
export, import |
--format <FORMAT> <OPTIONS> |
<FORMAT> <OPTIONS> (drop --format) |
| Export options: | ||
export --format dbt |
--format dbt |
dbt-models (format renamed) |
export --format great-expectations |
--sql-server-type <TYPE> |
--dialect <TYPE> |
export --format rdf |
--rdf-base <URI> |
--base <URI> |
export --format sql |
--sql-server-type <TYPE> |
--dialect <TYPE> |
export --format sql-query |
--sql-server-type <TYPE> |
--dialect <TYPE> |
| Import options: | ||
import --format bigquery |
--bigquery-[project|dataset|table] <NAME> |
--[project|dataset|table] <NAME> |
import --format dbt |
--dbt-model <NAME> |
--model <NAME> |
import --format glue |
--source <NAME>, --glue-table <NAME> |
--database <NAME>, --table <NAME> |
import --format iceberg |
--iceberg-table <NAME> |
--table <NAME> |
import --format unity |
--unity-table-full-name <NAME> |
--table <NAME> |
import --format spark |
--source <NAMES> |
--tables <NAMES> |
import |
--template |
dropped (was a no-op) |
The --schema option (referring to the ODCS JSON schema) was renamed to --json-schema to avoid confusion with --schema-name, which refers to the schema within the data contract to test for.
--debug (or set DATACONTRACT_CLI_DEBUG=1) to see the full traceback. (#1175)--help outputs (#1176)--schema a deprecated alias for --json-schema to (will be removed in v0.13.0)This release introduces several changes to improve the usability of datacontract-cli for AI Agents.
This release introduces several changes to improve the usability of datacontract-cli for AI Agents.
Fix in v0.12.1: re-added
--schemaas alias for the new--json-schema(will be removed in v0.13.0)
| Command | Old option | New option |
|---|---|---|
lint, test, ci, publish, catalog |
--schema <PATH> (still works until v0.13.0) |
--json-schema <PATH> |
export, import |
--format <FORMAT> <OPTIONS> |
<FORMAT> <OPTIONS> (drop --format) |
| Export options: | ||
export --format dbt |
--format dbt |
dbt-models (format renamed) |
export --format great-expectations |
--sql-server-type <TYPE> |
--dialect <TYPE> |
export --format rdf |
--rdf-base <URI> |
--base <URI> |
export --format sql |
--sql-server-type <TYPE> |
--dialect <TYPE> |
export --format sql-query |
--sql-server-type <TYPE> |
--dialect <TYPE> |
| Import options: | ||
import --format bigquery |
--bigquery-[project|dataset|table] <NAME> |
--[project|dataset|table] <NAME> |
import --format dbt |
--dbt-model <NAME> |
--model <NAME> |
import --format glue |
--source <NAME>, --glue-table <NAME> |
--database <NAME>, --table <NAME> |
import --format iceberg |
--iceberg-table <NAME> |
--table <NAME> |
import --format unity |
--unity-table-full-name <NAME> |
--table <NAME> |
import --format spark |
--source <NAMES> |
--tables <NAMES> |
import |
--template |
dropped (was a no-op) |
The --schema option (referring to the ODCS JSON schema) was renamed to --json-schema to avoid confusion with --schema-name, which refers to the schema within the data contract to test for.
--debug (or set DATACONTRACT_CLI_DEBUG=1) to see the full traceback. (#1175)--help outputs (#1176)Added --checks option to test command to selectively run check categories: schema, quality, servicelevel
--checks option to test command to selectively run check categories: schema, quality, servicelevel (#678)--schema-name option to test command to test a specific schema instead of all schemas (#1079,#1085 @kelsoufi-sanofi)precision/scale for number types from logicalTypeOptions to customProperties (#1145,#1160 @davidb-tada)--server name is not found (#1153,#1161 @Ai-chan-0411)Thanks to @kelsoufi-sanofi for the new --schema-name option on test, and to @Schokuroff, @Ai-chan-0411, and @davidb-tada for their contributions.
Added ci command for CI/CD-optimized test runs: multi-file support, GitHub Actions annotations and step summary, Azure DevOps annotations, --fail-on f
ci command for CI/CD-optimized test runs: multi-file support, GitHub Actions annotations and step summary, Azure DevOps annotations, --fail-on flag, --json output (#1114)changelog command and API endpoint (#1118 @davidb-tada)--all-errors mode for datacontract lint to report all JSON Schema validation errors, with matching all_errors support in the Python library and API (#1125 @jmbenedetto)--schema-name option to custom model export (#978 @AntoineGiraud)double/jsonb, MySQL bare varchar, missing Trino types (#1110)Special thanks to @davidb-tada for the outstanding contribution of the new changelog command and API endpoint! Also thanks to @barry0451 for multiple quality fixes across the SQL exporter and markdown export, and to @AntoineGiraud and @jmbenedetto for their feature contributions.
Escape single quotes in string values for SodaCL checks
logicalTypeOptions.format for SQL import from binary and uuid types (#790)Fix parser error for CSV / Parquet table names containing special characters
NUMERIC(18, 4) (#1083)--output-format json)Fix BigQuery import for repeated fields
<br> with <br /> (#1030)Made duckdb an optional dependency. Install with pip install datacontract-cli[duckdb] for local/S3/GCS/Azure file testing.
duckdb an optional dependency. Install with pip install datacontract-cli[duckdb] for local/S3/GCS/Azure file testing.fastparquet and numpy core dependencies.customProperties or parsing from physicalType (e.g., decimal(10,2)) (#996)Fix datacontract init to generate ODCS format instead of deprecated Data Contract Specification
datacontract init to generate ODCS format instead of deprecated Data Contract Specification (#984)type field by updating open-data-contract-standard to v3.1.2 (#971)required: false are no longer incorrectly treated as required during validation, enabling proper schema evolution where optional fields can be added to contracts without breaking validation of historical data files (#977)type: decimal will be mapped to decimal.Decimal instead of float.Add Impala engine support for Soda scans via ODCS impala server type.
impala server type.This is a major release with breaking changes: We switched the internal data model from Data Contract Specification to Open Data Contract Standard (OD…
This is a major release with breaking changes: We switched the internal data model from Data Contract Specification to Open Data Contract Standard (ODCS).
Not all features that were available are supported in this version, as some features are not supported by the Open Data Contract Standard, such as:
$ref (you can refer to external definitions via authoritativeDefinition)invalidValues)keys and values (use logical type map)scale and precision (define them in physicalType)The reason for this change is that the Data Contract Specification is deprecated, we focus on best possible support for the Open Data Contract Standard. We try to make this transition as seamless as possible. If you face issues, please open an issue on GitHub.
We continue support reading Data Contract Specification data contracts during v0.11.x releases until end of 2026. To migrate existing data contracts to Open Data Contract Standard use this instruction: https://datacontract-specification.com/#migration
--model option to --schema-name in the export command to align with ODCS terminology.*_converter.py to *_exporter.py for consistency (internal change).type: custom and engine: soda with raw SodaCL implementation.service_name attribute access to use ODCS field name serviceNamebreaking, changelog, and diff commands are now deleted (#925).terraform export format has been removed.The breaking, changelog, and diff commands are now deprecated and will be removed in a future version
expectation_suite_name to name in suite outputexpectation_type to type in expectationsdata_asset_type field from suite outputexpectation_type must update to use type{schema} and ${schema} placeholder in SQL quality checks to reference the server's database schema (#957)DATACONTRACT_SQLSERVER_DRIVER environment variable to specify the ODBC driver (#959)2022-01-15) are now kept as strings instead of being converted to datetime objects, fixing ODCS schema validationtable and view fields from custom server importJSON type instead of invalid empty STRUCT() for objects without defined properties (#940)breaking, changelog, and diff commands are now deprecated and will be removed in a future version (#925)### Added - Support for ODCS v3.1.0
Oracle DB: Client Directory for Connection Mode 'Thick' can now be specified in the DATACONTRACT_ORACLE_CLIENT_DIR environment variable
DATACONTRACT_ORACLE_CLIENT_DIR environment variable (#949)Support for Oracle Database (>= 19C)
import: Support for nested arrays in odcs v3 importer
DataContract().import_from_source() as an instance method is now deprecated. Use DataContract.import_from_source() as a class method instead.
DataContract().import_from_source() as an instance method is now deprecated. Use DataContract.import_from_source() as a class method instead.Export to DQX : datacontract export --format dqx
/test endpoint now supports publish_url parameter to publish test results to a URL. (#853)datacontract test now supports HTTP APIs.
datacontract test now supports HTTP APIs.datacontract test now supports Athena.Export to Excel: Convert ODCS YAML to Excel https://github.com/datacontract/open-data-contract-standard-excel-template
Import from Excel: Support the new quality sheet
Added support for Variant with Spark exporter, data_contract.test(), and import as source unity catalog
Excel Import should return ODCS YAML
  instead of   for tab in Markdown export.Support for Data Contract Specification v1.2.0
datacontract import --format json: Import from JSON filesdatacontract api [OPTIONS]: Added option to pass extra arguments for uvicorn.run()pytest tests\test_api.py: Fixed an issue where special characters were not read correctly from file.datacontract import --format json: Import from JSON filesdatacontract api [OPTIONS]: Added option to pass extra arguments for uvicorn.run()pytest tests\test_api.py: Fixed an issue where special characters were not read correctly from file.datacontract export --format mermaid: Fixed an issue where the mermaid export did not handle references correctlyImport anything to ODCS via the import --spec odcs flag
import --spec odcs flagexport --format htmlexport --format mermaidunity importer now supports more than a single table. You can use --unity-table-full-name multiple times to import multiple tables. And it will automatically add a server with the catalog and schema name.datacontract catalog [OPTIONS]: Added version to contract cards in index.html of the catalog (enabled search by version)unity importer no uses the native databricks types instead of relying on spark types. This allows for better type mapping and more accurate data contracts.datacontract export --format mermaid Export to Mermaid (#767, #725)
datacontract export --format mermaid Export to Mermaid (#767, #725)datacontract export --format html: Adding the mermaid figure to the html exportdatacontract export --format odcs: Export physical type to ODCS if the physical type is
configured in config objectdatacontract import --format spark: Added support for spark importer table level comments (#761)datacontract import respects --owner and --id flags (#753)datacontract export --format sodacl: Fix resolving server when using --server flag (#768)datacontract export --format dbt: Fixed DBT export behaviour of constraints to default to data tests when no model type is specified in the datacontract modelDatabricks: Add support for Variant type
datacontract export --format odcs: Export physical type if the physical type is configured in
config object (#757)datacontract export --format sql Include datacontract descriptions in the Snowflake sql export (
#756)Extracted the DataContractSpecification and the OpenDataContractSpecification in separate pip modules and use them in the CLI.
datacontract import --format excel: Import from Excel
template https://github.com/datacontract/open-data-contract-standard-excel-template (#742)Deprecated QualityLinter is now removed
datacontract test with DuckDB: Deep nesting of json objects in duckdb (#681)datacontract import --format csv produces more descriptive output. Replaced
using clevercsv with duckdb for loading and sniffing csv file.datacontract test --output-format junit --output TEST-datacontract.xml Export CLI test results to a file, in a standard format (e.g. JUnit) to improve
datacontract test --output-format junit --output TEST-datacontract.xml Export CLI test results
to a file, in a standard format (e.g. JUnit) to improve CI/CD experience (#650)
Added import for ProtoBuf
Code for proto to datacontract (#696)
dbt & dbt-sources export formats now support the optional --server flag to adapt the DBT column data_type to specific SQL dialects
Duckdb Connections are now configurable, when used as Python library (#666)
export to avro format add map type
isNullable and isUnique (#669)datacontract test --examples: This option was removed as it was not very popular and top-level examples section is deprecated in the Data Contract Spe…
datacontract test now also executes tests for service levels freshness and retention (#407)datacontract import --format sql is now using SqlGlot as importer.datacontract import --format sql --dialect <dialect> Dialect can now to defined when importing
SQL.datacontract test --examples: This option was removed as it was not very popular and top-level
examples section is deprecated in the Data Contract Specification v1.1.0 (#628)odcs_v2 (#645)Your coding agent can read these notes before it upgrades. Set up the MCP server →