NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #629 most downloaded on PyPI
Python APIs for using Delta Lake with Apache Spark
Last release 1 months ago
20 Aug 2026
Release timing varies
gaps range from 3 weeks to 8 months
Nearly every release is documented
notes for 28 of 29 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
34 releases · first in 2021
One column per quarter.
We are excited to announce the release of Delta Lake 4.4.0 . This release adds Apache Spark 4.2 support, expands integration with the UC Delta Table A
We are excited to announce the release of Delta Lake 4.4.0. This release adds Apache Spark 4.2 support, expands integration with the UC Delta Table API, and adds identity-column and generated-column support to Delta Spark.
CREATE TABLE now supports GENERATED ALWAYS AS IDENTITY and GENERATED BY DEFAULT AS IDENTITY.Delta Spark 4.4.0 is built for Apache Spark 4.2.0, Apache Spark 4.1.0, and Apache Spark 4.0.1. As with Apache Spark, Maven artifacts are published for Scala 2.13.
The key features of this release are:
CREATE TABLE now accepts Spark's GENERATED ALWAYS AS IDENTITY and GENERATED BY DEFAULT AS IDENTITY syntax, including custom start and increment values.SHOW PARTITIONS support (#6102): Users can inspect the partitions of a partitioned Delta table with the standard Spark SQL command.VOID column support (#6965, #6966): On Spark 4.1 and later, Delta reads now preserve VOID (NullType) columns instead of failing or silently dropping them. This includes time-travel and path-based reads.Delta Kernel is a set of Java libraries for building Delta connectors that read and write Delta tables without requiring each connector to implement the Delta protocol directly.
The key changes in this release are:
AddFile.partitionValues, allowing partitioned writes to name- and id-mapped tables to round-trip correctly.TIMESTAMP partition values use UTC ISO-8601 strings. Pre-epoch TIMESTAMP_NTZ values are also serialized correctly.Delta UniForm's delta-iceberg module keeps Apache Iceberg metadata synchronized with Delta commits, allowing Iceberg readers to query Delta tables without duplicating data.
Note: In Delta 4.4, delta-iceberg_2.13 supports Spark 4.1 and is not compatible with Spark 4.2.
The key changes in this release are:
saveAsTable overwrite.Delta Sharing is a Spark DataSource that supports batch, streaming, CDF, and time-travel reads on tables shared through the Delta Sharing protocol.
The key changes in this release are:
The Kernel-based delta-flink connector remains experimental. Delta 4.4.0 supports Apache Flink 2.0.2, 2.1.3, 2.2.1, and 2.3.0.
The key changes in this release are:
CREATE TABLE ... WITH ('write.mode' = 'upsert') uses the declared primary key to process Flink changelog rows. INSERT is treated as a new key, UPDATE_AFTER replaces an existing key, DELETE removes it, and UPDATE_BEFORE is ignored. A primary key is required in upsert mode.credentials.source=ambient to use workload identity, instance profiles, application-default credentials, or filesystem configuration instead of Unity Catalog credential vending. Table options beginning with fs. are now passed to the Kernel engine.VOID behavior (#6966): The protocol now defines how readers and writers handle VOID columns, complementing Spark 4.1+'s ability to preserve missing columns in query output.For the complete commit history, see the comparison between v4.3.1 and v4.4.0.
Adam Reeve, Ala Luszczak, Aleksei Shishkin, Alex Moschos, Alexandru Mihai, Amogh Jahagirdar, Annie Wang, Anoop Johnson, Ayush Raj, Bilal Akhtar, Brooks Walls, Charlene Lyu, Chen Wang, ChengJi, Chiin Luen Quah, Cuong Nguyen, Dhruv Arya, Eames Trinh, Eduard Tudenhoefner, Felipe Pessoto, foss-contributor, Gengliang Wang, Hao Jiang, Hao Sun, Hari Lamichhane, Itamar Turner-Trauring, Ivan Sadikov, Jason Chen, Jiayuan Chen, Jinhua Song, Johan Lasperas, Juliusz Sompolski, Kaiqi Jin, Lars Kroll, Leonid Lygin, littlegrasscao, Lukas Rupprecht, Miles Cole, Murali Ramanujam, ni-mi, Nimrod Ofek, Omar Elhadidy, openinx, OussamaSaoudi, Paddy Xu, PhilPlato, Prakhar Jain, Pratham Manja, Rajesh Parangi, Rakesh Veeramacheneni, Ryan Johnson, Ryan Liao, Sandro Sp, seewishnew, songhang, sotikoug83, Steven Yu, Sun Cao, Sushanth Sathish Kumar, Thang Long Vu, Thinh Bui, Timothy Wang, Uros Bojanic, Utkarsh, Xin Huang, Xintong (Oscar) Zhou, Yi Li, You Zhou, Yumingxuan Guo, yyanyy, Zhen Li, Zihao Xu, Ziya Mukhtarov
Nothing published for this version
We are pleased to announce the release of Delta Lake 4.3.1, a patch release on top of 4.3.0 with targeted bug fixes for OAuth authentication in the De
We are pleased to announce the release of Delta Lake 4.3.1, a patch release on top of 4.3.0 with targeted bug fixes for OAuth authentication in the Delta REST Catalog, S3A fast listing through FilterFileSystem wrappers, and UC managed-table metadata handling.
CaseInsensitiveStringMap.entrySet() lowercased camelCase OAuth config keys (e.g. oauth.clientId), breaking Delta REST Catalog authentication.Delta Spark 4.3.1 is built on Apache Spark 4.1.0 and Apache Spark 4.0.1. As with Apache Spark, we publish Maven artifacts for Scala 2.13.
Bug fixes in this release:
delta.enabledFastS3AListFrom with OSS UnityCatalog (>=0.4.1). Because the unitycatalog spark connector introduced CredScopedFileSystem, which relies on S3AFileSystem under the hood, this fix unwraps the internal filesystem of CredScopedFileSystem and handles the casting to enable fast S3A listing.is_managed_location into Delta table metadata: This reserved DSv2 catalog property was leaking into the committed Metadata.configuration at table creation. It is now filtered out at AbstractDeltaCatalog.createTable and synthesized at load for managed tables, matching the behaviour of Spark's V1Table.The Delta Kernel project is a set of Java libraries for building Delta connectors that read and write Delta tables without needing to understand the Delta protocol directly.
No changes to Delta Kernel in this patch release.
Delta UniForm's delta-iceberg and delta-hudi modules automatically keep Apache Iceberg and Apache Hudi metadata in sync with Delta commits, so Iceberg and Hudi readers can query Delta tables without data duplication.
No changes to Delta UniForm in this patch release.
Delta Sharing is a Spark DataSource that lets clients run batch, streaming, CDF, and time-travel reads on tables shared via the Delta Sharing protocol. The 2.13 suffix indicates Scala 2.13.
No changes to Delta Sharing in this patch release.
The Kernel-based delta-flink connector (experimental) is released as part of this patch.
No changes to Delta Flink in this patch release.
Daniel Wang, Murali Ramanujam, Zheng Hu, Rakesh Veeramacheneni, Tathagata Das, Timothy Wang, Vishnu Chandrashekhar, Xin Huang, Yi Li
We are excited to announce the release of Delta Lake 4.3.0, which delivers new features, performance improvements, and protocol updates across Delta S
We are excited to announce the release of Delta Lake 4.3.0, which delivers new features, performance improvements, and protocol updates across Delta Spark, Kernel, UniForm, Sharing, and Flink. See the highlights below for the marquee changes.
replaceOn and replaceUsing DataFrame APIs: Spark now supports selectively replacing table data with the result of a DataFrame. Use replaceUsing to replace rows that match on specified columns, or replaceOn to replace rows that satisfy a user-defined condition.Delta Spark 4.3.0 is built on Apache Spark 4.1.0 and Apache Spark 4.0.1. As with Apache Spark, we publish Maven artifacts for Scala 2.13.
The key features of this release are:
ALTER TABLE updates, are routed through the new API. Non-Delta tables and external tables, including name-based and path-based access, continue to use the legacy delegate.Other notable features and bug-fixes include:
The Delta Kernel project is a set of Java libraries for building Delta connectors that read and write Delta tables without needing to understand the Delta protocol directly.
The key features of this release are:
Other notable changes include:
u<path>@4 instead of u<path>@Optional[4], fixing DV-deduplication mismatches between Java Kernel and Spark/Scala writers.Delta UniForm's delta-iceberg and delta-hudi modules automatically keep Apache Iceberg and Apache Hudi metadata in sync with Delta commits, so Iceberg and Hudi readers can query Delta tables without data duplication.
UniForm now supports both Spark 4.0 and Spark 4.1 with Iceberg-spark 1.11.0.
The key features of this release are:
Other notable changes include:
Delta Sharing is a Spark DataSource that lets clients run batch, streaming, CDF, and time-travel reads on tables shared via the Delta Sharing protocol. Built on Spark and the delta-sharing-client library; the 2.13 suffix indicates Scala 2.13.
The key features of this release are:
The Kernel-based delta-flink connector (experimental) continues to evolve.
The key features of this release are:
Note: Review this section carefully before upgrading to delta 4.3.0.
Alden Lau, Alex Moschos, Amogh Jahagirdar, Anshul Baliga, Bilal Akhtar, Brooks Walls, ChengJi, Chirag Singh, Cuong Nguyen, Dhruv Arya, Divjot Arora, Eames Trinh, Eduard Tudenhoefner, Felipe Pessoto, GH-JamesD, Hao Jiang, Harsh Motwani, Hua Shi, Johan Lasperas, Kaiqi Jin, Leon Windheuser, Leonid Lygin, Liang-Chi Hsieh, Marco Kroll, Matthis Gördel, Murali Ramanujam, Omar Elhadidy, Pratham Manja, Sandro Sp, Sanuj Basu, Scott Sandre, Shivam Tiwari, Stevo Mitric, Thang Long Vu, Timothy Wang, Tom Zhu, Uros Bojanic, Wei Luo, Xin Huang, Yi Li, You Zhou, Yousof Hosny, Zihao Xu, Zikang Han, anniedde, littlegrasscao, Zheng Hu, Rakesh Veeramacheneni, Vishnu Chandrashekhar, songhang, yyanyy
Nothing published for this version
Deprecation of HMS Support: HMS does not support catalog-managed tables. Given the improvements to make metadata generation synchronous, UniForm is mo…
We are excited to announce the release of Delta Lake 4.2.0! This release includes significant new features, improved safety and compatibility, and important bug fixes.
startingTimestamp and skipChangeCommits.delta-flink connector that enables Apache Flink to read, write, and interact with catalog-managed Delta tables.Delta Spark 4.2.0 is built on Apache Spark 4.1.0 and Apache Spark 4.0.1. Similar to Apache Spark, we have released Maven artifacts for Scala 2.13.
The key features of this release are:
startingVersion, startingTimestamp, maxBytesPerTrigger, maxFilesPerTrigger, excludeRegex, skipChangeCommits, ignoreDeletes, ignoreChanges, and ignoreFileDeletion.INSERT ... BY NAME statements now support automatic schema evolution, adding missing columns to the target table when delta.schemaAutoMerge.enabled is set. This brings INSERT BY NAME behavior in line with INSERT SELECT with schema evolution.Other notable changes:
The Delta Kernel project is a set of Java and Rust libraries for building Delta connectors that can read and write to Delta tables without the need to understand the Delta protocol details.
The key features of this release are:
geometry and geography columns, including bounding-box data skipping via the StGeometryBoxesIntersect predicate.This release introduces a brand-new Kernel-based delta-flink connector as an experimental feature that enables Apache Flink to read, write, and interact with Catalog-managed Delta tables.
Maven artifacts:
The Delta Flink connector continues to evolve with Kernel-based implementations. The key features of this release are:
Delta UniForm's delta-iceberg and delta-hudi are the modules that automatically keep Apache Iceberg and Apache Hudi metadata in sync with Delta commits, enabling Iceberg and Hudi readers to query Delta tables without data duplication.
Delta UniForm continues to be supported only for Spark 4.0 in this release. For now, both Hudi and Iceberg remain incompatible with Spark 4.1, as support depends on upcoming releases from those projects providing Spark 4.1-compatible integration bundles. This is unchanged from Delta 4.1.0.
Notable changes:
Delta Sharing is a Spark DataSource that lets clients run batch, streaming, CDF, and time-travel reads on tables shared via the Delta Sharing protocol. Depends on Spark and the delta-sharing-client library. The 2.13 suffix means Scala 2.13.
The key features of this release are:
Alex Moschos, Andrei Tserakhau, Anoop Johnson, Bilal Akhtar, Brooks Walls, ChengJi, Chirag Singh, Cuong Nguyen, Dhruv Arya, Drake Lin, Eames Trinh, Fokko Driesprong, Gengliang Wang, Hao Jiang, Harsh Motwani, Johan Lasperas, Juliusz Sompolski, Kaiqi Jin, Lars Kroll, Leon Windheuser, Leonid Lygin, Liang-Chi Hsieh, Marko Ilić, Milan Stefanovic, Min Yang, Murali Ramanujam, Omar Elhadidy, Prakhar Jain, Rahul Potharaju, Scott Sandre, Sebastien Biollo, Shlok Jhawar, Tathagata Das, Thang Long Vu, Timothy Wang, Vitalii Li, Wei Luo, Xin Huang, Yi Li, You Zhou, Zhen Li, Zheng Hu, Zhipeng Mao, Zihao Xu, Zikang Han, Ziya Mukhtarov, anniedde, emkornfield, giovanni-sorice, littlegrasscao, openinx, richardc-db, seewishnew, songhang, yyanyy
We are excited to announce the release of Delta Lake 4.1.0\! This release includes significant new features, performance improvements, and important p
We are excited to announce the release of Delta Lake 4.1.0! This release includes significant new features, performance improvements, and important platform upgrades.
Delta Spark 4.1.0 is built on Apache Spark 4.1.0 and Apache Spark 4.0.1. Similar to Apache Spark, we have released Maven artifacts for Scala 2.13.
Starting in Delta 4.1.0, Maven artifacts include a Spark version suffix (e.g., delta-spark_4.1_2.13 instead of delta-spark_2.13), with backward compatibility preserved in this release but dependency updates recommended. Separate artifacts are now published for Spark 4.1 and Spark 4.0 so users can choose the version matching their Spark runtime.
Maven artifacts for Spark 4.1.0:
Maven artifacts for Spark 4.0.1:
Backward compatibility artifacts - no spark version in name and work with Spark 4.1.0:
Python artifacts: https://pypi.org/project/delta-spark/4.1.0/
The key features of this release are:
catalogManaged feature, enabling table creation, batch and streaming reads/writes (including time travel, and DML operations), history inspection, and OAuth-based authentication. This is still in preview and production usage is not recommended.MANAGED and EXTERNAL) is now fully atomic, working with UC 0.4.0. Other operations (including REPLACE TABLE, REPLACE TABLE AS SELECT, CREATE OR REPLACE TABLE, Dynamic Partition Overwrite) now fail fast instead of running in best-effort mode.always.Other notable changes:
AddFile stats as struct in checkpoints.The Delta Kernel project is a set of Java and Rust libraries for building Delta connectors that can read and write to Delta tables without the need to understand the Delta protocol details.
The key features of this release are:
Delta UniForm is enabled for Spark 4.0 in this release, restoring full Iceberg interoperability that was listed as a limitation in Delta 4.0.0. Both hudi and iceberg are currently not compatible with Spark 4.1, as support depends on upcoming releases providing Spark 4.1 compatible integration bundles.
The key features of this release are:
The key features of this release are:
VACUUM is blocked for catalog-managed tables, as data lifecycle must be managed through the catalog.Abdelrahman Orief, Ada Ma, AlSchlo, Alden Lau, Aleksei Shishkin, Alex Khakhlyuk, Allison Portis, Amanda Liu, Amogh Jahagirdar, Andreas Chatzistergiou, Anoop Johnson, Anudeep Konaboina, AnudeepKonaboina, Artur Owczarek, Bilal Akhtar, Calvin Qin, Carmen Kwan, ChengJi, ChengJi-db, Chirag Singh, Christos Stavrakakis, Cuong Nguyen, David, Dhruv Arya, Drake Lin, Eames Trinh, Felipe Pessoto, Fred Storage Liu, GH-JamesD, Gene Pang, Gengliang Wang, Hao Jiang, Harsh Motwani, Jake Bellacera, Jerry Zheng, Johan Lasperas, Juliusz Sompolski, Kaiqi Jin, Lars Kroll, Lennart Behme, Leon Windheuser, Lukas Rupprecht, Marco Kroll, Marius Grama, Mark Jarvin, Marko Ilić, Ming DAI, Murali Ramanujam, Noritaka Sekiyama, Omar Elhadidy, OussamaSaoudi, Paddy Xu, Philip Zhu, Qianru Lao, Qiyuan Dong, Rajesh Parangi, Rakesh Veeramacheneni, Robert Dillitz, Robert Pack, Sagar Mittal, Scott Sandre, Sebastian Baunsgaard, Shani Solomon, Sumeet Varma, Tathagata Das, Thang Long Vu, Tim Armstrong, Timothy Wang, Tom Zhu, Tom van Bussel, Tomasz Marzec, Venki Korukanti, Vitalii Li, Vrinda Jindal, Wei Luo, Wenchen Fan, Xi Liang, Xin Huang, Yi Li, Yingyi Bu, You Zhou, Yuchuan Huang, Yufa, Yumingxuan Guo, Yuya Ebihara, Zhen Li, Zhipeng Mao, Zihao Xu, Zikang Han, Ziya Mukhtarov, aleksandr-chernousov-db, anniedde, dengsh12, emkornfield, jiahao-db, ju-klein, littlegrasscao, mollyo-openai, openinx, richardc-db, rqureshi-openai, uros7251brick, yaoforx, Yan Yan
\[Spark\] Breaking change: rename managed table feature from catalogOwned-preview to catalogManaged; legacy ucTableId has also transitioned to the new…
We are excited to announce the release of Delta Lake 4.0.1! This release contains important bug fixes to 4.0.0 and it is recommended that users update to 4.0.1.
catalogOwned-preview to catalogManaged; legacy ucTableId has also transitioned to the new managed-table io.unitycatalog.tableIdauth.* configs; tokens are acquired and refreshed automatically; static tokens remain supported.NoSuchMethodError in REORG TABLE … APPLY (PURGE) when running with Spark 4.0.1.Component-specific bug fixes are detailed below.
Delta Spark 4.0.1 is built on Apache Spark™ 4.0.1. Similar to Apache Spark, we have released Maven artifacts for Scala 2.13.
The key features of this release are:
catalogOwned‑preview is standardized as catalogManaged. The associated Unity Catalog table ID property is updated accordingly (ucTableId → io.unitycatalog.tableId) ; calls that still send the legacy key are handled for compatibility during creation.REORG TABLE … APPLY (PURGE) to fail with NoSuchMethodError on Spark 4.0.1 by switching to the stable constructor and retrieving SQL configs via SparkSession.active.sessionState.conf in executorsOfficial compatibility with UC 0.3.1. Delta 4.0.1 is officially tested with UC integration tests validating end‑to‑end behavior with UC 0.3.1. See the UC 0.3.1 release for the corresponding connector capabilities and APIs.
Support Unity Catalog OAuth authentication. Use catalog‑scoped auth.* configuration; tokens are automatically acquired and refreshed, avoiding embedded static tokens. Legacy static‑token configs continue to work for backward compatibility.
Enable OAuth on a Spark catalog alias that points to UC:
# Point a Spark catalog alias at Unity Catalog
spark.sql.catalog.mycatalog = "io.unitycatalog.connectors.spark.UCSingleCatalog"
spark.sql.catalog.mycatalog.uri = "https://<your-workspace-host>"
# OAuth (dynamic tokens) — supported keys
spark.sql.catalog.mycatalog.auth.type = "oauth"
spark.sql.catalog.mycatalog.auth.oauth.uri = "https://<auth-server-endpoint>"
spark.sql.catalog.mycatalog.auth.oauth.clientId = "<client-id>"
spark.sql.catalog.mycatalog.auth.oauth.clientSecret = "<client-secret>"
# Static token (legacy-compatible)
spark.sql.catalog.mycatalog.auth.type = "static"
spark.sql.catalog.mycatalog.auth.token = "<personal-access-token>"
# Legacy key also supported: spark.sql.catalog.mycatalog.token
And run with Delta’s Spark extensions as usual:
--conf "spark.sql.extensions=io.delta.sql.DeltaSparkSessionExtension" \
--conf "spark.sql.catalog.spark_catalog=org.apache.spark.sql.delta.catalog.DeltaCatalog"
Support Unity Catalog Managed Delta table creation. You can now create UC‑managed Delta tables via standard CREATE TABLE on a UC‑backed Spark catalog; at creation time, Delta sends table properties to the UC server so the server is the source of truth.
Example:
CREATE TABLE mycatalog.my_schema.events (
id BIGINT,
ts TIMESTAMP,
data STRING
)
USING delta
TBLPROPERTIES (
'delta.feature.catalogManaged' = 'supported'
);
Allison Portis, Anudeep Konaboina, Dhruv Arya, Felipe Pessoto, Lukas Rupprecht, Oussama Saoudi, Tathagata Das, Timothy Wang, Yi Li, Hao Jiang, Zheng Hu
These connectors are in maintenance mode and, going forward, will only receive critical security fixes and high-severity bug patches in the 3.x series…
We are excited to announce the final release of Delta Lake 4.0.0! This release includes several exciting new features.
Details by each component.
Currently, Delta Standalone and its dependent connectors, including Delta Flink and Delta Hive, are no longer under active development. Starting in Delta 4.0 we will not be releasing these projects as part of the 4.x Delta releases. These connectors are in maintenance mode and, going forward, will only receive critical security fixes and high-severity bug patches in the 3.x series. We are committed to a full transition from Delta Standalone to Delta Kernel and a future Kernel-based Flink connector.
Delta Spark 4.0 is built on Apache Spark™ 4.0 . Similar to Apache Spark, we have released Maven artifacts for Scala 2.13.
The key features of this release are:
catalogOwned-preview feature enabled. This feature allows a catalog to broker all commits to the table it manages, giving the catalog the control and visibility it needs to prevent invalid operations (e.g. commits that violate foreign key constraints), enforce security and access controls, and opens the door for future performance optimizations. Currently write support includes INSERT, MERGE INTO, UPDATE, and DELETE operations.
catalogOwned-preview feature should not be enabled for production tables and tables created with this preview feature enabled may not be compatible with future Delta Spark releases.DROP FEATURE implementation allows dropping features instantly without truncating history. Dropping a feature introduces a new writer feature to the table, the checkpointProtection feature.
ALTER TABLE table_name DROP FEATURE feature_name
ALTER TABLE table_name DROP FEATURE feature_name TRUNCATE HISTORY
checkpointProtection feature can be dropped with history truncation.Other notable changes include:
deltaTable.dropFeatureSupport.deletionVector table feature.timestampdiff and timestampadd expressions for generated columns.spark.databricks.io.skipping.mdc.sortWithinPartitions (disabled by default) to improve data skipping at the Parquet level.UPDATE and MERGE to resolve struct fields by-name instead of by-position for structs nested inside map types during an update.MERGE APIs to return a DataFrame with the affected rows instead of Unit to align with SQL behavior.The Delta Kernel project is a set of Java and Rust libraries for building Delta connectors that can read and write to Delta tables without the need to understand the Delta protocol details.
The key features of this release are:
TransactionBuilder via calling TransactionBuilder.withLogCompactionInterval.txnBuilder.withClusteringColumns to create a clustered table or update existing clustering columns.deletionVectors, v2Checkpoint, and timestampNtz enabled.generateAppendActions and Kernel will serialize them and write them to the Delta log. These statistics are used in reads to prune files based on query filters.Other notable changes include:
Transaction and TransactionBuilder
txnBuilder.withTableProperties and txnBuilder.withTablePropertiesRemoved.getReadTableVersion.txnBuilder.withMaxRetries.delta.feature.<featureName> to “supported” in the table properties.DefaultExpressionHandler and for data skipping
operationParameters containing non-uniform values would throw an exception.Note: basic schema evolution support via providing an updated schema to the txnBuilder.withSchema method is close to completion and just missed the code cutoff for this release. Look out for this exciting change soon!
In this release of Delta Sharing Spark we have upgraded delta-sharing-client from 1.2.2 to 1.3.2. This enables the following changes:
In Delta Spark, UniForm with Iceberg is unavailable currently due to their lack of support for Spark 4.0. This will be enabled in a future release.
Ada Ma, Ala Luszczak, Alexey Shishkin, Allison Portis, Ami Oka, Amogh Jahagirdar, Andreas Chatzistergiou, Andrei Tserakhau, Andy Lam, Anoop Johnson, Anton Erofeev, Anurag Vaibhav, Bilal Akhtar, Carmen Kwan, Charlene Lyu, ChengJi-db, Chirag Singh, Christos Stavrakakis, Cuong Nguyen, Dhruv Arya, Dušan Tišma, Felipe Pessoto, FredLiu, Gene Pang, Hao Jiang, Harsh Motwani, Herman van Hovell, Jiaheng Tang, Johan Lasperas, Juliusz Sompolski, Jun, Kaiqi Jin, Lars Kroll, Lin Zhou, Livia Zhu, Lukas Rupprecht, Malte Sølvsten Velin, Marko Ilić, Ming Dai, Nick Lanham, Ole Sasse, Omar Elhadidy, Oussama Saoudi, Paddy Xu, Phil Plato, Qiyuan Dong, Rahul Shivu Mahadev, Rajesh Parangi, Rakesh Veeramacheneni, Scott Sandre, Slava Min, Stefan Kandic, Sumeet Varma, Thang Long Vu, Tom van Bussel, Venkata Sai Akhil Gudesa, Venki Korukanti, Vladimir Golubev, Wei Luo, Wenchen Fan, Xiaochong Wu, Xin Huang, Yumingxuan Guo, Ze'ev Maor, Zhipeng Mao, Zihao Xu, Ziya Mukhtarov, chenjian2664, emkornfield, jackierwzhang, kamcheungting-db, littlegrasscao, mozasaur, richardc-db
Previously they were written as INT96 which is an outdated and deprecated format for timestamps.
We are excited to announce the preview release of Delta Lake 4.0.0 on the preview release of Apache Spark 4.0.0! This release gives a preview of the following exciting new features.
Read below for more details. In addition, few existing artifacts are unavailable in this release that are listed at the end.
Delta Spark 4.0 preview is built on Apache Spark™ 4.0.0-preview1. Similar to Apache Spark, we have released Maven artifacts for Scala 2.13.
The key features of this release are:
ALTER TABLE t CHANGE COLUMN col TYPE type command or with schema evolution during MERGE and INSERT operations. See the type widening documentation for a list of all supported type changes and additional information. The table will be readable by Delta 4.0 readers without requiring the data to be rewritten. For compatibility with older versions, a rewrite of the data can be triggered using the ALTER TABLE t DROP FEATURE 'typeWidening' command.Other notable changes include:
CREATE TABLE LIKE with user provided properties. Previously any properties that were provided in the SQL command were ignored and only the properties from the source table were used.Literal expressions would cause an infinite loop when constructing data skipping filters.clock.currentTimeMillis() instead of System.nanoTime() for large commits since some systems return a very small number when System.nanoTime() is called.endOffset for reduced processing time.More features to come in the final release of Delta 4.0!
The Delta Kernel project is a set of Java and Rust libraries for building Delta connectors that can read and write to Delta tables without the need to understand the Delta protocol details.
This release of Delta Kernel Java contains the following changes:
INT64 physical format in Parquet in the DefaultParquetHandler. Previously they were written as INT96 which is an outdated and deprecated format for timestamps.DefaultExpressionHandler. Previously expressions would be eagerly evaluated for every row in the underlying vectors.LIKE in the DefaultExpressionHandler.DefaultParquetHandler.In addition to the above Delta Kernel Java changes, Delta Kernel Rust released its first version 0.1, which is available at https://crates.io/crates/delta_kernel.
The following features from Delta 3.2 are not supported in this preview release. We are working with the community to address the following gaps by the final release of Delta 4.0:
Abhishek Radhakrishnan, Allison Portis, Ami Oka, Andreas Chatzistergiou, Anish, Carmen Kwan, Chirag Singh, Christos Stavrakakis, Dhruv Arya, Felipe Pessoto, Fred Storage Liu, Hyukjin Kwon, James DeLoye, Jiaheng Tang, Johan Lasperas, Jun, Kaiqi Jin, Krishnan Paranji Ravi, Lin Zhou, Lukas Rupprecht, Ole Sasse, Paddy Xu, Prakhar Jain, Qianru Lao, Richard Chen, Sabir Akhadov, Scott Sandre, Sergiu Pocol, Sumeet Varma, Tai Le Manh, Tathagata Das, Thang Long Vu, Tom van Bussel, Venki Korukanti, Wenchen Fan, Yan Zhao, zzl-7
We are pleased to announce the release of Delta Lake 3.3.3, a patch release on top of 3.3.2 with targeted fixes for Delta Sharing cache correctness, t
We are pleased to announce the release of Delta Lake 3.3.3, a patch release on top of 3.3.2 with targeted fixes for Delta Sharing cache correctness, transaction log retention safety, and a Kernel Parquet-reader configuration bug, plus a Delta Sharing client upgrade.
Note: this patch release does not publish delta-iceberg (UniForm/Iceberg support) — see Delta UniForm below for details and follow-up plan.
delta-sharing-client 1.2.2 → 1.2.8: Picks up more resilient OAuth token parsing and improved retry/error-reporting behavior.RANDOMIZE_FILE_PREFIXES support: Wires up the previously-declared delta.randomizeFilePrefixes / delta.randomPrefixLength table properties on the write path, to spread S3 object keys for high-throughput workloads. Opt-in; default behavior unchanged.TRUNCATE TABLE: Spark 3.5.6 tightened V2Writes validation; Delta tables now correctly declare TRUNCATE/OVERWRITE_BY_FILTER capability.Configuration was not being passed when Kernel opened Parquet footers, so custom filesystem/credential settings could be silently ignored.Delta Spark 3.3.3 is built on Apache Spark 3.5.6. Maven artifacts are published for both Scala 2.12 and 2.13.
Fixes and changes in this release:
RANDOMIZE_FILE_PREFIXES support: delta.randomizeFilePrefixes / delta.randomPrefixLength table properties are now honored on writes, placing new data files under a random subdirectory prefix to spread S3 request load. Opt-in; default write path is unchanged.TRUNCATE TABLE: StagedDeltaTableV2 now implements SupportsTruncate and declares TRUNCATE/OVERWRITE_BY_FILTER capabilities, fixing TRUNCATE TABLE under Spark 3.5.6's stricter V2Writes validation.The Delta Kernel project is a set of Java libraries for building Delta connectors without needing to implement the Delta protocol directly.
Fixes in this release:
IllegalAccessError fix had dropped the Hadoop Configuration when constructing the ParquetReader used to read footers, so custom Hadoop filesystem settings could be ignored during Kernel reads.Delta UniForm's delta-iceberg and delta-hudi modules keep Apache Iceberg and Apache Hudi metadata in sync with Delta commits.
delta-iceberg is not published in 3.3.3. There will be no delta-iceberg artifact in this release. There are no changes to delta-iceberg from the previous version. Existing delta-iceberg_2.12/delta-iceberg_2.13 v3.3.2 artifacts remain available and are forward-compatible with delta-spark 3.3.3.No changes to Delta Hudi in this patch release (rebuilt against the updated build toolchain only).
Delta Sharing is a Spark DataSource for batch, streaming, CDF, and time-travel reads on tables shared via the Delta Sharing protocol.
Fixes and changes in this release:
delta-sharing-client to 1.2.8 (from 1.2.2): brings more resilient OAuth expires_in parsing, adaptive retry/backoff on retryable server errors, and clearer error reporting from streaming responses.No changes to Delta Standalone in this patch release.
No changes to the Delta Hive connector in this patch release.
The Kernel-based delta-flink connector (experimental).
No changes to Delta Flink in this patch release.
Daniel Mattos, Felipe Pessoto, Geeta Krishna Panda, littlegrasscao, Sun Cao, Venki Korukanti, Vishnu Chandrashekhar, Yingyi Bu
We are excited to announce the release of Delta Lake 3.3.2! This release contains several important bug fixes and improvements to the 3.3.1 release an
We are excited to announce the release of Delta Lake 3.3.2! This release contains several important bug fixes and improvements to the 3.3.1 release and it is recommended that users upgrade to 3.3.2.
Component specific bug fixes are detailed below.
Delta Spark 3.3.2 is built on Apache Spark™ 3.5.3. Similarly to Apache Spark, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
The key fixes in this release are:
The key fixes in this release are:
The key fixes in this release are:
Dhruv Arya, Prakhar Jain, Venkateshwar Korukanti, Scott Sandre
We are excited to announce the release of Delta Lake 3.3.1! This release contains a few bug fixes to the 3.3.0 release and it is recommended that user
We are excited to announce the release of Delta Lake 3.3.1! This release contains a few bug fixes to the 3.3.0 release and it is recommended that users upgrade to 3.3.1.
Component specific bug fixes are detailed below.
Delta Spark 3.3.1 is built on Apache Spark™ 3.5.3. Similarly to Apache Spark, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
The key fixes in this release are:
The key fixes in this release are:
No fixes or changes were made in the components below in this release but the corresponding artifacts are listed.
Wenchen Fan, Thang Long Vu
Previously they were written as INT96 which is an outdated and deprecated format for timestamps.
We are excited to announce the release of Delta Lake 3.3.0! This release includes several exciting new features.
Details by each component.
Delta Spark 3.3.0 is built on Apache Spark™ 3.5.3. Similarly to Apache Spark, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
The key features of this release are:
_metadata.row_id and _metadata.row_commit_version. Refer to the documentation on Row Tracking for more information and examples.Other notable changes include:
delta.enableChangeDataFeed=true) the protocol is upgraded to (1,7) and only the legacy feature is enabled. Previously the minimum protocol version would be selected and all preceding legacy features enabled.deltaTable.addFeatureSupport(...).spark.databricks.delta.skipping.partitionLikeFilters.enabled, applies arbitrary data skipping filters referencing Liquid clustering columns to files with the same min and max values on clustering columns. This may decrease the files scanned for selective queries on large Liquid tables.spark.databricks.delta.optimize.batchSize.delta.dataSkippingStatsColumns. Previously this would throw an exception._commit_timestamp column using In Commit Timestamps in CDF reads when ICT is enabled on a table.Literal expressions would cause an infinite loop when constructing data skipping filters.clock.currentTimeMillis() instead of System.nanoTime() for large commits since some systems return a very small number when System.nanoTime() is called._commit_timestamp column in Change Data Feed when reading from timezones other than UTC.You can now enable UniForm Iceberg on existing Delta tables without rewriting data files. You can then seamlessly read the table downstream in Iceberg clients such as Spark and Snowflake. See Enable by altering an existing table.
Other notable changes include:
The Delta Kernel project is a set of Java and Rust libraries for building Delta connectors that can read and write to Delta tables without the need to understand the Delta protocol details.
This release of Delta Kernel Java contains the following changes:
Snapshot.getTimestamp.DefaultParquetHandler.DefaultParquetHandler. Previously they were written as INT96 which is an outdated and deprecated format for timestamps.DefaultParquetHandler.DefaultExpressionHandler. Previously expressions would be eagerly evaluated for every row in the underlying vectors.DefaultExpressionHandler to use UTF8.DefaultExpressionHandler.Abhishek Radhakrishnan, Adam Binford, Alden Lau, Aleksei Shishkin, Alexey Shishkin, Allison Portis, Ami Oka, Amogh Jahagirdar, Andreas Chatzistergiou, Andrew Xue, Anish, Annie Wang, Avril Aysha, Bart Samwel, Burak Yavuz, Carmen Kwan, Charlene Lyu, ChengJi-db, Chirag Singh, Christos Stavrakakis, Cuong Nguyen, Dhruv Arya, Eduard Tudenhoefner, Felipe Pessoto, Fokko Driesprong, Fred Storage Liu, Hao Jiang, Hyukjin Kwon, Jacek Laskowski, Jackie Zhang, Jade Wang, James DeLoye, Jiaheng Tang, Jintao Shen, Johan Lasperas, Juliusz Sompolski, Jun, Jungtaek Lim, Kaiqi Jin, Kam Cheung Ting, Krishnan Paranji Ravi, Lars Kroll, Leon Windheuser, Lin Zhou, Liwen Sun, Lukas Rupprecht, Marko Ilić, Matt Braymer-Hayes, Maxim Gekk, Michael Zhang, Ming DAI, Mingkang Li, Nils Andre, Ole Sasse, Paddy Xu, Prakhar Jain, Qianru Lao, Qiyuan Dong, Rahul Shivu Mahadev, Rajesh Parangi, Rakesh Veeramacheneni, Richard Chen, Richard-code-gig, Robert Dillitz, Robin Moffatt, Ryan Johnson, Sabir Akhadov, Scott Sandre, Sergiu Pocol, Shawn Chang, Shixiong Zhu, Sumeet Varma, Tai Le Manh, Taiga Matsumoto, Tathagata Das, Thang Long Vu, Tom van Bussel, Tulio Cavalcanti, Venki Korukanti, Vishwas Modhera, Wenchen Fan, Yan Zhao, YotillaAntoni, Yumingxuan Guo, Yuya Ebihara, Zhipeng Mao, Zihao Xu, zzl-7
Write timestamp data to Parquet file as INT64 physical format instead of INT96 physical format. INT96 is a legacy physical format that is deprecated.
We are excited to announce the release of Delta Lake 3.2.1! This release contains important bug fixes to 3.2.0 and it is recommended that users upgrade to 3.2.1.
Details by each component.
Delta Spark 3.2.1 is built on Apache Spark™ 3.5.3. Similar to Apache Spark, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
The key changes of this release are:
The key changes of this release are:
The key changes of this release are:
The key changes of this release are:
For more information, refer to:
This release does not update Standalone. Standalone is being deprecated in favor of Delta Kernel, which supports advanced features in Delta tables.
Artifacts: delta-storage, delta-storage-s3-dynamodb
The key changes of this release are:
Abhishek Radhakrishnan, Allison Portis, Charlene Lyu, Fred Storage Liu, Jiaheng Tang, Johan Lasperas, Lin Zhou, Marko Ilić, Scott Sandre, Tathagata Das, Tom van Bussel, Venki Korukanti, Wenchen Fan, Zihao Xu
We are excited to announce the release of Delta Lake 3.2.0! This release includes several exciting new features.
We are excited to announce the release of Delta Lake 3.2.0! This release includes several exciting new features.
Delta Spark 3.2.0 is built on Apache Spark™ 3.5. Similar to Apache Spark, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
The key features of this release are:
clusterBy API in both Python and Scala to allow creating clustered tables using DeltaTable API. See the documentation and examples for more information.byte to short to integer using the ALTER TABLE t CHANGE COLUMN col TYPE type command or with schema evolution during MERGE and INSERT operations. The table remains readable by Delta 3.2 readers without requiring the data to be rewritten. For compatibility with older versions, a rewrite of the data can be triggered using the ALTER TABLE t DROP FEATURE 'typeWidening-preview’ command.
vacuumProtocolCheck ReaderWriter feature which ensures consistent application of reader and writer protocol checks during VACUUM operations, addressing potential protocol discrepancies and mitigating the risk of data corruption due to skipped writer checks._metadata.row_id and _metadata.row_commit_version.Other notable changes include:
spark.databricks.delta.delta.log.cacheSize) and retention duration (spark.databricks.delta.delta.log.cacheRetentionMinutes)Hudi is now supported by Delta Universal format in addition to Iceberg. Writing to a Delta UniForm table can generate Hudi metadata, alongside Delta. This feature is contributed by XTable.
Create a UniForm-enabled that automatically generates Hudi metadata using the following command:
CREATE TABLE T (c1 INT) USING DELTA TBLPROPERTIES ('delta.universalFormat.enabledFormats' = hudi);
See the documentation here for more details.
Other notable changes include:
The Delta Kernel project is a set of Java libraries (Rust will be coming soon!) for building Delta connectors that can read (and, soon, write to) Delta tables without the need to understand the Delta protocol details). In this release,e we improved the read support to make it production-ready by adding numerous performance improvements, additional functionality, and improved protocol support.
Support for time travel. Now you can read a table snapshot at a version id or snapshot at a timestamp.
Improved Delta protocol support.
checkpoint v2.timestamp partition type data column.timestamp_ntz.Improved table metadata read performance and reliability on very large tables with millions of files
LogStores from delta-storage module for faster listFrom calls._last_checkpoint checkpoint in case of transient failures. Loading the last checkpoint info from this file helps construct the Delta table state faster.Other notable changes include:
IS_NULL expression. Now the Predicate passed to Kernel ScanBuilder can include IS_NULL predicates.ParquetHandler implementations to multiple Parquet files in parallel. The current default implementation reads one file at a time, but the connectors can implement their own custom ParquetHandler to read the Parquet files in parallel.In this release we also added preview version of APIs that allows connectors to:
For more information, refer to:
Adam Binford, Ala Luszczak, Allison Portis, Ami Oka, Andreas Chatzistergiou, Arun Ravi M V, Babatunde Micheal Okutubo, Bo Gao, Carmen Kwan, Chirag Singh, Chloe Xia, Christos Stavrakakis, Costas Zarifis, Daniel Tenedorio, Davin Tjong, Dhruv Arya, Felipe Pessoto, Fred Storage Liu, Fredrik Klauss, Gabriel Russo, Hao Jiang, Hyukjin Kwon, Ian Streeter, Jason Teoh, Jiaheng Tang, Jing Zhan, Jintian Liang, Johan Lasperas, Jonas Irgens Kylling, Juliusz Sompolski, Kaiqi Jin, Lars Kroll, Lin Zhou, Miles Cole, Nick Lanham, Ole Sasse, Paddy Xu, Prakhar Jain, Rachel Bushrian, Rajesh Parangi, Renan Tomazoni Pinzon, Sabir Akhadov, Scott Sandre, Simon Dahlbacka, Sumeet Varma, Tai Le, Tathagata Das, Thang Long Vu, Tim Brown, Tom van Bussel, Venki Korukanti, Wei Luo, Wenchen Fan, Xupeng Li, Yousof Hosny, Gene Pang, Jintao Shen, Kam Cheung Ting, panbingkun, ram-seek, Sabir Akhadov, sokolat, tangjiafu
We are excited to announce the release of Delta Lake 3.1.0. This release includes several exciting new features.
We are excited to announce the release of Delta Lake 3.1.0. This release includes several exciting new features.
Details by each component.
Delta Spark 3.1.0 is built on Apache Spark™ 3.5. Similar to Apache Spark, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
The key features of this release are:
Other notable changes include:
This release of Delta adds a new module called delta-sharing-spark which enables reading Delta tables shared using the Delta Sharing protocol in Apache Spark™. It is migrated from https://github.com/delta-io/delta-sharing/tree/main/spark repository to https://github.com/delta-io/delta/tree/master/sharing repository. Last release version of delta-sharing-spark is 1.0.4 from the previous location. Next release of delta-sharing-spark is with the current release of Delta which is 3.1.0.
Supported read types are: read snapshot of the table, incrementally read the table using streaming or read the changes (Change Data Feed) between two versions of the table.
“Delta Format Sharing” is newly introduced since delta-sharing-spark 3.1, which supports reading shared Delta tables with advanced Delta features such as deletion vectors and column mapping.
Below is an example of reading a Delta table shared using the Delta Sharing protocol in a Spark environment. For more examples refer to the documentation.
import org.apache.spark.sql.SparkSession
val spark = SparkSession
.builder()
.appName("...")
.master("...")
.config(
"spark.sql.extensions",
"io.delta.sql.DeltaSparkSessionExtension"
).config(
"spark.sql.catalog.spark_catalog",
"org.apache.spark.sql.delta.catalog.DeltaCatalog"
).getOrCreate()
val tablePath = "<profile-file-path>#<share-name>.<schema-name>.<table-name>"
// Batch query
spark.read
.format("deltaSharing")
.option("responseFormat", "delta")
.load(tablePath)
.show(10)
Delta Universal Format (UniForm) allows you to read Delta tables from Iceberg and Hudi (coming soon) clients. Delta 3.1.0 provided the following improvements:
LIST and MAP data types and improves compatibility with popular Iceberg reader clients.The Delta Kernel project is a set of Java libraries (Rust will be coming soon!) for building Delta connectors that can read (and, soon, write to) Delta tables without the need to understand the Delta protocol details).
id mode. Now tables with column mapping id mode can be read by Kernel.slf4j loggingFor more information, refer to:
The key features of this release are
There are no updates to Standalone in this release.
Ala Luszczak, Allison Portis, Ami Oka, Amogh Akshintala, Andreas Chatzistergiou, Bart Samwel, BjarkeTornager, Christos Stavrakakis, Costas Zarifis, Daniel Tenedorio, Dhruv Arya, EJ Song, Eric Maynard, Felipe Pessoto, Fred Storage Liu, Fredrik Klauss, Gengliang Wang, Gerhard Brueckl, Haejoon Lee, Hao Jiang, Jared Wang, Jiaheng Tang, Jing Wang, Johan Lasperas, Kaiqi Jin, Kam Cheung Ting, Lars Kroll, Li Haoyi, Lin Zhou, Lukas Rupprecht, Mark Jarvin, Max Gekk, Ming DAI, Nick Lanham, Ole Sasse, Paddy Xu, Patrick Leahey, Peter Toth, Prakhar Jain, Renan Tomazoni Pinzon, Rui Wang, Ryan Johnson, Sabir Akhadov, Scott Sandre, Serge Rielau, Shixiong Zhu, Tathagata Das, Thang Long Vu, Tom van Bussel, Venki Korukanti, Vitalii Li, Wei Luo, Wenchen Fan, Xin Zhao, jintao shen, panbingkun
The repository https://github.com/delta-io/connectors is now deprecated.
We are excited to announce the final release of Delta Lake 3.0.0. This release includes several exciting new features and artifacts.
Here are the most important aspects of 3.0.0:
Unlike the initial preview release, Delta Spark is now built on top of Apache Spark™ 3.5. See the Delta Spark section below for more details.
Delta Universal Format (UniForm) will allow you to read Delta tables with Hudi and Iceberg clients. Iceberg support is available with this release. UniForm takes advantage of the fact that all table storage formats, such as Delta, Iceberg, and Hudi, actually consist of Parquet data files and a metadata layer. In this release, UniForm automatically generates Iceberg metadata and commits to Hive metastore, allowing Iceberg clients to read Delta tables as if they were Iceberg tables. Create a UniForm-enabled table using the following command:
CREATE TABLE T (c1 INT) USING DELTA TBLPROPERTIES (
'delta.universalFormat.enabledFormats' = 'iceberg');
Every write to this table will automatically keep Iceberg metadata updated. See the documentation here for more details, and the key implementations here and here.
The Delta Kernel project is a set of Java libraries (Rust will be coming soon!) for building Delta connectors that can read (and, soon, write to) Delta tables without the need to understand the Delta protocol details).
You can use this library to do the following:
Reading a Delta table with Kernel APIs is as follows.
TableClient myTableClient = DefaultTableClient.create() ; // define a client
Table myTable = Table.forPath(myTableClient, "/delta/table/path"); // define what table to scan
Snapshot mySnapshot = myTable.getLatestSnapshot(myTableClient); // define which version of table to scan
Predicate scanFilter = ... // define the predicate
Scan myScan = mySnapshot.getScanBuilder(myTableClient) // specify the scan details
.withFilters(scanFilter)
.build();
Scan.readData(...) // returns the table data
Full example code can be found here.
For more information, refer to:
This release of Delta contains the Kernel Table API and default TableClient API definitions and implementation which allow:
All previous connectors from https://github.com/delta-io/connectors have been moved to this repository (https://github.com/delta-io/delta) as we aim to unify our Delta connector ecosystem structure. This includes Delta-Standalone, Delta-Flink, Delta-Hive, PowerBI, and SQL-Delta-Import. The repository https://github.com/delta-io/connectors is now deprecated.
Delta Spark 3.0.0 is built on top of Apache Spark™ 3.5. Similar to Apache Spark, we have released Maven artifacts for both Scala 2.12 and Scala 2.13. Note that the Delta Spark maven artifact has been renamed from delta-core to delta-spark.
The key features of this release are:
Path.toUri.toString calls for each row in a table, resulting in a several hundred percent speed boost on DELETE operations (only when Deletion Vectors have been enabled on the table).DROP COLUMN and RENAME COLUMN have been used. This includes streaming support for Change Data Feed. See the documentation here for more details.delta.dataSkippingStatsColumns. Previously, Delta would only collect file-skipping statistics for the first N columns in the table schema (default to 32). Now, users can easily customize this.CONVERT TO DELTA. This feature was excluded from the Delta Lake 2.4 release since Iceberg did not yet support Apache Spark 3.4 (or 3.5). This command generates a Delta table in the same location and does not rewrite any parquet files.spark.sql.storeAssignmentPolicy instead of spark.sql.ansi.enabled.Other notable changes include
OPTIMIZE <catalog>.<db>.<tbl> will work.overwriteSchema when partitionOverwriteMode is set to dynamicdata subdirectory via the SQL configuration spark.databricks.delta.write.dataFilesToSubdir. This is used to add UniForm support on BigQuery.Delta-Flink 3.0.0 is built on top of Apache Flink™ 1.16.1.
The key features of this release are
Other notable changes include
The key features in this release are:
io.delta.standalone.checkpointing.enabled to false. This is only safe and suggested to do if another job will periodically perform the checkpointing.SHALLOW CLONEs and create Delta tables with external files.Adam Binford, Ahir Reddy, Ala Luszczak, Alex, Allen Reese, Allison Portis, Ami Oka, Andreas Chatzistergiou, Animesh Kashyap, Anonymous, Antoine Amend, Bart Samwel, Bo Gao, Boyang Jerry Peng, Burak Yavuz, CabbageCollector, Carmen Kwan, ChengJi-db, Christopher Watford, Christos Stavrakakis, Costas Zarifis, Denny Lee, Desmond Cheong, Dhruv Arya, Eric Maynard, Eric Ogren, Felipe Pessoto, Feng Zhu, Fredrik Klauss, Gengliang Wang, Gerhard Brueckl, Gopi Krishna Madabhushi, Grzegorz Kołakowski, Hang Jia, Hao Jiang, Herivelton Andreassa, Herman van Hovell, Jacek Laskowski, Jackie Zhang, Jiaan Geng, Jiaheng Tang, Jiawei Bao, Jing Wang, Johan Lasperas, Jonas Irgens Kylling, Jungtaek Lim, Junyong Lee, K.I. (Dennis) Jung, Kam Cheung Ting, Krzysztof Chmielewski, Lars Kroll, Lin Ma, Lin Zhou, Luca Menichetti, Lukas Rupprecht, Martin Grund, Min Yang, Ming DAI, Mohamed Zait, Neil Ramaswamy, Ole Sasse, Olivier NOUGUIER, Pablo Flores, Paddy Xu, Patrick Pichler, Paweł Kubit, Prakhar Jain, Pulkit Singhal, RunyaoChen, Ryan Johnson, Sabir Akhadov, Satya Valluri, Scott Sandre, Shixiong Zhu, Siying Dong, Son, Tathagata Das, Terry Kim, Tom van Bussel, Venki Korukanti, Wenchen Fan, Xinyi, Yann Byron, Yaohua Zhao, Yijia Cui, Yuhong Chen, Yuming Wang, Yuya Ebihara, Zhen Li, aokolnychyi, gurunath, jintao shen, maryannxue, noelo, panbingkun, windpiger, wwang-talend, sherlockbeard
Nothing published for this version
We are excited to announce the release of Delta Lake 2.4.0 on Apache Spark 3.4. Similar to Apache Spark™, we have released Maven artifacts for both Sc
We are excited to announce the release of Delta Lake 2.4.0 on Apache Spark 3.4. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
DELETE command. Previously, when deleting rows from a Delta table, any file with at least one matching row would be rewritten. With Deletion Vectors these expensive rewrites can be avoided. See What are deletion vectors? for more details.PURGE to remove Deletion Vectors from the current version of a Delta table by rewriting any data files with deletion vectors. See the documentation for more details.REPLACE WHERE expressions in SQL to selectively overwrite data. Previously “replaceWhere” options were only supported in the DataFrameWriter APIs.WHEN NOT MATCHED BY SOURCE clauses in SQL for the Merge command.INSERT INTO queries. Delta will automatically generate the values for any unspecified generated columns.TimestampNTZ data type added in Spark 3.3. Using TimestampNTZ requires a Delta protocol upgrade; see the documentation for more information.char or varchar column to a compatible type in the ALTER TABLE command. The new behavior is the same as in Apache Spark and allows upcasting from char or varchar to varchar or string.overwriteSchema with dynamic partition overwrite. This can corrupt the table as not all the data may be removed, and the schema of the newly written partitions may not match the schema of the unchanged partitions.DataFrame for Change Data Feed reads when there are no commits within the timestamp range provided. Previously an error would be thrown.Note: the Delta Lake 2.4.0 release does not include the Iceberg to Delta converter because iceberg-spark-runtime does not support Spark 3.4 yet. The Iceberg to Delta converter is still supported when using Delta 2.3 with Spark 3.3.
Alkis Evlogimenos, Allison Portis, Andreas Chatzistergiou, Anton Okolnychyi, Bart Samwel, Bo Gao, Carl Fu, Chaoqin Li, Christos Stavrakakis, David Lewis, Desmond Cheong, Dhruv Shah, Eric Maynard, Fred Liu, Fredrik Klauss, Haejoon Lee, Hussein Nagree, Jackie Zhang, Jintian Liang, Johan Lasperas, Lars Kroll, Lukas Rupprecht, Matthew Powers, Ming DAI, Ming Dai, Naga Raju Bhanoori, Paddy Xu, Prakhar Jain, Rahul Shivu Mahadev, Rui Wang, Ryan Johnson, Sabir Akhadov, Satya Valluri, Scott Sandre, Shixiong Zhu, Tom van Bussel, Venki Korukanti, Vitalii Li, Wenchen Fan, Xi Liang, Yaohua Zhao, Yuming Wang
We are excited to announce the release of Delta Lake 2.3.0 on Apache Spark 3.3. Similar to Apache Spark™, we have released Maven artifacts for both Sc
We are excited to announce the release of Delta Lake 2.3.0 on Apache Spark 3.3. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
CONVERT TO DELTA. This generates a Delta table in the same location and does not rewrite any parquet files. See the documentation for details.SHALLOW CLONE for Delta, Parquet, and Iceberg tables to clone a source table without copying the data files. SHALLOW CLONE creates a copy of the source table’s definition but refers to the source table’s data files.INSERT/DELETE/UPDATE/MERGE etc. operations using SQL configurations spark.databricks.delta.write.txnAppId and spark.databricks.delta.write.txnVersion.DeltaTable APIs. SQL Support will be added in Spark 3.4.CREATE TABLE LIKE to create empty Delta tables using the definition and metadata of an existing table or view.table_changes table-valued function.DROP COLUMN and RENAME COLUMN have been used. See the documentation for more details.delta.enableFastS3AListFrom to true to enable it.VACUUM operations in the transaction log. With this feature, VACUUM operations and their associated metrics (e.g. numDeletedFiles) will now show up in table history.MERGE for UPDATE SET <assignments> and INSERT (...) VALUES (...) actions. Previously, schema evolution was only supported for UPDATE SET * and INSERT * actions..show() support for COUNT(*) aggregate pushdown.df.saveAsTable for overwrite and append mode.trunc and date_trunc functions.date_format function with format yyyy-MM-dd.replaceWhere with the DataFrame V2 overwrite API to correctly evaluate less than conditions.INSERT OVERWRITE with complex data types when the source schema is read incompatible.VACUUM where sometimes the default retention period was used to remove files instead of the retention period specified in the table properties.deltaTable.details() Python/Scala/Java API.VACUUM table_name DRY RUN.Allison Portis, Andreas Chatzistergiou, Andrew Li, Bo Zhang, Brayan Jules, Burak Yavuz, Christos Stavrakakis, Daniel Tenedorio, Dhruv Shah, Felipe Pessoto, Fred Liu, Fredrik Klauss, Gengliang Wang, Haejoon Lee, Hussein Nagree, Jackie Zhang, Jiaheng Tang, Jintian Liang, Johan Lasperas, Jungtaek Lim, Kam Cheung Ting, Koki Otsuka, Lars Kroll, Lin Ma, Lukas Rupprecht, Ming DAI, Mitchell Riley, Ole Sasse, Paddy Xu, Prakhar Jain, Pranav, Rahul Shivu Mahadev, Rajesh Parangi, Ryan Johnson, Scott Sandre, Serge Rielau, Shixiong Zhu, Slim Ouertani, Tobias Fabritz, Tom van Bussel, Tushar Machavolu, Tyson Condie, Venki Korukanti, Vitalii Li, Wenchen Fan, Xinyi Yu, Yaohua Zhao, Yingyi Bu
We are excited to announce the release of Delta Lake 2.2.0 on Apache Spark 3.3. Similar to Apache Spark™, we have released Maven artifacts for both Sc
We are excited to announce the release of Delta Lake 2.2.0 on Apache Spark 3.3. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
The key features in this release are as follows:
LIMIT pushdown into Delta scan. Improve the performance of queries containing LIMIT clauses by pushing down the LIMIT into Delta scan during query planning. Delta scan uses the LIMIT and the file-level row counts to reduce the number of files scanned which helps the queries read far less number of files and could make LIMIT queries faster by 10-100x depending upon the table size.
Aggregate pushdown into Delta scan for SELECT COUNT(*). Aggregation queries such as SELECT COUNT(*) on Delta tables are satisfied using file-level row counts in Delta table metadata rather than counting rows in the underlying data files. This significantly reduces the query time as the query just needs to read the table metadata and could make full table count queries faster by 10-100x.
Support for collecting file level statistics as part of the CONVERT TO DELTA command. These statistics potentially help speed up queries on the Delta table. By default the statistics are collected now as part of the CONVERT TO DELTA command. In order to disable statistics collection specify NO STATISTICS clause in the command. Example: CONVERT TO DELTA table_name NO STATISTICS
Improve performance of the DELETE command by pruning the columns to read when searching for files to rewrite.
Fix for a bug in the DynamoDB-based S3 multi-cluster mode configuration. The previous version wrote an incorrect timestamp which was used by DynamoDB’s TTL feature to cleanup expired items. This timestamp value has been fixed and the table attribute renamed from commitTime to expireTime. If you already have TTL enabled, please follow the migration steps here.
Fix non-deterministic behavior during MERGE when working with sources that are non-deterministic.
Remove the restrictions for using Delta tables with column mapping in certain Streaming + CDF cases. Earlier we used to block Streaming+CDF if the Delta table has column mapping enabled even though it doesn’t contain any RENAME or DROP columns.
Other notable changes
where() calls in Optimize scala/python API. or _ in CONVERT TO DELTA command.MERGE INTO when there are multiple UPDATE clauses and one of the UPDATEs is with a schema evolution.SparkSession object is not found when using Delta APIslast_checkpoint file fails.AvailableNow trigger on a Delta table.Credits Abhishek Somani, Adam Binford, Allison Portis, Amir Mor, Andreas Chatzistergiou, Anish Shrigondekar, Carl Fu, Carlos Peña ,Chen Shuai, Christos Stavrakakis, Eric Maynard, Fabian Paul, Felipe Pessoto, Fredrik Klauss, Ganesh Chand, Hedi Bejaoui, Helge Brügner, Hussein Nagree, Ionut Boicu, Jackie Zhang, Jiaheng Tang, Jintao Shen, Jintian Liang, Joe Harris, Johan Lasperas, Jonas Irgens Kylling, Josh Rosen, Juliusz Sompolski, Jungtaek Lim, Kam Cheung Ting, Karthik Subramanian, Kevin Neville, Lars Kroll, Lin Ma, Linhong Liu, Lukas Rupprecht, Max Gekk, Ming Dai, Mingliang Zhu, Nick Karpov, Ole Sasse, Paddy Xu, Patrick Marx, Prakhar Jain, Pranav, Rajesh Parangi, Ronald Zhang, Ryan Johnson, Sabir Akhadov, Scott Sandre, Serge Rielau, Shixiong Zhu, Supun Nakandala, Thang Long Vu, Tom van Bussel, Tyson Condie, Venki Korukanti, Vitalii Li, Weitao Wen, Wenchen Fan, Xinyi, Yuming Wang, Zach Schuermann, Zainab Lawal, sherlockbeard (github id)
We are excited to announce the release of Delta Lake 2.1.1 on Apache Spark 3.3. This release contains important bug fixes to 2.1.0 and it is recommend
We are excited to announce the release of Delta Lake 2.1.1 on Apache Spark 3.3. This release contains important bug fixes to 2.1.0 and it is recommended that users update to 2.1.1. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
This release includes the following bug fixes and improvements:
commitTime to expireTime. If you already have TTL enabled, please follow the migration steps here.Credits Adam Binford, Allison Portis, Chen Shuai, Felipe Pessoto, Lars Kroll, Scott Sandre, Shixiong Zhu, Venki Korukanti
We are excited to announce the release of Delta Lake 2.1.0 on Apache Spark 3.3. Similar to Apache Spark™, we have released Maven artifacts for both Sc
We are excited to announce the release of Delta Lake 2.1.0 on Apache Spark 3.3. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
repartition(1) instead of coalesce(1) in Optimize for better performance when compacting many small files.DeltaTableBuilder to preserve table property case of non-delta properties when setting properties.replaceWhere option.Improvements to the benchmark framework (initial version added in version 1.2.0) including support for benchmarking arbitrary functions and not just SQL queries. We’ve also added Terraform scripts to automatically generate the infrastructure to run benchmarks on AWS and GCP.
Adam Binford, Allison Portis, Andreas Chatzistergiou, Andrew Vine, Andy Lam, Carlos Peña, Chang Yong Lik, Christos Stavrakakis, David Lewis, Denis Krivenko, Denny Lee, EJ Song, Edmondo Porcu, Felipe Pessoto, Fred Liu, Fu Chen, Grzegorz Kołakowski, Hedi Bejaoui, Hussein Nagree, Ionut Boicu, Ivan Sadikov, Jackie Zhang, Jiawei Bao, Jintao Shen, Jintian Liang, Jonas Irgens Kylling, Juliusz Sompolski, Junlin Zeng, KaiFei Yi, Kam Cheung Ting, Karen Feng, Koert Kuipers, Lars Kroll, Lin Zhou, Lukas Rupprecht, Max Gekk, Min Yang, Ming DAI, Nick, Ole Sasse, Prakhar Jain, Rahul Shivu Mahadev, Rajesh Parangi, Rui Wang, Ryan Johnson, Sabir Akhadov, Scott Sandre, Serge Rielau, Shixiong Zhu, Tathagata Das, Terry Kim, Thomas Newton, Tom van Bussel, Tyson Condie, Venki Korukanti, Vini Jaiswal, Will Jones, Xi Liang, Yijia Cui, Yousry Mohamed, Zach Schuermann, sherlockbeard, yikf
We are excited to announce the release of Delta Lake 2.0.2 on Apache Spark 3.2. This release contains important bug fixes and a few high-demand usabil
We are excited to announce the release of Delta Lake 2.0.2 on Apache Spark 3.2. This release contains important bug fixes and a few high-demand usability improvements over 2.0.1 and it is recommended that users update to 2.0.2. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
This release includes the following bug fixes and improvements:
numDeletedFiles) will now show up in table history.spark.databricks.delta.write.txnAppId and spark.databricks.delta.write.txnVersion.
Support passing Hadoop configurations via DeltaTable APIfrom delta.tables import DeltaTable
hadoop_config = {
"fs.azure.account.auth.type": "OAuth",
"fs.azure.account.oauth.provider.type": "...",
"fs.azure.account.oauth2.client.id": "...",
"fs.azure.account.oauth2.client.secret": "...",
"fs.azure.account.oauth2.client.endpoint": "..."
}
delta_table = DeltaTable.forPath(spark, <table-path>, hadoop_config)
DeltaTableBuilder:executeZOrderBy Java API which allows users to pass in varargs instead of a List._delta_log were malformed. For example, an add action with a missing } would be skipped. Now, queries will fail fast, preventing inaccurate results.Credits: Helge Brügner, Jiaheng Tang, Mitchell Riley, Ryan Johnson, Scott Sandre, Venki Korukanti, Jintao Shen, Yann Byron
We are excited to announce the release of Delta Lake 2.0.1 on Apache Spark 3.2. This release contains important bug fixes to 2.0.0 and it is recommend
We are excited to announce the release of Delta Lake 2.0.1 on Apache Spark 3.2. This release contains important bug fixes to 2.0.0 and it is recommended that users update to 2.0.1. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
This release includes the following bug fixes and improvements:
commitTime to expireTime. If you already have TTL enabled, please follow the migration steps here.Credits Adam Binford, Allison Portis, Chen Shuai, Lars Kroll, Scott Sandre, Shixiong Zhu, Venki Korukanti
We are excited to announce the release of Delta Lake 2.0.0 on Apache Spark 3.2.
We are excited to announce the release of Delta Lake 2.0.0 on Apache Spark 3.2.
Support Change Data Feed on Delta tables. Change Data Feed represents the row level changes between different versions of the table. When enabled, additional information is recorded regarding row level changes for every write operation on the table. See the documentation for more details.
Support Z-Order clustering of data to reduce the amount of data read. Z-Ordering is a technique to colocate related information in the same set of files. This data clustering allows column stats (released in Delta 1.2) to be more effective in skipping data based on filters in a query. See the documentation for more details.
Support for idempotent writes to Delta tables to enable fault-tolerant retry of Delta table writing jobs without writing the data multiple times to the table. See the documentation for more details.
Support for dropping columns in a Delta table as a metadata change operation. This command drops the column from metadata and not the column data in underlying files. See documentation for more details.
Support for dynamic partition overwrite. Overwrite only the partitions with data written into them at runtime. See documentation for details.
Experimental support for multi-part checkpoints to split the Delta Lake checkpoint into multiple parts to speed up writing the checkpoints and reading. See documentation for more details.
Python and Scala API support for OPTIMIZE file compaction and Z-order by.
Other notable changes
SimpleAWSCredentialsProvider or TemporaryAWSCredentialsProvider in S3 multi-cluster write supported LogStore.DataFrame to be written even if the column was nullable.Independent of this release, we have improved the framework for writing large scala performance benchmarks (initial version added in version 1.2.0), we have added support for running benchmarks on Google Compute Platform using Google Dataproc (in addition to the existing support for EMR on AWS)
Adam Binford, Alkis Evlogimenos, Allison Portis, Ankur Dave, Bingkun Pan, Burak Yilmaz, Chang Yong Lik, Chen Qingzhi, Denny Lee, Eric Chang, Felipe Pessoto, Fred Liu, Fu Chen, Gaurav Rupnar, Grzegorz Kołakowski, Hussein Nagree, Jacek Laskowski, Jackie Zhang, Jiaan Geng, Jintao Shen, Jintian Liang, John O'Dwyer, Junyong Lee, Kam Cheung Ting, Karen Feng, Koert Kuipers, Lars Kroll, Liwen Sun, Lukas Rupprecht, Max Gekk, Michael Mengarelli, Min Yang, Naga Raju Bhanoori, Nick Grigoriev, Nick Karpov, Ole Sasse, Patrick Grandjean, Peng Zhong, Prakhar Jain, Rahul Shivu Mahadev, Rajesh Parangi, Ruslan Dautkhanov, Sabir Akhadov, Scott Sandre, Serge Rielau, Shixiong Zhu, Shoumik Palkar, Tathagata Das, Terry Kim, Tyson Condie, Venki Korukanti, Vini Jaiswal, Wenchen Fan, Xinyi, Yijia Cui, Yousry Mohamed
Nothing published for this version
We are excited to announce the release of Delta Lake 1.2.1 on Apache Spark 3.2. Similar to Apache Spark™, we have released Maven artifacts for both Sc
We are excited to announce the release of Delta Lake 1.2.1 on Apache Spark 3.2. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
--packages mode. Previous release had a bug that resulted in user getting NullPointerException instead of proper error message when using Delta Lake with --packages mode either in pyspark or spark-shell (Fix, Test)pyspark to throw incorrect type of exceptions instead of expected AnalysisException. This issue is fixed. See issue #1086 for more details.--conf to not work for certain configuration parameters. This issue is fixed by having these configuration parameters begin with spark. See the updated documentation.LogStore implementation class config spark.delta.logStore.gs.impl from the scheme in the table path. See the updated documentation.Allison Portis, Chang Yong Lik, Kam Cheung Ting, Rahul Mahadev, Scott Sandre, Venki Korukanti
We are excited to announce the release of Delta Lake 1.2.0 on Apache Spark 3.2. Similar to Apache Spark™, we have released Maven artifacts for both Sc
We are excited to announce the release of Delta Lake 1.2.0 on Apache Spark 3.2. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13.
Support multi-cluster write in Delta Lake tables stored in S3. Users now have the option of specifying a new and experimental LogStore implementation that supports concurrent reads and writes to a single Delta Lake table in S3 from multiple Spark drivers. See the documentation for more details.
Support for compacting small files (optimize) into larger files in a Delta Lake table. Reduced number of data files improves read latency due to reduced metadata size and per-file overheads such as file-open overhead and file-close overhead. See the documentation for more details.
Support for data skipping using column statistics. Column statistics are collected for each file as part of the Delta Lake table writes. These statistics can be used during the reading of a Delta Lake table to skip reading files not matching the filters in the query. See the documentation for more details.
Support for restoring a Delta table to an earlier version. Restoring to an earlier version number or a version of a specific timestamp is supported using the SQL command, Scala APIs or Python APIs. See the documentation for more details.
Support for column renaming in a Delta Lake table without the need to rewrite the underlying Parquet data files. See the documentation for more details.
Support for arbitrary characters in column names in Delta tables. Before, the supported list of characters was limited by the support of the same in Parquet data format. Column names containing special characters such space, tab, ,, {, ( etc. are supported now. See the documentation for more details.
Support for automatic data skipping using generated columns. For any partition column that is a generated column, partition filters will be automatically generated from any data filters on its generating column(s), when possible.
Support for Google Cloud Storage is now generally available. See the documentation on how to read and write Delta Lake tables in Google Cloud Storage.
Other notable changes
delta-storage. This extracts out the LogStore interface and implementations in a separate module which is published as its own jar. This enables new implementations of LogStore without depending upon the complete Delta jars. See the migration guide here for more details.gettimestamp expression in generated columns.list calls to storageNullPointerException when trying to reference a DeltaLog created with a SparkContext that has stopped.Array.FileNotFoundException when reading Delta log files to distinguish between the corrupt log files and no files found.Independent of this release, we have also built a framework for writing large scale performance benchmarks on Delta tables using a real cluster. Currently, the framework provides a TPC-DS inspired benchmark to measure the ingestion time (e.g. time taken to create TPC-DS tables) and query times. But we encourage the community to contribute more benchmarks to measure performance of different real-world workloads on Delta tables.
Adam Binford, Alex Liu, Allison Portis, Anton Okolnychyi, Bart Samwel, Carmen Kwan, Chang Yong Lik, Christian Williams, Christos Stavrakakis, David Lewis, Denny Lee, Fabio Badalì, Fred Liu, Gengliang Wang, Hoang Pham, Hussein Nagree, Hyukjin Kwon, Jackie Zhang, Jan Paw, John ODwyer, Junlin Zeng, Jackie Zhang, Junyong Lee, Kam Cheung Ting, Kapil Sreedharan, Lars Kroll, Liwen Sun, Maksym Dovhal, Mariusz Krynski, Meng Tong, Peng Zhong, Prakhar Jain, Pranav, Ryan Johnson, Sabir Akhadov, Scott Sandre, Shixiong Zhu, Sri Tikkireddy, Tathagata Das, Tyson Condie, Vegard Stikbakke, Venkata Sai Akhil Gudesa, Venki Korukanti, Vini Jaiswal, Wenchen Fan, Will Jones, Xinyi Yu, Yann Byron, Yaohua Zhao, Yijia Cui
We are excited to announce the release of Delta Lake 1.1.0 on Apache Spark 3.2. Similar to Apache Spark™, we have released Maven artifacts for both Sc
We are excited to announce the release of Delta Lake 1.1.0 on Apache Spark 3.2. Similar to Apache Spark™, we have released Maven artifacts for both Scala 2.12 and Scala 2.13. The key features in this release are as follows.
Performance improvements in MERGE operation - On partitioned tables, MERGE operations will automatically repartition the output data before writing to files. This ensures better performance out-of-the-box for both the MERGE operation as well as subsequent read operations.
Support for passing Hadoop configurations via DataFrameReader/Writer options - You can now set Hadoop FileSystem configurations (e.g., access credentials) via DataFrameReader/Writer options. Earlier the only way to pass such configurations was to set Spark session configuration which would set them to the same value for all reads and writes. Now you can set them to different values for each read and write. See the documentation for more details.
Support for arbitrary expressions inreplaceWhere DataFrameWriter option - Instead of expressions only on partition columns, you can now use arbitrary expressions in the replaceWhere DataFrameWriter option. That is you can replace arbitrary data in a table directly with DataFrame writes. See the documentation for more details.
Improvements to nested field resolution and schema evolution in MERGE operation on array of structs - When applying the MERGE operation on a target table having a column typed as an array of nested structs, the nested columns between the source and target data are now resolved by name and not by position in the struct. This ensures structs in arrays have a consistent behavior with structs outside arrays. When automatic schema evolution is enabled for MERGE, nested columns in structs in arrays will follow the same evolution rules (e.g., column added if no column by the same name exists in the table) as columns in structs outside arrays. See the documentation for more details.
Support for Generated Columns in MERGE operation - You can now apply MERGE operations on tables having Generated Columns.
Fix for rare data corruption issue on GCS - Experimental GCS support released in Delta Lake 1.0 has a rare bug that can lead to Delta tables being unreadable due to partially written transaction log files. This issue has now been fixed (1, 2).
Fix for the incorrect return object in Python DeltaTable.convertToDelta() - This existing API now returns the correct Python object of type delta.tables.DeltaTable instead of an incorrectly-typed, and therefore unusable object.
Python type annotations - We have added Python type annotations which improve auto-completion performance in editors which support type hints. Optionally, you can enable static checking through mypy or built-in tools (for example Pycharm tools).
Other notable changes
DeltaTable.forName() for consistency with other APIsDeltaTableBuilder.partitionBy.userMetadata in the commit information when creating or replacing tables.Credits Abhishek Somani, Adam Binford, Alex Jing, Alexandre Lopes, Allison Portis, Bogdan Raducanu, Bart Samwel, Burak Yavuz, David Lewis, Eunjin Song, Feng Zhu, Flavio Cruz, Florian Valeye, Fred Liu, Guy Khazma, Jacek Laskowski, Jackie Zhang, Jarred Parrett, JassAbidi, Jose Torres, Junlin Zeng, Junyong Lee, KamCheung Ting, Karen Feng, Lars Kroll, Li Zhang, Linhong Liu, Liwen Sun, Maciej, Max Gekk, Meng Tong, Prakhar Jain, Pranav Anand, Rahul Mahadev, Ryan Johnson, Sabir Akhadov, Scott Sandre, Shixiong Zhu, Shuting Zhang, Tathagata Das, Terry Kim, Tom Lynch, Vijayan Prabhakaran, Vítor Mussa, Wenchen Fan, Yaohua Zhao, Yijia Cui, YuXuan Tay, Yuchen Huo, Yuhong Chen, Yuming Wang, Yuyuan Tang, Zach Schuermann, ericfchang, gurunath
We are excited to announce the release of Delta Lake 1.0.1 on Apache Spark™ 3.1, which back-ports bug fixes from Delta Lake 1.1.0 to Delta Lake 1.0.0.
We are excited to announce the release of Delta Lake 1.0.1 on Apache Spark™ 3.1, which back-ports bug fixes from Delta Lake 1.1.0 to Delta Lake 1.0.0.
The details of the fixed bugs are as follows:
Fix for rare data corruption issue on GCS - Experimental GCS support released in Delta Lake 1.0 has a rare bug that can lead to Delta tables being unreadable due to partially written transaction log files. This issue has now been fixed (1, 2).
Fix for the incorrect return object in Python DeltaTable.convertToDelta() - This existing API now returns the correct Python object of type delta.tables.DeltaTable instead of an incorrectly-typed, and therefore unusable object.
Fix for incorrect handling of special characters (e.g. spaces) in paths by MERGE/UPDATE/DELETE operations
Fix for Hadoop configurations not being used to write checkpoints
Improvements to DeltaTableBuilder API introduced in Delta 1.0.0
DeltaTableBuilder.partitionBy.Credits Jarred Parrett, Shixiong Zhu, Tathagata Das, Tom Lynch, Yijia Cui, Yaohua Zhao, gurunath
We are excited to announce the release of Delta Lake 1.0.0 on Apache Spark 3.1. The key features in this release are as follows.
We are excited to announce the release of Delta Lake 1.0.0 on Apache Spark 3.1. The key features in this release are as follows.
Unlimited MATCHED and NOT MATCHED clauses for merge operations in SQL - With the upgrade to Apache Spark 3.1, MERGE SQL command now supports any number of WHEN MATCHED and WHEN NOT MATCHED clauses (Scala, Java and Python APIs already support unlimited clauses since 0.8.0 on Spark 3.0). See the documentation on MERGE for more details.
New programmatic APIs to create tables - Delta Lake now allows you to directly create new Delta tables programmatically (Scala, Java, and Python) without using DataFrame APIs. We have introduced new DeltaTableBuilder and DeltaColumnBuilder APIs to specify all the table details that you can specify through SQL CREATE TABLE. See the documentation for details and examples.
Experimental support for Generated Columns - Delta Lake now supports Generated Columns which are a special type of columns whose values are automatically generated based on a user-specified function over other columns in the Delta table. You can use most built-in SQL functions in Apache Spark to generate the values of these generated columns. For example, you can automatically generate a date column (for partitioning the table by date) from the timestamp column; any writes into the table need only specify the data for the timestamp column. You can create Delta tables with Generated Columns using the new programmatic APIs to create tables. See the documentation for details.
Simplified storage configuration - Delta Lake can now automatically load the correct LogStore needed for common storage systems hosting the Delta table being read or written to. Users no longer need to explicitly configure the LogStore implementation if they are running Delta Lake on AWS S3, Azure blob stores, and HDFS. This also allows the same application to simultaneously read and write to Delta tables on different cloud storage systems. The scheme of the Delta table path is used to dynamically load the necessary LogStore implementation. Using storage systems other than the ones listed above still needs explicit configuration. See the documentation on storage configuration for details.
Experimental support for additional cloud storage systems - Delta Lake now has experimental support for Google Cloud Storage, Oracle Cloud Storage, IBM Cloud Object Storage. You will have to add an additional maven artifact delta-contribs to access the LogStores corresponding to them, and explicitly configure the LogStore names corresponding to the relevant path schemes. See the documentation on storage configuration for details. In addition, we have also defined a more stable LogStore API for building custom implementations.
Public APIs for catching exceptions due to conflicts - The exceptions thrown on conflict between concurrent operations have now been converted to public APIs. This allows you to catch those exceptions and retry your write operations. See the API documentation for details.
PyPI release - Delta Lake can now be installed from PyPI with pip install delta-spark. However, along with pip installation, you also have to configure the SparkSession. See the documentation for details.
Other notable changes
delta-contribs which contain contributions from the community that are still experimental and need more testing before being packaged in the main artifact delta-core.Delta Sharing In relation to this release, we have also introduced a new Delta Sharing project which is an open protocol for secure real-time exchange of large datasets, which enables organizations to share data in real-time regardless of which computing platforms they use. It is a simple REST protocol that securely shares access to part of a cloud dataset and leverages modern cloud storage systems, such as S3, ADLS, or GCS, to reliably transfer data. See the project repository and the release notes for details.
Credits Alex Ott, Ali Afroozeh, Antonio, Bruno Palos, Burak Yavuz, Christopher Grant, Denny Lee, Gengliang Wang, Guy Khazma, Howard Xiao, Jacek Laskowski, Joe Widen, Jose Torres, Lars Kroll, Linhong Liu, Meng Tong, Prakhar Jain, Pranav Anand, R. Tyler Croy, Rahul Mahadev, Ranu Vikram, Sabir Akhadov, Shixiong Zhu, Stefan Zeiger, Tathagata Das, Tom van Bussel, Vijayan Prabhakaran, Vivek Bhaskar, Wenchen Fan, Yijia Cui, Yingyi Bu, Yuchen Huo, Brenner Heintz, fvaleye, Herman van Hovell, Liwen Sun, Mahmoud Mahdi, Sabir Akhadov, Yaohua Zhao
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →