NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1770 most downloaded on PyPI
dlt is an open-source python-first scalable data loading library that does not require any backend to run.
Last release today
04 Oct 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 59 of the last 60 stable releases
13 versions withdrawn
withdrawn after publishing
9 years old
184 releases · first in 2018
fix(deployment): write launcher output as UTF-8 so job results and tr…
fix(deployment): write launcher output as UTF-8 so job results and tr…
clarifies color/no color output for agent loop
clarifies color/no color output for agent loop
One column per quarter.
…every retry writes its exception message. See Breaking Changes for the behavior change this implies.
Failed load packages no longer auto-abort (#3557 @anuunchin) — auto_abort_on_terminal_error defaults to False. With raise_on_failed_jobs=True the job is now retried and you get a LoadClientJobTerminalRetry instead of LoadClientJobFailed, and the package stays pending. Pipeline.drop_pending_packages() and the drop-pending-packages CLI command are deprecated in favour of abort_packages / abort-packages. This has no effect on ephemeral storage, which auto-aborts by wiping the working directory. Set the flag to True to restore the old behavior.
Filesystem layouts that omit {ext} now get the extension appended (#4220 @mattfaltyn) — the documented extension fallback was unreachable, so these layouts wrote extension-less files. They now write mocked-table.jsonl, and .jsonl.gz stays intact for compressed jobs. Pipelines already running an ext-less layout will start writing different file names.
Table prefix keeps its separator (#4283 @AnnasMazhar) — with layout="{table_name}" the prefix is now event. rather than event, which is what stops a replace of event from deleting events. A new warn_unsafe_layout_separators option (default True) warns when the separator around {table_name} can also occur inside table names; set it to False to keep a layout as it is.
Cross-destination joins (#4286 @rudolfix) — join datasets that live on different destinations. dlt attaches the foreign dataset into the query engine of the destination you read from, so the join runs in a single engine:
users = duckdb_pipeline.dataset().table("users")
orders = s3_pipeline.dataset().table("orders")
joined = users.join(orders, on="users.id = orders.user_id")Supported for duckdb, motherduck, ducklake, lance, lancedb and filesystem, with eager and lazy materialization
Snowflake nested types (#4276 @rudolfix) — native nested-type support behind a new use_nested_types flag: automatic for parquet with nested fields, and for jsonl/dicts with an upfront columns definition via a pyarrow schema. Schema and data-type evolution both work. The user-facing interface piggy-backs on the json data type and is experimental for now.
Input/output lineage in traces (#4289 @rudolfix) — traces now carry lineage in a relational form close to OpenLineage but ingestible by dlt itself. Inputs (per resource and data location) land on the extract step, and dlt derives the outputs onto the load step.
Manual load package abort (#3557 @anuunchin) — abort_packages records what happened instead of silently deleting, and gives you the chance to inspect or retry a failed job first. Adds list_pending_retry_jobs_in_package(), fail_pending_job() and retry_failed_job() on Pipeline, a dlt pipeline <name> load-package <load-id> fail-job <job> CLI command, and an .exceptions folder in the load package where every retry writes its exception message. See Breaking Changes for the behavior change this implies.
Retryable schema migrations (#4266 @rudolfix) — a retry_schema_update helper for use with tenacity, so a failing schema migration retries with jitter and backoff instead of failing the load:
from tenacity import retry, retry_if_exception, stop_after_attempt, wait_random_exponential
from dlt.pipeline.helpers import retry_schema_update
@retry(
stop=stop_after_attempt(5),
wait=wait_random_exponential(multiplier=1, max=30),
retry=retry_if_exception(retry_schema_update()),
reraise=True,
)
def load():
return pipeline.run(chess_source(...))_dlt tables on databricks, which made CREATE TABLE non-atomic, and adds an option to disable comment creation altogether via TBLPROPS.session_timezone can be set on clickhouse, databricks, duckdb, ducklake, postgres, redshift, snowflake and others, guarded by a new supports_session_timezone capability. Defaults are unchanged: UTC on duckdb and snowflake, where snowflake's was previously hardcoded. duckdb's TimeZone moves from global to per-connection config, so dlt no longer changes the global timezone of a duckdb connection you pass in; set session_timezone to an empty string to keep the duckdb default of the machine timezone.add_limit on the source factory (#4223 @richacode007-byte) — @dlt.source factories accept .add_limit() before instantiation, for both sync and async sources, and return the factory for chaining. Factory clones preserve the limit.dev.secrets.toml), vault identity or AWS account, region and prefix. See Breaking Changes for the trace shape change.instance key in the job require spec (#4262 @tetelio) — a job can declare its runner instance requirements as an open dict, with instance.size (for example {"size": "medium"}) read today. The legacy machine key still works and is still serialized, but now emits a DltDeprecationWarning, so existing manifests and deployments are unaffected.__repr__ for WorkspaceRunContext (#4022 @sohamwaghe) — shows the run dir and profile instead of an opaque object in a REPL or debugger.initialize_storage in schema-update error handling (#4325 @rudolfix) — CREATE SCHEMA was excluded from the retry-only-the-schema-update path.get_table_prefix_layout (#4283 @AnnasMazhar) — stops a replace of event from deleting events. See Breaking Changes for the delta and iceberg table move and the new warn_unsafe_layout_separators option.Schema.from_stored_schema, which skips migrate_schema and apply_defaults. dlt.dataset() and the dashboard schema-version loader therefore could not read schemas written by older dlt versions, raising KeyError: 'name' on any naming normalization.sge.Boolean (#4231 @chuenchen309) — bool fell into the int/float branch and produced a malformed numeric literal that crashed annotate_types() with decimal.InvalidOperation in Relation.where() and .filter().remove_columns as a sequence (#4241 @chuenchen309) — add_dlt_load_id_column silently dropped RecordBatch columns whose name is a substring of _dlt_load_id (id, load_id, load, dlt, …), because membership against a bare str is a substring test.work_group to primary (#4282 @burnash) — newer dbt-athena passes an empty value straight through.extra_placeholders in get_table_prefix (#4226 @axelray-dev) — when extra_placeholders overrode {schema_name}, replace-truncation and table listing matched zero files and silently left old data in place.{ext} (#4220 @mattfaltyn) — see Breaking Changes for the resulting file name change.OAuthJWTAuth.scopes is optional and defaults to None, but __post_init__ passed it to str.join, raising TypeError before authentication could begin.int(total) raises TypeError, not ValueError, when total_path resolves to an object or array, so pagination failed opaquely instead of raising the paginator's actionable error.has_more=true could override a stop already decided by maximum_offset, maximum_page or the response total, or re-enable a cursor paginator with no valid cursor. has_more=false can still stop pagination early.ModuleType.__dict__ usage infringing PEP 562 (#4215 @zilto)uvx dlthub-init@latest.guest and workspace developer roles, how invites and roles combine, and the safety rules for managing members.hub/getting-started/platform-tutorial.md and repoints the redirects at pipeline-operations/deployments so old URLs keep resolving.rest_api_source doing the full describe → load → dataset().pokemon.df() round trip, plus sql_database, filesystem | read_csv_duckdb and direct pandas/Polars/Arrow. Meta sections are kept intact.help wanted or good first issue./position example in an issue comment matched the workflow's trigger and the run began writing the wrong page.documentation and agent together sends two labeled webhooks and both matched. agent now arms the run, and the /position comment arm is removed.UV_SYNC_ARGS and stops uv run from resyncing prepared CI environments via UV_NO_SYNC. Also raises the GitPython minimum and keeps the SQLGlot compatibility test working across supported versions.pytest-rerunfailures filters for mssql, synapse and fabric ODBC transients and for databricks DatabaseTransientException. Benign assertion failures still fail fast.gather_metrics test on an early sleep return (#4236 @burnash)auto enables orphan removal when merge key is set
auto enables orphan removal when merge key is set
Nothing published for this version
{"size": "medium"} ). The legacy machine key is deprecated — it still works and serializes, but now emits a DltDeprecationWarning pointing to instance…
instance requirement spec for jobs (#4262 @tetelio) — Jobs can now declare runner instance requirements via require.instance (an open Dict[str, Any], e.g. {"size": "medium"}). The legacy machine key is deprecated — it still works and serializes, but now emits a DltDeprecationWarning pointing to instance. Fully backward compatible.has_more=true (#4227 @mattfaltyn) — An API returning has_more=true can no longer re-enable a paginator that already hit a stop condition (maximum_offset/maximum_page, response total, or a missing cursor), preventing requests past a hard limit or without a valid cursor. Fixes #4225.total_path resolving to a JSON object or array now raises the paginator's clear "not an integer" ValueError instead of escaping as an opaque TypeError.OAuthJWTAuth with the default scopes=None no longer raises TypeError; the scope claim is omitted when no scopes are configured. Fixes #4234.remove_columns as a sequence (#4241 @chuenchen309) — add_dlt_load_id_column no longer drops RecordBatch columns whose name is a substring of _dlt_load_id (e.g. id, load), which previously produced silent NULLs or a crash. Fixes #4240.UV_SYNC_ARGS, prevents uv run from resyncing prepared CI environments (UV_NO_SYNC), bumps the GitPython minimum, and keeps the SQLGlot compatibility test green across versions.pytest-rerunfailures filters so transient ODBC/Azure SQL and Databricks connection errors get retried instead of failing the run, while benign assertion failures still fail fast.ClickHouse staging-optimized replace strategy ( #3927 @filipesilva ) — ClickHouse now supports the staging-optimized replace strategy, performing atom
staging-optimized replace strategy (#3927 @filipesilva) — ClickHouse now supports the staging-optimized replace strategy, performing atomic table swaps via EXCHANGE TABLES so a full replace is instantaneous and never leaves the destination in a half-loaded state.Relation.join() (#3868 @burnash) — The dataset relation API gains explicit join support, giving you control over join conditions when composing relations.physical_location() accessor and can_join_with rules, so dlt knows when relations that live on different destinations (or datasets) can actually be joined together.Features
enable_atomic_replace config flag makes the truncate-and-insert strategy do a single-job, metadata-preserving WRITE_TRUNCATE_DATA load from GCS staging, giving BigQuery an atomic, transactional full replace.DATA_WRITER__COMPRESSION or ParquetFormatConfiguration(compression=...).ArrowToParquetWriter (#3896 @AyushPatel101) — Safe Arrow type promotions (e.g. float32 → float64) now work across flush batches instead of crashing once data spans multiple batches.FILLRECORD) to the generated COPY command for staged loads.dlt→dlthub command proxying in workspace context (#4037 @rudolfix) — Instead of silently running a dlthub command, dlt now shows a note and points to the correct dlthub counterpart, avoiding misleading behavior.Fixes
add_dlt_id (#4187 @nizar-zerrad) — Adding _dlt_id to an empty Arrow table no longer fails with ArrowInvalid.items_count).decompose="parallel" / parallel-isolated runs so concurrent tasks no longer truncate each other's staging files.log_period (#3962 @cbility) — Stops resetting the log timer on every counter update, so logs respect the configured log_period.jsonpath-ng 1.8.0, dropping the archived ply transitive dependency flagged by security scanners (CVE)./commit skill (#4098 @lis365b) — Claude Code tooling to keep commit messages Conventional-Commit clean (no footers/emojis).Destination factory (#4170 @rudolfix)This is a patch release that allows 1.28.x dlt to use future versions of dlthub-client
This is a patch release that allows 1.28.x dlt to use future versions of dlthub-client
Dropped Python 3.9 support ( #4074 @Travior ) — Python 3.9 reached end-of-life on 2025-10-31. dlt no longer tests against or advertises support for 3.
ruff/mypy target versions were bumped accordingly. Users still on 3.9 must upgrade to 3.10+ to use this release.timestamp[ns] and times as time64[ns] (via the arrow_stream return type) instead of date64 (ms). dlt now normalizes all connectorx temporal columns to microsecond precision (the renamed cast_connectorx_temporal_columns), truncating ns→us losslessly since source DBs carry no nanoseconds.detect_datetime_format so week-date cursors round-trip (#4061 @yashs33244) — Week-date formats (YYYY-Www) were detected as %Y-W%W, which disagrees with ISO weeks at year boundaries. "2026-W01" parsed to 2025-12-29 and re-rendered as "2025-W52", silently corrupting saved incremental cursors. Now emits %G-W%V / %G-W%V-%u.0x0 characters from postgres INSERT strings (#4086 @rudolfix) — Postgres escaping allowed null (0x0) characters through; they are now stripped in the escape function. INSERT statements are also split by \n only.sql_database (#4067 @rudolfix) — User-passed metadata is now correctly consulted for cached tables in both eager and deferred reflection. Fixes #4066.ai init (#4072 @lis365b) — Codex silently drops skills whose description frontmatter exceeds 1024 chars. The Codex install path now truncates the installed SKILL.md description to the cap (with a warning) so workbench skills like dlthub-router load correctly.dlthub-init vs dlthub-start consistently in onboarding docs (#4080 @bjoaquinc)test_common.yml into 27 parallel focused jobs (3 per OS/Python combo) so wall-clock is roughly max(suite) instead of sum(suites), moves runners to blacksmith-* variants, and enables Python 3.14 in CI (running the full suite, with 3.14-gated deps snowflake-connector-python>=4.4.0 and cffi>=1.18). Also includes job/workflow renames; one functional reduction — lance_s3 non-essential tests now run only on full-suite triggers.refresh is now the recommended way to do a full refresh — the replace switch is deprecated. Also fixes #3998 and #4017 .
refresh="drop_data" on Delta and persistent-catalog Iceberg no longer frees storage (#4051 @rudolfix) — Truncation is now a transactional delete that keeps the table, schema, version history, and data files (retained for time travel until vacuum). Previously the files were deleted. This corrects prior erroneous behavior, but pipelines relying on drop_data to reclaim disk space will no longer see storage freed without an explicit vacuum.replace now fully truncates empty and orphaned tables (#4010 @rudolfix) — Tables belonging to a replace resource that receive no data in a run (including nested tables, dynamic-name and variant tables) are now consistently truncated. Previously these tables could be left orphaned with stale rows surviving the reload, leaving the dataset in an inconsistent state. Pipelines that implicitly relied on that leftover data will now see those tables emptied.LanceNamespace + lance.Session across job clients; atomic single-commit-per-table writes uncommitted fragments in parallel and commits them in one version (Append/Overwrite/upsert). replace is now a single Overwrite commit so readers never see a partially-replaced table, and the namespace pool rebuilds handles on credential rotation to avoid ExpiredToken on long loads. Also fixes #3800 (Iceberg 409 Table already exists after drop_sources).replace / refresh truncation (#4010 @rudolfix) — replace resources now consistently truncate all participating tables even when a load carries no data (nested tables, dynamic names, variants included), and drop_data refresh truncates correctly on append and survives non-existent tables. refresh is now the recommended way to do a full refresh — the replace switch is deprecated. Also fixes #3998 and #4017.write_encoding option lets you choose the encoding of CSV files dlt writes (default utf-8), e.g. utf-8-sig for Excel BOM or latin-1/cp1252 for legacy importers. Set via [normalize.data_writer] write_encoding="latin-1".ExpiredToken failures on long-held connections (#4003).Features
zip path as good as pandas.--asset-url. Launcher path resolution uses find_spec instead of import_module so notebooks aren't executed before marimo run / streamlit run (~1–1.5s faster port readiness).Fixes
credential_chain secrets to survive temp-token expiry (#4021 @0ywfe) — Adds REFRESH auto so long-held sql_client connections no longer die with ExpiredToken once temporary AWS tokens rotate. Fixes #3987.Retry-After: 0 no longer triggers an immediate retry loop (#4043 @AstrakhantsevaAA) — Values ≤ 0 are treated as no actionable hint, letting tenacity's exponential backoff take over. Fixes #4036.create_indexes is enabled (#4011 @burnash) — Avoids Unity Catalog UC_REFERENTIAL_CONSTRAINT_DOES_NOT_EXIST failures when the matching primary/unique key isn't created.META_TYPE 'sqlite' (#3871 @Analect) — Splits the duckdb/sqlite branch so a duckdb:///catalog.duckdb catalog URI attaches cleanly instead of failing on PRAGMA journal_mode=WAL.__ in generated CLI docs to prevent markdown bold (#3993 @burnash)__deployment__.py file name in CLI docs (#3988 @nuetu)pytest-rerunfailures, scoped to just those two transient destinations via --only-rerun.test_multi_schema_selection flaking on slow CI.Hotfix: fixes #3998 ( merge with empty data after replace on incremental truncates the destination table). Upgrade for anyone on 1.27.0/1.27.1.
Hotfix: fixes #3998 (merge with empty data after replace on incremental truncates the destination table).
Upgrade for anyone on 1.27.0/1.27.1.
Nothing published for this version
…mcp , ai ) into a dedicated dlt[hub] plugin. See Breaking Changes above.
workspace extra removed and dlthub command split out (#3929 @rudolfix) — The workspace extra is gone; users should install marimo, pyarrow, ibis, fastmcp, and other dependencies directly. Part of dev tooling moved to a plugin: dlt dashboard, dlt pipeline ... show, dlt pipeline ... mcp now require pip install dlt[hub]. dlt ai was moved to dlthub ai.@dlt.resource without manual conversion. Auto-detected and routed through the Arrow extraction pipeline; LazyFrames are auto-collected before conversion.databricks_adapter(my_resource, insert_api="zerobus") enables loading via the Databricks Zerobus SDK into Delta tables. API mirrors bigquery_adapter; supported for append write disposition.dlt.Relation (#3889 @burnash) — Apply incremental filters directly on dlt.Relation, enabling efficient incremental reads from datasets.dlthub command split (#3929 @rudolfix) — Reorganizes _workspace modules and splits dev tooling (dashboard, mcp, ai) into a dedicated dlt[hub] plugin. See Breaking Changes above.dlt.Relation (#3889 @burnash) — See Highlights.dlthub command split (#3929 @rudolfix) — See Highlights / Breaking Changes.lance destination REST Namespace support (experimental) (#3908 @jorritsandbrink) — Adds experimental REST Namespace support to the lance destination, currently only validated against an ephemeral in-memory Lance REST Namespace proxy.-y / --yes flag to bypass non-interactive prompts (#3910 @anuunchin) — New flag auto-accepts all confirm() prompts via a dedicated ALWAYS_CONFIRM flag in echo.py, without leaking into prompt() or text_input(). Resolves #3592.RootModel handling and preserves root-level Annotated[...] metadata after Pydantic 2.13 moved the discriminator off root_field.annotation.scd2 unmapped insert for nested tables (#3812 @deschman) — Generates an explicit column list (excluding SCD2 from/to metadata) for the staging-to-destination insert, fixing DATATYPE_MISMATCH errors on Spark/Databricks ANSI mode when schemas drift. Fixes #3811.ReplicatedMergeTree / SharedMergeTree for merge delete temp tables, preventing duplicates on replicated/shared MergeTree deployments. Closes #3797._user_agent_entry (#3935 @xodn348) — Switches DatabricksCredentials.to_connector_params() to the non-prefixed user_agent_entry, silencing the deprecation warning printed on every connection. Closes #3934.with_load_id breaking on non-root sibling branches (#3878 @Travior)max_length param (#3826 @aditypan) — Removes spurious length appending in schema naming when max_length is provided. Fixes #3816.dataset_name to job metrics in load step (#3808 @rudolfix)join usage (#3902 @Travior)fork_tests_with_secrets.yml workflow gated by the ci from fork label and a maintainer approval in the fork-ci environment. Resolves #3850.fruit_pipeline fixture to module to avoid duckdb lock race (#3932 @burnash) — Stabilizes the dashboard e2e suite.Nothing published for this version
Nothing published for this version
Incremental external scheduler now raises instead of silently warning (#3877 @rudolfix) — Untyped/non-coercible cursor values now raise JoinSchedulerE
JoinSchedulerError; missing intervals raise ExternalSchedulerNotAvailable. Resources with allow_external_schedulers=True that previously fell back to dlt state will now fail. This is a bugfix that corrects previously incorrect behavior.dlt.Relation.join(...) (#3590 @Travior) — Adds a join() method on dlt.Relation based on the normalizer and table references, enabling fluent relational composition over datasets.TJobQueryTags is generalized to TQueryTags with a new operation field (with a compatibility export).dlt.current.interval() returns the active (start, end) interval or None, backed by an injectable TimeIntervalContext with optional allow_external_schedulers override and auto-detection from env vars / Airflow.dlt.Relation.join(...) (#3590 @Travior) — see Highlights.dlt/extract/incremental/sql.py (to_sqlglot_filter) honors timestamp_precision, supports_tz_aware_datetime_in_cast, and sqlite quirks; works on bound and unbound incrementals.start_value persisted in incremental state (#3877 @rudolfix) — Only written when rows actually arrive, so it is no longer advanced silently on empty runs.uuid_to_string PyArrow fast path (#3877 @rudolfix) — Numpy-vectorized with a pure-Python fallback; pyarrow ≥ 24 arrow.uuid extension arrays are coerced to canonical strings, and UUID columns under pyarrow < 24 also take the fast path.dlt.current.resource_metrics() counters are no longer dropped when every item is filtered out.NotRequired[T] (#3877 @rudolfix) — Via __required_keys__."dremio" dialect literal (#3877 @rudolfix) — Added to TSqlGlotDialect.Schema.unify_schemas() (#3898 @burnash) — The naming-convention check in Schema.unify_schemas() is now opt-in; also drops the max_length tests workaround.dlt.attach() when pipeline cannot be restored (#3890 @bjoaquinc) — Rewrites CannotRestorePipelineException messages to name required inputs, show a concrete dlt.attach(...) example, and offer dlt.dataset() as a lighter alternative; suppresses a redundant inner exception in tracebacks.iter_std (#3877 @rudolfix) — Reader threads swallow ValueError/OSError and always close the queue.ConfigFieldMissingException (#3877 @rudolfix)pipeline.md (#3885 @ShreyasGS) — Tense, contractions, and grammar cleanups via Harper + Vale Google Developer Docs style.Nothing published for this version
Nothing published for this version
Multischema datasets (#3770 @burnash) — See Breaking Changes above. Enables sidecar schemas (e.g. data-quality quarantine tables) to live alongside th…
dataset() method and still go back to single-schema dataset by providing pipeline.default_schema when creating dataset.lance destination (#3810 @jorritsandbrink) — New destination for the Lance table format with optional vector embedding generation via lancedb. Supports local storage and s3/az/gs, uses the Lance Directory Namespace V2 spec, and supports branching. Complements the existing lancedb destination (which targets LanceDB Cloud).lance destination (#3810 @jorritsandbrink) — See Highlights.ducklake: metadata_schema ATTACH option (#3763 @sangwookWoo) — Adds metadata_schema to DuckLakeCredentials so the DuckLake metadata schema can be configured independently from ducklake_name.__main__ in orchestrators (#3784 @rudolfix) — Closes #3586._dlt_id requirement when merging arrow tables without nested tables on ClickHouse.aws_session_token to staging s3() table function (#3769 @anuunchin) — Temporary AWS credentials now work for ClickHouse staging.read_csv (#3743 @biefan) — File is opened with the requested encoding so SFTP/paramiko stacks no longer pre-decode as UTF-8.trace.asdict() now retains pipelines that fail in the sync step before extract.lance extension promotion to built-in in duckdb 1.5.PermissionError in rename_tree (#3853 @burnash) — Resolves intermittent Windows CI failures during normalize→loaded rename.llm-native-workflow.md.hf login command (#3781 @julien-c)mypy configs to pyproject.toml (#3780 @zilto) — Partially resolves #3346.Custom resource metrics now stored as tables (#3718 @rudolfix) — Incremental metrics in the trace are now represented in table format. This changes th
insert-only merge strategy that performs idempotent, key-based appending: inserts records whose primary key doesn't exist in the destination while silently skipping duplicates. No updates or deletes. Supported across all SQL destinations, Delta Lake, and Iceberg.parallel and parallel-isolated decompose modes, all source components now fan out concurrently from a shared start node. Previously the first source had to complete before others could begin, adding unnecessary wall-clock time. This release also adds basic Airflow 3 support with smoke tests.replacing_merge_tree table engine type for ClickHouse that enables native deduplication and soft deletes via dedup_sort and hard_delete column hints.arrow_concat_promote_options can now be set to "default" or "permissive" instead of the hardcoded "none", enabling automatic type promotion when yielding multiple Arrow tables with slightly different inferred types.dlt pipeline info/show no longer crashes with UnknownDestinationModule on pipelines using @dlt.destination.primary_key=() to Incremental to disable deduplication is no longer silently overwritten by the resource's own primary key.md:) now raise a clear configuration error instead of a confusing connection failure.mask_columns() function supporting all sql_database backends.Streamlit dashboard removed (#3674 @rudolfix) — The legacy Streamlit-based pipeline dashboard (dlt pipeline show) has been removed. It was a dead code
Streamlit dashboard removed (#3674 @rudolfix) — The legacy Streamlit-based pipeline dashboard (dlt pipeline show) has been removed. It was a dead code for a long time.
New sources.<name>.<key> configuration lookup path (#3626 @rudolfix) — Source configuration now supports a compact layout. When a source's section name differs from its resource/source name, dlt now also looks up sources.<name>.<key> in addition to the full sources.<section>.<name>.<key> path. For example, for a source registered under section chess_com with name chess:
# Before (still works): full qualified path
[sources.chess_com.chess]
api_key = "secret"
# New (also works now): compact path using just the source name
[sources.chess]
api_key = "secret"
# Credentials follow the same pattern:
# Full: sources.chess_com.chess.credentials.api_key
# Compact: sources.chess.credentials.api_key
This is breaking if you previously had values at sources.<name> that were unrelated to this source — they will now be resolved where they were previously ignored.
AI Workbench (#3674 @rudolfix) — New dlt ai CLI command group that turns dlt workspaces into AI-assisted development environments. Includes toolkit system for installing curated skill/rule bundles, pluggable MCP server architecture with composable features (pipeline, workspace, toolkit, secrets), and multi-agent support (Claude Code, Cursor, Codex).
Relational normalizer optimization (#3626 @rudolfix) — Major performance improvements to JSON data normalization and schema evolution: 5x faster on flat data, ~2x on nested REST API data, ~1.8x on wide nested data. ISO timestamp parsing improved 2-3x by removing timezone conversions.
Iceberg table properties (#3699 @rudolfix) — Adds support for setting Iceberg table and namespace properties via the adapter and configuration.
override_data_path option to DuckLake ATTACH (#3709 @udus122) — New override_data_path configuration option that appends OVERRIDE_DATA_PATH true to the ATTACH statement, allowing the current DATA_PATH to override the path stored in catalog metadata.PageNumberPaginatorConfig, OffsetPaginatorConfig, and JSONResponseCursorPaginatorConfig.os.path.commonprefix() with os.path.commonpath() in FileStorage.is_path_in_storage() to correctly validate path containment using path segments instead of characters.dev_mode flag in pipeline local state so it persists across dlt.attach() calls. Detects dev→non-dev transitions and resets working folder cleanly.HF_ENDPOINT env var for card operations.start_out_of_range flag with range_start="open" (#3708 @AyushPatel101) — Correctly sets start_out_of_range=True when a row's cursor value equals start_value with range_start="open", fixing delayed can_close() in descending-order pipelines.dataset_name=None (#3710 @Travior) — Handles the case where dataset_name is None in LanceDBSqlClient.create_view, preventing None prefix in view names.Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Remove outdated Motherduck troubleshooting (#3683 @elviskahoro) — Removed read-only database troubleshooting section for deprecated DuckDB versions.
hf protocol support to the filesystem destination, enabling direct loading to Hugging Face datasets. Closes #1227.feat(workspace): add default exclude patterns for file selector (#3661 @canassa) — WorkspaceFileSelector now ships with DEFAULT_EXCLUDES (.git/, .venv
WorkspaceFileSelector now ships with DEFAULT_EXCLUDES (.git/, .venv/, __pycache__/, node_modules/, etc.) so well-known non-deployable paths are always excluded, even without a .gitignore.ignore_file_found attribute to WorkspaceFileSelector (#3663 @canassa) — Consumers can now check whether the configured ignore file (e.g. .gitignore) was actually found.utils.py and dlt_dashboard.py into focused modules with simplified UI across all sections.sse for http-stream transport for built-in MCP servers and annotates pipeline trace schema.sources.<name>.api_key..to_mermaid() now handles columns missing the data_type field instead of crashing.select_sequential_consistency to fix flaky tests caused by ClickHouse's eventual consistency model.dlthub changes.shutil docs.dlthub metrics section; update checks (#3641 @zilto)docs/ linting in one make command (#3666 @anuunchin) — Introduces an overarching lint target in the docs Makefile. Resolves #3642.source.md file.make test-load-local-p (#3645 @tetelio) — Convenience make install target for local load tests on duckdb and filesystem.Pydantic v1 support removed (#3572 @anuunchin) — All Pydantic v1 compatibility code has been removed. The codebase now requires Pydantic v2 only.
data_type contract semantic change (#3572 @anuunchin @rudolfix) — The data_type contract now applies to full data type (ie. precision, nullability), not only to variant columns (data type change). Users with data_type: freeze who relied on changing nullable/precision/scale on existing columns will now be blocked.merge_columns now removes compound properties (#3431 @anuunchin) — Previously merge_columns was purely additive, which caused compound properties like merge_key to be incorrectly replaced rather than properly merged. The function now correctly removes compound properties that should be removed.RootModel types (validation of event streams with various event types), schema contracts properly separate resource-defined vs data-derived hints, Pydantic model columns bypass contract checks when authoritative. Supports Pydantic models on arrow and model items with full schema contract enforcement. Prepares for Pydantic v3.ALTER TABLE ... SWAP for staging-optimized replace strategy on Snowflake, eliminating table downtime during data replacement.sql_database (#3595 @rudolfix) — Register custom TableLoader implementations as named backends. ConnectorX backend ported as PoC; ADBC and paginated loader implemented as test cases.llms.txt and Markdown docs generation (#3635 @rudolfix) — Generates llms.txt index and Markdown versions of docs pages with a "View Markdown" navigation option, making the docs LLM-friendly.rest_api: parallelized dependent resources (#3574 @Shadesfear) — Add parallelized flag to dependent resources (transformers) so child resource fetches run concurrently.dlt.Relation: filter by load_id (#3547 @zilto) — Filter dataset relations by load ID (experimental).dlt.Relation: flatten logic and improve typing (#3578 @zilto) — Remove dynamic methods; explicit return types for .df(), .arrow(), etc.SourceFactory (#3636 @rudolfix) — Add preprocessor hooks to dlt.source factory for modifying source instances.engine_kwargs for sql_database/sql_table sources (#3414 @tetelio) — Pass SQLAlchemy engine arguments directly to create_engine() for sources.DECFLOAT columns via the SQLAlchemy backend.query_result_bucket now optional (#3566 @arel) — Omit or set to None when using Athena's managed results bucket.extra_credentials for S3 (#2888 @warje) — Adds extra_credentials config for role-based S3 authentication._dlt_load_id written as dict on MSSQL + ADBC (#3584 @rudolfix)CREATE OR REPLACE for merge temp tables (#3589 @rudolfix)read_csv_duckdb respects filename=True (#3606 @karlanka)sql_database (#3638 @rudolfix)dlt init (#3615 @rudolfix)ibis-framework, remove sqlglot constraint (#3621 @Travior)pytest-xdist with fully isolated workers..zed/ directory (#3633 @zilto)Nothing published for this version
This release adds several interesting improvements and many bugfixes. Lancedb destination now uses duckdb extension to let you query lance tables with
This release adds several interesting improvements and many bugfixes. Lancedb destination now uses duckdb extension to let you query lance tables with SQL, ibis or sqlglot via our standard .dataset() interface. We introduced several iceberg-relates improvements (catalog support, s3 tables for Athena, advanced partitioning). There's also new fabric destination and additional options in `clickhouse_adapter. Finally: we have test environment for Oracle and we stared to fix Oracle related bugs.
SqlClientBase - query lance tables with ibis, sqlglot or raw SQL by @zilto in https://github.com/dlt-hub/dlt/pull/3527athena destination by @jorritsandbrink in https://github.com/dlt-hub/dlt/pull/3434fabric destination (by @mattiasthalen) by @jorritsandbrink in https://github.com/dlt-hub/dlt/pull/3535clickhouse_adapter extensions by @jorritsandbrink in https://github.com/dlt-hub/dlt/pull/3511Full Changelog: https://github.com/dlt-hub/dlt/compare/1.20.0...1.21.0
feat: implement ConfigurationFileSelector by @ivasio in https://github.com/dlt-hub/dlt/pull/3418
filesystem by @ivasio in https://github.com/dlt-hub/dlt/pull/3339JSONResponseCursorPaginator by @segetsy in https://github.com/dlt-hub/dlt/pull/3374Full Changelog: https://github.com/dlt-hub/dlt/compare/1.19.1...1.20.0
Nothing published for this version
Nothing published for this version
fixes arrow import in sql_database by @rudolfix in https://github.com/dlt-hub/dlt/pull/3411
Full Changelog: https://github.com/dlt-hub/dlt/compare/1.19.0...1.19.1
Feat: support return_type = arrow_stream for connectorx backend by @ivasio in https://github.com/dlt-hub/dlt/pull/3218
Schema.to_mermaid() by @zilto in https://github.com/dlt-hub/dlt/pull/3364snowflake clustering key modifications by @jorritsandbrink in https://github.com/dlt-hub/dlt/pull/3365with_table_name and other functions available through `dlt.pip… by @hello-world-bfree in https://github.com/dlt-hub/dlt/pull/3318query_adapter_callback by @dat-a-man in https://github.com/dlt-hub/dlt/pull/3253data_quality concept page by @zilto in https://github.com/dlt-hub/dlt/pull/3341dlt/extract/hints.py module by @luqmansen in https://github.com/dlt-hub/dlt/pull/3332Full Changelog: https://github.com/dlt-hub/dlt/compare/1.18.2...1.19.0
resolves "default" limit via ibis options by @rudolfix in https://github.com/dlt-hub/dlt/pull/3273
Full Changelog: https://github.com/dlt-hub/dlt/compare/1.18.1...1.18.2
fixes git import and enables tests by @rudolfix in https://github.com/dlt-hub/dlt/pull/3262
Full Changelog: https://github.com/dlt-hub/dlt/compare/1.18.0...1.18.1
A few cool feats in this release that also need your attention if you use them:
A few cool feats in this release that also need your attention if you use them:
databricks destination now supports iceberg table format, cluster and partition hints. note that previously they were ignored and mixing of cluster and partition is not allowed. in super rare cases you can get error messages that were not there previouslydlt to shut down load pools gracefully (previously we were raising exceptions). with console attached double CTRL-C will raise immediately. we also support graceful shutdowns for pipelines that run in thread pools. this is a behavioral change compared to 1.17.0dlt.pipeline(destination="warehouse") and via dlt.destination which now serves both as decorator and a factory, overall super helpful when you go from dev to production but in rare cases you may get linter and runtime error on custom destinations @dlt.destination that were not using kwargs to pass options to decoratorRelation got to_ibis method that works both on tables and queries and which uses dlt as a backed for ibis so you can read data from them_workspace modules where all code for cli, dashboards, mcp (and other things that you typically not use in runtime). there were changes in internal private interfaces. as a user you do not to worry about itdlt.Relation API and create bound Ibis tables by @zilto in https://github.com/dlt-hub/dlt/pull/3179pandas deps by @zilto in https://github.com/dlt-hub/dlt/pull/3157_workspace module by @rudolfix in https://github.com/dlt-hub/dlt/pull/3215TTableReference by @zilto in https://github.com/dlt-hub/dlt/pull/3093cluster by in bigquery on alter statements by @adrian-173 in https://github.com/dlt-hub/dlt/pull/3239.table(..., table_type="ibis") with .to_ibis() by @zilto in https://github.com/dlt-hub/dlt/pull/3225_workspace module by @rudolfix in https://github.com/dlt-hub/dlt/pull/3215pyproject.toml and reduce verbosity by @zilto in https://github.com/dlt-hub/dlt/pull/3205Full Changelog: https://github.com/dlt-hub/dlt/compare/1.17.1...1.18.0
Nothing published for this version
This patch release mostly addresses bugs and inconsistencies found in new ducklake destination. The most significant change was to rename `catalog_nam
This patch release mostly addresses bugs and inconsistencies found in new ducklake destination. The most significant change was to rename catalog_name to ducklake_name in destination.ducklake configuration in https://github.com/dlt-hub/dlt/pull/3153
postgresql by postgres in ATTACH (ducklake) by @zilto in https://github.com/dlt-hub/dlt/pull/3148read_parquet(use_arrow: bool) by @zilto in https://github.com/dlt-hub/dlt/pull/3149CONTRIBUTING.md by @jorritsandbrink in https://github.com/dlt-hub/dlt/pull/3135test_track_anon_event test by @rudolfix in https://github.com/dlt-hub/dlt/pull/3152Full Changelog: https://github.com/dlt-hub/dlt/compare/1.17.0...1.17.1
dashboard: fixes file opener on WSL by @rudolfix in https://github.com/dlt-hub/dlt/pull/3076
None if child table exists by @sh-rp in https://github.com/dlt-hub/dlt/pull/3048BaseOperator in airflow_helper.py (#2601) by @ianedmundson1 in https://github.com/dlt-hub/dlt/pull/3043filesystem via Incremental sort_order https://github.com/dlt-hub/dlt/pull/2737sql_client.raise_database_error creates circular __cause__ dependency by @anuunchin in https://github.com/dlt-hub/dlt/pull/3111ducklake destination (all buckets and catalog combinations supported) by @zilto in https://github.com/dlt-hub/dlt/pull/3015json and timestamp without timezone in pg_replication @anuunchin https://github.com/dlt-hub/verified-sources/pull/657sql_database and filesystem with examples and additional tests by @rudolfix in https://github.com/dlt-hub/dlt/pull/2737ducklake destination documentation by @rudolfix in https://github.com/dlt-hub/dlt/pull/3015Full Changelog: https://github.com/dlt-hub/dlt/compare/1.16.0...1.17.0
Improved timestamp handling, please carefully read https://dlthub.com/docs/general-usage/schema#handling-of-timestamp-and-time-zones. For some edge ca
dlt.Relation and dlt.Datasetdlt dashboard to see it. Learn more about how it works here: https://dlthub.com/docs/general-usage/dashboardMissingDependencyException should inherit ImportError by @zilto in https://github.com/dlt-hub/dlt/pull/2977dlt.Schema.to_dot() graphviz export by @zilto in https://github.com/dlt-hub/dlt/pull/2959dlt.Pipeline.__repr__ by @zilto in https://github.com/dlt-hub/dlt/pull/3022dlt.Dataset and dlt.Relation by @zilto in https://github.com/dlt-hub/dlt/pull/3059ruff check for linting by @zilto in https://github.com/dlt-hub/dlt/pull/2967__marimo__/ folders by @zilto in https://github.com/dlt-hub/dlt/pull/3008Full Changelog: https://github.com/dlt-hub/dlt/compare/1.15.0...1.16.0
This version will add .gz extensions to files that are compressed. That includes filesystem destinations, internal working directory and staging locat
This version will add .gz extensions to files that are compressed. That includes filesystem destinations, internal working directory and staging locations used to feed other destinations. A few practical hints:
filesystem destination will continue storing files without gz extension and they are not affected by the change (existing datasets will retain their behavior where this extension is not added for backwards compatibility).gz extension, also if dlt is configured to keep data in stageparquet files.has_more boolean flag logic to RESTClient OffsetPaginator by @michaelconan in https://github.com/dlt-hub/dlt/pull/2817__repr__ for @dlt.transformation by @zilto in https://github.com/dlt-hub/dlt/pull/2940Schema.to_dbml(), auto export schemas in dbml format by @zilto in https://github.com/dlt-hub/dlt/pull/2929arrow2 with arrow backend for connectorx, enables newest connectorx versions by @zilto in https://github.com/dlt-hub/dlt/pull/2933streamed_exec in delta merge upsert by @anuunchin in https://github.com/dlt-hub/dlt/pull/2961Full Changelog: https://github.com/dlt-hub/dlt/compare/1.14.1...1.15.0
Breaking Changes If you used pipeline.dataset() and used ibis syntax to write queries please read below:
Breaking Changes
If you used pipeline.dataset() and used ibis syntax to write queries please read below:
pipeline.dataset() - with Ibis Expressions has slightly changed, you will need to update your existing codebase if you are using Ibis Expressions. If you were using the Dataset without having Ibis installed, or without using any Ibis features, no changes are needed.
Protocols by @zilto in https://github.com/dlt-hub/dlt/pull/2870Full Changelog: https://github.com/dlt-hub/dlt/compare/1.12.3...1.14.1
Nothing published for this version
Nothing published for this version
Extend CSV quoting options in CsvWriter by @burnash in https://github.com/dlt-hub/dlt/pull/2810
Full Changelog: https://github.com/dlt-hub/dlt/compare/1.12.3...1.13.0
Nothing published for this version
(feat) allows to add SQL statements to schema migration executed after tables were created/altered by @rudolfix in https://github.com/dlt-hub/dlt/pull
str and repr to dataset and relation by @rudolfix in https://github.com/dlt-hub/dlt/pull/2796@dlt.transformation to __all__ by @zilto in https://github.com/dlt-hub/dlt/pull/2797Nothing published for this version
This is a prelease of dlt and our first build with the uv package manager.
This is a prelease of dlt and our first build with the uv package manager.
Quality of Life (fixing annoying little things)
Quality of Life (fixing annoying little things)
import dlt by @djudjuu in https://github.com/dlt-hub/dlt/pull/2707LIMIT env variable which was skipped before@dlt.source
def source():
@dlt.resource(write_disposition="merge", primary_key="_id")
def documents(access_token=dlt.secrets.value, limit=10):
yield from generate_json_like_data(access_token, limit)
return documents
⚠️ Still we do not recommend to define parametrized inner resources.
dlt always wraps resources in generators so your return will be converted to yield.DltResource from a resource function you must explicitly type the return value:@dlt.resource
def rv_resource(name: str) -> DltResource:
return dlt.resource([1, 2, 3], name=name, primary_key="value")
wrap and unwrap of functions. our decorators preserve both typing and runtime signature of decorated functions. makefun got removed.Incremental initializes from another Incremental as native value, it copies original type correctlydlt.resource can define configuration section (also using lambdas)Bugfixes and improvements
delete-insert merge to decrease query cost in https://github.com/dlt-hub/dlt/pull/2721Chores & tech debt
We switch to uv in the coming days and:
__repr__() for public interface by @zilto in https://github.com/dlt-hub/dlt/pull/2630dataset for testing now)load_id col in _dlt_loads table by @zilto in https://github.com/dlt-hub/dlt/pull/2729🧪 Upgrades to data access
x-annotation hints are propagatedSqlModel represent SQL query and is processed in extract, normalize and loaded in load stepscalar() on data access expressions ie.# get latest processed package id
max_load_id = pipeline.dataset()._dlt_loads.load_id.max().scalar()
🧪 Cool experimental stuff:
Check out our new embedded pipeline explorer app
dlt pipeline <name> show --marimo
dlt pipeline <name> show --marimo --edit
use edit option to enable Notebook/edit mode in Marimo + very cool Ibis dataset explorer
We updated contribution guidelines
Full Changelog: https://github.com/dlt-hub/dlt/compare/1.11.0...1.12.0
Nothing published for this version
Nothing published for this version
> We yanked this release from PyPI after discovering that the minimum allowed version of sqlglot could prevent dlt from being imported. This release h
[!IMPORTANT]
We yanked this release from PyPI after discovering that the minimum allowed version ofsqlglotcould preventdltfrom being imported. This release has been replaced by version1.12.1, which includes the same release notes.
Quality of Life (fixing annoying little things)
import dlt by @djudjuu in https://github.com/dlt-hub/dlt/pull/2707LIMIT env variable which was skipped before@dlt.source
def source():
@dlt.resource(write_disposition="merge", primary_key="_id")
def documents(access_token=dlt.secrets.value, limit=10):
yield from generate_json_like_data(access_token, limit)
return documents
⚠️ Still we do not recommend to define parametrized inner resources.
dlt always wraps resources in generators so your return will be converted to yield.DltResource from a resource function you must explicitly type the return value:@dlt.resource
def rv_resource(name: str) -> DltResource:
return dlt.resource([1, 2, 3], name=name, primary_key="value")
wrap and unwrap of functions. our decorators preserve both typing and runtime signature of decorated functions. makefun got removed.Incremental initializes from another Incremental as native value, it copies original type correctlydlt.resource can define configuration section (also using lambdas)Bugfixes and improvements
Chores & tech debt
We switch to uv in the coming days and:
__repr__() for public interface by @zilto in https://github.com/dlt-hub/dlt/pull/2630dataset for testing now)load_id col in _dlt_loads table by @zilto in https://github.com/dlt-hub/dlt/pull/2729🧪 Upgrades to data access
x-annotation hints are propagatedSqlModel represent SQL query and is processed in extract, normalize and loaded in load stepscalar() on data access expressions ie.# get latest processed package id
max_load_id = pipeline.dataset()._dlt_loads.load_id.max().scalar()
🧪 Cool experimental stuff:
Check out our new embedded pipeline explorer app
dlt pipeline <name> show --marimo
dlt pipeline <name> show --marimo --edit
use edit option to enable Notebook/edit mode in Marimo + very cool Ibis dataset explorer
We updated contribution guidelines
Full Changelog: https://github.com/dlt-hub/dlt/compare/1.11.0...1.12.0
Nothing published for this version
Nothing published for this version
feat: adds iceberg table properties configuration for athena (#2546) by @olexanderos in https://github.com/dlt-hub/dlt/pull/2555
Full Changelog: https://github.com/dlt-hub/dlt/compare/1.10.0...1.11.0
feat(paginator): enhance JSONResponseCursorPaginator to support cursor placement in request JSON body by @kang8 in https://github.com/dlt-hub/dlt/pull
dlt ai setup $IDE add cursor rules to REST API source to your dlt project (and more) @zilto in https://github.com/dlt-hub/dlt/pull/2503catalog_sync by @burnash in https://github.com/dlt-hub/dlt/pull/2494iceberg_mode and base_location templating by @burnash in https://github.com/dlt-hub/dlt/pull/2524We support dlt ai command via https://github.com/dlt-hub/verified-sources/tree/master/ai
Full Changelog: https://github.com/dlt-hub/dlt/compare/1.9.0...1.10.0
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →