NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3353 most downloaded on PyPI
Kedro-Datasets is where you can find all of Kedro's data connectors.
Last release 1 months ago
07 Aug 2026
Ships fairly regularly
a new release about every 2 months
Most releases are documented
notes for 38 of 45 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
45 releases · first in 2022
One column per quarter.
Breaking changes to experimental datasets
vectorstore_base.AbstractVectorStoreDataset and vectorstore_base.VectorStoreHandle, backend-agnostic abstract base classes for vector store datasets.| Type | Description | Location |
|---|---|---|
weaviate.WeaviateVectorStoreDataset |
A dataset that loads a handle for adding, searching, and deleting entries in Weaviate vector database collections. | kedro_datasets_experimental.weaviate |
faiss.FAISSVectorStoreDataset |
A dataset that loads a handle for adding, searching, and deleting entries in a FAISS vector store. | kedro_datasets_experimental.faiss |
feast.FeastDataset |
A dataset that handles storing and retrieving features from Feast. | kedro_datasets_experimental.feast |
pyspark version for the spark-base and spark-local extras is now 3.3 — users still on pyspark<3.3 must upgrade before using these extras.chromadb.ChromaDBDataset to the VectorStoreHandle approach. load_args/save_args are removed; the extras group is renamed from chromadb-chromadbdataset to chromadb-dataset.spark.SparkHiveDataset.exists() failing on Spark Connect sessions (e.g. Databricks Connect V2) by replacing the JVM-only _jsparkSession call with the PySpark Catalog.tableExists API.MLRunModel so user-supplied load_args are now passed to joblib.load() (previously silently dropped). Added a deserialization warning to the docstring.TensorFlowModelDataset: safe_mode=True is now the default for load_model() to prevent arbitrary code execution from untrusted model files. Fixed a bug where tf_device was lost from load_args after the first load call.SparkJDBCDataset JDBC URLs through credentials.os.PathLike support for NetworkX datasets.os.PathLike support for SVMLightDataset.Major features and improvements
send_individually option to APIDataset to send list items as individual requests instead of batched arrays.pytorch.PyTorchDataset: weights_only=True is now enforced by default on load to block arbitrary code execution from untrusted .pt files, user-supplied load_args and save_args are now correctly passed to torch.load and torch.save (previously silently dropped), and the misleading "pickle-safe" docstring was corrected.darts-torch-model-dataset optional dependency to point at the real PyPI package u8darts[all].polars.PolarsDatabaseDataset end-to-end and added a full test suite for it.opik.TraceDataset so credentials.project_name is now passed to configure() and persisted to Opik's session configuration.os.PathLike support for Spark datasets.ibis.FileDataset to support remote filepaths (e.g. s3://, abfss://, hf://) and added an fs_args argument to authenticate the filesystem used for version discovery.Breaking changes to experimental datasets
ArrowDataset, ParquetDataset, JSONDataset, CSVDataset.tensorflow.TensorFlowModelDataset and geopandas.GenericDataset.ibis.FileDataset.| Type | Description | Location |
|---|---|---|
opik.EvaluationDataset |
A dataset for managing Opik evaluation datasets. | kedro_datasets_experimental.opik |
PromptDataset, EvaluationDataset, TraceDataset) into a common opik._common module.PromptDataset, EvaluationDataset, TraceDataset) into a common langfuse._common module.os.PathLike support for plotly, matplotlib and pandas datasets.checkpoint.filepath validation for IncrementalDataset.README.md file for Opik experimental datasets and added information on opik.TraceDataset.opencv-python to ~=4.13.0.92 so experimental_test resolves on Python 3.14 (the old ~=4.12.0.88 capped numpy<2.3.0, which has no Windows cp314 wheel).pyproject.toml extra names for langfuse, opik, and langchain experimental datasets. The redundant package-family prefix has been dropped:
langfuse.LangfusePromptDataset → langfuse.PromptDatasetlangfuse.LangfuseTraceDataset → langfuse.TraceDatasetlangfuse.LangfuseEvaluationDataset → langfuse.EvaluationDatasetopik.OpikPromptDataset → opik.PromptDatasetopik.OpikTraceDataset → opik.TraceDatasetlangchain.LangChainPromptDataset → langchain.PromptDatasetlangfuse-langfusepromptdataset → langfuse-promptdatasetopik-opiktracedataset → opik-tracedatasetlangchain-langchainpromptdataset → langchain-promptdatasetMany thanks to the following Kedroids for contributing PRs to this release:
Major features and improvements
ibis-materialize and ibis-singlestoredb extras for the backends added in Ibis 12.0.ibis.TableDataset (available on backends that support MERGE INTO since Ibis 12.0).| Type | Description | Location |
|---|---|---|
langfuse.LangfuseEvaluationDataset |
A dataset for managing Langfuse evaluation datasets. | kedro_datasets_experimental.langfuse |
LangfuseTraceDataset documentation to the Langfuse README and restructured the page with a table of contents.databricks.ManagedTableDataset upsert write mode failing with [CONFIG_NOT_AVAILABLE] on Databricks Spark Connect runtimes by replacing spark.conf.set variable substitution with direct f-string interpolation in the MERGE SQL statement.ibis.TableDataset exists method to account for database (i.e. the collection of tables, or schema).gcsfs upper-bound pins (previously capped below 2023.7).os.PathLike support for the following dataset groups (text, json, yaml, pickle, geopandas, polars, openXML, holoview, biosequence, email and geopandas)Many thanks to the following Kedroids for contributing PRs to this release:
Priyanka
Akumawavez
Joris
Bas-commits
oomenn
Celina
Juanchodpg2
Major features and improvements
autogen mode to LangfuseTraceDataset for tracing AutoGen agent conversations with OpenTelemetry integration.api.APIDataset now stores the response received from a PUT or POST request via the response_dataset parameter.autogen mode to OpikTraceDataset for tracing AutoGen agent conversations with OpenTelemetry integration.| Type | Description | Location |
|---|---|---|
mlrun.MLRunAbstractDataset |
A base dataset for MLRun integration, can be used directly for generic artifacts | kedro_datasets_experimental.mlrun |
mlrun.MLRunModel |
A dataset for saving and loading ML models via MLRun with framework metadata | kedro_datasets_experimental.mlrun |
mlrun.MLRunDataframeDataset |
A dataset for saving and loading pandas DataFrames as MLRun artifacts | kedro_datasets_experimental.mlrun |
mlrun.MLRunResult |
A dataset for logging scalar results and metrics to MLRun | kedro_datasets_experimental.mlrun |
ibis.TableDataset exists method to account for database (i.e. the collection of tables, or schema).OpikTraceDataset and LangfuseTraceDataset now receive openai credentials as base_url and api_key, instead of openai_api_base and openai_api_key.Bump lxml version for xmldataset requirements if Python version is 3.13 and above.
Major features and improvements
Added support for Python 3.13.
Added the following new datasets:
| Type | Description | Location |
|---|---|---|
spark.SparkDatasetV2 |
A Spark dataset with Spark Connect, Databricks Connect support, and automatic Pandas-to-Spark conversion | kedro_datasets.spark |
| Type | Description | Location |
|---|---|---|
chromadb.ChromaDBDataset |
A dataset for loading and saving data to ChromaDB vector database collections | kedro_datasets_experimental.chromadb |
pandas.DeltaTableDataset to be compatible with deltalake version 1.x.plotly.JSONDataset encoding errors by defaulting thesave encoding to UTF-8.Removed the deprecated MatplotlibWriter datset. Matplotlib objects can now be handled using MatplotlibDataset.
MatplotlibWriter datset. Matplotlib objects can now be handled using MatplotlibDataset.mode save argument to ibis.TableDataset, supporting "append", "overwrite", "error"/"errorifexists", and "ignore" save modes. The deprecated overwrite save argument is mapped to mode for backward compatibility and will be removed in a future release. Specifying both mode and overwrite results in an error.ibis.TableDataset.| Type | Description | Location |
|---|---|---|
openxml.PptxDataset |
A dataset for loading and saving .pptx files (Microsoft PowerPoint) using python-pptx |
kedro_datasets.openxml |
| Type | Description | Location |
|---|---|---|
langchain.ChatOpenAIDataset |
Kedro dataset for loading a ChatOpenAI LangChain model. | kedro_datasets.langchain |
langchain.OpenAIEmbeddingsDataset |
Kedro dataset for loading an OpenAIEmbeddings model. | kedro_datasets.langchain |
langchain.ChatAnthropicDataset |
A dataset for loading a ChatAnthropic LangChain model. | kedro_datasets.langchain |
langchain.ChatCohereDataset |
A dataset for loading a ChatCohere LangChain model. | kedro_datasets.langchain |
| Type | Description | Location |
|---|---|---|
langfuse.LangfuseTraceDataset |
Kedro dataset to provide Langfuse tracing clients and callbacks | kedro_datasets_experimental.langfuse |
langchain.LangChainPromptDataset |
Kedro dataset for loading LangChain prompts | kedro_datasets_experimental.langchain |
pypdf.PDFDataset |
Kedro dataset to read PDF files and extract text using pypdf | kedro_datasets_experimental.pypdf |
langfuse.LangfusePromptDataset |
Kedro dataset for managing Langfuse prompts | kedro_datasets_experimental.langfuse |
opik.OpikPromptDataset |
A dataset to provide Opik integration for handling prompts | kedro_datasets_experimental.opik |
opik.OpikTraceDataset |
Kedro dataset to provide Opik tracing clients and callbacks | kedro_datasets_experimental.opik |
StudyDataset to properly propagate a RDB password through the dataset's credentials.Many thanks to the following Kedroids for contributing PRs to this release:
Added the following new experimental datasets:
| Type | Description | Location |
|---|---|---|
polars.PolarsDatabaseDataset |
A dataset to load and save data to a SQL backend using Polars | kedro_datasets_experimental.polars |
use_pyarrow=True save_args for LazyPolarsDataset partitioned parquet files.Make kedro-datasets compatible with Kedro 1.0.0.
kedro-datasets compatible with Kedro 1.0.0.| Type | Description | Location |
|---|---|---|
openxml.DocxDataset |
A dataset for loading and saving .docx files (Microsoft Word) using python-docx |
kedro_datasets.openxml |
PartitionedDataset to reliably load newly created partitions, particularly with ParallelRunner, by ensuring load() always re-scans the filesystem .encoding inside the dataset SQLQueryDataset to choose the encoding format of the query.APIDataset docstring to clarify that request parameters should be passed via load_args, not as top-level arguments.kedro-datasets now requires Kedro 1.0.0 or higher.Many thanks to the following Kedroids for contributing PRs to this release:
Renamed MatplotlibWriter to MatplotlibDataset for consistency with other dataset naming conventions. MatplotlibWriter is deprecated and will be remove…
PartitionedDataset.ibis-athena and ibis-databricks extras for the backends added in Ibis 10.0.MatplotlibWriter to MatplotlibDataset for consistency with other dataset naming conventions. MatplotlibWriter is deprecated and will be removed in a future release.| Type | Description | Location |
|---|---|---|
optuna.StudyDataset |
A dataset for saving and loading Optuna studies. | kedro_datasets_experimental.optuna |
darts.DartsTorchModelDataset |
A dataset for securely saving and loading Darts Torch Forecasting Models. | kedro_datasets_experimental.darts |
polars.CSVDataset save method on Windows using utf-8 as default encoding.table_name a keyword argument in the ibis.FileDataset implementation to be compatible with Ibis 10.0.snowflake.SnowflakeTableDataset implementation.pandas.GBQQueryDataset and pandas.GBQTableDataset.tracking.MetricsDataset and tracking.JSONDataset.Many thanks to the following Kedroids for contributing PRs to this release:
Added deprecation warning for tracking.MetricsDataset and tracking.JSONDataset.
database to ibis.TableDataset for load and save operations..csv ingestion.snowflake.SnowflakeTableDataset.ibis.FileDataset and ibis.TableDataset instances, thereby allowing nodes to save data loaded by one to the other (as long as they share the same connection configuration).| Type | Description | Location |
|---|---|---|
databricks.ExternalTableDataset |
A dataset for accessing external tables in Databricks. | kedro_datasets_experimental.databricks |
safetensors.SafetensorsDataset |
A dataset for securely saving and loading files in the SafeTensors format. | kedro_datasets_experimental.safetensors |
pandas.GBQTableDataset. In practice, this means that a dataset's connection details aren't used (or validated) until the dataset is accessed. On the plus side, the cost of connection isn't incurred regardless of when or whether the dataset is used. Furthermore, this makes the dataset object serializable (e.g. for use with ParallelRunner), because the unserializable client isn't part of it.pandas.GBQQueryDataset. This makes the dataset object serializable (e.g. for use with ParallelRunner) by removing the unserializable object.tracking.MetricsDataset and tracking.JSONDataset.kedro-catalog JSON schemas from Kedro core to kedro-datasets.video.VideoDataset from core to experimental dataset.ibis.TableDataset. Use ibis.FileDataset to load and save files with an Ibis backend instead.Many thanks to the following Kedroids for contributing PRs to this release:
Added the following new core datasets:
| Type | Description | Location |
|---|---|---|
ibis.FileDataset |
A dataset for loading and saving files using Ibis's backends. | kedro_datasets.ibis |
Fixed deprecated load and save approaches of GBQTableDataset and GBQQueryDataset by invoking save and load directly over pandas-gbq lib
| Type | Description | Location |
|---|---|---|
pytorch.PyTorchDataset |
A dataset for securely saving and loading PyTorch models | kedro_datasets_experimental.pytorch |
prophet.ProphetModelDataset |
A dataset for Meta's Prophet model for time series forecasting | kedro_datasets_experimental.prophet |
| Type | Description | Location |
|---|---|---|
plotly.HTMLDataset |
A dataset for saving a plotly figure as HTML |
kedro_datasets.plotly |
fs_args defaults in the same way as load_args and save_args and not have hardcoded values in the save methods.TensorFlowModelDataset.pandas-gbq libpandas optional dependencyload and save publicly for each dataset. This requires Kedro version 0.19.7 or higher.geopandas.GeoJSONDataset with geopandas.GenericDataset to support parquet and feather file formats.Many thanks to the following Kedroids for contributing PRs to this release:
Improved partitions.PartitionedDataset representation when printing.
partitions.PartitionedDataset representation when printing.ibis.TableDataset to make sure credentials are not printed in interactive environment.Added the following new experimental datasets:
| Type | Description | Location |
|---|---|---|
langchain.ChatAnthropicDataset |
A dataset for loading a ChatAnthropic langchain model. | kedro_datasets_experimental.langchain |
langchain.ChatCohereDataset |
A dataset for loading a ChatCohere langchain model. | kedro_datasets_experimental.langchain |
langchain.OpenAIEmbeddingsDataset |
A dataset for loading a OpenAIEmbeddings langchain model. | kedro_datasets_experimental.langchain |
langchain.ChatOpenAIDataset |
A dataset for loading a ChatOpenAI langchain model. | kedro_datasets_experimental.langchain |
rioxarray.GeoTIFFDataset |
A dataset for loading and saving geotiff raster data | kedro_datasets_experimental.rioxarray |
netcdf.NetCDFDataset |
A dataset for loading and saving "*.nc" files. | kedro_datasets_experimental.netcdf |
| Type | Description | Location |
|---|---|---|
dask.CSVDataset |
A dataset for loading a CSV files using dask |
kedro_datasets.dask |
yaml.YAMLDataset.metadata parameter for a few datasetsnetcdf.NetCDFDataset moved from kedro_datasets to kedro_datasets_experimental.Many thanks to the following Kedroids for contributing PRs to this release:
Removed arbitrary upper bound for s3fs.
s3fs.NetCDFDataset support for NetCDF4 via engine="netcdf4" and engine="h5netcdf"Many thanks to the following Kedroids for contributing PRs to this release:
Added the following new datasets:
| Type | Description | Location |
|---|---|---|
netcdf.NetCDFDataset |
A dataset for loading and saving *.nc files. |
kedro_datasets.netcdf |
ibis.TableDataset |
A dataset for loading and saving using Ibis's backends. | kedro_datasets.ibis |
. characters have been replaced with - in the optional dependencies names. Note that this might be breaking for some users. For example, users should now install optional dependencies for pandas.ParquetDataset from kedro-datasets like this:pip install kedro-datasets[pandas-parquetdataset]
setup.py and move to pyproject.toml completely for kedro-datasets.load_args:params will be typecasted as tuple.connection_args argument optional when calling create_connection() in sql_dataset.py.Many thanks to the following Kedroids for contributing PRs to this release:
Added MatlabDataset which uses scipy to save and load .mat files.
MatlabDataset which uses scipy to save and load .mat files.pandas.HDFDataset extra dependenciesMany thanks to the following Kedroids for contributing PRs to this release:
Removed Dataset classes ending with "DataSet", use the "Dataset" spelling instead.
huggingface.HFDataset and huggingface.HFTransformerPipelineDataset.s3fs to latest calendar-versioned release.PartitionedDataset and IncrementalDataset now both support versioning of the underlying dataset.TensorFlowModelDataset.Many thanks to the following Kedroids for contributing PRs to this release:
Added a deprecation warning when using polars.GenericDataSet or polars.GenericDataset that these have been renamed to polars.EagerPolarsDataset
PartitionedDataSet and IncrementalDataSet from the core Kedro repo to kedro-datasets and renamed to PartitionedDataset and IncrementalDataset.polars.LazyPolarsDataset, a GenericDataSet using polars's Lazy API.polars.GenericDataSet to polars.EagerPolarsDataset to better reflect the difference between the two dataset classes.polars.GenericDataSet or polars.GenericDataset that these have been renamed to polars.EagerPolarsDatasetpandas.SQLTableDataset, pandas.SQLQueryDataset, and snowflake.SnowparkTableDataset. In practice, this means that a dataset's connection details aren't used (or validated) until the dataset is accessed. On the plus side, the cost of connection isn't incurred regardless of when or whether the dataset is used.PickleDataset to explicitly mention cloudpickle support.Many thanks to the following Kedroids for contributing PRs to this release:
Renamed dataset and error classes, in accordance with the Kedro lexicon. Dataset classes ending with "DataSet" are deprecated and will be removed in 2…
tables version on kedro-datasets for Python < 3.8.Added polars.GenericDataSet, a GenericDataSet backed by polars, a lightning fast dataframe package built entirely using Rust.
polars.GenericDataSet, a GenericDataSet backed by polars, a lightning fast dataframe package built entirely using Rust.Many thanks to the following Kedroids for contributing PRs to this release:
## Major features and improvements * Added support for Python 3.11.
Made databricks.ManagedTableDataSet read-only by default.
databricks.ManagedTableDataSet read-only by default.
write_mode to allow save on the data set.api.APIDataSet where the sent data was doubly converted to json
string (once by us and once by the requests library).kedro-datasets optional dependencies, revert to setup.pyFixed problematic kedro-datasets optional dependencies.
kedro-datasets optional dependencies.Fixed problematic docstrings in pandas.DeltaTableDataSet causing Read the Docs builds on Kedro to fail.
pandas.DeltaTableDataSet causing Read the Docs builds on Kedro to fail.Implemented lazy loading of dataset subpackages and classes.
pandas.SQLQueryDataSet or pandas.SQLTableDataSet) if you load a different pandas dataset (e.g. pandas.CSVDataSet).pillow.ImageDataSet to be passed to save().pandas.DeltaTableDataSet.from kedro_datasets.pandas import SQLQueryDataSet or from kedro_datasets.pandas import SQLTableDataSet would result in ImportError: cannot import name 'SQLTableDataSet' from 'kedro_datasets.pandas'. Now, the same imports raise the more helpful and intuitive ModuleNotFoundError: No module named 'sqlalchemy'.Many thanks to the following Kedroids for contributing PRs to this release:
Fixed documentations of GeoJSONDataSet and SparkStreamingDataSet
GeoJSONDataSet and SparkStreamingDataSetFixed missing pickle.PickleDataSet extras in setup.py.
pickle.PickleDataSet extras in setup.py.Fixed problematic docstrings of APIDataSet.
SparkStreamingDataSet.APIDataSet.Added SQLAlchemy 2.0 support (and dropped support for versions below 1.4).
APIDataSet by replacing most arguments with a single constructor argument load_args. This makes it more consistent with other Kedro DataSets and the underlying requests API, and automatically enables the full configuration domain: stream, certificates, proxies, and more.>=0.16metadata attribute to all existing datasets. This is ignored by Kedro, but may be consumed by users or external plugins.ManagedTableDataSet for managed delta tables on Databricks.delta-spark upper bound to allow compatibility with Spark 3.1.x and 3.2.x.polars version to 0.17.TensorFlowModelDataset to TensorFlowModelDataSet to be consistent with all other plugins in kedro-datasets.Many thanks to the following Kedroids for contributing PRs to this release:
Added fsspec resolution in SparkDataSet to support more filesystems.
fsspec resolution in SparkDataSet to support more filesystems._preview method to the Pandas ExcelDataSet and CSVDataSet classes.SQLQueryDataSet as part of the Sphinx revamp on Kedro.Fixed problematic docstrings causing Read the Docs builds on Kedro to fail.
Added the following new datasets:
| Type | Description | Location |
|---|---|---|
polars.CSVDataSet |
A CSVDataSet backed by polars, a lighting fast dataframe package built entirely using Rust. |
kedro_datasets.polars |
snowflake.SnowparkTableDataSet |
Work with Snowpark DataFrames from tables in Snowflake. | kedro_datasets.snowflake |
mssql backend to the SQLQueryDataSet DataSet using pyodbc library.SparkDataSet on Databricks without specifying a file path with the /dbfs/ prefix.Change reference to kedro.pipeline.Pipeline object throughout test suite with kedro.modular_pipeline.pipeline factory.
kedro.pipeline.Pipeline object throughout test suite with kedro.modular_pipeline.pipeline factory.PyArrow range in line with PandasFixed doc string formatting in VideoDataSet causing the documentation builds to fail.
VideoDataSet causing the documentation builds to fail.Sync datasets with kedro.extras.datasets by @noklam in https://github.com/kedro-org/kedro-plugins/pull/69
kedro.extras.datasets by @noklam in https://github.com/kedro-org/kedro-plugins/pull/69ParquetDataSet to load using pandas instead of parquet by @SajidAlamQB in https://github.com/kedro-org/kedro-plugins/pull/89kedro-datasets 1.0.0 by @ankatiyar in https://github.com/kedro-org/kedro-plugins/pull/90Full Changelog: https://github.com/kedro-org/kedro-plugins/compare/kedro-airflow-0.5.1...kedro-datasets-1.0.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →