NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3115 most downloaded on PyPI
Kedro helps you build production-ready data and analytics pipelines
Last release 6 days ago
28 Sep 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 59 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
73 releases · first in 2019
Security fix : Custom logging handler, filter or formatter classes referenced in conf/logging.yml are no longer imported unless their module is on the…
params query parameter to GET /snapshot for resolving ${runtime_params:...} interpolation per request, using the same format as kedro run --params.llm_context_node, LLMContextNode, LLMContext and tool from experimental to stable. They no longer emit a KedroExperimentalWarning.conf/logging.yml are no longer imported unless their module is on the logging allowlist. Only the logging standard-library package (including logging.handlers) and kedro.logging are trusted by default. Projects that point conf/logging.yml at their own logging classes must now list those modules in the KEDRO_LOGGING_MODULE_ALLOWLIST environment variable (comma-separated) to keep them working.One column per quarter.
Security fix : A catalog type field resolved with runtime_params supplied through the HTTP server's POST /run no longer accepts an AbstractDataset cla…
validator in their catalog entry, and the DataCatalog enforces it on every load and save.
DataValidationError (a DatasetError subclass), reporting every failed check at once.kedro[pandera-pandas] or kedro[pandera-polars] extras; custom validator classes and functions also work.DATASET_VALIDATION setting and the KEDRO_DATASET_VALIDATION environment variable switch validation off.validate_dataset and validate_catalog validate on demand, reporting structured ValidationResult outcomes instead of raising.--runner-params to kedro run, allowing runner constructor keyword arguments such as max_workers to be passed from the CLI.NodeSnapshot.func_name.NodeSnapshot.source) to inspection snapshots for displaying node code.runtime_params to get_project_snapshot()._ProjectPipelines where concurrent access could trigger duplicate pipeline loading or operate on a cleared cache; all shared state mutations are now protected by a reentrant lock.serving_mode argument to KedroServiceSession.create() that preloads all pipelines upfront.RecursionError when initialising a session with dynaconf-backed settings by converting settings.SESSION_STORE_ARGS to a plain dict before deepcopying it.kedro_benchmarks from the built wheel so benchmark tests are no longer shipped with the package.get_close_matches returning duplicate suggestions when several inputs matched the same target, and being able to exhaust a one-shot iterable passed as targets.kedro new --tools=docs so the Sphinx HTML build runs sphinx-apidoc and creates API docs.--async flag for kedro run in favour of --runner-params=is_async=True.ParallelRunner catalog validation that identifies cloud-backed datasets (e.g. S3, GCS, Azure Data Lake) failing to pickle and suggests using ThreadRunner or SequentialRunner instead.AbstractDataset.__repr__, DatasetError messages, and HTTP server /snapshot and /run responses.runtime_params-resolved catalog type fields, in the templating guide and the HTTP server guide.kedro-skills plugin.mdformat as a Markdown autoformatter, run via make lint and as a pre-commit hook, and reformatted the existing Markdown. Fixed several step-by-step guides where numbered lists were rendering as repeated "1." instead of counting up, because their content wasn't indented under the right list item.DaskRunner example on the Dask deployment page.open_args_load examples in the data catalog documentation to use text mode so the encoding option applies.type field resolved with runtime_params supplied through the HTTP server's POST /run no longer accepts an AbstractDataset class chosen by the request. runtime_params-driven type selection through other, trusted callers (for example, kedro run --params) is unaffected.settings.py entry RUNNER_MODULES_WHITELIST has been renamed to RUNNER_MODULE_ALLOWLIST, following the <COMPONENT>_ALLOWLIST naming convention for Kedro allowlists. Projects that set the old name must rename it, otherwise the entries are ignored and the HTTP server rejects those runner modules.Many thanks to the following Kedroids for contributing PRs to this release:
Major features and improvements
KedroServiceSession to execute pipelines from HTTP requests. It offers the following endpoints:
/run to run pipelines with optional runtime parameters./health to check the server's health status./snapshot to retrieve the current project snapshot.kedro server start to run the server.--pipelines is specified, only the requested pipeline modules are imported.AbstractDataset.from_config() error message for custom dataset classes that are still abstract, so it no longer suggests invalid constructor arguments when required dataset methods are missing.kedro new accepting project names whose derived package name shadows a Python standard library module or is a Python keyword (e.g. email, json, import), which silently produced a broken, unimportable project. Such names are now rejected at creation time with a clear message.kedro pipeline create accepting Python keywords (e.g. for, import, return) as pipeline names. Such names are now rejected at creation time with a clear error message.ParallelRunner.AWSBatchRunner usage, namespace-based job submission, and S3-backed catalog configuration.Many thanks to the following Kedroids for contributing PRs to this release:
Added security model documentation covering trust boundaries, user responsibilities, and framework vulnerabilities.
KedroServiceSession, a new session implementation that allows for multiple runs and data injection.%load_node to include same-module helper dependencies via AST extraction, with explicit fallback warnings when extraction degrades to function-only source loading.kedro run --pipelines=<name> runs one or more named pipelines, only those pipelines' node signatures are inspected for type hints, avoiding unnecessary work and spurious "conflicting type" warnings from pipelines that aren't executing.None values are preserved and do not trigger validation errors.review-kedro-pr agent skill (Cursor, GitHub Copilot) for Kedro-aware PR review with optional GitHub comment posting.TRANSCODING_SEPARATOR alias from kedro.pipeline.pipeline.pydantic dependency extra, allowing users to enable Pydantic support with pip install "kedro[pydantic]".mne.Epochs) before saving them. Streaming behaviour is now restricted to generator-function nodes, so custom business objects are passed to catalog.save unchanged.KedroServiceSession.kedro.framework.session to include the new KedroServiceSession and AbstractSession classes.advanced_configuration.md with custom loader info.Fixed AttributeError when node functions have non-Pydantic/dataclass type hints on params: inputs. The parameter validation framework now correctly sk
AttributeError when node functions have non-Pydantic/dataclass type hints on params: inputs. The parameter validation framework now correctly skips types it cannot validate.Optional[Model] support and multi-type union limitations in parameter validation.Fixed a path traversal vulnerability in versioned dataset loading that could allow unauthorized file access via unsanitized version strings.
list_versions() method for versioned datasets to list available dataset versions.pipelines_to_find parameter to find_pipelines(), allowing users to selectively run a subset of existing pipelines by modifying the pipeline registry.--checkout flag can now be used on a new Kedro project from the default template, without a starter.SESSION_CLASS as a configurable project setting, allowing users to define a custom KedroSession subclassDataCatalog.load() and DataCatalog.save() now raise a DatasetError that includes the dataset name for easier debugging.before_pipeline_run, after_pipeline_run, and on_pipeline_error and the schema specified in the hooks specs.cachetools dependency and replaced it with a lightweight internal caching implementation.outputs=None, clarifying that the return value is ignored.preserve_logging flag to configure_project() to prevent runtime-added logging handlers from being overwritten when configure_project() is called after custom handlers have been attached (e.g. in a long-running server process such as FastAPI).find_config_file() to handle different config file extensions (.yml, .yaml)kedro runCatalogConfigResolver splitting sqlalchemy URL during pattern resolution.Major features and improvements
@experimental decorator to mark unstable or early-stage public APIs.--pipelines CLI option and pipeline_names argument in KedroSession.run() method.spaceflights-pyspark starter to use the new SparkDatasetV2 integration, enabling local, Databricks-native, and remote Spark execution workflows.llm_context_node and LLMContextNode for assembling LLMs, prompts, and tools into a runtime LLMContext within Kedro pipelines.preview_fn argument to Node class to add support for user-injectable node preview functions.support-agent-langgraph starter, which supports the above experimental features. This starter contains pipelines that leverage LangGraph for agentic workflows and Langfuse or Opik for prompt management and tracing.raise_errors=True in find_pipelines() calls in the project template's pipeline_registry.py to ensure pipeline discovery errors are raised during project runs.uvx installation.Spark Connect and Unity Catalog – first workflows, and local-to-remote development.Many thanks to the following Kedroids for contributing PRs to this release:
Fixed project version mismatch error. The error is now only raised when the major version of the project and Kedro package differ, allowing minor and
Major features and improvements
ignore_hidden parameter to the OmegaConfigLoader.click dependency to support versions 8.2.0 and above.deindex-old-docs.js script on Kedro documentation pages, preventing double injection of noindex meta tags after the MkDocs migration.PartionedDataset.kedro-dagster plugin to the list of community plugins.Many thanks to the following Kedroids for contributing PRs to this release:
The kedro micropkg CLI command has been removed as part of the micro-packaging feature deprecation.
KedroDataCatalog has been renamed to DataCatalog and is now the default catalog implementation.CatalogConfigResolver property.Read more in the Kedro documentation.
--namespaces CLI option and namespaces argument in KedroSession.run() method.Node class, ensuring . characters are reserved to be used as part of a namespace.prefix_datasets_with_namespace argument to the Pipeline class which allows users to turn on or off the prefixing of the namespace to the node inputs, outputs, and parameters.ParallelRunner via the KEDRO_MP_CONTEXT environment variable.--only-missing-outputs CLI flag to kedro run. This flag skips nodes when all their persistent outputs exist.kedro registry describe to return the node name property instead of creating its own name for the node.pre-commit-hooks dependency for new project creation.kedro catalog create command has been removed.kedro catalog list, kedro catalog rank, and kedro catalog resolve commands have been replaced with kedro catalog describe-datasets, kedro catalog list-patterns and kedro catalog resolve-patterns commands, respectively.kedro run option --namespace has been removed and replaced with --namespaces.kedro micropkg CLI command has been removed as part of the micro-packaging feature deprecation._is_project and _find_kedro_project are changed to is_kedro_project and find_kedro_project.extra_params and _extra_params to runtime_params.modular_pipeline module and moved functionality to the pipeline module instead.ModularPipelineError to PipelineError.Pipeline.grouped_nodes_by_namespace() was replaced with group_nodes_by(group_by), which supports multiple strategies and returns a list of GroupedNodes, improving type safety and consistency for deployment plugin integrations.session_id parameter to run_id in all runner methods and hooks to improve API clarity and prepare for future multi-run session support.DataCatalog methods: _get_dataset(), add_all(), add_feed_dict(), list(), and shallow_copy().runner.run() and session.run() — it now always returns all pipeline outputs, regardless of catalog configuration.AbstractRunner.run_only_missing() method, an older and underused API for partial runs. Please use --only-missing-outputs CLI instead.mkdocs as the documentation engine.DataCatalog documentation with improved structure and detailed description of new features. Read the DataCatalog documentation here.Many thanks to the following Kedroids for contributing PRs to this release:
See the migration guide for 1.0.0 in the Kedro documentation.
Changed DataCatalog.__getitem__ to raise DatasetNotFoundError for missing datasets, aligning with expected dictionary behavior.
Changed DataCatalog.__getitem__ to raise DatasetNotFoundError for missing datasets, aligning with expected dictionary behavior.
Added --only-missing-outputs CLI flag to kedro run. This flag skips nodes when all their persistent outputs exist.
--only-missing-outputs CLI flag to kedro run. This flag skips nodes when all their persistent outputs exist.AbstractRunner.run_only_missing() method, an older and underused API for partial runs. Please use --only-missing-outputs CLI instead.Added stricter validation to dataset names in the Node class, ensuring . characters are reserved to be used as part of a namespace.
Node class, ensuring . characters are reserved to be used as part of a namespace.prefix_datasets_with_namespace argument to the Pipeline class which allows users to turn on or off the prefixing of the namespace to the node inputs, outputs, and parameters.ParallelRunner via the KEDRO_MP_CONTEXT environment variable.kedro registry describe to return the node name property instead of creating its own name for the node.DataCatalog documentation with improved structure and detailed description of new features._is_project and _find_kedro_project are changed to is_kedro_project and find_kedro_project.extra_params and _extra_params to runtime_params.modular_pipeline module and moved functionality to the pipeline module instead.ModularPipelineError to PipelineError.Pipeline.grouped_nodes_by_namespace() was replaced with group_nodes_by(group_by), which supports multiple strategies and returns a list of GroupedNodes, improving type safety and consistency for deployment plugin integrations.micropkg CLI command have been removed.session_id parameter to run_id in all runner methods and hooks to improve API clarity and prepare for future multi-run session support.DataCatalog methods: _get_dataset(), add_all(), add_feed_dict(), list(), and shallow_copy().kedro catalog create.runner.run() — it now always returns all pipeline outputs, regardless of catalog configuration.See the migration guide for 1.0.0 in the Kedro documentation.
Made deprecation warning appear only when the deprecated --namespace flag is used
--namespace flag is usedAdded execution time to pipeline completion log.
_describe() accessed self.__dict__.Many thanks to the following Kedroids for contributing PRs to this release:
Note: On March 20th, a security vulnerability, CVE-2024-12215, was identified in Kedro. This issue stems from the deprecated micropackaging functional…
pipeline() and Pipeline into a single module (kedro.pipeline), aligning with the node()/Node design pattern and improving namespace handling.main branch version of kedro-starters instead of the respective release version.confirms during pipeline creation to support IncrementalDataset.OmegaConfcause an error during config resolution with runtime parameters.inputs in Node when created from dictionary for better performance.DEBUG to speed up the execution of project runs.kedro catalog rank, kedro catalog list, kedro catalog resolve and the kedro catalog create command will be removed.KedroDataCatalog that will replace DataCatalog while adopting the original DataCatalog name.--namespace option for kedro run. It will be replaced with --namespaces option which will allow for running multiple namespaces together.modular_pipeline module is deprecated and will be removed in Kedro 1.0.0. Use the pipeline module instead.Note: On March 20th, a security vulnerability, CVE-2024-12215, was identified in Kedro. This issue stems from the deprecated micropackaging functionality, which is scheduled for removal in the upcoming Kedro 1.0 release. While we agree with the CVE assigned, this vulnerability only poses a risk if you pull a malicious micropackage from an untrusted source. If you're concerned, we recommend avoiding the micropackaging feature for now and upgrading to Kedro 1.0 once it's released.
Many thanks to the following Kedroids for contributing PRs to this release:
Added DataCatalog deprecation warning.
KedroDataCatalog.filter() to filter datasets by name and type.Pipeline.grouped_nodes_by_namespace property which returns a dictionary of nodes grouped by namespace, intended to be used by plugins to facilitate deployment of namespaced nodes together.--conf-source, allowing configuration to be loaded from remote locations such as S3.DataCatalog deprecation warning._LazyDataset representation when printing KedroDataCatalog.MemoryDataset to infer assign copy mode for Ibis Tables, which previously would be inferred as deepcopy.pipelines/__init__.py exists when creating new pipelines.SequentialRunner to not use an executor pool to ensure it's single threaded.%load_node magic command to work with Jupyter Notebook >=7.2.0.7: Kedro Viz from Kedro tools.Many thanks to the following Kedroids for contributing PRs to this release:
Implemented KedroDataCatalog.to_config() method that converts the catalog instance into a configuration format suitable for serialization.
KedroDataCatalog.to_config() method that converts the catalog instance into a configuration format suitable for serialization.trufflehog with detect-secrets for detecting secrets within a code base.%load_ext kedro.node import to the pipeline template.kedro-catalog JSON schema to kedro-datasets.Partitioned dataset lazy saving docs page.KedroDataCatalog mutation after pipeline run.KedroDataCatalog._datasets compatible with DataCatalog._datasets.Many thanks to the following Kedroids for contributing PRs to this release:
Note: KedroDataCatalog is an experimental feature and is under active development. Therefore, it is possible we'll introduce breaking changes to this…
KedroDataCatalog.KedroDataCatalog.pyproject.toml file, allowing Kedro projects to work with project management tools like uv, pdm, and rye.Note: KedroDataCatalog is an experimental feature and is under active development. Therefore, it is possible we'll introduce breaking changes to this class, so be mindful of that if you decide to use it already. Let us know if you have any feedback about the KedroDataCatalog or ideas for new features.
DatasetAlreadyExistsError for ThreadRunner when Kedro project run and using runner separately..parquet suffix in docs and tests.Many thanks to the following Kedroids for contributing PRs to this release:
Removed ShelveStore to address a security vulnerability.
KedroDataCatalog repeating DataCatalog functionality with a few API enhancements:
_FrozenDatasets and access datasets as properties;add_feed_dict() was simplified to only add raw data;from_config() method to the constructor.requirements.txt to the dedicated section in pyproject.toml for project template.Protocol abstraction for the current DataCatalog and adding new catalog implementations.kedro run and kedro catalog commands.DataCatalog to a separate component - CatalogConfigResolver. Updated DataCatalog to use CatalogConfigResolver internally.session.run() output to be used when running it in the interactive environment.OmegaConfigLoader configuration validation to detect duplicate keys at all parameter levels, ensuring comprehensive nested key checking.Note: KedroDataCatalog is an experimental feature and is under active development. Therefore, it is possible we'll introduce breaking changes to this class, so be mindful of that if you decide to use it already. Let us know if you have any feedback about the KedroDataCatalog or ideas for new features.
ThreadRunner.SharedMemoryDataset.exists would not call the underlying MemoryDataset.KedroContext._get_catalog() and resolve_patterns so that both use _get_config_credentials()ShelveStore to address a security vulnerability.Made default run entrypoint in __main__.py work in interactive environments such as IPyhon and Databricks.
__main__.py work in interactive environments such as IPyhon and Databricks._find_run_command() and _find_run_command_in_plugins() from __main__.py in the project template to the framework itself.%load_node breaks with multi-lines import statements.rich mark up logs stop showing since 0.19.7.Many thanks to the following Kedroids for contributing PRs to this release:
The utility method get_pkg_version() is deprecated and will be removed in Kedro 0.20.0.
load and save publicly for each dataset in the core kedro library, and enabled other datasets to do the same. If a dataset doesn't expose load or save publicly, Kedro will fall back to using _load or _save, respectively.DataCatalog.DataCatalog pretty printing.DataCatalog shallow_copy() method to ensure it returns the type of the used catalog and doesn't cast it to DataCatalog.OmegaConfigLoader is printed, there are few missing arguments.OmegaConfigLoader's keys return empty dictionary.get_pkg_version() is deprecated and will be removed in Kedro 0.20.0.Many thanks to the following Kedroids for contributing PRs to this release:
All micro-packaging commands (kedro micropkg pull, kedro micropkg package) are deprecated and will be removed in Kedro 0.20.0.
raise_errors argument to find_pipelines. If True, the first pipeline for which autodiscovery fails will cause an error to be raised. The default behaviour is still to raise a warning for each failing pipeline.rich installed.conf/logging.yml will be used if it exists and KEDRO_LOGGING_CONFIG is not set; otherwise, default_logging.yml will be used.kedro micropkg pull, kedro micropkg package) are deprecated and will be removed in Kedro 0.20.0.globals and runtime_params with the OmegaConfigLoaderMany thanks to the following Kedroids for contributing PRs to this release:
Fixed breaking import issue when working on a project with kedro-viz on python 3.8.
kedro-viz on python 3.8.kedro-sphinx-theme for documentation.Kedro commands now work from any subdirectory within a Kedro project.
kedro run--telemetry flag to kedro new, allowing the user to register consent to have user analytics collected at the same time as the project is created.Pipeline object creation and summing.toposort in favour of the built-in graphlib module.--verbose flag.kedro pipeline create and kedro pipeline delete to read the base environment from the project settings.kedro catalog resolve to read credentials properly.kedro pipeline create from <project root>/src/tests/pipelines/<pipeline name> to <project root>/tests/pipelines/<pipeline name>..gitignore to prevent pushing Mlflow local runs folder to a remote forge when using mlflow and git.node-creation allowing self-dependencies when using transcoding, that is datasets named like name@format._is_project and _find_kedro_project have been moved to kedro.utils. We recommend not using private methods in your code, but if you do, please update your code to use the new location.merge_strategy argument in OmegaConfigLoader.Many thanks to the following Kedroids for contributing PRs to this release:
Addressed arbitrary file write via archive extraction security vulnerability in micropackaging.
%load_node for Jupyter Notebook and Jupyter Lab.%load_node and minimal support for Databricks.%load_node.kedro catalog resolve to work with dataset factories that use PartitionedDataset._EPHEMERAL attribute to AbstractDataset and other Dataset classes that inherit from it.kedro-telemetry and the data collected by it.Many thanks to the following Kedroids for contributing PRs to this release:
Removed example pipeline requirements when examples are not selected in tools.
tools.source_dir explicitly in pyproject.toml for non-src layout project.MemoryDataset entries are now included in free outputs.ruff format.SequentiallRunner and ParallelRunner.bootstrap_project and configure_project.kedro run and hook execution order.Loosen bound for kedro-telemetry by @merelcht in https://github.com/kedro-org/kedro/pull/3417
Release 0.19.1
kedro-telemetry by @merelcht in https://github.com/kedro-org/kedro/pull/3417:rocket: Major Features and improvements
:rocket: Major Features and improvements
--starter flag.--conf-source option to %reload_kedro, allowing users to specify a source for project configuration._ProjectSettings. This enables the use of config loader as a standalone class without affecting existing Kedro Framework users.:beetle: Bug fixes and other changes
:boom: Breaking changes
kedro.io (import them from kedro-datasets instead)KEDRO_LOGGING_CONFIG.data_set and DataSet to dataset and Dataset everywhere.create_default_data_set() method in the Runner in favour of using dataset factories to create default dataset instances.:writing_hand: Documentation changes
New Contributors
Full Changelog: https://github.com/kedro-org/kedro/compare/0.18.14...0.19.0
:rotating_light: If you are upgrading from Kedro 0.18, have a look at the migration guide.
We welcome every community contribution, large or small. See what we're working on now and report bugs or suggest future features. Until next time, The Kedro Team :yellow_heart:
All dataset classes ending with DataSet are deprecated and will be removed in Kedro 0.19.0 and kedro-datasets 2.0.0. Instead, use the updated class na…
--template flag for kedro pipeline create or via template/pipeline folder.runtime_params resolver with OmegaConfigLoader.OmegaConfigLoader to handle paths containing dots outside of conf_source.settings.py optional.standalone-datacatalog starter into its README file.kedro.extras.datasets). Install and import them from the kedro-datasets package instead.DataSet are deprecated and will be removed in Kedro 0.19.0 and kedro-datasets 2.0.0. Instead, use the updated class names ending with Dataset.pandas-iris, pyspark-iris, pyspark, and standalone-datacatalog are deprecated and will be archived in Kedro 0.19.0.PartitionedDataset and IncrementalDataset have been moved to kedro-datasets and will be removed in Kedro 0.19.0. Install and import them from the kedro-datasets package instead.Many thanks to the following Kedroids for contributing PRs to this release:
Spark Dependency: We've set an upper version limit for pyspark at <3.4 due to breaking changes in 3.4.
OmegaConfigLoader features:
OmegaConfigLoader through CONFIG_LOADER_ARGS.OmegaConfigLoader.kedro catalog resolve CLI command that resolves dataset factories in the catalog with any explicit entries in the project pipeline.conf/ structure for modular pipelines, and accordingly, updated the kedro pipeline create and kedro catalog create command.OmegaConfigLoader.setup.py in new Kedro project template and Kedro starters to pyproject.toml and moved flake8 configuration
to dedicated file .flake8.conf/ structure.OmegaConfigLoader to ignore config from hidden directories like .ipynb_checkpoints.data section to restructure beginner and advanced pages about the Data Catalog and datasets.ConfigLoader and the TemplatedConfigLoader to the OmegaConfigLoader. The ConfigLoader and the TemplatedConfigLoader are deprecated and will be removed in the 0.19.0 release.pytables to 3.8.0 due to compatibility issues.pyspark at <3.4 due to breaking changes in 3.4.moto version now supports parallel test execution for Python 3.10, resolving previous issues.kedro.io; only the module where they are defined is listed as the location.| Type | Deprecated Alias | Location |
|---|---|---|
AbstractDataset |
AbstractDataSet |
kedro.io.core |
AbstractVersionedDataset |
AbstractVersionedDataSet |
kedro.io.core |
layer attribute at the top level is deprecated; it will be removed in Kedro version 0.19.0. Please move layer inside the metadata -> kedro-viz attributes.Thanks to Laíza Milena Scheid Parizotto and Jonathan Cohen.
ConfigLoader and TemplatedConfigLoader will be deprecated. Please use OmegaConfigLoader instead.
OmegaConfigLoader except for oc.env.kedro catalog rank CLI command that ranks dataset factories in the catalog by matching priority.pyproject.toml.kedro catalog list to show datasets generated with factories.ruff as the linter and removed mentions of pylint, isort, flake8.Thanks to Laíza Milena Scheid Parizotto and Chris Schopp.
ConfigLoader and TemplatedConfigLoader will be deprecated. Please use OmegaConfigLoader instead.Added databricks-iris as an official starter.
Rebrand across all documentation and Kedro assets.
kedro run --params now updates interpolated parameters correctly when using OmegaConfigLoader.
kedro run --params now updates interpolated parameters correctly when using OmegaConfigLoader.metadata attribute to kedro.io datasets. This is ignored by Kedro, but may be consumed by users or external plugins.kedro.logging.RichHandler. This replaces the default rich.logging.RichHandler and is more flexible, user can turn off the rich traceback if needed.OmegaConfigLoader will return a dict instead of DictConfig.OmegaConfigLoader does not show a MissingConfigError when the config files exist but are empty.kedro package does not produce .egg files anymore, and now relies exclusively on .whl files.Many thanks to the following Kedroids for contributing PRs to this release:
Added deprecation warnings about the removal of kedro.extras.datasets.
KEDRO_LOGGING_CONFIG environment variable, which can be used to configure logging from the beginning of the kedro process.kedro run CLI command to session store to improve run reproducibility using Kedro-Viz experiment tracking.flake8 configuration.kedro.extras.datasets.Added new Kedro CLI kedro jupyter setup to setup Jupyter Kernel for Kedro.
kedro jupyter setup to setup Jupyter Kernel for Kedro.kedro package now includes the project configuration in a compressed tar.gz file.OmegaConfigLoader to load configuration from compressed files of zip or tar format. This feature requires fsspec>=2023.1.0._ProjectPipeline.Fixed bug that didn't allow to read or write datasets with s3a or s3n filepaths
s3a or s3n filepaths--params flagKedro-Viz experiment trackingA regression introduced in Kedro version 0.18.5 caused the Kedro-Viz console to fail to show experiment tracking correctly. If you experienced this issue, you will need to:
0.18.6<project-path>/data/session_store.db.Thanks to Kedroids tomohiko kato, tsanikgr and maddataanalyst for very detailed reports about the bug.
Added safe extraction of tar files in micropkg pull to fix vulnerability caused by CVE-2007-4559.
NOTE: This version of Kedro introduced a bug such that the Kedro-Viz console to fail to show experiment tracking correctly. We recommend that you don't use it and prefer instead to use Kedro version
0.18.6.
OmegaConfigLoader which uses OmegaConf for loading and merging configuration.--conf-source option to kedro run, allowing users to specify a source for project configuration for the run.omegaconf syntax as option for --params. Keys and values can now be separated by colons or equals signs.yield instead of return.
yield before proceeding with next chunk.OmegaConfigLoader.--namespace flag to kedro run to enable filtering by node namespace.node for all four dataset hooks.kedro run flags --nodes, --tags, and --load-versions to replace --node, --tag, and --load-version.kedro run options which take a list of nodes as inputs (--from-nodes and --to-nodes).micropkg manifest section in pyproject.toml isn't recognised as allowed configuration.load_ipython_extension not to register the %reload_kedro line magic when called in a directory that does not contain a Kedro project.anyconfig's ac_context parameter to kedro.config.commons module functions for more flexible ConfigLoader customizations.kedro.pipeline.Pipeline object throughout test suite with kedro.modular_pipeline.pipeline factory.after_dataset_saved hook only to be called for one output dataset when multiple are saved in a single node and async saving is in use.WARNING to DEBUG.micropkg pull to fix vulnerability caused by CVE-2007-4559.kedro runMany thanks to the following Kedroids for contributing PRs to this release:
project_version will be deprecated in pyproject.toml please use kedro_init_version instead.kedro run flags --node, --tag, and --load-version in favour of --nodes, --tags, and --load-versions.kedro test and kedro lint will be deprecated.
kedro_datasets with higher priority than kedro.extras.datasets. kedro_datasets is the namespace for the new kedro-datasets python package.UserDict and the configuration is accessed through conf_loader['catalog'].settings.py without creating a custom config loader.| Type | Description | Location |
|---|---|---|
svmlight.SVMLightDataSet |
Work with svmlight/libsvm files using scikit-learn library | kedro.extras.datasets.svmlight |
video.VideoDataSet |
Read and write video files from a filesystem | kedro.extras.datasets.video |
video.video_dataset.SequenceVideo |
Create a video object from an iterable sequence to use with VideoDataSet |
kedro.extras.datasets.video |
video.video_dataset.GeneratorVideo |
Create a video object from a generator to use with VideoDataSet |
kedro.extras.datasets.video |
dask.ParquetDataSet to work with the dask.to_parquet API.kedro micropkg pull for packages on PyPI.format in save_args for SparkHiveDataSet, previously it didn't allow you to save it as delta format.TensorFlowModelDataset when used without versioning; previously, it wouldn't overwrite an existing model.tf.device in TensorFlowModelDataset.VersionNotFoundError to handle insufficient permission issues for cloud storage.local_ns rather than a global variable.ShelveStore to its own module to ensure multiprocessing works with it.kedro.extras.datasets.pandas.SQLQueryDataSet now takes optional argument execution_options.attrs upper bound to support newer versions of Airflow.setuptools dependency to <=61.5.1.kedro test and kedro lint will be deprecated.We are grateful to the following for submitting PRs that contributed to this release: jstammers, FlorianGD, yash6318, carlaprv, dinotuku, williamcaicedo, avan-sh, Kastakin, amaralbf, BSGalvan, levimjoseph, daniel-falk, clotildeguinard, avsolatorio, and picklejuicedev for comments and input to documentation changes
kedro jupyter convert, kedro build-docs, kedro build-reqs and kedro activate-nbstripout will be deprecated.
Implemented autodiscovery of project pipelines. A pipeline created with kedro pipeline create <pipeline_name> can now be accessed immediately without needing to explicitly register it in src/<package_name>/pipeline_registry.py, either individually by name (e.g. kedro run --pipeline=<pipeline_name>) or as part of the combined default pipeline (e.g. kedro run). By default, the simplified register_pipelines() function in pipeline_registry.py looks like:
def register_pipelines() -> Dict[str, Pipeline]:
"""Register the project's pipelines.
Returns:
A mapping from pipeline names to ``Pipeline`` objects.
"""
pipelines = find_pipelines()
pipelines["__default__"] = sum(pipelines.values())
return pipelines
The Kedro IPython extension should now be loaded with %load_ext kedro.ipython.
The line magic %reload_kedro now accepts keywords arguments, e.g. %reload_kedro --env=prod.
Improved resume pipeline suggestion for SequentialRunner, it will backtrack the closest persisted inputs to resume.
False value for rich logging show_locals, to make sure credentials and other sensitive data isn't shown in logs.rich.kedro run -n [some_node], if some_node is missing a namespace the resulting error message will suggest the correct node name.rich logging.delta-spark upper bound to allow compatibility with Spark 3.1.x and 3.2.x.gdrive to list of cloud protocols, enabling Google Drive paths for datasets.%load_ext kedro.extras.extensions.ipython; use %load_ext kedro.ipython instead.kedro jupyter convert, kedro build-docs, kedro build-reqs and kedro activate-nbstripout will be deprecated.Required cookiecutter>=2.1.1 to address a known command injection vulnerability.
abfss to list of cloud protocols, enabling abfss paths.conf/base/logging.yml is now optional. See our documentation for details.kedro.starters entry point. This enables plugins to create custom starter aliases used by kedro starter list and kedro new.kedro new prompts to just one question asking for the project name.pyyaml upper bound to make Kedro compatible with the pyodide stack.myst_parser instead of recommonmark.INFO to DEBUG for low priority messages.info.log/errors.log files are no longer created in your project root, and running Kedro on read-only file systems such as Databricks Repos is now possible.root logger is now set to the Python default level of WARNING rather than INFO. Kedro's logger is still set to emit INFO level messages.SequentialRunner now has consistent execution order across multiple runs with sorted nodes.kedro jupyter notebook/lab no longer reuses a Jupyter kernel.cookiecutter>=2.1.1 to address a known command injection vulnerability.getpass.getuser.AbstractDataSet and AbstractVersionedDataSet as well as typing to all datasets.kedro.config.default_logger no longer exists; default logging configuration is now set automatically through kedro.framework.project.LOGGING. Unless you explicitly import kedro.config.default_logger you do not need to make any changes.kedro.extras.ColorHandler will be removed in 0.19.0.Added a new hook after_context_created that passes the KedroContext instance as context.
after_context_created that passes the KedroContext instance as context.after_command_run.ParserError exception error message.SparkDataSet to specify a schema load argument that allows for supplying a user-defined schema as opposed to relying on the schema inference of Spark.CONFIG_LOADER_CLASS validation so that TemplatedConfigLoader can be specified in settings.py. Any CONFIG_LOADER_CLASS must be a subclass of AbstractConfigLoader.run_params dictionary used in pipeline hooks.Jinja2 syntax loading with TemplatedConfigLoader using globals.yml._active_session, _activate_session and _deactivate_session. Plugins that need to access objects such as the config loader should now do so through context in the new after_context_created hook.config_loader is available as a public read-only attribute of KedroContext.hook_manager argument optional for runner.run.kedro docs now opens an online version of the Kedro documentation instead of a locally built version.kedro docs will be removed in 0.19.0.Removed deprecated functions load_context and get_project_context.
Kedro 0.18.0 strives to reduce the complexity of the project template and get us closer to a stable release of the framework. We've introduced the full micro-packaging workflow 📦, which allows you to import packages, utility functions and existing pipelines into your Kedro project. Integration with IPython and Jupyter has been streamlined in preparation for enhancements to Kedro's interactive workflow. Additionally, the release comes with long-awaited Python 3.9 and 3.10 support 🐍.
kedro.config.abstract_config.AbstractConfigLoader as an abstract base class for all ConfigLoader implementations. ConfigLoader and TemplatedConfigLoader now inherit directly from this base class.ConfigLoader.get and TemplatedConfigLoader.get API and delegated the actual get method functional implementation to the kedro.config.common module.hook_manager is no longer a global singleton. The hook_manager lifecycle is now managed by the KedroSession, and a new hook_manager will be created every time a session is instantiated.pipeline() without the params: prefix.Pipeline.filter() (previously in KedroContext._filter_pipeline()) to filter parts of a pipeline.username to Session store for logging during Experiment Tracking.from my_package.__main__ import main
main(
["--pipleine", "my_pipeline"]
) # or just main() if no parameters are needed for the run
cli.py from the Kedro project template. By default, all CLI commands, including kedro run, are now defined on the Kedro framework side. You can still define custom CLI commands by creating your own cli.py.hooks.py from the Kedro project template. Registration hooks have been removed in favour of settings.py configuration, but you can still define execution timeline hooks by creating your own hooks.py..ipython directory from the Kedro project template. The IPython/Jupyter workflow no longer uses IPython profiles; it now uses an IPython extension.kedro run configuration environment names can now be set in settings.py using the CONFIG_LOADER_ARGS variable. The relevant keyword arguments to supply are base_env and default_run_env, which are set to base and local respectively by default.| Type | Description | Location |
|---|---|---|
pandas.XMLDataSet |
Read XML into Pandas DataFrame. Write Pandas DataFrame to XML | kedro.extras.datasets.pandas |
networkx.GraphMLDataSet |
Work with NetworkX using GraphML files | kedro.extras.datasets.networkx |
networkx.GMLDataSet |
Work with NetworkX using Graph Modelling Language files | kedro.extras.datasets.networkx |
redis.PickleDataSet |
loads/saves data from/to a Redis database | kedro.extras.datasets.redis |
partitionBy support and exposed save_args for SparkHiveDataSet.open_args_save in fs_args for pandas.ParquetDataSet.load and save operations for pandas datasets in order to leverage pandas own API and delegate fsspec operations to them. This reduces the need to have our own fsspec wrappers.pandas.AppendableExcelDataSet into pandas.ExcelDataSet.save_args to feather.FeatherDataSet.%load_ext kedro.extras.extensions.ipython and use the line magic %reload_kedro.kedro ipython launches an IPython session that preloads the Kedro IPython extension.kedro jupyter notebook/lab creates a custom Jupyter kernel that preloads the Kedro IPython extension and launches a notebook with that kernel selected. There is no longer a need to specify --all-kernels to show all available kernels.pandas to 1.3. Any storage_options should continue to be specified under fs_args and/or credentials.black dependency in the project template to a non pre-release version.RegistrationSpecs and its associated register_config_loader and register_catalog hook specifications in favour of CONFIG_LOADER_CLASS/CONFIG_LOADER_ARGS and DATA_CATALOG_CLASS in settings.py.load_context and get_project_context.CONF_SOURCE, package_name, pipeline, pipelines, config_loader and io attributes from KedroContext as well as the deprecated KedroContext.run method.PluginManager hook_manager argument to KedroContext and the Runner.run() method, which will be provided by the KedroSession.get_hook_manager() and replaced its functionality by _create_hook_manager().KedroSession. run_id has been renamed to session_id as a result.settings.py setting CONF_ROOT has been renamed to CONF_SOURCE. Default value of conf remains unchanged.ConfigLoader and TemplatedConfigLoader argument conf_root has been renamed to conf_source.extra_params has been renamed to runtime_params in kedro.config.config.ConfigLoader and kedro.config.templated_config.TemplatedConfigLoader.KedroContext and is now implemented in a ConfigLoader class (or equivalent) with the base_env and default_run_env attributes.pandas.ExcelDataSet now uses openpyxl engine instead of xlrd.pandas.ParquetDataSet now calls pd.to_parquet() upon saving. Note that the argument partition_cols is not supported.spark.SparkHiveDataSet API has been updated to reflect spark.SparkDataSet. The write_mode=insert option has also been replaced with write_mode=append as per Spark styleguide. This change addresses Issue 725 and Issue 745. Additionally, upsert mode now leverages checkpoint functionality and requires a valid checkpointDir be set for current SparkContext.yaml.YAMLDataSet can no longer save a pandas.DataFrame directly, but it can save a dictionary. Use pandas.DataFrame.to_dict() to convert your pandas.DataFrame to a dictionary before you attempt to save it to YAML.open_args_load and open_args_save from the following datasets:
pandas.CSVDataSetpandas.ExcelDataSetpandas.FeatherDataSetpandas.JSONDataSetpandas.ParquetDataSetstorage_options are now dropped if they are specified under load_args or save_args for the following datasets:
pandas.CSVDataSetpandas.ExcelDataSetpandas.FeatherDataSetpandas.JSONDataSetpandas.ParquetDataSetlambda_data_set, memory_data_set, and partitioned_data_set to lambda_dataset, memory_dataset, and partitioned_dataset, respectively, in kedro.io.networkx.NetworkXDataSet has been renamed to networkx.JSONDataSet.kedro install in favour of pip install -r src/requirements.txt to install project dependencies.--parallel flag from kedro run in favour of --runner=ParallelRunner. The -p flag is now an alias for --pipeline.kedro pipeline package has been replaced by kedro micropkg package and, in addition to the --alias flag used to rename the package, now accepts a module name and path to the pipeline or utility module to package, relative to src/<package_name>/. The --version CLI option has been removed in favour of setting a __version__ variable in the micro-package's __init__.py file.kedro pipeline pull has been replaced by kedro micropkg pull and now also supports --destination to provide a location for pulling the package.kedro pipeline list and kedro pipeline describe in favour of kedro registry list and kedro registry describe.kedro package and kedro micropkg package now save egg and whl or tar files in the <project_root>/dist folder (previously <project_root>/src/dist).kedro build-reqs to compile requirements from requirements.txt instead of requirements.in and save them to requirements.lock instead of requirements.txt.kedro jupyter notebook/lab no longer accept --all-kernels or --idle-timeout flags. --all-kernels is now the default behaviour.KedroSession.run now raises ValueError rather than KedroContextError when the pipeline contains no nodes. The same ValueError is raised when there are no matching tags.KedroSession.run now raises ValueError rather than KedroContextError when the pipeline name doesn't exist in the pipeline registry..tar.gz).Node and Pipeline, as well as the modules kedro.extras.decorators and kedro.pipeline.decorators.DataCatalog, as well as the modules kedro.extras.transformers and kedro.io.transformers.Journal and DataCatalogWithDefault.%init_kedro IPython line magic, with its functionality incorporated into %reload_kedro. This means that if %reload_kedro is called with a filepath, that will be set as default for subsequent calls.hook_impl of the register_config_loader and register_catalog methods from ProjectHooks in hooks.py (or custom alternatives).run_id in the after_catalog_created hook, replace it with save_version instead.run_id in any of the before_node_run, after_node_run, on_node_error, before_pipeline_run, after_pipeline_run or on_pipeline_error hooks, replace it with session_id instead.settings.py filekedro.config.TemplatedConfigLoader, alter CONFIG_LOADER_CLASS to specify the class and CONFIG_LOADER_ARGS to specify keyword arguments. If not set, these default to kedro.config.ConfigLoader and an empty dictionary respectively.DATA_CATALOG_CLASS to specify the class. If not set, this defaults to kedro.io.DataCatalog.conf), update CONF_ROOT to CONF_SOURCE and set it to a string with the expected configuration location. If not set, this defaults to "conf".For a given pipeline:
active_pipeline = pipeline(
pipe=[
node(
func=some_func,
inputs=["model_input_table", "params:model_options"],
outputs=["**my_output"],
),
...,
],
inputs="model_input_table",
namespace="candidate_modelling_pipeline",
)
The parameters should look like this:
-model_options:
- test_size: 0.2
- random_state: 8
- features:
- - engines
- - passenger_capacity
- - crew
+candidate_modelling_pipeline:
+ model_options:
+ test_size: 0.2
+ random_state: 8
+ features:
+ - engines
+ - passenger_capacity
+ - crew
params: prefix when supplying values to parameters argument in a pipeline() call.kedro pipeline pull my_pipeline --alias other_pipeline, now use kedro micropkg pull my_pipeline --alias pipelines.other_pipeline instead.kedro pipeline package my_pipeline, now use kedro micropkg package pipelines.my_pipeline instead.pyproject.toml, you should modify the keys to include the full module path, and wrapped in double-quotes, e.g:[tool.kedro.micropkg.package]
-data_engineering = {destination = "path/to/here"}
-data_science = {alias = "ds", env = "local"}
+"pipelines.data_engineering" = {destination = "path/to/here"}
+"pipelines.data_science" = {alias = "ds", env = "local"}
[tool.kedro.micropkg.pull]
-"s3://my_bucket/my_pipeline" = {alias = "aliased_pipeline"}
+"s3://my_bucket/my_pipeline" = {alias = "pipelines.aliased_pipeline"}
pandas.ExcelDataSet, make sure you have openpyxl installed in your environment. This is automatically installed if you specify kedro[pandas.ExcelDataSet]==0.18.0 in your requirements.txt. You can uninstall xlrd if you were only using it for this dataset.pandas.ParquetDataSet, pass pandas saving arguments directly to save_args instead of nested in from_pandas (e.g. save_args = {"preserve_index": False} instead of save_args = {"from_pandas": {"preserve_index": False}}).spark.SparkHiveDataSet with write_mode option set to insert, change this to append in line with the Spark styleguide. If you use spark.SparkHiveDataSet with write_mode option set to upsert, make sure that your SparkContext has a valid checkpointDir set either by SparkContext.setCheckpointDir method or directly in the conf folder.pandas~=1.2.0 and pass storage_options through load_args or savs_args, specify them under fs_args or via credentials instead.kedro.io.lambda_data_set, kedro.io.memory_data_set, or kedro.io.partitioned_data_set, change the import to kedro.io.lambda_dataset, kedro.io.memory_dataset, or kedro.io.partitioned_dataset, respectively (or import the dataset directly from kedro.io).pandas.AppendableExcelDataSet entries in your catalog, replace them with pandas.ExcelDataSet.networkx.NetworkXDataSet entries in your catalog, replace them with networkx.JSONDataSet.kedro pipeline package --version to use kedro micropkg package instead. If you wish to set a specific pipeline package version, set the __version__ variable in the pipeline package's __init__.py file.kedro run --runner=ParallelRunner rather than --parallel or -p.ConfigLoader or TemplatedConfigLoader directly, update the keyword arguments conf_root to conf_source and extra_params to runtime_params.KedroContext to access ConfigLoader, use settings.CONFIG_LOADER_CLASS to access the currently used ConfigLoader instead.Bumped the Pillow minimum version requirement to 9.0 (Python 3.7+ only) following CVE-2022-22817.
pipeline now accepts tags and a collection of Nodes and/or Pipelines rather than just a single Pipeline object. pipeline should be used in preference to Pipeline when creating a Kedro pipeline.pandas.SQLTableDataSet and pandas.SQLQueryDataSet now only open one connection per database, at instantiation time (therefore at catalog creation time), rather than one per load/save operation.micropkg, to replace kedro pipeline pull and kedro pipeline package with kedro micropkg pull and kedro micropkg package for Kedro 0.18.0. kedro micropkg package saves packages to project/dist while kedro pipeline package saves packages to project/src/dist.pandas<1.4 to maintain compatibility with xlrd~=1.0.Pillow minimum version requirement to 9.0 (Python 3.7+ only) following CVE-2022-22817.PickleDataSet to be copyable and hence work with the parallel runner.pip-tools, which is used by kedro build-reqs, to 6.5 (Python 3.7+ only). This pip-tools version is compatible with pip>=21.2, including the most recent releases of pip. Python 3.6 users should continue to use pip-tools 6.4 and pip<22.astro-iris as alias for astro-airlow-iris, so that old tutorials can still be followed.kedro pipeline pull and kedro pipeline package will be deprecated. Please use kedro micropkg instead.Deprecated the "Thanks for supporting contributions" section of release notes to simplify the contribution process; Kedro 0.17.6 is the last release t…
pipelines global variable to IPython extension, allowing you to access the project's pipelines in kedro ipython or kedro jupyter notebook.params in CLI, i.e. kedro run --params="model.model_tuning.booster:gbtree" updates parameters to {"model": {"model_tuning": {"booster": "gbtree"}}}.pandas.SQLQueryDataSet to specify a filepath with a SQL query, in addition to the current method of supplying the query itself in the sql argument.ExcelDataSet to support saving Excel files with multiple sheets.| Type | Description | Location |
|---|---|---|
plotly.JSONDataSet |
Works with plotly graph object Figures (saves as json file) | kedro.extras.datasets.plotly |
pandas.GenericDataSet |
Provides a 'best effort' facility to read / write any format provided by the pandas library |
kedro.extras.datasets.pandas |
pandas.GBQQueryDataSet |
Loads data from a Google Bigquery table using provided SQL query | kedro.extras.datasets.pandas |
spark.DeltaTableDataSet |
Dataset designed to handle Delta Lake Tables and their CRUD-style operations, including update, merge and delete |
kedro.extras.datasets.spark |
kedro new --config config.yml was ignoring the config file when prompts.yml didn't exist.kedro viz --autoreload.pickle interface to PickleDataSet.sum syntax for connecting pipeline objects.pip-tools, which is used by kedro build-reqs, to 6.4. This pip-tools version requires pip>=21.2 while adding support for pip>=21.3. To upgrade pip, please refer to their documentation.plotly requirement for plotly.PlotlyDataSet and the pyarrow requirement for pandas.ParquetDataSet.kedro pipeline package <pipeline> now raises an error if the <pipeline> argument doesn't look like a valid Python module path (e.g. has / instead of .).overwrite argument to PartitionedDataSet and MatplotlibWriter to enable deletion of existing partitions and plots on dataset save.kedro pipeline pull now works when the project requirements contains entries such as -r, --extra-index-url and local wheel files (Issue #913)._FrozenDatasets creations..coveragerc from the Kedro project template. coverage settings are now given in pyproject.toml.git.load_versions that are not found in the data catalog would silently pass.kedro.extras.decorators and kedro.pipeline.decorators are being deprecated in favour of Hooks.kedro.extras.transformers and kedro.io.transformers are being deprecated in favour of Hooks.--parallel flag on kedro run is being removed in favour of --runner=ParallelRunner. The -p flag will change to be an alias for --pipeline.kedro.io.DataCatalogWithDefault is being deprecated, to be removed entirely in 0.18.0.Deepyaman Datta, Brites, Manish Swami, Avaneesh Yembadi, Zain Patel, Simon Brugman, Kiyo Kunii, Benjamin Levy, Louis de Charsonville, Simon Picard
kedro pipeline list and kedro pipeline describe are being deprecated in favour of new commands kedro registry list and kedro registry describe.
registry, with the associated commands kedro registry list and kedro registry describe, to replace kedro pipeline list and kedro pipeline describe.requirements.txt is packaged, its dependencies are embedded in the modular pipeline wheel file. Upon pulling the pipeline, Kedro will append dependencies to the project's requirements.in. More information is available in our documentation.kedro pipeline package/pull --all and pyproject.toml.cli.py from the Kedro project template. By default all CLI commands, including kedro run, are now defined on the Kedro framework side. These can be overridden in turn by a plugin or a cli.py file in your project. A packaged Kedro project will respect the same hierarchy when executed with python -m my_package..ipython/profile_default/startup/ from the Kedro project template in favour of .ipython/profile_default/ipython_config.py and the kedro.extras.extensions.ipython.dill backend to PickleDataSet.kedro pipeline package and kedro pipeline pull time, so that aliasing a modular pipeline doesn't break it.| Type | Description | Location |
|---|---|---|
tracking.MetricsDataSet |
Dataset to track numeric metrics for experiment tracking | kedro.extras.datasets.tracking |
tracking.JSONDataSet |
Dataset to track data for experiment tracking | kedro.extras.datasets.tracking |
fsspec version to 2021.04.kedro install and kedro build-reqs flows when uninstalled dependencies are present in a project's settings.py, context.py or hooks.py (Issue #829).kedro pipeline package and kedro pipeline pull time, so that aliasing a modular pipeline doesn't break it.dynaconf to <3.1.6 because the method signature for _validate_items changed which is used in Kedro.kedro pipeline list and kedro pipeline describe are being deprecated in favour of new commands kedro registry list and kedro registry describe.kedro install is being deprecated in favour of using pip install -r src/requirements.txt to install project dependencies.Added the following new datasets:
| Type | Description | Location |
|---|---|---|
plotly.PlotlyDataSet |
Works with plotly graph object Figures (saves as json file) | kedro.extras.datasets.plotly |
ConfigLoader.get() now raises a BadConfigException, with a more helpful error message, if a configuration file cannot be loaded (for instance due to wrong syntax or poor formatting).run_id now defaults to save_version when after_catalog_created is called, similarly to what happens during a kedro run.kedro ipython and kedro jupyter notebook didn't work if the PYTHONPATH was already set.env and extra_params to reload_kedro similar to how the IPython script works.kedro info now outputs if a plugin has any hooks or cli_hooks implemented.PartitionedDataSet now supports lazily materializing data on save.kedro pipeline describe now defaults to the __default__ pipeline when no pipeline name is provided and also shows the namespace the nodes belong to.EmailMessageDataSet added to doctree.kedro pipeline package now only packages the parameter file that exactly matches the pipeline name specified and the parameter files in a directory with the pipeline name.model input tables in accordance with our Data Engineering convention.kedro pipeline package takes the pipeline package version, rather than the kedro package version. If the pipeline package version is not present, then the package version is used.Kedro plugins can now override built-in CLI commands.
before_command_run hook for plugins to add extra behaviour before Kedro CLI commands run.pipelines from pipeline_registry.py and register_pipeline hooks are now loaded lazily when they are first accessed, not on startup:from kedro.framework.project import pipelines
print(pipelines["__default__"]) # pipeline loading is only triggered here
TemplatedConfigLoader now correctly inserts default values when no globals are supplied.KEDRO_ENV environment variable had no effect on instantiating the context variable in an iPython session or a Jupyter notebook.bootstrap_project method.configure_project is invoked if a package_name is supplied to KedroSession.create. This is added for backward-compatibility purpose to support a workflow that creates Session manually. It will be removed in 0.18.0.ModuleNotFoundError if register_pipelines not found, so that a more helpful error message will appear when a dependency is missing, e.g. Issue #722.kedro new is invoked using a configuration yaml file, output_dir is no longer a required key; by default the current working directory will be used.kedro new is invoked using a configuration yaml file, the appropriate prompts.yml file is now used for validating the provided configuration. Previously, validation was always performed against the kedro project template prompts.yml file.kedro new now generates user prompts to obtain configuration rather than supplying empty configuration.after_dataset_loaded run would finish before a dataset is actually loaded when using --async flag.kedro.versioning.journal.Journal will be removed.kedro.framework.context.KedroContext will be removed:
io in favour of KedroContext.catalogpipeline (equivalent to pipelines["__default__"])pipelines in favour of kedro.framework.project.pipelinesAdded support for compress_pickle backend to PickleDataSet.
compress_pickle backend to PickleDataSet.KedroContext instance:from kedro.framework.project import pipelines
print(pipelines)
pipeline_registry.py rather than hooks.py.kedro runsettings.py is not importable, the errors will be surfaced earlier in the process, rather than at runtime.kedro pipeline list and kedro pipeline describe no longer accept redundant --env parameter.from kedro.framework.cli.cli import cli no longer includes the new and starter commands.kedro.framework.context.KedroContext.run will be removed in release 0.18.0.Added env and extra_params to reload_kedro() line magic.
env and extra_params to reload_kedro() line magic.pipeline() API to allow strings and sets of strings as inputs and outputs, to specify when a dataset name remains the same (not namespaced).default_config.yml as prompts.yml.env and extra_params arguments to register_config_loader hook.settings are loaded. You will now be able to run:from kedro.framework.project import settings
print(settings.CONF_ROOT)
SparkDataSet in the interactive workflow.pyproject.toml for a tool.kedro section before treating the project as a Kedro project.DataCatalog::shallow_copy now it should copy layers.kedro pipeline pull now uses pip download for protocols that are not supported by fsspec.jsonschema schema definition for the Kedro 0.17 catalog.kedro install now waits on Windows until all the requirements are installed.--to-outputs option in the CLI, throughout the codebase, and as part of hooks specifications.ParquetDataSet wasn't creating parent directories on the fly.kedro ipython and kedro jupyter workflows. To fix this, follow the instructions in the migration guide below.Note: If you're using the
ipythonextension instead, you will not encounter this problem.
You will have to update the file <your_project>/.ipython/profile_default/startup/00-kedro-init.py in order to make kedro ipython and/or kedro jupyter work. Add the following line before the KedroSession is created:
configure_project(metadata.package_name) # to add
session = KedroSession.create(metadata.package_name, path)
Make sure that the associated import is provided in the same place as others in the file:
from kedro.framework.project import configure_project # to add
from kedro.framework.session import KedroSession
Mariana Silva, Kiyohito Kunii, noklam, Ivan Doroshenko, Zain Patel, Deepyaman Datta, Sam Hiscox, Pascal Brokmeier
The Kedro 0.17.0 release contains some breaking changes. If you update Kedro to 0.17.0 and then try to work with projects created against earlier vers…
KedroSession which is responsible for managing the lifecycle of a Kedro run.kedro new --starter=mini-kedro. It is possible to use the DataCatalog as a standalone component in a Jupyter notebook and transition into the rest of the Kedro framework.DatasetSpecs with Hooks to run before and after datasets are loaded from/saved to the catalog.kedro catalog create. For a registered pipeline, it creates a <conf_root>/<env>/catalog/<pipeline_name>.yml configuration file with MemoryDataSet datasets for each dataset that is missing from DataCatalog.settings.py and pyproject.toml (to replace .kedro.yml) for project configuration, in line with Python best practice.ProjectContext is no longer needed, unless for very complex customisations. KedroContext, ProjectHooks and settings.py together implement sensible default behaviour. As a result context_path is also now an optional key in pyproject.toml.ProjectContext from src/<package_name>/run.py.TemplatedConfigLoader now supports Jinja2 template syntax alongside its original syntax.ConfigLoader or the DataCatalog used in a project. If no such Hook is provided in src/<package_name>/hooks.py, a KedroContextError is raised. There are sensible defaults defined in any project generated with Kedro >= 0.16.5.ParallelRunner no longer results in a run failure, when triggered from a notebook, if the run is started using KedroSession (session.run()).before_node_run can now overwrite node inputs by returning a dictionary with the corresponding updates.isort and pytest configuration from <project_root>/setup.cfg to <project_root>/pyproject.toml.KedroSession to KedroContext.pyspark requirements to allow for installation of pyspark 3.0.--fs-args option to the kedro pipeline pull command to specify configuration options for the fsspec filesystem arguments used when pulling modular pipelines from non-PyPI locations.fsspec version to 0.9.s3fs version to 0.5 (S3FileSystem interface has changed since 0.4.1 version).kedro.cli and kedro.context modules in favour of kedro.framework.cli and kedro.framework.context respectively.kedro.io.DataCatalog.exists() returns False when the dataset does not exist, as opposed to raising an exception.catalog.yml file is no longer automatically created for modular pipelines when running kedro pipeline create. Use kedro catalog create to replace this functionality.include_examples prompt from kedro new. To generate boilerplate example code, you should use a Kedro starter.--verbose flag from a global command to a project-specific command flag (e.g kedro --verbose new becomes kedro new --verbose).dataset_credentials key in credentials in PartitionedDataSet.get_source_dir() was removed from kedro/framework/cli/utils.py.get_config, create_catalog, create_pipeline, template_version, project_name and project_path keys by get_project_context() function (kedro/framework/cli/cli.py).kedro new --starter now defaults to fetching the starter template matching the installed Kedro version.kedro_cli.py to cli.py and moved it inside the Python package (src/<package_name>/), for a better packaging and deployment experience..kedro.yml from the project template and replaced it with pyproject.toml.KEDRO_CONFIGS constant (previously residing in kedro.framework.context.context).kedro pipeline create CLI command to add a boilerplate parameter config file in conf/<env>/parameters/<pipeline_name>.yml instead of conf/<env>/pipelines/<pipeline_name>/parameters.yml. CLI commands kedro pipeline delete / package / pull were updated accordingly.get_static_project_data from kedro.framework.context.KedroContext.static_data.KedroContext constructor now takes package_name as first argument.context property on KedroSession with load_context() method._push_session and _pop_session in kedro.framework.session.session to _activate_session and _deactivate_session respectively.CONTEXT_CLASS variable in src/<your_project>/settings.py.KedroContext.hooks attribute. Instead, hooks should be registered in src/<your_project>/settings.py under the HOOKS key.[\w\.-]+$.KedroContext._create_config_loader() and KedroContext._create_data_catalog(). They have been replaced by registration hooks, namely register_config_loader() and register_catalog() (see also upcoming deprecations).kedro.framework.context.load_context will be removed in release 0.18.0.kedro.framework.cli.get_project_context will be removed in release 0.18.0.DeprecationWarning to the decorator API for both node and pipeline. These will be removed in release 0.18.0. Use Hooks to extend a node's behaviour instead.DeprecationWarning to the Transformers API when adding a transformer to the catalog. These will be removed in release 0.18.0. Use Hooks to customise the load and save methods.Deepyaman Datta, Zach Schuster
Reminder: Our documentation on how to upgrade Kedro covers a few key things to remember when updating any Kedro version.
The Kedro 0.17.0 release contains some breaking changes. If you update Kedro to 0.17.0 and then try to work with projects created against earlier versions of Kedro, you may encounter some issues when trying to run kedro commands in the terminal for that project. Here's a short guide to getting your projects running against the new version of Kedro.
Note: As always, if you hit any problems, please check out our documentation:
To get an existing Kedro project to work after you upgrade to Kedro 0.17.0, we recommend that you create a new project against Kedro 0.17.0 and move the code from your existing project into it. Let's go through the changes, but first, note that if you create a new Kedro project with Kedro 0.17.0 you will not be asked whether you want to include the boilerplate code for the Iris dataset example. We've removed this option (you should now use a Kedro starter if you want to create a project that is pre-populated with code).
To create a new, blank Kedro 0.17.0 project to drop your existing code into, you can create one, as always, with kedro new. We also recommend creating a new virtual environment for your new project, or you might run into conflicts with existing dependencies.
pyproject.toml: Copy the following three keys from the .kedro.yml of your existing Kedro project into the pyproject.toml file of your new Kedro 0.17.0 project:[tools.kedro]
package_name = "<package_name>"
project_name = "<project_name>"
project_version = "0.17.0"
Check your source directory. If you defined a different source directory (source_dir), make sure you also move that to pyproject.toml.
Copy files from your existing project:
project/src/project_name/pipelines from existing to new projectproject/src/test/pipelines from existing to new projectrequirements.txt and/or requirements.in.conf folder. Take note of the new locations needed for modular pipeline configuration (move it from conf/<env>/pipeline_name/catalog.yml to conf/<env>/catalog/pipeline_name.yml and likewise for parameters.yml).data/ folder of your existing project, if needed, into the same location in your new project.src/<package_name>/hooks.py.Update your new project's README and docs as necessary.
Update settings.py: For example, if you specified additional Hook implementations in hooks, or listed plugins under disable_hooks_by_plugin in your .kedro.yml, you will need to move them to settings.py accordingly:
from <package_name>.hooks import MyCustomHooks, ProjectHooks
HOOKS = (ProjectHooks(), MyCustomHooks())
DISABLE_HOOKS_FOR_PLUGINS = ("my_plugin1",)
Migration for node names. From 0.17.0 the only allowed characters for node names are letters, digits, hyphens, underscores and/or fullstops. If you have previously defined node names that have special characters, spaces or other characters that are no longer permitted, you will need to rename those nodes.
Copy changes to kedro_cli.py. If you previously customised the kedro run command or added more CLI commands to your kedro_cli.py, you should move them into <project_root>/src/<package_name>/cli.py. Note, however, that the new way to run a Kedro pipeline is via a KedroSession, rather than using the KedroContext:
with KedroSession.create(package_name=...) as session:
session.run()
Copy changes made to ConfigLoader. If you have defined a custom class, such as TemplatedConfigLoader, by overriding ProjectContext._create_config_loader, you should move the contents of the function in src/<package_name>/hooks.py, under register_config_loader.
Copy changes made to DataCatalog. Likewise, if you have DataCatalog defined with ProjectContext._create_catalog, you should copy-paste the contents into register_catalog.
Optional: If you have plugins such as Kedro-Viz installed, it's likely that Kedro 0.17.0 won't work with their older versions, so please either upgrade to the plugin's newest version or follow their migration guides.
Added documentation with a focus on single machine and distributed environment deployment; the series includes Docker, Argo, Prefect, Kubeflow, AWS Ba
kedro new --starter spaceflights.TypeError when converting dict inputs to a node made from a wrapped partial function.PartitionedDataSet improvements:
jalapeño will be accessible as DataCatalog.datasets.jalapeño rather than DataCatalog.datasets.jalape__o.kedro install for an Anaconda environment defined in environment.yml..kedro.yml to use kedro lint and kedro jupyter notebook convert.TensorFlowModelDataset in the HDF5 format with versioning enabled.run_result argument in after_pipeline_run Hooks spec.00-kedro-init.py file.Deepyaman Datta, Bhavya Merchant, Lovkush Agarwal, Varun Krishna S, Sebastian Bertoli, noklam, Daniel Petti, Waylon Walker
Added the following new datasets.
| Type | Description | Location |
|---|---|---|
email.EmailMessageDataSet |
Manage email messages using the Python standard library | kedro.extras.datasets.email |
pyproject.toml to configure Kedro. pyproject.toml is used if .kedro.yml doesn't exist (Kedro configuration should be under [tool.kedro] section).pipeline.py, having been replaced by hooks.py.register_pipelines(), to replace _get_pipelines()register_config_loader(), to replace _create_config_loader()register_catalog(), to replace _create_catalog()
These can be defined in src/<package-name>/hooks.py and added to .kedro.yml (or pyproject.toml). The order of execution is: plugin hooks, .kedro.yml hooks, hooks in ProjectContext.hooks..kedro.yml (or pyproject.toml) configuration file..isort.cfg settings into setup.cfg.project_name, project_version and package_name now have to be defined in .kedro.yml for projects generated using Kedro 0.16.5+.Enabled auto-discovery of hooks implementations coming from installed plugins.
ParallelRunner on Windows.GBQTableDataSet to load customised results using customised queries from Google Big Query tables.Ajay Bisht, Vijay Sajjanar, Deepyaman Datta, Sebastian Bertoli, Shahil Mawjee, Louis Guitton, Emanuel Ferm
Nothing published for this version
Added the following new datasets.
| Type | Description | Location |
|---|---|---|
pandas.AppendableExcelDataSet |
Works with Excel file opened in append mode |
kedro.extras.datasets.pandas |
tensorflow.TensorFlowModelDataset |
Works with TensorFlow models using TensorFlow 2.X |
kedro.extras.datasets.tensorflow |
holoviews.HoloviewsWriter |
Works with Holoviews objects (saves as image file) |
kedro.extras.datasets.holoviews |
kedro install will now compile project dependencies (by running kedro build-reqs behind the scenes) before the installation if the src/requirements.in file doesn't exist.only_nodes_with_namespace in Pipeline class to filter only nodes with a specified namespace.kedro pipeline delete command to help delete unwanted or unused pipelines (it won't remove references to the pipeline in your create_pipelines() code).kedro pipeline package command to help package up a modular pipeline. It will bundle up the pipeline source code, tests, and parameters configuration into a .whl file.DataCatalog:
DataCatalog.list() method.__ in DataCatalog.datasets, for ease of access to transcoded datasets.spark.SparkHiveDataSet.spark.SparkDataSet.pyarrow table in pandas.ParquetDataSet.kedro build-reqs CLI command:
kedro build-reqs is now called with -q option and will no longer print out compiled requirements to the console for security reasons.kedro build-reqs command are now passed to pip-compile call (e.g. kedro build-reqs --generate-hashes).kedro jupyter CLI command:
kedro jupyter notebook, kedro jupyter lab or kedro ipython with Jupyter/IPython dependencies not being installed.%run_viz line magic for showing kedro viz inside a Jupyter notebook. For the fix to be applied on existing Kedro project, please see the migration guide.pillow.ImageDataSet entry to the documentation.%run_viz line magic in existing projectEven though this release ships a fix for project generated with kedro==0.16.2, after upgrading, you will still need to make a change in your existing project if it was generated with kedro>=0.16.0,<=0.16.1 for the fix to take effect. Specifically, please change the content of your project's IPython init script located at .ipython/profile_default/startup/00-kedro-init.py with the content of this file. You will also need kedro-viz>=3.3.1.
Miguel Rodriguez Gutierrez, Joel Schwarzmann, w0rdsm1th, Deepyaman Datta, Tam-Sanh Nguyen, Marcus Gawronsky
Fixed deprecation warnings from kedro.cli and kedro.context when running kedro jupyter notebook.
kedro.cli and kedro.context when running kedro jupyter notebook.catalog and context were not available in Jupyter Lab and Notebook.kedro build-reqs would fail if you didn't have your project dependencies installed.In addition, all the specific datasets like CSVLocalDataSet, CSVS3DataSet etc. were deprecated. Instead, you must use generalized datasets like CSVDat…
kedro catalog list to list datasets in your catalogkedro pipeline list to list pipelineskedro pipeline describe to describe a specific pipelinekedro pipeline create to create a modular pipelinegit-style.kedro.cli and kedro.context have been moved into kedro.framework.cli and kedro.framework.context respectively. kedro.cli and kedro.context will be removed in future releases.Hooks, which is a new mechanism for extending Kedro.load_context changing user's current working directory..kedro.yml.node(func, "params:a.b", None)| Type | Description | Location |
|---|---|---|
pillow.ImageDataSet |
Work with image files using Pillow |
kedro.extras.datasets.pillow |
geopandas.GeoJSONDataSet |
Work with geospatial data using GeoPandas |
kedro.extras.datasets.geopandas.GeoJSONDataSet |
api.APIDataSet |
Work with data from HTTP(S) API requests | kedro.extras.datasets.api.APIDataSet |
joblib backend support to pickle.PickleDataSet.MatplotlibWriter dataset.pip install "kedro[pandas.ParquetDataSet]".encoding or compression, for fsspec.spec.AbstractFileSystem.open() calls when loading/saving a dataset. See Example 3 under docs.namespace property on Node, related to the modular pipeline where the node belongs.SequentialRunner(is_async=True) and ParallelRunner(is_async=True) class.MemoryProfiler transformer.pandas>=1.0.pyspark is not fully-compatible with 3.8 yet.CONTRIBUTING.md - added Developer Workflow._exists method to MyOwnDataSet example in 04_user_guide/08_advanced_io.PartitionedDataSet and IncrementalDataSet were not working with s3a or s3n protocol.pandas.ParquetDataSet.functools.lru_cache with cachetools.cachedmethod in PartitionedDataSet and IncrementalDataSet for per-instance cache invalidation.SparkDataSet when running on Databricks.SparkDataSet not allowing for loading data from DBFS in a Windows machine using Databricks-connect.DataSetNotFoundError to suggest possible dataset names user meant to type.make test-no-spark.kedro lint --check-only).kedro.io.kedro.contrib and extras folders.CSVBlobDataSet and JSONBlobDataSet dataset types.invalidate_cache method on datasets private.get_last_load_version and get_last_save_version methods are no longer available on AbstractDataSet.get_last_load_version and get_last_save_version have been renamed to resolve_load_version and resolve_save_version on AbstractVersionedDataSet, the results of which are cached.release() method on datasets extending AbstractVersionedDataSet clears the cached load and save version. All custom datasets must call super()._release() inside _release().TextDataSet no longer has load_args and save_args. These can instead be specified under open_args_load or open_args_save in fs_args.PartitionedDataSet and IncrementalDataSet method invalidate_cache was made private: _invalidate_caches.KEDRO_ENV_VAR from kedro.context to speed up the CLI run time.Pipeline.name has been removed in favour of Pipeline.tag().Pipeline.transform() in favour of kedro.pipeline.modular_pipeline.pipeline() helper function.PARAMETER_KEYWORDS private, and moved it from kedro.pipeline.pipeline to kedro.pipeline.modular_pipeline.DataCatalog.Since all the datasets (from kedro.io and kedro.contrib.io) were moved to kedro/extras/datasets you must update the type of all datasets in <project>/conf/base/catalog.yml file.
Here how it should be changed: type: <SomeDataSet> -> type: <subfolder of kedro/extras/datasets>.<SomeDataSet> (e.g. type: CSVDataSet -> type: pandas.CSVDataSet).
In addition, all the specific datasets like CSVLocalDataSet, CSVS3DataSet etc. were deprecated. Instead, you must use generalized datasets like CSVDataSet.
E.g. type: CSVS3DataSet -> type: pandas.CSVDataSet.
Note: No changes required if you are using your custom dataset.
Pipeline.transform() has been dropped in favour of the pipeline() constructor. The following changes apply:
from kedro.pipeline import pipelineprefix argument has been renamed to namespacedatasets has been broken down into more granular arguments:
inputs: Independent inputs to the pipelineoutputs: Any output created in the pipeline, whether an intermediary dataset or a leaf outputparameters: params:... or parametersAs an example, code that used to look like this with the Pipeline.transform() constructor:
result = my_pipeline.transform(
datasets={"input": "new_input", "output": "new_output", "params:x": "params:y"},
prefix="pre"
)
When used with the new pipeline() constructor, becomes:
from kedro.pipeline import pipeline
result = pipeline(
my_pipeline,
inputs={"input": "new_input"},
outputs={"output": "new_output"},
parameters={"params:x": "params:y"},
namespace="pre"
)
Since some modules were moved to other locations you need to update import paths appropriately.
You can find the list of moved files in the 0.15.6 release notes under the section titled Files with a new location.
Note: If you haven't made significant changes to your
kedro_cli.py, it may be easier to simply copy the updatedkedro_cli.py.ipython/profile_default/startup/00-kedro-init.pyand from GitHub or a newly generated project into your old project.
KEDRO_ENV_VAR from kedro.context. To get your existing project template working, you'll need to remove all instances of KEDRO_ENV_VAR from your project template:
kedro_cli.py and .ipython/profile_default/startup/00-kedro-init.py: from kedro.context import KEDRO_ENV_VAR, load_context -> from kedro.framework.context import load_contextenvvar=KEDRO_ENV_VAR line from the click options in run, jupyter_notebook and jupyter_lab in kedro_cli.pyKEDRO_ENV_VAR with "KEDRO_ENV" in _build_jupyter_envcontext = load_context(path, env=os.getenv(KEDRO_ENV_VAR)) with context = load_context(path) in .ipython/profile_default/startup/00-kedro-init.pykedro build-reqsWe have upgraded pip-tools which is used by kedro build-reqs to 5.x. This pip-tools version requires pip>=20.0. To upgrade pip, please refer to their documentation.
@foolsgold, Mani Sarkar, Priyanka Shanbhag, Luis Blanche, Deepyaman Datta, Antony Milne, Panos Psimatikas, Tam-Sanh Nguyen, Tomasz Kaczmarczyk, Kody Fischer, Waylon Walker
Pinned fsspec>=0.5.1, <0.7.0 and s3fs>=0.3.0, <0.4.1 to fix incompatibility issues with their latest release.
Your coding agent can read these notes before it upgrades. Set up the MCP server →