NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI
OpenLineage integration with Airflow
Last release 9 months ago
11 Dec 2025
Ships fairly regularly
a new release about every 3 weeks
Most releases are documented
notes for 44 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
109 releases · first in 2021
One column per quarter.
Spec: Add arbitrary extra info to JobDependency in JobDependenciesRunFacet `#4189` @kacpermuda *Add support for arbitrary extra information in JobDepe
#4189 @kacpermuda
Add support for arbitrary extra information in JobDependency within JobDependenciesRunFacet.#4185 @kacpermuda
Add debug mode support to file transport for better troubleshooting.#4160 @harels
Add support for capturing dbt model owner information from meta.owner in OpenLineage events.#4151 @mobuchowski
Add DbtNodeJobFacet to provide additional dbt node information in job facets.#4161 @tnazarew
Add default name support to Hive catalog facet in Spark integration.#4134 @pawel-big-lebowski
Add support for fetching input statistics for single input RDD jobs.#4153 @kchledowski
Migrate SQL parser from fork to upstream version 0.59 for better maintenance and compatibility.#4178 @mobuchowski
Reduce aggressiveness of UUID normalization in Spark integration.#4186 @kacpermuda
Improve logging output in Python client.#4165 @dolfinus
Fix relation size calculation to ensure values are within reasonable bounds.#4154 @mobuchowski
Add missing job facet schema to specification.Python: re-add missing __version__ variables in top of releaseable modules `#4135` @mobuchowski *Fixes breaking change in version 1.40.0.*
#4135 @mobuchowski
Fixes breaking change in version 1.40.0.#4135 @mobuchowski
Fixes breaking change in version 1.40.0.Java: Fix CVE in commons-lang3 `#4084` @mandalbalmukund *Upgrade commons-lang3 version to fix CVE security vulnerability.*
#4109 @jakub-moravec
Add a standardized batch API endpoint to OpenLineage specification for handling multiple events in a single request.#4116 @mobuchowski
Add ordinal_position field to track the position of fields in schema (1-indexed).#4112 @kacpermuda
Introduce JobDependenciesRunFacet to track dependencies between jobs.#4103 @jakub-moravec
Add support for temporary datasets to enable job-to-job lineage tracking.#4075 @luke-hoffman1
Add fallback configuration for BigQuery project ID in Metastore integration.#4123 @kacpermuda
Include examples in Python generated classes for better documentation.#4077 @dolfinus
Add support for parsing jTDS JDBC URL format in Java client.#4066 @tnazarew
Add ParentRunFacet to Hive integration for tracking parent-child run relationships.#4097 @tnazarew
Add support for tracking LOAD and IMPORT operations in Hive.#4085 @tnazarew
Add support for tracking EXPORT operations in Hive.#4079 @tnazarew
Add START event emission support to Hive integration.#4121 @usamakunwar
Fix Spark dataset facet builders for input datasets.#4114 @kchledowski
Fix job name trimming logic in Spark integration.#4113 @pawel-big-lebowski
Fix putAll operation failing on immutable maps.#4108 @pawel-big-lebowski
Fix multiple issues with RDD job handling in Spark.#4102 @kchledowski
Fix JDBC dbtable parsing to support any FROM clauses.#4083 @pawel-big-lebowski
Fix Spark connector configuration for Databricks environments.#4099 @mobuchowski
Catch NoClassDefFoundError when buggy implementations exist on classpath.#4104 @mobuchowski
Fix Snowflake identifier parsing to handle quoted identifiers correctly.#4105 @mobuchowski
Strip quotes from Snowflake account names for proper handling.#4092 @fm100
Fix facet property names from snake_case to camelCase for consistency.#4111 @kacpermuda
Fix Python client facet generator after moving to UV build system.#4093 @antonlin1
Fix retry configuration default merge with user-defined config in HTTP transports.#4084 @mandalbalmukund
Upgrade commons-lang3 version to fix CVE security vulnerability.#4126 @dolfinus
Ensure START and STOP events share the same runId in Hive integration.Spark: Normalize dataset names with configurable trimmers `#3996` @pawel-big-lebowski *Add configurable dataset name normalization with support for da
#3996 @pawel-big-lebowski
Add configurable dataset name normalization with support for date patterns, key-value pairs, and S3 location detection to enable proper dataset subsetting.#4057 @kchledowski
Add missing input symlink facets for Databricks Unity Catalog tables.#4058 @kchledowski
Refactor column-level lineage dependency collector tests for better organization and maintainability.#4069 @fm100
Fix typo in IcebergCommitReportOutputDatasetFacet property name.#4061 @pawel-big-lebowski
Fix dataset name trimming for column-level lineage inputs.#4062 @kacpermuda
Remove unnecessary numpy import from Python client.#3844 @kacpermuda
Remove Dagster integration from the repository.Spec: Add subset dataset facets to spec `#4008` @pawel-big-lebowski *Add subset dataset facets to OpenLineage specification for representing dataset r
#4008 @pawel-big-lebowski
Add subset dataset facets to OpenLineage specification for representing dataset relationships.#3978 @heron--
Allow attaching dataset quality information outside of InputDatasetFacet.#4018 @tnazarew
Add support for Spark structured streaming microbatch source write operations.#4016 @ddebowczyk92
Add catalog properties support to Spark integration for better catalog metadata tracking.#4039 @ddebowczyk92
Enhance BigQuery integration with GCP project ID and location in catalog properties.#3972 @kchledowski
Add support for tracking COALESCE transformations in Spark jobs.#3982 @ddebowczyk92
Add catalog facet support for vanilla Hive table operations.#4013 @pawel-big-lebowski
Output statistics now available in complete events for better observability.#3977 @pawel-big-lebowski
Add output statistics tracking for Spark RDD-based jobs.#4050 @pawel-big-lebowski
Improve generated model classes with proper equals and hashcode implementations.#4022 @mobuchowski
Add support for capturing dbt tags in OpenLineage events.#4017 @mobuchowski
Add dbt Cloud account ID tracking to dbt run facets.#3987 @mobuchowski
Enhance DbtRunRunFacet with additional metadata for better observability.#4006 @ddebowczyk92
Add native Google Cloud Platform Lineage transport for Python client.#3983 @JDarDagran
Add fsspec filesystem support to FileTransport for broader filesystem compatibility.#3980 @kacpermuda
Automatically add OpenLineage client version as default tag in events.#3986 @gabrysiaolsz
Add GCP Cloud Composer environment metadata facets to Airflow integration.#4055 @mobuchowski
Use dbt model aliases when generating dataset names for more accurate lineage.#4029 @EugeneYushin
Serialize OpenLineage events to JSON format for improved debug logging.#4030 @EugeneYushin
Properly respect user-overridden application names in event emission.#4003 @kchledowski
Refactor column-level lineage expression dependency collector for better maintainability.#3994 @JDarDagran
Enhance logging for Iceberg input statistics collection.#3985 @pawel-big-lebowski
Optimize S3 operations by limiting external getFileStatus calls for large object sets.#3964 @kchledowski
Refactor TransformationInfo into shared Java client for cross-integration reuse.#4026 @dolfinus
Enhance logging capabilities in asynchronous HTTP transport.#4000 @JDarDagran
Support Python type aliases in client code generation.#3997 @JDarDagran
Improve code generation to properly handle nearly identical class definitions.#4014 @dolfinus
Fail fast with clear errors when custom token providers fail to load.#4015 @dolfinus
Improve error visibility by not silencing import errors in transport factory.#3968 @kacpermuda
Update import paths to use versioned facet and event modules.#4012 @JDarDagran
Improve thread pool management in Java client utilities.#3965 @JDarDagran
Migrate from pre-commit to prek for pre-commit hook management.#4053 @jsjasonseba
Fix incorrect Glue catalog detection due to always attempting ARN resolution.#4052 @kchledowski
Fix column-level lineage failures on Spark runtimes without spark-hive package.#4031 @kchledowski
Fix missing input datasets and column-level lineage for CreateDataSourceTableAsSelect and CreateHiveTableAsSelect commands.#4044 @EugeneYushin
Fix BigQuery intermediate job filtering by using bucket configuration.#4002 @MaciejGajewski
Add additional exception handling for TypeNotPresentException in Spark 3.0.2.#4034 @JDarDagran
Correct license field specification in Python package metadata.#4045 @kacpermuda
Support both naming conventions for API key configuration parameter.#4037 @EugeneYushin
Fix build issue causing empty sources JAR files to be generated.Python: Add Datadog transport with configurable async routing `#3950` @mobuchowski *Add Datadog transport with intelligent routing between sync/async
#3950 @mobuchowski#3860 @orthoxerox#3933 @mobuchowski#3923 @kyungryun#3904 @pawel-big-lebowski#3956 @dolfinus#3925 @SalvadorRomo#3946 @pawel-big-lebowski#3949 @ddebowczyk92#3934 @pawel-big-lebowski#3930 @yunchipang#3947 @ddebowczyk92#3953 @jroachgolf84#3943 @JDarDagran#3860 @orthoxerox#3933 @mobuchowski#3923 @kyungryun#3950 @mobuchowski#3904 @pawel-big-lebowski#3956 @dolfinus#3925 @SalvadorRomo#3946 @pawel-big-lebowski#3949 @ddebowczyk92#3934 @pawel-big-lebowski#3930 @yunchipang#3947 @ddebowczyk92#3953 @jroachgolf84#3943 @JDarDagranSpark: support Delta 4.0 and cover it with tests on Spark 4.0. `#3877` @pawel-big-lebowski *Fix failing tests for Spark 4.0. Make delta integration te
#3877 @pawel-big-lebowski#3914 @pawel-big-lebowski#3921 @pawel-big-lebowski#3890 @jroachgolf84#3918 @mobuchowski#3816 @ddebowczyk92#3907 @pawel-big-lebowski#3851 @dolfinus#3895 @dolfinus#3899 @Shadi#3869 @mobuchowski#3902 @pawel-big-lebowskiSqlExecutionRDDVisitor and LogicalRDDVisitor classes to avoid memory leak.#3909 @pawel-big-lebowski#3908 @pawel-big-lebowski#3911 @pawel-big-lebowski#3915 @fetta#3905 @kacpermuda#3916 @mobuchowski#3894 @mobuchowski#3889 @pawel-big-lebowski#3897 @dolfinus#3901 @kacpermudadbt: Fix deprecated configs `#3859` @kacpermuda *Replaces deprecated dbt configurations with current alternatives*
#3848 @dolfinus#3850 @ddebowczyk92#3880 @pawel-big-lebowskispark.openlineage.disabled entry to disable OpenLineage integration through Spark config parameters#3779 @pawel-big-lebowskibuildDatasetsTimePercentage and facetsBuildingTimePercentage in docs for more details#3812 @mobuchowski#3764 @dolfinus#3829 @kacpermuda#3789 @dolfinus#3863 @dolfinus#3819 @mobuchowski#3826 @ddebowczyk92#3775 @ddebowczyk92#3858 @dolfinus#3856 @ddebowczyk92#3811 @pawel-big-lebowski#3881 @dolfinus#3843 @dolfinus#3857 @dolfinus#3855 @dolfinus#3838 @dolfinus#3841 @dolfinus#3839 @dolfinus#3817 @mobuchowski#3854 @dolfinus#3799 @pan-siekierski#3796 @dolfinus#3836 @mobuchowski#3800 @dolfinus#3849 @dolfinus#3793 @mobuchowski#3859 @kacpermuda.db suffix in database/namespace location name for BigQueryMetastoreCatalog #3874 @ddebowczyk92#3835 @ddebowczyk92#3871 @pawel-big-lebowski#3861 @pawel-big-lebowski#3832 @pawel-big-lebowski#3853 @kacpermuda#3825 @mobuchowski#3806 @mobuchowski#3830 @mobuchowski#3887 @kacpermuda#3814 @mobuchowskiHive: Integration added. `#3555` @tnazarew with @ddebowczyk92, @jphalip *Added OpenLineage Hive integration*
#3555 @tnazarew with @ddebowczyk92, @jphalip#3691 @pawel-big-lebowskiUnionRdd and NewHadoopRDD, which makes dynamic frames docker based test passing.#3781 @dolfinus#3777 @dolfinus#3786 @dolfinus#3717 @tnazarew#3715 @pawel-big-lebowski#3760 @ddebowczyk92#3738 @dolfinus#3739 @dolfinus#3725 @dolfinus#3744 @dolfinus#3726 @dolfinus#3763 @pawel-big-lebowski#3748 @dolfinus#3669 @kacpermuda#3731 @dolfinus#3713 @dolfinus#3754 @dolfinus#3709 @dolfinus#3766 @mvitale#3751 @ddebowczyk92#3785 @ddebowczyk92#3776 @kacpermuda#3680 @mobuchowski#3773 @dolfinus#3722 @pawel-big-lebowski#3749 @mobuchowski#3724 @dolfinus#3728 @JDarDagran#3762 @ngorchakovaPython: add TransformTransport for Python client `#3697` @kacpermuda *Introduces the TransformTransport class for event transformations.*
#3697 @kacpermuda#3659 @mobuchowski#3695 @mobuchowski#3685 @shinabel#3706 @kacpermuda#3700 @dolfinus#3686 @luke-hoffman1#3696 @dolfinus#3707 @dolfinus#3688 @dolfinus#3681 @pawel-big-lebowskiAvro: support schema facet for Avro datasets `#3650` @pawel-big-lebowski *This PR adds support for schema facets in Avro datasets.*
#3650 @pawel-big-lebowski#3672 @martinovm#3682 @MassyB#3683 @MassyB#3676 @pawel-big-lebowski#3667 @luke-hoffman1#3673 @ddebowczyk92#3663 @ddebowczyk92Flink: enhance JDBC extractors with additional types `#3652` @HuangZhenQiu
#3652 @HuangZhenQiu#3648 @mobuchowski#3637 @dolfinus#3664 @mudakacper#3661 @mobuchowski#3644 @tnazarew#3645 @pawel-big-lebowski#3641 @ddebowczyk92#3652 @HuangZhenQiu#3648 @mobuchowski#3637 @dolfinus#3664 @kacpermuda#3661 @mobuchowski#3644 @tnazarew#3645 @pawel-big-lebowski#3641 @ddebowczyk92Java: added support for LDAP connection strings `#3612` @luke-hoffman1
#3612 @luke-hoffman1#3625 @mehdimld#3605 @kuba0221#3572 @mobuchowski#3622 @dolfinus#3624 @dolfinus#3621 @dolfinus#3582 @ddebowczyk92#3607 @ddebowczyk92#3586 @dolfinus#3594 @pawel-big-lebowski
Register simple micrometer registry when no other registry configured.#3612 @luke-hoffman1#3625 @mehdimld#3605 @kuba0221#3572 @mobuchowski#3622 @dolfinus#3624 @dolfinus#3621 @dolfinus#3582 @ddebowczyk92#3607 @ddebowczyk92#3586 @dolfinus#3594 @pawel-big-lebowski
Register simple micrometer registry when no other registry configured.#3580 @MassyB
Buffered writing could cause dbt to write broken log lines. This PR handles that case.#3576 @pawel-big-lebowski
This PR fixes case where event v2 job and run events did not get user tags.#3583 @ddebowczyk92
This PR fixes potential NullPointerException in SaveIntoDataSourceCommandVisitor.#3612 @luke-hoffman1#3625 @mehdimld#3605 @kuba0221#3572 @mobuchowski#3622 @dolfinus#3624 @dolfinus#3621 @dolfinus#3582 @ddebowczyk92#3607 @ddebowczyk92#3586 @dolfinus#3594 @pawel-big-lebowski
Register simple micrometer registry when no other registry configured.dbt: read structured log file incrementally `#3580` @MassyB *Buffered writing could cause dbt to write broken log lines. This PR handles that case.*
#3580 @MassyB
Buffered writing could cause dbt to write broken log lines. This PR handles that case.#3576 @pawel-big-lebowski
This PR fixes case where event v2 job and run events did not get user tags.#3583 @ddebowczyk92
This PR fixes potential NullPointerException in SaveIntoDataSourceCommandVisitor.#3580 @MassyB
Buffered writing could cause dbt to write broken log lines. This PR handles that case.#3576 @pawel-big-lebowski
This PR fixes case where event v2 job and run events did not get user tags.#3583 @ddebowczyk92
This PR fixes potential NullPointerException in SaveIntoDataSourceCommandVisitor.Java: remove deprecated configs: 'disabledFacets' and 'timeout'. `#3522` @pawel-big-lebowski *Configs have been replaced with: 'facets.facet-name.disa…
#3531 @pawel-big-lebowski
Similar to Flink 1 integration, Flink 2 integration will emit CheckpointFacet.#3528 @pawel-big-lebowski
Table and field comments are available within generated OL events.#3522 @pawel-big-lebowski
Configs have been replaced with: 'facets.facet-name.disabled=true' and 'timeoutInMillis'.#3538 @sakjung
Fixes support for metrics in Iceberg SparkSessionCatalog.#3515 @pawel-big-lebowski
Fixes support for metrics in Iceberg RESTCatalog.#3535 @MassyB
Fixes race condition that was happening when using structured logs output.#3545 @MassyB
Skipped nodes no longed cause exceptions.#3550 @pawel-big-lebowski
InputStatistics facet for Iceberg datasets no longer produces incorrect stats.#3548 @pawel-big-lebowski
Subquery alias no longer is duplicating the inputs.#3552 @mobuchowski
Fixes catching InaccessibleMethodException in Java 17 within SparkExtensionVisitor.#3471 @leogodin217
User-supplied tags will allow the client to inject new tags or override tags provided by the integrations for jobs and runs.#3531 @pawel-big-lebowski
Similar to Flink 1 integration, Flink 2 integration will emit CheckpointFacet.#3528 @pawel-big-lebowski
Table and field comments are available within generated OL events.#3522 @pawel-big-lebowski
Configs have been replaced with: 'facets.<name of disabled facet>.disabled=true' and 'timeoutInMillis'.#3538 @sakjung
Fixes support for metrics in Iceberg SparkSessionCatalog.#3515 @pawel-big-lebowski
Fixes support for metrics in Iceberg RESTCatalog.#3535 @MassyB
Fixes race condition that was happening when using structured logs output.#3545 @MassyB
Skipped nodes no longed cause exceptions.#3550 @pawel-big-lebowski
InputStatistics facet for Iceberg datasets no longer produces incorrect stats.#3548 @pawel-big-lebowski
Subquery alias no longer is duplicating the inputs.#3552 @mobuchowski
Fixes catching InaccessibleMethodException in Java 17 within SparkExtensionVisitor.Python: allow adding user-supplied tags facets from config `#3471` @leogodin217 *User-supplied tags will allow the client to inject new tags or overri
#3471 @leogodin217
User-supplied tags will allow the client to inject new tags or override tags provided by the integrations for jobs and runs.#3493 @mobuchowski
Enabled parsing tags from config in Java client and Spark conf.#3487 @pawel-big-lebowski
Properly name case where TaskQueueCircuitBreaker allows a configurable blocking time after submitting a callable.#3503 @pawel-big-lebowski
Native Flink integration is now isolated within circuit breaker call.#3483 @ddebowczyk92
ServiceLoader should not fail to load OpenLineageExtensionProvider implementations in certain configurations.#3486 @MarquisC
Null Flink Job Manager address will default to localhost#3488 @MassyB
Handle case for tests on sources which don't have the attached_node defined in the manifest.Java: enable specifying custom SSL context `#3444` @pawel-big-lebowski *Enable providing configuration for SSL context within HTTP transport.*
#3444 @pawel-big-lebowski
Enable providing configuration for SSL context within HTTP transport.#3442 @pawel-big-lebowski
Spark integration filters OpenLineage events for specific plan node classes. This can be now extended with extra config entries: allowedSparkNodes and deniedSparkNodes. See Spark Configuration documentation for more details.#3437 @aritrabandyo
This circuit breaker that executes task on a queue backed threadpool, gives up tasks if the queue is full, and keeps track of rejected tasks.#3429 @whitleykeith
This allows Trino integration to emit proper events containing Trino datasets.#3430 @ssanthanam185
Adds coverage for AlterTableRecoverPartitionsCommandVisitor, RefreshTableCommandVisitor, RepairTableCommandVisitor.#3425 @d-m-h
This presents no functional change to the listener, however it will allow for improved initialisation of the listener in the future.#3435 @pawel-big-lebowski
In case of unsupported classes, warn logs without a stacktrace should be produced.#3443 @d-m-h
This is an initial refactor to a larger code base change that will see the removal of direct access of the QueryExecution object. It has no functional change on the way the integration behaves.COMPLETE events. #3434 @pawel-big-lebowski
*Send input datasets in COMPLETE events while making sure version facet is attached on START only.#3432 @MassyB
Fixes incorrect structure of ParentRunFacet.Flink: Experimental version for flink native lineage listener. `#3099` @pawel-big-lebowski *New flink listener to extract lineage through native Flink
#3099 @pawel-big-lebowski
New flink listener to extract lineage through native Flink interfaces. Supports Flink SQL. Requires Flink 2.0.#3362 @MassyB
New option for dbt integration now can handle test and build commands too.#3379 @ssanthanam185
Events emitted from RDDExecutionContext now include custom facets that get loaded as part of InternalHandlerFactory.#3390 @JDarDagran
DatasetTypeDatasetFacet allows explicit declaration of type of the resulting dataset.#3391 @JDarDagran
Adds with_additonal_properties method that allows to create modified instance of facet with additional properties.#3379 @ssanthanam185
Events emitted from RDDExecutionContext now include custom facets that get loaded as part of InternalHandlerFactory.#3403 @ssanthanam185
Those events shouldn't be filtered outside Databricks/Delta ecosystem.#3368 @ddebowczyk92
Fixes ClassNotFoundException issue when using the openlineage-spark integration alongside a Spark connector that implements the spark-extension-interfaces due to class loader conflicts.#3368 @cisenbe
SQL parser won't error on Snowflake's LATERAL keyword.#3311 @dsaxton-1password
dbt integration won't fail when looking at tests on seeds.#3379 @ssanthanam185
Spark integration now correctly handles complex jobs that have cycles and nested RDD trees.json file extension. #3404 @kacpermuda
When append=False, the json file extension wasn't properly added before.dbt: Consume dbt structured logs and report progress in real time. `#3314` @MassyB *If --consume-structured-logs flag is set, dbt integration will con
#3314 @MassyB
If --consume-structured-logs flag is set, dbt integration will consume dbt structured logs and report execution progress in real time.transform transport to allow event modification. #3301 @pawel-big-lebowski
New transport type allows to modify the event based on the specified transformer class.#3305[#3305] @pawel-big-lebowski
Emit events in parallel for composite transport. Running in parallel is a default behaviour continueOnFailure set to true. Default value of continueOnFailure got changed from false to true.ScanReport and CommitReport in OpenLineage events when dealing with Iceberg tables. #3256 @pawel-big-lebowski
Collects additional Iceberg metrics for datasets read or written through the library. Visit Dataset Metrics docs for more details.#3280 @mobuchowski
Adds support for duckdb adapter for dbt integration.DatasetFactory to support Dataset creation. #3207 @pawel-big-lebowski
Adds DatasetFactory to support Dataset creation. This class is used to create Dataset instances for DatasetFactory.#3285 @pawel-big-lebowski
GCS path now has correctly stripped leading slash…marks some public developers' API methods as deprecated.*
#3264 @mayurmadnani
Dbt integration now uses SQL parser to add information about collected column-level lineage.#3240#3263 @pawel-big-lebowski
Fix issues related to existing output statistics collection mechanism and fetch input statistics. Output statistics contain now amount of files written, bytes size as well as records written. Input statistics contain bytes size and number of files read, while record count is collected only for DataSourceV2 sources.#3238 @pawel-big-lebowski#3244 @tnazarew
Excludes META-INF/*TransportBuilder to avoid version conflictsDatasetFactory #3207 @pawel-big-lebowskiDatasetFactory class, marks some public developers' API methods as deprecated.Spark: Add Dataproc run facet to include jobType property `#3167` @codelixir *Updates the GCP Dataproc run facet to include jobType property*
#3167 @codelixir#3186 @JDarDagran#3221 @JDarDagran#3142 @arturowczarek#3205 @mobuchowski#3219 @arturowczarek#3148 @codelixir#3215 @mobuchowski#3217 @arturowczarek#3208 @MassyB#3141 @pawel-leszczynski#3167 @codelixir#3186 @JDarDagran#3221 @JDarDagran#3142 @arturowczarek#3205 @mobuchowski#3219 @arturowczarek#3148 @codelixir#3215 @mobuchowski#3217 @arturowczarek#3208 @MassyB#3141 @pawel-leszczynskiNothing published for this version
Nothing published for this version
Java: added CompositeTransport `#3039` @JDarDagran *This allows user to specify multiple targets to which OpenLineage events will be emitted.*
#3039 @JDarDagran#3062 @Imbruced#3043 @ddebowczyk92#3077 @ddebowczyk92#3129 @arturowczarek#3094 @JDarDagran#3114 @JDarDagran#3116 @JDarDagran#3097 #3098 @arturowczarek#3122 @ddebowczyk92#3054 @JDarDagran#2962 @Imbruced#3068 @jonathanlbt1#3107 @Imbruced#3095 @ImbrucedSQL: add support for `USE` statement with different syntaxes `#2944` @kacpermuda *Adjusts our Context so that it can use the new support for this stat
USE statement with different syntaxes #2944 @kacpermuda#3044 @arturowczarek#3007 #3023 @pawel-big-lebowskiwebsite directory.SingleQuotedString in Identifier() #3035 @kacpermudaIDENTIFIER function instead of treating it like table name #2999 @kacpermuda#2918 @Imbruced#3020 @arturowczarekSpec: add GCP Dataproc facet `#2987` @tnazarew *Registers the Google Cloud Platform Dataproc run facet.*
#2987 @tnazarew
Registers the Google Cloud Platform Dataproc run facet.#2983 @kacpermuda#3001 @arturowczarek#2986 @Imbruced#2990 @arturowczarek#2984 @arturowczarektable/.#2937 @d-m-h
Previously, reading Iceberg datasets outside the configured Spark catalog prevented the datasets from being present in the inputs property of the RunEvent.Nothing published for this version
Nothing published for this version
Python: add `CompositeTransport` `#2925` @JDarDagran *Adds a CompositeTransport that can accept other transport configs to instantiate transports and
CompositeTransport #2925 @JDarDagranCompositeTransport that can accept other transport configs to instantiate transports and use them to emit events.#2828 @pawel-big-lebowski#2854 @pawel-big-lebowskiWindow #2901 @tnazarewWindow-type nodes of a logical plan.#2913 @ImbrucedQueryExecution to get the SQL query used from the SQL field with a BFS algorithm.#2887 @Imbruced#2912 @arturowczarekFacetConfig accept the disabled flag for any facet instead of passing them as a list.#2906 @ImbrucedDatasetIdentifier from extension LineageNode #2900 @ddebowczyk92LogicalRelation has a grandChild node that implements the LineageRelation interface.BaseRelation #2893 @ddebowczyk92DatasetIdentifier is now extracted from the underlying node of LogicalRelation.#2889 @jonathanlbt1marquez-web service to docker-compose.yml.#2880 @jonathanlbt1#2877 @jonathanlbt1for each batch method #2868 @Imbruced#2943 @arturowczarekIcebergHandler support Glue catalog tables and create the symlink using the code from PathUtils.#2917 @arturowczarek#2892 @arturowczarekLogicalPlanSerializer now returns <failed-to-serialize-logical-plan> for failed serialization instead of an empty string.#2883 @arturowczarekCustomCollectorsUtils for improved readability.foreach batch mode #2868 @ImbrucedDatasetIdentifier from SaveIntoDataSourceCommandVisitor options #2934 @ddebowczyk92DatasetIdentifier from command's options instead of relying on p.createRelation(sqlContext, command.options()), which is a heavy operation for JdbcRelationProvider.Nothing published for this version
Nothing published for this version
Airflow: add `log_url` to `AirflowRunFacet` `#2852` @dolfinus *Adds taskinstance's log_url field to AirflowRunFacet.*
log_url to AirflowRunFacet #2852 @dolfinuslog_url field to AirflowRunFacet.Generate #2856 @tnazarewGenerate-type nodes of a logical plan (e.g., explode operations).DerbyJdbcExtractor #2869 @dolfinusJdbcExtractor implementation for Derby database. As this is a file-based DBMS, its Dataset namespace is file and name is an absolute path to a database file.#2859 @pawel-big-lebowskiJarVerifier plugin to ensure all compiled classes have a bytecode version of Java 8 or lower.#2851 @d-m-h#2865 @kacpermudaColumnLevelLineageBuilder #2850 @tnazarewStreams dependency in ColumnLevelLineageBuilder causing a ClassNotFoundException.#2863 @dolfinus#2855 @ddebowczyk92PlanUtils3 so Dataset identifier information based on a Table's properties is also retrieved during the construction of column-level lineage.#2861 @arturowczarekspark.app.name was autogenerated by Glue and uses the Glue job name in such cases. Also, each job name provisioning strategy is now extracted to a separate provider.Spark: configurable integration test `#2755` @pawel-big-lebowski *Provides command line tool capable of running Spark integration tests that can be cr
#2755 @pawel-big-lebowski#2809 #2837 @ddebowczyk92#2743 @pawel-big-lebowski#2789 @tnazarewColumnLineageDatasetFacet creation.InsertIntoHadoopFsRelationCommand #2794 @dolfinusINSERT INTO command for tables created with USING $fileFormat syntax, like USING orc.PostgresJdbcExtractor #2806 @dolfinus
Adds the default 5432 port to Postgres namespaces.TeradataJdbcExtractor #2826 @dolfinus
Converts JDBC URLs like jdbc:teradata/host/DBS_PORT=1024,DATABASE=somedb to datasets with namespace teradata://host:1024 and name somedb.table.MySqlJdbcExtractor #2825 @dolfinus
Handles different formats of MySQL JDBC URL, and produces datasets with consistent namespaces, like mysql://host:port.OracleJdbcExtractor #2824 @dolfinus
Handles simple Oracle JDBC URLs, like oracle:thin:@//host:port/serviceName and oracle:thin@host:port:sid, and converts each to a dataset with namespace oracle://host:port and name sid.schema.table or serviceName.schema.table.#2822 @pawel-big-lebowski#2838 @pawel-big-lebowskiUNION queries.#2756 #2801 @Sheeri
Updates the customLineage facet test for the new syntax created in #2756.spark.sql.warehouse.dir as table namespace #2767 @dolfinusspark.sql.warehouse.dir or hive.metastore.warehouse.dir as table namespace, instead of duplicating the table's location.JdbcExtractors #2830 @dolfinus#2807 @Akash2351
Fixes Glue symlinks with config parsing for Glue catalogid.#2800 @dolfinus
Fixes the DBFS namespace format.#2766 @dolfinus#2797 @dolfinusfile:/some/path/database.table uses file:/some/path/database/table. For dataset TABLE symlink, uses warehouse location instead of database location.#2827 @pawel-big-lebowski
Fixes an error caused by a recent upgrade of Spark versions that did not break existing tests.JdbcLocation #2831 @dolfinustransformationType and transformationDescription are marked as deprecated.*
#2720 @pawel-big-lebowski#2758 @tnazarew#2643 @codelixirGCPRunFacetBuilder and GCPJobFacetBuilder to report additional facets when running on Google Cloud Platform.#2773 @dolfinus#2698 @pawel-big-lebowskishadowJar content and prevent reported issues. These are hard to prevent currently and require manual verification of manually unpacked jar content.#2756 @tnazarewColumnLineageDatasetFacet. transformationType and transformationDescription are marked as deprecated.#2729 @harels#2740 @ngorchakovalocalServerId option from Kafka config #2738 @dolfinuslocalServerId from Kafka config, deprecated since 1.13.0.Transport.emit(String) #2737 @dolfinusTransport.emit(String) support, deprecated since 1.13.0.spark-interfaces-scala module #2781 @ddebowczyk92spark-interfaces-scala interfaces with new ones decoupled from the Scala binary version. Allows for improved integration in environments where one cannot guarantee the same version of openlineage-java.#2769 @algorithmy1namespace.name as Avro complex field type #2763 @dolfinusnamespace.name is now used as Avro "type" of complex fields (record, enum, fixed).#2776 @kacpermudadrop table for Spark 3.4 and above #2745 @pawel-big-lebowski @savannavalgi#2782 @dolfinuss3a:// and s3n:// schemes to s3://.#2761 @dolfinus$SPARK_CONF_DIR/hive-site.xml.#2749 @pawel-big-lebowskicur.getDependencies() is not null before adding dependencies.OpenLineageRunEventBuilder #2754 @pawel-big-lebowskiOpenLineageRunEventBuilder::buildRun.historyUrl format #2741 @dolfinushistoryUrl format in spark_applicationDetails.#2753 @mobuchowskiselect * from test_orders as test_orders are now parsed properly.Nothing published for this version
Spark: add `jobType` facet to Spark application events `#2719` @dolfinus *Adds jobType facet to runEvents emitted by SparkListenerApplicationStart.*
jobType facet to Spark application events #2719 @dolfinus
Adds jobType facet to runEvents emitted by SparkListenerApplicationStart.jobType facet to Spark application events #2719 @dolfinusjobType facet to runEvents emitted by SparkListenerApplicationStart.#2720 @pawel-big-lebowskiHostListNamespaceResolver, PatternNamespaceResolver, PatternMatchingGroupNamespaceResolver or custom implementation loaded with ServiceLoader. Feature is useful to resolve hostnames into cluster identifiers.#2735 @JDarDagran#2727 @mobuchowskijobType facet to Spark application events #2719 @dolfinusjobType facet to runEvents emitted by SparkListenerApplicationStart.#2735 @JDarDagran#2727 @mobuchowskiPython: suppress warning on importing v1 module in __init__.py. `#2713` @JDarDagran *Suppresses the deprecation warning when v1 facets are used.*
#2706 @dolfinusSchemaDatasetFacet with nested fields for Iceberg tables with list, map and struct columns.#2711 @dolfinusSchemaDatasetFacet with nested fields for Avro schemas with complex types (union, record, map, array, fixed).#2677 @dolfinusExecutionContext interface.SchemaDatasetFieldsFacet #2689 @dolfinusSchemaDatasetFieldsFacet. Also include field comment as description.SparkApplicationDetailsFacet #2688 @dolfinusSparkApplicationDetailsFacet to runEvents emitted on Spark application start.#2710 @kacpermuda#2693 @JDarDagran#2665 @pawel-big-lebowskiAthenaExtractor #2700 @kacpermudaSchemaDatasetFacet for Protobuf repeated primitive types #2685 @dolfinus#2653 @kacpermudaTokenizerErrors, PanicException #2703 @mobuchowski#2713 @JDarDagran#2686 #2687 @dolfinusrunEvents. The new UUID version produces monotonically increasing values, which leads to more performant queries on the OL consumer side. Note: UUID version is an implementation detail and can be changed in the future.Spark: drop `SparkVersionFacet` `#2659` @dolfinus *Drops the SparkVersion facet, deprecated since 1.2.0 and planned for removal since 1.4.0.*
#2674 @surisimran#2668.#2482 @pawel-big-lebowskiarray type, map type, oneOf and any.#2663 @julienledem#2652 @pawel-big-lebowskijobType property of JobTypeJobFacet to either SQL_JOB or RDD_JOB.#2646 @mobuchowskispark_jobDetails facet #2662 @dolfinus
Adds a SparkJobDetailsFacet, capturing information about Spark application jobs -- e.g. jobId, jobDescription, jobGroup, jobCallSite. This allows for tracking an OpenLineage RunEvent with a specific Spark job in SparkUI.ParentRunFacet key #2660 @dolfinus
Changes the integration to use the parent key for ParentFacet, dropping the outdated parentRun.SparkVersionFacet #2659 @dolfinus
Drops the SparkVersion facet, deprecated since 1.2.0 and planned for removal since 1.4.0.#2679 @JDarDagranParentRunFacet key #2661 @dolfinusParentRunFacet with the property name parent but the Great Expectations integration created a lineage event with parentRun. This renames ParentRunFacet key from parentRun to parent. For backwards compatibility, keep the old name.#2658 @blacklight
Includes profile and models in the dbt job name to make it more unique.org.apache.commons.lang3 instead of org.apache.commons.lang #2676 @harelsThe v2 interface introduces some breaking changes: facets are put into separate modules per JSON Schema spec file, some names are changed, and several…
#2609 @pawel-big-lebowskiDataSetEvent and JobEvent in Transport.emit #2611 @dolfinusTransport.emit(OpenLineage.DatasetEvent) and Transport.emit(OpenLineage.JobEvent), reusing the implementation of Transport.emit(OpenLineage.RunEvent). Please note: Transport.emit(String) is now deprecated and will be removed in 1.16.0.GZIP compression to HttpTransport #2603 #2604 @dolfinuscompression option to HttpTransport config in the Java and Python clients, with gzip implementation.#2571 #2597 #2598 @dolfinusmessageKey option to KafkaTransport config in the Python and Java clients, as well as the Proxy. This option replaces the localServerId option, which is now deprecated. Default value is generated using the run id (for RunEvent), job name (for JobEvent) or dataset name (for DatasetEvent). This value is used by the Kafka producer to distribute messages along topic partitions, instead of sending all the events to the same partition. This allows for full utilization of Kafka performance advantages.#2633 @mobuchowski
Adds a mechanism for forwarding metrics to any Micrometer-compatible implementation for Flink as has been implemented for Spark. Included: MeterRegistry, CompositeMeterRegistry, SimpleMeterRegistry, and MicrometerProvider.#2520 @JDarDagrandatamodel-code-generator for parsing JSON Schema and generating Pydantic or dataclasses classes, etc. In order to use attrs (a more modern version of dataclasses) and overcome some limitations of the tool, a number of steps have been added in order to customize code to meet OpenLineage requirements. Included: updated references to the latest base JSON Schema spec for all child facets. Please note: newly generated code creates a v2 interface that will be implemented in existing integrations in a future release. The v2 interface introduces some breaking changes: facets are put into separate modules per JSON Schema spec file, some names are changed, and several classes are now kw_only.#2583 @pawel-big-lebowskiSparkOpenlineageConfig and FlinkOpenlineageConfig for a more uniform configuration experience for the user. Renames OpenLineageYaml to OpenLineageConfig and modifies the code to use only OpenLineageConfig classes. Includes a doc update to mention that both ways can be used interchangeably and final documentation will merge all values provided.#2613 @tnazarew
Adds a TokenProviderTypeIdResolver to handle both FQCN and (for backward compatibility) api_key types in spark.openlineage.transport.auth.type.SparkConf & FlinkConf approaches #2583 @pawel-big-lebowskiSparkConf/FlinkConf). Allows each integration to have its own config entries.#2533 @pawel-big-lebowski
Enables configuration entries specifying ownership of the job that will result in an OwnershipJobFacet being attached to job facets.partitionKey format with Kafka implementation #2620 @dolfinus
Changes the format of Kinesis partitionKey from {jobNamespace}:{jobName} to run:{jobNamespace}/{jobName} to match the Kafka transport implementation.load_config return an empty dict instead of None when file empty #2596 @kacpermuda
utils.load_config() now returns an empty dict instead of None in the case of an empty file to prevent an OpenLineageClient crash.#2614 @dolfinus
Fixes rendering of javadoc for methods generated by lombok annotations by adding a delombok step.#2599 @mobuchowski…maintained and recent releases have introduced breaking changes.*
lineage_job_namespace and lineage_job_name macros #2582 @dolfinuslineage_job_namespace(), lineage_job_name(task) that return an Airflow namespace and Airflow job name, respectively.SchemaDatasetFacet #2548 @dolfinusSchemaDatasetFacet.#2588 @pawel-big-lebowskipmdTestScala212 of warnings that clutter the logs.#2591 @blacklightdbt-ol now propagates the exit code of the underlying dbt process even if no lineage events are emitted.#2585 @harelsjava.lang.StringIndexOutOfBoundsException.HashSet in column-level lineage instead of iterating through LinkedList #2584 @mobuchowskiHashSet for collection.#2572 @dolfinuspkg_resources dependency and replaces it with the packaging lib.airflow.macros.lineage_parent_id #2578 @blacklightlineage_parent_id Airflow macro and simplifies the format of the lineage_parent_id and lineage_run_id macros.#2579 @JDarDagran
Adds an upper limit on supported versions of Dagster as the integration is no longer actively maintained and recent releases have introduced breaking changes.lineage_job_namespace and lineage_job_name macros #2582 @dolfinus
Adds new Airflow macros lineage_job_namespace(), lineage_job_name(task) that return an Airflow namespace and Airflow job name, respectively.SchemaDatasetFacet #2548 @dolfinus
Allows nested fields support to SchemaDatasetFacet.airflow.macros.lineage_parent_id #2578 @blacklightlineage_parent_id Airflow macro and simplifies the format of the lineage_parent_id and lineage_run_id macros.#2591 @blacklightdbt-ol now propagates the exit code of the underlying dbt process even if no lineage events are emitted.#2579 @JDarDagran#2585 @harelsjava.lang.StringIndexOutOfBoundsException.#2624 @pawel-big-lebowskipkg_resources module on Python 3.12 #2572 @dolfinuspkg_resources dependency and replaces it with the packaging lib.HashSet in column-level lineage instead of iterating through LinkedList #2584 @mobuchowskiHashSet for collection.Common: add support for `SCRIPT`-type jobs in BigQuery `#2564` @kacpermuda In the case of SCRIPT-type jobs in BigQuery, no lineage was being extracted
SCRIPT-type jobs in BigQuery #2564 @kacpermuda
In the case of SCRIPT-type jobs in BigQuery, no lineage was being extracted because the SCRIPT job had no lineage information - it only spawned child jobs that had that information. With this change, the integration extracts lineage information from child jobs when dealing with SCRIPT-type jobs.#2272 @pawel-big-lebowski
This PR adds a spark-interfaces-scala package that allows lineage extraction to be implemented within Spark extensions (Iceberg, Delta, GCS, etc.). The Openlineage integration, when traversing the query plan, verifies if nodes implement defined interfaces. If so, interface methods are used to extract lineage. Refer to the README for more details.#2496 @mobuchowskiMeterRegistryyFactory, MicrometerProvider, StatsDMetricsBuilder, metrics config in OpenLineage config, and a Java client implementation.#2528 @mobuchowski#2556 @mobuchowski
Adds support for the Spark-BigQuery connector's query input option, which executes a query directly on BigQuery, storing the result in an intermediate dataset, bypassing Spark's computation layer. Due to this, the lineage is retrieved using the SQL parser, similarly to JDBCRelation.SparkPropertyFacetBuilder to support recording Spark runtime #2523 @Ruihua98SparkPropertyFacetBuilder to capture the RuntimeConfig of the Spark session because the existing SparkPropertyFacet can only capture the static config of the Spark context. This facet will be added in both RDD-related and SQL-related runs.fileCount to dataset stat facets #2562 @dolfinusfileCount field to DataQualityMetricsInputDatasetFacet and OutputStatisticsOutputDatasetFacet specification.dbt-ol should transparently exit with the same exit code as the child dbt process #2560 @blacklightdbt-ol transparently exit with the same exit code as the child dbt process.#2531 @HuangZhenQiuopenlineage-flink.jar.#2507 @pawel-big-lebowski
Fixes the class not found issue when checking for Cassandra classes. Also fixes the Maven pom dependency on subprojects..emit() method logging & annotations #2539 @dolfinus#2547 @dolfinusOpenLineageSql class could not load a native library, if returned None for all operations. But because the error message was suppressed, the user could not determine the reason.#2510 @mobuchowski#2535 @pawel-big-lebowskiIllegalStateException is always caught when accessing SparkSession.#2537 @pawel-big-lebowskiClassNotFoundError occurring on Databricks runtime and extends the integration test to verify DatabricksEnvironmentFacet.#2565 @d-m-h
The JobMetricsHolder#cleanUp(int) method now correctly purges unneeded state from both maps.UnknownEntryFacetListener #2557 @pawel-big-lebowski
Prevents storing the state when a facet is disabled, purging the state after populating run facets.JDBCOptions(table=...) containing subquery #2546 @dolfinusopenlineage-spark from producing datasets with names like database.(select * from table) for JDBC sources.#2563 @mobuchowskiIllegalStateException when accessing SparkSession #2535 @pawel-big-lebowski
IllegalStateException was not being caught.Nothing published for this version
Nothing published for this version
A new config param timeoutInMillis has been added. the Existing timeout has been removed from docs and will be deprecated in 1.13.*
#2518 @JDarDagran
Adds the new provider required by the latest version of Dagster.#2491 @HuangZhenQiu
Adds support for hybrid source lineage for users of Kafka and Iceberg sources in backfill usecases.#2472 @HuangZhenQiu
Bumps the Flink JDBC connector version to 3.1.2-1.18 for Flink 1.18.OpenLineageClientUtils#loadOpenLineageJson(InputStream) and change OpenLineageClientUtils#loadOpenLineageYaml(InputStream) methods #2490 @d-m-h
This improves the explicitness of the methods. Previously, loadOpenLineageYaml(InputStream) wanted the InputStream to contain bytes that represented JSON.#2486 @davidjgoss
Adds the status code and body as properties on the thrown exception when a non-success response is encountered in the HTTP transport.#2478 @mattiabertorello#2524 @kacpermuda
Refines the operator's attribute inclusion logic in facets to include only those known to be important or compact, ensuring that custom operator attributes with substantial data do not inflate the event size.task_instance copy fails #2492 @kacpermuda
Airflow will now proceed without rendering templates if task_instance copy fails in listener.on_task_instance_running.HttpTransport timeout #2475 @pawel-big-lebowski
The existing timeout config parameter is ambiguous: implementation treats the value as double in seconds, although the documentation claims it's milliseconds. A new config param timeoutInMillis has been added. the Existing timeout has been removed from docs and will be deprecated in 1.13.#2515 @pawel-big-lebowski
Adds a check for a null context before executing end(jobEnd).#2507 @pawel-big-lebowski
Fixes the class not found issue when checking for Cassandra classes. Also fixes the Maven POM dependency on subprojects.#2512 @HuangZhenQiu
Enables the JDBC table name with a schema prefix.#2508 @pawel-big-lebowski
For JDBC, the Flink integration is not adjusted to the Openlineage naming convention. There is code that extracts the dataset namespace/name from the JDBC connection url, but it's in the Spark integration. As a solution, this code has to be extracted into the Java client and reused by the Spark and Flink integrations.#2507 @pawel-big-lebowski
Flink is failing when no Cassandra classes are present on the class path. This is happening because of CassandraUtils class which has a static hasClasses method, but it imports Cassandra-related classes in the header. Also, the Flink subproject contains an unnecessary maven-publish plugin.#2504 @HuangZhenQiu
The shadow jar of Flink is not minimized, so some internal jars are listed as runtime dependences. This removes them from the final pom.xml file in the Flink module.#2479 @HuangZhenQiu
Following the namespace definition, we should use cassandra://host:port.Nothing published for this version
Nothing published for this version
Airflow: add support for `JobTypeJobFacet` properties `#2412` @mattiabertorello *Adds support for Job type properties within the Airflow Job facet.*
JobTypeJobFacet properties #2412 @mattiabertorelloJobTypeJobFacet properties #2411 @mattiabertorello
Adds support for Job type properties within the DBT Job facet.#2372 @pawel-big-lebowski
Adds support for multi-topic Kafka sinks. Limitations: recordSerializer needs to implement KafkaTopicsDescriptor. Please refer to the limitations sections in documentation.#2436 @HuangZhenQiu
Adds support for use cases that employ this connector.ServiceLoader #2435 @pawel-big-lebowski
Loads the circuit breaker builder with ServiceLoader as an addition to a list of implemented builders available within the existing package.#2371 @mobuchowski
Previously, the Spark event model described only single actions, potentially linked only to some parent run. Closes #1672.DataSourceV2Relation #2394 @pawel-big-lebowski
Enables built-in lineage extraction within from DataSourceV2Relation lineage nodes.JobTypeJobFacet properties #2410 @mattiabertorello
Adds support for Job type properties within the Spark Job facet.spark.LogicalPlan facet by default #2433 @pawel-big-lebowski
spark.LogicalPlan has been added to default value of spark.openlineage.facets.disabled.#2407 @pawel-big-lebowski
Introduces a circuit breaker mechanism to prevent effects of over-instrumentation. Implemented within Java client, it serves both the Flink and Spark integration. Read the Java client README for more details.openlineage-spark #2446 @d-m-h
Adds the capability to publish Scala 2.12 and 2.13 variants of openlineage-sparkapp module to be compiled with Scala 2.12 and Scala 2.13 variants of Apache Spark (https://github.com/OpenLineage/OpenLineage/pull/2432) @d-m-hspark.binary.version and spark.version properties control which variant to build.app module #2432 @d-m-h
Enables the app module to be built using both Scala 2.12 and Scala 2.13 variants of various Apache Spark versions, and enables the CI/CD pipeline to build and test them.UnknownEntryFacet creation #2431 @mobuchowski
Failure to generate UnknownEntryFacet was resulting in the event not being sent.#2405 @mattiabertorello
Creates a vendor folder to isolate Snowflake-specific code from the main Spark integration, enhancing organization and flexibility.#2403 @HuangZhenQiu
Resolves the PMD rule violation warnings in the Flink integration module.#2468 @d-m-h
The 'isReleaseVersion' property was removed from the build, preventing the Flink integration from being released.#2447 @kacpermuda
FileConfig was creating an additional file when not in append mode. Closes #2439.#2441 @kacpermuda
FileConfig was ignoring the append key in YAML config. Closes #2440#2379 @algorithmy1
In the case of symlinked Glue Catalog Tables, the parsing method was producing dataset names identical to the namespace.IcebergSourceWrapper for Iceberg connector 1.17 #2409 @ensctom
In Flink 1.17, the Iceberg catalogloader was loading the catalog in the open function, causing the loadTable method to throw a NullPointerException error.spark35, spark3, shared modules to produce Scala 2.12 and Scala 2.13 variants #2390 #2385#2384 @d-m-h
Migrates the three modules to use the refactored Gradle plugins. Also splits some tests into Scala 2.12- and Scala 2.13-specific versions.spark2 module to the new build process #2391 @d-m-hNoSuchMethodErrors were being thrown when running the openlineage-spack connector in an Apache Spark runtime compiled using Scala 2.13.Nothing published for this version
Flink: support Flink 1.18 `#2366` @HuangZhenQiu *Adds support for the latest Flink version with 1.17 used for Iceberg Flink runtime and Cassandra Conn
#2366 @HuangZhenQiu
Adds support for the latest Flink version with 1.17 used for Iceberg Flink runtime and Cassandra Connector as these do not yet support 1.18.#2376 @d-m-hLogicalPlan implementation #2361 @mattiabertorello
In the LogicalPlanSerializerTest class, the implementation of the LogicalPlan interface is different between Scala 2.12 and Scala 2.13. In detail, the IndexedSeq changes package from the scala.collection to scala.collection.immutable. This implements both of the methods necessary in the two versions.#2357 @mattiabertorello
This initial step is to start supporting compilation for Scala 2.13 in the 3.2+ Spark versions. Scala 2.13 changed the default collection to immutable, the methods to create an empty collection, and the conversion between Java and Scala. This causes the code to not compile between 2.12 and 2.13. This replaces the usage of direct Scala collection methods (like creating an empty object) and conversions utils with ScalaConversionUtils methods that will support cross-compilation.MERGE INTO queries on Databricks #2348 @pawel-big-lebowskiMERGE INTO queries on Databricks runtime.#2283 @nataliezeller1#2365 @mattiabertorello#2364 @kacpermuda#2358 @kacpermuda#2373 @kacpermuda#2350 @pawel-big-lebowski#2360 @mattiabertorelloremovePathPattern feature #2350 @pawel-big-lebowskiremovePathPattern if configured to do so.#2377 @d-m-h#2383 @d-m-h_COMPATIBILITY NOTICE_ Starting in 1.7.0, the Airflow integration will no longer support Airflow versions >=2.8.0. Please use the OpenLineage Airflow
COMPATIBILITY NOTICE
Starting in 1.7.0, the Airflow integration will no longer support Airflow versions >=2.8.0.
Please use the OpenLineage Airflow Provider instead.
COMPLETE and FAIL events in Airflow integration #2320 @kacpermuda
Adds a parent run facet to all events in the Airflow integration.#2316 #2318 @kacpermuda
Some scripts were not working well on MacOS. This adjusts them.run_id for FAIL event in Airflow 2.6+ #2305 @kacpermuda
The Run_id in a FAIL event was different than in the START event for Airflow 2.6+.TableLoader before loading a table #2314 @pawel-big-lebowski
Fixes a potential NullPointerException in 1.17 when dealing with Iceberg sinks.#2321 @pawel-big-lebowski
Adds a kafka:// prefix to Kafka topic datasets' namespaces.JobTypeJobFacet #2325 @pawel-big-lebowski
Fixes properties assignment in the Flink visitor.commons-logging relocate in target jar #2319 @pawel-big-lebowski
Avoids relocating a dependency that was getting excluded from the jar.#2315 @davidjgoss
Amends the Authority format for consistency with other references in the same section.#2330 @kacpermuda
To encourage use of the Provider, this removes the listener from the plugin if the Airflow version is >=2.8.0.COMPATIBILITY NOTICE
Starting in 1.7.0, the Airflow integration will no longer support Airflow versions >=2.8.0.
Please use the OpenLineage Airflow Provider instead.
COMPLETE and FAIL events in Airflow integration #2320 @kacpermuda#2316 #2318 @kacpermudarun_id for FAIL event in Airflow 2.6+ #2305 @kacpermudaRun_id in a FAIL event was different than in the START event for Airflow 2.6+.TableLoader before loading a table #2314 @pawel-big-lebowskiNullPointerException in 1.17 when dealing with Iceberg sinks.#2321 @pawel-big-lebowskikafka:// prefix to Kafka topic datasets' namespaces.JobTypeJobFacet #2325 @pawel-big-lebowskicommons-logging relocate in target jar #2319 @pawel-big-lebowski#2315 @davidjgossAuthority format for consistency with other references in the same section.#2330 @kacpermuda>=2.8.0, the Airflow integration's plugin does not import the integration's listener, disabling the external integration.Spark: update Jackson dependency to resolve `CVE-2022-1471` `#2185` @pawel-big-lebowski *Updates Gradle for Spark and Flink to 8.1.1. Upgrade Jackson…
#2220 @tsungchih
Gets event records for each target Dagster event type to support Dagster version 0.15.0+.dbt-ol send-events to send metadata of the last run without running the job #2285 @sophiely
Adds a new command to send events to OpenLineage according to the latest metadata generated without running any dbt command.#2229 @ensctom
Adds option for the Flink job listener to read jobnames and namespaces from Flink conf.#2284 @mobuchowski
Adds support for dbtable, enables lineage in the case of single input columns, and improves dataset naming.JobTypeJobFacet to contain additional job related information#2241 @pawel-big-lebowski
New JobTypeJobFacet contains the processing type such as BATCH|STREAMING, integration via SPARK|FLINK|... and job type in QUERY|COMMAND|DAG|....#2259 @JDarDagran
Adds quote information from sqlparser-rs.CVE-2022-1471 #2185 @pawel-big-lebowski
Updates Gradle for Spark and Flink to 8.1.1. Upgrade Jackson 2.15.3.#2296 @pawel-big-lebowski
Removes usage of Guava ImmutableList.commons-logging transitive dependency from published jar #2297 @pawel-big-lebowski
Ensures commons-logging is not shipped as this can lead to a version mismatch on the user's side.Nothing published for this version
Nothing published for this version
Flink: add Flink lineage for Cassandra Connectors `#2175` @HuangZhenQiu *Adds Flink Cassandra source and sink visitors and Flink Cassandra Integration
#2175 @HuangZhenQiu
Adds Flink Cassandra source and sink visitors and Flink Cassandra Integration test.rdd and toDF operations available in Spark Scala API #2188 @pawel-big-lebowskiExternalRddVisitor and adds support for extracting inputs from MapPartitionsRDD and ParallelCollectionRDD plan nodes.#2185 @pawel-big-lebowski
Modifies the Spark integration to support the latest Databricks Runtime version.#2107 @JDarDagran
Lowers the version requirements for attrs and requests and removes an unnecessary dependency.#2221 @JDarDagran
Don't render each entry in yaml files at start.#2167 @sophiely
Replaces the dataset and namespace with the data's physical location for more complete lineage across integrations.#2177 @JDarDagran
Redacted fields in ColumnLineageDatasetFacetFieldsAdditionalInputFields are now skipped.#2181 @pawel-big-lebowski
Use the same mechanism for RDD jobs to extract dataset identifier as used for Spark SQL.START and a single COMPLETE event are sent #2103 @pawel-big-lebowski
For Spark SQL at least four events are sent triggered by different SparkListener methods. Each of them is required and used to collect facets unavailable elsewhere. However, there should be only one START and COMPLETE events emitted. Other events should be sent as RUNNING. Please keep in mind that Spark integration remains stateless to limit the memory footprint, and it is the backend responsibility to merge several Openlineage events into a meaningful snapshot of metadata changes.Client: allow setting client's endpoint via environment variable `#2151` @mars-lan *Enables setting this endpoint via environment variable because cre
#2151 @mars-lan
Enables setting this endpoint via environment variable because creating the client manually in Airflow is not possible.#2149 @HuangZhenQiuFlinkIcebergSource and FlinkIcebergTableSource for Flink Iceberg lineage.#2147 @pawel-big-lebowskispark.openlineage.debugFacet=enabled needs to be set to include the facet. By default, the debug facet is disabled.#2165 @julwinAirflow: add some basic stats to the Airflow integration `#1845` @harels *Uses the statsd component that already exists in the Airflow codebase and wr
#1845 @harelsairflow.lineage.Table (if defined) #2138 @erikalfthanairflow.lineage.Table inlets/outlets to the OpenLineage Dataset.#2136 @erikalfthan#2118 @pawel-big-lebowskidelta and iceberg are not supported for Spark 3.5 at this time.#2139 @JDarDagran#2141 @JDarDagranairflow.providers.openlineage and adds more graceful logging to fix a corner case.#2142 @d-m-hprepareDatasetIdentifierFromDefaultTablePath method would override the scheme with the value of "file" when constructing a dataset identifier. It now uses the scheme of the CatalogTable's URI for this. Thank you @pawel-big-lebowski for the quick triage and suggested fix.Nothing published for this version
SQL: remove sqlparser dependency from iface-java and iface-py `#2090` @JDarDagran *Removes the dependency due to a breaking change in the latest relea…
ProcessingEngineRunFacet as part of the normal operation of the OpenLineageSparkEventListener #2089 @d-m-hProcessEngineRunFacet alongside the custom SparkVersionFacet (for now).
The SparkVersionFacet is deprecated and will be removed in a future release.spark.databricks.clusterUsageTags.clusterAllTags variable from databricks environment #2099 @Anirudh181001spark.databricks.clusterUsageTags.clusterAllTags to the list of environment variables captured from databricks.#2106 @tatianaDbtLocalArtifactProcessor in dbt projects that do not declare target-path.#2091 @harels#2044 @xli-1026apiKey if loading it from env variables @2029 @mobuchowskiapi_key to apiKey in create_token_provider.#2039 @pawel-big-lebowskirunning events after job completes #2075 @pawel-big-lebowski#2083 @pawel-big-lebowski#2076 @pawel-big-lebowski#2090 @JDarDagranNothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →