NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1874 most downloaded on PyPI
OpenLineage common python library for integrations
Last release 17 days ago
01 Sep 2026
Ships on a steady schedule
a new release about every 4 weeks
Most releases are documented
notes for 50 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
123 releases · first in 2021
One column per quarter.
Airflow: add support for `JobTypeJobFacet` properties `#2412` @mattiabertorello *Adds support for Job type properties within the Airflow Job facet.*
JobTypeJobFacet properties #2412 @mattiabertorelloJobTypeJobFacet properties #2411 @mattiabertorello
Adds support for Job type properties within the DBT Job facet.#2372 @pawel-big-lebowski
Adds support for multi-topic Kafka sinks. Limitations: recordSerializer needs to implement KafkaTopicsDescriptor. Please refer to the limitations sections in documentation.#2436 @HuangZhenQiu
Adds support for use cases that employ this connector.ServiceLoader #2435 @pawel-big-lebowski
Loads the circuit breaker builder with ServiceLoader as an addition to a list of implemented builders available within the existing package.#2371 @mobuchowski
Previously, the Spark event model described only single actions, potentially linked only to some parent run. Closes #1672.DataSourceV2Relation #2394 @pawel-big-lebowski
Enables built-in lineage extraction within from DataSourceV2Relation lineage nodes.JobTypeJobFacet properties #2410 @mattiabertorello
Adds support for Job type properties within the Spark Job facet.spark.LogicalPlan facet by default #2433 @pawel-big-lebowski
spark.LogicalPlan has been added to default value of spark.openlineage.facets.disabled.#2407 @pawel-big-lebowski
Introduces a circuit breaker mechanism to prevent effects of over-instrumentation. Implemented within Java client, it serves both the Flink and Spark integration. Read the Java client README for more details.openlineage-spark #2446 @d-m-h
Adds the capability to publish Scala 2.12 and 2.13 variants of openlineage-sparkapp module to be compiled with Scala 2.12 and Scala 2.13 variants of Apache Spark (https://github.com/OpenLineage/OpenLineage/pull/2432) @d-m-hspark.binary.version and spark.version properties control which variant to build.app module #2432 @d-m-h
Enables the app module to be built using both Scala 2.12 and Scala 2.13 variants of various Apache Spark versions, and enables the CI/CD pipeline to build and test them.UnknownEntryFacet creation #2431 @mobuchowski
Failure to generate UnknownEntryFacet was resulting in the event not being sent.#2405 @mattiabertorello
Creates a vendor folder to isolate Snowflake-specific code from the main Spark integration, enhancing organization and flexibility.#2403 @HuangZhenQiu
Resolves the PMD rule violation warnings in the Flink integration module.#2468 @d-m-h
The 'isReleaseVersion' property was removed from the build, preventing the Flink integration from being released.#2447 @kacpermuda
FileConfig was creating an additional file when not in append mode. Closes #2439.#2441 @kacpermuda
FileConfig was ignoring the append key in YAML config. Closes #2440#2379 @algorithmy1
In the case of symlinked Glue Catalog Tables, the parsing method was producing dataset names identical to the namespace.IcebergSourceWrapper for Iceberg connector 1.17 #2409 @ensctom
In Flink 1.17, the Iceberg catalogloader was loading the catalog in the open function, causing the loadTable method to throw a NullPointerException error.spark35, spark3, shared modules to produce Scala 2.12 and Scala 2.13 variants #2390 #2385#2384 @d-m-h
Migrates the three modules to use the refactored Gradle plugins. Also splits some tests into Scala 2.12- and Scala 2.13-specific versions.spark2 module to the new build process #2391 @d-m-hNoSuchMethodErrors were being thrown when running the openlineage-spack connector in an Apache Spark runtime compiled using Scala 2.13.Nothing published for this version
Flink: support Flink 1.18 `#2366` @HuangZhenQiu *Adds support for the latest Flink version with 1.17 used for Iceberg Flink runtime and Cassandra Conn
#2366 @HuangZhenQiu
Adds support for the latest Flink version with 1.17 used for Iceberg Flink runtime and Cassandra Connector as these do not yet support 1.18.#2376 @d-m-hLogicalPlan implementation #2361 @mattiabertorello
In the LogicalPlanSerializerTest class, the implementation of the LogicalPlan interface is different between Scala 2.12 and Scala 2.13. In detail, the IndexedSeq changes package from the scala.collection to scala.collection.immutable. This implements both of the methods necessary in the two versions.#2357 @mattiabertorello
This initial step is to start supporting compilation for Scala 2.13 in the 3.2+ Spark versions. Scala 2.13 changed the default collection to immutable, the methods to create an empty collection, and the conversion between Java and Scala. This causes the code to not compile between 2.12 and 2.13. This replaces the usage of direct Scala collection methods (like creating an empty object) and conversions utils with ScalaConversionUtils methods that will support cross-compilation.MERGE INTO queries on Databricks #2348 @pawel-big-lebowskiMERGE INTO queries on Databricks runtime.#2283 @nataliezeller1#2365 @mattiabertorello#2364 @kacpermuda#2358 @kacpermuda#2373 @kacpermuda#2350 @pawel-big-lebowski#2360 @mattiabertorelloremovePathPattern feature #2350 @pawel-big-lebowskiremovePathPattern if configured to do so.#2377 @d-m-h#2383 @d-m-h_COMPATIBILITY NOTICE_ Starting in 1.7.0, the Airflow integration will no longer support Airflow versions >=2.8.0. Please use the OpenLineage Airflow
COMPATIBILITY NOTICE
Starting in 1.7.0, the Airflow integration will no longer support Airflow versions >=2.8.0.
Please use the OpenLineage Airflow Provider instead.
COMPLETE and FAIL events in Airflow integration #2320 @kacpermuda
Adds a parent run facet to all events in the Airflow integration.#2316 #2318 @kacpermuda
Some scripts were not working well on MacOS. This adjusts them.run_id for FAIL event in Airflow 2.6+ #2305 @kacpermuda
The Run_id in a FAIL event was different than in the START event for Airflow 2.6+.TableLoader before loading a table #2314 @pawel-big-lebowski
Fixes a potential NullPointerException in 1.17 when dealing with Iceberg sinks.#2321 @pawel-big-lebowski
Adds a kafka:// prefix to Kafka topic datasets' namespaces.JobTypeJobFacet #2325 @pawel-big-lebowski
Fixes properties assignment in the Flink visitor.commons-logging relocate in target jar #2319 @pawel-big-lebowski
Avoids relocating a dependency that was getting excluded from the jar.#2315 @davidjgoss
Amends the Authority format for consistency with other references in the same section.#2330 @kacpermuda
To encourage use of the Provider, this removes the listener from the plugin if the Airflow version is >=2.8.0.COMPATIBILITY NOTICE
Starting in 1.7.0, the Airflow integration will no longer support Airflow versions >=2.8.0.
Please use the OpenLineage Airflow Provider instead.
COMPLETE and FAIL events in Airflow integration #2320 @kacpermuda#2316 #2318 @kacpermudarun_id for FAIL event in Airflow 2.6+ #2305 @kacpermudaRun_id in a FAIL event was different than in the START event for Airflow 2.6+.TableLoader before loading a table #2314 @pawel-big-lebowskiNullPointerException in 1.17 when dealing with Iceberg sinks.#2321 @pawel-big-lebowskikafka:// prefix to Kafka topic datasets' namespaces.JobTypeJobFacet #2325 @pawel-big-lebowskicommons-logging relocate in target jar #2319 @pawel-big-lebowski#2315 @davidjgossAuthority format for consistency with other references in the same section.#2330 @kacpermuda>=2.8.0, the Airflow integration's plugin does not import the integration's listener, disabling the external integration.Spark: update Jackson dependency to resolve `CVE-2022-1471` `#2185` @pawel-big-lebowski *Updates Gradle for Spark and Flink to 8.1.1. Upgrade Jackson…
#2220 @tsungchih
Gets event records for each target Dagster event type to support Dagster version 0.15.0+.dbt-ol send-events to send metadata of the last run without running the job #2285 @sophiely
Adds a new command to send events to OpenLineage according to the latest metadata generated without running any dbt command.#2229 @ensctom
Adds option for the Flink job listener to read jobnames and namespaces from Flink conf.#2284 @mobuchowski
Adds support for dbtable, enables lineage in the case of single input columns, and improves dataset naming.JobTypeJobFacet to contain additional job related information#2241 @pawel-big-lebowski
New JobTypeJobFacet contains the processing type such as BATCH|STREAMING, integration via SPARK|FLINK|... and job type in QUERY|COMMAND|DAG|....#2259 @JDarDagran
Adds quote information from sqlparser-rs.CVE-2022-1471 #2185 @pawel-big-lebowski
Updates Gradle for Spark and Flink to 8.1.1. Upgrade Jackson 2.15.3.#2296 @pawel-big-lebowski
Removes usage of Guava ImmutableList.commons-logging transitive dependency from published jar #2297 @pawel-big-lebowski
Ensures commons-logging is not shipped as this can lead to a version mismatch on the user's side.Nothing published for this version
Nothing published for this version
Flink: add Flink lineage for Cassandra Connectors `#2175` @HuangZhenQiu *Adds Flink Cassandra source and sink visitors and Flink Cassandra Integration
#2175 @HuangZhenQiu
Adds Flink Cassandra source and sink visitors and Flink Cassandra Integration test.rdd and toDF operations available in Spark Scala API #2188 @pawel-big-lebowskiExternalRddVisitor and adds support for extracting inputs from MapPartitionsRDD and ParallelCollectionRDD plan nodes.#2185 @pawel-big-lebowski
Modifies the Spark integration to support the latest Databricks Runtime version.#2107 @JDarDagran
Lowers the version requirements for attrs and requests and removes an unnecessary dependency.#2221 @JDarDagran
Don't render each entry in yaml files at start.#2167 @sophiely
Replaces the dataset and namespace with the data's physical location for more complete lineage across integrations.#2177 @JDarDagran
Redacted fields in ColumnLineageDatasetFacetFieldsAdditionalInputFields are now skipped.#2181 @pawel-big-lebowski
Use the same mechanism for RDD jobs to extract dataset identifier as used for Spark SQL.START and a single COMPLETE event are sent #2103 @pawel-big-lebowski
For Spark SQL at least four events are sent triggered by different SparkListener methods. Each of them is required and used to collect facets unavailable elsewhere. However, there should be only one START and COMPLETE events emitted. Other events should be sent as RUNNING. Please keep in mind that Spark integration remains stateless to limit the memory footprint, and it is the backend responsibility to merge several Openlineage events into a meaningful snapshot of metadata changes.Client: allow setting client's endpoint via environment variable `#2151` @mars-lan *Enables setting this endpoint via environment variable because cre
#2151 @mars-lan
Enables setting this endpoint via environment variable because creating the client manually in Airflow is not possible.#2149 @HuangZhenQiuFlinkIcebergSource and FlinkIcebergTableSource for Flink Iceberg lineage.#2147 @pawel-big-lebowskispark.openlineage.debugFacet=enabled needs to be set to include the facet. By default, the debug facet is disabled.#2165 @julwinAirflow: add some basic stats to the Airflow integration `#1845` @harels *Uses the statsd component that already exists in the Airflow codebase and wr
#1845 @harelsairflow.lineage.Table (if defined) #2138 @erikalfthanairflow.lineage.Table inlets/outlets to the OpenLineage Dataset.#2136 @erikalfthan#2118 @pawel-big-lebowskidelta and iceberg are not supported for Spark 3.5 at this time.#2139 @JDarDagran#2141 @JDarDagranairflow.providers.openlineage and adds more graceful logging to fix a corner case.#2142 @d-m-hprepareDatasetIdentifierFromDefaultTablePath method would override the scheme with the value of "file" when constructing a dataset identifier. It now uses the scheme of the CatalogTable's URI for this. Thank you @pawel-big-lebowski for the quick triage and suggested fix.Nothing published for this version
SQL: remove sqlparser dependency from iface-java and iface-py `#2090` @JDarDagran *Removes the dependency due to a breaking change in the latest relea…
ProcessingEngineRunFacet as part of the normal operation of the OpenLineageSparkEventListener #2089 @d-m-hProcessEngineRunFacet alongside the custom SparkVersionFacet (for now).
The SparkVersionFacet is deprecated and will be removed in a future release.spark.databricks.clusterUsageTags.clusterAllTags variable from databricks environment #2099 @Anirudh181001spark.databricks.clusterUsageTags.clusterAllTags to the list of environment variables captured from databricks.#2106 @tatianaDbtLocalArtifactProcessor in dbt projects that do not declare target-path.#2091 @harels#2044 @xli-1026apiKey if loading it from env variables @2029 @mobuchowskiapi_key to apiKey in create_token_provider.#2039 @pawel-big-lebowskirunning events after job completes #2075 @pawel-big-lebowski#2083 @pawel-big-lebowski#2076 @pawel-big-lebowski#2090 @JDarDagranNothing published for this version
Nothing published for this version
Flink: create Openlineage configuration based on Flink configuration `#2033` @pawel-big-lebowski *Flink configuration entries starting with openlineag
#2033 @pawel-big-lebowskiopenlineage.* are passed to the Openlineage client.#2004 @julienledem#2036 @pawel-big-lebowskispark.openlineage.jobName.appendDatasetName to false.
Unifies job names generated on the Databricks platform (using a dot job part separator instead of an underscore). The default behaviour can be altered with spark.openlineage.jobName.replaceDotWithUnderscore.#2057 @pawel-big-lebowski#2023 @mobuchowskiNone in TablesHierarchy to skip filtering on the schema level in the information schema query.KafkaSink #2042 @pentium3KafkaSinkVisitor by changing the KafkaSinkWrapper to catch schemas of type AvroSerializationSchema.CreateView events #1968#1987 @pawel-big-lebowskiCreateView nodes as root.MERGE INTO for delta tables identified by physical locations #2026 @pawel-big-lebowski#2035 @mobuchowskiadaptive_spark_plan in Databricks #2061 @algorithmy1adaptive_spark_plan from the excludedNodes in DatabricksEventFilter.Airflow: convert lineage from legacy `File` definition `#2006` @mobuchowski *Adds coverage for File entity definition to enhance backwards compatibili
File definition #2006 @mobuchowski
Adds coverage for File entity definition to enhance backwards compatibility.#1997 @JDarDagranDEBUG when extractor isn't found #2012 @kaxilWARNING to DEBUG when an extractor is not available.#2010 @mobuchowski#2025 @mobuchowskiHttpTransport by disabling automatic requests session reuse and not running SnowflakeExtractor again on job completion.#2001 @mars-lanHttpTransport in the case of a null URL.Flink: support Iceberg sinks `#1960` @pawel-big-lebowski *Detects output datasets when using an Iceberg table as a sink.*
#1960 @pawel-big-lebowskimerge into on delta tables #1958 @pawel-big-lebowskimerge into on Delta tables. Also refactors column-level lineage to deal with multiple Spark versions.merge into on Iceberg tables #1971 @pawel-big-lebowskimerge into on Iceberg tables.#1963 @juancappirest to the existing options of hive and hadoop in IcebergHandler.getDatasetIdentifier() to add support for Iceberg's RestCatalog.#1934 @mobuchowskiOPENLINEAGE_AIRFLOW_ENABLE_DIRECT_EXECUTION environment variable exists.openlineage-sql-java #1981 @davidjgoss#1975 @julienledem{ _deleted: true } object that can take the place of any job or dataset facet (but not run or input/output facets, which are valid only for a specific run).#1891 @AlexkuvaFileTransport and its configuration classes supporting append mode or write-new-file mode, which is especially useful when an object store does not support append mode, e.g. in the case of Databricks DBFS FUSE.#1999 @JDarDagranOPENLINEAGE_DISABLED to true if the provider is installed.config to config_class #1998 @mobuchowskiconfig class variable to config_class to avoid potential conflict with the config instance.#1959 @mobuchowski#1973 @pawel-big-lebowski#1968 @pawel-big-lebowskiProject node as root.openlineage.* logging levels via environment variables #1974 @JDarDagranOPENLINEAGE_{CLIENT/AIRFLOW/DBT}_LOGGING environment variables that can be set according to module logging levels and cleans up some logging calls in openlineage-airflow.Nothing published for this version
dbt: fix security vulnerabilities `#1945` @JDarDagran *Fixes vulnerabilities in the dbt integration and integration tests.*
#1947 @pawel-big-lebowski#1790 @pawel-big-lebowski#1928 @pawel-big-lebowski#1880 @pawel-big-lebowskiDatasetEvent and JobEvent types to the spec, along with support for the new types in the Python client.#1926 @mobuchowski#1950 @mobuchowski#1947 @pawel-big-lebowski#1790 @pawel-big-lebowski#1928 @pawel-big-lebowski#1880 @pawel-big-lebowskiDatasetEvent and JobEvent types to the spec, along with support for the new types in the Python client.#1926 @mobuchowski#1950 @mobuchowskidbt: add Databricks compatibility `#1829` @Ines70 *Enables launching OpenLineage with a Databricks profile.*
#1829 @Ines70#1829 Ines70#1913 gaborbernatschemaURL to run event #1917 gaborbernat
Adds the missing schemaURL to the client's RunState class.#1829 @Ines70#1913 @gaborbernatschemaURL to run event #1917 @gaborbernatschemaURL to the client's RunState class.Python client: deprecate `client.from_environment`, do not skip loading config `#1908` @mobuchowski *Deprecates the OpenLineage.from_environment metho…
client.from_environment, do not skip loading config #1908 @mobuchowskiOpenLineage.from_environment method and recommends using the constructor instead.client.from_environment, do not skip loading config #1908 @mobuchowskiOpenLineage.from_environment method and recommends using the constructor instead.Python client: add emission filtering mechanism and exact, regex filters `#1878` @mobuchowski *Adds configurable job-name filtering to the Python clie
#1878 @mobuchowski#1867 @pawel-big-lebowski[ and ] in Snowflake URIs #1883 @JDarDagran[ or ] were causing urllib.parse.urlparse to fail.#1878 @mobuchowski#1867 @pawel-big-lebowski[ and ] in Snowflake URIs #1883 @JDarDagran[ or ] were causing urllib.parse.urlparse to fail.Proxy: Fluentd proxy support (experimental) `#1757` @pawel-big-lebowski *Adds a Fluentd data collector as a proxy to buffer Openlineage events and sen
#1757 @pawel-big-lebowski#1856 @gaborbernat#1855 @pawel-big-lebowskispark.read.csv('file.csv').logicalPlan serialization issue on Databricks #1858 @pawel-big-lebowskispark_unknown facet by default to turn off serialization of logicalPlan.Spark: add Spark/Delta `merge into` support `#1823` @pawel-big-lebowski *Adds support for merge into queries.*
merge into support #1823 @pawel-big-lebowskimerge into queries.#1808 @nataliezeller1
Makes query handling more tolerant of variations in syntax and formatting.#1830 @pawel-big-lebowskiDeltaEventFilter class to filter events in cases where rewritten queries in adaptive Spark plans generate extra events.#1844 @Anirudh181001OpenLineageRunEventBuilder when it cast the Spark scheduler's ShuffleMapStage to boolean.#1840 @pawel-big-lebowski
Enriches Flink events so that missing eventTime, runId and job elements no longer produce errors.merge into support #1823 @pawel-big-lebowskimerge into queries.#1808 @nataliezeller1#1830 @pawel-big-lebowskiDeltaEventFilter class to filter events in cases where rewritten queries in adaptive Spark plans generate extra events.#1844 @Anirudh181001OpenLineageRunEventBuilder when it cast the Spark scheduler's ShuffleMapStage to boolean.#1840 @pawel-big-lebowski
Enriches Flink events so that missing eventTime, runId and job elements no longer produce errors.Support custom transport types `#1795` @nataliezeller1 *Adds a new interface, TransportBuilder, for creating custom transport types without having to
#1795 @nataliezeller1TransportBuilder, for creating custom transport types without having to modify core components of OpenLineage.#1418 @howardyoo#1796 @pawel-big-lebowski
It is a common scenario to write Spark output datasets with a location path ending with /year=2023/month=04. The Spark parameter spark.openlineage.dataset.removePath.pattern introduced here allows for removing certain elements from a path with a regex pattern.#1798 @pawel-big-lebowski
This mostly happens when getting table details on START event while the table is still not created.#1792 @pawel-big-lebowskiLogicalPlanSerializer to make use of non-shaded Jackson classes in order to serialize LogicalPlans. Note: class names are no longer serialized.#1801 @pawel-big-lebowski#1795 @nataliezeller1TransportBuilder, for creating custom transport types without having to modify core components of OpenLineage.#1418 @howardyoo#1796 @pawel-big-lebowski
It is a common scenario to write Spark output datasets with a location path ending with /year=2023/month=04. The Spark parameter spark.openlineage.dataset.removePath.pattern introduced here allows for removing certain elements from a path with a regex pattern.#1830 @pawel-big-lebowski
When spark plan is optimized, it is rewritten into adaptive plan which lead to duplicate Openlineage events: per normal and per adaptive plan. This changes filters the latter one.#1798 @pawel-big-lebowski
This mostly happens when getting table details on START event while the table is still not created.#1792 @pawel-big-lebowskiLogicalPlanSerializer to make use of non-shaded Jackson classes in order to serialize LogicalPlans. Note: class names are no longer serialized.#1801 @pawel-big-lebowskiSQL: parser improvements to support: `copy into`, `create stage`, `pivot` `#1742` @pawel-big-lebowski *Adds support for additional syntax available in
copy into, create stage, pivot #1742 @pawel-big-lebowski#1787 @JDarDagran#1788 @pawel-big-lebowskiCustomColumnLineageVisitor interface public to support custom column lineage.JobMetricsHolder #1786 @pawel-big-lebowskiput to fix a NPE occurring in JobMetricsHolder#1783 @pawel-big-lebowskiTableFactor::TableFunction to support queries containing table functions.#1785 @pawel-big-lebowskivisitor.rs.pass from several extract_on_complete methods #1771 @JDarDagran
Removes the code from three extractors.Spark: remove deprecated configs `#1711` by @tnazarew *Removes support for deprecated configs.*
#1717 by @tnazarewalter, truncate and drop statements #1695 by @pawel-big-lebowski#1727 by @JDarDagran.pyi public interface file for providing typing hints.#1718 by @tnazarewHttpTransport and the Spark integration.#1745 by @JDarDagranfrom_dict method to the Python client to support creating it from a dictionary.#1708 by @tnazarewOPENLINEAGE_DISABLED case-insensitive #1705 by @jedcunningham#1698 by @pawel-big-lebowskiUnsupportedDbtCommand when finding unsupported entry in args.which #1724 by @JDarDagrandbt-ol script to detect DBT commands in run_results.json only.#1717 @tnazarewalter, truncate and drop statements #1695 @pawel-big-lebowski#1727 @JDarDagran.pyi public interface file for providing typing hints.#1718 @tnazarewHttpTransport and the Spark integration.#1745 @JDarDagranfrom_dict method to the Python client to support creating it from a dictionary.#1708 @tnazarewOPENLINEAGE_DISABLED case-insensitive #1705 @jedcunningham#1698 @pawel-big-lebowskiUnsupportedDbtCommand when finding unsupported entry in args.which #1724 @JDarDagrandbt-ol script to detect DBT commands in run_results.json only.#1700 @pawel-big-lebowskiOneRowRelation and LocalRelation nodes.#1711 @tnazarewSpark: add support for a deprecated config `#1586` by @tnazarew *Maps the deprecated spark.openlineage.url to spark.openlineage.transport.url.*
DEBUG logging of events to transports #1633 by @mobuchowski
Ensures that the DEBUG loglevel on properly configured loggers will always log events, regardless of the chosen transport.CustomEnvironmentFacetBuilder class #1545 by New contributor @Anirudh181001AlterTableAddPartitionCommandVisitor and AlterTableSetLocationCommandVisitor #1629 by New contributor @nataliezeller1AlterTableAddPartitionCommand and AlterTableSetLocationCommand. The intended use case is a custom transport for the OpenMetadata lineage API.#1636 by @tnazarew#1664 by @mobuchowski#1631 by New contributor @rinzooltable.schema field or the operator default if the field is None.seed to the list of dbt-ol events #1649 by New contributor @pohek321dbt-ol test no longer fails when run against an event seed.#1634 by @pawel-big-lebowski#1586 by @tnazarewspark.openlineage.url to spark.openlineage.transport.url.#1590 by @tnazarew#1650 by @pawel-big-lebowskiLogicalRelation plans #1668 by @pawel-big-lebowskiselect col1, col2 from my_db.my_table that do not write output,
the Spark plan contained just a single node, which was wrongly treated as both
an input and output dataset.#1613 by @sekiknJobIdMapping and update macros to better support Airflow version 2+ #1645 by @JDarDagranOpenLineageAdapter's method to generate deterministic run UUIDs because using the JobIdMapping utility is incompatible with Airflow 2+.DEBUG logging of events to transports #1633 @mobuchowskiDEBUG loglevel on properly configured loggers will always log events, regardless of the chosen transport.CustomEnvironmentFacetBuilder class #1545 New contributor @Anirudh181001AlterTableAddPartitionCommandVisitor and AlterTableSetLocationCommandVisitor #1629 New contributor @nataliezeller1AlterTableAddPartitionCommand and AlterTableSetLocationCommand. The intended use case is a custom transport for the OpenMetadata lineage API.#1636 @tnazarew#1664 @mobuchowski#1631 New contributor @rinzooltable.schema field or the operator default if the field is None.seed to the list of dbt-ol events #1649 New contributor @pohek321dbt-ol test no longer fails when run against an event seed.#1634 @pawel-big-lebowski#1586 @tnazarewspark.openlineage.url to spark.openlineage.transport.url.#1590 @tnazarew#1650 @pawel-big-lebowskiLogicalRelation plans #1668 @pawel-big-lebowskiselect col1, col2 from my_db.my_table that do not write output,
the Spark plan contained just a single node, which was wrongly treated as both
an input and output dataset.#1613 @sekiknJobIdMapping and update macros to better support Airflow version 2+ #1645 @JDarDagranOpenLineageAdapter's method to generate deterministic run UUIDs because using the JobIdMapping utility is incompatible with Airflow 2+.Nothing published for this version
Airflow: add new extractor for FTPFileTransmitOperator `#1603` @sekikn *Adds a new extractor for this Airflow operator serving legacy systems.*
FTPFileTransmitOperator #1603 @sekikn#1601 @JDarDagran#1599 @mobuchowskiinclude_section argument for the Jinja render method to include only one profile if needed.compiled_code optional #1595 @JDarDagrancompiled_code optional for manifest > v7.Common: add explicit SQL dependency `#1532` @mobuchowski *Addresses 0.19.2 breaking change to GE integration by including SQL dependency explicitly.*
GCSToGCSOperator #1495 @sekikn#1522 @pawel-big-lebowski#1469 @fm100ruff instead of flake8, isort, etc., for linting and formatting #1526 @mobuchowskiruff package, which combines several linters and formatters into one fast binary.#1572 @JDarDagran#1532 @mobuchowskitqdm logging in dbt-ol #1549 @JDarDagrantqdm to show the correct number of iterations and adds START events for parent runs.#1493 @denimalpaca#1527 @mobuchowski#1556 @mobuchowskiKafkaTransport in the Java client and adds an exception if the required confluent-kafka module is missing from the Python client.#1507 @Varunvaruns9#1557 @mobuchowskiHadoopMapReduceWriteConfigUtil; makes the integration access BigQueryUtil and getTableId using reflection, which supports all BigQuery versions; makes logs provide the full serialized LogicalPlan on debug.Airflow: add Trino extractor https://github.com/OpenLineage/OpenLineage/pull/1288 @sekikn *Adds a Trino extractor to the Airflow integration.*
S3FileTransformOperator extractor https://github.com/OpenLineage/OpenLineage/pull/1450 @sekikn
Adds an S3FileTransformOperator extractor to the Airflow integration.NominalTimeRunFacet and OwnershipJobFacet https://github.com/OpenLineage/OpenLineage/pull/1410 @JDarDagran
Adds nominalEndTime and OwnershipJobFacet fields to the Airflow integration.ExtractionErrorRunFacet https://github.com/OpenLineage/OpenLineage/pull/1442 @mobuchowski
Adds a facet to the spec to reflect internal processing errors, especially failed or incomplete parsing of SQL jobs.collect_ignore, add flags to Pytest for cleaner output https://github.com/OpenLineage/OpenLineage/pull/1437 @JDarDagran
Removes the extractors directory from the ignored list, improving unit testing.#1288 @sekiknS3FileTransformOperator extractor #1450 @sekiknS3FileTransformOperator extractor to the Airflow integration.#1413 @JDarDagranNominalTimeRunFacet and OwnershipJobFacet #1410 @JDarDagrannominalEndTime and OwnershipJobFacet fields to the Airflow integration.#1417 @julienledem#1439 #1420 @fm100#1086 @wslulciucExtractionErrorRunFacet #1442 @mobuchowski#1432 #1461 @mobuchowski @StarostaGit#1383 @tnazarewcollect_ignore, add flags to Pytest for cleaner output #1437 @JDarDagranextractors directory from the ignored list, improving unit testing.Nothing published for this version
Airflow: support SQLExecuteQueryOperator `#1379` @JDarDagran *Changes the SQLExtractor and adds support for the dynamic assignment of extractors based
SQLExecuteQueryOperator #1379 @JDarDagranSQLExtractor and adds support for the dynamic assignment of extractors based on conn_type.SFTPOperator #1263 @sekikn#1136 @fhodaSagemakeProcessingOperator and SagemakerTransformOperator.#1166 @fhodaS3CopyObject in the Airflow integration.ExternalQueryRunFacet #1262 @howardyoo#1303 @merobi-hub#1330 @pawel-big-lebowski#1377 @mobuchowskieventTime field in Python client #1355 @pawel-big-lebowski
Validates the eventTime of a RunEvent within the client library.DbFsUtils constructor #1351 @wjohnsonDatabricksEnvironmentFacetBuilder and environment-properties facet by looking at the number of parameters in the DbFsUtils constructor to determine the runtime version.SQLExecuteQueryOperator #1379 @JDarDagranSQLExtractor and adds support for the dynamic assignment of extractors based on conn_type.SFTPOperator #1263 @sekikn#1136 @fhodaSagemakerProcessingOperator and SagemakerTransformOperator.#1166 @fhodaS3CopyObject in the Airflow integration.#1286 @mobuchowskiExternalQueryRunFacet #1262 @howardyoo#1303 @merobi-hub#1383 @tnazarew
#1330 @pawel-big-lebowski#1377 @mobuchowskieventTime field in Python client #1355 @pawel-big-lebowskieventTime of a RunEvent within the client library.DbFsUtils constructor #1351 @wjohnsonDatabricksEnvironmentFacetBuilder and environment-properties facet by looking at the number of parameters in the DbFsUtils constructor to determine the runtime version.…@JDarDagran *Uses deprecated resolver and constraints files provided by Airflow to avoid potential issues caused by pip's new resolver.*
task_instance argument to get_openlineage_facets_on_complete https://github.com/OpenLineage/OpenLineage/pull/1269 @JDarDagran--no-namespace-packages argument to the Mypy command and adjusts code to PEP 484..last_spec_commit_id.HttpTransport.Builder in favor of HttpConfig https://github.com/OpenLineage/OpenLineage/pull/1287 @collado-mikeBuilder in favor of HttpConfig only and replaces the existing Builder implementation by delegating to the HttpConfig.#1183 @pawel-big-lebowski#1200 @yogayang#1271 @pawel-big-lebowski#1233 @pawel-big-lebowski#1172 @StarostaGit @mobuchowski#1068 @wslulciuc#1249 @pawel-big-lebowskifacets definition with a list of available facets.#1300 @rossturk#1270 @rossturk#1295 @rossturkrelease.sh script.#1302 @JDarDagran#1238 @sekikntask_instance argument to get_openlineage_facets_on_complete #1269 @JDarDagrantask_instance argument to DefaultExtractor.#1290 @harels#1264 @JDarDagran--no-namespace-packages argument to the Mypy command and adjusts code to PEP 484.last_spec_commit_id, not just HEAD~1 #1298 @rossturk.last_spec_commit_id.#1287 @collado-mikeAirflow: add dag_run information to Airflow version run facet https://github.com/OpenLineage/OpenLineage/pull/1133 @fm100 *Adds the Airflow DAG run ID
dag_run information to Airflow version run facet #1133 @fm100taskInfo facet, making this additional information available to the integration.LoggingMixin to extractors #1149 @JDarDagranLoggingMixin class to the custom extractor to make the output consistent with general Airflow and OpenLineage logging settings.#1162 @mobuchowskiDefaultExtractor to support the default implementation of OpenLineage for external operators without the need for custom extractors.on_complete argument in DefaultExtractor #1188 @JDarDagranextract_on_complete.#1167 @StarostaGit @mobuchowskiget_connection_uri as extractor's classmethod #1169 @JDarDagranget_connection_uri method allowed for too many params, resulting in unnecessarily long URIs. This changes the logic to whitelisting per extractor.get_openlineage_facets_on_start/complete behavior #1201 @JDarDagranSqlJobFacet as a string #1143 @mobuchowskiquery from array to string to an fix error in the RedshiftSQLOperator.__extra__ case when filtering URI query params #1144 @JDarDagranconn.EXTRA_KEY in the get_connection_uri method to avoid exposing secrets in URIs via the __extra__ key.SQLCheckExtractors #1159 @denimalpaca_is_uppercase_names property to determine if the column should be upper cased in the SQLColumnCheckExtractor's _get_input_facets() method.#1180 @pawel-big-lebowskischemFacet is null.#1194 @denimalpaca#1128 @mobuchowskiAirflow: improve development experience https://github.com/OpenLineage/OpenLineage/pull/1101 @JDarDagran *Adds an interactive development environment
#1101 @JDarDagranoverwriteName to appName #1130 @tnazarewspark.openlineage.url and changes overwriteName to appName for clarity.#1116 @rossturk#1119 @mobuchowski#1056 @collado-mike#1111 @pawel-big-lebowski#1069 @pawel-big-lebowskiopenlineage.timeout is not provided.Init OpenLineageContext to DEBUG #1064 @varuntestaz#1090 @TheSpeeddingOPENLINEAGE_URL to be set #1107 @mobuchowskiOPENLINEAGE_URL in the dbt integration.#1126 @mobuchowski#1131 @mobuchowskiFix Spark integration issues including error when no openlineage.timeout https://github.com/OpenLineage/OpenLineage/pull/1069 @pawel-big-lebowski *Ope
openlineage.timeout https://github.com/OpenLineage/OpenLineage/pull/1069 @pawel-big-lebowski
OpenlineageSparkListener was failing when no openlineage.timeout was provided.openlineage.timeout #1069 @pawel-big-lebowskiOpenlineageSparkListener was failing when no openlineage.timeout was provided.openlineage.timeout #1069 @pawel-big-lebowskiOpenlineageSparkListener was failing when no openlineage.timeout was provided.Support ABFSS and Hadoop Logical Relation in Column-level lineage https://github.com/OpenLineage/OpenLineage/pull/1008 @wjohnson _Introduces an extrac
#1008 @wjohnsonextractDatasetIdentifier that uses similar logic to InsertIntoHadoopFsRelationVisitor to pull out the path on the HDFS compliant file system; tested on ABFSS and DBFS (Databricks FileSystem) to prove that lineage could be extracted using non-SQL commands.#939 @hmoazamKustoRelationVisitor to support lineage for Azure Kusto's Spark connector.#1020 @julienledem#935 @pawel-big-lebowskiSymlinkDatasetFacet in generated OpenLineage events.#1051 @mobuchowskicompiled_sql field to compiled_code to support Python models). Does not provide support for dbt's Python models.#1009 @mzareba382#1066 @mobuchowski#1050 @tnazarew#1049 @collado-mike#1065 @pawel-big-lebowskiRename all parentRun occurrences to parent from Airflow integration https://github.com/OpenLineage/OpenLineage/pull/1037 @fm100
parentRun occurrences to parent in Airflow integration 1037 @fm100parentRun property name to parent in the Airflow integration to match the spec.on_running event 1028 @JDarDagranon_running hook, which was changing the TaskInstance object along with the task attribute.parentRun occurrences to parent in Airflow integration 1037 @fm100parentRun property name to parent in the Airflow integration to match the spec.on_running event 1028 @JDarDagranon_running hook, which was changing the TaskInstance object along with the task attribute.Add BigQuery check support `#960` @denimalpaca
#960 @denimalpacaRUNNING EventType in spec and Python client #972 @mzareba382#974 @JDarDagran#995 @howardyooSymlinksDatasetFacet to spec #936 @pawel-big-lebowski#983 @hmoazam#1015 @conorbev#996 @julienledemRUNNING EventType in Flink integration for currently running jobs #985 @mzareba382#1018 @fm100#1025 @collado-mike#960 @denimalpacaBigQueryColumnCheckOperator and BigQueryTableCheckOperator.)RUNNING EventType in spec and Python client #972 @mzareba382RUNNING event state in the OpenLineage spec to indicate a running task and adds a RUNNING event type in the Python API.#974 @JDarDagraninformation_schema tables.)#995 @howardyooHttpLineageStream to forward a given OpenLineage event to any HTTP endpoint.SymlinksDatasetFacet to spec #936 @pawel-big-lebowskiSymlinksDatasetFacet, to support the storing of alternative dataset names.#983 @hmoazamRelationHandler, to support Spark data sources that do not have TableCatalog, Identifier, or TableProperties set, as is the case with the Azure Cosmos DB Spark connector.#1015 @conorbevairflow.lineage.entities.Table, which was then converted to an OpenLineage Dataset.)#996 @julienledemRUNNING EventType in Flink integration for currently running jobs #985 @mzareba382RUNNING event type in the Flink integration, changing events sent by Flink jobs from OTHER to this new type.#1018 @fm100to_json_encodable function in the Airflow integration to make task objects JSON-encodable.#1025 @collado-mikeAdd Spark 3.3.0 support `#950` @pawel-big-lebowski
#950 @pawel-big-lebowski#951 @mobuchowski#922 @pawel-big-lebowski#897 @mobuchowski#717 @denimalpaca#930 @JDarDagran#927 @tnazarew#905 @pawel-big-lebowski#914 @fenil25#917 @pawel-big-lebowski#942 @pawel-big-lebowski#950 @pawel-big-lebowski#951 @mobuchowski#922 @pawel-big-lebowski#897 @mobuchowski#717 @denimalpaca#930 @JDarDagran#927 @tnazarew#905 @pawel-big-lebowski#914 @fenil25#917 @pawel-big-lebowski#942 @pawel-big-lebowskiHTTP option to override timeout and properly close connections in openlineage-java lib. https://github.com/OpenLineage/OpenLineage/pull/909 @mobuchows
openlineage-java lib. https://github.com/OpenLineage/OpenLineage/pull/909 @mobuchowskiSqlExtractor to Airflow integration https://github.com/OpenLineage/OpenLineage/pull/907 @JDarDagranTaskListener in the Airflow integration https://github.com/OpenLineage/OpenLineage/pull/870 @mobuchowskiopenlineage-java lib. https://github.com/OpenLineage/OpenLineage/pull/855 @collado-mikeiceberg in Spark integration https://github.com/OpenLineage/OpenLineage/pull/856 @wslulciucopenlineage-java lib. #909 @mobuchowski#906 @JDarDagranSqlExtractor to Airflow integration #907 @JDarDagran#898 @merobi-hub#882 @denimalpacaTaskListener in the Airflow integration #870 @mobuchowskiopenlineage-java lib. #855 @collado-mike#891 @pawel-big-lebowskiiceberg in Spark integration #856 @wslulciucopenlineage-java lib. #909 @mobuchowski#906 @JDarDagranSqlExtractor to Airflow integration #907 @JDarDagran#898 @merobi-hub#882 @denimalpacaTaskListener in the Airflow integration #870 @mobuchowskiopenlineage-java lib. #855 @collado-mike#891 @pawel-big-lebowskiiceberg in Spark integration #856 @wslulciucAdd static code anlalysis tool mypy to run in CI for against all python modules (#802) @howardyoo
SaveIntoDataSourceCommandVisitor to extract schema from LocalRelaiton and LogicalRdd in spark integration (#794) @pawel-big-lebowskiInMemoryRelationInputDatasetBuilder for InMemory datasets to Spark integration (#818) @pawel-big-lebowskiSnowflakeOperatorAsync extractor support to Airflow integration #869 @denimalpacaFunctionRegistry.class serialization in Spark integration (#828) @mobuchowskirust-based SQL parser by default in Airflow integration (#835) @mobuchowskipytest and integration tests for Airflow integration (#851, #858) @denimalpacasqlalchemy lib for Great Expectations integration (#826) @pawel-big-lebowskiorg.apache.spark.sql.catalyst.plans.logical.CreateV2Table in Spark integration (#866) @pawel-big-lebowskiSpark: Column-level lineage introduced for Spark integration (https://github.com/OpenLineage/OpenLineage/pull/698, https://github.com/OpenLineage/Open
Added
Fixed
#802) @howardyooSaveIntoDataSourceCommandVisitor to extract schema from LocalRelaiton and LogicalRdd in spark integration (#794) @pawel-big-lebowskiInMemoryRelationInputDatasetBuilder for InMemory datasets to Spark integration (#818) @pawel-big-lebowski#755 @merobi-hubSnowflakeOperatorAsync extractor support to Airflow integration #869 @merobi-hub#889) @howardyooFunctionRegistry.class serialization in Spark integration (#828) @mobuchowskirust-based SQL parser by default in Airflow integration (#835) @mobuchowskipytest and integration tests for Airflow integration (#851,#858) @denimalpaca#881) @collado-mike#834, #890) @tnazarew @mobuchowskisqlalchemy lib for Great Expectations integration (#826) @pawel-big-lebowskiorg.apache.spark.sql.catalyst.plans.logical.CreateV2Table in Spark integration (#866) @pawel-big-lebowski#867,#874) @pawel-big-lebowskiopenlineage-airflow now supports getting credentials from Airflows secrets backend (https://github.com/OpenLineage/OpenLineage/pull/723) @mobuchowski
Added
Fixed
openlineage-airflow now supports getting credentials from Airflows secrets backend (#723) @mobuchowskiopenlineage-spark now supports Azure Databricks Credential Passthrough (#595) @wjohnsonopenlineage-spark detects datasets wrapped by ExternalRDDs (#746) @collado-mikePostgresOperator fails to retrieve host and conn during extraction (#705) @sekiknopenlineage-airflow now supports getting credentials from Airflows secrets backend (#723) @mobuchowskiopenlineage-spark now supports Azure Databricks Credential Passthrough (#595) @wjohnsonopenlineage-spark detects datasets wrapped by ExternalRDDs (#746) @collado-mikePostgresOperator fails to retrieve host and conn during extraction (#705) @sekiknAirflow integration uses new TaskInstance listener API for Airflow 2.3+ (#508) @mobuchowski
Added
openlineage-java lib (#480) @wslulciuc, @mobuchowskiFixed
openlineage-java lib (#480) @wslulciuc, @mobuchowskiPython implements Transport interface - HTTP and Kafka transports are available (https://github.com/OpenLineage/OpenLineage/pull/530) @mobuchowski
Added
Fixed
CI: add integration tests for Airflow's SnowflakeOperator and dbt-snowflake @mobuchowski
Added
Fixed
Catch possible failures when emitting events and log them @mobuchowski
Extract source code of PythonOperator code similar to SQL facet @mobuchowski
UnknownOperatorAttributeRunFacet to Airflow integration to record operators that don't produce lineage @collado-mikeUnknownOperatorAttributeRunFacet to Airflow integration to record operators that don't produce lineage @collado-mikeProxy backend example using Kafka @wslulciuc
Kafka @wslulciucKafka @wslulciucSupport for dbt-spark adapter @mobuchowski
backend to proxy OpenLineage events to one or more event streams 🎉 @mandy-chessell @wslulciucbackend to proxy OpenLineage events to one or more event streams 🎉 @mandy-chessell @wslulciucSeparated tests between Spark 2 & 3 @pawel-big-lebowski
fix import in spark3 visitor @mobuchowski
Spark3 support @OleksandrDvornik / @collado-mike
Add dbt v3 manifest support @mobuchowski
v3 manifest support @mobuchowskiv3 manifest support @mobuchowskiImplement OpenLineageValidationAction for Great Expectations @collado-mike
Default --project-dir argument to current directory in dbt-ol script @mobuchowski
--project-dir argument to current directory in dbt-ol script @mobuchowski--project-dir argument to current directory in dbt-ol script @mobuchowskiParse dbt command line arguments when invoking dbt-ol @mobuchowski. For example:
Parse dbt command line arguments when invoking dbt-ol @mobuchowski. For example:
$ dbt-ol run --project-dir path/to/dir
Set UnknownFacet for spark (captures metadata about unvisited nodes from spark plan not yet supported) @OleksandrDvornik
model from dbt job name @mobuchowskiopenlineage.spark.* to io.openlineage.spark.* @OleksandrDvornikParse dbt command line arguments when invoking dbt-ol @mobuchowski. For example:
$ dbt-ol run --project-dir path/to/dir
Set UnknownFacet for spark (captures metadata about unvisited nodes from spark plan not yet supported) @OleksandrDvornik
model from dbt job name @mobuchowskiopenlineage.spark.* to io.openlineage.spark.* @OleksandrDvornikYour coding agent can read these notes before it upgrades. Set up the MCP server →