NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2332 most downloaded on PyPI
OpenLineage common python library for integrations
Last release 1 months ago
01 Sep 2026
Ships on a steady schedule
a new release about every 4 weeks
Most releases are documented
notes for 50 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
123 releases · first in 2021
One column per quarter.
Java integrations: Upgrade Jackson to 2.18.8 for CVE-2026-54512 and CVE-2026-54513 #4765 @sulikismaylovv Updates bundled Jackson dependencies across J…
#4759 @tnazarew#4798 @tnazarewGcsTransport configuration instead of always using defaults.#4778 @mobuchowski#4812 @karthikchundi-commitsoracle://host:port dataset namespaces.#4747 @mobuchowski#4465 @kchledowski#4875 @mobuchowski#4861 @fmorillo7694#4878 @MSDehghan#4782 @mobuchowski with @tnazarew#4797 @tnazarew#4848 @mishrasangeeta87LOAD DATA INPATH as an input dataset alongside the target table output.#4804 @mobuchowski#4859 @mattfaltyn#4902 @dolfinusAsyncHttpTransport to the maintained httpx2 drop-in replacement and updates related configuration and tests.#4819 @mobuchowski#4612 @matveeysv#4838 @MSDehghan#4824 @mattfaltynAsyncHttpTransport usable after wait_for_completion() so later events are still delivered.#4882 @mattfaltyn#4900 @mattfaltynSTART event completes concurrently.#4509 @hcthakur2004#4808 @chuenchen309#4809 @chuenchen309dataQualityAssertions, matching the run-results processor.#4777 @kacpermuda#4846 @chuenchen309#4733 @chuenchen309bytes, restoring metrics and assertions.#4811 @chuenchen309OpenLineageValidationAction to construct on Great Expectations 1.x while retaining compatibility with versions that require the argument.#4765 @sulikismaylovv#4853 @Poojitha-R-Rao#4726 @zerafachris#4894 @mobuchowski#4884 @MSDehghan#4850 @mishrasangeeta87COPY INTO plan variants.#4849 @mishrasangeeta87#4815 @mishrasangeeta87#4835 @mishrasangeeta87#4898 @MSDehghan#4772 @MSDehghan#4896 @JDarDagranSparkCachedTableCatalog.#4779 @mobuchowskiCatalogManager from a class to an interface.#4775 @mattfaltyn#4791 @mattfaltyn#4865 @mattfaltynFILTER expressions.#4763 @mattfaltynARRAY subqueries in input lineage.#4785 @mattfaltynHAVING expressions.#4783 @mattfaltynJOIN ... ON conditions.#4855 @mattfaltynMERGE ON predicates.#4767 @mattfaltynUPDATE SET assignment expressions.#4788 @mattfaltynVALUES rows.#4852 @mattfaltynPIVOT wraps a derived table.Add generic OpenLineage context configuration for propagating parent-run info #4682 @kacpermuda Adds one JSON context payload for parent-run info, rep
#4682 @kacpermuda#4716 @himakolavennu#4705 @himakolavennu#4744 @wangxiaojing#4698 @maitraymukeshkumarmodi-aiml#4680 @kacpermuda#4711 @HeroCC#4743 @zerafachris#4731 @chuenchen309#4728 @mattfaltyn#4694 @mobuchowski#4729 @zerafachris#4722 @chuenchen309#4696 @hcthakur2004#4730 @chuenchen309--openlineage-dbt-job-name=value no longer breaks the dbt run.#4725 @chuenchen309~/.dbt/ as intended.#4732 @chuenchen309#4724 @chuenchen309#4723 @chuenchen309#4752 @mattfaltynDbt: Capture dbt model meta/config values #4653 @mobuchowski Capture per-model config values (e.g. materialized , access , owner , group ) and the use
#4653 @mobuchowskiconfig values (e.g. materialized, access, owner, group) and the user-defined meta map from the dbt manifest into a new dbt_model dataset facet, attached in both the legacy/local and structured-logs dbt processors.#3674 @arturowczarekUCSingleCatalog Spark catalog so output dataset facets are produced when using OSS Unity Catalog, which cannot reuse the existing DeltaHandler implementation.#4592 @jakub-moravec#4676 @mishrasangeeta87df.write.mode("append").insertInto(table) did not emit an output dataset, by extracting the target table directly from the AppendData logical plan, aligning append-mode handling with overwrite-mode.#4660 @mishrasangeeta87unity-catalog namespace instead of one derived from the underlying storage path, giving lineage consumers stable, catalog-qualified identifiers.#4652 @tnazarewio.openlineage.client.transports.gcs.shaded.x) instead of to a single shared shaded package, fixing Google library version conflicts when combining gcs and gcplineage transports in a composite transport.#4661 @arturowczarekabfs:// and wasb:// URIs the same as their TLS variants (abfss/wasbs) when deriving the object-storage namespace, since they point at the same underlying files and only differ in transport.#4666 @mobuchowskiCompositeTransport emit fails, so the real cause is visible instead of being lost.#4624 @hcthakur2004#4627 @hcthakur2004get_from_nullable_chain, keeping it reusable after lookup.#4679 @hcthakur2004#4586 @hcthakur2004#4656 @hcthakur2004file:// URIs before opening file transport paths so events can be appended through a file:// log path.#4665 @arturowczarekiceberg-spark SparkCatalog class rather than iceberg-core's Catalog, so the Iceberg handler no longer activates in environments where iceberg-core is on the classpath without iceberg-spark.Client/Java: Add timeout-aware run event emission support #4613 @jsingh-yelp Add a timeout-aware emit overload to the Java client and Transport API, a
#4613 @jsingh-yelpemit overload to the Java client and Transport API, allowing callers to provide a bounded wait when emitting RunEvents; Kafka transport uses the timeout to wait for producer acknowledgement before returning.#4596 @jsingh-yelpOpenLineageDetachedJobStatusChangedListener that initialises the Flink job ID from the REST API, bypassing the missing JobCreatedEvent on the JobManager side.#4615 @wangxiaojingJobStatus.CANCELED to EventType.ABORT in Flink 2 job status handling; previously all non-FINISHED terminal statuses were mapped to FAIL, causing user-canceled jobs to appear as failures in downstream OpenLineage consumers.#4635 @kacpermudaenvironmentVariables facet already set by the producer instead of replacing it; event-supplied values take precedence and a warning is logged on conflict.#4633 @fm100IcebergScanReport input dataset facet in CTAS/RTAS queries by reading the report directly from scanReportSupplier when available, rather than relying on OpenLineageMetricsReporter which is registered after the scan has already occurred.Java: Add CassandraJdbcExtractor #4610 @matveeysv Add CassandraJdbcExtractor to parse Cassandra JDBC URLs according to the driver specification, enabl
#4610 @matveeysvCassandraJdbcExtractor to parse Cassandra JDBC URLs according to the driver specification, enabling lineage tracking for Cassandra databases via JDBC.#4599 @fm100DbtStructuredLogsProcessor when using --consume-structured-logs option by attaching the column lineage facet to the output dataset on node finished events.#4591 @fm100externalQueryId in the externalQuery run facet when using the dbt-bigquery adapter; also adds externalQuery run facet support when using --consume-structured-logs.#4598 @codelixirdataproc_job_attempt_timestamp tag prefix when reading the job ID from Yarn tags in the GCP Dataproc facet, preventing the attempt timestamp from being reported as the job ID on retried jobs.#4611 @tnazarewV2SessionCatalogHandler used by DatasetBuilders.#4602 @mrpalash-amzdbtable path) by applying stripQuotes() normalization on both sides of identifier comparisons in ColumnLevelLineageBuilder and SqlCollector.Spark: Add iceberg s3 tables catalog support #4558 @mobuchowski Fix incorrect Glue table symlink attachment when using S3 Tables as the Iceberg catalo
#4558 @mobuchowski#4574 @tnazarew#4546 @adnanhemani#4536 @W-Ely#4552 @kacpermuda#4557 @mobuchowskiBaseCatalogTypeHandler.getIdentifier() to return Optional<DatasetIdentifier> instead of a nullable value, making the contract for Iceberg catalog type handlers more explicit and null-safe.#4587 @tnazarew#4561 @hcthakur2004refs/pull/<number>/head format (used in GitHub Actions pull_request events) in addition to the existing refs/pull/<number>/merge form.#4571 @mobuchowskiSnowflakeCatalogTypeHandler.getIdentifier() to return Optional<DatasetIdentifier> as required by BaseCatalogTypeHandler, correcting a type mismatch introduced when the Snowflake Iceberg REST catalog handler was first added.#4556 @dbathie-wtgsqlserver URLs to MsSqlDialect so SQL Server queries using bracketed identifiers ([schema].[table]) are parsed correctly during lineage extraction.Client/Go: Rewrite code generator `#4501` @tnazarew *Replace the quicktype-based text-manipulation generator with a structured pipeline that parses sp
#4501 @tnazarew
Replace the quicktype-based text-manipulation generator with a structured pipeline that parses spec files via go-jsonschema, resolves references, and renders facet classes — enabling extension of generated classes (e.g. byool resources). Also renames Run → RunWrapper and RunInfo → Run for consistency with Job and Dataset.#4542 @mobuchowski
Fix extract_adapter_type lookup for the Microsoft Fabric adapter by aligning the Adapter enum name with the dbt adapter type (fabric), while keeping the fabric-warehouse OpenLineage namespace unchanged.#4501 @tnazarew
Replace the quicktype-based text-manipulation generator with a structured pipeline that parses spec files via go-jsonschema, resolves references, and renders facet classes — enabling extension of generated classes (e.g. byool resources). Also renames Run → RunWrapper and RunInfo → Run for consistency with Job and Dataset.#4542 @mobuchowski
Fix extract_adapter_type lookup for the Microsoft Fabric adapter by aligning the Adapter enum name with the dbt adapter type (fabric), while keeping the fabric-warehouse OpenLineage namespace unchanged.Spark: Bump httpclient5 to 5.6.1 for CVE-2026-40542 `#4521` @Poojitha-R-Rao *Upgrade httpclient5 dependency to 5.6.1 to address security vulnerability…
#4485 @harels
Register dbt-fabric in the supported Adapter enum and add fabric:// namespace extraction, preventing NotImplementedError on Fabric profiles and enabling lineage tracking for Microsoft Fabric datasets.#4495 @och5351
Add ClickHouseJdbcExtractor to enable lineage tracking for ClickHouse JDBC connections, supporting both jdbc:clickhouse:// and jdbc:ch:// URL schemes with optional protocol prefix removal.#4506 @jsingh-yelp
Add ConfigFacetVisitor and SchemaFacetVisitor for non-Table API Flink 2 datasets, enabling schema and configuration metadata extraction for DataStream API connectors beyond the Table API.#4457 @kacpermuda
Extend DataQualityAssertionsDatasetFacet with additional fields (actual value, expected value, severity) aligned with TestRunFacet, bumping the facet spec to 1-1-0.#4500 @mobuchowski
Populate actual and expected fields on DataQualityAssertionsDatasetFacet and TestRunFacet with the dbt test failure count and error threshold, enabling consumers to distinguish passing tests from failures and understand configured tolerances.#4497 @och5351
Fix MySQL JDBC extractor to not prepend the URL database to already-qualified table names, preventing invalid 3-level identifiers like mydb.schema.table1 (MySQL treats DATABASE and SCHEMA as synonyms).#4520 @mobuchowski
Fix dbt-ol to recognize retry as a valid dbt command so lineage is captured when re-running failed nodes.#4523 @mobuchowski
Fix aggregate test events in the legacy consume_local_artifacts path that emitted COMPLETE (success) even when the embedded DataQualityAssertions facet showed error-severity assertion failures.#4472 @mobuchowski
Fix stale manifest.json from a prior invocation being loaded before dbt finishes its parse phase by subscribing to the ArtifactWritten structured log event (dbt ≥ 1.9) and lazy-loading for older versions.#4489 @hcthakur2004
Fix AttributeError: 'NoneType' object has no attribute 'startswith' crash when node_info.unique_id is absent from the event by adding an early None guard.#4515 @mobuchowski
Fix crash when the dbt owner metadata field is a list rather than a string, handling both single-owner and multi-owner configurations gracefully.#4499 @ichirotakami
Fix KeyError in extract_dataset_data() for dbt-fusion manifests where source nodes omit the description key by using .get() with an empty string default.#4522 @mobuchowski
Fix root=None in the parent facet for per-node events emitted via the legacy consume_local_artifacts path by propagating root_parent_* fields to dbt_run_metadata.#4503 @hcthakur2004
Fix file reading in the dbt provider to explicitly specify UTF-8 encoding, preventing UnicodeDecodeError on systems where the default locale encoding is not UTF-8.#4498 @1fanwang
Fix NullPointerException when processing Hive queries with 3+ way UNION ALL, INTERSECT, or EXCEPT set operations by recursively descending through intermediate QBExpr nodes instead of assuming leaf-only structure.#4439 @NETIZEN-11
Fix ClassCastException when Iceberg returns SparkChangelogTable instead of SparkTable by replacing the unsafe direct cast with instanceof checks and graceful fallback handling.#4521 @Poojitha-R-Rao
Upgrade httpclient5 dependency to 5.6.1 to address security vulnerability CVE-2026-40542.#4505 @creazyfrog
Fix incorrect use of Hadoop's 3-argument Path constructor that produced malformed S3 URIs like s3://bucket/prefix://database.db./table_name, which caused IllegalArgumentException and silently dropped all lineage events.#4427 @Poojitha-R-Rao
Upgrade AWS SDK version to address security vulnerability CVE-2026-33871 (netty-codec).Client/Go: Add Go client `#4358` @tnazarew *Add a new OpenLineage Go client with code generation from the spec, HTTP and GCP Lineage transport impleme
#4358 @tnazarew
Add a new OpenLineage Go client with code generation from the spec, HTTP and GCP Lineage transport implementations, and CI integration.#4398 @himakolavennu
Add sourceCodeLocation job facet emission for dbt events, with repo URL configurable via --openlineage-repo-url or OPENLINEAGE_REPO_URL, falling back to git remote get-url origin autodetection. URLs are normalized to host/org/repo format.#4449 @mobuchowski
Emit TestRunFacet from both the regular dbt processor path and the structured logs path for test result nodes, including outcome (pass/fail/warn/error), severity, and expected/actual values.#4436 @himakolavennu
Add optional pullRequestNumber field to SourceCodeLocationJobFacet (spec bumped to 1-1-0) with auto-detection from CI environment variables (GITHUB_REF for GitHub Actions, CI_MERGE_REQUEST_IID for GitLab CI).#4448 @mobuchowski
Add TestRunFacet JSON Schema for recording test outcomes (pass/fail/warn/error), severity, expected vs actual values, and test message, along with generated Python client code and facet registration.#4403 @himakolavennu
Make git autodetect for sourceCodeLocation opt-in by default (disabled=True) to avoid surprising subprocess calls during event emission, and reduce subprocess count from 4 to 2 by consolidating git calls.#4391 @mobuchowski
Remove support for Python 3.9, which reached end-of-life in November 2024. The minimum supported Python version is now 3.10.#4451 @bahram-cdt
Fix broken SPI directory path (META-INF.services/ → META-INF/services/), wrong class name in SPI file (missing .kinesis sub-package), and NullPointerException when properties config is omitted in the Kinesis transport.#4400 @mobuchowski
Fix the structured logs processor to correctly resolve parent datasets for dbt singular tests using parent_map, so test result assertions are now attached to the correct datasets as DataQualityAssertionsDatasetFacet.#4390 @orthoxerox
Fix column-level lineage for InMemoryRelation when it contains non-unique unqualified column names by using qualifiedName instead of name for correct column mapping.#4384 @usamakunwar
Fix dataset symlinks when using Glue catalog for Iceberg by detecting Glue-based catalogs via Spark configuration and generating correct ARNs even when not using the Hive metastore.#4366 @mobuchowski
Fix JDBC column lineage to skip extractInternalInputs when SQL-based column lineage is already available, preventing alias names from being incorrectly included as input fields alongside the original column names.Flink: Add DatasetConfigFacet support for Flink native listener `#4368` @jsingh-yelp *Add support for emitting DatasetConfigFacet in the Flink native
#4368 @jsingh-yelp
Add support for emitting DatasetConfigFacet in the Flink native listener, enabling configuration tracking for datasets processed by Flink jobs.#3747 @dolfinus
Introduce the HierarchyDatasetFacet to provide structured representation of dataset hierarchy levels (database, schema, table, etc.) without relying on dataset name parsing, enabling consistent handling across different database systems with varying hierarchy depths.#4383 @mobuchowski
Improve performance by switching from expensive semanticHash() calls to identity-based tracking using IdentityHashMap, eliminating hot path in large jobs during plan traversal.#4376 @mobuchowski
Optimize findDependentInputs by replacing LinkedList with HashSet for visited node tracking, improving lookup performance from O(n) to O(1).#4372 @mobuchowski
Fix handling of dbt singular tests when processing structured logs to ensure proper test result tracking and lineage extraction.#4357 @tnazarew
Fix the condition for adding project_id to catalog properties, ensuring proper BigLake catalog detection and handling.dbt: Attach ExtractionErrorRunFacet on metadata extraction failures `#4349` @harels *Attach ExtractionErrorRunFacet to run events when @handle_keyerro
#4349 @harels
Attach ExtractionErrorRunFacet to run events when @handle_keyerror-decorated extraction methods fail, making previously invisible extraction errors visible to downstream consumers instead of silently emitting incomplete events._get_model_node #4348 @harels
Fix exception type mismatch in _get_model_node() by raising KeyError instead of RuntimeError, allowing the @handle_keyerror decorator to catch it and return None gracefully when a node_id is not found in the manifest..get() for optional project version retrieval #4345 @zagoodman
Fix crash when version key is absent from dbt_project.yml, which became optional in dbt 1.5, by using .get() instead of direct key access.Client: Add JWT authentication support `#4313` @jakub-moravec *Add JWT authenticator for Java and Python clients, enabling token-based authentication
#4313 @jakub-moravec
Add JWT authenticator for Java and Python clients, enabling token-based authentication without requiring a custom authenticator implementation.#4283 @kchledowski
Enable extraction of input dataset symlinks from DataSourceRDD, providing richer lineage information for RDD-based Iceberg operations.#4329 @kchledowski
Disable column-level lineage extraction for LogicalRDD plans to prevent incorrect lineage caused by lost schema and transformation context.#4331 @kchledowski
Disable unreliable input schema extraction from LogicalRDD and instead extract schemas from Iceberg table metadata when reading via DataSourceRDD.#4285 @LegendPawel-Marut
Align schema definitions for dbt-run-run-facet and dbt-version-run-facet to fix validation inconsistencies.#4320 @ah12068
Handle missing profiles_dir key in run_results.json gracefully, falling back to default profile directory resolution.#4298 @gaurav-atlan
Fix the --target-path CLI argument not being parsed and passed to artifact processors, causing the default target path to always be used.#4312 @mobuchowski
Fix false Flink 2.x detection when modern V2-based connectors are used with Flink 1.x by using JobStatusChangedListenerFactory for version detection.#4282 @Lukas-Riedel
Send the Content-Encoding header when request body compression is enabled in the Java client, consistent with the Python client behavior.#4315 @mobuchowski
Add fast environment detection check to skip Databricks-specific event filtering on non-Databricks platforms, reducing overhead.#4311 @mobuchowski
Fix NullPointerException when processing Iceberg datasets with AWS Glue catalog by safely handling null Glue ARN values.#4316 @mobuchowski
Fix ClassCastException on DROP TABLE commands in Databricks Runtime 14.2+ by handling ResolvedIdentifier alongside ResolvedTable.build runtime dependency #4344 @mobuchowski
Remove the build package from runtime dependencies as it is only needed at build time and is already handled by the build system configuration.#4340 @tstrilka
Fix incorrect parent job name in ParentRunFacet for child events (SQL_JOB, RDD_JOB) on AWS Glue, where the raw spark.app.name was used instead of the resolved application name from platform-specific name resolvers.testAppendWithRDDTransformations and testAppendWithRDDProcessing on Java 8 + Spark 3.5 where the Iceberg vendor module is not compiled due to Iceberg 1.7 requiring Java 11+.dbt: Extract test severity from dbt tests `#4258` @mobuchowski *Add support for extracting and reporting severity information from dbt tests in OpenLi
#4258 @mobuchowski
Add support for extracting and reporting severity information from dbt tests in OpenLineage events.#4257 @mobuchowski
Extend the DataQualityAssertionsDatasetFacet with a severity field to indicate the importance level of data quality assertions.#4263 @kchledowski
Enable lineage extraction from DataSourceRDD when reading Iceberg data sources, supporting mixed RDD and DataFrame operations common in AWS Glue environments.#4262 @mobuchowski
Fix IndexError when running 'dbt-ol send-events' command without additional arguments by properly checking args length.#4264 @mobuchowski
Fix incorrect "namespace" value in BigQuery column-level lineage to match the format used by regular BigQuery dataset collection.#4268 @mobuchowski
Fix Delta detection failing when multiple Spark extensions are configured (comma-separated), enabling proper event filtering in environments like Azure Fabric or when using Gluten.#4243 @kchledowski
Fix service provider interface conflicts caused by package relocation not updating META-INF/services configuration files.Airflow: Remove Airflow integration from OpenLineage repository `#4212` @kacpermuda *The deprecated Airflow integration has been removed from the Open…
#4218 @RohithKayathi
Enable posting lineage events to DataZone domains in different regions from where data transformation jobs run.#4118 @kchledowski
Add new configuration option spark.openlineage.filter.rddEventsDisabled to selectively disable OpenLineage event emission for RDD operations while keeping SQL-based operations enabled.#4124 @kchledowski
Add schema and column-level lineage support for Snowflake datasets when using the Spark-Snowflake connector.#4215 @wslulciuc
Add support to override the application runID via the property spark.openlineage.applicationRunId.#4182 @jakub-moravec
Add a new facet to capture input parameters supplied to a job at the time of execution, enabling reproducibility, debugging, and richer lineage context.#3768 @tnazarew
Update GCP Lineage transport to use new version of the producer library with fixed dependency shading.#4220 @mobuchowski
Improve error messages to indicate which transport failed to create.#4207 @mobuchowski
Fix classloader conflicts with BigQuery connector by gating DEBUG toJSON() logging behind an additional flag and logging exceptions.#4197 @dolfinus
Fix type annotation for .with_additional_properties() method to correctly accept keyword arguments.#4192 @kchledowski
Fix BigQuery symlink namespace incorrectly having ".db" suffix in RUNNING and COMPLETE events by avoiding mutation of the Identifier object.#4229 @lawofcycles
Add fallback mechanism to retrieve AWS region from EC2 Instance Metadata Service when environment variables are unavailable in YARN cluster mode.#4222 @kchledowski
Fix missing inputs and column-level lineage when writing from AWS DynamicFrame by treating NewHadoopRDD as file-like.#4228 @RohithKayathi
Apply spark.openlineage.dataset.removePath.pattern to input field names in ColumnLineageFacet, and fix hashCode/equals methods to include additionalProperties.#4212 @kacpermuda
The deprecated Airflow integration has been removed from the OpenLineage repository.Spec: Add arbitrary extra info to JobDependency in JobDependenciesRunFacet `#4189` @kacpermuda *Add support for arbitrary extra information in JobDepe
#4189 @kacpermuda
Add support for arbitrary extra information in JobDependency within JobDependenciesRunFacet.#4185 @kacpermuda
Add debug mode support to file transport for better troubleshooting.#4160 @harels
Add support for capturing dbt model owner information from meta.owner in OpenLineage events.#4151 @mobuchowski
Add DbtNodeJobFacet to provide additional dbt node information in job facets.#4161 @tnazarew
Add default name support to Hive catalog facet in Spark integration.#4134 @pawel-big-lebowski
Add support for fetching input statistics for single input RDD jobs.#4153 @kchledowski
Migrate SQL parser from fork to upstream version 0.59 for better maintenance and compatibility.#4178 @mobuchowski
Reduce aggressiveness of UUID normalization in Spark integration.#4186 @kacpermuda
Improve logging output in Python client.#4165 @dolfinus
Fix relation size calculation to ensure values are within reasonable bounds.#4154 @mobuchowski
Add missing job facet schema to specification.Python: re-add missing __version__ variables in top of releaseable modules `#4135` @mobuchowski *Fixes breaking change in version 1.40.0.*
#4135 @mobuchowski
Fixes breaking change in version 1.40.0.#4135 @mobuchowski
Fixes breaking change in version 1.40.0.Java: Fix CVE in commons-lang3 `#4084` @mandalbalmukund *Upgrade commons-lang3 version to fix CVE security vulnerability.*
#4109 @jakub-moravec
Add a standardized batch API endpoint to OpenLineage specification for handling multiple events in a single request.#4116 @mobuchowski
Add ordinal_position field to track the position of fields in schema (1-indexed).#4112 @kacpermuda
Introduce JobDependenciesRunFacet to track dependencies between jobs.#4103 @jakub-moravec
Add support for temporary datasets to enable job-to-job lineage tracking.#4075 @luke-hoffman1
Add fallback configuration for BigQuery project ID in Metastore integration.#4123 @kacpermuda
Include examples in Python generated classes for better documentation.#4077 @dolfinus
Add support for parsing jTDS JDBC URL format in Java client.#4066 @tnazarew
Add ParentRunFacet to Hive integration for tracking parent-child run relationships.#4097 @tnazarew
Add support for tracking LOAD and IMPORT operations in Hive.#4085 @tnazarew
Add support for tracking EXPORT operations in Hive.#4079 @tnazarew
Add START event emission support to Hive integration.#4121 @usamakunwar
Fix Spark dataset facet builders for input datasets.#4114 @kchledowski
Fix job name trimming logic in Spark integration.#4113 @pawel-big-lebowski
Fix putAll operation failing on immutable maps.#4108 @pawel-big-lebowski
Fix multiple issues with RDD job handling in Spark.#4102 @kchledowski
Fix JDBC dbtable parsing to support any FROM clauses.#4083 @pawel-big-lebowski
Fix Spark connector configuration for Databricks environments.#4099 @mobuchowski
Catch NoClassDefFoundError when buggy implementations exist on classpath.#4104 @mobuchowski
Fix Snowflake identifier parsing to handle quoted identifiers correctly.#4105 @mobuchowski
Strip quotes from Snowflake account names for proper handling.#4092 @fm100
Fix facet property names from snake_case to camelCase for consistency.#4111 @kacpermuda
Fix Python client facet generator after moving to UV build system.#4093 @antonlin1
Fix retry configuration default merge with user-defined config in HTTP transports.#4084 @mandalbalmukund
Upgrade commons-lang3 version to fix CVE security vulnerability.#4126 @dolfinus
Ensure START and STOP events share the same runId in Hive integration.Spark: Normalize dataset names with configurable trimmers `#3996` @pawel-big-lebowski *Add configurable dataset name normalization with support for da
#3996 @pawel-big-lebowski
Add configurable dataset name normalization with support for date patterns, key-value pairs, and S3 location detection to enable proper dataset subsetting.#4057 @kchledowski
Add missing input symlink facets for Databricks Unity Catalog tables.#4058 @kchledowski
Refactor column-level lineage dependency collector tests for better organization and maintainability.#4069 @fm100
Fix typo in IcebergCommitReportOutputDatasetFacet property name.#4061 @pawel-big-lebowski
Fix dataset name trimming for column-level lineage inputs.#4062 @kacpermuda
Remove unnecessary numpy import from Python client.#3844 @kacpermuda
Remove Dagster integration from the repository.Spec: Add subset dataset facets to spec `#4008` @pawel-big-lebowski *Add subset dataset facets to OpenLineage specification for representing dataset r
#4008 @pawel-big-lebowski
Add subset dataset facets to OpenLineage specification for representing dataset relationships.#3978 @heron--
Allow attaching dataset quality information outside of InputDatasetFacet.#4018 @tnazarew
Add support for Spark structured streaming microbatch source write operations.#4016 @ddebowczyk92
Add catalog properties support to Spark integration for better catalog metadata tracking.#4039 @ddebowczyk92
Enhance BigQuery integration with GCP project ID and location in catalog properties.#3972 @kchledowski
Add support for tracking COALESCE transformations in Spark jobs.#3982 @ddebowczyk92
Add catalog facet support for vanilla Hive table operations.#4013 @pawel-big-lebowski
Output statistics now available in complete events for better observability.#3977 @pawel-big-lebowski
Add output statistics tracking for Spark RDD-based jobs.#4050 @pawel-big-lebowski
Improve generated model classes with proper equals and hashcode implementations.#4022 @mobuchowski
Add support for capturing dbt tags in OpenLineage events.#4017 @mobuchowski
Add dbt Cloud account ID tracking to dbt run facets.#3987 @mobuchowski
Enhance DbtRunRunFacet with additional metadata for better observability.#4006 @ddebowczyk92
Add native Google Cloud Platform Lineage transport for Python client.#3983 @JDarDagran
Add fsspec filesystem support to FileTransport for broader filesystem compatibility.#3980 @kacpermuda
Automatically add OpenLineage client version as default tag in events.#3986 @gabrysiaolsz
Add GCP Cloud Composer environment metadata facets to Airflow integration.#4055 @mobuchowski
Use dbt model aliases when generating dataset names for more accurate lineage.#4029 @EugeneYushin
Serialize OpenLineage events to JSON format for improved debug logging.#4030 @EugeneYushin
Properly respect user-overridden application names in event emission.#4003 @kchledowski
Refactor column-level lineage expression dependency collector for better maintainability.#3994 @JDarDagran
Enhance logging for Iceberg input statistics collection.#3985 @pawel-big-lebowski
Optimize S3 operations by limiting external getFileStatus calls for large object sets.#3964 @kchledowski
Refactor TransformationInfo into shared Java client for cross-integration reuse.#4026 @dolfinus
Enhance logging capabilities in asynchronous HTTP transport.#4000 @JDarDagran
Support Python type aliases in client code generation.#3997 @JDarDagran
Improve code generation to properly handle nearly identical class definitions.#4014 @dolfinus
Fail fast with clear errors when custom token providers fail to load.#4015 @dolfinus
Improve error visibility by not silencing import errors in transport factory.#3968 @kacpermuda
Update import paths to use versioned facet and event modules.#4012 @JDarDagran
Improve thread pool management in Java client utilities.#3965 @JDarDagran
Migrate from pre-commit to prek for pre-commit hook management.#4053 @jsjasonseba
Fix incorrect Glue catalog detection due to always attempting ARN resolution.#4052 @kchledowski
Fix column-level lineage failures on Spark runtimes without spark-hive package.#4031 @kchledowski
Fix missing input datasets and column-level lineage for CreateDataSourceTableAsSelect and CreateHiveTableAsSelect commands.#4044 @EugeneYushin
Fix BigQuery intermediate job filtering by using bucket configuration.#4002 @MaciejGajewski
Add additional exception handling for TypeNotPresentException in Spark 3.0.2.#4034 @JDarDagran
Correct license field specification in Python package metadata.#4045 @kacpermuda
Support both naming conventions for API key configuration parameter.#4037 @EugeneYushin
Fix build issue causing empty sources JAR files to be generated.Python: Add Datadog transport with configurable async routing `#3950` @mobuchowski *Add Datadog transport with intelligent routing between sync/async
#3950 @mobuchowski#3860 @orthoxerox#3933 @mobuchowski#3923 @kyungryun#3904 @pawel-big-lebowski#3956 @dolfinus#3925 @SalvadorRomo#3946 @pawel-big-lebowski#3949 @ddebowczyk92#3934 @pawel-big-lebowski#3930 @yunchipang#3947 @ddebowczyk92#3953 @jroachgolf84#3943 @JDarDagran#3860 @orthoxerox#3933 @mobuchowski#3923 @kyungryun#3950 @mobuchowski#3904 @pawel-big-lebowski#3956 @dolfinus#3925 @SalvadorRomo#3946 @pawel-big-lebowski#3949 @ddebowczyk92#3934 @pawel-big-lebowski#3930 @yunchipang#3947 @ddebowczyk92#3953 @jroachgolf84#3943 @JDarDagranSpark: support Delta 4.0 and cover it with tests on Spark 4.0. `#3877` @pawel-big-lebowski *Fix failing tests for Spark 4.0. Make delta integration te
#3877 @pawel-big-lebowski#3914 @pawel-big-lebowski#3921 @pawel-big-lebowski#3890 @jroachgolf84#3918 @mobuchowski#3816 @ddebowczyk92#3907 @pawel-big-lebowski#3851 @dolfinus#3895 @dolfinus#3899 @Shadi#3869 @mobuchowski#3902 @pawel-big-lebowskiSqlExecutionRDDVisitor and LogicalRDDVisitor classes to avoid memory leak.#3909 @pawel-big-lebowski#3908 @pawel-big-lebowski#3911 @pawel-big-lebowski#3915 @fetta#3905 @kacpermuda#3916 @mobuchowski#3894 @mobuchowski#3889 @pawel-big-lebowski#3897 @dolfinus#3901 @kacpermudadbt: Fix deprecated configs `#3859` @kacpermuda *Replaces deprecated dbt configurations with current alternatives*
#3848 @dolfinus#3850 @ddebowczyk92#3880 @pawel-big-lebowskispark.openlineage.disabled entry to disable OpenLineage integration through Spark config parameters#3779 @pawel-big-lebowskibuildDatasetsTimePercentage and facetsBuildingTimePercentage in docs for more details#3812 @mobuchowski#3764 @dolfinus#3829 @kacpermuda#3789 @dolfinus#3863 @dolfinus#3819 @mobuchowski#3826 @ddebowczyk92#3775 @ddebowczyk92#3858 @dolfinus#3856 @ddebowczyk92#3811 @pawel-big-lebowski#3881 @dolfinus#3843 @dolfinus#3857 @dolfinus#3855 @dolfinus#3838 @dolfinus#3841 @dolfinus#3839 @dolfinus#3817 @mobuchowski#3854 @dolfinus#3799 @pan-siekierski#3796 @dolfinus#3836 @mobuchowski#3800 @dolfinus#3849 @dolfinus#3793 @mobuchowski#3859 @kacpermuda.db suffix in database/namespace location name for BigQueryMetastoreCatalog #3874 @ddebowczyk92#3835 @ddebowczyk92#3871 @pawel-big-lebowski#3861 @pawel-big-lebowski#3832 @pawel-big-lebowski#3853 @kacpermuda#3825 @mobuchowski#3806 @mobuchowski#3830 @mobuchowski#3887 @kacpermuda#3814 @mobuchowskiHive: Integration added. `#3555` @tnazarew with @ddebowczyk92, @jphalip *Added OpenLineage Hive integration*
#3555 @tnazarew with @ddebowczyk92, @jphalip#3691 @pawel-big-lebowskiUnionRdd and NewHadoopRDD, which makes dynamic frames docker based test passing.#3781 @dolfinus#3777 @dolfinus#3786 @dolfinus#3717 @tnazarew#3715 @pawel-big-lebowski#3760 @ddebowczyk92#3738 @dolfinus#3739 @dolfinus#3725 @dolfinus#3744 @dolfinus#3726 @dolfinus#3763 @pawel-big-lebowski#3748 @dolfinus#3669 @kacpermuda#3731 @dolfinus#3713 @dolfinus#3754 @dolfinus#3709 @dolfinus#3766 @mvitale#3751 @ddebowczyk92#3785 @ddebowczyk92#3776 @kacpermuda#3680 @mobuchowski#3773 @dolfinus#3722 @pawel-big-lebowski#3749 @mobuchowski#3724 @dolfinus#3728 @JDarDagran#3762 @ngorchakovaPython: add TransformTransport for Python client `#3697` @kacpermuda *Introduces the TransformTransport class for event transformations.*
#3697 @kacpermuda#3659 @mobuchowski#3695 @mobuchowski#3685 @shinabel#3706 @kacpermuda#3700 @dolfinus#3686 @luke-hoffman1#3696 @dolfinus#3707 @dolfinus#3688 @dolfinus#3681 @pawel-big-lebowskiAvro: support schema facet for Avro datasets `#3650` @pawel-big-lebowski *This PR adds support for schema facets in Avro datasets.*
#3650 @pawel-big-lebowski#3672 @martinovm#3682 @MassyB#3683 @MassyB#3676 @pawel-big-lebowski#3667 @luke-hoffman1#3673 @ddebowczyk92#3663 @ddebowczyk92Flink: enhance JDBC extractors with additional types `#3652` @HuangZhenQiu
#3652 @HuangZhenQiu#3648 @mobuchowski#3637 @dolfinus#3664 @mudakacper#3661 @mobuchowski#3644 @tnazarew#3645 @pawel-big-lebowski#3641 @ddebowczyk92#3652 @HuangZhenQiu#3648 @mobuchowski#3637 @dolfinus#3664 @kacpermuda#3661 @mobuchowski#3644 @tnazarew#3645 @pawel-big-lebowski#3641 @ddebowczyk92Java: added support for LDAP connection strings `#3612` @luke-hoffman1
#3612 @luke-hoffman1#3625 @mehdimld#3605 @kuba0221#3572 @mobuchowski#3622 @dolfinus#3624 @dolfinus#3621 @dolfinus#3582 @ddebowczyk92#3607 @ddebowczyk92#3586 @dolfinus#3594 @pawel-big-lebowski
Register simple micrometer registry when no other registry configured.#3612 @luke-hoffman1#3625 @mehdimld#3605 @kuba0221#3572 @mobuchowski#3622 @dolfinus#3624 @dolfinus#3621 @dolfinus#3582 @ddebowczyk92#3607 @ddebowczyk92#3586 @dolfinus#3594 @pawel-big-lebowski
Register simple micrometer registry when no other registry configured.#3580 @MassyB
Buffered writing could cause dbt to write broken log lines. This PR handles that case.#3576 @pawel-big-lebowski
This PR fixes case where event v2 job and run events did not get user tags.#3583 @ddebowczyk92
This PR fixes potential NullPointerException in SaveIntoDataSourceCommandVisitor.#3612 @luke-hoffman1#3625 @mehdimld#3605 @kuba0221#3572 @mobuchowski#3622 @dolfinus#3624 @dolfinus#3621 @dolfinus#3582 @ddebowczyk92#3607 @ddebowczyk92#3586 @dolfinus#3594 @pawel-big-lebowski
Register simple micrometer registry when no other registry configured.dbt: read structured log file incrementally `#3580` @MassyB *Buffered writing could cause dbt to write broken log lines. This PR handles that case.*
#3580 @MassyB
Buffered writing could cause dbt to write broken log lines. This PR handles that case.#3576 @pawel-big-lebowski
This PR fixes case where event v2 job and run events did not get user tags.#3583 @ddebowczyk92
This PR fixes potential NullPointerException in SaveIntoDataSourceCommandVisitor.#3580 @MassyB
Buffered writing could cause dbt to write broken log lines. This PR handles that case.#3576 @pawel-big-lebowski
This PR fixes case where event v2 job and run events did not get user tags.#3583 @ddebowczyk92
This PR fixes potential NullPointerException in SaveIntoDataSourceCommandVisitor.Java: remove deprecated configs: 'disabledFacets' and 'timeout'. `#3522` @pawel-big-lebowski *Configs have been replaced with: 'facets.facet-name.disa…
#3531 @pawel-big-lebowski
Similar to Flink 1 integration, Flink 2 integration will emit CheckpointFacet.#3528 @pawel-big-lebowski
Table and field comments are available within generated OL events.#3522 @pawel-big-lebowski
Configs have been replaced with: 'facets.facet-name.disabled=true' and 'timeoutInMillis'.#3538 @sakjung
Fixes support for metrics in Iceberg SparkSessionCatalog.#3515 @pawel-big-lebowski
Fixes support for metrics in Iceberg RESTCatalog.#3535 @MassyB
Fixes race condition that was happening when using structured logs output.#3545 @MassyB
Skipped nodes no longed cause exceptions.#3550 @pawel-big-lebowski
InputStatistics facet for Iceberg datasets no longer produces incorrect stats.#3548 @pawel-big-lebowski
Subquery alias no longer is duplicating the inputs.#3552 @mobuchowski
Fixes catching InaccessibleMethodException in Java 17 within SparkExtensionVisitor.#3471 @leogodin217
User-supplied tags will allow the client to inject new tags or override tags provided by the integrations for jobs and runs.#3531 @pawel-big-lebowski
Similar to Flink 1 integration, Flink 2 integration will emit CheckpointFacet.#3528 @pawel-big-lebowski
Table and field comments are available within generated OL events.#3522 @pawel-big-lebowski
Configs have been replaced with: 'facets.<name of disabled facet>.disabled=true' and 'timeoutInMillis'.#3538 @sakjung
Fixes support for metrics in Iceberg SparkSessionCatalog.#3515 @pawel-big-lebowski
Fixes support for metrics in Iceberg RESTCatalog.#3535 @MassyB
Fixes race condition that was happening when using structured logs output.#3545 @MassyB
Skipped nodes no longed cause exceptions.#3550 @pawel-big-lebowski
InputStatistics facet for Iceberg datasets no longer produces incorrect stats.#3548 @pawel-big-lebowski
Subquery alias no longer is duplicating the inputs.#3552 @mobuchowski
Fixes catching InaccessibleMethodException in Java 17 within SparkExtensionVisitor.Python: allow adding user-supplied tags facets from config `#3471` @leogodin217 *User-supplied tags will allow the client to inject new tags or overri
#3471 @leogodin217
User-supplied tags will allow the client to inject new tags or override tags provided by the integrations for jobs and runs.#3493 @mobuchowski
Enabled parsing tags from config in Java client and Spark conf.#3487 @pawel-big-lebowski
Properly name case where TaskQueueCircuitBreaker allows a configurable blocking time after submitting a callable.#3503 @pawel-big-lebowski
Native Flink integration is now isolated within circuit breaker call.#3483 @ddebowczyk92
ServiceLoader should not fail to load OpenLineageExtensionProvider implementations in certain configurations.#3486 @MarquisC
Null Flink Job Manager address will default to localhost#3488 @MassyB
Handle case for tests on sources which don't have the attached_node defined in the manifest.Java: enable specifying custom SSL context `#3444` @pawel-big-lebowski *Enable providing configuration for SSL context within HTTP transport.*
#3444 @pawel-big-lebowski
Enable providing configuration for SSL context within HTTP transport.#3442 @pawel-big-lebowski
Spark integration filters OpenLineage events for specific plan node classes. This can be now extended with extra config entries: allowedSparkNodes and deniedSparkNodes. See Spark Configuration documentation for more details.#3437 @aritrabandyo
This circuit breaker that executes task on a queue backed threadpool, gives up tasks if the queue is full, and keeps track of rejected tasks.#3429 @whitleykeith
This allows Trino integration to emit proper events containing Trino datasets.#3430 @ssanthanam185
Adds coverage for AlterTableRecoverPartitionsCommandVisitor, RefreshTableCommandVisitor, RepairTableCommandVisitor.#3425 @d-m-h
This presents no functional change to the listener, however it will allow for improved initialisation of the listener in the future.#3435 @pawel-big-lebowski
In case of unsupported classes, warn logs without a stacktrace should be produced.#3443 @d-m-h
This is an initial refactor to a larger code base change that will see the removal of direct access of the QueryExecution object. It has no functional change on the way the integration behaves.COMPLETE events. #3434 @pawel-big-lebowski
*Send input datasets in COMPLETE events while making sure version facet is attached on START only.#3432 @MassyB
Fixes incorrect structure of ParentRunFacet.Flink: Experimental version for flink native lineage listener. `#3099` @pawel-big-lebowski *New flink listener to extract lineage through native Flink
#3099 @pawel-big-lebowski
New flink listener to extract lineage through native Flink interfaces. Supports Flink SQL. Requires Flink 2.0.#3362 @MassyB
New option for dbt integration now can handle test and build commands too.#3379 @ssanthanam185
Events emitted from RDDExecutionContext now include custom facets that get loaded as part of InternalHandlerFactory.#3390 @JDarDagran
DatasetTypeDatasetFacet allows explicit declaration of type of the resulting dataset.#3391 @JDarDagran
Adds with_additonal_properties method that allows to create modified instance of facet with additional properties.#3379 @ssanthanam185
Events emitted from RDDExecutionContext now include custom facets that get loaded as part of InternalHandlerFactory.#3403 @ssanthanam185
Those events shouldn't be filtered outside Databricks/Delta ecosystem.#3368 @ddebowczyk92
Fixes ClassNotFoundException issue when using the openlineage-spark integration alongside a Spark connector that implements the spark-extension-interfaces due to class loader conflicts.#3368 @cisenbe
SQL parser won't error on Snowflake's LATERAL keyword.#3311 @dsaxton-1password
dbt integration won't fail when looking at tests on seeds.#3379 @ssanthanam185
Spark integration now correctly handles complex jobs that have cycles and nested RDD trees.json file extension. #3404 @kacpermuda
When append=False, the json file extension wasn't properly added before.dbt: Consume dbt structured logs and report progress in real time. `#3314` @MassyB *If --consume-structured-logs flag is set, dbt integration will con
#3314 @MassyB
If --consume-structured-logs flag is set, dbt integration will consume dbt structured logs and report execution progress in real time.transform transport to allow event modification. #3301 @pawel-big-lebowski
New transport type allows to modify the event based on the specified transformer class.#3305[#3305] @pawel-big-lebowski
Emit events in parallel for composite transport. Running in parallel is a default behaviour continueOnFailure set to true. Default value of continueOnFailure got changed from false to true.ScanReport and CommitReport in OpenLineage events when dealing with Iceberg tables. #3256 @pawel-big-lebowski
Collects additional Iceberg metrics for datasets read or written through the library. Visit Dataset Metrics docs for more details.#3280 @mobuchowski
Adds support for duckdb adapter for dbt integration.DatasetFactory to support Dataset creation. #3207 @pawel-big-lebowski
Adds DatasetFactory to support Dataset creation. This class is used to create Dataset instances for DatasetFactory.#3285 @pawel-big-lebowski
GCS path now has correctly stripped leading slash…marks some public developers' API methods as deprecated.*
#3264 @mayurmadnani
Dbt integration now uses SQL parser to add information about collected column-level lineage.#3240#3263 @pawel-big-lebowski
Fix issues related to existing output statistics collection mechanism and fetch input statistics. Output statistics contain now amount of files written, bytes size as well as records written. Input statistics contain bytes size and number of files read, while record count is collected only for DataSourceV2 sources.#3238 @pawel-big-lebowski#3244 @tnazarew
Excludes META-INF/*TransportBuilder to avoid version conflictsDatasetFactory #3207 @pawel-big-lebowskiDatasetFactory class, marks some public developers' API methods as deprecated.Spark: Add Dataproc run facet to include jobType property `#3167` @codelixir *Updates the GCP Dataproc run facet to include jobType property*
#3167 @codelixir#3186 @JDarDagran#3221 @JDarDagran#3142 @arturowczarek#3205 @mobuchowski#3219 @arturowczarek#3148 @codelixir#3215 @mobuchowski#3217 @arturowczarek#3208 @MassyB#3141 @pawel-leszczynski#3167 @codelixir#3186 @JDarDagran#3221 @JDarDagran#3142 @arturowczarek#3205 @mobuchowski#3219 @arturowczarek#3148 @codelixir#3215 @mobuchowski#3217 @arturowczarek#3208 @MassyB#3141 @pawel-leszczynskiNothing published for this version
Nothing published for this version
Java: added CompositeTransport `#3039` @JDarDagran *This allows user to specify multiple targets to which OpenLineage events will be emitted.*
#3039 @JDarDagran#3062 @Imbruced#3043 @ddebowczyk92#3077 @ddebowczyk92#3129 @arturowczarek#3094 @JDarDagran#3114 @JDarDagran#3116 @JDarDagran#3097 #3098 @arturowczarek#3122 @ddebowczyk92#3054 @JDarDagran#2962 @Imbruced#3068 @jonathanlbt1#3107 @Imbruced#3095 @ImbrucedSQL: add support for `USE` statement with different syntaxes `#2944` @kacpermuda *Adjusts our Context so that it can use the new support for this stat
USE statement with different syntaxes #2944 @kacpermuda#3044 @arturowczarek#3007 #3023 @pawel-big-lebowskiwebsite directory.SingleQuotedString in Identifier() #3035 @kacpermudaIDENTIFIER function instead of treating it like table name #2999 @kacpermuda#2918 @Imbruced#3020 @arturowczarekSpec: add GCP Dataproc facet `#2987` @tnazarew *Registers the Google Cloud Platform Dataproc run facet.*
#2987 @tnazarew
Registers the Google Cloud Platform Dataproc run facet.#2983 @kacpermuda#3001 @arturowczarek#2986 @Imbruced#2990 @arturowczarek#2984 @arturowczarektable/.#2937 @d-m-h
Previously, reading Iceberg datasets outside the configured Spark catalog prevented the datasets from being present in the inputs property of the RunEvent.Nothing published for this version
Nothing published for this version
Python: add `CompositeTransport` `#2925` @JDarDagran *Adds a CompositeTransport that can accept other transport configs to instantiate transports and
CompositeTransport #2925 @JDarDagranCompositeTransport that can accept other transport configs to instantiate transports and use them to emit events.#2828 @pawel-big-lebowski#2854 @pawel-big-lebowskiWindow #2901 @tnazarewWindow-type nodes of a logical plan.#2913 @ImbrucedQueryExecution to get the SQL query used from the SQL field with a BFS algorithm.#2887 @Imbruced#2912 @arturowczarekFacetConfig accept the disabled flag for any facet instead of passing them as a list.#2906 @ImbrucedDatasetIdentifier from extension LineageNode #2900 @ddebowczyk92LogicalRelation has a grandChild node that implements the LineageRelation interface.BaseRelation #2893 @ddebowczyk92DatasetIdentifier is now extracted from the underlying node of LogicalRelation.#2889 @jonathanlbt1marquez-web service to docker-compose.yml.#2880 @jonathanlbt1#2877 @jonathanlbt1for each batch method #2868 @Imbruced#2943 @arturowczarekIcebergHandler support Glue catalog tables and create the symlink using the code from PathUtils.#2917 @arturowczarek#2892 @arturowczarekLogicalPlanSerializer now returns <failed-to-serialize-logical-plan> for failed serialization instead of an empty string.#2883 @arturowczarekCustomCollectorsUtils for improved readability.foreach batch mode #2868 @ImbrucedDatasetIdentifier from SaveIntoDataSourceCommandVisitor options #2934 @ddebowczyk92DatasetIdentifier from command's options instead of relying on p.createRelation(sqlContext, command.options()), which is a heavy operation for JdbcRelationProvider.Airflow: add `log_url` to `AirflowRunFacet` `#2852` @dolfinus *Adds taskinstance's log_url field to AirflowRunFacet.*
log_url to AirflowRunFacet #2852 @dolfinuslog_url field to AirflowRunFacet.Generate #2856 @tnazarewGenerate-type nodes of a logical plan (e.g., explode operations).DerbyJdbcExtractor #2869 @dolfinusJdbcExtractor implementation for Derby database. As this is a file-based DBMS, its Dataset namespace is file and name is an absolute path to a database file.#2859 @pawel-big-lebowskiJarVerifier plugin to ensure all compiled classes have a bytecode version of Java 8 or lower.#2851 @d-m-h#2865 @kacpermudaColumnLevelLineageBuilder #2850 @tnazarewStreams dependency in ColumnLevelLineageBuilder causing a ClassNotFoundException.#2863 @dolfinus#2855 @ddebowczyk92PlanUtils3 so Dataset identifier information based on a Table's properties is also retrieved during the construction of column-level lineage.#2861 @arturowczarekspark.app.name was autogenerated by Glue and uses the Glue job name in such cases. Also, each job name provisioning strategy is now extracted to a separate provider.Spark: configurable integration test `#2755` @pawel-big-lebowski *Provides command line tool capable of running Spark integration tests that can be cr
#2755 @pawel-big-lebowski#2809 #2837 @ddebowczyk92#2743 @pawel-big-lebowski#2789 @tnazarewColumnLineageDatasetFacet creation.InsertIntoHadoopFsRelationCommand #2794 @dolfinusINSERT INTO command for tables created with USING $fileFormat syntax, like USING orc.PostgresJdbcExtractor #2806 @dolfinus
Adds the default 5432 port to Postgres namespaces.TeradataJdbcExtractor #2826 @dolfinus
Converts JDBC URLs like jdbc:teradata/host/DBS_PORT=1024,DATABASE=somedb to datasets with namespace teradata://host:1024 and name somedb.table.MySqlJdbcExtractor #2825 @dolfinus
Handles different formats of MySQL JDBC URL, and produces datasets with consistent namespaces, like mysql://host:port.OracleJdbcExtractor #2824 @dolfinus
Handles simple Oracle JDBC URLs, like oracle:thin:@//host:port/serviceName and oracle:thin@host:port:sid, and converts each to a dataset with namespace oracle://host:port and name sid.schema.table or serviceName.schema.table.#2822 @pawel-big-lebowski#2838 @pawel-big-lebowskiUNION queries.#2756 #2801 @Sheeri
Updates the customLineage facet test for the new syntax created in #2756.spark.sql.warehouse.dir as table namespace #2767 @dolfinusspark.sql.warehouse.dir or hive.metastore.warehouse.dir as table namespace, instead of duplicating the table's location.JdbcExtractors #2830 @dolfinus#2807 @Akash2351
Fixes Glue symlinks with config parsing for Glue catalogid.#2800 @dolfinus
Fixes the DBFS namespace format.#2766 @dolfinus#2797 @dolfinusfile:/some/path/database.table uses file:/some/path/database/table. For dataset TABLE symlink, uses warehouse location instead of database location.#2827 @pawel-big-lebowski
Fixes an error caused by a recent upgrade of Spark versions that did not break existing tests.JdbcLocation #2831 @dolfinustransformationType and transformationDescription are marked as deprecated.*
#2720 @pawel-big-lebowski#2758 @tnazarew#2643 @codelixirGCPRunFacetBuilder and GCPJobFacetBuilder to report additional facets when running on Google Cloud Platform.#2773 @dolfinus#2698 @pawel-big-lebowskishadowJar content and prevent reported issues. These are hard to prevent currently and require manual verification of manually unpacked jar content.#2756 @tnazarewColumnLineageDatasetFacet. transformationType and transformationDescription are marked as deprecated.#2729 @harels#2740 @ngorchakovalocalServerId option from Kafka config #2738 @dolfinuslocalServerId from Kafka config, deprecated since 1.13.0.Transport.emit(String) #2737 @dolfinusTransport.emit(String) support, deprecated since 1.13.0.spark-interfaces-scala module #2781 @ddebowczyk92spark-interfaces-scala interfaces with new ones decoupled from the Scala binary version. Allows for improved integration in environments where one cannot guarantee the same version of openlineage-java.#2769 @algorithmy1namespace.name as Avro complex field type #2763 @dolfinusnamespace.name is now used as Avro "type" of complex fields (record, enum, fixed).#2776 @kacpermudadrop table for Spark 3.4 and above #2745 @pawel-big-lebowski @savannavalgi#2782 @dolfinuss3a:// and s3n:// schemes to s3://.#2761 @dolfinus$SPARK_CONF_DIR/hive-site.xml.#2749 @pawel-big-lebowskicur.getDependencies() is not null before adding dependencies.OpenLineageRunEventBuilder #2754 @pawel-big-lebowskiOpenLineageRunEventBuilder::buildRun.historyUrl format #2741 @dolfinushistoryUrl format in spark_applicationDetails.#2753 @mobuchowskiselect * from test_orders as test_orders are now parsed properly.Nothing published for this version
Spark: add `jobType` facet to Spark application events `#2719` @dolfinus *Adds jobType facet to runEvents emitted by SparkListenerApplicationStart.*
jobType facet to Spark application events #2719 @dolfinus
Adds jobType facet to runEvents emitted by SparkListenerApplicationStart.jobType facet to Spark application events #2719 @dolfinusjobType facet to runEvents emitted by SparkListenerApplicationStart.#2720 @pawel-big-lebowskiHostListNamespaceResolver, PatternNamespaceResolver, PatternMatchingGroupNamespaceResolver or custom implementation loaded with ServiceLoader. Feature is useful to resolve hostnames into cluster identifiers.#2735 @JDarDagran#2727 @mobuchowskijobType facet to Spark application events #2719 @dolfinusjobType facet to runEvents emitted by SparkListenerApplicationStart.#2735 @JDarDagran#2727 @mobuchowskiPython: suppress warning on importing v1 module in __init__.py. `#2713` @JDarDagran *Suppresses the deprecation warning when v1 facets are used.*
#2706 @dolfinusSchemaDatasetFacet with nested fields for Iceberg tables with list, map and struct columns.#2711 @dolfinusSchemaDatasetFacet with nested fields for Avro schemas with complex types (union, record, map, array, fixed).#2677 @dolfinusExecutionContext interface.SchemaDatasetFieldsFacet #2689 @dolfinusSchemaDatasetFieldsFacet. Also include field comment as description.SparkApplicationDetailsFacet #2688 @dolfinusSparkApplicationDetailsFacet to runEvents emitted on Spark application start.#2710 @kacpermuda#2693 @JDarDagran#2665 @pawel-big-lebowskiAthenaExtractor #2700 @kacpermudaSchemaDatasetFacet for Protobuf repeated primitive types #2685 @dolfinus#2653 @kacpermudaTokenizerErrors, PanicException #2703 @mobuchowski#2713 @JDarDagran#2686 #2687 @dolfinusrunEvents. The new UUID version produces monotonically increasing values, which leads to more performant queries on the OL consumer side. Note: UUID version is an implementation detail and can be changed in the future.Spark: drop `SparkVersionFacet` `#2659` @dolfinus *Drops the SparkVersion facet, deprecated since 1.2.0 and planned for removal since 1.4.0.*
#2674 @surisimran#2668.#2482 @pawel-big-lebowskiarray type, map type, oneOf and any.#2663 @julienledem#2652 @pawel-big-lebowskijobType property of JobTypeJobFacet to either SQL_JOB or RDD_JOB.#2646 @mobuchowskispark_jobDetails facet #2662 @dolfinus
Adds a SparkJobDetailsFacet, capturing information about Spark application jobs -- e.g. jobId, jobDescription, jobGroup, jobCallSite. This allows for tracking an OpenLineage RunEvent with a specific Spark job in SparkUI.ParentRunFacet key #2660 @dolfinus
Changes the integration to use the parent key for ParentFacet, dropping the outdated parentRun.SparkVersionFacet #2659 @dolfinus
Drops the SparkVersion facet, deprecated since 1.2.0 and planned for removal since 1.4.0.#2679 @JDarDagranParentRunFacet key #2661 @dolfinusParentRunFacet with the property name parent but the Great Expectations integration created a lineage event with parentRun. This renames ParentRunFacet key from parentRun to parent. For backwards compatibility, keep the old name.#2658 @blacklight
Includes profile and models in the dbt job name to make it more unique.org.apache.commons.lang3 instead of org.apache.commons.lang #2676 @harelsThe v2 interface introduces some breaking changes: facets are put into separate modules per JSON Schema spec file, some names are changed, and several…
#2609 @pawel-big-lebowskiDataSetEvent and JobEvent in Transport.emit #2611 @dolfinusTransport.emit(OpenLineage.DatasetEvent) and Transport.emit(OpenLineage.JobEvent), reusing the implementation of Transport.emit(OpenLineage.RunEvent). Please note: Transport.emit(String) is now deprecated and will be removed in 1.16.0.GZIP compression to HttpTransport #2603 #2604 @dolfinuscompression option to HttpTransport config in the Java and Python clients, with gzip implementation.#2571 #2597 #2598 @dolfinusmessageKey option to KafkaTransport config in the Python and Java clients, as well as the Proxy. This option replaces the localServerId option, which is now deprecated. Default value is generated using the run id (for RunEvent), job name (for JobEvent) or dataset name (for DatasetEvent). This value is used by the Kafka producer to distribute messages along topic partitions, instead of sending all the events to the same partition. This allows for full utilization of Kafka performance advantages.#2633 @mobuchowski
Adds a mechanism for forwarding metrics to any Micrometer-compatible implementation for Flink as has been implemented for Spark. Included: MeterRegistry, CompositeMeterRegistry, SimpleMeterRegistry, and MicrometerProvider.#2520 @JDarDagrandatamodel-code-generator for parsing JSON Schema and generating Pydantic or dataclasses classes, etc. In order to use attrs (a more modern version of dataclasses) and overcome some limitations of the tool, a number of steps have been added in order to customize code to meet OpenLineage requirements. Included: updated references to the latest base JSON Schema spec for all child facets. Please note: newly generated code creates a v2 interface that will be implemented in existing integrations in a future release. The v2 interface introduces some breaking changes: facets are put into separate modules per JSON Schema spec file, some names are changed, and several classes are now kw_only.#2583 @pawel-big-lebowskiSparkOpenlineageConfig and FlinkOpenlineageConfig for a more uniform configuration experience for the user. Renames OpenLineageYaml to OpenLineageConfig and modifies the code to use only OpenLineageConfig classes. Includes a doc update to mention that both ways can be used interchangeably and final documentation will merge all values provided.#2613 @tnazarew
Adds a TokenProviderTypeIdResolver to handle both FQCN and (for backward compatibility) api_key types in spark.openlineage.transport.auth.type.SparkConf & FlinkConf approaches #2583 @pawel-big-lebowskiSparkConf/FlinkConf). Allows each integration to have its own config entries.#2533 @pawel-big-lebowski
Enables configuration entries specifying ownership of the job that will result in an OwnershipJobFacet being attached to job facets.partitionKey format with Kafka implementation #2620 @dolfinus
Changes the format of Kinesis partitionKey from {jobNamespace}:{jobName} to run:{jobNamespace}/{jobName} to match the Kafka transport implementation.load_config return an empty dict instead of None when file empty #2596 @kacpermuda
utils.load_config() now returns an empty dict instead of None in the case of an empty file to prevent an OpenLineageClient crash.#2614 @dolfinus
Fixes rendering of javadoc for methods generated by lombok annotations by adding a delombok step.#2599 @mobuchowski…maintained and recent releases have introduced breaking changes.*
lineage_job_namespace and lineage_job_name macros #2582 @dolfinuslineage_job_namespace(), lineage_job_name(task) that return an Airflow namespace and Airflow job name, respectively.SchemaDatasetFacet #2548 @dolfinusSchemaDatasetFacet.#2588 @pawel-big-lebowskipmdTestScala212 of warnings that clutter the logs.#2591 @blacklightdbt-ol now propagates the exit code of the underlying dbt process even if no lineage events are emitted.#2585 @harelsjava.lang.StringIndexOutOfBoundsException.HashSet in column-level lineage instead of iterating through LinkedList #2584 @mobuchowskiHashSet for collection.#2572 @dolfinuspkg_resources dependency and replaces it with the packaging lib.airflow.macros.lineage_parent_id #2578 @blacklightlineage_parent_id Airflow macro and simplifies the format of the lineage_parent_id and lineage_run_id macros.#2579 @JDarDagran
Adds an upper limit on supported versions of Dagster as the integration is no longer actively maintained and recent releases have introduced breaking changes.lineage_job_namespace and lineage_job_name macros #2582 @dolfinus
Adds new Airflow macros lineage_job_namespace(), lineage_job_name(task) that return an Airflow namespace and Airflow job name, respectively.SchemaDatasetFacet #2548 @dolfinus
Allows nested fields support to SchemaDatasetFacet.airflow.macros.lineage_parent_id #2578 @blacklightlineage_parent_id Airflow macro and simplifies the format of the lineage_parent_id and lineage_run_id macros.#2591 @blacklightdbt-ol now propagates the exit code of the underlying dbt process even if no lineage events are emitted.#2579 @JDarDagran#2585 @harelsjava.lang.StringIndexOutOfBoundsException.#2624 @pawel-big-lebowskipkg_resources module on Python 3.12 #2572 @dolfinuspkg_resources dependency and replaces it with the packaging lib.HashSet in column-level lineage instead of iterating through LinkedList #2584 @mobuchowskiHashSet for collection.Common: add support for `SCRIPT`-type jobs in BigQuery `#2564` @kacpermuda In the case of SCRIPT-type jobs in BigQuery, no lineage was being extracted
SCRIPT-type jobs in BigQuery #2564 @kacpermuda
In the case of SCRIPT-type jobs in BigQuery, no lineage was being extracted because the SCRIPT job had no lineage information - it only spawned child jobs that had that information. With this change, the integration extracts lineage information from child jobs when dealing with SCRIPT-type jobs.#2272 @pawel-big-lebowski
This PR adds a spark-interfaces-scala package that allows lineage extraction to be implemented within Spark extensions (Iceberg, Delta, GCS, etc.). The Openlineage integration, when traversing the query plan, verifies if nodes implement defined interfaces. If so, interface methods are used to extract lineage. Refer to the README for more details.#2496 @mobuchowskiMeterRegistryyFactory, MicrometerProvider, StatsDMetricsBuilder, metrics config in OpenLineage config, and a Java client implementation.#2528 @mobuchowski#2556 @mobuchowski
Adds support for the Spark-BigQuery connector's query input option, which executes a query directly on BigQuery, storing the result in an intermediate dataset, bypassing Spark's computation layer. Due to this, the lineage is retrieved using the SQL parser, similarly to JDBCRelation.SparkPropertyFacetBuilder to support recording Spark runtime #2523 @Ruihua98SparkPropertyFacetBuilder to capture the RuntimeConfig of the Spark session because the existing SparkPropertyFacet can only capture the static config of the Spark context. This facet will be added in both RDD-related and SQL-related runs.fileCount to dataset stat facets #2562 @dolfinusfileCount field to DataQualityMetricsInputDatasetFacet and OutputStatisticsOutputDatasetFacet specification.dbt-ol should transparently exit with the same exit code as the child dbt process #2560 @blacklightdbt-ol transparently exit with the same exit code as the child dbt process.#2531 @HuangZhenQiuopenlineage-flink.jar.#2507 @pawel-big-lebowski
Fixes the class not found issue when checking for Cassandra classes. Also fixes the Maven pom dependency on subprojects..emit() method logging & annotations #2539 @dolfinus#2547 @dolfinusOpenLineageSql class could not load a native library, if returned None for all operations. But because the error message was suppressed, the user could not determine the reason.#2510 @mobuchowski#2535 @pawel-big-lebowskiIllegalStateException is always caught when accessing SparkSession.#2537 @pawel-big-lebowskiClassNotFoundError occurring on Databricks runtime and extends the integration test to verify DatabricksEnvironmentFacet.#2565 @d-m-h
The JobMetricsHolder#cleanUp(int) method now correctly purges unneeded state from both maps.UnknownEntryFacetListener #2557 @pawel-big-lebowski
Prevents storing the state when a facet is disabled, purging the state after populating run facets.JDBCOptions(table=...) containing subquery #2546 @dolfinusopenlineage-spark from producing datasets with names like database.(select * from table) for JDBC sources.#2563 @mobuchowskiIllegalStateException when accessing SparkSession #2535 @pawel-big-lebowski
IllegalStateException was not being caught.Nothing published for this version
Nothing published for this version
A new config param timeoutInMillis has been added. the Existing timeout has been removed from docs and will be deprecated in 1.13.*
#2518 @JDarDagran
Adds the new provider required by the latest version of Dagster.#2491 @HuangZhenQiu
Adds support for hybrid source lineage for users of Kafka and Iceberg sources in backfill usecases.#2472 @HuangZhenQiu
Bumps the Flink JDBC connector version to 3.1.2-1.18 for Flink 1.18.OpenLineageClientUtils#loadOpenLineageJson(InputStream) and change OpenLineageClientUtils#loadOpenLineageYaml(InputStream) methods #2490 @d-m-h
This improves the explicitness of the methods. Previously, loadOpenLineageYaml(InputStream) wanted the InputStream to contain bytes that represented JSON.#2486 @davidjgoss
Adds the status code and body as properties on the thrown exception when a non-success response is encountered in the HTTP transport.#2478 @mattiabertorello#2524 @kacpermuda
Refines the operator's attribute inclusion logic in facets to include only those known to be important or compact, ensuring that custom operator attributes with substantial data do not inflate the event size.task_instance copy fails #2492 @kacpermuda
Airflow will now proceed without rendering templates if task_instance copy fails in listener.on_task_instance_running.HttpTransport timeout #2475 @pawel-big-lebowski
The existing timeout config parameter is ambiguous: implementation treats the value as double in seconds, although the documentation claims it's milliseconds. A new config param timeoutInMillis has been added. the Existing timeout has been removed from docs and will be deprecated in 1.13.#2515 @pawel-big-lebowski
Adds a check for a null context before executing end(jobEnd).#2507 @pawel-big-lebowski
Fixes the class not found issue when checking for Cassandra classes. Also fixes the Maven POM dependency on subprojects.#2512 @HuangZhenQiu
Enables the JDBC table name with a schema prefix.#2508 @pawel-big-lebowski
For JDBC, the Flink integration is not adjusted to the Openlineage naming convention. There is code that extracts the dataset namespace/name from the JDBC connection url, but it's in the Spark integration. As a solution, this code has to be extracted into the Java client and reused by the Spark and Flink integrations.#2507 @pawel-big-lebowski
Flink is failing when no Cassandra classes are present on the class path. This is happening because of CassandraUtils class which has a static hasClasses method, but it imports Cassandra-related classes in the header. Also, the Flink subproject contains an unnecessary maven-publish plugin.#2504 @HuangZhenQiu
The shadow jar of Flink is not minimized, so some internal jars are listed as runtime dependences. This removes them from the final pom.xml file in the Flink module.#2479 @HuangZhenQiu
Following the namespace definition, we should use cassandra://host:port.Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →