NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
A tool for converting ONNX files to LiteRT/TFLite/TensorFlow, PyTorch native code (nn.Module), TorchScript (.pt), state_dict (.pt), Exported Program (.pt2), and Dynamo ONNX. It also supports direct conversion from LiteRT to PyTorch.
Last release 21 days ago
14 Sep 2026
Release timing varies
gaps range from 8 days to 2 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
629 releases · first in 2022
One column per quarter.
With --tflite_backend tf_converter , a Conv whose input has C == H == W and explicit padding can come out wrong while -cotof prints Matches for every
With --tflite_backend tf_converter, a Conv whose input has C == H == W and explicit padding can come out wrong while -cotof prints Matches for every OP.
When all non-batch axes of a Conv input have the same size, strict mode picks the input layout by trying every permutation that keeps the batch axis: transpose the validation data from the previous layer, run a dummy Conv on it, compare with the ONNX dummy inference, keep the smallest error. With explicit pads that can't be expressed as SAME, input_tensor is padded before that search but the validation data is not. For a 3x3 stride-2 Conv on 8x8 every candidate outputs (1,3,3,8) against ONNX's (1,8,4,4), all of them score the sentinel error, and min_abs_err_perm_1 silently stays at identity. If the input is still in ONNX order, e.g. straight out of a Reshape, the Conv uses the wrong axis as channels. The per-OP check does not catch it because it tolerates layout ambiguity when the dims are equal.
onnx2tf/ops/Conv.py only; the padding decision is unchanged. I keep a reference to the unpadded tensor, apply the winning permutation to that, and pad afterwards. Each transposed validation candidate gets the same padding before the dummy Conv is built. Both go through apply_explicit_padding, which returns its argument unchanged unless padded is set, so every other path behaves exactly as before. Ties still resolve to identity (strict less-than, identity is the first candidate).
Smallest model graph that shows it: X(1,512) -> Add(const) -> Reshape(1,8,8,8) -> Conv(k3, s2, pads=[1,1,1,1]) -> Y(1,8,4,4). The Reshape leaves the Conv input in ONNX order; the Add makes the input non-constant, since the dummy inference feeds np.ones and a constant input gives every permutation the same output.
import numpy as np, onnx
from onnx import TensorProto, helper, numpy_helper
rng = np.random.default_rng(0); N = 8
init = [numpy_helper.from_array(rng.standard_normal((N, N, 3, 3)).astype(np.float32), "W"),
numpy_helper.from_array(rng.standard_normal((N,)).astype(np.float32), "B"),
numpy_helper.from_array(rng.standard_normal((N*N*N,)).astype(np.float32), "bias512"),
numpy_helper.from_array(np.array([1, N, N, N], dtype=np.int64), "shape")]
g = helper.make_graph(
[helper.make_node("Add", ["X", "bias512"], ["A"]),
helper.make_node("Reshape", ["A", "shape"], ["R"]),
helper.make_node("Conv", ["R", "W", "B"], ["Y"], kernel_shape=[3,3], strides=[2,2], pads=[1,1,1,1])],
"conv_padded_same_axes",
[helper.make_tensor_value_info("X", TensorProto.FLOAT, [1, N*N*N])],
[helper.make_tensor_value_info("Y", TensorProto.FLOAT, [1, N, N//2, N//2])], init)
m = helper.make_model(g, opset_imports=[helper.make_opsetid("", 13)]); m.ir_version = 8
onnx.save(m, "conv_padded_same_axes.onnx")onnx2tf -i conv_padded_same_axes.onnx -o out -tb tf_converter -nuo -cotof
main (d6cd715), with a temporary print in the search loop:
perm=(0, 1, 2, 3) err=9223372036854775807 tf_out=(1, 3, 3, 8) onnx=(1, 8, 4, 4)
perm=(0, 1, 3, 2) err=9223372036854775807 tf_out=(1, 3, 3, 8) onnx=(1, 8, 4, 4)
perm=(0, 2, 1, 3) err=9223372036854775807 tf_out=(1, 3, 3, 8) onnx=(1, 8, 4, 4)
perm=(0, 2, 3, 1) err=9223372036854775807 tf_out=(1, 3, 3, 8) onnx=(1, 8, 4, 4)
perm=(0, 3, 1, 2) err=9223372036854775807 tf_out=(1, 3, 3, 8) onnx=(1, 8, 4, 4)
perm=(0, 3, 2, 1) err=9223372036854775807 tf_out=(1, 3, 3, 8) onnx=(1, 8, 4, 4)
chosen perm=[0, 1, 2, 3]
INFO: onnx_output_name: Y shape: (1, 8, 4, 4) dtype: float32 validate_result: Matches
ONNX/TFLite output check complete! (out/conv_padded_same_axes_accuracy_report.json) trigger=-cotof(auto) source=tf_converter target=float32 max_abs=57.625 rmse=16.7022 cosine=0.154714 pass=False
this PR:
perm=(0, 1, 2, 3) err=41.81199645996094 tf_out=(1, 4, 4, 8)
perm=(0, 1, 3, 2) err=30.937366485595703 tf_out=(1, 4, 4, 8)
perm=(0, 2, 1, 3) err=27.101259231567383 tf_out=(1, 4, 4, 8)
perm=(0, 2, 3, 1) err=3.814697265625e-06 tf_out=(1, 4, 4, 8)
chosen perm=[0, 2, 3, 1]
ONNX/TFLite output check complete! (out/conv_padded_same_axes_accuracy_report.json) trigger=-cotof(auto) source=tf_converter target=float32 max_abs=0 rmse=0 cosine=1 pass=True
Exported float32 TFLite against the ONNX model on the same random input, TF output transposed to NCHW:
| permutation search | -cotof output check |
float32 TFLite vs ONNX | |
|---|---|---|---|
| main (d6cd715) | all 6 candidates hit the sentinel error, identity picked | max_abs=57.6 cosine=0.15 pass=False |
SNR -2.98 dB, max abs error 46.0 |
| this PR | [0,2,3,1] err 3.8e-06, identity err 41.8 |
max_abs=0 cosine=1 pass=True |
SNR 342 dB, bit-identical |
Regression check: a (1,8,8,8) input fed straight into the Conv with pads [1,1,1,1] and with [1,1,0,0], and the Reshape model without the Add (constant input), all give a float32 TFLite that is bit-identical to what main produces, and the per-OP check still says Matches. The constant-input Reshape model is wrong on both, since np.ones dummy data cannot separate the permutations; that is unrelated to this change. The same model with -dsm still pads eagerly and outputs (1,4,4,8), and the default flatbuffer_direct backend is untouched. Python 3.12.11, tensorflow 2.21.0, onnx 1.20.1, onnxruntime 1.26.0, macOS arm64.
None.
Full Changelog: 2.6.8...2.6.9
Conv.make_node() reads input_tensor_shape before the workaround that transposes an NCHW input to NHWC, and never refreshes it.
Conv.make_node() reads input_tensor_shape before the workaround that transposes
an NCHW input to NHWC, and never refreshes it.
The depthwise test compares group against input_tensor_shape[-1]. For a transposed
input this compares the channel count against a spatial dimension, the test fails, and
the node falls through to the grouped-conv path — producing a filter of shape
[C, K, K, 1] instead of a DepthwiseConv2D.
Symptom: in a MobileNet/ConvNeXt-style network only the first depthwise conv of each
stage converts correctly (it consumes a conv output directly), while the following ones
are silently degraded (they consume a residual Add whose input required the transpose).
The ONNX model is valid — depthwise is Conv with group=C and weight [C, 1, K, K].
Refresh input_tensor_shape from input_tensor right before the depthwise test.
Full Changelog: 2.6.7...2.6.8
This pull request continues the characterize-first refactoring of the TensorFlow-free flatbuffer_direct backend. It restructures late shape, layout, t
This pull request continues the characterize-first refactoring of the TensorFlow-free flatbuffer_direct backend. It restructures late shape, layout, topology, fallback, and terminal-cleanup orchestration without changing public conversion policy or using newly recorded evidence to alter graph-rewrite control flow.
The main outcome is a smaller and more reviewable central lowerer whose phase boundaries are explicit, whose repeated orchestration sequences have focused owners, and whose observation-only results are retained in a bounded conversion-session store under stable phase IDs.
The branch also includes the full sequential Tier 0–4 TFLite/native-PyTorch regression evidence for the final checkpoint and updates the package version metadata to 2.6.7.
The direct FlatBuffer path has accumulated many interacting shape, layout, topology, compatibility, and model-recovery rules. Even when each rule is locally correct, several structural properties made maintenance difficult:
This branch makes those contracts explicit before attempting more aggressive scan elimination or control-flow optimization.
ConversionSession now owns a bounded, conversion-local phase-result store.
This removes observation-only lowerer locals while giving future optimization work a stable basis for differential testing.
Repeated multi-step sequences now have focused pass-module owners, including:
The owners preserve the original arguments, guards, ordering, cycle behavior, metadata updates, and temporary-state lifetime.
Large groups of previously adjacent lowerer calls were extracted into narrowly scoped orchestration owners. The extracted families include:
These owners return the same child mappings or strict summaries in the same source order. Existing lowerer compatibility wrappers are retained where instrumentation or callback injection depends on them.
The central lower_from_onnx2tf.py no longer manually coordinates each extracted cluster. It delegates orchestration while retaining the existing phase order and records only the bounded result required by the conversion session.
The refactoring deliberately preserves:
flatbuffer_direct remains the default backend.-cotof do not import TensorFlow.Each extraction was split into two checkpoints:
When an ownership move intentionally invalidated a structural source assertion, the initial failure was captured first. Tests were then updated to assert the new owner boundary while preserving runtime behavior. This avoids masking behavioral regressions as mechanical test maintenance.
The branch contains detailed implementation and checkpoint records in:
docs/fb_refactor8_improvements.mddocs/flatbuffer_direct_handoff_fb_refactor8.mddocs/fb_refactor8_pull_request_description.mdThe final branch checkpoint was validated against every eligible root-level ONNX model in Tiers 0–4.
VmSwap.fb-refactor7.Raw component counts remained exactly equal to the previous branch:
| Component | Pass | Fail | Missing report | Conversion error |
|---|---|---|---|---|
| TFLite | 347 | 20 | 10 | 2 |
| Native PyTorch | 138 | 46 | 193 | 2 |
The strict runner intentionally retains inherited missing reports and accuracy failures. Those entries were compared model-by-model rather than counted as new regressions.
The only numeric difference was randnlike4.onnx, whose native-PyTorch output is intentionally nondeterministic. Its classification was unchanged. The two DEIM models retain the user-approved near-tied TopK-index disposition. The two DPT-LeViT conversion errors are inherited and unchanged.
Total per-model conversion time decreased from 7,689.33 seconds to 7,376.09 seconds (-4.07%). Every Tier 0–4 aggregate total and median was below the fb-refactor7 observation. The slowest model completed in 208.95 seconds, well below the 600-second boundary.
Pass-efficiency metrics were exactly equal between the two branches:
Full evidence is recorded in:
docs/flatbuffer_direct_tier0_4_full_regression_2026-07-20.mddocs/baselines/flatbuffer_direct_tier0_4_tflite_native_pytorch_fb8_b35c36a0.jsonAll inference and conversion tests were run sequentially under uv.
The branch includes focused contracts for:
The final bulk-runner and corpus-manifest gate reports 44 passed tests. Numerous focused and affected suites were run at each characterize/extract checkpoint; their commands and exact counts are preserved in the branch documentation.
The full corpus run retained logs, JSON reports, pass metrics, and text diagnostics while deleting only reproducible artifacts after their model entry was atomically persisted. It reclaimed 51,216,397,607 bytes (47.70 GiB), then removed the uniquely named temporary run directory after condensed evidence validation.
The package version, lockfile entry, and documented container tags are updated consistently from 2.6.6 to 2.6.7. uv lock --check and an import-time metadata consistency assertion both pass.
This PR does not use the new phase evidence to eliminate graph scans or skip passes. Such optimization should be a separate change with mutation-positive and stable-path differential evidence for operator order, layout state, cycle handling, ModelIR digest, serialized artifacts, and numerical output.
The broader flatbuffer_direct refactor is therefore not declared complete by this PR. This checkpoint provides explicit ownership, bounded evidence, strong regression characterization, and a safer foundation for subsequent optimization without changing current public behavior.
Full Changelog: 2.6.6...2.6.7
This branch continues the staged refactoring of the TensorFlow-free flatbuffer_direct backend. It makes pass ownership, mutation evidence, graph index
This branch continues the staged refactoring of the TensorFlow-free
flatbuffer_direct backend. It makes pass ownership, mutation evidence, graph
index reuse, reconciliation decisions, and terminal validation substantially
more explicit while preserving the existing conversion policy and public
interfaces.
The central goal of this checkpoint is not to introduce a new converter or to
change the order of the existing compatibility rules. It is to make the
current behavior safer to evolve:
ModelIRGraphIndex;The release metadata is updated consistently from 2.6.5 to 2.6.6 in the
package, project metadata, lockfile, and documented container tags.
The direct FlatBuffer path has accumulated many interacting layout, shape,
quantization, recovery, and compatibility rules. Historically, a number of
those rules were invoked through raw calls in the central lowerer, returned
useful information that callers discarded, rebuilt producer/consumer state at
nearby boundaries, or triggered reconciliation even when the preceding owner
made no graph mutation.
Those patterns make apparently local changes difficult to reason about. A
counter can be incomplete because cleanup happened outside the counter, a
zero result can be mistaken for graph stability, or two adjacent passes can
independently rescan the same graph. This branch establishes explicit contracts
for these boundaries before using them for conservative scan elimination or
index sharing.
run_model_ir_pass_group() no longer scans the complete and growing
diagnostic history for every emitted event. One scan initializes the global
event count, maximum group sequence, and per-pass invocation counts. Those
counters are then updated as events are appended.
ConversionSession retains this state across pass groups. Ordinary
caller-owned lists remain supported through a safe one-scan fallback. The
external diagnostic fields, numbering semantics, list identity, skip/cycle
behavior, and invariant-failure reporting are unchanged.
Late, very-late, terminal, recovery, attention, quantized, binary-layout,
Conv1D, and safety-fallback boundaries now return or stage fixed-schema result
dictionaries instead of dropping child results. Multi-pass orchestration
preserves child order and shared pass-state scopes.
Observation-only results remain intentionally unconsumed when their counters
are not yet safe control-flow inputs. Where a stability decision is made, the
owner contract accounts for every relevant mutation, including cleanup-only
tensor pruning. Non-mutating iteration counters are excluded from mutation
summaries.
This gives later work auditable evidence without adding ModelIR copies,
fingerprints, or extra graph walks.
Static-shape reconciliation can now opt in to a complete mutation count. In
addition to ordinary tensor-shape updates, it records operator-option changes,
constant shape-parameter writes, and direct tensor metadata updates performed
during the existing fixed-point walk.
The default result schema remains compatible. Only characterized call sites
request the additional evidence, and their original guards, pruning behavior,
layout synchronization, and pass order are preserved.
This complete evidence is retained across primary, fallback, very-late, and
post-split reconciliation boundaries, including Conv-input, mixed-Concat,
Concat-axis, high-rank BatchMatMul, Pad, placeholder-MatMul, singleton Reshape,
SE/FC/Gather, PReLU, and SiNet paths.
Broad reconciliation is skipped only when all declared mutation evidence from
the immediately preceding owner is zero. This applies to selected singleton
Reshape, PReLU, SiNet, SE/Gather, late binary, shared late, final-shape,
post-fusion, placeholder-MatMul, and bounded convergence paths.
Mutation-positive behavior is unchanged. Cleanup-only owners use clamped net
tensor-count deltas where required, so a rewrite counter of zero is never
silently treated as stability when pruning may still have changed the graph.
Binary-layout convergence can also stop after a stable round instead of
running the remaining bounded rounds. The maximum round count and all
mutation-positive behavior remain unchanged.
run_indexed_prune_reconcile_cleanup() shares one ModelIRGraphIndex across
dead-operator pruning and static-shape reconciliation at three repeated phase
boundaries. It performs exactly those two existing operations and does not add
dynamic-reshape resolution, retries, or a new convergence branch.
run_indexed_binary_layout_adapter_cleanup() similarly shares one index
between the exact rank-four binary adapter and the singleton-broadcast adapter
at four repeated boundaries. The exact adapter still runs first. Both owners
now use indexed candidate lookup, operator insertion, and input replacement,
while preserving names, tensor metadata, quantization cloning, pruning, result
schemas, and existing guard inputs.
Fallback unbound-input repair also avoids one redundant reconciliation because
its indexed compatibility wrapper already performs the required positive-path
reconciliation with the live graph index.
Primary and fallback terminal layout validation now observe the graph after
the applicable terminal mutations and final topological ordering. When the
terminal graph is valid, validation clears only stale validation errors
inherited from recursive lowering. The diagnostic schema and actual error
conditions are unchanged.
The branch retains the complete terminal mutation dictionaries used to make
this boundary reviewable, including bounded binary convergence, high-rank
binary coalescing, boundary-signature realignment, high-rank BatchMatMul, Pad,
Conv-input, mixed-Concat, Concat-axis, stale-binary, and SiNet repairs.
Several responsibilities were moved behind focused internal interfaces:
Constant lowering lives in its op-family module;NodeView value-list annotations and ONNX graph-input typing are clearer toPrivate compatibility wrappers remain where existing structural callers rely
on them. The central lowerer keeps the same production call positions and
ordering unless a characterized shared runner replaces an exactly adjacent
pair.
Each behavior-changing implementation unit was preceded by a focused contract
that froze the relevant result schema, no-op path, mutation path, order,
arguments, state-scope ownership, and surrounding phase boundary. The strict
expected failure was removed only after the corresponding implementation and
the broader sequential gate passed.
This branch adds extensive focused coverage for result propagation,
cleanup-only mutations, deterministic ordering, idempotence, shared-index
construction, layout recovery, attention and quantized paths, fallback
relowering, terminal validation, and architectural ownership.
flatbuffer_direct remains the default backend.-cotof do not import or executeAt the final implementation checkpoint before the corpus run:
The broad gate includes the changed indexed owners, late and terminal
orchestration, static-shape reconciliation, shared ModelIR pass context, core
contracts, pass-efficiency checks, architecture constraints, and the optional
TensorFlow import boundary.
The corpus runner and manifest tests were also repeated after the full run:
44 passed.
All 379 active managed Tier 0-4 models were converted strictly one at a time at
converter checkpoint 0bf1bab4. The scope excludes only previously recorded
timeouts and explicit user exclusions.
fb-refactor6:The combined classifications exactly match the authoritative fb-refactor6
comparison:
| Classification | Models |
|---|---|
pass |
137 |
missing_pytorch_report |
187 |
pytorch_fail |
31 |
both_fail |
12 |
missing_both_reports |
6 |
missing_tflite_report |
4 |
conversion_error |
2 |
| Total | 379 |
The run requested TFLite and the native PyTorch package only. It did not
request TorchScript, Dynamo ONNX, ExportedProgram, or TensorFlow artifacts.
This full-corpus run is a single regression observation, not the broader
three-run warm-median performance gate.
+1.945%);+0.109%);+5.332%);+10% of the comparison run.Both runs recorded exactly 196,325 pass events. state_build_count,
snapshot_count, and fingerprint_count were unchanged, while
preflight_operators_visited decreased by 16. This provides no evidence of an
added scan or state-build regression.
This PR does not claim that every corpus model has numerical parity. It claims
that the branch introduces no new result or failure-signature regression
relative to the authoritative previous checkpoint.
randnlike4.onnx retains the same native-PyTorch known-failure class. Its1.828334451 to1.811047643; this is not treated as a converter regression.docs/fb_refactor7_improvements.mddocs/flatbuffer_direct_architecture.mddocs/flatbuffer_direct_tier0_4_full_regression_2026-07-18.mddocs/baselines/flatbuffer_direct_tier0_4_tflite_native_pytorch_fb7_0bf1bab4.jsondocs/flatbuffer_direct_handoff_fb_refactor7.mdThe generated model artifacts were intentionally not committed. After the
evidence was persisted and validated, 48.109 GiB of reproducible temporary
artifacts and the remaining run-owned /tmp files were removed.
This is a coherent refactoring checkpoint, not completion of the entire
long-term flatbuffer_direct redesign. Later work may continue the fixed
ConversionRequest/ConversionSession contract, phase/pass manager,
centralized layout planning, unified lowering registry, exporter separation,
quantization/split refresh, and PyTorch-family restructuring. Those items are
not silently mixed into this PR.
Full Changelog: 2.6.5...2.6.6
This pull request is the sixth staged checkpoint of the flatbuffer_direct refactor. It moves the remaining large groups of already-supported ModelIR c
This pull request is the sixth staged checkpoint of the flatbuffer_direct
refactor. It moves the remaining large groups of already-supported ModelIR
compatibility logic out of the central ONNX lowerer and into focused,
characterized pass and orchestration modules.
The central implementation in lower_from_onnx2tf.py is reduced from 25,647
lines on main to 5,585 lines. The branch adds 79 focused pass/orchestration
modules and updates or adds 92 test files. The large net addition is primarily
the extracted owners, synthetic characterization fixtures, architecture
contracts, and the recorded Tier 0-4 regression evidence; it is not a second
conversion pipeline.
The main goals of this checkpoint are:
-cotof outside the TensorFlow importNo dependency was added. Project and container metadata are synchronized at
version 2.6.5.
The direct FlatBuffer backend accumulated shape repair, layout recovery,
quantization compatibility, fusion, and late cleanup rules in one central
module. Many rules were individually valid, but their implicit ordering and
shared mutation made them difficult to understand or modify safely. A guard
that failed after an early tensor/layout mutation could leave the graph in a
partially rewritten state, and nested recovery helpers repeatedly assembled
nearly identical ModelIR/LayoutState/diagnostic contexts.
This branch follows a deliberately conservative sequence for each family:
This separates ownership without attempting a broad semantic rewrite at the
same time.
The branch extracts shape, layout, quantization, attention, binary, affine,
Concat, and terminal recovery behavior into op-family-oriented modules under
onnx2tf/tflite_builder/passes/.
Representative extracted families include:
The original call sites and their relative order remain visible in the
lowerer. Existing private names remain available through compatibility
wrappers, so downstream internal imports are not forced to change in this
checkpoint.
Several rule families previously performed metadata or graph mutations before
all rejection conditions had been evaluated. The branch adds preflight
planning or snapshot-based rollback behavior before extracting those owners.
The corrected families include Conv input adapters, Conv/Pool passthrough,
mixed-layout Concat, stale binary adapters, QLinear output propagation,
QLinear SiLU prefix recovery, Mean/MaxPool/Concat recovery, Softmax transpose
canonicalization, Concat bridge families, Reshape/Transpose collapse, and
attention cleanup/rank-lift patterns.
The tests explicitly verify that rejected candidates leave operator lists,
tensor metadata, quantization records, LayoutState, graph outputs, and
diagnostics unchanged. Successful rewrites retain deterministic candidate
priority, fixed-point restart behavior, unique naming, pruning order, and the
historical statistics keys.
Extracted owners use ModelIRGraphIndex for producer/consumer lookup and
differential updates where the family contract permits it. Nested pass groups
share one pass-state scope so graph indexes can be reused across compatible
owners instead of being rebuilt for each helper.
Architecture and efficiency tests cover:
Repeated frozen context dataclasses are consolidated around
ModelIRPassContext, owned by ConversionSession. Twenty-five historical
context names remain compatible internal aliases, while main-model consumers
reuse the Session-owned ModelIR, LayoutState, and diagnostics identities.
Callback-bearing recovery contexts now compose the shared context while
retaining their exact callback objects and argument contracts. Target-specific
and child-model conversions still receive fresh context where isolation is
required. No orchestration dataclass now duplicates the common core identity
fields.
NodeView.inputs and NodeView.outputs previously created a new dynamic class
for every tensor reference. Runtime attribute access worked, but static
analysis could not prove that .name, .onnx_name, .shape, and .dtype
existed. They now use lightweight typed instances with the same mutable
metadata and remapping behavior. This removes the Pylance attribute-access
error without suppressing diagnostics.
The bulk runner now publishes its increasingly large state JSON through a
same-directory temporary file followed by os.replace(). Concurrent readers
therefore observe either the previous complete state or the new complete
state, never the truncated interval created by direct open(..., "w") output.
The change preserves filenames, JSON schema, indentation, file mode, resume
behavior, and classification. Serialization or replacement failure removes
the unpublished temporary file and leaves the previous state intact. No
per-update fsync was added, avoiding unnecessary I/O cost for state files
that reached approximately 17 MiB.
The pre-correction evidence was committed before changing the writer, so the
observed failure and the correction remain independently reviewable.
flatbuffer_direct remains the default backend.-cotof do not import or execute TensorFlow.uv environment.The final managed run covered every active Tier 0-4 model that was below the
ten-minute policy boundary, did not trigger SWAP, and had not been explicitly
excluded by prior validation policy.
| Metric | Result |
|---|---|
| Active models completed | 379 / 379 |
| Timeout | 0 |
| Nonzero process-tree SWAP | 0 |
| Median per-model time | 8.746 s |
| Maximum per-model time | 191.394 s |
| Sum of per-model time | 7,542.595 s |
The component observations were:
| Result | TFLite | Native PyTorch |
|---|---|---|
| Accuracy pass | 347 | 138 |
| Accuracy fail | 20 | 46 |
| Missing report | 10 | 195 |
| Pre-report conversion error | 2 | 2 |
These are raw observations, not regression counts. The existing PyTorch
failures and missing reports remain visible and are not classified as
successful conversions.
Compared with the finalized fb-refactor5 baseline:
yolox_nano.onnx and rtdetrv4_s.onnx remain improved passes;fb-refactor6-specific TFLite or native-PyTorch regression was confirmed.The full model-level result, timing, normalized signatures, exclusions, and
baseline comparison are recorded in:
docs/flatbuffer_direct_tier0_4_full_regression_2026-07-17.md;docs/baselines/flatbuffer_direct_tier0_4_tflite_native_pytorch_fb6_13fcf7ce.json.The refactor was developed as many small characterization/correction/extraction
checkpoints. Focused owner/wrapper, architecture, efficiency, core lowering,
and TensorFlow-import-blocked gates were run at each boundary. The detailed
command history and checkpoint results are retained in
docs/fb_refactor6_pull_request_description.md and
docs/flatbuffer_direct_handoff_2026-07-14.md.
Latest relevant gates include:
uv lock --check, versiongit diff --check: passed.The authoritative 379-model conversion run is tied to converter checkpoint
13fcf7ce. The later atomic-writer and NodeView typing commits do not alter
lowering, export policy, scheduling, or model inference and were covered by
their focused gates rather than another multi-hour corpus run.
Full Changelog: 2.6.4...2.6.5
This PR continues the flatbuffer_direct refactoring after the fb-refactor4 work merged in #949 . It moves a large remaining set of layout, shape, quan
This PR continues the flatbuffer_direct refactoring after the fb-refactor4 work merged in #949. It moves a large remaining set of layout, shape, quantization, fusion, and graph-repair rules out of the central lowerer and into focused, indexed ModelIR passes while preserving the existing CLI, Python API, default backend, artifact names, report contracts, and optional-exporter behavior.
The branch contains 168 commits affecting 230 files. Its largest structural change is the addition of 71 focused pass modules and 83 focused test modules. The large line count is primarily the explicit extraction of production rules and their characterization tests, rather than a new parallel conversion path.
The direct TFLite path had accumulated many order-sensitive graph rewrites in a very large lowering function. Those rewrites repeatedly scanned the whole ModelIR, inferred layout locally, and mutated shared state through implicit conventions. This made individual fixes expensive to reason about and increased the probability that a change for one model family would disturb another.
This PR makes rule ownership and state boundaries explicit. The intention is not to remove supported behavior: it is to retain that behavior behind smaller semantic passes, indexed graph access, typed request/artifact policy, and focused regression contracts.
ConversionRequest once and stops propagating raw kwargs into lower layers.ArtifactPlan.These changes retain the legacy public keys, defaults, artifact order, filenames, failure messages, and return dictionary.
ModelIRPassStateScope when there is no intervening raw ModelIR mutation.ModelIRGraphIndex owns producer/consumer and operator-type lookups; LayoutState is kept synchronized through differential mutation helpers.This removes repeated index construction across many mean/attention, gate, concat, shuffle, boundary, slice, singleton-reshape, QDQ, SPP, and terminal-cleanup clusters without changing their historical execution order.
Seventy-one new production pass modules replace central-lowerer rule blocks. The extracted families include:
The central dispatch remains responsible for orchestration, while match/guard/rewrite behavior and its invariants live beside focused tests. Existing rule order was preserved during extraction.
Six fast-precanonicalization policies were moved out of the large exporter into pytorch_fast_precanonicalize_policy.py:
The managed bulk runner also gained an opt-in native-only PyTorch mode. It emits -fdopt without implicitly enabling TorchScript, Dynamo ONNX, or ExportedProgram, while the runner's existing default behavior remains unchanged.
/proc; any model that generates SWAP is excluded from future managed validation according to the documented policy.vit_b_encoder.onnx and superpoint_lightglue_end2end_fused_cpu.onnx as timeout_after_600s exclusions.The full current-branch run initially exposed four genuine TFLite regressions. Each was recorded, reproduced against detached main, bisected to its first bad extraction commit, and corrected with a generic guard rather than a model-name exception:
(1,) has a scalar, one-element NumPy backing, then normalize it to the declared shape.MUL only when the concat is one operand and the other operand is proven to depend on the matched global-pool Conv output. Arbitrary fan-out remains rejected.Final sequential validation passed all four repaired models with zero SWAP and maximum absolute errors of 0.000294536, 4.67896e-06, 2.72691e-06, and 0.00802997, respectively.
uv environment.-tb flatbuffer_direct conversion and -cotof remain behind the TensorFlow import blocker.py_compile checks passed.git diff --check passed.onnx2tf.__version__, pyproject.toml, and README Docker examples at 2.6.4.The authoritative sanitized-environment run processed all 381 then-active models sequentially with the exact command shape -tb flatbuffer_direct -cotof -fdopt:
yolox_nano.onnx and rtdetrv4_s.onnx);After correction, the four confirmed branch regressions passed exact sequential real-model revalidation. The next managed profile contains 379 active and 41 excluded Tier 0-4 records.
The run generated and evaluated the native PyTorch package for every active model. Because no managed pre-fb-refactor5 native-PyTorch baseline exists, non-passes were not automatically labeled as regressions. Eight representative high-signal failures were rerun sequentially on detached main; all eight produced the same behavior on both implementations. No fb-refactor5-specific native PyTorch regression was confirmed by that comparison.
tests/test_pytorch_exporter.py run was intentionally stopped at 89% when one exported-program archive test spent several minutes recompiling. At interruption it had 942 passes and 81 failures; this is characterization, not an acceptance result. Representative fast failures reproduced identically on detached main.Full Changelog: 2.6.3...2.6.4
This pull request is the accumulated fb-refactor4 checkpoint for the flatbuffer_direct backend. It restructures the converter around explicit internal
This pull request is the accumulated fb-refactor4 checkpoint for the flatbuffer_direct backend. It restructures the converter around explicit internal contracts, indexed graph state, ordered and validated passes, request-driven artifact generation, and a substantially decomposed PyTorch export stack.
The public CLI and Python API remain compatible. The default backend remains flatbuffer_direct, normal direct TFLite conversion and -cotof do not import TensorFlow or tf-keras, no new third-party dependency is introduced, and TensorFlow remains available only behind explicitly requested TensorFlow-family artifacts.
The package version is updated from 2.6.2 to 2.6.3 in the package metadata, lock file, and documented container examples.
The previous direct backend had accumulated a very large set of interacting graph rules and repeated whole-graph scans. This made small changes expensive to reason about, allowed ordering assumptions to remain implicit, duplicated producer/consumer and layout knowledge, and made optional artifact generation perform work that had not been requested.
The PyTorch exporter had the same problem at a different layer: native code generation, source rewriting, state-dict handling, runtime wrappers, fallback selection, TorchScript/Dynamo ONNX/ExportedProgram export, naming, shape policy, and layout policy were concentrated in a monolithic module with many implicit compatibility bindings.
This branch turns those implicit relationships into bounded internal interfaces while preserving the established conversion behavior.
The direct path is organized around this stage order:
ConversionRequest and ArtifactPlan;ConversionSession, GraphIndex, and LayoutState;PassPhase order;ArtifactPlan;ConversionResult back to the legacy return dictionary.Raw option dictionaries no longer need to propagate through the new internals. Artifact controls, split thresholds, reporting options, and quantization calibration controls are resolved only when the associated artifact is requested.
GraphIndex and ModelIRGraphIndex provide shared producer, consumer, duplicate-producer, operator-position, and operator-type views. Rewriters update those indexes differentially instead of rebuilding maps or repeatedly scanning the entire graph.
Passes have stable IDs, phases, priorities, maximum iteration counts, explicit change results, and graph fingerprints for deterministic cycle termination. Risky rewrites can run transactionally and roll back when shape, dtype, layout, unresolved-tensor, duplicate-producer, or public-output invariants fail.
LayoutState is carried through the session and reconciled with graph changes so layout knowledge is no longer reconstructed independently by every rule.
Large rule families were moved out of the central lowerer into focused pass modules. This includes dynamic reshape, graph cleanup, high-rank binary and MatMul handling, rank-4/rank-5 Concat families, quantized layout families, Pad, split fallback, channel shuffle, multi-branch gates, NDHWC gates, cost-volume/ScatterND patterns, PyTorch compatibility, recurrent/control-flow preparation, and layout validation.
The extracted passes retain compatibility wrappers where existing imports require them. Production call sites use registered runners, indexed root enumeration, explicit guards, shared pruning utilities, and transactional layout reconciliation. Model-name checks are not introduced; repairs are expressed as semantic graph patterns.
The artifact pipeline now avoids entering unrequested exporters, quantization paths, split planning, and report writers. Constant buffers and graph indexes are shared where possible, precision clones have bounded lifetimes, and redundant ModelIR copying and repeated validation-index construction are reduced.
The direct backend releases legacy graph objects before ModelIR lowering when they are no longer needed. Sequential ONNX/TFLite accuracy checking materializes large prepared ONNX graphs once in a managed temporary file rather than transferring serialized model bytes through a multiprocessing pipe.
The PyTorch stack is split into dedicated, mostly Torch-free policy and implementation owners, including:
Native codegen reuses ModelIRGraphIndex, and compatibility bindings required by the dynamically compiled legacy body are explicit and covered by architecture tests. Torch is imported only after a PyTorch-family artifact has actually been requested.
TorchScript, Dynamo ONNX, and ExportedProgram generation share the generated native package and common artifact support. TFLite-backed and SavedModel-backed fallbacks reuse the same package scaffolding without pulling TensorFlow into ordinary direct conversion.
The latest checkpoints repair several inherited PyTorch exporter failures without broad layout-policy changes:
torch.reshape.The last group removes all five inherited NoneType.group() CONCAT canonicalization crashes. Three of those tests now pass completely; the two remaining Yolov7 cases reach older structural assertions rather than crashing.
The repository now contains a managed Tier 0-4 corpus workflow, immutable quick-run manifests/results, normalized failure classifications, timeout exclusions, and SWAP detection for the active converter process tree.
Validation is always sequential. The runner does not use a process pool or parallel inference workers. A model that causes SWAP is classified and added to the managed exclusion policy before a later run.
The 2,000 threshold refers to ONNX graph node count for Tier classification. It is not a source-file line limit.
-cotof do not import TensorFlow or tf-keras.uv.A fixed short-runtime selection of 54 Tier 0-4 models was run sequentially:
| Result | Count |
|---|---|
| Pass | 46 |
| Expected accuracy failure | 4 |
| Missing report | 2 |
| 60-second quick-run timeout | 2 |
| Conversion error | 0 |
| SWAP detected | 0 |
No regression specific to fb-refactor4 was confirmed. Focused comparison showed that silero_vad.onnx fails identically on fb-refactor3, while d3net_dnn_double_44.onnx reaches the same quick ceiling on both branches. nchw.onnx passes direct execution on both branches with byte-identical artifacts and metrics; only the instrumented bulk run consumes the timeout headroom.
Apart from the explicitly accepted DEIM result, the largest maximum absolute error among passing models was 0.0443115234375, below the required 1e-1 ceiling. DEIM remains accepted according to its recorded near-tied TopK policy.
The complete central PyTorch exporter suite contains 1,120 tests. At checkpoint 3127e39d, with seven previously classified long-running cases deselected, the exact result was:
| Classification | Count |
|---|---|
| Passed | 1,019 |
| Non-timeout failures | 94 |
| Long-running exclusions | 7 |
All 94 non-timeout failures were reproduced against fb-refactor3 with the same exception type and normalized first message. Six belong to the explicitly excluded bread model family. Therefore, no non-timeout regression specific to fb-refactor4 was confirmed in that comparison.
The latest focused CONCAT/NMS work additionally passed 18 relevant tests, and the architecture suite passed 97 tests. The two remaining Yolov7 assertions are documented rather than hidden: one is a local cv64_in versus cv64_in_cf spelling expectation for the same channel-first tensor, and the other compares an unsanitized direct Dynamo ONNX graph with the helper path that intentionally applies the common ONNX sanitizer/optimizer.
A sequential rfdn_64x64.onnx integration gate generated and validated float32/float16 TFLite, a native PyTorch package and state dict, TorchScript, Dynamo ONNX, and ExportedProgram artifacts.
0.00003719330.0000433922The branch has also run the modular PyTorch policy/exporter tests, architecture/ownership checks, bulk-runner tests, TensorFlow import blockers, public-contract characterization, pass-level fixtures, quantization/split/report tests, and targeted compatibility-binding checks. The exact commands and checkpoint-specific results are retained in:
docs/flatbuffer_direct_architecture.mddocs/flatbuffer_direct_quick_regression_2026-07-14.mddocs/flatbuffer_direct_pytorch_regression_2026-07-14.mddocs/flatbuffer_direct_handoff_2026-07-14.mdThe final version update was checked with uv lock --check and an import assertion confirming onnx2tf.__version__ == "2.6.3".
This PR is a substantial checkpoint, not the end of the complete refactor plan.
d3net_dnn_double_44.onnx and instrumentation-sensitive nchw.onnx are unsuitable for the 60-second quick profile.Full Changelog: 2.6.2...2.6.3
This PR substantially refactors the TensorFlow-free flatbuffer_direct conversion path around deterministic, indexed, transactional ModelIR passes whil
This PR substantially refactors the TensorFlow-free flatbuffer_direct conversion path around deterministic, indexed, transactional ModelIR passes while preserving the existing CLI/Python API and compatibility entry points.
The main goal is to make converter behavior easier to reason about and safer to extend. Previously, many layout and cleanup rules lived in the central lowerer, repeatedly scanned the complete graph, and depended on implicit call order. This branch gives those rules explicit ownership, stable ordering, shared graph/layout state, transactional validation, focused characterization tests, and observable pass diagnostics.
The branch contains 202 commits and changes 104 files. Much of the apparent diff size comes from mechanically extracting existing rule implementations and their large inline tests from the central lowerer into dedicated pass and test modules.
ModelIRPassState so repeated pass groups reuse one ModelIRGraphIndex and one LayoutState instead of rebuilding graph maps for every rule.LayoutState and synchronized it through indexed graph mutations.lower_from_onnx2tf.py while moving implementation ownership to focused pass-family modules.The branch extracts and/or migrates the following families from the central lowerer into dedicated modules under onnx2tf/tflite_builder/passes/:
Each migrated family follows a staged process:
tflite_builder/reporting.py, retaining thin compatibility wrappers in the lowerer.evaluation_pass rule or the hard max_abs_error <= 1e-1 contract;-tb flatbuffer_direct conversion and -cotof remain TensorFlow-free. Optional TensorFlow exporters remain behind their existing optional boundary.2.6.1 to 2.6.2; uv.lock remains consistent.All local commands were run in the uv environment. Model conversion and inference were executed sequentially with at most one inference subprocess at a time.
1170 passed, 5 deselected, 2 warnings in 162.39suv lock --check passed.onnx2tf.__version__ == "2.6.2".git diff --check passed.birdnet.onnx: max_abs_error=6.4849853515625e-05imageclassifier.onnx: 6.67572021484375e-06model_resnet15x224_swish-072.onnx: 7.43865966796875e-05resnet18-v1-7.onnx: 7.152557373046875e-06keras_rnn.onnx: 6.258487701416016e-07pose_estimation_mediapipe_2023mar.onnx: 0.05517578125pose_estimation_mediapipe_2023mar_int8bq.onnx: 0.0443115234375evaluation_pass=true, and satisfied max_abs_error <= 1e-1.Formal tier-wide warm-up/three-run median timing and peak-RSS gates are not claimed as complete in this PR.
881b714);double_gru.onnx (known first-bad commit 94af81b);model1.onnx and fast_acvnet_generalization_opset16_192x320.onnx;Detailed implementation and validation history is recorded in:
docs/flatbuffer_direct_architecture.mddocs/flatbuffer_direct_handoff_2026-07-12.mddocs/flatbuffer_direct_handoff_2026-07-13.mdFull Changelog: 2.6.1...2.6.2
This pull request continues the flatbuffer_direct refactoring after #945 . It focuses on making the TensorFlow-free ONNX-to-TFLite path more reliable,
This pull request continues the flatbuffer_direct refactoring after #945. It focuses on making the TensorFlow-free ONNX-to-TFLite path more reliable, easier to reason about, and less expensive to maintain while preserving the existing CLI, Python API, artifact naming, report schemas, and conversion behavior.
The branch contains 51 commits and changes 76 files. The work is intentionally incremental: it improves the current conversion pipeline and records the remaining architectural work, but it does not claim that the full long-term refactoring plan is complete.
The former combined op_builders/quantized.py implementation has been removed. Quantized lowering is now organized into dedicated modules:
dynamic_quantize.pyquantize_linear.pyqlinear_binary.pyqlinear_activation.pyqlinear_concat.pyqlinear_conv.pyqlinear_fc.pyqlinear_pool.pyconv_integer.pyquantized_common.py for shared quantization, shape/signature, padding, and requantization primitivesquantized_common.py contains no build_* entry point. Public builder imports and registry dispatch remain compatible.
Each mechanical extraction is guarded by a normalized ModelIR fingerprint covering operators, tensors, constant buffers, options, versions, shapes, shape signatures, dtypes, and quantization metadata. The fingerprint serializer is shared across the family tests instead of being duplicated in each test module.
DynamicQuantizeLinear now uses nearest-even rounding and performs ROUND(x / scale) before adding the integer zero point. This avoids both the previous half-up behavior and FLOAT32 precision loss when a large zero point was added before rounding.
QLinearAdd, QLinearSoftmax, QLinearConv, and related Q/DQ paths retain explicit ONNX-compatible requantization where LiteRT fixed-point tie handling could otherwise differ by one quantum and become amplified downstream.
The direct MatMulInteger/QLinearMatMul implementation continues to follow portable INT32 ONNX semantics. Host-specific ONNX Runtime U8S8 saturation behavior is documented rather than emulated in the converter.
This branch adds or strengthens generic handling for:
The changes use semantic graph properties rather than model-name-specific rules.
Dynamic rank-4 GridSample is lowered with builtin LiteRT operators. The implementation supports bilinear/linear and nearest interpolation, zeros and border padding, both align-corners modes, runtime spatial dimensions, and ONNX Runtime-compatible NaN coordinate handling.
This recovered the large encoder.onnx path while keeping all 24 GridSample nodes builtin and avoiding a TensorFlow dependency or unresolved custom op.
The direct branch now releases the unused legacy GraphSurgeon graph before ModelIR lowering, avoids unnecessary lifetime extension of large serialization objects, and passes a managed ONNX model path to the isolated evaluator instead of pickling a complete protobuf through a multiprocessing pipe.
Evaluation remains sequential. At most one isolated inference subprocess is active at a time; no ProcessPool or parallel model worker was introduced.
Static-input evaluation may use the standard LiteRT delegate. If delegate preparation fails, the evaluator retries once with the builtin interpreter, still sequentially. Dynamic-input crash isolation remains delegate-free.
The changes recover or promote multiple Tier 0-4 models, including:
onnx_dense_optimized.onnx and its byte-identical counterpart;modnet_old.onnx;rf-detr-nano.onnx;LibreRFDETRn.onnx;Remaining active non-passes have normalized causes instead of unexplained skips. These include invalid source geometry/ranks, documented ONNX Runtime U8S8 kernel divergence, quantization-boundary amplification, and discontinuous TopK/GridSample sensitivity. The two DEIM variants are recorded as accepted successes per explicit project direction.
No global tolerance relaxation, model-name lowering rule, or forced index ordering was added.
The default flatbuffer_direct conversion path and direct accuracy evaluation remain TensorFlow-free. Architecture tests cover the new operation-family and shared quantization modules. TensorFlow remains optional and is required only for explicitly requested TensorFlow/SavedModel/Keras-style outputs or the tf_converter backend.
No new dependency was added. All implementation and verification work uses the existing uv environment.
The number 2,000 is an ONNX graph node/operation threshold, not a source-file line limit:
There is no 2,000-line acceptance gate for production or test modules. Modules are split according to ownership, coupling, and reviewability.
The managed Tier 0-4 profile currently records:
missing_tflite_report entries;tflite_fail entries;All 26 active non-passes have an explicit normalized cause or expected invalid/custom-runtime classification.
Tier 5 is retained as historical/final-phase coverage and is not used as an early refactoring gate.
The latest full core-environment regression was executed with one pytest process and no parallel workers:
985 passed, 5 deselected, 2 warnings in 121.63s
Additional focused verification includes:
26 passed;12 passed, 760 deselected;12 passed, 761 deselected;-cotof runs for recovered and diagnosed models;uv lock --check and an import-time version assertion for the 2.6.1 version update.The five deselected tests require optional environments unavailable in the core uv setup:
The two warnings are expected FLOAT16 overflow warnings in existing edge-case tests.
The architecture and handoff documents record:
The next planned implementation step is to extract coverage/correspondence reporting from lower_from_onnx2tf.py, followed by staged ownership of its large layout-pass collection. That work is deliberately not partially included in this PR.
The package, lockfile, and Docker examples are updated consistently from 2.6.0 to 2.6.1.
Full Changelog: 2.6.0...2.6.1
This PR substantially refactors and hardens the flatbuffer_direct backend while preserving the existing CLI/Python API and artifact conventions. It tu
This PR substantially refactors and hardens the flatbuffer_direct backend while preserving the existing CLI/Python API and artifact conventions. It turns the direct converter into a more explicit, testable pipeline, improves shape/layout/control-flow handling across a broad ONNX corpus, expands TensorFlow-free validation, and modularizes the PyTorch export path.
The branch contains 126 focused commits and changes 120 files. The implementation intentionally stays within the existing dependency set and keeps normal flatbuffer_direct conversion and -cotof evaluation independent of TensorFlow.
The direct backend had accumulated a large number of interacting, order-sensitive rules. A local fix could easily invalidate unrelated models, producer/consumer relationships were repeatedly reconstructed, layout assumptions were spread across individual operator lowerers, and large characterization tests consumed significant review and maintenance effort.
This work establishes clearer internal contracts and adds regression evidence before and alongside behavioral changes. The goal is not merely to make individual sample models pass, but to make future changes easier to reason about and safer to validate.
onnx2tf/tflite_builder/core/.The resulting flow is substantially closer to:
request -> ONNX analysis/normalization -> conversion session and graph index -> ordered passes -> operator lowering -> ModelIR validation/cleanup -> requested exporters
flatbuffer_direct correctness improvements-kat axis preservation.If branches.Squeeze sees a non-singleton extent from the inactive path.Where lowering so a rank-1 singleton condition uses SELECT_V2 when prefix-style SELECT broadcasting is invalid.-cotof evaluation TensorFlow-free.If, integer MatMul, optional tensors, unknown-rank Conv, Microsoft contrib operators, and graph ordering.Y, Y_h, and Y_c outputs.Representative recovered regressions include:
superpoint_lightglue_end2end_fused_cpu.onnx: maximum absolute error 1.946091651916504e-05.silero_vad.onnx with -kat input state sr: evaluation_pass=true, maximum absolute error 1.375097781419754e-06, RMSE 1.222495367934982e-07.-ois, and -kat options.flatbuffer_direct backend selection behavior.2.6.0.The final relevant local suite was executed with uv:
uv run pytest -q \
tests/test_tflite_builder_direct.py \
tests/test_accuracy_evaluator_seeded_input.py \
tests/test_onnx_reference_compat.py \
tests/test_flatbuffer_direct_bulk_runner.py \
-k 'not test_tflite_backend_matrix_add and not test_tflite_backend_matrix_hardswish_rewrite_on_off and not test_tf_converter_resize_cubic_avoids_flex_resize_bicubic and not test_tf_converter_resize_cubic_honors_cubic_coeff_a and not test_flatbuffer_direct_group_norm_alias_builtin_conversion'
796 passed, 5 deselected, 2 warnings
Additional focused validation included:
-cotof verification of the recovered Silero model.2.6.0.The five deselected tests are environment-specific rather than code regressions:
tf_converter tests require the optional TensorFlow/tf-keras environment, which is intentionally absent from the core environment._PyCode_GetExtra).The two warnings are existing NumPy float16 overflow warnings from tests that intentionally cast extreme values.
silero_vad (1).onnx references 14 lexical captures that do not exist in the serialized model, including all four LSTM weight/bias captures. It cannot produce a trustworthy ONNX reference result and remains documented as an input-model defect.inverse_11.onnx contains a 224x224 matrix inverse. The current exact direct lowering is intentionally limited to matrices up to 16x16; larger matrices fall back to ONNX_INVERSE, which stock LiteRT cannot execute. The observed reference matrices are nearly singular, so a low-order approximation would not satisfy the required absolute-error bound and has not been introduced.docs/flatbuffer_direct_architecture.md for scope and design.onnx2tf/tflite_builder/core/ for contracts, session state, graph/layout indexing, pass management, and validation.onnx2tf/tflite_builder/op_families/, op_registry.py, and representative op_builders/ changes.onnx2tf/utils/onnxruntime_compat.py, onnx_graph_repair.py, and accuracy_evaluator.py for TensorFlow-free validation.126 commits
120 files changed
34,855 insertions
11,788 deletions
Full Changelog: 2.5.2...2.6.0
This PR improves flatbuffer_direct strict integer quantization so more mixed-dtype and detection-postprocess-adjacent graphs can be converted without
This PR improves flatbuffer_direct strict integer quantization so more mixed-dtype and detection-postprocess-adjacent graphs can be converted without producing invalid full-integer artifacts. It also keeps the int8 and full-int8 outputs available when optional int16-activation variants are not supported by LiteRT kernels.
GATHER_ND, SCATTER_ND, TOPK_V2, ARG_MAX, EXP, CAST, LESS, LOGICAL_NOT, WHERE, and SHAPE.SCATTER_ND, because LiteRT does not support INT16 updates for that kernel. The int8 and full-int8 variants still validate and are emitted.Before:
-oiqt command when *_integer_quant_with_int16_act.tflite validation reached SCATTER_ND:
Updates of type 'INT16' are not supported by scatter_nd.After:
python -m onnx2tf -i /tmp/onnx2tf_yolox_quant/yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx -o /tmp/onnx2tf_yolox_quant/pre_nms_quant_out_after_skip -oiqt -cind images calib_data_416x416_n200.npy "[[[[0.0,0.0,0.0]]]]" "[[[[1.0,1.0,1.0]]]]" completed successfully.yolox_nano_ti_lite_26p1_41p8_pre_nms_integer_quant.tfliteyolox_nano_ti_lite_26p1_41p8_pre_nms_full_integer_quant.tfliteSCATTER_ND: LiteRT SCATTER_ND does not support INT16 updates.Validation run locally:
pytest -q tests/test_strict_integer_quantization.py tests/test_tensor_buffer_builder.py tests/test_tflite_builder_direct.py::test_flatbuffer_direct_integer_quantized_smokeruff check onnx2tf/tflite_builder/quantization.py onnx2tf/tflite_builder/__init__.py tests/test_strict_integer_quantization.py tests/test_tflite_builder_direct.pynpx --yes pyright onnx2tf/tflite_builder/quantization.py onnx2tf/tflite_builder/__init__.py tests/test_strict_integer_quantization.py tests/test_tflite_builder_direct.pypython -m py_compile onnx2tf/onnx2tf.py onnx2tf/tflite_builder/__init__.py onnx2tf/tflite_builder/quantization.py tests/test_strict_integer_quantization.pyRelated issue: #929
onnx2tf \
-i yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx \
-coion \
-oiqt \
-cind "images" "./calib_data_416x416_n200.npy" "[[[[0.485,0.456,0.406]]]]" "[[[[0.229,0.224,0.225]]]]"
onnx2tf \
-i yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx \
-coion \
-oiqt \
-qt per-tensor \
-cind "images" "./calib_data_416x416_n200.npy" "[[[[0.485,0.456,0.406]]]]" "[[[[0.229,0.224,0.225]]]]"
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.5.1...2.5.2
This PR fixes and hardens flatbuffer_direct integer quantization behavior.
This PR fixes and hardens flatbuffer_direct integer quantization behavior.
Related context:
flatbuffer_direct integer quantization previously used fixed / pseudo I/O qparams and could emit artifacts that were not strict full-integer quantized models.The changes make flatbuffer_direct strict integer quantization calibration-based, validate generated quantized TFLite files with LiteRT, and expand strict quantized operator coverage.
flatbuffer_direct strict integer quantization and collect tensor ranges using LiteRT Interpreter.SOFTMAX, LOGISTIC, and TANH.full_integer_quant artifacts.PRELU, and TRANSPOSE_CONV.tf_keras -> tensorflow.compat.v2 import chain in test_tflite_builder_direct.py.test_tflite_builder_direct.py.Before:
flatbuffer_direct -oiqt could use fixed qparams such as uint8 scale=1/255, zero_point=128.test_tflite_builder_direct.py imported common_functions only for external-data detection, which could trigger tf_keras import and fail when tensorflow.compat was unavailable.After:
flatbuffer_direct -oiqt requires representative calibration data and writes strict quantized artifacts only after LiteRT allocation validation succeeds.test_tflite_builder_direct.py no longer triggers the tf_keras import chain for check_model_has_external_data.Validation performed:
pytest -q tests/test_strict_integer_quantization.py tests/test_tensor_buffer_builder.py tests/test_tflite_builder_direct.py::test_flatbuffer_direct_integer_quantized_smoke
ruff check tests/test_tflite_builder_direct.py tests/test_strict_integer_quantization.py onnx2tf/tflite_builder/quantization.py
npx --yes pyright tests/test_strict_integer_quantization.py tests/test_tflite_builder_direct.py
#941 #942
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.5.0...2.5.1
Update Python dependency pins for TensorFlow 2.21, tf-keras 2.21, Keras 3.15, ONNX Runtime 1.26, OpenCV 4.13, NumPy 2.2.6, ml-dtypes 0.5.4, h5py 3.14,
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.4.3...2.5.0
stop re-transposing CONV_3D filters in the SavedModel exporter; flatbuffer-direct ModelIR already stores them as [d, h, w, in, out], which matches tf.
CONV_3D filters in the SavedModel exporter; flatbuffer-direct ModelIR already stores them as [d, h, w, in, out], which matches tf.nn.conv3dCONV_3D_TRANSPOSE filters for the same reason; ModelIR stores [d, h, w, out, in], matching tf.nn.conv3d_transpose.venv/bin/python -m py_compile onnx2tf/tflite_builder/saved_model_exporter.py tests/test_tflite2sm_phase1.pyCONV_3D and CONV_3D_TRANSPOSE SavedModel export/load/inference, both matching TensorFlow native outputs with np.testing.assert_allclose(..., rtol=1e-5, atol=1e-5)Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.4.2...2.4.3
Support for ScatterElements with reduction="add". Added support in both conversion paths.
Support for ScatterElements with reduction="add".
Added support in both conversion paths.
tf_converter
tf.tensor_scatter_nd_add when reduction="add"tf.tensor_scatter_nd_update path for reduction="none"flatbuffer_direct
reduction="add" as output = data + SCATTER_ND(indices, updates, data_shape)reduction="none"Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.4.1...2.4.2
prefer the identity permutation for NHWC-preserved InstanceNormalization inputs when candidate permutation errors are effectively tied
Refs #930
The issue attachment was downloadable and the reported command path was exercised locally. In this scratch environment, that full model still fails later in a downstream Conv shape mismatch, so this PR focuses on the InstanceNormalization permutation tie-break described in the issue rather than claiming full model conversion success.
A synthetic InstanceNormalization conversion completed in the side-run scratch environment before review. A follow-up tf_converter smoke after review reached TensorFlow/TFLite converter logging but hung in this local environment, so it was terminated without changing the repo.
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.4.0...2.4.1
This PR switches the default TFLite backend from tf_converter to flatbuffer_direct and aligns the surrounding API, CLI, tests, helper scripts, and doc
This PR switches the default TFLite backend from tf_converter to flatbuffer_direct and aligns the surrounding API, CLI, tests, helper scripts, and docs with that behavior.
tflite_backend in onnx2tf.convert() and the CLI to flatbuffer_directtf_converter available as an explicit compatibility pathtf_converter selectionflatbuffer_direct as the current defaulttf_converter coverageThis makes the faster direct path the out-of-the-box experience, reduces accidental dependency on TensorFlow-backed conversion, and makes the remaining legacy path explicit. It also removes ambiguity in docs and tests around which backend is responsible for SavedModel generation.
pytest -q tests/test_optional_tensorflow.pypytest -q tests/test_tflite2sm_phase1.py -k "flatbuffer_direct_output_saved_model_validation or tflite_direct_input_validation or tflite_direct_input_new_conflict_validation or tflite_direct_input_rejects_mixed_onnx_and_tflite_input"python tests/test_model_convert.py --helpFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.19...2.4.0
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This PR manually ports the optional dependency work from the experimental branch into feat-torch15 and updates the current codebase so that TensorFlow and PyTorch are no longer forced at base install time.
The main goal is to make the default onnx2tf installation lighter and safer for LiteRT / flatbuffer_direct users while keeping TensorFlow-backed and PyTorch-backed features available through explicit extras.
The branch also removes the remaining implicit TensorFlow dependency from -cotof when --tflite_backend flatbuffer_direct is used.
Important compatibility note:
TensorFlow is now fully optional. Users who need SavedModel export, H5 export, Keras v3 export, TFv1 PB export, or tf_converter must install the TensorFlow extra explicitly.
Recommended commands:
uv pip install -U "onnx2tf[tensorflow]"
or
uv sync --extra tensorflow
or
uv sync --all-extras
Without that extra, TensorFlow-backed features now fail fast with an explicit install hint instead of importing TensorFlow eagerly at module import time.
Additional note for reviewers: on the current flatbuffer_direct path, a root SavedModel is still written by default unless --disable_model_save is specified. --flatbuffer_direct_output_saved_model is only needed when the intent is to request/guarantee the explicit flatbuffer_direct SavedModel export behavior, especially for split output handling. In all such cases, TensorFlow extra installation is still required because SavedModel generation itself is TensorFlow-backed.
This PR includes the following functional changes.
tensorflow and tf-keras were moved out of the base dependency set into onnx2tf[tensorflow].import onnx2tf and python -m onnx2tf --help no longer require TensorFlow.onnx2tf[torch].torch==2.11.0 and uses the CPU wheel index in the uv configuration.flatbuffer_direct -cotof no longer depends on TensorFlow.--tflite_backend flatbuffer_direct, -cotof now means the TensorFlow-free report path only.-cotof flow, so -cotof itself no longer forces TensorFlow there.tf_converter keeps the existing TensorFlow-dependent behavior.onnx2tf[tensorflow].uv-based install commands.PYTHONPATH / LD_LIBRARY_PATH contamination now return a clearer diagnostic instead of a raw traceback.flatbuffer_direct -cotof behavior.Before:
import onnx2tf and some CLI entrypoints could fail early if TensorFlow / PyTorch were not available.flatbuffer_direct -cotof still implicitly required TensorFlow because it triggered SavedModel validation.After:
onnx2tf[tensorflow].onnx2tf[torch].import onnx2tf and CLI help work without TensorFlow / PyTorch.flatbuffer_direct -cotof works without TensorFlow and stays on the TensorFlow-free evaluation/report path.Validation executed on this branch:
pytest -q tests/test_optional_tensorflow.py tests/test_optional_pytorch.py
pytest -q tests/test_flatbuffer_direct_op_error_report.py tests/test_pytorch_bulk_runner.py
pytest -q tests/test_tflite_builder_direct.py -k 'flatbuffer_direct_accuracy_report_generation or split_accuracy_report_fail_on_threshold'
pytest -q tests/test_optional_tensorflow.py
pytest -q tests/test_tflite2sm_phase1.py -k 'cotof and (flatbuffer_direct_output_saved_model or tflite_direct_input_saved_model)'
N/A
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.18...2.3.19
remove conversion-test from the CodeQL workflow
conversion-test from the CodeQL workflow2.3.18uv.lock with the cooldown policy in placePYPI_API_TOKENFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.17...2.3.18
remove all remaining download_test_image_data call sites
download_test_image_data call sites-oiqt integer quantization2.3.17download_test_image_data import and deleted the obsolete helper from common_functions.py--test_data_nhwc_path was not provideddummy_onnx_inference / dummy_tf_inference-oiqt so it now exits with a clear error unless --quant_calib_input_op_name_np_data_path (-cind) is provided--test_data_nhwc_path
-oiqt
-cindpython -m py_compile onnx2tf/onnx2tf.py onnx2tf/utils/common_functions.pyrg checks confirming that download_test_image_data and the old auto-download wording no longer remain in the code/docshttps://github.com/PINTO0309/onnx2tf/issues/921
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.15...2.3.16
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This PR improves the flatbuffer_direct PyTorch export path with a focus on feature-level robustness rather than new surface area.
The branch tightens several failure-prone code generation and evaluation paths that were still blocking reliable use of the native PyTorch package flow in more mixed-layout models.
The largest part of this PR hardens native PyTorch code generation for mixed channel-first / channel-last graphs.
In particular, it improves how the exporter rewrites:
This is implemented mainly in onnx2tf/tflite_builder/pytorch_exporter.py and onnx2tf/tflite_builder/_pytorch_exporter_native_codegen_pipeline.py.
The practical effect is that generated PyTorch packages are less likely to produce invalid shape alignment code when the graph mixes NHWC-style public layout flow with internal channel-first execution.
This PR also fixes a pathological case in the raw export canonicalization path.
The affected logic could spend too long in rewrite passes while processing shape-matched channel-first binary and batch-normalization style patterns. The branch adds explicit forward progress handling and more targeted parsing so the canonicalizer no longer gets stuck on these cases.
This includes:
The TFLite accuracy evaluator previously treated allclose mismatches as hard failures even when the numeric metric thresholds were still within acceptable bounds.
This branch changes the evaluation contract so that pass/fail gating matches the PyTorch evaluator more closely:
allclose is stricterThis applies to both the standard evaluator and split-manifest evaluation.
These fixes improve the reliability of the flatbuffer_direct PyTorch workflow in real export scenarios:
Added / updated coverage includes:
2.3.152.3.15This PR is intentionally focused on PyTorch export quality and evaluation consistency. It does not introduce new user-facing flags or broaden the public API.
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.14...2.3.15
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This PR significantly improves the feat-torch10 native PyTorch export path and the flatbuffer_direct workflow, with an emphasis on replacing model-specific fixes with generalized structural logic.
The branch expands native package generation so that difficult ONNX graphs can stay on the native PyTorch backend more reliably, finish export faster, and preserve accuracy across a broad set of real-world models.
A large portion of the work in this branch moves the exporter away from brittle model-name- or op-name-based routing and toward shape-, layout-, and dataflow-based repair rules.
Representative improvements include:
Add, Concat, Resize, Reshape, Slice, pooling, channel-shuffle patterns, and public output bridgesThe practical effect is that many previously fragile native PyTorch conversions now succeed through common structural rules instead of accumulating one-off exceptions.
The branch adds broader fast-path routing and avoid-model-ir detection for native package export. This is especially important for models that are small but were previously taking an unexpectedly long time in write pytorch because they fell into expensive raw canonicalization.
Examples of the improvements in this area include:
A representative case is version-RFB-640.onnx, where the write pytorch phase now returns to a practical runtime while preserving native execution and accuracy.
This branch also improves supporting tooling around export validation:
flatbuffer_direct_bulk_runner utility and its test coverageThe test suite was significantly expanded, especially around the native PyTorch exporter and structural repair logic. This includes:
The main goal of this branch is not just to fix isolated regressions, but to improve the maintainability and reproducibility of the native PyTorch export path.
Instead of continuing to accumulate exact-name patches, the branch pushes the exporter toward a more defensible strategy:
This should reduce future regression risk and make new model support less dependent on ad hoc exceptions.
The branch was validated with native PyTorch output comparison on the following regression set, all passing:
age_googlenet.onnxalike_t_opset11_192x320.onnxAtan_11.onnxbaseline_simplified.onnxbread_180x320.onnxbread_nonfm_180x320.onnxdeeplabv3_mobilenet_v3_large.onnxdetpth_to_space_17.onnxdetr_demo.onnxdigits.onnxefficientformer_l1.onnxFastestDet.onnxhuman_segmentation_pphumanseg_2021oct.onnxiat_llie_180x320.onnxmobilenetv2-10.onnxnanodet-plus-m_416.onnxpidnet_S_cityscapes_192x320.onnxreducemax_softmax_workaround.onnxresnet18-v1-7.onnxrfdn_64x64.onnxshadowformer_istd_160x240.onnxsinet_320_op.onnxswinir-m_64x64_12.onnxts_ad_model.onnxversion-RFB-640.onnxyolox_s.onnxIn addition, targeted exporter tests covering the restored version-RFB-640 and yolox_s fixes were re-run and passed.
Because this branch contains a long sequence of generalization work, the most useful review lens is by subsystem rather than by individual commit:
onnx2tf/tflite_builder/pytorch_exporter.pyonnx2tf/tflite_builder/_pytorch_exporter_native_codegen_pipeline.pyonnx2tf/tflite_builder/accuracy_evaluator.pyonnx2tf/utils/flatbuffer_direct_bulk_runner.pytests/test_pytorch_exporter.pytests/test_flatbuffer_direct_bulk_runner.pytests/test_accuracy_evaluator_seeded_input.pyRecommended focus areas:
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.13...2.3.14
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This PR improves flatbuffer_direct conversion stability, native PyTorch package generation robustness, and constant-folding quality for several previously failing or regressing models.
The main goals of this branch were:
write pytorch regressions in the fast pathThis branch expands the fast precanonicalization repair logic in pytorch_exporter.py to handle several classes of broken generated code more reliably.
Key improvements include:
DepthToSpace-adjacent gather patterns in NHWC/CF bridge codenanodet-plus-m_416 stage-0 max-pool layout corruption, which also removes a severe write pytorch slowdownThese changes were validated against the regressions reported during branch development, including:
alike_t_opset11_192x320efficientformer_l1nanodet-plus-m_416FastestDetage_googlenetThis branch adds new constant-folding coverage in both the TFLite-side lowering flow and the generated Dynamo ONNX sanitization flow.
Highlights:
ScatterND evaluation/foldingReshape foldingAdd/Sub/Mul/DivScatterND and follow-up binary chains on cloned export IR only, so the original IR remains safe for PyTorch package generationThese changes simplify generated artifacts and remove unnecessary constant computation chains in ALIKE-derived models.
lower_from_onnx2tf.py now avoids over-aggressive reciprocal-multiply lowering for precision-sensitive paths that later feed integer Cast consumers.
This preserves exact division semantics where needed, which was necessary to fix descriptor/indexing accuracy regressions in ALIKE without weakening the general constant-division optimization path.
This branch also tightens AveragePool(count_include_pad=0) handling so that:
SAME cases that should already behave correctlyefficientformer_l1This separation was important to fix TFLite accuracy while avoiding PyTorch-package regressions.
The test suite has been extended with focused regression tests for:
ScatterND / constant Reshape / constant binary foldingcount_include_pad=0 handling across SAME, explicit pad, and ceil-mode casesIn addition to the new unit tests, the branch was checked with real model conversions on the flatbuffer_direct path.
Confirmed as passing on this branch:
alike_t_opset11_192x320: ONNX/TFLite pass=True, ONNX/PyTorch pass=Trueefficientformer_l1: ONNX/TFLite pass=True, ONNX/PyTorch pass=TrueFastestDet: ONNX/TFLite pass=True, ONNX/PyTorch pass=Trueage_googlenet: ONNX/TFLite pass=True, ONNX/PyTorch pass=Truenanodet-plus-m_416: ONNX/PyTorch pass=True, and the severe write pytorch slowdown was removedFor nanodet-plus-m_416, the native PyTorch package regression was fixed and the generation time returned to a practical range, but ONNX/TFLite still reports a remaining numerical mismatch on this branch. That issue is not introduced by this PR; the work here focuses on removing the package-generation/runtime regression and preserving previously validated models.
2.3.12 to 2.3.13Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.12...2.3.13
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This branch focuses on improving the practical stability of the flatbuffer_direct pipeline and the native PyTorch export path for transpose-heavy and layout-sensitive models.
The main motivation was a set of real regression cases that were already known to convert successfully in the past, but started to fail or degrade after recent changes. The affected patterns were not isolated to one model family: they spanned PIDNet, SiNet, DepthToSpace-based super-resolution models, detection heads such as FastestDet, and Dynamo ONNX artifacts generated from exported native PyTorch packages.
From a feature-improvement perspective, this PR is not just a collection of one-off fixes. It expands the optimizer and exporter so that more difficult graph topologies can stay on the fast path while preserving layout semantics, output shapes, and numerical parity.
Extended flatbuffer_direct graph optimization coverage for SiNet.
Hardened the native PyTorch export canonicalization and fast-repair path.
pag4 binary alignment regression in generated native packages.DepthToSpace(mode=CRD) NHWC gather repair so the repaired gather slices the channel axis instead of corrupting the spatial height.Improved exporter artifact consistency.
Expanded regression protection.
tests/test_pytorch_exporter.py and tests/test_tflite_builder_direct.py for PIDNet, SiNet, DepthToSpace, FastestDet alias propagation, and Dynamo ONNX output-shape restoration.Updated package versioning.
2.3.11 to 2.3.12.uv.lock accordingly.Representative regression cases addressed by this branch:
rfdn_64x64
ONNX/PyTorch output check failed. reason=Evaluation output shape mismatch after layout alignment. onnx_shape=(1, 3, 256, 256) tflite_shape=(1, 3, 192, 256)ONNX/PyTorch output check complete! ... max_abs=3.8147e-05 rmse=4.11336e-06 cosine=1 pass=TrueFastestDet
max_abs=0.514514 rmse=0.0449929 cosine=0.984069 pass=Falsemax_abs=1.37091e-05 rmse=3.60957e-07 cosine=1 pass=Trueage_googlenet
max_abs=0.787734 rmse=0.28074 cosine=0 pass=Falsemax_abs=2.11e-05 rmse=9.26852e-06 cosine=1 pass=True[] to [1, 8].Validation performed during the branch work included:
pytest coverage for the new exporter and optimizer regressions.age_googlenet, FastestDet, and rfdn_64x64 to confirm that previously successful conversions did not regress.N/A
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.11...2.3.12
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This PR brings the feat-torch7 improvements into main as the 2.3.11 update for the PyTorch export path.
The main focus of this branch is exporter robustness around layout bridges, rank-3 reshape/transpose edge cases, and Dynamo ONNX post-processing. In practice, it makes the generated PyTorch package, TorchScript / ExportedProgram artifacts, and sanitized Dynamo ONNX outputs much more stable for models that mix channel-first runtime values with channel-last logical layouts.
_rewrite_generated_model_source_for_exported_program() from undoing already-correct materializations for tensors that must stay in channel-last form before reshape.model.py semantically aligned with the original ONNX / lowered IR instead of reintroducing layout mismatches during artifact export.Mul + Add + Clip -> HardSigmoid style patterns.This branch adds focused tests for:
yolox_s output-shape preservation and existing bread helper behavioryolox_s.onnxThe generated *_dynamo.onnx could lose its final output shape metadata after sanitization. This branch preserves the expected output shape and adds regression coverage for it.
iat_llie_180x320.onnxThe generated PyTorch package could drift badly from ONNX because a required NHWC materialization was rewritten back into a raw channel-first alias before a rank-3 reshape.
After the fix, the conversion command below now passes the ONNX/PyTorch comparison:
python -m onnx2tf -i iat_llie_180x320.onnx -o /tmp/iat_llie_fix2_run -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoep
Observed result after the fix:
max_abs=4.76837e-07rmse=1.40854e-07cosine=1pass=TrueI validated the branch with targeted exporter / sanitizer regressions, including:
pytest -q tests/test_pytorch_exporter.py -k 'runtime_cf_alias_before_rank3_reshape or exported_program_preserves_rank3_reshape_materialize_without_model_ir or preserves_transpose_rank3_output_shape or preserves_atan_rank4_output_shape'
pytest -q tests/test_pytorch_exporter.py -k 'exported_program_preserves_rank3_reshape_materialize_without_model_ir or runtime_cf_alias_before_rank3_reshape or preserves_transpose_rank3_output_shape or preserves_atan_rank4_output_shape or preserves_yolox_output_shape_when_model_is_available or helper_postprocess_is_noop_for_bread_nonfm_when_model_is_available'
Both validation sets passed locally, and the branch also includes the version bump to 2.3.11 in:
onnx2tf/__init__.pypyproject.tomlREADME.mdMost of the code volume in this PR comes from two areas:
onnx2tf/tflite_builder/pytorch_exporter.pytests/test_pytorch_exporter.pyThe intention is not to change the overall export architecture, but to make the existing PyTorch export and Dynamo ONNX cleanup pipeline much more reliable for real-world bridge-heavy models.
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.10...2.3.11
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This branch hardens the native PyTorch export path so the raw-export canonicalization logic remains consistent across multiple previously fragile model families instead of fixing them one-by-one and reintroducing regressions elsewhere.
The main goal of this PR is to make the flatbuffer_direct + -cotof + -fdopt + -fdots + -fdodo + -fdoep workflow stable on a mixed regression set that repeatedly exposed layout-bridging, reshape, resize, transpose-conv, and output-boundary issues in the generated PyTorch package.
pytorch_exporter.py for channel-first/channel-last alias handling, including:
tests/test_pytorch_exporter.py and tests/test_tflite_builder_direct.py to lock in the previously fragile canonicalization paths.The recent failures were not isolated bugs in a single model. They came from the same class of canonicalization decisions being correct for one topology and incorrect for another. This PR treats the problem as a cross-model consistency issue and tightens the rewrite rules so they are applied only when the exporter has enough evidence about the intended logical layout and shape.
That makes the native package generation path more predictable and reduces the cycle of model-specific hotfixes that accidentally break another regression target.
Targeted regression tests were updated/added and the generated source was revalidated with py_compile.
In addition, the following end-to-end conversions were re-run with native export enabled:
python -m onnx2tf -i yolox_s.onnx -o tmp_yolox_s -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoeppython -m onnx2tf -i human_segmentation_pphumanseg_2021oct.onnx -o tmp_human_segmentation_pphumanseg_2021oct -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoeppython -m onnx2tf -i ts_ad_model.onnx -o tmp_ts_ad_model -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoeppython -m onnx2tf -i pidnet_S_cityscapes_192x320.onnx -o tmp_pidnet_S_cityscapes_192x320 -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoeppython -m onnx2tf -i version-RFB-640.onnx -o tmp_version-RFB-640 -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoeppython -m onnx2tf -i iat_llie_180x320.onnx -o tmp_iat_llie_180x320 -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoeppython -m onnx2tf -i bread_180x320.onnx -o tmp_bread_180x320 -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoeppython -m onnx2tf -i bread_nonfm_180x320.onnx -o tmp_bread_nonfm_180x320 -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoepFor all eight models:
ONNX/TFLite checks passedONNX/PyTorch checks passedThe largest changes are in the native PyTorch exporter and its regression coverage. The accompanying test updates are intentionally extensive because the core issue here is regression prevention across multiple model topologies, not just a single bug fix.
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.9...2.3.10
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This PR improves the feat-torch5 line of work before merging back into main, with a focus on flatbuffer_direct lowering quality and generated PyTorch package reliability.
The branch combines structural cleanup in the PyTorch native codegen path with a broad set of behavior-preserving export fixes and graph optimizations that were validated against real-world models such as PIDNet, YOLOX, MobileBERT, SwinIR, ts_ad_model, human_segmentation_pphumanseg, iat_llie, and bread.
pytorch_exporter.py into private pipeline/common modules.onnx2tf/tflite_builder/pytorch_exporter.py as the stable entrypoint while moving the heavy orchestration logic out of the main file.flatbuffer_direct lowering improvementsRESHAPE -1 dimensions when the input shape is fully known.0 copy-dim semantics when allowZero=False, while keeping allowZero=True behavior unchanged.RESHAPE chains more aggressively after late shape reconciliation.Div(variable, constant) patterns into Mul(variable, reciprocal_constant) when safe.Mul chains into a single Mul.Transpose/Reshape chains in multiple real models.load_state_dict typing so it matches the PyTorch base class contract and avoids Pylance override warnings.These changes make flatbuffer_direct outputs cleaner and more predictable, while also improving the reliability of the optional generated PyTorch artifacts (jit.pt, dynamo.onnx, ep.pt2).
In practice, this branch reduces redundant graph structure, resolves several export-time crashes, and improves maintainability of the PyTorch native codegen path without changing the public entrypoint.
Representative validation on this branch included:
pytest -q tests/test_tflite_builder_direct.py -k flatbuffer_directpytest -q tests/test_pytorch_exporter.py -k 'not when_model_is_available and not convert_flatbuffer_direct'tests/test_pytorch_exporter.pytests/test_tflite_builder_direct.pypidnet_S_cityscapes_192x320.onnxyolov7_tiny_head_0.768_post_480x640.onnxyolox_s.onnxlite_model_mobilebert_1_metadata_1.onnxhuman_segmentation_pphumanseg_2021oct.onnxswinir-m_64x64_12.onnxts_ad_model.onnxiat_llie_180x320.onnxbread_180x320.onnxflatbuffer_direct lowering, and PyTorch artifact/export cleanup.pytorch_exporter.py entrypoint remains intact even though the native codegen internals were reorganized.Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.8...2.3.9
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This PR continues the feat-torch* line of work by tightening the native export path around flatbuffer_direct and the generated PyTorch artifacts.
The branch focuses on practical reliability improvements for the direct ModelIR-based path that is intended to become the main conversion flow:
finder.onnxflatbuffer_direct lowering and validation behavior for shape/index/control-flow related edge cases2.3.8From a feature-improvement perspective, the main outcome is that flatbuffer_direct now behaves more like a first-class backend end-to-end: direct export, native PyTorch package generation, accuracy checking, and related validation/reporting are more consistent and substantially more robust on non-trivial graphs.
Key improvements included in this branch:
Fixed native torch export regressions affecting flatbuffer_direct output packages.
GridSample-adjacent accuracy regression where coordinate helper reshapes were reordered incorrectly in the generated packageStrengthened flatbuffer_direct lowering and runtime shape handling.
Expanded test coverage for direct export and native artifact generation.
flatbuffer_direct tests covering direct lowering, validation, and numerical/runtime edge casesUpdated release and CI metadata.
2.3.8A representative real-model regression fixed in this branch was the finder.onnx direct export flow.
Command:
onnx2tf -i finder.onnx -o tmp_finder -tb flatbuffer_direct -cotof -fdopt -fdots -fdodo -fdoep
Before:
ONNX/TFLite: passONNX/PyTorch: failmax_abs=0.043901rmse=0.01257cosine=0.999059After:
ONNX/TFLite: passONNX/PyTorch: passmax_abs=6.84522e-08rmse=2.1783e-08cosine=1Validation performed for this branch:
flatbuffer_direct focused test run:
pytest -q tests -k flatbuffer_direct685 passed, 393 deselected, 3 warningsNo related issue was linked for this change set.
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This PR brings the feat-torch3 work to main with a focus on two practical improvements:
flatbuffer_direct correctness, graph cleanup, and validation diagnosticsBefore these changes, several recurrent models still fell back to non-native execution backends, which meant TorchScript / Dynamo ONNX / ExportedProgram artifacts were skipped even when the generated package was otherwise usable. In parallel, flatbuffer_direct still had a few correctness gaps, including recurrent alias repair issues, layout-sensitive flatten behavior, redundant transpose chains, and a major AveragePool border-semantics mismatch that could produce visible accuracy regressions on real models.
The goal of this branch was not just to patch isolated failures, but to make the native export path and the direct flatbuffer path more robust as product features.
REVERSE_V2-based reverse-direction recurrent paths.saved_model / tflite execution backends by keeping supported recurrent graphs on the native backend instead of forcing unrolled primitive fallback in those cases.flatbuffer_direct, including orphan final-step tensor aliases that could leave unbound internal inputs in GRU-derived graphs.flatbuffer_direct, especially the split/concat patterns exposed by res2net50_48w_2s_Opset16.onnx, to reduce unnecessary NCHW/NHWC round-trips.AveragePool(count_include_pad=1) border semantics in flatbuffer_direct by materializing zero padding explicitly before VALID pooling when needed, instead of relying on TFLite SAME behavior that does not match ONNX exactly at the padded border.flatbuffer_direct runtime/lowering tests.Native export behavior before this branch:
execution_backend=saved_model or execution_backend=tfliteNative export behavior after this branch:
execution_backend=nativeflatbuffer_direct accuracy before the AveragePool semantic fix:
onnx2tf -i res2net50_48w_2s_Opset16.onnx -cotof -tb flatbuffer_directmax_abs=0.22682 rmse=0.0500467 cosine=0.999387 pass=Falseflatbuffer_direct accuracy after the fix:
onnx2tf -i res2net50_48w_2s_Opset16.onnx -cotof -tb flatbuffer_directmax_abs=4.76837e-06 rmse=8.87476e-07 cosine=1 pass=TrueValidation executed on this branch:
pytest -q tests/test_tflite_builder_direct.py -k "average_pool_include_pad or average_pool_exclude_pad"
4 passedpytest -q tests -k flatbuffer_direct
668 passed, 352 deselected, 3 warnings> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
This PR advances the feat-torch2 line by making flatbuffer_direct-generated native PyTorch packages much more useful in Torch 2 workflows and by tightening native export/runtime behavior on models that previously fell back, misreported shapes, or failed secondary artifact generation.
From a functional perspective, the branch now goes beyond generating a PyTorch package and TorchScript artifact. It can also emit Torch 2-oriented artifacts directly from the generated native package, while keeping the flatbuffer_direct path aligned with real-model parity and validation needs.
<img width="1390" height="680" alt="image" src="https://github.com/user-attachments/assets/8a78b9f7-1557-4e5a-954d-cff9a4a74d04" />
--flatbuffer_direct_output_dynamo_onnx (-fdodo) so a generated native PyTorch package can emit <model>_dynamo.onnx via torch.onnx.export(..., dynamo=True).--flatbuffer_direct_output_exported_program (-fdoep) so the same package can emit <model>_ep.pt2 via torch.export.save.flatbuffer_direct_output_dynamo_onnx and flatbuffer_direct_output_exported_program.flatbuffer_direct_output_pytorch when any secondary PyTorch artifact is requested, keeping the UX consistent with existing TorchScript behavior.metadata.json, including example input shapes, dynamic-input state, and export errors.shape_hints, test_data_nhwc_path, and per-input custom example data.TopK output shape/signature handling in the flatbuffer_direct lowering path, including cases where k is constant, stale, or runtime-derived.README.md, including dynamic-input guidance and how shape_hints, -cind, and --test_data_nhwc_path apply to the new export modes.2.3.6.onnxscript==0.6.2, which is required by the Dynamo ONNX export path.TopK, parity validation, and related flatbuffer_direct behavior.pytest -q tests -k flatbuffer_direct
651 passed, 317 deselected, 3 warningspidnet_S_cityscapes_192x320.onnx now emits both:
pidnet_S_cityscapes_192x320_dynamo.onnxpidnet_S_cityscapes_192x320_ep.pt2
instead of only writing warning entries to the generated package metadata.Before this branch, the flatbuffer_direct Torch path was useful primarily for package generation and, in some cases, TorchScript export. After this change, the generated native package becomes a practical bridge into broader Torch 2 ecosystems: users can generate Dynamo ONNX and ExportedProgram artifacts directly from the same conversion output, while the lowering/runtime fixes reduce the gap between synthetic smoke tests and real deployment models.
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.5...2.3.6
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct.
This PR significantly expands the feat-torch work around the flatbuffer_direct pipeline, with a primary focus on native PyTorch package generation, direct evaluation, and end-to-end validation.
The branch moves PyTorch export from an experimental wrapper-heavy path toward a practical direct-codegen flow built on ModelIR, with better output readability, stronger validation, and broader compatibility for both ONNX-input and TFLite-input workflows.
ModelIRAdded a path to generate PyTorch model source code from a LiteRT model without going through ONNX. This is a very experimental implementation and has only been tested on a few models.
<img width="1425" height="617" alt="image" src="https://github.com/user-attachments/assets/54e0a68f-2e00-484c-9d35-91484ed5b842" />
flatbuffer_direct ModelIR.onnx2tf/tflite_builder/pytorch_exporter.pyonnx2tf/tflite_builder/pytorch_package_runtime.pytorch.nn.Module-based model.py instead of always falling back to a wrapper package.Model(...)load_model(...)state_dict.pth loading through standard load_state_dict(...)state_dict.pth format was updated to be compatible with standard PyTorch state_dict() keys for native packages.<details><summary>A sample PyTorch model model.py generated from LiteRT</summary>
from __future__ import annotations
from pathlib import Path
from typing import Any, Callable, Dict, Optional
import torch
import torch.nn.functional as F
from .runtime import (
_align_tensor_to_target_shape,
_apply_module_transpose_conv2d,
_normalize_axes,
_reduce_mean,
_resolve_named_input_value,
_shape_list,
load_generated_weights,
)
PACKAGE_DIR = Path(__file__).resolve().parent
INPUT_NAMES = ['onnx____Reshape_0']
OUTPUT_NAMES = ['63']
class _Conv2dBlock(torch.nn.Module):
def __init__(
self,
*,
in_channels: int,
out_channels: int,
kernel_size: tuple[int, int],
stride: tuple[int, int],
padding: tuple[int, int],
dilation: tuple[int, int],
groups: int,
bias: bool,
pad: Optional[list[int]] = None,
activation: str = 'none',
negative_slope: float = 0.2,
pad_mode: str = 'constant',
pad_value: float = 0.0,
) -> None:
super().__init__()
self.conv = torch.nn.Conv2d(
in_channels=in_channels,
out_channels=out_channels,
kernel_size=kernel_size,
stride=stride,
padding=padding,
dilation=dilation,
groups=groups,
bias=bias,
)
self.pad = pad
self.activation = str(activation)
self.negative_slope = float(negative_slope)
self.pad_mode = str(pad_mode)
self.pad_value = float(pad_value)
def forward(self, x: torch.Tensor) -> torch.Tensor:
if self.pad is not None:
x = F.pad(x, self.pad, mode=self.pad_mode, value=self.pad_value)
x = self.conv(x)
if self.activation == 'leaky_relu':
return F.leaky_relu(x, negative_slope=self.negative_slope)
if self.activation == 'relu':
return torch.relu(x)
if self.activation == 'relu6':
return torch.clamp(x, min=0.0, max=6.0)
if self.activation == 'relu_n1_to_1':
return torch.clamp(x, min=-1.0, max=1.0)
if self.activation == 'relu_0_to_1':
return torch.clamp(x, min=0.0, max=1.0)
if self.activation == 'tanh':
return torch.tanh(x)
if self.activation == 'sigmoid':
return torch.sigmoid(x)
return x
class Model(torch.nn.Module):
const_ConvTranspose_12_transpose_conv_bias: torch.Tensor
const_ConvTranspose_15_transpose_conv_bias: torch.Tensor
const_ConvTranspose_18_transpose_conv_bias: torch.Tensor
const_decoder_decoder_lin_0_bias: torch.Tensor
const_onnx__MatMul_73: torch.Tensor
const_onnx__MatMul_74: torch.Tensor
def __init__(self, *, device: str | None = None, eval_mode: bool = True, load_weights: bool = True):
super().__init__()
self.input_names = list(INPUT_NAMES)
self.output_names = list(OUTPUT_NAMES)
self.conv_block_0 = _Conv2dBlock(
in_channels=1,
out_channels=64,
kernel_size=(1, 9),
stride=(1, 1),
padding=(0, 4),
dilation=(1, 1),
groups=1,
bias=True,
pad=None,
activation='relu',
negative_slope=0.2,
pad_mode='constant',
pad_value=0.0,
)
self.conv_block_1 = _Conv2dBlock(
in_channels=64,
out_channels=128,
kernel_size=(1, 5),
stride=(1, 1),
padding=(0, 2),
dilation=(1, 1),
groups=1,
bias=True,
pad=None,
activation='relu',
negative_slope=0.2,
pad_mode='constant',
pad_value=0.0,
)
self.conv_block_2 = _Conv2dBlock(
in_channels=128,
out_channels=256,
kernel_size=(1, 3),
stride=(1, 1),
padding=(0, 0),
dilation=(1, 1),
groups=1,
bias=True,
pad=[2, 2, 0, 0, 0, 0, 0, 0],
activation='none',
negative_slope=0.2,
pad_mode='constant',
pad_value=0.0,
)
self.conv_transpose2d_0 = torch.nn.ConvTranspose2d(
in_channels=256,
out_channels=128,
kernel_size=(1, 3),
stride=(1, 1),
padding=(0, 0),
bias=False,
)
self.conv_transpose2d_1 = torch.nn.ConvTranspose2d(
in_channels=128,
out_channels=64,
kernel_size=(1, 5),
stride=(1, 1),
padding=(0, 0),
bias=False,
)
self.conv_transpose2d_2 = torch.nn.ConvTranspose2d(
in_channels=64,
out_channels=64,
kernel_size=(1, 9),
stride=(1, 1),
padding=(0, 0),
bias=False,
)
self._init_constants()
if load_weights:
load_generated_weights(
model=self,
package_dir=PACKAGE_DIR,
device=device,
)
elif device is not None:
self.to(device)
if eval_mode:
self.eval()
def _init_constants(self) -> None:
self.register_buffer('const_ConvTranspose_12_transpose_conv_bias', torch.zeros([1, 1, 1, 128], dtype=torch.float32), persistent=True)
self.register_buffer('const_ConvTranspose_15_transpose_conv_bias', torch.zeros([1, 1, 1, 64], dtype=torch.float32), persistent=True)
self.register_buffer('const_ConvTranspose_18_transpose_conv_bias', torch.zeros([1, 1, 1, 64], dtype=torch.float32), persistent=True)
self.register_buffer('const_decoder_decoder_lin_0_bias', torch.zeros([1, 66, 1], dtype=torch.float32), persistent=True)
self.register_buffer('const_onnx__MatMul_73', torch.zeros([66, 2], dtype=torch.float32), persistent=True)
self.register_buffer('const_onnx__MatMul_74', torch.zeros([2, 66], dtype=torch.float32), persistent=True)
def _device(self) -> torch.device:
for parameter in self.parameters():
return parameter.device
for buffer in self.buffers():
return buffer.device
return torch.device('cpu')
def _max_pool2d_same(self, x: torch.Tensor, *, kernel_size: tuple[int, int], stride: tuple[int, int]) -> torch.Tensor:
pad_h_total = max(int(kernel_size[0]) - int(stride[0]), 0)
pad_w_total = max(int(kernel_size[1]) - int(stride[1]), 0)
pad_top = pad_h_total // 2
pad_bottom = pad_h_total - pad_top
pad_left = pad_w_total // 2
pad_right = pad_w_total - pad_left
x = F.pad(x, [pad_left, pad_right, pad_top, pad_bottom], mode='constant', value=float('-inf'))
return F.max_pool2d(x, kernel_size=kernel_size, stride=stride)
def forward(self, *args: torch.Tensor, **kwargs: torch.Tensor) -> Any:
if len(args) > 0 and len(kwargs) > 0:
raise RuntimeError('Use either positional inputs or keyword inputs, not both.')
if len(kwargs) > 0:
onnx____Reshape_0 = _resolve_named_input_value(kwargs, 'onnx____Reshape_0')
else:
if len(args) != 1:
raise RuntimeError(f'Input arity mismatch. expected={len(self.input_names)} actual={len(args)}')
onnx____Reshape_0 = args[0]
input = torch.reshape(onnx____Reshape_0, [int(v) for v in _shape_list(torch.as_tensor([-1, 1, 64], dtype=torch.int32, device=self._device()))])
Conv_2_conv1d_input_nchw2d = torch.reshape(input, [int(v) for v in _shape_list(torch.as_tensor([1, 1, 1, 64], dtype=torch.int32, device=self._device()))])
Conv_2_input = torch.reshape(Conv_2_conv1d_input_nchw2d, [int(v) for v in _shape_list(torch.as_tensor([1, 1, 64, 1], dtype=torch.int32, device=self._device()))])
Conv_4_input = self.conv_block_0(Conv_2_input.permute(0, 1, 3, 2).contiguous())
Conv_6_input = self.conv_block_1(Conv_4_input)
Conv_6_output = self.conv_block_2(Conv_6_input)
input_31 = torch.reshape(Conv_6_output.permute(0, 3, 1, 2).contiguous(), [int(v) for v in [1, 66, 256]])
input_35 = torch.relu(input_31)
_tmp_x_11 = input_35
_tmp_y_11 = self.const_onnx__MatMul_73
_tmp_x_11 = _tmp_x_11.transpose(-1, -2)
onnx____Add_49 = _align_tensor_to_target_shape(torch.matmul(_tmp_x_11, _tmp_y_11), [1, 256, 2])
onnx____MatMul_50 = torch.add(torch.as_tensor([0.05643776059150696, -0.00908423587679863], dtype=torch.float32, device=self._device()).reshape([1, 1, 2]), onnx____Add_49)
_tmp_x_13 = self.const_onnx__MatMul_74
_tmp_y_13 = onnx____MatMul_50
_tmp_x_13 = _tmp_x_13.transpose(-1, -2)
_tmp_y_13 = _tmp_y_13.transpose(-1, -2)
onnx____Add_52 = torch.matmul(_tmp_x_13, _tmp_y_13)
onnx____ConvTranspose_53 = torch.add(self.const_decoder_decoder_lin_0_bias, onnx____Add_52)
ConvTranspose_12_convtranspose1d_input_nchw2d = torch.reshape(onnx____ConvTranspose_53, [int(v) for v in _shape_list(torch.as_tensor([1, 1, 66, 256], dtype=torch.int32, device=self._device()))])
ConvTranspose_12_output = _apply_module_transpose_conv2d(self.conv_transpose2d_0, ConvTranspose_12_convtranspose1d_input_nchw2d, target_shape=[1, 1, 68, 128], fallback_shape=[1, 1, 68, 128], fused='NONE')
ConvTranspose_12_output_nhwc_cropped = ConvTranspose_12_output[0:1, 0:1, 2:66, 0:128]
ConvTranspose_12_output_nhwc_bias = torch.add(ConvTranspose_12_output_nhwc_cropped, self.const_ConvTranspose_12_transpose_conv_bias)
ConvTranspose_15_input = torch.relu(ConvTranspose_12_output_nhwc_bias)
ConvTranspose_15_output = _apply_module_transpose_conv2d(self.conv_transpose2d_1, ConvTranspose_15_input, target_shape=[1, 1, 68, 64], fallback_shape=[1, 1, 68, 64], fused='NONE')
ConvTranspose_15_output_nhwc_cropped = ConvTranspose_15_output[0:1, 0:1, 2:66, 0:64]
ConvTranspose_15_output_nhwc_bias = torch.add(ConvTranspose_15_output_nhwc_cropped, self.const_ConvTranspose_15_transpose_conv_bias)
ConvTranspose_18_input = torch.relu(ConvTranspose_15_output_nhwc_bias)
ConvTranspose_18_output = _apply_module_transpose_conv2d(self.conv_transpose2d_2, ConvTranspose_18_input, target_shape=[1, 1, 72, 64], fallback_shape=[1, 1, 72, 64], fused='NONE')
ConvTranspose_18_output_nhwc_cropped = ConvTranspose_18_output[0:1, 0:1, 4:68, 0:64]
ConvTranspose_18_output_nhwc_bias = torch.add(ConvTranspose_18_output_nhwc_cropped, self.const_ConvTranspose_18_transpose_conv_bias)
onnx____ReduceMean_60 = torch.reshape(ConvTranspose_18_output_nhwc_bias, [int(v) for v in _shape_list(torch.as_tensor([1, 64, 64], dtype=torch.int32, device=self._device()))])
onnx____Squeeze_61 = _reduce_mean(onnx____ReduceMean_60, _normalize_axes(torch.as_tensor([2], dtype=torch.int32, device=self._device()), onnx____ReduceMean_60.ndim), True)
t_63 = torch.reshape(onnx____Squeeze_61, [int(v) for v in _shape_list(torch.as_tensor([1, 64], dtype=torch.int32, device=self._device()))])
return t_63
def forward_named(self, *args: torch.Tensor, **kwargs: torch.Tensor) -> Dict[str, torch.Tensor]:
return {'63': self.forward(*args, **kwargs)}
def load_model(device: str | None = None, eval_mode: bool = True) -> Model:
return Model(device=device, eval_mode=eval_mode)
</details>
model.py more PyTorch-likemodel.pytorch / torch.nn.functional expressionsforward() methods into stage/helper structure where appropriatefused_conv_relu.yolov7_tiny_head_0.768_post_480x640yolox_slite_model_mobilebert_1_metadata_1target_models.txt.GATHER, GATHER_NDSLICE, STRIDED_SLICERESIZE_*LOGISTIC, SOFTMAXDEPTH_TO_SPACE, SPACE_TO_DEPTHARG_MAX, ARG_MINAVERAGE_POOL_2D, MAX_POOL_2DNON_MAX_SUPPRESSION_V4/V5onnx2tf/tflite_builder/pytorch_accuracy_evaluator.py-tb flatbuffer_direct -fdopt -cotof now supports same-input comparison reporting for:
ONNX ↔ TFLiteONNX ↔ PyTorch-it + -fdopt support--input_tflite_file_path together with:
-tb flatbuffer_direct-fdopt-cotof-it + -fdopt + -cotof, the reference comparison is now TFLite ↔ PyTorch instead of ONNX ↔ PyTorch.onnx2tf-pytorch-bulkonnx2tf/utils/pytorch_bulk_runner.pytarget_models.txtflatbuffer_direct SavedModel/TFLite bridge behaviorRANDOM_UNIFORMREVERSE_SEQUENCEflatbuffer_direct validation and report generation aligned across ONNX-input and TFLite-input paths.README.md extensively for the new PyTorch export flow.flatbuffer_direct PyTorch export-it + -fdopt)-cotofThis branch makes flatbuffer_direct much more practical as a direct export pipeline rather than only a TFLite writer. In particular, it enables:
Verified with focused tests during implementation and with the flatbuffer_direct suite:
pytest -q tests/test_tflite2sm_phase1.py tests/test_pytorch_exporter.py tests/test_model_convert.py tests/test_flatbuffer_direct_op_error_report.py tests/test_tflite_builder_direct.py
Result:
793 passed4 warningspyproject.toml; this branch only bumps the package version.Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.4...2.3.5
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct.
This PR significantly expands flatbuffer_direct builtin support and tightens runtime safety for the direct TFLite export path.
The main goal of this branch is to move a large set of README-listed ONNX operators from partial or custom-only handling into native builtin lowering, while also making the generated LiteRT models materially safer to execute.
In practical terms, this branch:
flatbuffer_directDeformConv so the builtin path no longer aborts LiteRT on invoke()The flatbuffer_direct backend is most valuable when it can emit TFLite builtins instead of relying on TensorFlow conversion fallback or custom ops. This branch improves that value proposition in two ways:
DeformConv now uses a LiteRT-safe lowering strategy for the supported standard case, instead of producing models that could serialize and allocate but still abort at runtime.This branch adds or completes builtin lowering coverage across multiple builder families, including signal, reduction, elementwise, indexing, normalization, pooling, attention, and loss paths.
Notable areas improved include:
Bernoulli, BlackmanWindow, HammingWindow, HannWindow, RandomNormal, RandomUniform, RandomUniformLikeLeakyRelu, Mean, ReduceLogSum, ReduceLogSumExp, ReduceSumSquare, ThresholdedRelu, IsInf, IsNaN, ShrinkAffineGrid, CenterCropPad, Compress, ReverseSequence, Scatter, TensorScatterGroupNormalization, LpPool, GlobalLpPool, DetAttention, DFT, STFT, MelWeightMatrix, RotaryEmbedding, NegativeLogLikelihoodLoss, SoftmaxCrossEntropyLossMaxRoiPool, MaxUnpool, DeformConvThese additions are backed by registry-level dispatch/validation work and corresponding builder implementations.
DeformConv LiteRT runtime safetyDeformConv received a dedicated runtime hardening pass.
The new builtin lowering is intentionally constrained to the standard 2D float pattern:
group=1offset_group=1FLOAT16/FLOAT32The previous path could generate a model that converted successfully but aborted the LiteRT interpreter during invoke(). The new path fixes that by:
GATHER(batchDims=1) dependency from the sampling pathUnsupported grouped DeformConv patterns are still preserved as explicit custom-op candidates when custom ops are enabled and allowlisted.
The branch updates the source of truth for builtin support in op_registry.py, then propagates that state into:
The README builtin summary count is also refreshed to the current table contents.
Package and container-facing version references are updated from 2.3.3 to 2.3.4 to match the functional expansion in this branch.
The branch includes substantial test expansion for flatbuffer_direct, including:
A notable addition is subprocess-isolated runtime verification for DeformConv, so native LiteRT abort regressions fail safely inside tests instead of taking down the full pytest worker.
pytest -q tests/test_tflite_builder_direct.py tests/test_tflite_builder_op_coverage.py
Result:
646 passed, 2 warnings in 129.48sThe core changes are concentrated in:
onnx2tf/tflite_builder/op_registry.pyonnx2tf/tflite_builder/op_builders/conv.pyonnx2tf/tflite_builder/op_builders/shape.pyonnx2tf/tflite_builder/op_builders/reduce.pyonnx2tf/tflite_builder/op_builders/elementwise.pyonnx2tf/tflite_builder/op_builders/pool.pyonnx2tf/tflite_builder/op_builders/index.pyonnx2tf/tflite_builder/op_builders/norm.pyonnx2tf/tflite_builder/op_builders/recurrent.pytests/test_tflite_builder_direct.pytests/test_tflite_builder_op_coverage.pyREADME.mdThis PR intentionally favors explicit constrained builtin support over broad but unsafe lowering.
The clearest example is DeformConv: the supported builtin path is now narrower, but it is meaningfully more correct because it survives actual LiteRT execution. Patterns outside that safe envelope still retain the existing custom-op fallback policy.
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.3...2.3.4
> Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`. With the v2.3.3 upd…
[!IMPORTANT] Starting with onnx2tf v2.4.0,
tf_converterwill be deprecated and the default backend will be switched toflatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated toflatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them intoflatbuffer_direct.
This PR improves the flatbuffer_direct backend by adding ONNX-input support for several legacy conversion options that were previously only effective on the tf_converter path, while intentionally keeping input_tflite_file_path direct-import validation unchanged.
The main goal is to close backend behavior gaps without weakening the direct backend's architecture or introducing hidden fallback behavior.
flatbuffer_direct ONNX inputThis PR wires the following options through the flatbuffer_direct export and lowering path:
disable_suppression_flextransposedisable_suppression_flexstridedsliceoptimization_for_gpu_delegatereplace_argmax_to_reducemax_and_indices_is_int64replace_argmax_to_reducemax_and_indices_is_float32replace_argmax_to_fused_argmax_and_indices_is_int64replace_argmax_to_fused_argmax_and_indices_is_float32fused_argmax_scale_rationot_use_onnxsim was already effective for ONNX-input conversion before backend branching, so this PR treats it as an existing capability and adds documentation/tests to make that explicit.
flatbuffer_direct now honors:
disable_suppression_flextransposedisable_suppression_flexstridedsliceBehavioral effect:
TRANSPOSE with the original rank instead of forcing rank compression/decomposition.SLICE/STRIDED_SLICE instead of rank-compressing the operation.This PR adds direct-path handling for optimization_for_gpu_delegate in places where the TensorFlow path already had special-case behavior.
Key additions:
BATCH_MATMUL fallback path.This keeps the implementation inside the ModelIR/direct-lowering pipeline rather than introducing any backend fallback.
flatbuffer_directThis PR adds direct backend support for the four legacy ArgMax replacement modes.
For:
replace_argmax_to_reducemax_and_indices_is_int64replace_argmax_to_reducemax_and_indices_is_float32The direct path now lowers ArgMax into a reduction-based subgraph that matches the legacy TensorFlow-path semantics, including:
keepdims supportint64 or float32For:
replace_argmax_to_fused_argmax_and_indices_is_int64replace_argmax_to_fused_argmax_and_indices_is_float32The direct path now supports fused handling for ONNX Resize -> ArgMax 4D patterns.
The implementation:
This intentionally does not attempt to add unrelated ScaleAndTranslate support.
This PR keeps existing validation philosophy intact.
Specifically:
input_tflite_file_path still rejects these ONNX-lowering options because there is no ONNX graph stage in direct-import mode.fused_argmax_scale_ratio validation is preserved.tf_converter is introduced.The implementation includes:
convert() into export_tflite_model_flatbuffer_direct()This PR adds and updates tests covering:
not_use_onnxsim=True on ONNX-input flatbuffer_directFull flatbuffer_direct-focused test execution was also run:
tests/test_tflite_builder_direct.pytests/test_tflite2sm_phase1.pytests/test_flatbuffer_direct_op_error_report.pytests/test_model_convert.pyResult:
677 passedThis branch also includes the package version bump to 2.3.3 and the matching uv.lock update.
Before this change, flatbuffer_direct had a meaningful feature gap versus the legacy TensorFlow conversion path for several CLI options that users still rely on when debugging backend-specific issues, tuning delegate compatibility, or preserving legacy conversion behavior.
This PR narrows that gap while preserving the direct backend's core principles:
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.2...2.3.3
Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`.
Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct.
This PR expands and hardens the flatbuffer_direct backend so that more conversion workflows stay on the direct ModelIR/FlatBuffer path instead of falling back to the legacy tf_converter path.
The branch focuses on improving functional parity, making interrupt/rewrite features work consistently in direct mode, and tightening the documentation/tests around the new behavior.
flatbuffer_direct now keeps the direct export flow for cases that previously forced a fallback or were rejected:
--output_h5--output_keras_v3--output_tfv1_pb--disable_model_save-it/--input_tflite_file_path direct-import workflowsInstead of dropping to tf_converter, these flows now use a ModelIR-derived SavedModel bridge when a non-TFLite artifact is required.
This means direct mode can now:
.h5, .keras, and TFv1 .pb artifacts without leaving the direct backend--disable_model_save while still performing internal staging/validation and leaving no final artifacts in the requested output directory-it input when appropriateValidation was also tightened so incompatible combinations are rejected explicitly rather than silently changing execution mode.
-inimc / -onimcflatbuffer_direct no longer depends on ONNX graph extraction for interrupt-based cropping.
Instead, it now crops the already lowered/imported top-level ModelIR using the requested boundary tensor names. This applies to:
-it/--input_tflite_file_path inputBehavioral impact:
sne4onnx for this path-dgc, -ebu, and -eruDirect mode now supports these options through shared ModelIR rewrites instead of forcing the TensorFlow conversion flow:
-dgc / --disable_group_convolution-ebu / --enable_batchmatmul_unfold-eru / --enable_rnn_unrollThe implementation unifies ONNX input and imported TFLite ModelIR handling by applying rewrites after lowering/import and before split planning / SavedModel bridging / final export.
In practice this adds:
CONV_2D decomposition in direct modeBATCH_MATMUL unfolding in direct modeIf a requested rewrite cannot be applied safely, conversion now fails explicitly instead of behaving like a no-op.
MeanVarianceNormalization with -me / --mvn_epsilonThis branch adds direct lowering support for ONNX MeanVarianceNormalization in flatbuffer_direct.
The new lowering expands MVN into builtin-friendly primitive ops:
MEANSUBMULMEANADDSQRTDIVmvn_epsilon is now threaded through the direct backend so -me / --mvn_epsilon affects both:
tf_converterflatbuffer_directfor ONNX input.
This removes one more feature gap between the two backends while keeping direct export fully self-contained.
This branch also updates the public surface and documentation so the new direct-path behavior is visible and consistent:
-esm for --eval_split_modelsflatbuffer_direct semantics for:
--disable_model_saveThis work reduces the number of cases where users must understand or work around backend switches.
Before this branch, enabling certain artifact outputs or graph-manipulation options could unexpectedly move execution away from flatbuffer_direct, or reject flows that were conceptually still compatible with direct export.
After this branch, the direct backend is much more coherent:
-it input behave more consistentlyI validated the branch with focused and backend-wide tests, including:
Most recent backend-wide run:
pytest -q tests -k flatbuffer_direct
Result:
599 passed141 deselected2 warningsThe warnings were non-failing and already known:
tf_kerasfloat16 cast overflow warning in an existing negative-infinity broadcast testFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.1...2.3.2
Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`.
Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct.
This PR substantially expands the flatbuffer_direct split workflow and makes the direct path more coherent across export, validation, TFLite input import, and SavedModel output.
The main goal of this branch is to turn split handling into a first-class ModelIR-based feature for flatbuffer_direct, rather than relying on the legacy ONNX/GraphSurgeon split flow.
enable_auto_split_modelflatbuffer_direct now uses a single ModelIR-based split planner for direct export.
Key behavior changes:
enable_auto_split_model=True is the public split entry point for flatbuffer_direct.ModelIR and produces the manifest-based artifact set.flatbuffer_direct.This makes split behavior deterministic and consistent with the direct backend’s internal representation.
onnx2tf \
-i deim_hgnetv2_s_wholebody28_ft_1250query_fixed.onnx \
-cotof \
-tb flatbuffer_direct \
-easm \
-asms 5MB # KB or MB or GB
<img width="827" height="554" alt="image" src="https://github.com/user-attachments/assets/22477679-7a70-4225-a0d5-66b132517325" />
<img width="2192" height="888" alt="image" src="https://github.com/user-attachments/assets/d2de788d-47f6-4558-9ca9-c97659c069c4" />
The partition builder was tightened so split artifacts are more faithful and easier to inspect.
Improvements include:
This specifically addresses the issue where split partitions appeared to have many disconnected constants despite being otherwise executable.
The direct split pipeline now supports partition-level SavedModel output.
New capabilities:
flatbuffer_direct + enable_auto_split_model + flatbuffer_direct_output_saved_model exports partition SavedModelsThis closes the gap between split TFLite output and SavedModel output for the direct backend.
-it / TFLite-input direct split supportTFLite-import (-it) can now participate in the direct split flow.
That means:
ModelIRThis makes the direct backend’s split functionality available beyond ONNX-originated conversions.
Split evaluation and direct validation behavior were simplified and clarified.
Changes include:
eval_split_models is now the single split evaluation interface and directly encodes the reference mode (onnx or unsplit_tflite)eval_split_reference option was removed-cotof no longer emits a SavedModel inference warning when the workflow did not actually produce a SavedModelThese changes reduce ambiguity in both CLI usage and generated reports.
The branch also hardens ONNX handling by enforcing a repository-wide runtime policy:
onnx.shape_inference.infer_shapes is not run when an ONNX model uses external_dataThis avoids unsafe shape-inference calls on external-data models while preserving the existing behavior for regular in-memory ONNX models.
Notable interface changes:
auto_split_tflite_by_size was removed as a public flatbuffer_direct split entry pointenable_auto_split_model is the public split trigger for flatbuffer_directauto_split_max_size remains the split target-size controleval_split_reference was removedeval_split_models now takes the reference mode directly (onnx or unsplit_tflite)These changes intentionally reduce redundant split/evaluation options and make the direct backend easier to operate.
I ran the flatbuffer-direct-focused regression suite:
pytest -q \
tests/test_tflite_builder_direct.py \
tests/test_tflite_builder_op_coverage.py \
tests/test_flatbuffer_direct_op_error_report.py \
tests/test_accuracy_evaluator_input_layout.py \
tests/test_accuracy_evaluator_name_map.py \
tests/test_accuracy_evaluator_seeded_input.py \
tests/test_accuracy_evaluator_subprocess.py \
tests/test_tflite_split_planner.py \
tests/test_tflite_builder_gridsample_validation.py
Result:
625 passed, 1 warningThe remaining warning is an existing float16 cast overflow warning during one direct-path test and does not fail the suite.
Before this branch, flatbuffer_direct split support was fragmented across multiple partially overlapping flags and code paths. This branch consolidates the split workflow around the direct backend’s own ModelIR, aligns TFLite and SavedModel outputs, improves artifact correctness, and removes several confusing validation/reporting edge cases.
As a result, feat-split turns split export from an experimental side path into a much more coherent and reviewable feature set for flatbuffer_direct.
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.3.0...2.3.1
Starting with onnx2tf v2.4.0, `tf_converter` will be deprecated and the default backend will be switched to `flatbuffer_direct`.
Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct.
<img width="1869" height="770" alt="image" src="https://github.com/user-attachments/assets/d984cf61-062f-46f4-8472-3253e220f44c" />
This PR introduces a complete SavedModel path for the direct LiteRT backend and aligns the CLI/runtime behavior with a no-fallback flatbuffer_direct policy.
Historically, SavedModel export behavior in direct-mode workflows depended on fallback/legacy conversion behavior. This branch makes SavedModel generation explicit and deterministic by using a dedicated ModelIR -> SavedModel exporter and by adding a TFLite -> ModelIR importer for direct TFLite input use cases.
The branch also adds machine-readable validation/reporting and a bulk verification runner to support large-scale model sweeps.
flatbuffer_direct SavedModel exportonnx2tf/tflite_builder/saved_model_exporter.py.tf.Module + @tf.function serving signature generation.shape_signature (-1 -> None) and TensorIR dtype mapping.ModelIR operators.WHILE).CUSTOM ops with detailed per-node diagnostics.onnx2tf/tflite_builder/__init__.py and return payload now includes saved_model_path when generated.--flatbuffer_direct_fallback_to_tf_converter usage from CLI/runtime path..tflite -> SavedModel only)input_tflite_file_path parameter to convert() and new CLI option:
-it, --input_tflite_file_pathonnx2tf/tflite_builder/tflite_importer.py:
ModelIR from TFLite FlatBuffer (main graph + subgraphs).dtype, shape, shape_signature, is_variable, quant fields, constant data).op_type, version, builtin options normalization, CUSTOM metadata preservation).serving_default SignatureDef boundary names when available.onnx2tf/onnx2tf.py:
input_onnx_file_path, input_tflite_file_path, and onnx_graph are mutually exclusive.disable_model_save=True is rejected in TFLite direct mode.tflite_backend must be flatbuffer_direct for this mode.-cotof behavior)<model_name>_saved_model_validation_report.json-cotof now runs SavedModel inference + comparison against input TFLite.onnx2tf/utils/tflite2sm_bulk_runner.py with CLI entrypoint behavior.bulk_status.json state).bulk_summary.json and bulk_summary.md.deformable_detr_one_input_simplemosaic-9accuracy_evaluator.py:
--flatbuffer_direct_output_saved_model / -fdosm descriptions in CLI and script-option sections.--input_tflite_file_path / -it usage and examples.FLATBUFFER_DIRECT_MIGRATION_GUIDE.md for no-fallback migration guidance.2.3.0 (onnx2tf/__init__.py, pyproject.toml, uv.lock).Added/updated test coverage across:
tests/test_tflite2sm_phase1.pytests/test_tflite2sm_bulk_runner.pytests/test_accuracy_evaluator_seeded_input.pyValidated in this branch with:
pytest -q tests/test_tflite2sm_phase1.py tests/test_tflite2sm_bulk_runner.py tests/test_accuracy_evaluator_seeded_input.py37 passedflatbuffer_directconvert() by re-raising NotImplementedError from the flatbuffer_direct fast path (instead of always wrapping as generic RuntimeError).ij,jk->ik-style equations.ij,jk->kj (no ONNX_EINSUM custom op).ii,jk->kjpytest -q tests -k flatbuffer_direct564 passed, 108 deselected, 1 warning.tflite direct-input mode for SavedModel generation.ModelIR in flatbuffer_direct flow (-fdosm), and.tflite input mode (-it).-cotof can produce strict SavedModel-vs-TFLite validation reports.N/A
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.2.2...2.3.0
This PR strengthens flatbuffer_direct conversion quality for DAMO-style graphs and broadens built-in operator coverage while keeping model contracts s
This PR strengthens flatbuffer_direct conversion quality for DAMO-style graphs and broadens built-in operator coverage while keeping model contracts stable at graph boundaries.
clone_model_ir_with_float32 to promote FP16 tensors/options into FP32.prune_identity_cast_operators and optimize_redundant_transpose_operators to simplify generated graphs before writing TFLite._resolve_dynamic_reshape_shapes now supports prefer_runtime_inferable_from_onnx_raw=True to preserve ONNX runtime-inferable -1 templates instead of stale concretized shapes.RANGE outputs with runtime-dependent lengths.preserveDynamicShape=True is set.SPLIT input dtypes:
SPLIT into SLICE chains (with optional cast in/out) so conversion succeeds for dtype combinations not accepted by LiteRT SPLIT._optimize_transpose_flatten_globalnorm_pad_prepost_nhwc_chains removes redundant pre/post transpose wrappers and rewrites affected shape/pad metadata into NHWC-consistent form.Inverse:
NxN (up to 16x16) by Gauss-Jordan-style tensor-op decomposition.GatherND:
batch_dims > 0 by flattening batch prefix, synthesizing batch indices (RANGE + TILE + CONCAT), gathering, then reshaping back.Gemm/MatMul -> FULLY_CONNECTED path:
LSTM path:
Slice shape inference:
GridSample:
Inverse, GatherND, GridSample) to match new lowering capabilities and constraints.GatherND and shape/index helper ops (RANGE, SHAPE) required by new lowerings.tests/test_tflite_builder_direct.py with targeted regression tests covering:
Inverse static 8x8 loweringGatherND(batch_dims>0) and GridSample unknown-shape scenariosSPLIT fallback rewriting-1 preference behaviorpytest -q tests/test_tflite_builder_direct.py568 passed, 1 warning2.2.2 (pyproject.toml, onnx2tf/__init__.py).2.2.2.Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.2.1...2.2.2
This PR delivers a broad flatbuffer_direct upgrade focused on three goals:
This PR delivers a broad flatbuffer_direct upgrade focused on three goals:
It includes the two commits in fix-sgsch:
d20dba8 Optimize redundant transpose chains in flatbuffer_directe7641fa Update flatbuffer_direct converters, metadata handling, and testsenable_accumulation_type_float16 into flatbuffer-direct export calls.reduced_precision_support in direct-written TFLite flatbuffers.fp16accfp16 when -eatfp16 is enabled,fp16accfp32 otherwise.*_float16.tflite metadata stayed fp16accfp32 despite -eatfp16.build_cumprod_op) using a TFLite decomposition:
RANGE, LESS/LESS_EQUAL, RESHAPE, TILE, FILL, SELECT_V2, REDUCE_PROD, optional REVERSE_V2.build_unique_op) with validation constraints.build_castlike_op).UNIQUE options serialization support in model writer (UniqueOptions).indices/updates,INT64 -> INT32 and UINT64 -> UINT32 normalization in critical paths.Where/control-flow mux dtype normalization for signed/unsigned integer outputs.INT64->INT32 CAST chain cleanup in post-lowering optimization.newShape=[] intent when required,0 / multiple -1 cases,onnxRawNewShape with existing concrete newShape.EXPAND_DIMS / SQUEEZE to RESHAPE:
Added and integrated multiple strict/final-stage optimization passes in lower_from_onnx2tf.py, including:
mul/add/mul forms),These reduce redundant transposes/casts and stabilize final graph topology after late rewrites.
padding_mode in {zeros, border}),If NMS-guard subpattern variant.If branch mux dtype normalization.Loop state output remapping/casting logic for dtype/shape consistency.onnx2tf package version 2.2.0 -> 2.2.1Executed full direct-builder regression:
pytest -q tests/test_tflite_builder_direct.pyAlso added/updated extensive tests for:
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.2.0...2.2.1
This PR consolidates the fix-rtmdet branch improvements focused on functional reliability, direct FlatBuffer export robustness, and conversion/typing
This PR consolidates the fix-rtmdet branch improvements focused on functional reliability, direct FlatBuffer export robustness, and conversion/typing stability across the ONNX->TFLite pipeline.
Compared to main, this branch contains 3 commits and broad updates across operator converters, the flatbuffer_direct backend, tests, and tooling.
flatbuffer_direct lowering correctnessflatbuffer_direct semantics.Impact:
onnx2tf/ops/* to improve type/shape safety and runtime consistency.Impact:
flatbuffer_direct coverage and validation alignmentrequires_constant_input for dynamic reduce axes where applicable).Impact:
benchmark_tflite.py for CPU/CUDA-oriented TFLite latency measurement workflows.onnx2tf.py, __init__.py, README.md, etc.).Impact:
tests/test_tflite_builder_direct.py with additional scenarios around dynamic pad/flatten/reshape behavior and related shape-signature expectations.Executed targeted flatbuffer_direct suites after fixes:
python -m pytest -q \
tests/test_tflite_builder_direct.py \
tests/test_tflite_builder_preprocess.py \
tests/test_tflite_builder_op_coverage.py \
tests/test_tflite_builder_gridsample_validation.py \
tests/test_flatbuffer_direct_op_error_report.py \
tests/test_tflite_split_planner.py
Result:
onnx2tf/ops/*.flatbuffer_direct updates in:
onnx2tf/tflite_builder/op_builders/shape.pyonnx2tf/tflite_builder/op_registry.pyfix-rtmdetRTMDet export paths heavily depend on stable handling of dynamic tensor metadata, reshape/pad/reduce behavior, and consistent converter typing assumptions. This branch improves those exact stability points, reducing conversion-time ambiguity and increasing confidence in direct TFLite generation paths.
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.1.5...2.2.0
This PR expands built-in coverage in flatbuffer_direct and improves stability for models with control flow and dynamic shapes.
This PR expands built-in coverage in flatbuffer_direct and improves stability for models with control flow and dynamic shapes.
In particular, it removes the ONNX_SLICE CUSTOM fallback that occurred while converting encoder-epoch-99-avg-1.onnx, allowing conversion to complete using built-in operators only.
It also includes version updates for the 2.1.5 release (pyproject.toml, onnx2tf/__init__.py, and Docker tags in README.md).
If in flatbuffer_direct.If branch lowering (arithmetic, comparison, shape ops, Slice, Conv, LSTM, etc.).value_info hint propagation to improve shape/dtype recovery.Size (SHAPE + REDUCE_PROD (+ CAST))._NodeWrap and branch wrapping to preserve empty optional inputs ("").Slice support for dynamic starts with constant ends/axes/steps on single-axis patterns.Slice starts/ends length checks so dynamic-start cases are validated correctly.GatherElements handling for unknown-rank placeholders (shape=[1], signature<0) in both validator and builder.Expand and Unsqueeze for unresolved placeholder ranks.LSTM initial state (initial_h/initial_c) validation to be shape-signature aware.LSTM state inputs.test_tflite_builder_direct.py for dynamic Slice, unknown-rank GatherElements, unknown-rank Unsqueeze, generic If branch mux, Size, and related paths.Before:
onnx2tf -i encoder-epoch-99-avg-1.onnx -cotof -tb flatbuffer_direct -kat x x_lensSlice_699.ONNX_SLICE CUSTOM ops were introduced.After:
flatbuffer_direct path.ONNX_SLICE CUSTOM op is introduced.allow_custom_ops=False, lower_onnx_to_ir completes with CUSTOM op count = 0.Validation:
pytest -q tests/test_tflite_builder_direct.py499 passed, 1 warningFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.1.4...2.1.5
Currently, `flatbuffer_direct` has better conversion stability than `tf_converter` mode.
Currently, flatbuffer_direct has better conversion stability than tf_converter mode.
This PR is a broad flatbuffer_direct upgrade focused on three goals:
Einsum
bchw,bnc->bnhw.SUM, TRANSPOSE, RESHAPE, BATCH_MATMUL, CAST).MatMul
Pooling
EXPAND_DIMS -> AVERAGE_POOL_2D -> SQUEEZE).ConvTranspose
TopK
k is scalar-constant, it is now constantized directly to INT32 for TOPK_V2.CAST -> SQUEEZE.Reduce axis robustness
-1) when rank metadata is unreliable, preventing incorrect axis remapping.Reshape allowzero semantics
allowzero=0 dim-copy semantics are preserved in TFLite lowering.Reshape-chain safety
0/-1 behavior).SHAPE consumers (not only reshape-related patterns).File name too long) in:
Gather index normalization (without -rtpo)Gather indices when the positions tensor is scalar-like:
[1] for runtime compatibility, while preserving ONNX output-rank semantics by inserting a final shape restore where needed.Full flatbuffer_direct direct-lowering test suite:
pytest -q tests/test_tflite_builder_direct.py492 passed, 1 warning.Added/updated targeted regression tests, including:
k constantization,SHAPE input repair,Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.1.3...2.1.4
This PR focuses on functional improvements in the flatbuffer_direct backend: broader builtin-op conversion coverage, stronger layout optimization (tra
This PR focuses on functional improvements in the flatbuffer_direct backend: broader builtin-op conversion coverage, stronger layout optimization (transpose elimination), and more stable evaluation/reporting behavior.
The goal is to reduce CUSTOM-op fallbacks, improve conversion fidelity on real models, and keep output checks reliable.
This PR adds and stabilizes several builtin conversion paths so models that previously required CUSTOM lowering can stay on builtin TFLite operators.
Highlights:
ADD chaining.best-352 style chains identified during debugging).Net effect:
TRANSPOSE ops,flatbuffer_direct evaluation-helper behavior for edge cases in temporary input handling and intermediate comparison flow.-dgc--disable_group_convolution is applied in both tf_converter and flatbuffer_direct flows.-cotof -tb flatbuffer_direct.best-352) and output checks remained passing after optimization.Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.1.2...2.1.3
In the flatbuffer_direct path, we had recurring issues in practical use:
In the flatbuffer_direct path, we had recurring issues in practical use:
TopK and MultiHeadAttention were not lowered as built-ins and ended up as custom ops, which made compatibility and accuracy validation unstable.shape_signature propagation and side effects from layout optimizations caused RESHAPE prepare errors and intermediate tensor comparison failures.-cotof validation, memory usage could spike heavily, leading to SWAP thrashing and near-freeze behavior.com.microsoft domain arrived as default domain, reducing runtime compatibility.This PR addresses these issues together to improve built-in coverage, stability, and memory behavior in flatbuffer_direct.
Expanded built-in lowering coverage
TopK.largest=1 (descending) and largest=0 (ascending-equivalent) using TOPK_V2 plus auxiliary ops.sorted=0: TFLite always returns descending-sorted output, so index order may not exactly match ONNX reference.MultiHeadAttention with strict validation (3-input Q/K/V form, rank-3, FLOAT16/FLOAT32, num_heads constraints, etc.).Slice.DepthToSpace CRD mode lowering in built-in form (channel reordering + DEPTH_TO_SPACE).Stabilized shape_signature and layout optimization behavior
index.py, shape.py, shared.py, and model_writer.py.shape_signature inconsistency risk around TopK, Slice, and Reshape.lower_from_onnx2tf.py.Transpose -> Reshape rewrites in axis-semantic cases (e.g., rank-3), avoiding meaning drift in LogSoftmax/Transpose/Add paths.Strengthened low-memory validation flow (memmap)
--disable_onnxruntime_output_memmap).--onnxruntime_output_memmap_dir to control memmap storage location.accuracy_evaluator to pass worker inputs/outputs via memmap, reducing large IPC tensor copies.Temp directory management and cleanup on abnormal termination
onnx2tf/utils/tempdir_cleanup.py.atexit and signal handlers (SIGTERM/SIGINT/SIGHUP/SIGQUIT).Improved OP error helper operability
| / - \) during helper execution.Improved ORT compatibility (Microsoft domain)
com.microsoft domain for selected contrib ops (including MultiHeadAttention) during conversion.Added standalone low-memory inference script
lowmem_test_infer.py.Test expansion
tests/test_tflite_builder_direct.py.test_flatbuffer_direct_min_topk_dynamic_k_loweringtest_flatbuffer_direct_matmul_vector_rhs_loweringtest_flatbuffer_direct_slice_dynamic_end_prefix_rank2_loweringtest_flatbuffer_direct_slice_dynamic_start_end_single_axis_loweringtest_flatbuffer_direct_logsoftmax_loweringBefore (examples)
Custom ops lowered: ... op_types=[ONNX_TOPK] ...Custom ops lowered: ... op_types=[ONNX_MULTIHEADATTENTION] ...tflite/kernels/reshape.cc:94 num_input_elements != num_output_elements ...OP error report generation was skipped ... helper process timed outAfter (behavior with this PR)
TopK and MultiHeadAttention now have built-in lowering paths.TopK(sorted=0) now emits an explicit warning for potential index-order mismatch.N/A
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.1.1...2.1.2
Enhanced flatbuffer_direct graph rewrite passes to remove redundant transpose/reshape chains across a wide set of real-model patterns.
flatbuffer_direct graph rewrite passes to remove redundant transpose/reshape chains across a wide set of real-model patterns.ConvInteger in flatbuffer_direct (with validator/registry wiring), reducing unresolved custom-op failures.tests/test_tflite_builder_direct.py for new rewrite and lowering patterns.2.1.1 (pyproject.toml, onnx2tf/__init__.py, README.md).ConvInteger; previously this path could fail as unresolved custom op.pytest -q tests/test_tflite_builder_direct.py
451 passed, 1 warningflatbuffer_direct behavior/coverage and associated regression hardening.Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.1.0...2.1.1
This PR significantly accelerates the flatbuffer_direct conversion path and improves operational visibility/stability.
This PR significantly accelerates the flatbuffer_direct conversion path and improves operational visibility/stability.
--tflite_backend flatbuffer_direct that avoids per-node TensorFlow graph building (op.make_node() loop).PrependByte loops with Builder.CreateByteVector fast-path for large constant buffers.Measured example (*_float32.tflite, same model):
tf_converter: ~24.947sflatbuffer_direct: ~0.239s(Actual speedup varies by model/options/environment.)
https://github.com/user-attachments/assets/846f9c57-93ac-43a0-8cc2-9c989923c320
flatbuffer_direct fast path in convert()
output_h5, output_keras_v3, output_tfv1_pb, auto-split-model, disable_model_save).--flatbuffer_direct_fallback_to_tf_converter is enabled.Removed runtime dependency on flatc for onnx2tf conversion flows (including -coion).
onnx2tf/tflite_builder/schema/schema.fbsonnx2tf/tflite_builder/schema/schema_generated.pyflatc normal operation.FlatBuffer writer performance overhaul
tflite_builder/model_writer.py.ONNX2TF_FLATBUFFER_DIRECT_SERIALIZERONNX2TF_FLATBUFFER_DIRECT_SERIALIZER_FALLBACK_TO_OBJECT_PACKProgress and timing visibility
dynamic_ncols=True.Report generation control
-cotof (--check_onnx_tf_outputs_elementwise_close_full) is enabled.Lowering/operator coverage and robustness
Col2Im support in flatbuffer_direct dispatch (validation + builder path).Tests and docs
tests/test_tflite_builder_direct.py with new optimization test cases.2.1.0; added tqdm dependency.pytest -q tests/test_tflite_builder_direct.py
381 passed, 1 warningFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.27...2.1.0
This PR improves fix-yolo conversion quality and stability, with a focus on flatbuffer_direct parity and layout-optimization coverage.
This PR improves fix-yolo conversion quality and stability, with a focus on flatbuffer_direct parity and layout-optimization coverage.
Inverse builtin lowering in flatbuffer_direct (SLICE/MUL/SUB/ADD/NEG/CONCATENATION/DIV) for square 2x2/3x3 float matrices.Inverse behavior to pseudo-lower by default (to avoid FlexMatrixInverse emission) and only keep tf.linalg.inv when -rtpo Inverse is explicitly set.GroupNorm, Inverse) so ORT graph validation remains consistent after graph rewrites.serving_default_* and generic suffix matching), reducing false input-shape mismatches during accuracy checks.flatbuffer_direct layout optimization improvementsADD fanout chains (MUL(const)->ADD(const)->post-TRANSPOSE)UNARY->MUL(const)->ADD(const)->post-TRANSPOSE)MaxPool with indices: added constrained builtin lowering path and validation.Conv builder: added explicit compute-cast handling to stabilize dtype paths.Slice handling: fixed axes-based begin/end/strides expansion and negative-axis normalization.Executed flatbuffer_direct full test target:
uv run --with pytest pytest tests -k flatbuffer_direct -ra
Result:
373 passed69 deselected1 warningonnx2tf.py, ops/Inverse.py, ops/Slice.py, utils/common_functions.pytflite_builder/lower_from_onnx2tf.py, op_registry.py, op_builders/*, accuracy_evaluator.py, preprocess rulesREADME.md, __init__.py, pyproject.toml, uv.locktests/test_tflite_builder_direct.py, tests/test_tflite_builder_preprocess.py, tests/test_microsoft_domain_supplement.py, tests/test_accuracy_evaluator_name_map.pyFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.26...2.0.27
This PR improves flatbuffer_direct conversion stability and builtin coverage, with a focus on string/optional handling and redundant layout-wrapper re
This PR improves flatbuffer_direct conversion stability and builtin coverage, with a focus on string/optional handling and redundant layout-wrapper reduction.
fix-strStringNormalizer under constrained runtime patterns.
OptionalHasElement cases to reduce custom-op fallback.Unsqueeze builtin axis normalization with output-rank semantics (covers opset8-style patterns).Pad builtin path to support non-zero constant padding via PADV2 (for non-quantized tensors).STRING).Transpose -> Reshape -> Transpose (including NHWC/NCHW channel-tail singleton variants)Squeeze -> Reshape identity chainspytest -q tests/test_tflite_builder_direct.py -k "string_normalizer or optional_has_element or unsqueeze or pad or transpose_se_conv_mul_prepost_nhwc_chain or squeeze_reshape_identity_chains_removed or transpose_reshape_transpose_to_expanddims_nhwc"
28 passedpytest -q tests/test_accuracy_evaluator_subprocess.py
2 passedFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.25...2.0.26
This PR improves flatbuffer_direct conversion reliability and coverage, with a focus on control-flow lowering and shape/runtime correctness.
This PR improves flatbuffer_direct conversion reliability and coverage, with a focus on control-flow lowering and shape/runtime correctness.
Loop and improved branch handling for If in flatbuffer_direct.onnx2tf.py (preserving graph metadata such as domain/IR version) and improved eval fallback behavior.RandomNormalLike, RandomUniformLike) to avoid None shape conversion failures.Concat, ConvTranspose, Slice, control-flow related lowering paths).2.0.25).Recent conversion failures and regressions were observed around:
ONNX_LOOP custom-op fallback,This PR addresses those issues by expanding built-in lowering coverage and tightening shape/runtime validation paths.
pytest -q tests/test_tflite_builder_direct.py -k "loop_"
4 passed.flatbuffer_direct behavior and diagnostics, and keeps compatibility with existing conversion paths.Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.24...2.0.25
This PR improves flatbuffer_direct conversion quality and operator coverage, with a strong focus on removing redundant transpose chains and reducing c
This PR improves flatbuffer_direct conversion quality and operator coverage, with a strong focus on removing redundant transpose chains and reducing custom-op fallbacks.
If lowering:
flatbuffer_direct, including:
Log, Max, ReduceMin, ScatterElements, RoiAlign, LayerNormalizationIf validation/dispatch rules for constrained patternsNON_MAX_SUPPRESSION_V4 or V5 via --switch_nms_version (-snms)lower_from_onnx2tf.py to remove redundant pre/post layout adapters and preserve required boundary adapters.Executed:
pytest -q tests/test_tflite_builder_direct.py tests/test_tflite_builder_op_coverage.pyResult:
351 passed, 1 warningWarning observed:
RuntimeWarning: overflow encountered in cast in float16 cast path during test_flatbuffer_direct_where_neg_inf_broadcast_no_nan.Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.23...2.0.24
This PR brings the latest fix-var branch improvements into main and focuses on making the flatbuffer_direct path more practical for real ONNX models.
This PR brings the latest fix-var branch improvements into main and focuses on making the flatbuffer_direct path more practical for real ONNX models.
op_registry, op_builders, and lowering passes).REDUCE_MAX) for flatbuffer_direct.pytest -q tests/test_tflite_builder_direct.py (full suite): 307 passedflatbuffer_direct models using --report_op_coverage.fix-var accumulated tightly related direct-backend changes that were validated together.flatbuffer_direct behavior and test/documentation alignment.Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.22...2.0.23
This PR improves flatbuffer_direct lowering coverage and removes redundant layout conversion chains.
This PR improves flatbuffer_direct lowering coverage and removes redundant layout conversion chains.
InstanceNormalization lowering to builtin TFLite ops (MEAN/SUB/MUL/ADD/SQRT/DIV chain)Dropout lowering (inference-time no-op + optional mask generation path)Reshape dynamic shape input (INT32/INT64, rank-1) in validation/loweringCONCAT + SHAPEtranspose_slice_logistic_concat_reshape_tail NHWC rewritetranspose_pre_unary_reshape_transpose_suffix rewritetranspose_logistic_muladd_prepost rewriteSLICE begin/size when input dimensions are dynamictests/test_tflite_builder_direct.py with new coverage for:
2.0.22 and sync dependency/lock updates (sne4onnx/sng4onnx 2.0.1)pytest -q tests/test_tflite_builder_direct.py -k "transpose_pre_unary_reshape_transpose_suffix_nhwc_chain_optimized or transpose_logistic_muladd_prepost_nhwc_chain_optimized"pytest -q tests/test_tflite_builder_direct.py -k "transpose_add_reshape_transpose_suffix_optimization or transpose_mulconst_add_reshape_transpose_suffix_optimization or transpose_pre_unary_reshape_transpose_suffix_nhwc_chain_optimized or transpose_pre_add_mul_add_prelu_nhwc_chain_optimized or transpose_logistic_muladd_prepost_nhwc_chain_optimized"onnx2tf -i shadowformer_istd_160x240_split.onnx -cotof -tb flatbuffer_direct --report_op_coverageFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.21...2.0.22
This PR expands builtin coverage in flatbuffer_direct and adds a broad set of transpose-chain optimizations to remove redundant NCHW/NHWC layout adapt
This PR expands builtin coverage in flatbuffer_direct and adds a broad set of transpose-chain optimizations to remove redundant NCHW/NHWC layout adapters.
Min -> MINIMUM (multi-input is lowered as chained binary ops)TopK -> TOPK_V2 (with explicit constraints for axis/largest/sorted/k shape/indices dtype)DepthToSpace -> DEPTH_TO_SPACE (DCR) / RESHAPE+TRANSPOSE+RESHAPE (CRD)HardSwish -> direct lowering to HARD_SWISH (removes pseudo-op expansion)>1 without --output_nms_with_argmax via class-wise NON_MAX_SUPPRESSION_V4onnx2tf/ops/NonMaxSuppression.pyPad builtin enhancement
pads (rank-1, length 2*rank) via CAST+RESHAPE+TRANSPOSEAveragePool builtin enhancement
SAME_LOWER / count_include_pad in {0,1}count_include_pad=0 with non-zero effective paddingflatbuffer_direct rewrites)
CASTMEAN, PRELU, and HardSigmoid-expanded chains (including MUL+ADD+RELU_0_TO_1)CONV_2D -> MUL(const) -> ADD(const)MAXIMUM(0.0) -> MINIMUM(1.0) to RELU_0_TO_12.0.21pytest -q -k flatbuffer_direct
254 passed, 53 deselectedonnx2tf -i sinet_320_op.onnx -cotof -tb flatbuffer_direct --report_op_coverageONNX/TFLite output check is pass=TrueFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.20...2.0.21
This PR now includes all current changes from fix-resize-cubic (not only Resize).
This PR now includes all current changes from fix-resize-cubic (not only Resize).
RESHAPE + BATCH_MATMUL + RESHAPE + BATCH_MATMULcoordinate_transformation_mode (align_corners, asymmetric, half_pixel, pytorch_half_pixel)cubic_coeff_aexclude_outsidelower_from_onnx2tf.py to reduce redundant transpose insertion and improve NHWC propagation robustness.tests/test_tflite_builder_direct.pytests/test_tflite_builder_op_coverage.pytests/test_microsoft_domain_supplement.pyfix-resize-cubic -> mainFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.19...2.0.20
This PR improves flatbuffer_direct conversion coverage and robustness in onnx2tf, with a strong focus on layout-safe lowering, grouped convolution beh
This PR improves flatbuffer_direct conversion coverage and robustness in onnx2tf, with a strong focus on layout-safe lowering, grouped convolution behavior, and additional builtin operator support.
convert():
output_nms_with_argmax when NonMaxSuppression class-dimension constraints are detected.disable_group_convolution into direct lowering.Tile, Erf, ScatterND, and GlobalAveragePool.Einsum routing (FULLY_CONNECTED for const RHS, BATCH_MATMUL for non-const RHS).UNIDIRECTIONAL_SEQUENCE_LSTM lowering for forward LSTM direction.starts/ends attrs).SOFTMAX / pre-ARG_MAX transpose handling and preserve layout boundaries when needed.pytest -q tests/test_tflite_builder_direct.py tests/test_tflite_builder_op_coverage.py tests/test_tflite_builder_preprocess.py tests/test_accuracy_evaluator_input_layout.py227 passed2.0.19 (pyproject.toml, onnx2tf/__init__.py, README.md, uv.lock).Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.18...2.0.19
Expand flatbuffer_direct builtin lowering coverage and remove multiple CUSTOM fallbacks by adding direct dispatch/builders/validators for additional O
Expand flatbuffer_direct builtin lowering coverage and remove multiple CUSTOM fallbacks by adding direct dispatch/builders/validators for additional ONNX ops.
_DISPATCH_REGISTRY builtin path:
Abs, Acos, Acosh, And, ArgMin, Asin, Asinh, Atan, Atanh, BitShift, BitwiseAnd, BitwiseNot, BitwiseOr, BitwiseXor, Ceil, Celu, Cos, Cosh, Elu, Equal, EyeLike, Floor, GRU, GatherND, Gelu, Greater, Hardmax, Less, LessOrEqual, Mish, NonMaxSuppression, NonZero, Not, Or, RNN, Range, ReduceL1, ReduceL2, Round, Selu, Sign, Sin, Sinh, Softplus, Softsign, Tan, Trilu, Where, Xor.ArgMin, GatherND, Hardmax, NonZero, builtin NonMaxSuppressionRNN, GRUReduceL1, ReduceL2EyeLike, Range, TriluGRU: supports linear_before_reset={0,1} and direction in {forward, reverse, bidirectional} under flatbuffer_direct constraints.RNN: direct builtin path with explicit constraints/validation.model_writer.py:
ArgMinOptions, NonMaxSuppressionV4/V5Options, SequenceRNNOptionsoutput_nms_with_argmax through flatbuffer_direct lowering/report contexts.2.0.17 -> 2.0.18.update-builder.md, update-builder2.md.python -m py_compile passed for all changed python modules.Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.17...2.0.18
This PR expands flatbuffer_direct builtin coverage and stabilizes int8 model conversion flows used on the imp-int8 branch.
This PR expands flatbuffer_direct builtin coverage and stabilizes int8 model conversion flows used on the imp-int8 branch.
DynamicQuantizeLinearShapeConstantOfShapeFusedMatMul (including Microsoft-domain flow)OneHotMatMulIntegerPowReciprocalLRNLogSoftmaxLSTM path (BIDIRECTIONAL_SEQUENCE_LSTM lowering)
<img width="454" height="569" alt="image" src="https://github.com/user-attachments/assets/cb4efac5-7acf-4b98-ab57-a248c0fad3e1" />Resize builtin path:
sizes input (in addition to constant scales/sizes)ONNX_RESIZE custom-op in affected models (e.g. FCN-resnet50 int8 case)CAST, NEG, FLOOR_MOD, MAXIMUM, MINIMUM, BATCH_MATMULpytest -q tests/test_tflite_builder_op_coverage.pypytest -q tests/test_tflite_builder_direct.py -k "logsoftmax_lowering or lrn_lowering or pow_lowering or reciprocal_lowering or onehot_lowering or matmul_integer_lowering or dynamic_quantize_linear_lowering or shape_lowering or constant_of_shape_lowering or fused_matmul_lowering or resize_dynamic_sizes_lowering or reconcile_cast_propagates_output_shape_signature or reconcile_neg_propagates_output_shape_signature or reconcile_floor_mod_propagates_output_shape_signature or reconcile_maximum_propagates_output_shape_signature or reconcile_minimum_propagates_output_shape_signature or reconcile_batch_matmul_propagates_output_shape_signature"Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.16...2.0.17
Updated opcode emission for builtin codes >127 compatibility (deprecatedBuiltinCode handling).
This PR is focused on regression recovery after the 2.0.15 series, while keeping the TF backend stable and improving flatbuffer_direct robustness.
onnx2tf/ops/Conv.py and onnx2tf/ops/Transpose.py that caused channel/layout mismatch regressions.onnx2tf/ops/NonMaxSuppression.py to normalize:
boxes: [batch, num_boxes, 4]scores: [batch, num_classes, num_boxes]
This addresses reported shape mismatch failures.onnx2tf/onnx2tf.py:
ArgMax, Cast, Expand, GatherElements, Mod, ReduceMaxDiv(x, const) -> reciprocal precompute + MulCAST -> MUL -> CASTConv/DepthwiseConv + (RELU/RELU6/TANH) fusion into fusedActivationFunction in flatbuffer_direct IR.deprecatedBuiltinCode handling).2.0.16.onnx2tf -i convnext-det.onnx -cotof -tb flatbuffer_direct --report_op_coverage
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.15...2.0.16
This PR significantly improves the flatbuffer_direct path, focusing on conversion stability, quantization-chain optimizations, shape/NHWC propagation,
This PR significantly improves the flatbuffer_direct path, focusing on conversion stability, quantization-chain optimizations, shape/NHWC propagation, and redundant transpose reduction.
main): 9281976fix-var): 2d9ca0aonnx2tf \
-i text_detection_en_ppocrv3_2023may_int8.onnx \
-cotof \
-tb flatbuffer_direct \
--report_op_coverage
| INT8 ONNX | INT8 Flatbuffer |
|---|---|
| <img width="298" height="634" alt="image" src="https://github.com/user-attachments/assets/c253a826-081e-4439-aba2-bf1444bfe19d" /> | <img width="335" height="622" alt="image" src="https://github.com/user-attachments/assets/cefe6304-acb1-4bb9-8055-28b76cddd2c7" /> |
| <img width="1500" height="176" alt="image" src="https://github.com/user-attachments/assets/5142e47c-d005-4a71-8783-7aaa7849722d" /> |
shape / shape_signature propagation, mainly in lower_from_onnx2tf.py.Transpose-chain elimination around NCHW/NHWC bridges.Resize*, Conv, DepthwiseConv, ConvTranspose, Mean, Squeeze, and Concat.Dequantize -> HardSigmoid -> QuantizeDequantize -> TransposeConv -> QuantizeDequantize -> Reshape -> QuantizeDequantize -> MaxPool -> QuantizeDequantize -> Softmax -> QuantizeQLinearConv, QLinearConcat, QLinearAdd, QLinearMul, and QLinearSigmoid.bn_fold_wave0 (preprocess/rules/bn_fold.py)
Conv/ConvTranspose + BatchNormalization foldDequantize + BatchNormalizationConv -> Mul -> Add affine chainscleanup_unused_initializers_z9 (preprocess/rules/cleanup_unused_initializers.py)
Captures: lines during SavedModel export.ONNX2TF_SUPPRESS_TF_STARTUP_STDERR (default: enabled)flatbuffer_direct OP error and tensor-correspondence reporting.__init__.py, pyproject.toml).Validated mainly with real conversion runs (examples):
text_detection_en_ppocrv3_2023may_int8.onnxhuman_segmentation_pphumanseg_2023mar_int8.onnxpose_estimation_mediapipe_2023mar_int8bq.onnxface_recognition_sface_2021dec_int8.onnxAlso confirmed:
python -m compileall on key modified modulesbn_fold_wave0 and cleanup_unused_initializers_z9 statisticsFull Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.14...2.0.15
This PR consolidates recent fix-lp work for flatbuffer_direct INT8 conversion quality, diagnostics, and reportability.
This PR consolidates recent fix-lp work for flatbuffer_direct INT8 conversion quality, diagnostics, and reportability.
*_tensor_correspondence_report.json).*_op_error_report.json*_op_error_report.csvtools/flatbuffer_direct_op_error_report.pyonnx2tf/utils/flatbuffer_direct_op_error_report.pygenerate_op_error_report(...) to reuse report generation from onnx2tf.py.QLinearAdd, QLinearConcat, QLinearGlobalAveragePool, Reshape, Resizeop_builders/pool.py, op_builders/quantized.py, op_builders/shape.pyQLinearSoftmax.2.0.14 and README container tag references updated.<img width="1027" height="213" alt="image" src="https://github.com/user-attachments/assets/cbafca0f-73d5-4d02-a24d-4241a95b680f" />
Full Changelog: https://github.com/PINTO0309/onnx2tf/compare/2.0.13...2.0.14
Your coding agent can read these notes before it upgrades. Set up the MCP server →