NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2987 most downloaded on PyPI
Local-first SDK for clinical extraction and de-identification workflows on hardware you control.
Last release 19 days ago
15 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 41 of 44 stable releases
Nothing withdrawn
no release was ever pulled
1 years old
50 releases · first in 2025
One column per month.
The static Python API comparison against v2.3.0 reports zero removed symbols, zero narrowed signatures, and zero new deprecations . The REST API retai…
OpenMed 2.5 brings clinical privacy and extraction previews, richer FHIR and OMOP exports, safer document intake, and new privacy-policy and audit tools to the Python SDK. It preserves the v2.3 family-based model-registry API while adding an explicit API for managing separate model tiers and formats.
This release covers changes since v2.3.0. Python 3.10+ remains supported, and the core dependency set is unchanged.
ClinicalPrivacyProcessor combines explicit language, category, and role controls with document-level review status. The clinical_preserve profile targets identifiers while retaining clinical context such as diagnoses, medications, doses, and measurements. Full-document ONNX batching preserves source offsets and input order, with bounded cancellation and per-document outcomes.fa pack adds Persian patterns, Eastern-digit normalization, Jalali date handling, and Iranian national-ID surrogate support. Persian is also available in the REST language enum. Its default model route remains a compatibility placeholder; this release does not supply newly qualified dedicated Persian weights.Clinical-preserving processing is opt-in and remains a preview. No language/model route is clinically qualified by default; applications must establish qualification for their intended data, language, policy, and runtime.
to_fhir() facade dispatches supported entities to Condition, Observation, MedicationStatement, and Procedure resources, assembles deterministic Bundles, and reports unmapped labels. A conservative DiagnosticReport exporter adds explicit status handling and reference normalization.measurement, procedure_occurrence, visit_occurrence, observation_period, and note_nlp rows, preserving supported numeric values, units, dates, and source offsets. Cohort validation checks keys, relationships, vocabulary references, and note provenance.Synthetic FHIR/OMOP conformance coverage includes the official HL7 R4 validator, a malformed-resource negative control, and OMOP column, key, and reference checks. The Java validator and restricted vocabulary content are not bundled.
preflight_asset brings manifest, media-type, profile, resource-limit, and digest checks into one ordered accept-or-abstain report. An unevaluable check is reported as an abstention rather than a successful validation.
MOBILE_V1 and DESKTOP_V1 profiles bound bytes, dimensions, pages, frames, and duration before decoding. Batch summaries detect duplicate identifiers and digests and validate aggregate limits.New local tools make privacy decisions and their supporting evidence easier to inspect:
These reporting APIs use bounded metadata, counts, offsets, hashes, and provenance instead of retaining protected source values. Deletion, upload, trace cleanup, and other external effects remain explicit operations.
The optional Snowpark adapter adds caller-managed warehouse de-identification. The separate privacy-proxy application requires an injected transport, keeps mappings scoped to each request, and supports placeholder restoration across streaming responses.
Agent integrations gain strict provider-result and run-summary schemas, opaque correlation identifiers, typed governance identifiers, and artifact references with digest and size metadata.
Federated-training helpers add scheduling windows, round-status summaries, validated update metadata, and aggregate metrics with clipping declarations and minimum-group suppression. Retraining tools rank aggregate evidence, score trigger signals, and prepare recipe proposals. Compute, cost, energy, and carbon accounting adds explicit budget reports.
These additions provide orchestration and evidence contracts. They do not automatically train, deploy, or qualify a model.
Existing v2.3 registry callers keep their family-based API. RegistryService, family CLI selectors, unambiguous model aliases, and caller-owned schema-v1 files remain supported. An SDK upgrade does not require migrating those files.
Applications that need separate channels for model tiers or formats can opt into SlotRegistryService and schema-v2 slots keyed by family::tier::format. Stored slot versions are independent of version-like tokens in model repository names. The compatibility adapter retains existing schema-v2 files and refuses ambiguous multi-slot family operations without writing.
The static Python API comparison against v2.3.0 reports zero removed symbols, zero narrowed signatures, and zero new deprecations. The REST API retains its existing paths and component schemas, with Persian added to the language enum.
Other fixes include compatible Swift dependency resolution, stricter ONNX metadata validation, request-scoped privacy-proxy restoration, Windows artifact-deletion identity checks, and refreshed documentation and container dependencies.
The clinical, governance, and orchestration additions described above are Python capabilities. Matching npm, Swift, and Android version numbers do not imply that every new Python API is implemented on those platforms.
Use these coordinates once the v2.5.0 packages and tag are published:
# Python SDK
pip install --upgrade "openmed==2.5.0"
# Optional Hugging Face and FHIR integrations
pip install --upgrade "openmed[hf,fhir]==2.5.0"
# JavaScript package
npm install openmed@2.5.0Swift Package Manager:
.package(url: "https://github.com/maziyarpanahi/openmed.git", from: "2.5.0")Android through JitPack:
implementation("com.github.maziyarpanahi:openmed:v2.5.0")Container:
docker pull ghcr.io/maziyarpanahi/openmed:v2.5.0For Helm deployments, select the release image with image.tag=v2.5.0. Review the migration guide before adopting slot registries, new serialized schemas, or the clinical-preserving preview.
OpenMed is assistive software. Clinical outputs require qualified review and must not automatically determine diagnosis, treatment, or billing. Structural validation and synthetic regression tests do not establish clinical efficacy, deployment-specific privacy performance, or regulatory compliance.
Before deployment, evaluate identifier recall, critical leakage, source-span integrity, clinical-text preservation, date handling, and surrogate consistency on the intended languages and runtimes. Model artifacts and their existing release targets are unchanged by this SDK release.
Versioned links become available with the v2.5.0 tag:
Full Changelog: v2.3.0...v2.5.0
OpenMed 2.5 adds clinical privacy and extraction previews, local privacy and audit controls, FHIR and OMOP validation, bounded multimodal intake, and registry and training orchestration. This release compares against v2.3.0; there is no intervening v2.4.0 tag. See the release notes and migration guide.
Added a functional-status zero-shot NER domain with ADL, assistance, mobility, functional-scale, assistive-device, and cognitive-status labels, synthetic span fixtures, and offline per-label coverage reporting (#911).
Added complete detection of bounded German postal-address fields and fragment protection inside known clinical phrases, with person-name counterexamples and independent mask/remove/replace regression checks.
Added preview clinical-preserving privacy processing with explicit language, category and role controls, full-document ONNX tensor batching, bounded cancellation and per-document review status. Clinical protection, source offsets and output policy remain consistent across the safety sweep.
Added German clinical context, temporal and quantity extraction regressions, memory-streamed Tesseract OCR and PDF reading-order/redaction checks. These preview capabilities require independent task and language qualification.
Added bounded BMP CORE/INFO header geometry preflight with explicit limits, value-free errors, and synthetic file-level regression tests (#3114).
Added bounded GIF logical-screen and bounded global-color-table preflight with explicit limits, value-free errors, and synthetic file-level regression tests (#3115).
Added tiny RF64 refusal fixtures that pin WAV envelope rejection, short-read boundaries, and stream restoration without changing the parser (#3117).
Added bounded VP8, VP8L, and VP8X WebP geometry preflight with explicit limits, value-free errors, and synthetic file-level regression tests (#3116).
Added structural locale normalization and explicit alias duplicate/unsupported-format checks with synthetic file-level regression fixtures (#3119).
Added a privacy-safe multimodal preflight report (preflight_asset) that
runs manifest validation, bounded media-type detection, modality profile
checks, limit-profile evaluation, and a bounded digest pass in a fixed order
and returns one accept-or-abstain PreflightReport with ordered, allowlisted
findings, a preflight AbstentionRecord, and byte-stable JSON; unevaluable
checks, including a PDF's pixel rules, abstain rather than accept (#2980).
Added immutable pre-decode limit profiles for multimodal assets (MOBILE_V1,
DESKTOP_V1) with inclusive ceilings for bytes, pages, pixels per unit, total
pixels, frames, and audio duration, evaluated over the privacy-safe asset
manifest into deterministic LimitFinding records; unevaluable rules,
including a PDF's pixel rules, are reported as insufficient_metadata rather
than assumed safe (#2956).
Added an exact schema version and strict, bounded dictionary and JSON parsers for privacy-safe agent run summaries, including duplicate-field, non-finite number, unknown-field, and unsupported-version rejection (#3038).
Added strict, content-free multimodal provider result envelopes with bounded counts and timing, deterministic serialization, and value-free failures (#3006).
Added an optional Snowpark adapter and generated Python UDF SQL for in-warehouse text de-identification with lazy dependency loading, compatible Snowpark registration, and escaped SQL literals (#2369).
Added an immutable, provenance-aware local terminology cache keyed by exact vocabulary releases, with deterministic response fingerprints, stale-release refusal, response-free reports, and bounded value-free validation (#2400).
Added a bounded, deterministic, offline policy-migration checker with fail-closed schema and protection-type changes, privacy-safe reports, and a report-bound human acknowledgement gate for weakening changes (#2407).
Added an in-memory no-PHI telemetry exporter with closed counter families, bounded dimensions and totals, atomic event validation, exception-type-only categorization, deterministic JSON and Prometheus rendering, and no mandatory network transport (#2414).
Added a bounded, deterministic tabular re-identification risk report with aggregate-only JSON and Markdown renderers, immutable report state, fail-closed consistency checks, and locally derived threshold outcomes (#2411).
Added a deterministic local pre-push privacy scanner that checks every new commit blob for direct identifiers, secrets, and sensitive structured fields; emits value-free reports; supports narrowly versioned synthetic-fixture allowlists; and installs atomically while preserving existing hooks (#2298).
Added a versioned, bounded privacy policy-as-data schema for jurisdiction, recall floors, de-identification actions, surrogate strategy, and privacy-safe audit retention, with deterministic local-only loading, strict duplicate and alias handling, and value-free validation failures (#2406).
Added a bounded, thread-safe privacy budget ledger for named aggregate release contexts with atomic epsilon/delta charging, counts-only evidence, immutable configuration views, and value-free failures (#2410).
Added caller-owned HMAC-SHA256 audit-report key rotation with bounded key material, key-ID based current and retained-key verification, canonical mapping checks, fail-closed provider handling, and value-free failures (#2408).
Added a bounded, deterministic minimum-necessary structured field selector with caller-declared purpose mappings, policy allowlists and denylists, fail-closed unknown declarations, value-free decision explanations, and projection restricted to selector-approved fields (#2412).
Added a bounded, counts-only audit-artifact retention planner with explicit disposition rules, deletion evidence, remaining-set verification, strict input fields, and fail-closed future timestamps (#2409).
Added a deterministic, offline CycloneDX 1.6 evidence generator for the base runtime dependency closure, with source and manifest hashes, bounded local inputs, atomic output, and no embedded URLs or build paths (#2416).
Added a deterministic offline dependency risk report that correlates local locked versions with caller-supplied advisory snapshots, emits bounded value-free risk summaries, and performs no package-manager or network calls (#2417).
Added deterministic counts-only trace privacy audit artifacts with canonical policy and file hashes, immutable category counts, value-free JSON and Markdown renderings, stable file fingerprinting, and private atomic writes (#2302).
Added deterministic, PHI-free local release compute, cost, energy, and carbon tracking with orchestrator-linked stage timings, per-run and rolling budget verdicts, family/tier/workload breakdowns, optional advisory queue throttling, and hash-verified ledger replay (#1244).
Added a bounded, deterministic nested-resource redaction contract with explicit scalar paths and actions, stable arrays and identifiers, closed policy validation, and raw-value-free reports and failures (#2413).
Added declarative field-level FHIR and OMOP de-identification policies with fail-closed identifier handling, patient-consistent date shifting, schema linting, CSV/Parquet support, and resumable FHIR NDJSON integration (#2187).
Added the canonical grounded-span to_fhir() facade with label-driven
Condition, Observation, MedicationStatement, and Procedure dispatch,
deterministic Bundle assembly, PHI-free exported/unmapped label counts, and
graceful skipping for labels without an exporter. The facade remains the
same callable as the established grounded exporter and never synthesizes a
Patient resource.
Added a deterministic, local key-custody metadata validator for synthetic signing and surrogate workflows, with lifecycle transition checks, purpose/algorithm compatibility, digest-only reports, and fail-closed rejection of bytes, secret-like, or unknown fields (#2648).
Added bounded local deletion verification for fingerprinted sensitive artifacts, with symlink, alias, and hard-link refusal, independent recovery copies, commit-stage rollback, and counts-only evidence (#2418).
Added a bounded, manifest-driven deletion impact planner with deterministic counts-only reports, reverse-dependency analysis, ownership checks, and explicit plan-bound confirmation before injected local execution (#2529).
Added a deterministic, offline OMOP cohort export validator for key, relationship, vocabulary, and NOTE/NOTE_NLP provenance invariants, with aggregate counts and content-derived row fingerprints instead of source values (#2402).
Added a bounded, deterministic FHIR R5 Bundle round-trip fidelity diff with stable entry matching, explicit serializer-difference declarations, and value-free reports containing structural paths, types, and SHA-256 digests (#2401).
Added bounded, deterministic structured access reviews that compare workflow read and export declarations with resource schemas and deny policies while keeping schema values out of JSON, Markdown, and validation failures (#2419).
Added bounded, policy-aware diffs for aggregate redaction summaries, with closed value-free inputs, deterministic policy fingerprints, and structured action, category, and count changes (#2426).
Added bounded privacy-policy composition with explicit scope-overrides, deterministic scope precedence and inheritance, and validated value-free decision traces (#2522).
Added a bounded, metadata-only synthetic privacy regression corpus manifest with deterministic fixture hashes, policy and severity coverage validation, immutable invariants, and atomic local persistence (#2420).
Added a bounded, local evidence-bundle integrity verifier with file and manifest hashes, policy and provenance checks, and value-free reports (#2427).
Added a bounded tabular schema-drift privacy gate with counts-only evidence, conservative stable-ID matching, and release blocking for unsafe role or structural drift (#2524).
Added a bounded, deterministic referential-integrity auditor for surrogate maps with cardinality, collision, orphan, and cross-table consistency checks, closed input schemas, and counts-only value-free reports (#2538).
Added a bounded, deterministic nested structured-redaction idempotence checker for comparing shape, action, surrogate, policy, and count evidence across synthetic FHIR- and OMOP-shaped passes without retaining protected values (#2523).
Added a bounded, offline privacy evidence replay verifier with counts-only synthetic manifests, stable policy/environment/result fingerprints, and privacy-safe schema, environment, policy, and result drift reports (#2527).
Added an offline manifest-coherence regenerator and CI drift gate for the runtime model registry, PII language defaults, governed README counts, registry model cards, and generated model and benchmark documentation (#77).
Added exact OMOP CDM v5.4 visit_occurrence, observation_period, and
note_nlp exporters with deterministic local keys, bounded clinical dates,
source offsets, and assertion-derived NLP term fields (#2360).
Added exact OMOP CDM v5.4 measurement and procedure_occurrence row
exporters with shared Athena concept resolution, deterministic unmapped
fallback, and preservation of numeric lab values, units, and ranges (#275).
Added dependency-free US Core 9.0.0 conformance checks for exported Condition, laboratory Observation, MedicationRequest, and AllergyIntolerance resources, including base-R4-first validation, must-support warnings, required-binding errors, canonical profile resolution, and compact CC0 constraint metadata (#2366).
Completed the synthetic grounding/export conformance suite with fail-closed
out-of-process HL7 FHIR R4 validation, an official-validator malformed
resource negative control, expanded ACHILLES-style OMOP column/key/reference
checks, and paired JSON/Markdown BenchmarkReport artifacts (#2359).
Added a versioned federated update metadata envelope with coordinator-owned parameter expectations, bounded exact shape arithmetic, deterministic JSON, clipping declarations, and value-free rejection of unknown or identifying fields (#3010).
Added typed, canonical governance identifiers for capabilities, purposes, policies, workflows, and tools, with shared validation and value-free diagnostics (#3042).
Added a versioned no-PHI exception taxonomy for telemetry and audit records, with owner-free approval metadata, bounded digest-only evidence, explicit UTC expiry checks, deterministic serialization, and value-free validation failures (#2528).
Added a bounded, deterministic audit-envelope parser with redacted payload metadata, canonical fingerprints, strict schema and signature validation, and value-free diagnostics (#2594).
Added a deterministic, local-only privacy exception budget gate that counts bounded synthetic waiver metadata by severity, scope, expiry, and policy fingerprint, failing closed on exceeded or unbounded exceptions (#2591).
Added local, deterministic FHIR ValueSet expansion over caller-loaded free
vocabulary snapshots plus explicit FHIR $expand and ECL delegation to a
caller-supplied terminology endpoint. Results include versioned provenance;
caching is user-controlled, and restricted member codes are never persisted
without a second explicit policy opt-in (#926).
Added closed, versioned federated aggregate metric envelopes with finite clipping bounds, minimum-group suppression, coarse participant bands, controlled privacy mechanisms, confidence intervals, deterministic JSON, and value-free rejection of client-level or unknown fields (#3011).
Added dependency-free base FHIR R4 structural validation for eight exported clinical resource types, including deterministic structured findings, cardinality and primitive datatype checks, fixed required bindings, Bundle aggregation, and a compact CC0-derived constraint table (#2364).
Added typed, 128-bit opaque correlation identifiers for agent runs and actions, with strict kind-aware parsing, deterministic metadata-only JSON, parent-action validation, and value-free failures (#2973).
Added strict, content-free agent artifact references with opaque identifiers, a closed artifact-kind vocabulary, versioned schema IDs, digest and size metadata, deterministic JSON, and value-free validation failures, including oversized integers and deeply nested JSON (#2999).
Added a conservative, deterministic FHIR DiagnosticReport exporter with
R4/R5 union allowlisting (32-field), explicit unknown status, type-gated
scalars and Reference normalization, effective[x] mutual exclusivity,
deep-copy evidence preservation, field-name-only value-free errors, and
no network or clock dependency (#2566).
Added privacy-safe multimodal asset batches with opaque batch identifiers, canonical asset ordering, duplicate identifier and digest detection, a bounded asset count, derived byte, page, frame, and duration totals, and sorted value-free findings for invalid, oversized, overflowing, or inconsistent batches (#3002).
Rekeyed the committed model-registry state to schema v2: sparse
family::tier::format release-channel slots (the baseline_key
convention shared by gates/baseline.json, gates/rollout_state.json,
and the release ledger), created only by coordinate-matched RELEASABLE
promotions, with assigned per-slot SemVer that is validated as stored
state and never recomputed from repo-id version tokens. Ships a
fail-closed one-time v1 migration (registry_ctl.py migrate) that maps
pointers through committed baseline coordinate evidence and leaves the
file unchanged on any ambiguity (#1804).
Preserve the v2.3 family registry API, CLI selectors, serialized views, and
unambiguous aliases; expose slot operations through SlotRegistryService
with an explicit v2 state contract and fail-closed compatibility adapter.
Pin Swift tokenization to the validated 0.1.24 release so clean package resolution cannot select an incompatible MLX dependency graph.
Add SDK-only readiness evidence for unchanged model artifacts and pointer targets while retaining signed model gates for model releases.
Validate ONNX label metadata before importing optional runtimes; malformed labels now fail at the metadata boundary.
Keep local privacy-proxy request mappings scoped to one request and reject unknown, duplicate, or malformed placeholders on inbound restoration.
Refresh Debian certificate and OpenSSL package pins used by the container build and validate release artifact size budgets against measured growth.
Fixed verified artifact deletion and rollback on Windows Python 3.12 by comparing explicit creation timestamps across pathname and descriptor stat results, while retaining identity and in-read mutation checks.
OpenMed 2.5 expands the Python SDK with clinical privacy and extraction previews, local privacy and audit controls, bounded multimodal intake, FHIR and OMOP validation, and metadata-only training orchestration. The release baseline is v2.3.0; there is no intervening v2.4.0 release.
This page describes the candidate. Registry publication, immutable package tags, and deployed documentation are verified separately after publication.
The release adds local policy schemas, composition and migration checks, minimum-necessary field selection, field-level FHIR/OMOP policies, tabular risk and schema-drift reports, surrogate integrity checks, and structured redaction and idempotence evidence.
Audit tooling covers access scope and expiry, key custody and rotation, evidence integrity, replay, lineage and freshness, report size and cardinality, waiver lifecycle, exception budgets, policy coverage and simulation, and release-evidence aggregation. Reports use declared identifiers, counts, offsets, hashes, and bounded metadata rather than protected source values.
Encrypted surrogate mappings, local deletion planning and verification, retention cleanup, pre-push and CI privacy scanning, dataset-upload guards, and session-end trace scrubbing are explicit local operations. The optional Snowpark adapter supports caller-managed warehouse processing; it does not turn the core package into a network service.
Read the 2.3-to-2.5 migration guide, including
registry compatibility and the versioned model-state migration. Python 3.10+
remains supported. The core dependency set stays unchanged; Snowpark is
optional and the grounding-validate extra does not bundle the Java validator.
After publication, the intended coordinates are:
pip install --upgrade "openmed==2.5.0"
pip install --upgrade "openmed[hf,fhir]==2.5.0"
npm install openmed@2.5.0
.package(url: "https://github.com/maziyarpanahi/openmed.git", from: "2.5.0")
implementation("com.github.maziyarpanahi:openmed:v2.5.0")
Container consumers select ghcr.io/maziyarpanahi/openmed:v2.5.0.
The Helm application version and default tag are 2.5.0; explicit image
selection uses image.tag=v2.5.0.
The new clinical, privacy-governance, and orchestration APIs described above are Python capabilities. Updating the npm, Apple, and Android package versions does not imply these new APIs run natively on every platform.
Synthetic regression tests and package checks establish the behavior they exercise. They do not establish clinical efficacy, deployment-specific identifier recall, regulatory compliance, or clinical-device suitability. Clinical outputs require qualified review and must not automatically determine diagnosis, treatment, or billing.
This SDK candidate preserves the existing model pointer targets and model artifacts. The registry state representation changes independently of those targets. A model promotion requires its own staged golden and public SHIELD benchmark evidence, signed gates, and authorization under the release-stream policy.
The complete change record is in the changelog.
Package, artifact, Android, Pages, and container budgets are explicit and blocking. Container digest policy, vulnerability scanning, SBOM generation,…
OpenMed v2.3.0 expands the stable local-first healthcare AI SDK across privacy-safe agents and traces, bounded multimodal intake, clinical evidence, local training, cross-platform runtimes, deployment adapters, and release hardening without removing a public Python symbol.
This release adds deterministic asset manifests, streaming digests, media-type detection, typed abstention, document and email redaction, privacy-safe agent outcomes, trace sharding and recovery, clinical evidence tables, patient-record filtering, training reproducibility records, WebGPU and Android acceleration, iOS extensions, Electron and Tauri bridges, TensorRT and GGUF runtimes, search and batch adapters, self-hosted services, Kubernetes operations, and stable CLI and schema contracts.
The committed ancestry since v2.2.0 contains 253 commits and 652 changed files. Commit subjects reference 127 unique issue or pull-request identifiers, while GitHub-generated release notes associate 129 pull requests with the complete ancestry before the final commit. The static public Python inventory grows from 37,735 to 41,729 symbols with 3,994 additions, no removals or renamed symbols, no narrowed callable signatures, and no new deprecations. The REST contract remains compatible at 19 paths and 17 component schemas.
Install or upgrade after the immutable tag and tag-driven package workflows complete:
pip install --upgrade "openmed==2.3.0"
pip install --upgrade "openmed[hf,fhir]==2.3.0"
pip install --upgrade "openmed[mlx]==2.3.0"
npm install openmed@2.3.0Swift Package Manager:
.package(url: "https://github.com/maziyarpanahi/openmed.git", from: "2.3.0")Android through JitPack:
implementation("com.github.maziyarpanahi:openmed:v2.3.0")Container and Helm after publication:
docker pull ghcr.io/maziyarpanahi/openmed:v2.3.0
helm upgrade --install openmed-service deploy/helm/openmed-service --set image.tag=v2.3.0Release documentation after the tag is available: https://github.com/maziyarpanahi/openmed/blob/v2.3.0/docs/release/v2.3.0.md
Migration guide after the tag is available: https://github.com/maziyarpanahi/openmed/blob/v2.3.0/docs/migration/2.2-to-2.3.md
Full changelog: v2.2.0...v2.3.0
OpenMed 2.3 introduces a dependency-free, versioned asset manifest with bounded metadata, strict digest and media validation, deterministic serialization, and value-free rejection of unsafe paths, URLs, free text, and unknown fields. Streaming digest helpers preserve caller-owned stream positions, and prefix inspection recognizes PDF, PNG, JPEG, TIFF, DICOM, and WAV without trusting filename extensions.
Typed abstention records and deterministic image, PDF, DICOM, and audio profiles let applications reject malformed or unsupported assets before opening or decoding them. New archive-extraction, resource-path portability, file-sharding, artifact-inventory, and export-filename policies preserve counts, hashes, offsets, and provenance without retaining source values.
Document processing adds conservative multi-column PDF reading order, local burned-in PDF redaction with source-text removal verification and masked fidelity evidence, HTML visible-text extraction with source offsets, EML redaction, an explicitly isolated optional MSG bridge, RTF and fixed-width text intake, and privacy-safe notebook-cell handling. These paths remain bounded, local-first, and fail closed on malformed or unbounded input.
Agent runs gain a closed outcome vocabulary, deterministic summaries, monotonic timing records, consent verification results, bounded tool-call metadata, and content-free failure reasons. These records are designed for stable machine consumption without accepting arbitrary free-text reasons or source clinical values.
Trace tooling adds metadata-only store discovery, JSONL and structured tool-call redaction, secret detection, bounded streaming, deterministic multiprocess sharding, local inventories, transactional in-place replacement with verified backups, crash recovery, and structural-fidelity proof. Training adapters cover role-message, preference-pair, and columnar schemas while preserving non-content structure and emitting hashed diagnostics.
Local training adds an audited teacher-ensemble registry, federated-round lifecycle records, and reproducibility-hash recomputation and verification. These are additive evidence contracts and do not make training output clinically authoritative.
OpenMed 2.3 adds deterministic clinical evidence tables with source offsets, controlled assertion and review metadata, optional protected-value hashes, and value-free JSON and Markdown rendering. Scoped substance-use SDOH extraction and experiencer-aware patient-record filtering separate patient-eligible, family, other, hypothetical, negated, and refuted evidence with auditable reasons.
Consent receipt verification is immutable and non-throwing, with stable content-free outcome codes. Clinical evidence, SDOH extraction, record filtering, terminology mappings, and consent results remain assistive software requiring qualified review; they must not automatically trigger diagnosis, treatment, billing, data release, or publication.
Browser support adds typed batched WebGPU token classification, an audited WGSL head, deterministic local WASM fallback, parity and recall gates, and a zero-upload privacy playground. Android gains QNN and NNAPI provider selection with deterministic CPU fallback and release-AAR size and cold-start budgets. Apple platforms gain reusable Share and Action extension modules for bounded local redaction. Electron and Tauri integrations preserve renderer isolation, bounded IPC, offline enforcement, and safe worker recovery.
The npm package exports DEFAULT_MODEL_ID as OpenMed/OpenMed-PII-ClinicalE5-Small-33M-v1-onnx-android and routes this model family through the root INT8 loader. extractPii() retains O tokens and reconstructs missing character offsets before BIO decoding. alignTokenOffsets() is available for custom pipelines. Final spans use JavaScript UTF-16 offsets, preserve decomposed accents and supplementary Unicode letters, and fail with a content-free error when tokens cannot be aligned. The real public model passed a synthetic local redaction smoke test with matching ESM and CommonJS results; this is functional evidence, not a clinical recall qualification.
Model runtimes add device-specific TensorRT export, local Q4_K_M GGUF grounding through caller-provided llama.cpp, expanded MLX architecture coverage and decoding performance, and explicit edge-SBC benchmark profiles. Model files and external runtimes remain caller-supplied, pinned, checksum-verified trust boundaries.
Opt-in integration contracts now cover OpenSearch, Elasticsearch, Spark, Beam, Airflow, LlamaIndex, PostgreSQL, dbt, BigQuery-compatible remote functions, KServe, Triton, a namespaced Kubernetes model operator, HPA guidance, a hardened self-hosted Compose bundle, and a loopback-only local redaction service. No external process, database, cluster, browser, credential, or telemetry path becomes mandatory for core local PHI processing.
The release adds deterministic openmed init scaffolding, local file redaction, stable JSON result envelopes, shell completion, CLI help drift checks, configuration and schema snapshots, integration capability discovery, offline installation diagnostics, and portable Agent Skills bundles and validation.
The canonical contributor workflow uses the committed uv lock, with Pixi and Nix paths for reproducible environments. Package, artifact, Android, Pages, and container budgets are explicit and blocking. Container digest policy, vulnerability scanning, SBOM generation, provenance tests, secret scanning, repository policy, license policy, immutable lock verification, and deterministic install checks remain part of the release boundary.
Hosted model conversion and publication automation was removed. Model conversion, evaluation, and publication are explicit local maintainer operations and remain separate from this SDK release.
The static Python API comparison against v2.2.0 records 3,994 additions, no removed or renamed public symbols, no narrowed callable signatures, and no newly deprecated symbols. The REST contract remains at 19 paths and 17 component schemas.
Swift, Kotlin/Android, JavaScript, CLI, configuration, trace schemas, evidence records, model artifacts, browser contracts, and deployment surfaces expand. npm consumers retain the v2.2 required numeric offsets on TokenClassificationEntity, TokenClassificationPipeline, and model-loader output. The additive RawTokenClassificationEntity, RawTokenClassificationPipeline, and RawTransformersRuntime types accept offset-less runtime input; alignTokenOffsets() converts it to aligned entities. Model loaders preserve runtime metadata and resource disposal. Final OpenMedSpan offsets remain required. Treat alignment errors as failed scans, never as evidence of PII-free content. Applications should refresh generated clients and schema snapshots, explicitly handle new closed vocabularies, re-qualify enabled document formats and runtimes, and repeat privacy acceptance testing on the actual deployment platform.
See the complete migration checklist: https://github.com/maziyarpanahi/openmed/blob/v2.3.0/docs/migration/2.2-to-2.3.md
OpenMed 2.3.0 is an SDK release and does not promote a model pointer. The committed PII/latest and PII/last_green pointers both remain OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1-mlx, and PII/canary remains unset.
No new model-candidate report is claimed for this SDK-only release. Any future change to canary, latest, or last_green still requires a real staged candidate, public SHIELD and golden evidence, signed extraction gates, and a final readiness decision of exactly READY.
f7ebc7eaeb81b0a76f5fd5784830a63895662578; the working tree is clean and current origin/master (0e07ba71b7895255d8d14d4a25642e5cb37209d8) is an ancestor. All 19 inspected release surfaces resolve to 2.3.0 or v2.3.0; the target tag is unused locally and on origin.f7ebc7eaeb81b0a76f5fd5784830a63895662578: 44 checks passed, 3 expected conditional jobs were skipped, and none failed or remain pending. All seven Python unit jobs, Nix, Pages/browser tests, Android, OpenMedKit/iOS/watchOS/visionOS, amd64/arm64 container smoke, image SBOM, browser extension, and the image/lockfile vulnerability scan passed. Expected skips were Pages deployment, container manifest publication, and the unchanged-path CJK/Indic job.CVE-2026-81726 in nltk==3.10.3, with no fixed version. NLTK is present only through the optional agents, llamaindex, quickumls, and scrubadub dependency trees and is absent from the service image. OpenMed does not call the affected model-artifact APIs. The maintainer approved a seven-day exception restricted to this CVE, the nltk package, and the uv.lock target, expiring on 2026-09-11. The scanner threshold and fixed-version enforcement remain unchanged, and the exception cannot suppress a future finding that names a fixed version. Upstream advisory: GHSA-8mgp-746c-j5xpopenmed package passed installation, zero-finding npm audit, all 30 tests, ESM/CommonJS builds and declarations, strict legacy-consumer type compilation, and an 18-file package dry run. Real Transformers.js 3.8.1 inference on synthetic text passed with ESM/CommonJS parity and successful resource disposal. The browser extension passed its build, typecheck, and end-to-end browser test.pypi and npm GitHub environments have their expected token secret names and no protection rules or branch restrictions. PyPI and npm both return HTTP 404 for version 2.3.0, so the version remains available. The same workflow's latest tag run, for v2.2.0 on 2026-08-21, succeeded in all six jobs: npm verification, Python build/attestation/verification, PyPI publication, npm publication, signed release evidence, and release SBOM. Credential values were neither read nor changed; their current validity is exercised only by the real tag-driven publication.master triggers the production Pages deployment and the normal branch container build. Pushing the immutable v2.3.0 tag triggers Python provenance and PyPI publication, npm verification and publication with provenance, the GitHub release with distribution evidence and SBOMs, the multi-architecture GHCR image with provenance and follow-on signing, image SBOM generation, and Android release validation. The tag is the JitPack release coordinate. Maven Central upload is an optional Android path and will be skipped because its four signing/Sonatype secrets are not configured; this does not fail the tag workflow or affect the documented JitPack release.OpenMed keeps local processing as the default, but no de-identification, structured privacy, or clinical evidence system can guarantee zero residual risk. Validate direct-identifier recall, critical leakage, span integrity, language and script coverage, format handling, policy behavior, quantized-model deltas, and device behavior on deployment-specific synthetic fixtures before production use.
Clinical extraction, evidence, filtering, mapping, and model-generated output are assistive software, not a medical device or a source of clinical ground truth. Outputs require qualified review and must not automatically trigger diagnosis, treatment, billing, data publication, or other clinical decisions.
Licensed vocabularies, real-patient records, DUA-gated corpora, UMLS, SNOMED CT, CPT, MIMIC, i2b2, and n2c2 assets are not bundled. Restricted integrations require caller-supplied rights, keys, snapshots, model weights, or out-of-process bridges.
GitHub-generated notes associate 129 pull requests with the complete release ancestry and identify 17 contributors: @janithcd, @maziyarpanahi, @AadityaAnand, @alberthammerich, @dpersek, @Hitesh-XS, @anudit, @DrVelvetFog, @mohamedhossammohamed, @PouyanJay, @Mr-Neutr0n, @Rahul-pamula, @wessim852, @krudo-taco, @aayushi-sing, @kkkhs, and @09Catho.
New contributors identified by GitHub-generated notes: @janithcd, @AadityaAnand, @Hitesh-XS, @anudit, @Rahul-pamula, @wessim852, @aayushi-sing, @kkkhs, and @09Catho.
Associated pull requests: #2697, #2907, #2896, #2898, #2909, #2675, #2912, #2910, #2902, #2903, #2913, #2901, #2897, #2900, #2906, #2908, #2904, #2905, #2899, #2914, #2214, #2309, #2310, #2314, #2311, #2317, #2318, #2325, #2319, #2320, #2329, #2328, #2330, #2327, #2464, #2332, #2493, #2431, #2436, #2334, #2333, #2331, #2335, #2336, #2337, #2494, #2919, #2491, #2445, #2497, #2447, #2498, #2480, #2487, #2486, #2921, #2927, #2492, #2500, #2602, #2667, #2502, #2580, #2455, #2883, #2461, #2254, #2930, #2933, #2928, #2932, #2929, #2931, #2616, #2508, #2255, #2257, #2503, #2489, #2506, #2505, #2259, #2258, #2504, #2507, #2277, #2633, #2946, #2962, #2268, #2252, #2437, #2452, #2454, #2893, #2935, #2936, #2937, #2938, #2939, #2940, #2941, #2942, #2943, #2944, #2945, #2947, #2965, #2890, #2967, #2988, #2634, #2972, #2915, #2969, #2882, #2881, #2917, #2916, #2892, #2884, #3015, #2985, #2984, #2970, #2963, #2983, #2987, and #2960.
The GitHub-generated PR set differs from the 127 identifiers extracted from commit subjects because generated release notes associate merged and preserved contributor work through GitHub metadata, while the subject audit counts only identifiers literally present in commit subjects.
f7ebc7eaeb81b0a76f5fd5784830a63895662578 and its exact-head checks remain green.master commit, all version surfaces, checks, and unused v2.3.0 tag.v2.3.0 tag only after maintainer approval.Note truncated.
OpenMed 2.3 expands the stable v2 contract across privacy-safe agent and trace
workflows, multimodal asset intake, clinical evidence, local training,
cross-platform runtimes, deployment adapters, and release hardening. The final
audited v2.2.0..v2.3.0 release-branch range contains 252 commits and 651 changed
files.
The static public Python surface grows from 37,735 to 41,729 symbols with 3,994 additions, zero removals or narrowed signatures, and zero new deprecations. Python, Swift, Kotlin/Android, JavaScript, REST, CLI, configuration, serialized evidence, and deployment contracts are reviewed in the 2.2-to-2.3 migration guide.
Added a dependency-free, versioned multimodal asset manifest with strict media and digest validation, bounded metadata-only fields, deterministic JSON serialization, and value-free rejection of paths, URLs, free text, and unknown fields (#2954).
Added bounded streaming SHA-256 asset digests with caller-owned stream position restoration and value-free limit and read failures (#2979).
Added bounded, dependency-free detection for PDF, PNG, JPEG, TIFF, DICOM, and WAV prefixes, with stable match, mismatch, and unknown validation results that do not log source bytes or trust filename extensions (#2955).
Added strict, deterministic multimodal abstention records with typed pipeline stages, stable reason codes, metadata-only JSON, and value-free validation failures (#2977).
Added image, PDF, DICOM, and audio profiles that validate canonical manifest metadata into deterministic field-and-reason findings without opening or decoding an asset (#2978).
Added a closed, JSON-safe agent outcome vocabulary with success, abstention, reviewer-handoff, policy-denial, and failure classes, deterministic serialization, and value-free rejection of unknown codes or free-text reasons (#2950).
Added bounded, deterministic agent-run summaries for closed outcomes, workflow identifiers, tool-call counts, durations, and artifact digests, with direct-construction invariants and value-free privacy failures (#2951).
Added deterministic monotonic timing metadata records for agent runs and actions with exact integer durations and value-free validation failures (#2974).
Added deterministic clinical evidence tables with source offsets, controlled assertion and review metadata, optional protected-value hashes, and value-free JSON and Markdown rendering (#2567).
Added immutable, non-throwing consent receipt verification results with stable content-free outcome codes while preserving one-time receipt consumption.
Added an audited teacher-ensemble registry for weak labeling with manifest-resolved PII and Privacy Filter members, bounded weights and agreement thresholds, checksum-validator policies, and fail-closed runtime source matching (#284).
Added an experiencer-aware patient-record span filter
(openmed.clinical.filter_patient_record) that partitions per-span
ClinicalAssertion records into patient-record eligible and excluded sets.
Non-patient experiencers (family and other) and hypothetical spans are
excluded with auditable reasons; negated patient spans are retained and
marked refuted. Includes a medical-device-style advisory disclaimer that
the filter is a record-construction aid, not a clinical decision (#2251).
Added a deterministic Jupyter notebook cell redaction helper that preserves code sources and execution structure, applies explicit markdown and output policies, removes unredacted binary MIME data, and emits counts-only, value-free summaries and failures (#2561).
Added a license-quarantined MedCAT/CogStack subprocess bridge
(openmed/interop/bridges/medcat.py) that shells out to a user-provided
MedCAT process and maps its {cui, name, score} concept output onto
OpenMed span-code fields ({system, code, score}). MedCAT is Elastic
License 2.0 and is never imported in-process or bundled; invocation is
blocked until the caller acknowledges the license via
OPENMED_ACCEPT_MEDCAT_LICENSE or an interactive prompt. Added an empty
interop-gpl extra documenting that it installs nothing (#1789).
Added a deterministic offline resource-path portability audit with bounded inputs; traversal, root, reserved-name, normalization, and case-fold checks; immutable hash-only reports; and value-free failures (#2637).
Added a bounded, metadata-only archive extraction safety policy with cross-platform traversal and link rejection, normalized duplicate detection, expansion limits, and immutable counts-only decisions (#2635).
Added a deterministic export filename policy derived from validated artifact metadata, schema versions, and short provenance fingerprints, with path, raw-identifier, clock-derived, and value-leaking input rejection (#2584).
Added a deterministic offline artifact inventory with bounded safe-path handling, byte counts, media types, SHA-256 fingerprints, and aggregate-only JSON and Markdown reports (#2581).
Added a canonical CLI result envelope with bounded counters, artifact fingerprints, and remediation codes; strict JSON parsing; immutable state; and free-text-free failures (#2636).
Added a deterministic CLI help-surface drift checker with canonical command, option, argument, and default snapshots plus machine-readable compatibility reports (#2583).
Added a deterministic structured-schema snapshot compatibility checker with versioned field-path, type, and optionality rules plus value-free change evidence and canonical JSON output (#2582).
Added fail-fast JSON Schema validation for OpenMedConfig, TOML files, and
custom profiles, with aggregated value-free diagnostics, an installed schema
path helper, and complete remote-backend field coverage (#2264).
Added a dependency-free OpenSearch ingest redaction processor with validated local policies, explicitly selected fields, immutable document copies, cache-only defaults, and aggregate value-free diagnostics (#2389).
Added a dependency-free Elasticsearch ingest redaction processor with explicit static field rules, deterministic pipeline serialization, injected local redaction, and counts-only value-free diagnostics (#2388).
Added device-specific TensorRT engine export for ONNX token classifiers with bounded dynamic shape profiles, FP16 and fail-closed INT8 calibration, per-family G4 recall evidence, finite synthetic parity checks, rollback-safe engine and metadata publication, trusted-engine logits inference, and device-tier benchmark records (#834).
Added configurable Android QNN and NNAPI execution-provider selection with deterministic CPU fallback, per-family operator-coverage reporting, and bounded PHI-free latency, span-parity, and recall evidence (#851).
Added a dependency-optional Apache Beam redaction transform with explicit schema metadata, bounded record and byte state, capped retries, deterministic serialization, cache-only defaults, and aggregate value-free reports (#2387).
Added a dependency-optional Spark redaction transform with immutable, pickle-safe configuration, partition-local workers, deterministic retry behavior, bounded serialization, and stable value-free failures (#2386).
Added a locked Pixi Python 3.12 workflow for Linux x86_64, Intel macOS, and Apple Silicon macOS, with environments mirroring the development, documentation, Hugging Face, service, and MLX extras (#2348).
Added parser-derived Bash, Zsh, and Fish completion scripts and documented the stable machine-readable CLI output workflow (#2347).
Added a sender-authorized Electron de-identification bridge with bounded IPC, a shared serialized utility-process model cache, Node- and Electron-stack offline enforcement, renderer-safe span projection, and timeout-safe worker recovery (#824).
Added a cross-browser Manifest V3 PHI guard that detects and masks text locally, fails closed on unscanned form submissions, persists per-site policy controls without raw text, and verifies zero detection-time network egress with a synthetic unpacked-extension test (#820).
Added a dependency-free local capability probe for injected optional integrations, with deterministic availability counts, provider fingerprints, safe missing-extra classification, and exception-text-free JSON reports (#2585).
Added a deterministic integration capability matrix covering supported adapters, optional requirements, policy boundaries, documentation, and offline test evidence, with local source and dependency validation (#2390).
Added a deterministic offline file-sharding planner that balances declared local file metadata under byte and file-count limits, fingerprints normalized paths, rejects duplicates, and emits counts-only plans without reading files (#2639).
Added crash-safe transactional trace redaction with a value-free recovery journal, fingerprint-verified bounded resume and rollback, transaction-owned staging cleanup, and idempotent completed recovery (#2559).
Added a deterministic cost-versus-cloud benchmark with measured local throughput amortization, cited dated AWS and Azure paid-price tiers, breakeven math, JSON/Markdown CLI output, and fail-closed citation checks (#2342).
Added lazy runtime wiring for validated anonymizer-provider plugins and the
openmed.providers registrar compatibility group, with canonical-label and
locale routing, deterministic Faker access, idempotent discovery, PHI-safe
failure warnings, and built-in-generator fallback (#2341).
Added lazy async wrappers for PII extraction, de-identification, and text analysis, plus ordered batch execution with an optional hard concurrency bound that keeps synchronous work off the event-loop thread (#2338).
Added a Kubernetes HPA reference for aggregate queue-depth and in-flight request metrics, with a concurrent CPU signal, exact load-to-replica guidance, Prometheus Adapter wiring, bounded queue labels, and PHI-safe metric tests (#831).
Added openmed redact-files for local-only text and line-delimited file
redaction with atomic output, PHI-free JSON summaries, consistent surrogate
replacement, and no source overwrite (#2278).
Added reusable iOS Share and Action extension modules for bounded plain-text redaction with bundled policy selection, local-only Nano Core ML assets, fail-closed tokenizer loading, guaranteed runtime-cache cleanup, and host-returnable output that preserves original span offsets (#835).
Added a bounded offline JSON-lines de-identification sidecar with a typed Tauri host and frontend bridge, model pinning, serialized process reuse, renderer-safe errors, strict response validation, and synthetic termination and egress coverage (#823).
Added a local Q4_K_M GGUF grounding runtime with private stdin prompt transport, subprocess-only llama.cpp integration, deterministic top-k recall certification, artifact-bound SHA-256 evidence, and fail-closed loading (#904).
Added a bounded, dependency-free browser network-egress proof harness with exact or path-scoped model-asset allowlists, immediate raw-URL disposal, source-safe digest reports, and fail-closed local trace validation (#2374).
Added a deterministic offline installation smoke check with a clean temporary home, selected-environment entry-point and package-version proof, bundled-manifest validation, repeatable synthetic redaction hashes, and value-free failure reports (#2378).
Added a bounded zero-upload browser privacy playground with deterministic local rules, trusted same-origin adapter support, aggregate-only status, source-safe labels, and explicit network-boundary controls (#2373).
Added a canonical, lockfile-backed uv contributor workflow with an explicitly pinned CI frontend, frozen optional-extra installs, uv-native package builds, and documented pip and Nix fallback paths (#2339).
Added deterministic counts-only comparator reports with fixed metric definitions, bounded aggregate failure accounting, hashed custom identifiers, immutable sanitized state, environment fingerprints, and value-free JSON, Markdown, and write errors (#2380).
Added a standard-library Agent Skills exporter for deterministic ZIP and tar.gz bundles with per-file SHA-256 manifests, source revision provenance, data-driven host and topical-pack selection, portable source-path checks, and rollback-safe overwrite handling (#2307).
Added a deterministic offline Agent Skills validation gate for frontmatter, identifiers, local references, pack membership, and executable-helper help and test contracts, with symlink and local-path containment, path-only diagnostics, scratch-isolated helper probes, and a dedicated CI workflow that runs every focused skill test (#2306).
Added the local-first setup-openmed skill and versioned de-identification
policy template for collecting five bounded privacy decisions, producing a
deterministic atomically written review draft with path-free status output,
and stopping at an explicit human approval gate before the policy can control
a run (#2305).
Added an offline-first self-hosted Compose bundle with loopback-only default publishing, a hardened non-root runtime, persistent cache and read-only model mounts, an internal network, bounded logs and processes, a readiness probe, and opt-in-only remote integrations (#2372).
Added a local-only self-hosted redaction service with explicit text and UTF-8 file workflows, deterministic offline defaults, counts-only review state, loopback Host and request-size guards, content-free errors, and an accessible aggregate-status page (#2371).
Added a deterministic synthetic-only de-identification comparator harness with explicit fail-closed fixture provenance, enforced offline execution, bounded inputs, aggregate privacy metrics, resource budgets, and source-safe reports (#2379).
Added an opt-in bundled-model manifest and offline bootstrap for the small English PII model, with registry checksum and license pins, mandatory cached artifact-integrity proof, concurrency-safe socket guarding, and no silent network fallback (#2375).
Added deterministic offline bootstrap diagnostics for cache readiness, integrity manifests, optional dependencies, and local-only configuration, with stable exit codes and value-free human and JSON reports (#2376).
Added a deterministic standalone local-redactor manifest with a synchronized package/dependency boundary, permissive-license enforcement, explicit opt-in integrations, and excluded restricted dependencies and assets (#2377).
Added metadata-only local agent trace-store discovery with platform-aware defaults, explicit opt-out, no content reads or symlink following, PHI-free store labels, and aggregate counts and byte sizes (#2279).
Added deterministic spawn-backed parallel trace-file sharding with fresh per-file stores, stable input-order merging, safe sequential fallback, and PHI-minimized aggregate failure metadata (#2285).
Added a local registry for training-conversation schemas with collision-safe aliases, recursive format detection, fail-closed validation, and hashed value-free diagnostics (#2286).
Added a role-message training schema adapter with recursive content-path redaction, deterministic structure preservation, hashed path diagnostics, and fail-closed handling for cycles and unknown parts (#2287).
Added a preference-pair training schema adapter with structure-preserving redaction, bounded span reconciliation, validated schema-version reports, and privacy-safe labels and diagnostics (#2288).
Added a local-first, schema-preserving columnar trace-batch adapter with bounded iteration, nested text-path redaction, deterministic defaults, unchanged labels and metadata, and hashed value-free diagnostics (#2289).
Added a streaming, schema-preserving JSONL agent-trace content walker and rewriter with explicit string paths, value-free errors, duplicate-key rejection, same-file overwrite protection, and caller-supplied local transforms (#2280).
Added structure-aware tool-call trace redaction for JSON objects and encoded payloads, with caller-controlled content paths, deterministic serialization, hashed path-only reports, and a local-only default de-identifier (#2281).
Added local credential and secret-token detection for authorization headers, environment values, provider tokens, and private keys, with bounded scanning and value-free, hashed diagnostics (#2283).
Added bounded-memory streaming redaction for structured trace records and NDJSON, with independent record and byte limits, deterministic pseudonyms, aggregate-only progress, and local cancellation (#2284).
Added a deterministic, read-only local trace privacy inventory with counts-only store, category, and file aggregates; byte ranges; file-status totals; hashed caller-supplied labels; and value-free renderers (#2290).
Added local-only transactional in-place trace redaction with sibling temporary files, source-consistency checks, exclusive backups, metadata preservation, atomic replacement, cleanup, and value-free errors (#2291).
Added a deterministic offline trace-fidelity verifier that limits changes to declared content fields; preserves order, linkage, identifiers, timestamps, labels, scalar types, and structure; and emits hashed value-free diagnostics (#2292).
Added versioned topical agent-skill packs for privacy, interoperability, coding, evaluation, and research, with an offline deterministic builder, membership and size-budget validation, canonical relative links, selection-only output, and fail-closed output preflight (#2303).
Added the deterministic ask-openmed workflow router skill, with a
fail-closed privacy override for ambiguous or negated safety statements,
fixed intake-to-verification handoff ordering, canonical links to existing
skills, and PHI-free route diagnostics (#2304).
Added standard-library HTML/HTM visible-text extraction with source character offsets and markup-preserving redaction write-back (#278).
Added an offline, versioned key-lifecycle helper and operator guide for audit-key rotation, retired-key verification, surrogate-vault re-keying, environment isolation, and file-permission hygiene without serializing keys.
Added conservative two- and three-column PDF reading-order reconstruction, preserving source word bboxes and character-span projection while leaving single-column extraction byte-for-byte compatible with the source-order path.
Added deterministic, local redacted-PDF rendering with burned-in opaque rectangles, clean non-PHI text-layer reconstruction, global source-text removal verification, masked page-layout fidelity reports, synthetic fixtures, enforceable regression gates, bounded raster budgets, Type 3 font rejection, and plaintext-free serialized evidence with sanitized render errors.
Added a rooted, backward-compatible public error taxonomy with stable machine-readable codes, actionable PHI-safe diagnostics, REST/MCP mappings, synthetic contract fixtures, and API documentation.
Added a production browser token-classification runtime with typed batched WebGPU inference, deterministic local WASM fallback, an audited WGSL classification head, Python-reference parity and recall gates, per-device warm/cold benchmark records, and real headless-browser coverage.
Added local EML header, plain-text, HTML, and attachment PHI redaction with
decoded source-offset maps, deterministic safety sweeps, image-only PDF
attachment output, and an explicit isolated extract-msg bridge extra for
optional Outlook MSG input.
Added committed Android OpenMedKit release-AAR and offline cold-start budgets, with blocking Gradle/CI gates and measured values in the Android job summary.
Added a Triton ONNX model-repository generator and configuration-selected KServe V2 HTTP/gRPC inference backend with local tokenization and decoding, mocked local/remote span-parity coverage, and no bundled serving runtime.
Added a Kopf-based Kubernetes model operator with the namespaced
OpenMedModel CRD, manifest-pointer warm-pool rollouts, lifecycle conditions
and Events, retained-version rollback, least-privilege RBAC, hardened
deployment assets, operator documentation, and a synthetic fake-API reconcile
suite.
Added a BigQuery-compatible warehouse remote-function handler that validates
batched row envelopes, groups policy-specific calls through process_batch,
emits PHI-safe error replies, and ships synthetic tests, container deployment
guidance, and registration DDL (#839).
Added a deterministic, fully offline openmed init project scaffold with
researcher, app-developer, and data-engineer presets, bundled OpenMedConfig
schema validation, synthetic starter pipelines, and collision-safe reruns.
Added opt-in, no-PHI OpenTelemetry spans and aggregate histograms for all ten
core privacy-pipeline stages, with lazy optional imports, no exporter by
default, shared Timer measurements, synthetic leakage regression tests, and
an otel installation extra.
Added a minimal local-artifact edge-sbc ONNX Runtime profile, native ARM64
Raspberry Pi and Jetson synthetic benchmark workflow, aggregate cold-start,
token-throughput, install-size, and peak-RSS records, plus fail-closed
footprint budgets and archived ARM64 proxy evidence.
The openmed npm package now defaults to the public
OpenMed/OpenMed-PII-ClinicalE5-Small-33M-v1-onnx-android repository
(exported as DEFAULT_MODEL_ID) and routes -onnx-android model ids through
loadOnnxModel() instead of the unavailable former default.
The openmed npm package aligns Transformers.js token-classification output
back to the source text before BIO decoding through alignTokenOffsets();
extractPii() requests ignore_labels: [] to retain the full sequence. The
documented loadOnnxModel() to deidentify() path no longer silently returns
zero spans solely because the runtime omits character offsets.
Token alignment preserves decomposed accents and supplementary Unicode letters. Unalignable tokens fail with a content-free error instead of silently producing incomplete redaction; custom pipelines can supply exact source offsets.
Preserved the v2.2 numeric-offset TypeScript contracts while adding raw-token input types. Model loaders align output and retain runtime metadata and resource disposal; the browser extension remains source-compatible.
Removed obsolete Debian vulnerability exceptions after the current image report confirmed they no longer apply; security thresholds are unchanged.
nltk package, and uv.lock, and
expires on 2026-09-11; it cannot suppress a fixed upstream release.OpenMed 2.3.0 expands the stable v2 SDK across privacy-safe agent and trace
workflows, multimodal asset intake, clinical evidence, local training,
cross-platform runtimes, deployment adapters, and release hardening without
removing a public Python symbol.
Release date: 2026-09-04.
The audited v2.2.0..v2.3.0 release-branch range contains 252 commits and 651
changed files. Commit subjects reference 127 unique issue or pull-request
identifiers across that complete ancestry range.
The static public Python comparison grows from 37,735 to 41,729 symbols: 3,994 additions, zero removals or narrowed signatures, and zero new deprecations. The REST surface remains at 19 paths and 17 component schemas.
Swift, Kotlin/Android, JavaScript, CLI, configuration, trace schemas, evidence records, and deployment contracts expand. npm callers should account for the corrected default model, additive raw-token input types, and explicit alignment failures described in the 2.2-to-2.3 migration guide for the cross-surface review and upgrade checklist.
The immutable tag drives package and image publication. Registry availability can lag the source release while tag-triggered workflows complete, so automation should verify the required coordinate before deployment.
pip install --upgrade "openmed==2.3.0"
pip install --upgrade "openmed[hf,fhir]==2.3.0"
pip install --upgrade "openmed[mlx]==2.3.0"
npm install openmed@2.3.0
The web package remains unscoped and provides ESM and CommonJS exports. Install
@huggingface/transformers to use the default ONNX model, or inject a local
token-classification pipeline. See the migration guide for the offset contract.
dependencies: [
.package(url: "https://github.com/maziyarpanahi/openmed.git", from: "2.3.0"),
]
dependencies {
implementation("com.github.maziyarpanahi:openmed:v2.3.0")
}
JitPack resolves the immutable v2.3.0 repository tag. OpenMedKit performs
local inference and declares no Android INTERNET permission.
After publication, the immutable container coordinate is
ghcr.io/maziyarpanahi/openmed:v2.3.0. The Helm chart's app version and
default image tag are synchronized to 2.3.0; explicit release-image
selection uses image.tag=v2.3.0.
OpenMed keeps local processing as the default. Model downloads and optional remote integrations are explicit boundaries; after required artifacts are present, core PHI processing does not require a cloud service. Telemetry stays off by default, and release evidence uses hashes, counts, offsets, and provenance rather than raw identifiers.
No de-identification or structured privacy system can guarantee zero residual risk. Validate direct-identifier recall, leakage, span integrity, language and format coverage, policy behavior, and quantized-model deltas on the exact deployment path.
Clinical extraction, evidence, filtering, and mapping are assistive software, not a medical device or a source of clinical ground truth. Outputs require qualified review and must not automatically trigger diagnosis, treatment, billing, data release, or other clinical decisions.
The optional agents, llamaindex, quickumls, and scrubadub dependency
trees include NLTK 3.10.3. Its model-artifact path-security advisory
CVE-2026-81726
has no published fixed version as of 2026-09-14. OpenMed does not call the
affected model-file APIs, and the service image does not install NLTK.
Applications using these optional dependencies must not expose NLTK model
import/export paths to untrusted input. This is a known upstream limitation,
not a claim that NLTK is patched. The repository's waiver is restricted to
this CVE, the nltk package, and the uv.lock target, and expires on
2026-09-28. A fixed upstream release cannot be waived by this policy.
The exact release commit is required to pass:
v2.2.0 tree.OpenMed 2.3.0 does not promote a model pointer. Model conversion, evaluation,
and publication remain explicit local maintainer operations. A future model
promotion still requires a real candidate, public SHIELD and golden evidence,
signed extraction gates, and a final decision of exactly READY.
Hosted checks must refer to the exact release commit; successful jobs on an earlier source head do not qualify the release. Package publication, container publication, release assets, and registry verification remain tag-driven follow-up actions.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v2.2.0...v2.3.0
JavaScript: npm audits reported zero vulnerabilities; the web package passed build, typecheck, nine runtime tests, and 15-file dry-run packaging, whil…
This release adds checksum-pinned terminology snapshots, ranked grounding and calibration, FHIR R4 patient summaries and clinical documents, local profile validation, FHIR-to-OMOP CDM 5.4 mapping, deterministic form and key/value extraction, cross-format offset projection, PDF table reconstruction, XLSX/PPTX/ODT intake, HL7 v2 and X12 837 privacy handling, and structured release-risk controls.
OpenMed 2.2 also adds native Maple clinical task support and Compass vision-language inference across Python and OpenMedKit, a four-part synthetic notebook gallery, service grounding and streaming de-identification routes, GraphQL, mTLS, HMAC replay protection, prompt-injection guards, MCP authorization boundaries, Android network-denial guarantees, and PHI-safe diagnostic descriptions.
The audited v2.1.0..v2.2.0 range contains 111 commits and 571 changed files, with 38 PRs associated through GitHub-generated release notes. The public Python inventory grows from 31,619 to 37,735 symbols with 6,116 additions, no removals or narrowed callable signatures, and no new deprecations. The REST contract grows additively from 17 to 19 paths and from 15 to 17 component schemas.
Install or upgrade after the tag-driven package workflows complete:
pip install --upgrade "openmed==2.2.0"
pip install --upgrade "openmed[hf,fhir]==2.2.0"
pip install --upgrade "openmed[mlx]==2.2.0"
npm install openmed@2.2.0Swift Package Manager:
.package(url: "https://github.com/maziyarpanahi/openmed.git", from: "2.2.0")Android through JitPack:
implementation("com.github.maziyarpanahi:openmed:v2.2.0")Container:
docker pull ghcr.io/maziyarpanahi/openmed:v2.2.0Release documentation: https://github.com/maziyarpanahi/openmed/blob/v2.2.0/docs/release/v2.2.0.md
Migration guide: https://github.com/maziyarpanahi/openmed/blob/v2.2.0/docs/migration/2.1-to-2.2.md
Full changelog: v2.1.0...v2.2.0
OpenMed 2.2 introduces a local-first terminology workbench with checksum-pinned snapshots, ranked grounding, calibration, abstention, section context, caller-supplied Athena and crosswalk support, and privacy-safe provenance. Licensed terminology remains caller-supplied and is never bundled.
FHIR support now covers R4 patient summaries and clinical documents, explicit R4/R5 boundaries, local profile validation, Bundle reference-integrity reporting, SDC privacy projection, OperationOutcome helpers, and resumable Bulk Data checkpoints that retain digests rather than source records.
The FHIR-to-OMOP bridge targets OMOP CDM 5.4 and keeps vocabulary resolution caller-controlled. Generated mappings remain assistive evidence requiring downstream clinical and terminology review.
Document processing adds deterministic clinical form and key/value extraction, cross-format offset projection, PDF table reconstruction, spreadsheet, presentation, and OpenDocument intake, HL7 v2 narrative extraction, X12 837 redaction, and fail-closed MIME quarantine.
Structured privacy adds k-anonymity, l-diversity, t-closeness, membership-inference self-tests, aggregate-only differential privacy, and qualified-review evidence for release decisions. These tools report risk and policy evidence; they do not guarantee that a dataset is anonymous or suitable for release.
Optional adapters cover Arrow Flight, SQLAlchemy, PostgreSQL PL/Python, executable UDFs, distributed SQL, Dataflow, Dagster, Ray, pandas-on-Spark, search ingest, and stream processors. External services, credentials, databases, and runtimes remain explicit trust boundaries.
The service gains additive POST /ground and POST /pii/deidentify/stream operations, GraphQL, backpressure, batching, load-test assets, model-cache quotas, CPU INT8 token classification, and additive Go and TypeScript client coverage.
Security work adds mTLS, HMAC replay protection, prompt-injection guards, MCP protected-resource and OAuth-style authorization boundaries, consent receipts, upstream endpoint policy, and aggregate Part 11-oriented audit evidence.
Core PHI processing remains local after explicitly required artifacts are available. Telemetry remains off by default, and audit evidence uses hashes, counts, offsets, thresholds, and provenance rather than raw identifiers or source clinical text.
Python adds the MapleClinicalAssistant task API, MLX Maple export and runtime helpers, and the native openmed.mlx Compass vision-language runtime. OpenMedKit adds Maple request, response, parsing, and MLX runtime types plus native vision-language model loading and generation.
Model weights remain external and must be pinned and checksum-verified. Maple reasoning, entity, relation, and de-identification results and Compass image-grounded output remain assistive and require human review.
Android OpenMedKit retains its public method signatures while hardening local execution with no INTERNET permission, socket-denial tests, opt-in aggregate logging, and PHI-safe entity descriptions. EntityPrediction.description now emits label, Unicode-scalar offsets, confidence, and a SHA-256 digest instead of detected source text; applications that need a local preview must use the explicit text field and keep it out of logs, telemetry, crash reports, and remote diagnostics.
The static Python API comparison against v2.1.0 records 6,116 additions, no removed or renamed public symbols, no narrowed callable signatures, and no newly deprecated symbols.
The v2.1 import openmed.clinical.grounding.SnapshotManifest retains its snapshot-cache meaning. The new vocabulary manifest is exported as VocabularySnapshotManifest, and openmed.clinical.exporters.omop.ConceptResolver remains a public type alias after the OMOP exporter became a package.
Swift adds Maple and Compass APIs without removing an existing OpenMedKit API. Android public method signatures remain compatible; the diagnostic description hardening is the only called-out behavioral migration.
Applications should refresh generated REST clients before adopting the new grounding or streaming routes and re-qualify terminology snapshots, FHIR profiles, OMOP mappings, structured-data policies, model artifacts, and device-specific acceptance tests before enabling new workflows.
OpenMed 2.2.0 is an SDK release and does not promote a model pointer. The committed PII/latest and PII/last_green pointers remain OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1-mlx, and PII/canary remains unset.
Model promotion remains a separate fail-closed workflow requiring real staged golden and public SHIELD evidence, signed gates, and a final readiness decision of exactly READY. Missing model-candidate artifacts are not fabricated for this SDK release.
openmed/py.typed, and contain no bundled model binary or restricted vocabulary payload.OpenMed keeps local processing as the default, but no de-identification or structured privacy system can guarantee zero residual risk. Validate direct-identifier recall, critical leakage, span integrity, language and format coverage, policy behavior, quantized-model deltas, and device behavior against deployment-specific fixtures before production use.
Clinical extraction, grounding, mapping, and model-generated output are assistive software, not a medical device or a source of clinical ground truth. Outputs require qualified review and must not automatically trigger diagnosis, treatment, billing, data publication, or other clinical decisions.
Licensed vocabularies, real-patient records, DUA-gated corpora, UMLS, SNOMED CT, CPT, MIMIC, i2b2, and n2c2 assets are not bundled. Restricted integrations require caller-supplied rights, keys, snapshots, or out-of-process bridges.
GitHub-generated release notes associate 38 PRs with the complete release ancestry: #2228, #2230, #2237, #2239, #2241, #2243, #2244, #2245, #2541, #2543, #2548, #2549, #2550, #2551, #2678, #2679, #2680, #2681, #2682, #2685, #2686, #2687, #2688, #2689, #2690, #2691, #2692, #2693, #2694, #2695, #2696, #2698, #2699, #2700, #2885, #2886, #2887, and #2891.
The generated-note set is intentionally smaller than the complete 111-commit ancestry because maintainer integration batches preserve source contributor commits while presenting one reviewed integration PR per coherent subsystem.
Full Changelog: v2.1.0...v2.2.0
OpenMed 2.2 completes the trustworthy clinical-data-exchange milestone across
terminology grounding, document intake, FHIR, OMOP, structured privacy, MCP,
service security, local model runtimes, and offline release evidence. The final
audited v2.1.0..v2.2.0 range contains 111 commits and 571 changed files.
GitHub generated notes associate 38 PRs with that range, including the
contributor commits preserved by maintainer integration batches.
The static public Python surface grows from 31,619 to 37,735 symbols with
6,116 additions, zero removals or narrowed signatures, and zero new
deprecations. The REST surface grows additively from 17 to 19 paths and from
15 to 17 component schemas through POST /ground and
POST /pii/deidentify/stream. Swift adds public Maple and Compass local-model
runtimes without removing an existing package API. Android keeps its public
method signatures while making diagnostic descriptions and internal logging
PHI-safe by default.
fhir, dagster, and sqlalchemy extras, expanded the
multimodal and service extras, and added the openmed-executable-udf entry
point.EntityPrediction.description now emits label, offsets,
confidence, and a SHA-256 digest instead of raw detected text. Applications
that need a local UI preview must read the explicit text field and must not
send it to diagnostics or telemetry.2.2.0 / v2.2.0.master.openmed.clinical.grounding.SnapshotManifest
binding while exposing the new vocabulary manifest as
VocabularySnapshotManifest, and retained ConceptResolver as a public
type alias after the OMOP exporter became a package. The v2.1-to-v2.2 static
API gate now reports zero breaking symbols.OpenMed 2.2.0 completes the trustworthy clinical-data-exchange milestone
across terminology grounding, document intake, FHIR, OMOP, structured
privacy, MCP, service security, local model runtimes, Android local inference,
and offline release evidence without removing a public Python symbol.
Release date: 2026-08-21.
The audited v2.1.0..v2.2.0 range contains 111 commits and 571 changed files;
GitHub generated notes associate 38 PRs with the complete ancestry range.
The static public Python comparison grows from 31,619 to 37,735 symbols:
6,116 additions, zero removals or narrowed signatures, and zero new
deprecations. The REST surface grows additively from 17 to 19 paths and from
15 to 17 component schemas with POST /ground and
POST /pii/deidentify/stream.
The v2.1 root import openmed.clinical.grounding.SnapshotManifest keeps its
snapshot-cache meaning. The new vocabulary type is available as
VocabularySnapshotManifest. The public OMOP ConceptResolver type alias is
also retained.
Swift adds public Maple task/runtime types and Compass vision-language loading
and generation without removing an existing OpenMedKit API. Android public
method signatures are unchanged, but EntityPrediction.description now
contains a SHA-256 digest instead of raw detected text. Applications that need
a local UI preview must use the explicit text field and keep it out of logs,
telemetry, crash reports, and remote diagnostics.
See the 2.1-to-2.2 migration guide for the cross-surface review and upgrade checklist.
The immutable tag drives package and image publication. Registry availability can lag the source release while tag-triggered workflows complete, so automation should verify the required coordinate before deployment.
pip install --upgrade "openmed==2.2.0"
pip install --upgrade "openmed[hf,fhir]==2.2.0"
pip install --upgrade "openmed[mlx]==2.2.0"
npm install openmed@2.2.0
The web package remains unscoped and provides ESM and CommonJS exports.
dependencies: [
.package(url: "https://github.com/maziyarpanahi/openmed.git", from: "2.2.0"),
]
dependencies {
implementation("com.github.maziyarpanahi:openmed:v2.2.0")
}
JitPack resolves the immutable v2.2.0 repository tag. OpenMedKit performs
local inference and declares no Android INTERNET permission.
After publication, the immutable container coordinate is
ghcr.io/maziyarpanahi/openmed:v2.2.0. The Helm chart's app version and
default image tag are synchronized to 2.2.0; explicit release-image
selection uses image.tag=v2.2.0.
OpenMed keeps local processing as the default. Model downloads and optional remote integrations are explicit boundaries; after required artifacts are present, core PHI processing does not require a cloud service. Telemetry stays off by default, and release evidence uses hashes, counts, offsets, and provenance rather than raw identifiers.
No de-identification or structured privacy system can guarantee zero residual risk. Validate direct-identifier recall, leakage, span integrity, language and format coverage, policy behavior, and quantized-model deltas on the exact deployment path.
Clinical extraction and mapping are assistive software, not a medical device or a source of clinical ground truth. Outputs require qualified review and must not automatically trigger diagnosis, treatment, billing, data release, or other clinical decisions.
The OpenMed v2.2 tested standards matrix is generated from synthetic fixtures and records the exact FHIR, OMOP, and evidence hashes used by the conformance gate. It is implementation evidence, not certification by HL7, OHDSI, a regulator, or a standards body.
Licensed vocabularies, real-patient records, DUA-gated corpora, UMLS, SNOMED CT, CPT, MIMIC, i2b2, and n2c2 assets are not bundled. Restricted integrations require caller-supplied rights, keys, snapshots, or out-of-process bridges.
The exact release commit is required to pass:
v2.1.0 tree; andlatest equal to retained
last_green evidence.OpenMed 2.2.0 does not promote a model pointer. PII/latest and
PII/last_green remain
OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1-mlx, and PII/canary remains
unset. Model promotion remains a separate fail-closed workflow requiring a
real candidate and benchmark evidence. Missing candidate files block that
workflow but do not create a fictitious candidate for this SDK release.
Hosted checks must refer to the exact release commit; successful jobs on an earlier source head do not qualify the release. Package publication, container publication, release assets, and registry verification remain tag-driven follow-up actions.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v2.1.0...v2.2.0
Repository, license, secret, action-reference, dependency-vulnerability, SBOM, provenance, API compatibility, package-content, and deterministic test…
This release expands clinical section and note-type routing, temporal and coreference graphs, calibrated relations, radiology and discharge-summary structures, dosing checks, span-grounded fact recall, offline terminology grounding, OMOP loading, FHIR and OpenEHR export, cohort phenotype resolution, structured privacy, multimodal intake, and typed MCP workflows.
OpenMed 2.1 also strengthens the deployment surface: Android now follows the documented Unicode-scalar offset contract and ships a fail-closed 753-entry on-device catalog derived from the committed 2,266-entry public manifest; the browser and Node.js package remains available as openmed; and the Python, Swift, Helm, service, and container version surfaces are synchronized on 2.1.0.
The public Python inventory grows from 20,538 to 31,619 symbols with 11,081 additions, no removed or narrowed public symbols, and no new deprecations. The REST contract grows additively from 15 paths and 12 component schemas to 17 paths and 15 schemas.
Install or upgrade:
pip install --upgrade "openmed==2.1.0"
npm install openmed@2.1.0Swift Package Manager:
.package(url: "https://github.com/maziyarpanahi/openmed.git", from: "2.1.0")Android through JitPack:
implementation("com.github.maziyarpanahi:openmed:v2.1.0")Container:
docker pull ghcr.io/maziyarpanahi/openmed:v2.1.0Release notes: https://openmed.life/docs/release/v2.1.0/
Migration guide: https://openmed.life/docs/migration/2.0-to-2.1/
Full changelog: v2.0.0...v2.1.0
The current published v2.0.0 tag resolves to rewritten-history commit 94ace7d and is an ancestor of current master through merge boundary b9ab7a3d. Public API compatibility compares the tagged trees directly, while the release integration ledger follows changes after that boundary.
The integration ledger in CHANGELOG.md records 213 PR-associated integrations and seven direct integrations, plus the final release-hardening change set. It covers the Python package, tests, documentation, examples, website, Android, JavaScript, service, MCP, deployment, model manifest, evaluation, security, and release-engineering surfaces.
OpenMed 2.1 expands deterministic clinical document processing across section and document-type routing, coreference, temporal ordering, medication and problem relations, radiology findings, discharge structures, dosing checks, and span-grounded fact recall.
Terminology workflows accept explicit local snapshots and preserve version, code-system, match, abstention, and provenance evidence. OncoTree and other restricted or externally governed terminology inputs remain caller-supplied; the SDK does not silently download or bundle them.
Clinical MCP tools expose typed grounding, FHIR export, risk, signed-audit, model-search, and staged pipeline operations. Existing MCP clients should refresh cached schemas before enabling the new tools.
New local workflows cover OMOP loading, cohort phenotype resolution, FHIR and OpenEHR export, streaming tabular processing, declarative generalization, relational k-anonymity, aggregate-only differential privacy, and attacker-model risk reporting.
Apache Beam, Ray Data, PySpark, Haystack, LangGraph, and sdcMicro integration paths remain optional. The interop-gpl extra installs no GPL code; R and sdcMicro remain separate, out-of-process, user-managed software.
DUA-gated corpora and licensed terminologies are never bundled. Loaders require caller-supplied paths, credentials, or snapshots and preserve explicit local provenance rather than treating data availability as permission to redistribute.
Vietnamese is now a model-backed PII language route, while Urdu and additional Indic codes gain explicit routing and user-supplied-model behavior. Script segmentation is grapheme-aligned and keeps combining sequences, virama conjuncts, zero-width-joiner sequences, and regional-indicator pairs intact.
RTF extraction is stdlib-only and maps extracted characters back to source offsets. DICOM-SR, OCR layout, and other multimodal paths continue to require deployment-specific validation and optional dependencies where documented.
The release advertises 35 built-in PII routes: 33 are model-backed and the Russian and Tamil routes are explicit placeholders. Production Russian or Tamil extraction requires caller-qualified weights and deployment evidence.
The documented Python imports remain available, including:
from openmed import OpenMedConfig, analyze_text, deidentify, extract_piiNew CLI groups cover model-cache management, batch-run resume and reporting, OMOP loading, cohort resolution, OpenEHR export, registry lineage, and release rollback. Existing commands remain available.
REST adds POST /cohort/resolve and POST /omop/load without replacing an existing route. Applications that generate clients from OpenAPI can regenerate to expose the new operations; applications using only existing routes do not need a compatibility shim.
Android OpenMedKit now returns half-open Unicode-scalar offsets consistently from entity predictions, spans, token decoding, and policy de-identification. Applications that pass offsets to Kotlin UTF-16 string APIs for non-BMP text must convert with EntityPrediction.utf16SpanIn.
The Android AAR derives a 753-entry permissively licensed ONNX/TFLite catalog from the committed 2,266-entry public manifest and fails the build if that derivation is empty. Catalog metadata identifies discoverable artifacts; it is not downloaded model data or model-quality evidence.
Swift package sources are unchanged in this release range. The package version coordinates and demo bundle versions are synchronized to 2.1.0, and the existing Unicode-offset parity tests remain green.
The unscoped npm package openmed continues to ship ESM and CommonJS exports for browser and Node.js use. Helm chart metadata, default image selection, and the generated REST OpenAPI version are synchronized to 2.1.0.
The static public Python comparison against the published v2.0.0 tree records 11,081 additions, no removals or renames, no narrowed callable signatures, and no newly deprecated symbols.
Android Unicode offsets are the only called-out behavioral migration. ASCII and Basic Multilingual Plane-only text retains the same numeric offsets; code that applies offsets to strings containing emoji or other non-BMP characters must convert scalar offsets to Kotlin UTF-16 indices first.
Applications using the Android model catalog should re-evaluate pinned or filtered entries against the refreshed catalog. Applications relying on the former dedicated Tamil default must configure and qualify explicit weights.
See the complete migration guide at https://openmed.life/docs/migration/2.0-to-2.1/.
2.1.0, with 17 paths and 15 component schemas.twine check, passed content inspection, and installed with working imports and CLI entry points in a clean Python 3.11 environment.2.1.0.2.1.0, and the v2.1.0 tag was unused locally and on origin at validation time.The final merged release commit must still receive green hosted checks on that exact SHA. In particular, the Linux containerized browser job must compare the approved pixel snapshots, applicable Apple simulator jobs must pass, container build and smoke tests must pass on amd64 and arm64, and the secret-backed signed Android Central Portal bundle must be produced where configured.
Release readiness also requires real staged inputs at artifacts/release-candidate.json, artifacts/release-candidate-shield.json, and artifacts/staged-models.jsonl. Those inputs drive fresh golden and public SHIELD evaluation, signed release gates, evidence binding, and a final readiness decision of exactly READY. They are not fabricated or replaced by the synthetic unit-test results above.
Registry availability, immutable package and image coordinates, live documentation, checksums, attestations, and published release assets are verified only after the tag-driven workflows complete.
OpenMed keeps local processing as the default, but no de-identification system can guarantee zero residual risk. Validate direct-identifier recall, critical leakage, span integrity, language and script coverage, policy behavior, quantized-model deltas, and device behavior against deployment-specific fixtures before production use.
Clinical extraction is assistive software, not a medical device or a source of clinical ground truth. Outputs require qualified review and must not automatically trigger diagnosis, treatment, billing, or other clinical decisions.
Build hashes and package checksums establish artifact identity; they do not prove model quality. Model-specific evidence must be evaluated independently for each selected model, language, quantization, runtime, and deployment.
Thank you to every contributor whose work is included in this release, with a special welcome to first-time contributors @speedyk-005 and @josephkehan-prog.
The complete pull-request and commit inventory, including direct integrations and final release hardening, is recorded in CHANGELOG.md.
Note truncated.
OpenMed 2.1 is the first feature release on the stable v2 line. The audited
source scope covers every current-master change after the v2.0.0 integration
boundary at b9ab7a3d. The current published v2.0.0 tag resolves to the
rewritten-history commit 94ace7d and is an ancestor of master through that
boundary. Public API compatibility compares the tagged trees directly, while
the integration ledger below follows changes after the integration boundary.
The range adds clinical section, note-type, relation, temporal, coreference, radiology, discharge, medication, dosing, and fact-faithfulness workflows; offline terminology grounding, OMOP, FHIR, OpenEHR, cohort, and clinical MCP surfaces; structured generalization, relational privacy, differential-privacy, streaming, and attacker-model risk tools; multilingual, RTF, DICOM-SR, OCR, Android, Flutter, Beam, Ray, Spark, plugin, and model-cache adapters; and expanded deterministic, signed, rollback-safe evaluation and release gates.
The static Python API grows from 20,538 to 31,619 public symbols with 11,081
additions, zero breaking changes, and zero new deprecations. REST grows
additively from 15 paths and 12 component schemas to 17 paths and 15 schemas.
Android's offset implementation now matches the documented Unicode scalar
contract; callers that treated offsets as Kotlin UTF-16 indices for non-BMP
text should follow docs/migration/2.0-to-2.1.md.
Added dependency-free OpenDocument Text (.odt) extraction with paragraph
and list reading order, deterministic table linearization, character-offset
source maps, multimodal registry discovery, and usage documentation (#857).
Added a read-only Strawberry GraphQL endpoint for selective analysis and de-identification fields, canonical entity discovery, policy details, safe aggregate risk facets, introspection, and deterministic SDL export (#828).
Added versioned HMAC-SHA256 request signing over method, path, timestamp, nonce, and body digest, with client-side header helpers, bounded fail-closed replay protection, and verifier-compatible signatures on async job webhooks (#849).
Added a weekday-themed model release orchestrator that chains conversion, synthetic evaluation, signed release gates, model-card generation, publication, fresh-environment smoke checks, last-green rollback, quarantine reporting, and an append-only offline audit ledger (#1243).
Completed longitudinal document linking with exact caller-supplied patient boundaries, conservative cross-document entity de-duplication with complete hashed occurrence provenance, and summary-card/timeline adapters (#1284).
Added offline family-transfer adapter routing that prefers installed target adapters, falls back to compatible donor adapters with scored provenance, and returns explicit unsupported or unavailable routing failures (#1331).
Added stdlib-only RTF text extraction (openmed.multimodal.extract_rtf,
dispatched by redact_document for .rtf) with a character-offset map back
to the source. Destination groups such as \fonttbl, \colortbl, \info,
\pict, and \*-marked extensions are skipped; control words, control
symbols, \'hh codepage escapes (\ansicpg-aware), \uN Unicode escapes
with the group-scoped \ucN fallback count, and \bin payloads are handled
without leaking markup into the extracted text (#856).
Completed clinical temporal timeline composition with DCT/TIMEX anchors on every ordered event, transitively reduced public TLINK graphs, metric-ready edge keys, and retained/pruned privacy-safe decision provenance (#1253).
Added closure-aware temporal TLINK F1, PHI-safe transitive-closure consistency scoring, a zero-violation blocking gate, and synthetic discharge-summary gold with DCT, EVENT-TIMEX, EVENT-EVENT, reduction, and contradiction-trap coverage (#1309).
Added deterministic OncoTree tumor-type mapping
(openmed.clinical.load_oncotree, map_tumor_type) against a
caller-supplied local release snapshot (path / OPENMED_ONCOTREE_PATH and
version / OPENMED_ONCOTREE_VERSION; nothing is bundled or downloaded). The
snapshot must be a flat JSON list of tumor-type nodes; nested OncoTree tree
dumps are unsupported. Exact and normalized name/code lookup supports an
optional caller-supplied synonyms list and indexes history and revocation
aliases with current codes winning collisions; unmatched or ambiguous
mentions stay unmapped with a reason (no fuzzy/lexical fallback).
Results are version-stamped OncoTreeMapping values. Includes synthetic
golden fixtures and oncotree_top1_accuracy evaluation support.
Added an experimental yasbd sentence-segmentation backend selectable via
segment_text(..., backend="yasbd") and
analyze_text(..., sentence_backend="yasbd"), backed by the optional
yasbd-lib extra. The default routing and core dependency set remain
unchanged; opt-in spans are normalized to OpenMed's exact contiguous-offset
contract, with explicit errors for missing dependencies, unknown backends,
and conflicting preconstructed segmenters (#1848).
Added deterministic Urdu-versus-Arabic disambiguation for shared Arabic
script runs. urdu_language_evidence() scores the six Urdu-exclusive letters
(tteh, ddal, rreh, noon ghunna, heh doachashmee, yeh barree) and their sixteen
Arabic presentation forms, derived from single-character NFKC decompositions
so the Koranic stop-sign ligatures U+FDF0/U+FDF1 are excluded. Extended
Arabic-Indic digits reinforce an existing letter signal but never trigger one,
keeping Persian on the Arabic route. Evidence moves ur ahead of ar in the
run's candidate order, and runs report stdlib:urdu-cues when an Urdu pack is
registered or stdlib:arabic-fallback at a lower confidence when none is.
Script-run offsets and grapheme boundaries are unchanged (#1571).
Registered the Indic and Urdu routing candidates (mr, ne, bn, as,
ta, kn, ml, gu, pa, or, ur) across the public language
surfaces. Nepali and Urdu now have display names, model prefixes, and REST,
MCP, TypeScript, and Go language enums; Nepali resolves to Faker's native
ne_NP locale. Languages in USER_SUPPLIED_MODEL_LANGUAGES claim no bundled
default model and raise an actionable ValueError naming every user-supplied
code when model_name is omitted, while SUPPORTED_LANGUAGES stays
model-backed-only so documented model-backed counts are unchanged (#1569).
Promoted Vietnamese (vi) to a model-backed PII language pack routed to
OpenMed/OpenMed-PII-Vietnamese-SuperClinical-Small-44M-v1, taking
SUPPORTED_LANGUAGES to 35 codes. Adds Vietnamese month names, deterministic
locale PHI generation, vi_VN surrogate and CCCD provider coverage across the
REST, MCP, TypeScript, and Go surfaces, and a second synthetic golden i18n
fixture exercising a native ngày D tháng M năm YYYY date, an 0xx mobile,
a 12-digit CCCD, and a diacritic-bearing address (#263).
Added grapheme-aligned mixed-script run routing. segment_by_script now
yields ScriptRun, a tuple-compatible NamedTuple, and every run boundary
falls on an extended grapheme-cluster boundary, so a run can no longer split a
combining sequence, an Indic virama conjunct, a zero-width joiner sequence, or
a regional-indicator pair. Each cluster takes the script of its first
script-bearing code point, keeping a cross-script combining mark attached to
the base character it decorates. LanguageRun gained candidates,
normalizer, tokenizer, and numeral_set, and SCRIPT_NORMALIZERS,
SCRIPT_NUMERAL_SETS, normalizer_for_script, and numeral_set_for_script
expose the per-script routing tables (#1570).
Added decide_rollback() in openmed/eval/rollout.py, the pure decision
function mapping a gate diff to a rollback target. It diffs monitored
per-label recall and residual leakage against the committed last-green
baseline via eval/history.diff_against_baseline, applies the shared
G7_RECALL_DROP_LIMIT tolerance, and returns HOLD / ADVANCE /
ROLLBACK. A regression past tolerance rolls back to the committed
last_green pointer and never advances, even when the candidate's own gate
is RELEASABLE. The decision is side-effect-free and reproducible from the
report plus committed baseline and rollout state with no live API call, and
emits a PHI-free audit record carrying metric names, numeric deltas, store
keys and hashes only (#1803).
Added a read-only catalog coherence gate that checks every models.jsonl
canonical_labels value against openmed.core.labels.CANONICAL_LABELS,
resolving aliases (CHEM/SIMPLE_CHEMICAL -> CHEMICAL) while still
rejecting labels that only survive normalize_label's OTHER fallthrough;
exposed as openmed.core.labels.is_recognized_label,
openmed.core.catalog_coherence.manifest_label_errors, and a Catalog coherence workflow (#2246).
Separated fail-closed model promotion from tag-driven Library/SDK
publication so an SDK tag cannot accidentally attempt a pointer promotion
without a staged challenger, while retaining API compatibility and migration
enforcement in the tag-driven provenance job. Recalibrated the synthetic
Chinese and Indic throughput gate from six GitHub-hosted Ubuntu runs instead
of comparing hosted Linux against an Apple Silicon workstation baseline.
Also fixed Transformers 5 local-snapshot loading so local_files_only is not
forwarded twice to AutoConfig.
Refreshed the canonical public model snapshot from 1,520 to 2,266 entries and restored the Android AAR's generated on-device catalog with 753 permissively licensed ONNX/TFLite entries. Manifest refreshes now disable implicit Hub authentication, preserve audited metadata for retained and converted models, distinguish generative PII models from token-classification evidence, and retain MIT license metadata. Android packaging now fails closed instead of writing an empty catalog.
Replaced the Tamil default's authenticated-only checkpoint with the existing
public multilingual placeholder and classified Tamil alongside Russian as a
non-model-backed compatibility route. The stable
pii_ta_msuperclinical_large registry key now resolves to that placeholder;
production Tamil extraction still requires explicitly qualified weights.
Fixed quadratic script segmentation on text containing long combining-mark
runs whose marks carry a different script from their base. Such input passes
validate_pii_input because the combining and format-sequence guards reset on
each other's characters, and previously cost seconds per document in
segment_by_script, route_runs, and is_indic_text. Cluster starts are now
memoized so segmentation stays linear (#1570).
Fixed Pipeline.stage2_language_script rejecting national-ID-only and
user-supplied language codes that openmed.core.pii already accepted, so an
explicit lang is no longer refused one stage earlier (#1569).
Fixed the shared input gateway rejecting USER_SUPPLIED_MODEL_LANGUAGES
codes. openmed.utils.gateway.validate_language now includes them in its
default acceptance set, so the REST and MCP edges accept every code they
advertise on their language enums instead of returning unsupported_language
for ne and ur. include_national_id still toggles exactly
NATIONAL_ID_ONLY_LANGUAGES (#1569).
Fixed day-first date handling for Vietnamese so shifted, replacement, and
format-preserving date surrogates all render DD/MM/YYYY instead of
MM/DD/YYYY, matching the dmy locale contract already declared for vi
(#263).
Corrected the languages metadata on the 18 OpenMed-PII-Vietnamese-*
manifest rows from ["en"] to ["vi"], so Vietnamese PII checkpoints resolve
through get_pii_models_by_language("vi"). Those 34 registry keys move from
pii_vietnamese_* to pii_vi_* and, as with the Bengali, Chinese, and Tamil
reclassification, they no longer appear in
get_pii_models_by_language("en"), which drops from 219 to 185 entries
(#263).
Fixed the PySpark batch de-identification adapter so
make_deidentify_udf() supplies concrete pandas Series annotations during
UDF construction instead of failing with an unsupported Any signature
(#1942).
Fixed openmed risk discover, risk assess, and risk anonymize handling
of UTF-8 BOM-prefixed CSV and TSV schemas so the first column is classified
consistently, and added bounded validation causes to structured-release CLI
errors instead of replacing actionable TypeError and ValueError details
with a generic schema mismatch.
9b867bcc (nursing-care observation domain),
9b3fa7b4 (TNM extraction), 3c5dad71 (HGVS parsing), e41628df (NER
family label maps), 37d5817f (release run ledger), 544e75bf (private
marketplace owner email), and a6e10b6b (README maintenance).OpenMed 2.1.0 is the first feature release on the stable v2 line. It expands
local-first clinical extraction, structured privacy, terminology grounding,
interoperability, evaluation, and on-device tooling without removing a public
Python symbol.
Release date: 2026-08-12.
The release scope covers every current-master change after the v2.0.0
integration boundary at b9ab7a3d. The current published v2.0.0 tag resolves
to rewritten-history commit 94ace7d and is an ancestor of master through
that boundary. API compatibility compares the published tag tree directly
with the v2.1.0 tree; the integration ledger starts after the boundary.
The static public Python comparison grows from 20,538 to 31,619 symbols:
11,081 additions, zero removals or narrowed signatures, and zero new
deprecations. The REST surface grows additively from 15 to 17 paths and from
12 to 15 component schemas with POST /cohort/resolve and POST /omop/load.
Android now enforces the documented cross-platform Unicode scalar offset
contract. Applications that passed OpenMed offsets directly to Kotlin UTF-16
string APIs for text containing non-BMP characters must use
EntityPrediction.utf16SpanIn. See the
2.0-to-2.1 migration guide for details and the
complete cross-surface review.
The bundled Android model catalog is generated from public manifest metadata
only and contains 753 permissively licensed ONNX/TFLite entries in this
release. Its build now fails if that derivation is empty, preventing a
successful AAR from shipping a catalog that makes the documented
ModelCatalog.entries.first() flow unusable.
The prior Tamil default checkpoint is no longer present in the public Hub catalog. Tamil now uses the same explicit public placeholder convention as Russian, so this release advertises 33 model-backed PII languages across 35 supported routes. The compatibility alias remains available, but production Tamil extraction requires caller-qualified weights and evaluation.
The immutable tag drives package and image publication. Registry availability can lag the source release while those workflows complete, so automation should verify the required coordinate before deployment.
pip install --upgrade "openmed==2.1.0"
pip install --upgrade "openmed[hf,zh,indic]==2.1.0"
npm install openmed@2.1.0
The web package remains unscoped and provides ESM and CommonJS exports.
dependencies: [
.package(url: "https://github.com/maziyarpanahi/openmed.git", from: "2.1.0"),
]
Swift package sources are unchanged in this release range.
dependencies {
implementation("com.github.maziyarpanahi:openmed:v2.1.0")
}
After publication, the immutable container coordinate is
ghcr.io/maziyarpanahi/openmed:v2.1.0. The Helm chart's app version and
default image tag are synchronized to 2.1.0; explicit release-image
selection uses image.tag=v2.1.0.
OpenMed keeps local processing as the default. Model downloads and optional remote integrations are explicit boundaries; after required artifacts are present, core PHI processing does not require a cloud service. Telemetry stays off by default, and release evidence uses hashes, counts, offsets, and provenance rather than raw identifiers.
No de-identification system can guarantee zero residual risk. Validate direct identifier recall, leakage, span integrity, language coverage, policy behavior, and quantized-model deltas on the exact deployment path.
Clinical extraction is assistive software, not a medical device or a source of clinical ground truth. Outputs require qualified review and must not automatically trigger diagnosis, treatment, billing, or other clinical decisions.
The exact release commit is required to pass:
v2.0.0 tree; andlatest model
pointer equal to its retained last_green evidence.OpenMed 2.1.0 does not promote a model pointer. Model promotion remains a
separate fail-closed workflow: any future change to canary, latest, or
last_green requires real staged golden and public SHIELD reports, a signed
model gate, and a final readiness decision of exactly READY. Missing
candidate files block that model dispatch, but do not turn an unchanged model
pointer into a fictitious candidate during an SDK tag.
Hosted checks must refer to the exact release commit; successful jobs on an earlier source head do not qualify the release. Package publication, container publication, release assets, and registry verification remain tag-driven follow-up actions.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v2.0.0...v2.1.0
Your coding agent can read these notes before it upgrades. Set up the MCP server →