NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2987 most downloaded on PyPI
Local-first SDK for clinical extraction and de-identification workflows on hardware you control.
Last release 19 days ago
15 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 41 of 44 stable releases
Nothing withdrawn
no release was ever pulled
1 years old
50 releases · first in 2025
One column per month.
root SECURITY.md and private vulnerability reporting guidance
1.7.0 is the multimodal de-identification and evaluation-depth release.This release turns the v1.6 privacy assurance foundation into a broader local-first clinical data platform: multimodal document intake, OCR adapters, structured clinical extraction, FHIR/HL7/CDA/CSV de-identification, streaming and batch service controls, richer policy profiles, and typed clients for Python, Swift, and TypeScript.
It also deepens the evidence layer around releases: benchmark scorecards, threshold sweeps, leakage heatmaps, membership-inference probes, k-anonymity metrics, utility-loss reports, release-gate previews, evidence bundles, SBOMs, reproducible locks, and stronger supply-chain checks.
Release date: 2026-07-01.
$de-identify, FHIR Bulk NDJSON, deterministic FHIR bundles, OperationOutcome, Provenance, AuditEvent, and CodeableConcept validation/provenance helpers.AnalyzeResult, DeidentificationResult.to_dataframe(), redaction preview diffs, cross-document surrogate vaults, patient-keyed date shifting, format-preserving redaction, minimum-necessary strength selection, custom recognizers, streaming incremental de-identification, explain traces, section stamping, and per-document risk budgets.openmed deid, openmed fhir bundle, openmed models recommend, openmed models diff, openmed policy diff, openmed doctor, openmed gates preview, openmed gates bundle, openmed audit, openmed risk, model-card previews, a typed Python REST client, a TypeScript REST client, and Swift policy-profile support.OpenMed v1.7.0 is about broadening what can be safely de-identified and proving more of the surrounding workflow.
The v1.6 release established policy-aware spans, audit reports, risk scoring, and release gates. v1.7 applies that foundation to more input types and operational paths: scanned documents, PDFs, Markdown and AsciiDoc, structured tables, HL7 v2, CDA/C-CDA XML, FHIR operations, bulk NDJSON, chat logs, Kafka streams, and REST clients.
The evaluation story also moves from single aggregate scores toward release-grade evidence: per-language leakage heatmaps, per-section recall, dataset cards, fixture coverage, model scorecards, threshold sweeps, fairness and robustness reports, utility-loss metrics, and evidence bundles that make gate decisions easier to inspect.
The new multimodal layer gives OpenMed a shared contract for document ingestion and source-preserving redaction.
This release adds:
openmed.multimodal primitives for document content, source spans, handlers, and lazy registrationThe result is a cleaner path from scanned or mixed-format clinical material to audited text, coordinates, redactions, and PHI-safe summaries.
v1.7.0 adds a large set of healthcare interchange helpers.
New or expanded interop paths include:
$de-identify operation wrappersurn:uuid referencesOperationOutcome generationProvenance and AuditEvent emission from signed audit reportsThese additions make OpenMed more useful in pipelines that already speak health-data formats instead of plain text alone.
Clinical extraction is now broader and more context-aware.
This release adds or expands:
Together these changes reduce over-redaction risk and improve downstream structured clinical outputs.
v1.7.0 expands multilingual PII coverage and locale correctness.
New coverage includes:
The runtime and service layers are more production-oriented.
Notable additions:
/livez and /readyz probes with graceful shutdown behaviorThe evaluation system now produces more granular evidence for privacy and model quality.
New capabilities include:
v1.7.0 adds model export and training infrastructure across Apple, browser, PyTorch, and quantized paths.
New work includes:
The release also hardens project operations.
Additions include:
SECURITY.md and private vulnerability reporting guidanceuv.analyze_text(..., output_format="dict") now returns a frozen AnalyzeResult; to_dict() and mapping access preserve the legacy payload shape.OperationOutcome output emits R4 issue.expression; legacy issue.location is accepted on input but is not emitted.ServiceRuntime.get_loader() now returns the warm-pool proxy. Use get_model_loader() when raw loader access is required.languages parameter.OPENMED_SERVICE_TRUSTED_HOSTS; wildcard CORS and trusted-host settings are rejected.--quantized-output requires --quantize int8.format_preserve expands the de-identification action enum and schema surface.shift_dates boolean remains accepted, but new code should prefer method="shift_dates" with explicit date-shift options.Reviewed range: v1.6.0..origin/release/openmed-170
v1.6.0 at 1863350borigin/release/openmed-170 at 36ae8666e7f080b4 (revert grouped Actions update), 38e86172 (pin setup-uv action), cd3f2b9e (run manifest validation through uv), d1061d7b (README update), and 79d56791 (avoid MkDocs autorefs in doctest outputs)Major PR groups:
v1.6.0...v1.7.0 and keep the Fixes #399 issue reference out of the PR count.CHANGELOG.md should be expanded to include the 36 squash/direct PRs as explicit PR-number references, or whether the detailed release-note inventory is the source of truth for that level of traceability.RELEASE_NOTES_v1.6.0.md were used as the style reference.v1.6.0..origin/release/openmed-170; current reviewed head is 36ae8666.v1.7.0 resolve to 184 PR links. The extra #399 in a commit title is an issue reference, not a PR, and is not counted as a PR.1.6 search hits after the version bump are historical changelog entries, compare ranges, CycloneDX/spec/style/dependency values, or changelog-generator examples.python3 scripts/release/check_release_version.py --version 1.7.0 passed.python3 scripts/release/check_repo_policy.py passed.70 passed, 2 warnings for release changelog, FHIR provenance, audit/risk CLI, OpenAPI spec, and service API tests.make docs-build passed under MkDocs strict mode.git diff --check passed.Note truncated.
This release summarizes 148 pull requests merged into
release/openmed-170 after v1.6.0. The diff is additive overall: 483 files
changed, with no deleted or renamed files detected in the release range.
redact_document, image redaction, PDF span coordinate
projection, Markdown/AsciiDoc offset-preserving extraction, audit-safe image,
PDF, and DOCX metadata scrubbing, and JSONL chat-log de-identification with
speaker pseudonymization (#555, #567, #726, #745, #755, #758).$de-identify, FHIR Bulk NDJSON,
deterministic FHIR Bundle, FHIR OperationOutcome, FHIR Provenance /
AuditEvent, deterministic urn:uuid, code-system provenance,
CodeableConcept checks, and flat-table clinical entity export helpers (#566,
#642, #631, #629, #626, #625, #553, #705, #737, #777, #784, #689, #690).DeidentificationResult.to_dataframe,
redaction preview diffs, cross-document surrogate vaults, patient-keyed date
shifting, format-preserving identifier redaction, minimum-necessary strength
selection, streaming incremental de-identification, typed analyze results,
pipeline explain traces, section stamping, and per-document risk budgets
(#706, #695, #729, #704, #778, #779, #731, #611, #727, #785, #733).openmed deid, openmed fhir bundle,
openmed models recommend, openmed models diff, openmed policy diff,
openmed doctor, openmed gates preview, openmed gates bundle,
openmed audit, openmed risk, and active-learning queue management (#741,
#777, #721, #780, #771, #772, #775, #735, #787, #613).openmed_deid pipeline component (#372, #624).analyze_text(..., output_format="dict") now returns a frozen
AnalyzeResult; to_dict() and mapping access preserve the legacy dict shape
(#611).section metadata after section
stamping (#785).unknown rather than deriving a normal/high/low
result (#560)./health remains as a compatibility alias, while /livez and /readyz
expose split liveness/readiness state and shutdown drains in-flight
model-backed requests (#722).uv sync / uv run, with GitHub
Actions refs validated and Dependabot Actions updates limited to minor/patch
bumps (#185, #700).shift_dates documentation now describes patient-keyed stable date shifting;
the legacy boolean remains accepted but deprecated in favor of
method="shift_dates" (#704).python-dateutil and fallback paths,
including month-first English month-name dates, and aligned uv.lock with the
dev extra dependency set (#616, #649).ftfy, section detection, and
date-shift capabilities (#781).subprocess.run calls in reproducibility hash and
release-gate issue helpers (#1090).mask, remove, replace, hash,
shift_dates, and reidentify() examples (#409).SECURITY.md, private vulnerability disclosure guidance, security
issue-template routing, security docs, and README links (#648).make sbom, CI artifact upload, tagged
release SBOM attachment, and supply-chain docs (#720).OperationOutcome output emits R4 issue.expression; legacy
issue.location is accepted on input but is not emitted, and non-R4
severities such as info are rejected (#566).ServiceRuntime.get_loader() returns the warm-pool proxy; use
get_model_loader() when raw loader access is required (#632).--quantized-output requires --quantize int8 (#619).languages parameter
(#717).format_preserve expands the action enum/schema surface and updates schema
fingerprints (#778).OPENMED_SERVICE_TRUSTED_HOSTS; wildcard CORS/trusted-host settings are
rejected (#686).OpenMed 1.6.0 is the privacy assurance and release-governance release.
1.6.0 is the privacy assurance and release-governance release.This release turns de-identification from a single API call into a full, auditable privacy system: policy-aware pipelines, canonical span records, deterministic safety backstops, calibrated thresholds, signed audit reports, re-identification risk scoring, and release gates that fail closed when evidence is missing.
It also adds the infrastructure needed to ship model artifacts responsibly: manifest-driven model cards, Hugging Face publishing controls, dependency and secret-scanning gates, public benchmark reporting, and a richer evaluation suite.
OpenMedSpan records, a ten-stage privacy pipeline, detector arbitration, calibrated thresholds, deterministic safety sweeps, and six bundled policy profiles.models.jsonl manifest, manifest refresh workflow, manifest-driven Hugging Face model cards, and HF publishing for converted MLX/CoreML artifacts.openmed CLI surface with benchmark and calibration commands, plus a de-identification cookbook notebook and clinical NER families example.OpenMed v1.6.0 is about moving from "the model found PII" to "the system can prove what happened."
The new privacy stack records normalized spans, detector provenance, policy actions, thresholds, safety-sweep findings, risk metadata, and stable hashes. That makes de-identification outputs easier to audit, reproduce, benchmark, and review before they reach clinical, operational, or demo workflows.
The same theme shows up in release engineering. Model and package releases now have stronger evidence requirements: calibrated thresholds, last-green baselines, quantization deltas, release gate reports, dependency/license checks, and secret scanning.
deidentify() now routes through a staged privacy pipeline.
The pipeline includes:
Bundled policy profiles:
hipaa_safe_harborhipaa_expert_review_assistgdpr_pseudonymizationresearch_limited_datasetstrict_no_leakclinical_minimal_redactionExample:
from openmed import deidentify
result = deidentify(
"Patient Jane Rivera, MRN 123456, was seen on 2026-06-01.",
method="replace",
policy="hipaa_safe_harbor",
keep_mapping=True,
audit=False,
)
print(result.deidentified_text)Use audit=True when you need the signed audit-report path instead of the regular DeidentificationResult.
OpenMed now includes signed audit reports and re-identification risk scoring.
Audit reports can include:
Risk reporting adds leakage and re-identification views for text and tabular records, including singleton and quasi-identifier analysis. The new adversarial re-identification benchmark mode lets release candidates be tested against round-trip and surrogate-consistency failures.
The new evaluation framework is leakage-first. It adds BenchmarkReport, golden fixtures, SHIELD comparison support, weak labeling, public/gated dataset adapters, cold-start latency, bootstrap confidence intervals, and report rendering.
Release gates now cover G1a-G8 readiness checks with signed reports and fail-closed CI. The release-gate workflow expects evidence for:
This is not only for dashboards. It is meant to prevent a model artifact from being promoted without the evidence needed to defend the release.
v1.6.0 adds several clinical interoperability building blocks:
CardiacFinding, ECGFinding, EjectionFraction, CardiacProcedure, CardiacDevice, and AnatomyThe FHIR assembler produces deterministic transaction/batch Bundles with stable urn:uuid fullUrl values, internal reference rewriting, and request metadata for server ingestion.
No cardiology model is registered in v1.6.0. Cardiology routing is forward metadata and zero-shot label-map support; public model suggestions still fall back to existing general medical recommendations until a cardiology model ships.
The canonical model manifest is now a first-class artifact.
This release adds:
models.jsonlhf-publish environment policyHF_WRITE_TOKEN handling guidancePackaging now includes the manifest, gate baseline, policy/schema JSON, LICENSE, and NOTICE so release artifacts have the metadata they need.
The openmed CLI is now packaged and includes benchmark and calibration commands. The benchmark command supports SHIELD suite runs, multi-model JSON output, and adversarial re-identification mode.
New and refreshed documentation includes:
v1.6.0 expands release hardening:
pip-audit gate with time-boxed ignoresswift-format scripts and CI gateThe old duplicate .github/workflows/release.yml workflow was removed. PyPI publishing now goes through the guarded tag/manual publish.yml workflow.
ar, ja, and tr.trust_remote_code allowlist matching is case-insensitive.keep_year=True targets a non-leap year.OPENMED_SERVICE_MAX_TEXT_LENGTH and return 422 on oversized text.BatchProcessor.iter_process now honors batch_size while preserving order.remove mappings and repeated same-type placeholders now work with keep_mapping=True.deidentify(..., audit=True) returns an audit report. Leave audit=False when callers expect DeidentificationResult.deidentify(..., keep_mapping=True) can now produce unique repeated placeholders such as [NAME_2]. Snapshot tests may need updates. Without keep_mapping, legacy repeated mask placeholders are preserved.method="shift_dates" in new code. shift_dates remains available as a compatibility alias.OPENMED_SERVICE_MAX_TEXT_LENGTH characters now receive a 422 unless the limit is raised..github/workflows/publish.yml for PyPI release operations.HF_WRITE_TOKEN in the protected hf-publish environment.make format, make lint, and make format-check.make format-swift and make lint-swift.Reviewed range: v1.5.5..master
v1.5.5 at adff5feorigin/master at 3b62f2ef654531, .gitignore; 8f398cc, CHANGELOG.md; baad551, CHANGELOG.md; 0769007, docs/swift-openmedkit.md; 0190625, docs/website/index.html; 0b9bdd8, openmed/__about__.py; 3b62f2e, README.md)Major PR groups:
en, fr, de, it, es, nl, hi, te, pt, ar, ja, tr).ResourceType/id values; #559 / #513 shifts default-model lowercase date labels by canonical label.uv.lock through the docs extra, not imported by runtime code.origin/master on 2026-06-22 after PRs #559, #554, #550, and #556 plus the direct release-prep commits through 3b62f2e were merged. The latest release tag is still v1.5.5, reviewed head is 3b62f2e, and the first-parent range contains 76 commits: 69 merged PRs plus direct release-prep commits f654531, 8f398cc, baad551, 0769007, 0190625, 0b9bdd8, and 3b62f2e.gh pr list metadata, and this draft resolved to the same 64 merged PR numbers before PR #551. After PR #551, #559, #554, #550, and #556, live GitHub metadata and this draft resolve to the same 69 merged PR numbers, with no missing or extra PRs..gitignore commit.cardiology zero-shot labels plus private cardiology routing while public model suggestions retain the intended fallback behavior.method="shift_dates" now uses canonical date labels, de-identification audit signing rejects empty HMAC keys, FHIR Bundle assembly rejects duplicate ResourceType/id values, and empty/dangling FHIR Bundle edge cases now have regression coverage.origin/master CI for 3b62f2e passed.python3 scripts/release/check_release_version.py --version 1.6.0 passed after the local version bump.uv run --extra docs mkdocs build --strict passed.uv run --extra dev ruff check . passed.uv run --extra dev pytest tests/unit -q passed: 1,615 passed, 1 skipped.uv run --with build python -m build passed and produced 1.6.0 artifacts.uv run --with twine twine check dist/* passed.git diff --check passed after the changelog, version, and release-note edits.Note truncated.
OpenMedSpan schema contracts, a ten-stage Pipeline, detector arbitration/cascade routing, calibrated per-label/language/policy thresholds, deterministic safety sweep backstops, and six bundled policy profiles (hipaa_safe_harbor, hipaa_expert_review_assist, gdpr_pseudonymization, research_limited_dataset, strict_no_leak, clinical_minimal_redaction).openmed benchmark pii --attack reid.BenchmarkReport, synthetic golden de-identification fixtures, public/reference dataset adapters, DUA-gated corpus stubs, SHIELD comparison-suite support, weak labeling utilities, cold-start latency, and deterministic bootstrap confidence intervals.CardiacFinding, ECGFinding, EjectionFraction, CardiacProcedure, CardiacDevice, Anatomy) plus cardiology keyword routing metadata for future model registration. Public model suggestions continue to fall back to existing general medical models until a cardiology model is registered.models.jsonl manifest, manifest refresh workflow, manifest-driven Hugging Face model card generation, and HF publishing support for converted MLX/CoreML artifacts.openmed CLI surface with benchmark and calibration commands, plus a de-identification cookbook notebook and an offline clinical NER families example.deidentify() now routes through the staged policy pipeline and accepts policy, calibration, threshold, and audit controls. When audit=True, it returns an audit report rather than the regular DeidentificationResult.deidentify(..., keep_mapping=True) now emits unique placeholders for repeated entities of the same type, such as [NAME] and [NAME_2], so re-identification round trips can distinguish them.latency.cold_start_ms in reports.publish.yml workflow; the duplicate release workflow was removed.swift-format scripts, and CI now enforces the updated repo policy, lint, tests, security, secret-scan, Swift-format, and release-gate jobs.LICENSE, and NOTICE.method="shift_dates" to recognize canonical date labels before redaction, so lowercase date output from the default English PII model and date_of_birth labels are shifted instead of masked; keep_mapping no longer treats shifted dates as mask placeholders.ResourceType/id values instead of silently overwriting the earlier resource in the internal reference map. Duplicate resources raise a ValueError that names the colliding key, preventing downstream references from being rewritten to the wrong Bundle entry.ar, ja, and tr for the lang field. These languages have published PII models and are listed in SUPPORTED_LANGUAGES, but the lang Literal in openmed/service/schemas.py was never updated, so the service rejected them with a 422 even though the Python API and the models worked. The four lang annotations now share a single PIILanguage alias kept in sync with SUPPORTED_LANGUAGES (guarded by a regression test).trust_remote_code allowlist matching for first-party and environment-configured privacy-filter repositories.keep_year=True targets a non-leap year.OPENMED_SERVICE_MAX_TEXT_LENGTH (default 1_000_000 characters).BatchProcessor.iter_process so batch_size is honored while preserving output order.remove mappings and repeated entity-type re-identification round trips when keep_mapping=True.hf-publish environment and HF_WRITE_TOKEN policy for model publishing.pip-audit security gate with time-boxed ignores, and gitleaks CI/pre-commit secret scanning with a canary fixture.AuditReport.sign() and AuditReport.verify() require a non-empty HMAC key. None, empty strings, and empty byte strings now raise ValueError instead of producing or accepting weak signatures.shift_dates remains available as a compatibility alias; prefer method="shift_dates" in new code.OPENMED_SERVICE_MAX_TEXT_LENGTH characters now receive a 422 response unless the limit is raised.OpenMed 1.5.5 is a local-first PII, service lifecycle, and README refresh release.
1.5.5 is a local-first PII, service lifecycle, and README refresh release.This release adds batch PII extraction/de-identification, REST model unload and keep-alive controls, stronger Swift/OpenMedKit long-document PII handling, and a polished multilingual README/brand refresh.
BatchProcessor(operation="extract_pii") and BatchProcessor(operation="deidentify") for batch PII workflows.GET /models/loaded, POST /models/unload, request-level keep_alive, and OPENMED_SERVICE_KEEP_ALIVE.1.5.5.Full Changelog: v1.5.2...v1.5.5
Full Changelog: v1.5.2...v1.5.5
BatchProcessor(operation="extract_pii") and BatchProcessor(operation="deidentify"), including document-level batch_size chunking, shared loader/pipeline reuse, tests, docs, and a runnable example.GET /models/loaded, POST /models/unload, request-level keep_alive, OPENMED_SERVICE_KEEP_ALIVE, and model-loader cache release helpers.docs/brand/, and an animated on-device PII de-identification demo (docs/brand/openmed-pii-demo.gif).This release fixes the privacy-filter remote-code trust boundary, tightens model-name routing for the OpenAI/OpenMed Privacy Filter family, and brings
1.5.2 is a security and MLX conversion hardening release.This release fixes the privacy-filter remote-code trust boundary, tightens model-name routing for the OpenAI/OpenMed Privacy Filter family, and brings the public HuggingFace-to-MLX converter in line with the already-published OpenMed MLX Privacy Filter artifacts.
privacy-filter no longer route through the trusted remote-code path.PrivacyFilterTorchPipeline so trust_remote_code defaults to False.OPENMED_TRUSTED_REMOTE_CODE_MODELS for operators who need to trust controlled/private fine-tunes.1.5.2.The Privacy Filter dispatcher previously used broad substring matching for model identifiers. A model name such as attacker/foo-privacy-filter-bar could be treated as part of the privacy-filter family and reach a path where custom repository code may be loaded.
1.5.2 separates routing from trust:
trust_remote_code=True is only enabled for allowlisted first-party repos or operator-controlled local/env-configured models;PrivacyFilterTorchPipeline(..., trust_remote_code=True) calls now fail fast for untrusted model IDs.Trusted first-party remote-code models:
openai/privacy-filterOpenMed/privacy-filter-multilingualOpenMed/privacy-filter-nemotronCustom/private fine-tunes can be allowlisted with:
export OPENMED_TRUSTED_REMOTE_CODE_MODELS="my-org/my-privacy-filter-finetune"
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v1.5.1...v1.5.2
trust_remote_code=True for model identifiers outside an explicit allowlist of first-party OpenAI/OpenMed privacy-filter family models (openai/privacy-filter, OpenMed/privacy-filter-multilingual, OpenMed/privacy-filter-nemotron). Previously, any HuggingFace repository whose name contained the substring privacy-filter would be loaded with custom-code execution enabled, allowing remote code execution by anyone able to control the model_name parameter on /pii/extract or /pii/deidentify. Operators with custom fine-tunes of the privacy-filter family can extend the allowlist via the OPENMED_TRUSTED_REMOTE_CODE_MODELS environment variable (comma-separated repo IDs).PrivacyFilterTorchPipeline's trust_remote_code default from True to False. The first-party dispatcher (openmed.core.backends.create_privacy_filter_pipeline) opts in explicitly only for allowlisted models.1.5.2.openai/privacy-filter, OpenMed/privacy-filter-nemotron, and OpenMed/privacy-filter-multilingual) by casting BF16 tensors to float32 before NumPy conversion, remapping OPF/Nemotron checkpoints into the OpenMed MLX runtime layout, fusing Q/K/V projections, preserving classifier bias, and validating converted weight keys/shapes before artifact save.tests/unit/test_privacy_filter_security.py covering the identifier matcher, allowlist gate, env-var override, local-artifact trust, and dispatcher opt-in.tests/unit/service/test_api.py that POST the attacker-controlled model_name payload to /pii/extract and /pii/deidentify and verify the privacy-filter dispatcher is never reached.Patch release for OpenMed 1.5.1.
Patch release for OpenMed 1.5.1.
Changes:
This release tag is intended to trigger the tag-driven build and publish workflows.
1.5.1.OpenMed 1.5.0 expands multilingual PII extraction and MLX availability with Arabic, Japanese, and Turkish support across the Python SDK and preconvert
OpenMed 1.5.0 expands multilingual PII extraction and MLX availability with Arabic, Japanese, and Turkish support across the Python SDK and preconverted Apple MLX artifacts.
ar, ja, and tr.-mlx repositories.ar_EG, ja_JP, and tr_TR.Atatürk Caddesi 12.model_name or model_id without unintended Hub lookups.The supported Arabic, Japanese, and Turkish PII token-classification models were converted to MLX, uploaded to private OpenMed/<source>-mlx repos, and added to the Medical MLX Models collection.
Deferred models use architectures not covered by the current converter pass: ModernBERT, Qwen3, and Longformer.
analyze_text(..., model_id=...) now works as an alias for model_name, including existing local model directories.ModelLoader now preserves filesystem paths before registry/default-org expansion and keeps local-path Hugging Face loading local-only.extract_pii(..., lang="ar") now defaults to OpenMed/OpenMed-PII-Arabic-SnowflakeMed-Large-568M-v1.extract_pii(..., lang="ja") now defaults to OpenMed/OpenMed-PII-Japanese-BigMed-Large-560M-v1.extract_pii(..., lang="tr") now defaults to OpenMed/OpenMed-PII-Turkish-SuperClinical-Small-44M-v1.OpenMedConfig(backend="mlx") can resolve supported new-language PII checkpoints to uploaded -mlx artifacts directly.1.5.0.uv run pytest tests/unit/test_pii_i18n.py tests/unit/test_model_registry_multilingual.py tests/unit/mlx/test_mlx_inference.py — 211 passeduv run pytest tests/unit/test_pii_i18n.py — 150 passeduv run pytest tests/unit/test_core.py tests/unit/test_utils.py -q — 71 passed; uv run pytest tests/unit -q — 1078 passed, 1 skippeduv run python scripts/release/check_release_version.py --version 1.5.0 — passedgit diff --check — passedFull Changelog: https://github.com/maziyarpanahi/openmed/compare/v1.4.0...v1.5.0
ar), Japanese (ja), and Turkish (tr) PII extraction support in the Python SDK, including language defaults, localized regex patterns, fake replacement data, and anonymizer locale routing.-mlx repositories so OpenMedConfig(backend="mlx") can resolve uploaded artifacts directly.1.5.0.en_US with a warning if a requested locale is unavailable at runtime.Cadde İnönü 12 and common Turkish name-first forms such as Atatürk Caddesi 12.This release brings the OpenMed Multilingual Privacy Filter into the main OpenMed ecosystem across Python, MLX, OpenMedKit, the iOS Scan Demo, and the
This release brings the OpenMed Multilingual Privacy Filter into the main OpenMed ecosystem across Python, MLX, OpenMedKit, the iOS Scan Demo, and the web demo experience. The new family officially supports 16 languages and ships in PyTorch, MLX full-precision, and MLX 8-bit forms.
The headline: developers can now use the same extract_pii() / deidentify() API for the OpenAI baseline, OpenAI Nemotron Privacy Filter, and OpenMed Multilingual Privacy Filter, while Apple demos can showcase all three model choices without changing application code.
OpenMed/privacy-filter-multilingualOpenMed/privacy-filter-multilingual-mlxOpenMed/privacy-filter-multilingual-mlx-8bit1.4.0.OpenMed now documents and routes three Privacy Filter families:
| Variant | PyTorch | MLX full | MLX 8-bit |
|---|---|---|---|
| OpenAI Privacy Filter | openai/privacy-filter |
OpenMed/privacy-filter-mlx |
OpenMed/privacy-filter-mlx-8bit |
| OpenAI Nemotron Privacy Filter | OpenMed/privacy-filter-nemotron |
OpenMed/privacy-filter-nemotron-mlx |
OpenMed/privacy-filter-nemotron-mlx-8bit |
| OpenMed Multilingual Privacy Filter | OpenMed/privacy-filter-multilingual |
OpenMed/privacy-filter-multilingual-mlx |
OpenMed/privacy-filter-multilingual-mlx-8bit |
All three families use the OpenAI Privacy Filter architecture. The multilingual family uses OpenMed multilingual PII training data and officially supports 16 languages.
The public API stays the same:
from openmed import extract_pii, deidentify
text = "Patient Marie Dubois, nee le 14/03/1982, email marie.dubois@example.fr."
entities = extract_pii(
text,
model_name="OpenMed/privacy-filter-multilingual-mlx-8bit",
)
safe = deidentify(
text,
model_name="OpenMed/privacy-filter-multilingual-mlx-8bit",
method="replace",
consistent=True,
seed=42,
)
On Apple Silicon with MLX available, the MLX artifact runs through PrivacyFilterMLXPipeline. On other hosts, OpenMed substitutes the matching PyTorch checkpoint and emits a one-time warning:
OpenMed/privacy-filter-mlx* -> openai/privacy-filterOpenMed/privacy-filter-nemotron-mlx* -> OpenMed/privacy-filter-nemotronOpenMed/privacy-filter-multilingual-mlx* -> OpenMed/privacy-filter-multilingualThe iOS Scan Demo now presents three privacy engines cleanly:
The multilingual path uses OpenMed/privacy-filter-multilingual-mlx-8bit so the demo stays aligned with the 8-bit Apple artifact strategy. The sample controls now use compact EN, FR, and AR buttons, and switching language/sample clears previous annotations before the next run starts.
The multilingual web studio now uses a single top-to-bottom scan pass and redacts line by line during that pass, matching the original Privacy Filter Studio demo feel without looping the scan effect.
1.4.0.1.4.0.OpenMed/privacy-filter-multilingual-mlx and OpenMed/privacy-filter-multilingual-mlx-8bit are first-class model names in the MLX routing table.openmed-mlx.json; stale cached HTTP error bodies are no longer treated as manifests by the scan demo downloader.This release adds targeted unit coverage for multilingual Privacy Filter routing, MLX family alias dispatch, and family-aware fallback behavior. The OpenMed Scan Demo was also rebuilt after the multilingual 8-bit integration.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v1.3.0...v1.4.0
OpenMed/privacy-filter-multilingual — PyTorch / Transformers (CPU + CUDA).OpenMed/privacy-filter-multilingual-mlx — MLX full-precision (Apple Silicon).OpenMed/privacy-filter-multilingual-mlx-8bit — MLX 8-bit quantized (Apple Silicon and OpenMedKit demos).
These artifacts use the OpenAI Privacy Filter architecture and officially support 16 languages through the OpenMed multilingual PII corpus._MLX_MODEL_MAP entries for the full and 8-bit multilingual MLX repo IDs.privacy-filter-multilingual and multilingual-privacy-filter MLX family aliases, both resolving to the existing OpenAI Privacy Filter model class and BIOES decoder.OpenMed/privacy-filter-multilingual on non-MLX hosts instead of the OpenAI baseline.examples/privacy_filter_multilingual_studio/, a web demo comparing the OpenAI baseline, OpenAI Nemotron Privacy Filter, and OpenMed Multilingual Privacy Filter with English, French, and Arabic examples.OpenMed/privacy-filter-multilingual-mlx-8bit, a three-engine picker, EN/FR/AR sample buttons, and new French/Arabic scanned demo documents for screenshot-ready flows.1.4.0.openmed-mlx.json manifests after a public model becomes available.This release turns PII handling into a more complete cross-platform workflow: Faker-backed obfuscation, deterministic surrogates, a canonical PII labe
#OpenMed v1.3.0 is the privacy and anonymization release.
<img width="4112" height="2202" alt="CleanShot 2026-04-29 at 10 03 53@2x" src="https://github.com/user-attachments/assets/8c976be9-702c-4c3c-979e-91e6cbee08da" />
This release turns PII handling into a more complete cross-platform workflow: Faker-backed obfuscation, deterministic surrogates, a canonical PII label taxonomy, unified Privacy Filter routing across MLX and PyTorch, Nemotron-PII Privacy Filter artifacts, and a new interactive Privacy Filter Studio.
The headline: OpenMed can now detect, mask, remove, hash, date-shift, or realistically replace identifiers with locale-aware surrogates, while using the same extract_pii() / deidentify() API on Apple Silicon, Linux, Windows, and service deployments.
method="replace".PrivacyFilterTorchPipeline.pt) support to the REST API schemas.PII de-identification is only useful when the output is both safe and usable.
Simple masking is sometimes enough, but many clinical, operational, and demo workflows need text that still looks realistic: names that look like names, phone numbers that keep their separators, dates that keep their local ordering, and repeated mentions that resolve to the same fake person.
OpenMed v1.3.0 moves beyond static replacement lists. It gives developers a single privacy API that can:
That makes OpenMed more practical for clinical prototypes, privacy demos, evaluation harnesses, and local-first healthcare applications.
method="replace" now uses openmed.core.anonymizer.Anonymizer.
The anonymizer supports:
hashlib.blake2blocale="pt_BR" or locale="en_GB"register_label_generator()register_clinical_provider()Example:
from openmed import deidentify
text = "Patient Pedro Almeida, CPF 123.456.789-09, phone +351 912 345 678."
result = deidentify(
text,
method="replace",
lang="pt",
locale="pt_BR",
consistent=True,
seed=42,
)
print(result.deidentified_text)
Deterministic mode means the same (label, original value) pair maps to the same surrogate within a call. Passing seed= makes the output reproducible across runs.
OpenMed now includes custom Faker providers for clinical and national ID shapes where Faker's built-ins are missing or insufficient:
It also reuses Faker's locale-specific built-ins where they already validate against OpenMed's checksum logic:
pt_BR.cpf and pt_BR.cnpjnl_NL.ssn for BSNfr_FR.ssn for NIRit_IT.ssn for Codice Fiscalees_ES.nieopenmed.core.labels introduces CANONICAL_LABELS and normalize_label().
This gives downstream code one stable label vocabulary even when models emit different naming schemes:
first_nameFIRSTNAMEB-NAME, I-EMAIL, or S-PHONEThe anonymizer, replacement mapping, and Privacy Filter routes now use this normalization layer to reduce model-family-specific branching.
OpenMed v1.3.0 exposes two Privacy Filter checkpoint families through the same public API:
| Variant | PyTorch | MLX full | MLX 8-bit |
|---|---|---|---|
| OpenAI Privacy Filter | openai/privacy-filter |
OpenMed/privacy-filter-mlx |
OpenMed/privacy-filter-mlx-8bit |
| Nemotron-PII fine-tune | OpenMed/privacy-filter-nemotron |
OpenMed/privacy-filter-nemotron-mlx |
OpenMed/privacy-filter-nemotron-mlx-8bit |
Both families use the OpenAI Privacy Filter architecture. The Nemotron-PII artifacts are fine-tuned on the Nemotron PII dataset and reuse the existing Privacy Filter pipeline and model architecture.
Use the same API everywhere:
from openmed import extract_pii, deidentify
text = "Patient Sarah Connor, DOB 03/15/1985, MRN 4471882."
entities = extract_pii(
text,
model_name="OpenMed/privacy-filter-nemotron-mlx-8bit",
)
safe = deidentify(
text,
model_name="OpenMed/privacy-filter-nemotron-mlx-8bit",
method="replace",
consistent=True,
seed=42,
)
On Apple Silicon with MLX available, MLX artifacts run through PrivacyFilterMLXPipeline. On other hosts, MLX-only model names are automatically substituted with the matching PyTorch checkpoint:
OpenMed/privacy-filter-mlx* -> openai/privacy-filterOpenMed/privacy-filter-nemotron-mlx* -> OpenMed/privacy-filter-nemotronA one-time UserWarning explains the substitution.
openmed.torch.PrivacyFilterTorchPipeline loads the Privacy Filter family via Transformers:
trust_remote_code=True by default for the OpenAI Privacy Filter familyInstall:
pip install -U "openmed[hf]"
Run:
from openmed import extract_pii
result = extract_pii(
"Alice Smith emailed alice@example.com.",
model_name="openai/privacy-filter",
)
The Python MLX Privacy Filter runtime now shares decoding utilities with the PyTorch path:
TokenLabelInfobuild_label_infoviterbi_decodelabels_to_token_spanstrim_span_whitespacerefine_privacy_filter_spanThis keeps BIOES/Viterbi decoding consistent across backends.
The MLX model class also now honors classifier_bias / unembedding_bias in artifact configs. This keeps the original OpenAI Privacy Filter bias-less by default while allowing Nemotron-PII artifacts to load their biased classifier head correctly.
OpenMedKit also gained Privacy Filter classifier-head bias support.
The native MLX artifact loader now decodes classifier_bias / unembedding_bias and builds the Privacy Filter head with a learned bias when Nemotron-PII artifacts require it, while preserving the bias-less baseline path.
The OpenMed Scan Demo privacy-filter option now points at OpenMed/privacy-filter-nemotron-mlx-8bit and labels the engine as OpenAI Nemotron Privacy Filter throughout the picker, download events, and README.
This release adds examples/privacy_filter_studio/, an interactive two-pane web demo for PII de-identification.
It includes:
Run:
pip install -U "openmed[mlx]" # or "openmed[hf]" off Apple Silicon
uvicorn examples.privacy_filter_studio.app:app --reload --port 8770
Open:
http://127.0.0.1:8770
Override the model:
OPENMED_STUDIO_MODEL=OpenMed/privacy-filter-nemotron-mlx-8bit \
uvicorn examples.privacy_filter_studio.app:app --port 8770
New and updated docs/examples:
docs/anonymization.mdexamples/obfuscation_demo.pyexamples/privacy_filter_unified.pyexamples/privacy_filter_studio/examples/privacy_filter_book/app.pyThe anonymization guide covers deterministic surrogates, locale resolution, format preservation, custom generators, clinical ID providers, and the Privacy Filter routing model.
faker>=22.0 is now a required core dependency.method="replace" no longer uses the old small static fake-data lists. Downstream tests that asserted exact prior replacement strings should be updated.extract_pii() skips regex smart-merging by design, because the model already performs Viterbi-constrained BIOES span construction.Other de-identification methods are unchanged:
maskremovehashshift_datesInstall or upgrade:
pip install -U openmed
For PyTorch Privacy Filter support:
pip install -U "openmed[hf]"
For Apple Silicon MLX support:
pip install -U "openmed[mlx]"
Recommended checks for application upgrades:
method="replace" outputs, switch to seeded deterministic expectations or assert that originals are removed.classifier_bias or unembedding_bias in the artifact config when the classifier head has bias.Release-prep validation included:
git diff --check
.venv/bin/python -m compileall -q examples/privacy_filter_studio openmed/mlx/models/privacy_filter.py
.venv/bin/python -m pytest tests/unit/mlx/test_privacy_filter_mlx.py tests/unit/test_privacy_filter_routing.py
.venv/bin/python -m pytest tests/unit/core/test_anonymizer.py tests/unit/core/test_labels.py tests/unit/test_pii.py tests/unit/test_privacy_filter_routing.py tests/unit/test_pii_multilingual_regression.py tests/unit/mlx/test_privacy_filter_mlx.py tests/unit/service/test_api.py
Results captured during release prep:
20 passed, 8 skipped471 passed, 1 skipped, 11 warningsThe warnings are pre-existing span-validation warnings from multilingual PII regression fixtures.
OpenMed v1.3.0 is about making privacy work feel less like a demo trick and more like an actual developer surface: local when possible, portable when needed, deterministic when tests demand it, and realistic enough for useful clinical workflows.
Thank you to everyone testing the Privacy Filter artifacts, poking at de-identification edge cases, trying the OpenMedKit paths, and helping OpenMed move toward a more practical open-source healthcare AI stack.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v1.2.0...v1.3.0
openmed.core.anonymizer):
Anonymizer class with cached per-locale Faker instances, deterministic seeding (hashlib.blake2b), and label-keyed generator dispatch.AnonymizerConfig dataclass for advanced configuration.LANG_TO_LOCALE) covering all nine OpenMed languages; Telugu falls back to en_IN with a one-time UserWarning.AadhaarProvider (Verhoeff checksum), GermanSteuerIdProvider, MedicalRecordNumberProvider, NPIProvider. Faker's built-ins are reused for pt_BR.cpf/cnpj, nl_NL.ssn (BSN), fr_FR.ssn (NIR), it_IT.ssn (Codice Fiscale), and es_ES.nie after empirical verification against OpenMed's existing checksum validators.register_clinical_provider() and register_label_generator() for extending coverage.openmed.core.labels):
CANONICAL_LABELS set with 47 canonical labels in UPPER_SNAKE_CASE.normalize_label() maps English lowercase, the 52 Portuguese UPPERCASE labels, BIOES-tagged variants (B-NAME, I-DATE), and arbitrary mixed-case forms to a single canonical form.openmed.core.backends):
select_privacy_filter_backend(), resolve_privacy_filter_model(), and create_privacy_filter_pipeline() route privacy-filter requests to MLX on Apple Silicon and PyTorch elsewhere with a one-time UserWarning when an MLX-only artifact name (OpenMed/privacy-filter-mlx*) is substituted with openai/privacy-filter on non-Mac hosts.extract_pii() and deidentify() now route privacy-filter models through this dispatcher, skipping regex smart-merging since the model already does Viterbi-constrained BIOES decoding.openmed.torch.PrivacyFilterTorchPipeline):
openai/privacy-filter (or any compatible HuggingFace fine-tune) via transformers.AutoModelForTokenClassification with auto device selection (CUDA → CPU).openmed.core.decoding):
TokenLabelInfo, build_label_info, viterbi_decode, labels_to_token_spans, zero_viterbi_biases, VITERBI_BIAS_KEYS extracted from the MLX pipeline so the Torch wrapper reuses the same BIOES Viterbi decoder.trim_span_whitespace, refine_privacy_filter_span for span post-processing across both backends.deidentify() keyword arguments: consistent: bool, seed: Optional[int], locale: Optional[str] for deterministic, locale-overridable obfuscation. Passing seed= alone implies consistent=True.pt) accepted by REST API schemas in openmed/service/schemas.py (was previously library-only despite full core support).examples/obfuscation_demo.py — random vs deterministic surrogates, locale walkthrough, format-preserving phone numbers, pt_BR CPF generation with checksum verification.examples/privacy_filter_unified.py — same extract_pii() / deidentify() call works on Apple Silicon (MLX) and Linux (PyTorch); compares the OpenAI baseline against the Nemotron-PII fine-tune side-by-side.examples/privacy_filter_studio/ — interactive FastAPI + static web studio for two-pane PII masking/randomization with sample clinical notes, highlighted entities, backend/model status, and an explicit first-run download toggle.OpenMed/privacy-filter-nemotron — PyTorch / Transformers (CPU + CUDA).OpenMed/privacy-filter-nemotron-mlx — MLX full-precision (Apple Silicon).OpenMed/privacy-filter-nemotron-mlx-8bit — MLX 8-bit quantized (Apple Silicon).
These checkpoints are the OpenAI Privacy Filter architecture (gpt-oss-style sparse-MoE transformer with local attention, sink tokens, RoPE+YaRN, tiktoken o200k_base) fine-tuned on the Nemotron PII dataset. They reuse OpenAIPrivacyFilterForTokenClassification and PrivacyFilterMLXPipeline unchanged — no new architecture code needed._MLX_MODEL_MAP entries for the two new Nemotron MLX repo IDs in openmed.mlx.inference._SUPPORTED_TOKEN_CLASSIFICATION_MODEL_TYPES (privacy-filter-nemotron, nemotron-privacy-filter) — both resolve to the existing openai-privacy-filter family so a Nemotron-fine-tune MLX artifact can ship with either family identifier in its manifest and still dispatch correctly.openmed.core.backends:
_TORCH_FALLBACK_BY_FAMILY table and _torch_fallback_for() helper.OpenMed/privacy-filter-nemotron instead of the unrelated default openai/privacy-filter, so the user gets the training distribution they asked for. A one-time UserWarning names the substitute._TORCH_FALLBACK_BY_FAMILY.OpenAIPrivacyFilterForTokenClassification now honors classifier_bias / unembedding_bias in artifact configs, while keeping the original OpenAI checkpoint bias-less by default.classifier_bias / unembedding_bias and builds the Privacy Filter head with a learned bias when Nemotron-PII artifacts require it.method="replace" upgraded in place to use the new Faker-backed Anonymizer. Surrogates are now locale-aware (e.g. German names for lang="de", Portuguese phones for lang="pt"), format-preserving, and optionally deterministic. The previous tiny static LANGUAGE_FAKE_DATA lists are kept as a deprecated fallback used only when a Faker locale is unavailable.examples/privacy_filter_book/app.py) migrated to PrivacyFilterTorchPipeline for the CPU side, replacing the inline AutoTokenizer/AutoModelForTokenClassification/pipeline triple.openmed.core.decoding. Behavior unchanged.OpenMed/privacy-filter-nemotron-mlx-8bit and labels the engine as OpenAI Nemotron Privacy Filter throughout the picker, download events, and README.faker>=22.0 is now a required core dependency. Slim installs that skip the ML extras will still pull Faker (~3 MB).method="replace" outputs no longer come from the prior hardcoded list (["Jane Smith", "John Doe", "Alex Johnson", "Sam Taylor"], etc.). Any test or downstream code asserting on those exact strings must either pass consistent=True, seed=<value> and update expected output, or assert non-equality with the original. All other methods (mask, remove, hash, shift_dates) are unchanged.extract_pii() skips regex smart-merging by design. Users who previously chained the low-level MLX pipeline with merge_entities_with_semantic_units() manually may see different entity counts; the new path produces cleaner spans because the model's Viterbi decoder already enforces BIOES validity.tests/unit/core/test_labels.py (102), tests/unit/core/test_anonymizer.py (171, includes per-locale checksum validation across 100s of generated IDs), tests/unit/test_privacy_filter_routing.py (22 — backend selection, family-aware Torch fallback, dispatch, integration), Nemotron parametrisation of the existing privacy-filter MLX dispatch test (tests/unit/mlx/test_privacy_filter_mlx.py::test_dispatches_privacy_filter_pipeline), and Portuguese obfuscation regressions in tests/unit/test_pii_multilingual_regression.py (3).classifier_bias / unembedding_bias config decoding, Nemotron-biased Privacy Filter forward shape, and the baseline bias-less head.No intentional Python API breaking changes from v1.0.0.
This release takes the Apple and MLX foundation from v1.0.0 and turns it into a fuller local inference platform: Python MLX, Swift OpenMedKit, public Privacy Filter artifacts, experimental GLiNER-family extraction, and a redesigned iPhone scan demo that feels like a product instead of a debug console.
The headline: OpenMed can now run more of the clinical document workflow locally, from PII de-identification to clinical entity and relation extraction, across Python, macOS, and iOS.
OpenMed/privacy-filter-mlx and OpenMed/privacy-filter-mlx-8bit, including the smaller 8-bit artifact for Apple apps.Healthcare AI demos often look impressive until the document leaves the device.
OpenMed v1.2.0 moves in the opposite direction. The goal is a practical local-first workflow where a clinical note can be scanned, reviewed, de-identified, analyzed, and summarized without requiring inference through an external service.
That matters for:
Python MLX now understands OpenMed custom task artifacts through openmed-mlx.json manifests.
New or expanded runtime paths include:
token-classificationopenai-privacy-filterzero-shot-nerzero-shot-sequence-classificationzero-shot-relation-extractionThe Privacy Filter runtime includes:
weights.safetensors with weights.npz fallbackInstall:
pip install -U "openmed[mlx]"
Run the public 8-bit Privacy Filter artifact:
from huggingface_hub import snapshot_download
from openmed.mlx.inference import create_mlx_pipeline
model_path = snapshot_download("OpenMed/privacy-filter-mlx-8bit")
pipe = create_mlx_pipeline(model_path)
entities = pipe("Alice Smith emailed alice@example.com and called 415-555-0101.")
print(entities)
Run GLiNER-Relex relation extraction:
from huggingface_hub import snapshot_download
from openmed.mlx.inference import GLiNERRelexMLXPipeline
model_path = snapshot_download("OpenMed/gliner-relex-base-v1.0-mlx")
extractor = GLiNERRelexMLXPipeline(model_path)
result = extractor.inference(
"Aspirin was prescribed for headache after the patient reported migraine symptoms.",
labels=["medication", "condition", "symptom"],
relations=["treats", "associated with"],
threshold=0.5,
relation_threshold=0.9,
)
print(result["entities"])
print(result["relations"])
OpenMedKit now goes well beyond the first token-classification milestone.
New Swift runtime support includes:
OpenMedModelStoreNew public APIs:
OpenMedZeroShotNER
OpenMedZeroShotClassifier
OpenMedRelationExtractor
Swift Package Manager:
dependencies: [
.package(url: "https://github.com/maziyarpanahi/openmed.git", from: "1.2.0"),
]
Run Privacy Filter on device:
import OpenMedKit
let modelURL = try await OpenMedModelStore.downloadMLXModel(
repoID: "OpenMed/privacy-filter-mlx-8bit"
)
let openmed = try OpenMed(backend: .mlx(modelDirectoryURL: modelURL))
let entities = try openmed.extractPII(
"Alice Smith emailed alice@example.com and called 415-555-0101."
)
Run GLiNER-Relex relation extraction:
import OpenMedKit
let modelURL = try await OpenMedModelStore.downloadMLXModel(
repoID: "OpenMed/gliner-relex-base-v1.0-mlx"
)
let extractor = try OpenMedRelationExtractor(modelDirectoryURL: modelURL)
let result = try extractor.extract(
"Aspirin was prescribed for headache after migraine symptoms.",
entityLabels: ["medication", "condition", "symptom"],
relationLabels: ["treats", "associated with"],
threshold: 0.5,
relationThreshold: 0.9
)
The demo apps are now much closer to the release experience we want users to see.
swift/OpenMedDemo now includes the public 8-bit Privacy Filter artifact as a selectable MLX model for macOS and iOS testing.
swift/OpenMedScanDemo has been redesigned around a guided clinical workflow:
The scan demo also includes:
from: "1.2.0" after the release tag is published.Release-prep validation completed on April 24, 2026:
python -m pytest tests/unit/mlx/test_mlx_inference.py tests/unit/mlx/test_privacy_filter_mlx.py tests/unit/test_pii.py tests/unit/test_pii_entity_merger.py
cd swift/OpenMedKit && swift test
xcodebuild -project swift/OpenMedDemo/OpenMedDemo.xcodeproj -scheme OpenMedDemo -destination 'generic/platform=iOS' build
xcodebuild -project swift/OpenMedScanDemo/OpenMedScanDemo.xcodeproj -scheme OpenMedScanDemo -destination 'generic/platform=iOS' build
Results:
143 passed, 1 skipped47 passed, 12 skippedThe Swift skips are the existing SwiftPM CLI guard for tests that require real MLX runtime resources. Physical-device MLX smoke testing remains the acceptance path for those gated artifact tests.
OpenMed v1.2.0 is a release about making healthcare NLP feel more local, more practical, and more usable by real app developers.
Thank you to everyone testing the models, trying the Apple demos, filing sharp feedback, and pushing OpenMed toward a better open-source healthcare AI stack.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v1.1.0...v1.2.0
OpenMed/privacy-filter-mlx and OpenMed/privacy-filter-mlx-8bit artifacts.OpenMedZeroShotNEROpenMedZeroShotClassifierOpenMedRelationExtractortask/family manifests, tokenizer assets, weights.safetensors, and weights.npz fallback paths.OpenMed v1.1.0 adds first-class Portuguese PII detection and de-identification support with lang="pt".
OpenMed v1.1.0 adds first-class Portuguese PII detection and de-identification support with lang="pt".
OpenMed/OpenMed-PII-Portuguese-SnowflakeMed-Large-568M-v115/03/1985, 15-03-1985, 15 de março de 1985uv pip install "openmed[hf]"
from openmed import extract_pii
text = (
"Dr. Pedro Almeida, CPF: 123.456.789-09, "
"email: pedro@hospital.pt, tel: +351 912 345 678"
)
result = extract_pii(text, lang="pt")
print([(entity.label, entity.text) for entity in result.entities])
Expected detection includes name, CPF, email, and phone entities.
Tested with targeted unit, registry, multilingual regression, label-normalization, and smoke-example coverage.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v1.0.0...v1.1.0
lang="pt"
OpenMed/OpenMed-PII-Portuguese-SnowflakeMed-Large-568M-v1This version brings together the core Python runtime, Apple Silicon MLX support, a public Swift package, and a much clearer Apple-platform story.
This version brings together the core Python runtime, Apple Silicon MLX support, a public Swift package, and a much clearer Apple-platform story.
openmed package for Python workflowsOpenMedKit Swift package for macOS and iOSsafetensors-first MLX packaging with weights.npz fallbackOpenMed 1.0.0 includes MLX runtime support across:
OpenMedKit is now part of the public 1.0.0 developer story for Apple apps.
Use it from Xcode as a Swift package to integrate OpenMed into macOS and iOS projects.
The included Xcode demo now defaults to:
OpenMed/OpenMed-PII-LiteClinical-Small-66M-v11.0.0 milestone.Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.6.4...v1.0.0
openmed.mlx.models.bert_tc: Pure MLX BERT implementation with token-classification headopenmed.mlx.inference: MLX NER pipeline producing HuggingFace-compatible output formatopenmed.mlx.convert: CLI tool to convert HuggingFace token-classification models to MLX format with optional 4/8-bit quantizationsimple, first, average, and max aggregation strategiesopenmed.coreml.convert: CLI tool to convert HuggingFace models to CoreML .mlpackage formatct.RangeDim, float16/float32 precisionid2label mapping in model metadata for self-contained deploymentswift/OpenMedKit/)
NERPipeline: CoreML inference with softmax → BIO decoding → entity extractionPostProcessing: BIO tag grouping with first/average/max aggregation strategiesEntityPrediction: Swift equivalent of Python's EntityPrediction dataclassswift-transformers for HuggingFace-compatible tokenizationopenmed.core.backends)
InferenceBackend protocol with is_available() and create_pipeline() interfaceHuggingFaceBackend and MLXBackend implementationsget_backend() auto-detection with explicit override via config.backendpip install openmed[mlx] and pip install openmed[coreml]OpenMed-PII-SuperClinical-Small-44M-v1 as conversion and testing targetbackend field to OpenMedConfig (None/auto, "hf", "mlx")v1.0.0 releasev0.6.4 fixes the issues surfaced by the v0.6.3 quality gates: tokenizer span extension now handles Unicode combining marks and diacritics, pattern val
Release date: 2026-03-24
v0.6.4 fixes the issues surfaced by the v0.6.3 quality gates: tokenizer span extension now handles Unicode combining marks and diacritics, pattern validation failures are properly penalized in confidence scoring, and overly-permissive regex patterns for postal codes, phone numbers, and national IDs have been tightened across French, German, Hindi, and Telugu.
New Verhoeff checksum validator and pattern for Indian Aadhaar numbers:
आधार) and Telugu (ఆధార్) keywords_fix_entity_spans now correctly handles non-ASCII text:
.isalnum() with unicodedata.category check covering letters (L), combining marks (M), and numbers (N).strip() that created text-mismatch false positives in the quality gateWhitespace-only differences between text[start:end] and entity.text (common after span normalization) are now downgraded from WARNING to INFO level. Genuine text mismatches remain WARNING + SpanValidationWarning.
When a pattern's validator fails (e.g., SSN checksum, NIR key), merged confidence now uses a 90/10 model/pattern weight (instead of the normal 60/40). This prevents high-confidence scores on structurally-invalid entities.
| Pattern | Before | After | Impact |
|---|---|---|---|
| French postal code | \d{5} |
01-95 + 971-976 prefixes |
Rejects medical codes, invalid dept prefixes |
| German Steuer-ID | \d{11} |
[1-9]\d{10} |
Rejects leading-zero sequences |
| German postal code | \d{5} |
01xxx-99xxx |
Rejects 00xxx range |
| German phone | \d{2,4}[\s/-]?\d{3,8} |
\d{2,4}[\s/-]?\d{4,8} |
Rejects short (< 4 digit) suffixes |
| Pattern | Old base_score | New base_score | Rationale |
|---|---|---|---|
| French NIR | 0.40 | 0.55 | High structural specificity + validator |
| German Steuer-ID | 0.20 | 0.35 | Tightened pattern + validator |
| French postal code | 0.30 | 0.25 | Still ambiguous even with prefix filter |
| German postal code | 0.30 | 0.25 | Same reasoning |
New normalize_label() mappings:
bsn, dni, nie, aadhaar → national_idmedical_record_number, mrn → medical_recordaccount_number → accountcredit_debit_card, credit_card, debit_card → payment_card| Suite | Tests | Status |
|---|---|---|
| Span-boundary guards | 21 | All pass |
| PII accuracy | 28 | All pass |
| Multilingual regression | 33 | All pass |
| Label-map consistency | 46 | All pass |
| Full suite | 660 | All pass |
tests/unit/test_pii_accuracy.py — Confidence penalty, pattern tightening, and calibration testsopenmed/processing/outputs.py — Unicode-aware _fix_entity_spans with capped extensionopenmed/core/quality_gates.py — Relaxed text-mismatch with whitespace fallbackopenmed/core/pii_entity_merger.py — Validation flag in merging, expanded normalize_labelopenmed/core/pii_i18n.py — Aadhaar validator + patterns, tightened postal/phone/ID patterns, score calibrationopenmed/__about__.py — Version 0.6.3 → 0.6.4docs/website/index.html — softwareVersion → 0.6.4CHANGELOG.md — Added v0.6.4 sectionREADME.md — Updated version referencestests/unit/test_quality_gates.py — Combining-mark and whitespace-mismatch teststests/unit/test_pii_multilingual_regression.py — Aadhaar tests for Hindi/Telugutests/unit/ner/test_label_map_consistency.py — Expanded normalize_label coveragetests/unit/test_pii_entity_merger.py — Updated for 6-element semantic unit tuplesFull Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.6.3...v0.6.4
validate_aadhaar) for 12-digit Aadhaar numberstests/unit/test_pii_accuracy.py)
_fix_entity_spans now Unicode-aware — replaced .isalnum() with unicodedata.category check covering letters, combining marks, and numbers; capped forward extension at 10 characters; removed redundant .strip() that caused text-mismatch false positivesnormalize_label expanded with bsn, dni, nie, aadhaar → national_id; mrn → medical_record; account_number → account; credit_debit_card → payment_card\d{5} to range-constrained 01-95 + DOM-TOM 971-976 prefixes — reduces false positives from medical codes[1-9]\d{10}); base_score raised to 0.3500xxx rangev0.6.4 releasev0.6.3 hardens the PII/NER extraction pipeline with deterministic guardrails: span-boundary validation, multilingual regression coverage, and label-ma
Release date: 2026-03-19
v0.6.3 hardens the PII/NER extraction pipeline with deterministic guardrails: span-boundary validation, multilingual regression coverage, and label-map consistency checks. These quality gates catch tokenizer bugs, model drift, and label inconsistencies before they reach users.
New runtime validation module (openmed.core.quality_gates) that runs automatically after tokenizer repair and smart merging:
validate_entity_spans(entities, text) — checks every entity for:
start < end (no inverted or zero-length spans)start >= 0, end <= len(text))text[start:end] matches stored entity text (catches stale spans after merging)detect_overlapping_entities(entities) — returns overlapping span pairs for informational useSpanValidationWarning and tags entity.metadata["span_valid"], but never silently drops entitiesIntegrated into both OutputFormatter.format_predictions() (after _fix_entity_spans) and extract_pii() (after smart merging).
31 golden-input regression tests covering all 8 supported languages:
| Language | Tests | Entity types validated |
|---|---|---|
| English (en) | 5 | NAME, DATE, PHONE, SSN, merging |
| French (fr) | 4 | NAME, DATE, PHONE, NIR |
| German (de) | 4 | NAME, DATE, PHONE, merging |
| Italian (it) | 3 | NAME, DATE, PHONE |
| Spanish (es) | 4 | NAME, DATE, PHONE, confidence |
| Dutch (nl) | 3 | NAME, DATE, PHONE |
| Hindi (hi) | 4 | NAME, DATE, PHONE, merging |
| Telugu (te) | 4 | NAME, DATE, PHONE, confidence |
All tests use mocked model output for fast, deterministic execution.
32 tests validating configuration invariants:
defaults.json has at least 1 label, no case-insensitive duplicatesgeneric fallback domain always existsnormalize_label() is idempotent for all known label variantsis_more_specific()) agrees with documented hierarchiesentity_types in OPENMED_MODELS PII entries are recognized by normalize_label()19 unit tests covering valid entities, inverted/zero-length spans, negative start, out-of-bounds end, text mismatch, overlap detection (adjacent, nested, multiple), and integration with _fix_entity_spans output.
| Suite | Tests | Status |
|---|---|---|
| Span-boundary guards | 19 | All pass |
| Multilingual PII regression | 31 | All pass |
| Label-map consistency | 32 | All pass |
| Full suite | 600 | All pass |
openmed/core/quality_gates.py — Span validation + overlap detectiontests/unit/test_quality_gates.py — Guard unit teststests/unit/test_pii_multilingual_regression.py — Multilingual regression teststests/unit/ner/test_label_map_consistency.py — Label-map invariant testsopenmed/processing/outputs.py — Integrated span guard after _fix_entity_spansopenmed/core/pii.py — Integrated span guard after smart mergingopenmed/__about__.py — Version 0.6.2 → 0.6.3docs/website/index.html — softwareVersion → 0.6.3CHANGELOG.md — Added v0.6.3 sectionREADME.md — Updated version referencesFull Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.6.2...v0.6.3
openmed.core.quality_gates)
validate_entity_spans() checks start < end, in-bounds, text-match, and zero-length invariants for every entity after tokenizer repair and smart mergingdetect_overlapping_entities() returns pairs of overlapping character spans for informational useSpanValidationWarning emitted on violations — warn-only, never silently drops entitiesOutputFormatter.format_predictions() (after _fix_entity_spans) and extract_pii() (after smart merging)tests/unit/test_pii_multilingual_regression.py)
tests/unit/test_quality_gates.py)
_fix_entity_spanstests/unit/ner/test_label_map_consistency.py)
defaults.json domain invariants (at least 1 label per domain, no case-insensitive duplicates, generic domain exists)normalize_label() idempotency checks across all known label variantsis_more_specific()entity_types in OPENMED_MODELS recognized and idempotent under normalize_label()v0.6.3 releasev0.6.2 expands OpenMed's multilingual PII support and hardens the new REST service for production use.
v0.6.2 expands OpenMed's multilingual PII support and hardens the new REST service for production use.
OpenMed/OpenMed-PII-Dutch-SuperClinical-Large-434M-v1OpenMed/OpenMed-PII-Hindi-SuperClinical-Large-434M-v1OpenMed/OpenMed-PII-Telugu-SuperClinical-Large-434M-v1extract_pii() and deidentify() now support:
lang="nl" for Dutchlang="hi" for Hindilang="te" for TeluguFrench, German, Italian, and Spanish continue to expose the full 35-model multilingual family. Dutch, Hindi, and Telugu currently ship one flagship public checkpoint each.
The FastAPI service introduced in v0.6.1 now includes:
ServiceRuntime and ModelLoader reuse across requestsOPENMED_SERVICE_PRELOAD_MODELSshift_datesv0.6.2examples/pii_multilingual_new_languages.py306 passed in 0.44sFull Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.6.1...v0.6.2
extract_pii() and deidentify() now accept lang="nl", lang="hi", and lang="te"OpenMed/OpenMed-PII-Dutch-SuperClinical-Large-434M-v1OpenMed/OpenMed-PII-Hindi-SuperClinical-Large-434M-v1OpenMed/OpenMed-PII-Telugu-SuperClinical-Large-434M-v1nl, hi, and teexamples/pii_multilingual_new_languages.py for registry, regex, and live-model smoke coverageopenmed.service.runtime.ServiceRuntime for shared per-process config and model-loader reuseOPENMED_SERVICE_PRELOAD_MODELS to warm selected models at startuplang valuesget_pii_models_by_language() now returns sparse public releases for nl, hi, and te while keeping English filtering correctModelLoader.create_pipeline() now caches created pipelines for repeated requests with identical parametersshift_dates alias more strictlyuv pip install "openmed[hf]"dateutilv0.6.2 releaseOpenMed v0.6.1 introduces a Dockerized REST API MVP for deploying text analysis and PII workflows over HTTP.
OpenMed v0.6.1 introduces a Dockerized REST API MVP for deploying text analysis and PII workflows over HTTP.
openmed.serviceGET /health for service status, version, and active profilePOST /analyze mapped to analyze_text(..., output_format="dict")POST /pii/extract for multilingual PII extractionPOST /pii/deidentify for de-identification workflowsDockerfile and .dockerignoreservice dependencies with fastapi and uvicorn[standard]dev dependencies with fastapi and httpx for API testing0.6.1OPENMED_PROFILE.Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.6.0...v0.6.1
openmed.serviceGET /health endpoint for service status and active profile reportingPOST /analyze endpoint mapped to analyze_text(..., output_format="dict")POST /pii/extract endpoint mapped to extract_pii(...)POST /pii/deidentify endpoint mapped to deidentify(...)Dockerfile for service deployment.dockerignore for smaller build contextsservice dependency extra in pyproject.toml (fastapi, uvicorn[standard])dev extra with API test dependencies (fastapi, httpx)docs/rest-service.mdOpenMed v0.6.0 is a breaking, API-first release. This version removes terminal interfaces and streamlines the package around Python APIs for services,
OpenMed v0.6.0 is a breaking, API-first release. This version removes terminal interfaces and streamlines the package around Python APIs for services, notebooks, and production pipelines.
openmed console entrypoint from package metadata.openmed.cliopenmed.zero_shot.cliopenmed.tuicli_main from the top-level public API.cli, tui.analyze_text, BatchProcessor, PII APIs).docs/cli.md, docs/tui.md)..github/workflows/publish.yml now triggers on push tags (v*) and validates tag/version alignment.openmed/__about__.py as the source of truth.
scripts/release/release.py now bumps version only.python3 in Make/release scripts.If you were using CLI commands such as openmed analyze, openmed batch, openmed models, or openmed tui, migrate to Python APIs.
from openmed import analyze_text
result = analyze_text(
"Patient diagnosed with chronic myeloid leukemia and started imatinib.",
model_name="disease_detection_superclinical",
)
print(result.entities)
from openmed import BatchProcessor
processor = BatchProcessor(model_name="disease_detection_superclinical")
output = processor.process_texts([
"Text one",
"Text two",
])
print(output.summary())
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.5.8...v0.6.0
openmed console entrypoint from package metadataopenmed.cli and openmed.tui modulesopenmed.zero_shot.clicli_main from the top-level openmed public APIcli, tui)publish.yml)openmed/__about__.py as the version source of truthFeature/v0.5.8 fixes pii extraction by @maziyarpanahi in https://github.com/maziyarpanahi/openmed/pull/25
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.5.7...v0.5.8
first_name, last_name, dob, postal_code, etc.)FIRST_NAME, LAST_NAME, and ZIPCODE values across supported languagesextract_pii() and deidentify() now strip leading/trailing whitespace before inference so spans remain aligned with analyze_text() validation behaviorFix: Added _fix_entity_spans() in OutputFormatter that extends end forward while the next character is still alphanumeric. This runs before entity gro
Fix: Added _fix_entity_spans() in OutputFormatter that extends end forward while the next character is still alphanumeric. This runs before entity grouping, so all downstream consumers (extraction, de-identification, smart merging) get correct spans.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.5.6...v0.5.7
_fix_entity_spans() to correct tokenizer end-offset truncation and trim whitespace around predicted spansThis release adds full Spanish PII support and expands multilingual model coverage.
This release adds full Spanish PII support and expands multilingual model coverage.
lang="es" in:
extract_pii()deidentify()validate_spanish_dni()validate_spanish_nie()DD/MM/YYYY and 15 de enero de 2020 format)+34, mobile/landline formats)pii_biomed_bert_full (BiomedBERTFull-Base-110M)pii_lite_clinical_u (LiteClinicalU-Small-66M)Spanish- prefix).get_pii_models_by_language("es") now returns full Spanish model set.get_default_pii_model("es") now returns the Spanish default model.normalize_accents parameter in extract_pii() and deidentify()replace workflows:
0.5.6.Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.5.5...v0.5.6
Spanish PII Detection & De-identification: Full Spanish language support for PII extraction
extract_pii() and deidentify() now accept lang="es" for Spanish clinical textlang="es"Spanish National ID Validators: DNI and NIE document validation with checksum verification
validate_spanish_dni() — Spanish DNI 8-digit + check letter (mod-23 lookup table)validate_spanish_nie() — Spanish NIE with X/Y/Z prefix conversion and DNI algorithm2 New English Base Model Architectures: Expanded PII model coverage
pii_biomed_bert_full — BiomedBERTFull-Base-110M for comprehensive biomedical PII detectionpii_lite_clinical_u — LiteClinicalU-Small-66M for universal lightweight PII detectionExpanded Model Registry: 35 Spanish PII models + 8 new models across existing languages
get_pii_models_by_language("es") returns all 35 Spanish modelsget_default_pii_model("es") returns the recommended Spanish default modelAccent Normalization: Transparent accent stripping for models trained on accent-free text
normalize_accents parameter on extract_pii() and deidentify() (auto-enabled for Spanish)_strip_accents() helper preserves character count via NFC/NFD normalizationnormalize_accents=True) or disabled (normalize_accents=False) for any languageSpanish Locale Data: Culturally appropriate synthetic data for the replace method
Testing: Comprehensive Spanish PII test coverage
"es" to "ja" in unsupported language assertions_LANGUAGE_CONFIG in model registry now includes "es": {"name": "Spanish", "prefix": "Spanish-"}SUPPORTED_LANGUAGES expanded to include "es"_shift_date, _shift_date_basic, _format_date_like_original) now support SpanishOpenMed — v0.5.5: Multilingual PII Detection & De-identification
OpenMed — v0.5.5: Multilingual PII Detection & De-identification
OpenMed now supports multilingual PII detection and de-identification for clinical text in English, French, German, and Italian. The model catalog has expanded from 33 to 132+ PII models, and all core APIs remain fully backward compatible.
Multilingual PII Extraction & De-identification
Fully backward compatible. All existing English-only usage works unchanged without any code modifications.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.5.1...v0.5.5
Multilingual PII Detection & De-identification: Language-aware PII extraction for clinical text
extract_pii() and deidentify() now accept a lang parameter (ISO 639-1: en, fr, de, it)lang is specifiedNational ID Validators: Country-specific document validation with checksum verification
validate_french_nir() — French NIR/INSEE 15-digit social security numbers (mod-97 checksum)validate_german_steuer_id() — German 11-digit tax identification numbers (digit-frequency rules)validate_italian_codice_fiscale() — Italian 16-character alphanumeric fiscal codesLocale-Aware Date Handling: Language-appropriate date parsing and formatting
fr/de/it (DD/MM/YYYY, DD.MM.YYYY)en (MM/DD/YYYY)Culturally Appropriate De-identification: Language-specific synthetic data for the replace method
LANGUAGE_FAKE_DATA dictionary for English, French, German, and ItalianExpanded Model Registry: Multilingual model generation across all PII architectures
get_pii_models_by_language() — returns all PII models for a given languageget_default_pii_model() — returns the recommended default model for a languageNew Module: openmed/core/pii_i18n.py — Internationalization module
SUPPORTED_LANGUAGES, DEFAULT_PII_MODELS, LANGUAGE_PII_PATTERNS constantsget_patterns_for_language() — returns combined English + language-specific regex patternsLANGUAGE_MONTH_NAMES dictionary with month names in all 4 languagesDocumentation
Testing
test_pii_i18n.py — unit tests for the i18n module (373 lines)test_model_registry_multilingual.py — unit tests for multilingual model generation (202 lines)test_pii.py and test_pii_entity_merger.py with multilingual test cases_redact_entity() and _generate_fake_pii() now propagate lang parameter for language-appropriate replacementsnormalize_label() handles national ID variants (nir, insee, steuer_id, codice_fiscale) and postal code variants (postcode, zipcode, postal_code)national_id sub-types for cross-language entity resolutionCATEGORIES["Privacy"] dynamically includes all PII model keys (English + multilingual)__init__.py exports with multilingual PII support functionsContext-Aware PII Scoring: Presidio-inspired confidence scoring system
Context-Aware PII Scoring: Presidio-inspired confidence scoring system
PIIPattern dataclass extended with base_score, context_words, context_boost, and validator fieldsfind_context_words() - boosts confidence when keywords like "SSN:", "DOB:", "NPI:" appear near detected entitiesvalidate_ssn(), validate_luhn() (credit cards), validate_npi(), validate_phone_us()Website Updates
OpenMed-PII-SuperClinical-Small-44M-v1merge_entities_with_semantic_units() now supports context-aware pattern scoringmedical-tokenizer.md and pii-smart-merging.md to nav structurecli.md to PII notebook (now links to GitHub)pii-smart-merging.md to non-existent documentation pagesentity.text, entity.label, entity.confidence)Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.5.0...v0.5.1
PII Detection & De-identification: HIPAA-compliant PII extraction and de-identification
PII Detection & De-identification: HIPAA-compliant PII extraction and de-identification
extract_pii() function for detecting PII entities in clinical textdeidentify() function with 5 de-identification methods:
mask: Replace with placeholders ([NAME], [DATE], etc.)remove: Complete removal of PII entitiesreplace: Replace with synthetic datahash: Cryptographic hashing for record linkingshift_dates: Shift dates while preserving temporal relationshipsreidentify() function for reversing de-identification with stored mappingsPIIEntity and DeidentificationResult dataclassesSmart Entity Merging: Advanced post-processing to fix tokenization fragmentation
date_of_birth > date)PIIPattern classuse_smart_merging=True parametermerge_entities_with_semantic_units(), find_semantic_units(), calculate_dominant_label(), PII_PATTERNSPII CLI Commands: Comprehensive command-line interface for PII operations
openmed pii extract: Extract PII entities from text or filesopenmed pii deidentify: De-identify text or files with method selectionopenmed pii batch-extract: Batch PII extraction from directoriesopenmed pii batch-deidentify: Batch de-identification with method selection--date-shift-days) for temporal preservationPII TUI Mode: Interactive PII detection in the terminal interface
PII Model Registry: Added PII detection models
pii_detection_superclinical (434M parameters)Comprehensive Documentation
Testing
use_smart_merging=True)Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.4.0...v0.5.0
Interactive TUI (Terminal User Interface): Full-featured terminal workbench for clinical NER analysis
Interactive TUI (Terminal User Interface): Full-featured terminal workbench for clinical NER analysis
openmed or openmed tuiTUI Documentation: Comprehensive guide at docs/tui.md
Website Updates
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.3.0...v0.4.0
Batch Processing: Process multiple texts or files in a single operation
Batch Processing: Process multiple texts or files in a single operation
BatchProcessor class for configurable batch operationsBatchItem, BatchItemResult, BatchResult dataclassesprocess_batch() convenience functionbatch command with full feature supportConfiguration Profiles: Named configuration presets for different environments
dev, prod, test, fastOpenMedConfig.from_profile() and with_profile() methodslist_profiles(), get_profile(), save_profile(), delete_profile() functionsconfig profiles, profile-show, profile-use, profile-save, profile-delete--profile flag for config show commandPerformance Profiling: Built-in timing and metrics utilities
Timer context manager for measuring code blocksProfiler class for tracking metrics across multiple runs@profile decorator for easy function profilingProfilingMetrics dataclass for structured timing dataDocumentation
Testing
Medical-aware tokenization is now remap-only: the model always uses its own tokenizer/vocab; we tokenize separately for output and remap spans to clea
--use-medical-tokenizer and --medical-tokenizer-exceptions flagsMedical-aware pre-tokenizer now ships in-core and is on by default for fast HF tokenizers; it uses Bert-style splitting plus a curated list of biomedi
tokenizers isn’t installed or the tokenizer is slow.DEFAULT_MEDICAL_EXCEPTIONS = [
"COVID-19",
"SARS-CoV-2",
"IL-6",
"IL-2",
"TNF-alpha",
"BCR-ABL1",
"CAR-T",
"post-CAR-T",
"t(8;21)",
"t(15;17)",
]
def build_medical_pretokenizer(exceptions: Optional[Iterable[str]] = None):
...
def apply_medical_pretokenizer(tokenizer: Any, exceptions: Optional[Iterable[str]] = None) -> bool:
...
OpenMedConfig gains use_medical_tokenizer (default True) and medical_tokenizer_exceptions; CLI adds --use-medical-tokenizer/--no-medical-tokenizer and --medical-tokenizer-exceptions, with env overrides (OPENMED_USE_MEDICAL_TOKENIZER, OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS) honored.# Medical-aware pre-tokenizer toggle (fast tokenizers only)
use_medical_tokenizer: bool = True
# Optional list of hyphenated/clinical exceptions to keep intact
medical_tokenizer_exceptions: Optional[List[str]] = None
analyze_parser.add_argument(
"--use-medical-tokenizer",
dest="use_medical_tokenizer",
action="store_true",
default=None,
help="Force-enable the medical-aware pre-tokenizer (default from config).",
)
analyze_parser.add_argument(
"--no-medical-tokenizer",
dest="use_medical_tokenizer",
action="store_false",
default=None,
help="Disable the medical-aware pre-tokenizer and fall back to the model default.",
)
analyze_parser.add_argument(
"--medical-tokenizer-exceptions",
default=None,
help="Comma-separated extra terms to keep intact (e.g., MY-DRUG-123,ABC-001).",
)
if getattr(self.config, "use_medical_tokenizer", False):
...
applied = apply_medical_pretokenizer(
tokenizer,
exceptions=getattr(self.config, "medical_tokenizer_exceptions", None)
)
...
if getattr(self.config, "use_medical_tokenizer", False):
...
applied = apply_medical_pretokenizer(
ner_pipeline.tokenizer,
exceptions=getattr(self.config, "medical_tokenizer_exceptions", None),
)
infer_tokenizer_max_length and helper alignment/truncation utilities; CLI/registry introspection can surface sensible max lengths.docs/medical-tokenizer.md, README/CLI docs updates, runnable examples in examples/custom_tokenizer/ plus benchmark/demo notebooks.tests/test_medical_tokenizer.py, tests/test_cli_medical_tokenizer.py).tokenizers>=0.15 added to the hf extra in pyproject.toml; enable the medical tokenizer by installing with [hf].Notable behavior change: fast tokenizers now default to the medical pre-tokenizer; disable via config/CLI/env if you need raw tokenizer parity with prior runs or baseline comparisons.
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.2.0...v0.2.1
Sentence detection and segmentation
New features
Improvements
Documentation & packaging
Tests & fixtures
Upgrade / compatibility notes
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.1.10...v0.2.0
openmed command)analyze command for single text analysismodels list and models info commandsconfig show and config set commandsNew / changed functionality (high-level)
New / changed functionality (high-level)
Quality, docs & tests
Maintenance & developer experience
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.2.0-rc1...v0.2.0-rc2
Sentence detection and segmentation
New features
Improvements
Documentation & packaging
Tests & fixtures
Upgrade / compatibility notes
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.1.10...v0.2.0-rc1
Introduced the openmed command-line interface with analyze, models, and config subcommands, making it easy to run inference, discover models, and mana
openmed command-line interface with analyze, models, and config
subcommands, making it easy to run inference, discover models, and manage
settings from the terminal (includes new console script wiring and entry
points).models list now supports --include-remote Hugging
Face discovery, models info surfaces inferred max sequence length, and
configuration helpers honor XDG paths and environment overrides.ModelLoader.get_max_sequence_length, the top-level get_model_max_length,
and metadata propagation so analyze_text reports the effective context window
while keeping tokenizer settings in sync.truncation kwarg,
ensuring the CLI and library work with the updated transformers>=4.50 API.transformers>=4.50 / huggingface-hub>=0.30
and documented the Python ≥3.10 requirement for better upstream compatibility.scripts/reset_uv_env_and_run_tests.sh to automate nuking the environment,
recreating a UV-managed Python 3.11 venv, reinstalling OpenMed from source, and
running the full (slow-inclusive) pytest suite.Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.1.9...v0.1.10
analyze_text() one-call inference APIOpenMedConfigNothing published for this version
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.1.8...v0.1.9
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.1.8...v0.1.9
Nothing published for this version
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.1.7...v0.1.8
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.1.7...v0.1.8
Nothing published for this version
Nothing published for this version
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.1.6...v0.1.7
Full Changelog: https://github.com/maziyarpanahi/openmed/compare/v0.1.6...v0.1.7
Full Changelog: https://github.com/maziyarpanahi/openmed/commits/0.1.5
Full Changelog: https://github.com/maziyarpanahi/openmed/commits/0.1.5
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →