NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1831 most downloaded on PyPI
High-performance HTML to Markdown converter
Last release today
04 Oct 2026
Ships fairly regularly
a new release about every 9 days
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
198 releases · first in 2024
One column per month.
Rust core: clamp table colspan/rowspan to prevent pathological allocations on malformed HTML.
colspan/rowspan to prevent pathological allocations on malformed HTML.convert.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.10...v2.14.11
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.10...v2.14.11
ConvertWithMetadata() deserialization for metadata enums (link_type, image_type, data_type, text_direction) by honoring the JSON wire values.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.9...v2.14.10
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.9...v2.14.10
ThreadPoolExecutor parallelism doesn't regress performance, and always build the extension with metadata support (so convert_with_metadata is always available).chore(deps): bump actions/download-artifact from 5 to 7 by @dependabot[bot] in https://github.com/Goldziher/html-to-markdown/pull/155
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.7...v2.14.9
<script type="application/ld+json"> tags (including when placed in <head>), preserving the script contents for parsing.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.7...v2.14.8
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.7...v2.14.8
html-to-markdown-rs): enable the metadata feature by default so convert_with_metadata is available without extra Cargo features.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.6...v2.14.7
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.6...v2.14.7
.cargo/config.toml so Rustler can compile without requiring user-specific linker flags.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.5...v2.14.6
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.5...v2.14.6
ruby-platform gems when multiple CI jobs produce identical artifacts for the same version.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.4...v2.14.5
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.4...v2.14.5
rubygems-* artifacts into separate directories (no merge), and publishing gems recursively with an integrity check.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.3...v2.14.4
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.3...v2.14.4
osx-x64 native FFI library on macos-15-intel (macOS-13 runners are retired), unblocking NuGet publication.mix deps.get && mix test works outside this monorepo.chore(deps): bump actions/download-artifact from 6 to 7 by @dependabot[bot] in https://github.com/Goldziher/html-to-markdown/pull/151
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.2...v2.14.3
convert_with_metadata (no more ImportError on import).wrap=true.external, relative) to match cross-language expectations.metadata feature enabled (avoids missing html_to_markdown_convert_with_metadata when the workspace was built with --no-default-features).convert_with_metadata/3 + MetadataConfig backed by the Rust metadata extractor.convertWithMetadata.html_to_markdown_ffi libraries into the NuGet package under runtimes/*/native..../packages/go/v2), and docs/examples were updated accordingly..sdkmanrc for Java 25 + Maven 4; keep maven-source-plugin on 3.3.1 because 4.0.0-beta-1 is not compatible with Maven 4.0.0-rc-4.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.1...v2.14.2
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.14.1...v2.14.2
scripts/common/install-maven-latest.sh and applied repo-wide lint/format cleanups.Issue #147: Word wrap now works correctly in list items when using the -w/--wrap flag. List items with long text are properly wrapped while preserving
-w/--wrap flag. List items with long text are properly wrapped while preserving list structure and indentation for both ordered and unordered lists.strip_tags and preserve_tags options now correctly prevent <meta> and <title> tags from being extracted into YAML frontmatter when extract_metadata is enabled.strip_newlines=true no longer causes excessive whitespace around block elements. Structural whitespace is now properly normalized while still removing newlines within paragraph content.This release completes the metadata extraction feature across all language bindings with comprehensive documentation, critical bug fixes, and 100% lan
This release completes the metadata extraction feature across all language bindings with comprehensive documentation, critical bug fixes, and 100% language compliance.
--with-metadata flag with JSON output support--extract-document, --extract-headers, --extract-links, --extract-images, --extract-structured-data{"markdown": "...", "metadata": {...}}ConvertWithMetadata() function with typed structsconvertWithMetadata() method with Java recordsConvertWithMetadata() method with C# recordshtml_to_markdown_convert_with_metadata() C functionMETADATA.md files (TypeScript: 480 lines, Ruby: 228 lines)max_structured_data_size default (100KB → 1MB)--extract-document flag not being mapped to MetadataConfigmax_structured_data_size default (was 100KB, corrected to 1MB)DEFAULT_MAX_STRUCTURED_DATA_SIZE constant from Rust coreDEFAULT_MAX_STRUCTURED_DATA_SIZE: usize = 1_000_000 constant in Rust coreThis release includes complete metadata extraction support for:
All metadata extraction features are fully backward compatible. To use the new features:
# Extract all metadata as JSON
html-to-markdown input.html --with-metadata \
--extract-document --extract-headers --extract-links \
--extract-images --extract-structured-data
import "github.com/Goldziher/html-to-markdown/packages/go/htmltomarkdown"
result, err := htmltomarkdown.ConvertWithMetadata(html)
if err != nil {
log.Fatal(err)
}
fmt.Printf("Markdown: %s\n", result.Markdown)
fmt.Printf("Title: %s\n", result.Metadata.Document.Title)
import io.github.goldziher.htmltomarkdown.*;
MetadataExtraction result = HtmlToMarkdown.convertWithMetadata(html);
System.out.println("Markdown: " + result.markdown());
System.out.println("Title: " + result.metadata().document().title());
using HtmlToMarkdown;
var result = HtmlToMarkdownConverter.ConvertWithMetadata(html);
Console.WriteLine($"Markdown: {result.Markdown}");
Console.WriteLine($"Title: {result.Metadata.Document.Title}");
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.13.0...v2.14.0
--with-metadata flag with JSON output support for extracting document metadata, headers, links, images, and structured data from HTML documents.
--extract-document, --extract-headers, --extract-links, --extract-images, --extract-structured-data{"markdown": "...", "metadata": {...}}ConvertWithMetadata() function with typed structs for metadata extraction.
convertWithMetadata() method with Java records for metadata extraction.
ConvertWithMetadata() method with C# records for metadata extraction.
html_to_markdown_convert_with_metadata() C function for language-agnostic metadata extraction.
packages/typescript/METADATA.md (480 lines) and packages/ruby/METADATA.md (228 lines)max_structured_data_size default (100KB → 1MB)--extract-document flag not being mapped to MetadataConfig.
max_structured_data_size default (was 100KB, should be 1MB).
DEFAULT_MAX_STRUCTURED_DATA_SIZE constant from Rust coreDEFAULT_MAX_STRUCTURED_DATA_SIZE: usize = 1_000_000 constant in Rust coreNew convert_with_metadata API across all bindings (Python, TypeScript/Node, Ruby, PHP, WASM) returning Markdown + extracted metadata in one call.
convert_with_metadata API across all bindings (Python, TypeScript/Node, Ruby, PHP, WASM) returning Markdown + extracted metadata in one call.hasMetadataSupport() runtime detection, expanded docs.convert_with_metadata.convert_with_metadata wrapper and redundant ? types.convert_with_metadata() function returning both markdown and extracted metadata in a single pass.hasMetadataSupport(), and 600+ lines of documentation.convert_with_metadata and fixed redundant ? symbols in RBS type annotations.Escape literal | characters inside table cells while leaving pipes inside and untouched to avoid rendering backslashes in code spans/blocks (fixes #14
| characters inside table cells while leaving pipes inside <code> and <pre> untouched to avoid rendering backslashes in code spans/blocks (fixes #140).WebAssembly bundler target now supports Cloudflare Workers, Wrangler, and modern bundlers that provide WebAssembly.Module instead of WebAssembly.Insta
WebAssembly.Module instead of WebAssembly.Instance.examples/wasm-node: Node.js example using dist-node targetexamples/wasm-rollup: Browser example using dist-web target with Rollupexamples/wasm-cloudflare: Cloudflare Workers example using bundler target with WranglerWebAssembly.Module instances, building the proper import namespace for wasm-bindgen glue functions.Harden Node/WASM bundles to emit fully typed doc comments (no stray any) so the options parameter and inline image attributes stay aligned with WasmCo
any) so the options parameter and inline image attributes stay aligned with WasmConversionOptions.examples/wasm-rollup and document it in the WASM README to guide bundler integrations.task sync-versions.WasmConversionOptions typedef and emit typed doc comments (including typed inline-image attributes), so no any annotations leak into the published dist, dist-node, dist-web, or docs bundles.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.11.2...v2.11.3
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.11.2...v2.11.3
PanicException in the Python bindings when processing long anchors (resolves #139) and add a regression test to keep the truncation logic safe.CI: Fix aarch64-unknown-linux-gnu CLI binary builds by switching from cross-rs to native cross-compilation to avoid glibc version mismatch
This release fixes the publish workflow failures that prevented v2.11.0 from completing successfully.
<pre><code> blocks while safely dedenting whitespace across multibyte characters to avoid panics when leading spaces are non-ASCII; regression fixture added for issue #134. Thanks @bbeardsley for the contribution.Normalize whitespace inside link labels (collapse newlines and extra spaces) so anchors with messy HTML do not emit multi-line [] text.
UTF-8 safety (Fix #127): guard whitespace trimming against mid-codepoint truncation to eliminate multilingual panics; add fixture + regression test.
<img> elements with width/height now render as Markdown images instead of raw HTML; regression test covers inline-data URIs with dimensions.Fix issue #121 regressions: SPA shell and Hacker News samples now render full Markdown output (new fixtures/tests).
Deprecation Warnings - Updated CLI tests to use CARGO_BIN_EXE env var instead of deprecated cargo_bin method
All notable changes to html-to-markdown will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
html_to_markdown Hex package built with Rustler, exposing the Rust core converter to Elixir with configurable options plus convert/2 and convert!/2.e2e/wasm-wasmtime) plus task wasm:test:wasmtime to compile the html-to-markdown-wasm artefact for wasm32-unknown-unknown and execute it inside Wasmtime. CI now runs these tests to ensure the WASM package works outside the browser runtime.tl parser – The HTML parser dependency now points to the actively maintained astral-tl fork (still imported as tl) so comment parsing stays up to date with upstream fixes.Goldziher.HtmlToMarkdown to avoid clashing with an existing community package.Cargo.toml, fixing the “failed to load manifest” error in the publish workflow.<p>…</p><hr> now emits a blank line before --- while preserving blockquote spacing so the rule is never misinterpreted as a setext heading.<!----> comment nodes are normalized before parsing, so comment placeholders no longer cause the following content to disappear.uv sync invocation in CI and the Taskfile now runs with --no-install-workspace, ensuring Python dependencies are resolved without mutating editable installs before the subsequent build/test steps run.NuGet/login@v1 (OIDC → short-lived API key) before pushing artifacts, removing the dependency on long-lived secrets.mix hex.publish --yes from packages/elixir, with ex_doc bundled as a dev dependency so documentation generation works during release.scripts/sync_versions.py now updates Elixir @version declarations, the C# .csproj, and the Java pom.xml (alongside every npm/pyproject/Gemfile manifest). task sync-versions bumps the entire multi-language stack to 2.8.2 in one shot.elixir:update plus full java:{install,update,test,lint} tasks so task setup, task update, task test, and task lint cover every published runtime (Go, C#, Elixir, Java) just like the CI workflows.task bench:bindings harness and ship with comprehensive performance data in their READMEs. C# leads at ~1.4k ops/sec (≈171 MB/s), Go at ~1.3k ops/sec (≈165 MB/s), and Java at ~1.0k ops/sec (≈126 MB/s) on the 129 KB Wikipedia lists fixture.<nav>, <form>, and related elements (along with all their children) were dropped by default, causing important content inside these tags to be lost. Users who want preprocessing must now explicitly enable it via PreprocessingOptions { enabled: true, ... }. The CLI behavior is unchanged (preprocessing has always been opt-in with --preprocess).edition = "2024" and rust-version = "1.85" from the workspace to keep toolchain configuration centralized.cargo build matching ext-php-rs's proven approach, resolving LLVM 19 MMX header incompatibilities and Zend symbol linking errors.packages/go and examples/go-smoke modules, fixing "directory prefix does not contain main module" errors by running each check from within its Go module directory.examples/go-smoke/main.go fmt.Println call (detected by newly-working golangci-lint).<html>, <head>, or <body> wrappers when their classes resemble navigation chrome, so large Wikipedia fixtures once again emit full markdown (restoring the Vitest length/table expectations for Node bindings and keeping WASM conversions consistent).html-to-markdown-wasm now expose initWasm()/wasmReady so edge runtimes that instantiate WebAssembly modules asynchronously (Cloudflare Workers, Vite dev servers, etc.) can await initialization before calling convert(), eliminating the __wbindgen_start runtime error.<footer> content unless the element carries explicit navigation hints (role/class/id). Python and Rust conversions once again preserve footer copy while still stripping true navigation footers such as .site-footer menus.task bench:bindings throughput numbers so runtime documentation stays aligned with the shared fixtures.examples/{node,wasm,python,ruby,php,rust}-smoke plus an overview README to exercise both the published artifacts and local builds before a release.html-to-markdown-node, html-to-markdown-wasm) and provide correct import samples..gitignore now drops .venv/, vendor directories, and nested node_modules/ so smoke tests and language-specific toolchains don’t dirty the tree.extconf.rb uses a relative path for the embedded Cargo crate and the crate’s Cargo.toml now declares explicit edition, rust-version, and dependency pins, allowing gem install outside the workspace.scripts/sync_versions.py updates every html-to-markdown-rs dependency pin (workspace root plus downstream crates) to keep cross-language releases in lockstep.html-to-markdown/extension or html-to-markdown 2.7.1 before release.convertInlineImagesBuffer / convertBytesWithInlineImages, letting benchmark harnesses feed Buffer/Uint8Array data directly without creating intermediate JS strings.< escaping, script/style stripping) now happens in a single streaming pass that hands owned buffers straight to tl::parse_owned, cutting multiple allocations from every conversion.tools/runtime-bench/results/latest.json.2.7.0 via task sync-versions.json native extensions, preventing libruby incompatibility errors during task bench:bindings.<strong> Normalization (Fix #111) – The Rust converter now tracks when bold markup is already active, so nested <b>/<strong> combinations (including <mark>, <summary>, <legend>) no longer generate **** artifacts (<b>bo<b>ld</b>er</b> correctly becomes **bolder**). The CommonMark harness documents the four spec examples that expect stacked markers and skips them accordingly.<h1>…<h6> so pretty-printed HTML like <h2>Heading\n Text</h2> renders as a single Markdown heading line.<input>, <script>, empty <b>) no longer collapses surrounding spaces; fixtures like test_chomp, test_form_with_inputs_inline_mode, and checkbox/task-list rendering now match their expected double-space gaps.<!DOCTYPE …> declarations are stripped during preprocessing so they never leak as stray PUBLIC… text in the output, even when metadata extraction is enabled.html-to-markdown-rb crate under packages/ruby/ext/html-to-markdown-rb/native and pointed extconf.rb at that path so every published gem now contains the Cargo sources it needs to compile on install.html-to-markdown npm package and to consistently list our supported targets (Node, WASM, Python, Ruby, PHP, CLI).task update to upgrade Rust crates, npm packages, Bundler gems, Python requirements, and Composer dependencies across the monorepo.clippy::unnecessary-map-or in the converter and hOCR table builder by using .is_none_or, keeping inline-image filtering and column pruning logic clear while allowing cargo clippy -D warnings to pass.scripts/package_php_pie_source.sh now copies packages/ruby/.../native into the temporary workspace so the Ruby crate exists when PIE builds the PHP extension.is_tag output in publish workflow that caused all publishing jobs to be skippedoptionalDependencies to html-to-markdown-node package.json to properly link platform-specific binariesscripts/sync_versions.py) to maintain consistency across all package manifests (Rust, Node.js, Python, Ruby, WASM)task sync-versions command to Taskfile for easy version synchronization across the monorepoCARGO_BIN_EXE env var instead of deprecated cargo_bin methodcriterion::black_box with std::hint::black_box in benchmarksgoldziher/html-to-markdown) providing native HTML to Markdown conversion for PHP 8.2+
HTM2MD_* to HTMLTOMARKDOWN_* for improved clarity in Makefile.frag and config.m4goldziher/html-to-markdown, moved the Composer metadata to the repository root, and refreshed the documentation/badges for every language target..gem.html-to-markdown-rb) with CLI proxy and comprehensive specs.preserve_tags option - Preserve specific HTML tags in their original HTML form instead of converting them to Markdown. This is useful for complex elements like tables that may not convert well to Markdown. Fixes issue #95.
["table", "form"])strip_tags - can use both options togetherPreprocessingOptions.enabled default changed from False to True to ensure robust handling of malformed HTML. Users who want minimal preprocessing can explicitly set enabled=False.<input type="checkbox"> elements when remove_forms is enabled (default). Checkboxes are now preserved during preprocessing to enable proper task list conversion (- [x] / - [ ]).
input tag to allowed tags in all sanitization presets (minimal, standard, aggressive)type and checked attributes on input elementsdata: URLs from image src attributes. Base64-encoded inline images (data URIs) are now preserved during preprocessing.
data to allowed URL schemes in all sanitization presetsconvert_with_inline_images functionality for base64-encoded images<span class="ocrx_word"> elements in hOCR documents. Words now have proper spaces between them.
OcrxWord converter to insert space before each word if output doesn't end with whitespace or markdown formatting characters*text*, [alt](url), `code`)class attributes on all elements (required for detecting hOCR element types)<meta> tags with name and content attributes (required for hOCR metadata detection)<head> tags (container for meta tags)<head> container element. The extractor now finds orphaned meta tags anywhere in the document, not just inside <head> elements.preserve_tags functionality with preprocessing - Fixed preserve_tags not working when HTML preprocessing is enabled (the new default). The sanitizer now:
preserve_tags list and allows those tags through sanitizationid, class, style, title, etc.) on preserved tagsremove_forms from stripping form tags when they're in the preserve listsvg, circle, rect, path, line, polyline, polygon, ellipse, gwidth, height, viewBox, cx, cy, r, x, y, d, fill, strokeconvert_with_inline_images to capture inline SVG elements< or > characters appear in HTML text content (e.g., 1<2, mathematical comparisons). The converter now:
1<2, 1 < 2 < 3, and angle brackets at tag boundariesgetrandom backend for wasm32-unknown-unknown targetswasm_js backend and strip wasm-pack .gitignore files so published packages ship the compiled .wasm artifacts.pyo3) to their latest compatible releases and refreshed lockfiles.getrandom's wasm_js feature, restoring WebAssembly builds.files list so published tarballs now include compiled .node artifacts, CommonJS shims, and typings.InlineImage, InlineImageWarning, and InlineImageConfig) alongside convert_with_inline_images, with dedicated regression tests.--version test updated to assert the new release number.hocr_spatial_tables option on ConversionOptions (Rust, Python, CLI) with --no-hocr-spatial-tables flag to disable spatial table reconstruction when desired.--version output and package metadata now report version 2.2.0 consistently.ocr_table markers.convert_with_inline_images() function to extract embedded images during conversion
data:image/*)InlineImageConfig with options for:
HtmlExtraction with markdown, extracted images, and warningsParsingOptions class in favor of direct encoding parameter on ConversionOptionshocr_extract_tables option (always enabled for hOCR content)hocr_table_column_threshold option (uses built-in heuristics)hocr_table_row_threshold_ratio option (uses built-in heuristics)hocr/spatial.rs moduleocr_table elements, preventing false positives.exe extension on Windowsscripts/ directoryVersion 2.0.0 represents a complete rewrite of html-to-markdown with a high-performance Rust backend, delivering 10-30x performance improvements while maintaining full backward compatibility through a v1 compatibility layer.
V2 adopts CommonMark-compliant defaults for better interoperability:
| Option | V1 Default | V2 Default | Reason |
|---|---|---|---|
list_indent_width |
4 | 2 | CommonMark standard |
bullets |
"-" | "*+-" | Cycling bullets for nested lists |
escape_asterisks |
true | false | Minimal escaping |
escape_underscores |
true | false | Minimal escaping |
escape_misc |
true | false | Minimal escaping |
newline_style |
"backslash" | "spaces" | CommonMark two-space line breaks |
code_block_style |
"backticks" | "indented" | CommonMark 4-space indent |
heading_style |
"underlined" | "atx" | CommonMark # headings |
preprocessing.enabled |
false | false | No change (opt-in) |
Migration: If you relied on v1 defaults, explicitly set options to match v1 behavior.
The following v1 CLI flags are not supported in v2. The Python CLI proxy will raise helpful error messages when these flags are used:
| Removed Flag | Reason | Migration |
|---|---|---|
--strip |
Feature removed in v2 | Remove flag (feature no longer available) |
--convert |
Feature removed in v2 | Remove flag (feature no longer available) |
Note on Redundant Flags: The following v1 flags are redundant in v2 (they match the defaults) but are silently accepted for backward compatibility:
--no-escape-asterisks, --no-escape-underscores, --no-escape-misc (v2 defaults to minimal escaping)--no-wrap (v2 defaults to no wrapping)--no-autolinks (Rust CLI defaults to no autolinks)--no-extract-metadata (Rust CLI defaults to no metadata extraction)These flags can be safely removed from your commands, or you can leave them for compatibility.
Note: The Rust CLI only supports positive flags (e.g., --escape-asterisks, --autolinks, --wrap). Negative flags (--no-*) are only supported through the Python CLI proxy for v1 compatibility.
* Item 1\n\n + Nested\n* Item 1\n + Nested\nscraper and html5everconvert(html, options, preprocessing) - primary API entry pointConversionOptions - comprehensive conversion settings (now includes encoding)PreprocessingOptions - HTML cleaning configurationConversionOptions.pyi files)convert_to_markdown() function with all v1 kwargscargo-llvm-covconverters.py, processing.py, preprocessor.py)The following v1 features were removed in v2:
code_language_callback - Removed (use code_language option for default language)strip option - Removed (use preprocessing options instead)convert option - Removed (all supported tags are converted by default)convert_to_markdown_stream() - Removed (html5ever does not support streaming parsing)custom_converters - Planned for future release with Rust and Python callback supportIf you're using the v1 API, your code will continue to work:
from html_to_markdown import convert_to_markdown
# This still works in v2!
markdown = convert_to_markdown(html, heading_style="atx")
from html_to_markdown import convert, ConversionOptions
options = ConversionOptions(heading_style="atx")
markdown = convert(html, options)
V1 CLI flags are automatically translated to v2:
# V1 style (still works)
html-to-markdown --preprocess-html --escape-asterisks input.html
# V2 style (recommended)
html-to-markdown --preprocess input.html # escaping is default
Real-world performance improvements over v1 (Apple M4):
| Document Type | Size | V2 Latency | V2 Throughput | Speedup vs V1 (2.5 MB/s) |
|---|---|---|---|---|
| Lists (Timeline) | 129KB | 0.62ms | 208 MB/s | 83x |
| Tables (Countries) | 360KB | 2.02ms | 178 MB/s | 71x |
| Mixed (Python wiki) | 656KB | 4.56ms | 144 MB/s | 58x |
V2's Rust engine delivers 60-80x higher throughput than V1's Python/BeautifulSoup implementation across real-world documents.
crates/
├── html-to-markdown/ # Core conversion library
├── html-to-markdown-py/ # Python bindings (PyO3)
└── html-to-markdown-cli/ # Native CLI binary
html_to_markdown/
├── api.py # V2 API
├── options.py # V2 configuration dataclasses
├── v1_compat.py # V1 compatibility layer
├── cli_proxy.py # CLI argument translation
├── _rust.pyi # Rust binding type stubs
└── __init__.py # Public API exports
None if using v1 compatibility layer. If migrating to v2 API:
convert_to_markdown → convertConversionOptions)| Aspect | V1 | V2 |
|---|---|---|
| Primary API | convert_to_markdown(**kwargs) |
convert(html, options, preprocessing, parsing) |
| Configuration | Keyword arguments | Dataclasses (ConversionOptions, etc.) |
| Type Safety | Basic type hints | Full .pyi stubs + generics |
| Compatibility Layer | N/A | convert_to_markdown() with v1 kwargs |
| Document Type | V1 Throughput | V2 Throughput | Speedup |
|---|---|---|---|
| Lists (Timeline) | 2.5 MB/s | 208 MB/s | 83x |
| Tables (Countries) | 2.5 MB/s | 178 MB/s | 71x |
| Mixed (Python wiki) | 2.5 MB/s | 144 MB/s | 58x |
| Average | 2.5 MB/s | 177 MB/s | 71x |
| Component | V1 | V2 |
|---|---|---|
| HTML Parser | BeautifulSoup4 / lxml | html5ever (Rust) |
| Sanitizer | Custom Python | html5ever DOM filtering |
| Conversion | Pure Python (~3,850 lines) | Pure Rust (~4,800 lines) |
| Bindings | N/A | PyO3 |
| CLI | Python wrapper | Native Rust binary |
| Dependencies | bs4, lxml, soupsieve | None (statically linked) |
| HTML | V1 Output | V2 Output |
|---|---|---|
<ul><li>Item</li></ul> |
* Item (4 spaces) |
- Item (2 spaces) |
<h1>Title</h1> |
Title\n===== |
# Title |
Text*with*stars |
Text\*with\*stars |
Text*with*stars |
<br> |
Two trailing spaces | Backslash \ |
<pre>code</pre> |
```\ncode\n``` |
Indented 4 spaces |
These differences reflect v2's alignment with CommonMark specification.
html_to_markdown/converters.py (1220 lines)html_to_markdown/processing.py (1195 lines)html_to_markdown/preprocessor.py (404 lines)html_to_markdown/whitespace.py (293 lines)html_to_markdown/utils.py (37 lines).skipTotal: ~3,850 lines of Python code removed, replaced by ~4,800 lines of Rust
abi3 for Python 3.10+ wheel reuseFor changes in v1.x releases, see git history before the v2 rewrite.
Deterministic uv installs – All automation calls to uv sync now use --no-install-workspace, keeping editable installs untouched until explicit build/t
uv sync now use --no-install-workspace, keeping editable installs untouched until explicit build/test steps run.NuGet/login@v1 (OIDC → short-lived API key) before pushing packages, so no long-lived secrets are needed.mix hex.publish --yes from packages/elixir and ships ex_doc as a dev dependency so documentation generation succeeds on Hex.pm builds.gpg2.uv sync invocation in CI and the Taskfile now runs with --no-install-workspace, ensuring Python dependencies are resolved without mutating editable installs before the subsequent build/test steps run.NuGet/login@v1 (OIDC → short-lived API key) before pushing artifacts, removing the dependency on long-lived secrets.mix hex.publish --yes from packages/elixir, with ex_doc bundled as a dev dependency so documentation generation works during release.Unified Version Sync – scripts/sync_versions.py now keeps every manifest (Elixir, C#, Java, npm, PyPI, Ruby) in lockstep whenever we run task sync-ver
scripts/sync_versions.py now keeps every manifest (Elixir, C#, Java, npm, PyPI, Ruby) in lockstep whenever we run task sync-versions.elixir:update plus full java:{install,update,test,lint} helpers so task setup/update/test/lint cover all runtimes (Go, C#, Elixir, Java).erlangpack/github-action@v3 with HEX_TOKEN and publishes straight from packages/elixir.NUGET_API_KEY secrets and push automatically.WasiqB/maven-publish-action@v1 on Temurin JDK 22 (setup-java@v5).scripts/sync_versions.py now updates Elixir @version declarations, the C# .csproj, and the Java pom.xml (alongside every npm/pyproject/Gemfile manifest). task sync-versions bumps the entire multi-language stack to 2.8.2 in one shot.elixir:update plus full java:{install,update,test,lint} tasks so task setup, task update, task test, and task lint cover every published runtime (Go, C#, Elixir, Java) just like the CI workflows.Release Pipeline – Bumped all package manifests to v2.8.1 so the publish workflow can push fresh artifacts after the v2.8.0 smoke-test fixes (PyPI, np
Java, C#, and Go Bindings (First Release) – First public release of official Java (JNA), C# (.NET), and Go (CGO) language bindings. All three are inte
task bench:bindings harness and ship with comprehensive performance data in their READMEs. C# leads at ~1.4k ops/sec (≈171 MB/s), Go at ~1.3k ops/sec (≈165 MB/s), and Java at ~1.0k ops/sec (≈126 MB/s) on the 129 KB Wikipedia lists fixture.<nav>, <form>, and related elements (along with all their children) were dropped by default, causing important content inside these tags to be lost. Users who want preprocessing must now explicitly enable it via PreprocessingOptions { enabled: true, ... }. The CLI behavior is unchanged (preprocessing has always been opt-in with --preprocess).edition = "2024" and rust-version = "1.85" from the workspace to keep toolchain configuration centralized.cargo build matching ext-php-rs's proven approach, resolving LLVM 19 MMX header incompatibilities and Zend symbol linking errors.packages/go and examples/go-smoke modules, fixing "directory prefix does not contain main module" errors by running each check from within its Go module directory.examples/go-smoke/main.go fmt.Println call (detected by newly-working golangci-lint).Node/WASM Binding Regression – HTML preprocessing no longer drops , , or wrappers when their classes resemble navigation chrome, so large Wikipedia fi
<html>, <head>, or <body> wrappers when their classes resemble navigation chrome, so large Wikipedia fixtures once again emit full markdown (restoring the Vitest length/table expectations for Node bindings and keeping WASM conversions consistent).html-to-markdown-wasm now expose initWasm()/wasmReady so edge runtimes that instantiate WebAssembly modules asynchronously (Cloudflare Workers, Vite dev servers, etc.) can await initialization before calling convert(), eliminating the __wbindgen_start runtime error.<footer> content unless the element carries explicit navigation hints (role/class/id). Python and Rust conversions once again preserve footer copy while still stripping true navigation footers such as .site-footer menus.Full changelog: https://github.com/Goldziher/html-to-markdown/blob/main/CHANGELOG.md#272---2025-11-12
Language-specific READMEs now publish the latest benchmark tables for Node, WASM, Python, Ruby, PHP, and TypeScript.
examples/ to install every binding from npm/PyPI/RubyGems/Composer/crates.io and from the local workspace before tagging a release..gitignore filters nested node_modules/, .venv/, and vendor/ directories so example installs don't dirty the tree.extconf.rb at the embedded Cargo crate via a relative path, and the crate manifest declares explicit edition, rust-version, and dependency pins.scripts/sync_versions.py also bumps every html-to-markdown-rs dependency pin to keep downstream crates in sync with the workspace version.task bench:bindings throughput numbers so runtime documentation stays aligned with the shared fixtures.examples/{node,wasm,python,ruby,php,rust}-smoke plus an overview README to exercise both the published artifacts and local builds before a release.html-to-markdown-node, html-to-markdown-wasm) and provide correct import samples..gitignore now drops .venv/, vendor directories, and nested node_modules/ so smoke tests and language-specific toolchains don’t dirty the tree.extconf.rb uses a relative path for the embedded Cargo crate and the crate’s Cargo.toml now declares explicit edition, rust-version, and dependency pins, allowing gem install outside the workspace.scripts/sync_versions.py updates every html-to-markdown-rs dependency pin (workspace root plus downstream crates) to keep cross-language releases in lockstep.goldziher/html-to-markdown or html-to-markdown 2.7.1 before release.Added zero-copy inline-image APIs for Node (N-API) and WASM bindings so benchmark drivers can feed Buffer/Uint8Array payloads directly without re-enco
convertInlineImagesBuffer / convertBytesWithInlineImages, letting benchmark harnesses feed Buffer/Uint8Array data directly without creating intermediate JS strings.< escaping, script/style stripping) now happens in a single streaming pass that hands owned buffers straight to tl::parse_owned, cutting multiple allocations from every conversion.tools/runtime-bench/results/latest.json.2.7.0 via task sync-versions.json native extensions, preventing libruby incompatibility errors during task bench:bindings.<strong> Normalization (Fix #111) – The Rust converter now tracks when bold markup is already active, so nested <b>/<strong> combinations (including <mark>, <summary>, <legend>) no longer generate **** artifacts (<b>bo<b>ld</b>er</b> correctly becomes **bolder**). The CommonMark harness documents the four spec examples that expect stacked markers and skips them accordingly.<h1>…<h6> so pretty-printed HTML like <h2>Heading\n Text</h2> renders as a single Markdown heading line.<input>, <script>, empty <b>) no longer collapses surrounding spaces; fixtures like test_chomp, test_form_with_inputs_inline_mode, and checkbox/task-list rendering now match their expected double-space gaps.<!DOCTYPE …> declarations are stripped during preprocessing so they never leak as stray PUBLIC… text in the output, even when metadata extraction is enabled.chore(deps): bump actions/download-artifact from 5 to 6 by @dependabot[bot] in https://github.com/Goldziher/html-to-markdown/pull/117
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.6.5...v2.6.6
html-to-markdown-rb crate under packages/ruby/ext/html-to-markdown-rb/native and pointed extconf.rb at that path so every published gem now contains the Cargo sources it needs to compile on install.html-to-markdown npm package and to consistently list our supported targets (Node, WASM, Python, Ruby, PHP, CLI).task update to upgrade Rust crates, npm packages, Bundler gems, Python requirements, and Composer dependencies across the monorepo.clippy::unnecessary-map-or in the converter and hOCR table builder by using .is_none_or, keeping inline-image filtering and column pruning logic clear while allowing cargo clippy -D warnings to pass.scripts/package_php_pie_source.sh now copies packages/ruby/.../native into the temporary workspace so the Ruby crate exists when PIE builds the PHP extension.Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.6.4...v2.6.5
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.6.4...v2.6.5
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.6.3...v2.6.4
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.6.3...v2.6.4
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.6.2...v2.6.3
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v2.6.2...v2.6.3
is_tag output in publish workflow that caused all publishing jobs to be skippedoptionalDependencies to html-to-markdown-node package.json to properly link platform-specific binariesscripts/sync_versions.py) to maintain consistency across all package manifests (Rust, Node.js, Python, Ruby, WASM)task sync-versions command to Taskfile for easy version synchronization across the monorepoDeprecation Warnings - Updated CLI tests to use CARGO_BIN_EXE env var instead of deprecated cargo_bin method
CARGO_BIN_EXE env var instead of deprecated cargo_bin methodcriterion::black_box with std::hint::black_box in benchmarksNode.js Platform Packages - Fixed publishing of platform-specific npm packages. The workflow now correctly packs npm directories into .tgz files befor
This release adds official PHP extension support alongside existing Rust, Python, Node.js, Ruby, and WASM bindings.
This release adds official PHP extension support alongside existing Rust, Python, Node.js, Ruby, and WASM bindings.
goldziher/html-to-markdown) providing native HTML to Markdown conversionSee CHANGELOG.md for complete details.
goldziher/html-to-markdown) providing native HTML to Markdown conversion for PHP 8.2+
HTM2MD_* to HTMLTOMARKDOWN_* for improved clarity in Makefile.frag and config.m4Version bump to 2.5.6 across Rust, Python, Ruby, npm (node + wasm), and Homebrew distributions.
See CHANGELOG.md#256---2025-10-30 for details.
# Rust CLI
cargo install html-to-markdown-cli --version 2.5.6
# Python
pip install html-to-markdown==2.5.6
# Node.js / Bun
npm install html-to-markdown-node@2.5.6
# WebAssembly bundle
npm install html-to-markdown-wasm@2.5.6
# Ruby
gem install html-to-markdown -v 2.5.6
# Homebrew
brew tap goldziher/tap
brew install html-to-markdown
Prebuilt binaries and binding artifacts for every supported platform are attached below. Pick the archive matching your operating system and CPU architecture if you prefer manual downloads.
Version bump to 2.5.5 across every binding (Rust crates, npm packages, PyPI wheels/sdist, Ruby gems, Homebrew taps).
README.md, so the refreshed docs and performance guidance flow through to RubyGems.See CHANGELOG.md#255---2025-10-30 for the complete list of changes.
# Rust CLI
cargo install html-to-markdown-cli --version 2.5.5
# Python
pip install html-to-markdown==2.5.5
# Node.js / Bun
npm install html-to-markdown-node@2.5.5
# WebAssembly bundle
npm install html-to-markdown-wasm@2.5.5
# Ruby
gem install html-to-markdown -v 2.5.5
# Homebrew
brew tap goldziher/tap
brew install html-to-markdown
Prebuilt binaries and binding artifacts for every supported platform are attached below. Pick the file that matches your operating system/architecture if you prefer manual downloads.
Version bump to 2.5.4 across every binding (Rust crates, npm packages, PyPI wheel/sdist, Ruby gems, Homebrew taps).
See CHANGELOG.md#254---2025-10-30 for the detailed list of changes.
# Rust CLI
cargo install html-to-markdown-cli --version 2.5.4
# Python
pip install html-to-markdown==2.5.4
# Node.js / Bun
npm install html-to-markdown-node@2.5.4
# WebAssembly bundle
npm install html-to-markdown-wasm@2.5.4
# Ruby
gem install html-to-markdown -v 2.5.4
# Homebrew
brew tap goldziher/tap
brew install html-to-markdown
Prebuilt binaries and binding artifacts for every supported platform are attached below. Pick the file that matches your operating system/architecture if you prefer manual downloads.
Ship html-to-markdown 2.5.3 across every language binding (Rust crates, npm packages, PyPI wheel/sdist, Ruby gems, Homebrew binary).
scripts/prepare_ruby_gem.rb clears stale CLI binaries before packaging so each platform bundle contains the correct executable.Refer to CHANGELOG.md#253---2025-10-30 for the complete list of changes.
# Rust CLI
cargo install html-to-markdown-cli --version 2.5.3
# Python
pip install html-to-markdown==2.5.3
# Node.js / Bun
npm install html-to-markdown-node@2.5.3
# WebAssembly bundle
npm install html-to-markdown-wasm@2.5.3
# Ruby
gem install html-to-markdown -v 2.5.3
# Homebrew
brew tap goldziher/tap
brew install html-to-markdown
Prebuilt binaries and binding artifacts for every supported platform are attached below. Pick the file that matches your operating system/architecture if you prefer manual downloads.
.gem.Fix Ruby gem packaging to embed standalone Cargo manifest (no workspace inheritance) so installs compile out of tree successfully.
Magnus-based Ruby gem (html-to-markdown-rb) with CLI proxy and comprehensive specs.
html-to-markdown-rb) with CLI proxy and comprehensive specs.See the CHANGELOG for full details.
See the CHANGELOG for full details.
npm install html-to-markdown-node
npm install html-to-markdown-wasm
pip install html-to-markdown
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
Download the appropriate binary for your platform below.
preserve_tags option - Preserve specific HTML tags in their original HTML form instead of converting them to Markdown. This is useful for complex elements like tables that may not convert well to Markdown. Fixes issue #95.
["table", "form"])strip_tags - can use both options togetherPreprocessingOptions.enabled default changed from False to True to ensure robust handling of malformed HTML. Users who want minimal preprocessing can explicitly set enabled=False.<input type="checkbox"> elements when remove_forms is enabled (default). Checkboxes are now preserved during preprocessing to enable proper task list conversion (- [x] / - [ ]).
input tag to allowed tags in all sanitization presets (minimal, standard, aggressive)type and checked attributes on input elementsdata: URLs from image src attributes. Base64-encoded inline images (data URIs) are now preserved during preprocessing.
data to allowed URL schemes in all sanitization presetsconvert_with_inline_images functionality for base64-encoded images<span class="ocrx_word"> elements in hOCR documents. Words now have proper spaces between them.
OcrxWord converter to insert space before each word if output doesn't end with whitespace or markdown formatting characters*text*, [alt](url), `code`)class attributes on all elements (required for detecting hOCR element types)<meta> tags with name and content attributes (required for hOCR metadata detection)<head> tags (container for meta tags)<head> container element. The extractor now finds orphaned meta tags anywhere in the document, not just inside <head> elements.preserve_tags functionality with preprocessing - Fixed preserve_tags not working when HTML preprocessing is enabled (the new default). The sanitizer now:
preserve_tags list and allows those tags through sanitizationid, class, style, title, etc.) on preserved tagsremove_forms from stripping form tags when they're in the preserve listsvg, circle, rect, path, line, polyline, polygon, ellipse, gwidth, height, viewBox, cx, cy, r, x, y, d, fill, strokeconvert_with_inline_images to capture inline SVG elements< or > characters appear in HTML text content (e.g., 1<2, mathematical comparisons). The converter now:
1<2, 1 < 2 < 3, and angle brackets at tag boundariesgetrandom backend for wasm32-unknown-unknown targetsSee the CHANGELOG for full details.
See the CHANGELOG for full details.
npm install html-to-markdown-node
npm install html-to-markdown-wasm
pip install html-to-markdown
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
Download the appropriate binary for your platform below.
See the CHANGELOG for full details.
See the CHANGELOG for full details.
npm install html-to-markdown-node
npm install html-to-markdown-wasm
pip install html-to-markdown
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
Download the appropriate binary for your platform below.
wasm_js backend and strip wasm-pack .gitignore files so published packages ship the compiled .wasm artifacts.See the CHANGELOG for full details.
See the CHANGELOG for full details.
npm install html-to-markdown-node
npm install html-to-markdown-wasm
pip install html-to-markdown
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
Download the appropriate binary for your platform below.
pyo3) to their latest compatible releases and refreshed lockfiles.getrandom's wasm_js feature, restoring WebAssembly builds.files list so published tarballs now include compiled .node artifacts, CommonJS shims, and typings.See the CHANGELOG for full details.
See the CHANGELOG for full details.
npm install html-to-markdown-node
npm install html-to-markdown-wasm
pip install html-to-markdown
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
Download the appropriate binary for your platform below.
See the CHANGELOG for full details.
See the CHANGELOG for full details.
npm install html-to-markdown-node
npm install html-to-markdown-wasm
pip install html-to-markdown
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
Download the appropriate binary for your platform below.
InlineImage, InlineImageWarning, and InlineImageConfig) alongside convert_with_inline_images, with dedicated regression tests.--version test updated to assert the new release number.See the CHANGELOG for full details.
See the CHANGELOG for full details.
npm install @html-to-markdown/node
npm install @html-to-markdown/wasm
pip install html-to-markdown
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
Download the appropriate binary for your platform below.
hocr_spatial_tables option on ConversionOptions (Rust, Python, CLI) with --no-hocr-spatial-tables flag to disable spatial table reconstruction when de
hocr_spatial_tables option on ConversionOptions (Rust, Python, CLI) with --no-hocr-spatial-tables flag to disable spatial table reconstruction when desired.--version output and package metadata now report version 2.2.0 consistently.See the CHANGELOG for full details.
See the CHANGELOG for full details.
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
pip install html-to-markdown
Download the appropriate binary for your platform below.
See the CHANGELOG for full details.
See the CHANGELOG for full details.
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
pip install html-to-markdown
Download the appropriate binary for your platform below.
convert_with_inline_images() function to extract embedded images during conversion
data:image/*)InlineImageConfig with options for:
HtmlExtraction with markdown, extracted images, and warningsParsingOptions class in favor of direct encoding parameter on ConversionOptionshocr_extract_tables option (always enabled for hOCR content)hocr_table_column_threshold option (uses built-in heuristics)hocr_table_row_threshold_ratio option (uses built-in heuristics)hocr/spatial.rs moduleocr_table elements, preventing false positives.exe extension on Windowsscripts/ directorySee the CHANGELOG for full details.
See the CHANGELOG for full details.
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
pip install html-to-markdown
Download the appropriate binary for your platform below.
See the CHANGELOG for full details.
See the CHANGELOG for full details.
brew tap goldziher/tap
brew install html-to-markdown
cargo install html-to-markdown-cli
pip install html-to-markdown
Download the appropriate binary for your platform below.
Version 2.0.0 represents a complete rewrite of html-to-markdown with a high-performance Rust backend, delivering 10-30x performance improvements while maintaining full backward compatibility through a v1 compatibility layer.
V2 adopts CommonMark-compliant defaults for better interoperability:
| Option | V1 Default | V2 Default | Reason |
|---|---|---|---|
list_indent_width |
4 | 2 | CommonMark standard |
bullets |
"-" | "*+-" | Cycling bullets for nested lists |
escape_asterisks |
true | false | Minimal escaping |
escape_underscores |
true | false | Minimal escaping |
escape_misc |
true | false | Minimal escaping |
newline_style |
"backslash" | "spaces" | CommonMark two-space line breaks |
code_block_style |
"backticks" | "indented" | CommonMark 4-space indent |
heading_style |
"underlined" | "atx" | CommonMark # headings |
preprocessing.enabled |
false | false | No change (opt-in) |
Migration: If you relied on v1 defaults, explicitly set options to match v1 behavior.
The following v1 CLI flags are not supported in v2. The Python CLI proxy will raise helpful error messages when these flags are used:
| Removed Flag | Reason | Migration |
|---|---|---|
--strip |
Feature removed in v2 | Remove flag (feature no longer available) |
--convert |
Feature removed in v2 | Remove flag (feature no longer available) |
Note on Redundant Flags: The following v1 flags are redundant in v2 (they match the defaults) but are silently accepted for backward compatibility:
--no-escape-asterisks, --no-escape-underscores, --no-escape-misc (v2 defaults to minimal escaping)--no-wrap (v2 defaults to no wrapping)--no-autolinks (Rust CLI defaults to no autolinks)--no-extract-metadata (Rust CLI defaults to no metadata extraction)These flags can be safely removed from your commands, or you can leave them for compatibility.
Note: The Rust CLI only supports positive flags (e.g., --escape-asterisks, --autolinks, --wrap). Negative flags (--no-*) are only supported through the Python CLI proxy for v1 compatibility.
* Item 1\n\n + Nested\n* Item 1\n + Nested\nscraper and html5everconvert(html, options, preprocessing) - primary API entry pointConversionOptions - comprehensive conversion settings (now includes encoding)PreprocessingOptions - HTML cleaning configurationConversionOptions.pyi files)convert_to_markdown() function with all v1 kwargscargo-llvm-covconverters.py, processing.py, preprocessor.py)The following v1 features were removed in v2:
code_language_callback - Removed (use code_language option for default language)strip option - Removed (use preprocessing options instead)convert option - Removed (all supported tags are converted by default)convert_to_markdown_stream() - Removed (html5ever does not support streaming parsing)custom_converters - Planned for future release with Rust and Python callback supportIf you're using the v1 API, your code will continue to work:
from html_to_markdown import convert_to_markdown
# This still works in v2!
markdown = convert_to_markdown(html, heading_style="atx")
from html_to_markdown import convert, ConversionOptions
options = ConversionOptions(heading_style="atx")
markdown = convert(html, options)
V1 CLI flags are automatically translated to v2:
# V1 style (still works)
html-to-markdown --preprocess-html --escape-asterisks input.html
# V2 style (recommended)
html-to-markdown --preprocess input.html # escaping is default
Real-world performance improvements over v1 (Apple M4):
| Document Type | Size | V2 Latency | V2 Throughput | Speedup vs V1 (2.5 MB/s) |
|---|---|---|---|---|
| Lists (Timeline) | 129KB | 0.62ms | 208 MB/s | 83x |
| Tables (Countries) | 360KB | 2.02ms | 178 MB/s | 71x |
| Mixed (Python wiki) | 656KB | 4.56ms | 144 MB/s | 58x |
V2's Rust engine delivers 60-80x higher throughput than V1's Python/BeautifulSoup implementation across real-world documents.
crates/
├── html-to-markdown/ # Core conversion library
├── html-to-markdown-py/ # Python bindings (PyO3)
└── html-to-markdown-cli/ # Native CLI binary
html_to_markdown/
├── api.py # V2 API
├── options.py # V2 configuration dataclasses
├── v1_compat.py # V1 compatibility layer
├── cli_proxy.py # CLI argument translation
├── _rust.pyi # Rust binding type stubs
└── __init__.py # Public API exports
None if using v1 compatibility layer. If migrating to v2 API:
convert_to_markdown → convertConversionOptions)| Aspect | V1 | V2 |
|---|---|---|
| Primary API | convert_to_markdown(**kwargs) |
convert(html, options, preprocessing, parsing) |
| Configuration | Keyword arguments | Dataclasses (ConversionOptions, etc.) |
| Type Safety | Basic type hints | Full .pyi stubs + generics |
| Compatibility Layer | N/A | convert_to_markdown() with v1 kwargs |
| Document Type | V1 Throughput | V2 Throughput | Speedup |
|---|---|---|---|
| Lists (Timeline) | 2.5 MB/s | 208 MB/s | 83x |
| Tables (Countries) | 2.5 MB/s | 178 MB/s | 71x |
| Mixed (Python wiki) | 2.5 MB/s | 144 MB/s | 58x |
| Average | 2.5 MB/s | 177 MB/s | 71x |
| Component | V1 | V2 |
|---|---|---|
| HTML Parser | BeautifulSoup4 / lxml | html5ever (Rust) |
| Sanitizer | Custom Python | html5ever DOM filtering |
| Conversion | Pure Python (~3,850 lines) | Pure Rust (~4,800 lines) |
| Bindings | N/A | PyO3 |
| CLI | Python wrapper | Native Rust binary |
| Dependencies | bs4, lxml, soupsieve | None (statically linked) |
| HTML | V1 Output | V2 Output |
|---|---|---|
<ul><li>Item</li></ul> |
* Item (4 spaces) |
- Item (2 spaces) |
<h1>Title</h1> |
Title\n===== |
# Title |
Text*with*stars |
Text\*with\*stars |
Text*with*stars |
<br> |
Two trailing spaces | Backslash \ |
<pre>code</pre> |
```\ncode\n``` |
Indented 4 spaces |
These differences reflect v2's alignment with CommonMark specification.
html_to_markdown/converters.py (1220 lines)html_to_markdown/processing.py (1195 lines)html_to_markdown/preprocessor.py (404 lines)html_to_markdown/whitespace.py (293 lines)html_to_markdown/utils.py (37 lines).skipTotal: ~3,850 lines of Python code removed, replaced by ~4,800 lines of Rust
abi3 for Python 3.10+ wheel reusefeat: add hocr support by @Goldziher in https://github.com/Goldziher/html-to-markdown/pull/79
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v1.15.0...v1.16.0
feat: add support for custom navigation classes in HTML preprocessing by @tommyjs007 in https://github.com/Goldziher/html-to-markdown/pull/77
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v1.14.1...v1.15.0
fix: resolve ancestor caching issue causing escaping inside code/pre blocks by @Goldziher in https://github.com/Goldziher/html-to-markdown/pull/78
Full Changelog: https://github.com/Goldziher/html-to-markdown/compare/v1.14.0...v1.14.1
Added html5lib parser support for enhanced HTML5 compliance and standards conformance
--source-encoding CLI parameter and source_encoding API parametererrors="replace" for encoding issues--source-encoding optionconvert_to_markdown() (resolves #73)None - This release maintains full backward compatibility while adding new features.
html5lib - For enhanced HTML5 standards compliancepip install html-to-markdown[html5lib]# For best performance (recommended)
pip install html-to-markdown[lxml]
# For maximum standards compliance
pip install html-to-markdown[html5lib]
# For all parser options
pip install html-to-markdown[lxml,html5lib]
--source-encoding instead of separate encoding parametersFull Changelog: https://github.com/Goldziher/html-to-markdown/compare/v1.13.0...v1.14.0
Parser-agnostic testing system - A comprehensive fixture system that allows tests to run with both lxml and html.parser backends, ensuring compatibili
Your coding agent can read these notes before it upgrades. Set up the MCP server →