NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1831 most downloaded on PyPI
High-performance HTML to Markdown converter
Last release today
04 Oct 2026
Ships fairly regularly
a new release about every 9 days
Nearly every release is documented
notes for 59 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
198 releases · first in 2024
One column per month.
Nothing published for this version
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.16.0 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.16.0")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 1bcb58b6ba627a666d0d0bf2a1f5de805ea6073f37dfac31acf0d61b7b851e61
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{
.url = "https://github.com/xberg-io/html-to-markdown/releases/download/v3.16.0/html-to-markdown-rs-zig-v3.16.0.tar.gz",
.hash = "html_to_markdown_rs-3.16.0-QtXyW771xgEVxyysYDlovxp_V64MG34dJP_F1ozZGsZC",
},
},
ConversionOptions::inline_data_media chooses what to write for an image or embedded media
element whose address is an inline data: URL, so the payload no longer has to land in the
output. keep (the default) writes the URL as before. alt_text_only writes the alt text: the
title of an inline <svg>, the fallback content of <video> and <audio>, nothing for an
<iframe>. drop_element writes nothing. It covers <img>, <graphic>, inline <svg>,
<video>, <audio> and <iframe>. With either of the last two choices, a real address on the
element wins over the data: one: a lazy-load attribute or srcset candidate on <img>,
another address attribute on <graphic>, a nested <source> on <video> and <audio>, and
a <source> of the <picture> that holds an <img>. The document structure follows the
markdown: no image node for a dropped element, and no address when only the alt text is
written. A link whose only content the option removed is dropped with it, instead of turning
into a link labelled with its own address. A link whose own address is a data: URL gets the
same choice: alt_text_only and drop_element both write the link's text with no destination,
since a link has no caption separate from its text the way an image has alt text, and that text
is never dropped along with the address. Extracted images do not change. The CLI takes it as
--inline-data-media (#528).WASM: assigning null or undefined to five optional fields now throws. The fields are
WasmConversionOptions.visitor, WasmConversionOptionsUpdate.visitor and preprocessing,
WasmConversionResult.document and WasmImageMetadata.dimensions. They borrow the value you
assign instead of taking it, and an assignment of null cleared them before. To unset one, call
clearVisitor(), clearPreprocessing(), clearDocument() or clearDimensions().
The alef pin in alef.toml is now 0.101.0, the version that generated the committed bindings.
Regenerating with 0.97.0 dropped the Go binding's runtime.LockOSThread calls.
The FFI Symbols CI gate now fails when a detector matches no call site, and names the silent
language. Before, a detector that stopped matching read exactly like a clean pass, so a
restyled binding dropped out of the diff with every check green. Each detector must also match
in every one of its required roots: the binding package for C#, Java, Go and Zig, and both
e2e/c and test_apps/c for C. So the alef-generated e2e/zig tests cannot keep a silent Zig
binding looking covered, and the vendored Go copy of the C header cannot do the same for C. The
PHP comparison fails the same way once the extension exports a function and its probe root is
missing. The --json summary now carries the per-language counts, the per-root counts and the
silent detectors.
CI E2E now ends in one E2E result job that fails unless every other job in the workflow
passed or was skipped because its path filters did not match. Before, a skipped leg left the
run as green as a passing one, whatever the reason for the skip, and no single check covered
every leg. Each leg's path filter condition is now written once, as an output of the change
detection job, and both the leg and the result job read that output. A script test fails when a
job is added to the workflow without being listed in the result job.
Removed an unused copy of the <a> handler that no conversion path called. Links are
converted by the one live handler, as before; output does not change.
Text after a <center>, <search>, <hgroup> or <dialog> in a table cell joined it
(#692). <table><tr><td>a<center><h2>h</h2>x</center>y</td></tr></table> gave | a h xy |
in both converters, and the full converter wrote | a h x y | for a <dialog>. These convert
like a <div> but were not counted as blocks, so the text after them got no cell break. They
now give | a h x y |, or | a<br>h<br>x<br>y | with br_in_tables on, as a <div> does. In a
list item, text after one of them now starts a paragraph in the item, as after a <div>, instead
of continuing the container's last line or leaving the list.
Counting them as blocks changes three more outputs, again to match a <div>. With
newline_style: backslash, the hard break before one of them is dropped: <p>a<br><center>b</center>c</p>
gives a, b and c as paragraphs instead of a\ then b. A dialog in a link label is set off
by spaces, so <a href="u">l<dialog>b</dialog>m</a> gives [l b m](u) instead of [lb m](u).
Bold or italic around one of them is closed before it and opened again after it:
<p><b>a<center>b</center>c</b></p> gives **a**, **b** and **c** instead of one bold run
across three paragraphs. Visitors now get is_inline false for these four tags.
Text before a <dialog> ran into it (#692). <p>a<dialog>b</dialog>c</p> gave ab, then
c, and <td>a<dialog>b</dialog>c</td> gave | ab c | in the full converter and | a bc | in
the fast one: the dialog wrote no break before its content. A dialog now converts like a <div>, so the paragraph gives a,
b and c as three paragraphs, the cell gives | a b c | in both converters, and in a list
item the dialog content starts a paragraph in the item.
Plain text output joined a <center> or <dialog> to the text after it (#692).
<p><b>a<center>b</center>c</b></p> with output_format: plain gave abc on one line. Plain
output now starts both on their own line, as it does for a <div>: a, b and c.
A <details>, <summary>, <menu> or <legend> joined the text around it (#692). In a
table cell, <td>a<menu>b</menu>c</td> gave | ab c | and a summary gave | a**b** c |. In
a list item, the text after a details left the list, and a summary, menu or legend ran into the
text before it. In a paragraph, <p>a<summary>b</summary>c</p> gave a**b**, then c. A
details and a menu now convert like a <div>, and a summary or legend is placed the way a
<div> is placed and stays bold: the cell gives | a b c | or | a **b** c | in both
converters, the list item keeps b and c as paragraphs in the item, and the paragraph gives
a, b and c apart. As with the tags above, these four now count as blocks: the backslash
break before them is dropped, bold around a details, summary or menu is closed and reopened,
plain output starts a menu or legend on its own line, and visitors get is_inline false for them. A summary that holds
a list with a quote in it, inside a details in a list item, now stays in the item, so with the
default or tab list indent its bold markers no longer span the quote.
A block in a heading ran into the words around it. <h1>a<div>b</div>c</h1> gave # abc
in the full converter and # a b c in the fast one. Both now give # a b c, for a <div> and for
every tag that converts like one, a details and a menu included. In the fast converter, a block inside a summary, figcaption or
table caption now writes its break into that element's text, so a rustdoc heading such as
impl Any for T<div class="where">where T: ...</div> gives for T where instead of for Twhere.
In a heading in a table cell, the text after a block ran into it.
<td><h2>a<div>b</div>c</h2></td> gave | a bc |. The text after the block now gets the cell
break, so the cell gives | a b c |, or | a<br>b<br>c | with br_in_tables on. The fast
converter leaves this shape to the full converter.
In a link label in a table cell, the text after a block ran into it.
<td><a href="u"><span>a<div>b</div>c</span></a></td> gave | [a bc](u) |. The text after the
block now gets the cell break, as in a heading in a cell, so the cell gives | [a b c](u) |, or
| [a<br>b<br>c](u) | with br_in_tables on.
Plain text output left a list item marker alone on its line. <ol><li><div>a</div>b</li></ol>
with output_format: plain gave 1. on a line of its own, then a and b. A block that opens a
list item now starts on the marker line, so the item gives 1. a, then b.
The fast converter dropped the bold of a <legend> and split a link label at a <summary>.
<legend>x</legend> gave x where the full converter gives **x**, and
<a href="u">l<summary>b</summary>m</a> gave a link label across three paragraphs where the full
converter gives [l b m](u). The fast converter now writes a legend in bold, as it does a summary,
and leaves a summary inside a link to the full converter, as it already did for a <div>.
The WASM binding used up a visitor handle on assignment, and convert() ignored
options.visitor (#517). Assigning a WasmVisitorHandle to WasmConversionOptions.visitor
moved it into the options, so a second options object could not take the same handle. The
setter now borrows the handle. convert() now uses options.visitor when you pass no visitor
argument, and a visitor argument still wins over the options field.
A line break on its own source line became a paragraph break (#683). First\n<br>\nSecond
gave First\n\nSecond, two paragraphs: the newline before the <br> put its hard-break marker
on a line of its own, and an empty line ends a paragraph. The break now ends the line of the text
before it, First \nSecond, in both converters and with both newline styles. This also holds
when that text ends inside an element that writes no line end, such as <span>, <font> or a
custom element: <span>First\n</span><br>Second.
A table cell ignored escape_underscores and escape_asterisks (#638).
The full converter always escaped _ and * in a cell, so
<table><tr><td>sample_value</td></tr></table> gave sample\_value while
<p>sample_value</p> gave sample_value. Every escape option now acts the same in a cell as
anywhere else. A | in a cell is still escaped whatever the options say, and with
escape_ascii on and escape_misc off it is now escaped once, as \|, instead of as \\|,
which some renderers show with a stray backslash.
A | in a code span, link or image in a table cell broke the table.
<table><tr><td><code>a|b</code></td></tr></table> gave | `a|b` |. GFM splits a row on
a pipe in a code span too, so the row no longer matched the delimiter row and the whole table
became a paragraph. A pipe in a link destination or title, an image description, preformatted
text or a code span in a nested table did the same. In Markdown output every pipe in a cell is
now escaped as \|, which GFM reads as | in a code span too. In Djot output a pipe in a
link or an image in a cell is now escaped as \|, and a pipe in a verbatim span stays bare,
because Djot does not split a row there.
An empty task item lost its checkbox on GitHub. <ul><li><input type="checkbox"></li></ul>
gave - [ ]. GFM reads a checkbox only when content follows it, so cmark-gfm and markdown-it
showed the text [ ]. It now gives - [ ]  , a checkbox: the character reference renders
as a space. Djot output does not change.
Text next to a paragraph, heading or other block in a table cell joined it (#645).
<table><tr><td><p>a</p>b</td></tr></table> gave | ab |, and so did a heading, a <div>,
a list or a code block before the text. Text before a heading or a code block joined it too.
A block in a cell is now separated from the cell content before and after it by the cell break:
| a b |, or | a<br>b | with br_in_tables on. A block at the start of bold, a code span or
another inline element still joins the text before that element, and a block inside a heading
still joins the text next to it.
The two converters wrote a quote or a paragraph in a table cell differently (#647). A quote
in a cell now has no > marker in either converter, as a heading, a list and a code block in a
cell have none: <blockquote>a</blockquote>b gives | a b |. The fast converter wrote a
paragraph after other cell content as <br> with br_in_tables off; it now writes a space, as
the full converter does.
The fast converter dropped a line break that no element encloses (#679).
a<br>1) t gave a1) t, so the two lines joined. The fast converter now writes the break as
the full converter does, a \n1\) t, the same as when the input sits in <body>. The line
after a hard break also no longer keeps a leading space: <div>a<br> b</div> gives a \nb in
both converters.
Hard breaks that differed between the converters or lost their place.
A whitespace character reference after a break (a<br> b) no longer makes a paragraph
break in the fast converter. A break in a heading followed by a space
gives one space (## a b). With the backslash newline style, a line in a list item that holds
only a break keeps the item's indent (#681). Wrap no longer cuts a list item's text at a ===
line left of the item's column, which dropped the hard breaks after it (#680), and keeps a
hard break right before such a line.
Two ordered lists next to each other became one list (#666).
<ol><li>a</li></ol><ol><li>b</li></ol> gave 1. a\n\n1. b, which CommonMark and Djot read
as one list, because a blank line does not end a list. An ordered list that follows an ordered
list with only blank lines between them now writes the other delimiter, so it gives
1. a\n\n1) b, also when a section, article, figure or similar element wraps either list. A
list written as text, in a heading or between inline markers, keeps ..
In Djot output a nested list directly under its item's text was text (#670).
Djot needs a blank line before a list that follows text, so - a\n * b is one paragraph
there. A nested list after its item's text now starts after a blank line in Djot output:
- a\n\n * b.
An empty nested list item after text turned the text into a heading (#667).
<ol><li>a<ul><li></li></ul></li></ol> gave 1. a\n -\n: an empty item cannot interrupt a
paragraph, so the lone - read as a heading underline and the item was lost. A nested list whose
first item writes nothing on its marker line now starts after a blank line: 1. a\n\n -\n.
An empty item after text directly inside its list does the same: <ul>t<li></li></ul> gives
t\n\n-\n, not t\n-\n. A list between inline markers, such as a highlight, is text and
does not change. With custom bullets such as -, nested single-item lists that end in an
empty item wrote - - -, a thematic break. The empty item's marker now starts the next line,
at the column it had on the line of markers: - -\n -.
A task item in an ordered list lost its number (#659).
<ol><li><input type="checkbox">p</li></ol> gave - [ ] p, a bullet list. A task item now
writes the marker of its own list, so it gives 1. [ ] p, and its content column follows the
width of that number. Djot output keeps - [ ] p, because Djot has task items only in bullet
lists.
A nested ordered list that does not start at 1 joined the text before it (#662).
CommonMark lets only a list that starts at 1 interrupt a paragraph, so in
<ol start="10"><li>a<ol start="100"><li>q</li></ol></li></ol> the line 100. q was part of
the paragraph a. Such a list now starts after a blank line when a paragraph is open before
it: 10. a\n\n 100. q. The blank line makes the outer list loose.
Text after a line break became a list, quote or heading (#651). <p>a<br>1) t</p> gave
a \n1) t, so 1) t became a list item; in bold, in a list item, in a quote and at the top
level alike. Text that starts the line after a hard break and would interrupt the paragraph
(1., 1), 01., -, +, *, >, #, a rule, a fence or an underline) now has that
character escaped, 1\) t, in both converters. A line that cannot interrupt a paragraph, like
2. t, is left as it is.
Heading text that reads as a closing # or a link definition lost the heading (#661).
<h1>#</h1> gave # #, an empty heading, because CommonMark reads a # run at the end of the
line as the heading's closing sequence; a # and ## lost their # the same way. Such a run
now has its first # escaped, # \#, in both converters. With heading_style: underlined,
<h1>[a]: b</h1> gave [a]: b over its underline, which is a link reference definition, so
the heading was lost. Text that starts a definition now has its [ escaped, \[a]: b. A Djot
heading and a closed ATX heading keep their # run as it is, and plain heading text is
unchanged.
A table that starts a task item became the task's text (#630).
<ul><li><input type="checkbox"><table><tr><td>c</td></tr></table></li></ul> gave
- [ ] | c |, so the header row was the task's text and the table was lost. A table that is a
task item's first content now starts after a blank line at the content column, also inside a
<div> or <section>. A block in an element that writes nothing before it, like a <span> or
an <hgroup>, now also starts below the checkbox; in bold or a link it stays text on the
checkbox line.
The fast converter dropped a task item's checkbox (#632). With the fast converter forced,
<ul><li><input type="checkbox"><p>p</p></li></ul> gave - p. A list item that holds a
checkbox now goes to the full converter, which writes - [ ] p. The full converter also reads
type="CHECKBOX" as a checkbox now, as a browser does.
Content after an HTML block that preserve_tags keeps joined the block (#655).
CommonMark ends an HTML block only at a blank line, and none followed a preserved block
element, so in <ul><li><input type="checkbox"><div></div><h2>h</h2>t</li></ul> the heading
and the text were part of the HTML. A blank line now follows a preserved element that starts
an HTML block, in a list item and at the top level:
- [ ]  \n <div></div>\n\n ## h\n t.
A table whose cells held only rules was dropped (#628). The full converter took a table
with no text and no image for a blank spacer table and wrote nothing, so
<table><tr><td><ul><li><hr></li></ul></td></tr></table> gave an empty document. A rule now
counts as content, and both converters write the table as | --- | over its delimiter row. Text
in a cell that looks like a list marker no longer turns a rule after it into ___, and the fast
converter now trims the space before a rule in a cell as the full converter does.
A table whose only cell held a line break was dropped with br_in_tables on (#646).
<table><tr><td><br></td></tr></table> gave an empty document from the full converter while
the fast converter kept the table as | <br> |. With br_in_tables on, a <br> in a cell
writes a literal <br>, so it now counts as content, matching the fast converter; with the
option off it still collapses to a space the cell trims away, so the table is still dropped
there, as it was before.
A rule inside bold, italic, a summary or a caption split the emphasis markers (#603). The
rule was written after a blank line, which ended the paragraph between the markers, so the
markers showed as literal text: <summary>t<hr></summary> gave **t\n\n---**, and a
definition list that starts its definition with a rule did the same inside a caption, <b> or
<em>. Markdown has no rule inside emphasis, so the rule is now the text --- in the running
line, as it already was in a link: **t ---**. This covers <b>, <strong>, <em>, <i>,
<summary>, <figcaption>, a table caption, <del>, <s>, <strike>, <ins>, <mark>,
<var>, <dfn>, <q>, and <sub> and <sup> when they write a symbol.
A list item took the checkbox of an item in its nested list (#604).
<ul><li>X<ul><li><input type="checkbox" checked> A</li></ul>ZZ</li></ul> gave - [x] X AZZ:
the outer item became a task item and the nested list was joined into its text. A checkbox in a
nested list now belongs to that list's item, so the output is - X, the nested - [x] A and
then ZZ.
Text after a quote, in a list inside bold, italic or <q>, rendered inside the quote
(#615). <b><ul><li>x<blockquote>q</blockquote>t</li></ul></b> gave **- x\n > q\n t**,
so t** continued the quote's paragraph. The first item's marker follows the opening marker,
so that item is text and writes no content column, and the blank line after the quote went
with the column. A nested item's marker starts its own line, so it is a real list item, and
**- a\n * x\n > q\n t** kept t in the quote too. Text after a quote or a nested
list now starts after a blank line, at the column of the innermost real list item:
**- x\n > q\n\nt** and **- a\n * x\n > q\n\n t**. A quote 4 or more columns
past that item's column, or past the start of the line when no item is real, is the
paragraph's own text, so nothing changes there and the markers stay in one paragraph.
A quote that starts a list item rendered outside the item (#617).
<ul><li><blockquote>q</blockquote></li></ul> gave -\n> q, an empty item and a quote after
the list, also when the list is inside a quote. The quote now starts on the marker line,
- > q.
A quote that starts a task item became the task's text (#622). A checked task item whose
first content is <blockquote>q</blockquote> gave - [x] > q, which renders the text
[x] > q. The quote now starts on the next line at the item's content column,
- [x]  \n > q. The checkbox line ends in a space written as a character reference: GFM
reads a checkbox only when content follows it, so cmark-gfm shows a bare [x] line as text.
A nested list, a heading or a code block that starts a task item does the same, also inside a
<div> or a <section>, and also when the checkbox is in a <p> of its own. A rule there
gave - [ ] ---, the text [ ] ---; it now comes after a blank line,
- [ ]  \n\n ---, since --- right under the checkbox line makes that line a heading.
A rule that starts a list item rendered outside the item (#623).
<ul><li><hr></li></ul> gave -\n\n---, an empty item and a rule after the list. The rule is
now ___ on the marker line, - ___: - --- is a rule of its own.
Text inside a list before an item joined the item's marker (#625). <ul>how<li>do</li></ul>
gave how- do, one line of text, and the item was lost. The item now starts its own line,
how\n- do. When its marker cannot interrupt the text above it, as with 3., a blank line
comes first: how\n\n3. do.
Text after a quote stayed in the quote for a later item of a list inside bold or italic
(#633). <b><ul><li>a<ol><li>x</li><li>y<blockquote>q</blockquote>t</li></ol></li></ul></b>
gave **- a\n 1. x\n 2. y\n > q\n t**, and t** continued the quote. The marker
2. cannot interrupt a paragraph, but after the real item 1. x no paragraph is open, so 2.
starts an item. The same holds after a quote of the enclosing item. Text
after the quote now starts after a blank line: **- a\n 1. x\n 2. y\n > q\n\n t**.
Some first blocks of a task item still became the task's text (#634). An ordered list that
starts at a number other than 1 and a code block in the indented style cannot interrupt the
checkbox line, so they now start after a blank line: - [ ]  \n\n 3. x and
- [ ]  \n\n c.
The code block keeps its indent; before, - [ ]\n c lost the code. A quote after an empty
inline element, such as <span></span>, now starts on the next line like a quote right after
the checkbox. An image before the quote is still text on the checkbox line, and so is an empty
element that preserve_tags writes as HTML. An empty list, heading or code block writes
nothing, so text after it stays on the checkbox line: - [ ] t. A fenced code block that holds
only whitespace still writes its fences, so it starts on the next line and the text after it
stays out of the code.
A quote after a line break in a task item stayed on the checkbox line (#650).
<ul><li><input type="checkbox"><br><blockquote>q</blockquote></li></ul> gave - [ ] > q, and
the quote was text. The converter's own output now decides which element writes first, also
inside a <div> or a <section>. A line break, a , a <template>, a <noscript>, an
<input> that is not the checkbox, an empty <picture> or an image with skip_images writes
nothing there, so the quote now starts on the next line: - [ ]  \n > q.
An underlined heading ended its list item (#635). With heading_style set to underlined,
the underline of a heading in a list item was written at column 0: <ul><li><h2>q</h2></li></ul>
gave - q\n-, an item and an empty item. The underline now gets the item's content column,
- q\n --. After a line of the item, the heading starts after a blank line, so it does not
continue that line's paragraph: - a\n\n q\n --. A block after a heading of one letter starts
its own paragraph: <ul><li><h2>q</h2><p>t</p></li></ul> gives - q\n --\n\n t. In a list
item the underline has at least two dashes, since a lone - line reads as an empty item.
Many blocks in one list item took quadratic time (#649). Each block checked every earlier
line of the item to see whether the item was still open, and each item of a list inside bold or
italic checked every earlier item. <ul><li> with 5000 headings after text took seconds. Each
check now reads only the lines written since the last one, so the time grows linearly.
A heading after text in a quote joined the text (#640). <blockquote>a<h2>q</h2></blockquote>
gave > a## q, a paragraph, and with heading_style set to underlined it gave > aq\n> -,
one heading. The heading now starts after a blank quote line: > a\n>\n> ## q. An underlined
heading does the same after any line of text in the quote, such as a line that ends in a hard
break or the last item of a list.
An underlined heading whose text starts a block lost its heading (#653). With
heading_style set to underlined, <h2>-</h2> gave -\n-, two empty list items. The text of
an underlined heading is now escaped where it would start a block: a list marker (-, *,
1., also on its own), a quote, a # heading or a rule. <h2>-</h2> gives \-\n-.
The lines of a list item in a quote or under tab indent left the item (#654). A list inside
a quote in a list item counted the markers outside the quote too, so its lines sat further in
than the item, and the underline of a heading sat short of it. The quote now starts its content
as a container of its own, so a list in it counts only its own markers. A block after text in a
list item inside a quote now also starts its own line, as it does outside a quote. An
underlined heading in a quote in a list item now gets the one-dash underline it gets in a
quote elsewhere: <ul><li><blockquote><h2>q</h2></blockquote></li></ul> gives - > q\n > -,
where it gave - > q\n > --. Bold or italic around the quote no longer changes the lists in
it, so text after a nested quote in such a list leaves the nested quote. With
list_indent_type set to tabs, the lines of a nested item were one tab short of its content
column: - a\n\t* q\n\n\tt put t in the outer item. They now reach the column where the
item's text starts, - a\n\t* q\n\n\t\tt. With list_indent_width set to 4 or with tab
indent, a quote right after an opening bold marker, a summary's or a caption's, now writes a
nested quote or list of its list at the column of the nearest real list item, where they became
a code block. A list inside <mark> or <del> is now text after its marker, as inside bold,
so text after a quote in it no longer becomes a code block.
Wrap mode joined a rule or a heading underline to the text next to it (#607). With wrap
on, a --- line followed by text became one line of text, --- B, and the rule was lost. The
underline of an underlined heading was joined to the heading text (Heading -------), or cut
off from it by a blank line for =======, so the heading was lost too. A rule now stays on its
own line, and an underline stays right under its heading text, which is not reflowed, as with a
# heading. Both hold inside a quote too.
Wrap mode joined the keys of the frontmatter into one line. With wrap on and metadata
extraction on, the YAML frontmatter went through the reflow like body text, so
---\ntitle: My Page\n--- became --- title: My Page --- and the frontmatter was lost. Only
the text after the frontmatter is wrapped now.
Wrap mode folded list items, code fences, headings and table rows inside a quote into text.
With wrap on, <blockquote><ul><li>alpha</li><li>beta</li></ul></blockquote> gave
> - alpha - beta, one item, and a code block inside a quote lost its code. Outside a quote, a
~~~ code block was reflowed like a paragraph. Every line that starts a block now keeps its own
line in and out of a quote, and a code block ends only at a fence that closes it.
Wrap mode turned a tight list loose when an item went on to a second line (#616). A line
right under a list item that starts no block of its own, at column 0 or indented, belongs to
the item's text. With wrap on, the reflow wrote it as a paragraph of its own and put a blank
line before the next item, so the list rendered loose. It now joins the item's text and is
wrapped with it. For the same reason, a line that starts with a number such as 1990. or
57) stays in its paragraph, in a quote and in a list item: only a bullet or a number equal
to 1, such as 1. or 01), can end a paragraph and start a list. Under a list item, a number
line left of the item's text still ends the item.
Wrap mode dropped a hard line break (#613). With wrap on, <p>a<br>b</p> gave a b:
the reflow joined the line after a <br> to the line before it, in a paragraph, a quote and a
list item. A hard break is now a line end the reflow never joins across, so each side of it is
wrapped on its own and the break stays, with both newline styles.
Wrap mode could start a line with a list marker and turn text into a list (#614). With
wrap on, a break before a -, 1., # or > in running text started a new line with it,
which opened a list, a heading or a quote. A wrapped line now never starts with a word that
opens a block there, also in a run such as --- --- --- or * * *; the word stays at the
end of the line before it, which can then run past the wrap width. A number equal to 1 with
leading zeros, such as 01. or 001), opens a list like 1. does, so it is kept off a line
start too, and a link label or image alt line that starts with one is escaped. A number
followed by non-breaking spaces, such as Word's 1. Cut, is no longer read as a
list marker, and the reflow no longer breaks a line at a non-breaking space.
Wrap mode broke a link whose address holds a space. An address with a space is written in
angle brackets, [Share](<https://example.com/?text=a b>), and a line end inside the brackets
ends the link. With wrap on, the reflow broke the line there. It now keeps the address in
angle brackets on one line.
Wrap mode cut a nested list marker off from its text. A list item that holds a nested list
on its own line, such as - 3. [vote](...) title, wrapped to - 3. and the text on the next
line. - 3. alone is an empty nested item, and the text below it left the nested list. The
reflow now treats both markers as one, so the text stays in the nested item. An item whose
text starts a heading or a code fence, such as - ## Title, is no longer reflowed: the heading
kept only its first words, and the code lines were joined and wrapped like text.
A page whose bytes open with a mangled byte order mark lost its whole head. A real leading
U+FEFF is stripped before parsing, but one a wrong encoding guess mangles beyond recognition
reads as ordinary text by the time #527's head search sees it, and that search treated any such
text as the start of the body, so it gave up before it ever reached <head> and reported no
title, no meta tags and no base or canonical link. A browser discards anything ahead of the
document's own <html> tag without letting it block the real head inside, so the head search
now does the same: text directly in front of <html> no longer ends the search, on both tiers.
A head-only fragment with no <html> tag is unaffected: text ahead of <head> there still ends
the search, as #527 intended.
An <img> whose src spelled the data: scheme in upper or mixed case ignored its lazy-load
address. The check that sends a data: source to the data-src, data-lazy-src,
data-original and srcset fallbacks compared the scheme in lower case only, so
<img src="DATA:..." data-src="real.png"> kept the payload while the same image with data:
used real.png. The scheme now matches in any case.
An image whose data: scheme was written in upper or mixed case was not extracted, and the
metadata reported it as a relative image. Inline image extraction and the metadata image type
compared the scheme in lower case only. Every check now uses the one case-insensitive test the
converter uses for its markdown output.
Text right after a list or a table rendered inside it (#570, #571). Inline content that
directly followed a list or a table in the same container continued the block's last line.
After a list it became a lazy continuation of the last item, so
<div><ul><li>A</li></ul>para</div> rendered para inside the item; after a table it became
one more table row. Inline content after a block now starts its own paragraph after a blank
line, the same as text after a paragraph, in both rendering paths. This includes a list item
that ends in a line break, and a <br> right after a list with backslash line breaks, which
also put the text inside the item. Text after a horizontal rule gets the same blank line.
A block inside a list item, and the text after it, left the item (#583). A heading after
the item's text joined that text (<li>A<h3>H</h3>tail</li> gave - A### H), a rule ended the
list, and the text after a div, table, blockquote, nested list or definition list became a lazy
continuation of that block. The text after a code block or paragraph, a definition list, and
a section, article, header, footer, aside or main element fell out of the list. Inside a list
item, a block after other content of the item now starts on its own line at the item's content
column, and text after a block starts its own paragraph at that column. A paragraph or div
after the item's text is now its own paragraph instead of joining the text, which makes the
list loose. The column is written only while the item is still open, also inside a
definition list or a section that the item holds. Whitespace between the marker and the
item's first block no longer counts as content in strict whitespace mode, and with wrap a
paragraph inside an item keeps its indent, also inside a blockquote. The fast conversion
path hands these items to the full converter.
Text after a list or table at the end of an inline wrapper continued it (#585). In
<div><span><ul><li>A</li></ul></span>para</div>, para still continued the list's last item,
because the rule from #570 looked only at the element right before the text. Text after an
inline element whose last content is a block now starts its own paragraph too, in both
rendering paths.
A horizontal rule right after a line of text turned the text into a heading (#584). Inside
a paragraph, and at the start of a definition, the rule was written on the line right after
the text, so <p>t<hr>B</p> gave t\n--- and <dl><dt>t</dt><dd><hr></dd></dl> gave the same.
Markdown reads --- under text as a heading underline, so t became a heading and the rule
was lost. The rule now starts after a blank line there too, as it already did after text in a
<div>.
The nightly benchmark guardrail scored timings on hardware it was never calibrated on. The
runner pool moved from the AMD EPYC 9V74 the baseline was calibrated on to an EPYC 7763, and
every fixture read 15% to 45% slower. htmbench compare still scored each timing, printed 26
FAIL lines, and then passed the run as advisory, so the job was green while measuring nothing.
On a CPU other than the calibrated one no timing is scored now: the run reports
TIMINGS NOT SCORED with both CPUs and fails, and under --allow-host-mismatch (the nightly
job) it succeeds with a Benchmark timings not scored warning on the run page instead. The
fixture inventory stays fatal on any host, and a regression on the calibrated CPU still fails.
The benchmark baseline recorded stale output sizes for five fixtures. The 3.14.2
conversion fixes moved the Markdown output of gh-121-hacker-news, gh-127-issue,
gh-190/firsteigen, gh-190/rbloggers and wikipedia/small_html, and the baseline was never
updated. Each change was traced to the fix that made it and reviewed: images kept in layout rows
(5b26d732d, 31f2015b1), and whitespace no longer opening a line (c5b8d1baa), which also stops
two lines rendering as indented code blocks. Only output_bytes changes; the calibrated timings
stay as measured. gh-190/plusblog changed for a different reason, fixed below.
An <img> with no usable src could take its address from the middle of a srcset
candidate. The fallback split srcset and data-srcset on every comma, so a comma inside a
parenthesised descriptor or inside a URL started a new candidate: a.png (x, b.png 3x ), c.png 2x
produced b.png, and a.png?w=1,2 2x produced 2. Candidates are now split with the HTML
spec's srcset parsing steps, including the parentheses rule, and only the spec's five ASCII
whitespace characters separate a URL from its descriptor.
The srcset fallback could choose a candidate a browser never loads, or compare a width with a
density. Descriptors were read with Rust's float parsing and nothing else, so a.png foo stayed
eligible, a.png infx beat every other candidate, a first NaNx candidate could not be beaten,
+2x counted as 2x, and 900x beat 800w. Descriptors now go through the spec's descriptor parser: a candidate with an
unknown token, a number outside the spec's grammar, a zero width, a second descriptor of one kind,
or an h without a w is dropped, and a list with no valid candidate keeps src. Widths and
densities are not compared with each other, because a browser needs sizes and the viewport to
do that: when any candidate has a width, the largest width wins, and otherwise the largest density
wins, a candidate with no descriptor counting as 1x.
A <br> in a <span> after a list or a layout table pulled the next paragraph into the last
list item. A <span> removed the line break that ends the item's line, so the <br> became a
hard break at the end of the item and the paragraph after it rendered inside the item. Without a
<br>, the span's text was joined onto the item's last word, and a <span> right after a
horizontal rule was joined onto the ---. A <span> now leaves the line break before it in
place (#546).
Head metadata kept its character references encoded. The <base href>, the
<link rel="canonical"> href and the <title> text reached the frontmatter and the structured
metadata as written, so <base href="https://example.com/it's/"> produced
base: https://example.com/it's/. They now go through the same decoder as <meta content>
and body text, including the Windows-1252 mapping for numeric references 128-159.
A newline in a head value started a new frontmatter key. The frontmatter wrote each
key: value line as it was, so a <title> or <meta content> holding a newline, literal or
written as , could add or override a key. Each key and value is now one YAML scalar: a
value that plain YAML would misread (a newline, : , a space before #, a leading -, # or
@, a control character) is written in double quotes with YAML escapes. Other values stay
unquoted.
Legacy named references without a semicolon were not decoded. The spec lets about a hundred
names such as ©, & and é close without ;, and browsers decode them in text:
© 2024 is © 2024. The converter kept them as written. They now decode on both tiers,
with the spec's longest-name rule (¬it; is ¬it;). In an attribute value a legacy name
followed by = or a letter or digit stays as written, so ?a=1©=2 in an href is unchanged.
Frontmatter values that YAML reads as numbers, booleans, null or dates were not quoted. A
value such as 3, true, null or 2024-01-01 is a valid plain scalar, so a YAML reader
turned meta-algolia-search-order: 3 into the number 3. A value that the YAML 1.2 core schema
or a YAML 1.1 reader resolves to anything other than a string is now written in double quotes.
Numeric character references without a semicolon were not decoded. The spec decodes '
and ' without their ;, in text and in attribute values, so it's is it's in a
browser. The converter kept the reference as written. It now decodes on both tiers.
Tier 1 kept a reference encoded after an unknown name. When an unknown name such as &foo
had a ; a few bytes later, Tier 1 wrote the whole span as it was, so &foo & kept
& where Tier 2 wrote &. Tier 1 now writes the & alone and reads on, as Tier 2 does.
Tier 1's fallback message added a ; the page did not have. When Tier 1 handed a reference
without its ; to Tier 2, the log message showed ' for an input of '. The message
now shows the reference as written (#565).
Tier 1's fallback message called a known reference unknown. When Tier 1 handed a reference
without its ; to Tier 2, the log message called it an unknown HTML entity even when the
reference was one Tier 2's decoder knows, such as ' or ©. The message now says the
reference is missing its ; when the name is known, and keeps the unknown wording for names
that really are unknown (#586).
base_url could pick a <base href> that a browser ignores. The document base came from a
byte scan for the first <base tag in the source, so a <base> inside a comment, inside
<title>, <textarea>, <script>, <style> or another raw-text element, inside <template>
or SVG, or in a body that a <frameset> replaces, set the base for every relative link. The base
now comes from the first <base> with an href in tree order, read from an html5ever parse of
the document. The parse stops at the first <base href> once no later markup can place a node
in front of it or remove it, in <head> or in the body, and a page without a <base tag is not
parsed at all. Found while adopting base_url downstream, where the same two mistakes had
already been fixed once in a link pre-pass.
A data: or javascript: <base href> became the base for relative links. A browser
ignores such a base and resolves against the page's own URL, as the HTML "frozen base URL"
steps require. base_url now does the same, so relative links on such a page resolve against
the caller's base_url instead of failing to resolve.
The base and canonical metadata kept the last tag, not the first. The head extractor
overwrote each value when it met another tag, so the reported base could differ from the
base that base_url resolves against. The base metadata (and base_href in the document
metadata) is now the same first <base href> in tree order that the document base uses, read
once, and the first <link rel="canonical"> wins.
The <meta> metadata kept the last tag with a given name. Each <meta name> or
<meta property> overwrote the value an earlier tag with the same key had stored. The first
tag per key now wins on both tiers, as the base and canonical metadata do.
A page without a <head> tag reported no base metadata. The parser creates the head
itself, so <base href="https://a.example/"><p>x</p> sets the document base, but the head
extractor looked for a <head> tag in the source and found none. The base metadata now comes
from the same parsed document as the document base on both tiers, with or without a <head>
tag.
A > inside a <base> attribute value turned off the early stop of the base parse. The
parse is fed in pieces, and a piece could end inside the quoted value, so the <base> tag only
completed in the next piece, which was never checked. The parse now notes each <base href>
element when html5ever's tree builder creates it and checks only that element's ancestors, so
the parse stops wherever the tag bytes fall, and the cost of the check does not grow with the
size of the tree. The document base itself was always correct.
The title metadata kept the last <title>, not the first. A head with the titles First
and Second reported Second, while a browser shows First. The first title now wins on both
tiers, as the base, canonical and <meta> metadata already do. An empty first title also
wins, as in a browser, so the page reports no title.
<meta> names that differ only in letter case let the last tag win. Meta names do not
depend on letter case, but <meta name="Description"> followed by <meta name="description">
gave the second value in the document metadata, and the frontmatter printed both. The names are
now compared in any letter case, and the first tag wins in the frontmatter and the document
metadata.
A <head> tag inside the body gave the two tiers different metadata. The parser ignores a
<head> tag once the body has started. Tier 1 read such a stray head when the page had no head
of its own, and Tier 2 read it when the real head was empty. Both tiers now read only the first
head before the body. The body starts at a <body> tag, at text, or at any tag other than the
ones a head can hold, so <p>x</p><head><title>Stray</title></head> has no title. The
document structure gives a metadata block only for the head the metadata reads. A leading
UTF-8 byte order mark is dropped first, as a browser's decoder drops it, so it does not start
the body and no longer appears at the start of the output.
A <meta> tag named title, base or canonical replaced the document's title, base and
canonical link. <meta name="base" content="/meta/"> next to <base href="/real/"> set
base_href in the document metadata to /meta/, while the links resolved against /real/, and
<meta name="title"> replaced the <title> text. The base_href and canonical_url fields
now come only from the <base> element and <link rel="canonical">, and such a meta tag is an
ordinary entry in meta_tags. A <meta name="title"> still gives the title of a page without
a <title> element, but it no longer replaces the text of one.
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.15.1 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.15.1")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 71d0dd05b010b8790070fcecf28a136a359e32c38cca542dcb2c7f59febf8b2f
base_url did not resolve a <blockquote cite>, and Tier 1 dropped the citation entirely.
cite is a destination the converter renders — Tier 2 emits it as a trailing — <url> line —
but it was the one such destination base_url left alone, so a relative citation stayed
relative. Tier 1's blockquote handling read no attributes at all, so the same markup produced
the citation on one tier and nothing on the other; a cited blockquote now bails to Tier 2,
which is authoritative for it. Found while adopting base_url downstream, where a converter
that resolves every destination except this one forces the caller to keep a whole
link-rewriting pre-pass alive for it.
The Go binding could read a different OS thread's FFI error slot. Convert,
HeaderMetadata.IsValid and the visitor entry point now pin the goroutine for the duration of
the cgo call with runtime.LockOSThread. The FFI layer keeps its last-error state in a
thread_local! (crates/html-to-markdown-ffi/src/lib.rs:31), so without pinning the Go runtime
was free to reschedule the goroutine between the call that stamped the error and the
htm_last_error_code read that reports it — surfacing a nil error for a call that had in fact
failed. Emitted by alef 0.97.0; no Go API changed.
alef generator to 0.97.0 and regenerated every binding. Apart from the Go thread
pinning above, the only other generated change is the Kotlin Android Gradle wrapper moving from
9.7.1 to 9.8.0; everything else in the regeneration is version strings and provenance hashes.docs-site changelog mirror against CHANGELOG.md. The two are compared from
the first ## [ heading onward rather than over [Unreleased] alone, because this repo releases
straight out of [Unreleased] and leaves it empty on main — an Unreleased-only comparison
would compare zero lines and pass while a released section drifted.. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.15.0 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.15.0")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 7fbd75de8ea930de4800ca081c7fb64f221a3b15085ee131320c05f145065cdc
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{
.url = "https://github.com/xberg-io/html-to-markdown/releases/download/v3.15.0/html-to-markdown-rs-zig-v3.15.0.tar.gz",
.hash = "html_to_markdown_rs-3.15.0-QtXyW7WtswG1ksUAa5-D_8DD0NAHbFrsS3xhg1ym5HOB",
},
},
ConversionOptions::base_url resolves relative href/src destinations against a
caller-supplied base URL, honoring a document's own <base href> the way a browser does.
Defaults to None, so output is byte-identical for callers who do not set it. Resolution
happens identically in both the Tier 1 and Tier 2 rendering paths.rmcp to 3.4.1.alef generator to 0.96.5 and regenerated every binding. Generator drift
only; no public API of any binding changed.converter/utility/content.rs and converter/inline/link.rs (each over the
1000-line quality gate) into smaller modules, and reduced convert_table_row's cyclomatic
complexity by extracting its visitor-hook pre-pass into a separate function. No behavior
change; every existing import path is preserved via re-exports.The R binding compiles again. ConversionOptions::base_url is the struct's first
Option<String>, and alef's extendr backend assigned it a bare String, failing the whole R
package build with error[E0308]: expected Option<String>, found String. Fixed in alef 0.96.5.
It went unnoticed for two commits because two independent mechanisms each suppressed the leg.
On the commit that introduced the line, Test: R was cancelled: ci-e2e.yaml's
cancel-in-progress: true group is keyed on the branch, so the next push to main killed the
run before R finished. On the commit after that, Test: R was skipped: the leg is gated on
a packages/r/**/crates/html-to-markdown/** paths filter that commit did not match. Neither
state is a failure, so CI E2E reported success twice without ever compiling R.
CI's fixture-snippet validation now really validates the 341 Swift snippets. They had all
been reporting Unavailable with no such module 'HtmlToMarkdown': Swift 6.3 made
swiftbuild the default build system, whose bin path holds the .swiftmodule files
directly and emits no Modules/ directory, while alef points -I at <bin-path>/Modules.
The session now reconstructs that layout after building. The gate went red without any
change to this tree, when the runner's preinstalled Swift moved to 6.4.
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.14.3 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.14.3")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 4730d1e764c5a05be95af2a0e0e2a95a25cf5a881e6849702ae2442e2cda08ff
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{
.url = "https://github.com/xberg-io/html-to-markdown/releases/download/v3.14.3/html-to-markdown-rs-zig-v3.14.3.tar.gz",
.hash = "html_to_markdown_rs-3.14.3-QtXyW5ksoQFf0DSqr1e_aA0CRl9vVlYUILb_HMSz-muR",
},
},
<h2><strong>Alpha</strong><strong><em><br></em></strong>Beta</h2> rendered ## **Alpha**Beta
from 3.14.2 on, where 3.14.1 kept the space, and a paragraph or table cell joined the words the
same way. Issue #501 taught the shared wrapper emitter that a whitespace-only body at a line
start contributes nothing, keyed on the destination buffer being empty -- but the destination
can also be an enclosing wrapper's fresh scratch buffer, which is empty mid-line, so the inner
<em>'s one space was dropped there and the outer <strong> came out empty. The line-start
rule now fires only on the block's own buffer, using the address test the text-node fallback
already uses for the same distinction, so <mark>, <ins>, <del>, <sub> and <sup>
wrappers move with it. A <br> inside a wrapper inside a table cell had the same shape on its
own since before 3.14.2 -- <td>Alpha<em><br></em>Beta</td> -- and now keeps its space too; a
cell's own leading <br> still contributes nothing. Tier 1 bails on adjacent emphasis, so the
change is Tier-2 only.<p><i>Alpha</i><span><span>\n</span></span>Beta</p> rendered *Alpha*Beta where a browser
shows a space. Issues #430 and #491 taught the text-node fallback that a lone newline inside
an inline wrapper separates words when the wrapper is followed by inline content, but the check
looked one level up only: with a second wrapper the inner <span> is the last child of the
outer one and the newline was dropped. The check now climbs through every transparent inline
ancestor that has nothing after it and stops at the first block. Tier 1 already emitted the
space, so this was a live cross-tier divergence; the tiers now agree.<a href="/o"><div><a href="/i">Inner</a></div></a> rendered [](/o) and then
[](/o)[Inner](/i): the repair legitimately closes the outer <a> at the <div> and
reconstructs it inside, and the clone reached the renderer indistinguishable from an authored
element. A renderer rule keyed on shape would either drop a genuine empty anchor (CommonMark
example 484) or a deliberately authored duplicate, so the fix is at parse time: every <a>
start tag is stamped with a private origin id before the tree builder sees it, the clones
inherit it, and on the repaired tree the halves of a split anchor with no content of their own
are unwrapped in place. When no half has content the authored one is kept, so the destination
still appears once as [](/o); <a href="/o"><div>Text<a href="/i">Inner</a></div></a> keeps
the half that carries Text and renders [Text](/o)[Inner](/i). Input the repair never runs
on, and an anchor the repair leaves whole, are unchanged.. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.14.2 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.14.2")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: fffa609814de9160dd41b5baec568ca6e6ba64ec651a61c2185d23297429d6cd
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.14.2/html-to-markdown-rs-zig-v3.14.2.tar.gz\",\n .hash = \"html_to_markdown_rs-3.14.2-QtXyW-msoAHzb7cg9XA_Z65GtZlDZZdqN03SQJQ_ZCRh\",\n },\n},\n```\n
& that would read as a character reference in a link destination or title now stays
literal (#498). CommonMark
decodes entity and numeric character references inside destinations and titles, so the
decoded ?+ that <img src="?&plus;"> produced from 3.14.0 on re-parsed as ?+ --
a different URL. <a href> and title had carried the same defect since 3.12, when they
started decoding; only src was new to it. Such an & is now written back as &, the
form 3.13 emitted and the one non-CommonMark consumers of the URL also read correctly. Only a
reference the HTML5 decoder really recognises is touched: ?a&b and &foo; are unchanged.
Tier 1 wrote a link's href straight into the output with no escaping at all, so it now
shares Tier 2's append_url_destination; that also ends a divergence unrelated to entities,
where Tier 1 emitted [T](/a b(c) and [T](/a(b) against Tier 2's [T](</a b(c>) and
[T](/a\(b).[ in a link label or image alt is now escaped
(#499). CommonMark accepts a
bracket in a label only escaped or as a matched pair; an unmatched ] was already escaped,
an unmatched [ was not. <img alt="[A;B)" src="S"> rendered , which
re-parses as a dangling ! followed by a real [A;B)](S) link -- the image was gone. The
shared escape_link_label helper now escapes every opener left without a closer, so both
tiers and every caller (links, images, <graphic>, SVG titles, embedded media) move together.<p>P</p><div><span> </span><img alt="A" src="S"></div> rendered the image as
 indented by four columns, which CommonMark reads as an indented code block, so
the image was lost. Issue #460 dropped leading whitespace at the start of a
paragraph; a <div> never set that flag and the run fell through to the verbatim fallback.
Leading ASCII whitespace on a line is never Markdown content, so the text-node fallback now
drops it at a line start of the block's own buffer, the <div> handler measures its content
start the way <p> does so an inline wrapper's empty scratch buffer is not mistaken for one
(its single space must still reach the #481 handling), and a whitespace-only wrapper spliced
in at a line start contributes nothing -- <p>A</p><p><i> </i>B</p> no longer opens its
second paragraph with a stray space. A multi-space run mid-line collapses to one space, as it
already did between two inline siblings. A run carrying a decoded is untouched. Tier 1
already emitted every case this way.<b>Alpha</b><b><span>\n</span></b><b>Beta</b> rendered **AlphaBeta**, and
Alpha<i>\n</i>Beta rendered AlphaBeta, where a browser shows a space. Two drops of one
shape: a newline-only text node returned early because the wrapper's scratch buffer was empty
-- the buffer is fresh per wrapper, so its length says nothing about the document -- and a
<br> with nothing before it in that buffer left a bare newline that chomp_inline did not
count as a space. Both now surface as the single separating space issue #481 already gives a
literal <b> </b>, still suppressed after an existing space; a truly empty <b></b> still
emits nothing. Tier 1 bails on every one of these shapes, so the change is Tier-2 only.<table><tr><td>A</td><td>B</td></tr><tr><td>C</td></tr></table> rendered - A B / - C:
ragged row lengths alone classified a table as a layout table. A headerless table with a
short row is ordinary tabular data far more often than it is an email-signature grid, and the
regular renderer already pads a short row to the table's width (issue #13), so it now renders
| A | B |, | --- | --- |, | C | |. Layout still triggers on more than one nested table,
colspan/rowspan combined with border="0", a blank table, or a short table dense with
links. Tier 1 keeps its stricter bail on ragged rows -- it has no padding of its own -- which
only ever sends more input to the path that does. A layout row also keeps an image as
 instead of degrading it to alt text: the row is a list item, and list items and
data cells keep images by default; only headings degrade them. keepInlineImagesIn is no
longer needed for that (issue #433), though it still governs headings and links.<br> or a blank nested table. A <br> inside a
layout cell emitted a hard-break marker, and a nested table that rendered to nothing still
emitted the blank line meant to separate content; either put a bare newline inside a list
item's line and ended the item. Both were latent -- the Hacker News footer in the gh-121
fixture is <img><table>bar</table><br>links, and while the spacer image degraded to nothing
every separator stayed suppressed -- and keeping the image surfaced them. A <br> now follows
the settled cell rule its <div>/<p> continuations already use (issue #470), and an empty
table output writes no separator. Across the benchmark corpus this rejoined three split list
items in one fixture and changed nothing else; the other movements are images now kept in
layout rows and leading whitespace dropped at line starts (#501).[One](/one) and
[Two](/two) into text; only the outer destination survived. A layout cell is a list item's
text, not a link label, and it already holds a bare nested table on the lines after its
bullet, so the deferred table now lands there the same way. Headings and true inline labels
keep refusing. Tier 1 bails on a table opened inside a link, so this is Tier-2 only.. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.14.1 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.14.1")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 9c98097979c5f41dbc30f203458bffa6f7015623a511e929e43b41b905548684
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.14.1/html-to-markdown-rs-zig-v3.14.1.tar.gz\",\n .hash = \"html_to_markdown_rs-3.14.1-QtXyW4m6oAGO98CVG8cop9GlGXL7S_bSjcZTQjxF9anL\",\n },\n},\n```\n
A multi-line link label or image alt no longer has one of its lines read as block
structure (#496). CommonMark
parses block structure before inline structure (spec 0.31.2 appendix A), so a continuation
line that looks like a block opener ends the paragraph the label lives in and the [/![
never reaches its ].  produced no image at all -- it produced
<h2></p>, with the image gone. A hard line break did not help, because it
is inline and block parsing has already finished by then, so the reporter's escaping
workarounds could not work either. Measured against comrak, 14 of 18 opener shapes destroyed
the construct: setext - and =, all three thematic breaks, ATX headings, both code
fences, block quotes, all three bullet markers, both ordered-list delimiters, and HTML
blocks. The first non-blank character of a continuation line is now backslash-escaped when
that line would open a block that can interrupt a paragraph -- ordered lists at their
./) delimiter, since a digit cannot carry an escape. Lines that open nothing are
untouched: four columns of indent (indented code cannot interrupt a paragraph), a type-7
HTML block such as a bare <span>, 2. x, -x, seven #. The label's first line is never
escaped -- it is preceded on that same line by the caller's own [/B, losing the break
entirely. [A \n](H)B re-parses to exactly the original <a href="H">A<br /></a>B
(verified against comrak), so the break belongs inside the label. Both tiers dropped it for
the same reason expressed twice -- the whole label was whitespace-trimmed, and once
flattened a " \n" marker is indistinguishable from the incidental whitespace that really
does belong before a </a> -- and Tier 2 additionally never emitted a leading break at
all, because its "nothing on this line yet" test compares a fresh label buffer's length
against the enclosing block's start offset. A run of breaks at one edge still collapses to a
single break (two adjacent markers would put a blank line in the label, and a blank line
ends the paragraph, destroying the link), and a label of nothing but breaks still collapses
to empty. A heading and a pipe-table cell are single-line and still fold every break to a
space. The issue's second example asks for B[A \n](H), which moves the break to the far
side of the label text; the break is preserved where the <br> actually was instead --
B[ \nA](H), which comrak renders back to the input DOM.
Tier 1 no longer emits three spaces where Tier 2 emits one for a <br> inside a link inside
a table cell (| [A - B](H) | against | [A - B](H) |). Tier 1 folded its hard-break
marker late, in close_table_cell, which left the marker's two spaces behind; it now folds
in close_link, which is also what keeps the #496 escaping from firing on a label that is
about to become single-line anyway.
Generated Go and R e2e suites no longer assert against strings their fixtures never
specified (alef pin 0.90.0 to 0.91.5). Two independent generator defects were corrupting
fixture values on their way into test code: alef's Go emitter rendered a multi-line value as
a raw backtick literal and its own writer then trimmed the trailing whitespace off every
physical line, so the three assertions carrying a Markdown two-space hard break read
[Alpha\n](url)Beta where the fixture said [Alpha \n](url)Beta; and alef's R emitter ran
every plain string argument through the PascalCase-to-snake_case transform meant only for
enum wire values, so Alpha<span ...> was emitted as alpha<span ...> and
Beta<a ...><br>Alpha</a> as beta<a ...><br>_alpha</a>. The R defect had been silently
wrong since the 3.14.0 paragraph_whitespace_only_span_separates_words fixture landed; it
went unnoticed because the E2E workflow had skipped every language test job on the three
preceding commits, so "green" meant "nothing ran". Fixed upstream in alef 0.91.5 rather than
by trimming the fixtures. The pin bump also carries alef 0.91.0-0.91.3, which for this repo
is limited to a simpler argument-marshalling path in the Node visitor bridge (no API or
behaviour change) and dropping its unused tokio-util dependency.
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.14.0 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.14.0")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: fe726ca8efe0b85254e9bf71288dc63689a8b36d7ccf4340f8de9b23a937b1e0
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.14.0/html-to-markdown-rs-zig-v3.14.0.tar.gz\",\n .hash = \"html_to_markdown_rs-3.14.0-QtXyW3HsnwG1ntS4tMEUg15s715iH-Dhj46cz3pBoOio\",\n },\n},\n```\n
html5ever to 0.40.1, which fixes silent content loss on the HTML repair path.
0.40.0's serializer dropped the leading 0xC2 byte of a two-byte UTF-8 sequence, corrupting
every character in U+0080..U+00BF except NBSP -- §, ©, °, · among them. Because
repair_with_html5ever serializes the repaired tree back to a string for the primary parser to
re-read, a single such character made the whole re-parse fail and everything after it vanish.
This repo was pinned to 0.39.0 to avoid it; 0.40.1 carries the upstream fix, verified here by
building the same tree against both versions rather than by trusting the release note. A
regression test now pins the behaviour with non-ASCII fixtures -- the existing repair-path tests
are ASCII-only and passed on the broken release, which is why the defect was invisible.rmcp to 3.4.0. It deprecates the ServerInfo alias in favour of ServerConfig
(the same type, renamed because it collided with the protocol's own serverInfo field), so
the MCP server's get_info moved with it. No wire-visible change.Character references in attribute values are now decoded
(#494). <a href> was the only
attribute in the converter that called the entity decoder, so every other user-visible
attribute emitted its raw source text: title="A&B" reached the output as the literal
A&B, and alt="it's" as it's. Twelve sites were affected -- title on <a>,
<img>, <graphic> and <abbr>; alt on <img> and <graphic>; src on <img>,
<audio>, <video>, <iframe> and <source>; cite on <blockquote>; the language-
class on a code fence; and <meta content> in extracted metadata. They were found by probing
every attribute that reaches output, not by reading the ones that looked likely.
Attribute reads now go through one shared accessor so a new attribute cannot reopen the gap.
The Markdown escaping is unchanged and was never at fault -- a literal " in a title was
always escaped correctly; the decoded character simply never arrived. Tier 1 carried the same
defect independently and moves in lockstep.
Three Tier-1 divergences that entity decoding made reachable. Each was latent long before
it could be triggered, and each is now pinned by a parity test. An <img> in a heading was
emitted as  by the fast scanner whenever keepInlineImagesIn was empty, where
the DOM path correctly replaces it with its alt text -- an empty list names no heading, so it
permits nothing. A heading whose body merely ended in whitespace was never trimmed, because
the trim only ran for bodies containing a newline; a decoded therefore survived where
the DOM path dropped it. And a link title containing a quote was escaped as " rather
than \". The first two were found by the generated-corpus parity test over 3,000 documents,
the third by a sweep over every character decoding newly makes reachable.
A wrapper element no longer defeats the nested-table fix
(#488). A nested <table> inside a
<td> was detected with a single-node tag-name test over the cell's direct children, so
wrapping it in a <div> bypassed the 3.12.4 deferral, the pipe escaping and the row fold all at
once. The inner table's raw | characters then read as the outer row's cell boundaries on
reparse -- content loss, not a cosmetic diff. The nested-table counter has always descended
through wrappers; the two now agree. The same line fixes the sibling-cell shape, which was
corrupt in the same way.
Content after a table whose last row is never closed is no longer lost
(#489). In
<table>...<tr></table><p>Visible footer</p>, the primary parser discards a close tag that does
not match the top of its open-element stack, so </table> vanished and the paragraph was
adopted by the still-open <tr> -- where the cell collector, which keeps only td/th, dropped
it. Such a document now takes the same html5ever repair path that #336, #479 and #486 already
use. A row that yields no cells also stops emitting a phantom empty row, so the next real row
becomes the header, matching Tier 1. Two further shapes are fixed by the same gate: a second
<tr> opened without closing the first (its row was silently dropped) and a <td> placed
directly inside <tbody> (which produced no output at all). Text-only children are deliberately
excluded from the gate, so a <tr> </tr> spacer does not pay for a full re-parse.
An anchor wrapping a table renders the table, not an escaped link label
(#490). Any block content inside an
<a> became link-label content, so <a href="..."><table>...</table></a> crushed the whole
table into a single label -- and the label escaper then correctly escaped the markdown that
produced, leaving an unreadable run of \|. The escaping was never the bug; handing a table to
the label builder was. The anchor's direct children are now partitioned, the inline half forms
the label and the deferred half renders as blocks after it. Only a deferred subtree that
actually contains a <table> triggers this, and never inside a heading or an inline context,
so every other anchor shape is byte-identical.
A whitespace-only inline wrapper still separates the words around it
(#491).
Alpha<span style="white-space:pre">\n</span>13 rendered as Alpha13. The predicate added for
#430 asks whether the next sibling is an element, so a bare text node after the wrapper took
the failing path and the newline vanished. Tier 1 was already correct, so the two tiers
disagreed on this input; they are now pinned together by a parity test. Note that the reported
white-space: pre is incidental -- that property is not implemented, and the defect reproduced
without it, exactly as a browser collapses the newline to a space either way.
The wrapper contributes a separator only when the next word butts straight up against it;
text that already opens with whitespace supplies its own, and is left alone.
keepInlineImagesIn now means something for <a>
(#492). The option was consulted for
headings and for layout cells, but never for anchors, so an <img> inside a link that also held
a block element was replaced by its alt text (or dropped entirely when it had none) no matter
what the option said. Listing "a" now keeps the image as markdown, in both the block and inline
anchor paths, and for <graphic> as well as <img>. The change is purely additive -- it can
turn an image on, never off -- so output is byte-identical for anyone not naming "a" in the
option. The Tier-1 scanner was updated in lockstep.
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.13.0 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.13.0")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: e2d32613e83b802407ec33226abbe46f052c85dc3c2604348902305861a637b4
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.13.0/html-to-markdown-rs-zig-v3.13.0.tar.gz\",\n .hash = \"html_to_markdown_rs-3.13.0-QtXyW3H7ngHrW5ZT2Swijkf_J5kWLQzW_K0zapcnt8De\",\n },\n},\n```\n
BREAKING (java): enum constants are now SCREAMING_SNAKE_CASE. LinkStyle.Inline becomes
LinkStyle.INLINE, HeadingStyle.Atx becomes HeadingStyle.ATX, and so on across every
generated enum; builder default sites move with them. The JSON wire values are unchanged, so
no serialized payload moves -- only the Java-facing constant names.
BREAKING: the RDFa structured-data variant is spelled consistently across bindings. One
Rust variant previously generated seven different names. Each binding now uses the correct
acronym segmentation for its own convention: rdfa (python options.py, elixir, ruby,
swift), RDFA (python stub, java, kotlin), Rdfa (csharp, go). Dart keeps rdFa, which is
what flutter_rust_bridge itself declares. The wire value "rdfa" is unchanged everywhere.
BREAKING (python): the type stub declared an enum member the runtime does not have. The
.pyi emitted StructuredDataType.RD_FA while the extension declares RDFA, so a type
checker accepted the name that raises AttributeError and rejected the one that works. Both
now agree on RDFA.
BREAKING (elixir): VisitResult's payload key matches the NIF struct. %{type: :custom, value: ...} becomes %{type: :custom, custom: ...}, and :error likewise. The NIF struct has
always declared custom/error fields, so the previous shape never round-tripped.
BREAKING (ruby): VisitResult.from_hash reads the real wire key. It read _0; the core
enum is adjacently tagged with content = "output", so _0 never matched and the payload was
silently discarded in both directions.
<br> inside an inline code span no longer emits a raw newline inside the backticks
(#487). <code>A<br>B</code>
produced `A\nB`, which CommonMark does not read back as one code span. The span is now
split and joined by a real hard break -- `A` + the newline_style marker + `B` --
so the marker sits outside the span, where it is syntax rather than content, and <code>
behaves like <b> and <i>, which already emitted hard breaks here. The output is round-trip
stable through a CommonMark render; the previous form was not. <kbd> and <samp> had the
same defect and are fixed with it.
Both tiers were wrong and disagreed with each other -- Tier 1 emitted `A \nB`, leaking
the two-space marker into the code content -- and both are fixed, so the tiers now agree. A
pre-existing Tier-1 bug surfaced on the way: inside a link label the <br> branch ran before
the code-span check, trapping two spaces in the first segment.
Contexts that cannot carry a hard break fold to a single space instead of splitting: headings
and table cells. A link label splits, matching what a bare <br> in link text already did.
A literal line ending inside an inline code span folds to a space. <code>a\nb</code>
(a real newline in the source, no <br>) produced `a\nb`. CommonMark gives a line
ending inside a code span no hard-break meaning and renders it as a space, so the old output
was never round-trip stable. <pre> and fenced blocks are unaffected and keep every newline.
html5ever stays pinned at 0.39.0. 0.40.0's serializer drops the C2 lead byte for every
codepoint in U+0080-U+00BF except NBSP (§, ©, °, ±, », ...), emitting invalid UTF-8.
In this crate that makes repair_with_html5ever fail its from_utf8 check and silently skip
the repair pass on any document containing one of those characters -- it fails safe, but
degrades quality invisibly. Fixed upstream in servo/html5ever#784, which is merged but not yet
published; this crate will adopt 0.40.x once it is.
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.12.4 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.12.4")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 697bc6dd8380207b937c572f6a53a884708a73431710d7e070ea672840ac74d3
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.12.4/html-to-markdown-rs-zig-v3.12.4.tar.gz\",\n .hash = \"html_to_markdown_rs-3.12.4-QtXyW3nRngEhb98D0_wBXswCR21TMEOJZSnqtf50O6z6\",\n },\n},\n```\n
Render <strike> as strikethrough. The Tier-2 inline dispatch listed only del/s, so
<strike> fell through to a plain child walk and lost its ~~ markers entirely.
Merge adjacent inline elements that share a repeated-character delimiter, extending the
#483 fix from <em>/<strong>
to every other family. <del>A</del><del>B</del> produced ~~A~~~~B~~, which CommonMark
reparses as one strikethrough containing A~~~~B; two adjacent code spans produced
`AB``, reparsed as one span carrying the literal backticks.<ins>, <mark>, <var>and<dfn>` had the same defect.
Keep the word separator a whitespace-only inline element stands for, extending the
#481 fix to <ins>, <sub>,
<sup>, <var>, <dfn>, <abbr> and <q>, each of which joined the words either side.
Keep <code> </code> as a code span whose content is a space. It was dropped outright, and
the delimiter-space padding that an all-spaces body used to get turned one space into three
(CommonMark strips one space per end only when the content is not entirely spaces).
Stop <ins>, <kbd>, <samp>, <var> and <dfn> from emitting their markers inside a
code span or fenced block, where ==, a second backtick pair and * are literal content
rather than formatting. Every other inline handler already suppressed itself there.
Close 22 Tier-1/Tier-2 output divergences across the inline elements, found by sweeping both
converters over 26 tags. Tier-1 now reproduces <var>/<dfn>, bails on the shapes it cannot
(<q>, <mark>, the new delimiter merges, a whitespace-only body), and suppresses
<strong>/<em> markers inside a code span as Tier-2 does.
Treat adjacent duplicate Rust warning flags as equivalent in benchmark provenance checks while preserving timing gates.
Publish every native NuGet runtime package required by the managed package's runtime graph.
Upload Dart native archives and checksums required by the generated package downloader.
Publish Go installer archive aliases and required SHA-256 sidecars.
Verify the installed Homebrew CLI and FFI versions directly in the registry smoke task.
Stage original native archives through the shared Zig packager; the 3.12.3 source archive requires separate C FFI libraries.
Merge adjacent inline emphasis into a single delimiter run, so <i>A</i><i>B</i><i>C</i>
renders as *ABC* rather than the *A**B**C* CommonMark reparses as nested emphasis
(#483).
Emit exactly one space for a whitespace-only inline element; A<i> </i>B duplicated it and,
inside a paragraph, <p>A<i> </i>B</p> dropped it entirely
(#481).
Drive the musl Node cross-compile through the shared build action, whose per-leg artifact
staging replaces the napi artifacts call that @napi-rs/cli 3.9.1 made fail on any
single-target matrix leg.
Keep the visible text of a <tr> nested directly inside a <td>, a shape malformed
newsletter HTML produces; the row and its content were dropped silently. The same fix
repairs tbody-less tables with implicitly-closed cells, where
<table><tr><th>h1<th>h2<tr><td>a<td>b</table> emitted only a one-column header and
dropped both the second header and the entire data row
(#486).
Render a data table's nested single-cell table as its own table instead of flattening it into a cell of escaped pipes (#484).
Apply the WHATWG numeric character reference replacement table, so › decodes to
› rather than a raw C1 control character; null, surrogate and out-of-range
references now yield U+FFFD (#485).
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.12.3 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.12.3")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: bf2a6f51683d998c7ff08fc8b9bee06d47f40c00fe226af234de500f2b7f99ae
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.12.3/html-to-markdown-rs-zig-v3.12.3.tar.gz\",\n .hash = \"html_to_markdown_rs-3.12.3-QtXyW6p0AQARDHIsBKu1XVrpRf5J7ornocJZf2alzaq5\",\n },\n},\n```\n
<!-- zig-3123-native-linking -->
The Zig archive in 3.12.3 contains the Zig sources and requires the matching C FFI archive from this release. Pass the extracted `lib` and `include` directories as the dependency build options `ffi_path` and `ffi_include_path`; this configuration was verified with all 84 Zig consumer tests.
font-size: 0 wrappers (#476).nil for optional Ruby conversion fields.. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.12.2 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.12.2")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: b6286c80ee02da4dfa2d8050dc41a0adcc83eaba9d50c7bb30d2cbd27e84c951
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.12.2/html-to-markdown-rs-zig-v3.12.2.tar.gz\",\n .hash = \"html_to_markdown_rs-3.12.2-QtXyW6p0AQBk-GH2nSYgzG0f0fbEkw8paxk4nRKAhbXT\",\n },\n},\n```\n
bullets now applies to layout-table rows
(#472). A table with inconsistent
column counts and no <th>/<caption> is treated as a layout table and renders each row as a
list item — but that renderer hardcoded - and never read options.bullets, so configuring
the option had no effect on the markers actually emitted. Surfaced by the reporter of #470, who
set bullets to "*+-" in 3.12.0 and still got hyphens.
The marker now cycles through bullets by nesting depth the same way list items do, so a
layout table nested inside a list takes the next marker rather than repeating its parent's. The
prefix strip that prevents a doubled-up marker assumed a hyphen as well, and now accepts any
configured bullet.
Default output is unchanged — the default bullets string already starts with -. Tier-1
bails on layout tables, so this is a Tier-2 path with no parity mirror.
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.12.1 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.12.1")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: dd30529d1d5ac11e5a28990a725b3ed6aaa35737c549ce4935121bb4edc85774
Correctness release covering the four defects reported against 3.12.0. Three are conversion bugs in the core, one extends the hidden-element rule; all four were reported through the Java binding and reproduce identically in every language.
An uppercase void element no longer swallows the rest of the document
(#467). The bundled astral-tl
parser matches void elements against an all-lowercase table using a tag's raw source bytes, so
<META> missed it and was pushed onto the open-element stack as a container: the parent's
close tag could not pop it and every following sibling was absorbed as its child. For
<head><META ...></head><body> that left <body> a grandchild of <head>, out of reach of the
direct-child rescue in handle_head, and the whole document converted to an empty string. The
charset attribute in the report was incidental -- the crate has no charset handling, and every
HTML5 void element (<BR>, <IMG>, <HR>, <INPUT>, <LINK>, ...) reproduced it. Tier-2 now
lowercases void-element tag names during preprocessing; attribute names and values are left
byte for byte alone. Tier-1 already lowercased before lookup, so this also closes a tier
divergence.
A nested table inside a cell keeps its row boundaries under br_in_tables
(#469). A Markdown cell cannot hold a
nested table, so the inner table is flattened into the outer cell with its pipes escaped. Until
3.11.2, br_in_tables: true merely skipped the whole-cell newline fold, letting the inner rows
leak out of the cell as raw newlines -- malformed, but a GFM parser could still see two rows.
Making that fold unconditional (correctly, for issues #456 and #457: a raw newline between two
pipes splits the row across physical lines) collapsed the rows onto one line joined by spaces
and erased the boundaries. The fold stays unconditional; the flattened rows are now joined with
the literal <br> that br_in_tables already means everywhere else in a cell. A preceding
sibling is separated from the nested table too, which previously ran straight into its first
pipe (Before\| ID). This restores row boundaries, not table structure -- a real nested GFM
table remains impossible and the inner pipes stay escaped.
Adjacent paragraphs in a layout-table cell are separated
(#470). A table with inconsistent
column counts and no <th>/<caption> renders each row as a list item, and those cells convert
as inline. That suppressed the ordinary block separator while never reaching the table-cell
continuation rule, so <p><b>Alice Example</b></p><p><i>Customer Service</i></p> emitted
**Alice Example***Customer Service* -- merged words and invalid emphasis. Such cells now
follow the settled cell rule from issues #453/#454: a literal <br> under br_in_tables, a single
space otherwise. <div> continuations are covered by the same change. Layout rows are list
items, so they still do not take table-cell pipe or emphasis escaping. Not a regression -- this
predates 3.8.3.
font-size: 0 now marks an element as not rendered
(#468), joining display: none,
visibility: hidden and the hidden attribute. Reported against generated email banners whose
marker text reached the Markdown despite being invisible in a browser. Any exact zero length is
recognised (0, 0px, 0.0em, .0%, in any casing, with or without !important) and the CSS
last-declaration-wins cascade applies as it already does to display.
This drops content that previous versions emitted. One case is deliberately exempt: the same
declaration is the classic inline-block/email spacing hack, where the wrapper kills inter-child
whitespace and each child restores a readable size. When a descendant re-declares a non-zero
font-size, the subtree is kept. That guard is a one-level-of-inheritance heuristic, not a
cascade -- a size restored from a stylesheet is out of reach of a byte-level pass. Tier-1 has no
subtree awareness and bails on any font-size: 0, deferring to the tier that can see the
descendant.
Detection remains unconditional, as it has always been for the other three: no option governs it.
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.12.0 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.12.0")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 33854e506db4b72de2d9de44da9ef7f9ce6e975dafdc7d890d3f1ec969ea951e
Correctness release. A broad pass over converter correctness: defects in the shipped output, plus a run of cases where the Tier-1 fast scanner and the Tier-2 converter disagreed on the same input -- the library picks a tier automatically, so those meant one document could convert two ways.
Test coverage behind it: every one of the 652 CommonMark spec examples is now exercised through
a conversion-fixpoint oracle (the exact-match test compares against the spec's own rendering, so
it can only run on the 131 examples admitting a single valid form), a Tier-1/Tier-2 differential
oracle over the benchmark corpus and a generated document set, and a fuzz target.
Against that fixpoint oracle, 644 of the 652 examples now round-trip unchanged with escaping enabled and 629 with the shipped defaults, up from 607 and 597. Of the 8 that remain, three are inherent to the round trip -- adjacent block quotes, and adjacent lists sharing a bullet, merge under any compliant reparse whatever we emit -- and the other five differ only in the byte spelling of a link destination.
htm_conversion_options_update_visitor is gone from the C API. The getter documented a
non-null return as caller-owned data, but ConversionOptionsUpdate is only ever constructed
through htm_conversion_options_update_from_json, and its visitor field carries a full
serde(skip) -- so Deserialize could never populate it and no FFI setter existed. The symbol
could only ever return 0. Breaking for C consumers: it disappears from
html_to_markdown.h and from the compiled library, so a call that used to yield NULL at
runtime is now a link error. Nothing that worked stops working. ConversionOptions.visitor is
unaffected -- it carries the same attribute but is genuinely reachable, because
htm_options_set_visitor writes it directly, bypassing serde.ConversionOptions fields that the library and the MCP surface already
exposed but the CLI hardcoded: --exclude-selectors, --url-escape-style,
--max-image-size, --capture-svg, --no-infer-dimensions and --tier-strategy. Each
now defaults to the library's own default, so an invocation that omits them is unchanged.
The three image flags require --extract-inline-images, which they are inert without.escape_link_label returns Cow<'_, str> and skips its two allocations when the text
contains no [ or ], which is the common case for link labels, image alt text and
titles. The two-pass design is unchanged -- it exists to avoid an O(n^2) String::insert
shape.Uppercase attribute names no longer discard link destinations and image sources.
HTML attribute names are case-insensitive, but the Tier-2 converter matched them
byte-for-byte as written, so <a HREF="up.html">link</a> converted to bare link and
<img SRC="a.png" ALT="cat"> to  -- the destination and the alt text were
dropped, not merely reformatted. The Tier-1 fast scanner has always compared attribute
names case-insensitively, so these inputs were also a tier disagreement: the same
document converted two ways depending on which tier ran. The astral-tl 0.8.0 parser
upgrade lowercases attribute keys at parse time, which fixes both.
Tier-1 and Tier-2 disagreed on GFM autolinks. autolinks defaults to true but was
not gated in the Tier-1 router, and the Tier-1 scanner had no autolink branch, so
<a href="https://x.com">https://x.com</a> converted to <https://x.com> on Tier-2 and
[https://x.com](https://x.com) on Tier-1 -- the library picks a tier automatically, so
the same document converted two ways. Tier-1 now implements the autolink form rather than
being gated off it, since gating a default-true option would have made the fast path
unreachable for ordinary input.
Tier-1 could return the wrong output for an autolink whose label carried inline markup.
Tier-2 tests its autolink predicate against the tag-stripped text, so
<a href="https://x.com"><b>https://x.com</b></a> still autolinks there; Tier-1 compared
its rendered label (**https://x.com**), never matched, and emitted
[**https://x.com**](https://x.com). Tier-1 now bails to Tier-2 on this shape. The bail is
guarded by a subsequence precheck so it fires only when an autolink is actually possible:
on the benchmark corpus, none of the 889 scheme-href links that carry nested markup trigger
it, so decorated links keep the fast path.
<address>, <search>, <hgroup> and <center> no longer merge into their neighbours.
All four are block-level, but they reached a pass-through handler that emits no separator,
so <address>foo</address><address>bar</address> produced foobar -- nothing in the output
recorded that these were ever distinct blocks. <colgroup>, <col>, <base>, <html> and
<body> share the same internal classification and are deliberately unchanged: the first
three are table-internal or void metadata, and the last two wrap every document.
Two block containers in one table cell are separated by one space, not three. The fast
scanner emitted a hard line break between them regardless of br_in_tables, and the
table-cell finalizer turned it into a three-space run, where the converter emits a single
space under the default options.
Five Tier-1/Tier-2 divergences closed. <blockquote> omitted its trailing blank line
in the fast scanner, invisible to every existing test because they all placed the
blockquote last; fixing it exposed a <pre> fence stripping one trailing newline where the
converter strips all of them. A text node's leading space survived at the start of a bare
<span>/<u>, producing a double space across the Google Docs and WordPress fixtures.
<address>, <search>, <hgroup> and <center> emitted a block separator the converter
does not. Images with a lazy-load src and flattened nested-table pipes now agree as well.
An HTML comment no longer forces a redundant reparse of the whole document. Custom
elements are detected by looking for a hyphen in a tag name, but the scan treated the text
inside <!--...--> as a tag name -- !--c-- contains a hyphen -- so any document with any
comment was reparsed through the repair path for nothing. On a 200-element document
producing byte-identical output, a single comment cost 1.8x. Comments are near-universal in
real pages, so most documents were paying it.
A table nested inside another table's cell no longer destroys the inner cells. The
nested table's own row and separator syntax was flattened into the outer cell unescaped, so
its bare | characters were read as additional cell boundaries for the outer row. GFM
truncates a row to the header's column count, so the inner cells were dropped outright on
reparse. The flattened content is now pipe-escaped outside code spans.
A run of between two inline elements survives. A whitespace-only text node
between inline siblings was collapsed to a single plain space unconditionally. str::trim
is Unicode-aware and counts U+00A0 as whitespace, so the run was destroyed on the first
conversion, not merely on a round trip. The same collapse also ran without checking whether
the output already ended in a space, stacking a real inter-element space against the
synthetic one left behind by <style> removal into a literal double space.
Content written directly inside <table>, outside any row or cell, is no longer dropped.
HTML5's "in table" insertion mode foster-parents such content: it moves to just before the
table and survives. The parser this crate uses performs no such fixup and the table builder
recognised only caption/colgroup/col/thead/tbody/tfoot/tr there, so raw text
was silently discarded and stray elements went through a no-op handler --
<table>abc</table> converted to nothing at all.
It appeared to work whenever a comment happened to sit nearby, but that was coincidence: a
separate defect treats !--c-- as a tag name, sees a hyphen, and reroutes any
comment-bearing document through the html5ever repair path, which does implement foster
parenting. The text survived by accident. That path is now entered deliberately for this
shape. Comments themselves are excluded, since HTML5 keeps them as children of the table
and nothing is lost.
Lazy-loaded images resolve to their real address instead of converting to nothing.
Lazy-loading libraries leave src empty or holding a 1x1 data: placeholder and put the
actual URL in data-src, data-lazy-src, data-original, data-srcset or srcset, so
those images came out as ![alt]() or a base64 blob -- effectively invisible on a large
share of modern pages. The fallback applies only when src is already empty or a data:
URI, so a plain <img src="..."> is byte-identical to before. A populated non-data: src
is trusted even when it looks like a placeholder, since some pages carry the real photo
there while srcset holds only the lazy-load stand-in.
A heading inside <summary> or <figcaption> no longer splices its # prefix into
unrelated text. Tier-1 records a heading's content offset against whichever buffer is
active when the tag opens, and <summary>/<figcaption> accumulate into their own buffer.
Only table cells were special-cased, so a heading in a <summary> took a small
buffer-relative offset and used it to index the whole document output instead -- inserting
heading prefixes into the middle of already-emitted text. Rustdoc's
<details><summary><h3 class="code-header"> shape turned Sample into Samp#### #### le.
This corrupted adjacent content, not just the heading.
An unclosed <table> no longer discards its rows. Every element still open at end of
input is closed implicitly, but Tier-1's implicit-close path did nothing for <table>,
dropping the entire accumulated table -- fully-formed rows included -- where the explicit
</table> path rendered it correctly. Text sitting directly inside <table> outside any
row or cell is now routed to Tier-2 rather than silently discarded.
An empty title="" is treated as absent instead of rendered as (url ""). The empty
annotation carries no information and no Markdown serializer round-trips it, so converting
the re-rendered output produced different Markdown than the first pass. Applies to links,
images and graphics alike. A whitespace-only title is still a title; only a genuinely empty
attribute changed.
HTML5 bogus comments render as nothing instead of leaking their text. <?php echo 1; ?>
emitted ?php echo 1; ?>, <!bogus> emitted a stray >, and </3> emitted </3>. All
three are comment tokens under the tokenizer's tag-open, markup-declaration-open and
end-tag-open states, so they render as nothing -- as real <!-- --> comments already did.
Identical constructs were rendering differently based only on which tokenizer state they
happened to reach.
Most visibly this cleans up Microsoft Word HTML, whose downlevel-revealed conditional
comments (<![if !vml]> … <![endif]>, not wrapped in <!--) were surfacing as literal noise
around every image and footnote.
Real comments, CDATA, and doctypes are stepped over rather than scanned into. That matters
for downlevel-hidden conditional comments: <!--[if gte mso 9]> … <![endif]--> is a real
comment whose terminator is the --> at the end of <![endif]-->, so treating that
<![endif] as bogus would delete the comment's own terminator and swallow the rest of the
document. <"> had its alt silently become a real
nested link and lost the destination. Both ends of the pair are escaped, since escaping only
the closing bracket leaves the outer [ to be captured by a later ] and the image then fails
to form at all. A genuine nested  inside link text is left intact.
A link or image with no visible text no longer turns its own href into emphasis. Such an
element falls back to the href as its label, and that fallback bypassed the text escaper, so a
raw * or _ in the URL round-tripped into real emphasis. Two destination-escaping gaps
closed with it: a literal backslash before the paren escaping in an unbalanced-parens
destination was not itself escaped, and a raw line ending inside an angle-bracket-wrapped
destination is not valid CommonMark at all and is now folded to a space.
The Tier-1 fast scanner matches the Tier-2 converter on every fix above, and on a further run of divergences found by the differential oracle. Because the library selects a tier by input shape, each divergence was a case where one document could convert two ways. Three suppressions were also removed from the oracle's allow-list, so it now generates those shapes freely instead of avoiding them; the four that remain are each documented in place.
A hard line break inside a link's visible text survives a round trip. A <br> in a link
label was collapsed to a space, so converting, rendering back to HTML, and converting again
lost the break -- yet a hard break inside link text is perfectly legal CommonMark. Ordinary
soft newlines from wrapped source text still collapse to a space, and a break at the very
start or end of a label is still dropped, having no line to break to.
Four more places mistook inline content for a bare list marker. The check was a two-byte
suffix test, and a closing **bold** plus its space ends in the same two bytes as a real *
bullet plus its space. A block quote after inline text in a list item lost its continuation
indent and fell out of the item entirely on reparse, since CommonMark matches containers per
line; a fenced code block was glued onto the preceding inline line, where the fence is not a
valid opener; a <div> ran straight onto the previous text with no separator at all; and
Tier-1's paragraph handler ran the check without first confirming an open list item, so
top-level text merely ending in a hyphen and a space lost the blank line before the next
paragraph. Several of these also omitted +, the third bullet in the default cycle, so they
misbehaved at every third nesting level.
An unclosed <p> or <div> in a list item no longer swallows the next item.
<ul><li><div></li><li>foo</li></ul> converted to - - foo, turning two sibling items into
one item holding a nested list, and <ul><li><p>x</li><li>foo</li></ul> converted to
- x- foo, running both onto one line so the second stopped being a list item at all.
Closing the tag explicitly was already correct, so this hit precisely the shape the HTML5
parsing algorithm resolves with an implied end tag -- ordinary markup, not an edge case. The
primary parser nests the following item under the unclosed element; the misparse detector
that already re-parses such trees with the spec-compliant fallback only recognized a block
nested under an inline ancestor, so a div or p never triggered it. Genuinely nested
lists are unaffected.
Leading whitespace on soft-wrapped continuation lines no longer shrinks on every pass.
A run of spaces after a newline in block text collapsed to one space, but CommonMark
strips a line's leading whitespace entirely when assembling a paragraph, so a renderer
dropped the space that was kept and the next conversion saw one fewer -- the text never
stabilized. Such runs now collapse to nothing, which is what a round trip already forces.
Whitespace at a text node's own edge is untouched, since that is what keeps adjacent words
apart.
A literal backslash in a link destination is no longer swallowed. Two shapes lost data.
A backslash before ASCII punctuation was consumed as a CommonMark escape of that character
on reparse, so href="\*" came back as href="*". A backslash at the end of a destination
with no title after it merged with the closing parenthesis into an escaped \), so the
destination never terminated -- href="x\" reparsed as the literal text [t](x), losing
the href and the link structure and leaving raw brackets in the rendered output. Both are now
escaped; a trailing backslash followed by a title's space, which is harmless, still is not.
The Tier-1 fast scanner no longer emits broken output for block children of a list item.
A <div> inside an <li> came out with no continuation indent, so its content fell out of
the list on reparse; <blockquote>, <table>, <dl>, a <p> continuing existing text and a
<pre> as an item's first content were each wrong in their own way. Under the shipped
defaults Tier-1 is never reached for these, because metadata extraction and highlight styling
both force the DOM converter first -- but with those disabled, all six shapes diverged.
Rendering them correctly needs Tier-1 to defer its separator decisions the way the DOM
converter does, which also entangles the loose-list heuristic, so the scanner now declines
these shapes and falls back instead. No corpus coverage is lost: the new bail never fires
across either parity corpus.
A lone significant character after a <br> is no longer dropped. <p>a<br> </p>
kept its non-breaking space, but the same markup with a newline after the <br> lost it --
so whether content survived depended on nothing but whether the source HTML happened to be
pretty-printed. A rendered <br> is always followed by a literal newline, which flipped the
handler into a branch that discarded the run wholesale. Separately, a <br> inside a heading
emitted the two-space hard-break marker, which has no meaning on a single line: a renderer
collapses it, so the next conversion saw different bytes. It now emits one space.
br_in_tables: false is honoured for lists inside a table cell. Both converters emitted
a literal <br> between sibling <li> elements in a cell regardless of the option, so the
setting silently did nothing for the shape it most often applies to -- a list in a cell, as
MediaWiki sidebars produce. This also caused a round-trip instability: once a renderer
flattens the list, that same <br> took the ordinary option path on the second pass and
collapsed to a space, so the two passes disagreed.
A non-breaking space between two <br> tags is no longer dropped. The paragraph
pre-filter discarded a text node outright when trimming found it empty and both neighbours
were empty inline elements -- and trimming is Unicode-aware, so a lone U+00A0 counted as
empty even though it is visible content. It surfaced only on a second conversion, because the
source spells it (six ASCII bytes) while a renderer re-serializes it as the literal
character. Only genuine ASCII whitespace is dropped there now.
. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.11.6 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.11.6")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: ec1059f3bad2f58afe4a999d3c1d98fca3aaff35d8d9abb1c2e22949d92d6ef7
Re-release of 3.11.5, which never reached any registry: the publish run failed and crates.io still tops out at 3.11.4. Also carries the release-gate fixes below.
<br> runs emit one hard break each, and none trailing a block. A run of consecutive <br>
elements collapsed inconsistently, and a <br> immediately before a block boundary added a
stray break to the output.
packages/r/src/Makevars is tracked instead of gitignored. The file is alef-generated and
alef-owned, but packages/r/.gitignore discarded it, so alef verify --exit-code -- the
"Verify binding freshness" step of CI Rust -- failed on every run. src/Makevars.win stays
ignored.
The Dart e2e before hook installs the pinned flutter_rust_bridge_codegen. It hard-coded
2.12.0 while the project pinned flutter_rust_bridge 2.13.0, so CI Dart aborted on alef's
version-disagreement check. [crates.dart] frb_version now declares the pin explicitly next to
the hook that has to match it.
enum_module from the node and java e2e call overrides; neither emitter
reads it.. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.11.4 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.11.4")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 0dcf2deb57f153bdce75230cbc584f83781fb9a55988bef9fa7d4a23cbcef8e6
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.11.4/html-to-markdown-rs-zig-v3.11.4.tar.gz\",\n .hash = \"html_to_markdown_rs-3.11.4-QtXyW394AQD_FoWgOtRCLqffMWMRZHbeNkh0p3RZPItR\",\n },\n},\n```\n
alef.toml's [workspace.docs.snippets] timeout_secs was 120s, which also bounds every
session's before hook. On a cold checkout the kotlin_android (assembleDebug), wasm
(pnpm run build:all) and swift (swift build) sessions could not finish their before build
within that window, so their snippets were reclassified unresolved_dependency and reported
Unavailable — failing CI's alef snippets check --strict (Failed: 0 but a large
Unavailable count) even though no snippet itself was broken. Raised to 900s, matching the
precedent in tree-sitter-language-pack's alef.toml.
The benchmark regression gate no longer fails at random on the CPU the runner happened to draw.
htmbench compare compared the whole Provenance struct for equality, so a capture from an
Intel Xeon host could not be evaluated against a baseline calibrated on an AMD EPYC host: it
aborted with benchmark provenance mismatch before any timing was compared, even though the
timings themselves passed every threshold. GitHub's ubuntu-24.04 label spans both vendors and
the calibration campaign is workflow_dispatch-only, so which host each side drew was a coin
flip. cpu_model and cpu_count are now excluded from the provenance contract; every other
field — rustc, Cargo, profile, build flags, core features, measurement mode, tier and visitor
selection, iteration/warmup settings, runner image and class — still hard-fails on drift. A
differing CPU is reported instead as a host mismatch, and the new off-by-default
--allow-host-mismatch flag (passed only by the CI regression job) downgrades timing violations
to advisory on hardware the baseline was never calibrated on. Thresholds are unchanged and the
baseline was not re-cut: a regression measured on the calibrated CPU still fails the run.
The FFI Symbols CI gate is green again. It was failing on two independent findings.
The htm_register_html_visitor allowlist entry is retired. It claimed "visitor registration
never landed in the FFI crate", which is no longer true: registration landed as the vtable API
htm_visitor_create / htm_options_set_visitor / htm_visitor_free, and the alef 0.62.6 regen
(4765a4139) replaced the last phantom caller — a Java Panama SymbolLookup.find(...) in
NativeLib.java — with lookups of those real symbols. The checker reported the entry as
orphaned rather than resolved only because it matches on exact symbol names and the
replacement uses different ones. The C# visitor surface reaches the same exports through
NativeMethods.cs and HtmlToMarkdownConverter.Convert, and all 49 tests in
e2e/csharp/tests/VisitorTests.cs pass against the built native library, so the public
IHtmlVisitor / HtmlVisitorBridge API is wired end to end.
htm_node_type_from_json is allowlisted, replacing it. The alef 0.62.9 regen (dc27e0560)
re-added a C# P/Invoke for it, but NodeType is a fieldless Copy enum: it crosses the C ABI
as int32_t, so the FFI exports htm_node_type_from_i32 / htm_node_type_from_str and no
handle-returning from_json. Nothing calls NodeTypeFromJson, so the declaration is latent
rather than a live crash. This is a re-regression — the same symbol was allowlisted and retired
once before in 555a29d20. Fixed upstream in alef's C# backend, which now skips scalar-crossing
named types when emitting handle-lifecycle P/Invokes; the entry comes out with the first regen
on an alef release carrying that fix.
A failed per-language artifact build no longer aborts the release. In v3.11.3 four
Build PHP extension (... windows-x86_64) legs failed and nothing else did, yet the GitHub
release was never promoted out of draft, never finalized and never announced, and the Homebrew
bottles, cargo-binstall verification, Packagist trigger and asset audit never ran — 3.11.3
reached every language registry but its release page stayed a draft.
Two GitHub Actions behaviours caused it, and the publish workflow now accounts for both. A
failure propagates down the entire needs chain, so a job's implicit success() gate is false
even when its own direct needs succeeded; and the propagated failure also poisons the value of
needs.<job>.result, so upload-php-pie-release reported failure to downstream jobs despite
concluding success — which is why always() plus !contains(needs.*.result, 'failure') still
skipped verify-assets, release-finalize and announce-discord.
The release spine (promote-release → release-finalize → announce-discord, plus
verify-binstall and the Homebrew bottle jobs) now gates only on jobs whose whole ancestry is
release-blocking: version validation, crates.io, and the CLI and C FFI assets. Per-language
artifact jobs stay in needs for ordering — the draft still flips only after every asset upload
has settled — but their results no longer appear in any gate. promote-release publishes a
promoted output for downstream jobs to gate on, because job outputs carry no failure poison.
Report release outcome job closes the workflow. It reads the run's real job conclusions from
the API and writes a job summary stating whether the tag shipped and naming every artifact that
is missing, then fails the run if anything failed. A release that ships without one language's
binary is now red and explicitly labelled incomplete rather than silently green.Verify release assets is a reporting gate rather than a release gate. It runs on every real
tag and no longer guards itself with !contains(needs.*.result, 'failure'), which had skipped
the asset audit in exactly the situation the audit exists for. Nothing in the release spine
depends on its result.. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.11.3 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.11.3")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: ddbdd9e96628e32dedff06413fcfe03aa737e7d2954fa83b9aebb5a206c92a78
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.11.3/html-to-markdown-rs-zig-v3.11.3.tar.gz\",\n .hash = \"html_to_markdown_rs-3.11.3-QtXyW394AQBYIV_EqbMNRe7MJwNJSYL-dauh-eXn8H3N\",\n },\n},\n```\n
Leading normalized whitespace at the start of a paragraph is no longer emitted before an image, preventing CommonMark from interpreting the image as an indented code block (#460).
The Swift crate compiles again. Codegen emitted a bare EnumName::Variant path for
data-carrying variants when converting a Swift string back into a Rust enum, which rustc rejects
with E0533: expected value, found struct variant; the regenerated crate routes those through
per-enum string conversion helpers instead.
The e2e freshness gate no longer fails on formatting it cannot produce. It installs Elixir, so
alef's mix format pass runs there as it does locally — without it the job emitted unformatted
.exs and reported the difference as fixture staleness.
The crates.io skip guard reads a real output again. The publish workflow asked
check-registry for an extra-packages key (cli_exists), but a composite action propagates
only the outputs it declares, so that key arrived as an empty string and the derived all_exist
could never be true. The guard has therefore never once skipped an already-published version. It
now reads the action's declared all-exist output.
packages/java/pom.xml is alef-owned and regenerated. It had drifted out of alef's ownership
since roughly alef 0.60.x for want of a provenance marker, so this release lands the accumulated
template changes at once: the file is reindented to four spaces, gains <excludes> /
<sourceFileIncludes> blocks, and moves checkstyle to 13.11.0 and jackson-databind /
jackson-datatype-jdk8 to 2.22.2. Those versions are alef's pinned template defaults, not
html-to-markdown-specific choices.
All Rust dependencies taken to their latest versions (cargo upgrade --incompatible followed by
cargo update): rmcp / rmcp-macros 3.1.2 to 3.1.4 plus twelve transitive bumps. Fourteen
packages changed, none downgraded.
alef pinned to 0.62.8.
The benchmark harness now records nine timing samples with median and median absolute deviation, captures runner and toolchain provenance, and supports reviewed fixture-specific noise floors without changing the existing percentage regression thresholds (#461, #462).
Java's VisitResult is now a plain sealed interface. Its Jackson @JsonSerialize /
@JsonDeserialize annotations and the nested serializer and deserializer classes are gone; the
visitor bridge marshals the type directly and never routed it through Jackson.
IHtmlVisitor, HtmlVisitorAdapter and HtmlVisitorBridge.
They shipped alongside the live HtmlVisitor, VisitorBridge and VisitorHandle surface but
nothing in the binding referenced them, so an implementation of IHtmlVisitor was never
invoked. Implement HtmlVisitor instead.chore(release): substitute Swift checksum for 3.11.2
chore(release): substitute Swift checksum for 3.11.2
The R package now builds against the vendored core crate instead of failing to resolve it.
configure stripped the path = ... clause from src/rust/Cargo.toml and wrote a cargo source
replacement to src/rust/.cargo/config.toml, but cargo reads .cargo/config.toml only from the
working directory and its ancestors, and Makevars.in runs cargo from packages/r/src — so
src/rust/.cargo is a descendant and that config was never read on any build. With the path
clause gone the dependency fell through to crates.io, which did not have the workspace version.
The path is now rewritten to point at the vendored copy, which needs no cargo config, no
.cargo-checksum.json, and no complete vendor tree.
Newlines inside a table cell no longer leak into the rendered row when br_in_tables is enabled
(#456,
#457). The whole-cell newline backstop
was gated on !br_in_tables, so enabling br_in_tables disabled the last-resort guarantee that a
raw newline never reaches a cell. Two consequences: a <pre> inside a cell now renders inline,
dropping its fence, indentation and language info string, matching how headings and list items
already render in cells; and under WhitespaceMode::Strict a newline inside a cell now folds to a
single space. Every other whitespace byte is still preserved exactly under Strict, and text
outside table cells is unaffected.
The Tier-1 fast path now honours br_in_tables inside table cells. It previously emitted a
sentinel that always expanded to three spaces, ignoring the option, so a document routed through
Tier-1 could render <br> in a cell differently from the same document routed through Tier-2.
Tier-1 also folds a text node's trailing newline the way Tier-2 does, fixing cells that were
double-spaced against Tier-2 when the source HTML was pretty-printed across several lines.
A literal backslash in HTML prose is no longer silently lost when the Markdown is read back
(#458). CommonMark consumes a backslash
that precedes ASCII punctuation, so emitting 3\*4 from <p>3\*4</p> produced Markdown that
reparsed as 3*4 — the source character was gone. Such a backslash is now doubled, and so is one
that sits immediately before a line ending (where CommonMark would read it as a hard line break)
or at the end of a text run (where whatever the emitter appends next would become its escape
target). This changes output under default options: any document whose prose contains a backslash
in one of those three positions gains a second backslash there. A backslash before anything else
is already literal and is still emitted bare, so Windows paths such as C:\Users\Alice are
unchanged. Code spans, code blocks and link titles are also unaffected — the first two are
verbatim by design and the third already escaped backslashes through its own rule. The escape is
deliberately not gated behind escape_misc or escape_ascii, because it preserves a character
the source actually contained rather than neutralising Markdown syntax the way those flags do.
chore(release): substitute Swift checksum for 3.10.6
chore(release): substitute Swift checksum for 3.10.6
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.6")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 63dab7ab85743b5527de0149cb369029c67f1b55ad69897de09a3611186e4d3d
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.10.6/html-to-markdown-rs-zig-v3.10.6.tar.gz\",\n .hash = \"html_to_markdown_rs-3.10.6-QtXyW946AQBULKaX63V4lYQJddpr-rfMxGrlWJ-1NuKp\",\n },\n},\n```\n
linux-x64-musl and linux-arm64-musl packages really are published now. 3.10.5
added them, but both cross-compiles failed to link (cannot find libgcc_s.so.1) and, because the
npm publish job requires every matrix leg, no Node package reached the registry at all. The build
action exported CC/CARGO_TARGET_*_LINKER pointing at musl-gcc, and cargo-zigbuild only sets
those when they are unset, so the zig cross-compile was silently replaced by a host-arch musl-gcc
that cannot produce a musl cdylib. The musl legs now opt out of that export.. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.10.5 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.5")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: c0b82237d3061180f6d703fb63577c7846946f1c7a7b215f3f2dce0dad425845
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.10.5/html-to-markdown-rs-zig-v3.10.5.tar.gz\",\n .hash = \"html_to_markdown_rs-3.10.5-QtXyW946AQBwksYWmeh8ZYZxjdGOXhyRIhIco_mO_8sv\",\n },\n},\n```\n
linux-x64-musl and linux-arm64-musl packages are now built and published. They
were advertised in the main package's optionalDependencies but never produced, so Alpine and
other musl installs silently fell back to no native binding, and pnpm install --frozen-lockfile
could not resolve them.Object namespace, which
collided with unrelated libraries (notably the parser gem's Parser constant). Generated types
now stay namespaced under HtmlToMarkdown (tree-sitter-language-pack issue #173).Result<_, Error> instead of
Result<_, ConversionError>, the docs site shipped 102 empty language tabs across five pages, and
several bindings' snippets referenced helpers that no longer exist.. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.10.4 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.4")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 17b70ce9462a61fa531efe3af8d18ee7c131f0730f984f2a1ce4dceecb35bf64
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.10.4/html-to-markdown-rs-zig-v3.10.4.tar.gz\",\n .hash = \"html_to_markdown_rs-3.10.4-QtXyW4k5AQDWluuATcik-i20djpNBPoL9skYtduLI1pd\",\n },\n},\n```\n
import { convert } from "@xberg-io/html-to-markdown" now
resolve. The napi-generated index.js ended with module.exports = nativeBinding, which Node's
ESM↔CJS interop (cjs-module-lexer) cannot statically analyze, so named imports threw
SyntaxError: The requested module ... does not provide an export named 'convert'. The build now
appends explicit module.exports.<name> = nativeBinding.<name> re-exports for every public
runtime export; require() is unaffected (#450).. package ( url : " https://github.com/xberg-io/html-to-markdown " , from : " 3.10.3 " )
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.3")The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 373475b37682322e24c5908edf693d67b6e35b8cdb6a7bc73332967d4326e196
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.10.3/html-to-markdown-rs-zig-v3.10.3.tar.gz\",\n .hash = \"html_to_markdown_rs-3.10.3-QtXyW4k5AQBwi_ZfAeNFoqv13MaDLPXkXxbjDRufhmg_\",\n },\n},\n```\n
swift build. The RustBridgeC
target was header-only, so XCBuild failed to link (RustBridgeC.o was never emitted) for every
iOS/macOS consumer of the published SwiftPM package. It now ships a translation unit with an
anchor symbol so the object is always produced (#449).html-to-markdown-php crate alongside the other
in-crate bindings; the standalone packages/php layout has been removed. Composer consumers are
unaffected.`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.2") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.2")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 2a97cbff8829e5a321616404f2ef3eb0264846a4f8f318f1e5b264369ed79560
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.10.2/html-to-markdown-rs-zig-v3.10.2.tar.gz\",\n .hash = \"html_to_markdown_rs-3.10.2-QtXyW4k5AQB4oIu3B3SJJICdKDLFUpfmOfSYis50Nj-c\",\n },\n},\n```\n
cargo binstall html-to-markdown-cli support (#448) — prebuilt CLI binaries can now be
installed directly from GitHub Releases without compiling from source. Adds
[package.metadata.binstall] to the CLI crate plus a release-time verify-binstall CI
job that installs via cargo binstall and smoke-tests the binary.`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.1") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.1")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: e3323476d17a6d51d4d05097c2777886c6755fead20e67db8a0589a2f5afd65a
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.10.1/html-to-markdown-rs-zig-v3.10.1.tar.gz\",\n .hash = \"html_to_markdown_rs-3.10.1-QtXyW4k5AQClajQ_qj0_3Lu57j2HeeA5sdFHJgwLoxpH\",\n },\n},\n```\n
io.xberg:html-to-markdown-android) now bundles its JNI native libraries.
Previously the published AAR contained no .so files, so every consumer crashed at runtime with
UnsatisfiedLinkError: library "libhtm_jni.so" not found (#446). The publish workflow built
the wrong crate (the C-FFI html-to-markdown-ffi) and uploaded it from a path the build action
never wrote, staging nothing. It now builds the JNI crate (html-to-markdown-rs-jni →
libhtm_jni.so) and stages it for every ABI, and a regenerated Gradle guard (alef 0.48.16) fails
the build if a correctly-named lib*_jni.so is ever missing.tracing spans and events as a first-class observability
surface: an html_to_markdown::convert span (input_len, output_format, wrap,
extract_metadata, extract_images, tier_strategy fields) wraps every conversion, with
DEBUG events at parse/walk/render stage boundaries and WARN/ERROR events on recovered or
fatal failures. The library never installs a subscriber — attach one in your application to
observe it. The CLI now initializes a tracing-subscriber (respecting RUST_LOG, with
--debug raising the default level) and routes all diagnostics through tracing instead of raw
stdout/stderr writes.arm64-v8a, x86_64, armeabi-v7a, x86).base64 0.23 and rmcp 3.0.1.`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.0") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.10.0")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 947e2403c1e0ddc547e9986d429f26a2084cef7c26c1b5ebfe13e3114e53ad9c
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.10.0/html-to-markdown-rs-zig-v3.10.0.tar.gz\",\n .hash = \"html_to_markdown_rs-3.10.0-QtXyW843AQD_k5XBLDSSyk40MOYfbNa_zLY5viQBzRLu\",\n },\n},\n```\n
rmcp 3.0 (MCP 2026-07-28 specification). The server now advertises
protocol version 2026-07-28 and negotiates down for older clients, so existing integrations keep
working. Minimum supported Rust version is now 1.88.convert_html (with json:true) and extract_metadata now
return structuredContent alongside the text JSON, so clients can consume the result as data.prompts/list, resources/list,
resources/read): ttlMs of one hour with a public cache scope.`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.9.2") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.9.2")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: ef3f0182adaaa140194cb739980b4fe10fb8cf3da6f14c99467d16fd34d68803
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.9.2/html-to-markdown-rs-zig-v3.9.2.tar.gz\",\n .hash = \"html_to_markdown_rs-3.9.2-QtXyW803AQAMY_N-BJoGh2LD3RJSBLL4qlFsoGIyCbUJ\",\n },\n},\n```\n
runtime.json (from the newly generated
runtime.json.template) before dotnet pack via the xberg-io/actions/render-runtime-json step.`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.9.1") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.9.1")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 5fc61ee6316d2f3520ab0a11458212a0c8625ba824b60e69e825ee604e3d0e21
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.9.1/html-to-markdown-rs-zig-v3.9.1.tar.gz\",\n .hash = \"html_to_markdown_rs-3.9.1-QtXyW803AQB7_qht6ZGjN8F2CKrYtXec2o2r54P7wqbh\",\n },\n},\n```\n
DepthLimitExceeded message now reports the effective
limit value, and the warning is emitted exactly once when a deeply-nested DOM is truncated.
Thanks @br411 (#428).`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.9.0") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.9.0")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 295f17d8b80b18e96546a6a17903b25294482c0e0430461df9c11ca75caea5bc
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.9.0/html-to-markdown-rs-zig-v3.9.0.tar.gz\",\n .hash = \"html_to_markdown_rs-3.9.0-QtXyW803AQAaNYax4YGJJ0rCYPD7MXKAYXAxq5jGVZL-\",\n },\n},\n```\n
max_depth, honored up to an internal
backstop of 1024. Deeply-nested email HTML that previously lost content past depth 64 now converts.
When the limit does truncate a subtree, a DepthLimitExceeded warning is surfaced instead of the
content being dropped silently.<br> in table cells with br_in_tables (#429): a <br> inside a table cell now emits a literal
<br> (valid single-line GFM) instead of a physical newline that broke the row.<span> whose sole
content is a newline now collapses to a single separating space instead of being dropped, so adjacent
inline text no longer glues together.<br> hard
break ( \n / \\\n), preserving the line break.keep_inline_images_in inside layout-table cells (#433): images inside a td/th listed in
keep_inline_images_in now stay as markdown when a Tier-2 layout table converts cells as inline,
instead of reducing to alt text.`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.3") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.3")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: ebd87aede2038f72a93970be19f911840424a147dd048ef984d297ef77772749
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.8.3/html-to-markdown-rs-zig-v3.8.3.tar.gz\",\n .hash = \"html_to_markdown_rs-3.8.3-QtXyW803AQDTcLKjGkAj61xzY0YTcAAc3YMKJMY0s6GS\",\n },\n},\n```\n
ConversionOptions.builder().withVisitor(...) now works. The visitor
upcall FunctionDescriptors were generated with a JAVA_LONG return layout while the handleVisit*
bridge methods return int, so the Java Linker rejected every stub with IllegalArgumentException: Wrong method handle type: (MemorySegment×5)int — even a no-op visitor threw before any callback ran.
Fixed upstream in the Alef generator (0.34.4); the descriptor return layout is now JAVA_INT.alef:format, ruby:format, and csharp:format tasks
(which invoked alef fmt); task format / poly fmt --fix . is the single formatter.`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.2") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.2")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: b902de73c91fc7bf34441bdedd65a9904cb2763db5c94cfdeccb46da594f3c72
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.8.2/html-to-markdown-rs-zig-v3.8.2.tar.gz\",\n .hash = \"html_to_markdown_rs-3.8.2-QtXyW803AQDu_fgSpeirhdTWsvAh1U-XyMNST9BUVrlB\",\n },\n},\n```\n
`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.1") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.1")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 22a18ff610fb70d069124970ea8ac824d93731ff7140b13cb5e747153c1e1c58
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.8.1/html-to-markdown-rs-zig-v3.8.1.tar.gz\",\n .hash = \"html_to_markdown_rs-3.8.1-QtXyW803AQC8DSTcdTAfB2GelrmxNA92CwyVoyGLIQhV\",\n },\n},\n```\n
`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.0") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.0")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 63daa2b7e9a25fde7febbc89fae536bb3abad1e13dbea851eb010c0e557b94be
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.8.0/html-to-markdown-rs-zig-v3.8.0.tar.gz\",\n .hash = \"html_to_markdown_rs-3.8.0-QtXyWxI3AQDzs_o3jEePH3M0v6SWDUkUshO_vQwFx2YG\",\n },\n},\n```\n
Stable release promoting 3.8.0-rc.2 (fully published). Version-only bump synced across all manifests.
`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.0-rc.2") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.0-rc.2")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: cbb995e8dc22ddea598c5cd5e5be13e719c8c12f7fd17f5479a789b6133e52ea
`swift .package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.0-rc.1") `
Add to your Package.swift:
.package(url: "https://github.com/xberg-io/html-to-markdown", from: "3.8.0-rc.1")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 3d8c5b4883e6ccf68780ce1d8143bd4318534af171c4673ad2de1ea50aae4a6c
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/xberg-io/html-to-markdown/releases/download/v3.8.0-rc.1/html-to-markdown-rs-zig-v3.8.0-rc.1.tar.gz\",\n .hash = \"html_to_markdown_rs-3.8.0-rc.1-QtXyWxc3AQAUNCiNxqxjuYS8lZk1IMAocjA_xp5z2NcV\",\n },\n},\n```\n
@xberg-io/html-to-markdown (the NAPI-RS crate itself — the separate -node package and the
TypeScript wrapper under packages/typescript/ are removed) with platform packages
@xberg-io/html-to-markdown-<platform>; WASM and CLI move to @xberg-io/html-to-markdown-wasm and
@xberg-io/html-to-markdown-cli. Java/Kotlin Maven coordinates and namespace move to io.xberg
(io.xberg:html-to-markdown[-android], JNI symbols Java_io_xberg_android_…), and the C# NuGet id to
XbergIo.HtmlToMarkdown. The GitHub org (github.com/xberg-io), Homebrew tap
(xberg-io/homebrew-tap), publisher GitHub App, sponsors links, and all docs/badges follow. The legal
entity name Kreuzberg, Inc. is unchanged.release/swift/<version> branch carrying the substituted
XCFramework checksum. The alef-generated Swift e2e/test-app pins
.package(url: …, branch: "release/swift/<version>"), but the publish workflow only force-moved
the v<version> tag and never created that branch, so SwiftPM could not resolve the package. The
checksummed commit is now also pushed to refs/heads/release/swift/<version>.
(.github/workflows/publish.yaml)$NF instead of $2. The smoke harness'
download_ffi.sh read the version from $2, which only holds in the block require (…) go.mod form;
the inline require <path> <version> form (emitted by the v3.7.2 regen) shifts the version to $NF,
so the script built a malformed module-cache path and failed with "Binding directory not found". This
is test-harness only — the published Go package is unaffected.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.7.2") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.7.2")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: f02a6cb0fb98942cca063d2357d97b182f4812efb851793cc944e1a327a7cfba
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.7.2/html-to-markdown-rs-zig-v3.7.2.tar.gz\",\n .hash = \"html_to_markdown_rs-3.7.2-QtXyWxU3AQDv862xcu9C78TqO91fqpIAMEc8XI2kBqyo\",\n },\n},\n```\n
alef.toml.alef_version to 0.26.6 and regenerate every binding. Fixes the dart
wrapper crate failing to build under cargo build --no-default-features: alef 0.26.x had regressed
the 0.25.33 fix and re-emitted #[cfg(feature = ...)] on the generated lib.rs mirror struct /
opaque-wrapper declarations, their From conversions, and from_json bridge fns, while
frb_generated.rs references those types/functions unconditionally (E0425: cannot find type VisitorHandle / create_html_metadata_from_json). alef 0.26.6 keeps those declarations
unconditional again. Verified: the regenerated dart crate compiles under --no-default-features,
default, and --all-features. (alef 0.26.6)kreuzberg/openssl-vendored. The
shared build-swift-artifactbundle action hardcoded that feature for its Linux (cargo-zigbuild)
targets, but it only exists in the kreuzberg core swift crate — so html-to-markdown's swift bundle
build failed (the package '…' does not contain this feature), which cascaded to skip
release-finalize (and with it the packages/go/vX.Y.Z Go module tag). The action now takes a
linux-features input (left empty here). (shared xberg-io/actions/build-swift-artifactbundle)ci(homebrew): drop the Intel-macOS `sonoma` bottle from the publish matrix. The formula intentionally has no x86_64-apple-darwin url (Apple-Silicon on
sonoma bottle from the publish matrix. The formula intentionally has no x86_64-apple-darwin url (Apple-Silicon only), but the bottle matrix still built a sonoma (Intel) bottle, which failed with formula requires at least a URL. Because release-finalize gates on publish-homebrew-bottles with if: !contains(needs.*.result, 'failure'), that one chronic failure silently skipped release-finalize for three releases (v3.6.20, v3.6.21, v3.7.0) — and with it the packages/go/vX.Y.Z Go module tag, leaving go get …/packages/go/v3@vX.Y.Z unresolvable. Removing the Intel bottle restores release-finalize (and the Go tag) for every future release. (.github/workflows/publish.yaml)cargo on PATH in the test before-hook. The "Run E2E tests (Windows)" step overrode PATH via env: with ${{ env.PATH }}, which omits cargo's $GITHUB_PATH additions under Git Bash, so the cargo build … html-to-markdown-ffi before-hook failed with cargo: command not found. Prepend the target dirs to the live $PATH at runtime instead. (.github/workflows/ci-e2e.yaml)MIX_ENV=test the force_build: … or Mix.env() in [:dev] clause does not apply, so mix test tried to download a precompiled NIF for the current (unreleased) version and failed with the precompiled NIF file does not exist in the checksum file. Set RUSTLER_PRECOMPILED_FORCE_BUILD_ALL=1 so the test/doc compile builds the NIF locally. (scripts/ci/elixir/run-tests.sh)Package.swift links. The swift e2e before-hook built only html-to-markdown-rs-swift, but the generated manifest links html_to_markdown_ffi, so swift test failed with library 'html_to_markdown_ffi' not found. Build html-to-markdown-ffi in the before-hook too (matching Go/C#/C). (alef.toml)--locked from the PHP test before-hook. The PHP e2e job's extension build rewrites the native package deps to the published registry version and runs cargo update, mutating the workspace Cargo.lock; the subsequent cargo build --locked -p html-to-markdown-php then failed with cannot update the lock file … --locked. (alef.toml)--no-frozen-lockfile for the NAPI binding install. build-node-napi ran a frozen pnpm install, but the napi platform packages in optionalDependencies are pinned to the unpublished release version and can never be in pnpm-lock.yaml, so the install failed with ERR_PNPM_OUTDATED_LOCKFILE across every Node build/e2e. (shared kreuzberg-dev/actions/build-node-napi, alef.toml)pie install over composer require. The PHP package is a native ext-php-rs extension that composer require cannot load (the cause of #420); every PHP install snippet now leads with pie install kreuzberg-dev/html-to-markdown. (docs/, readme_templates/)verify-release-assets. A dropped cell now fails the release instead of silently shipping a partial PIE matrix (#333). (.github/workflows/publish.yaml)alef.toml.alef_version to 0.26.5 and regenerate every binding, e2e suite, README, and API doc. (alef 0.26.5)<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.7.1/html-to-markdown-rs-zig-v3.7.1.tar.gz\",\n .hash = \"html_to_markdown_rs-3.7.1-QtXyWxU3AQC5VboClItxYGlHr9zwZbNiB_JuxgFAXM5z\",\n },\n},\n```\n
`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.7.0") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.7.0")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 38734253518dde8bf666d44fb4c8affe1acd68af295ac6f7949df81d96ef7851
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.7.0/html-to-markdown-rs-zig-v3.7.0.tar.gz\",\n .hash = \"html_to_markdown_rs-3.7.0-QtXyWxU3AQAU6AEplSD117Nl7Fs_7Y6-UTw7GJy4f9Z-\",\n },\n},\n```\n
convert_html tool now accepts a typed config object covering every settable ConversionOptions field (heading/list/escaping/whitespace/wrapping, preprocessing, image extraction, output format, tier strategy, …) instead of an opaque untyped JSON blob, so MCP clients discover all options through the tool's generated inputSchema; enum options are accepted as case-insensitive strings parsed by the core parsers. A new extract_metadata tool returns structured <head>/<meta> metadata (title, Open Graph, Twitter Card, JSON-LD/microdata, headers, links, images) as JSON. Both tools carry the full MCP annotation set (title, read_only_hint=true, idempotent_hint=true, destructive_hint=false, open_world_hint=false). The typed ConvertConfig mirror is implemented MCP-side (no changes to the alef-tracked core option types) and guarded by a drift test that fails if a core option is added without being mirrored. The mcp feature now implies metadata. (crates/html-to-markdown/src/mcp/)convert_to_markdown, extract_main_content, inspect_metadata) that drive the tools, with arguments; resources — htmltomarkdown://options-schema (the JSON Schema of every conversion option) and htmltomarkdown://output-formats (the markdown/djot/plain guide); and completions — argument autocompletion for prompt arguments (e.g. output_format → markdown/djot/plain). (crates/html-to-markdown/src/mcp/catalog.rs)`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.21") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.21")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: c063867f51327d7692a84c34483c3293d4758c01b28700bcd1b59545e9aa8c3c
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.21/html-to-markdown-rs-zig-v3.6.21.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.21-QtXyWxY3AQBLYrrrYs7DVvSH2JbbveqX8cC3K121BXL4\",\n },\n},\n```\n
alef.toml.alef_version to 0.25.60 and regenerate every binding, e2e suite, README, and API doc. Folds in the 0.25.59–0.25.60 generator fixes: the R/extendr by-reference DTO rework now emits a non-optional Named param following an optional param as &T (passed by reference via an owned name_core binding) instead of the non-compiling Nullable<&T>::into_option() path; Kotlin/Kotlin-Android content-union text() accessors reference the actual data-class payload property (value) instead of the non-existent field0; generated binding rustdoc de-links core intra-doc references (e.g. [`Error::LanguageNotFound`]) to plain code spans so rustdoc -D rustdoc::broken-intra-doc-links passes; and sync-versions now runs the same format_generated pass as alef all, so version-bumped manifests (package.json, composer.json, Package.swift) are byte-identical to the generate path and no longer trip freshness gates. (alef 0.25.60)`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.20") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.20")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: bd510aebc5cd41d3b742fb5df32d7de62e3e7a3ac67c298e9f2658e1d0ac3365
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.20/html-to-markdown-rs-zig-v3.6.20.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.20-QtXyWxY3AQAAiIzL1plM-zpynxmxL23d_JqELIak6jIK\",\n },\n},\n```\n
8a5203b7d) enabled several strict hooks that ran against alef-generated code and failed on it. All per-language code under packages/<lang>/, e2e/, and test_apps/ is generated and ships as alef emits it, so each hook now carries the same generated-tree exclusion the other formatters already use: yard-coverage excludes packages/ruby/ (its native.rb holds undocumentable Sorbet-sig RBI defs); air-check/air-format exclude packages/ (alef formats R with styler, not air); palantir-java-format excludes packages/, e2e/, test_apps/, and the vendored .mvn/ wrapper. The lintr hook runs lintr::lint_dir(".") and ignores pre-commit file scoping, so its exclusion lives in a new root .lintr (exclusions: list("e2e", "packages", "test_apps")); tsc-typecheck runs tsc --noEmit from the repo root, so a root tsconfig.json scopes type-checking to packages/typescript/src. .yardoc/ (yard's cache side-effect) is gitignored. (.pre-commit-config.yaml, .lintr, tsconfig.json, .gitignore)alef verify output-sensitive and the v2.3.0 hookset added alef-docs-fresh (which runs alef verify), so any prek formatter that rewrites a generated file differently than alef emits it now breaks verify. Fourteen generated files hit this because the hooks' pinned tools differ from alef's: ruff/ruff-format (line-wrapping generated Python), pyproject-fmt (alef runs no pyproject-fmt pass), oxfmt (alef formats composer.json via npm oxfmt; the hook uses the oxc binary with a different style), shfmt (generated download_ffi.sh/install.sh/gradlew), cargo-sort (the workspace-excluded R and Elixir NIF Cargo.toml), and end-of-file-fixer (the Dart frb_generated.dart and test_apps __init__.py). Each now carries the generated-tree exclusion the other tools already use, so the files ship exactly as alef emits them and alef verify stays clean. (.pre-commit-config.yaml)alef.toml.alef_version to 0.25.58 and regenerate every binding, e2e suite, README, and API doc. Folds in the 0.25.55–0.25.58 generator fixes that apply here: the generated Go cmd/download_ffi/main.go tool now carries the standard auto-generated by alef marker so golangci-lint's generated: lax skips it (0.25.x had dropped its //go:build ignore guard, exposing inherent-to-a-downloader gosec/errcheck findings); the generate/all up-to-date skip now also compares against on-disk output (not just the side cache), so out-of-band drift (a git restore, a hand-edit, an interrupted write) is regenerated rather than silently retained; extract OR-merges cfg gates for same-named types with disjoint gates; R/extendr forwards cfg-gated core features into the generated crate's [features] table so it builds under -D warnings. The visible churn is the generator's .d.ts/TS-bridge reformatting (tabs→2-space, double→single quotes). Verified locally: alef verify is clean and the 3.6.20 bump is synced across all manifests. (alef 0.25.58)`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.19") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.19")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 99ab6fe1c30809ef3a65768fa02a4b3442af58500a94561e848c10194f125be5
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.19/html-to-markdown-rs-zig-v3.6.19.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.19-QtXyWxY3AQB1EZz9fEW2CPe1LBO1Zhzn9LIRKwOv_LnI\",\n },\n},\n```\n
pkg:version syntax. Both generated PHP test_app installers (test_apps/php/install.sh, test_apps/php_ext/run_tests.sh) ran pie install --version "$VERSION" <pkg>, but PIE parses --version/-V as "print PIE's own version" and exits without installing — the pinned extension version was never fetched, so the registry-mode PHP smoke validated whatever stale build happened to be present (or nothing). Use pie install "<pkg>:$VERSION" so the targeted release is actually installed. (alef 0.25.54)node-bindings job runs setup-node-workspace (which installs with --no-frozen-lockfile) and then build-node-napi, whose own pnpm install --filter re-enforces CI's default frozen-lockfile. The napi platform optionalDependencies are pinned to the not-yet-published release version, so they cannot be in the lockfile — pnpm 11.6.0 treated the redundant install as a no-op, but the 11.8.0 bump made it fail (ERR_PNPM_OUTDATED_LOCKFILE), dropping every Node build and the npm publish. Pass install-deps: false so the already-installed workspace is reused. (.github/workflows/publish.yaml)x86_64-apple-darwin (Apple Silicon only), but scripts/publish/html-to-markdown.rb.tmpl + homebrew.json still referenced cli-x86_64-apple-darwin.tar.gz, so "Update Homebrew formulas" failed (no assets match the file pattern) and the tap stayed on the prior version. Remove the macOS on_intel block; macOS is now arm-only. (libhtml-to-markdown is unchanged — the C FFI still ships an x86_64-apple-darwin archive.)alef.toml.alef_version to 0.25.54 and regenerate every binding, e2e suite, README, and API doc. Folds in the 0.25.51–0.25.54 generator fixes: the PHP PIE pkg:version install fix above; Go — unit/newtype-tuple enum constants emit serde wire values, a parameter named result no longer collides with the codegen variable, and Option<&[u8]> returns generate correctly; R/extendr — cfg-variant dedup, registration entries drop the stray #[cfg(...)], and opaque-method/Option/Vec/error returns convert properly; C# — corrected host-capsule native method name; Swift — Package.swift dependency argument order and e2e harness_extras product override; test_apps runner — declared [crates.e2e.env] vars are exported to the run command. Verified locally: both PIE fixes are present in the regenerated test_apps and the 3.6.19 bump is synced across all manifests.oxfmt/oxlint away from generated e2e/+test_apps/. alef has no e2e/test_apps JS/TS formatter (alef.toml [crates.e2e.format] covers go/python/rust/c only), so oxfmt was the lone tool reformatting generated test JS, producing churn on every regenerate. Add the ^(e2e/|test_apps/) exclude that go-fmt, ruff-format, and golangci-lint already carry; generated test JS now ships as alef emits it, matching e2e/ruby, e2e/php, etc. (.pre-commit-config.yaml)`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.18") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.18")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 43a13bc905dc7afabc227feb3fe3d777b7d567d5a4ca03bda0c8b72b44baf2f4
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.18/html-to-markdown-rs-zig-v3.6.18.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.18-QtXyWxY3AQB5BUAdeLk5OXKTGaGRaxJ-Kf72SFcp-ExQ\",\n },\n},\n```\n
build-python-wheels action source-built libheif 1.23.0 (with libde265/x265 codec headers) for all consumers because kreuzberg links libheif-sys. html-to-markdown does not use libheif, but inherited the build — which fails on the manylinux2014 (CentOS 7 EOL) base whose yum repos no longer carry libde265-devel/x265-devel (Error: Not tolerating missing names on install). The action now gates the libheif build behind an opt-in build-libheif input (default off); h2m no longer builds it, so all Python wheels build again. (xberg-io/actions@v1)php8.5 + macos-arm64, restoring PHP publishing. v3.6.17 tried to build the php8.5 Apple-Silicon PIE asset on the macos-14 runner (#333), but shivammathur/setup-php cannot provision PHP 8.5 on arm64 on any macOS runner (the install leaves empty paths — sed: : No such file, /php.ini: No such file). That failing cell skipped upload-php-pie-release, dropping every PHP PIE asset (a v3.6.15-class regression). Restoring the exclusion republishes the PHP extension for all supported targets (PHP 8.2–8.5 on linux/macos-x86_64/windows, plus arm64-darwin for 8.2–8.4). #333 stays open pending an upstream php@8.5 arm64 formula. (.github/workflows/publish.yaml)alef.toml.alef_version to 0.25.50 and regenerate every binding, e2e suite, README, and API doc. Folds in the 0.25.50 cross-language visitor-codegen fixes that turn the chronically-red e2e visitor suites green: Node — the bridge now reads the externally-tagged { Custom: ... }/{ Error: ... } payload key (was lowercased to custom/error, yielding [object Object]); Ruby — visitor callbacks take named params and interpolate {placeholder} templates; PHP — bare-string returns map to Custom for multi-payload result enums; Elixir — unit-variant atoms match the snake_case wire name (:skip no longer leaks as literal "Skip"); Swift — visitBlockquote's depth is UInt, so the override satisfies the protocol. Verified locally: every language e2e suite passes and alef verify is clean.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.16") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.16")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: e2938481938d3f16077e672f37d878a3d35371ff9ee0ac1783e331a22c16005a
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.16/html-to-markdown-rs-zig-v3.6.16.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.16-QtXyWxY3AQCAtiJ7keG3srFaaRpSBRbbRy7g8MitNVCd\",\n },\n},\n```\n
php8.5 + macos-arm64 from the PHP extension matrix, restoring PHP publishing. v3.6.15 removed this exclusion to chase the Apple-Silicon PIE asset for #333, but shivammathur/setup-php still cannot provision PHP 8.5 on macos-arm64 — the homebrew tap has no php@8.5 arm64 formula, so the cell fails at the setup step with "Could not setup PHP 8.5". Because upload-php-pie-release does not run when needs.php-extension.result == 'failure', that single failing cell dropped every PHP PIE asset from the v3.6.15 release (a regression from v3.6.14's full set). Restoring the exclusion returns the matrix to all-green and republishes the PHP extension for every supported target (PHP 8.2–8.5 on linux/macos-x86_64/windows, plus arm64-darwin for 8.2–8.4). #333 stays open until upstream ships a php@8.5 arm64 formula. (.github/workflows/publish.yaml)publish-rubygems because the Build Ruby gem (windows-x64) cell hung ~3h on setup-rust and was cancelled, failing the needs.ruby-gem.result == 'success' gate. This is a transient runner fault, not a code defect; cutting v3.6.16 re-runs the publish so the gems ship.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.15") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.15")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 2bfed7f642b479cdbc298719dc38570a8a3c0a1994ec6e58f7f732dee106cb35
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.15/html-to-markdown-rs-zig-v3.6.15.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.15-QtXyWxY3AQDBJVMum8A8vFPEjN4UzVel2A0j1DR4wh4g\",\n },\n},\n```\n
php8.5 on macos-latest, but shivammathur/setup-php cannot install PHP 8.5 on arm64 (no php@8.5 arm64 homebrew formula), so the cell failed at setup and — because upload-php-pie-release skips on a failed matrix — dropped all PHP PIE assets from this release. See the 3.6.16 entry. (.github/workflows/publish.yaml)HtmlToMarkdownApi class in all PHP examples. The README quick-start, API reference, and docs snippets used the HtmlToMarkdown userland wrapper, which is only autoloaded via Composer's PSR-4 and is absent on the PIE install path (extension-only) — so HtmlToMarkdown::convert() raised "Class not found" for users who installed via pie. Examples now call HtmlToMarkdownApi::convert(), the native class registered by the extension itself, which works on both the PIE and Composer install paths (#415). (readme_templates/partials/, docs/snippets/php/, docs/language-guides.md)alef.toml.alef_version to 0.25.44 (was 0.25.40) and regenerate every binding, e2e suite, README, and API doc. Folds in the 0.25.41–0.25.44 cross-language generator fixes. Verified locally: alef verify is clean, cargo check --workspace --all-features compiles, and prek run --all-files (including the alef freshness hook) passes.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.14") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.14")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 97eafa47ebefbaf3dba90f374fbf7f5f733e6638abb9ae0313105b833d974774
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.14/html-to-markdown-rs-zig-v3.6.14.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.14-QtXyWxY3AQALNK2X6LpqtnVhv3vPx0qFP5ejWjqyU3m5\",\n },\n},\n```\n
alef.toml.alef_version to 0.25.40 (was 0.25.36) and regenerate every binding. Advances the pin past the [Unreleased] 0.25.33 note to a tagged, published alef revision and folds in the 0.25.34–0.25.40 cross-language fixes. Binding-visible changes: Java drops the throwing UnsupportedOperationException DTO/enum-method stubs (NodeContext.withOwnedAttributes/intoOwned, ConversionOptions.defaultInstance, PreprocessingOptions.defaultInstance) — there is no JNI symbol for these yet, so the stubs compiled but threw at runtime; absence is safer than a misleading throw until DTO marshaling lands (boxed-Boolean serde defaults also restored). Swift emits a matching pub fn visitor_handle_noop definition for the bridge no-op declaration so the swift rust crate compiles (was E0425 under 0.25.36's decl-without-def), alongside the upstream streaming-owner declaration rework. Node (NAPI) map-returning functions now convert borrowed maps to owned HashMap instead of returning them bare (fixes a would-be E0308 from the new DTO-method emission). Python (pyo3) drops a redundant # type: ignore[arg-type] on the visitor assignment that mypy --strict (warn_unused_ignores) rejected, and None-guards the coercion comprehension for optional Vec<enum> fields. Verified locally: task alef:generate is fresh (alef verify), every workspace binding crate compiles, and prek run --all-files is clean.upload-php-pie-release now hard-errors if zero php-package-* artifacts reach the aggregator instead of producing an empty PIE upload. The download step adds if-no-files-found: error, so a future regression that drops all PHP build outputs surfaces immediately instead of silently shipping a release with no PIE assets — the failure mode behind v3.6.11's missing-asset reports (see #333). The matrix-cell tolerance from v3.6.12 (needs.php-extension.result != 'cancelled' && != 'skipped') is unchanged. (.github/workflows/publish.yaml)html-to-markdown-rs-dart in the --no-default-features --workspace sweep. The provisional exclusion added earlier in this [Unreleased] cycle is now unnecessary: alef 0.25.33 + the queued [Unreleased] follow-up drop #[cfg(feature = ...)] from every dart-wrapper mirror declaration and the bidirectional From impls, so the wrapper crate compiles cleanly under default, --no-default-features, and --all-features. Verified locally against a fresh task alef:generate. (.task/languages/rust.yml)alef.toml.alef_version to 0.25.33 and regenerate every binding. Sweeps in the cross-language fixes accumulated through 0.25.30–0.25.33 plus the (still-[Unreleased]) dart cfg follow-up that the regen incorporates locally: pyo3 convert(...) / ConversionOptions(__init__) visitor kwarg widened to HtmlVisitor | object | None (resolves #403); swift options-field factory exported as make{Trait}Handle (matches docs snippets + e2e generator); java record + builder PMD cleanup (per-instance non-final fields, redundant component docs removed); kotlin file-level @file:Suppress extended with ReturnCount + NestedBlockDepth on the shared emitter; codegen Vec<core::T> → Vec<wrapper::T> conversion on non-opaque method returns (pyo3, magnus, extendr); wasm Option<Vec<UnitEnum>> getter/setter emission; php untagged-data-enum delegation via serde_json::from_value; extract honors #[cfg_attr(alef, alef(skip))] on impl blocks (no more duplicate #[no_mangle] symbols on builder + field name collisions); cfg-gated public function re-exports treated as binding surface; extendr extendr_module! registration entries gated on the source FunctionDef cfg; dart mirror enum + struct declarations + bidirectional From impls all unconditional now so frb_generated.rs resolves against the local crate regardless of the dart wrapper's feature set. Pin will advance again to a tagged alef revision once the [Unreleased] dart follow-up + extendr fix ship in 0.25.34.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.13") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.13")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 743af85f3475584026e68cd9313fe920813acc75b8c8f0ba5c741441546e98fb
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.13/html-to-markdown-rs-zig-v3.6.13.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.13-QtXyWxY3AQAOuSMCutFyCJ0h7wa33iC6cK0G87TcTLyC\",\n },\n},\n```\n
shfmt-normalize download_ffi.sh so CI Lint's shfmt hook stops rejecting the v3.6.12 baseline. The hand-written cgo library stager shipped with v3.6.12 used inline case-arm bodies; CI's shfmt -i 2 -ci -bn -s profile splits them onto separate lines. CI Lint failed on every v3.6.12 commit until the file was normalized.binding_excluded fields and emit ..Default::default() so core types with custom Default impls (e.g. anything that pulls config from the environment) are no longer shadowed by zeroed field defaults across Python/Node/Ruby/WASM/extendr/PHP method bodies and From-impls; FFI same-name function dedup moved out of the shared extractor pass into a backend-local pass (backends::ffi::gen_bindings::functions::cfg_dedup::dedup_same_name_functions) so every other backend and the e2e call-export validator see the original multi-entry surface untouched (v0.25.26 had over-collapsed the shared surface and stripped alef(skip)-tagged siblings from every backend); FRB Dart bridge functions wrapping core functions that return primitives now emit the cross-type cast (e.g. .map(|v| v as i64)) instead of the redundant .map(|v| v) that v0.25.25–0.25.27 emitted; generated Kotlin Android files add ReturnCount to the @file:Suppress list so detekt no longer fails sealed-class deserializers with 3+ variants.make{Trait}Handle (matching the documented spec, the docs snippets under docs/snippets/swift/visitor/, and the e2e test generator) instead of make{Trait}{TypeAlias} (which silently doubled the trait stem to makeHtmlVisitorVisitorHandle); generated Swift e2e visitor closures qualify the context type with the host module (HtmlToMarkdown.NodeContext) so the unqualified reference is no longer ambiguous against RustBridge.NodeContext that the test file also imports.SsrfPolicy.denyPrivate = false when the binding exposes the field, since WASM has no std::env::var and SsrfPolicy::from_env() always falls back to deny_private = true, which rejected the localhost mock-server requests every WASM e2e fixture targets.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.12") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.12")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 8d2b0d455eda725eeb10d9c9f4eb08e5c046b6bf75c9801bb3724551613dec0a
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.12/html-to-markdown-rs-zig-v3.6.12.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.12-QtXyWxY3AQBpMnJOX-eeLWy5vBwNKdKuM3KMEJDmT3tM\",\n },\n},\n```\n
.lib/<platform>/ so smoke tests link cleanly. The published Go binding declares // #cgo LDFLAGS: -L${SRCDIR}/.lib/<platform>/ -lhtml_to_markdown_ffi, but ${SRCDIR} resolves to the binding's source directory under GOMODCACHE, which is empty after go mod download. The package's own cmd/download_ffi is //go:build ignore-tagged and only invocable via the binding's own go generate, which test_apps doesn't run. Fix: test_apps/go/download_ffi.sh reads MODULE_VERSION from go.mod, downloads the matching html-to-markdown-rs-ffi-v<version>-<rust-triple>.tar.gz artifact from the GitHub release, caches it under ${XDG_CACHE_HOME:-~/.cache}/html-to-markdown-ffi/, makes the binding's module-cache subtree writable, and copies the library into <binding>/.lib/<platform>/. The smoke task invokes the script before go test. Restores Go to the registry-mode smoke matrix.upload-php-pie-release with cancelled/skipped guards so partial matrix success still publishes the available PIE archives. The job was gated on needs.php-extension.result == 'success', which meant a single transient matrix-cell failure (e.g. the ENOTFOUND artifact-upload flake observed on the v3.6.11 publish run) cascaded into skipping the entire PIE upload step. With the matrix already declared fail-fast: false and the in-job actions/download-artifact@v4 step configured with pattern: php-package-* + merge-multiple: true, the upload now proceeds whenever at least one matrix cell produced an artifact. Transient flakes no longer block the entire release surface.import { A as B } → CJS { A: B } rename translation (so oxfmt no longer aborts on Expected ',' or '}' but found 'as'); Go FFI copyLibraryToBindingPackage no-op when source aliases destination (avoids 0-byte libhtml_to_markdown_ffi.dylib after go generate); Swift Package.swift v__ALEF_SWIFT_VERSION__ placeholder substitution at scaffold-write time so SwiftPM consumers don't 404 against the release asset URL; Homebrew run_tests.sh formula-installed CLI preflight using the parameterized binary name; alef post-generation formatter pipeline aligned with downstream prek hooks (ktfmt for Kotlin format, gofmt+goimports for Go, oxfmt for TS/JSON, shfmt for shell scripts, php-cs-fixer for PHP, cargo sort for emitted Cargo.toml). Resolves CI Lint formatter divergence that had been red on every commit since v3.6.11.install.sh emits pie install --version "$VERSION" "<pkg>" instead of the deprecated pie install "<pkg>:$VERSION" form that PIE 1.4.5+ rejects with Unable to find an installable package <pkg> for version <ver>. Restores PHP to the registry-mode smoke matrix.[crates.dart] excluded_default_features / [crates.swift] excluded_default_features keep optional cargo features out of the wrapper's default = [...] array while still declaring them as opt-in forwarding entries, preventing target-conditional cross-compile activation of features whose system deps aren't cross-compile-ready.cargo update + cargo metadata on crates.io registry-index propagation lag (up to 6 attempts × 30 s) so per-language build jobs don't hard-fail in the first few minutes after Publish Rust crates completes.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.11") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.11")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 7ed93d95480c6f4fb8be896c4d0d681f824b4f7b0c05214be90808d1b00f5648
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.11/html-to-markdown-rs-zig-v3.6.11.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.11-QtXyW2U3AQDx1Y78dxpYAvaoqCWhHU00g15dLSF9wXWy\",\n },\n},\n```\n
Table layout pre-pass: skip nested-table rendering during column-width measurement (issue #406, residual fix). The MAX_CELL_WIDTH = 200 cap shipped in v3.6.10 (4013a6864) bounded the discarded output but not the measurement CPU. For deeply nested layout HTML (e.g. the reporter's Outlook digest: 393 <table> tags, 851 KB), cell_text_content still recursively dispatched handle_table_with_context on every nested table during the outer measurement pre-pass, triggering its own pre-pass on every descendant cell — combinatorial explosion (cells × nested_cells × ...) unbounded at greater nesting depth. The reproducer still ran for tens of minutes on v3.6.10. The fix threads measure_width_only: bool through the conversion Context; walk_node's "table" arm short-circuits to descendant text content when set, keeping the pre-pass linear in descendant character count. The reproducer now converts in 0.04 s. Regression test nested_layout_tables_convert_within_wall_clock_budget covers a synthetic 4×4 nested-cell fixture with a 10s wall-clock budget. Tier-1 vs Tier-2 separator-row dash counts now diverge on the nested-table fallback (Tier-1 still measures the rendered cell text); existing tier-divergence tests were updated to assert outer-row content equality instead of byte-equality. Resolves the residual #406 report from hobofan on 2026-06-16.
(via alef 0.25.19) Generated Swift app harness migrates fixtures JSON from triple-quote multi-line string literal to chunked-array [...].joined(), avoiding Swift's multi-line string literal content must begin on a new line error when the Jinja whitespace-trim placed content on the same line as the opening """.
(via alef 0.25.19) Generated FFI service-API codegen clones borrowed opaque-pointer params at the call site so consuming Rust APIs that take T by value receive an owned value while the C caller retains the original handle for its own _free call.
(via alef 0.25.19) Generated csharp e2e csproj template branches <RuntimeIdentifier> on OSArchitecture, picking osx-arm64 / linux-arm64 / win-arm64 on arm64 runners instead of hardcoded *-x64.
(via alef 0.25.19) Generated magnus binding Cargo.toml emits a cfg-feature forwarding [features] block so any #[cfg(feature = "X")] arms in the binding compile under -D warnings.
list[X] / dict, TypeScript X[] / Record, Go []X / map, etc.), applied via map_non_code_lines so code fences stay intact.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.10") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.10")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 4e0b4f17c6db1b49fefd2bfad7577e6e073722438419ade1b32f3bf49579ee1e
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.10/html-to-markdown-rs-zig-v3.6.10.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.10-QtXyW5A3AQD9bNUXwRl8QUCX273nTQ1FjyjdtRITStl5\",\n },\n},\n```\n
Table column-width: cap per-cell measurement at 200 characters (issue #406). Converting Microsoft-style HTML emails with hundreds of nested layout tables produced runaway output (73 MB / 2.2 s for the reporter's 851 KB reproducer) because collect_row_cell_widths used the rendered width of each cell — including any inner table's separator rows — as the outer column width. The outer table's separator was then emitted with that many dashes, which fed back into the grandparent's measurement, doubling at every nesting level. Capping the per-cell measurement at MAX_CELL_WIDTH = 200 bounds total output to O(cells × 200) and keeps both the Tier-1 fast path and the Tier-2 walker numerically consistent. The reproducer now converts in 0.47 s producing 218 KB of markdown. Bisected by the reporter to alef-driven Arc<Mutex<>> visitor migration in 3.4.1 (7f6178f25); the v3.6 series already fixed the visitor-mutex hot loop in 4863d1ab6, but the underlying exponential-width feedback remained. Regression test added: deeply_nested_layout_tables_do_not_produce_runaway_output.
Bump alef to 0.25.18. Picks up: (a) the e2e visitor codegen fixes for Node and C# — Node tests now pass { visitor: V as any } through options.visitor rather than as a silently-dropped third positional argument, C# tests merge into the existing new ConversionOptions() literal rather than appending a fourth arg, restoring every Custom-substitution visitor test. (b) The Zig e2e build.zig fix dropping the duplicated _run suffix in test-sequencing dependOn calls (conversion_run_run → conversion_run), unblocking Test: Zig. (c) The Swift box-delegate cast fix removing the spurious Int(...) wrap on usize/isize args so the bridge call matches the UInt protocol declaration, unblocking Test: Swift. (d) The Ruby Rakefile modifier-if fix for cross_compile_versions assignment, removing the Style/IfUnlessModifier rubocop offense that broke all 4 Test: Ruby matrix jobs. (e) The PHP PIE URL {OSLower} placeholder fix so pie install xberg-io/html-to-markdown resolves the correct asset name (resolves h2m #333). (f) The binding-crate feature passthrough block emission in FFI/Node/PHP/Wasm Cargo.toml — each binding crate now declares metadata, visitor, inline-images, testkit as passthrough features forwarding to html-to-markdown-rs/X, fixing the unexpected cfg condition value errors under RUSTFLAGS="-D warnings" for code paths gated by core features. (g) The swift-bridge owner-type extern block + cfg-gating fix preventing proc-macro expansion failures on cfg-disabled types. (h) The JNI trait-use emission fix so Tier-B Rust-public trait extension points are reachable from the generated JNI shims.
`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.9") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.9")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 386b88c8eb69a6e9e452d096bba8f5ee946e4e5c1302c8413bbe03ad321938c0
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.9/html-to-markdown-rs-zig-v3.6.9.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.9-QtXyW483AQB2QTeQXfPHxVUY89_eaHJfU-JGQdwcGEpe\",\n },\n},\n```\n
unreachable_patterns allow attribute for the crate-root, which was missing from the dart gen_rust_crate scaffold. v3.6.8's CI Rust failed with unreachable pattern errors at packages/dart/rust/src/lib.rs:1739 and :2102: the dart enum-conversion path emits an _ => unreachable!("cfg-gated variant ... not active in this build") catch-all so the match remains exhaustive when a #[cfg(feature = "X")]-gated variant is compiled out, but the dart binding's [features] table forwards testkit unconditionally, so the cfg-gated arm IS compiled in and the catch-all is unreachable — -D warnings turned the lint into an error. alef 0.25.17 adds unreachable_patterns to the existing #![allow(unused_variables, unreachable_code)] crate-root attribute, matching the swift backend's already-correct allow list.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.8") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.8")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 7157b8597555299de67406e321da830b8d1f9ab40e625d9b9c750dd6fe68522d
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.8/html-to-markdown-rs-zig-v3.6.8.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.8-QtXyW483AQAFezpI-LCw9upmkZNh29k25kjs5Qvc_Osx\",\n },\n},\n```\n
manifest_abs canonicalize fix (resolves Elixir NIF / PHP extension / Ruby gem / Python sdist alef publish prepare failures from the v3.6.7 publish run), the Dart cfg(testkit) whitespace fix (unblocks Validate: Rust and Test: Dart on main), the rustler .clone() skip for reference parameters, the JNI workspace-target host fix, the extendr enum type-path resolver, the swift From<core> cfg-gating for variant arms, the FFI visitor enum-context i32 emission, the visitor_result bare-string Custom routing, the e2e csharp SDK-AssemblyInfo suppression, prerelease set-version support (0.25.10–11); the swift cfg-union postprocessing fixes for default-build wrapper-type emission and the default = [<features>] Cargo.toml emission (0.25.12–15); and the drop of cfg propagation on enum From-impl match arms — binding crates do not declare gated features themselves but pull the core dependency with them enabled, so propagating #[cfg(feature = "testkit")] to binding-side arms produced unexpected cfg condition value: testkit errors under -D warnings on py/node bindings (0.25.16).<Version>. Deleted the hand-committed packages/csharp/HtmlToMarkdown/Properties/AssemblyInfo.cs (carrying a stale AssemblyVersion("3.4.0")) and switched the scaffold to let the .NET SDK derive AssemblyVersion, AssemblyFileVersion, and AssemblyInformationalVersion from <Version> plus the new <Company> / <Product> MSBuild properties. The published 3.4.0 assembly identity was breaking task test-apps:smoke:csharp against every NuGet package since the original ship.rewrite-native-deps. The workaround in 3.6.7 (a8c7fe40c) disabled rewrite-native-deps on PHP and Ruby build-* actions to dodge the alef 0.25.9 canonicalize bug. With 0.25.11's fix the action works correctly again, so the four rewrite-native-deps: "false" overrides are removed.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.6") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.6")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 82b01ca2391062fbf5ec7605b71cf45945798cab93d1ac0a4332ec4ca4832aaf
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.6/html-to-markdown-rs-zig-v3.6.6.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.6-QtXyWyUzAQDih9OVg4DwLQntyJ9TEXHr0hK1sjdMQcxj\",\n },\n},\n```\n
alef publish prepare for non-workspace-member NIF crates (Ruby gem, Elixir NIF). alef 0.25.4–0.25.6 had a bug where cargo update --locked and cargo metadata --locked would fail for binding crates that are not workspace members (like the Rustler NIF and Magnus gem). The seed lockfile lacked a [[package]] entry for the binding crate, causing cargo to reject it. alef 0.25.7 dropped --locked from the metadata validation step and set .env_remove("CARGO_BUILD_LOCKED") to allow cargo to resolve these crates from the registry. This fix restores Ruby (all platforms) and Elixir NIF (all platforms + Hex package) support.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.5") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.5")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 5ba69988fdcedb97a1a44a8314cfdc844d453d7f37e0ed57bfe4a690a4b71a00
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.5/html-to-markdown-rs-zig-v3.6.5.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.5-QtXyWyUzAQAevO8UVL_nNfNdQr5r_o-1WI6KQi4mqB2U\",\n },\n},\n```\n
Elixir Hex/NIF + Ruby gem publish: bump alef pin to 0.25.4 so alef publish prepare produces a valid binding lockfile. alef 0.25.1's vendor::scrub_or_regenerate_lock strict-mode path failed on every v3.6.4 Elixir NIF and macOS/Linux Ruby gem build with cargo update -p <lockfile> (or final cargo metadata validation) failed (exit code 101) ... cannot update the lock file ... because --locked was passed — the seed lockfile's workspace-member path entries collided with the registry-source entries the rewrite added, and the final cargo metadata --locked validation could not reconcile them. alef 0.25.3 added strip_workspace_member_entries to drop the path-source entries from the seed before per-member cargo update -p runs, plus a full registry-URL package-id spec (registry+https://github.com/rust-lang/crates.io-index#NAME@VERSION) to disambiguate the per-member update. 0.25.4 also escapes Rust reserved keywords in extendr struct fields (relevant to R bindings consuming flat data enums with serde(tag = "type")).
Python sdist on Alpine/musl: actions/rewrite-native-deps v1.8.69 now strips path = "..." from [workspace.dependencies] entries. The 3.6.4 sdist's root Cargo.toml shipped with [workspace.dependencies] html-to-markdown-rs = { version = "3.6.4", path = "crates/html-to-markdown" }, and cargo eagerly validates every workspace-dep entry on pip install — bailing with failed to read .../crates/html-to-markdown/Cargo.toml because the workspace crate is not bundled in the sdist. v1.8.66 added [patch.*] stripping for sdist consumers (resolved #390) but missed [workspace.dependencies]. The new step drops path from every workspace-dependency entry, leaving the version so the dep resolves from crates.io on consumer install. Resolves #402.
`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.4") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.4")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 9f4ea09b97c299e2ea9d9f02492e761939ce7973fc1b8c3e58126bc80f5da502
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.4/html-to-markdown-rs-zig-v3.6.4.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.4-QtXyWyUzAQDUyryV0fgsWUTgIpA_Eb2eerg7cNK8EZnB\",\n },\n},\n```\n
CLI: drop reqwest brotli feature to keep alloc-no-stdlib on 2.x. The transitive brotli crate (pulled in via async-compression) jumped to 8.x which expects alloc-no-stdlib 3.x, but the workspace resolves the sibling at 2.x. The resulting trait drift broke cargo build on stable rustc with the trait bound 'StandardAlloc: alloc::Allocator<...>' is not satisfied, blocking the ubuntu-latest Python wheel build during the v3.6.3 publish run. The CLI's HTTP client now uses gzip+deflate only.
Publish workflow: dispatch publish-pubdev with the release tag, not the branch ref. When publish.yaml was triggered by release: published, github.ref_name resolved to the branch where the release was authored (e.g. main). The child workflow's OIDC token then carried refType=branch, which pub.dev rejects with publishing is only allowed from 'tag' refType. Dispatch now uses needs.prepare.outputs.tag so the child runs against the tag and the OIDC token is accepted.
Taskfile: test-apps:test:zig now grep'd the root Cargo.toml for the workspace version. The previous grep targeted crates/html-to-markdown/Cargo.toml, which uses version.workspace = true and has no literal version line — H2M_VERSION resolved to empty and the zig-fetch URL collapsed to …/download/v/html-to-markdown-rs-zig-v.tar.gz (404).
Track Cargo.lock for reproducible builds. Cargo.lock was previously gitignored, defeating cargo build --locked everywhere it was used. Every CI runner resolved deps fresh, allowing semver-compatible drift to silently introduce broken transitives (the brotli/alloc-no-stdlib mismatch above). The root workspace lockfile plus the per-package nested lockfiles (R, Ruby, Elixir NIF, e2e/rust, test_apps/rust) are now tracked.
Pass --locked to cargo build in CI, publish, scripts, and alef.toml hooks. Now that the lockfile is tracked, every CI/publish cargo build runs with --locked so a dirty index can't silently substitute newer transitive deps. Local-dev tasks (.task/languages/rust.yml) intentionally remain without --locked so contributors can pick up dep updates.
Regenerated cross-language bindings with alef 0.25.1. Pulls in: csharp e2e codegen using KreuzbergConverter facade; swift e2e codegen no longer emits ? chains on non-Optional RustString returns; C e2e codegen panics on missing fields_c_types keys instead of silently miscompiling; FFI/NAPI/PyO3/Magnus/Rustler test surface cleanup; vendor::scrub_or_regenerate_lock preserves workspace lockfile pins via per-member cargo update + cargo metadata --locked validation; NAPI strips the readonly keyword from emitted service.cjs.
Python `.pyi`: `ConversionOptions.visitor` attribute type now resolves to `HtmlVisitor | None`. The class-attribute annotation pass in alef's pyo3 stu
Python .pyi: ConversionOptions.visitor attribute type now resolves to HtmlVisitor | None. The class-attribute annotation pass in alef's pyo3 stub emitter previously printed the opaque VisitorHandle concrete type for Option<dyn HtmlVisitor> fields, while __init__ parameters and convert(...) already resolved to the HtmlVisitor Protocol. Field-style assignment (options.visitor = MyVisitor()) is now type-clean under pyright/pylance. Closes #403.
NAPI convert: visitor callbacks now fire when passed via options.visitor. The generated convert(html, options) shim previously declared a third standalone visitor parameter and used it directly while ignoring options.visitor. The TypeScript surface advertises a single uniform entry point — convert(html, { visitor: { … } }) — so visitor callbacks routed through options were silently dropped. Fixed in alef v0.24.16; the standalone parameter is gone and options.visitor is the sole source. Closes #395.
NAPI options-field bridge: drop unused mut on the closure binding. The generated convert(html, options) shim's Option<JsConversionOptions>.map(|o| …) closure moved o straight into o.into() without mutating the binding, so mut o triggered warn(unused_mut) on crates/html-to-markdown-node/src/lib.rs. Fixed in alef v0.24.17 by dropping mut from the napi-side closure binding; the wasm-side template is unchanged because it does mutate o.visitor before the into() call.
Ruby: precompiled-gem ABI fallback to source-gem. The cross_compile_versions list in the Magnus Rakefile template now targets Ruby 3.5/3.4/3.3/3.2 (dropping 3.1, adding 3.5), and the gemspec required_ruby_version bounds the upper edge with >= 3.2.0, < 4.0. On Ruby 4.0+ / 4.1.0dev, RubyGems now refuses the precompiled platform gem and falls back to the source gem, eliminating incompatible ABI version of binary load failures on prerelease ABIs. Closes #405, #409.
C# / Kotlin Android: wrapper class renamed HtmlToMarkdownRs → HtmlToMarkdownConverter. The previous …Rs suffix was Rust-implementation-bleed in the public binding surface. The new name is idiomatic for both ecosystems and matches the names already used in our hand-written docs. BREAKING for downstream C# / Kotlin-Android consumers: apply s/HtmlToMarkdownRs/HtmlToMarkdownConverter/g to source code that imported or invoked the wrapper. Closes #408.
PHP PIE: macOS install resolves the extension at the archive root. pie install xberg-io/html-to-markdown on macOS arm64 previously failed because the staged extension was named html_to_markdown.so while PIE's UnixBuild probes for <extname>.<dylib_ext> on macOS. The publish workflow now stages the extension as html_to_markdown.dylib on macOS targets and .so on Linux, restoring PIE installs on Apple Silicon and Intel Macs. Closes #334.
Python wheel: HtmlVisitor trait-bridge runtime-import resolved. alef's pyo3 trait-marker-class emitter creates _Trait{Name}Marker types in the DTO module (visitor.rs) to satisfy TypeScript/Ruby/etc. downstream trait surface generation. The PyO3 bridge uses these markers as import-path sentinels, but h2m's ConversionOptions visitor field only became public in v3.6.1 after its prior conditional #[cfg(feature = "visitor")] gate was removed. Wheels shipped without the marker import working, causing ImportError: cannot import name '_TraitHtmlVisitorMarker' at runtime when user code called convert(..., visitor). Fixed in alef v0.21.0 (marker-import path resolved from the actual DTO module).
Java JAR: NativeLib RID alignment with publish-workflow classifier names. The publish workflow (publish.yaml) uploads native libraries with RID classifiers: osx-aarch64, osx-x86_64, linux-x86_64, linux-aarch64, windows-x64. The NativeLib.scala.kt loader looked for osx-arm64 and osx-x86-64 (mismatched case and infix), breaking native library resolution at JAR initialization. Fixed in alef v0.20.5; RID names now match the publish assets exactly.
R tarball: GitHub-release install fixed. The CRAN tarball sources from GitHub release tags when users install from .tar.gz source (non-CRAN edge case). The configure script ran cargo with CARGO_HOME=... override to redirect vendor offline access, but cargo vendor itself does not honor CARGO_HOME (only cargo build does), so the offline build failed with "could not find html-to-markdown in the registry". extendr 0.18.1 correctly passes --registry crates-io --offline through to cargo vendor, restoring offline builds. The configure script also fixed inbound str_as_str imports from extendr that the v0.18.0 landing initially broke.
Homebrew: brew trust prerequisite documented. Homebrew 6.0+ (released 2026-05) requires brew trust xberg-io/tap before installing from the third-party tap. Updated docs/installation.md, README.md, and test_apps/homebrew/run_tests.sh with the trust step and explanatory note. Smoke test now calls brew trust "$TAP" || true before brew bundle install, making CI-tested on Homebrew 6.0+ environments.
date-released: auto-stamped. alef sync-versions --release-date YYYY-MM-DD now stamps the release date directly into the workspace-generated CITATION.cff, eliminating the prior hand-edit step at release-cut time. The release date for v3.6.2 is 2026-06-13. Closes #327.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.1") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.1")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 14c8c053556462a759750ddbe4cdad68c042237c6891261a9fe927150028c58e
<!-- zig-fetch -->
Add to your build.zig.zon:
.dependencies = .{
.html-to-markdown-rs-zig = .{\n .url = \"https://github.com/kreuzberg-dev/html-to-markdown/releases/download/v3.6.1/html-to-markdown-rs-zig-v3.6.1.tar.gz\",\n .hash = \"html_to_markdown_rs-3.6.1-QtXyW_UxAQBfzcbtn8qRyDF4Damjof-cx9rh30fzM4hk\",\n },\n},\n```\n
Ruby gem: Rakefile reverted to v3.5.5 working pattern. v3.6.0's alef 0.21.0 regen introduced RbSys::ExtensionTask with manual Dir.chdir nesting that broke nested file "Cargo.lock" task declarations. Reverts to the v3.5.5 Rake::ExtensionTask pattern for all 5 platforms (macOS arm64/x86_64, Linux x86_64/aarch64, Windows x64). Resolves bundle exec compilation failures.
Zig: curl HTTP/2 support. Zig's vendored curl bindings required +h2 feature flag for HTTPS URL fetch operations. Updated packages/zig/build.zig.zon to enable the feature, restoring GitHub-release downloads on Zig consumers.
Homebrew bottle ordering stabilized. alef v0.20.5 sort order for bottle entries now matches brew audit expectations; formulae no longer fail the Homebrew official-tap checklist.
`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.0") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.6.0")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: cba4c2ab228670451e3bc59687f4a19d0aca1fd9f22c61336c4afc34cd03b66c
Tiered HTML-to-Markdown conversion architecture. Clean HTML inputs now have an opt-in
fast path through a Tier-1 single-pass byte scanner (converter/tier1/); on anything the
scanner cannot prove byte-equivalent to the existing Tier-2 DOM walker, it returns a
structured bail and the dispatcher falls back to Tier-2 (tl::parse + walk_node)
transparently. Output is always byte-equal to what Tier-2 would have produced for the
same input — verified by 116 oracle snapshots and a per-fixture byte-equality integration
test (tests/tier1_byte_equality_test.rs). Tier-3 (html5ever repair) remains the
fallback for truly malformed HTML.
ConversionOptions::tier_strategy (TierStrategy::Auto | Tier2 | Tier1)
for runtime tier selection. Auto (default) lets the classifier decide based on the
prescan signals and options shape; Tier2 forces the existing DOM-walk path; Tier1 is
testkit-only (#[cfg(any(test, feature = "testkit"))]) for debugging and benchmarking.
Exposed to Node (tierStrategy: "auto" | "tier2") and Wasm (WasmTierStrategy enum).
CLI and Python pick up the field via the standard options-mapping flow.
Tier-1 router style-option gate. The classifier forces Tier-2 when any of these
deviate from Tier-1's hardcoded value: heading_style, code_block_style,
strong_em_symbol, bullets, list_indent_*, whitespace_mode, newline_style,
escape_* flags, output_format, link_style, url_escape_style, compact_tables,
default_title, sub_symbol, sup_symbol, highlight_style. 47 integration tests
enforce the gate.
Tier-1 conservative bail set. Tier-1 returns Err(BailReason::*) rather than emit
potentially-wrong output on: custom elements, CDATA, unescaped <, inline SVG, HTML
5 optional-close edge cases, <table> with rowspan/colspan/block-children-in-cells/
caption/mixed-section-order, nested lists (Tier-2 cycles bullets by depth), <pre>
with non-Indented code_block_style, named HTML entities outside the 45-entry
zero-alloc table, table cells containing |, and <br> inside table cells. Tests
cover every variant.
Benchmark harness (tools/benchmark-harness/, binary htmbench) with
run / compare / oracle / oracle:bless / survey / mdream subcommands.
Per-group regression guardrails (baselines/baseline.json, guardrails.json).
Wired under task bench:* namespace.
Bench regression gate in CI — every PR that touches the Rust core or bench harness now
runs task bench:oracle && task bench:run && task bench:compare on ubuntu-24.04-arm,
failing on any fixture that exceeds its per-group threshold in
tools/benchmark-harness/guardrails.json (5% clean_large, 8% clean_medium, 10% other groups,
30% adversarial). Baselines are blessed deliberately by humans, not automatically by CI.
Publish workflow now authenticates via the kreuzberg-dev-publisher GitHub App (org-level
BOT_APP_ID / BOT_APP_PRIVATE_KEY) for all writes (own-repo pushes, tag updates, release-asset
uploads, homebrew-tap commits, Discord dedup marker). OIDC publisher jobs (PyPI/npm/hex/Maven/NuGet/crates)
are unaffected. HOMEBREW_TOKEN is no longer used.
BREAKING: NodeContext::attributes is no longer a public field. Access attributes via
the new attributes() method (fn attributes(&self) -> &BTreeMap<String, String>).
Struct-literal construction of NodeContext is no longer possible; use the provided
constructors: with_borrowed_attributes, with_owned_attributes (public), or
with_lazy_attributes (pub(crate), for internal use only). Attributes are now
lazily materialized from the DOM on first access when constructed via the internal path,
eliminating the per-element BTreeMap allocation on the visitor hot path. Measured
throughput improvement with visitor enabled (100-iteration harness, Apple Silicon):
small_html −29%, medium_python −10%, large_rust −24%, tables_countries −24%.
DomContext::parent_tag_name now returns Option<&str> instead of Option<String>.
This is an internal API (pub(crate)); external consumers are unaffected.
Tier-1 byte scanner activated in production (tier_strategy = Auto). The classifier
decides per-input whether Tier-1's single-pass byte scanner runs; on bail, the dispatcher
falls back to Tier-2 transparently. Measured throughput on the harness corpus (29 fixtures,
6.4 MB total; Apple Silicon, cargo build --release):
| Fixture | Size | ms (best) | Throughput |
|---|---|---|---|
real-world/wikipedia/medium_python.html |
1.24 MB | 62.58 ms | 19.0 MB/s |
real-world/wikipedia/large_rust.html |
1.07 MB | 37.17 ms | 27.3 MB/s |
mdream/github-markdown-complete.html |
430 KB | 10.57 ms | 38.7 MB/s |
mdream/react-learn.html |
265 KB | 12.11 ms | 20.9 MB/s |
mdream/wikipedia-small.html |
166 KB | 5.63 ms | 28.1 MB/s |
real-world/issues/gh-121-hacker-news.html |
57 KB | 1.08 ms | 50.3 MB/s |
mdream/nuxt-example.html |
3.6 KB | 0.029 ms | 116.1 MB/s |
Per-group regression thresholds (5–30%) are enforced on every PR via task bench:compare (see Added section).
memchr-driven text scan: decode_and_collapse_into and decode_entities_into use
memchr::memchr3 / memchr::memchr to skip ahead to the next special byte (<, &,
whitespace boundary) and bulk-copy plain text runs in a single push_str. Replaces a
byte-by-byte conditional inner loop and closes a substantial portion of the gap to main's
heavily-optimized Tier-2 path on Wikipedia-scale documents (e.g., +32% on wikipedia-small).
Tier-1 dispatcher reuses normalized input on Tier-2 fallback. When the Tier-1 scanner
bails or the classifier routes to Tier-2, the dispatcher threads the already-computed
normalize_input Cow<str> through to the Tier-2 path instead of recomputing it. Eliminates
one full-input pass on every bail-fallback or Tier-2-routed call.
collapse_excess_blank_lines in-place. Replaced the fresh-String rewrite with
String::retain, eliminating an output-sized allocation on every Tier-1 success whose
output contains \n\n\n. Measured wins (best-of-3, --force-tier1, vs prior commit):
wikipedia/lists_timeline -3.81%, gh-121-hacker-news -4.23%, github-markdown-complete
-2.99%, wikipedia/medium_python -2.94%, mdream/wikipedia-small -2.68%, mdn-array
-2.12%, gh-190/firsteigen -2.55%. No fixture regressed.
htmbench --force-tier2 flag mirroring --force-tier1, for clean head-to-head
benchmarks now that the Auto router activates Tier-1 on most fixtures and so cannot be
used as a Tier-2 control.
convert() accepts options as a bare ConversionOptions in addition to
Option<ConversionOptions> (resolves #398). The second parameter now bounds
impl Into<Option<ConversionOptions>>, so convert(html, opts),
convert(html, Some(opts)), convert(html, None), and
convert(html, ConversionOptions::default()) are all valid Rust call shapes.
Existing callers continue to compile unchanged — this is purely additive
flexibility, not a breaking change. The ~250-line conversion body is held in a
private non-generic convert_inner so the generic wrapper monomorphises exactly
once per call site rather than per Into impl chosen.
url_escape_style option (UrlEscapeStyle::Angle | UrlEscapeStyle::Percent). When set
to Percent, link and image destinations are percent-encoded instead of wrapped in angle
brackets, producing output that all Markdown parsers handle correctly even when the URL
contains <, >, spaces, or parentheses (resolves #392).
Non-deterministic SVG attribute serialization in converter/media/svg.rs:
serialize_element iterated tag.attributes() over astral-tl's internal HashMap,
which has non-deterministic iteration order. SVG data: URIs therefore differed across
runs. Fixed by sorting attributes by name before emission, restoring determinism.
spurious blank lines after frontmatter and lists (MD012) (resolves #399). Block-level emission now collapses runs of three or more consecutive newlines into exactly two, so the frontmatter→body and list→next-block transitions no longer produce extra blank lines that violate markdownlint MD012.
autolinks: bare paths and filenames are no longer wrapped as autolinks (resolves #397).
Per GFM §6.5, autolinks require an absolute URI with a scheme — but the previous check only
compared the link text to the href, so <a href="foobar.png">foobar.png</a> became the
invalid <foobar.png> (which parsers read as a literal HTML tag). Added has_uri_scheme
helper that validates the RFC 3986 scheme grammar (ALPHA *( ALPHA / DIGIT / "+" / "-" / "." )
followed by :). Bare paths, fragments, and filenames now render as [text](href). https://,
mailto:, ftp://, data:, and other schemed URLs continue to autolink as before.
node binding: visitor callbacks (visitText, visitLink, visitHeading, …) now fire
(resolves #395). JsHtmlVisitorBridge now stores a persistent napi::bindgen_prelude::ObjectRef<false>
obtained via Object::create_ref() and materializes a fresh local Object handle per callback
through obj_ref.get_value(&env). The previous bridge stored raw napi_value pointers extracted
via transmute_copy and reconstructed via Object::from_raw — but a napi_value is a local handle
tied to the HandleScope active at construction time, and by the time visitor methods fired deep
inside convert() the scope was no longer active, so get_named_property("visitText") silently
returned Err and the bridge fell through to VisitResult::Continue without dispatching. A
Drop impl on the bridge calls obj_ref.unref(&env) so the JS object can be GC'd after
conversion. Verified against the issue repro: visitText/visitLink/visitHeading all fire.
(alef commit 1ffdaafe4)
code block: CodeBlockStyle::Backticks and Tildes no longer emit a trailing blank
line inside the fence and now insert a blank line after the closing fence (resolves #396).
The fenced emitter pushed content verbatim (which already ends in \n for any trailing-newline
source) plus an extra \n before the closing fence, producing …\n\n```\n. It also closed with a
single \n, so following block content butted up against the closing fence with no blank line
separator — Indented style was unaffected because its content path strips trailing newlines
via .lines().join("\n") and already emits \n\n after. The Backticks/Tildes path now trims
trailing newlines from inner content (content.trim_end_matches('\n')) before re-emitting a
single \n, and pushes \n\n after the closing fence to match Indented.
php test_app: PIE install.sh now installs xberg-io/html-to-markdown instead of
the non-existent xberg-io/html-to-markdown-rs (resolves #98 / smoke regression). The
Packagist project only publishes the un-suffixed name; alef.toml was renamed and the regen
picks it up in both install.sh and the inline comment.
r binding: conversion_options() helper exported in NAMESPACE. alef's R emitter
generated the R/options.R helper for ergonomic ConversionOptions construction but never
exported it, so htmltomarkdown::conversion_options(...) raised could not find function
at runtime — which downstream convert(..., options=) callers then surfaced as the
extendr-api 0.9.0 options must be a named list validation error. Resolves #99 / smoke
regression. (alef commit 160d504f3)
java binding: NativeLib.<clinit> no longer throws NoSuchElementException when an
optional FFI symbol is missing. The error-context handle lookups (LAST_ERROR_CODE,
LAST_ERROR_CONTEXT) used Optional.orElseThrow() mid-static-init, escalating any partial
symbol set into a hard NoClassDefFoundError; they now fall back to null and the rest of
the loader's null-checks handle the absence gracefully. Resolves #100 / smoke regression.
(alef commit d296bfdeb)
c FFI test_app Makefile now embeds an LC_RPATH so the smoke binary can find the
dylib at runtime on macOS without DYLD_LIBRARY_PATH. Adds -Wl,-rpath,$(FFI_LIB_DIR)
to the LDFLAGS path. Resolves #101 / smoke regression.
csharp binding: trait-bridge catch blocks no longer emit unused Exception ex variables.
The generated TraitBridges.cs had 25+ catch (Exception ex) blocks where ex was never
referenced; with <TreatWarningsAsErrors>true</TreatWarningsAsErrors> the build failed
with CS0168, which in turn blocked NuGet publish for 3 releases (root cause of #104). alef
now emits bare catch (Exception) when the exception variable isn't actually used.
(alef commit 16e81b20d)
.npmrc at workspace root: minimum-release-age=0 so the npm Build Node bindings and
Build WASM package jobs stop failing at pnpm install time with
ERR_PNPM_MINIMUM_RELEASE_AGE_VIOLATION for transitively-recent dep pins. The v3.5.5 fix
only patched the test command; this covers every pnpm install call repo-wide. Resolves
#96.
Go subtag publish (packages/go/v3.5.x) is already wired through the
xberg-io/actions finalize-release job (publish.yaml:2418 passes
go-module-path: packages/go/v3); no action needed here. The reason v3.5.2–v3.5.7 didn't
push subtags is upstream main-release-job failures cascading. Resolves #97.
`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.5.7") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.5.7")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 12a52b6c918a2647573aac50b22edc87ed8de3fee4f45a54d1de9c0180d45646
Cargo.toml version. v3.5.6's release commit (2c37942f8) claimed to bump the workspace version 3.5.5 → 3.5.6 but the change never actually landed in the commit. Every binding manifest (Cargo.toml, pyproject.toml, gemspec, mix.exs, pom.xml, .csproj, package.json, …) stayed pinned at 3.5.5, so the v3.5.6 publish workflow built v3.5.5 binaries against the v3.5.6 tag and uploaded them as *-v3.5.5-*.tar.gz assets to the v3.5.6 GitHub Release. crates.io / PyPI / Hex / RubyGems / Maven publishes were either no-ops (3.5.5 already exists) or rejected; npm + WASM + NuGet + Homebrew steps also failed. v3.5.6 is skipped; v3.5.7 is the corrected ship of every fix originally rolled into v3.5.6.alef sync-versions).`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.5.5") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.5.5")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 1d3bc13444ca07cf0170dcec43f93725b7136bf14c321f66825e584ea9fabbdf
windows-x64 gem via bundle install inside the rb-sys-dock container. The previous matrix entry ran rb-sys-dock --platform ... -- bundle exec rake "native[$target]" gem, which invoked bundle exec inside the container without first materialising the lockfile against the container's Ruby version. The rb-sys-dock container ships its own Ruby toolchain (4.0.2 preview) while the host Gemfile.lock pins 3.3.11, so bundler failed with "Could not find …" and hung for 47 minutes. The command now wraps the container invocation in bash -c "bundle install --jobs=4 --retry=3 && bundle exec rake …" so the lockfile materialises against the container's Ruby before rake runs. (.github/workflows/publish.yaml).dylib.tar.gz to "match the platform-native extension," but rustler_precompiled 0.9.0 (the latest version on Hex; no .dylib-aware version exists) hardcodes .so for every non-Windows consumer download URL in lib_name_with_ext/2 and ignores all caller overrides. h2m's hand-maintained publish.yaml had already normalised darwin uploads to .so, so h2m's own Hex publish kept working; the bug only surfaced in sibling polyglot repos. The alef v0.20.5 revert (.dylib → .so) restores polyrepo alignment and is picked up by this regen. (alef.toml pin → v0.20.5, xberg-io/actions/generate-elixir-checksums@v1)download_ffi.sh now hits the correct release asset prefix. The [crates.e2e.registry.packages.c] name was html-to-markdown-ffi, but publish.yaml uploads C FFI tarballs with asset-prefix: html-to-markdown-rs-ffi-. Drift caused task test-apps:smoke:c to 404 on the GH release tarball. (alef.toml)mvnw, mvnw.cmd, .mvn/wrapper/maven-wrapper.properties). Previously missing because mvnw emission landed in alef v0.20.2 and h2m was pinned at v0.20.1. (alef.toml pin → v0.20.5)*.sh (e.g. run_tests.sh, download_ffi.sh) is now created with the +x bit set. alef's write_scaffold_files_with_overwrite() skipped the shebang-chmod helper that write_files() applied, so every alef-generated shell script in e2e/ and test_apps/ landed as -rw-r--r--, breaking task test-apps:smoke:homebrew with "permission denied" on ./run_tests.sh. Fixed by alef v0.20.3's shared apply_shebang_chmod() helper called from both writers. (alef.toml pin → v0.20.5)ffi_skip_methods honouring in I{Trait} interface emission. (alef.toml pin → v0.20.5)packages/elixir/native/html_to_markdown_nif/Cargo.toml)Handle::try_current() before WORKER_RUNTIME.block_on(...) so the bridge is safe to call from within an outer tokio runtime. (alef.toml pin → v0.20.5)--config.minimumReleaseAge=0 to pnpm test so freshly-published RC packages clear pnpm's supply-chain policy check. (alef.toml pin → v0.20.5)packages/java/pom.xml)crates/html-to-markdown-cli/Cargo.toml)package.json)mix hex.outdated non-zero exits before mix deps.update --all. (.task/languages/elixir.yml)`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.5.4") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.5.4")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: 7a339c0bd84922788404b09d6b82abf3bf6b952c613ec084d464b332198f83e9
index.d.ts after napi build in the node-typescript-defs publish job. The napi-rs macros deliberately filter out the *Update DTO types (ConversionOptionsUpdate, PreprocessingOptionsUpdate), so the d.ts regenerated by napi build lacks them; the wrapper at packages/typescript/src/index.ts re-exports those names and tsc fails with TS2724 on every publish. The publish job now git checkout HEAD -- crates/html-to-markdown-node/index.d.ts between napi build and the artifact copy, preserving alef's augmented d.ts.windows-x64 gem via rake-compiler-dock from the Linux runner. The previous matrix entry's cargo build --target x86_64-pc-windows-gnu against the ucrt64 MSYS2 toolchain corrupted a linker argument and failed every release with cannot find -l■. Linux-host docker-based rake-compiler-dock is rb-sys's recommended cross-build path and matches what every other rb-sys gem uses for Windows.RustlerPrecompiled now resolves the correct platform extension on macOS. v0.20.0's rustler emit lets RustlerPrecompiled's built-in mapping pick .dylib for *-apple-darwin targets and .so for *-linux-*, so mix deps.get no longer 404s against *.so.tar.gz on darwin consumers.alef:hash: now computed over generation inputs, not emitted content. Post-format whitespace drift no longer invalidates alef verify, so CI Lint stops reporting bindings as stale after every prek --all-files cycle.-Djava.library.path=.../target/release in registry dep_mode. The Maven-Central JAR bundles natives under /natives/{rid}/ and alef's loader extracts them at startup; overriding the library path to a non-existent local directory broke UnsatisfiedLinkError on smoke runs. Override is now emitted only in local dep_mode.target/release/ build to resolve the FFI. download_ffi.sh is skipped only when both the FFI header and shared library are present locally; otherwise the published tarball is fetched. Smoke runs from a clean checkout work without a prior cargo build.install.sh before composer install. Removes the implicit "ext-html_to_markdown must be pre-installed" requirement.name() dedup.`swift .package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.5.3") `
Add to your Package.swift:
.package(url: "https://github.com/kreuzberg-dev/html-to-markdown", from: "3.5.3")
The Swift binding is distributed as a pre-built artifact bundle. No Rust toolchain required.
Artifact: HtmlToMarkdown-rs.artifactbundle.zip
Checksum: a8a6266351763f619dc24bee592c98a989e215ee9a8316758c71b0b330d4616d
<td/>, <br/>, etc.) embedded in EPUB-derived HTML. The bundled astral-tl parser treats / as an identifier character, so <td/> was parsed as a tag literally named "td/" and subsequent siblings nested under it — table rows collapsed into a single cell and any content after the broken table was dropped. convert_api::normalize_input now preprocesses <tag/> → <tag /> before tokenization so the trailing slash is read as a self-closing marker. Regression covered by tests/issue_391_xhtml_self_closing.rs. Closes #391.docs/reference/api-python.md documented a class MyVisitor(VisitorHandle): subclass pattern with VisitResult.Continue enum returns, but VisitorHandle is #[pyclass(unsendable, from_py_object)] (sealed; cannot be subclassed) and VisitResult exposes only property getters (VisitResult.Continue raises AttributeError). The trait rustdoc on HtmlVisitor now leads with the plain-class + string/dict return idiom that the magnus-style bridge actually accepts. Closes #389.crates/html-to-markdown/ directory because a stale .gitmodules entry pointing at a deleted homebrew-tap submodule caused cargo package --list to fail silently during the CI maturin run, so maturin only bundled the PyO3 crate and shipped a sdist that could not build. Removed the stale .gitmodules; maturin's path-dependency auto-bundling now works again and the resulting sdist contains all 177 core-crate files. Closes #390.html-to-markdown-rs-swift. Stale html-to-markdown-swift reference broke the Swift bundle step on every release since v3.5.0.[tool.maturin] include directive (maturin rejects archive paths with ..).Your coding agent can read these notes before it upgrades. Set up the MCP server →