NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #866 most downloaded on PyPI
Convert Word documents from docx to simple and clean HTML and Markdown
Last release 8 days ago
26 Sep 2026
Release timing varies
gaps range from 2 weeks to 9 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
13 years old
86 releases · first in 2013
Read the children of w:customXml elements.
Read the children of w:customXml elements.
Handle paragraphs and runs that have been moved and tracked as a revision.
Improve escaping in Markdown writer. This avoids some cases where source documents could inject arbitrary HTML into the converted document.
Avoid excessive backtracking when parsing an unterminated string with many escape sequences. The previous behaviour would allow maliciously crafted do
Avoid excessive backtracking when parsing an unterminated string with many escape sequences. The previous behaviour would allow maliciously crafted documents to cause a denial of service.
Note that it is still strongly recommended to process untrusted documents in a separate thread with a timeout to avoid potential similar issues.
Handle complex field separator and end characters without corresponding start characters.
One column per quarter.
Fix: on Windows, when an image's content type includes a backslash in the subpart, files may be written outside of the directory set by --output-dir.
Fix: on Windows, when an image's content type includes a backslash in the subpart, files may be written outside of the directory set by --output-dir.
Detect and ignore numbering levels that use numStyleLink to refer to themselves.
Handle hyperlinked wp:anchor and wp:inline elements.
Handle hyperlinked wp:anchor and wp:inline elements.
Handle hyperlink complex fields with unquoted hrefs.
Ignore style definitions using a style ID that has already been used.
Ignore style definitions using a style ID that has already been used.
Fix conversion of unmerged table cells.
Disable external file accesses by default. External file access can be enabled using the external_file_access argument.
Handle numbering levels defined without an index.
Add "Heading" and "Body" styles, as found in documents created by Apple Pages, to the default style map.
Add "Heading" and "Body" styles, as found in documents created by Apple Pages, to the default style map.
Handle structured document tags representing checkboxes wrapped in other elements, such as table cells. Previously, the wrapping elements would have been ignored.
Ignore deleted table rows.
Add notes on security.
Ignore AlternateContent elements when there is no Fallback element.
Detect checkboxes, both as complex fields and structured document tags, and convert them to checkbox inputs.
Detect checkboxes, both as complex fields and structured document tags, and convert them to checkbox inputs.
Ignore AlternateContent elements when there is no Fallback element.
Add style mapping for highlights.
Switch the precedence of numbering properties in paragraph properties and the numbering in paragraph styles so that the numbering properties in paragr
Support attributes in HTML paths in style mappings.
Support attributes in HTML paths in style mappings.
Improve error message when failing to find the body element in a document.
Drop support for Python 2.7, Python 3.5 and Python 3.6.
Add support for the strict document format.
Support merged paragraphs when revisions are tracked.
Add a pyproject.toml to add an explicit build dependency on setuptools.
Only use the alt text of image elements as a fallback. If an alt attribute is returned from the function passed to mammoth.images.img_element, that va
Ignore w:u elements when w:val is missing.
Emit warning instead of throwing exception when image file cannot be found for a:blip elements.
When extracting raw text, convert tab elements to tab characters.
When extracting raw text, convert tab elements to tab characters.
Handle internal hyperlinks created with complex fields.
Handle w:num with invalid w:abstractNumId.
Convert symbols in supported fonts to corresponding Unicode characters.
Support numbering defined by paragraph style.
Add style mapping for all caps.
Handle underline elements where w:val is "none".
* Read font size for runs. * Support soft hyphens.
Update supported Python versions to 2.7 and 3.4 to 3.8.
Improve list support by following w:numStyleLink in w:abstractNum.
* Preserve empty table rows.
Always write files as UTF-8 in the CLI.
Fix: default style mappings caused footnotes, endnotes and comments containing multiple paragraphs to be converted into a single paragraph.
Read the children of v:rect elements.
Read part paths using relationships. This improves support for documents created by Word Online.
Parse paragraph indents.
Read part paths using relationships. This improves support for documents created by Word Online.
Add style mapping for small caps.
Add style mapping for small caps.
Add style mapping for tables.
Read children of v:group elements.
Read w:noBreakHyphen elements as non-breaking hyphen characters.
Extract the default data URI image converter to the images module.
Extract the default data URI image converter to the images module.
Add anchor on hyperlinks as fragment if present.
Convert target frames on hyperlinks to targets on anchors.
Detect header rows in tables and convert to thead > tr > th.
Handle complex fields that do not have a "separate" fldChar.
* Add transforms.run.
Read children of w:object elements.
Read children of w:object elements.
Add support for document transforms.
Handle hyperlinks created with complex fields.
Handle absolute paths within zip files. This should fix an issue where some images within a document couldn't be found.
Allow style names to be mapped by prefix. For instance:
Allow style names to be mapped by prefix. For instance:
r[style-name^='Code '] => code
Add default style mappings for Heading 5 and Heading 6.
Allow escape sequences in style IDs, style names and CSS class names.
Allow a separator to be specified when HTML elements are collapsed.
Add include_embedded_style_map argument to allow embedded style maps to be disabled.
Include embedded styles when explicit style map is passed.
Ignore bold, italic, underline and strikethrough elements that have a value of false or 0.
Ignore v:imagedata elements without relationship ID with warning.
Use alt text title as alt text for images when the alt text description is blank or missing.
Handle comments without author initials.
Handle comments without author initials.
Change numbering of comments to be global rather than per-user to match the behaviour of Word.
* Add support for comments.
Add support for w:sdt elements. This allows the bodies of content controls, such as bibliographies, to be converted.
Add support for table cells spanning multiple rows.
Add support for table cells spanning multiple columns.
Improve script installation on Windows by using entry_points instead of scripts in setup.py.
Remove deprecated convert_underline argument.
Remove deprecated convert_underline argument.
Officially support ID prefixes.
Generated IDs no longer insert a hyphen after the ID prefix.
The default ID prefix is now the empty string rather than a random number followed by a hyphen.
Rename mammoth.images.inline to mammoth.images.img_element to better reflect its behaviour.
Improve collapsing of similar non-fresh HTML elements.
Allow bold and italic style mappings to be configured.
Handle references to missing styles when reading documents.
Handle XML where the child nodes of an element contains text nodes.
Always use mc:Fallback when reading mc:AlternateContent elements.
Remove duplicate messages from results.
Remove duplicate messages from results.
Read v:imagedata with r:id attribute.
Read children of v:roundrect.
Ignore office-word:wrap, v:shadow and v:shapetype.
Continue with warning if external images cannot be found.
Continue with warning if external images cannot be found.
Add support for embedded style maps.
* Fix Python 3 support.
Generate warnings for not-understood style mappings and continue, rather than stopping with an error.
Generate warnings for not-understood style mappings and continue, rather than stopping with an error.
Support file objects without a name attribute again (broken since 0.3.20).
Ignore w:numPr elements without w:numId or w:ilvl children.
Your coding agent can read these notes before it upgrades. Set up the MCP server →