413 Commits
Author SHA1 Message Date
Nefelibataandafourney 58783e3272 fix(docx): handle unknown math functions in OMML converter without crashing (#2268)
* fix(doc-intel): default api_version to None in DocumentIntelligenceConverter

Do not set a default api_version string in DocumentIntelligenceConverter. If api_version is omitted or None, avoid passing api_version kwarg to DocumentIntelligenceClient so Azure SDK uses its native default version.

Fixes #1904

* fix(docx): handle unknown math functions in OMML converter without crashing

Add common standard math functions to FUNC dictionary in latex_dict.py and fallback to \operatorname{} for unknown functions in omml.py instead of throwing NotImplementedError.

Fixes #1982

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 15:50:58 -07:00
hanhan761 dd2df72913 fix: ImageConverter gracefully handles LLM API failures (fixes #1942) (#1948) 2026-09-03 15:33:16 -07:00
Lazizbek Ergashevandafourney 697a2a0e0b fix(rss): treat Atom text content and summary as plain text (#2374)
* fix(rss): treat Atom text content and summary as plain text
* Parse content types more carefully.
* Don't pass text through the html converter.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 15:01:34 -07:00
hanhan761andafourney 94a84834a2 chore: remove unused mammoth import from PlainTextConverter (#1951) (#1953)
* fix: handle LLM API errors gracefully in ImageConverter (#1942)

If the OpenAI API call fails (auth error, rate limit, network timeout, etc.), the exception propagates and crashes the entire conversion. Wrap the call in try/except to gracefully skip the LLM caption, matching the pattern already used by PptxConverter.

* chore: remove unused mammoth import from PlainTextConverter (#1951)

The mammoth import was copied from DocxConverter but PlainTextConverter does not need it. Remove the dead code and unused _dependency_exc_info variable.

* Revert unrelated changes.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 12:58:47 -07:00
c870d49284 fix: suppress pydub RuntimeWarning when ffmpeg is missing (fixes #1685) (#1985)
* fix: suppress pydub RuntimeWarning when ffmpeg is missing

Closes #1685

Importing pydub emits a RuntimeWarning when ffmpeg/avconv is not
installed on the system. This warning appears even when audio
transcription is not actively being used.

_transcribe_audio.py already wraps the imports in a
warnings.catch_warnings block that suppresses DeprecationWarning and
SyntaxWarning. Add a filter for the ffmpeg RuntimeWarning as well.

* Modify RuntimeWarning filter for ffmpeg

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: afourney <adamfo@microsoft.com>
Co-authored-by: afourney <adam.fourney@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-03 12:41:55 -07:00
Mukunda Rao Katta 7de6ce4943 Support short YouTube URLs in converter (#1882) 2026-09-03 12:22:34 -07:00
a5ddc8efd5 fix: handle DOCX files with inconsistent ZIP filename casing (#1812) (#2016)
* fix: handle DOCX files with inconsistent ZIP filename casing (#1812)

Some document generators (e.g. certain Microsoft Word versions, legal
document systems) produce .docx files where the central directory records
one casing (e.g. 'customXml/item2.xml') but the local file headers record
another (e.g. 'customXML/item2.xml'). Python's zipfile module raises
BadZipFile when reading such files.

Add _fix_zip_filename_casing() to patch local file header filenames to
match the central directory before any ZIP processing occurs.

* test: register DOCX zip casing test in __main__ runner list

* fix: encode central directory filenames with the ZIP's actual charset

ZIP filenames are stored as cp437 unless general purpose flag bit 11
(0x800) marks them UTF-8, which is how zipfile decoded orig_filename.
Re-encoding unconditionally as UTF-8 produced the wrong bytes for
non-ASCII cp437 names, so a local header could be skipped or patched
with mismatched bytes.

* fix(ocr): pre-process the DOCX before reading it in DocxConverterWithOCR

convert() read the caller's stream twice before repairing it: once for the
embedded style map, and once to extract images for OCR. A .docx whose ZIP
local file headers disagree with the central directory on casing would then
fail those reads. The image extraction swallows every exception, so the
result was an empty OCR map and silently missing OCR blocks -- output
identical to the no-OCR path, with no error.

Call pre_process_docx() once at the top and feed the repaired stream to
every read, matching the ordering the base DocxConverter already uses. This
also drops a redundant second rebuild of the archive on both branches.

* Improve name comparison for casing mismatch

* test: assert non-case ZIP filename mismatches are still rejected

Equal encoded length does not imply two filenames differ only in case, so
the local file header repair must not wave through archives whose local and
central directory names genuinely disagree. Cover an unrelated same-length
name, a single differing character, and a difference outside the cased
characters, alongside the case-only mismatch that should still be repaired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: lyydsheep <lyydsheep@lyydsheepdeMacBook-Pro.local>
Co-authored-by: Adam Fourney <adamfo@microsoft.com>
2026-09-03 12:18:46 -07:00
hanhan761andafourney 07cbd135df fix: guard against missing oMath element in DOCX math converter (#1979) (#1995)
* fix: guard against missing oMath element in DOCX math converter (#1979)

Add a None check in _convert_omath_to_latex() to return an empty
string when the math element is missing, instead of passing None to
oMath2Latex() which raises TypeError.

Fixes #1979

* Formatting.

* Change assertions to validate markdown output
* Added regression test.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 11:27:56 -07:00
hanhan761andafourney 589362f099 fix: WikipediaConverter renders hash None heading when page has no title (#1990)
* fix: WikipediaConverter renders # None heading when page has no title
* Fix assertion to check markdown output instead of text
* Treat whitespace titles as empty.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 11:02:57 -07:00
2d9c49c4e1 Fix exiftool JSON decoding to use UTF-8 (#2067)
* Fix exiftool JSON decoding to use UTF-8
* Removed unused import.

---------

Co-authored-by: Vedant <vedantved55555@gmail.com>
Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 10:48:06 -07:00
王艺霖andafourney 5deed1e896 fix: catch OSError when exiftool binary is missing (#1960) (#2082)
* fix: catch OSError when exiftool binary is missing (#1960)

When exiftool_path points to a binary that does not exist, the
subprocess.run calls raised FileNotFoundError, which propagated
unhandled through exiftool_metadata() and crashed the whole
conversion.

Add OSError to the except clauses at both subprocess.run sites
and wrap the error in a RuntimeError that includes the path,
giving users a clear, actionable error message.

Add a regression test for the missing-binary case and a sanity
check that the exiftool_path=None early return still works.

* Differentiate errors.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 10:36:39 -07:00
aoright ff9015652e fix(pptx): prevent crash in PptxConverter when chart title lacks a text frame (#2194)
* fix(pptx): prevent crash in PptxConverter when chart title lacks a text frame

Chart title text_frame can be None if the title is set directly or programmatically in python-pptx. Added a None check before attempting to read text.
2026-09-03 10:12:33 -07:00
Alvin Tangandalvinttang cfe54d24cc fix: PptxConverter tolerates None shape.text and notes text (#2059)
shape.text and notes_text_frame.text can return None for runs with no
<a:t> child or certain third-party decks. Treat None as empty string so
one malformed shape does not fail the whole file.

Refs #1808

Co-authored-by: alvinttang <alvin@pm.me>
2026-09-03 09:56:09 -07:00
Lubrsy 70033462f0 fix(pptx): ignore empty llm captions (#1886)
* fix(pptx): ignore empty llm captions
* Refactor alt_text creation in _pptx_converter.py
2026-09-03 09:34:15 -07:00
Yufeng Heandafourney b752951ec7 fix(docx): preserve underlined text (#2017)
* fix(docx): preserve underlined text
* Update assertion to check markdown instead of text content
* fix(docx): apply default underline style map in OCR converter too
* fix(docx): do not let the underline default override embedded style maps
* Update embedded style map reading from file stream

---------

Co-authored-by: afourney <adam.fourney@gmail.com>
2026-09-03 08:40:25 -07:00
αI a85dff7fc1 fix(ipynb): preserve leading # in notebook heading titles (#2371)
* fix(ipynb): preserve leading # in notebook heading titles

Use removeprefix instead of lstrip so title text that starts with '#'
or spaces is not over-stripped. Fixes #2367.
2026-09-02 21:06:57 -07:00
dependabot[bot]andafourney 42b98ec256 Bump actions/setup-python from 5.6.0 to 7.0.0 (#2321)
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 5.6.0 to 7.0.0.
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](https://github.com/actions/setup-python/compare/a26af69be951a213d495a4c3e4e4022e16d87065...5fda3b95a4ea91299a34e894583c3862153e4b97)

---
updated-dependencies:
- dependency-name: actions/setup-python
  dependency-version: 7.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-02 20:59:08 -07:00
dependabot[bot]andafourney f89f5b56b0 Bump actions/checkout from 5.1.0 to 7.0.1 (#2320)
Bumps [actions/checkout](https://github.com/actions/checkout) from 5.1.0 to 7.0.1.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/fbc6f3992d24b796d5a048ff273f7fcc4a7b6c09...3d3c42e5aac5ba805825da76410c181273ba90b1)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: 7.0.1
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-02 20:55:17 -07:00
Henry Suandafourney 12b0552821 fix(youtube): handle missing title metadata without raising AssertionError (#2238)
* fix(youtube): handle missing title metadata without raising AssertionError

* Empty titles would still raise assertion errors. Fixed.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-02 20:07:18 -07:00
Sheroy Cooper 20d06b6c85 docs: use canonical markdown result property (#2259) 2026-09-02 16:27:03 -07:00
Yufeng Heandafourney 0f26ef9fbc fix(xlsx): tolerate legacy showZeroes sheet views (#2064)
* fix(xlsx): tolerate legacy showZeroes sheet views

* Narrow the find/replace of showZeros

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-02 16:10:59 -07:00
Lucas Ma 22db17026a fix: normalize data URI parameter case (#2120) 2026-09-02 11:18:53 -07:00
Lucas MaandLucas Ma 05dfe17716 fix: handle URI schemes case-insensitively (#2121)
* fix: handle URI schemes case-insensitively

---------

Co-authored-by: Lucas Ma <7184042+pony-maggie@users.noreply.github.com>
2026-09-02 11:04:07 -07:00
Guillermo Dolsandafourney 25c1485477 fix(outlook): read .msg string properties saved in the non-Unicode format (#2295)
* fix: read .msg string properties saved in the non-Unicode format

Every string property in a .msg lives under a stream whose name ends in its
MAPI type: 001F for PT_UNICODE (UTF-16LE) or 001E for PT_STRING8, written in
the message's code page. Outlook writes one or the other for a given message,
never both, so a message saved in the legacy non-Unicode format carries no
001F streams at all.

The converter addressed only the 001F names. Such a message therefore came
out as bare scaffolding -- "# Email Message" followed by "## Content" -- with
From, To, Subject and the body all silently dropped, and no error raised.

Each property is now read from the 001F stream and, failing that, from its
001E counterpart. PT_STRING8 streams record no encoding of their own, so the
charset is detected with charset_normalizer, as is already done for other
8-bit sources in the codebase. The Unicode path is unchanged.

* Read code page info once.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-02 10:59:10 -07:00
Nefelibata 83ce26def9 fix(doc-intel): default api_version to None in DocumentIntelligenceConverter (#2267)
* fix(doc-intel): default api_version to None in DocumentIntelligenceConverter

Do not set a default api_version string in DocumentIntelligenceConverter. If api_version is omitted or None, avoid passing api_version kwarg to DocumentIntelligenceClient so Azure SDK uses its native default version.

Fixes #1904
2026-09-02 07:58:35 -07:00
e57e33291f fix: truncate uppercase data image URIs (#2122)
* fix: truncate uppercase data image URIs

---------

Co-authored-by: Lucas Ma <7184042+pony-maggie@users.noreply.github.com>
Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-02 00:00:08 -07:00
149f8b2114 fix: ZipConverter renders '(unknown)' instead of literal 'None' when stream has no source info (#2134)
* fix: ZipConverter emits '(unknown)' instead of 'None' when stream has no source info

When MarkItDown.convert_stream() is called with a ZIP stream that has no
associated URL, local_path, or filename (e.g. a raw io.BytesIO), the
ZipConverter header read:

    Content from the zip file `None`:

because stream_info.url, stream_info.local_path, and stream_info.filename
were all None and Python f-strings render None as the literal string 'None'.

Fix: fall back to '(unknown)' when all three source-info fields are absent,
producing the more descriptive:

    Content from the zip file `(unknown)`:

Add a regression test in test_module_misc.py that verifies the output does
not contain the literal string 'None' in this scenario.

* Fixed formatting.

---------

Co-authored-by: JSap0914 <JSap0914@users.noreply.github.com>
Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-01 23:49:19 -07:00
Sai Medavarapu 36618531a5 fix: initialize md_text in _parse_rss_type to prevent UnboundLocalError (#2164) 2026-09-01 22:51:18 -07:00
will wangandafourney ee6fe25e3e fix(rss): preserve Atom XHTML content (#2297)
* fix(rss): preserve Atom XHTML content
* Support xhtml namespace prefixes.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-01 22:32:27 -07:00
Mo Huiandafourney bebea8bddd fix: buffer CLI stdin before format detection on Windows (#2351)
* fix: buffer CLI stdin before format detection
* Fix test to check stream contents instead of comparing references.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-01 22:01:55 -07:00
5294fa3a5c fix: preserve strikethrough semantics for <strike>, line-through CSS, and w:dstrike (#2356)
* fix: preserve strikethrough semantics for w:dstrike in DOCX pre-processing

Normalize <w:dstrike> to <strike> during DOCX pre-processing so double
strikethrough runs keep their strikethrough semantics in the output.

* Apply pre-processing steps independently

---------

Co-authored-by: Wasim <wasim@example.com>
Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-01 21:48:04 -07:00
0b3a45234e fix: remove extra closing brace from caron and ring-above accents (#2279)
`CHR["\u030c"]` and `CHR["\u030a"]` each carried a fourth `}`, so the
templates read `\check{{{0}}}}` and `\ocirc{{{0}}}}`. Every other accent
in the table is `\name{{{0}}}`. `oMath2Latex.do_acc` applies these with
`latex_s.format(c_dict["e"])`, and an unbalanced template raises
`ValueError: Single '}' encountered in format string`.

The damage is not limited to the accented equation. `pre_process_docx`
runs `_pre_process_math` over the whole of `word/document.xml` inside a
blanket `except Exception` and, on error, writes the *original*
unprocessed XML back. Mammoth does not render OMML, so a single caron or
ring-above accent silently removes **every** equation in the document,
with no error and a zero exit code.

Both are standard entries in Word's Equation > Accent gallery
(U+030C COMBINING CARON, U+030A COMBINING RING ABOVE).

Measured on the repo's own `equations.docx` fixture: 2 equations
recovered normally, 0 after appending one caron-accented equation, 3
with this fix applied.

Adds `test_docx_math_accents.py`: a table-wide guard asserting every
`{0}` template survives `.format()`, plus direct coverage of the two
affected accents and an unaffected control.

Co-authored-by: AndrewAvery7 <AndrewAvery7@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 21:07:33 -07:00
Sonai Biswas 5e1f3813c5 fix: preserve percent-encoded href path bytes (#2173) 2026-09-01 20:55:39 -07:00
S Mandafourney f9b0dcb423 fix(docx): ignore malformed styles missing type (#2190)
* fix(docx): ignore malformed styles missing type
* Small fixes, defaulting to paragraph.
---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-01 20:07:14 -07:00
hanhan761 d08976de08 fix: IpynbConverter.accepts() catches UnicodeDecodeError on non-decodable content (fixes #1894) (#1929)
* fix: IpynbConverter.accepts() raises UnicodeDecodeError on non-ASCII files

Wrap the decode block in accepts() with try/except (LookupError,
ValueError) so that non-ASCII files (e.g. French PDFs) whose MIME type
starts with 'application/json' do not crash the conversion pipeline.

Fixes #1894
2026-09-01 19:16:35 -07:00
S1MS4 556695de0c fix: handle math run with no text child in OMML->LaTeX conversion (#2189)
do_r() called elm.findtext("./m:t") and iterated over the result
directly. When a math run (<m:r>) has no <m:t> text child (e.g. a
run that only carries formatting properties, produced by some Word
equation editors), findtext() returns None and iterating over it
raises TypeError: 'NoneType' object is not iterable.

Because equation pre-processing is applied at the whole-document.xml
level with a blanket try/except, this single malformed run aborts
LaTeX conversion for every equation in the document, silently
dropping all native Word equations from the output.
2026-09-01 18:37:12 -07:00
Henry Su a139677f94 fix(epub): safely extract metadata text without crashing on None nodeValue or nested elements (#2247)
* fix(epub): safely extract metadata text without crashing on None nodeValue or nested elements
2026-09-01 18:20:36 -07:00
af56495dbb fix: escape pipes and newlines in CSV values (#2266)
* fix: escape pipes and newlines in CSV values

CsvConverter joins cells with " | " and passes values through unescaped, so
two characters that are legal inside a CSV field silently corrupt the table:
a pipe is read as a column separator, and a newline inside a quoted field ends
the row early.

The pipe case loses data rather than just looking wrong. A row with an
unescaped pipe declares more columns than the header, and renderers discard
the surplus -- 'cheap | fast' renders as 'cheap'.

Escape pipes as \| and turn embedded CR/LF into <br> so the record stays on
one line.

* Replaced <br> with ' ', and don't blindly escape '|' if they are already escaped.

Co-authored-by: Asjad Abbas <215788583+asjad3@users.noreply.github.com>
Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-01 17:51:21 -07:00
Zhewen Tan f1df5ac4ab Fix percent-encoded Windows drive paths in file URIs (#2315)
* Fix percent-encoded Windows file URIs
2026-09-01 16:59:41 -07:00
Gyanu Mayankandafourney be631e1fc0 Preserve strikethrough from <strike> and CSS line-through (#2342)
* Keep strikethrough for <strike> and CSS line-through.

<s> and <del> already became ~~text~~. The obsolete strike element and inline text-decoration: line-through were flattened to plain text, so deleted wording looked current.

* Don't handle css style tags yet.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-01 16:20:29 -07:00
Ruiming Zhaoandafourney d04418ec9a fix(csv): strip the UTF-8 BOM and skip blank rows before building the table (#2303)
* fix(csv): strip the UTF-8 BOM and skip blank rows before building the table

Two things broke real-world CSV conversion: Excel prepends a UTF-8 BOM that landed inside the first header cell, and a single leading or trailing blank line parsed as an empty row and shifted every column (a leading blank line destroyed the table entirely, since the empty row became the header). Strip the BOM and drop empty rows before deciding what the table looks like.

* Moved tests into the misc. file, and keep internal blank lines in tables.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-01 13:23:20 -07:00
Sinan Olsson-Pasic f618f478a7 feat(cli): support MARKITDOWN_CU_ENDPOINT / MARKITDOWN_DOCINTEL_ENDPOINT (#2358)
* feat(cli): read endpoints from MARKITDOWN_CU_ENDPOINT / MARKITDOWN_DOCINTEL_ENDPOINT. Lets an operator configure the Azure endpoint once in the environment, so callers only need --use-cu or -d. Explicit flags still take precedence, and an empty variable is treated as unset. Fixes #2326.

* Add a leading space between string literals.
2026-09-01 09:41:20 -07:00
Guillermo Dolsandafourney bd7b77e2c4 fix: correct \underleftarrow macro and map math italic h in equation conversion (#2293)
U+20EE COMBINING LEFT ARROW BELOW mapped to \underledtarrow, which is not a
LaTeX macro -- the name is a scrambled \underleftarrow. The two neighbouring
entries in the same table show the intent: U+20D6 -> \overleftarrow and
U+20EF -> \underrightarrow. The template still formats, so nothing raises;
the equation just ends up with an undefined control sequence.

T normalizes the Mathematical Alphanumeric Symbols back to ASCII, but skipped
math italic small h. That letter has no codepoint of its own: U+1D455 is
permanently reserved because Unicode unifies it with U+210E PLANCK CONSTANT.
Following the contiguous block left h as the only Latin letter to survive
untranslated, so an equation reading h(x)=g(x) converted to "ℎ(x)=g(x)".

Co-authored-by: afourney <adamfo@microsoft.com>
2026-08-31 15:30:40 -07:00
Tyyyy a01f1cbfca fix(cli): allow Content Understanding conversion from stdin (#2318) 2026-08-31 15:24:10 -07:00
Jeremy Schoemaker 4752813fa5 fix: fall back to plain text when RSS item content triggers RecursionError (#2333)
* fix: fall back to plain text when RSS item content triggers RecursionError

Deeply nested HTML inside an RSS item's description or content:encoded
previously hit the broad except in _parse_content and returned the raw
unconverted HTML, silently embedding HTML tags in the markdown output.

Catch RecursionError specifically and fall back to BeautifulSoup's
iterative get_text(), mirroring the HTML converter fix from #1644.
Includes a deterministic regression test.

* Honor strict=true, similar to HTML
2026-08-31 14:38:57 -07:00
Rāna(Bass Ver.) 76b965ad52 fix: support extended content-disposition filenames (#2045)
* fix: support extended content-disposition filenames

* Added a test to verify decoding.
2026-08-31 13:59:59 -07:00
Fabio Catalano 4243c26b3a Mitigate UnicodeDecodeError due to wrong ASCII charset guess for long files (#2360)
* Add json-bug test

* Harden ipyn_ converter accepts

* Extend the window used to guess the charset

* Add help message in case of UnicodeDecodeError

* Fix style

* Remove leftover print
2026-08-31 13:23:05 -07:00
Dan Fiedler 9dc0d6579b Pin GitHub Actions to full-length commit SHAs (#2316) 2026-08-19 12:34:00 -07:00
afourney fd239d5d2b Bump version to 0.1.7 (#2258) v0.1.7 2026-07-29 11:16:03 -07:00
afourney 007ea023b9 Fix omml template bugs. (#2257) 2026-07-29 10:08:30 -07:00