* fix(doc-intel): default api_version to None in DocumentIntelligenceConverter
Do not set a default api_version string in DocumentIntelligenceConverter. If api_version is omitted or None, avoid passing api_version kwarg to DocumentIntelligenceClient so Azure SDK uses its native default version.
Fixes#1904
* fix(docx): handle unknown math functions in OMML converter without crashing
Add common standard math functions to FUNC dictionary in latex_dict.py and fallback to \operatorname{} for unknown functions in omml.py instead of throwing NotImplementedError.
Fixes#1982
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix(rss): treat Atom text content and summary as plain text
* Parse content types more carefully.
* Don't pass text through the html converter.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: handle LLM API errors gracefully in ImageConverter (#1942)
If the OpenAI API call fails (auth error, rate limit, network timeout, etc.), the exception propagates and crashes the entire conversion. Wrap the call in try/except to gracefully skip the LLM caption, matching the pattern already used by PptxConverter.
* chore: remove unused mammoth import from PlainTextConverter (#1951)
The mammoth import was copied from DocxConverter but PlainTextConverter does not need it. Remove the dead code and unused _dependency_exc_info variable.
* Revert unrelated changes.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: suppress pydub RuntimeWarning when ffmpeg is missing
Closes#1685
Importing pydub emits a RuntimeWarning when ffmpeg/avconv is not
installed on the system. This warning appears even when audio
transcription is not actively being used.
_transcribe_audio.py already wraps the imports in a
warnings.catch_warnings block that suppresses DeprecationWarning and
SyntaxWarning. Add a filter for the ffmpeg RuntimeWarning as well.
* Modify RuntimeWarning filter for ffmpeg
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: afourney <adamfo@microsoft.com>
Co-authored-by: afourney <adam.fourney@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* fix: handle DOCX files with inconsistent ZIP filename casing (#1812)
Some document generators (e.g. certain Microsoft Word versions, legal
document systems) produce .docx files where the central directory records
one casing (e.g. 'customXml/item2.xml') but the local file headers record
another (e.g. 'customXML/item2.xml'). Python's zipfile module raises
BadZipFile when reading such files.
Add _fix_zip_filename_casing() to patch local file header filenames to
match the central directory before any ZIP processing occurs.
* test: register DOCX zip casing test in __main__ runner list
* fix: encode central directory filenames with the ZIP's actual charset
ZIP filenames are stored as cp437 unless general purpose flag bit 11
(0x800) marks them UTF-8, which is how zipfile decoded orig_filename.
Re-encoding unconditionally as UTF-8 produced the wrong bytes for
non-ASCII cp437 names, so a local header could be skipped or patched
with mismatched bytes.
* fix(ocr): pre-process the DOCX before reading it in DocxConverterWithOCR
convert() read the caller's stream twice before repairing it: once for the
embedded style map, and once to extract images for OCR. A .docx whose ZIP
local file headers disagree with the central directory on casing would then
fail those reads. The image extraction swallows every exception, so the
result was an empty OCR map and silently missing OCR blocks -- output
identical to the no-OCR path, with no error.
Call pre_process_docx() once at the top and feed the repaired stream to
every read, matching the ordering the base DocxConverter already uses. This
also drops a redundant second rebuild of the archive on both branches.
* Improve name comparison for casing mismatch
* test: assert non-case ZIP filename mismatches are still rejected
Equal encoded length does not imply two filenames differ only in case, so
the local file header repair must not wave through archives whose local and
central directory names genuinely disagree. Cover an unrelated same-length
name, a single differing character, and a difference outside the cased
characters, alongside the case-only mismatch that should still be repaired.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: lyydsheep <lyydsheep@lyydsheepdeMacBook-Pro.local>
Co-authored-by: Adam Fourney <adamfo@microsoft.com>
* fix: guard against missing oMath element in DOCX math converter (#1979)
Add a None check in _convert_omath_to_latex() to return an empty
string when the math element is missing, instead of passing None to
oMath2Latex() which raises TypeError.
Fixes#1979
* Formatting.
* Change assertions to validate markdown output
* Added regression test.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: WikipediaConverter renders # None heading when page has no title
* Fix assertion to check markdown output instead of text
* Treat whitespace titles as empty.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: catch OSError when exiftool binary is missing (#1960)
When exiftool_path points to a binary that does not exist, the
subprocess.run calls raised FileNotFoundError, which propagated
unhandled through exiftool_metadata() and crashed the whole
conversion.
Add OSError to the except clauses at both subprocess.run sites
and wrap the error in a RuntimeError that includes the path,
giving users a clear, actionable error message.
Add a regression test for the missing-binary case and a sanity
check that the exiftool_path=None early return still works.
* Differentiate errors.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix(pptx): prevent crash in PptxConverter when chart title lacks a text frame
Chart title text_frame can be None if the title is set directly or programmatically in python-pptx. Added a None check before attempting to read text.
shape.text and notes_text_frame.text can return None for runs with no
<a:t> child or certain third-party decks. Treat None as empty string so
one malformed shape does not fail the whole file.
Refs #1808
Co-authored-by: alvinttang <alvin@pm.me>
* fix(docx): preserve underlined text
* Update assertion to check markdown instead of text content
* fix(docx): apply default underline style map in OCR converter too
* fix(docx): do not let the underline default override embedded style maps
* Update embedded style map reading from file stream
---------
Co-authored-by: afourney <adam.fourney@gmail.com>
* fix(ipynb): preserve leading # in notebook heading titles
Use removeprefix instead of lstrip so title text that starts with '#'
or spaces is not over-stripped. Fixes#2367.
* fix: read .msg string properties saved in the non-Unicode format
Every string property in a .msg lives under a stream whose name ends in its
MAPI type: 001F for PT_UNICODE (UTF-16LE) or 001E for PT_STRING8, written in
the message's code page. Outlook writes one or the other for a given message,
never both, so a message saved in the legacy non-Unicode format carries no
001F streams at all.
The converter addressed only the 001F names. Such a message therefore came
out as bare scaffolding -- "# Email Message" followed by "## Content" -- with
From, To, Subject and the body all silently dropped, and no error raised.
Each property is now read from the 001F stream and, failing that, from its
001E counterpart. PT_STRING8 streams record no encoding of their own, so the
charset is detected with charset_normalizer, as is already done for other
8-bit sources in the codebase. The Unicode path is unchanged.
* Read code page info once.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix(doc-intel): default api_version to None in DocumentIntelligenceConverter
Do not set a default api_version string in DocumentIntelligenceConverter. If api_version is omitted or None, avoid passing api_version kwarg to DocumentIntelligenceClient so Azure SDK uses its native default version.
Fixes#1904
* fix: ZipConverter emits '(unknown)' instead of 'None' when stream has no source info
When MarkItDown.convert_stream() is called with a ZIP stream that has no
associated URL, local_path, or filename (e.g. a raw io.BytesIO), the
ZipConverter header read:
Content from the zip file `None`:
because stream_info.url, stream_info.local_path, and stream_info.filename
were all None and Python f-strings render None as the literal string 'None'.
Fix: fall back to '(unknown)' when all three source-info fields are absent,
producing the more descriptive:
Content from the zip file `(unknown)`:
Add a regression test in test_module_misc.py that verifies the output does
not contain the literal string 'None' in this scenario.
* Fixed formatting.
---------
Co-authored-by: JSap0914 <JSap0914@users.noreply.github.com>
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: buffer CLI stdin before format detection
* Fix test to check stream contents instead of comparing references.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: preserve strikethrough semantics for w:dstrike in DOCX pre-processing
Normalize <w:dstrike> to <strike> during DOCX pre-processing so double
strikethrough runs keep their strikethrough semantics in the output.
* Apply pre-processing steps independently
---------
Co-authored-by: Wasim <wasim@example.com>
Co-authored-by: afourney <adamfo@microsoft.com>
`CHR["\u030c"]` and `CHR["\u030a"]` each carried a fourth `}`, so the
templates read `\check{{{0}}}}` and `\ocirc{{{0}}}}`. Every other accent
in the table is `\name{{{0}}}`. `oMath2Latex.do_acc` applies these with
`latex_s.format(c_dict["e"])`, and an unbalanced template raises
`ValueError: Single '}' encountered in format string`.
The damage is not limited to the accented equation. `pre_process_docx`
runs `_pre_process_math` over the whole of `word/document.xml` inside a
blanket `except Exception` and, on error, writes the *original*
unprocessed XML back. Mammoth does not render OMML, so a single caron or
ring-above accent silently removes **every** equation in the document,
with no error and a zero exit code.
Both are standard entries in Word's Equation > Accent gallery
(U+030C COMBINING CARON, U+030A COMBINING RING ABOVE).
Measured on the repo's own `equations.docx` fixture: 2 equations
recovered normally, 0 after appending one caron-accented equation, 3
with this fix applied.
Adds `test_docx_math_accents.py`: a table-wide guard asserting every
`{0}` template survives `.format()`, plus direct coverage of the two
affected accents and an unaffected control.
Co-authored-by: AndrewAvery7 <AndrewAvery7@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix: IpynbConverter.accepts() raises UnicodeDecodeError on non-ASCII files
Wrap the decode block in accepts() with try/except (LookupError,
ValueError) so that non-ASCII files (e.g. French PDFs) whose MIME type
starts with 'application/json' do not crash the conversion pipeline.
Fixes#1894
do_r() called elm.findtext("./m:t") and iterated over the result
directly. When a math run (<m:r>) has no <m:t> text child (e.g. a
run that only carries formatting properties, produced by some Word
equation editors), findtext() returns None and iterating over it
raises TypeError: 'NoneType' object is not iterable.
Because equation pre-processing is applied at the whole-document.xml
level with a blanket try/except, this single malformed run aborts
LaTeX conversion for every equation in the document, silently
dropping all native Word equations from the output.
* fix: escape pipes and newlines in CSV values
CsvConverter joins cells with " | " and passes values through unescaped, so
two characters that are legal inside a CSV field silently corrupt the table:
a pipe is read as a column separator, and a newline inside a quoted field ends
the row early.
The pipe case loses data rather than just looking wrong. A row with an
unescaped pipe declares more columns than the header, and renderers discard
the surplus -- 'cheap | fast' renders as 'cheap'.
Escape pipes as \| and turn embedded CR/LF into <br> so the record stays on
one line.
* Replaced <br> with ' ', and don't blindly escape '|' if they are already escaped.
Co-authored-by: Asjad Abbas <215788583+asjad3@users.noreply.github.com>
Co-authored-by: afourney <adamfo@microsoft.com>
* Keep strikethrough for <strike> and CSS line-through.
<s> and <del> already became ~~text~~. The obsolete strike element and inline text-decoration: line-through were flattened to plain text, so deleted wording looked current.
* Don't handle css style tags yet.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix(csv): strip the UTF-8 BOM and skip blank rows before building the table
Two things broke real-world CSV conversion: Excel prepends a UTF-8 BOM that landed inside the first header cell, and a single leading or trailing blank line parsed as an empty row and shifted every column (a leading blank line destroyed the table entirely, since the empty row became the header). Strip the BOM and drop empty rows before deciding what the table looks like.
* Moved tests into the misc. file, and keep internal blank lines in tables.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* feat(cli): read endpoints from MARKITDOWN_CU_ENDPOINT / MARKITDOWN_DOCINTEL_ENDPOINT. Lets an operator configure the Azure endpoint once in the environment, so callers only need --use-cu or -d. Explicit flags still take precedence, and an empty variable is treated as unset. Fixes#2326.
* Add a leading space between string literals.
U+20EE COMBINING LEFT ARROW BELOW mapped to \underledtarrow, which is not a
LaTeX macro -- the name is a scrambled \underleftarrow. The two neighbouring
entries in the same table show the intent: U+20D6 -> \overleftarrow and
U+20EF -> \underrightarrow. The template still formats, so nothing raises;
the equation just ends up with an undefined control sequence.
T normalizes the Mathematical Alphanumeric Symbols back to ASCII, but skipped
math italic small h. That letter has no codepoint of its own: U+1D455 is
permanently reserved because Unicode unifies it with U+210E PLANCK CONSTANT.
Following the contiguous block left h as the only Latin letter to survive
untranslated, so an equation reading h(x)=g(x) converted to "ℎ(x)=g(x)".
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: fall back to plain text when RSS item content triggers RecursionError
Deeply nested HTML inside an RSS item's description or content:encoded
previously hit the broad except in _parse_content and returned the raw
unconverted HTML, silently embedding HTML tags in the markdown output.
Catch RecursionError specifically and fall back to BeautifulSoup's
iterative get_text(), mirroring the HTML converter fix from #1644.
Includes a deterministic regression test.
* Honor strict=true, similar to HTML
* Add json-bug test
* Harden ipyn_ converter accepts
* Extend the window used to guess the charset
* Add help message in case of UnicodeDecodeError
* Fix style
* Remove leftover print