194 Commits
Author SHA1 Message Date
Machen John 1f9530a6d3 fix(wikipedia): drop the site suffix from italicised titles (#2458)
Titles rendered with markup have no mw-page-title-main span, so the
converter fell through to the document <title>, suffix and all.
2026-09-30 22:24:20 -07:00
Amit Anand 44f3a0ed71 Use frozenset for OMML direct tag lookup (#1919)
Replace __direct_tags tuple with frozenset in oMath2Latex for O(1) membership checks in process_unknow(). No behavioral change.
2026-09-30 22:16:01 -07:00
Yasir Alibrahem f70e637af2 optimize non-seekable stream buffering using shutil.copyfileobj (#1510) 2026-09-30 22:11:23 -07:00
hchmch 5610ced394 Add language parameter for the audio transcription (#1295)
* Add language parameter for the audio transcription
* style: format audio language option with Black
2026-09-30 22:00:00 -07:00
Octopus 53f38ea102 fix: handle None sys.stdout.encoding in CLI output (#1656)
* fix: handle None sys.stdout.encoding in CLI output (fixes #1597)
2026-09-30 16:20:37 -07:00
Gopal Bagaswar 533ea574b3 fix(wikipedia): accept mobile (m.wikipedia.org) URLs in WikipediaConverter (#1850) 2026-09-30 16:09:02 -07:00
Yasir Alibrahem 199fe8b185 increase chunk size from 512 bytes to 64 KB for HTTP response reading (#1511) 2026-09-30 16:00:48 -07:00
afourney b8f79c57eb Bump versions or markitdown and markitdown-ocr (#2545) 2026-09-21 14:02:08 -07:00
Nithin JambulaandAdam Fourney d51937de3a fix: bump youtube-transcript-api to >=1.2.3 for Python 3.14 support (#2407)
* fix: bump youtube-transcript-api to >=1.2.3 for Python 3.14 support
* Add 3.14 to test matrix, and limit advertised compatibility.

---------

Co-authored-by: Adam Fourney <adamfo@microsoft.com>
2026-09-21 09:47:07 -07:00
afourney 945314a45d Refactor Office OCR converters to reuse core conversion pipelines (#2506)
* Add _image_to_html helper for docx ocr conversion.
* Ensure compatibility with core markitdown version
* Have pptx, xlsx and docx expose an _image_to_html semi-private method, overridable by plugins.
* Fix conditional check for llm_description
* Update relationships function calls for drawings and images
* Forward stream_info to LLMVisionOCRService
2026-09-16 10:23:08 -07:00
afourney eb31b5c945 bump_versions (#2490) 2026-09-14 09:48:43 -07:00
afourney cc0ca9edd8 Preserve whitepace underlines (#2477)
* Preserve whitespace underlines
2026-09-12 13:25:31 -07:00
afourney 3a6ce44358 Throw an error when llm client fails with image converter. (#2476) 2026-09-12 12:11:30 -07:00
afourney 5640da7145 fix(youtube): fall back to HTML when no video content is extracted (#2469) 2026-09-11 15:52:59 -07:00
afourney 485a7b93b6 fix(docx): preserve namespaces when repairing stylesheets (#2467)
* fix(docx): preserve namespaces when repairing stylesheets
* Fixed d:strike as well
2026-09-11 15:05:05 -07:00
afourney 75e7114598 fix: avoid splitting UTF-8 characters during charset detection (#2466) 2026-09-11 14:35:07 -07:00
afourney 2277fff52d Improve efficiency of escaping pipes in CSVs (#2464) 2026-09-11 12:24:26 -07:00
dickbown f0c5d01fbe perf(csv): Batch-remove blank rows to eliminate the quadratic overhead in table conversion 2026-09-11 11:12:55 -07:00
afourney 40658cdcc5 Fix ANSI Outlook MSG decoding for Japanese code pages and padded strings (#2462) 2026-09-11 11:05:33 -07:00
afourney 9480644d9c Added file paths tests. (#2454) 2026-09-10 15:52:54 -07:00
afourney 73a26dac09 Expand test matrix to include windows-latest (#2453) 2026-09-10 15:36:19 -07:00
afourney e99a726687 Reject UNC and Windows device paths (similar to how netlocs are already rejected). 2026-09-10 15:02:35 -07:00
Çağdaş Yürekli 57a481e1aa fix(mcp): migrate to MCP SDK 2.x so 2026-07-28 clients can connect (#2363) 2026-09-10 12:20:11 -07:00
kevin 075c05bd4d fix(pptx): do not emit a heading for a slide with an empty title (#2442)
* fix(pptx): do not emit a heading for a slide with an empty title
2026-09-10 09:46:52 -07:00
liyrds 6270920c28 fix(zip): preserve content of entries with duplicate filenames (#2434) 2026-09-09 23:54:16 -07:00
Machen John b59e0643dc fix: prefer data-src over placeholder data URI in img src (#2417)
* fix: prefer data-src over placeholder data URI in img src
* Keep data URIs is keep_data_uris is True
2026-09-09 23:30:55 -07:00
freetg71527152-ui f217a288a7 Fix --list-plugins help text to reference correct --use-plugins flag (#2381) 2026-09-09 22:49:11 -07:00
Sushant Lokhande fef3a5de16 fix(epub): resolve percent-encoded manifest hrefs to ZIP entries (#2413) 2026-09-09 22:37:57 -07:00
SpongeBob f9042b4227 Fix PPTX shape sorting treating top-zero as missing (#2408) 2026-09-09 16:34:11 -07:00
kevinandafourney 24b9e122ea Do not convert an undecodable file to the word "None" (#2418)
* Do not convert an undecodable file to the word "None"
* Inline the fallback conversions.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-09 16:28:27 -07:00
SpongeBob 0504d7d50f Fix Zip converter forwarding kwargs to nested conversions (#2409)
* Fix Zip converter forwarding kwargs to nested conversions
* Add regression test for Zip converter kwargs forwarding
* Fixed args construction.
* test(zip): cover DOCX member conversion with and without an archive URL
2026-09-09 16:08:01 -07:00
Erwin Kersten 83c7730093 fix(docker): move base image off EOL bullseye (#2421) 2026-09-09 15:32:41 -07:00
kevin 5dc7bafcaf fix(epub): forward conversion options to the HTML converter (#2426) 2026-09-09 15:30:01 -07:00
afourney 75cc34a364 fix(rss): preserve complete feed content and resolve relative link (#2432)
* fix(rss): preserve complete feed content and resolve relative link
* Fix extra whitespace, failing tests.
2026-09-09 15:20:57 -07:00
kevin 2e71e117ea fix(ipynb): strip the UTF-8 BOM so a notebook is not emitted as raw JSON (#2425) 2026-09-09 11:33:59 -07:00
kevin 481f22a9a9 fix(pptx): do not emit a Notes heading for a slide with no notes (#2427) 2026-09-09 11:29:47 -07:00
kevin a04fd8bd19 fix(rss): drop layout whitespace from feed and entry titles (#2428)
* fix(rss): drop layout whitespace from feed and entry titles
* Refactor return value handling in RSS converter
* Do not flatten descriptions.
2026-09-09 11:23:30 -07:00
Xing Zheng cb785cb2f7 fix(rss): support namespace-prefixed Atom feeds (#2429) 2026-09-09 09:49:30 -07:00
Kim Do Yeon ab6792caca Fix CSV parsing for CR-only line endings (#2412)
Enable universal newline handling without translation in the CSV text stream. Add public API regression coverage for LF, CRLF, and CR records, optional final newlines, and quoted multiline cells.

Fixes #2411.
2026-09-09 09:36:37 -07:00
afourney 21887b3bf0 Preserve all CSV columns by matching the widest row (#2430) 2026-09-09 08:40:53 -07:00
kapil971390 4459ed0115 fix(pptx): chart_title.text_frame is never None -- use has_text_frame (#2385)
The #2194 guard (chart.chart_title.text_frame is not None) doesn't
check what it's meant to. python-pptx's ChartTitle.text_frame is
documented as destructive -- it creates a text frame if one isn't
already present -- so it can never return None against the real,
installed python-pptx (1.0.2, unpinned in pyproject.toml).

has_text_frame is the property that actually reflects presence,
without the auto-creating side effect. Swaps both converters to use
it, updates the existing mock-based tests to match the real API
contract (has_text_frame=False instead of text_frame=None), and adds
a positive-path test for each so both branches are covered.

Verified against the real, installed python-pptx (not just source
reading): with has_title=True and no title text ever set,
chart_title.text_frame is None is False and .text_frame.text is ''
-- confirming the existing guard is a no-op.
2026-09-04 10:04:57 -07:00
afourney a035350629 Remove uv lockfile that was accidentally committed. (#2380) 2026-09-03 20:57:50 -07:00
afourney bc90c5d7a5 Test markitdown-ocr in CI against the sibling markitdown (#2378)
* Test markitdown-ocr in CI against the sibling markitdown
* Run the full matrix.
2026-09-03 20:28:53 -07:00
afourney 2dffd8bb6e Bump versions. (#2377) 2026-09-03 17:21:53 -07:00
Nefelibataandafourney 58783e3272 fix(docx): handle unknown math functions in OMML converter without crashing (#2268)
* fix(doc-intel): default api_version to None in DocumentIntelligenceConverter

Do not set a default api_version string in DocumentIntelligenceConverter. If api_version is omitted or None, avoid passing api_version kwarg to DocumentIntelligenceClient so Azure SDK uses its native default version.

Fixes #1904

* fix(docx): handle unknown math functions in OMML converter without crashing

Add common standard math functions to FUNC dictionary in latex_dict.py and fallback to \operatorname{} for unknown functions in omml.py instead of throwing NotImplementedError.

Fixes #1982

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 15:50:58 -07:00
hanhan761 dd2df72913 fix: ImageConverter gracefully handles LLM API failures (fixes #1942) (#1948) 2026-09-03 15:33:16 -07:00
Lazizbek Ergashevandafourney 697a2a0e0b fix(rss): treat Atom text content and summary as plain text (#2374)
* fix(rss): treat Atom text content and summary as plain text
* Parse content types more carefully.
* Don't pass text through the html converter.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 15:01:34 -07:00
hanhan761andafourney 94a84834a2 chore: remove unused mammoth import from PlainTextConverter (#1951) (#1953)
* fix: handle LLM API errors gracefully in ImageConverter (#1942)

If the OpenAI API call fails (auth error, rate limit, network timeout, etc.), the exception propagates and crashes the entire conversion. Wrap the call in try/except to gracefully skip the LLM caption, matching the pattern already used by PptxConverter.

* chore: remove unused mammoth import from PlainTextConverter (#1951)

The mammoth import was copied from DocxConverter but PlainTextConverter does not need it. Remove the dead code and unused _dependency_exc_info variable.

* Revert unrelated changes.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 12:58:47 -07:00
c870d49284 fix: suppress pydub RuntimeWarning when ffmpeg is missing (fixes #1685) (#1985)
* fix: suppress pydub RuntimeWarning when ffmpeg is missing

Closes #1685

Importing pydub emits a RuntimeWarning when ffmpeg/avconv is not
installed on the system. This warning appears even when audio
transcription is not actively being used.

_transcribe_audio.py already wraps the imports in a
warnings.catch_warnings block that suppresses DeprecationWarning and
SyntaxWarning. Add a filter for the ffmpeg RuntimeWarning as well.

* Modify RuntimeWarning filter for ffmpeg

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: afourney <adamfo@microsoft.com>
Co-authored-by: afourney <adam.fourney@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-03 12:41:55 -07:00
Mukunda Rao Katta 7de6ce4943 Support short YouTube URLs in converter (#1882) 2026-09-03 12:22:34 -07:00