413 Commits
Author SHA1 Message Date
Machen John 1f9530a6d3 fix(wikipedia): drop the site suffix from italicised titles (#2458)
Titles rendered with markup have no mw-page-title-main span, so the
converter fell through to the document <title>, suffix and all.
2026-09-30 22:24:20 -07:00
Amit Anand 44f3a0ed71 Use frozenset for OMML direct tag lookup (#1919)
Replace __direct_tags tuple with frozenset in oMath2Latex for O(1) membership checks in process_unknow(). No behavioral change.
2026-09-30 22:16:01 -07:00
Yasir Alibrahem f70e637af2 optimize non-seekable stream buffering using shutil.copyfileobj (#1510) 2026-09-30 22:11:23 -07:00
hchmch 5610ced394 Add language parameter for the audio transcription (#1295)
* Add language parameter for the audio transcription
* style: format audio language option with Black
2026-09-30 22:00:00 -07:00
Octopus 53f38ea102 fix: handle None sys.stdout.encoding in CLI output (#1656)
* fix: handle None sys.stdout.encoding in CLI output (fixes #1597)
2026-09-30 16:20:37 -07:00
Gopal Bagaswar 533ea574b3 fix(wikipedia): accept mobile (m.wikipedia.org) URLs in WikipediaConverter (#1850) 2026-09-30 16:09:02 -07:00
Yasir Alibrahem 199fe8b185 increase chunk size from 512 bytes to 64 KB for HTTP response reading (#1511) 2026-09-30 16:00:48 -07:00
afourney b8f79c57eb Bump versions or markitdown and markitdown-ocr (#2545) v0.1.8 2026-09-21 14:02:08 -07:00
Nithin JambulaandAdam Fourney d51937de3a fix: bump youtube-transcript-api to >=1.2.3 for Python 3.14 support (#2407)
* fix: bump youtube-transcript-api to >=1.2.3 for Python 3.14 support
* Add 3.14 to test matrix, and limit advertised compatibility.

---------

Co-authored-by: Adam Fourney <adamfo@microsoft.com>
2026-09-21 09:47:07 -07:00
afourney 945314a45d Refactor Office OCR converters to reuse core conversion pipelines (#2506)
* Add _image_to_html helper for docx ocr conversion.
* Ensure compatibility with core markitdown version
* Have pptx, xlsx and docx expose an _image_to_html semi-private method, overridable by plugins.
* Fix conditional check for llm_description
* Update relationships function calls for drawings and images
* Forward stream_info to LLMVisionOCRService
2026-09-16 10:23:08 -07:00
afourney eb31b5c945 bump_versions (#2490) v0.1.8b2 2026-09-14 09:48:43 -07:00
afourney cc0ca9edd8 Preserve whitepace underlines (#2477)
* Preserve whitespace underlines
2026-09-12 13:25:31 -07:00
afourney 3a6ce44358 Throw an error when llm client fails with image converter. (#2476) 2026-09-12 12:11:30 -07:00
afourney 5640da7145 fix(youtube): fall back to HTML when no video content is extracted (#2469) 2026-09-11 15:52:59 -07:00
afourney 485a7b93b6 fix(docx): preserve namespaces when repairing stylesheets (#2467)
* fix(docx): preserve namespaces when repairing stylesheets
* Fixed d:strike as well
2026-09-11 15:05:05 -07:00
afourney 75e7114598 fix: avoid splitting UTF-8 characters during charset detection (#2466) 2026-09-11 14:35:07 -07:00
afourney 2277fff52d Improve efficiency of escaping pipes in CSVs (#2464) 2026-09-11 12:24:26 -07:00
dickbown f0c5d01fbe perf(csv): Batch-remove blank rows to eliminate the quadratic overhead in table conversion 2026-09-11 11:12:55 -07:00
afourney 40658cdcc5 Fix ANSI Outlook MSG decoding for Japanese code pages and padded strings (#2462) 2026-09-11 11:05:33 -07:00
afourney 9edea6f5f6 Pin cryptography to 46.0.3 (in the CI, only) to temporarily enable ARM tests of the rest of the suite (#2460) 2026-09-11 10:22:13 -07:00
afourney 8c11f50275 Add arm to matrix. Simplify matrix. (#2459)
* Add arm to matrix. Simplify matrix.
* Fixed a broken powershell CLI, which caused the test suite to exit without running tests.
2026-09-11 10:11:29 -07:00
afourney 9480644d9c Added file paths tests. (#2454) 2026-09-10 15:52:54 -07:00
afourney 73a26dac09 Expand test matrix to include windows-latest (#2453) 2026-09-10 15:36:19 -07:00
afourney e99a726687 Reject UNC and Windows device paths (similar to how netlocs are already rejected). 2026-09-10 15:02:35 -07:00
Çağdaş Yürekli 57a481e1aa fix(mcp): migrate to MCP SDK 2.x so 2026-07-28 clients can connect (#2363) 2026-09-10 12:20:11 -07:00
kevin 075c05bd4d fix(pptx): do not emit a heading for a slide with an empty title (#2442)
* fix(pptx): do not emit a heading for a slide with an empty title
2026-09-10 09:46:52 -07:00
liyrds 6270920c28 fix(zip): preserve content of entries with duplicate filenames (#2434) 2026-09-09 23:54:16 -07:00
Machen John b59e0643dc fix: prefer data-src over placeholder data URI in img src (#2417)
* fix: prefer data-src over placeholder data URI in img src
* Keep data URIs is keep_data_uris is True
2026-09-09 23:30:55 -07:00
freetg71527152-ui f217a288a7 Fix --list-plugins help text to reference correct --use-plugins flag (#2381) 2026-09-09 22:49:11 -07:00
Sushant Lokhande fef3a5de16 fix(epub): resolve percent-encoded manifest hrefs to ZIP entries (#2413) 2026-09-09 22:37:57 -07:00
SpongeBob f9042b4227 Fix PPTX shape sorting treating top-zero as missing (#2408) 2026-09-09 16:34:11 -07:00
kevinandafourney 24b9e122ea Do not convert an undecodable file to the word "None" (#2418)
* Do not convert an undecodable file to the word "None"
* Inline the fallback conversions.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-09 16:28:27 -07:00
SpongeBob 0504d7d50f Fix Zip converter forwarding kwargs to nested conversions (#2409)
* Fix Zip converter forwarding kwargs to nested conversions
* Add regression test for Zip converter kwargs forwarding
* Fixed args construction.
* test(zip): cover DOCX member conversion with and without an archive URL
2026-09-09 16:08:01 -07:00
Erwin Kersten 83c7730093 fix(docker): move base image off EOL bullseye (#2421) 2026-09-09 15:32:41 -07:00
kevin 5dc7bafcaf fix(epub): forward conversion options to the HTML converter (#2426) 2026-09-09 15:30:01 -07:00
afourney 75cc34a364 fix(rss): preserve complete feed content and resolve relative link (#2432)
* fix(rss): preserve complete feed content and resolve relative link
* Fix extra whitespace, failing tests.
2026-09-09 15:20:57 -07:00
kevin 2e71e117ea fix(ipynb): strip the UTF-8 BOM so a notebook is not emitted as raw JSON (#2425) 2026-09-09 11:33:59 -07:00
kevin 481f22a9a9 fix(pptx): do not emit a Notes heading for a slide with no notes (#2427) 2026-09-09 11:29:47 -07:00
kevin a04fd8bd19 fix(rss): drop layout whitespace from feed and entry titles (#2428)
* fix(rss): drop layout whitespace from feed and entry titles
* Refactor return value handling in RSS converter
* Do not flatten descriptions.
2026-09-09 11:23:30 -07:00
Xing Zheng cb785cb2f7 fix(rss): support namespace-prefixed Atom feeds (#2429) 2026-09-09 09:49:30 -07:00
dependabot[bot] a2a7a5293d chore(deps): bump actions/setup-python from 5.6.0 to 7.0.0 (#2431)
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 5.6.0 to 7.0.0.
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](https://github.com/actions/setup-python/compare/v5.6.0...5fda3b95a4ea91299a34e894583c3862153e4b97)

---
updated-dependencies:
- dependency-name: actions/setup-python
  dependency-version: 7.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
2026-09-09 09:40:50 -07:00
Kim Do Yeon ab6792caca Fix CSV parsing for CR-only line endings (#2412)
Enable universal newline handling without translation in the CSV text stream. Add public API regression coverage for LF, CRLF, and CR records, optional final newlines, and quoted multiline cells.

Fixes #2411.
2026-09-09 09:36:37 -07:00
afourney 21887b3bf0 Preserve all CSV columns by matching the widest row (#2430) 2026-09-09 08:40:53 -07:00
afourney b6e8bbdce6 Updated README (#2393)
* MarkItDown is now a standalone project, and the README is updated to reflect this.
2026-09-06 21:58:12 -07:00
kapil971390 4459ed0115 fix(pptx): chart_title.text_frame is never None -- use has_text_frame (#2385)
The #2194 guard (chart.chart_title.text_frame is not None) doesn't
check what it's meant to. python-pptx's ChartTitle.text_frame is
documented as destructive -- it creates a text frame if one isn't
already present -- so it can never return None against the real,
installed python-pptx (1.0.2, unpinned in pyproject.toml).

has_text_frame is the property that actually reflects presence,
without the auto-creating side effect. Swaps both converters to use
it, updates the existing mock-based tests to match the real API
contract (has_text_frame=False instead of text_frame=None), and adds
a positive-path test for each so both branches are covered.

Verified against the real, installed python-pptx (not just source
reading): with has_title=True and no title text ever set,
chart_title.text_frame is None is False and .text_frame.text is ''
-- confirming the existing guard is a no-op.
2026-09-04 10:04:57 -07:00
afourney a035350629 Remove uv lockfile that was accidentally committed. (#2380) 2026-09-03 20:57:50 -07:00
afourney e153a92a18 Be sure to run the proper Python version. (#2379) v0.1.8b1 2026-09-03 20:36:11 -07:00
afourney bc90c5d7a5 Test markitdown-ocr in CI against the sibling markitdown (#2378)
* Test markitdown-ocr in CI against the sibling markitdown
* Run the full matrix.
2026-09-03 20:28:53 -07:00
afourney 2dffd8bb6e Bump versions. (#2377) 2026-09-03 17:21:53 -07:00
afourney 0e9ede3e05 Update README to help scope PRs. (#2376)
* Update README to help scope PRs.
2026-09-03 16:26:27 -07:00