* fix: bump youtube-transcript-api to >=1.2.3 for Python 3.14 support
* Add 3.14 to test matrix, and limit advertised compatibility.
---------
Co-authored-by: Adam Fourney <adamfo@microsoft.com>
* Add _image_to_html helper for docx ocr conversion.
* Ensure compatibility with core markitdown version
* Have pptx, xlsx and docx expose an _image_to_html semi-private method, overridable by plugins.
* Fix conditional check for llm_description
* Update relationships function calls for drawings and images
* Forward stream_info to LLMVisionOCRService
* Fix Zip converter forwarding kwargs to nested conversions
* Add regression test for Zip converter kwargs forwarding
* Fixed args construction.
* test(zip): cover DOCX member conversion with and without an archive URL
Enable universal newline handling without translation in the CSV text stream. Add public API regression coverage for LF, CRLF, and CR records, optional final newlines, and quoted multiline cells.
Fixes#2411.
The #2194 guard (chart.chart_title.text_frame is not None) doesn't
check what it's meant to. python-pptx's ChartTitle.text_frame is
documented as destructive -- it creates a text frame if one isn't
already present -- so it can never return None against the real,
installed python-pptx (1.0.2, unpinned in pyproject.toml).
has_text_frame is the property that actually reflects presence,
without the auto-creating side effect. Swaps both converters to use
it, updates the existing mock-based tests to match the real API
contract (has_text_frame=False instead of text_frame=None), and adds
a positive-path test for each so both branches are covered.
Verified against the real, installed python-pptx (not just source
reading): with has_title=True and no title text ever set,
chart_title.text_frame is None is False and .text_frame.text is ''
-- confirming the existing guard is a no-op.
* fix(doc-intel): default api_version to None in DocumentIntelligenceConverter
Do not set a default api_version string in DocumentIntelligenceConverter. If api_version is omitted or None, avoid passing api_version kwarg to DocumentIntelligenceClient so Azure SDK uses its native default version.
Fixes#1904
* fix(docx): handle unknown math functions in OMML converter without crashing
Add common standard math functions to FUNC dictionary in latex_dict.py and fallback to \operatorname{} for unknown functions in omml.py instead of throwing NotImplementedError.
Fixes#1982
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix(rss): treat Atom text content and summary as plain text
* Parse content types more carefully.
* Don't pass text through the html converter.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: handle LLM API errors gracefully in ImageConverter (#1942)
If the OpenAI API call fails (auth error, rate limit, network timeout, etc.), the exception propagates and crashes the entire conversion. Wrap the call in try/except to gracefully skip the LLM caption, matching the pattern already used by PptxConverter.
* chore: remove unused mammoth import from PlainTextConverter (#1951)
The mammoth import was copied from DocxConverter but PlainTextConverter does not need it. Remove the dead code and unused _dependency_exc_info variable.
* Revert unrelated changes.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: suppress pydub RuntimeWarning when ffmpeg is missing
Closes#1685
Importing pydub emits a RuntimeWarning when ffmpeg/avconv is not
installed on the system. This warning appears even when audio
transcription is not actively being used.
_transcribe_audio.py already wraps the imports in a
warnings.catch_warnings block that suppresses DeprecationWarning and
SyntaxWarning. Add a filter for the ffmpeg RuntimeWarning as well.
* Modify RuntimeWarning filter for ffmpeg
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: afourney <adamfo@microsoft.com>
Co-authored-by: afourney <adam.fourney@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>