Production 2026 MT pipeline. NLLB-200 / mBART / Marian by use case.
Working BLEU + chrF computation in pure Python (with the exact
sacrebleu caveat about tokenization inconsistency being the most
common silent benchmark bug).
Names the five production failure modes research flags: hallucination,
off-target generation, terminology drift, formality mismatch, length
explosion on short input. Each with a concrete mitigation.
Fine-tuning recipe for domain adaptation. Decision table for picking
between NLLB-200, mBART, Marian, and LLM-based translation.
Ship artifact: mt-evaluator skill with the 5-point shipping checklist
and refusal rules (no ship without language-ID check, no reference-free
scoring without explicit opt-in).
~75 minutes. Prerequisites lesson 05/10 and 05/04.