Findings Of The Wmt25 Multilingual Instruction Shared Task Persistent Hurdles In Reasoning Generation And Evaluation 2025 10 29
Captured source
source ↗Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation North Mini Code. Cohere's first model for developers. Learn more
Oct 29, 2025 Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation The WMT25 Multilingual Instruction Shared Task (MIST) introduces a benchmark to evaluate large language models across 30 languages, covering five problem types, and highlights limitations in automatic evaluation while providing a standardized framework for future progress.
read the paper
Authors
Tom Kocmi, Ekaterina Artemova, Eleftherios Avramidis, Eleftheria Briakou, Pinzhen Chen, Marzieh Fadaee, Markus Freitag, Roman Grundkiewicz, Yupeng Hou, Philipp Koehn, Julia Kreutzer, Saab Mansour, Stefano Perrella, Lorenzo Proietti, Parker Riley, Eduardo Sánchez, Patrícia Schmidtová, Mariya Shmatova, Vilém Zouhar
Abstract
The WMT25 Multilingual Instruction Shared Task (MIST) introduces a benchmark to evaluate large language models (LLMs) across 30 languages. The benchmark covers five types of problems: machine translation, linguistic reasoning, open-ended generation, cross-lingual summarization, and LLM-as-a-judge. We provide automatic evaluation and collect human annotations, which highlight the limitations of automatic evaluation and allow further research into metric meta-evaluation. We run on our benchmark a diverse set of open- and closed-weight LLMs, providing a broad assessment of the multilingual capabilities of current LLMs. Results highlight substantial variation across sub-tasks and languages, revealing persistent challenges in reasoning, cross-lingual generation, and evaluation reliability. This work establishes a standardized framework for measuring future progress in multilingual LLM development.
multilingual Evaluation Reasoning
Related works
Research The Culture Funnel: You can’t align what isn’t in the data
Read
Research Tiny Aya: Bridging Scale and Multilingual Depth
Read
Research Unlocking Reasoning Capability on Machine Translation in Large Language Models
Read
Notability
notability 5.0/10Shared task findings, substantive but not a model release.