WritingCohereCoherepublished Oct 29, 2025seen Jun 26

Findings Of The Wmt25 Multilingual Instruction Shared Task Persistent Hurdles In Reasoning Generation And Evaluation 2025 10 29

Open original ↗

Captured source

source ↗

Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation North Mini Code. Cohere's first model for developers. Learn more

Oct 29, 2025 Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation The WMT25 Multilingual Instruction Shared Task (MIST) introduces a benchmark to evaluate large language models across 30 languages, covering five problem types, and highlights limitations in automatic evaluation while providing a standardized framework for future progress.

read the paper

Authors

Tom Kocmi, Ekaterina Artemova, Eleftherios Avramidis, Eleftheria Briakou, Pinzhen Chen, Marzieh Fadaee, Markus Freitag, Roman Grundkiewicz, Yupeng Hou, Philipp Koehn, Julia Kreutzer, Saab Mansour, Stefano Perrella, Lorenzo Proietti, Parker Riley, Eduardo Sánchez, Patrícia Schmidtová, Mariya Shmatova, Vilém Zouhar

Abstract

The WMT25 Multilingual Instruction Shared Task (MIST) introduces a benchmark to evaluate large language models (LLMs) across 30 languages. The benchmark covers five types of problems: machine translation, linguistic reasoning, open-ended generation, cross-lingual summarization, and LLM-as-a-judge. We provide automatic evaluation and collect human annotations, which highlight the limitations of automatic evaluation and allow further research into metric meta-evaluation. We run on our benchmark a diverse set of open- and closed-weight LLMs, providing a broad assessment of the multilingual capabilities of current LLMs. Results highlight substantial variation across sub-tasks and languages, revealing persistent challenges in reasoning, cross-lingual generation, and evaluation reliability. This work establishes a standardized framework for measuring future progress in multilingual LLM development.

multilingual Evaluation Reasoning

Related works

Research The Culture Funnel: You can’t align what isn’t in the data

Read

Research Tiny Aya: Bridging Scale and Multilingual Depth

Read

Research Unlocking Reasoning Capability on Machine Translation in Large Language Models

Read

Notability

notability 5.0/10

Shared task findings, substantive but not a model release.