North Small Translate
Captured source
source ↗Skip to content AI for Empowerment: Your freedom. Your focus. See how AI gives you more time for what truly moves you. Explore now
Products
Solutions
Resources
Blog
Research
Company
Sign in
Request a demo
Platform North
Enterprise-ready AI for business
Compass
Intelligent search and discovery
Models Command
Generative language models
Transcribe New
Speech recognition model
North Small Translate NEW
Machine translation model
Parse New
Document parsing model
Embed
Semantic representation model
Rerank
Retrieval optimization model
Models Overview
Product Products Overview
Total Cost of AI Ownership
Pricing
Featured Command: High-performance generative AI models for real-world applications
Deploy Model Vault
Dedicated model inference platform
Private Deployments
On-prem or isolated VPCs
Security
Protect your data at every stage
See deployment options
By Industry Financial Services
Public Sector
Technology
Telecommunications
Energy and Utilities
Healthcare and Life Sciences
Manufacturing
Featured Model Vault provides fully-isolated, performant inference with Saas simplicity
Insights Customer Stories
For Developers Developers
Models Overview
Docs
Discord
LLM University
Connect Partners
Events
Webinars
Merch Store
Featured How CoreWeave used Cohere North to transform its customer support in 90 days
Blog
The latest news, launches, and insights
Read more
The state of sovereign AI adoption in 2026
Cohere and the University of Waterloo launch partnership to strengthen Canada’s AI talent pipeline
Introducing North Automations: Intelligent workflow orchestration
Research Cohere Labs
Cohere’s ML research lab
Explorations Future(s) of Work
How will AI change the way we work?
Aya Models
Multilingual AI at scale
All Papers
Initiatives Research Scholars
Finding the new generation of ML talent
Open Science Community
Championing global, open science
Catalyst Grants
Supporting impactful ML endeavors
Resources Blog
Hugging Face
Events
Featured The future of work debate has an evidence problem
About
Careers
Newsroom
Sep 10, 2026
4 minute read
Introducing North Small Translate: A leading sovereign open-weight machine translation model Outsized performance, right-sized footprint — translation built for speed and cost efficiency.
Today, we're releasing North Small Translate, a mixture-of-experts machine translation model with strong performance across 50+ languages. Across WMT26 benchmarks,¹ North Small Translate achieves an 83.6 score across all languages, outperforming proprietary models like DeepL and Google Translate, as well as open-weight alternatives such as Gemma 4 31B (off), GLM 5.2, and Mistral Large 3.
North Small Translate marks a significant milestone as Cohere's first translation model in the North model family. It builds on our multilingual and translation lineage — from the Tiny Aya model family to Command A Translate — and represents a clear next step in Cohere’s commitment to offering high-quality machine translation wherever it is needed.
Now available for research and non-commercial use under a CC BY-NC 4.0 license, North Small Translate advances Cohere’s mission to make sovereign AI a technological reality.
Visit Hugging Face to download the weights - available in several near-lossless quantizations - explore our HuggingFace space to demo the model, and read our implementation guides . Snapshot
Model North-Small-Translate-1.0
License Open-weights, non-commercial
Architecture MoE
Model size 218B total; 25B active
Context length 16k input, 16k output
Input modalities Text
Output modalities Text
Languages Supports 50+ languages. Full list
Optimized for Machine translation
Hardware (minimum) 1× B200 @ W4A4 2× H100s @ W4A4
Leading translation quality in the open ecosystem North Small Translate outperforms similarly sized open-weight models under 1T parameters and API-based translation models in various dimensions of machine translation on average. In WMT model evaluations, North Small Translate leads with an WMT26 All Languages benchmark score of 83.60, compared with 81.56 for Qwen 3.5 397B A17B, 76.50 for GLM 5.2 FP8, 81.37 for DeepL NextGen, 79.46 for Gemma 4 31B (on), and 68.20 for Google Translate. North Small Translate (Agentic) — which can find errors and fix errors in translation — scores even higher, at 84.36.² Image 1: Overall WMT26 scores across all languages across major models, benchmarked for WMT 2026 using GPT-5.6-Sol as a judge.
North Small Translate performs strongly across 32 high-resource languages and 18 additional languages. North Small Translate is the most consistent performer across the full spread of regions, without the sharp regional drop-offs seen in other models of its size. On average across all languages, it is the best-performing dedicated machine translation model in this evaluation — open or closed. Image 2: Overall WMT scores by regional average across DeepL, Google Translate, and Gemma 31B (on), and North Small Translate using GPT-5.6-Sol as a judge. At the regional level, North Small Translate punches above its weight, beating Gemma 4 31B (on) outright in Europe (82.2 vs. 73.9) while running essentially even with it in South Asia (86.2 vs. 86.7).
At the regional level, both North Small Translate and its Agentic counterpart beat Gemma 4 31B (on) outright across Europe — EU languages (82.74 Agentic / 82.17 standard vs. 72.73) and non-EU European languages (81.52 / 81.24 vs. 75.90) — while running essentially even with it in South Asia (87.13 / 86.16 vs. 88.04).
Both versions also outperform DeepL NextGen across every non-European region tested — MENA, South Asia, Southeast Asia, and East Asia — with the largest advantage in South Asia and MENA (roughly 8–10 points ahead of DeepL), a moderate edge in Southeast Asia (about 4–5 points), and the narrowest edge in East Asia (about 1-3 points, with the standard model closing in on DeepL's 85.41 score).
Increased throughput for faster workflows North Small Translate is built for high-throughput generation, prioritizing raw output speed even as concurrency scales.
In our testing, North Small Translate achieved up to 1.4x higher output throughput than Gemma 4 31B TP1 (1 x GPU) under identical concurrency levels and hardware configurations — 112 vs. 81 Output Tokens per Second (TOPS) at low concurrency and 39 vs. 30 TOPS at high concurrency. In practical terms, that's 30-38% more tokens...
Excerpt shown — open the source for the full document.
Notability
notability 4.0/10Cohere small translation model, very low HN traction