explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionaryagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • Quick reference: what people are asking
  • What Cohere shipped
  • WMT26 scores: what the numbers actually mean
  • Agentic translation: the 84.36 workflow
  • Throughput, long documents, and cost
  • Sovereign AI and the RWS partnership
  • What this means for what you build or pay
  • Honest limitations
  • Getting started
  • Related reading
← Back to blog

explainx / blog

Cohere North Small Translate: WMT26 Leader at 83.6 (Open Weights)

Cohere, Machine Translation, Open Weights, WMT26, Sovereign AI, MoE Models

Cohere released North Small Translate on Sept 10, 2026 — a 218B MoE translation model scoring 83.60 on WMT26, beating DeepL and Qwen 3.5. CC BY-NC weights on HF.

Sep 11, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
Cohere North Small Translate: WMT26 Leader at 83.6 (Open Weights)

On September 10, 2026, Cohere released North Small Translate — a 218-billion-parameter mixture-of-experts model built exclusively for machine translation across 50+ languages. In Cohere's own WMT26 evaluation, judged by GPT-5.6-Sol, it scores 83.60 across all languages — ahead of DeepL NextGen (81.37), Qwen 3.5 397B A17B (81.56), and Google Translate (68.20). With Cohere's agentic multi-pass workflow — a second pass that finds and fixes translation errors — the score climbs to 84.36.

That makes North Small Translate the third major open-weight release from Cohere's North family in four months, after Command A+ (May 2026, Apache 2.0 enterprise model) and North Mini Code (June 2026, Apache 2.0 coding model). The translation model shares Command A+'s 218B total / 25B active MoE footprint but ships under CC BY-NC 4.0 — research and non-commercial only on Hugging Face, with commercial deployment routed through RWS's Language Weaver platform.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

Quick reference: what people are asking

table · 2 cols
QuestionAnswer
When did it launch?September 10, 2026 — Cohere blog announcement
What is the WMT26 score?83.60 standard; 84.36 with agentic multi-pass (Cohere-evaluated, GPT-5.6-Sol judge)
Is it commercially open?No — CC BY-NC 4.0 on Hugging Face; commercial via RWS Language Weaver
How big is the model?218B total params, 25B active MoE; 16k input / 16k output context
Minimum hardware?1× B200 or 2× H100 at W4A4 quantization
How does it compare to DeepL?Cohere reports beating DeepL NextGen in every non-European region tested, with 8–10 point leads in South Asia and MENA
Same family as North Mini Code?Same "North" branding, different size and license — Mini Code is 30B Apache 2.0 for coding; Translate is 218B CC BY-NC for MT
Official WMT26 winner?Not yet — WMT26 uses human evaluation and has not published automatic leaderboard results

What Cohere shipped

Cohere positions North Small Translate as "outsized performance, right-sized footprint" — a dedicated translation model rather than a general LLM asked to translate as a side task. The model continues Cohere's multilingual lineage from Tiny Aya through Command A Translate, but this is the first North-family model optimized end-to-end for translation throughput and cost.

Model snapshot

table · 2 cols
SpecValue
Model IDNorth-Small-Translate-1.0
ArchitectureSparse MoE
Total parameters218B
Active parameters25B
Context16k input, 16k output
ModalitiesText in, text out
Languages50+ (32 high-resource + 18 additional)
License (HF weights)CC BY-NC 4.0 (research / non-commercial)
Minimum hardware1× B200 @ W4A4 or 2× H100 @ W4A4
Hugging FaceCohereLabs/North-Small-Translate-1.0 (+ w4a16 quant)
DemoHugging Face Space (linked from Cohere blog)

The hardware requirement mirrors Command A+: Cohere's W4A4 lossless quantization lets a model this large run on two H100s instead of a full multi-node cluster. For teams already running Command A+ on-premises for RAG or agent workflows, North Small Translate is designed to slot into the same sovereign infrastructure stack.

WMT26 scores: what the numbers actually mean

Cohere's headline number — 83.60 on WMT26 All Languages — needs context before you paste it into a procurement deck.

Cohere's scoring scale

Cohere defines its WMT benchmark score ranges as:

  • 0–20: not acceptable
  • 20–40: borderline
  • 40–60: acceptable
  • 60–80: good with major errors
  • 80–100: perfect or with minor errors

At 83.60, Cohere places North Small Translate in the top band — "perfect or with minor errors" — across the full language spread. The agentic variant at 84.36 adds roughly 0.76 points through error detection and correction passes.

Head-to-head comparison (Cohere-evaluated)

table · 2 cols
SystemWMT26 All Languages score
North Small Translate (Agentic)84.36
North Small Translate83.60
Qwen 3.5 397B A17B81.56
DeepL NextGen81.37
Gemma 4 31B (on)79.46
GLM 5.2 FP876.50
Google Translate68.20

Cohere also breaks results by region. North Small Translate and its agentic variant beat Gemma 4 31B outright across European language groups — EU languages score 82.17 standard / 82.74 agentic versus Gemma's 72.73 — while running essentially even with Gemma in South Asia (86.16 vs 88.04 for Gemma). Against DeepL NextGen, Cohere reports advantages in every non-European region: roughly 8–10 points in South Asia and MENA, 4–5 in Southeast Asia, and 1–3 in East Asia.

The caveat: vendor eval vs official WMT26

This is where explainx.ai's read diverges from a press-release recap. The WMT26 shared task page states organizers will not release automatic preliminary results and that systems are evaluated with humans. Submission deadlines ran through July 2026; official human evaluation papers follow the main conference timeline.

Cohere's 83.60 figure is from Cohere's own benchmark harness using GPT-5.6-Sol as an automatic judge on WMT26-style test data — not an official WMT26 human score. That does not make the comparison useless; vendor-run LLM-judge evals are now standard for rapid model comparison (see explainx.ai's guide on how to read AI benchmarks). But it does mean:

  1. Do not cite 83.60 as an official WMT26 ranking until human evaluation results publish.
  2. Compare like-for-like — Cohere judged all systems in the same harness, so relative ordering within their table is more meaningful than the absolute number.
  3. Run your own eval on your language pairs and domains before switching production traffic.

If you need a framework for interpreting vendor benchmark claims more broadly, the AI benchmarks complete guide covers contamination, judge bias, and when automatic scores diverge from human preference.

Agentic translation: the 84.36 workflow

The North Small Translate (Agentic) variant at 84.36 is not a separate model weights file — it is a multi-pass workflow where the system translates, identifies errors, and re-translates corrected segments. Cohere describes it as finding and fixing translation mistakes rather than a single-shot prompt.

For builders already running agentic pipelines — whether with North Mini Code in OpenCode or document ingestion through Cohere Parse 5 — the pattern is familiar: specialization beats asking one general model to do everything in one pass. Translation quality gains of ~0.76 points on Cohere's scale may sound modest, but in the 80–100 band where errors are already minor, that delta can matter for regulated content, legal localization, and brand voice consistency.

A minimal agentic translation loop you can prototype locally:

python
import os
from cohere import ClientV2

co = ClientV2(api_key=os.environ["CO_API_KEY"])

def agentic_translate(text: str, source: str, target: str) -> str:
    # Pass 1: initial translation
    draft = co.chat(
        model="north-small-translate-1.0",
        messages=[{
            "role": "user",
            "content": f"Translate from {source} to {target}:\n\n{text}",
        }],
    )
    translation = draft.message.content[0].text

    # Pass 2: error detection + correction (agentic refinement)
    refined = co.chat(
        model="north-small-translate-1.0",
        messages=[{
            "role": "user",
            "content": (
                f"Review this {source}→{target} translation. "
                f"Fix any errors, awkward phrasing, or meaning drift. "
                f"Return only the corrected translation.\n\n"
                f"Source:\n{text}\n\nDraft:\n{translation}"
            ),
        }],
    )
    return refined.message.content[0].text

Exact API model IDs and agentic endpoints may differ from this sketch — check Cohere's implementation guides linked from the announcement. The point is architectural: quality is a workflow property, not just a weights property, and Cohere is productizing that for translation the same way multi-pass parsing evals work for document AI.

Throughput, long documents, and cost

Speed vs Gemma 4 31B

Cohere reports North Small Translate achieves up to 1.4× higher output throughput than Gemma 4 31B on identical hardware — 112 vs 81 output tokens per second at low concurrency, and 39 vs 30 TOPS at high concurrency. That is a 30–38% tokens-per-second advantage, which compounds on batch translation jobs spanning thousands of documents.

Long-context translation

Long documents break many translation systems. Cohere's long-context evaluation — translating two book chapters in a single call, scored per paragraph via xComet-XL — puts North Small Translate at 48.9, more than double Google Translate (21.3) and Gemma 4 31B (19.4). With 16k input and 16k output windows, the model is sized for chapter-scale content without the quality collapse Cohere attributes to general-purpose LLMs on long inputs.

Cost per task

Cohere's commercial pricing comparison (for enterprises evaluating licensed deployment) shows North Small Translate at $0.000676 per task averaging 661 tokens, versus $0.038928 per task for Gemini 3.1 Pro Preview (high) — a 5,762% cost gap on Cohere's published per-task math. Similarly sized open models like Qwen 3.5 397B A17B land around $0.004525 per task; Command A+ at $0.005158.

Those figures come from Cohere's pricing page token/character rates applied to average task sizes in their eval — not a guarantee for your corpus. Legal, medical, and marketing copy with heavy terminology will consume more tokens per segment. Still, for high-volume localization pipelines, a dedicated MT model at sub-millisecond per-task economics is a different budget conversation than routing everything through a frontier reasoning model.

Sovereign AI and the RWS partnership

Cohere frames North Small Translate under its sovereign AI mission — organizations running translation on their own hardware without sending text to US hyperscaler APIs. That aligns with Command A+'s Apache 2.0 enterprise positioning, though the license story is stricter here: CC BY-NC 4.0 blocks commercial self-hosting without a separate deal.

For production localization, Cohere partnered with RWS (Language Weaver). RWS works with more than 80% of the world's top 100 brands on translation and localization. Language Weaver gives enterprises security, scalability, and a dedicated translation platform — the path Cohere recommends when open-weight research access is not enough.

The split mirrors Cohere's broader product map:

table · 2 cols
NeedPath
Research, academic eval, non-commercial prototypingHugging Face weights (CC BY-NC 4.0)
On-prem sovereign deployment (negotiated license)Cohere Model Vault + commercial terms
Full localization platform with TMS integrationRWS Language Weaver

Teams building multilingual RAG pipelines may combine North Small Translate for query/document translation with Parse 5 for document ingestion — two specialized models instead of one general LLM doing both jobs poorly. See RAG vs agentic RAG for when translation quality in the retrieval layer affects downstream answer faithfulness.

What this means for what you build or pay

If you run localization at scale: North Small Translate is the first open-weight MT model from Cohere that claims to beat DeepL and Qwen 3.5 on a broad multilingual eval — worth a pilot on your language pairs before the next contract renewal with a proprietary API.

If you self-host sovereign AI: The 2× H100 footprint matches Command A+. If you already deployed Command A+ for enterprise RAG, adding translation is an infrastructure-compatible extension — but check CC BY-NC terms before production use.

If you build agents: The agentic multi-pass pattern (draft → critique → refine) is reusable beyond translation. Same architectural lesson as multi-pass document parsing evals and coding agent review loops.

If you compare open vs closed: North Mini Code is Apache 2.0 for coding; North Small Translate is CC BY-NC for translation. Cohere is not using one license for the whole North family — read the card before you assume commercial freedom. Closed vs open-weight alternatives still applies.

Honest limitations

  • License: CC BY-NC 4.0 on Hugging Face is not Apache 2.0. Commercial products need RWS or a Cohere commercial agreement.
  • Benchmark independence: 83.60 is Cohere-judged with GPT-5.6-Sol, not WMT26 official human scores.
  • Hardware bar: Two H100s at W4A4 is "modest" by frontier standards but still enterprise-grade — not a laptop model.
  • 16k context ceiling: Long book chapters fit; entire books do not — plan chunking for book-length workflows.
  • No vision/multimodal: Text only. Image or PDF translation requires OCR upstream (Parse 5, Mistral OCR, or similar).
  • Regional variance: South Asia scores run even with Gemma 4 31B — not a universal blowout on every geography.
  • Agentic overhead: The 84.36 score costs extra inference passes — faster single-shot at 83.60 vs higher quality at 2× latency.

Getting started

  1. Download weights: CohereLabs/North-Small-Translate-1.0 on Hugging Face (plus w4a16 quantized variant).
  2. Try the demo: Hugging Face Space linked from Cohere's announcement.
  3. Read implementation guides: Cohere documentation for deployment specs and API examples.
  4. Commercial path: Contact RWS Language Weaver for licensed production deployment.

Related reading

  • Cohere North Mini Code — Apache 2.0 agentic coding model
  • Cohere Command A+ — 218B MoE enterprise model on 2 H100s
  • Cohere Parse 5 — document parsing at $1.50/1k pages
  • How to read AI benchmarks — vendor vs independent evals
  • AI benchmarks complete guide 2026
  • Closed source vs local open-weight alternatives
  • RAG vs agentic RAG
  • Choose open-weight vs closed AI models

Primary sources: Introducing North Small Translate · Hugging Face weights · WMT26 shared task


Model specs, WMT26 scores, throughput figures, and pricing comparisons are as published by Cohere on September 10, 2026. WMT26 official human evaluation results were not available at publication time — treat Cohere's 83.60 figure as a vendor-run automatic evaluation, not an official conference ranking. Re-run benchmarks on your language pairs before production deployment. Follow @explainx_ai for model launch coverage.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Aug 22, 2026

Inkling Is Free on OpenRouter With a 1M Context Window — the Fine Print

A headline reading "Thinky Machines makes Inkling MoE models free on OpenRouter with 1M context" is trending — and it's a garbled reference to Thinking Machines Lab, whose 975B-parameter Inkling has carried a free, rate-limited OpenRouter endpoint since its July 17, 2026 launch. Here's what "free" actually means, and what changed versus five weeks ago.

Aug 3, 2026

Inkling-Small: Thinking Machines’ 12B-Active MoE Matches Inkling at 1/4 Size

Two weeks after Inkling, Thinking Machines Lab released full weights for Inkling-Small — a 276B MoE with 12B active that beats its larger sibling on several reasoning and agentic benches, while Inkling keeps the knowledge lead.

Jul 16, 2026

Inkling: Thinking Machines Lab Open-Weights MoE for Customization (July 2026)

Thinking Machines Lab shipped Inkling on July 15, 2026 — a 975B-parameter MoE with full weights on Hugging Face, controllable thinking effort, native audio and vision, and a self-finetuning demo via Tinker and OpenCode. explainx.ai explains what it is good for, what it is not, and how it compares to Kimi, Nemotron, and closed frontier models.