AGENTIC COMMONSAI industry briefings

繁中EN

TOPICAI Coding & Developer ToolsPUBLISHED 2026-10-07

All English articlesAI Engineering

Three Data Pillars That Teach Small Models to Reason in Your Language

If you ask today’s reasoning model a question in Swahili or Arabic, you’ll usually get the answer in your language and the thinking in English. That’s a real product problem: users can’t verify steps they can’t read, and knowledge that lives most naturally in the source language — cultural context, idioms, local nuance — gets flattened in an internal translation step.

Cohere’s research team published a data-mixing approach on October 6, 2026 that attacks this directly. Their result, Tiny Aya L2-Thinker (3.35B parameters, built on the Tiny Aya base model), reasons in the prompt language over 93% of the time across 60 languages, with mostly small accuracy costs. They’ve released the model weights and the multilingual reasoning data.

Why training, not prompting, is the fix

Cohere tested the cheap workaround first. Asking Qwen3.5-4B to think in the user’s language barely changed its behavior. Forcefully prefilling the opening words of the reasoning trace worked better but cost accuracy — and it’s not something you can ask end users to do. Their conclusion: reliable in-language reasoning has to come from the training data.

That matters for anyone evaluating models for multilingual products. A vendor’s claim that a model “supports” a language says nothing about which language the reasoning trace appears in. If inspectability of the thinking is part of your product’s value — for compliance review, tutoring, or debugging — that’s a question to ask explicitly.

The three-pillar data mix

The approach combines three data types, each doing a different job:

  • English reasoning data: about 1.7 million step-by-step solutions generated by gpt-oss-120b. This teaches problem-solving. Alone, it produces a model that reasons in the user’s language only 12.8% of the time.
  • Multilingual reasoning data: English reasoning traces translated into 44 languages, only ~5,000 examples per language. Adding this small set lifts the in-language rate from 12.8% to 86.1%, and accuracy improves too.
  • Multilingual non-reasoning data: plain Q&A pairs in target languages, no traces. This is the transfer mechanism — it lets the ability carry to languages the model never saw reasoning examples in.

The builder’s takeaway here is the asymmetry. The expensive, scarce resource (translated reasoning traces) is only needed in small quantities. The widely available resource (non-reasoning multilingual data) does the heavy lifting for generalization. You don’t have to build reasoning capability language by language.

What the trade-offs actually look like

Cohere compared the model against a twin trained identically except for English-only reasoning. On five of six benchmarks, accuracy dropped by at most two to three points; open-ended writing actually improved slightly. The exception is PolyMath, a competition-level math benchmark, where accuracy fell more noticeably — the authors suspect the missing reinforcement learning stage is the cause.

Low-resource languages held up well. Grouping evaluation languages by available web text into four tiers, Tiny Aya’s in-language rate stayed at 94%+ even on the lowest tier (languages like Swahili, Zulu, and Welsh). Larger comparison models degraded badly: Magistral-Small-24B collapsed from 72% to 5%, and M-Thinker-7B’s traces roughly doubled in length. Tiny Aya also averaged under 5,000 thinking tokens, which the team attributes to its multilingually optimized tokenizer.

One more finding worth internalizing: across the models tested, long reasoning traces correlated with repetition, not deliberation. Qwen3.5-4B routinely burned 10,000–25,000 tokens cycling through the same steps without reaching an answer. Token count is not a proxy for effort.

Why this fits Cohere’s enterprise push

This research is a natural companion to Cohere’s agent platform work — their North 2 release made the case that enterprises want agents they can govern, and a readable reasoning trace is a governance feature. A reviewer in São Paulo or Nairobi who can audit every reasoning step has a very different relationship with the system than one staring at an English monologue.

The limitations are honest: competition math still suffers, and the 3.35B size means this is evidence about data mixing more than a production-grade reasoner. But if you’re building for non-English users, the actionable next step is to start measuring the in-language reasoning rate of the models you deploy — accuracy benchmarks alone will hide the gap.

Sources

AGENTIC COMMONSOperated by PHLEGON LABS
SHAREXEMAIL
Support us

Related reading

  1. Jev's Decision-Only Model: Where It Fits in Your Stack

    TypeSafe's Jev returns typed decisions with calibrated probabilities in 70–500 ms, changing where you can afford to put judgment calls.

    AI

  2. What Fyxer's 53% Draft Acceptance Rate Changes for How You Build Trustworthy AI Assistants

    Fyxer's specialized-model email system shows how fine-tuning on real assistant workflows and user edits builds AI trust.

    AI

  3. Claude Opus 5: Near-Frontier Intelligence at Half the Price, with Judgment as the Real Edge

    Anthropic's Claude Opus 5 offers near-Fable 5 performance at half cost, excelling in coding, knowledge work, and self-verification. Explore benchmarks, real-world use cases, and…

    AI