Best Local LLM for Multilingual Use (October 2026)

Which open-weight model actually handles non-English work, what each provider claims about language coverage, and the tradeoffs between breadth and depth.

Updated October 8, 2026

The four billion people who speak languages with weak model support get worse answers from hosted AI than English speakers do. Local inference is one way to even that out, but only if the model you pick actually handles your language well - and the headline language counts on Hugging Face cards are not a quality guarantee.

The practical choice is between four open-weight families available through Ollama today. They differ on licence, gating, the size they ship in, and the quality they deliver outside English. The table below is built from what each model’s card and the Hugging Face API actually say, not from the language count claim alone.

What the language count on a card actually means

A model card often opens with a large number (“201 languages”, “140+ languages”, “24 languages”). That number tells you how many languages the training set at least touched, not how well the model handles any one of them. A model that lists 201 languages will still type-mix or romanise a low-resource language it rarely saw during pretraining.

Two practical signals to check before trusting a number:

  • Tokenisation fragmentation. SentencePiece or BPE vocabularies built primarily on English split a Hindi word into four sub-tokens where a model trained on Hindi might split it into one. The more sub-tokens per word, the more context a non-English message eats per real word, and the harder the model finds the next-token prediction. Phi-4-mini’s card lists 23 named European and Asian languages with full tokeniser support; Qwen3.5’s card explicitly trains on 201.
  • Benchmark depth. Qwen3.5-9B publishes multilingual-specific numbers: MMMLU 81.2 across many languages, WMT24++ 72.6 over 55 languages via XCOMET-XXL, PolyMATH 57.3. Phi-4-mini publishes Multilingual MMLU 49.3 and MGSM 63.9 over its 23 named languages. Gemma 3 4B’s card claims broad training coverage but the published benchmarks above are vision-language rather than text-only multilingual.

The four models worth comparing

All four are available through ollama pull or docker model pull, and the licence and gating flags below come from the Hugging Face API (cardData.license, gated).

ModelParamsLicenceGated?Ollama tag (default q4_K_M)Native contextLanguages claimedMultilingual benchmark
Qwen3.5-9B9BApache-2.0Noqwen3.5:9b (6.6 GB)262K (extensible to 1,010K via YaRN)“201 languages and dialects”MMMLU 81.2, WMT24++ 72.6 (55 langs, XCOMET-XXL), PolyMATH 57.3
Mistral Small 4 119B A6B119B total / 6.5B activatedApache-2.0Nomistral-small:latest (14 GB)256K (--max-model-len 262144)24 languagesAA LCR 0.72 at 1.6K chars vs Qwen 5.8-6.1K
Gemma 3 4B IT4Bgemma (custom)Manual - must accept Google’s terms at ai.google.dev/gemma/termsgemma3:4b (3.3 GB)128K in / 8K out”over 140 languages”Card publishes vision-language rather than text-only multilingual scores
Phi-4-mini-instruct3.8BMITNophi4-mini:3.8b (2.5 GB)128K23 named languages (Arabic, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, Hebrew, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish, Swedish, Thai, Turkish, Ukrainian)Multilingual MMLU 49.3, MGSM 63.9

The licence matters here. Apache-2.0 and MIT both allow commercial use with attribution and a patent grant. Gemma’s custom licence requires a separate acceptance on the ai.google.dev terms page before Hugging Face gated access unlocks (gated: "manual" in the API), and it adds a prohibition list on use cases that is narrower than the OSI-approved alternatives. Phi-4’s MIT grant is the cleanest of the four.

Pick by task

Cross-language chat and translation, with one machine. Qwen3.5-9B on a 12 GB GPU is the broadest fit. Its 9B parameters are small enough to run at q4_K_M (6.6 GB) on consumer hardware, and the WMT24++ 72.6 result over 55 languages is the strongest published multilingual signal of the four models. The 262K native context can hold a long bilingual transcript in a single prompt. Internal link: see best-local-llm-runners for the Ollama vs llama.cpp vs LM Studio question, and local-ai-power-and-thermals for what a 9B model at q4_K_M draws from the wall.

Production summarisation where output length is metered. Mistral Small 4 119B A6B is a mixture-of-experts model - 119B total, 6.5B active per token - that publishes an AA LCR (Length-Controlled Reasoning score of 0.72 at 1.6K characters versus 5.8-6.1K for the comparable Qwen models), per the Mistral model card. At 14 GB on the Ollama mistral-small:latest tag it needs ~16 GB of system RAM to stay resident, and it is a real architecture step up from Qwen3.5-9B on English reasoning. Internal link: pair it with moe-vs-dense-local-ai for the parameter math behind “active per token.”

Region-locked low-resource languages with a 4-8 GB GPU. Gemma 3 4B IT. The card claims “over 140 languages” in training data and ships a vision-language model that handles text and image input. The catch is the licence: gated: "manual" on Hugging Face means a one-time click-through to Google’s terms at ai.google.dev/gemma/terms, and the published benchmarks above the language list are vision rather than pure-text multilingual. If your text workload is image-heavy (signs, scanned forms, multilingual captions) Gemma 3 has no rival at 3.3 GB. If it is text-only and you would rather skip the gating step, Qwen3.5-9B is the next step down.

Edge hardware and MIT-licensed deployment. Phi-4-mini-instruct. The 23 named languages map closely to Western European plus Japanese, Korean, Russian, Arabic, Hebrew and Thai - not as broad as Qwen3.5 or Gemma, but enough for most European and East Asian business workflows. At 2.5 GB the q4_K_M build fits a Raspberry Pi 5 with 8 GB of RAM (see run-llm-locally-on-raspberry-pi) and the MIT licence means there is no gating, no separate terms page, and no use-case prohibition list to read. MGSM 63.9 (multilingual grade-school math) is the strongest published signal of the three “real-world European” candidates if your workload leans on reasoning in non-English prompts.

Practical setup and the catches

There are three recurring pitfalls when running these models for non-English work.

  • System prompt language sets the response language. A model that handles both English and Japanese will still drift to whichever language the system prompt or the first user turn uses. For bilingual deployments, set the system prompt in the target language and front-load an example turn in the same language; do not rely on a single mid-prompt switch.
  • Quantisation quality varies by language. A Q4_K_M build of a 9B model loses about 1 to 2 MMLU points in English versus BF16, and the gap is wider in less-trained languages. If your workload is a single European language plus English, Q4_K_M is fine. If it includes low-resource languages with light training data, drop to a higher quant (Q5_K_M, Q8_0) or run BF16 if you have the VRAM.
  • Tokeniser cost adds up at long context. A Hindi sentence that takes 80 sub-tokens under a Western-trained tokenizer will eat far more KV cache than the equivalent English. For very long-document Q&A over non-English material, best-local-models-by-context-window explains what each family’s KV-cache math actually looks like.

Bottom line

  • Need one model that handles the most languages on consumer hardware: Qwen3.5-9B at qwen3.5:9b (6.6 GB, Apache-2.0, no gating, 262K context, multilingual benchmarks over 55 languages).
  • Need the highest published non-English quality and can spare the RAM: Mistral Small 4 119B A6B at mistral-small:latest (14 GB, Apache-2.0, 119B/6.5B MoE, 256K context).
  • Image-heavy multilingual workloads on small hardware, and you can accept the gating: Gemma 3 4B IT at gemma3:4b (3.3 GB, custom licence, 140+ languages, vision-language).
  • Edge deployment in 23 named Western/East Asian languages, MIT clean: Phi-4-mini-instruct at phi4-mini:3.8b (2.5 GB, MIT, no gating, 128K context).
  • Avoid choosing on the headline language count alone. The published multilingual benchmarks (MMMLU, WMT24++, MGSM, PolyMATH, AA LCR) and the specific named-language list tell you more about whether your real workload will work.