What's New in Ollama 0.40, and Should You Upgrade?

v0.40.0-rc0 ships MLX-by-default on Apple Silicon, transformers runs llama.cpp quants, v0.35.0 adds decision models. What changed and how to upgrade.

Updated October 2, 2026

Ollama shipped a version jump that confused a lot of readers in late September 2026: v0.34.4 was the last stable in the 0.34 line, then v0.40.0-rc0 appeared on 25 September, then v0.35.0 shipped as a stable on 28 September, and v0.35.1-rc2 followed a day later. The version number is real, the rc0 tag is still a pre-release, and the upgrade question depends on which channel you are on. This page tracks what each tag actually changed and how to move without breaking a working setup.

What the version numbers actually mean

The 0.34 → 0.40 jump is not a typo or a re-numbering. The GitHub releases page lists v0.34.4 (stable, 23 September 2026), then v0.40.0-rc0 (pre-release, 25 September), then v0.35.0 (stable, 28 September), then v0.35.1-rc2 (pre-release, 29 September). All four tags are real releases pulled from the Ollama GitHub repository; the 0.40 stream and the 0.35 stream are being prepared in parallel, with the 0.35 line on a faster bug-fix cadence and 0.40 reserved for the next behavior-changing cut.

Practically: most people should stay on the latest stable tag (v0.35.0 as of 2 October 2026). v0.40.0-rc0 is the tag to watch if you specifically want the Apple-Silicon MLX-by-default change before it ships as stable.

What v0.34.4 → v0.40.0-rc0 actually changed

The visible rc0 release body has one headline item: “Models run on MLX on Apple Silicon by default”. The Ollama release notes for v0.40.0-rc0 state that on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX, and that additional model support will be enabled during the pre-release window. This is the same direction the v0.34 line was moving in (v0.34.4 lists “Updated llama.cpp, MLX, and XGrammar” as one of its bullets, alongside “Qwen 3.8 prompt processing is faster on Apple Silicon” and “Gemma 4 on Apple Silicon now picks the best image resolution per image”), but the rc0 build makes MLX the default dispatch path rather than an opt-in. The trade-off: if a model is not on Ollama’s MLX-supported list yet, rc0 falls back to the llama.cpp engine on the same Apple Silicon hardware.

The commit comparison for v0.34.4 → v0.40.0-rc0 also surfaces a quieter change worth flagging: PR #18627 (the api: deprecate typical_p follow-up) switches the Ollama API from rejecting requests that contain typical_p to logging a warning. Model creation still rejects the parameter, but server-side inference now tolerates it. A separate commit in the same range (“llama-server: prepare to remove compatibility patch”) lays groundwork for manifest-list storage so the runner can coexist with multiple engine builds and lazily convert legacy GGUFs into llama.cpp-compatible children. None of these are user-visible in a default install today, but the deprecation is the kind of change that breaks third-party clients if you do not know it is coming.

v0.35.0 and the decision-models API

The v0.35.0 stable shipped on 28 September 2026 with a feature the rc0 stream does not have: support for decision models via /v1/systemone, based on TypeSafe’s Jev API. The release notes describe these as returning “choices, probabilities, and scores instead of text,” aimed at ticket triage, model routing, and content classification. Two models are listed as available at launch: Nimble from Bespoke Labs and Tev1 from Together AI, both pullable with ollama pull nimble (and ollama pull tev1 respectively, per the release body). Three question types are supported: choice, noul, and score.

The other v0.35.0 changes are quality-of-life: Settings opens without waiting for model discovery, the macOS update menu no longer lies about available updates at startup, stalled MLX model downloads no longer hang indefinitely, and API requests containing the deprecated typical_p parameter now log a warning instead of failing.

transformers now loads GGUF quants

On 22 September 2026, Hugging Face published a blog post titled “Transformers now runs llama.cpp quants” that is worth knowing about even if you never touch Ollama 0.40. The blog describes a packed-inference path in the transformers library that loads GGUF checkpoints via from_pretrained and runs them on Apple Silicon (MPS) by reusing ggml kernels through the kernels library. Three things matter for the Ollama reader:

  1. Architectures covered at launch are narrow. Qwen3.5 dense and MoE plus compatible Qwen3.8 checkpoints. Expansion is gradual.
  2. The packed inference path is MPS-only. Linux and Windows CUDA users get GGUF import via dequantization, not the packed path.
  3. Padding and batching are still rough. Unpadded inputs benefit from a mask optimization; padded batches do not. This is not a drop-in llama.cpp replacement for production serving.

The HF blog benchmarks on a MacBook Pro M2 Max with 32 GB unified memory under macOS 26.6 with PyTorch 2.12.1 and kernels 0.17.0, comparing three GGUF checkpoints across transformers and llama.cpp. It reports transformers is “close to llama.cpp across all three checkpoints” but notes the transformers measurement includes prefill while llama-bench reports decode-only throughput. The exact throughput numbers are not in the blog body (the chart is an image) and the install path is pinned to transformers main until the next release.

What this means in practice: if you run an Ollama workflow that wants Python introspection, fine-tuning from a quantized checkpoint (GgufConfig(dequantize=True)), or a single Python API surface for both HF and GGUF models, the new transformers path is the option to know about. For day-to-day Ollama serving, llama.cpp remains what the HF team itself recommends.

Upgrade checklist

A safe upgrade sequence as of 2 October 2026:

  1. Confirm the tag you are moving between. Run ollama --version and record it. If you are on v0.34.x, both v0.35.0 and v0.40.0-rc0 are valid destinations. If you are on v0.33.x or earlier, you are far enough behind that reading the v0.34 release notes first is worth the ten minutes.
  2. Back up your model directory. The default location is $OLLAMA_MODELS or ~/.ollama/models; the Ollama FAQ documents both. The manifest-list storage and lazy migration work is targeted at making future upgrades painless, but it does not yet eliminate the rare model-store repair.
  3. Pick your channel. The Ollama install page at ollama.com/download ships the current stable (v0.35.0 at last check). To get v0.40.0-rc0 specifically, pull the rc binary directly from the GitHub release assets for your platform (ollama-darwin.tgz, ollama-linux-amd64.tar.zst, or the platform-specific MLX, ROCm, and JetPack variants).
  4. Run ollama list before and after. If a model you depend on is missing from the output, the model file is still in the directory and you can re-link it rather than re-pulling.
  5. Decide whether you want MLX-by-default. v0.40.0-rc0 flips it on for Apple Silicon. If you have a specific llama.cpp-only model you are returning under Ollama today, retest that workload on rc0 before committing to it.
  6. Watch the typical_p warning. If a third-party client sends typical_p, v0.35.0 logs a warning instead of failing. v0.34.4 still rejects. If your client logs the warning at high volume, that is a signal to migrate off typical_p.
  7. Skip the version number anxiety. v0.35.0 is the current stable. v0.40.0-rc0 is the pre-release for the next behavior-changing cut. You do not have to pick v0.40 just because the number is higher.

Bottom line

Ollama 0.40 is real, but it is currently a release candidate, not a stable release, and the headline change (MLX-by-default on Apple Silicon) only matters if you are on Apple Silicon and want the new default dispatch path. v0.35.0 is the stable to install today, and its decision-models API via /v1/systemone is the more generally useful new feature. The transformers + GGUF news from Hugging Face is independent of Ollama but lands in the same week and is worth knowing if you want to fine-tune or introspect GGUF models in Python without leaving the transformers API.