Provisionally Supported

Ollama v0.32.1 Improves Gemma 4 Tool Calling and Fixes MLX Cache Leak

Published: 16 July 2026 Last checked: 17 July 2026 Source: primary

Summary

Ollama released version 0.32.1 with improvements to Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations. The update also fixes a recurrent MLX model cache leak that could increase memory use across requests and improves cache snapshot performance.

Why this matters

This patch resolves a memory leak in MLX caching that could affect long-running inference sessions and improves tool-calling reliability for Gemma 4 users.

provisionally_supported

Available evidence supports this claim, but verification is incomplete.

Why this rating?
Source class primary

Primary source — direct from official documentation, research paper, or specification.

Source tier Not assessed

Claim-level source tier has not yet been determined from reviewed evidence records.

Corroboration Not assessed

Independent corroboration has not yet been determined from reviewed claim assertions.

Independent verification Not assessed

Independent verification has not yet been determined from reviewed claim assertions.

Conflict of interest Low risk

No obvious commercial conflict of interest identified.

Timeliness 40 days ago

Last checked 40 days ago — information may be outdated.

Reproducibility Not assessed

Reproducibility has not yet been determined from reviewed claim assertions.

Sources

Claim-level evidence

No claim-level evidence has been publicly resolved for this story yet. The source links above are references, not a claim-level corroboration count.