Provisionally Supported
Ollama v0.32.1 Improves Gemma 4 Tool Calling and Fixes MLX Cache Leak
Summary
Ollama released version 0.32.1 with improvements to Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations. The update also fixes a recurrent MLX model cache leak that could increase memory use across requests and improves cache snapshot performance.
Why this matters
This patch resolves a memory leak in MLX caching that could affect long-running inference sessions and improves tool-calling reliability for Gemma 4 users.