Ollama 0.32.15 adds model metadata caching
The release adds a model metadata cache designed to reduce per-request overhead for local model serving.
Verify official source →Run open models locally with a simple developer-first runtime.
A free local AI runtime for downloading, running, and integrating language and multimodal models on your own hardware. It is designed for developers and privacy-conscious users who want a simple local model workflow rather than a hosted chatbot service.
Opens official website
Live change feed
Shared with the homepage and tool pages
The release adds a model metadata cache designed to reduce per-request overhead for local model serving.
Verify official source →GitHub is rolling out persistent memory, access to local models through Ollama, additional enterprise controls, improved chat workflows, and MCP reliability fixes in its JetBrains integration.
Verify official source →Ollama 0.32.5 fixes an MLX Metal issue that could reduce output quality for NVFP4 models, with the release notes specifically highlighting Laguna.
Verify official source →Ollama 0.32.4 adds Laguna support through the MLX engine for Apple GPUs and includes fixes for quantized-model decoding and speculative decoding.
Verify official source →A free local AI runtime for downloading, running, and integrating language and multimodal models on your own hardware. It is designed for developers and privacy-conscious users who want a simple local model workflow rather than a hosted chatbot service.
Compare before you choose
Use these comparisons and guides to decide whether Ollama fits your workflow, budget, and output needs.
Compare LM Studio and Ollama for local model downloads, privacy, desktop use, APIs, CLI workflows, MCP, hardware, and developer automation.
guideUse the AI Linkbase decision map to choose cost controls, a fallback, automation, or a local AI route.