Ollama 0.32.15 adds model metadata caching
The release adds a model metadata cache designed to reduce per-request overhead for local model serving.
Verify official source →Open models locally and in the cloud through one runtime and API
Open-model runtime for local and cloud inference with CLI, API, desktop apps and individual, team and enterprise plans.
Opens official website
Live change feed
Shared with the homepage and tool pages
The release adds a model metadata cache designed to reduce per-request overhead for local model serving.
Verify official source →GitHub is rolling out persistent memory, access to local models through Ollama, additional enterprise controls, improved chat workflows, and MCP reliability fixes in its JetBrains integration.
Verify official source →Ollama 0.32.5 fixes an MLX Metal issue that could reduce output quality for NVFP4 models, with the release notes specifically highlighting Laguna.
Verify official source →Ollama 0.32.4 adds Laguna support through the MLX engine for Apple GPUs and includes fixes for quantized-model decoding and speculative decoding.
Verify official source →Ollama runs open models locally and through optional Ollama Cloud. Local models remain unlimited on a user's own hardware. Free includes starter cloud credits; Pro is $20 monthly or $200 annually with $60 monthly usage credits, larger pro models and multiple concurrent requests. Max is $100 monthly with $300 monthly usage credits, early access and ten concurrent requests. Team is Early Access at $500 monthly for unlimited users with $1,000 shared monthly credits. Enterprise is custom.
Run models locally, receive starter cloud usage credits and add credits to unlock all cloud models. No service fee.
$60 of monthly usage credits, access to larger pro models and multiple concurrent cloud requests.
Billed annually; includes the Pro cloud capabilities and monthly usage credits.
$300 of monthly usage credits, early access to newest models and ten concurrent requests.
Unlimited users, $1,000 shared monthly usage credits, centralized administration and priority support.
Volume pricing, custom security review, support and cost controls.
Prices are public list prices before tax and may vary by country, billing term, and workspace. Last verified 2026-09-01. Official pricing source →
💰 Pricing Breakdown
Official pricing verified 2026-09-01. Free $0; Pro $20/month or $200/year with $60 monthly usage credits; Max $100/month with $300 monthly usage credits; Team Early Access $500/month for unlimited users with $1,000 shared monthly credits; Enterprise custom. Local models on your hardware are unlimited. Extra usage can be added and model token rates vary.
View current pricing →Compare before you choose
Use these comparisons and guides to decide whether Ollama fits your workflow, budget, and output needs.
Compare LM Studio and Ollama for local model downloads, privacy, desktop use, APIs, CLI workflows, MCP, hardware, and developer automation.
guideUse the AI Linkbase decision map to choose cost controls, a fallback, automation, or a local AI route.
Tabnine
Private code assistance and agents built for governed engineering organizations
View Tool →Blackbox AI
Coding agents and multi-model inference across IDE, CLI and API
View Tool →Codeium
Free code completion across 70+ languages
View Tool →LiteLLM
Put models, agents, and MCP servers behind one governed AI gateway.
View Tool →