what stackpulse tracks
Ollama releases from GitHub
StackPulse watches Ollama release notes and keeps the original source link close to every summary.
Get up and running with large language models locally StackPulse turns upstream changelogs into scannable summaries with risky changes, deprecations, migration notes, and source links.
what stackpulse tracks
StackPulse watches Ollama release notes and keeps the original source link close to every summary.
upgrade risk
Risky changes are separated from normal feature notes so you can scan upgrade impact before changing production dependencies.
migration notes
Migration steps and recommended actions are only shown when the upstream release notes support them.
This release changes the default repeat_penalty to 1.0 (off) for models that don't set it, matching other engines and speeding up speculative decoding. It also improves prefill performance on NVFP4 MLX models by 7-8% for specific models.
Users with models that relied on the default repeat_penalty of 1.1 may see different behavior and should set the parameter explicitly if needed.
Update models to explicitly set repeat_penalty if they relied on the previous default behavior.
This release introduces support for NVIDIA Nemotron 3.5 Lightning, a 30B mixture-of-experts model, and includes fixes for the Muse Glimmer function calling parser.
Users interested in NVIDIA Nemotron 3.5 Lightning model or working with Muse Glimmer function calling parser are affected.
Users can try the new Nemotron 3.5 Lightning model by running 'ollama run nemotron-3.5-lightning'.
This release introduces Muse Glimmer, a powerful model for coding agents and personal assistants, now available across all platforms. It leverages Ollama's MLX engine for optimized performance on Apple Silicon.
Developers and users leveraging coding agents or personal assistants will benefit from the expanded platform support and optimized performance.
Download and run Muse Glimmer locally using the provided commands to integrate it with your preferred coding agent or personal assistant.
This release introduces initial support for Muse Glimmer, Meta's newest open model, via Ollama's MLX engine on Apple Silicon. It enables local execution of agent workloads and personal assistants.
Developers and users leveraging Ollama for local agent workloads and personal assistants on Apple Silicon are primarily affected.
Download and run Muse Glimmer locally using the provided commands.
This release improves performance on Apple GPUs, enhances OpenAI API compatibility, and introduces cloud-only model support. It temporarily removes experimental image generation.
Users relying on experimental image generation are affected and should stay on version 0.32.5.
Users needing image generation should remain on version 0.32.5; others can upgrade for performance and compatibility improvements.
This release improves performance for Qwen3.5 on Apple GPUs and enhances OpenAI API compatibility, while temporarily removing experimental image generation support.
Users relying on experimental image generation will need to stay on v0.32.5.
Continue using 0.32.5 if image generation support is required.
This release primarily includes an update to the MLX component, as part of ongoing maintenance and improvements.
Users relying on the MLX component may need to test compatibility with this update.
Test the release candidate in your environment to ensure compatibility.
This release addresses a bug in MLX Metal that could reduce output quality for NVFP4 models, particularly Laguna.
Users of NVFP4 models, especially Laguna, may have experienced reduced output quality prior to this fix.
Update to v0.32.5 to resolve the MLX Metal bug affecting NVFP4 models.
This release focuses on quantization improvements, bug fixes, and enhancements to model handling and testing. Notable changes include fixes for data races, quantization handling, and memory management for loaded models.
Developers and users working with quantization and model handling in Ollama are primarily affected.
This release introduces support for Laguna on Apple GPUs via the MLX engine and improves quantization for speculative-decoding drafts. Additionally, it fixes Qwen3 MoE decoding issues and enhances performance on M5 Max devices.
Users leveraging Apple GPUs or speculative-decoding drafts will benefit from the new features and performance improvements.
Upgrade to v0.32.4 to take advantage of the new GPU support and performance enhancements.
This release focuses on bug fixes, improved integrations, and expanded GPU support. Key changes include fixes for model downloads, improved support for various models, and updates to MLX and llama.cpp engines.
Users with GPU setups and those using specific models like Laguna 2.1 will benefit most from these changes.
Update to this version to benefit from improved GPU support and model compatibility fixes.
This release focuses on agent improvements, UX/DX cleanup, and updates to dependencies like llama.cpp and MLX. It also includes changes to Claude Code channels and Hermes integration.
Users relying on the standalone agent command or working with cloud models and Claude Code channels will be affected.
Update to this version to benefit from agent improvements and ensure compatibility with the latest changes.
This release focuses on improving the agent's functionality, cleaning up code semantics, and updating dependencies like llama.cpp and MLX. It also introduces a skills system for the agent and enhances the Claude Code channels.
Users relying on the standalone agent command will need to adjust their workflows.
Update to the latest version to benefit from new features and improvements.
This release improves Gemma 4 tool calling and multi-turn reasoning, fixes a memory leak in MLX model cache, and enhances agent functionality.
Users of Gemma 4 and MLX models may see improved performance and reliability.
Introduces a new interactive agent experience and simplifies integration selection while adding deprecation warnings for older models.
Users of older models will see deprecation warnings, and those using the Codex App integration will need to switch to ChatGPT.
Update integrations to use ChatGPT instead of Codex App and consider upgrading from deprecated models.
This release introduces a new agent UI and adds support for the qwen3.5 parser and renderer for Qwen3.5/Next models. It also includes warnings for old agent models.
Users working with Qwen3.5/Next models or using agent functionality may be affected.
Update to access new features and ensure compatibility with Qwen3.5/Next models.
This pre-release includes improvements to CUDA toolkit lookup, ROCm device support cleanup, and UTF-8-safe file handling. It also removes deprecated features and updates dependencies.
Users relying on removed ROCm devices or the client2 experimental feature will be affected.
Update to this version if encountering issues with CUDA or ROCm devices, and check for removed experimental features.
This release includes improvements to CUDA toolkit lookup, GPU support, and code cleanup. It also updates documentation and removes no longer supported devices.
Users with CUDA CC 6.x GPUs will benefit from new features, while those using deprecated configurations will need to update.
Update GPU drivers if using CUDA CC 6.x GPUs and check compatibility for ROCm devices.
This update expands GPU compatibility with Flash Attention support for older NVIDIA GPUs and improved iGPU handling for vision models. It also includes model loading fixes and engine updates.
Users with older NVIDIA GPUs or integrated GPUs running vision models benefit most from these changes.
Update to take advantage of improved GPU support and model handling.
This release includes fixes for CUDA toolkit compatibility, updates to ROCM support, and general cleanup of deprecated code. It also removes unsupported devices and improves UTF-8 file handling.
Users relying on the removed ROCM devices may need to update their configurations.
Update your configurations and check compatibility with ROCM devices.