Hoppa till innehåll
VibekollenBETAVibekollen

Källa

Ollama

69 referat i flödet. Sammanfattningarna är skrivna av Vibekollen — innehållet tillhör Ollama.

ReleaseOllama

v0.40.0

What's Changed Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX. ollama pull qwen3.8 ollama run qwen3.8 Additional models include gemma4, qwen3.6 and qwen3.5 Decision models are now available on MLX as well: Nimble tev1 clef clef-flash We will continue testing and enabling additional models. Full Changelog: v0.34.4...v0.40.0-rc2

för 9 tim sedan github.com

ReleaseOllama

v0.35.0

What's Changed app: sync macOS update menu and icon at startup by @hoyyeva in #18622 app: defer Settings model discovery by @ParthSareen in #18598 app: isolate cloud-setting tests from Windows user config by @drifkin in #18626 api: deprecate typical_p by @dhiltgen in #18627 bench: use HumanEval patch prompts by @dhiltgen in #17480 feat: add System One scoring API by @ParthSareen in #18606 mlx: bound pull stall retries and let the watchdog interrupt them by @dhiltgen in #18625 Full Changelog: v0.34.4...v0.35.0-rc1

28 sep. github.com

ReleaseOllama

v0.34.4

What's Changed server: fix intermittent "model not found" errors. by @rick-github in #18438 server: apply structured outputs in a single pass on thinking models by @jessegross in #18479 app: avoid System Events for ChatGPT/Codex detection by @hoyyeva in #18601 llama.cpp: version update by @dhiltgen in #18577 MLX: version bump by @dhiltgen in #18576 mlx: select Gemma 4 image resolution dynamically by @dhiltgen in #18603 mlx: speed up Qwen 3.8 prompt processing by @dhiltgen in #18550 Full Changelog: v0.34.3...v0.34.4-rc0

24 sep. github.com

ReleaseOllama

v0.34.4

What's Changed server: fix intermittent "model not found" errors. by @rick-github in #18438 server: apply structured outputs in a single pass on thinking models by @jessegross in #18479 app: avoid System Events for ChatGPT/Codex detection by @hoyyeva in #18601 llama.cpp: version update by @dhiltgen in #18577 MLX: version bump by @dhiltgen in #18576 mlx: select Gemma 4 image resolution dynamically by @dhiltgen in #18603 mlx: speed up Qwen 3.8 prompt processing by @dhiltgen in #18550 Full Changelog: v0.34.3...v0.34.4-rc0

23 sep. github.com

ReleaseOllama

v0.34.1

What's Changed app: fix ChatGPT model selector spacing mlxrunner: Evict prefix cache snapshots from the active conversation mlxrunner: check system free memory and wait for evicted runners before loading the next MLX model llm: raise token repeat limit to 100 and return error instead of incomplete result mlx: scope array lifetimes instead of pinning and sweeping llm: keep gemma3n projector off the CPU app: refresh Apps layout and command copy feedback MLX and llama.cpp updates Full Changelog: v0.34.0...v0.34.1-rc1

15 sep. github.com

ReleaseOllama

v0.34.1

What's Changed app: fix ChatGPT model selector spacing mlxrunner: Evict prefix cache snapshots from the active conversation mlxrunner: check system free memory and wait for evicted runners before loading the next MLX model llm: raise token repeat limit to 100 and return error instead of incomplete result mlx: scope array lifetimes instead of pinning and sweeping llm: keep gemma3n projector off the CPU app: refresh Apps layout and command copy feedback MLX and llama.cpp updates Full Changelog: v0.34.0...v0.34.1-rc1

15 sep. github.com

ReleaseOllama

v0.33.3: gemma4: image and audio input support

Safetensors gemma4 imports served by the MLX engine now answer image and audio chats. Images run through both vision architectures: the transformer tower (26B, 31B, e-series) and the 12B's encoder-free unified embedder. Audio arrives through the same intake the ollama API already accepts for gemma4 GGUFs — WAV bytes in the images field, OpenAI input_audio parts, and /v1/audio/transcriptions uploads — with the e2b/e4b checkpoints running clips through their conformer audio encoder and the 12b unified checkpoint embedding the raw waveform directly. Clips longer than 30 seconds are split evenly into chunks of at most 30 seconds, cut at pauses, and encoded independently. Each modality serves only checkpoints that carry it: 26B/31B have no audio config and reject audio input, and checkpoints with an unrecognized vision architecture still load as text-only models and reject image requests. The server previously hid the vision and audio capabilities for gemma4 safetensors because the engine served neither. Both suppressions are removed, and existing imports start advertising the capabilities without re-importing since import already records them.

3 sep. github.com

ReleaseOllama

v0.33.0

Ollama v0.33.0 förbättrar integreringen med Claude Desktop genom att låta dig växla mellan lokala Ollama-modeller direkt från menyraden. Det finns en ny Apps-vy för att hantera integrationer, och flera tekniska fel åtgärdas — bland annat problem med caching där avbrutna beräkningar nu sparar sin framsteg istället för att börja om från början. Versionen fixar också installationsproblem på Linux och Windows, och förbättrar användarens första gång med programmet.

26 aug. github.com

ReleaseOllama

v0.33.0

Ollama v0.33.0 är en ny version av verktyget som låter dig köra AI-modeller lokalt på din dator. Uppdateringen innehåller flera förbättringar: fixes för att mlx (en snabb ML-motor) fungerar rätt på Linux och Windows, integrering med Claude Desktop-appen, förbättrad användarupplevelse i onboarding-flödet, och nya funktioner för att koppla samman olika appar. Versionen är ännu en release candidate (testversion innan officiell släpp).

22 aug. github.com

ReleaseOllama

v0.32.13: qwen3.8: support developer instructions (#17749)

qwen3.8: support developer instructions Qwen3.8 does not define a developer role, while OpenAI-compatible coding agents commonly send developer instructions before user messages. Fold the leading system/developer instruction prefix into a single system turn before Qwen3.8 validation, preserving instruction precedence without changing Qwen3.5 or other renderer behavior. Add streaming tool-call integration coverage for the native Ollama, OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages request shapes. Each case exercises prior assistant tool calls, tool results, follow-up rendering, and parsed tool-call output. Add Qwen3.8 to the release tools sweep. Removes an unnecessary unit test that should not have been included in the original 3.8 PR. review comments

14 aug. github.com

ReleaseOllama

v0.32.12: qwen3.8: add renderer and MLX import support

Qwen3.8 keeps the Qwen3.5 model architecture and parser, but its chat template adds reasoning-effort and preserved-thinking semantics. Detect those template markers during safetensors import, select the qwen3.8 renderer, and cover thinking, tools, continuation, and malformed parser input. Make indexed safetensors imports use the weight map's shard names instead of independently filtering files by the model-* convention. Reject unsafe shard paths, ignore unindexed tensors, and fail when an indexed weight is missing or stored in a different shard. Retain the conservative model-* scan when no index is present. Treat Classification.Quantize as the effective tensor format and pass it to the manifest writer. This records file_type for automatic block-FP8-to-MXFP8 conversion and recognized prequantized inputs, preserves requested quantization and base-plus-draft behavior, and avoids claiming one type for mixed or unknown formats. Normalize both supported convolution weight layouts with an explicit reshape.

14 aug. github.com

ReleaseOllama

v0.32.10

Ollama v0.32.10 är en uppdatering av det verktyg som låter dig köra AI-modeller lokalt på din dator. Ändringar inkluderar att modeller nu per standard är inställda på att inte upprepa sig själva (repeat_penalty ändras från 1.1 till 1.0), vilket gör allt snabbare. Det finns också en hastighetshöjning på 7–8 procent för vissa GPU-accelererade modeller som Qwen3.6 och Muse Glimmer. En bugg reparerades där verifieringen av modellfiler ibland hoppasss över.

13 aug. github.com