Local AI, 28 Sept – 4 Oct 2026

35 posts · All weeks

Sunday, 4 October 2026

Aleph Alpha’s Kolibri 78.1B MoE model launches, delivering multilingual inference with just 3.46B active parameters.

Saturday, 3 October 2026

llama.cpp now runs Cloudflare’s Clef decision models, enabling local multimodal decision inference alongside text generation.

Friday, 2 October 2026

Cloudflare unveiled Clef, an open‑source decision‑model library with RL fine‑tuning for local deployment.

Thursday, 1 October 2026

Ollama v0.35.0 adds Decision Models via /v1/systemone API, while Reflex engine outpaces llama.cpp and vLLM in cold‑start latency.

Wednesday, 30 September 2026

AMD unveiled an agentic PC that runs 300‑billion‑parameter models on‑device, eliminating cloud reliance.

Tuesday, 29 September 2026

OPPO's ColorOS 17 now runs on‑device LLMs, bringing offline Xiaobou AI assistant to smartphones.

Monday, 28 September 2026

llama.cpp’s new speculative decoding speeds prompt lookup, while Perplexity runs 27B models on AMD Ryzen AI Max.