Local AI, 7 Sept – 13 Sept 2026
Monday, 7 September 2026
Alibaba's Qwen3.8 Flash Next launches lightweight on-device inference, previewing the upcoming Qwen4 architecture.
-
Alibaba Releases Qwen3.8 Flash Next for Local Deployment
Alibaba's Qwen3.8 Flash Next provides a lightweight, optimized model for on-device inference with previews of the more capable Qwen4 architecture.
-
Apple's New Mac Mini and Studio Bet Big on On-Device AI
Apple positions its updated Mac Mini and Studio models as premium on-device AI platforms, signaling major hardware improvements for local LLM inference.
-
IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B
IFM has released the K2 Horizon series with six openly-licensed models spanning 0.9B to 375B parameters, providing diverse options for local deployment across different hardware constraints.
-
Ollama v0.34.0: ChatGPT Desktop Integration and Apple Silicon Improvements
Ollama releases v0.34.0 with ChatGPT Desktop integration, improved structured output performance on Apple Silicon, and enhanced model management features for local deployment.
-
Speculative Decoding in vLLM on AMD GPUs
vLLM now supports speculative decoding on AMD GPUs, enabling significant inference speed improvements for local LLM deployment on AMD hardware.