Local AI, 23 Feb – 1 Mar 2026

120 posts · All weeks

Major stories this week include the release of Elastic's best-in-class embedding models for high-performance semantic search and the achievement of GLM-5 as the top open-weights model on the Extended NYT Connections benchmark with an 81.8 score. Additionally, Qwen3.5-35B-A3B emerged as a highly efficient model for local deployment.

Standout posts include "Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference", which explores practical techniques for maximizing local LLM inference speed, and "The Complete Developer's Guide to Running LLMs Locally: From Ollama to Production", a comprehensive resource for deploying LLMs locally.

Sunday, 1 March 2026

AgentLens provides open-source observability tools for local LLM agent deployments.

Saturday, 28 February 2026

Qwen3.5-35B runs on Raspberry Pi 5 at 3+ tokens/second with effective prompt engineering.

Friday, 27 February 2026

Qualcomm's Snapdragon 8 Elite Gen 5 enhances on-device AI inference on Samsung Galaxy S26 series.

Thursday, 26 February 2026

Qwen3.5 122B achieves 25 tokens/second on a 72GB VRAM setup with three 3090s.

Wednesday, 25 February 2026

Mirai secures $10M to optimize on-device AI performance with Qwen3.5 models.

Tuesday, 24 February 2026

Anthropic reveals distillation attacks on Claude models by DeepSeek and Moonshot AI labs.

Monday, 23 February 2026

GLM-5 surpasses Kimi K2.5 Thinking on the Extended NYT Connections benchmark.