Local AI, 23 Mar – 29 Mar 2026

104 posts · All weeks

Major stories this week include the release of Qwen 3.5 models and the announcement of Alibaba's commitment to continuous open-sourcing of Qwen and Wan models, as well as the demonstration of a 400B-parameter language model running on an iPhone.

Standout posts include "Building a Production AI Receptionist" and "Powerful AI Search Engine Built on Single GeForce RTX 5090", which showcase practical applications of local LLM deployment.

Sunday, 29 March 2026

TurboQuant optimizes local LLM inference on Linux with OLED displays and Nvidia RTX 5070 graphics.

Saturday, 28 March 2026

CERN deploys custom AI models on silicon chips for Large Hadron Collider data filtering.

Friday, 27 March 2026

Mistral AI's Voxtral model outperforms ElevenLabs on local hardware.

Thursday, 26 March 2026

Google introduces TurboQuant for efficient local LLM deployment.

Wednesday, 25 March 2026

Llama.cpp benchmarks compare RTX 5090 performance against AMD AI395 in local inference scenarios.

Tuesday, 24 March 2026

FlashAttention-4 delivers 2.7x faster inference on NVIDIA B200 GPUs.

Monday, 23 March 2026

Alibaba open-sources Qwen and Wan models for local LLM deployment.