All Posts
Have something to share? Submit a post
The archive is organised by week. Each week has its own page carrying that week's summary and every post published in it.
-
Local AI, 28 Sept – 4 Oct 2026
35 local AI stories, 28 Sept – 4 Oct 2026. Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model. New in Llama.cpp: Decision Models.
-
Local AI, 21 Sept – 27 Sept 2026
30 local AI stories, 21 Sept – 27 Sept 2026. Llama.cpp Fork Achieves 2-4x MultiGPU Speedup for MoE Models Larger Than VRAM.
-
Local AI, 14 Sept – 20 Sept 2026
34 local AI stories, 14 Sept – 20 Sept 2026. Apple's A20 Pro Chip Doubles Speed for 27B Parameter Models vs A19 Pro.
-
Local AI, 7 Sept – 13 Sept 2026
32 local AI stories, 7 Sept – 13 Sept 2026. Run 744B MoE Models on a Laptop With Disk Streaming, No GPU Needed.
-
Local AI, 31 Aug – 6 Sept 2026
37 local AI stories, 31 Aug – 6 Sept 2026. GGUF Quantization: Shrink LLMs 72% in 12 Steps. llama.cpp 0.4.0: Qwen3.8-Flash-Next and On-Demand Tensor Reading.
-
Local AI, 24 Aug – 30 Aug 2026
32 local AI stories, 24 Aug – 30 Aug 2026. Efficient Decode Context Parallelism with vLLM for Long Context Workloads.
-
Local AI, 17 Aug – 23 Aug 2026
33 local AI stories, 17 Aug – 23 Aug 2026. Google COSMO Leak Reveals Gemini Nano and On-Device AI Skills. llama.cpp Adds CUDA Pool Operations Support.
-
Local AI, 10 Aug – 16 Aug 2026
44 local AI stories, 10 Aug – 16 Aug 2026. HackerNoon Compares 7 Best Self-Hosted Inference Servers for Open-Source Models.
-
Local AI, 3 Aug – 9 Aug 2026
59 local AI stories, 3 Aug – 9 Aug 2026. Chrome and Edge Now Require 20GB Free Space for AI Models. DeepSeek V4 Flash Achieves 82.7% on Terminal-Bench 2.1.
-
Local AI, 27 Jul – 2 Aug 2026
61 local AI stories, 27 Jul – 2 Aug 2026. AI Efficiency Layer Cuts Energy Use and Expands Server Capacity on Existing Hardware.
-
Local AI, 20 Jul – 26 Jul 2026
59 local AI stories, 20 Jul – 26 Jul 2026. AMD Advancing AI 2026: Enterprise AI Architecture Basics for Startup Founders.
-
Local AI, 13 Jul – 19 Jul 2026
64 local AI stories, 13 Jul – 19 Jul 2026. AI Coding Agents Should Optimize for Less Owned Code. AI Inference Costs: Build vs. Rent.
-
Local AI, 6 Jul – 12 Jul 2026
63 local AI stories, 6 Jul – 12 Jul 2026. CEO Calls for Lower AI Pricing to Enable Practical Labor Automation Deployment.
-
Local AI, 29 Jun – 5 Jul 2026
58 local AI stories, 29 Jun – 5 Jul 2026. If You Can Write Acceptance Criteria, You Can Write an AI Routing Policy.
-
Local AI, 22 Jun – 28 Jun 2026
55 local AI stories, 22 Jun – 28 Jun 2026. GEEKOM A9 Max Delivers 32GB RAM and Native LLM Support in Compact Form Factor.
-
Local AI, 15 Jun – 21 Jun 2026
60 local AI stories, 15 Jun – 21 Jun 2026. Agentic Systems Course: Learn to Build AI Agents with Live AI Coding.
-
Local AI, 8 Jun – 14 Jun 2026
64 local AI stories, 8 Jun – 14 Jun 2026. It Is Beginning: AI Improves Itself. Ask HN: What Problem Did AI Create at Your Company That Didn't Exist Before?.
-
Local AI, 1 Jun – 7 Jun 2026
65 local AI stories, 1 Jun – 7 Jun 2026. Community Survey: AI Coding Tools Usage Patterns and Local Deployment Preferences.
-
Local AI, 25 May – 31 May 2026
59 local AI stories, 25 May – 31 May 2026. Why Chinese AI Labs Went Open and Will Remain Open. Chrome Quietly Downloads 4GB AI Model Without User Permission.
-
Local AI, 18 May – 24 May 2026
60 local AI stories, 18 May – 24 May 2026. Redditor Successfully Runs 1 Trillion Parameter LLM Using Cheap Intel Optane DIMMs.
-
Local AI, 11 May – 17 May 2026
70 local AI stories, 11 May – 17 May 2026. A Cheap Fix That Saves the AI $400M Dollars a Year and Brings 4B People Online. A Lo-Fi Rebellion Against A.I.
-
Local AI, 4 May – 10 May 2026
69 local AI stories, 4 May – 10 May 2026. DistillFast: AI Cost Optimization Tool for Model Efficiency.
-
Local AI, 27 Apr – 3 May 2026
65 local AI stories, 27 Apr – 3 May 2026. How to Test AI Agents When They Never Give the Same Answer Twice.
-
Local AI, 20 Apr – 26 Apr 2026
65 local AI stories, 20 Apr – 26 Apr 2026. Blueprint: AI Hardware Design. Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop.
-
Local AI, 13 Apr – 19 Apr 2026
85 local AI stories, 13 Apr – 19 Apr 2026. Gemma 4 Just Replaced My Whole Local LLM Stack.
-
Local AI, 6 Apr – 12 Apr 2026
95 local AI stories, 6 Apr – 12 Apr 2026. I Gave My AI Shell Access and Felt Uneasy – So I Sandboxed It.
-
Local AI, 30 Mar – 5 Apr 2026
90 local AI stories, 30 Mar – 5 Apr 2026. Apple Research Shows Self-Distillation Significantly Improves Local Code Generation.
-
Local AI, 23 Mar – 29 Mar 2026
Major stories this week include the release of Qwen 3.5 models and the announcement of Alibaba's commitment to continuous open-sourcing of Qwen and Wan models, as well as the demonstration of a 400B-parameter language...
-
Local AI, 16 Mar – 22 Mar 2026
Major stories this week include AMD's declaration that on-device AI inference has reached a critical point, and Apple's on-device AI raising privacy concerns in the British Parliament. Other notable developments include...
-
Local AI, 9 Mar – 15 Mar 2026
Nemotron 9B and Qwen 3.5 models were highlighted for large-scale local inference. Nota AI showcased on-device AI optimization. Posts like "Fine-Tuned Qwen SLMs" and "Qwen 3.5 Ultra-Compact Models" stood out for local AI...
-
Local AI, 2 Mar – 8 Mar 2026
Alibaba's CoPaw AI agent and AMD's Ryzen AI 400 series were major stories, with Apple's Neural Engine also being reverse-engineered for local model training. Don't miss "Qwen 3.5 27B Achieves 100+ Tokens/s Decode" and...
-
Local AI, 23 Feb – 1 Mar 2026
Major stories this week include the release of Elastic's best-in-class embedding models for high-performance semantic search and the achievement of GLM-5 as the top open-weights model on the Extended NYT Connections...
-
Local AI, 16 Feb – 22 Feb 2026
Alibaba unveiled a major AI model upgrade ahead of DeepSeek's release, and Cohere released Tiny Aya, a 3.3B parameter multilingual model. Standout posts include "I broke into my own AI system in 10 minutes" and "I...
-
Local AI, 9 Feb – 15 Feb 2026
Big stories include GLM-5, a 744B parameter MoE model, and MiniMax M2.5, a 230B parameter model. Don't miss "Community Member Builds 144GB VRAM Local LLM Powerhouse" and "NVIDIA's Dynamic Memory Sparsification" for...