Tagged "llama-cpp"
-
Liquid AI Releases LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation
-
llama.cpp b10524 Makes MoE Expert Scatter Deterministic in OpenCL
-
GGUF Quantization Deep Dive: Q4_K_M vs IQ4_XS vs IQ4_NL Performance
-
AMD EPYC ZenDNN Accelerates llama.cpp Prompt Processing 4.5x
-
Native vLLM and ROCm 7.15 Support for AMD RDNA2 GPUs on Windows
-
What If Local LLM Inference Is Using Consumer Hardware Wrong?
-
GGUF Quantization Compared: Q4_K_M vs. IQ4_XS vs. IQ4_NL Performance Analysis
-
Llama.cpp Release b10485: GGML Sync with Platform-Specific Optimizations
-
DeepSeek V4 Flash Shrunk to 57GB for Local macOS Inference with Compiler Generation
-
Llama-macOS – Agentic and MCP Native macOS Front End for Llama.cpp
-
Unsloth Releases Qwen 3.8 27B GGUF Quantised Weights
-
HackerNoon Compares 7 Best Self-Hosted Inference Servers for Open-Source Models
-
Hugging Face State of Open Models: Summer 2026 Observations
-
7 Best Self-Hosted Inference Servers for Open-Source Models Compared (2026)
-
Running DeepSeek's 284B LLM on a Laptop: Quantisation and GGUF Optimization
-
llama.cpp Improves Muse Glimmer Tool Calling with Latest Update
-
llama.cpp Updates Tool Call Detection for Muse Glimmer
-
DEF CON 34 Exposes 10 Critical Vulnerabilities in Local AI Systems
-
llama.cpp Adds Tool Isolation Support via Docker
-
llama.cpp Improves CUDA Performance with Kernel Fusion
-
Llama.cpp Fixes Metal NORM Operations for Apple Silicon
-
Llama.cpp Adds LRU Scheduler for Multi-Model Serving
-
Llama.cpp B10327 Fixes CUDA Quantized Copy Kernel Performance
-
llama.cpp Build b10301: CUDA Optimization and Compiler Warning Fixes
-
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems
-
llama.cpp b10298: Multi-Token Multi-Dimension Chunk Serialization Support
-
LFM2.5-2.6B: On-Device Agentic Model With 128K Context and Tool Calling
-
llama.cpp Adds DeepSeek V4 Flash Chat Template Support
-
llama.cpp b10256 – SYCL SDPA Extended to Quantized KV Caches
-
llama.cpp Build b10258: Sampling Architecture Refinements
-
llama.cpp Release b10257 – Vulkan LLVMpipe Fixes
-
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
-
Q4 vs Q6 vs Q8: The Quantization Decision Framework for Local LLMs
-
Tim Cook Called Apple's On-Device AI a 'Competitive Weapon' in Final Earnings Call as CEO
-
Your Smartwatch Now Detects a Heart Irregularity in Milliseconds – Without Ever Touching the Cloud
-
GPU Half-Idle: The Hundred-Billion-Dollar Race to Squeeze 10x Efficiency from Silicon
-
4 Reasons I'm Canceling My ChatGPT Subscription for Local AI
-
Ask HN: What are you using for LLM inference in production?
-
Open-Weights AI Models Have Become Good Enough
-
CliffordNet: All You Need Is Geometric Algebra
-
Titan Transients and LLM Scalability
-
Gemma 4's Quantized Models Finally Made Local AI Practical in Homelab
-
Running Local LLMs on Raspberry Pi: Exploring Edge Inference Boundaries
-
Deploying 1-Bit Bonsai-27B with PrismML and llama.cpp for Local Inference
-
AMD Ryzen AI MAX+ 395 Discussed for Local AI Deployment
-
From Foldables to Smart Glasses, Samsung's Galaxy AI Push Moves Beyond the Cloud
-
Ruff v0.16.0: 413 Default Rules for Code Quality in AI Development
-
Edge AI Is Coming to Creative Production and It Will Change Everything
-
Don't Buy an Uncensored AI on a Flash Drive: What You Can Do Instead
-
Run the Mythos Enhanced Coding Model Locally with llama.cpp and Pi
-
Microsoft Strikes Multibillion-Dollar Deal with French AI Firm Mistral
-
llama.cpp b10075 Packs Four Local AI Runtime Upgrades
-
On-Device AI Ignites WAIC 2026: How Compute-in-Memory Chips Are Stuffing 100-Billion-Parameter LLMs Into Your Pocket
-
Sunday Reboot: Shrinking Models and an On-Device AI Future
-
This Open-Source Extension Lets You Rewrite Your X Algorithm Using a Local LLM, and It Healed My Timeline
-
Jan: Open, Cross-Platform AI App with Useful Proprietary Models
-
AI Inference Costs: Build vs. Rent
-
NVIDIA's On-Device AI Gains Japan's Manufacturing Giants' Backing
-
AMD Ryzen 7 7700X3D Linux Performance Review
-
Apple in Talks with PrismML to Shrink AI Models 15x for iPhone Deployment
-
llama.cpp's 4.26× Intel Gain Has a Narrow Catch
-
7 Python Frameworks for Orchestrating Local AI Agents
-
On-Device AI That Respects Your Privacy Gains Traction
-
Show HN: AITerm – a macOS Terminal with an AI Command Loop and a Safety Gate
-
Python 3.15's Ultra-Low Overhead Interpreter Profiling Mode – Ken Jin's Blog
-
ConlangCrafter: Constructing Languages with a Multi-Hop LLM Pipeline
-
Don't Sleep on BitNet (2025)
-
Google's LiteRT.js Enables On-Device AI Inference in Web Browsers
-
Indian Companies Look to Chinese LLMs as AI Costs Bite
-
Show HN: Turn Meeting Recordings into Searchable Transcripts. All Local
-
Show HN: Call to Control AI Agents via the Web
-
A Font That Humans Can Read But AI Cannot
-
Study: Cerebellum Helps AI Ignore the Ordinary for More Efficient Computing
-
Developer Ditches Ollama for llama.cpp's WebUI: A Practical Comparison
-
Record and Replay: Teach AI Agents Desktop Workflows by Showing Them Once
-
Show HN: OpenVole 4.5 Is Out
-
CorvinOS – Self-Hosted OS for AI Agents with Compliance Built Into Runtime
-
Exploiting Sparsity for Long Context Inference: Million Token on Commodity GPUs
-
The Triage Is the Product: Running AI Agents Against Ethereum's Protocol Code
-
Relm – Local LLMs as Base-R Objects with Interpretability
-
Apple's MacBook Lineup Overhaul Features M7 Chip for Enhanced Local AI
-
Local LLM Performance Gap With Frontier Models Smaller Than Expected
-
Amazon Developing Custom On-Device AI Chips for Echo and Fire TV Lineups
-
3 Local LLM Workflows That Actually Save Me Time
-
Article Compares Continuous and Static Batching in LLM Inference
-
Transcribe.cpp – ggml speech-to-text inference engine
-
I Quantized a Local LLM on My Home Server and Ditched Cloud AI for Smart Home Control Entirely
-
llama.cpp Tutorial: Run a Local LLM in 12 Steps
-
You Can Now Run Max AI Models on Apple Silicon
-
GEEKOM A9 Max Delivers 32GB RAM and Native LLM Support in Compact Form Factor
-
PewDiePie's Open-Source AI Workspace Gains Traction as Practical Local Deployment Platform
-
Liquid AI Ships LFM2.5-230M with Broad Framework Support for On-Device Inference
-
DEEPX and Sixfab Launch AI HAT for Raspberry Pi Edge Inference
-
Developer Replaces Entire Browser Extension Stack With Single Local LLM
-
Mac Mini Emerges as Top Choice for Local On-Device AI Deployment
-
Qwable: New Free Local Model Brings Claude-like Capabilities to Edge Devices
-
NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
-
Developers Run Local LLMs on Windows 11
-
What else is included in the 'GGUF' file format used by llama.cpp for AI language models, besides weights?
-
Qualcomm Launches Snapdragon START to Speed AI Smart Glasses to Market
-
FlashRT: Execution State for Latency-First AI
-
My Self-Hosted LLMs Are a Lot More Than Just a Chat Replacement – Here's How They Boost My Productivity
-
On-Device AI Market Projected to Reach $75.5 Billion by 2033
-
Intel Core Ultra X7 Panther Lake Performance Benchmarked on Linux
-
Companies Question Cost of AI as Token Maximization Spending Adds Up
-
Stop Guessing Which Local AI Models Fit Your Hardware — This Free Tool Does It for You
-
Most People Use Ollama or llama.cpp for Local LLMs, but These Are the Tools I Switch to When It Gets Serious
-
Repo-Slopscore: Detecting AI Contributions in Git Repositories via Commit Analysis
-
Ask HN: What Problem Did AI Create at Your Company That Didn't Exist Before?
-
Building Smart Home Analytics with Local LLMs: A Practical Setup Guide
-
AMD claims 256-core Zen 6 'Venice' CPU beats Nvidia Vera by 3.3x
-
Qualcomm Launches Dragonwing MBM Silicon with Advanced On-Device AI Capabilities
-
Google Chrome Quietly Deploys 4GB Local AI Model; Users Can Now Disable or Remove It
-
Prefill Once, Fan Out: KV Snapshot Sharing for Multi-Agent LLM Pipelines
-
Apple Unveils AFM 3 Core Advanced with 20 Billion Parameters for On-Device AI
-
TokenTamer: A Proxy That Reduces LLM Token Usage Through Context Compression
-
Developer Switches from LM Studio to llama.cpp, Citing Performance and Simplicity
-
AI bills can be as big as a postdoc salary. Is the cost worth it?
-
Google AI Edge Gallery Launches on macOS With Offline Gemini Models
-
Pizx – zx and Pi AI = shell scripting with 15 AI agent patterns
-
Google Introduces Gemma 4 QAT for Ultra-Low Memory Local Inference
-
Apple iPad Air with M4 Chip Drops to $1349; Powerful On-Device LLM Inference Now More Accessible
-
NVIDIA Unveils First PC Chips at Computex 2026; CEO Jensen Huang Details New Hardware
-
Best Local LLM Setup for RTX 5090: llama.cpp Fork with TurboQuant
-
Google's New Gemma 4 12B AI Model Is Built for Laptops
-
Running Infinite Context Lengths on 8GB GPU Without Out Of Memory
-
Show HN: CLI for Scoring OpenAPI for LLM Legibility
-
Show HN: Lowfat – Pluggable CLI Filter Saving 91.8% of LLM Tokens
-
Run Llama.cpp In-Process from Java with Project Panama FFM
-
WSL 3 Brings Near-Native GPU and NPU Passthrough for Local AI on Windows
-
NVIDIA RTX Spark Superchip Delivers 6,144 CUDA Cores for Consumer Local AI Inference
-
Tether AI Upgrades QVAC SDK With TurboQuant for Data Center-Sized Memory on Everyday Devices
-
NVIDIA and Microsoft Team Up to Bring Secure On-Device AI Agents to Windows PCs
-
JetBrains Releases Mellum2: A 12B MoE Model for Fast, Specialized Tasks
-
Phison and Intel Roll Out aiDAPTIV to Boost Local AI on Intel AI PC Platforms
-
Meet Memory OS: A 6-Layer Open-Source Memory Stack Built on Hermes Agent
-
Two LLM UI Patterns That Aren't Chat
-
Netflix Wiz Creates App to Slash AI Bills, Then Open Sources It
-
Nvidia Enters Windows Laptop Market, Taking on Intel and AMD
-
NVIDIA Levels Up Local AI Agents Across RTX PCs and DGX Spark
-
NVIDIA Launches N1X/N1 CPU-GPU SoC for PC Market, Targeting Heavy On-Device AI Users
-
Snapdragon C Specs Revealed: 6nm Process, On-Device AI Engine for Budget Laptops
-
Microsoft and Nvidia to Unveil First Windows PCs with Nvidia CPUs and AI Capabilities
-
Liquid AI Unveils Edge-Focused LFM2.5 Model for On-Device AI Agents
-
Mistral AI Launches Mistral Vibe
-
Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM Inference
-
llama.cpp GGUF Parser Flaws: Critical Integer Overflow Enables Arbitrary Reads in Every Local AI Stack
-
DeepSeek's Flagship V4 Pro Model Drops to 75% Lower Pricing, Increasing Competitive Pressure on Local Inference Economics
-
Dell Launches 14 Plus Laptop with Intel Core Ultra 9 and 32GB RAM at $1,499.99, Enabling Local Model Inference
-
Samsung's Exynos 2800 Brings HBM Memory to Mobile AI, Enabling Faster Local Model Inference
-
Developer Switches from LM Studio to llama.cpp, Reports No Performance Downgrade
-
Users Report Superior Performance Switching from LM Studio to llama.cpp
-
Gemma 4: A New Budget-Focused Model in Posit AI
-
Google Chrome Raises Privacy Questions with 4GB AI Model Download
-
How to Self-Host LibreChat with Docker
-
AMD Unveils Ryzen AI Halo Developer Platform for On-Device AI Workloads
-
Google Makes Gemini 3.5 Flash the Default AI Model for Billions of Users
-
llama.cpp MTP Leak Fix Stabilizes Local AI Agents
-
llama.cpp Checkpoint Fix Accelerates Local Coding Agents
-
User Migration from LM Studio/Ollama to llama.cpp Shows Growing Preference
-
AI Token Streaming Isn't About SSE vs. WebSockets
-
Chrome Is Quietly Downloading a 4GB AI Model Without Your Permission
-
I Stopped Trying to Replace My Cloud LLMs, and Local Models Finally Made Sense
-
llama.cpp Adds Multi-Token Prediction, Doubles Qwen 3.6B Throughput for Local Inference
-
Running Large Language Models on Single-Board Computer Clusters: Creative Edge Deployment
-
Samsung's Exynos 2800 Brings Significant On-Device AI Capabilities
-
Local LLMs Offer Unique Advantages That Cloud AI Services Cannot Match
-
The Time Bomb Went Off: AI's All-You-Can-Eat Era Just Ended in Real Time
-
Local LLMs Enable Intelligent Smart Camera Control Without Cloud Dependency
-
Linux 7.1-rc4 Released: Kernel Updates Relevant to Local LLM Inference
-
Google Limits Gemini Intelligence to New Flagships—Hardware Requirements for Local Deployment
-
Towards Local Plug-and-Play AI
-
Chrome Quietly Downloads 4GB AI Model Without User Permission
-
A Lo-Fi Rebellion Against A.I
-
Offline Voice-to-Text and AI Keyboard App for Local Processing
-
Local LLM Integration Enables Replacement of Paid Subscription Services
-
Chrome Silently Downloads 4GB Gemini Nano Model Without User Consent
-
Orthrus Reshapes Economics of Local AI Inference with New Optimization Approach
-
SynapseKit: A New Production Framework for Deploying LLMs
-
llama.cpp Delivers Sharp Performance Gains for AMD RDNA3 Users
-
AI, open code and vulnerability risk in the public sector
-
Running Local AI LLMs on Mini PCs Without NVIDIA GPUs
-
Local LLM Persistent Context Prevents Repetitive Mistakes
-
Running a Local LLM on a 12-Year-Old Raspberry Pi
-
How I Used a Local LLM to Organize the Store on My NAS
-
BT Explainer: Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Lucebox Brings Faster Local AI Inference to AMD Strix Halo
-
Mass NPM Supply Chain Attack Hits TanStack, Mistral AI, and 170 Packages
-
I Think I Figured Out What an AI IDE Looks Like
-
Running a Local LLM on a 12-Year-Old Raspberry Pi: Practical Edge Inference
-
Microsoft Researchers Find AI Models and Agents Can't Handle Long-Running Tasks
-
LLM Hallucinations in the Wild
-
DFlash Speculative Decoding Delivers 8.5x Speed Improvement for LLM Inference
-
$200 NVIDIA V100 Server GPU Mod Beats RTX 3060 in Local LLM Test
-
Cotypist – AI Autocomplete for Mac
-
Lython: Experimental Python Compiler Toolchain Based on LLVM
-
Mlx-serve: Run LLMs Natively on Your Mac
-
Chrome Is Secretly Downloading 4GB Gemini Nano Model Without User Consent
-
How to Run LLMs Locally on Your Laptop for Free: A Beginner's Guide
-
Google Releases Gemma 4 Multi-Token Prediction Drafters To Accelerate AI Inference
-
Google Removes Privacy Assurances After Stuffing Devices With Their AI Model
-
Google Chrome Downloads 4GB Gemini Nano Model Silently Without User Consent
-
Microsoft VibeVoice C++ Port Enables Local Voice AI on CPU and GPU Without Python
-
llama.cpp Now Supports Multi-Token Prediction in Beta
-
Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Supercharging LLM Inference on Google TPUs: Achieving 3X Speedups With Diffusion-Style Speculative Decoding
-
Gemma 4 Just Replaced My Whole Local LLM Stack
-
Google Drops COSMO: Experimental On-Device AI Assistant for Android
-
PFlash Claims 10x Prefill Speedup Over llama.cpp
-
Local LLMs Work Best When You're Not Loyal to Just One
-
Building a Raspberry Pi-Based Local LLM Server for Remote Access
-
How to Make SSE Token Streams Resumable, Cancellable, and Multi-Device
-
Ubuntu is Going All In on Generative AI and Other Linux Distros Might Follow
-
Linux Setup for Local LLMs Takes Minutes Compared to Windows Hours
-
Running Capable Local LLMs Without Expensive GPU Hardware
-
Show HN: Arkloop – Open-Source, Local-First Agent Client
-
How Much "Brain Damage" Can an LLM Tolerate?
-
Estimating Black-Box LLM Parameter Counts via Factual Capacity
-
Intel N150 Mini PC Runs Local LLM for Home Assistant
-
Llama.cpp Runs on SGI Power Challenge from 1995 with MIPS R8000 Kernel
-
NVIDIA Nemotron 3 Nano Omni Powers Multimodal Agent Reasoning in a Single Efficient Open Model
-
Grokfeed: Terminal Feed Reader for HN, Reddit, and Lobste.rs Using Claude Code
-
Picking Your First Local LLM Is Easier Than the Internet Makes It Sound
-
An Update on GitHub Availability: Infrastructure Lessons for Hosted LLM Tools
-
Hipfire: A Rust-Native AMD Inference Engine That Outperforms llama.cpp
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
Linux Crushes Windows on llama.cpp Inference by Double Digits
-
Run a Local LLM Server on Raspberry Pi with Remote Access Capabilities
-
Building Real-World On-Device AI with LiteRT and NPU
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results
-
Intel OpenVINO 2026.1 Integrates llama.cpp with Wildcat Lake and Arc Pro B70
-
Llama.cpp's Auto Fit Feature Quietly Reshapes Local AI Inference on Consumer Hardware
-
The Open-Source AI Ecosystem Keeps Treating llama.cpp Like a Second-Class Citizen
-
Malicious GGUF Models Could Trigger Remote Code Execution on SGLang Servers
-
AI Quota Inflation Is No Token Effort. It's Baked In
-
Bun v1.3.13
-
llama.cpp Merges Speculative Checkpointing for Major Inference Speed Boost
-
LlaMa.cpp Robot Wars
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
Kilo is the VS Code Extension That Actually Works with Every Local LLM
-
Unweight: Lossless MLP Weight Compression for LLM Inference
-
Sorting 1M u64 KV-Pairs in 20ms on i9-13980HX Using Branchless Rust Implementation
-
Kilo Is the VS Code Extension That Actually Works With Every Local LLM I Throw at It
-
ChatMCP – Connect your AI browser chats to your coding agents
-
The 'Ollama' Tool Has Numerous Problems, and Some Argue That Llama.cpp Is Better
-
Project Glasswing and the ASF: Open-Source's Chance to Win the AI Era
-
DotLLM – Building an LLM Inference Engine in C#
-
Dynamic Expert Cache in llama.cpp Achieves 27% Faster Inference on Large MoE Models
-
Sovereign AI: Why the Next GPT Will Be Born in Our Living Rooms
-
Qwen 3.5 Small – On-Device Multimodal Models Released
-
ASUS Malaysia to Bring UGen300 USB AI Accelerator in Q2 for Portable On-Device AI Inferencing
-
Self-Hosted LLM Took Personal Knowledge Management System to the Next Level
-
Speculative Decoding Achieves 29% Speed Boost for Gemma-4 31B
-
Qwen3 Audio and Vision Support Now Available in llama.cpp
-
Audio Processing Support Lands in llama.cpp with Gemma-4
-
MiniMax M2.7 Open-Sources Globally as Industry's First Self-Improving Model
-
MiniMax M2.7 Is Now Open Source
-
Users Report Significant Performance Improvements After Migrating from Ollama to llama.cpp
-
Critical Unsloth Gemma-4 Chat Template Updates for Tool Calling
-
Intel Arc Pro B70 32GB Achieves 12 Tokens/Sec on Qwen 3.5-27B
-
Tether Launches QVAC SDK for Cross-Platform Local AI Development
-
Ollama's Limitations for Production Local LLM Deployments
-
Gemma 4 Template Improvements Enhance Tool Use and Dialog Compliance
-
Gemini-CLI, Llama.cpp, and Qwen3.5 Running on NVIDIA Jetson TK1
-
Intel Releases OpenVINO 2026.1 With Backend For Llama.cpp, New Hardware Support
-
Gemma 4 Support Stabilized in Llama.cpp
-
Gemma 4 GGUF Models Updated with Critical Quantization Fixes
-
EXAONE 4.5 33B Model Released with Multiple Quantization Formats
-
Speculative Decoding Made My Local LLM Actually Usable
-
Ollama is Still the Easiest Way to Start Local LLMs, But It's the Worst Way to Keep Running Them
-
MemPalace, the Highest-Scoring AI Memory System Ever Benchmarked
-
TurboQuant-Optimized llama.cpp Fork Delivers GFX906 GPU Acceleration
-
TurboQuant in Llama.cpp Achieves 6X Smaller KV Cache
-
Google AI Edge Gallery Tops App Store Charts with On-Device Gemma 4
-
Quantization Strategy Comparison: Balancing Quality and Speed on Consumer Laptops
-
Context Window Optimization: Extending Gemma 4 Context Length Through Efficient Projection Quantization
-
GPU Memory for LLM Inference (Part 1)
-
Qwen 3.6 Free Model Available via OpenRouter
-
Vektor – Local-First Associative Memory for AI Agents
-
Unpaved: Audit Toolkit for AI Developer Tool Bias in Global South Contexts
-
Microsoft Quantum Development Kit Ported to Rust: 100x Faster and Smaller
-
Apple Research Shows Self-Distillation Significantly Improves Local Code Generation
-
GPUs vs. TPUs: Decoding the Powerhouses of AI
-
Gemma 4 KV Cache Memory Issues Fixed in llama.cpp
-
Gemma 4 2B Successfully Runs on Raspberry Pi 5
-
VRAM Optimization Technique Cuts Gemma 4 Memory Usage by 3x
-
Google Gemma 4 Released with GGUF Quantizations
-
OpenUMA – Apple-Style Unified Memory for x86 AI Inference
-
Intel's $949 GPU Has 32GB of VRAM for Local AI, but Software is Why Nvidia Keeps Winning
-
Show HN: Extra-Platforms, Python Library to Detect OS, Arch, Shell, CI, AI
-
SmolLM2-360M Running on Samsung Galaxy Watch 4 with 74% Memory Reduction
-
Gemini CLI – Open-Source AI Agent for Terminal Integration
-
Llama.cpp Merging TurboQuant Lite (attn-rot) with Major Performance Gains
-
ROCm Integration in Ubuntu 26.04 Advances Linux GPU Inference
-
Local AI Ecosystem Extends Far Beyond Ollama
-
Closed Source AI = Neofeudalism
-
Samsung launches Galaxy Book6 series in India with Nvidia RTX 5070 graphics and on-device AI
-
Intel's $949 GPU has 32GB of VRAM for local AI, but the software is why Nvidia keeps winning
-
DeepSeek V3 Complete Guide: Deploy and Optimize Local AI in 2026
-
Local AI Ecosystem Extends Far Beyond Ollama
-
Unsloth Studio Beta Ships 50+ New Features for Local Model Training and Inference
-
Introduction to Nyreth v1.0
-
HP Launches Copilot+ PCs in India with On-Device AI Capabilities for Local Inference
-
TurboQuant KV Cache Compression Achieves 22.8% Faster Decoding at 32K Context
-
Coding Implementation to Run Qwen3.5 Reasoning Models Distilled With Claude-Style Thinking Using GGUF and 4-Bit Quantization
-
TurboQuant Benchmarked in Llama.cpp: Google's Extreme Compression Research Tested in Practice
-
RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
-
Quantization Reveals Outliers Impacting LLM Accuracy
-
Nota AI and SiMa.ai Partner on Physical AI Technology for Local Deployment
-
Pluggable's TBT5-AI: First Thunderbolt Dock Explicitly Targeting Local LLM Workstations
-
Google's TurboQuant: The Unsexy AI Breakthrough Worth Watching
-
Apple Plans Slimmed-Down Gemini Models for Local iPhone AI Features
-
OmniCoder v2 Released: Improved Code Generation for Local Deployment
-
Researcher Successfully Runs Local LLMs on Legacy "Dead" GPU With Surprising Results
-
Show HN: Open Agent Spec – Treat AI Agents Like Typed Functions, Not Prompt Chains
-
Llama.cpp Benchmark: RTX 5090 vs Enterprise Systems Compared
-
I built Rubric, an open source Sentry for AI. Looking for beta testers
-
Llama.cpp ROCm 7 vs Vulkan Performance Benchmarks on AMD Mi50
-
LM Studio Releases Reworked Plugins with Fully Local Web Research
-
Automating Read-It-Later Workflows with Local LLMs for Overnight Summarization
-
Setting Up a Private AI Brain on Windows: Complete Guide to Local LLM Deployment
-
Careless Whisper – Personal Local Speech to Text
-
Rust Project Perspectives on AI
-
ik_llama.cpp Fork Delivers 26x Faster Prompt Processing on Qwen 3.5 27B
-
What AI Augmentation Means for Technical Leaders
-
Qualcomm and Samsung's 30-Year AI Alliance Enters a New Phase as On-Device AI Chip Race Heats Up
-
Community Converges on Optimal KV Cache Quantization Strategies for Qwen 3.5 Models
-
LMCache Dramatically Accelerates LLM Inference on Oracle Data Science Platform
-
Kilo Is the VS Code Extension That Actually Works With Every Local LLM I Throw At It
-
You're Using Your Local LLM Wrong If You're Prompting It Like a Cloud LLM
-
Hugging Face Releases One-Liner for Automatic Hardware Detection and Model Selection
-
LucidShark – Local-first, open-source quality and security gate
-
Unsloth Studio: Open-Source Web UI for Training and Running LLMs Locally
-
I Switched to a Local LLM for These 5 Tasks and the Cloud Version Hasn't Been Worth It Since
-
On-Device AI: Tether's QVAC Fabric Enables Local Training
-
Run LLMs Locally with Llama.cpp
-
I Ran Local LLMs on a 'Dead' GPU, and the Results Surprised Me
-
Mistral Releases Small 4 Open-Source Model Under Apache 2.0
-
How I Used Lima for an AI Coding Agent Sandbox
-
Kimi Introduces Attention Residuals: 1.25x Compute Performance at <2% Overhead
-
Practical Fix for Qwen 3.5 Overthinking in llama.cpp
-
Apple's On-Device AI Raises Privacy Alarms Across British Parliament
-
This External GPU Enclosure Tries to Break Cloud Dependence for Local AI Inference
-
AMD Declares 'AI on the PC Has Crossed an Important Line' – Agent Computers as Next Breakthrough
-
OpenClaw vs Eigent vs Claude Cowork: Comparing Open-Source AI Collaboration Platforms
-
Running Qwen3.5-27B Across Multiple GPUs Over LAN Achieves Practical Speed for Local Inference
-
AMD Launches Agent System Optimized for Local AI Inference With Ryzen and Radeon
-
How to Run Local LLMs in 2026: The Complete Developer's Guide
-
Intel OpenVINO Backend Support Now Available in llama.cpp
-
Memory Should Decay: Implementing Temporal Memory Decay in Local LLM Systems
-
AgentArmor: Open-Source 8-Layer Security Framework for AI Agents
-
3-Path Agent Memory: 8 KB Recurrent State vs. 156 MB KV Cache at 10K Tokens
-
Llama.cpp Adds True Reasoning Budget Support
-
Quantization Explained: Q4_K_M vs AWQ vs FP16 for Local LLMs
-
Nvidia Releases Nemotron 3 Super: 120B MoE Model for Local Deployment
-
Cutile.jl Brings Nvidia CUDA Tile-Based Programming to Julia
-
SK Hynix Completes Qualification for LPDDR6 Memory Optimized for AI Inference
-
LMF – LLM Markup Format
-
NVIDIA Jetson Brings Open Models to Life at the Edge
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Llama.cpp Celebrates Major Milestone: From Leak to Industry Standard
-
Mnemos: Persistent Memory System for Local AI Agents
-
FreeBSD 14.4 Released: Implications for Local LLM Deployment
-
Community Survey: AI Content Automation Stacks in 2026
-
M5 Max and M5 Ultra Chipsets Demonstrate Significant Bandwidth Improvements for Local LLM Inference
-
Qwen 3.5 Ultra-Compact Models Enable On-Device AI from Watches to Gaming
-
Strix Halo (Ryzen AI Max+ 395) Achieves Strong Local Inference Performance with ROCm 7.2
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
HP Refreshes Lineup with AI-Focused Workstations
-
Llama.cpp Prompt Processing Optimization: Ubatch Size Configuration Guide
-
Llama.cpp Merges Automatic Parser Generator to Mainline
-
Turning Your Linux Terminal into a Local AI Assistant
-
llama.cpp Merges Agentic Loop and MCP Client Support
-
Apple Unveils MacBook Pro with M5 Pro and M5 Max Featuring On-Device AI
-
Apple Unveils MacBook Pro With M5 Pro and M5 Max for On-Device AI
-
OpenWrt 25.12.0 – Stable Release
-
AMD Launches Copilot+ Desktop Chips to Compete in On-Device AI Market
-
ÆTHERYA Core – Deterministic Policy Engine for Governing LLM Actions
-
Qwen 3.5 0.8B Successfully Deployed on 7-Year-Old Samsung S10E Using llama.cpp
-
Qwen 3.5 Small Models Released: 0.8B to 9B Parameters Optimized for On-Device Inference
-
Framework Choice Critical: llama.cpp and vLLM Outperform Ollama for Qwen 3.5 Testing
-
C7: Pipe Up-to-Date Library Docs Into Any LLM From the Terminal
-
GitDelivr: A Free CDN for Git Clones Built on Cloudflare Workers and R2
-
Critical: Qwen 3.5 Requires BF16 KV Cache, Not FP16 for Accurate Inference
-
Huawei's SuperPoD Portfolio Creates New Option for Global Computing at MWC Barcelona 2026
-
5 Useful Docker Containers for Agentic Developers
-
Unsloth Dynamic 2.0 GGUFs
-
Seco Launches Edge AI System-on-Module at Embedded World 2026
-
Arduino and Qualcomm Bring On-Device AI Learning to Indian Schools
-
DeepSeek Releases DualPath: Addressing Storage Bandwidth Bottlenecks in Agentic Inference
-
DeepSeek Paper – DualPath: Breaking the Bandwidth Bottleneck in LLM Inference
-
Qwen3.5 Thinking Mode Can Be Disabled for Production Inference Optimization
-
Qwen3.5-27B Identified as Sweet Spot for Mid-Range Local Deployment
-
Mirai Announces $10M to Advance On-Device AI Performance for Consumer Devices
-
How AI is Redefining Price and Performance in Modern Laptops
-
Advanced Quantization Techniques Show Surprising Performance Gains Over Standard Methods
-
Apple Accelerates U.S. Manufacturing with Mac Mini Production
-
Show HN: A Ground Up TLS 1.3 Client Written in C
-
Enterprise Infrastructure Guide: Running Local LLMs for 70-150 Developers
-
nanollama: Open-Source Framework for Training Llama 3 from Scratch with One-Command GGUF Export
-
Open-Source llama.cpp Finds Long-Term Home at Hugging Face
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
Ouro 2.6B Thinking Model GGUFs Released with Q8_0 and Q4_K_M Quantization
-
Open-Source + AI: ggml Joins Hugging Face, llama.cpp Stays Open—Local AI's Long-Term Home
-
I Thought I Needed a GPU to Run AI Until I Learned About These Models
-
Strix Halo Performance Benchmarks: Minimax M2.5, Step 3.5 Flash, Qwen3 Coder
-
GGML.AI Acquired by Hugging Face
-
Kitten TTS V0.8 Released: New State-of-the-Art Super-Tiny TTS Model Under 25 MB
-
PaddleOCR-VL Now Integrated into llama.cpp for Multilingual OCR
-
Enhanced Quantization Visualization Methods for Understanding LLM Compression Trade-offs
-
Qwen 3.5-397B-A17B Now Available for Local Inference with Aggressive Quantisation
-
Self-Hosted AI: A Complete Roadmap for Beginners
-
Meet Sarvam Edge: India's AI Model That Runs on Phones and Laptops With No Internet
-
Open-Source Models Now Comprise 4 of Top 5 Most-Used Endpoints on OpenRouter
-
GNOME's AI Assistant Newelle Adds llama.cpp Support and Command Execution
-
Scaling llama.cpp On Neoverse N2: Solving Cross-NUMA Performance Issues
-
MiniMax-M2.5 230B MoE Model Released with GGUF Support for Local Deployment
-
SnowBall Technique Addresses Context Window Limitations in Local LLMs
-
MiniMax Releases M2.5 Model with SOTA Coding and Agent Capabilities
-
Context Management Identified as Real Bottleneck in AI-Assisted Coding
-
Optimal llama.cpp Settings Found for Qwen3 Coder Next Loop Issues
-
GitHub Announces Support for Open Source AI Project Maintainers
-
New Header-Only C++ Benchmark Tool for Predictive Models on Raw Binary Streams
-
Developer Switches from Ollama and LM Studio to llama.cpp for Better Performance