Tagged "quantization"
-
Quantization-Aware Healing: 4-Bit Models Outperform Full-Precision Originals
-
FreeToken: Edge-Native MoE Serving Engine Runs 753B GLM-5.2 on Single Workstation GPU
-
Google COSMO Leak Reveals Gemini Nano and On-Device AI Skills
-
Local LLM Generates Dynamic UIs on $30 ESP32 Display
-
Strong Domain Adaptation Results with Qwen 3 4B Fine-Tuning
-
Ollama v0.33.0 Adds Claude Desktop Integration and App Management
-
llama.cpp Build b10581 Adds DSpark Support for Faster Local Inference
-
Liquid AI Releases DSpark Version of Compact LFM2.5 Models with Up to 2.67x Speedup
-
Self-Hosting AI Models on a Raspberry Pi 5: A Complete Guide to Free, Private, Local AI Inference
-
Qwen3.8-27B: Running a Frontier-class Open Model on Your Local GPU
-
Liquid AI Releases LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation
-
Ollama Runs Free AI Models Locally on Mac, Windows and Linux
-
Qwen3.8-27B Matches Claude Opus 4.6 on Coding, Runs on Consumer GPUs
-
What If Local LLM Inference Is Using Consumer Hardware Wrong?
-
GGUF Quantization Deep Dive: Q4_K_M vs IQ4_XS vs IQ4_NL Performance
-
DeepSeek V4 Flash Shrunk to 57GB for Local macOS Inference with Compiler Generation
-
Qwen3.8-27B Surpasses 1 Million Downloads, Overseas Developers Race to Maximize Local Deployment
-
GGUF Quantization Compared: Q4_K_M vs. IQ4_XS vs. IQ4_NL Performance Analysis
-
AMD Adds Day 0 Qwen3.8 Support, Radeon AI PRO R9700 Hits 51.8 Tokens per Second
-
Google Pixel 11 Launches With Faster On-Device Gemini at $899 Starting Price
-
The Qwen MLX Challenge
-
How an $8 ESP32 S3 Microcontroller Runs a 28.9M Parameter Local LLM
-
Show HN: I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac
-
Meta's Muse Glimmer Achieves Fast On-Device Agentic AI with ExecuTorch
-
HackerNoon Compares 7 Best Self-Hosted Inference Servers for Open-Source Models
-
Unsloth Releases Qwen 3.8 27B GGUF Quantised Weights
-
Qwen 3.8 27B Successfully Runs on 16GB RAM Using LM Studio
-
Ollama Adds Qwen 3.8 27B with Optimised Apple Silicon Support
-
AMD Optimizes Qwen 3.8 27B for Ryzen AI Max and Radeon GPUs
-
Hugging Face State of Open Models: Summer 2026 Observations
-
7 Best Self-Hosted Inference Servers for Open-Source Models Compared (2026)
-
DeepX's DX-M1 On-Device AI Chip Achieves $13M in Orders
-
Liquid AI Releases LFM2.5-VL-3B: Compact Vision-Language Model for Edge Inference
-
Running DeepSeek's 284B LLM on a Laptop: Quantisation and GGUF Optimization
-
LFM2.5-VL-3B: Lightweight Vision-Language Model Optimized for Edge Deployment
-
Ollama v0.32.10: Faster Prefill Performance on NVFP4 Models with System Config Support
-
How to Run Local LLMs for Free on Slow Laptops: A Practical Guide
-
Benchmarking Local LLMs on Consumer Hardware: Real-World Performance Data
-
Minisforum N5 Max: Running Qwen 27B Locally with Open WebUI and Ollama
-
NVIDIA Enables Local Agentic AI Workflows with Meta's Muse Glimmer
-
Meta's Muse Glimmer – Local, Agentic, Multimodal, and Open Source
-
Meta Releases Muse Glimmer: 30B Open-Source LLM for Local Deployment
-
How to Install Ollama on Windows 11 for Local AI Inference
-
DEF CON 34 Exposes 10 Critical Vulnerabilities in Local AI Systems
-
vLLM v0.27.0rc2 Release Candidate Available
-
On-Device AI Market Combines AI Operations With Local Processing
-
llama.cpp Improves CUDA Performance with Kernel Fusion
-
How to Run a Local LLM With Ollama: 13 Steps, 90 Min
-
Running AI Agents on Mobile: Phone Transformed Into Self-Installing LLM Agent
-
Chrome's On-Device AI Model Requires 20GB Storage Space
-
Llama.cpp B10327 Fixes CUDA Quantized Copy Kernel Performance
-
Chrome and Edge Browsers Quietly Deploy Up to 20GB AI Models on Windows 11
-
Optimizing Qwen 3.6 for Local Development: A Developer's Guide
-
Show HN: Local Multi-Agent AI Running on Android Phone
-
Liquid AI LFM2.5-2.6B: Open-Weights Agentic Model With 128K Context and Tool Calling
-
Shrinking an AI Model 86% Doesn't Make It 86% Dumber: Compression Breakthroughs
-
Self-Hosted LLM Costs 2026: Comprehensive Pricing Comparison
-
Show HN: Benchmark Local LLMs Fit for Your Device Specs
-
llama.cpp Build b10301: CUDA Optimization and Compiler Warning Fixes
-
Google Chrome Reveals Storage Requirements for Integrated Local AI Models
-
LFM2.5-2.6B: On-Device Agentic Model With 128K Context and Tool Calling
-
Voice Notes Shouldn't Cross the Ocean – Keep Your Thoughts Private
-
PrismML's Bonsai 27B Brings On-Device AI to Apple iPhone 17 Pro
-
ASUS Vivobook S16 Arrives with 45 TOPS NPU and OLED Display
-
Homebench: Comprehensive Benchmarking Tool for Local LLMs
-
llama.cpp b10256 – SYCL SDPA Extended to Quantized KV Caches
-
K-EXAONE 2.0 Brings 262K Context to Frontier AI
-
28.9M-Parameter LLM Runs Locally on ESP32-S3 at 9 Tokens/s
-
Kioxia Is Coming for Samsung and SK Hynix With UFS 5.0 and PCIe 6.0 AI NAND
-
Oppo Reno16 Pro 5G Pairs On-Device AI With a 6,700mAh Battery for Creators
-
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
-
The KV Cache Survival Guide: Why Your GPU Runs Out of Memory with Local LLMs
-
NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework
-
The KV Cache Survival Guide: Why Your GPU Runs Out of Memory with Local LLMs
-
Q4 vs Q6 vs Q8: The Quantization Decision Framework for Local LLMs
-
Tim Cook Called Apple's On-Device AI a 'Competitive Weapon' in Final Earnings Call as CEO
-
Your Smartwatch Now Detects a Heart Irregularity in Milliseconds – Without Ever Touching the Cloud
-
GPU Half-Idle: The Hundred-Billion-Dollar Race to Squeeze 10x Efficiency from Silicon
-
4 Reasons I'm Canceling My ChatGPT Subscription for Local AI
-
Samsung's Newest Foldable Phones Use Google's Gemini Nano 4 On-Device AI Model
-
Phi-4 Mini vs Gemma 3 vs Llama 3.2: 128K vs 32K Context Window Comparison
-
Ask HN: What are you using for LLM inference in production?
-
Kioxia UFS 5.0 Embedded Flash Memory Enables On-Device AI with Advanced Storage Architecture
-
I Built a Free AI Curriculum from Philosophy to LLMs
-
EU Opens Call for Seven 'Gigafactories' to Train Next-Generation AI
-
CliffordNet: All You Need Is Geometric Algebra
-
Kioxia's UFS 5.0 Embedded Flash Enables Practical On-Device AI
-
How Much Does a Local LLM Actually Cost to Run? Energy Costs Measured on Apple Silicon
-
Run a Local LLM on Raspberry Pi's Bare Metal—Linux Not Necessary
-
Can a 2.8T Model Run on a Single Node of Nvidia B300 X8?
-
Gemma 4's Quantized Models Finally Made Local AI Practical in Homelab
-
Running Local LLMs on Raspberry Pi: Exploring Edge Inference Boundaries
-
Deploying 1-Bit Bonsai-27B with PrismML and llama.cpp for Local Inference
-
AMD Ryzen AI MAX+ 395 Discussed for Local AI Deployment
-
Brief notes on the OpenAI/Hugging Face incident
-
CPU vs GPU vs NPU: Which Semiconductor Does What?
-
Show HN: Agent Console – A Local Dashboard for Codex and Claude Code
-
OPPO Launches Xiaobu Next Beta, Debuts On-Device Multi-Agent System on Smartphones
-
GitHub Copilot With Ollama: Run Local AI Models In VS Code Offline
-
No Wi-Fi, No Data Transfer, Tablets Can Now Summarise Sensitive Documents Locally
-
From Foldables to Smart Glasses, Samsung's Galaxy AI Push Moves Beyond the Cloud
-
Claude Code Cut System Prompt by 80%: Implications for Small Local Models
-
Build Self-Scaling OCR Pipeline with Qwen 3.5 and Kubernetes
-
Edge AI Is Coming to Creative Production and It Will Change Everything
-
Nota AI Joins AMD Robotics Partner Network to Expand On-Device AI Optimisation
-
Round-Trip Correctness: New Metric for Generative AI Process Modeling
-
SK hynix 3D-Stacked DRAM-on-Logic Architecture Could Solve On-Device AI Memory Constraints
-
Gemini Nano 4 Arrives with Samsung's Latest Foldables, Bringing LLMs to Mobile Edge
-
Google's Gemma AI Runs Locally on a $300 Mini PC, and It Replaced ChatGPT
-
Arm China Unveils "Tianxuan" CPU and Xingchen 300 Platform, Targeting Ubiquitous AIoT with On-Device AI Portfolio
-
My Local LLM Struggles with Big Questions—Here's What It's Actually Good At
-
AI Inference is Rewriting the GPU Buying Playbook
-
Microsoft Strikes Multibillion-Dollar Deal with French AI Firm Mistral
-
AMD Acquires FastFlowLM to Accelerate On-Device AI Inferencing
-
llama.cpp b10075 Packs Four Local AI Runtime Upgrades
-
On-Device AI vs Cloud AI: Which One Should Power Your Next Phone?
-
Sunday Reboot: Shrinking Models and an On-Device AI Future
-
LLM Wiki Implementation: Community Resource for Local Deployment
-
AI Data Center Power Constraints Are the Real 2026 Bottleneck
-
This Open-Source Extension Lets You Rewrite Your X Algorithm Using a Local LLM, and It Healed My Timeline
-
Nubia Announces AI Agent Smartphone with On-Device AI Processing
-
Qualcomm's Xu Hao: Agentic AI Phones Surge as On-Device AI Shifts from Passive Response to Proactive Service
-
Samsung Galaxy Watch 9 to Feature Snapdragon Wear Elite Chip: Report
-
Qwen 3.8 with 2.4T Parameters Going Open-Weight Soon
-
Jan: Open, Cross-Platform AI App with Useful Proprietary Models
-
Apple in Early Talks With PrismML on AI Compression Tech
-
Host Private Local AI on NVIDIA DGX Spark Using Ollama and Open WebUI
-
Google Demonstrates New On-Device AI Features for Pixel 10
-
I Thought My Local AI Would Replace My Claude Subscription — Then I Tried Automating My PC
-
AMD Ryzen 7 7700X3D Linux Performance Review
-
Google Gemma 4 Debuts for Pixel 10 With Powerful On-Device AI Features
-
Apple in Talks with PrismML to Shrink AI Models 15x for iPhone Deployment
-
Mira Murati's Thinking Machines Launches Open-Weight AI Model
-
llama.cpp's 4.26× Intel Gain Has a Narrow Catch
-
Google expands on-device AI for Pixel phones with Gemma 4
-
Apple Boosts On-Device AI, Partners With PrismML to Enable Running Large Models Locally on iPhone
-
Don't Sleep on BitNet (2025)
-
Indian Companies Look to Chinese LLMs as AI Costs Bite
-
The 5 Coolest Open-Source Projects I've Discovered in 2026
-
Show HN: GGUFun, Play Snake and a Simple Maze on Ollama Using Hand Crafted GGUFs
-
CEO Calls for Lower AI Pricing to Enable Practical Labor Automation Deployment
-
Google Pixel Implements Local AI for Screenshot Analysis With Privacy Controls
-
WSL Transforms Windows Into a Viable Local LLM Development Platform
-
Edge AI Brings On-Device Intelligence and Health Monitoring to Smartwatches
-
Companies Are Scrambling to Curtail Soaring AI Costs
-
Cost vs. Accuracy in CursorBench 3.1: The Effect of Family and Spend
-
AMD ZenDNN 6.0 Boosts AI Inference on EPYC CPUs With FP16 and MoE Acceleration
-
Study: Cerebellum Helps AI Ignore the Ordinary for More Efficient Computing
-
Developer Ditches Ollama for llama.cpp's WebUI: A Practical Comparison
-
GitHub Copilot With Ollama: Run Local AI Models In VS Code Offline & Free
-
Making AI Code Review Measurable
-
What Every AI Builder Learns the Hard Way
-
Show HN: Tarit – Self-host Sandbox Cloud and Hypervisor for AI Agents
-
Viability of Local Models for Coding
-
Ollama is the Easiest Way to Start Local LLMs, But These 6 Alternatives Are Also Worth Trying
-
Ollama Runs 32B Local AI Models on a $599 Mac via Quantization for Free
-
Syntiant Files for IPO on Momentum of Low-Power On-Device AI Chip Demand
-
Samsung UFS 5.0 Storage Interface Optimizes On-Device AI Performance and Latency
-
NIS2 Compliance Drives European Office Software Toward Local AI Solutions
-
Edge AI Transformation Coming to Creative Production Workflows
-
Google Rolls Out Android 17 and Gemma 4 with Advanced On-Device AI
-
Compressor V2: Three Compression Layers for 50% LLM Agent Cost Cut
-
VisionAId: On-Device Vision for the Visually Impaired
-
Ollama is the Open-Source App That Finally Made Free Local AI Useful on My PC
-
Open Source 1B LLM Trained from Scratch for $315 with Weights and Data Released
-
Open Source AI Must Win: A Call to Action for the Local LLM Community
-
Local LLM Performance Gap With Frontier Models Smaller Than Expected
-
Theoretical Bottlenecks for Scaling LLM Inference to Achieve Higher Token per Second
-
Practitioner Quantized Local LLM for Smart Home Control, Eliminating Cloud Dependency
-
Article Compares Continuous and Static Batching in LLM Inference
-
Transcribe.cpp – ggml speech-to-text inference engine
-
I Quantized a Local LLM on My Home Server and Ditched Cloud AI for Smart Home Control Entirely
-
Samsung Unveils UFS 5.0 Solution for Next-Gen On-Device AI Applications
-
How to Choose Between Small and Frontier Models
-
Samsung Presents UFS 5.0 Storage Targeted at On-Device AI Performance
-
Reachy Mini Adds Local Conversational AI
-
Google's Gemma AI Runs Locally on a $300 Mini PC, and It Replaced ChatGPT for More Than Expected
-
llama.cpp Tutorial: Run a Local LLM in 12 Steps
-
Tiny LLM Benchmark: Jetson Orin Nano Super 8GB
-
Qualcomm AI Hub Expands to 1,500 Optimized Models for Edge Deployment
-
Liquid AI Ships LFM2.5-230M with Broad Framework Support for On-Device Inference
-
TriAttention Solves KV Cache Memory Bottleneck in Local LLM Inference
-
A Guide on How to Run Nemotron 3 Super 120B Thinking on 2 Nvidia DGX Spark
-
Hermes MoA Virtual Models: 8% Higher Than Opus 4.8, 11% Higher Than GPT 5.5
-
Qualcomm Brings Data Center AI Technology to Smartphones for Enhanced On-Device Capabilities
-
DEEPX and Sixfab Launch AI HAT for Raspberry Pi Edge Inference
-
Building Tool-Using Agents With Local LLMs
-
Developer Replaces Entire Browser Extension Stack With Single Local LLM
-
I Ran a Local LLM on My Underpowered Chromebook, and It Actually Works
-
Samsung Unveils UFS 5.0 Storage Solution Optimized for On-Device AI
-
Qualcomm Acquires Modular AI in $3.9 Billion Deal to Accelerate On-Device AI
-
Mac Mini Emerges as Top Choice for Local On-Device AI Deployment
-
ORA: Smaller Models. Same Intelligence
-
NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
-
Samsung Develops UFS 5.0 Flash Storage for On-Device AI with 10.8GB/s Speeds
-
Why Small Local AI Models Get More Use Than Claude or Gemini
-
Developers Run Local LLMs on Windows 11
-
Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
-
2026 On-Device AI Market Intensifies: Apple, Google, and Samsung Compete for Local AI Dominance
-
Mac Mini Positioned as Premier On-Device AI Computer for Local LLM Inference
-
Samsung's UFS 5.0 Addresses Critical Memory Bandwidth Bottleneck in Mobile AI Inference
-
Snapdragon Reality Elite: What is it, new devices announced, and more
-
Xiaomi vs Huawei On-Device AI: Decoding the AI Strategies of 8 Major Smartphone Giants
-
Offline Raspberry Pi Voice Assistant Runs Local LLM
-
Form Before Data: Addressing the Real Bottleneck in Physical AI Systems
-
DeepSWE Benchmark Updated with GLM 5.2 and Expanded Model Comparisons
-
What else is included in the 'GGUF' file format used by llama.cpp for AI language models, besides weights?
-
Qualcomm Launches Snapdragon START to Speed AI Smart Glasses to Market
-
Switching AI Tools Mid-Sprint Cost Us a Day (and What We Learned)
-
Show HN: I built an 11-LLM consensus engine to detect AI hallucination
-
Free Tool Helps Match Local AI Models to Your Hardware
-
Gaming PC vs Phone Local LLM Deployment: Only One Remains in Daily Use
-
Developer Replaces Entire Browser Extension Stack with Single Local LLM
-
Qualcomm Debuts Snapdragon Reality Elite XR Platform with On-Device AI
-
On-Device AI Market Projected to Reach $75.5 Billion by 2033
-
Intel Core Ultra X7 Panther Lake Performance Benchmarked on Linux
-
Tryll Engine Raises $600K to Deploy On-Device AI Characters in Games
-
Genesis AI Launches Eno General-Purpose Robot with Embedded AI
-
An End-to-End Machine Learning Pipeline on Time-Series Data
-
Qwen and Fable: Open-Weights 35B Mixture-of-Experts Agentic Coding Model
-
Ollama Emerges as Leading Open-Source Local AI Platform
-
Tensordyne Napier AI Processor Announced with Logarithmic Math
-
Brick: State-of-the-Art LLM Routing
-
Samsung's Exynos 2600 Doubles On-Device AI Performance in MLPerf Benchmarks
-
Stop Guessing Which Local AI Models Fit Your Hardware — This Free Tool Does It for You
-
Most People Use Ollama or llama.cpp for Local LLMs, but These Are the Tools I Switch to When It Gets Serious
-
Why Tool Calling is More Important Than Model Size for Local LLMs
-
General-Purpose Large Language Models Outperform Specialized Clinical AI
-
Ask HN: What Problem Did AI Create at Your Company That Didn't Exist Before?
-
Building Smart Home Analytics with Local LLMs: A Practical Setup Guide
-
Brilliant Labs Halo: Open-Source AI Glasses for On-Device Intelligence
-
Qualcomm Snapdragon 8 Gen 4: Flagship Chip Powering the Next Wave of Premium Android Phones
-
RTX 5080 and RTX 3090 Setup Achieves 80 Tok/s on Qwen 3.6 27B Q8
-
vLLM vs Ollama 2026: 793 vs 41 TPS Performance Benchmark
-
Google Chrome Quietly Deploys 4GB Local AI Model; Users Can Now Disable or Remove It
-
DeepSeek V4 Performance Analysis: 1.6T Day 0 to Day 43 Scaling Trends
-
Apple Rebuilt Its On-Device AI Stack at WWDC 2026
-
Qualcomm Unveils Dragonwing MBM Silicon with Integrated On-Device AI and Connectivity
-
Google Releases Gemma 4 QAT Models with Reduced Memory Requirements for Mobile and Laptop Deployment
-
Developer Reports Ollama Setup Takes Minutes Compared to Hours with LM Studio
-
Apple Enhances Siri With On-Device AI for Faster, Private Voice Responses
-
Show HN: Veritrooper – find what your AI gets wrong about your own docs
-
AI bills can be as big as a postdoc salary. Is the cost worth it?
-
Ask HN: What is the AI setup for an experienced dev starting on a new project?
-
Google Introduces Gemma 4 QAT for Ultra-Low Memory Local Inference
-
NVIDIA Unveils First PC Chips at Computex 2026; CEO Jensen Huang Details New Hardware
-
Best Local LLM Setup for RTX 5090: llama.cpp Fork with TurboQuant
-
Google Releases Gemma 4 QAT Models for Local AI Deployment
-
Running Local AI Models on Old Laptops Without GPU
-
Running Infinite Context Lengths on 8GB GPU Without Out Of Memory
-
Qualcomm Snapdragon C Specifications Revealed: 6nm Process with Dedicated On-Device AI Engine
-
Show HN: Lowfat – Pluggable CLI Filter Saving 91.8% of LLM Tokens
-
Google Launches AI Edge Gallery on macOS for Running Gemini Models Locally
-
Longsys Redefines On-Device AI with Groundbreaking Edge Memory Solutions
-
Reducing GPU Costs for AI Inference: FP8, FP4, and vLLM Optimization Techniques
-
Google Releases Gemma 4 12B: Encoder-Free Multimodal Model for 16GB Laptops
-
Snapdragon C Processor Brings On-Device AI Engine to Wearables and Edge Devices
-
NVIDIA RTX Spark Superchip Delivers 6,144 CUDA Cores for Consumer Local AI Inference
-
A Cinematic Landing-Page Hero for 80 Cents (GPT Image 2 and Veo 3.1)
-
Tether AI Upgrades QVAC SDK With TurboQuant for Data Center-Sized Memory on Everyday Devices
-
Phison and Intel Roll Out aiDAPTIV to Boost Local AI on Intel AI PC Platforms
-
Qualcomm Reveals Snapdragon C with Advanced On-Device AI Engine
-
Fine-tuning an LLM to Write Docs Like It's 1995
-
Chrome Quietly Downloads 4GB AI Model for Local Processing
-
How to Run LLM Locally Without Falling for the Hype
-
Snapdragon C Specs Revealed: 6nm Process, On-Device AI Engine for Budget Laptops
-
What Apple Knows About AI That Silicon Valley Won't Admit
-
Oracle APEX 26.1 Expands AI Choice with Out-of-the-Box Support for Major AI Providers
-
Snapdragon C Debuts with 6nm Process and Dedicated On-Device AI Engine
-
MediaTek Dimensity 7500 Brings On-Device AI and Enhanced Power Efficiency to Mid-Range Phones
-
Zoho-Backed Netrasemi Launches 12nm AI Chip, Mass Production Begins This Year
-
Apple Doubles Down on On-Device AI at WWDC 2026, Setting Privacy-First Strategy
-
MediaTek Launches Dimensity 8550 4nm SoC with Integrated On-Device AI Focus
-
Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
-
Tweaking Local Language Model Settings with Ollama
-
Alibaba Cloud Joins PyTorch Foundation as Platinum Member
-
MediaTek Dimensity 8550 Shifts Focus to Gemini Nano V3 and On-Device AI on Phones
-
Privacy-Focused Raspberry Pi Zero 2W DIY Security Camera with On-Device AI and End-to-End Encryption
-
The Anatomy of an LLM
-
Local LLM Setup: How to Use RAG and an Embedding Model to Stop Wasting Context
-
llama.cpp GGUF Parser Flaws: Critical Integer Overflow Enables Arbitrary Reads in Every Local AI Stack
-
DeepSeek's Flagship V4 Pro Model Drops to 75% Lower Pricing, Increasing Competitive Pressure on Local Inference Economics
-
Dell Launches 14 Plus Laptop with Intel Core Ultra 9 and 32GB RAM at $1,499.99, Enabling Local Model Inference
-
Samsung's Exynos 2800 Brings HBM Memory to Mobile AI, Enabling Faster Local Model Inference
-
Developer Switches from LM Studio to llama.cpp, Reports No Performance Downgrade
-
Anker Soundcore Liberty 5 Pro Earbuds Feature Dedicated On-Device AI Chip with Touch Screen
-
Apple's 2026 AI Strategy Prioritizes On-Device Model Deployment
-
Maker Demonstrates Portable AI with Suitcase-Integrated Jetson Orin Setup
-
Users Report Superior Performance Switching from LM Studio to llama.cpp
-
Show HN: An Open-Source Interactive AI Engineering Syllabus (1,100 Papers)
-
Qualcomm's AI-Device Strategy Reflects Growing Market Momentum in On-Device Intelligence
-
Why AI Hardware Is a Chip Layer Problem
-
Redditor Successfully Runs 1 Trillion Parameter LLM Using Cheap Intel Optane DIMMs
-
How to Self-Host LibreChat with Docker
-
AMD Unveils Ryzen AI Halo Developer Platform for On-Device AI Workloads
-
New 8B Local LLM Design Marks Biggest Shift Since DeepSeek R1
-
Self-Hosting LLMs Reveals Local AI Has a Friction Problem, Not a Quality Problem
-
Google Makes Gemini 3.5 Flash the Default AI Model for Billions of Users
-
Show HN: Interactive and Stylized AI Chat Chrome Extension
-
A/B Tested Gemini 3.1 Pro vs. Claude Opus 4.6 – Usage Quota and Quality Comparison
-
110 Tokens/Second on RTX 4070 Super with Qwen 3.6 35B
-
The Brain vs. Deep Learning Part I: Computational Complexity Analysis
-
Benchmarking a Portable AI Workstation: Lenovo ThinkPad P16 Gen 3, Part 2
-
Meta Plans Agentic AI on Smartphones and Wearables by 2026
-
On-Device AI to Be in 80% of Wearables by 2032
-
Running Large Language Models on Single-Board Computer Clusters: Creative Edge Deployment
-
Samsung's Exynos 2800 Brings Significant On-Device AI Capabilities
-
Ansede-static: Offline SAST Tool Demonstrates Value of Local AI Tools
-
The Time Bomb Went Off: AI's All-You-Can-Eat Era Just Ended in Real Time
-
The AI Layoff Receipts: Market Consolidation Accelerates Open-Source Model Adoption
-
Local LLMs Enable Intelligent Smart Camera Control Without Cloud Dependency
-
Linux 7.1-rc4 Released: Kernel Updates Relevant to Local LLM Inference
-
A Cheap Fix That Saves the AI $400M Dollars a Year and Brings 4B People Online
-
Local LLM Takes Control of Video Doorbell—The Future of Smart Cameras
-
Google Limits Gemini Intelligence to New Flagships—Hardware Requirements for Local Deployment
-
Towards Local Plug-and-Play AI
-
MegaTrain: Full Precision Training of 100B+ Parameter LLMs on a Single GPU
-
Offline Voice-to-Text and AI Keyboard App for Local Processing
-
Local LLM Integration Enables Replacement of Paid Subscription Services
-
DwarfStar 4: Native Inference Engine Optimized for DeepSeek V4 Flash
-
Orthrus Reshapes Economics of Local AI Inference with New Optimization Approach
-
AI/ML Benchmark Tool for Local LLM Inference and XGBoost Training
-
Arm and Google Collaborate on On-Device AI Optimization Techniques
-
Show HN: Find the best local LLM for your hardware, ranked by benchmarks
-
Chrome Automatically Downloads 4GB AI Model for Local Processing
-
Running AI Models Locally on M4 Processors with 24GB Memory
-
Running Local AI LLMs on Mini PCs Without NVIDIA GPUs
-
Running a Local LLM on a 12-Year-Old Raspberry Pi
-
How I Used a Local LLM to Organize the Store on My NAS
-
I Stopped Paying for ChatGPT and Switched to a Local LLM That Runs on My Laptop
-
BT Explainer: Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Mainline Linux 6.12 on Annapurna Labs Alpine V2 (Ubiquiti UNVR, UDM-Pro)
-
Running a Local LLM on a 12-Year-Old Raspberry Pi: Practical Edge Inference
-
DFlash Speculative Decoding Delivers 8.5x Speed Improvement for LLM Inference
-
Ollama Out-of-Bounds Read Vulnerability Allows Remote Process Memory Leak
-
One LM Studio Setting Makes Local LLMs Competitive With Cloud Models
-
DistillFast: AI Cost Optimization Tool for Model Efficiency
-
Small On-Device AI Model Beats Claude Sonnet 4.5 and GPT-5
-
Chrome's On-Device AI Features Consuming 4GB of Storage for Gemini Nano
-
How I Used a Local LLM to Organize the Store on My NAS
-
Lemonade Gives AMD Startups a Wider Path to Local Inference
-
How to Run LLMs Locally on Your Laptop for Free: A Beginner's Guide
-
Google Releases Gemma 4 Multi-Token Prediction Drafters To Accelerate AI Inference
-
Local LLM Rewrites Resume Better Than ChatGPT, and It's Not Even Close
-
Perplexity Brings On-Device AI Workflow to Macs with 'Personal Computer' Feature
-
Nota AI Partners with Mobilint to Accelerate On-Device AI on Domestic NPU Infrastructure
-
Building a Local LLM News Brief Taught Me the Real Problem Wasn't the Sources, It Was the Apps
-
Enterprise Workplace AI: Questions on Standardizing Local vs Cloud Models
-
Improving Code Quality with Local Claude and Codex Models
-
I Replaced ChatGPT and Claude With This Powerful Local LLM and Saved Over $20 a Month While Gaining Full Control
-
A 49-Line Physics Classifier That Beats kNN on 76% of Benchmarks
-
5 Things I Wish Someone Had Told Me Before I Tried Self-Hosting a Local LLM
-
Major Smartphone Brands Introduce Advanced On-Device AI Features
-
Anker's Thus Chip Puts AI On-Device, Promising Faster Responses And Better Privacy
-
I Put a Local LLM on My Phone and Stopped Needing Cloud AI for Most Tasks
-
Running a Serious AI Model on a Consumer GPU Just Got Easier and That Matters More Than the Benchmark
-
NIST's CAISI Evaluation of DeepSeek V4 Pro Finds It On Par with GPT-5
-
Anker's New 'Thus' Chip Brings 150x AI Power to Earbuds
-
Single-Command Setup Tool Automates Claude AI Workstation Configuration
-
New Open-Source Tool Automatically Matches Local LLMs to Your PC Hardware
-
Building a Raspberry Pi-Based Local LLM Server for Remote Access
-
Building a Remote-Accessible Local LLM Server on Raspberry Pi
-
Running Capable Local LLMs Without Expensive GPU Hardware
-
IBM Introduces Granite 4.1 Family of Models for Local Deployment
-
Google's Gemma 4 Brings Powerful AI Capabilities to Phones and Laptops
-
How Much "Brain Damage" Can an LLM Tolerate?
-
Estimating Black-Box LLM Parameter Counts via Factual Capacity
-
Wipeout Clone Runs Native on ESP32-S3, Pushing Edge Hardware to Its Limits
-
Intel N150 Mini PC Runs Local LLM for Home Assistant
-
NVIDIA Nemotron 3 Nano Omni Powers Multimodal Agent Reasoning in a Single Efficient Open Model
-
Grokfeed: Terminal Feed Reader for HN, Reddit, and Lobste.rs Using Claude Code
-
Picking Your First Local LLM Is Easier Than the Internet Makes It Sound
-
Stop Guessing: Open-Source Tool Predicts Which Local LLMs Run on Your PC
-
Building a Local AI Stack: Five Docker Containers to Replace ChatGPT Subscriptions
-
Google's Gemma 4: Powerful AI Models Optimized for Your Phone and Laptop
-
Economic Implications of AI Adoption: Why Local Deployment Matters for Cost Control
-
Why the Same LLM Gives Different Answers in Different Environments
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Linux Crushes Windows on llama.cpp Inference by Double Digits
-
Show HN: Phonetic Formatter – Offline English Text to IPA on iPhone and iPad
-
Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Run a Local LLM Server on Raspberry Pi with Remote Access Capabilities
-
Google's Gemma 4 Brings Powerful On-Device AI to Phones and Laptops
-
How to Make Sense of AI
-
Building Real-World On-Device AI with LiteRT and NPU
-
Netherlands Reaches Deal to Cut Reliance on U.S. Cloud Tech
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results
-
AI Agent Designs a RISC-V CPU Core from Scratch
-
Anker Unveils 'Thus' Chip to Bring On-Device AI Across Product Line
-
Llama 4 Scout on MLX: The Complete Apple Silicon Guide (2026)
-
10GB VRAM Local LLM: The Complete Setup Guide (2026)
-
Show HN: We built an OCR server that can process 270 dense images/s on a 5090
-
Externalization in LLM Agents: Unified Review of Memory and Harness Engineering
-
Llama.cpp's Auto Fit Feature Quietly Reshapes Local AI Inference on Consumer Hardware
-
The Open-Source AI Ecosystem Keeps Treating llama.cpp Like a Second-Class Citizen
-
Malicious GGUF Models Could Trigger Remote Code Execution on SGLang Servers
-
Intel Extends AI PC Reach With New Core Ultra Series 3 Launch
-
Running DeepSeek R1 Locally: Your Complete Setup Guide
-
Controlling the Secondary Fan on Minisforum AI Pro HX 370
-
Minisforum Launches N5 Max AI NAS with OpenClaw
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
115 TOPS in 0.67L: CHUWI AuBox X Packs On-Device AI Power Into a Palm-Sized Mini PC
-
Unweight: Lossless MLP Weight Compression for LLM Inference
-
Laimark – 8B LLM That Self-Improves on Consumer GPUs
-
Project Glasswing and the ASF: Open-Source's Chance to Win the AI Era
-
Building a Voice AI Wearable in a Casio F91W with Whisper and BLE
-
Prefill Is Compute-Bound, Decode Is Memory-Bound: Optimizing GPU Utilization for LLM Inference
-
Bonsai 1.7B in the Browser: A 290MB 1-bit LLM on WebGPU
-
Running Gemma 4 on an iPhone 13 Pro
-
SigMap – Shrink AI Coding Context 97% with Auto-Scaling Token Budget
-
MiniMax M2.7 GGUF Investigation Reveals NaN Issues Affecting 21-38% of Hugging Face Conversions
-
Fine-Tuned Qwen3.5-0.8B for OCR Outperforms Previous 2B Release
-
Minisforum N5 MAX AI NAS Delivers 126 TOPS with 200TB Storage for Local LLM Workloads
-
Sovereign AI: Why the Next GPT Will Be Born in Our Living Rooms
-
Qwen 3.5 Small – On-Device Multimodal Models Released
-
MiniMax M2.7 Achieves SOTA Performance Under 64GB on Mac with TQ Quantization
-
Show HN: SkillCompass – Open-Source Quality Evaluator for Your AI Skills
-
Speculative Decoding Achieves 29% Speed Boost for Gemma-4 31B
-
Qwen3 Audio and Vision Support Now Available in llama.cpp
-
MiniMax-M2.7 Delivers Exceptional Performance on Consumer Hardware
-
Learn LLM Internals
-
Researchers Achieve 1-Bit Quantization of OLMo-3 7B Using Distillation
-
The Best Local AI Model for Home Assistant Isn't Always the Biggest One
-
Universal Knowledge Store and Grounding Layer for AI Reasoning Engines
-
MiniMax M2.7 Is Now Open Source
-
Unsloth Completes Comprehensive MiniMax M2.7 GGUF Quantization Suite
-
On-Device AI: Achieving Powerful AI Capabilities Without Internet Connectivity
-
MiniMax M2.7 Released: New Model Available for Local Deployment
-
Google's Gemma 4 Brings Free Agentic AI to Your Phone With Zero Data Leaving the Device
-
Users Report Significant Performance Improvements After Migrating from Ollama to llama.cpp
-
Critical Unsloth Gemma-4 Chat Template Updates for Tool Calling
-
Google's Gemini Nano 4 Offers Faster, Smarter Local Inference Capabilities
-
AI PC Market Projected to Reach $235B by 2032, Driven by On-Device Computing Adoption
-
Intel Arc Pro B70 32GB Achieves 12 Tokens/Sec on Qwen 3.5-27B
-
Gemma 4 31B vs Qwen 3.5 27B: Comprehensive Long Context Benchmark
-
LLM Wiki v2: Extended Knowledge Base for LLM Practitioners
-
5 Open-Source Projects Running Transformers on CPUs to GPUs in Pure Java
-
CarryAI's Serverless Vision-Language Models Enable On-Device Multimodal AI
-
Energy Consumption: The Final Frontier for AI and Local Inference
-
Building Offline AI Companions on Severely Constrained Hardware (8GB RAM)
-
Gemini-CLI, Llama.cpp, and Qwen3.5 Running on NVIDIA Jetson TK1
-
Run Qwen3.5 on an Old Laptop: A Lightweight Local Agentic AI Setup Guide
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results — and It Wasn't About the Parameters
-
Intel Releases OpenVINO 2026.1 With Backend For Llama.cpp, New Hardware Support
-
Gemma 4 Support Stabilized in Llama.cpp
-
Gemma 4 GGUF Models Updated with Critical Quantization Fixes
-
EXAONE 4.5 33B Model Released with Multiple Quantization Formats
-
Speculative Decoding Made My Local LLM Actually Usable
-
Running a 1.7B Parameters LLM on an Apple Watch
-
Google AI Edge Gallery Showcases Offline Inference with Gemma 4
-
Comprehensive Benchmark: 37 LLMs Tested on MacBook Air M5 With Open-Source Tool
-
Quansloth Using Google's Turboquant Breaks the VRAM Wall for Local LLMs
-
Gemma 4 26B Achieves Impressive Local Performance With Proper Configuration
-
TurboQuant-Optimized llama.cpp Fork Delivers GFX906 GPU Acceleration
-
CricketBrain: Neuromorphic Signal Processor in Rust (0.175us/step, 944 bytes)
-
TurboQuant in Llama.cpp Achieves 6X Smaller KV Cache
-
Show HN: Lightweight LLM Tracing Tool with CLI
-
Google AI Edge Gallery Tops App Store Charts with On-Device Gemma 4
-
Gemma 4 31B Achieves Exceptional Performance on Local Hardware
-
Quantization Strategy Comparison: Balancing Quality and Speed on Consumer Laptops
-
Context Window Optimization: Extending Gemma 4 Context Length Through Efficient Projection Quantization
-
GPU Memory for LLM Inference (Part 1)
-
Qualcomm Snapdragon Innovations Enable Advanced On-Device AI for Wearables
-
Qwen 3.6 Free Model Available via OpenRouter
-
GMKtec NucBox K17 Launches with 97 TOPS AI Performance for Local Inference
-
Gemma 4 31B Achieves Third Place on FoodTruck Bench, Beating Larger Models
-
Gemma 4 26B MoE Emerges as Optimal All-Around Local Model for Consumer Hardware
-
Unpaved: Audit Toolkit for AI Developer Tool Bias in Global South Contexts
-
Qwen 3.5 397B Reduced to 35% Parameters With Usable Quality on 96GB GPU
-
DGX Spark Hardware Limitations: Missing NVFP4 Support Undermines Local AI Value Proposition
-
Apple Research Shows Self-Distillation Significantly Improves Local Code Generation
-
Ollama Gets Blazing Fast on Macs with Full MLX Support and 2× Speedups
-
Mixed Precision Quantization on MLX with TurboQuant Implementation
-
Nex Life Logger: Local Activity Tracker with AI Agent Integration
-
Gemma 4 26B A4B Outperforms Qwen 3.5 35B on Apple Silicon
-
Gemma 4 2B Successfully Runs on Raspberry Pi 5
-
Google Gemma 4 Released with GGUF Quantizations
-
Gemma 4 on Arm: Optimized On-Device AI for Mobile and Edge Deployment
-
Qwen 3.6-Plus Released
-
Bonsai 1-Bit Models Deliver Exceptional Local Inference Performance
-
TurboQuant Enables Qwen 3.5-27B on 16GB Consumer GPUs
-
ByteShape Releases Qwen 3.5 9B Quantisations with Hardware-Matched Tuning Guide
-
Satcove – Query 5 AI Models Simultaneously and Get Structured Verdicts
-
Llama.cpp Merging TurboQuant Lite (attn-rot) with Major Performance Gains
-
PrismML Announces 1-Bit Bonsai: First Commercially Viable 1-Bit LLMs
-
Local AI Ecosystem Extends Far Beyond Ollama
-
Does RAG Help AI Coding Tools?
-
Local AI didn't replace my subscriptions, but it did take over these 6 tasks
-
Running AI on a Raspberry Pi, Part 2: Running AI on a Pi in Under 5 minutes
-
Intel's $949 GPU has 32GB of VRAM for local AI, but the software is why Nvidia keeps winning
-
Ollama Launches Pi: The Minimal Coding Agent That Powers OpenClaw Is Now Yours to Customize
-
Select the Right Hardware for Your Local LLM Deployment with This Online Guide
-
Samsung Launches Galaxy Book6 Series in India with NVIDIA RTX 5070 Graphics and On-Device AI
-
ESP32-S31: 320MHz 2-Core Microcontroller with 512KB SRAM and Networking
-
Google's TurboQuant Shows Memory Constraints Remain Critical for Local LLM Inference
-
DaVinci-MagiHuman: Open-Source AI Model for Realistic Video Generation
-
TurboQuant: Understanding the Quantization Breakthrough
-
OLED Emerges as the Display Standard for Energy-Efficient AI Systems
-
Mixed KV Cache Quantization: Performance Risks and Pitfalls
-
Introduction to Nyreth v1.0
-
Samsung Galaxy Book6 Series Brings Intel Core Ultra Chips for On-Device LLM Inference
-
TurboQuant KV Cache Compression Achieves 22.8% Faster Decoding at 32K Context
-
Qwen3 512k Context via TurboQuant on Mac mini
-
Apple Gets Full Gemini Access and Uses Distillation to Build Lightweight On-Device AI
-
This Wearable Runs an On-Device AI With 2-Week Battery Life
-
Coding Implementation to Run Qwen3.5 Reasoning Models Distilled With Claude-Style Thinking Using GGUF and 4-Bit Quantization
-
Hold on to Your Hardware: Implications for Local LLM Deployment
-
TurboQuant Benchmarked in Llama.cpp: Google's Extreme Compression Research Tested in Practice
-
RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
-
Qwen 3.5 27B Achieves 1.1M Tokens/Second on B200 GPUs with Optimized vLLM Config
-
Quantization Reveals Outliers Impacting LLM Accuracy
-
RF-DETR Nano and YOLO26 Enable On-Device Object Detection on Smartphones
-
Samsung Galaxy A37 and A57 5G Launch with On-Device AI Capabilities in India
-
Show HN: Beforeyouship – Pre-Build Tool to Estimate LLM Cost
-
Intel Launches Arc Pro B70/B65 with 32GB VRAM for Local AI Inference
-
Google's TurboQuant: The Unsexy AI Breakthrough Worth Watching
-
NVIDIA Releases GPT-OSS-Puzzle-88B, a Deployment-Optimized Model
-
Apple Plans Slimmed-Down Gemini Models for Local iPhone AI Features
-
Private Brain LLM Setup on Windows PC Eliminates Need for Paid Cloud Services
-
OmniCoder v2 Released: Improved Code Generation for Local Deployment
-
Researcher Successfully Runs Local LLMs on Legacy "Dead" GPU With Surprising Results
-
Google TurboQuant: Extreme Compression for Local LLM Deployment
-
Running an Open-Weight LLM Locally on an Apple Watch
-
.APKs Are Just .ZIPs: Semi-Legally Hacking Software for Orphaned Hardware
-
Ultra-Large 400B-Class LLM Runs on iPhone in Test
-
Four Raspberry Pi AI Tools You Can Try This Week Beyond OpenClaw
-
LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language
-
KV Cache Quantization Levels Benchmarked on SWE-bench: Practical Trade-offs for Local Inference
-
FOMOE: Running 397B Parameter Qwen3.5 MoE at 5-9 tok/s on $2,100 Desktop Hardware
-
FlashAttention-4 Delivers 2.7x Faster Inference with 1613 TFLOPs/s on Blackwell GPUs
-
Open-Source Tool Helps Determine Which Local LLMs Run on Your PC
-
Chinese LLM Ecosystem Landscape: ByteDance Doubao, Alibaba, and Open-Source Competition
-
Qt 6.11 Released with Enhanced Cross-Platform Deployment Capabilities
-
Korea to Deploy Domestic AI Chips in Smart Cities as NPU Trials Scale Up
-
How to Build a Self-Hosted AI Server with LM Studio: Step-by-Step Guide
-
Powerful AI Search Engine Built on Single GeForce RTX 5090
-
BrowserOS 0.44.0 Release: Advances in Local AI Integration for Web-Based Applications
-
Setting Up a Private AI Brain on Windows: Complete Guide to Local LLM Deployment
-
Ditching Paid AI Services: Building Self-Hosted LLM Solutions as ChatGPT, Claude, and Gemini Alternatives
-
Qwen 3.5 122B Uncensored (Aggressive) Released with New K_P Quantisations
-
Why You Should Use Both ChatGPT and Local LLMs: A Practical Hybrid Approach
-
AI Playground for Developers Built in Vite and Python
-
Rust Project Perspectives on AI
-
ik_llama.cpp Fork Delivers 26x Faster Prompt Processing on Qwen 3.5 27B
-
Running an AI Agent on a 448KB RAM Microcontroller
-
Qwen 3.5 397B emerges as top-performing local coding model
-
Apple M5 Max 128GB real-world performance benchmarks for local inference
-
ASUS ExpertCenter PN55 Mini PC Combines AMD AI CPU and 55 TOPS NPU
-
Qwen 3.5 Emerges as Top Performer for Local Deployment with Extensive Quantization Options
-
Repurpose Old GPUs as Dedicated AI Inference Accelerators
-
Ultra-Compact 28M Parameter Models Show Promise for Specialized Domain Tasks
-
Community Converges on Optimal KV Cache Quantization Strategies for Qwen 3.5 Models
-
Multiverse Computing Targets On-Device AI With Compressed Models and New API Portal
-
Tether's QVAC Introduces Cross-Platform Bitnet LoRA Framework for On-Device AI Training
-
You're Using Your Local LLM Wrong If You're Prompting It Like a Cloud LLM
-
Hugging Face Releases One-Liner for Automatic Hardware Detection and Model Selection
-
Browser-Based Transcription Tools
-
Unsloth Studio: Open-Source Web UI for Training and Running LLMs Locally
-
Snapdragon 8 Elite Gen 5 Hands the Galaxy S26 the AI Upgrade We've Been Waiting For
-
Run LLMs Locally with Llama.cpp
-
I Ran Local LLMs on a 'Dead' GPU, and the Results Surprised Me
-
Qwen 3.5 4B Outperforms Nvidia Nemotron 3 4B in Local Benchmarks
-
Mistral Small 4 119B Released with NVFP4 Quantisation Support
-
Mistral Releases Small 4 Open-Source Model Under Apache 2.0
-
Researcher Discovers Universal "Danger Zone" in Transformer Model Architecture at 50% Depth
-
Kimi Introduces Attention Residuals: 1.25x Compute Performance at <2% Overhead
-
OpenClaw Isn't the Only Raspberry Pi AI Tool—Here Are 4 Others You Can Try This Week
-
OmniCoder-9B: Efficient Coding Model for 8GB GPUs
-
AMD Declares 'AI on the PC Has Crossed an Important Line' – Agent Computers as Next Breakthrough
-
Nota Added to Three Technology and Growth ETFs in a Row – Market Recognition for AI Efficiency
-
I made Karpathy's Autoresearch work on CPU
-
Cicikus v3 Prometheus 4.4B – An Experimental Franken-Merge for Edge Reasoning
-
Running Qwen3.5-27B Across Multiple GPUs Over LAN Achieves Practical Speed for Local Inference
-
Two Local Models Prove Competitive Enough to Replace ChatGPT, Gemini, and Copilot
-
India's Mobile-First AI Strategy Could Accelerate Local Inference Adoption in Emerging Markets
-
Qwen3.5-397B Achieves 282 tok/s on 4x RTX PRO 6000 Blackwell Through Custom CUTLASS Kernel
-
Best Local LLM Models 2026: Developer Comparison
-
Intel Updates LLM-Scaler-vLLM With Support For More Qwen3/3.5 Models
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Qwodel – An Open-Source Unified Pipeline for LLM Quantization
-
Quantization Explained: Q4_K_M vs AWQ vs FP16 for Local LLMs
-
Nvidia Releases Nemotron 3 Super: 120B MoE Model for Local Deployment
-
Comprehensive MoE Backend Benchmarks for Qwen3.5-397B: Real Numbers vs Hype
-
Show HN: Detect When an LLM Silently Changes Behavior for the Same Prompt
-
SK Hynix Completes Qualification for LPDDR6 Memory Optimized for AI Inference
-
NVIDIA Jetson Brings Open Models to Life at the Edge
-
Experiment: 0.8B Model Self-Improvement on MacBook Air Yields Surprising Results
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Qwen 3.5-35B Uncensored GGUF Models Now Available
-
Llama.cpp Celebrates Major Milestone: From Leak to Industry Standard
-
.ispec: Runtime Specification Validation for AI System Consistency
-
FreeBSD 14.4 Released: Implications for Local LLM Deployment
-
Community Survey: AI Content Automation Stacks in 2026
-
Qwen 3.5 Ultra-Compact Models Enable On-Device AI from Watches to Gaming
-
HP OMEN MAX 16 Review: Is Local AI on a Laptop Viable in 2026?
-
Qwen 3.5 Family Benchmark Comparison Shows Strong Performance Across Smaller Models
-
Qwen 3.5 Derestricted Model Available for Local Deployment
-
Qwen 3.5 Small Expands On-Device AI to Phones and IoT with Offline Support
-
Nota AI to Showcase End-to-End On-Device AI Optimization at Embedded World 2026
-
How to Run Your Own Local LLM — 2026 Edition
-
Samsung Opens Registration for Vision AI QLED and OLED Television Integration
-
Snapdragon Wear Elite Unveiled at MWC 2026, Advancing Wearable AI Inference
-
HP Refreshes Lineup with AI-Focused Workstations
-
Qwen 3.5 27B Achieves Strong Local Inference Performance
-
Apple Launches MacBook Neo with A18 Pro Chip for Affordable Local AI Inference
-
Student Researcher Achieves 42x Model Compression Through Novel Architecture
-
IBM Granite 4.0 1B Speech Model Released for Multilingual Speech Recognition
-
Windows 11 Notepad Gets On-Device AI Text Generation Without Subscription
-
Mojo: Creating a Programming Language for an AI World with Chris Lattner
-
OPPO and MediaTek Highlight On-Device AI Innovations at MWC 2026
-
Final Qwen3.5 Unsloth GGUF Update with Improved Size/Quality Tradeoffs
-
Alibaba Releases Qwen 3.5 AI Model with On-Device AI Support
-
Show HN: TLDR – Free Chrome Extension for AI-Powered Article Summarization
-
MediaTek Advances Omni Model for Efficient Smartphone Inference
-
Apple Unveils MacBook Pro with M5 Pro and M5 Max Featuring On-Device AI
-
Kakao Launches Kanana AI for On-Device Schedule and Recommendation Management
-
Unity Showcases Manufacturing AI Workflow at Smart Factory Expo
-
Apple Unveils MacBook Pro With M5 Pro and M5 Max for On-Device AI
-
Quantifying Cost Savings with Local LLMs for Development
-
OpenWrt 25.12.0 – Stable Release
-
On-Device AI Laptop Lineups Become Standard Across Major Manufacturers
-
Qwen 3.5-27B Q4 Quantization Comparison and Analysis
-
Qualcomm Snapdragon Wear Elite Brings On-Device AI to Smartwatches
-
Alibaba's Qwen 3.5 Small Model Runs Directly on iPhone 17
-
Qwen 3.5 27B Achieves 100+ Tokens/s Decode on Dual RTX 3090s with 170K Context
-
Qualcomm Launches Snapdragon Wear Elite for On-Device AI on Wearables
-
Local LLM Performance Improvements: A Year of Progress Since DeepSeek R1 Moment
-
HP ZBook Ultra 14 G1a Workstation Reclaims Local AI Workflows for Professionals
-
Running Local AI Models on Mac Studio 128GB: 4B, 20B & 120B Tested
-
Critical: Qwen 3.5 Requires BF16 KV Cache, Not FP16 for Accurate Inference
-
Qwen 3.5-35B-A3B Emerges as Efficient Daily Driver, Replacing 120B Models
-
Apple Intelligence, Galaxy AI, Gemini: Why Your AI-Powered Phone Is Worth Repairing
-
How to Run High-Performance LLMs Locally on the Arduino UNO Q
-
Qwen3.5-35B Successfully Runs on Raspberry Pi 5 at 3+ Tokens/Second
-
Galaxy S26 Debuts AI-Powered Scam Detection in Bold Security Push
-
Arduino, Qualcomm Bring On-Device AI and Robotics Learning to Indian School Systems
-
Unsloth Dynamic 2.0 GGUFs
-
Accuracy vs. Speed in Local LLMs: Finding Your Sweet Spot
-
Qwen3.5-35B Unsloth Dynamic GGUFs Achieve SOTA Across Nearly All Quantisation Levels
-
Qwen3.5-35B RTX 5080 Experiments Confirm KV q8_0 as Free Lunch, Q4_K_M Remains Optimal
-
Meta Reveals AI-Packed Smartwatch In 2026 – Why Wearables Shift Now
-
On-Device AI in Mobile Apps: What Should Run on the Phone vs the Cloud (A 2026 Decision Guide)
-
Arduino and Qualcomm Bring On-Device AI Learning to Indian Schools
-
Android Phones Are Getting Smarter Without Internet — On-Device AI as the Next Shift
-
Android Phones Are Getting Smarter Without Internet — Here's Why On-Device AI Is the Next Big Shift
-
Arduino, Qualcomm Bring On-Device AI and Robotics Learning to Indian School Systems
-
5 Useful Docker Containers for Agentic Developers
-
Snapdragon 8 Elite Gen 5 for Galaxy Official: 5 Key Improvements that Push the Boundaries
-
Running LLMs on Raspberry Pi and Edge Devices: A Practical Guide
-
Qwen3.5 122B Achieves 25 tok/s on 72GB VRAM Setup
-
Qwen 3.5 MoE Delivers 100K Context Window at 40+ TPS on RTX 5060 Ti
-
New Era of On-Device AI Driven by High-Speed UFS 5.0 Storage
-
Qwen3.5 Series Releases Comprehensive Model Lineup Across All Tiers
-
Qwen3.5-27B Identified as Sweet Spot for Mid-Range Local Deployment
-
PyTorch Foundation Announces New Members as Agentic AI Demand Grows
-
Mirai Announces $10M to Advance On-Device AI Performance for Consumer Devices
-
Show HN: 100% LLM Accuracy–No Fine-Tuning, JSON Only
-
How AI is Redefining Price and Performance in Modern Laptops
-
Advanced Quantization Techniques Show Surprising Performance Gains Over Standard Methods
-
What Breaks When AI Agent Frameworks Are Forced Into <1MB RAM and Sub-ms Startup
-
No, Local LLMs Can't Replace ChatGPT or Gemini — I Tried
-
Kioxia Sampling UFS 5.0 Embedded Flash Memory for Next-Generation Mobile Applications
-
Elastic Introduces Best-in-Class Embedding Models for High Performance Semantic Search
-
Show HN: Dypai – Build Backends from Your IDE Using AI and MCP
-
Enhanced Interface Speed Enables High-Performance On-Device AI Features in Smartphones
-
Anthropic Has Never Open-Sourced an LLM: Implications for Local Deployment Strategy
-
Enterprise Infrastructure Guide: Running Local LLMs for 70-150 Developers
-
Future of Mobile AI: What On-Device Intelligence Means for App Developers
-
How Do You Know Which SKILL.md Is Good?
-
nanollama: Open-Source Framework for Training Llama 3 from Scratch with One-Command GGUF Export
-
Qwen3-Code-Next Proves Practical for Local Development: Real-World Coding Tasks on Mac Studio
-
Open-Source llama.cpp Finds Long-Term Home at Hugging Face
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
Custom Portable Workstation Optimized for Local AI Inference Builds
-
Open-Source Framework Achieves Gemini 3 Deep Think Level Performance Through Local Model Scaffolding
-
Nvidia Could Launch Its First Laptops With Its Own Processors
-
The Complete Stack for Local Autonomous Agents: From GGML to Orchestration
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
DietPi Released a New Version v10.1
-
GGML Joins Hugging Face: What This Means for Local Model Optimization
-
At India AI Impact Summit, Intel Showcases AI PCs and Cost-Efficient Frugal AI
-
CPU-Trained Language Model Outperforms GPU Baseline After 40 Hours
-
AI PCs Explained: 7 Critical Truths About NPUs and Privacy
-
Ouro 2.6B Thinking Model GGUFs Released with Q8_0 and Q4_K_M Quantization
-
I Thought I Needed a GPU to Run AI Until I Learned About These Models
-
Strix Halo Performance Benchmarks: Minimax M2.5, Step 3.5 Flash, Qwen3 Coder
-
GGML.AI Acquired by Hugging Face
-
Taalas Etches AI Models onto Transistors to Rocket Boost Inference
-
Qwen3 Coder Next Remains Effective at Aggressive Quantization Levels
-
[Release] Ouro-2.6B-Thinking: ByteDance's Recurrent Model Now Runnable Locally
-
Kitten TTS V0.8 Released: New State-of-the-Art Super-Tiny TTS Model Under 25 MB
-
VaultAI – 42 AI Models on a Portable SSD, Works Offline for $399
-
Mirai Secures $10M to Optimize On-Device AI Amid Cloud Cost Surge
-
The Path to Ubiquitous AI (17k tokens/sec)
-
Enhanced Quantization Visualization Methods for Understanding LLM Compression Trade-offs
-
Sarvam Brings AI to Feature Phones, Cars, and Smart Glasses
-
Running Local LLMs and VLMs on Arduino UNO Q with yzma
-
Hardware Economics Shift: DDR5 RDIMM Pricing Now Comparable to GPUs for Local Inference
-
Local Vision-Language Models for Document OCR and PII Detection in Privacy-Critical Workflows
-
Qualcomm Ventures Positions India as Blueprint for Affordable On-Device AI Infrastructure
-
Same INT8 Model Shows 93% to 71% Accuracy Variance Across Snapdragon Chipsets
-
Ask HN: What is the best bang for buck budget AI coding?
-
Qwen 3.5-397B-A17B Now Available for Local Inference with Aggressive Quantisation
-
Alibaba Unveils Major AI Model Upgrade Ahead of DeepSeek Release
-
NVIDIA's Dynamic Memory Sparsification Cuts LLM Inference Costs by 8x
-
MiniMax-M2.5 230B MoE Model Released with GGUF Support for Local Deployment
-
GPT-OSS 120B Uncensored Model Released in Native MXFP4 Precision
-
MiniMax Releases M2.5 Model with SOTA Coding and Agent Capabilities
-
Ring-1T-2.5 Released with SOTA Deep Thinking Performance
-
GitHub Announces Support for Open Source AI Project Maintainers
-
Running Mistral-7B on Intel NPU Achieves 12.6 Tokens/Second
-
Samsung's REAM: Alternative Model Compression Technique
-
New Header-Only C++ Benchmark Tool for Predictive Models on Raw Binary Streams
-
GLM-5 Released: 744B Parameter MoE Model Targeting Complex Tasks
-
Community Member Builds 144GB VRAM Local LLM Powerhouse