Tagged "model-compression"
-
Quantization-Aware Healing: 4-Bit Models Outperform Full-Precision Originals
-
Google COSMO Leak Reveals Gemini Nano and On-Device AI Skills
-
llama.cpp Build b10581 Adds DSpark Support for Faster Local Inference
-
Liquid AI Releases LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation
-
DeepSeek V4 Flash Shrunk to 57GB for Local macOS Inference with Compiler Generation
-
The Qwen MLX Challenge
-
How an $8 ESP32 S3 Microcontroller Runs a 28.9M Parameter Local LLM
-
Show HN: I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac
-
Qwen 3.8 27B Successfully Runs on 16GB RAM Using LM Studio
-
Meta's Muse Glimmer on ExecuTorch Enables Fast On-Device Agentic AI
-
DeepX's DX-M1 On-Device AI Chip Achieves $13M in Orders
-
Running DeepSeek's 284B LLM on a Laptop: Quantisation and GGUF Optimization
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Local Agent Deployment
-
Chrome's On-Device AI Model Requires 20GB Storage Space
-
Chrome and Edge Browsers Quietly Deploy Up to 20GB AI Models on Windows 11
-
Shrinking an AI Model 86% Doesn't Make It 86% Dumber: Compression Breakthroughs
-
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems
-
LFM2.5-2.6B: On-Device Agentic Model With 128K Context and Tool Calling
-
PrismML's Bonsai 27B Brings On-Device AI to Apple iPhone 17 Pro
-
28.9M-Parameter LLM Runs Locally on ESP32-S3 at 9 Tokens/s
-
Kioxia Is Coming for Samsung and SK Hynix With UFS 5.0 and PCIe 6.0 AI NAND
-
Q4 vs Q6 vs Q8: The Quantization Decision Framework for Local LLMs
-
Your Smartwatch Now Detects a Heart Irregularity in Milliseconds – Without Ever Touching the Cloud
-
Samsung's Newest Foldable Phones Use Google's Gemini Nano 4 On-Device AI Model
-
CliffordNet: All You Need Is Geometric Algebra
-
Gemma 4's Quantized Models Finally Made Local AI Practical in Homelab
-
Deploying 1-Bit Bonsai-27B with PrismML and llama.cpp for Local Inference
-
From Foldables to Smart Glasses, Samsung's Galaxy AI Push Moves Beyond the Cloud
-
Nota AI Joins AMD Robotics Partner Network to Expand On-Device AI Optimisation
-
Multiverse Computing's CompactifAI Models Now Fully Compatible with Intel Xeon 6 Processors
-
SK hynix 3D-Stacked DRAM-on-Logic Architecture Could Solve On-Device AI Memory Constraints
-
Sunday Reboot: Shrinking Models and an On-Device AI Future
-
AI Data Center Power Constraints Are the Real 2026 Bottleneck
-
Nubia Announces AI Agent Smartphone with On-Device AI Processing
-
Apple in Early Talks With PrismML on AI Compression Tech
-
Apple in Talks with PrismML to Shrink AI Models 15x for iPhone Deployment
-
Apple Boosts On-Device AI, Partners With PrismML to Enable Running Large Models Locally on iPhone
-
Don't Sleep on BitNet (2025)
-
Google Pixel Implements Local AI for Screenshot Analysis With Privacy Controls
-
Edge AI Brings On-Device Intelligence and Health Monitoring to Smartwatches
-
Study: Cerebellum Helps AI Ignore the Ordinary for More Efficient Computing
-
Apple Explores Running Larger AI Models on iPhone with On-Device Compression
-
Edge AI Smartwatch Shipments Jump 70% as Apple Leads Health-Focused Boom
-
Edge AI Transformation Coming to Creative Production Workflows
-
Compressor V2: Three Compression Layers for 50% LLM Agent Cost Cut
-
Local LLM Performance Gap With Frontier Models Smaller Than Expected
-
Google's Gemma AI Runs Locally on a $300 Mini PC, and It Replaced ChatGPT for More Than Expected
-
TriAttention Solves KV Cache Memory Bottleneck in Local LLM Inference
-
ORA: Smaller Models. Same Intelligence
-
Why Small Local AI Models Get More Use Than Claude or Gemini
-
Xiaomi vs Huawei On-Device AI: Decoding the AI Strategies of 8 Major Smartphone Giants
-
Gaming PC vs Phone Local LLM Deployment: Only One Remains in Daily Use
-
Genesis AI Launches Eno General-Purpose Robot with Embedded AI
-
Tensordyne Napier AI Processor Announced with Logarithmic Math
-
Ask HN: What Problem Did AI Create at Your Company That Didn't Exist Before?
-
Brilliant Labs Halo: Open-Source AI Glasses for On-Device Intelligence
-
Google Releases Gemma 4 QAT Models with Reduced Memory Requirements for Mobile and Laptop Deployment
-
A Cinematic Landing-Page Hero for 80 Cents (GPT Image 2 and Veo 3.1)
-
Tether AI Upgrades QVAC SDK With TurboQuant for Data Center-Sized Memory on Everyday Devices
-
Mistral AI Launches Mistral Vibe
-
DeepSeek's Flagship V4 Pro Model Drops to 75% Lower Pricing, Increasing Competitive Pressure on Local Inference Economics
-
Apple's 2026 AI Strategy Prioritizes On-Device Model Deployment
-
Maker Demonstrates Portable AI with Suitcase-Integrated Jetson Orin Setup
-
The Brain vs. Deep Learning Part I: Computational Complexity Analysis
-
Meta Plans Agentic AI on Smartphones and Wearables by 2026
-
Google Tensor SDK Beta with LiteRT Enables Efficient On-Device AI
-
On-Device AI to Be in 80% of Wearables by 2032
-
Local LLMs Enable Intelligent Smart Camera Control Without Cloud Dependency
-
Google Limits Gemini Intelligence to New Flagships—Hardware Requirements for Local Deployment
-
MegaTrain: Full Precision Training of 100B+ Parameter LLMs on a Single GPU
-
Chrome Automatically Downloads 4GB AI Model for Local Processing
-
Running a Local LLM on a 12-Year-Old Raspberry Pi
-
DistillFast: AI Cost Optimization Tool for Model Efficiency
-
Chrome's On-Device AI Features Consuming 4GB of Storage for Gemini Nano
-
Perplexity Brings On-Device AI Workflow to Macs with 'Personal Computer' Feature
-
Anker's Thus Chip Puts AI On-Device, Promising Faster Responses And Better Privacy
-
Building a Raspberry Pi-Based Local LLM Server for Remote Access
-
How Much "Brain Damage" Can an LLM Tolerate?
-
Google's Gemma 4: Powerful AI Models Optimized for Your Phone and Laptop
-
Building Real-World On-Device AI with LiteRT and NPU
-
Anker Unveils 'Thus' Chip to Bring On-Device AI Across Product Line
-
10GB VRAM Local LLM: The Complete Setup Guide (2026)
-
Unweight: Lossless MLP Weight Compression for LLM Inference
-
Bonsai 1.7B in the Browser: A 290MB 1-bit LLM on WebGPU
-
SigMap – Shrink AI Coding Context 97% with Auto-Scaling Token Budget
-
Researchers Achieve 1-Bit Quantization of OLMo-3 7B Using Distillation
-
On-Device AI: Achieving Powerful AI Capabilities Without Internet Connectivity
-
CarryAI's Serverless Vision-Language Models Enable On-Device Multimodal AI
-
Quansloth Using Google's Turboquant Breaks the VRAM Wall for Local LLMs
-
CricketBrain: Neuromorphic Signal Processor in Rust (0.175us/step, 944 bytes)
-
Google AI Edge Gallery Tops App Store Charts with On-Device Gemma 4
-
Quantization Strategy Comparison: Balancing Quality and Speed on Consumer Laptops
-
Qwen 3.5 397B Reduced to 35% Parameters With Usable Quality on 96GB GPU
-
Mixed Precision Quantization on MLX with TurboQuant Implementation
-
Bonsai 1-Bit Models Deliver Exceptional Local Inference Performance
-
TurboQuant Enables Qwen 3.5-27B on 16GB Consumer GPUs
-
Claw64 – Full Agentic Loop in <4KB on Commodore 64
-
Google's TurboQuant Shows Memory Constraints Remain Critical for Local LLM Inference
-
TurboQuant: Understanding the Quantization Breakthrough
-
CERN Embeds Tiny AI Models in Silicon Chips for Real-Time LHC Data Filtering
-
Apple Gets Full Gemini Access and Uses Distillation to Build Lightweight On-Device AI
-
Coding Implementation to Run Qwen3.5 Reasoning Models Distilled With Claude-Style Thinking Using GGUF and 4-Bit Quantization
-
TurboQuant Benchmarked in Llama.cpp: Google's Extreme Compression Research Tested in Practice
-
RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
-
Quantization Reveals Outliers Impacting LLM Accuracy
-
Nota AI and SiMa.ai Partner on Physical AI Technology for Local Deployment
-
Samsung Galaxy A37 and A57 5G Launch with On-Device AI Capabilities in India
-
Google's TurboQuant: The Unsexy AI Breakthrough Worth Watching
-
NVIDIA Releases GPT-OSS-Puzzle-88B, a Deployment-Optimized Model
-
Apple Plans Slimmed-Down Gemini Models for Local iPhone AI Features
-
Google TurboQuant: Extreme Compression for Local LLM Deployment
-
Running an Open-Weight LLM Locally on an Apple Watch
-
.APKs Are Just .ZIPs: Semi-Legally Hacking Software for Orphaned Hardware
-
Ultra-Large 400B-Class LLM Runs on iPhone in Test
-
LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language
-
Running an AI Agent on a 448KB RAM Microcontroller
-
Multiverse Computing Targets On-Device AI With Compressed Models and New API Portal
-
Researcher Discovers Universal "Danger Zone" in Transformer Model Architecture at 50% Depth
-
Nota Added to Three Technology and Growth ETFs in a Row – Market Recognition for AI Efficiency
-
Nota AI to Showcase End-to-End On-Device AI Optimization at Embedded World 2026
-
Student Researcher Achieves 42x Model Compression Through Novel Architecture
-
ETH Zurich Research Challenges Context-Length Assumptions in LLM Agents
-
OPPO and MediaTek Highlight On-Device AI Innovations at MWC 2026
-
On-Device AI Laptop Lineups Become Standard Across Major Manufacturers
-
Qualcomm Snapdragon Wear Elite Brings On-Device AI to Smartwatches
-
Meta Reveals AI-Packed Smartwatch In 2026 – Why Wearables Shift Now
-
Arduino and Qualcomm Bring On-Device AI Learning to Indian Schools
-
Mirai Announces $10M to Advance On-Device AI Performance for Consumer Devices
-
Kioxia Sampling UFS 5.0 Embedded Flash Memory for Next-Generation Mobile Applications
-
Enhanced Interface Speed Enables High-Performance On-Device AI Features in Smartphones
-
At India AI Impact Summit, Intel Showcases AI PCs and Cost-Efficient Frugal AI
-
Sarvam Brings AI to Feature Phones, Cars, and Smart Glasses
-
NVIDIA's Dynamic Memory Sparsification Cuts LLM Inference Costs by 8x
-
Samsung's REAM: Alternative Model Compression Technique