Tagged "release"
-
vLLM-iOS Achieves 88% Faster Multi-Agent Inference on Mobile Devices
-
JetBrains Releases Junie Local: On-Device Coding Agent for macOS
-
Run Open Models on Claude Desktop via Ollama Integration
-
Ollama v0.33.0 Adds Claude Desktop Integration and Improved Caching
-
Llama.cpp Build 10620: Continued Optimization for Local Inference
-
Qwen 3.6 Now Easier to Run Locally on Mac with JetBrains Integration
-
Xiaomi Unveils Xring O3, O100 and D100 Chips for On-Device AI and Smart Infrastructure
-
llama.cpp Build 10605: Mamba2 GEMM Optimization Improves State-Space Model Performance
-
FreeToken: Edge-Native MoE Serving Engine Runs 753B GLM-5.2 on Single Workstation GPU
-
llama.cpp Adds CUDA Pool Operations Support
-
Ollama v0.33.0 Adds Claude Desktop Integration and App Management
-
vLLM's Disaggregated Serving Cuts GPU Interference, Delivering 2.5x Higher Goodput
-
llama.cpp Build b10581 Adds DSpark Support for Faster Local Inference
-
Liquid AI Releases DSpark Version of Compact LFM2.5 Models with Up to 2.67x Speedup
-
Ollama v0.33.0 Release Candidate Adds Claude Desktop Integration and Performance Improvements
-
Liquid AI Releases LFM2.5-DSpark Draft Models with 3.18x Faster Decoding
-
llama.cpp b10549: Tensor Parallelism Support for LFM2/LFM2MOE Models
-
Ollama v0.32.15: Time-to-First-Token Cut in Half with Metadata Caching
-
llama.cpp b10524 Makes MoE Expert Scatter Deterministic in OpenCL
-
Liquid AI Releases LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation
-
Ollama v0.32.15 Adds Model Metadata Cache to Reduce Per-Request Overhead
-
Native vLLM and ROCm 7.15 Support for AMD RDNA2 GPUs on Windows
-
Llama.cpp Release b10485: GGML Sync with Platform-Specific Optimizations
-
Llama-macOS – Agentic and MCP Native macOS Front End for Llama.cpp
-
Unsloth Releases Qwen 3.8 27B GGUF Quantised Weights
-
Ollama Adds Qwen 3.8 27B with Optimised Apple Silicon Support
-
Ollama Adds Qwen 3.8 27B with Apple Silicon Optimizations
-
Liquid AI Releases LFM2.5-VL-3B: Compact Vision-Language Model for Edge Inference
-
Ollama 0.32.11: DeepSeek Harness and Meta's Muse Code Integration
-
LFM2.5-VL-3B: Lightweight Vision-Language Model Optimized for Edge Deployment
-
AMD Launches Gorgon Halo and ROCm.AI for Local AI Inference with Workstation Hardware
-
Ollama v0.32.10: Faster Prefill Performance on NVFP4 Models with System Config Support
-
vLLM v0.27.0 Brings Major Performance Improvements and New Model Support
-
vLLM v0.27.0 Released with 561 Commits and Expanded Model Support
-
llama.cpp Improves Muse Glimmer Tool Calling with Latest Update
-
llama.cpp Updates Tool Call Detection for Muse Glimmer
-
vLLM v0.27.0 Released with Major Kernel Improvements and New Model Support
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Agent Execution
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Local Agent Deployment
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Agent Execution
-
Meta's Muse Glimmer Now Available Across All Platforms in Ollama
-
Meta's Muse Glimmer Now Available Across All Platforms in Ollama
-
Meta's Muse Glimmer Now Available Across All Platforms via Ollama
-
vLLM v0.27.0 – Kimi K3 Support and 561 Commits from 242 Contributors
-
NVIDIA Magpie TTS – Open-Weights Multilingual Voice Agents with Full Deployment Control
-
Meta's Muse Glimmer – Local, Agentic, Multimodal, and Open Source
-
Meta Releases Muse Glimmer: 30B Open-Source LLM for Local Deployment
-
Muse Glimmer Now Available on Ollama – Meta's Open Multimodal Agent Model
-
ShoutFlow Launches Pay-Once, On-Device AI Dictation App for the Mac
-
llama.cpp Adds Tool Isolation Support via Docker
-
vLLM v0.27.0rc2 Release Candidate Available
-
llama.cpp Improves CUDA Performance with Kernel Fusion
-
Llama.cpp Fixes Metal NORM Operations for Apple Silicon
-
MSI Crosshair A16 HX: Professional Gaming Laptop Built for AI and Gaming
-
Llama.cpp Adds LRU Scheduler for Multi-Model Serving
-
Llama.cpp B10327 Fixes CUDA Quantized Copy Kernel Performance
-
NeuronAI: First Free Unified TTS, STT, and LLM Platform
-
Liquid AI LFM2.5-2.6B: Open-Weights Agentic Model With 128K Context and Tool Calling
-
Liquid AI Releases LFM2.5-2.6B: Powerful Agentic Model for Raspberry Pi and Edge Devices
-
llama.cpp Build b10301: CUDA Optimization and Compiler Warning Fixes
-
vLLM v0.27.0rc1: Latest Release Candidate for High-Performance Inference
-
llama.cpp b10298: Multi-Token Multi-Dimension Chunk Serialization Support
-
Ollama v0.32.6: Faster Apple GPU Inference with Speculative Decoding
-
Seeed Studio's reCamera Pro Makes On-Device AI Faster and Easier
-
DeepSeek V4 Flash Optimized for Single AMD MI300X GPU
-
PrismML's Bonsai 27B Brings On-Device AI to Apple iPhone 17 Pro
-
ASUS Vivobook S16 Arrives with 45 TOPS NPU and OLED Display
-
Homebench: Comprehensive Benchmarking Tool for Local LLMs
-
llama.cpp Adds DeepSeek V4 Flash Chat Template Support
-
Gainz.fast – Local Inference, Faster
-
llama.cpp b10256 – SYCL SDPA Extended to Quantized KV Caches
-
llama.cpp Build b10258: Sampling Architecture Refinements
-
llama.cpp Release b10257 – Vulkan LLVMpipe Fixes
-
K-EXAONE 2.0 Brings 262K Context to Frontier AI
-
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
-
NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework
-
NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework
-
A local-first grid of grids for notes (similar to treesheets)
-
Simple Open WebUI Alternative for Running Ollama Models in Web Browser
-
Kioxia UFS 5.0 Embedded Flash Memory Enables On-Device AI with Advanced Storage Architecture
-
Tether Data Releases VisionPsy-Nano: Open Source Edge Visual Language Model
-
NightRun UEFI Application Boots Local LLM on Raspberry Pi 5 and x86 PCs Without an OS
-
Kioxia's UFS 5.0 Embedded Flash Enables Practical On-Device AI
-
Enprompta: Prompt Registry, LLM Evals, and Observability for Production AI Apps
-
Triton Control: Open-Source Control Plane for Nvidia Triton on Kubernetes
-
faster-enhancer.c: C Library for Stable Real-Time On-Device Denoising
-
NVIDIA Releases Molt: Agentic RL Training Framework Scaling to Trillion-Parameter Models
-
OPPO Launches Xiaobu Next Beta, Debuts On-Device Multi-Agent System on Smartphones
-
Ruff v0.16.0: 413 Default Rules for Code Quality in AI Development
-
Apertus 1.5: Swiss Open-Weight, Open-Source LLM Released
-
A New Way of Debugging Open-Weight Models - IBM
-
MSI Pro Max Edge AI+ Mini PC Runs 120B Local AI Models With 128GB RAM
-
Transept: AI Translation Workspace Prioritizing Human-Centric Design
-
Grok Launches Excel AI Add-in for Integrated Model Access
-
Round-Trip Correctness: New Metric for Generative AI Process Modeling
-
Mozilla Firefox 153 ESR Adds On-Device AI Capabilities for Enterprise Deployment
-
Apertus 1.5 Released with Local AI Improvements
-
Multiverse Computing's CompactifAI Models Now Fully Compatible with Intel Xeon 6 Processors
-
Shanghai Droi Technology Launches DroiClaw AI Operating System with Hybrid Edge-Cloud Architecture
-
Gemini Nano 4 Arrives with Samsung's Latest Foldables, Bringing LLMs to Mobile Edge
-
Arm China Unveils "Tianxuan" CPU and Xingchen 300 Platform, Targeting Ubiquitous AIoT with On-Device AI Portfolio
-
llama.cpp b10075 Packs Four Local AI Runtime Upgrades
-
Deterministic Arena: Testing and Comparing AI Agents Through Code Execution
-
LLM Wiki Implementation: Community Resource for Local Deployment
-
Nubia Announces AI Agent Smartphone with On-Device AI Processing
-
Shikigami: Run AI Coding Agents in Parallel Using Git Worktrees
-
Qwen 3.8 with 2.4T Parameters Going Open-Weight Soon
-
Mozilla AI Releases Llamafile 0.10.4 With New Transcribefile Built On Transcribe.cpp
-
Google Gemma 4 Debuts for Pixel 10 With Powerful On-Device AI Features
-
Mira Murati's Thinking Machines Launches Open-Weight AI Model
-
Vivo Unveils Security Solution for On-Device AI at AI for Good Global Summit 2026
-
Stop Paying for Search APIs—This Self-Hosted Tool Lets Your Local LLM Search the Web for Free
-
Google's LiteRT.js Enables On-Device AI Inference in Web Browsers
-
Qualcomm Unveils Snapdragon Reality Elite for On-Device AI and Spatial Computing
-
DolphinDB v3.00.6 and v2.00.19: Introducing DolphinX for Enterprise AI Agents
-
Runeward: Sandboxing AI Agents with Policy Gates
-
Grinta – A Local-First Coding Agent Built for Long Autonomous Runs
-
AgentKindergarten – Daycare for Your AI Coding Agents
-
AMD ZenDNN 6.0 Boosts AI Inference on EPYC CPUs With FP16 and MoE Acceleration
-
Record and Replay: Teach AI Agents Desktop Workflows by Showing Them Once
-
Show HN: OpenVole 4.5 Is Out
-
Intel-Scaler-vLLM 0.21.0-b1 Brings Latest Features for vLLM on Intel GPUs
-
Exploiting Sparsity for Long Context Inference: Million Token on Commodity GPUs
-
Relm – Local LLMs as Base-R Objects with Interpretability
-
Opendray – Run Claude Code/Codex Agents on Your Own Box
-
Tencent Open-Sources Hy3 295B MoE Model Built for STEM Reasoning
-
Show HN: Trace – Open-source, Self-organizing Memory for LLM Agents
-
Samsung UFS 5.0 Storage Interface Optimizes On-Device AI Performance and Latency
-
Off-Grid AI Launches Emergency Preparedness Platform Powered by Local LLM Inference
-
AMD Ryzen AI Halo Mini PC Delivers Powerful Local Inference With Open-Source Stack
-
Google Rolls Out Android 17 and Gemma 4 with Advanced On-Device AI
-
Compressor V2: Three Compression Layers for 50% LLM Agent Cost Cut
-
code-on-incus: Isolated Machine Environments for AI Agents
-
Microsoft's Intelligent Terminal Works Seamlessly with Local LLMs
-
LongCat-2.0 Released
-
Ollama's New MLX Engine Delivers Significant Performance Gains on Mac
-
PewDiePie Releases Open-Source Odysseus AI Workspace
-
Open Source 1B LLM Trained from Scratch for $315 with Weights and Data Released
-
Transcribe.cpp – ggml speech-to-text inference engine
-
Apple Updates Creator Studio with AI Video Editing, Image Generation, and Logic Pro Enhancements
-
Samsung Unveils UFS 5.0 Solution for Next-Gen On-Device AI Applications
-
Meet EverOS: An Open Source Markdown-First Agent Memory Runtime With Hybrid BM25 + Vector Retrieval
-
LLM-Free, Layout-Aware PDF Chunker in Pure Rust
-
Qualcomm AI Hub Expands to 1,500 Optimized Models for Edge Deployment
-
You Can Now Run Max AI Models on Apple Silicon
-
GEEKOM A9 Max Delivers 32GB RAM and Native LLM Support in Compact Form Factor
-
Liquid AI Ships LFM2.5-230M with Broad Framework Support for On-Device Inference
-
DEEPX and Sixfab Launch AI HAT for Raspberry Pi Edge Inference
-
NeoEyes NE503 Brings 20 TOPS of On-Device AI to Industrial Cameras
-
DEEPX and Sixfab Launch 'DEEPX AI HAT' to Drive Edge Physical AI on Raspberry Pi
-
Samsung Unveils UFS 5.0 Storage Solution Optimized for On-Device AI
-
Qwable: New Free Local Model Brings Claude-like Capabilities to Edge Devices
-
Local AI Orchestrator with Computer and Browser Access
-
ORA: Smaller Models. Same Intelligence
-
PipeVoice: The Free Local Alternative to Whisper Flow
-
DeepSWE v1.1 – Updated Execution and Grading for Software Engineering Tasks
-
Show HN: Agnes AI – Free Multimodal API (Text, Image, Video), OpenAI-Compatible
-
Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
-
Samsung Unveils UFS 5.0 Storage Optimized for On-Device AI Applications
-
Founders OS – Self-Hosted AI with Real Business Context
-
MCP Server Enables Claude to Automate Mac Tasks and Self-Correct
-
Qualcomm Launches Snapdragon START to Speed AI Smart Glasses to Market
-
Qualcomm Launches Snapdragon Reality Elite for AI-Powered Spatial Computing
-
PageToMD – A CLI tool to turn web pages into clean Markdown for AI agents
-
Free Tool Helps Match Local AI Models to Your Hardware
-
Ray Serve LLM Achieves 24x Performance Improvement in Distributed Inference
-
Qualcomm Debuts Snapdragon Reality Elite XR Platform with On-Device AI
-
Unreal Engine 5.8 Adds MCP Server for AI Agents
-
TongFlow: Free Open-Source Multi-Modal AI Workflow Studio
-
Qualcomm Snapdragon Reality Elite Brings 48 TOPS AI to XR Devices
-
Genesis AI Launches Eno General-Purpose Robot with Embedded AI
-
Qwen and Fable: Open-Weights 35B Mixture-of-Experts Agentic Coding Model
-
Google's DiffusionGemma Brings Novel Text Generation to Local LLMs
-
ProData AI – 14 MCP Tools for Automated Data Science
-
Hermes Agent Transforms Local LLMs Into Executable Agents
-
CoreMCP – MCP Server for On-Prem Databases
-
Local-First TypeScript Guard for Runaway AI-Agent Costs
-
Tensordyne Napier AI Processor Announced with Logarithmic Math
-
Repo-Slopscore: Detecting AI Contributions in Git Repositories via Commit Analysis
-
Docfai.app Launches With Free Trial for Local Document Processing
-
Brilliant Labs Halo: Open-Source AI Glasses for On-Device Intelligence
-
Contrail Compute AIX: First RISC-V AI Execution Platform
-
AMD PACE: New vLLM Plugin Enables Efficient CPU-Based Inference
-
Google's DiffusionGemma Achieves 4x Faster Text Generation for Local Deployment
-
Outpost – Capability-based API access for AI agents
-
AMD's Lemonade SDK Adds NVIDIA CUDA Support for Cross-Platform Local AI
-
Qualcomm Launches Dragonwing MBM Silicon with Advanced On-Device AI Capabilities
-
Apple Unveils AFM 3 Core Advanced with 20 Billion Parameters for On-Device AI
-
Qualcomm Unveils Dragonwing MBM Silicon with Integrated On-Device AI and Connectivity
-
CoAnalyst360: Multi-Agent AI Platform for Investigative Questions
-
Google Releases Gemma 4 QAT Models with Reduced Memory Requirements for Mobile and Laptop Deployment
-
Qualcomm Unveils Dragonwing IQ10 RRD Platform for Rapid Edge AI Deployment
-
Tinytasktree – Behavior-tree-style task orchestration for LLM agents
-
Google AI Edge Gallery Launches on macOS With Offline Gemini Models
-
DockSec: Open-Source AI-Powered Container Security Scanner for Self-Hosted Deployments
-
Google Introduces Gemma 4 QAT for Ultra-Low Memory Local Inference
-
Qualcomm's Dragonwing IQ10 RRD Fast-Tracks Robots From Prototype to Production Deployment
-
NVIDIA Unveils First PC Chips at Computex 2026; CEO Jensen Huang Details New Hardware
-
Google's New Gemma 4 12B AI Model Is Built for Laptops
-
Google Releases Gemma 4 QAT Models for Local AI Deployment
-
Running Infinite Context Lengths on 8GB GPU Without Out Of Memory
-
LLM Checker Tool Helps Identify Models for Your PC
-
Sawtooth – An Async, Multi-Tiered Memory Framework for LLM Agents
-
NVIDIA Dynamo Snapshot Accelerates AI Inference Startup on Kubernetes
-
Google Launches AI Edge Gallery on macOS for Running Gemini Models Locally
-
Google Releases Gemma 4 12B Model for Local Inference on 16GB Enterprise Laptops
-
Run Llama.cpp In-Process from Java with Project Panama FFM
-
Longsys Redefines On-Device AI with Groundbreaking Edge Memory Solutions
-
Bosgame Launches VTA-439 Mini PC with 86 TOPS for Practical Local AI
-
Google Launches AI Edge Gallery on macOS for Running Gemini Models Locally
-
Google Releases Gemma 4 12B: Encoder-Free Multimodal Model for 16GB Laptops
-
Microsoft Expands On-Device AI Models in Edge Browser with New APIs for Local Inference
-
Perplexity Unveils Hybrid Local-Cloud Inference System for Intelligent Task Distribution
-
Snapdragon C Processor Brings On-Device AI Engine to Wearables and Edge Devices
-
WSL 3 Brings Near-Native GPU and NPU Passthrough for Local AI on Windows
-
NVIDIA RTX Spark Superchip Delivers 6,144 CUDA Cores for Consumer Local AI Inference
-
Tether AI Upgrades QVAC SDK With TurboQuant for Data Center-Sized Memory on Everyday Devices
-
NVIDIA and Microsoft Team Up to Bring Secure On-Device AI Agents to Windows PCs
-
JetBrains Releases Mellum2: A 12B MoE Model for Fast, Specialized Tasks
-
Phison and Intel Roll Out aiDAPTIV to Boost Local AI on Intel AI PC Platforms
-
Qualcomm Reveals Snapdragon C with Advanced On-Device AI Engine
-
Netflix Wiz Creates App to Slash AI Bills, Then Open Sources It
-
NVIDIA Levels Up Local AI Agents Across RTX PCs and DGX Spark
-
NVIDIA Launches N1X/N1 CPU-GPU SoC for PC Market, Targeting Heavy On-Device AI Users
-
Snapdragon C Specs Revealed: 6nm Process, On-Device AI Engine for Budget Laptops
-
Liquid AI Launches Edge-Focused LFM2.5 Model to Power On-Device AI Agents
-
Oracle APEX 26.1 Expands AI Choice with Out-of-the-Box Support for Major AI Providers
-
Snapdragon C Debuts with 6nm Process and Dedicated On-Device AI Engine
-
MediaTek Dimensity 7500 Brings On-Device AI and Enhanced Power Efficiency to Mid-Range Phones
-
Rsync 3.4.3 Features Hundreds of Claude Commits
-
Zoho-Backed Netrasemi Launches 12nm AI Chip, Mass Production Begins This Year
-
MediaTek Launches Dimensity 8550 4nm SoC with Integrated On-Device AI Focus
-
Liquid AI Unveils Edge-Focused LFM2.5 Model for On-Device AI Agents
-
Google Launches Tiny Board for Running Gemma 3 Locally
-
Superpowers: An Agentic Skills Framework for AI Coding Workflows
-
Mistral AI Launches Mistral Vibe
-
Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM Inference
-
Dell Launches 14 Plus Laptop with Intel Core Ultra 9 and 32GB RAM at $1,499.99, Enabling Local Model Inference
-
LM Studio 0.4 Introduces Headless Deployment for Local LLM APIs
-
Gemma 4: A New Budget-Focused Model in Posit AI
-
Show HN: An Open-Source Interactive AI Engineering Syllabus (1,100 Papers)
-
From Source Code to LLM Constraints: A Semantic Extractor for Python, SwiftUI, Lua
-
AMD Unveils Ryzen AI Halo Developer Platform for On-Device AI Workloads
-
PLLuM: Poland's Ministry of Digital Affairs Releases Open Models on HuggingFace
-
llama.cpp MTP Leak Fix Stabilizes Local AI Agents
-
llama.cpp Checkpoint Fix Accelerates Local Coding Agents
-
Intel llm-scaler-vllm 1.4 Released With Updated Components and Arc Pro B70 Support
-
AMD's New Ryzen AI Max Pro 400 with 192GB LPDDR5X Memory
-
Adobe Photoshop Update Brings On-Device AI Processing
-
Google Tensor SDK Beta with LiteRT Enables Efficient On-Device AI
-
Open Source Local Audio Stem Separation Tool Released
-
llama.cpp Adds Multi-Token Prediction, Doubles Qwen 3.6B Throughput for Local Inference
-
eXo MCP Server Enables Secure AI Agent Access to Workplace Tools
-
Samsung's Exynos 2800 Brings Significant On-Device AI Capabilities
-
AMD's Lemonade SDK Advances macOS Support for Local AI Inference with ROCm 7.13
-
N8n-MCP: AI Assistants Can Now Build and Search n8n Workflows
-
DwarfStar 4: Native Inference Engine Optimized for DeepSeek V4 Flash
-
AI/ML Benchmark Tool for Local LLM Inference and XGBoost Training
-
SynapseKit: A New Production Framework for Deploying LLMs
-
RelaxAI – UK sovereign LLM inference at 80% cheaper than OpenAI/Claude
-
Hedy AI Launches Privacy-First On-Device AI Processing Platform
-
Berget AI Announces Berget Code for European Teams Powered by Kimi K2.6
-
BT Explainer: Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
AMD's vLLM-ATOM Plugin Supercharges DeepSeek-R1 and Kimi-K2 Inference on MI350/MI400
-
Mlx-serve: Run LLMs Natively on Your Mac
-
LibreOffice 26.4 Beta Integrates Local AI Writing Features
-
Google Releases Gemma 4 Multi-Token Prediction Drafters To Accelerate AI Inference
-
Perplexity Brings On-Device AI Workflow to Macs with 'Personal Computer' Feature
-
Zed Editor Integrates AI Features with Local Deployment Focus
-
Sarvam Edge: Indian-Built AI Models Run Offline on Phones and Laptops Without Internet
-
llama.cpp Now Supports Multi-Token Prediction in Beta
-
Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
NordVPN Adds On-Device AI Voice Detector to Chrome Extension to Identify Synthetic Audio
-
Anker's Thus Chip Puts AI On-Device, Promising Faster Responses And Better Privacy
-
Thoth – Open-Source Local-First AI Assistant
-
SQL Server 2025 Adds Built-in Chunking and Vector Support
-
Google Drops COSMO: Experimental On-Device AI Assistant for Android
-
ScopeGuard 0.0.7: Go Linter with Model Context Protocol Support
-
PFlash Claims 10x Prefill Speedup Over llama.cpp
-
Anker's New 'Thus' Chip Brings 150x AI Power to Earbuds
-
New Open-Source Tool Automatically Matches Local LLMs to Your PC Hardware
-
IBM Introduces Granite 4.1 Family of Models for Local Deployment
-
Google's Gemma 4 Brings Powerful AI Capabilities to Phones and Laptops
-
Show HN: Arkloop – Open-Source, Local-First Agent Client
-
Google's Gemma 4: Powerful AI Models Optimized for Your Phone and Laptop
-
Pocket LLM v1.5.0 Brings Multimodal AI to Android with No Cloud Required
-
Unsloth's Custom Kernels Make LLM Fine-Tuning Viable on Consumer GPUs
-
Pluggable's TBT5-AI: First Thunderbolt Dock Explicitly Targeting Local LLM Workstations
-
Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Google's Gemma 4 Brings Powerful On-Device AI to Phones and Laptops
-
Mathesar 0.10.0
-
Seed3D 2.0
-
Anker Unveils 'Thus' Chip to Bring On-Device AI Across Product Line
-
Intel OpenVINO 2026.1 Integrates llama.cpp with Wildcat Lake and Arc Pro B70
-
Intel LLM-Scaler vLLM 0.14.0 Released With Official Arc Pro B70 Support
-
Tesseron: New API Framework for AI Agents with Developer-Defined Configuration
-
Sarvam Edge: India's Offline AI Model Runs on Phones and Laptops Without Internet
-
go-AI: New Inference API Library for Go Released
-
Google's Gemma 4 Finally Makes Local LLM Deployment Compelling for Practitioners
-
Bun v1.3.13
-
Intel Extends AI PC Reach With New Core Ultra Series 3 Launch
-
Minisforum Launches N5 Max AI NAS with OpenClaw
-
Web Agent Bridge: Open-Source OS for AI Agents
-
Build a More Secure, Always-On Local AI Agent with OpenClaw and NVIDIA NemoClaw
-
115 TOPS in 0.67L: CHUWI AuBox X Packs On-Device AI Power Into a Palm-Sized Mini PC
-
Laimark – 8B LLM That Self-Improves on Consumer GPUs
-
DotLLM – Building an LLM Inference Engine in C#
-
Google's Gemma 4 Brings Game-Changing Performance to Local Laptop Inference
-
DFlash Doubles Token Generation Speed of Qwen3.5 27B on Mac M5 Max
-
Fine-Tuned Qwen3.5-0.8B for OCR Outperforms Previous 2B Release
-
Minisforum N5 MAX AI NAS Delivers 126 TOPS with 200TB Storage for Local LLM Workloads
-
Qwen 3.5 Small – On-Device Multimodal Models Released
-
OpenNebula 7.2 "Dark Horse" Released with Enhanced Infrastructure Support
-
ASUS Malaysia to Bring UGen300 USB AI Accelerator in Q2 for Portable On-Device AI Inferencing
-
Qwen3 Audio and Vision Support Now Available in llama.cpp
-
Audio Processing Support Lands in llama.cpp with Gemma-4
-
MiniMax M2.7 Open-Sources Globally as Industry's First Self-Improving Model
-
Defender – Local Prompt Injection Detection for AI Agents
-
Rapidly Scaffold Agents, MCP Servers, APIs, Websites on AWS
-
MiniMax M2.7 Is Now Open Source
-
Unsloth Completes Comprehensive MiniMax M2.7 GGUF Quantization Suite
-
MiniMax M2.7 Released: New Model Available for Local Deployment
-
MiniMax M2.7 Advances Scalable Agentic Workflows on NVIDIA Platforms for Complex AI Applications
-
Google's Gemma 4 Brings Free Agentic AI to Your Phone With Zero Data Leaving the Device
-
Critical Unsloth Gemma-4 Chat Template Updates for Tool Calling
-
Aisbf (AI Should Be Free) Proxy 0.99.18 Released
-
Google's Gemini Nano 4 Offers Faster, Smarter Local Inference Capabilities
-
ASUS ExpertBook P1 Integrates On-Device AI for Enterprise Collaboration
-
DMax: New Parallel Decoding Paradigm for Diffusion Language Models
-
LLM Wiki v2: Extended Knowledge Base for LLM Practitioners
-
CarryAI's Serverless Vision-Language Models Enable On-Device Multimodal AI
-
Tether Launches QVAC SDK for Cross-Platform Local AI Development
-
VoxCPM2: New Open-Source TTS Model with Voice Cloning and Design
-
Intel Releases OpenVINO 2026.1 With Backend For Llama.cpp, New Hardware Support
-
Gemma 4 GGUF Models Updated with Critical Quantization Fixes
-
EXAONE 4.5 33B Model Released with Multiple Quantization Formats
-
Docsie Launches On-Premise AI Platform for Regulated Industries
-
GitHub Copilot CLI Adds Support for BYOK and Local Model Deployment
-
Google's Gemma 4 Brings Powerful On-Device AI to Android and iOS
-
Octopoda: Open Source Memory Layer for Fully Offline AI Agents
-
Google Launches Offline AI Dictation App for iOS with Gemma
-
TurboQuant-Optimized llama.cpp Fork Delivers GFX906 GPU Acceleration
-
Lenovo Korea Launches AI-Powered Industrial Edge Solutions
-
Show HN: Lightweight LLM Tracing Tool with CLI
-
Google Previews Gemini Nano 4 for Android AICore with On-Device Capabilities
-
Qwen 3.6 Free Model Available via OpenRouter
-
GMKtec NucBox K17 Launches with 97 TOPS AI Performance for Local Inference
-
Ollama Gets Blazing Fast on Macs with Full MLX Support and 2× Speedups
-
Mixed Precision Quantization on MLX with TurboQuant Implementation
-
Netflix Open-Sources VOID Model for Video Object Deletion
-
Samsung Launches Galaxy Book6 Series with NVIDIA RTX 5070 and On-Device AI
-
Google Launches Gemma 4 For Advanced On-Device AI
-
Gemma 4 KV Cache Memory Issues Fixed in llama.cpp
-
NVIDIA Accelerates Gemma 4 for Local Agentic AI on RTX GPUs
-
Google Gemma 4 Released with GGUF Quantizations
-
Google Launches Gemma 4 Open Models for Local On-Device AI
-
AMD Provides Day 0 Support for Gemma 4 on Ryzen AI Processors and GPUs
-
Gemma 4 on Arm: Optimized On-Device AI for Mobile and Edge Deployment
-
Qwen 3.6-Plus Released
-
TinyGPU Adds Mac Support for External Nvidia GPU Acceleration
-
Bonsai 1-Bit Models Deliver Exceptional Local Inference Performance
-
Apple Silicon Macs Run Local AI Faster with Ollama's New MLX Support
-
Gemini CLI – Open-Source AI Agent for Terminal Integration
-
ByteShape Releases Qwen 3.5 9B Quantisations with Hardware-Matched Tuning Guide
-
Llama.cpp Merging TurboQuant Lite (attn-rot) with Major Performance Gains
-
PrismML Announces 1-Bit Bonsai: First Commercially Viable 1-Bit LLMs
-
Samsung launches Galaxy Book6 series in India with Nvidia RTX 5070 graphics and on-device AI
-
Ollama Launches Pi: The Minimal Coding Agent That Powers OpenClaw Is Now Yours to Customize
-
Samsung Launches Galaxy Book6 Series in India with NVIDIA RTX 5070 Graphics and On-Device AI
-
Dell Technologies Unveils 10 AI PC Models for Business, from Ultralight Laptops to Ultracompact Desktops
-
ESP32-S31: 320MHz 2-Core Microcontroller with 512KB SRAM and Networking
-
IBM Granite 4.0 3B Vision: Compact Enterprise-Grade Document AI
-
Scion: Running Concurrent LLM Agents with Isolated Identities and Workspaces
-
Acer TravelMate AI Laptops Launch in UAE for Business On-Device Inference
-
Unsloth Studio Beta Ships 50+ New Features for Local Model Training and Inference
-
Introduction to Nyreth v1.0
-
HP Launches Copilot+ PCs in India with On-Device AI Capabilities for Local Inference
-
GLM-5.1 Model Weights Launching Early April for Local Deployment
-
Mistral AI Releases Voxtral: Open-Source TTS Model Beating ElevenLabs on Local Hardware
-
Samsung Galaxy A37 and A57 5G Launch with On-Device AI Capabilities in India
-
Pluggable's TBT5-AI: First Thunderbolt Dock Explicitly Targeting Local LLM Workstations
-
Intel Launches Arc Pro B70/B65 with 32GB VRAM for Local AI Inference
-
Google's TurboQuant: The Unsexy AI Breakthrough Worth Watching
-
NVIDIA Releases GPT-OSS-Puzzle-88B, a Deployment-Optimized Model
-
Meta Releases HyperAgents: Self-Improving AI
-
New Open-Weight Models Released: GigaChat-3.1-Ultra and Lightning Variants
-
HP Launches IQ On-Device AI Assistant, Advancing Enterprise AI Adoption on PCs
-
OmniCoder v2 Released: Improved Code Generation for Local Deployment
-
Google TurboQuant: Extreme Compression for Local LLM Deployment
-
Lemonade 10.0.1 Improves Setup Process For Using AMD Ryzen AI NPUs On Linux
-
Qt 6.11 Released with Enhanced Cross-Platform Deployment Capabilities
-
Self-Hostable AI Agents and Internal Software Framework Released
-
MiniMax M2.7 Model to Be Released as Open Weights
-
LM Studio Releases Reworked Plugins with Fully Local Web Research
-
Velr: Embedded Property-Graph Database for Local LLM Applications
-
BrowserOS 0.44.0 Release: Advances in Local AI Integration for Web-Based Applications
-
Qwen 3.5 122B Uncensored (Aggressive) Released with New K_P Quantisations
-
Nvidia Nemotron Cascade 2 30B Emerges as Powerful Alternative to Qwen Models
-
Atuin v18.13 – Better Search, a PTY Proxy, and AI for Your Shell
-
Pydantic-Deep: Production Deep Agents for Pydantic AI
-
ASUS ExpertCenter PN55 Mini PC Combines AMD AI CPU and 55 TOPS NPU
-
NVIDIA Nemotron Cascade 2 30B Delivers 120B-Class Performance in Compact Form Factor
-
Llamafile 0.10 Released with GPU Support and Rebuilt Core
-
NVIDIA Nemotron 3 Nano 4B Enables On-Device Inference Directly in Web Browsers via WebGPU
-
Cybersecurity Skills for AI Agents – agentskills.io Standard Implementation
-
Dell Pro Max 16 Plus Launches With Enterprise-Grade Discrete NPU for On-Device AI
-
Multiverse Computing Targets On-Device AI With Compressed Models and New API Portal
-
Meet Sarvam Edge: India's AI Model That Runs on Phones and Laptops With No Internet
-
Tether's QVAC Introduces Cross-Platform Bitnet LoRA Framework for On-Device AI Training
-
Hugging Face Releases One-Liner for Automatic Hardware Detection and Model Selection
-
Unsloth Studio: Open-Source Web UI for Training and Running LLMs Locally
-
MiniMax-M2.7: New Compact Model Announced for Local Deployment
-
On-Device AI: Tether's QVAC Fabric Enables Local Training
-
Mamba 3: State Space Model Architecture Optimized for Inference
-
Mistral Small 4 119B Released with NVFP4 Quantisation Support
-
Mistral Releases Small 4 Open-Source Model Under Apache 2.0
-
Mistral Releases Leanstral: First Open-Source Code Agent for Lean 4 Proof Assistant
-
Kimi Introduces Attention Residuals: 1.25x Compute Performance at <2% Overhead
-
NVIDIA Updates Nemotron 3 122B License, Removes Deployment Restrictions
-
OmniCoder-9B: Efficient Coding Model for 8GB GPUs
-
StepFun Releases SFT Dataset Used to Train Step 3.5 Flash for Community Fine-Tuning
-
Cicikus v3 Prometheus 4.4B – An Experimental Franken-Merge for Edge Reasoning
-
Nvidia's Nemotron 3 Super: Understanding the Significance for Local LLM Deployment
-
AMD Launches Agent System Optimized for Local AI Inference With Ryzen and Radeon
-
Intel OpenVINO Backend Support Now Available in llama.cpp
-
Lemonade v10 Brings Linux NPU Support and Multi-Modal Capabilities
-
Intel Updates LLM-Scaler-vLLM With Support For More Qwen3/3.5 Models
-
Llama.cpp Adds True Reasoning Budget Support
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Qwodel – An Open-Source Unified Pipeline for LLM Quantization
-
Nvidia Releases Nemotron 3 Super: 120B MoE Model for Local Deployment
-
SK Hynix Completes Qualification for LPDDR6 Memory Optimized for AI Inference
-
Texas Instruments Launches NPU-Powered MCUs for Low-Power Edge AI
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Qwen 3.5-35B Uncensored GGUF Models Now Available
-
Kali Linux Integrates Local Ollama and MCP for AI-Driven Penetration Testing
-
Gloss: Open-Source, Local-First RAG Alternative to NotebookLM Built in Rust
-
FreeBSD 14.4 Released: Implications for Local LLM Deployment
-
Fish Audio Open-Sources S2: Expressive Text-to-Speech with Natural Language Control and 100ms Latency
-
SK Hynix Develops 1c LPDDR6 DRAM to Boost On-Device AI Performance in Mobile Devices
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Qwen 3.5 Derestricted Model Available for Local Deployment
-
Qwen 3.5 Small Expands On-Device AI to Phones and IoT with Offline Support
-
Engram – Open-Source Persistent Memory for AI Agents
-
Snapdragon Wear Elite Unveiled at MWC 2026, Advancing Wearable AI Inference
-
HP Refreshes Lineup with AI-Focused Workstations
-
Apple Launches MacBook Neo with A18 Pro Chip for Affordable Local AI Inference
-
Sarvam AI Releases 30B and 105B Open-Source Models Trained from Scratch
-
Open WebUI Adds Native Terminal Tool Calling with Qwen3.5 35B Support
-
Jse v2.0 AI Output Specification
-
IBM Granite 4.0 1B Speech Model Released for Multilingual Speech Recognition
-
Llama.cpp Merges Automatic Parser Generator to Mainline
-
Alibaba Releases Qwen 3.5 AI Model with On-Device AI Support
-
Building PyTorch-Native Support for IBM Spyre Accelerator
-
Final Qwen3.5 Unsloth GGUF Update with Improved Size/Quality Tradeoffs
-
Alibaba Releases Qwen 3.5 AI Model with On-Device AI Support
-
llama.cpp Merges Agentic Loop and MCP Client Support
-
Apple Unveils MacBook Pro with M5 Pro and M5 Max Featuring On-Device AI
-
Kakao Launches Kanana AI for On-Device Schedule and Recommendation Management
-
Apple Unveils MacBook Pro With M5 Pro and M5 Max for On-Device AI
-
RunAnywhere Launches Production-Grade On-Device AI Platform for Enterprise Scale
-
Apple M5 Pro and M5 Max: 4× Faster LLM Processing
-
OpenWrt 25.12.0 – Stable Release
-
AMD Launches Copilot+ Desktop Chips to Compete in On-Device AI Market
-
Qualcomm Snapdragon Wear Elite Brings On-Device AI to Smartwatches
-
ÆTHERYA Core – Deterministic Policy Engine for Governing LLM Actions
-
Qualcomm Snapdragon Wear Elite: 2B Parameter NPU for Personal AI Wearables
-
Apple M4 iPad Air Targets AI Users with Double M1 Speed Performance
-
Alibaba's Qwen 3.5 Small Model Runs Directly on iPhone 17
-
Qwen 3.5 Small Models Released: 0.8B to 9B Parameters Optimized for On-Device Inference
-
AMD Ryzen AI 400 Series Desktop Processors Launch with Integrated 60 TOPS NPU
-
Qualcomm Launches Snapdragon Wear Elite for On-Device AI on Wearables
-
GitDelivr: A Free CDN for Git Clones Built on Cloudflare Workers and R2
-
AMD Expands Ryzen AI 400 Series Portfolio for Consumer and Enterprise AI PC Options
-
Jan Releases Code-Tuned 4B Model for Efficient Local Code Generation and Development Tasks
-
Alibaba's Open-Source CoPaw AI Agent Now Compatible with MCP and ClawHub Skills
-
ParseHive – AI-Powered Invoice Data Extraction for Windows and Mac
-
Huawei's SuperPoD Portfolio Creates New Option for Global Computing at MWC Barcelona 2026
-
DeepSeek V4 Multimodal Model Coming Next Week With Image and Video Generation
-
Qwen 3.5-35B-A3B Emerges as Efficient Daily Driver, Replacing 120B Models
-
Unsloth Dynamic 2.0 GGUFs
-
The ML.energy Leaderboard
-
LLmFit: One-Command Hardware-Aware Model Selection Across 497 Models and 133 Providers
-
Qwen3.5-35B Unsloth Dynamic GGUFs Achieve SOTA Across Nearly All Quantisation Levels
-
Krasis: Hybrid CPU/GPU MoE Runtime Achieves 3,324 Tokens/Second Prefill on RTX 5080
-
Seco Launches Edge AI System-on-Module at Embedded World 2026
-
Snapdragon 8 Elite Gen 5 Powers Galaxy S26 Series With Enhanced On-Device AI
-
On-Device Function Calling in Google AI Edge Gallery
-
Apple: Python bindings for access to the on-device Apple Intelligence model
-
DeepSeek Releases DualPath: Addressing Storage Bandwidth Bottlenecks in Agentic Inference
-
DeepSeek Paper – DualPath: Breaking the Bandwidth Bottleneck in LLM Inference
-
Qwen3.5 Thinking Mode Can Be Disabled for Production Inference Optimization
-
Qwen3.5 Series Releases Comprehensive Model Lineup Across All Tiers
-
Red Hat Launches AI Enterprise for Hybrid AI Deployments
-
Qwen3.5-35B-A3B Emerges as Game-Changer for Agentic Coding Tasks
-
Kioxia Sampling UFS 5.0 Embedded Flash Memory for Next-Generation Mobile Applications
-
Meta's OpenClaw Release Raises Questions About Open-Source Model Safety and Alignment
-
Elastic Introduces Best-in-Class Embedding Models for High Performance Semantic Search
-
Elastic Introduces Best-in-Class Embedding Models for High Performance Semantic Search
-
Making Wolfram Technology Available as Foundation Tool for LLM Systems
-
Asus ExpertBook B3 G2 with 50 TOPS AI Sets New Enterprise Standard
-
DietPi Released a New Version v10.1
-
Google Open-Sources NPU IP, Synaptics Implements It for Hardware Acceleration
-
Ouro 2.6B Thinking Model GGUFs Released with Q8_0 and Q4_K_M Quantization
-
Ollama 0.17 Released With Improved OpenClaw Onboarding
-
Vellium v0.3.5: Major Writing Mode Overhaul and Native KoboldCpp Support
-
Claude Code Open – AI Coding Platform with Web IDE and Agents
-
[Release] Ouro-2.6B-Thinking: ByteDance's Recurrent Model Now Runnable Locally
-
Kitten TTS V0.8 Released: New State-of-the-Art Super-Tiny TTS Model Under 25 MB
-
SanityBoard Adds 27 New Model Evaluations Including Qwen 3.5 Plus, GLM 5, and Gemini 3.1 Pro
-
PaddleOCR-VL Now Integrated into llama.cpp for Multilingual OCR
-
NVIDIA Releases Dynamo v0.9.0: Infrastructure Overhaul With FlashIndexer and Multi-Modal Support
-
Free ASIC-Accelerated Llama 3.1 8B Inference at 16,000 Tokens/Second
-
LayerScale Launches Inference Engine Faster Than vLLM, SGLang, and TRT-LLM
-
Aegis.rs: Open Source Rust-Based LLM Security Proxy Released
-
OpenClaw Refactored in Go, Runs on $10 Hardware
-
AMD Announces Day 0 Support for Qwen 3.5 LLM on Instinct GPUs
-
Tailscale Releases New Tool to Prevent Sensitive Data Leakage to Cloud AI Services
-
Sarvam AI Launches Edge Model to Challenge Major AI Players with Local-First Approach
-
Alibaba's Qwen3.5-397B Achieves #3 Position in Open Weights Model Rankings
-
GLM-5 Technical Report: DSA Innovation Reduces Training and Inference Costs
-
Cloudflare Releases Agents SDK v0.5.0 with Rust-Powered Infire Engine for Edge Inference
-
Asus ExpertBook B3 G2 Laptop Features Ryzen AI 9 HX 470 CPU in 1.41kg Ultraportable Form Factor
-
Cohere Releases Tiny Aya: Efficient 3.3B Multilingual Model for 70+ Languages
-
ASUS Zenbook 14 Launches in India with AI-Capable Hardware, Starting at Rs 1,15,990
-
Qwen 3.5-397B-A17B Now Available for Local Inference with Aggressive Quantisation
-
Meet Sarvam Edge: India's AI Model That Runs on Phones and Laptops With No Internet
-
InitRunner: YAML-Based AI Agent Framework with RAG and Memory
-
GPU-Accelerated DataFrame Library for Local Inference Workloads
-
Alibaba Unveils Major AI Model Upgrade Ahead of DeepSeek Release
-
GPT-OSS 20B Now Runs 100% Locally in Browser via WebGPU
-
GNOME's AI Assistant Newelle Adds llama.cpp Support and Command Execution
-
NVIDIA's Dynamic Memory Sparsification Cuts LLM Inference Costs by 8x
-
MiniMax-M2.5 230B MoE Model Released with GGUF Support for Local Deployment
-
LLaDA2.1 Introduces Token Editing for Massive Speed Gains in Local Inference
-
GPT-OSS 120B Uncensored Model Released in Native MXFP4 Precision
-
ByteDance Releases Seed2.0 LLM with Complex Real-World Task Improvements
-
MiniMax Releases M2.5 Model with SOTA Coding and Agent Capabilities
-
WinClaw: Windows-Native AI Assistant with Office Automation
-
Ring-1T-2.5 Released with SOTA Deep Thinking Performance
-
MiniMax M2.5: 230B Parameter MoE Model Coming to HuggingFace
-
Ming-flash-omni-2.0: 100B MoE Omni-Modal Model Released
-
Student Releases Dhi-5B: Multimodal Model Trained for Just $1,200
-
GitHub Announces Support for Open Source AI Project Maintainers
-
Microsoft MarkItDown: Document Preprocessing Tool for LLMs
-
ByteDance Releases Seedance 2.0 AI Development Platform
-
OpenClaw with vLLM Running for Free on AMD Developer Cloud
-
Memio Launches AI-Powered Knowledge Hub for Android with Local Processing
-
Samsung's REAM: Alternative Model Compression Technique
-
New Header-Only C++ Benchmark Tool for Predictive Models on Raw Binary Streams
-
GLM-5 Released: 744B Parameter MoE Model Targeting Complex Tasks
-
Godot MCP Gives AI Assistants Full Access to Game Engine Editor
-
Arm SME2 Technology Expands CPU Capabilities for On-Device AI
-
Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts
-
DeepSeek Launches Model Update with 1M Context Window