Tagged "apple-silicon"
-
vLLM-iOS Achieves 88% Faster Multi-Agent Inference on Mobile Devices
-
JetBrains Releases Junie Local: On-Device Coding Agent for macOS
-
Llama.cpp Build 10620: Continued Optimization for Local Inference
-
Qwen 3.6 Now Easier to Run Locally on Mac with JetBrains Integration
-
llama.cpp Build b10581 Adds DSpark Support for Faster Local Inference
-
Ollama v0.33.0 Release Candidate Adds Claude Desktop Integration and Performance Improvements
-
Llama.cpp Release b10485: GGML Sync with Platform-Specific Optimizations
-
DeepSeek V4 Flash Shrunk to 57GB for Local macOS Inference with Compiler Generation
-
The Qwen MLX Challenge
-
Llama-macOS – Agentic and MCP Native macOS Front End for Llama.cpp
-
Show HN: I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac
-
Ollama Adds Qwen 3.8 27B with Optimised Apple Silicon Support
-
Meta's Muse Glimmer on ExecuTorch Enables Fast On-Device Agentic AI
-
Ollama Adds Qwen 3.8 27B with Apple Silicon Optimizations
-
Local Model Performance Benchmarks on MacBook Pro M5 Max: Real-World Inference Metrics
-
Apple's On-Device AI Strategy Focuses on Privacy and Latency, Not ChatGPT Competition
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Local Agent Deployment
-
Meta's Muse Glimmer Now Available Across All Platforms in Ollama
-
Meta's Muse Glimmer Now Available Across All Platforms in Ollama
-
Meta's Muse Glimmer Now Available Across All Platforms via Ollama
-
MacPaw and Liquid AI: Complete On-Device AI Stack for macOS
-
Meta Releases Muse Glimmer: 30B Open-Source LLM for Local Deployment
-
Muse Glimmer Now Available on Ollama – Meta's Open Multimodal Agent Model
-
ShoutFlow Launches Pay-Once, On-Device AI Dictation App for the Mac
-
Llama.cpp Fixes Metal NORM Operations for Apple Silicon
-
MacPaw Partners With Liquid AI to Deploy On-Device AI Across Mac Ecosystem
-
llama.cpp Build b10301: CUDA Optimization and Compiler Warning Fixes
-
Ollama v0.32.6: Faster Apple GPU Inference with Speculative Decoding
-
PrismML's Bonsai 27B Brings On-Device AI to Apple iPhone 17 Pro
-
Apple's Hardware Is Ready for On-Device AI and PrismML Just Delivered a Real Breakthrough
-
How Much Does a Local LLM Actually Cost to Run? Energy Costs Measured on Apple Silicon
-
Wisprkey – 100% Free and Local Voice Typing for Mac
-
Odysseus - PewDiePie's Self-Hosted AI Finally Runs Fast on Mac
-
Apple in Early Talks With PrismML on AI Compression Tech
-
Apple Boosts On-Device AI, Partners With PrismML to Enable Running Large Models Locally on iPhone
-
Show HN: AITerm – a macOS Terminal with an AI Command Loop and a Safety Gate
-
Apple's M6, M7, and M8 Chip Roadmap Shifts Focus Toward AI
-
Apple's Failed Self-Driving Car Program Left a Legacy of Powerful AI Chips
-
Running Local AI on Mac With Home Assistant Integration
-
Apple Explores Running Larger AI Models on iPhone with On-Device Compression
-
Ollama Runs 32B Local AI Models on a $599 Mac via Quantization for Free
-
Apple's MacBook Lineup Overhaul Features M7 Chip for Enhanced Local AI
-
Ollama's New MLX Engine Delivers Significant Performance Gains on Mac
-
Apple Updates Creator Studio with AI Video Editing, Image Generation, and Logic Pro Enhancements
-
Asahi Linux 7.1 Progress Report
-
You Can Now Run Max AI Models on Apple Silicon
-
Apple's M7 Chip Delivers 56% Memory Bandwidth Increase for On-Device AI
-
The Mac Mini is the Best On-Device AI Computer You Can Buy: Here's Why
-
Mac Mini Emerges as Top Choice for Local On-Device AI Deployment
-
Mac Mini Positioned as Premier On-Device AI Computer for Local LLM Inference
-
MCP Server Enables Claude to Automate Mac Tasks and Self-Correct
-
Apple unveils Core AI for on-device generative models
-
Most People Use Ollama or llama.cpp for Local LLMs, but These Are the Tools I Switch to When It Gets Serious
-
Show HN: 11 Model Families Ported to Apple's CoreAI On-Device Framework
-
Apple Unveils AFM 3 Core Advanced with 20 Billion Parameters for On-Device AI
-
Apple Rebuilt Its On-Device AI Stack at WWDC 2026
-
Apple Enhances Siri With On-Device AI for Faster, Private Voice Responses
-
Google AI Edge Gallery Launches on macOS With Offline Gemini Models
-
Apple iPad Air with M4 Chip Drops to $1349; Powerful On-Device LLM Inference Now More Accessible
-
Google Launches AI Edge Gallery on macOS for Running Gemini Models Locally
-
Google Launches AI Edge Gallery on macOS for Running Gemini Models Locally
-
What Apple Knows About AI That Silicon Valley Won't Admit
-
Apple Doubles Down on On-Device AI at WWDC 2026, Setting Privacy-First Strategy
-
Samsung's Exynos 2800 Brings HBM Memory to Mobile AI, Enabling Faster Local Model Inference
-
Apple's 2026 AI Strategy Prioritizes On-Device Model Deployment
-
Why AI Hardware Is a Chip Layer Problem
-
AMD Unveils Ryzen AI Halo Developer Platform for On-Device AI Workloads
-
M5 Max MacBook Runs Local Large Language Models Efficiently
-
Auditing Apple's DifferentialPrivacy.framework: Bugs, Misconfig, Practical Risks
-
Samsung's Exynos 2800 Brings Significant On-Device AI Capabilities
-
AMD's Lemonade SDK Advances macOS Support for Local AI Inference with ROCm 7.13
-
Apple's M5 MacBook Air Advances On-Device AI with Redesigned Hardware
-
Running AI Models Locally on M4 Processors with 24GB Memory
-
Lucebox Brings Faster Local AI Inference to AMD Strix Halo
-
Cotypist – AI Autocomplete for Mac
-
Mlx-serve: Run LLMs Natively on Your Mac
-
Perplexity Brings On-Device AI Workflow to Macs with 'Personal Computer' Feature
-
On-Device AI Market Poised for Explosive Growth as Major Tech Companies Invest Heavily
-
Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Google's Gemma 4 Brings Powerful AI Capabilities to Phones and Laptops
-
Show HN: Phonetic Formatter – Offline English Text to IPA on iPhone and iPad
-
Llama 4 Scout on MLX: The Complete Apple Silicon Guide (2026)
-
Running Gemma 4 on an iPhone 13 Pro
-
DFlash Doubles Token Generation Speed of Qwen3.5 27B on Mac M5 Max
-
oMLX Framework Implements DFlash Attention for Optimized Inference
-
MiniMax M2.7 Achieves SOTA Performance Under 64GB on Mac with TQ Quantization
-
DFlash Speculative Decoding Achieves 3.3x Speedup on Apple Silicon
-
Parakeet Streaming ASR on Apple Silicon via CoreML
-
AIYO Wisper: Local Voice-to-Text for macOS Using WhisperKit
-
On-Device Apple Intelligence Vulnerable to Prompt Injection Attacks
-
Running a 1.7B Parameters LLM on an Apple Watch
-
Comprehensive Benchmark: 37 LLMs Tested on MacBook Air M5 With Open-Source Tool
-
Google Launches Offline AI Dictation App for iOS with Gemma
-
Apple Brings Enhanced On-Device AI Features to iPhone
-
Real-time Multimodal AI on Apple Silicon: Gemma E2B Demo Shows Practical Edge Deployment
-
Gemma 4 26B MoE Emerges as Optimal All-Around Local Model for Consumer Hardware
-
Ollama Gets Blazing Fast on Macs with Full MLX Support and 2× Speedups
-
Mixed Precision Quantization on MLX with TurboQuant Implementation
-
Kokoro TTS Achieves 20× Realtime Speed on CPU-Only On-Device Inference
-
Samsung Launches Galaxy Book6 Series with NVIDIA RTX 5070 and On-Device AI
-
Gemma 4 KV Cache Memory Issues Fixed in llama.cpp
-
Gemma 4 26B A4B Outperforms Qwen 3.5 35B on Apple Silicon
-
Apfel – The Free AI Already on Your Mac
-
Google Gemma 4 Released with GGUF Quantizations
-
Gemma 4 Makes Local AI Agents Practical
-
April 2026 TLDR Setup for Ollama and Gemma 4 26B on a Mac mini
-
TinyGPU Adds Mac Support for External Nvidia GPU Acceleration
-
Apple Silicon Macs Run Local AI Faster with Ollama's New MLX Support
-
Is Anyone Working on an AI Operating System?
-
Ollama Adopts Apple's MLX Framework for Faster Local AI on Mac
-
Select the Right Hardware for Your Local LLM Deployment with This Online Guide
-
M5 Max Delivers 1.7x Faster Inference Than M3 Max on Qwen 3.5 Models
-
TurboQuant KV Cache Compression Achieves 22.8% Faster Decoding at 32K Context
-
Qwen3 512k Context via TurboQuant on Mac mini
-
Apple Gets Full Gemini Access and Uses Distillation to Build Lightweight On-Device AI
-
mlx-Code: Run Claude Code Locally with MLX-LM
-
RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
-
Liquid AI's LFM2-24B Achieves 50 Tokens/Second in Web Browser via WebGPU
-
Apple Plans Slimmed-Down Gemini Models for Local iPhone AI Features
-
Running an Open-Weight LLM Locally on an Apple Watch
-
Ultra-Large 400B-Class LLM Runs on iPhone in Test
-
Ditching Paid AI Services: Building Self-Hosted LLM Solutions as ChatGPT, Claude, and Gemini Alternatives
-
Multi-Token Prediction support coming to MLX-LM for Qwen 3.5
-
DeepSeek R1 RTX 4090 vs Apple M3 Max: Benchmark & Performance Guide
-
Apple M5 Max 128GB real-world performance benchmarks for local inference
-
NVIDIA Nemotron 3 Nano 4B Enables On-Device Inference Directly in Web Browsers via WebGPU
-
Dictare – Open-source Voice Layer for AI Coding Agents (100% Local)
-
Startup Transforms Mac Mini Into Full-Powered AI Inference System With External GPU
-
Local LLMs on Apple Silicon Mac 2026: M1 M2 M3 Guide
-
Apple M5 Max 128GB Benchmark Results for Local LLM Inference
-
Experiment: 0.8B Model Self-Improvement on MacBook Air Yields Surprising Results
-
M5 Max and M5 Ultra Chipsets Demonstrate Significant Bandwidth Improvements for Local LLM Inference
-
Apple Launches MacBook Neo with A18 Pro Chip for Affordable Local AI Inference
-
Real-World Qwen 3.5 9B Agent Performance on M1 Pro Validates Edge Deployment
-
MediaTek Advances Omni Model for Efficient Smartphone Inference
-
Apple Unveils MacBook Pro with M5 Pro and M5 Max Featuring On-Device AI
-
Apple Unveils MacBook Pro With M5 Pro and M5 Max for On-Device AI
-
Apple M5 Pro and M5 Max: 4× Faster LLM Processing
-
AMD Launches Copilot+ Desktop Chips to Compete in On-Device AI Market
-
Apple M4 iPad Air Targets AI Users with Double M1 Speed Performance
-
Alibaba's Qwen 3.5 Small Model Runs Directly on iPhone 17
-
VibeWhisper – macOS Voice-to-Text with 100% Local Processing Option
-
Running Local AI Models on Mac Studio 128GB: 4B, 20B & 120B Tested
-
Apple Neural Engine Reverse-Engineered for Local Model Training on Mac Mini M4
-
Show HN: Caret – Tab to Complete at Any App on Your Mac
-
Apple: Python bindings for access to the on-device Apple Intelligence model
-
Researchers Develop Persistent Memory System for Local LLMs—No RAG Required
-
Apple Accelerates U.S. Manufacturing with Mac Mini Production
-
Qwen3-Code-Next Proves Practical for Local Development: Real-World Coding Tasks on Mac Studio
-
Nvidia Could Launch Its First Laptops With Its Own Processors
-
AI-Powered Reverse-Engineering of Rosetta 2 for Linux
-
Apple Researchers Develop On-Device AI Agent That Interacts With Apps for You
-
PaddleOCR-VL Now Integrated into llama.cpp for Multilingual OCR
-
Complete Offline AI System: Voice Control and Smart Home via Local LLM and Radio Without Internet
-
GPT4All Replaces Ollama On Mac After Quick Trial
-
Meet Sarvam Edge: India's AI Model That Runs on Phones and Laptops With No Internet
-
Sourdine: Open-Source macOS App for 100% Local AI Transcription
-
MiniMax-M2.5 230B MoE Model Released with GGUF Support for Local Deployment
-
MiniMax Releases M2.5 Model with SOTA Coding and Agent Capabilities