Tagged "performance"
-
AMD EPYC ZenDNN Accelerates llama.cpp Prompt Processing 4.5x
-
vLLM v0.27.0 Brings Major Performance Improvements and New Model Support
-
Benchmarking Local LLMs on Consumer Hardware: Real-World Performance Data
-
vLLM v0.27.0rc2 Release Candidate Available
-
llama.cpp Improves CUDA Performance with Kernel Fusion
-
Gainz.fast – Local Inference, Faster
-
Kioxia Is Coming for Samsung and SK Hynix With UFS 5.0 and PCIe 6.0 AI NAND
-
Squeezing Silicon Limits: Effective Strategies to Eliminate GPU Idle Time and Maximize GPU Utilization
-
Q4 vs Q6 vs Q8: The Quantization Decision Framework for Local LLMs
-
Testing Top Local LLMs Against ChatGPT and Claude Reveals Performance Gaps
-
Kioxia's UFS 5.0 Embedded Flash Enables Practical On-Device AI
-
Titan Transients and LLM Scalability
-
Gemma 4's Quantized Models Finally Made Local AI Practical in Homelab
-
Sol-5.6 and Opus 5 Models Demonstrate Strong One-Shot Game Performance
-
Testing Local LLMs on Real Tasks: Honest Assessment of Practical Utility
-
Gemini Notebook: On-Device AI in Action
-
Code Mode Can Help Smaller LLM Models
-
MSI Pro Max Edge AI+ Mini PC Runs 120B Local AI Models With 128GB RAM
-
llama.cpp b10075 Packs Four Local AI Runtime Upgrades
-
On-Device AI Ignites WAIC 2026: How Compute-in-Memory Chips Are Stuffing 100-Billion-Parameter LLMs Into Your Pocket
-
GPT-5.6 Sol vs. Claude Fable 5 in CNC Red Alert 2 Benchmark
-
My Local LLM Struggles With Big Questions—Here's What It's Actually Good At
-
AMD Ryzen 7 7700X3D Linux Performance Review
-
Python 3.15's Ultra-Low Overhead Interpreter Profiling Mode – Ken Jin's Blog
-
Cost vs. Accuracy in CursorBench 3.1: The Effect of Family and Spend
-
Study: Cerebellum Helps AI Ignore the Ordinary for More Efficient Computing
-
Apple's MacBook Lineup Overhaul Features M7 Chip for Enhanced Local AI
-
Local LLM Performance Gap With Frontier Models Smaller Than Expected
-
Theoretical Bottlenecks for Scaling LLM Inference to Achieve Higher Token per Second
-
Samsung Presents UFS 5.0 Storage Targeted at On-Device AI Performance
-
TriAttention Solves KV Cache Memory Bottleneck in Local LLM Inference
-
Apple's M7 Chip Delivers 56% Memory Bandwidth Increase for On-Device AI
-
The Mac Mini is the Best On-Device AI Computer You Can Buy: Here's Why
-
Samsung Unveils UFS 5.0 Storage Solution Optimized for On-Device AI
-
ORA: Smaller Models. Same Intelligence
-
DeepSWE Benchmark Updated with GLM 5.2 and Expanded Model Comparisons
-
FlashRT: Execution State for Latency-First AI
-
Ray Serve LLM Achieves 24x Performance Improvement in Distributed Inference
-
RTX 5080 and RTX 3090 Setup Achieves 80 Tok/s on Qwen 3.6 27B Q8
-
I Replaced Cloud LLMs with Local Models Running Off a Proxmox LXC, and the Performance Trade-Off Was Worth It
-
Longsys Redefines On-Device AI with Groundbreaking Edge Memory Solutions
-
Reducing GPU Costs for AI Inference: FP8, FP4, and vLLM Optimization Techniques
-
Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM Inference
-
Users Report Superior Performance Switching from LM Studio to llama.cpp
-
vLLM vs Ollama 2026: Performance Benchmark Reveals 9x Throughput Gap
-
llama.cpp Checkpoint Fix Accelerates Local Coding Agents
-
110 Tokens/Second on RTX 4070 Super with Qwen 3.6 35B
-
llama.cpp Adds Multi-Token Prediction, Doubles Qwen 3.6B Throughput for Local Inference
-
Bito's AI Architect Improves Claude Opus Task Success Rate by 35%
-
Orthrus Reshapes Economics of Local AI Inference with New Optimization Approach
-
ROCm 7.2.3 Delivers Performance Improvements Over 7.0.0 on AMD Radeon AI PRO
-
Open-Source Local LLM Emerges as Viable Cloud AI Competitor
-
Lucebox Brings Faster Local AI Inference to AMD Strix Halo
-
DFlash Speculative Decoding Delivers 8.5x Speed Improvement for LLM Inference
-
One LM Studio Setting Change Makes Local LLMs Competitive With Cloud Models
-
Lython: Experimental Python Compiler Toolchain Based on LLVM
-
Bun's Experimental Rust Rewrite Achieves 99.8% Test Compatibility on Linux
-
Sarvam Edge: Indian-Built AI Models Run Offline on Phones and Laptops Without Internet
-
Microsoft VibeVoice C++ Port Enables Local Voice AI on CPU and GPU Without Python
-
Google Accelerates Gemma 4 Inference Speed 3x With Multi-Token Prediction Drafters
-
llama.cpp Now Supports Multi-Token Prediction in Beta
-
NIST's CAISI Evaluation of DeepSeek V4 Pro Finds It On Par with GPT-5
-
Linux Setup for Local LLMs Takes Minutes Compared to Windows Hours
-
Hipfire: A Rust-Native AMD Inference Engine That Outperforms llama.cpp
-
Show HN: We built an OCR server that can process 270 dense images/s on a 5090
-
llama.cpp Merges Speculative Checkpointing for Major Inference Speed Boost
-
Gemma 4 Just Replaced My Whole Local LLM Stack
-
Sorting 1M u64 KV-Pairs in 20ms on i9-13980HX Using Branchless Rust Implementation
-
Intel's $949 GPU Has 32GB of VRAM for Local AI, but the Software Is Why Nvidia Keeps Winning
-
The 'Ollama' Tool Has Numerous Problems, and Some Argue That Llama.cpp Is Better
-
SigMap – Shrink AI Coding Context 97% with Auto-Scaling Token Budget
-
DFlash Doubles Token Generation Speed of Qwen3.5 27B on Mac M5 Max
-
Running Same Prompts Through Claude and Local LLM Revealed Unexpected Results
-
Google Gemma 4 Delivers Exceptional Speed and Accuracy for Local Inference
-
DFlash Speculative Decoding Achieves 3.3x Speedup on Apple Silicon
-
Speculative Decoding Made My Local LLM Actually Usable
-
MemPalace, the Highest-Scoring AI Memory System Ever Benchmarked
-
Your Next Assistant is Your PC: How On-Device AI is Transforming Work, One Workflow at a Time
-
CricketBrain: Neuromorphic Signal Processor in Rust (0.175us/step, 944 bytes)
-
Microsoft Quantum Development Kit Ported to Rust: 100x Faster and Smaller
-
Ollama Gets Blazing Fast on Macs with Full MLX Support and 2× Speedups
-
Gemma 4 31B Outperforms GLM 5.1 in Real-World Testing
-
Google's TurboQuant Shows Memory Constraints Remain Critical for Local LLM Inference
-
TurboQuant: Understanding the Quantization Breakthrough
-
Mixed KV Cache Quantization: Performance Risks and Pitfalls
-
Linux Significantly Outperforms Windows for Local LLM Inference
-
TurboQuant KV Cache Compression Achieves 22.8% Faster Decoding at 32K Context
-
RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
-
Qwen 3.5 27B Achieves 1.1M Tokens/Second on B200 GPUs with Optimized vLLM Config
-
Liquid AI's LFM2-24B Achieves 50 Tokens/Second in Web Browser via WebGPU
-
Google's TurboQuant: The Unsexy AI Breakthrough Worth Watching
-
Google TurboQuant: Extreme Compression for Local LLM Deployment
-
Rust Project Perspectives on AI
-
Multi-Token Prediction support coming to MLX-LM for Qwen 3.5
-
Snapdragon 8 Elite Gen 5 Hands the Galaxy S26 the AI Upgrade We've Been Waiting For
-
P-EAGLE: Faster LLM Inference with Parallel Speculative Decoding in vLLM
-
Memory Should Decay: Implementing Temporal Memory Decay in Local LLM Systems
-
3-Path Agent Memory: 8 KB Recurrent State vs. 156 MB KV Cache at 10K Tokens
-
Quantization Explained: Q4_K_M vs AWQ vs FP16 for Local LLMs
-
Cutile.jl Brings Nvidia CUDA Tile-Based Programming to Julia
-
FreeBSD 14.4 Released: Implications for Local LLM Deployment
-
Mojo: Creating a Programming Language for an AI World with Chris Lattner
-
The Emerging Role of SRAM-Centric Chips in AI Inference
-
Apple M5 Pro and M5 Max: 4× Faster LLM Processing
-
Qwen 3.5 vs Qwen 3 Benchmark Analysis: Generational Performance Improvements Visualized
-
Accuracy vs. Speed in Local LLMs: Finding Your Sweet Spot
-
Snapdragon 8 Elite Gen 5 Powers Galaxy S26 Series With Enhanced On-Device AI
-
Qwen3.5 122B Achieves 25 tok/s on 72GB VRAM Setup
-
Qwen 3.5 MoE Delivers 100K Context Window at 40+ TPS on RTX 5060 Ti
-
Qwen 3.5 Underperforms on Hard Coding Tasks—APEX Benchmark Analysis
-
New Era of On-Device AI Driven by High-Speed UFS 5.0 Storage
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
I Thought I Needed a GPU to Run AI Until I Learned About These Models
-
Taalas Etches AI Models onto Transistors to Rocket Boost Inference
-
24 Simultaneous Claude Code Agents on Local Hardware