Tagged "performance-benchmark"
13 articles tagged performance-benchmark, 22 February 2026 to 5 October 2026. Newest first.
-
GLM-5.3-Flash: 13-Step Guide to API vs Self-Hosted Deployment
A practical deployment guide comparing API-based and self-hosted options for GLM-5.3-Flash, covering the complete setup process for local inference across different hardware configurations.
-
Prefill Concurrency in SGLang: Consistent TTFT Under Multi-Tenant Load
SGLang's new prefill concurrency feature addresses head-of-line blocking in multi-tenant LLM serving, maintaining consistent time-to-first-token even under variable request loads. This improves the viability of shared local LLM deployments.
-
Husky: Model-Specific Inference Engine Achieves 4.5x Speedup Over Apple MLX
A new inference engine optimised for Apple Silicon demonstrates dramatic performance improvements over existing solutions, achieving up to 4.5x faster inference than MLX for specific model architectures.
-
What Actually Fits on Dual RTX 3090s: Qwen 27B and the KV Cache Math
Why a 27B model holds 262K context on 48GB — 48 of its 64 layers have no KV cache at all — and what decode speed you should honestly expect from two 3090s.
-
Ollama v0.32.15: Time-to-First-Token Cut in Half with Metadata Caching
Ollama's latest release dramatically improves time-to-first-token by caching resolved model metadata, reducing startup latency from 995ms to 524ms in benchmarks.
-
Nvidia Boosts Token Throughput 5x With Software Optimizations, Reshaping AI Inference Economics
Nvidia achieves a 5x improvement in token throughput for LLM inference through software optimizations in vLLM, dramatically improving the economics of local and self-hosted model deployment. This breakthrough demonstrates that software efficiency can match or exceed hardware upgrades for inference workloads.
-
Article Compares Continuous and Static Batching in LLM Inference
A detailed analysis comparing continuous and static batching strategies for LLM inference, helping local deployment practitioners optimize throughput and latency trade-offs on resource-constrained hardware.
-
Google's Gemma AI Runs Locally on a $300 Mini PC, and It Replaced ChatGPT for More Than Expected
A real-world deployment report showing that Google's Gemma model, running on modest consumer hardware, can handle practical AI tasks that previously required cloud-based services.
-
Comparison of Two Frameworks: 40% Token Efficiency Improvement
A detailed comparison shows that Wasp achieves the same application functionality with 2.5M tokens versus 4.0M tokens in Next.js, highlighting the importance of framework choice for optimizing local LLM inference costs.
-
FlashAttention-4 Delivers 2.7x Faster Inference with 1613 TFLOPs/s on Blackwell GPUs
FlashAttention-4, written in Python, achieves near-matmul-speed attention kernels with 71% GPU utilization on NVIDIA B200, delivering 2.1-2.7x faster inference than Triton. This breakthrough optimizes the attention bottleneck for local LLM deployment.
-
Qwen 3.5 27B on Dual RTX 3090s: 170K Context Holds, 100+ Tokens/s Claim Disputed
A widely shared r/LocalLLaMA video reported Qwen 3.5 27B running at 100+ tokens/second decode with a 170K context window on dual RTX 3090s. The context claim holds and is in fact understated — 262K fits. The decode figure is contradicted by independent benchmarks measuring 41.4 t/s on the same model and hardware, and the original video has never been independently verified.
-
How AI is Redefining Price and Performance in Modern Laptops
Modern laptops are increasingly optimized for local AI inference through improved hardware accelerators, specialized chips, and software frameworks. This shift is creating more capable platforms for running quantized language models without cloud dependency.
-
Asus ExpertBook B3 G2 with 50 TOPS AI Sets New Enterprise Standard
Asus announces the ExpertBook B3 G2, an enterprise laptop featuring 50 TOPS of AI compute, establishing new performance benchmarks for business-class local inference devices.