Tagged "memory-bandwidth-optimization"
3 articles tagged memory-bandwidth-optimization, 26 February 2026 to 15 April 2026. Newest first.
-
Dynamic Expert Cache in llama.cpp Achieves 27% Faster Inference on Large MoE Models
A new optimization technique for llama.cpp improves CPU+GPU token generation speed by 27% on Qwen3.5-122B through dynamic expert caching, raising practical inference rates from 15 to 23 tokens per second.
-
Gemma 4 on Arm: Optimized On-Device AI for Mobile and Edge Deployment
Arm releases optimizations for Gemma 4 enabling efficient deployment on Arm-based processors for mobile devices and edge endpoints, bringing enterprise-grade AI to mobile platforms.
-
DeepSeek Paper – DualPath: Breaking the Bandwidth Bottleneck in LLM Inference
DeepSeek researchers present DualPath, a novel approach to address bandwidth limitations during LLM inference. This work tackles one of the primary performance bottlenecks in local and edge LLM deployment.