Tagged "llm-inference"
6 articles tagged llm-inference, 22 February 2026 to 29 May 2026. Newest first.
-
MediaTek Launches Dimensity 8550 4nm SoC with Integrated On-Device AI Focus
MediaTek has introduced the Dimensity 8550, a 4nm mobile system-on-chip featuring dedicated AI processing capabilities and support for Gemini Nano, enabling efficient on-device LLM inference on mid-range smartphones.
-
Supercharging LLM Inference on Google TPUs: Achieving 3X Speedups With Diffusion-Style Speculative Decoding
Google researchers have demonstrated 3x inference speedups on TPUs using diffusion-style speculative decoding, a novel optimization technique that could influence local inference strategies. The breakthrough shows how advanced decoding methods can dramatically reduce latency on specialized hardware.
-
How to Make SSE Token Streams Resumable, Cancellable, and Multi-Device
A practical guide to improving server-sent event (SSE) token streaming for LLM inference, enabling better user experiences with resumable downloads and multi-device support in local deployments.
-
Unweight: Lossless MLP Weight Compression for LLM Inference
Cloudflare Research presents a new lossless weight compression technique for MLP layers in language models, enabling faster inference and reduced memory footprint without quality degradation. A breakthrough for memory-constrained local deployments.
-
Cutile.jl Brings Nvidia CUDA Tile-Based Programming to Julia
Cutile.jl enables tile-based CUDA programming in Julia, offering improved GPU utilization and performance optimization capabilities for compute-intensive workloads including LLM inference.
-
AI Is Stress Testing Processor Architectures and RISC-V Fits the Moment
RISC-V architecture emerges as a compelling alternative for AI workloads as traditional processor designs face thermal and efficiency challenges under LLM inference loads, opening new possibilities for local deployment on custom silicon.