Tagged "batch-inference"
11 articles tagged batch-inference, 19 February 2026 to 9 September 2026. Newest first.
-
vLLM v0.29.0 Advances with Model Runner V2 as Default
vLLM's latest release makes Model Runner V2 the default for all models, featuring CUDA graph memory profiling and improved performance across deployment scenarios.
-
HackerNoon Compares 7 Best Self-Hosted Inference Servers for Open-Source Models
A comprehensive 2026 comparison of leading self-hosted inference servers evaluates deployment options for running open-source models locally, covering performance, ease of use, and feature parity across major frameworks.
-
Most People Use Ollama or llama.cpp for Local LLMs, but These Are the Tools I Switch to When It Gets Serious
An experienced practitioner compares advanced local LLM deployment tools beyond the popular Ollama and llama.cpp, highlighting specialized frameworks for production scenarios.
-
Bosgame Launches VTA-439 Mini PC with 86 TOPS for Practical Local AI
Bosgame has released the VTA-439 mini PC featuring 86 TOPS of AI compute in a compact form factor, specifically designed for accessible local LLM deployment and practical everyday use cases.
-
Linux 7.1-rc4 Released: Kernel Updates Relevant to Local LLM Inference
Latest Linux kernel release candidate includes optimizations impacting edge LLM deployment on commodity hardware. Performance improvements for memory management and CPU scheduling affect local inference efficiency.
-
Intel LLM-Scaler vLLM 0.14.0 Released With Official Arc Pro B70 Support
A new vLLM release brings production-ready support for Intel's Arc Pro B70 GPU, enabling optimized batch inference and high-throughput local LLM serving on Intel discrete graphics.
-
Free AI Video Clipper Using Scene and Speech-Based Segmentation
An open-source project provides local AI-powered video segmentation and automatic clipping based on scene changes and speech patterns. This tool demonstrates practical multimedia processing with on-device inference, eliminating cloud API dependencies.
-
LMCache Dramatically Accelerates LLM Inference on Oracle Data Science Platform
Oracle integrates LMCache, a cutting-edge prompt caching and KV cache optimization technique, into their cloud data science platform to accelerate LLM inference and reduce computational overhead.
-
AMD Launches Agent System Optimized for Local AI Inference With Ryzen and Radeon
AMD announces a new integrated system designed specifically for local AI workloads, combining Ryzen CPUs with Radeon GPU acceleration for efficient inference.
-
Intel Arc Pro B70 Workstation GPU Confirmed via vLLM AI Release Notes
Intel's Arc Pro B70 discrete GPU receives official support in vLLM release notes, expanding local LLM inference options for professional workstations. The BMG-G31 architecture targets professional AI computing workflows.
-
Hardware Economics Shift: DDR5 RDIMM Pricing Now Comparable to GPUs for Local Inference
Analysis shows DDR5 RDIMM memory costs have reached parity with high-end GPUs like RTX 3090s on a per-gigabyte basis, forcing local LLM builders to reconsider their hardware stacking strategies.