Tagged "inference-serving"
3 articles tagged inference-serving, 15 April 2026 to 11 September 2026. Newest first.
-
Cambricon Adapts DeepSeek-V4.1-Flash on vLLM Stack for Efficient Inference
Cambricon's Day-0 project successfully adapts DeepSeek-V4.1-Flash within the vLLM inference stack, demonstrating practical optimization of large open models for deployment. This work bridges advanced open models with production-grade serving infrastructure.
-
vLLM v0.27.0 Released with 561 Commits and Expanded Model Support
vLLM v0.27.0 brings significant improvements including Kimi K3 model support with full-stack integration, new kernel optimizations, and contributions from 242 developers. This major release advances the inference serving infrastructure for local and on-premises deployments.
-
DGX Spark Setup Guide: Running vLLM and PyTorch for Local LLM Inference Backend
A developer details their setup process for NVIDIA DGX Spark hardware running vLLM with Hugging Face models as a local API backend for education and analytics applications while maintaining privacy.