Tagged "multi-model-serving"
4 articles tagged multi-model-serving, 18 March 2026 to 8 August 2026. Newest first.
-
Llama.cpp Adds LRU Scheduler for Multi-Model Serving
Llama.cpp B10313 introduces an LRU (Least Recently Used) scheduler for its router, enabling better resource management when serving multiple models simultaneously. This enhancement improves request handling and model eviction policies for local inference servers.
-
Triton Control: Open-Source Control Plane for Nvidia Triton on Kubernetes
A new open-source project providing a control plane for managing Nvidia Triton Inference Server deployments on Kubernetes, streamlining multi-model serving infrastructure.
-
Can IBM's RITS Platform and vLLM Reset the Bar for Enterprise AI Access?
IBM's RITS platform combined with vLLM is positioning local and on-premises LLM deployment as a viable enterprise alternative, with improved accessibility and control.
-
Custom GPU Multiplexer Achieves 0.3ms Model Switching on Legacy Hardware
A developer built a custom Linux kernel module that multiplexes six GPUs through a single PCIe slot, enabling model hot-swapping in under 0.3 milliseconds using repurposed Bitcoin mining hardware.