Tagged "inference-servers"
4 articles tagged inference-servers, 1 July 2026 to 19 September 2026. Newest first.
-
Ollama v0.34.3: Model Thinking Controls and Expanded Apple Silicon Support
Ollama releases v0.34.3 with new thinking level controls for models and expanded Apple Silicon support, including Nemotron H vision models on Mac hardware.
-
Benchmarking Local LLM Servers: Llama.cpp, Llamafile, LM Studio, and Ollama
A practical benchmark comparison of four major local LLM serving frameworks, measuring performance across speed, memory usage, and ease of deployment on consumer hardware.
-
HackerNoon Compares 7 Best Self-Hosted Inference Servers for Open-Source Models
A comprehensive 2026 comparison of leading self-hosted inference servers evaluates deployment options for running open-source models locally, covering performance, ease of use, and feature parity across major frameworks.
-
Article Compares Continuous and Static Batching in LLM Inference
A detailed analysis comparing continuous and static batching strategies for LLM inference, helping local deployment practitioners optimize throughput and latency trade-offs on resource-constrained hardware.