Tagged "self-hosted-llm"
17 articles tagged self-hosted-llm, 21 February 2026 to 1 October 2026. Newest first.
-
Ollama for Managing Local Language Models: A KDnuggets Cheat Sheet
KDnuggets publishes a comprehensive cheat sheet for Ollama, the popular tool for running and managing language models locally, providing practical guidance for developers deploying LLMs on-device.
-
Open-Weight AI on Kubernetes: Comparing vLLM and KubeAI for Local Deployment
A comprehensive guide examines vLLM and KubeAI as competing solutions for deploying open-weight models on Kubernetes clusters, helping teams choose the right inference framework for self-hosted LLM workloads.
-
Netflix Details Its In-House LLM Serving Platform with Triton and vLLM
Netflix has published details about its production LLM serving infrastructure, combining NVIDIA Triton and vLLM for efficient model deployment. This real-world case study demonstrates battle-tested patterns for scaling LLM inference at enterprise scale.
-
Multiverse Computing's CompactifAI Models Now Fully Compatible with Intel Xeon 6 Processors
All CompactifAI optimised models have achieved compatibility with Intel Xeon 6 processors, enabling efficient inference on enterprise server hardware and expanding deployment options for self-hosted local LLM infrastructure. This compatibility expands the practical deployment platforms for optimised models.
-
Open-Source AI on OCI: Serving LLMs on Kubernetes with vLLM, Qdrant, and Terraform
Oracle publishes a comprehensive guide for deploying open-source LLMs on Kubernetes clusters using vLLM for inference optimization, Qdrant for vector search, and Terraform for infrastructure as code. This practical approach enables scalable self-hosted LLM deployments on enterprise infrastructure.
-
Nvidia Boosts Token Throughput 5x With Software Optimizations, Reshaping AI Inference Economics
Nvidia achieves a 5x improvement in token throughput for LLM inference through software optimizations in vLLM, dramatically improving the economics of local and self-hosted model deployment. This breakthrough demonstrates that software efficiency can match or exceed hardware upgrades for inference workloads.
-
If You Can Write Acceptance Criteria, You Can Write an AI Routing Policy
An article demonstrating how acceptance criteria frameworks can be applied to define AI routing policies for local multi-model deployments. This provides practical guidance for orchestrating multiple LLMs in self-hosted environments.
-
Using mirrord to Verify AI-SRE Fixes Against Staging Clusters
MetalBear demonstrates practical SRE techniques using mirrord to test AI-powered infrastructure fixes against staging environments without full redeployment. This approach reduces friction when deploying local and self-hosted AI systems.
-
DeepSWE Benchmark Updated with GLM 5.2 and Expanded Model Comparisons
The DeepSWE software engineering benchmark has been updated with new results for GLM 5.2 and other models, providing fresh performance data for evaluating local LLM deployments on code generation tasks. This comprehensive benchmark helps practitioners select appropriate models for their infrastructure.
-
Show HN: SpadeBox – Sandboxed tools and JavaScript runtime for AI agents
SpadeBox provides a sandboxed JavaScript runtime environment specifically designed for local AI agent execution. Enables secure tool use and code execution without compromising the host system.
-
Due to DMA, Siri AI Delayed in EU for iOS 27 and iPadOS 27
Apple announced that its new on-device AI features for Siri will be delayed in the European Union due to compliance requirements under the Digital Markets Act.
-
Running Espressif's OpenClaw-Inspired AI Agent on ESP32 with Self-Hosted LLM Works in Practice
A developer successfully deployed an AI agent on ESP32 microcontroller hardware using a self-hosted LLM backend, demonstrating the feasibility of edge AI at the microcontroller level. This achievement showcases practical integration of local inference across diverse hardware platforms.
-
Running DeepSeek R1 Locally: Your Complete Setup Guide
SitePoint publishes a comprehensive guide for setting up and running DeepSeek R1 on local hardware, covering installation, configuration, and optimization tips for self-hosted inference.
-
I Built a Local AI Stack with 5 Docker Containers, and Now I'll Never Pay for ChatGPT Again
A practical guide demonstrating how to assemble a complete local AI stack using five Docker containers, eliminating dependency on cloud API services. This showcases end-to-end self-hosted LLM infrastructure design.
-
Self-Hosted LLM Elevates Personal Knowledge Management Systems to New Levels
A practitioner shares how deploying a self-hosted LLM transformed their personal knowledge management workflow, highlighting practical benefits and implementation strategies for local AI deployment.
-
LoKI – Local AI Assistant for Linux and WSL
LoKI is a new local AI assistant purpose-built for Linux and Windows Subsystem for Linux environments, providing self-hosted conversational capabilities without external API dependencies.
-
Search and Analyze Documents from the DOJ Epstein Files Release with Local LLM
A practical demonstration of deploying local LLMs for large-scale document analysis, using the newly released DOJ files as a case study. This project showcases real-world applications of self-hosted language models for sensitive document processing.