Tagged "deployment-guide"
25 articles tagged deployment-guide, 31 March 2026 to 27 August 2026. Newest first.
-
Benchmarking Qwen3.8 27B Quantizations: 4-bit Shows Strong Performance, 1-bit Collapses
Detailed quantization benchmarks for Qwen3.8 27B reveal that 4-bit quantization maintains strong performance while 1-bit variants suffer significant degradation, providing practical guidance for local deployment scenarios.
-
How to Run a Local LLM With Ollama: 13 Steps, 90 Min
A comprehensive step-by-step guide for setting up and running local LLMs using Ollama, covering the entire process from installation to inference in approximately 90 minutes.
-
How To Run Kimi K3 Moonshot AI In Ollama
A tutorial covering both command-line and desktop application setup for running the Kimi K3 Moonshot model locally via Ollama.
-
Run Ollama Locally on Windows 11: Setup Guide
A practical walkthrough for deploying Ollama on Windows 11, lowering barriers for mainstream users to run local language models on consumer hardware.
-
LLM Wiki Implementation: Community Resource for Local Deployment
A new GitHub project provides comprehensive documentation and implementation guides for deploying language models locally, serving as a centralized wiki for the local LLM community.
-
How to Run an LLM Locally: 13 Steps, 90 Min
A comprehensive practical guide for setting up and running large language models on your own hardware in under 90 minutes. Perfect for beginners looking to get started with local LLM deployment.
-
Running OpenClaw with Ollama: Practical Guide to Local LLM Deployment
KDnuggets published a practical guide demonstrating how to run OpenClaw models with Ollama, providing step-by-step instructions for developers seeking to deploy specialized models locally.
-
Off-Grid AI Launches Emergency Preparedness Platform Powered by Local LLM Inference
Off-Grid AI demonstrates practical real-world deployment of local LLM inference by building an emergency response system that operates without cloud connectivity, eliminating latency and dependency issues.
-
The Hitchhiker's Guide to Agentic AI
A comprehensive guide published on arXiv provides foundational knowledge and practical insights for building and deploying agentic AI systems. This resource is essential reading for developers scaling from simple LLM inference to complex agent orchestration.
-
Running AI Locally, Part 2: From VMware Context to Hands-On Tools
The second installment in a series covering practical approaches to running AI workloads locally, including virtualization context and hands-on tooling recommendations for self-hosted inference.
-
A Guide on How to Run Nemotron 3 Super 120B Thinking on 2 Nvidia DGX Spark
Practical deployment guide for running NVIDIA's large reasoning model (120B parameters) on a two-node DGX Spark cluster with distributed inference techniques.
-
Developers Run Local LLMs on Windows 11
Guide demonstrating how developers can set up and run local LLMs directly on Windows 11, expanding accessibility of on-device AI inference beyond specialized Linux and Mac environments.
-
I Built a Bedside AI Assistant That Reads Me the News Without Touching the Cloud
A practical demonstration of building a completely local AI assistant that delivers personalized news without any cloud connectivity, showcasing real-world on-device LLM deployment techniques.
-
Best VPS for Ollama 2026 and Setup Guide
A comprehensive guide covering the best virtual private servers for running Ollama in 2026, including configuration recommendations and performance considerations for different use cases.
-
Qwen3-Coder-Next Local Deployment: Complete Developer Guide for 2026
A comprehensive guide for deploying Qwen3-Coder-Next, a state-of-the-art coding model optimized for local environments. The guide covers setup, configuration, and practical deployment strategies for developers.
-
How I Used a Local LLM to Organize the Store on My NAS
A practical case study demonstrating how local LLMs can be deployed on Network Attached Storage systems for practical applications like file organization and metadata management without cloud connectivity.
-
How to Run LLMs Locally on Your Laptop for Free: A Beginner's Guide
A comprehensive beginner's guide covering the fundamentals of running language models locally without cloud dependencies, including tools, hardware requirements, and practical setup instructions.
-
5 Things I Wish Someone Had Told Me Before I Tried Self-Hosting a Local LLM
A practical guide sharing key lessons learned from self-hosting local LLMs, covering pitfalls and best practices that can accelerate the learning curve for practitioners new to on-device inference. The article distills common mistakes and recommendations from real-world deployment experience.
-
Building a Remote-Accessible Local LLM Server on Raspberry Pi
A practical guide demonstrating how to deploy and access a local LLM server running on a Raspberry Pi from anywhere, combining edge deployment with convenient remote access.
-
10GB VRAM Local LLM: The Complete Setup Guide (2026)
A comprehensive guide covering practical methods to run capable local LLMs with just 10GB of VRAM, including quantization techniques, model selection, and optimization strategies for resource-constrained systems.
-
Gemma 4 Makes Local AI Agents Practical
Google's Gemma 4 26B model demonstrates significant capabilities for running autonomous AI agents on consumer hardware, marking a milestone for practical local LLM deployment.
-
April 2026 TLDR Setup for Ollama and Gemma 4 26B on a Mac mini
A community-contributed quick-start guide documents practical steps for deploying Gemma 4 on Mac mini hardware using Ollama, providing a reference implementation for local inference setup.
-
A Journey to a Reliable and Enjoyable Locally Hosted Voice Assistant
An in-depth guide documenting the development and deployment of a fully local voice assistant, covering the complete stack from speech recognition to language understanding and synthesis without cloud dependencies.
-
Local AI didn't replace my subscriptions, but it did take over these 6 tasks
A practical analysis of which specific workflows and tasks are most effective for local AI tools, helping practitioners identify high-impact use cases for self-hosted deployment.
-
Running AI on a Raspberry Pi, Part 2: Running AI on a Pi in Under 5 minutes
A practical guide demonstrating how to deploy and run AI models on Raspberry Pi hardware in minimal time, making edge inference accessible to developers and hobbyists.