Tagged "vllm"
-
Native vLLM and ROCm 7.15 Support for AMD RDNA2 GPUs on Windows
-
What If Local LLM Inference Is Using Consumer Hardware Wrong?
-
Qwen3.8-27B Surpasses 1 Million Downloads, Overseas Developers Race to Maximize Local Deployment
-
HackerNoon Compares 7 Best Self-Hosted Inference Servers for Open-Source Models
-
Hugging Face State of Open Models: Summer 2026 Observations
-
7 Best Self-Hosted Inference Servers for Open-Source Models Compared (2026)
-
vLLM v0.27.0 Brings Major Performance Improvements and New Model Support
-
vLLM v0.27.0 Released with 561 Commits and Expanded Model Support
-
vLLM v0.27.0 Released with Major Kernel Improvements and New Model Support
-
vLLM v0.27.0 – Kimi K3 Support and 561 Commits from 242 Contributors
-
DEF CON 34 Exposes 10 Critical Vulnerabilities in Local AI Systems
-
vLLM v0.27.0rc2 Release Candidate Available
-
vLLM v0.27.0rc1: Latest Release Candidate for High-Performance Inference
-
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems
-
LFM2.5-2.6B: On-Device Agentic Model With 128K Context and Tool Calling
-
AMD's MI355X Undercuts Nvidia's B300 on Cost to Run China's Kimi K3
-
Squeezing Silicon Limits: Effective Strategies to Eliminate GPU Idle Time and Maximize GPU Utilization
-
GPU Half-Idle: The Hundred-Billion-Dollar Race to Squeeze 10x Efficiency from Silicon
-
Ask HN: What are you using for LLM inference in production?
-
Titan Transients and LLM Scalability
-
Open-Weight AI on Kubernetes: Comparing vLLM and KubeAI for Local Deployment
-
K3 Model Achieves 20 Tokens/Second on 80x RTX 5090 Cluster
-
Netflix Details Its In-House LLM Serving Platform with Triton and vLLM
-
Ruff v0.16.0: 413 Default Rules for Code Quality in AI Development
-
Don't Buy an Uncensored AI on a Flash Drive: What You Can Do Instead
-
AI Inference Costs: Build vs. Rent
-
NVIDIA's On-Device AI Gains Japan's Manufacturing Giants' Backing
-
Open-Source AI on OCI: Serving LLMs on Kubernetes with vLLM, Qdrant, and Terraform
-
Nvidia Boosts Token Throughput 5x With Software Optimizations, Reshaping AI Inference Economics
-
Show HN: Call to Control AI Agents via the Web
-
Show HN: OpenVole 4.5 Is Out
-
Intel-Scaler-vLLM 0.21.0-b1 Brings Latest Features for vLLM on Intel GPUs
-
Exploiting Sparsity for Long Context Inference: Million Token on Commodity GPUs
-
The Triage Is the Product: Running AI Agents Against Ethereum's Protocol Code
-
Critical GPU Memory Leak Vulnerability Discovered in vLLM
-
Edge AI Transformation Coming to Creative Production Workflows
-
Local LLM Performance Gap With Frontier Models Smaller Than Expected
-
Article Compares Continuous and Static Batching in LLM Inference
-
GEEKOM A9 Max Delivers 32GB RAM and Native LLM Support in Compact Form Factor
-
PewDiePie's Open-Source AI Workspace Gains Traction as Practical Local Deployment Platform
-
Liquid AI Ships LFM2.5-230M with Broad Framework Support for On-Device Inference
-
NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
-
Intel Core Ultra X7 Panther Lake Performance Benchmarked on Linux
-
Companies Question Cost of AI as Token Maximization Spending Adds Up
-
Most People Use Ollama or llama.cpp for Local LLMs, but These Are the Tools I Switch to When It Gets Serious
-
Why Tool Calling is More Important Than Model Size for Local LLMs
-
Docfai.app Launches With Free Trial for Local Document Processing
-
Ask HN: What Problem Did AI Create at Your Company That Didn't Exist Before?
-
AMD PACE: New vLLM Plugin Enables Efficient CPU-Based Inference
-
vLLM vs Ollama 2026: 793 vs 41 TPS Performance Benchmark
-
Prefill Once, Fan Out: KV Snapshot Sharing for Multi-Agent LLM Pipelines
-
TokenTamer: A Proxy That Reduces LLM Token Usage Through Context Compression
-
DockSec: Open-Source AI-Powered Container Security Scanner for Self-Hosted Deployments
-
NVIDIA Unveils First PC Chips at Computex 2026; CEO Jensen Huang Details New Hardware
-
Reducing GPU Costs for AI Inference: FP8, FP4, and vLLM Optimization Techniques
-
NVIDIA Levels Up Local AI Agents Across RTX PCs and DGX Spark
-
Microsoft and Nvidia to Unveil First Windows PCs with Nvidia CPUs and AI Capabilities
-
Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM Inference
-
vLLM vs Ollama 2026: Performance Benchmark Reveals 9x Throughput Gap
-
How to Self-Host LibreChat with Docker
-
AMD Unveils Ryzen AI Halo Developer Platform for On-Device AI Workloads
-
Deploying Hermes Agent for Free on AMD Developer Cloud with Open Models and vLLM
-
Intel llm-scaler-vllm 1.4 Released With Updated Components and Arc Pro B70 Support
-
Local LLMs Offer Unique Advantages That Cloud AI Services Cannot Match
-
The AI Layoff Receipts: Market Consolidation Accelerates Open-Source Model Adoption
-
AMD's Lemonade SDK Advances macOS Support for Local AI Inference with ROCm 7.13
-
Linux 7.1-rc4 Released: Kernel Updates Relevant to Local LLM Inference
-
Google Limits Gemini Intelligence to New Flagships—Hardware Requirements for Local Deployment
-
Orthrus Reshapes Economics of Local AI Inference with New Optimization Approach
-
SynapseKit: A New Production Framework for Deploying LLMs
-
Local LLM Persistent Context Prevents Repetitive Mistakes
-
Lucebox Brings Faster Local AI Inference to AMD Strix Halo
-
AMD's vLLM-ATOM Plugin Supercharges DeepSeek-R1 and Kimi-K2 Inference on MI350/MI400
-
DFlash Speculative Decoding Delivers 8.5x Speed Improvement for LLM Inference
-
Supercharging LLM Inference on Google TPUs: Achieving 3X Speedups With Diffusion-Style Speculative Decoding
-
Ubuntu is Going All In on Generative AI and Other Linux Distros Might Follow
-
Linux Setup for Local LLMs Takes Minutes Compared to Windows Hours
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
Can IBM's RITS Platform and vLLM Reset the Bar for Enterprise AI Access?
-
Elastic KV Cache Memory Breakthrough Enables Efficient Bursty LLM Serving and GPU Sharing
-
Build Your Own Local AI Stack with 5 Docker Containers and Eliminate ChatGPT Subscriptions
-
I Built a Local AI Stack With 5 Docker Containers, and Now I'll Never Pay for ChatGPT Again
-
Intel LLM-Scaler vLLM 0.14.0 Released With Official Arc Pro B70 Support
-
AI Quota Inflation Is No Token Effort. It's Baked In
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
Unweight: Lossless MLP Weight Compression for LLM Inference
-
Sorting 1M u64 KV-Pairs in 20ms on i9-13980HX Using Branchless Rust Implementation
-
Researcher Discovers 221 Bugs in vLLM Stemming From Single Root Cause
-
Prefill Is Compute-Bound, Decode Is Memory-Bound: Optimizing GPU Utilization for LLM Inference
-
DotLLM – Building an LLM Inference Engine in C#
-
DGX Spark Setup Guide: Running vLLM and PyTorch for Local LLM Inference Backend
-
OpenNebula 7.2 "Dark Horse" Released with Enhanced Infrastructure Support
-
Intel Arc Pro B70 32GB Achieves 12 Tokens/Sec on Qwen 3.5-27B
-
Warp Decode vs. vLLM's Triton Kernel: Performance Crossover Analysis
-
Ollama's Limitations for Production Local LLM Deployments
-
Hugging Face Moves Safetensors Under PyTorch Foundation
-
Speculative Decoding Made My Local LLM Actually Usable
-
Ollama is Still the Easiest Way to Start Local LLMs, But It's the Worst Way to Keep Running Them
-
GPU Memory for LLM Inference (Part 1)
-
Satsgate: Monetize AI Agents and APIs with Lightning L402 Protocol
-
5 Useful Docker Containers for Agentic Developers
-
GPUs vs. TPUs: Decoding the Powerhouses of AI
-
NVIDIA Accelerates Gemma 4 for Local Agentic AI on RTX GPUs
-
OpenUMA – Apple-Style Unified Memory for x86 AI Inference
-
Intel's $949 GPU Has 32GB of VRAM for Local AI, but Software is Why Nvidia Keeps Winning
-
Gemini CLI – Open-Source AI Agent for Terminal Integration
-
Is Anyone Working on an AI Operating System?
-
ROCm Integration in Ubuntu 26.04 Advances Linux GPU Inference
-
Local AI Ecosystem Extends Far Beyond Ollama
-
Samsung launches Galaxy Book6 series in India with Nvidia RTX 5070 graphics and on-device AI
-
Intel's $949 GPU has 32GB of VRAM for local AI, but the software is why Nvidia keeps winning
-
RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
-
Qwen 3.5 27B Achieves 1.1M Tokens/Second on B200 GPUs with Optimized vLLM Config
-
Pluggable's TBT5-AI: First Thunderbolt Dock Explicitly Targeting Local LLM Workstations
-
Researcher Successfully Runs Local LLMs on Legacy "Dead" GPU With Surprising Results
-
Developer Builds Fully Local Multi-Agent System Using vLLM and Parallel Inference
-
Build a $1,500 AI Server with DeepSeek-R1 on RTX 4090
-
Community Converges on Optimal KV Cache Quantization Strategies for Qwen 3.5 Models
-
LMCache Dramatically Accelerates LLM Inference on Oracle Data Science Platform
-
Kimi Introduces Attention Residuals: 1.25x Compute Performance at <2% Overhead
-
OpenClaw vs Eigent vs Claude Cowork: Comparing Open-Source AI Collaboration Platforms
-
AMD Launches Agent System Optimized for Local AI Inference With Ryzen and Radeon
-
P-EAGLE: Faster LLM Inference with Parallel Speculative Decoding in vLLM
-
Runpod Report: Qwen Has Overtaken Meta's Llama As The Most-Deployed Self-Hosted LLM
-
Intel Updates LLM-Scaler-vLLM With Support For More Qwen3/3.5 Models
-
How to Install OpenClaw with Ollama (Step-by-Step Tutorial)
-
Nvidia Pushes Jetson as Edge Hub for Open AI Models
-
Cutile.jl Brings Nvidia CUDA Tile-Based Programming to Julia
-
Show HN: Aver – a Language Designed for AI to Write and Humans to Review
-
Nemotron 9B Powers Large-Scale Local Inference: Patent Classification and Real-Time Applications
-
HP Refreshes Lineup with AI-Focused Workstations
-
Intel Arc Pro B70 Workstation GPU Confirmed via vLLM AI Release Notes
-
Framework Choice Critical: llama.cpp and vLLM Outperform Ollama for Qwen 3.5 Testing
-
AMD Expands Ryzen AI 400 Series Portfolio for Consumer and Enterprise AI PC Options
-
Huawei's SuperPoD Portfolio Creates New Option for Global Computing at MWC Barcelona 2026
-
DeepSeek Releases DualPath: Addressing Storage Bandwidth Bottlenecks in Agentic Inference
-
DeepSeek Paper – DualPath: Breaking the Bandwidth Bottleneck in LLM Inference
-
Enterprise Infrastructure Guide: Running Local LLMs for 70-150 Developers
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
LayerScale Launches Inference Engine Faster Than vLLM, SGLang, and TRT-LLM
-
High Bandwidth Flash Memory Could Alleviate VRAM Constraints in Local LLM Inference
-
Self-Hosted AI: A Complete Roadmap for Beginners
-
Open-Source Models Now Comprise 4 of Top 5 Most-Used Endpoints on OpenRouter
-
Critical vLLM RCE Vulnerability Allows Remote Code Execution via Video Links
-
OpenClaw with vLLM Running for Free on AMD Developer Cloud
-
Heaps Do Lie: Debugging a Memory Leak in vLLM
-
Mistral AI Debugs Critical Memory Leak in vLLM Inference Engine