Tagged "nvidia"
-
llama.cpp Build 10605: Mamba2 GEMM Optimization Improves State-Space Model Performance
-
llama.cpp Adds CUDA Pool Operations Support
-
AMD EPYC ZenDNN Accelerates llama.cpp Prompt Processing 4.5x
-
Native vLLM and ROCm 7.15 Support for AMD RDNA2 GPUs on Windows
-
AMD Adds Day 0 Qwen3.8 Support, Radeon AI PRO R9700 Hits 51.8 Tokens per Second
-
HackerNoon Compares 7 Best Self-Hosted Inference Servers for Open-Source Models
-
AMD Optimizes Qwen 3.8 27B for Ryzen AI Max and Radeon GPUs
-
DeepX's DX-M1 On-Device AI Chip Achieves $13M in Orders
-
AMD Launches Gorgon Halo and ROCm.AI for Local AI Inference with Workstation Hardware
-
vLLM v0.27.0 Released with Major Kernel Improvements and New Model Support
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Agent Execution
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Local Agent Deployment
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Agent Execution
-
Meta's Muse Glimmer Now Available Across All Platforms in Ollama
-
vLLM v0.27.0 – Kimi K3 Support and 561 Commits from 242 Contributors
-
NVIDIA Enables Local Agentic AI Workflows with Meta's Muse Glimmer
-
NVIDIA Magpie TTS – Open-Weights Multilingual Voice Agents with Full Deployment Control
-
Meta Releases Muse Glimmer: 30B Open-Source LLM for Local Deployment
-
How to Install Ollama on Windows 11 for Local AI Inference
-
Muse Glimmer Now Available on Ollama – Meta's Open Multimodal Agent Model
-
llama.cpp Improves CUDA Performance with Kernel Fusion
-
Llama.cpp B10327 Fixes CUDA Quantized Copy Kernel Performance
-
llama.cpp Build b10301: CUDA Optimization and Compiler Warning Fixes
-
DeepSeek V4 Flash Optimized for Single AMD MI300X GPU
-
llama.cpp Release b10257 – Vulkan LLVMpipe Fixes
-
AMD's MI355X Undercuts Nvidia's B300 on Cost to Run China's Kimi K3
-
NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework
-
NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework
-
The KV Cache Survival Guide: Why Your GPU Runs Out of Memory with Local LLMs
-
Squeezing Silicon Limits: Effective Strategies to Eliminate GPU Idle Time and Maximize GPU Utilization
-
Building a Dual V100 AI Workstation for Local LLMs
-
Nvidia Accelerates Chip Engineering with AI Agents
-
Triton Control: Open-Source Control Plane for Nvidia Triton on Kubernetes
-
Can a 2.8T Model Run on a Single Node of Nvidia B300 X8?
-
NVIDIA Releases Molt: Agentic RL Training Framework Scaling to Trillion-Parameter Models
-
K3 Model Achieves 20 Tokens/Second on 80x RTX 5090 Cluster
-
AMD Ryzen AI MAX+ 395 Discussed for Local AI Deployment
-
Netflix Details Its In-House LLM Serving Platform with Triton and vLLM
-
Nvidia Isn't the Only Choice for Local LLMs Anymore, and AMD Test Proves It
-
AI Inference is Rewriting the GPU Buying Playbook
-
AMD Acquires FastFlowLM to Accelerate On-Device AI Inferencing
-
NVIDIA's On-Device AI Gains Japan's Manufacturing Giants' Backing
-
Host Private Local AI on NVIDIA DGX Spark Using Ollama and Open WebUI
-
Nvidia Showcases Nemotron Models for Japanese AI Development
-
Mira Murati's Thinking Machines Launches Open-Weight AI Model
-
Nvidia Boosts Token Throughput 5x With Software Optimizations, Reshaping AI Inference Economics
-
Rapid Rise of Open Source Models in the U.S.: Nvidia Nemotron Ultra Grows Quickly on Ollama
-
Intel-Scaler-vLLM 0.21.0-b1 Brings Latest Features for vLLM on Intel GPUs
-
AMD Lemonade Enables Local AI Portability With New Nvidia Support
-
Tiny LLM Benchmark: Jetson Orin Nano Super 8GB
-
A Guide on How to Run Nemotron 3 Super 120B Thinking on 2 Nvidia DGX Spark
-
DEEPX and Sixfab Launch AI HAT for Raspberry Pi Edge Inference
-
NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
-
Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
-
Getting Started With NVIDIA DGX Spark: Unboxing, First Boot, Dashboard, and Running Gemma Locally
-
Best VPS for Ollama 2026 and Setup Guide
-
Building 8 AI Tools With Zero API Costs Using Nvidia NIM
-
RTX 5080 and RTX 3090 Setup Achieves 80 Tok/s on Qwen 3.6 27B Q8
-
Show HN: LiveHere – AI Videos with Self-Hosted Nvidia Cosmos on H200 GPUs
-
AMD claims 256-core Zen 6 'Venice' CPU beats Nvidia Vera by 3.3x
-
AMD's Lemonade SDK Adds NVIDIA CUDA Support for Cross-Platform Local AI
-
DeepSeek V4 Performance Analysis: 1.6T Day 0 to Day 43 Scaling Trends
-
NVIDIA Unveils First PC Chips at Computex 2026; CEO Jensen Huang Details New Hardware
-
NVIDIA Dynamo Snapshot Accelerates AI Inference Startup on Kubernetes
-
NVIDIA Joins Windows on Arm Ecosystem, Driving Arm-Based AI Notebook Adoption to 34.2% by 2029
-
Apple's Overhauled Siri Will Reportedly Run on Nvidia's Blackwell Chips
-
NVIDIA RTX Spark Superchip Delivers 6,144 CUDA Cores for Consumer Local AI Inference
-
NVIDIA and Microsoft Team Up to Bring Secure On-Device AI Agents to Windows PCs
-
Nvidia Enters Windows Laptop Market, Taking on Intel and AMD
-
NVIDIA Levels Up Local AI Agents Across RTX PCs and DGX Spark
-
NVIDIA Launches N1X/N1 CPU-GPU SoC for PC Market, Targeting Heavy On-Device AI Users
-
Microsoft and Nvidia to Unveil First Windows PCs with Nvidia CPUs and AI Capabilities
-
Zoho-Backed Netrasemi Launches 12nm AI Chip, Mass Production Begins This Year
-
Maker Demonstrates Portable AI with Suitcase-Integrated Jetson Orin Setup
-
Deploying Hermes Agent for Free on AMD Developer Cloud with Open Models and vLLM
-
Nvidia Raises Video Encoder Limit to 12 on Consumer GPUs
-
AMD's Lemonade SDK Advances macOS Support for Local AI Inference with ROCm 7.13
-
Local LLM Takes Control of Video Doorbell—The Future of Smart Cameras
-
Maker Builds Offline Jetson-Powered Chatbot Suitcase
-
llama.cpp Delivers Sharp Performance Gains for AMD RDNA3 Users
-
Kog AI – Building a Real-Time Inference Stack on AMD Instinct GPUs
-
Running Local AI LLMs on Mini PCs Without NVIDIA GPUs
-
AMD's vLLM-ATOM Plugin Supercharges DeepSeek-R1 and Kimi-K2 Inference on MI350/MI400
-
$200 NVIDIA V100 Server GPU Mod Beats RTX 3060 in Local LLM Test
-
Lemonade Gives AMD Startups a Wider Path to Local Inference
-
On-Device AI Market Poised for Explosive Growth as Major Tech Companies Invest Heavily
-
Local AI Just Got Easier on Windows and the Implications Go Beyond the Benchmark
-
AMD Posts HDMI 2.1 FRL Patches for Amdgpu Linux Driver
-
Linux Setup for Local LLMs Takes Minutes Compared to Windows Hours
-
NVIDIA Nemotron 3 Nano Omni Powers Multimodal Agent Reasoning in a Single Efficient Open Model
-
Hipfire: A Rust-Native AMD Inference Engine That Outperforms llama.cpp
-
Unsloth's Custom Kernels Make LLM Fine-Tuning Viable on Consumer GPUs
-
NVIDIA Adds Day-0 DeepSeek V4 Blackwell Support
-
Intel OpenVINO 2026.1 Integrates llama.cpp with Wildcat Lake and Arc Pro B70
-
Intel LLM-Scaler vLLM 0.14.0 Released With Official Arc Pro B70 Support
-
Build a More Secure, Always-On Local AI Agent with OpenClaw and NVIDIA NemoClaw
-
Intel's $949 GPU Has 32GB of VRAM for Local AI, but the Software Is Why Nvidia Keeps Winning
-
Google's Gemma 4 Brings Game-Changing Performance to Local Laptop Inference
-
DGX Spark Setup Guide: Running vLLM and PyTorch for Local LLM Inference Backend
-
MiniMax M2.7 Advances Scalable Agentic Workflows on NVIDIA Platforms for Complex AI Applications
-
Intel Arc Pro B70 32GB Achieves 12 Tokens/Sec on Qwen 3.5-27B
-
Gemini-CLI, Llama.cpp, and Qwen3.5 Running on NVIDIA Jetson TK1
-
AMD Announces Day 0 Support for Google Gemma 4 Across Processors and GPUs
-
PyTorch Foundation Welcomes Helion as a Foundation-Hosted Project to Standardize Open, Portable, and Accessible AI Kernel Authoring
-
DGX Spark Hardware Limitations: Missing NVFP4 Support Undermines Local AI Value Proposition
-
Ollama Gets Blazing Fast on Macs with Full MLX Support and 2× Speedups
-
Samsung Launches Galaxy Book6 Series with NVIDIA RTX 5070 and On-Device AI
-
NVIDIA and Google Optimize Gemma 4 AI Models for Local RTX Deployment
-
GPUs vs. TPUs: Decoding the Powerhouses of AI
-
AMD Rolls Out Gemma 4 Model Support Across Full Range of GPUs & CPUs
-
NVIDIA Accelerates Gemma 4 for Local Agentic AI on RTX GPUs
-
Google Launches Gemma 4 Open Models for Local On-Device AI
-
AMD Provides Day 0 Support for Gemma 4 on Ryzen AI Processors and GPUs
-
Chinese Chipmakers Claim Nearly Half of Local Market as Nvidia's Lead Shrinks
-
Lotte Innovate and DeepX Collaborate on Mass Production of Domestic AI Semiconductors
-
TinyGPU Adds Mac Support for External Nvidia GPU Acceleration
-
Intel's $949 GPU Has 32GB of VRAM for Local AI, but Software is Why Nvidia Keeps Winning
-
Intel's Arc GPU Offers 32GB VRAM for Local AI, But Software Ecosystem Lags Behind
-
Is Anyone Working on an AI Operating System?
-
ROCm Integration in Ubuntu 26.04 Advances Linux GPU Inference
-
Samsung launches Galaxy Book6 series in India with Nvidia RTX 5070 graphics and on-device AI
-
Intel's $949 GPU has 32GB of VRAM for local AI, but the software is why Nvidia keeps winning
-
Select the Right Hardware for Your Local LLM Deployment with This Online Guide
-
Samsung Launches Galaxy Book6 Series in India with NVIDIA RTX 5070 Graphics and On-Device AI
-
Samsung Galaxy Book6 Brings Consumer-Grade On-Device AI Hardware to Market
-
mlx-Code: Run Claude Code Locally with MLX-LM
-
RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
-
Intel Launches Arc Pro B70/B65 with 32GB VRAM for Local AI Inference
-
NVIDIA Releases GPT-OSS-Puzzle-88B, a Deployment-Optimized Model
-
Researcher Successfully Runs Local LLMs on Legacy "Dead" GPU With Surprising Results
-
FlashAttention-4 Delivers 2.7x Faster Inference with 1613 TFLOPs/s on Blackwell GPUs
-
Llama.cpp ROCm 7 vs Vulkan Performance Benchmarks on AMD Mi50
-
Nvidia Nemotron Cascade 2 30B Emerges as Powerful Alternative to Qwen Models
-
DeepSeek R1 RTX 4090 vs Apple M3 Max: Benchmark & Performance Guide
-
Build a $1,500 AI Server with DeepSeek-R1 on RTX 4090
-
Repurpose Old GPUs as Dedicated AI Inference Accelerators
-
NVIDIA Nemotron Cascade 2 30B Delivers 120B-Class Performance in Compact Form Factor
-
Llamafile 0.10 Released with GPU Support and Rebuilt Core
-
NVIDIA Nemotron 3 Nano 4B Enables On-Device Inference Directly in Web Browsers via WebGPU
-
I Ran Local LLMs on a 'Dead' GPU, and the Results Surprised Me
-
Qwen 3.5 4B Outperforms Nvidia Nemotron 3 4B in Local Benchmarks
-
Mistral Small 4 119B Released with NVFP4 Quantisation Support
-
NVIDIA Updates Nemotron 3 122B License, Removes Deployment Restrictions
-
Nvidia's Nemotron 3 Super: Understanding the Significance for Local LLM Deployment
-
Running Qwen3.5-27B Across Multiple GPUs Over LAN Achieves Practical Speed for Local Inference
-
Startup Transforms Mac Mini Into Full-Powered AI Inference System With External GPU
-
Open-Source GreenBoost Driver Augments NVIDIA GPU VRAM With System RAM and NVMe Storage
-
AMD Launches Agent System Optimized for Local AI Inference With Ryzen and Radeon
-
Qwen3.5-397B Achieves 282 tok/s on 4x RTX PRO 6000 Blackwell Through Custom CUTLASS Kernel
-
Intel OpenVINO Backend Support Now Available in llama.cpp
-
Linux 7.0 AMDGPU Fixing Idle Power Issue For RDNA4 GPUs After Compute Workloads
-
How to Install OpenClaw with Ollama (Step-by-Step Tutorial)
-
Nvidia Pushes Jetson as Edge Hub for Open AI Models
-
Nvidia Releases Nemotron 3 Super: 120B MoE Model for Local Deployment
-
Comprehensive MoE Backend Benchmarks for Qwen3.5-397B: Real Numbers vs Hype
-
Cutile.jl Brings Nvidia CUDA Tile-Based Programming to Julia
-
NVIDIA Jetson Brings Open Models to Life at the Edge
-
Intel Arc Pro B70 Workstation GPU Confirmed via vLLM AI Release Notes
-
Qwen3.5-27B Identified as Sweet Spot for Mid-Range Local Deployment
-
Nvidia Could Launch Its First Laptops With Its Own Processors
-
Google Is Exploring Ways to Use Its Financial Might to Take on Nvidia
-
NVIDIA Releases Dynamo v0.9.0: Infrastructure Overhaul With FlashIndexer and Multi-Modal Support
-
LayerScale Launches Inference Engine Faster Than vLLM, SGLang, and TRT-LLM
-
AMD Announces Day 0 Support for Qwen 3.5 LLM on Instinct GPUs
-
NVIDIA's Dynamic Memory Sparsification Cuts LLM Inference Costs by 8x
-
Mistral AI Debugs Critical Memory Leak in vLLM Inference Engine
-
Community Member Builds 144GB VRAM Local LLM Powerhouse