Tagged "llama"
-
Ollama Runs Free AI Models Locally on Mac, Windows and Linux
-
Google Pixel 11 Launches With Faster On-Device Gemini at $899 Starting Price
-
Llama-macOS – Agentic and MCP Native macOS Front End for Llama.cpp
-
The KV Cache Survival Guide: Why Your GPU Runs Out of Memory with Local LLMs
-
Run Ollama Locally on Windows 11: Setup Guide
-
Phi-4 Mini vs Gemma 3 vs Llama 3.2: 128K vs 32K Context Window Comparison
-
AMD Ryzen AI MAX+ 395 Discussed for Local AI Deployment
-
MSI Pro Max Edge AI+ Mini PC Runs 120B Local AI Models With 128GB RAM
-
AMD Acquires FastFlowLM to Accelerate On-Device AI Inferencing
-
Claude Plus a Local LLM Cuts AI Costs in Half, and I'm Never Going Back to Cloud-Only
-
Ollama Secures $65M Series B Funding to Grow its Open-source AI Platform
-
'AI Code Is Insane Trash' – David Gerard on Code Generation Quality
-
South Korea Building Sovereign Cybersecurity AI After US Export Controls
-
Nvidia Showcases Nemotron Models for Japanese AI Development
-
AI-Assisted Development Exhaustion Highlights Need for Better Local Tooling
-
Ollama Closes $65M Series B, Reaches 8.9M Developers on Local Open-Weight AI
-
GitHub Copilot With Ollama: Run Local AI Models In VS Code Offline & Free
-
Viability of Local Models for Coding
-
Open Source AI Must Win: A Call to Action for the Local LLM Community
-
Qualcomm AI Hub Expands to 1,500 Optimized Models for Edge Deployment
-
Qualcomm Acquires Modular AI in $3.9 Billion Deal to Accelerate On-Device AI
-
Why Small Local AI Models Get More Use Than Claude or Gemini
-
Mac Mini Positioned as Premier On-Device AI Computer for Local LLM Inference
-
Samsung's UFS 5.0 Addresses Critical Memory Bandwidth Bottleneck in Mobile AI Inference
-
Why Tool Calling is More Important Than Model Size for Local LLMs
-
General-Purpose Large Language Models Outperform Specialized Clinical AI
-
Show HN: 11 Model Families Ported to Apple's CoreAI On-Device Framework
-
AI bills can be as big as a postdoc salary. Is the cost worth it?
-
Google's New Gemma 4 12B AI Model Is Built for Laptops
-
Bosgame Launches VTA-439 Mini PC with 86 TOPS for Practical Local AI
-
Nvidia Enters Windows Laptop Market, Taking on Intel and AMD
-
Lenovo Bets on On-Device AI to Lift Business PC Upgrades
-
Developer Switches from LM Studio to llama.cpp, Reports No Performance Downgrade
-
Show HN: I Built a Debugging Challenge for the AI Coding Age
-
AgentSlice – Make AI Coding Agents Ask Before They Edit
-
New 8B Local LLM Design Marks Biggest Shift Since DeepSeek R1
-
A/B Tested Gemini 3.1 Pro vs. Claude Opus 4.6 – Usage Quota and Quality Comparison
-
Hardware LLM Taalas Reaches >14,000 TPS on Llama 3.1 8B
-
AMD's New Ryzen AI Max Pro 400 with 192GB LPDDR5X Memory
-
The AI Layoff Receipts: Market Consolidation Accelerates Open-Source Model Adoption
-
Safety Paradox: How RLHF Creates the AI Psychosis Problem It's Meant to Prevent
-
I Stopped Paying for ChatGPT and Switched to a Local LLM That Runs on My Laptop
-
Gemma 4 Replaces Entire Local LLM Stack for Many Practitioners
-
I Think I Figured Out What an AI IDE Looks Like
-
LLM Hallucinations in the Wild
-
Continue.dev for Developers: Complete Local AI Coding Assistant Setup
-
Local LLM Rewrites Resume Better Than ChatGPT, and It's Not Even Close
-
Perplexity Brings On-Device AI Workflow to Macs with 'Personal Computer' Feature
-
Claude Code with a Local LLM Running Offline Is the Hybrid Setup I Didn't Know I Needed
-
AI Coding Tools Are Silently Disagreeing with Each Other
-
Local LLMs Work Best When You're Not Loyal to Just One
-
AMD Posts HDMI 2.1 FRL Patches for Amdgpu Linux Driver
-
New Open-Source Tool Automatically Matches Local LLMs to Your PC Hardware
-
Meta Just Killed Open-Source AI
-
IBM Introduces Granite 4.1 Family of Models for Local Deployment
-
Wipeout Clone Runs Native on ESP32-S3, Pushing Edge Hardware to Its Limits
-
Grokfeed: Terminal Feed Reader for HN, Reddit, and Lobste.rs Using Claude Code
-
Picking Your First Local LLM Is Easier Than the Internet Makes It Sound
-
Unsloth's Custom Kernels Make LLM Fine-Tuning Viable on Consumer GPUs
-
Using a Local LLM as a Zero-Shot Classifier
-
Llama 4 Scout on MLX: The Complete Apple Silicon Guide (2026)
-
Show HN: I Can't Write Python. It Works Anyway – Local LLM Automation
-
Copilot Rate-Limiting Issues Highlight Cloud AI Service Limitations
-
Developer Shares Golden Stack for Local Coding Assistant Integration Directly Inside Code Editors
-
Running Same Prompts Through Claude and Local LLM Revealed Unexpected Results
-
Google's Gemini Nano 4 Offers Faster, Smarter Local Inference Capabilities
-
LiteLLM Integrates with Ollama to Simplify Running 100+ Models Locally
-
Ollama Gets Blazing Fast on Macs with Full MLX Support and 2× Speedups
-
Apple Silicon Macs Run Local AI Faster with Ollama's New MLX Support
-
Claude Code Source Leaked: Community Extracts Multi-Agent Orchestration Framework
-
PrismML Announces 1-Bit Bonsai: First Commercially Viable 1-Bit LLMs
-
Closed Source AI = Neofeudalism
-
GLM-5.1 Model Weights Launching Early April for Local Deployment
-
Homelab Consolidation: Replacing 3 Models with Single 122B MoE Model on AMD Ryzen AI MAX+
-
Private Brain LLM Setup on Windows PC Eliminates Need for Paid Cloud Services
-
Llama.cpp Benchmark: RTX 5090 vs Enterprise Systems Compared
-
Qwen 3.5 Models: Optimal Settings and Reduced Overthinking Configuration
-
Ditching Paid AI Services: Building Self-Hosted LLM Solutions as ChatGPT, Claude, and Gemini Alternatives
-
Qwen 3.5 122B Uncensored (Aggressive) Released with New K_P Quantisations
-
Nvidia Nemotron Cascade 2 30B Emerges as Powerful Alternative to Qwen Models
-
Llama 8B Matches 70B Performance on Multi-Hop QA Using Structured Prompting
-
Cursor's Composer 2 model attribution dispute highlights open-source licensing concerns
-
Qwen 3.5 397B emerges as top-performing local coding model
-
Apple M5 Max 128GB real-world performance benchmarks for local inference
-
Qwen 3.5 Emerges as Top Performer for Local Deployment with Extensive Quantization Options
-
NVIDIA Nemotron Cascade 2 30B Delivers 120B-Class Performance in Compact Form Factor
-
Ultra-Compact 28M Parameter Models Show Promise for Specialized Domain Tasks
-
MiniMax-M2.7: New Compact Model Announced for Local Deployment
-
Local Qwen Models Master Browser Automation Through Iterative Replanning
-
Researcher Discovers Universal "Danger Zone" in Transformer Model Architecture at 50% Depth
-
Practical Fix for Qwen 3.5 Overthinking in llama.cpp
-
Open-Source LLMs Rapidly Displacing Proprietary SOTA Models
-
NVIDIA Updates Nemotron 3 122B License, Removes Deployment Restrictions
-
OmniCoder-9B: Efficient Coding Model for 8GB GPUs
-
Qwen 3.5 122B Demonstrates Exceptional Reasoning for Local Deployment
-
Running Qwen3.5-27B Across Multiple GPUs Over LAN Achieves Practical Speed for Local Inference
-
Open-Source GreenBoost Driver Augments NVIDIA GPU VRAM With System RAM and NVMe Storage
-
Qwen3.5-397B Achieves 282 tok/s on 4x RTX PRO 6000 Blackwell Through Custom CUTLASS Kernel
-
Fine-Tuned 14B Model Outperforms Claude Opus 4.6 on Ada Code Generation
-
Runpod Report: Qwen Has Overtaken Meta's Llama As The Most-Deployed Self-Hosted LLM
-
Intel Updates LLM-Scaler-vLLM With Support For More Qwen3/3.5 Models
-
Llama.cpp Adds True Reasoning Budget Support
-
Comprehensive MoE Backend Benchmarks for Qwen3.5-397B: Real Numbers vs Hype
-
Local AI Coding Assistant: Complete VS Code + Ollama + Continue Setup
-
NVIDIA Jetson Brings Open Models to Life at the Edge
-
Experiment: 0.8B Model Self-Improvement on MacBook Air Yields Surprising Results
-
Llama.cpp Celebrates Major Milestone: From Leak to Industry Standard
-
Simple Layer Duplication Technique Achieves Top Open LLM Leaderboard Performance
-
8 Local LLM Settings Most People Never Touch That Fixed My Worst AI Problems
-
M5 Max and M5 Ultra Chipsets Demonstrate Significant Bandwidth Improvements for Local LLM Inference
-
HP OMEN MAX 16 Review: Is Local AI on a Laptop Viable in 2026?
-
Fine-Tuned Qwen SLMs (0.6–8B) Demonstrate Competitive Performance Against Frontier LLMs on Specialized Tasks
-
Qwen 3.5 Derestricted Model Available for Local Deployment
-
Qwen 3.5 27B Achieves Strong Local Inference Performance
-
Reverse engineering a DOS game with no source code using Codex 5.4
-
OpenSpec: Spec-driven development (SDD) for AI coding assistants
-
Benchmark: Local Open-Source LLMs Competitive in Real-Time Trading Applications
-
ETH Zurich Research Challenges Context-Length Assumptions in LLM Agents
-
Qwen3-Coder-Next Achieves Top Ranking on SWE-bench at Pass@5
-
Open WebUI Adds Native Terminal Tool Calling with Qwen3.5 35B Support
-
llama.cpp Merges Agentic Loop and MCP Client Support
-
llama-swap Emerges as Superior Alternative to Ollama and LM-Studio
-
Quantifying Cost Savings with Local LLMs for Development
-
Apple M5 Pro and M5 Max: 4× Faster LLM Processing
-
Qwen 3.5-35B-A3B Achieves 37.8% on SWE-bench Verified Hard
-
4 Free Tools to Run Powerful AI on Your PC Without a Subscription
-
Qwen3.5 122B Achieves 25 tok/s on 72GB VRAM Setup
-
Qwen 3.5 MoE Delivers 100K Context Window at 40+ TPS on RTX 5060 Ti
-
Qwen 3.5 Underperforms on Hard Coding Tasks—APEX Benchmark Analysis
-
Researchers Develop Persistent Memory System for Local LLMs—No RAG Required
-
DeepSeek Releases DualPath: Addressing Storage Bandwidth Bottlenecks in Agentic Inference
-
Apple Accelerates U.S. Manufacturing with Mac Mini Production
-
Anthropic Reveals Industrial-Scale Distillation Attacks by Chinese AI Labs
-
Comparing Manual vs. AI Requirements Gathering: 2 Sentences vs. 127-Point Spec
-
Anthropic Has Never Open-Sourced an LLM: Implications for Local Deployment Strategy
-
Show HN: Agora – AI API Pricing Oracle with X402 Micropayments
-
nanollama: Open-Source Framework for Training Llama 3 from Scratch with One-Command GGUF Export
-
Kitten TTS V0.8 Released: New State-of-the-Art Super-Tiny TTS Model Under 25 MB
-
SanityBoard Adds 27 New Model Evaluations Including Qwen 3.5 Plus, GLM 5, and Gemini 3.1 Pro
-
PaddleOCR-VL Now Integrated into llama.cpp for Multilingual OCR
-
Free ASIC-Accelerated Llama 3.1 8B Inference at 16,000 Tokens/Second
-
Enhanced Quantization Visualization Methods for Understanding LLM Compression Trade-offs
-
Ask HN: What is the best bang for buck budget AI coding?
-
Qwen 3.5-397B-A17B Now Available for Local Inference with Aggressive Quantisation
-
Self-Hosted AI: A Complete Roadmap for Beginners
-
Meet Sarvam Edge: India's AI Model That Runs on Phones and Laptops With No Internet
-
Open-Source Models Now Comprise 4 of Top 5 Most-Used Endpoints on OpenRouter
-
GNOME's AI Assistant Newelle Adds llama.cpp Support and Command Execution
-
Scaling llama.cpp On Neoverse N2: Solving Cross-NUMA Performance Issues
-
NVIDIA's Dynamic Memory Sparsification Cuts LLM Inference Costs by 8x
-
MiniMax-M2.5 230B MoE Model Released with GGUF Support for Local Deployment
-
LLaDA2.1 Introduces Token Editing for Massive Speed Gains in Local Inference
-
GPT-OSS 120B Uncensored Model Released in Native MXFP4 Precision
-
SnowBall Technique Addresses Context Window Limitations in Local LLMs
-
MiniMax Releases M2.5 Model with SOTA Coding and Agent Capabilities
-
Context Management Identified as Real Bottleneck in AI-Assisted Coding
-
Optimal llama.cpp Settings Found for Qwen3 Coder Next Loop Issues
-
Student Releases Dhi-5B: Multimodal Model Trained for Just $1,200
-
GitHub Announces Support for Open Source AI Project Maintainers
-
New Header-Only C++ Benchmark Tool for Predictive Models on Raw Binary Streams
-
Developer Switches from Ollama and LM Studio to llama.cpp for Better Performance