Tagged "gguf"
- Qwen3.8-Flash-Next Added to llama.cpp with GGUF Support
- Liquid AI Releases LFM2.5-DSpark Draft Models with 3.18x Faster Decoding
- Liquid AI Releases LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation
- GGUF Quantization Deep Dive: Q4_K_M vs IQ4_XS vs IQ4_NL Performance
- DeepSeek V4 Flash Shrunk to 57GB for Local macOS Inference with Compiler Generation
- Qwen3.8-27B Surpasses 1 Million Downloads, Overseas Developers Race to Maximize Local Deployment
- GGUF Quantization Compared: Q4_K_M vs. IQ4_XS vs. IQ4_NL Performance Analysis
- Unsloth Releases Qwen 3.8 27B GGUF Quantised Weights
- Qwen 3.8 27B Successfully Runs on 16GB RAM Using LM Studio
- AMD Optimizes Qwen 3.8 27B for Ryzen AI Max and Radeon GPUs
- Hugging Face State of Open Models: Summer 2026 Observations
- Running DeepSeek's 284B LLM on a Laptop: Quantisation and GGUF Optimization
- llama.cpp Improves Muse Glimmer Tool Calling with Latest Update
- Show HN: GGUFun, Play Snake and a Simple Maze on Ollama Using Hand Crafted GGUFs
- What else is included in the 'GGUF' file format used by llama.cpp for AI language models, besides weights?
- llama.cpp GGUF Parser Flaws: Critical Integer Overflow Enables Arbitrary Reads in Every Local AI Stack
- Ollama Out-of-Bounds Read Vulnerability Allows Remote Process Memory Leak
- Malicious GGUF Models Could Trigger Remote Code Execution on SGLang Servers
- MiniMax M2.7 GGUF Investigation Reveals NaN Issues Affecting 21-38% of Hugging Face Conversions
- Fine-Tuned Qwen3.5-0.8B for OCR Outperforms Previous 2B Release
- Unsloth Completes Comprehensive MiniMax M2.7 GGUF Quantization Suite
- ByteShape Releases Qwen 3.5 9B Quantisations with Hardware-Matched Tuning Guide
- Coding Implementation to Run Qwen3.5 Reasoning Models Distilled With Claude-Style Thinking Using GGUF and 4-Bit Quantization
- Qwen 3.5 122B Uncensored (Aggressive) Released with New K_P Quantisations
- Qwen 3.5-35B Uncensored GGUF Models Now Available
- Final Qwen3.5 Unsloth GGUF Update with Improved Size/Quality Tradeoffs
- Qwen 3.5-27B Q4 Quantization Comparison and Analysis
- Unsloth Dynamic 2.0 GGUFs
- Qwen3.5-35B Unsloth Dynamic GGUFs Achieve SOTA Across Nearly All Quantisation Levels
- nanollama: Open-Source Framework for Training Llama 3 from Scratch with One-Command GGUF Export
- Ouro 2.6B Thinking Model GGUFs Released with Q8_0 and Q4_K_M Quantization