Tagged "mixture-of-experts"
28 articles tagged mixture-of-experts, 12 February 2026 to 2 October 2026. Newest first.
-
Allen Institute Releases Olmo-Core 3: Open Training Infrastructure for Large Mixture-of-Experts Models
Allen Institute has released Olmo-Core 3, an open-source training infrastructure designed for large-scale mixture-of-experts (MoE) models, enabling community-driven development of efficient models suitable for local deployment.
-
Llama.cpp Fork Achieves 2-4x MultiGPU Speedup for MoE Models Larger Than VRAM
A community fork of llama.cpp enables efficient distributed inference for Mixture-of-Experts models that exceed single GPU VRAM capacity, achieving 2-4x speedup improvements across multiple GPUs.
-
Diffusion Reads: 24 Answers in One Forward Pass, and Why I Stopped Batching Them
A discrete diffusion model answers a whole canvas of questions in one denoise step. Batching 24 questions into that canvas ran 8.5x faster and halved the scores. The pod spec, both servers, and the run that produced AUC 0.984 on 49,536 questions for about $12.
-
llama.cpp Broadens MoE Optimization Heuristics for AMD RDNA3.5
llama.cpp release b10997 improves Mixture-of-Experts performance on AMD's latest architecture with refined tile heuristics and verified correctness on Ryzen AI MAX+.
-
MoE Expert Offload: What a 35B Model Actually Costs on a 12GB Card
92.9% of Qwen3.6-35B-A3B is routed expert weights, and only 3.1% of them are read per token — which is why a 19 GiB model runs on 12 GB at all. The roofline arithmetic for expert offload, and why the same sum that permits 50 tok/s at 8K refuses it at 128K.
-
Llama.cpp B10758: Hexagon MUL_MAT Fusion and MoE Optimizations for Qualcomm Hardware
Latest llama.cpp release adds Qualcomm Hexagon MUL_MAT and MUL_MAT_ID fusion optimizations, enabling efficient inference on Qualcomm processors used in edge devices and Android hardware. This expands local inference support beyond traditional server/desktop GPUs.
-
FreeToken: Edge-Native MoE Serving Engine for Consumer Hardware
FreeToken is a mixture-of-experts serving engine aimed at running frontier-scale open-weight models on consumer hardware, using CPU-GPU co-execution rather than a multi-GPU cluster.
-
llama.cpp b10524 Makes MoE Expert Scatter Deterministic in OpenCL
llama.cpp releases build b10524 with deterministic MoE expert scatter operations in OpenCL backend, improving reliability for Mixture of Experts models on GPU acceleration. This optimization is crucial for consistent inference behavior.
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Local Agent Deployment
NVIDIA's new 30B mixture-of-experts model with only 3B active parameters is now available in Ollama, optimized for building always-on agents with minimal resource requirements. The model is designed for agent frameworks like OpenClaw and Hermes.
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Agent Execution
NVIDIA's new 30B mixture-of-experts model with 3B active parameters is now available in Ollama v0.32.9, optimized for agent workloads and on-device execution. The model is designed for frameworks like OpenClaw and Hermes, bringing efficient MoE inference to local deployments.
-
Qwen and Fable: Open-Weights 35B Mixture-of-Experts Agentic Coding Model
A new open-weights 35B Mixture-of-Experts model combining Qwen and Fable for agentic coding tasks, optimized for local deployment with improved efficiency through sparse computation patterns.
-
JetBrains Releases Mellum2: A 12B MoE Model for Fast, Specialized Tasks
JetBrains introduces Mellum2, a 12-billion parameter mixture-of-experts model designed for efficient local inference in multi-model AI pipelines. The model balances performance and resource consumption for on-device deployment scenarios.
-
Homelab Consolidation: Replacing 3 Models with Single 122B MoE Model on AMD Ryzen AI MAX+
A homelabber consolidated their inference setup from three separate models down to a single 122B mixture-of-experts model on consumer hardware (Ryzen AI MAX+ 395 with 128GB RAM), providing detailed benchmarks and practical insights on model consolidation strategy.
-
FOMOE: Running 397B Parameter Qwen3.5 MoE at 5-9 tok/s on $2,100 Desktop Hardware
Fast Opportunistic Mixture of Experts (FOMOE) enables inference of massive 397-billion parameter models using Q4_K_M quantization on dual $500 consumer GPUs with 32GB RAM, solving the memory bottleneck of MoE models through intelligent flash-backed weight streaming.
-
Nvidia Releases Nemotron 3 Super: 120B MoE Model for Local Deployment
Nvidia has released Nemotron 3 Super, a 120B mixture-of-experts model with only 12B active parameters, designed as an open-source alternative for agentic reasoning tasks. The hybrid Mamba-Transformer architecture offers competitive performance with reduced computational requirements.
-
Comprehensive MoE Backend Benchmarks for Qwen3.5-397B: Real Numbers vs Hype
A detailed benchmark of every major MoE backend for Qwen3.5-397B NVFP4 on workstation GPUs reveals actual sustained performance of 50.5 tok/s, significantly lower than commonly cited claims. The analysis uncovers kernel issues in Nvidia's own CUTLASS implementation.
-
Krasis: Hybrid CPU/GPU MoE Runtime Achieves 3,324 Tokens/Second Prefill on RTX 5080
New open-source runtime optimises mixture-of-experts models by splitting prefill to GPU and decode to CPU, enabling larger MoE models to run on single consumer GPUs with dramatic throughput improvements.
-
Qwen 3.5 MoE Delivers 100K Context Window at 40+ TPS on RTX 5060 Ti
Qwen3.5's mixture-of-experts variant achieves exceptional throughput with 100,000 token context window on a single mid-range GPU, reaching 41+ tokens per second using the Vulkan backend. This demonstrates practical feasibility of ultra-long context models on consumer hardware.
-
Qwen3.5 Series Releases Comprehensive Model Lineup Across All Tiers
Alibaba released the complete Qwen3.5 model family including 27B, 35B-A3B, and 122B-A10B variants, each optimized for different deployment scenarios and providing extensive benchmark comparisons.
-
Qwen3.5-35B-A3B Emerges as Game-Changer for Agentic Coding Tasks
The newly released Qwen3.5-35B-A3B model with MoE architecture is delivering exceptional performance for coding agents on consumer hardware, with users reporting impressive results running on a single RTX 3090.
-
Alibaba's Qwen3.5-397B Achieves #3 Position in Open Weights Model Rankings
Alibaba's newly released Qwen3.5-397B mixture-of-experts model ranks #3 in the Artificial Analysis Intelligence Index among open-weight models, offering a powerful option for large-scale local deployment.
-
Qwen3-Next 80B MoE Achieves 39 Tokens/Second on RTX 5070/5060 Ti Dual-GPU Setup
A community member has optimised Qwen3-Next 80B mixture-of-experts to run at 39 tokens/second on dual RTX 50-series GPUs with 32GB total VRAM, sharing previously undiscovered configuration solutions for consumer-grade hardware.
-
Qwen 3.5-397B-A17B Now Available for Local Inference with Aggressive Quantisation
Alibaba's Qwen 3.5-397B mixture-of-experts model is now available on HuggingFace with multiple quantisation options, including a 113GB IQ2_XS variant that fits on consumer hardware. Early benchmarks show performance competitive with Gemini 3 Pro and GPT-5.2 on spatial reasoning tasks.
-
MiniMax-M2.5 230B MoE Model Released with GGUF Support for Local Deployment
MiniMax-M2.5, a 230B parameter mixture-of-experts model, is now available in GGUF format for local deployment with impressive performance benchmarks on consumer hardware.
-
GPT-OSS 120B Uncensored Model Released in Native MXFP4 Precision
An uncensored version of GPT-OSS 120B has been released featuring native MXFP4 precision training, offering 117B parameters with MoE architecture for efficient local deployment.
-
MiniMax M2.5: 230B Parameter MoE Model Coming to HuggingFace
MiniMax officially confirms open-source release of M2.5, a 230B parameter MoE model with only 10B active parameters, showing impressive SWE-Bench performance at 80.2%.
-
Ming-flash-omni-2.0: 100B MoE Omni-Modal Model Released
Ant Group releases Ming-flash-omni-2.0, a 100B MoE model with 6B active parameters supporting unified speech, SFX, music generation alongside image, text, and video processing.
-
GLM-5 Released: 744B Parameter MoE Model Targeting Complex Tasks
Zhipu AI releases GLM-5, a massive 744B parameter MoE model with 32B active parameters, designed for complex systems engineering and long-horizon agentic tasks with significant performance improvements over GLM-4.5.