Tagged "deepseek-v4-1-flash"
3 articles tagged deepseek-v4-1-flash, 11 September 2026 to 5 October 2026. Newest first.
-
vLLM v0.31.0: DeepSeek-V4.1-Flash with FlashMLA Mega Attention and Sparse MQA Logits
vLLM releases v0.31.0 with major performance optimizations for DeepSeek-V4.1-Flash including FlashMLA mega attention with NVFP4 compressed KV cache and sparse MQA logits, contributed by 307 contributors across 717 commits.
-
vLLM v0.30.0 Released With DeepSeek-V4.1 and Advanced Optimizations
vLLM v0.30.0 brings 762 commits including support for DeepSeek-V4.1-Flash with MXFP8 quantization and async prefetch optimizations for improved throughput on local hardware.
-
Cambricon Adapts DeepSeek-V4.1-Flash on vLLM Stack for Efficient Inference
Cambricon's Day-0 project successfully adapts DeepSeek-V4.1-Flash within the vLLM inference stack, demonstrating practical optimization of large open models for deployment. This work bridges advanced open models with production-grade serving infrastructure.