RTX 5080 and RTX 3090 Setup Achieves 80 Tok/s on Qwen 3.6 27B Q8

1 min read
Hacker Newspublisher

A developer has published compelling benchmark results showing how to achieve 80 tokens per second when running Qwen 3.6 27B in Q8 quantization across an RTX 5080 and RTX 3090 GPU setup. This represents a significant performance milestone for local inference, demonstrating that multi-GPU configurations can deliver production-grade throughput for moderately-sized quantized models.

For practitioners deploying local LLMs, these results are particularly valuable because they provide concrete evidence of what dual mid-to-high-end consumer GPUs can achieve. The Q8 quantization level offers a sweet spot between model quality and inference speed, making this configuration practical for many real-world applications. The benchmark details help inform hardware purchasing decisions and capacity planning for on-device inference systems.

Read the full benchmark and technical details on the author's blog to understand the exact setup, optimization techniques used, and performance characteristics across different batch sizes.


Source: Hacker News · Relevance: 9/10