A dual-GPU setup combining RTX 5080 and RTX 3090 cards reaches 80 tokens per second when running Qwen 3.6 27B in Q8 quantization. The configuration demonstrates significant inference speed improvements for large language models.
The RTX 5080 and RTX 3090 pairing delivers 80 tokens per second on the 27-billion parameter Qwen 3.6 model with 8-bit quantization. This throughput level makes the setup viable for practical LLM deployments requiring sustained generation speeds.
The dual-card approach leverages NVIDIA's latest RTX 5080 architecture alongside the established RTX 3090, balancing performance and cost efficiency. Q8 quantization reduces model size while maintaining reasonable accuracy, a critical factor for fitting large models across consumer-grade GPUs.
The benchmark has gained traction in developer communities, with 117 points and 43 comments on Hacker News indicating strong interest in multi-GPU inference configurations. Such setups address the gap between consumer hardware capabilities and production-scale LLM serving requirements.
The achievement highlights ongoing optimization efforts in the open-source ML community, where combining older and newer generations of GPUs offers practical alternatives to enterprise-class hardware for inference workloads.
Samsung's latest Galaxy Z Fold 8, Z Fold 8 Ultra, and Z Flip 8 are tracking to break all previous Z-series pre-order records, with demand up 30% year-over-year.
D-Wave, known for quantum annealers, is now developing gate-based quantum hardware using dual-rail qubits and has successfully demonstrated entanglement in the new architecture.
SpaceX plans to increase compute capacity more than fivefold by end of 2027, potentially requiring over two million Nvidia Rubin GPUs. The ambitious expansion underscores the company's AI infrastructure strategy as its AI segment generated $2.56 billion in Q2 revenue.
Uber and UK startup Wayve have secured private hire vehicle licenses in London for supervised autonomous taxi operations. The move marks a significant step toward commercial robotaxi deployment in Europe.