Self-Hosted LLM Inference on 2× RTX PRO 6000 Blackwell GPUs
We tested 12 model families, from 2 B to 122 B parameters, across more than 85 configurations and 435 benchmark runs. That covered dense and mixture-of-experts (MoE) models, BF16, FP8, NVFP4 and INT4, tensor-parallel, data-parallel and expert-parallel layouts, with and without speculative decoding, all through a standardized vLLM benchmark suite. Here are the results.
Read article

















