Skip to content

DataPacket blog

Latest article

Self-Hosted LLM Inference on 2× RTX PRO 6000 Blackwell GPUs

We tested 12 model families, from 2 B to 122 B parameters, across more than 85 configurations and 435 benchmark runs. That covered dense and mixture-of-experts (MoE) models, BF16, FP8, NVFP4 and INT4, tensor-parallel, data-parallel and expert-parallel layouts, with and without speculative decoding, all through a standardized vLLM benchmark suite. Here are the results.

Read article

All articles

DataPacket newsletter

Get updates on new hardware, locations and interconnections