· 20 min read
We recently started offering servers with 2× NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. 96 GB of GDDR7 per GPU, 192 GB in total, in an ordinary air-cooled rack server with no need for liquid cooling. Before you put one in production (or in your closet), you probably want to know what it can actually do with modern LLMs.
We tested 12 model families, from 2 B to 122 B parameters, across more than 85 configurations and 435 benchmark runs. That covered dense and mixture-of-experts (MoE) models, BF16, FP8, NVFP4 and INT4, tensor-parallel, data-parallel and expert-parallel layouts, with and without speculative decoding, all through a standardized vLLM benchmark suite. Every run used the same server hardware configuration that’s available from us.
This post summarizes what we learned. The headline numbers, the surprises, the traps… and what any of it means if you're deciding to run it yourself.
We wanted to compare single- and dual-GPU performance and understand the economics of self-hosting. This is not a deployment guide nor an attempt to tune every model to its limit. We used vLLM because the author knows it best; SGLang, llama.cpp and other engines may behave differently.
You want to know why this is an interesting proposition for inference. Data-center-class Blackwells (B300 and friends) deliver astonishing throughput, but they have high data center demands: ~1,400 W per GPU, liquid cooling, and serious facility work before the first token is generated (See B300 power requirements, Supermicro B300 supercluster datasheet). There’s the other branch of the family (RTX PRO 6000): 600 W power cap, air-cooled, happy in a standard rack with standard power. Two of them give you 192 GB of fast memory (96 GB per card) and enough compute to serve very capable models without having to rebuild your electrical room.
Correct, the RTX PRO 6000 is based on the consumer Blackwell die (SM120), not the data-center SM100 (See SM100 vs. SM120). In practice this means some data-center kernels and features don't carry over. There is also no NVLink. The two GPUs talk over PCIe, which matters more than you'd think and required a workaround in our setup (See Gotchas section).
The following determine what an inference server can actually do for you. Conflating them can be the root cause of bad capacity planning:
Everything ran on the target hardware itself, deployed via Kubernetes with one vLLM engine per configuration (described with kustomize overlays). Prometheus was scraping vLLM metrics. The benchmarks used the vLLM’s own bench sweep tooling against the OpenAI-compatible completions endpoint with random-token prompts (See vLLM benchmarking docs).
Our setup was five synthetic benchmarks per configuration:

The environment was Linux 6.18.38, vLLM 0.25.1 (0.26.0 for the DeepSeek-V4-Flash and DFlash experiments, 0.28.0 for Qwen3.8-27B, and a current dev build for Qwen3.8-Flash-Next), CUDA 13.0, driver 580.126.20. A representative TP2 run showed both GPUs at 100% utilization:
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.126.20 Driver Version: 580.126.20 CUDA Version: 13.0 |
|-----------------------------------------+------------------------+----------------------+
| 0 NVIDIA RTX PRO 6000 Blac... On | 00000000:01:00.0 Off | 0 |
| N/A 46C P0 161W / 600W | 92351MiB / 97887MiB | 100% |
+-----------------------------------------+------------------------+----------------------+
| 1 NVIDIA RTX PRO 6000 Blac... On | 00000000:E1:00.0 Off | 0 |
| N/A 45C P0 160W / 600W | 92365MiB / 97887MiB | 100% |
+-----------------------------------------+------------------------+----------------------+The cross-model picture, feat. every model in its best-performing configuration:
[CHART – headline throughput: Two-panel horizontal bar, one bar per model in its configuration, sorted by decode. Prefill 4K (tok/s) on the left, concurrent decode (tok/s) on the right. Visually distinguished MoE and dense models.]

1. Architecture beats size. On the same 2-GPU node, the MoE Gemma4-26B-A4B delivers 7,740 tok/s of aggregate decode. The similarly sized dense Gemma4-31B manages 2,200 tok/s. This is a 3.5× gap, with prefill nearly 6× faster for the MoE model. The like-for-like BF16 TP2 comparison shows a ~3× decode gap (5,737 vs. 1,935 tok/s). On a single GPU, it is again roughly treefold (3,758 vs. 1,202 tok/s), using the dense model's clean INT4 run. With only a few billion active parameters per token, the MoE models were a genuinely different class of workload on this hardware, not a marginal improvement (AMD MoE serving guide covers this distinction in more detail).
2. The sweet spot is a mid-size MoE. Qwen3.6-35B-A3B, with 3 B active parameters, was the star of the whole exercise: nearly 12,000 tok/s of aggregate decode, ~45,000 tok/s of prefill, and a comfortable 180–400 tok/s single-stream, depending on configuration. The 120B-class MoEs (Nemotron-3-Super-120B-A12B, Qwen3.5-122B-A10B) also fit; they ran surprisingly well once quantized.
3. 192 GB goes further than it sounds. Quantized 120B-class MoEs leave a huge amount of room for KV cache. The Nemotron TP2 NVFP4 configuration addresses 17.8 million KV cache tokens, enough for 67 concurrent sessions at full 256K context. That's the kind of number that makes agentic workloads with long conversations actually practical.
But let’s look into the full capacity picture in the next section.
Throughput tells you how fast tokens flow; KV cache capacity tells you how many sessions the server can hold at all. Every active conversation keeps its context – prompt plus everything generated so far – resident in GPU memory as the KV cache. When the cache fills up, new and preempted requests wait, and the users experience that as randomly stalled generations.
Capacity is whatever is left of the 192 GiB after model weights and runtime overhead, so quantization and parallelism layout matter as much as the model itself. The architecture matters as well. You see, models store KV per attention head, and the per-token cost ranged from ~5 KB (Nemotron) to ~140 KB (Gemma4-31B) across the families we tested – a 25× difference that had nothing to do with model size.
The KV-maximizing configuration we tested per model on two GPUs; capacities as reported by vLLM at startup (after weights, activations and CUDA graphs):
[CHART – KV capacity: Line chart, one line per model (KV-maximizing configuration): concurrent sessions on a log scale vs. context length (32K / 128K / 256K). Shows the capacity ranking and how steeply session counts fall as context grows — Nemotron's flat line vs. DeepSeek-V4-Flash falling off a cliff.]

Please read the session counts as "maximum concurrent sessions if every session holds a full context" – sessions with shorter contexts fit proportionally more.
Quantization buys capacity twice. Smaller weights leave more room for KV. On a single GPU, Qwen3.6-35B-A3B in BF16 leaves 19 GiB (~1.0 M tokens); in NVFP4 it leaves 60 GiB (~2.7 M tokens). On two GPUs it's the difference between 5.2 M and 7.4 M tokens.
Architecture can make a model, regardless of its speed, impractical. DeepSeek-V4-Flash posts respectable throughput, but at ~55 KB/token it holds barely three 128K sessions. On this hardware, it's built for short-context traffic. Gemma4-31B pays an eye-watering ~140 KB/token, and even with 152 GiB of free KV it only reaches 1.2 M.
KV exhaustion is what "falling over" looks like. The Gemma4-31B BF16 single-GPU run failed 36% of our 1,000-request decode benchmark. It needed ~1.2 M tokens of KV and had 206k available. The engine spent the run preempting and recomputing.
Nemotron is the capacity king. ~5.5 KB per token plus small NVFP4 weights add up to 17.8 M tokens: 135 concurrent 128K-context sessions on a single two-GPU node.
Oh, and regarding DP2 layouts: total capacity was the same, but it was split across two independent replicas, and a single session could only use one replica's share. This is relevant mostly for very long contexts.
Since Qwen3.6-35B-A3B was the workhorse candidate, here is the full configuration matrix for it: parallelism layout, quantization, and speculative decoding, all on two GPUs unless noted:
[CHART – throughput vs. latency trade-off: Scatter plot of the matrix above: x = concurrent decode (tok/s), y = single-request decode (tok/s), one point per configuration, labeled by quantization and speculative decoding (MTP / DFlash). The two clusters – shared-endpoint configs bottom-right, single-session configs (MTP, DFlash) top-left – make the configuration choice visible at a glance.]

DP2+EP beats TP2 for MoE. Splitting the model across both GPUs (TP2) roughly doubles decode throughput versus a single GPU. Running two independent data-parallel replicas with expert parallelism (DP2+EP) does better still: 8,760 vs. 6,820 tok/s in BF16, plus it wins prefill outright at 55,900 tok/s on 4K prompts with FP8, versus 41,400 tok/s for TP2 with the same FP8 quantization. At least for MoE, two smaller engines beat one big engine. The trade-off is single-request speed, which drops to single-GPU levels, because one request lives on one replica.
Quantization is nearly free performance. FP8 and NVFP4 improve prefill and concurrent decode while shrinking weights enough to grow the KV cache from 5.2 M to 7.4 M tokens. We didn’t evaluate quality loss, that's a separate – and harder – question. Throughput-wise, though, there’s very little reason to run BF16 on this hardware.
MTP (multi-token prediction speculative decoding) is a single-user feature. With num_speculative_tokens=3, single-request decode nearly doubles from 181 → 348 tok/s in BF16, and up to 403 tok/s with NVFP4. It hurts under heavy concurrency, since batch sizes are already large and the speculative overhead eats the gains (6,820 → 6,290 tok/s in BF16, and NVFP4+MTP drops concurrent decode to 4,880 tok/s). All in all, If your workload looks like "many users, medium speed each,” leave speculative decoding off. If it looks like "one user, fast,” turn it on.
DFlash speculative decoding is a niche superpower. With a dedicated draft model and 15 speculative tokens, single-request decode hits 956 tok/s, which is comfortably interactive. Concurrent decode collapses to 2,230 tok/s, and in our DFlash experiments the stability issues showed up on Qwen3.6-27B, which dropped ~19% of concurrent decode requests (See Caveats section). A great fit for a dedicated single-session endpoint, a terrible idea for a shared one.
The Qwen3.6-generation numbers above show our benchmarks when we started. However, the Qwen3.8 release changed the picture enough to deserve its own section: we ran the same suite against the dense Qwen3.8-27B and the hybrid-attention Qwen3.8-Flash-Next. 20 configurations, 100 benchmark runs, and the 27B is, in our opinion, the best all-round fit for this hardware to date.
The reasoning is… not raw throughput. The 35B MoE still reigns in raw throughput. It's what you get per unit of capability: a dense 27B is a full-quality model, and this node serves it at ~4,000 tok/s aggregate decode, ~15,700 tok/s prefill, 180 tok/s single-stream with MTP, 2.3 M tokens of KV cache and a 256K context window. Every stable configuration completed 100% of our benchmark requests. If you're picking one engine for a general-purpose endpoint on two RTX PRO 6000s, this is the way.

TP2+NVFP4 is the default recipe. The 2nd GPU doubles aggregate decode from 1,490 → 3,210 tok/s and triples KV capacity. NVFP4 wins outright every column: prefill +46% over BF16, concurrent decode +23%, single-stream +80%. It also leaves the most room for cache. As with the MoE models, there’s no throughput case for BF16 here.
MTP is the interactive switch. Three speculative tokens take single-stream decode from 45 → 112 tok/s in BF16 and from 81 → 180 tok/s in NVFP4 at a ~25% aggregate-throughput cost. Toggle it off for shared endpoints, toggle it on for single users; keep in mind the same trade-off as on the 35B.
DFlash2 turns the 27B into a coding buddy. With a dedicated draft model and 7 speculative tokens, single-stream hits 243 tok/s in BF16 and 360 tok/s in NVFP4; that’s genuinely snappy for a dense 27B. Aggregate decode roughly halves. Keep it on a dedicated single-session endpoint.
One GPU is a legitimate deployment. A single RTX PRO 6000 runs the NVFP4 27B at 2,710 tok/s of aggregate decode with 0.96 M tokens of cache, and, amusingly, at higher prefill throughput than TP2 (17,100 vs. 15,700 tok/s), since two-GPU prefill pays the PCIe all-reduce tax. So TP2 wins where it matters for a shared endpoint: decode and capacity.
Flash-Next is the other branch of the family: an MoE where most layers use linear attention (gated delta net) and where only a few use full attention. For serving, that means one thing: KV cache costs ~26 KB per token instead of the 27B's ~65 KB, so context is 2.5× cheaper per GiB. It ships FP8-quantized, keeps the 256K context window, and currently needs a vLLM dev build (vllm/vllm-openai:qwen38-flash-next). The NVFP4 checkpoints failed to load outright on it, so treat Flash-Next as FP8-only… for now.

Flash-Next matches the 27B's prefill and lands ~20% below its concurrent decode. MTP, though, is a poor trade on this architecture: aggregate decode halves (3,130 → 1,540 tok/s) for a modest single-stream gain (94 → 127). Unlike on the 27B, we can't recommend it.
KV capacity per configuration as reported by vLLM at startup; sessions assume every session holds a full context (shorter contexts fit proportionally more):

In this generation, the 27B's ~65 KB/token is the price of dense attention – nearly twice that of the older Qwen3.6-27B's ~34 KB. Quantization pays for it: NVFP4 frees 141.5 GiB of cache versus 115.5 GiB in BF16, good for 69 concurrent 32K sessions or 8 full 256K-context sessions.
Flash-Next fits 1.70 M tokens into just 42.5 GiB of KV – within 25% of the 27B's total capacity from less than a third of the memory. For long-context, RAG and agentic traffic, it's the capacity play.
Editorial footnote about single-GPU: one 96 GB card runs the 27B NVFP4 with 29 concurrent 32K sessions, leaving the second GPU free for a second engine (an MTP twin for interactive users, say) or batch work.
The formula = monthly cost ÷ (throughput × seconds per month) – at the illustrative $2,800/month:

If we provisioned one of these nodes today: Qwen3.8-27B, TP2, NVFP4 as the shared endpoint; its MTP twin (or a single-GPU deployment, with the second card spared) for interactive work; Flash-Next for long-context and agentic traffic. The 35B MoE remains the answer when tokens-per-second-per-dollar is the only variable.
Let’s be real, the tables above are best-case numbers from synthetic, homogeneous load. Real traffic is not homogeneous, and this is where naive capacity planning falls apart.
A single GPU cannot efficiently prefill and decode at the same time: processing a prompt is compute-bound, generating tokens is bandwidth-bound, and a prefill (every user message, every tool call in an agentic session) interrupts all pending decodes. Here is what that looks like in practice: 16 concurrent coding sessions running against Qwen3.6-35B-A3B BF16 TP2 with MTP over about two hours of real usage.
The same configuration that sustains 6,300+ tok/s under pure concurrent decode, and ~350 tok/s on a single request, averages only ~240 tok/s of aggregate generation across these sessions — peaks barely reach 500 tok/s. Prompt processing interruptions on every request bring the averages down dramatically. Note that under this load the GPUs are not fully saturated — but the headroom left is only some, and filling it would take more concurrent requests, each of which brings its own prompt processing that interrupts everyone else's decodes. The problem isn't raw capacity — it's that decode keeps getting interrupted by prefill.
The industry-standard solution is P/D (prefill/decode) disaggregation: deploy the model at least twice, one instance dedicated to prompt processing and one to token generation, with the KV cache handed off between them. Then tune the P:D ratio to your traffic — prompt-heavy document workloads want more P instances, chat and code generation want more D. vLLM has first-class support for this, with transfer layers such as NIXL to move KV cache between instances. With 2× RTX PRO 6000 nodes, disaggregation pushes you toward smaller and/or quantized models — but as the tables above show, that's a perfectly livable place to be.
Generalizing from our results, self-hosting strategies fall into three buckets:
1. Run the largest, most capable model possible, regardless of token speed. With llama.cpp's CPU+GPU hybrid inference for MoE models, large models like GLM-5.2 or Kimi K2.6 are possible at ~16 tok/s generation, which amounts to ~$100 per 1M output tokens. Prompt processing speeds, however, are horrendous. This is a "Sunday afternoon deep-research" rig, not a production one.
2. Run a hardware-adequate model for internal needs. Small (<50 B) or medium (100–300 B) parameter models, serving a few workflows or a handful of users. This is the 2× RTX PRO 6000's comfort zone: one node, one engine, no disaggregation, Qwen3.6-35B-A3B-class models at 150–400 tok/s per session.
3. Run a hardware-adequate model for customer-facing or heavy agentic workloads. Disaggregation is a must. On 2× RTX PRO 6000 nodes this means smaller and/or quantized models per instance. Multiple nodes, each running a P/D pair of, say, NVFP4-quantized 35B-class MoEs, is a genuinely competitive building block.
And to pick the right model, understand your workload's shape:
Raw throughput only matters once you convert it to cost per token. The formula is simple:
cost per 1M tokens = monthly server cost × 1,000,000 / (tokens per second × 2,592,000)(2,592,000 = seconds in a 30-day month.) Illustratively, at $2,800/month for the 2× RTX PRO 6000 server, at 100% utilization:
[CHART – cost vs. utilization: Grouped bar chart on a log scale: cost per 1M output tokens for the rows above at 100% / 50% / 30% utilization. Shows utilization, not hardware price, as the dominant term in self-hosting economics.]

Even the least efficient configuration lands in the same ballpark as, or below, typical API pricing for comparable models, and prompt tokens are nearly free.
But note the asterisk on the whole table: that's 100% utilization, which nobody achieves. Real fleets with request gaps, idle periods and prefill/decode interference run at a fraction of synthetic peak – at ~30% effective utilization, multiply the numbers by ~3.3 (e.g. the flagship NVFP4 setup lands around $0.30 per 1M tokens). Utilization, not hardware price, is the dominant term in self-hosting economics. P/D disaggregation and sensible autoscaling exist precisely to keep that number up.
If we had to compress all of the above into a few rules of thumb:
Benchmark stability wasn’t universal. Gemma4-31B (dense) was the main offender: its single-GPU BF16 run failed 36% of concurrent decode requests, because the KV cache simply can’t support the offered concurrency, and requests got preempted or aborted, and its MTP variants were unstable wherever we ran them, failing 46% (BF16, single GPU), 51% (INT4, single GPU) and 32% (BF16, TP2) of concurrent decode requests. Its non-MTP INT4 and TP2 configurations complete cleanly.
Nemotron-3-Super-120B-A12B TP2 with MTP+NVFP4 failed all decode requests and 32% of 4K prefill. Treat that combination as broken until proven otherwise. Qwen3.6-27B TP2 with DFlash dropped ~19% of concurrent decode requests. Qwen3.8-27B with DFlash2 on a single GPU (BF16) dropped 31%; every other Qwen3.8-27B configuration completed cleanly. The Qwen3.8-Flash-Next NVFP4 checkpoints failed to load on the current dev build (a missing weight_scale on the n-gram embedding), so Flash-Next results were FP8-only.
All throughput numbers come from synthetic random-token prompts. Real prompts with cacheable prefixes will behave differently, usually better on prefill, thanks to prefix caching.
We report throughput and latency, not quality. Quantization changes model behavior, and the right way to evaluate that on your tasks is a separate project.
There are topics we deliberately didn't cover here, but at this scale, they matter, too:
Founded in 2014, DataPacket is a dedicated server provider operating a global low-latency network. With a footprint of 67 locations across 6 continents, DataPacket helps businesses–including gaming and video streaming companies–to deliver great online experiences.