768 GB of GPU memory. A 400 GB model with a 400,000-token context, in one box, no cloud.
“8 RTX PRO 6000's wasn't what I expected”
How far can you push AI workloads on a single liquid-cooled server?
Alex Ziskind explores the performance limits of a liquid-cooled server equipped with 8x NVIDIA RTX Pro 6000 GPUs and 768 GB of VRAM. This deep dive examines how such massive hardware configurations handle concurrent LLM requests, complex agentic workflows, and high-token context windows without relying on cloud infrastructure.
Comino Grando RTX PRO 6000 Review: 768GB of VRAM in a Liquid-Cooled 4U Chassis
Configuration #23 in the Grando configurator. Both reviews tested the same unit: eight RTX PRO 6000 Server Edition, AMD EPYC, 512 GB DDR5.
Eight GPUs at full load put out the heat of a small furnace. The Grando carries it away in its own sealed liquid loop, so there is nothing to build around it: no server room, no facility water, no data-centre cooling. All it needs is a room and electricity. Site requirements are listed in the product sheet.
Direct liquid cooling on all eight GPUs and the CPU: no hot spots, no throttling under sustained AI load, and far quieter than an air-cooled 8-GPU server. An office corner or a small server room is enough.
Hot-swap modules, each on its own cord: one module can fail and the server keeps running, and the load can be spread across ordinary circuits instead of one dedicated feed.
Dripless quick disconnects on every GPU: swap a card without draining the loop. Filled, leak-checked and fully tested at the factory. StorageReview: “the maintenance simplicity of air-cooled systems to a high-performance liquid-cooled platform”.
Eight cards, 96 GB each, 768 GB in total. A model larger than one card is split across GPUs by the inference engine (tensor parallelism in vLLM). A model fits if its weights leave enough space for the KV cache at your required context size. Weights are in the format the model uses.
● measured in the review
○ internal benchmarks and estimations
Reference figures from one independent review, not the ceiling of the hardware. Stock vLLM, NVFP4, tensor parallel over 8 GPUs, no speculative decoding, 300 W per-GPU cap. Decode: tokens per second per user. Prefill: prompt reading speed. TTFT: wait for the first token, at 2,000 and 128,000-token prompts.
Configuration #23 in the Grando configurator is the build both reviewers tested: eight RTX PRO 6000 with 768 GB of GPU memory. Need a different CPU, memory or storage? Talk to an engineer.