768 GB of GPU memory. A 400 GB model with a 400,000-token context, in one box, no cloud.

Comino Grando Server
8× RTX PRO 6000

Independently reviewed by StorageReview and Alex Ziskind · 2026
Comino Grando Server
Independent reviews

Two reviews, one 8× RTX PRO server

“It's one of two boxes that I've recently tested where I don't feel like I'm limited by how many agents I could run.”Video timestamp · 25:24
“The hottest GPU I saw was 62 ºC and the CPU peaked at 65 ºC. So the liquid cooling here is not a gimmick.”Video timestamp · 7:57
Alex Ziskind · YouTube · 16/09/2026

“8 RTX PRO 6000's wasn't what I expected”

How far can you push AI workloads on a single liquid-cooled server?

Alex Ziskind explores the performance limits of a liquid-cooled server equipped with 8x NVIDIA RTX Pro 6000 GPUs and 768 GB of VRAM. This deep dive examines how such massive hardware configurations handle concurrent LLM requests, complex agentic workflows, and high-token context windows without relying on cloud infrastructure.
Comino Grando Server, rear panel, photographed by StorageReview
206.7 tok/saggregate output with eight Claude Code sessions on MiniMax M2.5, served locally (Claude Code is the client; the model is MiniMax, not Claude)
38.7 tok/sper developer at eight sessions, above the ~37 tok/s the article measured for Claude Opus 4.6 via the OpenRouter API
70+ dBat full fan speed, against 90 dB or more for air-cooled servers with the same GPU density
3–38 °Cambient operating range: an office corner or a small server room, no data centre required
StorageReview · Dylan Dougherty · 16/04/2026

Comino Grando RTX PRO 6000 Review: 768GB of VRAM in a Liquid-Cooled 4U Chassis

“The liquid cooling is not a bolt-on addition; it is the architecture.”StorageReview, on the design
“The system runs quietly enough to sit in a startup office, a small machine room, or a dedicated corner of an open workspace.”StorageReview, on noise and placement
“Teams of four to eight engineers can operate at near-optimal throughput without perceptible degradation in responsiveness.”StorageReview, on the Claude Code test
Read the review
Configuration

What you order

Configuration #23 in the Grando configurator. Both reviews tested the same unit: eight RTX PRO 6000 Server Edition, AMD EPYC, 512 GB DDR5.

Grando Server4U, rack or desktop, direct liquid cooling on every GPU and the CPU
8× NVIDIA RTX PRO 600096 GB each, 768 GB of GPU memory in total across eight cards
AMD EPYCsingle socket, 9004 / 9005 series
512 GB – 2 TB DDR5system memory, sized to your models
Redundant powerhot-swap power supply system
Why liquid cooling

No infrastructure needed

Eight GPUs at full load put out the heat of a small furnace. The Grando carries it away in its own sealed liquid loop, so there is nothing to build around it: no server room, no facility water, no data-centre cooling. All it needs is a room and electricity. Site requirements are listed in the product sheet.

Runs where servers usually can't

Direct liquid cooling on all eight GPUs and the CPU: no hot spots, no throttling under sustained AI load, and far quieter than an air-cooled 8-GPU server. An office corner or a small server room is enough.

Redundant power supply system

Hot-swap modules, each on its own cord: one module can fail and the server keeps running, and the load can be spread across ordinary circuits instead of one dedicated feed.

Maintainable liquid loop

Dripless quick disconnects on every GPU: swap a card without draining the loop. Filled, leak-checked and fully tested at the factory. StorageReview: “the maintenance simplicity of air-cooled systems to a high-performance liquid-cooled platform”.

Memory, measured in the review

768 GB. What fits.

Eight cards, 96 GB each, 768 GB in total. A model larger than one card is split across GPUs by the inference engine (tensor parallelism in vLLM). A model fits if its weights leave enough space for the KV cache at your required context size. Weights are in the format the model uses.

● measured in the review
○ internal benchmarks and estimations

ModelParametersWeightsNative contextFits in 768 GB
GLM 5.2 · NVFP4753B total / 32B active433 GB ●1MYes, 738 GB used at 400K ctx
Kimi K2.5 · INT41,027B / 32B595 GB ○256KYes, short context
GLM 5.3 Flash · NVFP4321B / 18B185 GB ●1MYes
DeepSeek V4 Flash · FP4291B / 10B160 GB ○1MYes, ran in the review
Qwen3 235B-A22B · NVFP4235B / 22B144 GB ●256KYes
Qwen 3.8 Flash-Next · NVFP4180B / 6B101 GB ○256KYes, ran in the review
GLM 5.2 · FP8753B / 32B768 GB ○1MNo room for KV cache
The largest load tested, GLM 5.2, used 738 GB of the 784 GB reported by the monitoring tools in the video (tool-reported totals differ slightly from the nominal 8 × 96 GB) and required 4.5 minutes to load, with a 400,000-token context window configured. Larger contexts require additional KV-cache memory. Please verify your requirements in the configurator. Configurator
Speed, measured in the review

Tokens per second

Reference figures from one independent review, not the ceiling of the hardware. Stock vLLM, NVFP4, tensor parallel over 8 GPUs, no speculative decoding, 300 W per-GPU cap. Decode: tokens per second per user. Prefill: prompt reading speed. TTFT: wait for the first token, at 2,000 and 128,000-token prompts.

Live conversation100+ tokens per second for one user: voice agents, chat, support that answers faster than you read
RankModelWeights1 user8 users, eachPrefillTTFT 2K / 128KFeels like
1Qwen 3.8 Flash-Next101 GB126 tok/s82 tok/s12,6000.25 s / 14.5 sInstant
2GLM 5.3 Flash needed a vLLM patch185 GB104 tok/s—8,3000.6 s / ~18 sInstant
3DeepSeek V4 Flash160 GB102 tok/s1,000 agents: 6–7K total8,7000.6 s / 26 sInstant
Deep reasoningthe biggest models the memory allows: research, analysis, agents that think for minutes, batch jobs
RankModelWeights1 userMany usersPrefillTTFT 2K / 128KFeels like
1Qwen3 235B-A22B144 GB85 tok/s254 total at 644,7790.5 s / crashed at 128K, fine to 32KFaster than reading
2GLM 5.2 753B433 GB48 tok/s111 total at 322,2001.0 s / 50 sReading pace
Agentic codingmeasured by StorageReview on the same eight-GPU build: concurrent Claude Code sessions served locally on MiniMax M2.5 via vLLM. Claude Code is only the client: the model is MiniMax M2.5, not Claude. The 37 tok/s baseline is Claude Opus 4.6 via OpenRouter, a speed reference, not a quality comparison. Per-developer and aggregate figures are averaged differently in the source and do not multiply exactly
Developers at oncePer developerAggregateExperience
1 session67.3 tok/s64.7 tok/s
2 sessions57.4 tok/s95.1 tok/s
4 sessions49.2 tok/s177.2 tok/sHighly responsive interactive coding
8 sessions38.7 tok/s206.7 tok/sSweet spot: above the ~37 tok/s Claude Opus 4.6 API baseline (OpenRouter) per developer
16 sessions31.1 tok/s105.8 tok/sPushing the limits of the 230B model
Batch servingpeak aggregate tokens per second at 256 concurrent requests, measured by StorageReview on the same eight-GPU build with vLLM. MiniMax M2.5 prefill-heavy peaked at 128 concurrent requests (7,357 tok/s); the table shows its 256-request result
ModelPrecisionEqual 256 in / 256 outPrefill-heavy 8K / 1KDecode-heavy 1K / 8K
GPT-OSS 120Bnative MXFP411,72621,6367,570
GPT-OSS 20Bnative MXFP417,28032,06111,187
Llama 3.1 8B InstructFP812,10920,1377,353
Qwen3 Coder 30B A3BFP810,98516,6594,907
Mistral Small 3.1 24BBF168,92511,8464,975
MiniMax M2.5, 230Bnative5,7537,1412,555
Which model for which job?

Four ways to use it

Coding agents, offlineDeploy a team of agents in your repository, each sending the full 50,000-token context with every interaction, while keeping your code securely on-site.
From the reviewsMiniMax M2.5 serving eight Claude Code sessions: 38.7 tok/s each, 206.7 tok/s in total, above the ~37 tok/s Claude Opus 4.6 API baseline (StorageReview). Claude Code is the client; the model served locally is MiniMax M2.5, not Claude. Qwen 3.8 Flash-Next: 82 tok/s per agent with eight agents (Ziskind).
Live voice and chatThis solution is for sales assistants, support engineers, and speakers to use during calls. It needs the first token in under one second and around 100 tokens per second per user. In the review this held for a single user (126 tok/s); with eight concurrent users the figure was 82 tok/s each.
From the reviewsQwen 3.8 Flash-Next: 126 tok/s for one user and 0.25 s to the first token on a 2,000-token prompt. GLM 5.3 Flash and DeepSeek V4 Flash follow at 104 and 102 tok/s (Ziskind).
Deep reasoningThis is the largest model supported by available memory, intended for research and analysis where answer quality is prioritized over speed.
From the reviewsGLM 5.2, 753B parameters: 48 tok/s decode for one user, with 738 GB in use and a 400,000-token context window configured; first token after ~1 s on a 2,000-token prompt and ~50 s on a 128,000-token prompt (Ziskind).
Video generationText-to-video and video-editing models run on the same GPUs. NVIDIA rates the RTX PRO 6000 Blackwell Server Edition at 3.3× the text-to-video generation speed of the previous-generation L40S, and models such as WAN and LTX Video run on RTX PRO Blackwell with FP4 support for over 2× performance at half the VRAM. Each card also carries four NVENC and four NVDEC engines with 4:2:2 support.
From NVIDIANot covered by the two reviews. Figures above are NVIDIA's: “NVIDIA Blackwell Universal Data Center GPU” (GTC, March 2025) and “RTX Blackwell GPUs Accelerate Video Editing” (RTX AI Garage, June 2025). 96 GB per card and eight cards let you render batches in parallel, on-site.
Ready when you are

Order the 8× RTX PRO server

Configuration #23 in the Grando configurator is the build both reviewers tested: eight RTX PRO 6000 with 768 GB of GPU memory. Need a different CPU, memory or storage? Talk to an engineer.