Hardware · model fit


Memory first. Then throughput, power and support.

The smallest machine that holds a model is not automatically the machine that serves your team. We separate weight residency from context, concurrency, latency and operational headroom.

Official specificationsPlanning calculation

Does the model fit?

Choose a model and the aggregate accelerator-memory class. The result is a planning gate, never a performance promise.

Calculated locally

Runs in this browser. No selection is transmitted or stored.

Planning result

Select a model and memory class.

Reference components, not anonymous tiers

These specifications anchor the architecture. Final systems add CPU, RAM, storage, fabric, redundancy, warranty and serving validation.

Verified 26 July 2026
96 GBRTX PRO 6000 Blackwell · ECC GDDR7 · 600 W
128 GBDGX Spark · coherent unified memory · proof-node class
141 GBH200 SXM · HBM3e · enterprise fabric
288 GBMI355X · HBM3e · 1.4 kW accelerator

A DGX Spark can hold large quantised weights in unified memory, but that does not make it equivalent to a high-bandwidth datacentre accelerator for latency or concurrency.

Four delivery envelopes

24–32 GB

Compact Node

< 1 kW

  • Apertus 8B
  • gpt-oss-20b
  • Qwen 27B INT4

Pilot, local RAG, one principal or small team.

192–384 GB

Vault

2–8 kW

  • larger multimodal models
  • higher concurrency
  • redundant storage and service paths

Multi-user archive system.

8 accelerators

Cell

8–40 kW

  • Kimi K2.5 validated INT4
  • DeepSeek V3 candidates
  • separate development and production

Regulated team and large-model system.

fabric programme

Sovereign Rack

50–200+ kW

  • Kimi K2 and K3 official serving envelopes
  • K3 starts at 8× GB300 or MI350X-class; production may be multi-node
  • liquid cooling and heat reuse

A facility engagement, never an appliance. Exact power and performance remain site- and benchmark-specific.

What the calculator deliberately does not promise

Context headroom

Longer context increases KV-cache and runtime demand. A weight fit at 2K can fail at 128K.

Concurrency

One fast user and twenty simultaneous users are different architectures even when the model is identical.

Quality

Hardware fit says nothing about groundedness, language quality or refusal behaviour on your archive.

Procurement rule: no named component becomes a client configuration until current Swiss availability, warranty path, lead time, exact chassis fit and benchmark results are recorded.

Fit is the first gate. Measurement is the second.

See the exact benchmark contract before treating a configuration as production-ready.

Open the performance lab View planning prices →