Compact Node
< 1 kW
- Apertus 8B
- gpt-oss-20b
- Qwen 27B INT4
Pilot, local RAG, one principal or small team.
Hardware · model fit
The smallest machine that holds a model is not automatically the machine that serves your team. We separate weight residency from context, concurrency, latency and operational headroom.
Choose a model and the aggregate accelerator-memory class. The result is a planning gate, never a performance promise.
Select a model and memory class.
These specifications anchor the architecture. Final systems add CPU, RAM, storage, fabric, redundancy, warranty and serving validation.
A DGX Spark can hold large quantised weights in unified memory, but that does not make it equivalent to a high-bandwidth datacentre accelerator for latency or concurrency.
< 1 kW
Pilot, local RAG, one principal or small team.
0.6–1.5 kW
The normal starting point for serious local production.
2–8 kW
Multi-user archive system.
8–40 kW
Regulated team and large-model system.
50–200+ kW
A facility engagement, never an appliance. Exact power and performance remain site- and benchmark-specific.
Longer context increases KV-cache and runtime demand. A weight fit at 2K can fail at 128K.
One fast user and twenty simultaneous users are different architectures even when the model is identical.
Hardware fit says nothing about groundedness, language quality or refusal behaviour on your archive.
Procurement rule: no named component becomes a client configuration until current Swiss availability, warranty path, lead time, exact chassis fit and benchmark results are recorded.
See the exact benchmark contract before treating a configuration as production-ready.