Apertus v1.5 8B
Swiss, multilingual and compact, with text, image and audio input. A candidate for controlled RAG after exact task evaluation.
Model intelligence
A model is deployable only when its weights, licence, serving engine, memory envelope and behaviour all fit. This catalogue separates those facts — and dates every conclusion.
Filter by the system you want to own, not by a leaderboard headline.
12 models shown
Swiss, multilingual and compact, with text, image and audio input. A candidate for controlled RAG after exact task evaluation.
A Swiss large-model candidate with text, image and audio input; exact multimodal and long-context workloads require measurement.
A compact reasoning and tool-use candidate. Active parameters affect compute, not the full weight-residency requirement.
The official single-80-GB fit is a weight/runtime claim, not a multi-user throughput promise. Concurrency is benchmarked separately.
A compact multimodal option for document and vision workflows, subject to exact quantisation and context evaluation.
Interesting for long-context private archives. Published cache and latency advantages remain vendor/research evidence until measured on the delivery stack.
Rack-class despite 32B active parameters: the expert weights must remain resident. The official deployment guide is the sizing anchor.
Multimodal and more memory-efficient than FP8, but still a multi-GPU system whose context and concurrency must be proven.
A current multimodal checkpoint with image input. Local video is not claimed: the official card describes video support as experimental and API-only.
A coding-agent derivative with preserved reasoning traces and image input. It is evaluated as a specialist, not assumed to be the default archive model.
A large MoE candidate where the custom model licence and validated multi-GPU serving path fit the client's governance boundary.
We do not advertise K3 as an on-prem model. Eligibility begins only after official weights, licence, hashes, engine support and a reproducible benchmark exist.
The Kimi name spans very different deployment realities. Linear can fit a professional node when quantised. K2 through K2.7 are large-system programmes with distinct general, multimodal and coding roles. K3 remains a watch item, not an authorised local deployment claim.
Parameter honesty: active parameters estimate compute per token. Total resident weights, KV cache, runtime buffers, context and concurrency determine memory.
A source link does not constitute an ANULUM quality endorsement. Model selection follows evaluation on the client's tasks and constraints.
We test candidate models against your documents, language, refusal rules and infrastructure boundary.