Apertus v1.5 8B
Swiss, multilingual and compact, with text, image and audio input. A candidate for controlled RAG after exact task evaluation.
Model intelligence
A model is deployable only when its weights, licence, serving engine, memory envelope and behaviour all fit. This catalogue separates those facts — and dates every conclusion.
Compare the wider open-model candidate programme, use cases and benchmark gates →
Filter by the system you want to own, not by a leaderboard headline.
12 models shown
Swiss, multilingual and compact, with text, image and audio input. A candidate for controlled RAG after exact task evaluation.
A Swiss large-model candidate with text, image and audio input; exact multimodal and long-context workloads require measurement.
A compact reasoning and tool-use candidate. Active parameters affect compute, not the full weight-residency requirement.
The official single-80-GB fit is a weight/runtime claim, not a multi-user throughput promise. Concurrency is benchmarked separately.
A compact multimodal option for document and vision workflows, subject to exact quantisation and context evaluation.
Interesting for long-context private archives. Published cache and latency advantages remain vendor/research evidence until measured on the delivery stack.
Rack-class despite 32B active parameters: the expert weights must remain resident. The official deployment guide is the sizing anchor.
Multimodal and more memory-efficient than FP8, but still a multi-GPU system whose context and concurrency must be proven.
A current multimodal checkpoint with image input. Local video is not claimed: the official card describes video support as experimental and API-only.
A coding-agent derivative with preserved reasoning traces and image input. It is evaluated as a specialist, not assumed to be the default archive model.
A large MoE candidate where the custom model licence and validated multi-GPU serving path fit the client's governance boundary.
Official weights now unblock a sovereign evaluation and rack-architecture programme. Production remains gated by exact engine, context, concurrency, tool-schema, security, quality and licence acceptance; current public serving guidance supports image input, not video or audio.
The Kimi name spans very different deployment realities. Linear can fit a professional node when quantised. K2 through K2.7 are large-system programmes. K3 is now an open-weight, rack-scale candidate: available for controlled evaluation and architecture work, but not pre-validated production.
Parameter honesty: active parameters estimate compute per token. Total resident weights, KV cache, runtime buffers, context and concurrency determine memory.
This page uses open weight for downloadable checkpoints under publisher terms and reserves open source for exact versions whose artefacts and licences support that description. Licence, checkpoint provenance and intended use are reviewed before deployment.
Benchmark status: publisher scores remain publisher evidence. No card on this page carries an ANULUM quality score unless a reproducible ANULUM run is linked.
A source link does not constitute an ANULUM quality endorsement. Model selection follows evaluation on the client's tasks and constraints.
We test candidate models against your documents, language, refusal rules and infrastructure boundary.