Apertus
Multilingual, Swiss-developed candidates for controlled RAG, document and multimodal work where local provenance matters.
Use: confidential archives, multilingual retrieval, compact and large nodes.
Open-model programme · reviewed 29 July 2026
ANULUM evaluates serious open-source and open-weight candidates across compact nodes, professional systems and rack-scale infrastructure. Kimi K3 is the current flagship open-weight reference for maximum-capability evaluation—not the automatic answer to every workload.
Open source is reserved here for models whose exact artefacts and terms support that description. Open weight means weights can be obtained under stated terms, which may be custom, gated or use-restricted. Every engagement reviews the exact checkpoint, licence version, intended use and redistribution path.
Catalogue rule: inclusion means “worth evaluating”, not “endorsed”, “safe”, “compliant” or “production-ready”. Third-party quantisations enter only after provenance and integrity checks.
The roster is broad enough to avoid vendor lock-in and narrow enough to remain source-reviewable.
Multilingual, Swiss-developed candidates for controlled RAG, document and multimodal work where local provenance matters.
Use: confidential archives, multilingual retrieval, compact and large nodes.
Reasoning, tool use and research-transparent or permissively licensed paths at very different system scales.
Use: agents, reasoning, coding, long context and controlled research.
Smaller model families for document extraction, local assistants, vision and cost-sensitive private deployment.
Use: OCR, classification, office workflows, edge and single-node systems.
Multimodal and agentic candidates between professional multi-GPU systems and rack-class deployments.
Use: multilingual assistants, visual documents, tools and higher-concurrency services.
Purpose-shaped candidates belong in specialist lanes instead of silently becoming the default general model.
Use: software engineering, policy checks, guardrails and narrow automation.
Official weights and serving guidance make K3 a credible maximum-capability sovereign-evaluation candidate. It remains a rack programme with client-specific proof gates.
Measure: retrieval recall, claim support, citation correctness, abstention and reviewer time on a frozen document set.
Measure: factual defects, instruction adherence, source fidelity, redaction behaviour and human revision burden.
Measure: held-out task completion, test pass rate, tool-schema validity, permission adherence, rollback and receipt completeness.
Measure: OCR and table accuracy, layout grounding, image-supported claims and failure visibility across real client formats.
Measure: German, English, French and Italian task quality, terminology stability, named-entity preservation and citation fidelity.
Measure: time to first token, sustained output, concurrency, memory, context cost, energy and recovery on the declared system.
Score status
Official cards and reports establish publisher-reported benchmarks. ANULUM publishes a measured result only with the checkpoint hash, engine and settings, declared hardware, dataset or task version, run record and acceptance threshold.
Sources reviewed 29 July 2026. Model inventories and terms change; the exact checkpoint and licence are re-verified at engagement start.
Start with the workload, custody boundary, languages, latency target and hardware you are willing to operate.