Model intelligence


Choose the model after the workload.

A model is deployable only when its weights, licence, serving engine, memory envelope and behaviour all fit. This catalogue separates those facts — and dates every conclusion.

Official model factsCalculated memory plansRegistry verified 29 July 2026 · Kimi K3 review by 12 August 2026

Compare the wider open-model candidate programme, use cases and benchmark gates →

Explore the deployment catalogue

Filter by the system you want to own, not by a leaderboard headline.

12 models shown

Swiss AI

Apertus v1.5 8B

Deployable
Total / active
8B / 8B
Context
262K
Licence
Apache 2.0 + policy
Planning memory
16 GB quantised

Swiss, multilingual and compact, with text, image and audio input. A candidate for controlled RAG after exact task evaluation.

Swiss AI

Apertus v1.5 70B

Deployable
Total / active
70B / 70B
Context
262K
Licence
Apache 2.0 + policy
Planning memory
48 GB INT4

A Swiss large-model candidate with text, image and audio input; exact multimodal and long-context workloads require measurement.

OpenAI

gpt-oss-20b

Deployable
Total / active
21B / 3.6B
Context
131K
Licence
Apache 2.0
Planning memory
24 GB

A compact reasoning and tool-use candidate. Active parameters affect compute, not the full weight-residency requirement.

OpenAI

gpt-oss-120b

Deployable
Total / active
117B / 5.1B
Context
131K
Licence
Apache 2.0
Official fit
one 80 GB GPU

The official single-80-GB fit is a weight/runtime claim, not a multi-user throughput promise. Concurrency is benchmarked separately.

Qwen

Qwen3.5 27B

Deployable
Total / active
27B / 27B
Context
262K
Licence
Apache 2.0
Planning memory
24 GB INT4

A compact multimodal option for document and vision workflows, subject to exact quantisation and context evaluation.

Moonshot AI

Kimi Linear 48B-A3B

Deployable
Total / active
48B / 3B
Context
1M
Licence
MIT
Planning memory
48 GB INT4

Interesting for long-context private archives. Published cache and latency advantages remain vendor/research evidence until measured on the delivery stack.

Moonshot AI

Kimi K2

Large-system
Total / active
1T / 32B
Context
128K
Licence
Modified MIT
Official FP8 path
16 GPUs at 128K

Rack-class despite 32B active parameters: the expert weights must remain resident. The official deployment guide is the sizing anchor.

Moonshot AI

Kimi K2.5

Large-system
Total / active
1T / 32B
Context
256K
Licence
Modified MIT
Format
native INT4

Multimodal and more memory-efficient than FP8, but still a multi-GPU system whose context and concurrency must be proven.

Moonshot AI

Kimi K2.6

Large-system
Total / active
1T / 32B
Context
256K
Licence
Modified MIT
Format
native INT4

A current multimodal checkpoint with image input. Local video is not claimed: the official card describes video support as experimental and API-only.

Moonshot AI

Kimi K2.7 Code

Coding specialist
Total / active
1T / 32B
Context
256K
Licence
Modified MIT
Format
native INT4

A coding-agent derivative with preserved reasoning traces and image input. It is evaluated as a specialist, not assumed to be the default archive model.

DeepSeek

DeepSeek V3

Large-system
Total / active
671B / 37B
Context
128K
Licence
Model licence
Planning class
Cell / Rack

A large MoE candidate where the custom model licence and validated multi-GPU serving path fit the client's governance boundary.

Moonshot AI

Kimi K3

Rack programmeUnmeasured by ANULUM
Total / active
2.8T / 104B
Context
1M
Licence
Kimi K3 License
Official starting path
8× GB300 or MI350X-class

Official weights now unblock a sovereign evaluation and rack-architecture programme. Production remains gated by exact engine, context, concurrency, tool-schema, security, quality and licence acceptance; current public serving guidance supports image input, not video or audio.

Kimi without the fog

The Kimi name spans very different deployment realities. Linear can fit a professional node when quantised. K2 through K2.7 are large-system programmes. K3 is now an open-weight, rack-scale candidate: available for controlled evaluation and architecture work, but not pre-validated production.

Parameter honesty: active parameters estimate compute per token. Total resident weights, KV cache, runtime buffers, context and concurrency determine memory.

A catalogue is not a claim that every model is open source.

This page uses open weight for downloadable checkpoints under publisher terms and reserves open source for exact versions whose artefacts and licences support that description. Licence, checkpoint provenance and intended use are reviewed before deployment.

Benchmark status: publisher scores remain publisher evidence. No card on this page carries an ANULUM quality score unless a reproducible ANULUM run is linked.

Bring the workload, not a preferred logo.

We test candidate models against your documents, language, refusal rules and infrastructure boundary.

Define the evaluation Check hardware fit →