Performance lab
A token-per-second number is not a system benchmark.
Users feel time to first token, reading cadence and queueing. Operators feel aggregate throughput, tail latency, energy and failures. Compliance teams need quality and refusal evidence. We measure all three views.
Build the benchmark contract
Generate the scenario that belongs in a proposal and acceptance test. This creates a protocol, not a performance estimate.
Protocol output
Choose a system and workload.
Public result status
We would rather publish an empty measured table than fill it with vendor numbers labelled as ours.
Current boundary: no absolute ANULUM throughput result is published yet. Vendor claims may inform planning, but they never enter the green measured class.
Seven acceptance scenarios
What every result must disclose
Exact artefact
Model revision and hash, licence, quantisation, serving engine and version.
Exact machine
Accelerators, memory, CPU, RAM, fabric, power limits, software and driver versions.
Exact workload
Input and output length, concurrency, cold/warm state, sampling and repetition count.
User latency
TTFT p50/p95 and inter-token latency p50/p95 — not averages alone.
Operator capacity
Per-request and aggregate output tokens/s, queueing, failures and energy per 1,000 output tokens.
Answer quality
Groundedness, citation correctness, unsupported-claim refusal and access-control tests.
Evidence legend
| Class | Meaning | May support a client promise? |
|---|---|---|
| Official | A source-controlled vendor or model-publisher specification. | Only for the specification it states. |
| Measured | ANULUM result with reproducible configuration and raw evidence. | Yes, within the tested envelope. |
| Calculated | Capacity or cost derived from explicit assumptions. | Planning only. |
| Provisional | Commercial or technical value awaiting quotation or validation. | No. |