Questions & answers
The difficult questions belong before procurement.
Private AI combines infrastructure, information governance and probabilistic software. These answers state what can be engineered, what must be measured and what remains a client or adviser decision.
01 · Data custody
Where data goes, who can reach it
Does our archive ever need to leave our chosen boundary?
No. Air-gapped and on-premises architectures can keep documents, indexes, prompts and responses inside the client-controlled boundary. Dedicated Swiss hosting is a different custody choice, and a policy-gated hybrid explicitly defines which non-sensitive requests may leave. The selected architecture and its exceptions are documented before implementation.
Does ANULUM retain a copy of client data?
Not by default. The client archive belongs on the client-owned or expressly contracted environment. Support access, diagnostic exports and temporary transfer paths are separately scoped, time-bound and logged. A contract must state retention and deletion duties for any exceptional copy.
Will documents or prompts be used to train a model?
Not unless the client explicitly commissions and authorises that separate activity. Retrieval over an archive does not require training the base model on the archive. Indexing, adapter training and full model training are different processes and must not be blurred.
Who has administrator access?
The architecture names every administrator class, authentication method and approval path. It can be designed with no standing ANULUM access: the client opens a maintenance window, authorises a named engineer, records the session and closes access afterwards. Break-glass access, if any, is separately controlled and tested.
02 · Models
Open weights without open-ended claims
What is the difference between open source and open weights?
Open weights means the trained parameter files are available under stated terms. It does not automatically mean the training data, complete training code or every component is open source. We record the licence and usage policy of each material component and have the client obtain legal interpretation where needed.
Do we need the largest available model?
Usually not. Retrieval quality, document parsing, permissions, citations and evaluation often matter more than parameter count. A smaller model that meets the acceptance threshold on the real workload can be faster, cheaper, easier to recover and easier to keep within one security boundary.
What do quantisation and context length change?
Quantisation reduces weight memory but can change quality and engine compatibility. Longer context increases cache memory and can reduce throughput; an advertised maximum is not a promise at useful concurrency. We test the exact model, format, engine, prompt shape and user load proposed for production.
Can the model be replaced later?
Yes, if the system preserves portable documents, indexes or rebuild procedures, prompts, evaluation cases and version records. A replacement is a governed change: licences, behaviour, citations, refusals, latency, memory and rollback are checked before it becomes the accepted baseline.
Can ANULUM deploy Kimi K3?
Not as a confirmed local product today. Our dated catalogue keeps Kimi K3 on readiness watch because an official local checkpoint, licence and reproducible self-hosting path have not been established. Kimi Linear and K2-family checkpoints have different, documented deployment realities. See the model catalogue.
03 · Hardware and performance
Capacity is a workload result
How fast will the system be?
A credible answer names time to first token, generation rate, end-to-end answer latency, input length, output length and concurrency. Public values are planning evidence only. The proof node measures representative client questions on the candidate production stack and reports the distribution, not only the fastest run.
How many simultaneous users can one node serve?
That depends on model, quantisation, context, output length, batching, retrieval and the service-level target. We load-test scenarios rather than multiply a single-user tokens-per-second number. Capacity planning also reserves memory and thermal headroom for degraded conditions.
Can existing server hardware be reused?
Sometimes. We check GPU support, memory, PCIe topology, power, cooling, storage endurance, warranty and serving-engine compatibility. The internal HP ML350 laboratory platform is useful for reproducible engineering, but it is not represented as current production hardware or as a customer reference.
What happens if a GPU, disk or index fails?
The design assigns a recovery path: degraded service or spare capacity, encrypted backups, configuration capture, index reconstruction and tested restore. Recovery time and recovery point targets are agreed according to business impact. A backup is not accepted merely because a job reported success; restore evidence matters.
Can the waste heat be reused?
Potentially for dense, steady deployments. Useful heat depends on temperature level, load profile, distance, building demand and local engineering. ANULUM develops the IT-side heat and telemetry assumptions; licensed HVAC and electrical partners validate and execute building works after a site survey.
04 · Security and governance
On-premises is a location, not a control framework
Does on-premises deployment make the system compliant?
No. Location can reduce some transfer and third-party risks, but lawful purpose, transparency, data minimisation, permissions, security, human responsibility, retention and individual rights still require design. ANULUM supplies technical facts and controls; the client and its advisers determine the applicable legal position.
When is a data-protection impact assessment needed?
That depends on the jurisdiction, data, purpose and risk. The Swiss FDPIC states that the technology-neutral Federal Data Protection Act applies to AI-supported processing and identifies impact assessment for high-risk cases. Liechtenstein operates under its own applicable data-protection framework. The architecture review produces facts for counsel or the data-protection officer; it is not legal advice.
How are air-gapped systems updated?
Through a documented offline release: approved source, malware scan, hashes and signatures, controlled media, two-person verification where required, staged installation, acceptance tests and retained rollback artefacts. The update path is part of the threat model because removable media is itself a boundary crossing.
How are prompt injection and permission leakage tested?
The evaluation corpus includes adversarial documents, conflicting instructions, cross-role retrieval attempts and requests for unsupported conclusions. We test the full chain — parser, index, access filter, prompt construction, model and user interface — because a safe model cannot compensate for an unsafe retrieval layer.
Can AI output replace professional judgement?
No. The system can retrieve, summarise, compare and draft with citations, but legal, financial and clinical decisions remain with authorised professionals. The accepted use policy states prohibited reliance and escalation conditions, and the interface should make sources and uncertainty easy to inspect.
05 · Commercial and operating model
Know what you own and what continues
What does the architecture review deliver?
A fixed-scope decision pack: current-state and custody map, use-case boundary, threat and dependency observations, candidate deployment class, indicative bill of materials and operating cost, proof-node proposal, principal risks and an explicit stop/revise/proceed recommendation.
Are the public prices quotations?
No. They are dated, non-binding planning ranges with stated inclusions and exclusions. A quotation follows discovery of workload, location, warranty, integration, support and tax conditions. Hardware prices and availability can move quickly, so validity periods are explicit. See the current ranges.
Who owns the hardware, data and deployment artefacts?
The intended default is client ownership of client data and client-procured hardware. Software licences, third-party components, ANULUM pre-existing tools and project-specific deliverables have different rights, which the contract itemises. Custody should never depend on an ambiguous ownership assumption.
What ongoing support is available?
Support can cover health review, backup checks, security and dependency updates, capacity monitoring, incident response, controlled model changes and periodic evaluation. Exact hours, response targets, access model, exclusions and client responsibilities are defined for the accepted baseline.
Can we leave or operate without ANULUM?
That is a design requirement. The exit package can include data and index exports, model identifiers and licences, configuration, dependency inventory, runbooks, backup/restore procedure, risk register and knowledge transfer. Any component that cannot be transferred is identified before commitment.
Primary regulatory reading
- Swiss FDPIC: AI and data protection ↗
- FINMA Guidance 08/2024: AI governance and risk management ↗
- Liechtenstein Data Protection Authority: artificial intelligence ↗
Links are supplied as primary reading, not as a compliance certification or substitute for advice specific to the institution and use case. Source review: 26 July 2026.