Guardian Cloud · Module
specialised models, self-hosted on GPU
Guardian doesn’t call one general model for everything. It runs a fleet of specialised models — each trained or tuned for a single job — self-hosted on GPU in your region, so your data never leaves the contour.
Detection is fast, reasoning is deep, command generation is platform-aware, and source-code work is kept apart from system administration. The right model answers each question — and a frontier cloud model is consulted only on the highest-risk changes.
Every model is validated on real GPUs before it touches a server — and keeps learning through our daily briefings.
A model gateway routes each task to the model built for it — officers reason, shields detect, specialists generate commands, coders write code, the embedder retrieves doctrine. A frontier cloud model is consulted only on the highest-risk changes.
Model fleet — gateway routes each task to its specialist
Every model in the fleet is exercised on real GPUs before it touches a server. These are the documented numbers for the models we run today — nothing rounded up.
| Model | Role | Result |
|---|---|---|
| Audit Officer · Qwen3-30B-A3B-Thinking-2507 | Initial server audit | LOW / CRITICAL — correct |
| ITDR Officer · Gemma-4-26B-A4B | Intrusion verdict (ROE) | 100 / 100 · safety gate 100% |
| Cloud AI · Gemma-4-26B-A4B-it | Conversational interface | live |
| Detection shields · 3× Qwen3-4B-Instruct-2507 | First-line detection | 3 / 3 in 726–952 ms |
| All-Platform specialist · Qwen3-Coder-30B-A3B-Instruct | Command generation | 97.0% |
| Azure specialist | Command generation | 94.0% |
| GCP specialist | Command generation | 91.0% |
| AWS specialist | Command generation | in retraining |
| Coding · Qwen3-Coder gen + Qwen3-Thinking review | Source-code changes | generate → security review |
| Doctrine embedder · Qwen3-VL-Embedding-8B | Knowledge retrieval | semantic read — PASS |
| GLM-5.2 · self-hosted | Main brain · validator · initial testing | in-contour · high-risk review |
Validation score by model (%)
Every model in this fleet was trained by us on tens of thousands of high-quality instructions, and tested on hundreds of real work examples and incidents — each one for its own specialization.
| Model | Overall | In-distribution | Out-of-distribution | JSON-valid |
|---|---|---|---|---|
| All-Platform specialist | 97.0% | 49 / 50 | 48 / 50 | 97.0% |
| Azure specialist | 94.0% | 47 / 50 | 47 / 50 | 94.0% |
| GCP specialist | 91.0% | 45 / 50 | 46 / 50 | 95.0% |
| AWS specialist | in retraining | — | — | — |
100 tests per spec — 50 in-distribution + 50 out-of-distribution, AWQ on A100.
Open the raw test files for review
The AWS specialist is being retrained on a corrected dataset; its number lands here once measured. For closed, on-prem data centers, the Enterprise tier runs a lightweight fleet entirely on local hardware — validated at 100% across 400 scenarios.
Separation is a safety and quality decision, not an accident of history.
Get started
Guardian Cloud takes over administration and defence of your infrastructure. Connecting takes minutes.
Guardian’s fleet is built on the work of leading model labs. Our thanks to the teams whose models we fine-tune and self-host — each adapted for its own specialization.
Qwen Team — Alibaba
Qwen3-Coder-30B-A3B-Instruct · Qwen3-30B-A3B-Thinking-2507 · Qwen3-4B-Instruct-2507 · Qwen3-VL-Embedding-8B · Qwen3-VL-Reranker-2B
— cloud specialists, audit reasoning, detection shields, doctrine embedder & reranker
Google DeepMind
Gemma-4-26B-A4B · Gemma-4-26B-A4B-it
— the ITDR officer & Cloud AI
Z.ai
GLM-5.2
— the main brain and validator of the council: frontier-grade open weights under the MIT licence, self-hosted inside the perimeter
Anthropic
This platform was engineered together with Claude — Opus 4.7 and Opus 4.8. Our special thanks to Anthropic; their models helped us shape the architecture, write and harden the code. In production, all validation runs on our self-hosted GLM-5.2 — your code never leaves the contour.