OAK Assure is the assurance layer for defence AI. It grades a model against a known truth — performance, calibration, conformal coverage, adversarial robustness and drift — and returns a PASS / FAIL verdict with the evidence. And it now goes further: a provable robustness radius, a data flywheel that finds where your data is thin or dirty, and a sovereign, grounded AI copilot that cites, abstains and logs. AI you can field — because you can prove it.
OAK Assure now spans the whole life of a trustworthy defence-AI model — from a provable robustness guarantee, to the data flywheel that keeps it healthy, to a sovereign AI assistant you can actually field.
Defence wants an AI assistant — but a black box that confabulates, cites nothing and keeps no record is unusable in an ops room or a bid. OAK's copilot is the opposite: grounded in your own corpus, cites every source, abstains when it isn't sure, and logs every answer. An assistant you can trust.
Not an empirical number that moves every time the attacker does — a provable L2 robustness radius for any classifier, via randomized smoothing. A guarantee, not a hope.
Three questions, answered: coverage analysis shows where training data is thin; label quality finds which labels are dirty (confident learning); data-Shapley values what each example is worth. Collect and clean the data that actually moves the model.
A domain-agnostic engine that fuses any set of sources over any hypotheses — emitter ID, modulation ID, EO/IR target ID, multi-INT tracks — with a calibrated, abstaining verdict.
Class-conditional conformal prediction, temporal fusion and source-reliability estimation — coverage that holds per class and over time, not just on average.
Every suite classifier ships a model card — identity, provenance, metrics and an assurance grade — reproducible evidence behind a reliable-AI / TRL claim.
Each dimension is scored and gated against a configurable threshold. The result is a per-gate PASS / FAIL — illustrative defaults, set per programme.
| Dimension | Metric | Gate (default) |
|---|---|---|
| Performance | clean accuracy + per-class precision / recall / F1 + confusion | accuracy ≥ 0.80 |
| Calibration | Expected Calibration Error, Brier score, reliability curve | ECE ≤ 0.10 |
| Conformal coverage | split-conformal prediction-set coverage & average set size | no under-coverage |
| Adversarial robustness | accuracy on the adversarial split — spoof / decoy / noise | drop ≤ 0.25 |
| Input drift | Population Stability Index, train vs test | mean PSI ≤ 0.25 |
OAK Assure's five-gate battery is proven on a model trained on a real public dataset — not only the suite's synthetic records — the credibility step from demonstrator toward relevant-environment validation.
On the DroneAudioDataset — 360 drone + 600 background real clips — Assure grades a drone-vs-background detector across all five gates: accuracy, calibration, conformal coverage, adversarial robustness and drift.
The real-data model returns a clean PASS (acc 0.85 · ECE 0.048 · coverage 0.90 · robustness-drop 0.02 · drift 0.04), and the harden-and-re-grade path holds the verdict on real audio.
On a smaller eval the same model sits on the calibration line and Assure correctly FAILs it — the gate catches a marginally-miscalibrated real model, proving the verdict is real, not a rubber stamp.
--report mode for CI / scheduled assurance runs
OAK Assure doesn't only grade a model; it can harden it and re-grade, honestly, so a borderline FAIL becomes a defensible PASS.
OAK Assure is the second half of the reliable-AI loop: OAK Synthetic Environment makes the labelled data — with a known truth and an adversarial split — and this harness grades the model on it. The sim makes the data; the harness grades the AI.
It reuses the EW SUITE's AI libraries
(ai_core classifiers, ai_calibration metrics) and reads any JSONL
with the SynthEnv record schema. It produces assurance evidence; it is not an accreditation.
A verdict is where assurance starts, not where it ends. Five capabilities take a graded model the rest of the way: package it for a reviewer, keep it honest in production, audit the copilot, say what it doesn't know, and prove it hasn't been tampered with.
One step from a passing grade to the package a reviewer actually asks for: a Goal-Structuring-Notation assurance case and a crosswalk to NIST AI RMF, ISO 42001, ISO/IEC 24029, Canada's TBS Algorithmic Impact Assessment and the NATO / DoD AI principles — with the impact-tier obligations that come with it.
A grade is a snapshot; the field moves. A live monitor watches input drift, confidence, abstain rate, class mix and calibration, raises an alarm when any drifts, and tells you when the model is due for re-validation.
The assistant gets graded too. Every answer is scored for groundedness and faithfulness to its cited sources, checked for appropriate refusal, and scanned for prompt-injection — an audit of the copilot, not just a demo of it.
Not one number but two: how much uncertainty is irreducible noise in the data versus the model out of its depth — plus a novelty detector that flags inputs unlike anything it trained on. So "I'm not sure" means something.
The security half of assurance: scans for backdoor triggers and poisoned training labels, and estimates exposure to model extraction and membership-inference — signed into a tamper-evident, hash-chained integrity manifest.
Bring a classifier and a labelled set — or generate one — and we'll run the assurance pass with you.
Request a demo