AI Assurance · Test, Evaluation, Verification & Validation

Would you field this model?

OAK Assure is the assurance layer for defence AI. It grades a model against a known truth — performance, calibration, conformal coverage, adversarial robustness and drift — and returns a PASS / FAIL verdict with the evidence. And it now goes further: a provable robustness radius, a data flywheel that finds where your data is thin or dirty, and a sovereign, grounded AI copilot that cites, abstains and logs. AI you can field — because you can prove it.

UNCLASSIFIED · representational // T&E tooling — produces assurance evidence, not an accreditation
Certified
provable robustness radius
Grounded
assured AI copilot — cites & abstains
Flywheel
find thin & dirty training data
PASS/FAIL
a verdict, with the evidence
New — the assurance layer, expanded

Beyond the verdict: prove it, ground it, improve it

OAK Assure now spans the whole life of a trustworthy defence-AI model — from a provable robustness guarantee, to the data flywheel that keeps it healthy, to a sovereign AI assistant you can actually field.

Sovereign assured AI copilot

Defence wants an AI assistant — but a black box that confabulates, cites nothing and keeps no record is unusable in an ops room or a bid. OAK's copilot is the opposite: grounded in your own corpus, cites every source, abstains when it isn't sure, and logs every answer. An assistant you can trust.

Certified robustness

Not an empirical number that moves every time the attacker does — a provable L2 robustness radius for any classifier, via randomized smoothing. A guarantee, not a hope.

The data flywheel

Three questions, answered: coverage analysis shows where training data is thin; label quality finds which labels are dirty (confident learning); data-Shapley values what each example is worth. Collect and clean the data that actually moves the model.

Assured fusion

A domain-agnostic engine that fuses any set of sources over any hypotheses — emitter ID, modulation ID, EO/IR target ID, multi-INT tracks — with a calibrated, abstaining verdict.

Deeper conformal guarantees

Class-conditional conformal prediction, temporal fusion and source-reliability estimation — coverage that holds per class and over time, not just on average.

TEVV model cards

Every suite classifier ships a model card — identity, provenance, metrics and an assurance grade — reproducible evidence behind a reliable-AI / TRL claim.

The core verdict

Five dimensions, five gates, one verdict

Each dimension is scored and gated against a configurable threshold. The result is a per-gate PASS / FAIL — illustrative defaults, set per programme.

DimensionMetricGate (default)
Performanceclean accuracy + per-class precision / recall / F1 + confusionaccuracy ≥ 0.80
CalibrationExpected Calibration Error, Brier score, reliability curveECE ≤ 0.10
Conformal coveragesplit-conformal prediction-set coverage & average set sizeno under-coverage
Adversarial robustnessaccuracy on the adversarial split — spoof / decoy / noisedrop ≤ 0.25
Input driftPopulation Stability Index, train vs testmean PSI ≤ 0.25
Validated on real data — not just synthetic

Grades a model trained on real public data

OAK Assure's five-gate battery is proven on a model trained on a real public dataset — not only the suite's synthetic records — the credibility step from demonstrator toward relevant-environment validation.

Real acoustic model

On the DroneAudioDataset — 360 drone + 600 background real clips — Assure grades a drone-vs-background detector across all five gates: accuracy, calibration, conformal coverage, adversarial robustness and drift.

Full battery, real verdict

The real-data model returns a clean PASS (acc 0.85 · ECE 0.048 · coverage 0.90 · robustness-drop 0.02 · drift 0.04), and the harden-and-re-grade path holds the verdict on real audio.

A genuine gate

On a smaller eval the same model sits on the calibration line and Assure correctly FAILs it — the gate catches a marginally-miscalibrated real model, proving the verdict is real, not a rubber stamp.

TEVV Workbench

Load, grade, and read the verdict

  • Load or generate a dataset, pick a model and target, run the assessment
  • A PASS / FAIL verdict with the per-gate table and per-perturbation robustness
  • Reliability & robustness diagnostics; compare all models side by side
  • Export a self-contained HTML assurance report — evidence behind a reliable-AI / TRL claim
  • A headless --report mode for CI / scheduled assurance runs
OAK Assure — TEVV workbench: PASS verdict with the per-gate table
Close the loop

Harden a borderline model — and re-grade it

OAK Assure doesn't only grade a model; it can harden it and re-grade, honestly, so a borderline FAIL becomes a defensible PASS.

  • Calibrate — post-hoc temperature scaling on a held-out split. Monotonic, so clean accuracy and the argmax are unchanged; it only honestens the confidences → lower ECE
  • Augment — robust training on decoy / noise / spoof-style perturbations of the training split (never the held-out test split) → higher adversarial robustness
  • A worked example: a naïve model FAILs the calibration gate and is weak to decoys; hardening drives ECE and decoy-robustness back into gate — with clean accuracy, coverage and drift preserved
Reliable-AI, end to end

The other half of the story

OAK Assure is the second half of the reliable-AI loop: OAK Synthetic Environment makes the labelled data — with a known truth and an adversarial split — and this harness grades the model on it. The sim makes the data; the harness grades the AI.

It reuses the EW SUITE's AI libraries (ai_core classifiers, ai_calibration metrics) and reads any JSONL with the SynthEnv record schema. It produces assurance evidence; it is not an accreditation.

Govern & sustain — the assurance lifecycle

From a passing grade to an accreditation — and staying assured in the field

A verdict is where assurance starts, not where it ends. Five capabilities take a graded model the rest of the way: package it for a reviewer, keep it honest in production, audit the copilot, say what it doesn't know, and prove it hasn't been tampered with.

Accreditation dossier

One step from a passing grade to the package a reviewer actually asks for: a Goal-Structuring-Notation assurance case and a crosswalk to NIST AI RMF, ISO 42001, ISO/IEC 24029, Canada's TBS Algorithmic Impact Assessment and the NATO / DoD AI principles — with the impact-tier obligations that come with it.

Runtime assurance

A grade is a snapshot; the field moves. A live monitor watches input drift, confidence, abstain rate, class mix and calibration, raises an alarm when any drifts, and tells you when the model is due for re-validation.

Copilot assurance

The assistant gets graded too. Every answer is scored for groundedness and faithfulness to its cited sources, checked for appropriate refusal, and scanned for prompt-injection — an audit of the copilot, not just a demo of it.

Uncertainty, decomposed

Not one number but two: how much uncertainty is irreducible noise in the data versus the model out of its depth — plus a novelty detector that flags inputs unlike anything it trained on. So "I'm not sure" means something.

Model integrity

The security half of assurance: scans for backdoor triggers and poisoned training labels, and estimates exposure to model extraction and membership-inference — signed into a tamper-evident, hash-chained integrity manifest.

Grade your model

Bring a classifier and a labelled set — or generate one — and we'll run the assurance pass with you.

Request a demo
Precision engineering for defence, aerospace & space — Kanata, Ontario, Canada