VERDICT · Testing & assurance · under Frontier Lab

MEAIOW VERDICT

You cannot declare a system works. You assure it does.

VERDICT is MEAIOW’s AI testing and assurance framework. Testing AI is not testing conventional software — AI behaves probabilistically, learns from shifting data and can act autonomously. VERDICT replaces assertion with evidence: seven disciplines, four quality themes, nine testing modules and a continuous lifecycle, applied from before go-live through the whole life of every system.

Validated · Evidenced · Risk-Calibrated · Documented · Independent · Continuous · Trusted

VERDICT · Version 1.0 · The fifth dimension of FIRST AI-D · British AI, built for the world

Testing, evaluation, assurance

Three activities. One chain of evidence.

Testing checks the complete AI-enabled system — infrastructure, APIs, integrations, interfaces and data flows. Evaluation assesses the model itself against agreed quality attributes. Assurance is the confidence that follows: testing generates the evidence, evaluation interprets it, assurance proves the system is safe, fair and accountable.

AI types covered

Rule-based / deterministic

Predictable logic. Testing verifies every rule and boundary case is correct and complete.

Machine learning

Trained on data. Testing targets data quality, accuracy, generalisation and bias.

Generative

Open-ended output. Testing emphasises appropriateness, factual accuracy and control of harmful content.

Agentic / autonomous

Probabilistic and adaptive. Testing targets safe autonomy, novel situations, reward hacking and human override.

The seven disciplines

V·E·R·D·I·C·T

Not seven tests passed in sequence — seven disciplines forming one continuous assurance cycle across the life of every system.

V

Verified Performance

Performance is verified against documented, reproducible benchmarks before go-live. A system that has not been verified has not been assured.

E

Evidence-Led Assurance

Every statement rests on primary evidence VERDICT generates — never vendor assertion or self-reported compliance. No evidence, no assurance.

R

Risk-Calibrated Testing

Test rigour is proportionate to each deployment’s risk. Higher risk demands more rigorous, more adversarial, more frequent testing.

D

Documented Audit Trail

Every assessment produces a tamper-evident record regulators, boards and investigators can examine independently.

I

Independent Oversight

The team that builds a system does not assess it. VERDICT operates structurally apart from delivery, through an independent review line.

C

Continuous Assurance

Assurance does not end at go-live. It continues through monitoring, scheduled testing and audit cycles set by risk classification.

T

Trust Made Measurable

Trust is not a feeling. It is a product — quantified, evidence-backed, independently verified and sustained over the operational life of every system.

Core quality attributes

What we test for, grouped into four themes.

These guide both system-level testing and model-level evaluation across every deployment.

Safety & Ethics

Avoiding harm, respecting rights, non-discrimination — ethics treated as testable risk.

Openness & Trust

Explainability, transparency and traceability of decisions to those they affect.

Performance & Resilience

Accuracy, robustness under stress, graceful and safe failure.

User & Context Fit

Fitness for the people and the real-world conditions the system serves.

Continuous defensive assurance

Assurance is a lifecycle, not a one-time event.

VERDICT defines testing activities and deliverables at every phase — quality built in from the start, not retrofitted before launch.

01

Planning & Design

02

Data Collection & Prep

03

Model Development

04

Validation & Verification

05

Operational Readiness

06

Deployment

07

Monitoring & Continuous Assurance

The nine testing modules

The lifecycle says when. The modules say what.

Nine modular building blocks of AI testing, emphasised or de-emphasised by AI type and risk level. Full metrics and thresholds are in the framework document.

MODULE 01

Data & Input Validation

Inputs are valid, representative and compliant — schema checks, statistical profiling, PII redaction and lineage.

MODULE 02

Model Functionality

The model works as intended across typical inputs, edge cases and failure modes — re-tested after every change.

MODULE 03

Bias & Fairness

Outputs tested by group, not just in aggregate — counterfactual, demographic parity and equal-opportunity analysis.

MODULE 04

Explainability & Transparency

Decisions can be understood, traced and explained to overseers and to the people affected.

MODULE 05

Robustness & Adversarial

Behaviour under stress and deliberate attack — perturbations, jailbreaks, prompt injection, fail-safe verification.

MODULE 06

Performance & Efficiency

Speed, scale and resource use under realistic and peak load, including latency and throughput targets.

MODULE 07

Integration & System

The AI works as part of the wider system — end-to-end workflows, dependencies and graceful degradation.

MODULE 08

User Acceptance & Ethical Review

It works for the people it serves and clears a formal ethical sign-off before deployment.

MODULE 09

Continuous Monitoring

After go-live: dashboards, drift alerts, incident response and scheduled re-testing as data and risk shift.

The assurance cycle

Governance without independent verification is self-certification.

—Shakil Siddiqui, Founder & CEO of MEAIOW

Before go-live — six domains, all must pass

A structured pre-deployment assessment. A single failing domain is a hold, not a partial pass.

Verified task executionPolicy & procedure adherenceRisk-calibrated tool & accessRobustness & adversarial resilienceIndependent oversight validationContinuous-assurance infrastructure

Standard classification

Annual audit

Quarterly performance monitoring reviews. Full seven-discipline audit producing a signed assurance report.

Medium-risk classification

Semi-annual audit

Monthly monitoring reviews plus real-time anomaly detection.

High-risk classification

Continuous + 6-monthly board review

Full regulatory-alignment verification and adversarial red-team testing.

Independence by design. VERDICT is the fifth dimension of FIRST AI-D — Govern, Map, Measure, Manage, Verify. It is engineered within the Frontier Lab but does not report to it. Every audit report is signed by the lead assessor and independently reviewed, separate from delivery, before it reaches the Accountable Owner. The people who build a system never clear it.

Built for two contexts

Two editions. One standard of evidence.

VERDICT is calibrated for each context — a Government & Public Sector edition and a Private Sector & Commercial edition — sharing the same disciplines, modules and lifecycle.

Aligned across the UK, EU, US, Middle East and Asia-Pacific — NIST AI RMF, EU AI Act, ISO 42001, UK DSIT principles, OECD AI Principles and GDPR.

Private Sector & Commercial Edition

For commercial organisations deploying AI in competitive, regulated, high-stakes environments.

Where FIRST AI-D defines what must be in place for an AI system to be considered governed, VERDICT determines whether it actually is — the evidence base on which every commercial assurance statement rests.

Government & Public Sector Edition

A shared approach to testing, evaluation and assurance of AI across the public sector.

Designed to help departments test whole systems, evaluate AI models, and provide assurance of quality, trustworthiness and proportionate risk management — for rule-based, machine-learning, generative and agentic systems alike.

Both editions download as print-ready files — use your browser’s Print → Save as PDF to produce a PDF copy.

Turn assurance into evidence.

We run a VERDICT assessment against a real agent in your sector — six domains, signed off, independently reviewed.

No obligation. No sales process. Just a clear picture of what is possible.