VERDICT · Testing & assurance · under Frontier Lab
MEAIOW VERDICT
You cannot declare a system works. You assure it does.
VERDICT is MEAIOW’s AI testing and assurance framework. Testing AI is not testing conventional software — AI behaves probabilistically, learns from shifting data and can act autonomously. VERDICT replaces assertion with evidence: seven disciplines, four quality themes, nine testing modules and a continuous lifecycle, applied from before go-live through the whole life of every system.
Validated · Evidenced · Risk-Calibrated · Documented · Independent · Continuous · Trusted
Testing, evaluation, assurance
Three activities. One chain of evidence.
Testing checks the complete AI-enabled system — infrastructure, APIs, integrations, interfaces and data flows. Evaluation assesses the model itself against agreed quality attributes. Assurance is the confidence that follows: testing generates the evidence, evaluation interprets it, assurance proves the system is safe, fair and accountable.
AI types covered
Rule-based / deterministic
Predictable logic. Testing verifies every rule and boundary case is correct and complete.
Machine learning
Trained on data. Testing targets data quality, accuracy, generalisation and bias.
Generative
Open-ended output. Testing emphasises appropriateness, factual accuracy and control of harmful content.
Agentic / autonomous
Probabilistic and adaptive. Testing targets safe autonomy, novel situations, reward hacking and human override.
The seven disciplines
V·E·R·D·I·C·T
Not seven tests passed in sequence — seven disciplines forming one continuous assurance cycle across the life of every system.
V
Verified Performance
Performance is verified against documented, reproducible benchmarks before go-live. A system that has not been verified has not been assured.
E
Evidence-Led Assurance
Every statement rests on primary evidence VERDICT generates — never vendor assertion or self-reported compliance. No evidence, no assurance.
R
Risk-Calibrated Testing
Test rigour is proportionate to each deployment’s risk. Higher risk demands more rigorous, more adversarial, more frequent testing.
D
Documented Audit Trail
Every assessment produces a tamper-evident record regulators, boards and investigators can examine independently.
I
Independent Oversight
The team that builds a system does not assess it. VERDICT operates structurally apart from delivery, through an independent review line.
C
Continuous Assurance
Assurance does not end at go-live. It continues through monitoring, scheduled testing and audit cycles set by risk classification.
T
Trust Made Measurable
Trust is not a feeling. It is a product — quantified, evidence-backed, independently verified and sustained over the operational life of every system.
Core quality attributes
What we test for, grouped into four themes.
These guide both system-level testing and model-level evaluation across every deployment.
Safety & Ethics
Avoiding harm, respecting rights, non-discrimination — ethics treated as testable risk.
Openness & Trust
Explainability, transparency and traceability of decisions to those they affect.
Performance & Resilience
Accuracy, robustness under stress, graceful and safe failure.
User & Context Fit
Fitness for the people and the real-world conditions the system serves.
Continuous defensive assurance
Assurance is a lifecycle, not a one-time event.
VERDICT defines testing activities and deliverables at every phase — quality built in from the start, not retrofitted before launch.
01
Planning & Design
02
Data Collection & Prep
03
Model Development
04
Validation & Verification
05
Operational Readiness
06
Deployment
07
Monitoring & Continuous Assurance
The nine testing modules
The lifecycle says when. The modules say what.
Nine modular building blocks of AI testing, emphasised or de-emphasised by AI type and risk level. Full metrics and thresholds are in the framework document.
Data & Input Validation
Inputs are valid, representative and compliant — schema checks, statistical profiling, PII redaction and lineage.
Model Functionality
The model works as intended across typical inputs, edge cases and failure modes — re-tested after every change.
Bias & Fairness
Outputs tested by group, not just in aggregate — counterfactual, demographic parity and equal-opportunity analysis.
Explainability & Transparency
Decisions can be understood, traced and explained to overseers and to the people affected.
Robustness & Adversarial
Behaviour under stress and deliberate attack — perturbations, jailbreaks, prompt injection, fail-safe verification.
Performance & Efficiency
Speed, scale and resource use under realistic and peak load, including latency and throughput targets.
Integration & System
The AI works as part of the wider system — end-to-end workflows, dependencies and graceful degradation.
User Acceptance & Ethical Review
It works for the people it serves and clears a formal ethical sign-off before deployment.
Continuous Monitoring
After go-live: dashboards, drift alerts, incident response and scheduled re-testing as data and risk shift.
The assurance cycle
Governance without independent verification is self-certification.
—Shakil Siddiqui, Founder & CEO of MEAIOW
A structured pre-deployment assessment. A single failing domain is a hold, not a partial pass.
Verified task executionPolicy & procedure adherenceRisk-calibrated tool & accessRobustness & adversarial resilienceIndependent oversight validationContinuous-assurance infrastructure
Standard classification
Annual audit
Quarterly performance monitoring reviews. Full seven-discipline audit producing a signed assurance report.
Medium-risk classification
Semi-annual audit
Monthly monitoring reviews plus real-time anomaly detection.
High-risk classification
Continuous + 6-monthly board review
Full regulatory-alignment verification and adversarial red-team testing.
Independence by design. VERDICT is the fifth dimension of FIRST AI-D — Govern, Map, Measure, Manage, Verify. It is engineered within the Frontier Lab but does not report to it. Every audit report is signed by the lead assessor and independently reviewed, separate from delivery, before it reaches the Accountable Owner. The people who build a system never clear it.
Built for two contexts
Two editions. One standard of evidence.
VERDICT is calibrated for each context — a Government & Public Sector edition and a Private Sector & Commercial edition — sharing the same disciplines, modules and lifecycle.
Aligned across the UK, EU, US, Middle East and Asia-Pacific — NIST AI RMF, EU AI Act, ISO 42001, UK DSIT principles, OECD AI Principles and GDPR.
For commercial organisations deploying AI in competitive, regulated, high-stakes environments.
Where FIRST AI-D defines what must be in place for an AI system to be considered governed, VERDICT determines whether it actually is — the evidence base on which every commercial assurance statement rests.
A shared approach to testing, evaluation and assurance of AI across the public sector.
Designed to help departments test whole systems, evaluate AI models, and provide assurance of quality, trustworthiness and proportionate risk management — for rule-based, machine-learning, generative and agentic systems alike.
Both editions download as print-ready files — use your browser’s Print → Save as PDF to produce a PDF copy.
Turn assurance into evidence.
We run a VERDICT assessment against a real agent in your sector — six domains, signed off, independently reviewed.
No obligation. No sales process. Just a clear picture of what is possible.