Assessments
Find out if your AI agent is ready for production
A fixed-scope assessment of an AI agent or AI application you are building or already running. We test how it behaves, where it fails and whether it stays within its permissions, then give you a findings report and a prioritised remediation plan.
What we assess
Reliability
Does it complete the intended task across realistic scenarios?
Output quality
Accuracy, factuality and consistency where they matter.
Tool use
Correct tool selection, parameters and sequencing.
Boundaries
Can it be pushed outside its intended authority?
Security
Prompt injection, data leakage and unsafe actions.
Human oversight
Do escalation and approval paths work?
Traceability
Can you reconstruct what it did and why?
Failure handling
What happens when inputs, tools or models misbehave?
What you get
- Findings ranked by severity
- Evidence for each finding
- Mapping to recognised frameworks, where useful
- A remediation plan your team can act on
Optional: the evaluation suite we built, so you can re-run it.
How it works
Step 1: Scoping call
Agree the system, its purpose, environments and access.
Step 2: Set-up
Access to a test environment and documentation.
Step 3: Evaluation
Scenario, adversarial and boundary testing.
Step 4: Readout
Findings walkthrough with your team and the remediation plan.
Our assessments and reports support your risk and compliance work. They are not a certification, audit opinion or legal advice.
Frequently asked questions
What does the Agent Assurance Assessment cover?
Eight areas: reliability, output quality, tool use, boundaries, security, human oversight, traceability and failure handling.
What do we receive?
An Agent Assurance Report with findings ranked by severity, evidence for each finding, mapping to recognised frameworks where useful, and a remediation plan your team can act on. Optionally, the evaluation suite we built, so you can re-run it.
How does the assessment work?
Four steps: a scoping call to agree the system and access, set-up of a test environment, scenario, adversarial and boundary testing, and a readout of the findings and remediation plan with your team.
Can you assess an agent that is already in production?
Yes. The assessment covers AI agents and AI applications you are building or already running.
Request an Agent Assurance Assessment
Tell us about the system. We’ll arrange a scoping call.