Home / Services / AI system audit

AI system audit

Your assistant answers. The question is whether it answers well, and you don't find that out by poking at it for an afternoon: you measure it. Case suite, adversarial testing and a failure report with a prioritised fix plan.

The problem

What fails is rarely the model

Almost no AI system fails because the model is bad. It fails because a prompt changed and nobody re-checked what already worked, because retrieval brings back the wrong document, because a user's negation gets lost along the way, or because the system answers confidently about something it does not know.

These failures share an uncomfortable trait: nothing breaks. No error in the logs, no alert on the dashboard. The answer arrives, looks reasonable, and is wrong. They only surface when a customer hits one, and by then weeks have passed.

The most common case: the system displays a confidence percentage nobody ever validated. In a project of my own, that "92%" turned out to be a saturated formula measuring absolutely nothing. We found it by auditing, not by using it.
What's included

What actually gets done

  • Inventory of expected behaviour. What the system should do, what it must never do, and where the decisions with consequences live.
  • Representative case suite. Real and synthetic cases with the correct answer annotated, runnable in one go. Yours to keep, protecting every future change.
  • Adversarial testing. Cases written to break it: false positives and negatives, hallucinations, bias, out-of-scope answers, ignored negations, missed red flags.
  • System security. Prompt injection, context leaking between conversations, and what information ends up in the logs without anyone deciding it should.
  • Reasoning traceability. If the system operates in a critical domain, being able to reconstruct why it answered what it answered stops being a luxury.
  • Report and plan. Failures classified by severity and frequency with reproducible examples, and fixes ordered by impact and effort.
How it works

Two weeks, three deliverables

The express audit is designed to give a useful answer fast, without blocking your team.

  1. Days 1 to 3. Context with your team, access to a test environment and construction of the case suite.
  2. Days 4 to 8. Execution, adversarial testing and analysis of the failures found, with reproducible examples.
  3. Days 9 to 10. Report, prioritised fix plan and a conclusions session with whoever decides.

If you then want me to implement the fixes, that is quoted separately. The audit commits you to nothing: the report is yours and your team can execute it without me.

Experience

Why me

Because I haven't read about this, I have lived it. I am co-founder and CTO of a clinical decision support platform in production, a domain where a wrong answer has real consequences.

  • A regression suite of nearly 300 clinical cases that runs on every change to the reasoning.
  • Periodic adversarial audits of my own system, hunting false positives, false negatives and misfired red flags.
  • Experience in regulated domains: healthcare under GDPR, legal with confidential documentation, and emotional wellbeing with guardrails.

I am also independent: if your system was built by another vendor, I have no interest in softening the diagnosis.

Questions

Frequently asked

How long does it take?

The express version is two weeks: one to build the suite and run it, another for analysis, report and conclusions. Large systems extend by module, but the first useful report always lands in those two weeks.

Does it work with the OpenAI, Anthropic or Google API?

Yes. The complete system is audited (prompts, retrieval, rules, guardrails and final answer), not the model in isolation. Who serves the model is irrelevant: what gets measured is what your user sees.

What do you need from us?

Access to a test environment, documentation of what the system should do, and half a day from one or two people. Anonymised real conversation history raises precision considerably.

What do we keep?

The executable case suite, the report of classified failures with reproducible examples, and the fix plan prioritised by impact and effort.

What if the conclusion is that it's fine?

Then you keep a test suite that protects the system from now on, and the confidence of having measured it. That is a perfectly possible outcome and it gets said plainly.

Want to know if your AI answers well?

Tell me what system you have and what you use it for. I'll tell you what I would audit, how long it takes and what to expect from the report.