CriterIAQA · New service
Would you put your customers in a car without an inspection?
Your AI agents talk to real customers, quote prices, make promises, and make decisions. CriterIAQA is roadworthiness testing for your AI agents: we inspect them before they reach production and give you a clear verdict, ready or not ready, backed by evidence, not promises.
CriterIAQA Report
AGT-0192 · Booking Assistant
- Response testing
- Escalation rules
- Traceability (evidence)
- Authorized data
Ready for production
CriterIAQA verdict · illustrative example
What it is
The quality control almost no one runs on their agents today
A car doesn't go on the road without its annual inspection. An AI agent, today, does reach production without anyone inspecting it. CriterIAQA closes that gap: we test your agent before it talks to a real customer and hand you a verdict, not an opinion.
Without QA
The agent goes straight from development to production. No one knows how it responds under pressure, whether it hallucinates, whether it follows business rules, or whether a customer can manipulate it. The first real test is an upset customer or a complaint.
With CriterIAQA
The agent goes through a 4-point checklist, with documented evidence for each test. You come out with a clear verdict: green, yellow, or red. If needed, you also get a concrete correction plan before launch.
How it works
The CriterIAQA 4-point checklist
The checklist is grounded in recognized practices for evaluating agents: task completion, tool use, traceability, and instruction-following. We apply it with human judgment and evidence for your business.
Response testing
Does the agent finish the task, or stall halfway through? We test the full flow, not just the happy path.
Escalation rules
Does it use the right tools, or fail invisibly? We validate when it should resolve on its own and when it should escalate to a human.
Traceability (evidence)
Can every response be traced back to a real source, without hallucinating? We document the origin of each claim the agent makes.
Authorized data
Does it respect the rules and data it's actually allowed to use? We verify it doesn't expose or invent what it shouldn't.
What you get
A clear verdict: ready or not ready for production
At the end of the inspection you get a readiness signal, not a 40-page report you have to interpret on your own.
Green
Ready for production. The agent passed the 4-point checklist with documented evidence.
Yellow
Fix before launch. We identify the exact points to adjust and the plan to reach green.
Red
Not fit for real customers. The agent should not go to production in its current state.
CriterIAQA Report: every agent evaluated gets documented evidence, ready to show leadership, legal, or your own end client.
Why it matters
When no one ran QA before production
These cases have been publicly reported. None involved cutting-edge technology: a basic checklist before going to production is directly related to what went wrong.
Air Canada
Feb. 2024The chatbot invented a refund policy that didn't exist. A tribunal ordered the airline to honor it.
Related control: TraceabilityChevrolet (dealership)
Dec. 2023A chatbot agreed to sell a car for US$1 after a prompt-injection attack.
Related control: Response testingDPD
Jan. 2024A user manipulated the chatbot into insulting its own company. It went viral.
Related control: Response testingLawyer vs. Avianca (U.S.)
Jun. 2023A lawyer submitted court filings citing legal cases invented by a generative AI. They never existed.
Related control: TraceabilityThe risk of skipping it
Why do QA, and what happens if you don't
The same cases, seen through the risk that could have been avoided.
Binding legal liability
A tribunal ordered Air Canada to honor a policy its own chatbot invented.
CriterIAQA aims to catch failures before they reach the customer, not after the crisis, and gives you a traceable, defensible basis before leadership, legal, and regulators.
Viral reputational damage
DPD's chatbot insulted its own company on social media. It went viral.
Turns trust into a verifiable process, not a personal bet by whoever approves the launch.
Direct financial loss
A dealership chatbot "sold" a car for US$1 after a prompt attack.
A basis for moving from pilot to production with clear exit criteria, not the hope that no one tests it.
Legal exposure from hallucinations
A lawyer submitted court filings citing legal cases invented by a generative AI.
We document evidence of control over the agent's responses before they reach a customer or a court.
Frameworks like the NIST AI RMF, ISO/IEC 42001 and, where applicable, the EU AI Act help structure controls, traceability, and oversight. This information is general and does not replace legal advice.
Schedule your inspection
Before your agent talks to its next customer
An initial session to review your agent against the CriterIAQA checklist and define the scope of the full assessment.
