AI agent testing for chat and voice
Your AI agent talks to customers every day: by chat, by phone, in their own language. AltaSec tests it the way your customers will, and the way an attacker will, and tells you whether it is ready to go live. You get one go / no-go verdict for every task, language and channel, with the evidence behind it.
The two questions that decide whether an agent is ready
Does it answer right? That is quality. The agent must give the right answer from your rules and your knowledge base, never invent a policy or a price, use the right tone and formality for each language, get numbers, dates and names right, and reply in the customer's own language.
Can it be tricked? That is security. Customers and attackers will try to talk it into breaking your rules, to pull out personal data or its own instructions, or to make it promise what you never allowed. The same attacks arrive typed or spoken.
Most teams check these separately, once, in one language. We test both together, in every language you serve, and give one verdict.
Why every language matters
An agent that behaves well in English can behave differently in Romanian or Portuguese: it may switch to the wrong formality, mishear a number, drift into another variety of the language, or give in to an attack that it would have refused in English. That is why every language we test is checked by native speakers. Today we test English, Romanian and Portuguese; more European languages are next.
Chat and voice
A chat agent reads and writes. A voice agent also has to hear and speak, and each of those can fail on its own. We test the three parts of a voice agent one by one and as the whole call, so a failure points to the part you need to fix. See voice agent testing.
How an engagement works
- You set the rules. The tasks your agent handles, the languages and channels, and the failures you can't afford. You sign a written authorization and we agree the scope and the limits.
- We test. Simulated customers and attackers talk to your agent, typed and spoken. Fixed rules check the facts, AI judges review every conversation, and native speakers check what machines can't.
- You get the Passport. Our go / no-go verdict for every task, language and channel, with the evidence, a fix per finding, and a list of what we could not test.
- We re-test every change. Every failure we find becomes a permanent test that runs again after each update and at every release.
The details are in our methodology.
What you get: the Passport
| Verdict | What it means |
|---|---|
| Go | Ready for this task, language and channel: every check met |
| Review | It works, with a problem to fix first |
| No-go | A failure you can't afford. Some failures are never averaged away: money moved without permission, another customer's data, a false "done", a missed emergency. One case is a no-go |
| Not tested | Not tested, with the reason stated. Never counted as a pass |
Every finding comes with its evidence (the transcript or the recording, and the check that caught it) and a concrete fix. The Passport is our verdict from our tests, not a certificate: no test can prove an AI is right 100% of the time, and the Passport says what was not tested.
When to test
Before launch, after every change to the model, the prompt or the knowledge base, and at every release. Each failure we find becomes a test that runs again next time, so a fixed problem stays fixed. Monitoring live conversations after launch is on our roadmap.
Frequently asked questions
Do you need access to our code or model?
No. We need access to your agent (an endpoint, a test chat or a test phone number) and your written authorization. We test it from the outside, the way a customer reaches it.
Which languages do you test?
English, Romanian and Portuguese today, each checked by native speakers. More European languages are next: recordings with native speakers are under way.
Is the Passport a certificate?
No. It is our verdict with evidence, never a certificate, seal or badge. Results map to GDPR, the EU AI Act and the OWASP Top 10 for LLM Applications as evidence for your own assessment, not as a compliance verdict.
Where is our data processed?
In the EU only, by processors named in our data processing agreement with you. Your data is pseudonymised, and evidence is deleted when the engagement closes, per the retention we agree with you.
