Service · AI red-teaming

AI red-teaming for chatbots and AI assistants

AI red-teaming means testing an AI system the way real users and attackers will. AltaSec is an EU-based AI red-teaming service: we test LLM chatbots and AI assistants for prompt injection, jailbreaks, system-prompt leakage and data leaks, check every answer against your own rules, and deliver an evidence report that also states what could not be tested.

What is AI red-teaming?

Red-teaming is an old security practice: a team plays the attacker so you find the weaknesses before someone else does. AI red-teaming applies it to systems built on large language models (LLMs). Instead of probing servers and code, it probes behaviour: what the model says and does when someone talks to it with intent.

For a chatbot or AI assistant, that means sending it the messages a curious user, a competitor or an attacker would send — jailbreaks, prompt injection, attempts to pull out data — and checking every reply against what the assistant is allowed to do.

Why chatbots need it

Classic pentests and scanners look at infrastructure and code. They do not show how a language model behaves in conversation, and that is where most chatbot incidents start:

  • It leaks what it knows. System prompts, API keys, personal data and internal documents can come out in an answer when someone asks the right way.
  • It gets talked into things. Prompt injection and jailbreaks push it outside the job you gave it.
  • It breaks your rules. It gives answers your own policy forbids — and you are the one who has to explain them under GDPR and the EU AI Act.

What AltaSec tests today

Every engagement today covers two checks.

1. Red-team. We run automated adversarial tests against your text chatbot or assistant through its API: jailbreaks, prompt injection and attempts to extract data. Every reply is checked for:

  • leaked secrets and API keys;
  • personal data in Dutch, German and English;
  • planted canary data turning up where it must not;
  • risky output: script injection, javascript: links, markdown-image exfiltration and phishing links.

2. Your rules, as tests. You tell us in writing what your AI is for and what it must never share or do. We turn that policy into targeted tests and report on every rule. Rule packs map the results to GDPR, the EU AI Act and the OWASP Top 10 for LLM Applications — as evidence for your own assessment, not a compliance verdict.

Multi-turn attacks, testing through the live chat widget on your website, and scanning the data you feed the AI are on our roadmap; voice agents and tool-using agents come later. Only what is listed above is part of an engagement today.

How an engagement works

  1. Scope and authorize. You sign the rules of engagement and a written authorization. Together we agree what is tested and the request limits.
  2. Test. We run automated attacks against your chatbot's endpoint, within the agreed limits. No code access is needed.
  3. Judge and review. An LLM judge, kept isolated from the attack content, scores each reply against your rules. A person reviews every finding.
  4. Report. You get an HTML report and a machine-readable JSON export.

The details of each step are in our methodology.

What you get

  • Evidence transcripts — the exact conversation behind each finding, so your team can reproduce it.
  • A fix per finding — a concrete recommendation for your developers.
  • Coverage statement — what was tested against which rule, and what was not assessed, with the reason.
  • Explicit limitations — results are point-in-time and not exhaustive, and the report says so.
  • JSON export — the same findings in machine-readable form, for your tracker.

What AI red-teaming cannot tell you

A red-team shows how your assistant behaved against the tests that were run, at the time they were run. It cannot prove that no other attack will work, and models, prompts and data change. That is why our report separates a verified pass (backed by a positive check, such as a planted canary that was not disclosed) from no violation found (judged, not verified) and from not assessed (not tested, with the reason). Not tested is never counted as a pass. We issue reports with evidence — never certificates, seals or badges.

Frequently asked questions

Do you need access to our code or model?

No. We need your chatbot's endpoint and your written authorization. We test it from the outside, the way a user would reach it.

Which chatbots can you test?

Text chatbots and AI assistants that we can reach through an API. Testing through a live chat widget on your website is next on our roadmap.

How is AI red-teaming different from a jailbreak test?

A jailbreak test is one part of it. AI red-teaming also covers prompt injection, data extraction and every rule in your own policy, and reports what was not tested. See prompt injection vs jailbreaking.

Where is our data processed?

In the EU only, by processors named in our data processing agreement with you. Your data is pseudonymised, and evidence is deleted when the engagement closes, per the retention we agree with you.