Methodology

How AltaSec tests and reports

AltaSec's methodology is built around one principle: a result has to be backed by evidence you can check. We test your chatbot through its API, check every reply with detectors and an isolated LLM judge, have a person review every finding, and report exactly what was and was not tested. Not tested is never counted as a pass.

1. Scope and authorization

Nothing is tested without your written authorization. Before any test runs:

  • you sign the rules of engagement, and we agree what is tested and the request limits (how many requests we send, and how fast);
  • you give us your chatbot's endpoint — no access to your code or model is needed;
  • you describe in writing what your AI is for and what it must never share or do. This policy is what we test against.

2. Test design

We build the test plan from two sources:

  • Your policy. Each rule becomes a target with its own tests.
  • Attack techniques. Jailbreaks, prompt injection and data-extraction attempts, aimed at your rules.

Where it fits the engagement, we agree canaries with you: unique marker values placed where your assistant can see them. A canary that appears in a reply proves disclosure; a canary that never appears is what lets us verify a pass instead of just judging it.

3. Test execution

Tests are sent automatically to your chatbot's endpoint, within the agreed limits. Every request and reply is recorded as evidence. If the target stops answering or errors pile up, the run halts rather than filling the report with noise — and the affected rules are reported as not assessed, with the reason.

4. Detection and judging

Every reply goes through two kinds of checks:

  • Detectors look for concrete signals: secrets and API keys, personal data in Dutch, German and English, planted canaries, and risky output such as script injection, javascript: links, markdown-image exfiltration and phishing links.
  • An LLM judge scores each reply against your rules. The judge is kept isolated from the attack content, so an injection aimed at your chatbot cannot steer the verdict.

5. Human review

A person reviews every finding before it reaches the report. Automated checks find candidates; a human confirms them.

6. The four outcomes

Every rule ends in exactly one outcome:

Outcome What it means What backs it
Violation A reply broke the rule The transcript, the detector or judge result, and human review
Verified pass No violation, and a positive check proves it For example: a planted canary was not disclosed
No violation found No violation was seen, but nothing could verify it The judge's assessment only — reported as judged, not verified
Not assessed The rule could not be tested The reason, for example the target timed out

A pass is only "verified" when a positive check backs it, and "not assessed" is its own outcome. It is never shown as a pass.

7. The report

You receive an HTML report and a machine-readable JSON export containing:

  • evidence transcripts for every finding, so your team can reproduce it;
  • a fix per finding, as a concrete recommendation for your developers;
  • a coverage statement: what was tested against which rule, and what was not assessed, with the reason;
  • explicit limitations: results are point-in-time and not exhaustive.

Rule packs map results to GDPR, the EU AI Act and the OWASP Top 10 for LLM Applications — as evidence for your own assessment, not a compliance verdict. We issue reports with evidence, never certificates, seals or badges.

8. Data handling

Engagement data is processed in the EU only, by processors named in our data processing agreement with you. Your data is pseudonymised, and evidence is deleted when the engagement closes, per the retention we agree with you.

What comes next

Re-testing with a fixed / regressed / new comparison between runs, multi-turn attacks, and testing through a live website chat widget are in development. See the roadmap.