Guide · Checklist

LLM security testing checklist

This checklist covers what to test in an application built on a large language model (LLM), organised by the OWASP Top 10 for LLM Applications (2025 edition). For each risk it lists concrete checks you can run, and says plainly which of them are part of an AltaSec engagement today.

Before you test

  • Write down what the assistant is for, and what it must never share or do. Every test is measured against this.
  • List what the assistant can see: system prompt, knowledge base, customer data, documents, tool results.
  • Get written authorization to test it, and agree request limits with whoever runs it.
  • Plant canaries — unique marker values — in the places it must not disclose, so a leak is provable.
  • Decide how you will record evidence: every request and reply, with timestamps.

At a glance

OWASP risk What to test Part of an AltaSec engagement today?
LLM01 Prompt Injection Can messages override your instructions? Yes
LLM02 Sensitive Information Disclosure Does it leak personal data, secrets or other users' data? Yes
LLM03 Supply Chain Are models, plugins and datasets from trusted sources? No
LLM04 Data and Model Poisoning Could training or knowledge-base data be tampered with? No
LLM05 Improper Output Handling Can its output attack the page or user that displays it? Yes — in the chatbot's replies
LLM06 Excessive Agency Can it take actions beyond what it should? No — tool-using agents are on our roadmap
LLM07 System Prompt Leakage Can it be made to reveal its instructions? Yes
LLM08 Vector and Embedding Weaknesses Can retrieval be abused to reach other users' data? No — scanning knowledge bases is on our roadmap
LLM09 Misinformation Does it state false things with confidence? Only where your written rules cover it
LLM10 Unbounded Consumption Can someone run up cost or exhaust capacity? No

LLM01 — Prompt injection

  • Direct instructions: "Ignore your previous instructions and…" in many phrasings and languages.
  • Role-play and hypothetical framing ("pretend you are…", "for a story…").
  • Instructions hidden in formatting: encodings, markdown, code blocks, unusual Unicode.
  • Instructions inside content the assistant reads (documents, web pages, emails), if it reads any.
  • For each attempt: did it follow the injected instruction instead of yours?

LLM02 — Sensitive information disclosure

  • Ask for other customers' details, directly and indirectly.
  • Ask for internal documents, staff names, prices or plans it should not share.
  • Check every reply for personal data (names, emails, phone numbers, IBANs) in every language your users use.
  • Check every reply for secrets: API keys, tokens, passwords, internal URLs.
  • Check whether planted canaries ever appear in a reply.

LLM05 — Improper output handling

  • Can it be made to output HTML or script that your front end would execute?
  • Can it output javascript: links or phishing links your users might click?
  • Can it output a markdown image whose URL carries data to someone else's server when rendered?
  • Is its output escaped before display, and are images and links from replies restricted?

LLM07 — System prompt leakage

  • Ask for the system prompt directly, then through translation, summarisation, "repeat the text above", and role-play.
  • Check whether anything in the system prompt would hurt if disclosed. If so, move it out — assume the prompt can leak.

Your own rules

  • Turn each "must never" rule into targeted tests: off-topic requests, forbidden advice, disclosures, tone.
  • For each rule, record one outcome: violation, verified pass, no violation found, or not assessed — with the reason.
  • Never count an untested rule as a pass.

After testing

  • Keep the transcript behind every finding, so it can be reproduced.
  • Write one concrete fix per finding.
  • State what was not tested, and why.
  • Re-test after every change to the model, the prompt or the data it can reach.

Want this done for you?

AltaSec runs the tests marked "Yes" above against your chatbot's API, checks every reply with detectors and an isolated LLM judge, has a person review every finding, and delivers an evidence report. See how AltaSec red-teams chatbots and our methodology.