Direct answer: The OWASP Top 10 for LLM Applications 2026 was released on 4 August 2026. Prompt Injection stays first and Sensitive Information Disclosure second. Excessive Agency rises to third, reflecting how many applications now let a model call tools and take actions. Unbounded Consumption and Misinformation move up, Improper Output Handling moves to tenth, and a new entry, Hidden Context Exposure, appears at eighth where System Prompt Leakage sat in 2025. Each risk can be tested, and most of the tests are specific to how your application is built rather than to the model you use.
The 2026 list
- LLM01 Prompt Injection
- LLM02 Sensitive Information Disclosure
- LLM03 Excessive Agency
- LLM04 Supply Chain
- LLM05 Data and Model Poisoning
- LLM06 Unbounded Consumption
- LLM07 Misinformation
- LLM08 Hidden Context Exposure
- LLM09 Vector and Embedding Weaknesses
- LLM10 Improper Output Handling
What changed from 2025
Compared with the 2025 list, the movement tells a clear story about where LLM applications have gone in a year:
- Excessive Agency moved from sixth to third. Agents with tool access, write permissions and the ability to chain actions are now ordinary. The damage from a manipulated model is set by what it is allowed to do.
- Unbounded Consumption moved from tenth to sixth and Misinformation from ninth to seventh. Cost and reliability failures are now treated as security problems, not just quality ones.
- Supply Chain moved from third to fourth, Data and Model Poisoning from fourth to fifth, and Vector and Embedding Weaknesses from eighth to ninth. Still present, relatively lower.
- Improper Output Handling moved from fifth to tenth. It has not become safe. It is a well-understood class with well-understood fixes, and it still causes real injection bugs when output is trusted.
- Hidden Context Exposure is new at eighth, in the position System Prompt Leakage held in 2025. Read literally, it is the wider version of that risk: anything placed in the model's context that the user was never meant to see. That includes system prompts, but also retrieved documents, tool results, other users' data and memory. Read OWASP's own entry for its precise scope.
The full text of each entry is on the OWASP GenAI Security Project site. What follows is how we test for each one.
Free weekly email
Get The Compliance Brief every Tuesday
One email a week from Jacob Masse: the security and compliance stories that changed something that week, and what each one means if you sell software to enterprise buyers. Five stories, a take on each, five minutes to read.
Free. Unsubscribe in one click, and replies reach Jacob directly. Read the latest issue or browse the archive.
How to test for each risk
LLM01 Prompt Injection
Test both direct injection, where the user types the attack, and indirect injection, where the attack arrives in content the model reads: a web page, an uploaded document, an email, a support ticket, a tool result. Indirect is usually the more serious, because the user who triggers it may be the victim. Build test documents that carry instructions and check whether the model follows them, especially where it can act. Measure the result by what happened, not by what the model said.
LLM02 Sensitive Information Disclosure
Map what sensitive data can reach the model: training or fine-tuning data, retrieved documents, tool outputs, conversation history. Then try to get it out, as an unauthenticated user, as a low-privilege user, and as one tenant asking about another. Test that redaction happens before data enters the context rather than after the model writes its answer.
LLM03 Excessive Agency
List every tool the model can call, the permissions behind each, and whether a human confirms the consequential ones. Then try to make the model call tools it should not, with arguments it should not use, in sequences nobody intended. The fixes are architectural: least-privilege credentials per tool, actions scoped to the current user's own rights, confirmation for anything irreversible, and rate limits on actions. If the model's service account can do more than the user can, that gap is the vulnerability.
LLM04 Supply Chain
Inventory the models, model hosts, fine-tuning datasets, plugins, agent frameworks and libraries in the application. Check where each comes from, how it is pinned and verified, and what happens when a provider changes a model behind the same name. Review the licences and terms of models and datasets, which can restrict use in ways the team has not noticed.
LLM05 Data and Model Poisoning
Ask who can influence what the model learns from or retrieves: fine-tuning data, feedback loops, documents added to a knowledge base, shared memory. Then test whether a low-trust user can insert content that later changes answers for other users. The controls are provenance, review before content enters a trusted store, and separation between user-contributed and curated sources.
LLM06 Unbounded Consumption
Test what an attacker or a runaway agent can make you pay for. Very long inputs, requests that force long outputs, recursive agent loops, expensive tool calls, and automated extraction of model behaviour through high request volumes. Check per-user and per-tenant limits, maximum context and output sizes, loop and step limits on agents, and spend alerts that someone actually receives.
LLM07 Misinformation
Testing here is about the application, not about the model's general knowledge. Where the product gives answers users act on, check whether answers are grounded in your sources, whether citations point to real content that supports the claim, and whether the interface makes uncertainty visible. Build a set of questions with known answers from your own domain and run it on every model or prompt change.
LLM08 Hidden Context Exposure
Everything in the context window should be treated as potentially visible to the user. Test whether system prompts, retrieved documents the user lacks permission to see, tool results, internal identifiers, other sessions or other users' memory can be extracted, directly or through indirect injection. The fix is upstream: do not put secrets or another user's data in the context in the first place, and enforce permissions at retrieval rather than asking the model to keep a secret.
LLM09 Vector and Embedding Weaknesses
For retrieval-augmented applications, test access control in the vector store. Can one tenant's query retrieve another tenant's chunks? Are permissions applied at query time, or only when documents were ingested? Can a user insert content designed to be retrieved for other users' questions? Check, too, whether embeddings of sensitive data are treated with the same protection as the data itself.
LLM10 Improper Output Handling
Treat model output as untrusted input to whatever consumes it. Test for script injection where output is rendered as HTML, injection into SQL or shell where output feeds a query or command, server-side request forgery where output becomes a URL, and unsafe deserialization where it becomes structured data. The fixes are the ordinary ones: encoding, parameterization, allow-lists and validation. They just have to be applied to the model as if it were a user.
Assessment or red teaming?
An AI security assessment tests your application systematically against the list above and your architecture: what can reach the model, what the model can reach, and where the boundaries fail. It suits a team shipping a feature that needs to be checked before customers rely on it.
LLM red teaming is objective-based and adversarial: a team tries to achieve a specific harmful outcome, such as exfiltrating another customer's data through an agent, by any route. It suits applications already in production with significant agency or sensitive data, and it is most useful after the systematic issues are fixed.
Our AI security assessment is from $4,000, scoped to what you built. LLM red teaming is from $9,000 for an objective-based campaign. Hands-on testing for both is delivered with our testing partners, and the findings land in one standard report with severity, reproduction steps and fixes. For a first look at your own risk, the AI security risk assessment is free.
Frequently asked questions
Does the 2026 list replace the 2025 one?
Yes, for anyone mapping tests or controls to the OWASP LLM Top 10. Update the references in your test plans and security questionnaires, and note that the numbering changed for almost every entry.
Is prompt injection fixable?
Not fully at the model level today. It is containable at the application level: limit what the model can reach and do, treat its output as untrusted, and require confirmation for consequential actions. Testing measures how well that containment holds.
We use a major model provider. Is the provider responsible for these risks?
Partly, for the model itself. Most of the list is about how your application uses the model: what goes into the context, what tools it can call, how output is handled. Those are yours.
How often should we test?
Before launch, and again when you add tools, data sources or a new model. A change in what the model can reach changes the risk more than any change in the model.
Need your AI feature tested against the 2026 list? We assess what can reach the model, what the model can reach, and where the boundaries fail.
AI security assessmentOr LLM red teaming