Compliance

Phase 1, Phase 2, then keep it running

A fixed-price gap analysis, remediation through to your audit, and upkeep after it. All in a workspace you keep.

All compliance →
Security

Testing, review and leadership

Led by a published security researcher with five CVEs. One standard report, letters for your buyers, and retests of your fixes.

All security →
Who we help

Prove you are secure

To the people you sell to, raise from or answer to.

All industries →
Resources

Learn the space

Original research, free tools, and plain-language guides on security and compliance, from a published security researcher.

Read the blog →
Security

OWASP Top 10 for LLM Applications 2026: What Changed and How to Test Each Risk

Direct answer: The OWASP Top 10 for LLM Applications 2026 was released on 4 August 2026. Prompt Injection stays first and Sensitive Information Disclosure second. Excessive Agency rises to third, reflecting how many applications now let a model call tools and take actions. Unbounded Consumption and Misinformation move up, Improper Output Handling moves to tenth, and a new entry, Hidden Context Exposure, appears at eighth where System Prompt Leakage sat in 2025. Each risk can be tested, and most of the tests are specific to how your application is built rather than to the model you use.

The 2026 list

  1. LLM01 Prompt Injection
  2. LLM02 Sensitive Information Disclosure
  3. LLM03 Excessive Agency
  4. LLM04 Supply Chain
  5. LLM05 Data and Model Poisoning
  6. LLM06 Unbounded Consumption
  7. LLM07 Misinformation
  8. LLM08 Hidden Context Exposure
  9. LLM09 Vector and Embedding Weaknesses
  10. LLM10 Improper Output Handling

What changed from 2025

Compared with the 2025 list, the movement tells a clear story about where LLM applications have gone in a year:

  • Excessive Agency moved from sixth to third. Agents with tool access, write permissions and the ability to chain actions are now ordinary. The damage from a manipulated model is set by what it is allowed to do.
  • Unbounded Consumption moved from tenth to sixth and Misinformation from ninth to seventh. Cost and reliability failures are now treated as security problems, not just quality ones.
  • Supply Chain moved from third to fourth, Data and Model Poisoning from fourth to fifth, and Vector and Embedding Weaknesses from eighth to ninth. Still present, relatively lower.
  • Improper Output Handling moved from fifth to tenth. It has not become safe. It is a well-understood class with well-understood fixes, and it still causes real injection bugs when output is trusted.
  • Hidden Context Exposure is new at eighth, in the position System Prompt Leakage held in 2025. Read literally, it is the wider version of that risk: anything placed in the model's context that the user was never meant to see. That includes system prompts, but also retrieved documents, tool results, other users' data and memory. Read OWASP's own entry for its precise scope.

The full text of each entry is on the OWASP GenAI Security Project site. What follows is how we test for each one.

Free weekly email

Get The Compliance Brief every Tuesday

One email a week from Jacob Masse: the security and compliance stories that changed something that week, and what each one means if you sell software to enterprise buyers. Five stories, a take on each, five minutes to read.

Free. Unsubscribe in one click, and replies reach Jacob directly. Read the latest issue or browse the archive.

How to test for each risk

LLM01 Prompt Injection

Test both direct injection, where the user types the attack, and indirect injection, where the attack arrives in content the model reads: a web page, an uploaded document, an email, a support ticket, a tool result. Indirect is usually the more serious, because the user who triggers it may be the victim. Build test documents that carry instructions and check whether the model follows them, especially where it can act. Measure the result by what happened, not by what the model said.

LLM02 Sensitive Information Disclosure

Map what sensitive data can reach the model: training or fine-tuning data, retrieved documents, tool outputs, conversation history. Then try to get it out, as an unauthenticated user, as a low-privilege user, and as one tenant asking about another. Test that redaction happens before data enters the context rather than after the model writes its answer.

LLM03 Excessive Agency

List every tool the model can call, the permissions behind each, and whether a human confirms the consequential ones. Then try to make the model call tools it should not, with arguments it should not use, in sequences nobody intended. The fixes are architectural: least-privilege credentials per tool, actions scoped to the current user's own rights, confirmation for anything irreversible, and rate limits on actions. If the model's service account can do more than the user can, that gap is the vulnerability.

Shipping an AI feature? We assess it against the OWASP LLM Top 10 and your own architecture, with findings your engineers can fix. AI security assessment

LLM04 Supply Chain

Inventory the models, model hosts, fine-tuning datasets, plugins, agent frameworks and libraries in the application. Check where each comes from, how it is pinned and verified, and what happens when a provider changes a model behind the same name. Review the licences and terms of models and datasets, which can restrict use in ways the team has not noticed.

LLM05 Data and Model Poisoning

Ask who can influence what the model learns from or retrieves: fine-tuning data, feedback loops, documents added to a knowledge base, shared memory. Then test whether a low-trust user can insert content that later changes answers for other users. The controls are provenance, review before content enters a trusted store, and separation between user-contributed and curated sources.

LLM06 Unbounded Consumption

Test what an attacker or a runaway agent can make you pay for. Very long inputs, requests that force long outputs, recursive agent loops, expensive tool calls, and automated extraction of model behaviour through high request volumes. Check per-user and per-tenant limits, maximum context and output sizes, loop and step limits on agents, and spend alerts that someone actually receives.

LLM07 Misinformation

Testing here is about the application, not about the model's general knowledge. Where the product gives answers users act on, check whether answers are grounded in your sources, whether citations point to real content that supports the claim, and whether the interface makes uncertainty visible. Build a set of questions with known answers from your own domain and run it on every model or prompt change.

LLM08 Hidden Context Exposure

Everything in the context window should be treated as potentially visible to the user. Test whether system prompts, retrieved documents the user lacks permission to see, tool results, internal identifiers, other sessions or other users' memory can be extracted, directly or through indirect injection. The fix is upstream: do not put secrets or another user's data in the context in the first place, and enforce permissions at retrieval rather than asking the model to keep a secret.

LLM09 Vector and Embedding Weaknesses

For retrieval-augmented applications, test access control in the vector store. Can one tenant's query retrieve another tenant's chunks? Are permissions applied at query time, or only when documents were ingested? Can a user insert content designed to be retrieved for other users' questions? Check, too, whether embeddings of sensitive data are treated with the same protection as the data itself.

LLM10 Improper Output Handling

Treat model output as untrusted input to whatever consumes it. Test for script injection where output is rendered as HTML, injection into SQL or shell where output feeds a query or command, server-side request forgery where output becomes a URL, and unsafe deserialization where it becomes structured data. The fixes are the ordinary ones: encoding, parameterization, allow-lists and validation. They just have to be applied to the model as if it were a user.

Assessment or red teaming?

An AI security assessment tests your application systematically against the list above and your architecture: what can reach the model, what the model can reach, and where the boundaries fail. It suits a team shipping a feature that needs to be checked before customers rely on it.

LLM red teaming is objective-based and adversarial: a team tries to achieve a specific harmful outcome, such as exfiltrating another customer's data through an agent, by any route. It suits applications already in production with significant agency or sensitive data, and it is most useful after the systematic issues are fixed.

Our AI security assessment is from $4,000, scoped to what you built. LLM red teaming is from $9,000 for an objective-based campaign. Hands-on testing for both is delivered with our testing partners, and the findings land in one standard report with severity, reproduction steps and fixes. For a first look at your own risk, the AI security risk assessment is free.

Frequently asked questions

Does the 2026 list replace the 2025 one?

Yes, for anyone mapping tests or controls to the OWASP LLM Top 10. Update the references in your test plans and security questionnaires, and note that the numbering changed for almost every entry.

Is prompt injection fixable?

Not fully at the model level today. It is containable at the application level: limit what the model can reach and do, treat its output as untrusted, and require confirmation for consequential actions. Testing measures how well that containment holds.

We use a major model provider. Is the provider responsible for these risks?

Partly, for the model itself. Most of the list is about how your application uses the model: what goes into the context, what tools it can call, how output is handled. Those are yours.

How often should we test?

Before launch, and again when you add tools, data sources or a new model. A change in what the model can reach changes the risk more than any change in the model.

Need your AI feature tested against the 2026 list? We assess what can reach the model, what the model can reach, and where the boundaries fail.

AI security assessmentOr LLM red teaming

Before you go

Want the rest of this by email?

If this was useful, I send a few short notes on AI and LLM security. Unsubscribe in one click, and replies reach me directly.

From Jacob Masse, principal of traztech: the files by email, then a few short notes over the next month. No spam, unsubscribe in one click.

Want a second opinion on where you stand?

We run SOC 2, ISO 27001 and the rest of the compliance stack for startups and SMEs, and the security testing that sits behind it. The first call is free, and we will tell you if you are not ready to start yet.

Book a free call

Track record

Who is actually doing the work

5
Published CVEs, including a CVSS 9.1
Zero
Exceptions on a SOC 2 Type II built from nothing in-house

Published vulnerability research

Five published CVEs. CVE-2024-45163 (CVSS 9.1) is a flaw in the Mirai botnet itself, which gave defenders a way to shut down attacker infrastructure. CVE-2026-42626 takes HP ENVY 5000 printers offline from any unauthenticated device on the same network.

A SOC 2 Type II built from nothing

At Humera, a venture-backed US security company, Jacob built the compliance programme in-house from nothing: no report, no policies, no documented controls. It ended in a Type II attestation with zero exceptions.