OWASP LLM Top 10 · MITRE ATLAS · Art. 15 EU AI Act

AI Security: Penetration Testing and Red Teaming for AI and LLM Systems

We test your AI features the way an attacker would attack them: prompt injection, data exfiltration, manipulated RAG and overridden agents. You receive clearly documented attack paths, remediation recommendations and audit-ready evidence for customers, auditors and with a view to the future requirements of Art. 15 EU AI Act.

Project-based · fixed price by scope OWASP Top 10 for LLM Applications First findings usually within 1–2 weeks

Free and non-binding · 30 minutes · confidential under NDA · no tooling installation required

AI security, penetration testing and red teaming for LLM, RAG and agent systems
Law und Technology from a single mandate, EU AI Act classification and OWASP / MITRE ATLAS testing under one roof.

Working for regulated industries and SMEs

Dr. iur., LL.M., CIPP/E Findings that hold up before the board and supervisory authority
OWASP LLM Top 10 · MITRE ATLAS Recognised testing frameworks for AI systems
Confidentiality & data residency in Switzerland NDA, duty of confidentiality, data in Switzerland
Advisory and audit kept separate credible towards auditors
Oliver Stutz

Technical lead of this mandate

Oliver Stutz

Co-Founder · InfoSec & Cybersecurity Expert · Bachelor of IT (Bond University)

Oliver Stutz leads the technical testing of your AI systems per OWASP, MITRE ATLAS and NIST AI RMF and is responsible for the SIDD development centre for technical information security, cloud security and the integration of generative AI. The legal classification under the EU AI Act and the revised FADP (Swiss DSG) is provided by Dr. Dominic Staiger (CIPP/E, Attorney at Law). This produces findings that are both technically robust and defensible before the supervisory authority, customers and the board of directors.

LinkedIn profile

What is AI Security?

Brief definition

AI Security is the technical security testing of AI, LLM, RAG and agentic systems. AI Security, AI security testing, LLM pentest and AI red teaming all refer to the same service in our case. Unlike a classic penetration test, which examines deterministic applications and infrastructure, AI Security addresses the specific attack surfaces of language-processing and semi-autonomous systems: prompt injection, disclosure of sensitive information, insecure output handling, excessive agency of agents, and the poisoning of training and retrieval data. Testing is carried out against recognised frameworks: the OWASP Top 10 for LLM Applications (2025), MITRE ATLAS and the NIST AI Risk Management Framework.

An AI system does not behave deterministically. The same input can produce different outputs, and attacks hide in seemingly harmless text, for example in a document that your retrieval system ingests or in a web page that your agent fetches (indirect prompt injection). A one-off test therefore only provides a snapshot. We work with semi-automated, repeatable campaigns that systematically run through typical attack paths and document them in a traceable way, as a basis for remediation, re-testing and proof towards third parties.

Do you need an AI security assessment?

AI security becomes relevant as soon as an AI system has an external footprint or gains access to sensitive data and functions. The following four triggers are the most common in practice; if one of them applies to you, an initial consultation is worthwhile.

1

Go-live of an AI feature: You are putting a customer-facing chatbot, a RAG search or an agent into production and need to know before launch whether it can be manipulated or abused for data exfiltration.

2

Auditor or customer requirement: An EU customer, a procurement questionnaire or an ISO 27001 audit requires evidence that your AI systems have been tested and secured.

3

High-risk classification under the EU AI Act: If your system falls under the high-risk categories (Art. 6 et seq., Annex III), Art. 15 EU AI Act will in future require evidence of accuracy, robustness and cybersecurity, including protection against, among others, data poisoning, model poisoning and adversarial examples (Art. 15(5)). Following the provisional Digital AI Omnibus agreement, the applicability of these obligations is expected to be postponed; until publication in the Official Journal, the original deadlines remain authoritative.

4

Incident or near-miss: A tester, a customer or a researcher has reported a prompt injection, a data leak or misdirected agent behaviour, and the board of directors wants to understand the scope.

What makes us unique

Law, software engineering and in-house AI development, in a single team

At our firm, these three disciplines interlock within a single team. Neither a pure law firm nor a conventional consultancy can deliver this combination. This is how we assess your AI systems with the knowledge of those who build AI themselves, while also placing every finding in its legal context.

Deep legal expertise

Lawyers with doctorates holding CIPP/E, Attorney at Law (NY) and Solicitor (UK) qualifications. We classify AI with legal certainty under the EU AI Act, the FADP and the GDPR, drawing on ongoing mandates rather than textbooks.

Our own software engineers

Our own development centre and the Priverion Platform. We know how AI systems are built today, and we examine what a system technically really does instead of relying on self-declarations.

In-house AI development

We develop and operate our own LLM and AI systems. As a result, we know the best practices and pitfalls firsthand, from RAG to prompt security.

What do we take care of?

We assess your AI system across the entire attack surface, from the model through the input and output layers to tools, data sources and the supply chain. The scope is defined in writing before we begin and is tailored to your system.

  • Threat modelling for LLM, RAG and agentic architectures: inventory of models, tools, datasets, identities and trust boundaries.
  • Prompt injection testing, both direct (jailbreaks) and indirect via manipulated content in RAG sources, documents and retrieved web pages (OWASP LLM01).
  • Testing for disclosure of sensitive information and system prompt leakage (OWASP LLM02, LLM07).
  • Insecure output handling: Testing for XSS, SSRF and code execution through unfiltered model outputs (OWASP LLM05).
  • Excessive agency of agents: Assessment of agents and tool calls with write, transaction or system privileges for abuse and missing constraints (OWASP LLM06; OWASP Agentic Security Initiative).
  • RAG and vector DB security: Poisoning of the knowledge base (data store poisoning), cross-tenant data leaks and weaknesses in embeddings (OWASP LLM04, LLM08).
  • Model and MLOps supply chain: Evaluation of model provenance, version control, update validation and SBOMs for AI components (OWASP LLM03).
  • Uncontrolled resource and cost consumption (unbounded consumption / "denial of wallet"): Testing for unbounded consumption (OWASP LLM10).
  • Framework alignment: Mapping of findings to NIST AI RMF, MITRE ATLAS and the security controls of ISO/IEC 42001:2023.
  • Evidence documentation: traceably documented findings with severity, remediation recommendation and evidence, with a view to Art. 15 EU AI Act and ISO 27001 Annex A; on request, a re-test after remediation.

Effort and depth of assessment

The effort depends on the architecture and attack surface; a single LLM chat application is assessed more quickly than an agentic system with multiple tools, external data sources and transaction privileges. We define the scope in advance and price on a project basis. Because AI systems are probabilistic, we recommend repeated campaigns rather than a one-off test.

10

Risk categories: full run-through of the OWASP Top 10 for LLM Applications (2025).

MITRE ATLAS

Coverage of the tactics and techniques relevant to your architecture, according to scope.

2 re-tests

typically included, to demonstrate the effectiveness of the remediation.

Service packages

We work on a project basis. Instead of monthly flat rates, you choose a testing depth that suits your architecture and your evidence requirements. All prices are indicative; the binding fixed-price offer follows after the scoping.

LLM Application Test

CHF 8'000–15'000 Indicative · 1 application

Focused security assessment of a single LLM or chatbot application.

  • Threat modelling of the application
  • Prompt injection (direct/indirect) and jailbreak testing
  • Testing for sensitive information disclosure and system-prompt leakage
  • Insecure output handling
  • Findings report with severity rating and remediation recommendations, one re-test

AI Red Teaming Programme

from CHF 30,000 recurring campaigns

Ongoing, adversarial testing campaigns for high-risk systems or to prepare for the high-risk requirements (Art. 15 EU AI Act).

  • Repeated, partially automated red-team campaigns
  • Full OWASP LLM and agentic coverage across multiple cycles
  • Evidence package geared towards Art. 15 EU AI Act and ISO 27001 Annex A
  • Integration into your ISMS / your ISO 42001 governance
  • Support for the board of directors and supervisory bodies by Dr Staiger

Not sure which testing depth is right for you? Arrange an initial consultation, we define the scope together.

How we proceed

The process is designed for reproducibility and defensibility; each phase delivers documented evidence.

Scoping and classification

We clarify the architecture, data, tools and the legal role (provider or deployer, EU AI Act risk class) and define the scope of the assessment in writing.

Threat modelling

We map attack surfaces, trust boundaries and critical paths along the OWASP LLM Top 10 and MITRE ATLAS.

Active testing

Semi-automated and manual attacks, prompt injection, data exfiltration, output handling, agent/RAG abuse.

Analysis and report

Each finding with severity, reproduction steps and a concrete remediation recommendation; mapped to the EU AI Act, ISO 42001 and ISO 27001.

Remediation support and re-test

We verify the implemented measures and document their effectiveness.

Evidence and repetition

You receive an evidence package for auditors and clients; for probabilistic systems we recommend recurring campaigns.

Why SIDD

What the three disciplines above mean for the security of your AI systems:

We know how LLMs break

Because we build LLM and AI systems ourselves, we know the vulnerabilities from development, not just from checklists, from prompt injection to insecure agents.

Recognised testing frameworks

OWASP Top 10 for LLM Applications, MITRE ATLAS and NIST AI RMF, mapped to ISO/IEC 42001 and ISO 27001 Annex A.

Finding plus legal classification

Every technical finding is also classified legally (Art. 15 EU AI Act), so the results stand up before clients, auditors and the board of directors.

Confidential, data in Switzerland

NDA and confidentiality obligation; findings and data remain in Switzerland. Where Dr Staiger provides the legal classification, attorney-client professional secrecy (Art. 321 SCC) additionally applies.

Advisory and audit kept separate

We keep advisory and testing strictly separate, which makes our findings credible to auditors.

Evidence instead of slides

AI inventory, risk and findings audit-ready on the Priverion Platform, on request with a re-test after remediation.

Classic penetration test and AI security compared

Both assessments have their place. An existing cloud LLM provider secures its model, not your application around it. The following table shows why a classic penetration test does not cover the AI-specific risks.

AspectClassic penetration testAI Security / LLM penetration test
Subject of the assessmentWeb, network, API, infrastructureLLM, RAG, agents, model supply chain
System behaviourDeterministicProbabilistic / stochastic
Guiding attacksOWASP Top 10 (Web)Prompt injection, poisoning, excessive agency
Reference frameworkOWASP, PTES, OSSTMMOWASP LLM Top 10, MITRE ATLAS, NIST AI RMF
Indirect attacksRarely relevantCentral (prompt injection via RAG/web)
Testing cadenceRepeatable on an ad hoc basisRepeated campaigns recommended
Compliance relevanceISO 27001 Annex AISO 27001 Annex A, Art. 15 EU AI Act, ISO 42001

Classic penetration testing and AI security complement each other; we also carry out the classic penetration test as well as the Vulnerability Scan as well.

References

Regulated industrial group

RAG assistant for internal documentation

Ahead of the internal rollout of a RAG-based knowledge assistant, we tested the system for indirect prompt injection and cross-tenant data access. Three critical paths to the disclosure of non-released documents were closed before go-live and confirmed in the re-test. (Anonymised, regulated industry.)

SaaS provider in the healthcare sector

Customer-facing AI agent

For a provider with a customer-facing, tool-using agent, we produced a threat model and an evidence package that makes the robustness and cybersecurity aspects (with a view to Art. 15 EU AI Act) demonstrable to customers and auditors. (Anonymised, health SaaS.)

Our tool: LexCommand

Why we work with LexCommand, our own Swiss legal AI

LexCommand is our in-house, citation-backed legal AI for the law of Switzerland, Germany, Austria and the EU. Developed and run sovereignly in Switzerland by Priverion GmbH, the company behind SIDD. We do not just preach data sovereignty and provability, we built them into our own tool, alongside the Priverion Platform.

Sovereign in Switzerland

The AI runs self-hosted on Swiss infrastructure, with no external cloud LLMs. As an independent Swiss company with no foreign parent, we process your documents in an environment we control.

No citation, no claim

Every legal statement traces back to a retrievable primary source, or it does not appear at all. That makes our recommendations auditable and verifiable, instead of merely sounding plausible.

From effort to judgement

LexCommand takes over searching, cross-checking and sourcing. That shortens turnaround times and frees our senior advisors for judgement and client dialogue, with no loss of diligence.

Three disciplines, one picture

We look at data protection, information security and AI security on a shared source base with a framework crosswalk. So you see overlapping obligations in one consolidated picture, instead of three isolated analyses.

For the AI security assessment, concretely, LexCommand grounds the legal classification of every finding against Art. 15 EU AI Act, ISO 27001 Annex A and ISO/IEC 42001 with an exact source in the evidence pack, and given the shifting AI Act deadlines it applies the version of each norm in force on your chosen reference date.

Temporally deterministic (as of today or any reference date), with jurisdiction isolation (CH/DE/AT/EU) and a citation verifier at the end of every answer.

Frequently asked questions

What does an AI security assessment cost?

The assessment is project-based. A single LLM application typically ranges from CHF 8'000–15'000, RAG and agent systems from CHF 15'000–30'000, and ongoing red-team programmes from CHF 30'000. The binding fixed price follows after scoping.

Is a one-off assessment enough?

For an initial assessment, yes. Because AI systems are probabilistic and models, prompts and data sources change, we recommend repeated campaigns. An assessment is a snapshot, not a lasting guarantee.

Our LLM comes from a large cloud provider, isn't that already secure?

The provider secures the model, not your application. Prompt injection, insecure output handling, excessive agent permissions and manipulated RAG arise in your integration, and that is precisely where our assessment focuses.

How long does a project take?

A single LLM application is usually assessed and reported within one to two weeks. RAG and agent systems take longer. We schedule the re-test once the findings have been remediated.

Does this make us EU AI Act compliant?

No, guaranteed compliance does not exist. We provide technical findings and evidence that contribute to meeting the (future) requirements of Art. 15 EU AI Act (accuracy, robustness, cybersecurity) and ISO 27001 Annex A that make your due diligence demonstrable. This does not replace case-specific legal advice.

From when do the EU AI Act obligations apply to high-risk systems?

Prohibitions (Art. 5) and AI literacy (Art. 4) have applied since 02.02.2025, and GPAI obligations since 02.08.2025. Most high-risk obligations (Annex III) were originally scheduled to apply from 02.08.2026. Following the provisional agreement on the "Digital AI Omnibus", they are expected to be postponed to 02.12.2027 (Annex III) and 02.08.2028 (Annex I) respectively. As of 18.06.2026 this agreement had not yet been published in the EU Official Journal and is therefore not final; until formal adoption, the original deadlines formally apply. As a precaution, we plan for both scenarios. The requirements of Art. 15 (accuracy, robustness, cybersecurity) follow the high-risk timeline. We update this information on an ongoing basis.

How does this dovetail with our CISO or DPO?

We work on a shared-responsibility basis. The AI security findings fit into an existing ISMS and into the work of the CISO und Data Protection Officer (DPO) ; the fundamental rights impact assessment (FRIA, Art. 27 EU AI Act), where applicable to your system, can be conducted alongside the DPIA (Art. 35 GDPR). On request, we combine the assessment with our services ISMS/ISO 27001, Penetration test and the AI Officer role.

Last reviewed: 18.06.2026 by Dr. Dominic Staiger. This page does not replace case-specific legal advice.

Make the security of your AI systems demonstrable

We assess your LLM, RAG and agent systems in line with OWASP, MITRE ATLAS and with a view to Art. 15 EU AI Act, and deliver findings that hold up before clients, auditors and the board of directors.