llmaudit.eu

LLM security audit · European Union

Scope an audit

Annex 03

EU AI Act for companies deploying chatbots, RAG and agents

Updated 12 min read llmaudit.eu

Most guidance published before August 2026 is now wrong about dates, and a good deal of it was written for AI vendors rather than for the companies that buy their products. This annex is for the second group: organizations that deploy a chatbot, a retrieval assistant or an agent, and need to know what binds them, when, and what evidence answers it.

Which regulation, and which version

The instrument is Regulation (EU) 2024/1689, the Artificial Intelligence Act, adopted on 13 June 2024, published in the Official Journal on 12 July 2024 and in force since 1 August 2024. It has since been amended by Regulation (EU) 2026/1744, the Digital Omnibus on AI, adopted on 8 July 2026, published in the Official Journal on 24 July 2026 and in force from the third day after publication, which the European Commission states as 27 July 2026.

That second instrument is the reason a lot of published advice is stale. It is adopted law, not a proposal, and it moved two of the three deadlines that matter to most deployers while leaving the third exactly where it was.

Are you a provider or a deployer

Nearly everything else follows from this. Under Article 3, a provider develops an AI system, or has one developed, and places it on the market or puts it into service under its own name or trademark. A deployer uses an AI system under its own authority, except where the use is a personal, non-professional activity. Buying a license to a commercial assistant and switching it on for your staff makes you a deployer. Building the assistant, or shipping someone else's model under your brand, makes you a provider.

The trap is Article 25. A distributor, importer, deployer or other third party becomes the provider, with the full Article 16 obligations, in any of three cases: they put their name or trademark on a high-risk system already on the market; they make a substantial modification to a high-risk system; or they change the intended purpose of a system, including a general-purpose one, so that it becomes high-risk. White-labeling a vendor's chatbot as your own customer service agent, and repurposing a general assistant to screen job applications, both sit inside that article.

The Commission's FAQ on transparency adds a useful clarification for large organizations: where the deployer is a legal person, individual employees acting under its instructions are not separate deployers, and the legal person remains the deployer even when contractors or freelancers operate the system on its behalf.

The calendar, as amended

AI Act application dates after Regulation (EU) 2026/1744
DateWhat appliesStatus
1 Aug 2024Regulation (EU) 2024/1689 enters into forceIn force
2 Feb 2025Chapters I and II: prohibited practices, AI literacyIn force
2 Aug 2025General-purpose AI model obligations, governance, penaltiesIn force
27 Jul 2026Regulation (EU) 2026/1744 enters into forceIn force
2 Aug 2026General application, including Article 50 transparencyIn force
2 Dec 2026Article 50(2) marking grace ends; two added prohibitions applyPending
2 Dec 2027High-risk obligations, Article 6(2) and Annex IIIPending
2 Aug 2028High-risk obligations, Article 6(1) and Annex IPending
Dates taken from the Official Journal texts of both regulations. Where a secondary source disagrees, prefer the Official Journal.

Two of these deserve a note. The Annex III move from 2 August 2026 to 2 December 2027 was justified in the recitals by the delayed availability of harmonized standards, common specifications and national competent authorities. And the 2 December 2026 entry is narrow: it is a four-month transitional period for providers of generative systems that were already on the market before 2 August 2026, and it covers the machine-readable marking obligation only.

Article 4: the duty that already applies to everyone

Since 2 February 2025, providers and deployers have had an AI literacy duty. Regulation (EU) 2026/1744 rewrote it: the obligation is now to take measures to support the development of AI literacy among staff and other people operating the system on the organization's behalf, taking account of their technical knowledge, experience, education and the context of use. The amended text says explicitly that this does not require guaranteeing any specific level of AI literacy for any individual.

In practice this is a documentation exercise rather than a certification one: what the people who use the assistant were told, when, and what they were told about its limits. Named threat scenarios from a security assessment make unusually good material for it, because they are specific to the system your staff actually use.

Article 50 in practice

This is the part that is already binding on ordinary deployments, and it splits by role. Paragraphs 1 and 2 bind providers; paragraphs 3 and 4 bind deployers; paragraph 5 sets the timing for all of them.

  • 50(1), providers. Systems intended to interact directly with natural persons must be designed so that those persons are informed they are interacting with an AI system, unless it is obvious to a reasonably well-informed, observant and circumspect person. The Commission guidance says this exception should be read restrictively.
  • 50(2), providers. Systems generating synthetic audio, image, video or text must mark their output in a machine-readable format as artificially generated or manipulated, effectively, interoperably, robustly and reliably as far as technically feasible.
  • 50(3), deployers. Deployers of emotion recognition or biometric categorization systems must inform the people exposed to them of the operation of the system.
  • 50(4), deployers. Deep fakes must be disclosed. AI-generated or manipulated text published to inform the public on matters of public interest must be labeled, unless it went through human review and someone holds editorial responsibility.
  • 50(5). The information must be clear and distinguishable and provided at the latest at the time of the first interaction or exposure, and must meet the applicable accessibility requirements.

The Commission has published guidelines on these obligations and has assessed a Code of Practice on Transparency of AI-Generated Content as an adequate compliance route for the marking and labeling duties in 50(2), 50(4) and 50(5). Signing the code is voluntary; organizations that do not sign must demonstrate compliance by other adequate means and, per the Commission's FAQ, should expect more requests for information. For 50(1) and 50(3) there is no code, and operators determine adequate measures themselves with the guidelines in view.

Enforcement sits mainly with national market surveillance authorities, with the AI Office competent for systems built on general-purpose models where the same entity provides both, and for systems integrated into very large online platforms or search engines. The European Data Protection Supervisor covers EU institutions. The Commission put the ceiling for these breaches at EUR 15 million or 3 percent of total worldwide turnover, with proportionality taken into account for SMEs and small mid-caps.

The Annex III self-check

High-risk classification is where deployers most often guess. Article 6(2) makes any system listed in Annex III high-risk. The areas most relevant to ordinary corporate deployments are area 4, employment and workforce management, and area 5, access to essential private and public services.

  • Area 4 covers recruitment and selection, including targeted job advertising, analyzing and filtering applications and evaluating candidates; and decisions on promotion, termination, task allocation based on behavior or personal traits, and monitoring or evaluating performance.
  • Area 5 covers eligibility for essential public assistance benefits and services, evaluating creditworthiness or establishing a credit score with an exception for financial-fraud detection, risk assessment and pricing for life and health insurance, and emergency call triage and dispatch.
  • Other areas cover biometrics, critical infrastructure safety components, education and vocational training, law enforcement, migration and border control, and the administration of justice and democratic processes.

Article 6(3) provides a derogation where a listed system does not pose a significant risk of harm because it performs a narrow procedural task, improves the result of a previously completed human activity, detects decision-making patterns without replacing human assessment, or performs a preparatory task. That derogation has a hard limit written into the same paragraph: a system that performs profiling of natural persons is always high-risk.

Article 6(4) then requires a provider relying on the derogation to document that assessment before the system is placed on the market or put into service, and to register under Article 49(2). A deployer that becomes a provider through Article 25 inherits that duty. In other words, the conclusion "we decided it is not high-risk" is itself a documented artifact, not an internal opinion.

Article 26: what a deployer of a high-risk system must do

From 2 December 2027 for Annex III systems, Article 26 sets out the deployer duties. Twelve paragraphs, of which these matter most to a security or engineering function.

  • Use the system in accordance with the instructions for use, with appropriate technical and organisational measures.
  • Assign human oversight to natural persons who have the necessary competence, training, authority and support.
  • Where the deployer's controls input data, ensure it is relevant and sufficiently representative for the intended purpose.
  • Monitor operation, inform the provider and the market surveillance authority without undue delay where use may present a risk, and suspend use; report serious incidents immediately.
  • Keep the automatically generated logs, where they are under the deployer's control, for a period appropriate to the purpose and at least six months, unless other Union or national law says otherwise.
  • Inform workers representatives and affected workers before putting a high-risk system into service at the workplace.
  • Use the provider's Article 13 information to carry out a data protection impact assessment where one is required.
  • Inform natural persons that they are subject to the use of an Annex III high-risk system where it makes or assists decisions about them.
  • Cooperate with competent authorities.

The six-month log floor is the requirement that most often fails a technical check. Assistant deployments frequently keep seven or fourteen days of application logs, and the model provider retains its own for a different period under a different contract. Reconciling those is cheap before a deadline and expensive after an incident.

Article 15, and what a security audit actually evidences

Article 15 requires high-risk systems to be designed and developed to achieve an appropriate level of accuracy, robustness and cybersecurity and to perform consistently in those respects throughout their lifecycle. Paragraph 5 is the one that describes this work: high-risk systems must be resilient against attempts by unauthorized third parties to alter their use, outputs or performance by exploiting system vulnerabilities, and the technical solutions must, where appropriate, prevent, detect, respond to, resolve and control for attacks that manipulate the training data set, poison pre-trained components, use adversarial examples or model evasion, or mount confidentiality attacks.

A security audit produces evidence for a defined subset of the obligations. It does not produce a compliance opinion, and any provider claiming otherwise is overselling.

What technical testing can and cannot evidence
ObligationEvidence a security audit producesStill needed from elsewhere
Art. 15 robustness and cybersecurityAdversarial test record with payloads, run counts and outcomesAccuracy metrics and the declared accuracy levels
Art. 12 and 26(6) loggingCoverage and retention check against the six-month floorRetention policy and contractual terms with the provider
Art. 50 transparencyDisclosure check at first interaction and after session resetLegal review of the obvious-interaction exception
Art. 6(3) and 6(4) classificationWhether profiling actually occurs in the deployed data flowThe legal classification memo itself
Art. 26(2) human oversightWhether the confirmation step shows the real action and argumentsRole definitions, training and authority to intervene
Art. 9 risk managementNamed, tested threat scenarios with residual risk statedThe risk management system as a documented process
The left column is an obligation. The middle column is an input to demonstrating it. They are not the same thing.

Penalties

Article 99 sets three tiers. Breach of the Article 5 prohibitions carries up to EUR 35 million or 7 percent of total worldwide annual turnover, whichever is higher. Breach of provider obligations under Article 16, deployer obligations under Article 26 or the Article 50 transparency duties carries up to EUR 15 million or 3 percent. Supplying incorrect, incomplete or misleading information to notified bodies or national authorities carries up to EUR 7.5 million or 1 percent. For SMEs, including start-ups, each fine is capped at the lower of the percentage or the fixed amount.

ISO/IEC 42001 and prEN 18286

A management-system certificate is not compliance. ISO/IEC 42001 is a voluntary, certifiable standard for an AI management system, useful for organizing governance and increasingly asked for in procurement; the AI Act is binding law with obligations that a management system does not by itself discharge.

The bridge being built is prEN 18286, a European standard on a quality management system for AI Act regulatory purposes, drafted by CEN-CENELEC JTC 21. According to the Cloud Security Alliance research note on its public enquiry, it operationalizes the Article 17 quality management requirement through thirteen elements and includes an annex mapping to the ISO/IEC 42001 Annex A controls; once cited in the Official Journal it would give a presumption of conformity under Article 40. It is not published yet, so for now it is a planning input rather than a compliance route.

A deployer checklist

  • Write down, for each assistant, whether you are a provider or a deployer, and re-check it against Article 25 whenever it is rebranded or repurposed.
  • Record what your staff were told about the system and its limits, to evidence the Article 4 duty.
  • Check the AI disclosure appears at first interaction, including after a session reset and on any deep link into the assistant.
  • For generated media and published text, check the marking and labeling survive export, copy and download.
  • Run the Annex III self-check honestly, and if you conclude the system is not high-risk, write the Article 6(4) assessment down now rather than in 2027.
  • Measure your log retention against the six-month floor, including the logs held by the model provider rather than by you.
  • Commission adversarial testing that produces a written record, and keep the record with the risk assessment rather than in a ticket.
  • Re-test after a model version change, a system prompt change, a new retrieval source or a new tool. Each of those invalidates part of the previous result.

For the technical half of this list, the OWASP LLM Top 10 audit checklist sets out what each test produces, and the obligation finder on the home page turns a role and a use case into the specific articles and dates.

Sources

  1. Regulation (EU) 2024/1689 (Artificial Intelligence Act), consolidated Official Journal text EUR-Lex · 2024 Articles 3, 4, 6, 12, 15, 25, 26, 50 and 99, and Annex III as summarized above.
  2. Regulation (EU) 2026/1744 (Digital Omnibus on AI) EUR-Lex · 2026 Amended Article 4 wording, the amended Article 113 dates, the four-month Article 50(2) transition and the small mid-cap definition.
  3. AI Act: regulatory framework and implementation timeline European Commission · 2026
  4. Transparency obligations under Article 50: frequently asked questions European Commission · 2026
  5. Guidelines on transparency obligations for AI systems European Commission · 2026
  6. Safer and more transparent AI European Commission · 2026
  7. Research note on prEN 18286 and ISO/IEC 42001 Cloud Security Alliance · 2026

Questions

Related questions

We only use a commercial assistant. Does any of this apply to us?
Yes, in two ways. You are a deployer, so the Article 4 AI literacy duty has applied since 2 February 2025. And if you put the assistant in front of customers under your own name, Article 3 and Article 25 can make you its provider, which brings the Article 50(1) disclosure duty with it. Whether anything else applies depends on what the assistant is used for.
Does the AI Act apply to companies outside the EU?
It applies to providers placing systems on the EU market and to deployers using them under their authority in the EU, regardless of where the organization is established. The Commission's FAQ also notes that providers located outside the EU are covered where the output of their system is used in the EU. Establishment is not the test; the market and the use are.
Is a penetration test enough to satisfy Article 15?
No single artifact satisfies an article. Adversarial testing produces the record that Article 15(5) contemplates, and it is the part most organizations are missing, but the article also covers accuracy, declared accuracy metrics, resilience to errors and faults, and feedback loops in systems that keep learning. Treat the test as the evidence for one paragraph, not as the compliance answer.