The question of whether an AI application qualifies as high risk sounds like a technical qualification. In practice, it is a classification that depends on a number of concrete factors, and that classification then determines how much oversight, documentation and demonstrability is expected of you. The current legal definitions and thresholds are not on this page; you will find those elsewhere. Here we explain what the classification depends on and what changes in your organization if that classification shifts.
There is not one single characteristic that makes an application high risk. It is a combination of factors that together form the picture: the domain in which the application is deployed, the extent to which a decision has direct consequences for a person, and the question of whether there is a human who can actually assess the AI's advice or outcome and correct it if necessary before it takes effect. An application that only summarizes text is generally positioned differently in this classification than an application that factors into a decision about an individual. The context in which an application is used also matters: the same underlying technology can be considered low risk in one domain and high risk in another, depending on what happens with the outcome.
This is one of the reasons why a list of approved software is not sufficient. The classification does not depend on the label on the package, but on how it is actually used and where it sits in the decision-making process.
If an application is assessed as higher risk, something changes in what is expected of the organization: more documentation about how the application works and has been tested, a clearer oversight mechanism for the outcomes, and a way to show afterwards who paid attention to what. This is not a one-time action but an ongoing obligation: the documentation and oversight must remain current for as long as the application is in use.
The classification is also not static. An application that qualifies as limited risk today may be deployed differently tomorrow — in a different domain, with a greater impact on individuals, or with less human intervention than before. Those who determine the classification once and never revisit it run the risk that the classification no longer matches actual use. This applies equally to applications that are modified after acquisition; exactly what changes when you adjust a model yourself depends on the nature of that adjustment, as described on the page about adjustments to a model.
The classification of an application is only meaningful if it is applied to what is actually being used. A large part of the use within an organization runs outside the official inventory: teams that use a language model for a draft email, an analyst who uses an external model for a first version of a report, a department that has acquired a tool without reporting it. This use does not arise from ill will. It arises because the official route is slower than the need, or because no one knew that a route needed to be followed at all.
The consequence is that the IT list rarely gives a complete picture. Those who want to know what is really needed in terms of risk classification must map out the actual use themselves — and that requires asking questions of the people who do the work, without any repercussions following from it. Anyone who senses in the very first conversation that an honest answer will have consequences will not answer honestly. That renders the inventory unusable before it is even completed.
Not every application of AI falls within the framework in which risk classification is relevant. Some applications fall outside the scope of the regulation by their nature or purpose, and it is wise to make that distinction early, before spending time on classification that is not needed; which applications this concerns is described on the page about applications that fall outside the scope of the obligations. The role your organization plays — as a provider of a system or as a user of it — also influences which obligations are attached to a risk classification, and that role is not always clear in advance; what determines this is explained on the page about the distinction between provider and user. That same role can, incidentally, change as soon as you modify a system or offer it under your own name, something explained in more detail on the page about the circumstances under which your role changes.
A classification by risk level almost always raises the follow-up question of what is urgent and what can still be planned. That distinction is not the same for every organization and depends on which applications are already in production and which are still in development. An overview of what generally requires attention first and what can follow later is available on the page about what needs to happen now and what you can plan.
Once it is clear which applications within your organization qualify as high risk, another question naturally follows: how much of the underlying work is actually done by AI, and how much by people who check or supplement the outcome. That question falls outside the scope of this scan, but is answered by the work scan of FTE TO AI, which calculates per task what portion of the work can reasonably be taken over by AI.
Vraag maar. Governance begint bij weten wat er draait — ook wat niemand heeft goedgekeurd.
Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.