re-ai-gov Join the waiting list

Kennisbank

What the risk level of an AI application depends on

A classification is not a label, it is an outcome

The question of whether an AI application within your organisation qualifies as high risk is not a question you answer once and then tick off. It is an outcome of a number of factors that together determine how heavily an application weighs. If one of those factors changes, the classification may change as well. That makes risk classification less a matter of filling in a form and more a matter of keeping track.

This page describes the mechanism behind that classification: what it depends on, and what happens when the situation changes. The current legal definitions and thresholds are covered elsewhere; here the focus is on what you need to be able to see and demonstrate within your own organisation.

What influences the classification

A number of factors structurally play a role in determining whether an application weighs heavily.

The first is the domain in which the application is deployed. Some domains — think of decisions about people, their opportunities or their rights — weigh more heavily than applications that are purely internal and supportive. What exactly the application does within that domain makes a difference here too: a system that advises weighs differently than a system that decides automatically.

The second factor is the role your organisation plays in relation to the system. Do you build it yourself, purchase it, or use a service that has AI embedded in it without you having procured it as such? The role you play — provider or user — partly determines which obligations fall on you and which part of the responsibility remains with a supplier.

The third factor is what happens to the model after its initial deployment. A system that is used unchanged as delivered is a different situation than a system that you fine-tune, reconfigure, or have learn on your own data. What changes when you adapt a model yourself is therefore a question that must be answered separately — adaptation can shift a system from one category to another.

The fourth factor is scope: how many people are affected, at what frequency, and how reversible is a mistake. An application that is used occasionally by a small team weighs differently than an application that influences decisions for thousands of customers on a daily basis.

Why the same application can shift classification

These factors are not static. An application that qualifies as limited risk today may no longer be so tomorrow — not because the rules changed, but because the application itself changed. A pilot that scales up to the whole organisation, an internal tool that starts affecting a customer process, a model that is retrained on new data: each of these steps can tip the classification.

This means that classification is not a one-off exercise. It is a question that must be asked again each time something changes about what a system does, who it affects, or who manages it. A governance structure that does not track this falls behind actual practice.

What does not fall within this classification

Not every AI application needs to go through this entire process. Part of what runs within an organisation falls outside the scope of risk classification — for example because it does not affect decisions about people or because it has only a supportive, non-determining function. Exactly where that line lies and which applications fall outside scope depends on the same factors described above: domain, role, adaptation and scope. So it is not a separate list, but the other side of the same assessment.

From classification to action

A classification by itself changes nothing. It only determines which follow-up steps are relevant, and with what urgency. Some actions are needed immediately, others can be scheduled — and what needs to happen now and what can be scheduled again depends on how heavily the application weighs and how many people it affects.

To make those steps visible and traceable, decisions about classification and follow-up actions need to be recorded somewhere — not as loose notes, but as part of a record of oversight decisions that shows who decided what and on the basis of which information. Without that record, a classification is a snapshot that no one can reconstruct afterwards.

Why this cannot be done without a complete picture

This classification question can only be answered properly if it is known what is actually running. An IT list of approved tools is not enough for this: a large part of usage arises outside that list, in teams that started using a tool without reporting it anywhere. Anyone who wants to map that usage must ask about it — and do so only without any consequences attached, otherwise the answer will not come.

The next question: how much work is it, and where

Once it is clear which applications exist and how they weigh, a follow-up question arises that is no longer about risk, but about work: which part of a task can an AI system take over, and which part remains human work. That question is answered by the work scan from FTE TO AI, which calculates per task which part of the work can be taken over — a different angle than risk classification, but one that builds on the same inventory.

This page describes the mechanism behind the classification. The Responsible AI Scan itself is under development; anyone who wants to use its results once they become available can sign up for the waiting list.

Andrewde assistent van de Responsible AI Scan

Vraag maar. Governance begint bij weten wat er draait — ook wat niemand heeft goedgekeurd.

Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.