re-ai-gov Join the waiting list

Kennisbank

The pilot that was never officially stopped

How a test setup comes into being

A team wanted to try something out. A chatbot for customer questions, a script that summarises reports, a connection to a language model to sort emails. No major decision was needed: someone had an account, an API key or a free trial, and within an afternoon something was running. The pilot worked, or worked well enough, and no one had a reason to switch it off anymore.

This is not an exception. It is the ordinary way AI enters an organisation. Not through a tender or an approved budget, but through a test setup that outlives its own trial status. The test phase is never formally closed, because there was never a formal beginning to close either.

Why it does not disappear on its own

A pilot that works gets used. And what gets used becomes a dependency. The team that built the report summariser may have moved on, but the summary still comes in every week. No one has been instructed to switch it off, and no one wants to take the risk of something ceasing to work without knowing what will replace it.

On top of that, a test setup usually has no owner in the sense that a governance structure expects. No risk classification has been made, no data analysis carried out, no decision taken on who is responsible if things go wrong. The pilot exists in a middle space: too much in use to ignore, too informal to manage. That is exactly the pattern that leads to employees using a tool that no one has approved — only here at team level instead of individual level.

Why assigning blame makes the problem worse

The reflex on discovering an unauthorised test setup is often: who allowed this, and why did we not know. That question is understandable, but it works against you. Anyone who senses in the first conversation that a blame question is coming will not say anything the next time. And the next test setup — which will undoubtedly appear — will disappear from view just as easily as this one did.

An inventory built on trust yields more than an audit built on control. Not because people would want to hide something, but because the information you are looking for sits with the user, and that user only talks if he is not held to account for what he finds.

What you can do with a test setup

Once a test setup is in view, the first question is not whether it may continue to exist, but exactly what it does. What data goes into it, who uses the output, and what happens if the output is wrong. Those are the same questions that apply to any other AI application, and the answer determines whether something is an informal tool or a risk that deserves attention.

Classification follows: does the application fit into a risk category that already exists, or does it call for a new assessment. Then: what needs to be recorded to demonstrate that the organisation knows what is running and why. That is precisely what what you need to record per application addresses — not as extra bureaucracy, but as the minimum documentation needed to distinguish a pilot from a risk.

A test setup that passes this check can continue to exist, now with an owner and a classification. A test setup that fails the check must be phased out — but that is a controlled process, not a dismissal of the person who once started it.

The broader context: shadow AI is the rule, not the exception

The test setup that was never switched off is one form of a broader pattern. Think of the browser extension with access to your mail that someone once installed to save time, or the company data that ends up in a free chat window because it was faster than the official tool. All these situations share one characteristic: they arose from a practical need, not a policy choice, and they continue to exist as long as no one asks about them.

The IT list of approved software is therefore not the starting point of an inventory — it is the starting point of a search for what falls outside that list. Anyone who wants to know what that search looks like will find an approach in how you build an AI inventory.

From inventory to insight into the work itself

A test setup that comes to light usually raises a follow-up question that goes beyond governance: why did this actually work so well that no one dared to stop it? That is a question about the work itself, not just about the risk. The work scan from FTE TO AI calculates per task what proportion of the work can be taken over by AI, thereby making visible what a test setup already implicitly showed: that part of the work can be organised differently. Where the Responsible AI Scan maps out what is running and under what risk, the work scan shows where that use comes from and what it structurally means for the division of tasks.

The current state of this component

The Responsible AI Scan, with the inventory, classification and governance set described above, is under development. Anyone already dealing with this now and who wants to be informed as soon as the instrument becomes available can sign up for the waiting list.

Andrewde assistent van de Responsible AI Scan

Vraag maar. Governance begint bij weten wat er draait — ook wat niemand heeft goedgekeurd.

Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.