Skip to content

Decision criteria

Sizing your first AI pilot: small, measurable, with the ending written down

How to size a first AI pilot: a small scope, success conditions written before you start, a clear point where you stop, and a plan for what happens if it works.

6 MIN READ

You've already decided AI makes sense for your company — or at least that it deserves a test. The next decision is more concrete and gets fumbled far more often: how much pilot to build. Too big, and it isn't a pilot: it's a project in disguise, carrying a project's risk with a test's improvisation. Too vague, and nobody will be able to say afterwards whether it worked — which was the only reason for doing it. This article is about sizing that first pilot: what's in and what stays out, which success conditions get written down before you start, when you stop, and what happens next. If you're still on the previous decision — whether applied AI belongs in your business at all — that's a different question; here we take the answer as given and design the first step.

A pilot is not a small project (or a demo)

It helps to separate three things that get confused. A demo shows something is possible; it doesn't touch your operation or your real data, which is why it answers no business question. A project builds something meant to stay; it's justified once you already know the thing works and what remains is doing it properly. A pilot sits in between and has a purpose of its own: answering a question with real data and contained risk. The question is usually some variant of "does this save time or reduce errors in our specific case, with our data and our people?".

Everything about sizing flows from that purpose. A pilot isn't sized to impress or to "keep momentum going": it's sized so the question gets answered with the minimum investment and the minimum risk. Anything that doesn't help answer it is excess. And one uncomfortable but healthy consequence: a pilot that cannot fail is not a pilot — it's theatre. If the outcome will be declared a success whatever happens, save yourself the effort.

Scope: one task, one team, one slice of reality

The temptation when sizing is to add: while we're at it, let's cover this too, and let that other department try it as well. Every addition makes the pilot more expensive, slower and — this is the serious part — harder to interpret: if something goes wrong, you won't know which part failed.

A healthy scope is defined by subtracting, along three cuts:

  • One task, not one area. Not "AI in admin", but a task you can name: classifying what arrives in one inbox, extracting the data from one type of document, drafting the first version of one type of reply. If the task doesn't fit in a sentence, it's still too big.
  • One small, willing team. The people who do that task today and want to try. A pilot is not the moment to convert sceptics: it's the moment to learn fast with people who are on board. The sceptics come later, once you have results in hand.
  • One slice of reality, not all of it. One document type, one intake channel, one segment of cases. Real enough for the result to mean something; contained enough that a mistake doesn't hurt.

And one precondition no amount of trimming fixes: the material the pilot works on has to be reasonably in order. Piloting AI on chaotic data doesn't answer "does AI work here?"; it answers "is our data a mess?" — and you already had that answer.

Success conditions get written before you start

This is the part that separates a pilot from a pastime, and it happens before launch, not at the end. Once the results are in, any outcome can be argued either way; written down beforehand, the paper decides. Three things in writing:

  • What we measure. The measure comes from the pilot's question: time the task takes, errors that slip through, cases resolved first time. And its obligatory companion: the starting point. If you don't know what the task costs today, measure that before switching anything on — without a before, there is no after to compare.
  • What result would make us continue. You don't need an invented numerical threshold: an honest sentence will do, along the lines of "we continue if the team that used it would rather not go back to the old way and the measures back that up". What matters is that it's written in advance and that a specific, named person will read it at the end and decide.
  • Which lines don't get crossed. What data may go in and what may not, which cases always go to a person, who reviews what. Limits aren't a brake on the pilot: they're what lets you hand it to a real team without nasty surprises.

The ending is designed too: stopping is a valid result

A pilot needs an end date agreed up front — a short one; if it needs a long time to prove anything, the scope is too big — and, above all, it needs stopping to be a dignified exit. If nobody has planned how it stops, the pilot will never die: it will stay half switched on, undecided, consuming attention and leaving behind a vague sense that "the AI didn't quite work out" without anyone able to say why.

Designing the exit is simple: on the agreed date, the person who signed the success conditions re-reads them and picks one of three — continue and grow, adjust and run once more, or stop. A pilot stopped on time with a clear conclusion is a success of the method: you bought learning cheaply. The alternative — zombie initiatives nobody kills — is the most expensive way of not deciding.

What happens if it works: the question almost nobody asks in advance

It's the quietest sizing mistake of all: the pilot goes well and there's no path prepared. What was a test stays in production de facto — no owner, no support, held up by the goodwill of the team that tried it. Months later it's a critical dependency nobody maintains: the worst of both worlds, the risks of a production system with the guarantees of an experiment.

Before launch, sketch the "what if it works" in broad strokes: what it would take for the whole team to use this (not in detail — in order of magnitude), who inside the company would own it, and which parts of the pilot get thrown away. That last one is worth saying out loud from day one: a pilot is built to learn, not to last, and getting serious usually means calmly rebuilding what was assembled in a hurry. Turning a pilot into a system the company can actually operate is project territory, with a different setup and different guarantees — and that's fine: that's what the pilot was for.

Five questions before you launch the pilot

  1. What specific question will this pilot answer?
  2. Does the scope fit in one sentence: one task, one team, one slice of reality?
  3. Are the success conditions — measure, starting point, who decides — written down before you start?
  4. Does it have an end date, and is stopping an acceptable result?
  5. If it works, do you know roughly what comes next and who would own it?

If all five have answers, your pilot is well sized: small in scope and serious in design. That combination — not enthusiasm, not budget — is what makes a first step with AI teach you something worth knowing.

After reading

Does this sound like your case?

If this describes something sitting on your desk, tell us about it. We'll come back with a first read before proposing anything.

Tell us about your case

A first 30-minute call with direct senior interlocution — no commitment and no sales pitch.