An AI pilot has done its job when it supports a decision. Continuing, changing, stopping and resolving a specific blocker are all legitimate outcomes.
The problem begins when experimentation becomes the permanent operating model. Tools are demonstrated, licences are bought and interested people test prompts, yet nobody can say which task changed, what was learned or what the firm decided.
This can feel prudent because deployment carries genuine risk. A pilot without a decision question or review date is less cautious than it appears. It spends time, data, money and employee trust without controlling what follows.
Define what “commit” means
Commitment does not have to mean a firm-wide rollout. It means making an explicit operational decision with resources, ownership, controls and a route to reassessment.
A firm might commit to using an approved tool for one internal drafting task with mandatory review. It might fund a bounded production service for one team. It might decide that a promising use case must wait until access controls are repaired. It might retire the tool because the benefit does not justify its risk and operating cost.
Each is more useful than describing another cycle of exploration.
The scale of commitment should match the consequence of the task. Automating a low-consequence internal classification, assisting regulated advice and exposing client-facing output need different evidence and governance.
Start with evidence from a defined task
A pilot should test a claim such as: “This tool can produce a usable first draft of this recurring internal document while reducing preparation time, without increasing review effort or material errors.”
That statement identifies a user, task, intended benefit and important counterweight. Before testing, define the baseline, sample, permitted data, evaluation method and decision threshold.
Measure the whole task. A faster first draft can be a poor result if checking takes longer, errors become harder to detect or employees move sensitive information into an unapproved system. Adoption is useful context and cannot substitute for value. High use can reflect novelty or a weak control; low use can mean the task is infrequent or the tool is unnecessary.
Record variation. If experienced staff benefit and new joiners do not, an average may hide an important safety issue. If one document type works and another fails, narrow the use case instead of declaring the tool successful or useless.
Four readiness questions
The source proposed four universal conditions. They remain useful questions, with thresholds tailored to the use case.
Has the use case produced decision-quality evidence?
The pilot should have pre-agreed measures, representative work and documented exceptions. Vendor demonstrations and enthusiastic anecdotes are insufficient. A result below the target may still support a narrower or redesigned commitment if the team understands why.
Check whether the environment resembles production. A curated test set, unusually attentive reviewers or manual support from the supplier can overstate normal performance.
Are the data and access controls sufficient?
The required data should be accurate enough for the task, lawfully and securely available, and restricted to appropriate users and purposes. Map what enters the system, what the provider retains, where outputs go and how access is removed.
Perfection is unnecessary. A known limitation can be acceptable when it is contained, visible to users and reflected in review. A hidden data-quality problem that can scale into client or operational harm is a reason to pause.
Can the firm operate the control framework?
Policies matter only when translated into the workflow. Name who approves the use, reviews outputs, handles incidents, monitors the supplier and decides changes. Define prohibited uses, human accountability, record keeping and escalation.
Apply relevant legal, privacy, security, professional and client obligations. Use the NIST AI Risk Management Framework as a useful governance reference and the ICO's AI and data-protection guidance for UK personal-data questions. Neither replaces advice on the specific service.
Is there an accountable sponsor and an operational owner?
The sponsor protects the business decision and resolves consequential trade-offs. The operational owner manages the service, evidence and controls. One person may hold both roles for a small use case; a high-impact deployment will involve several accountable functions.
Visible enthusiasm is not allocated capacity. Put responsibilities into objectives and governance, and ensure somebody can stop the service when controls fail.
If the answer is “not yet”
A blocker should produce a specific next decision.
For data, that may be a short assessment of the affected fields and access routes. For governance, it may be testing an output-review process against realistic volume. For value, it may be expanding the sample or comparing the tool with a simpler workflow change. For sponsorship, it may be deciding that the use case lacks enough commercial importance to continue.
Give the work an owner, evidence requirement and date. Avoid arbitrary urgency. Eight weeks may be too long for a simple control and far too short for a complex data remediation. The next review should follow the work needed and the cost of waiting.
Stopping can be the responsible result. Close accounts, remove integrations, deal with retained data, preserve useful learning and confirm contractual consequences. An abandoned pilot that remains connected is an unmanaged production risk.
What operational commitment requires
A funded service boundary
Show licence, integration, assurance, support, monitoring, training and exit costs. A separate budget line can improve visibility, though small uses may reasonably sit within an existing budget. The requirement is traceability and an owner, rather than a particular accounting structure.
Capacity to run it
Identify who supports users, reviews exceptions, maintains knowledge, assesses updates and responds to incidents. Do not base the return on volunteer labour or assume the vendor will perform client-side governance.
Measures agreed before launch
Carry forward the baseline and decide which measures indicate benefit, harm and use. Include quality, rework, exceptions and downstream effect alongside time or volume. Specify who collects them and what action each threshold prompts.
A review cadence
Book the first operational review and define earlier triggers. Ninety days may suit a regular internal workflow; a high-impact or rapidly changing service may need much earlier review. Rare use cases may need longer to generate evidence.
The review should allow continue, adjust, narrow, pause, replace or stop. It should examine incidents and near misses, changes to models or terms, user workarounds and whether the original task still matters.
A controlled release
Move by a deliberate increment: one team, one document type, limited data or a parallel run. Define who is included, how users are trained, what remains outside scope and how rollback works. Expansion follows evidence rather than a calendar promise.
The costs of permanent piloting
Pilot spend is the most visible cost. The larger cost may be organisational learning that leads nowhere.
Employees who repeatedly test tools and never hear a decision learn that participation has little value. That is not resistance to AI; it is a response to weak follow-through. Close each pilot with a short account of what was tested, what was found, what happens next and why.
Unmanaged personal use can also grow while the formal programme deliberates. Give people a current route for approved experimentation, data boundaries and questions. A slow central decision does not freeze behaviour.
Competitive claims require care. Operational experience can build capability. Rushed deployment can instead build incidents and distrust. The strategic advantage comes from making better, faster decisions about suitable tasks, rather than maximising the number of tools in production.
Partnership decisions need a better paper
In a partnership, commitment may require consent across leaders with different exposure to the work and risk. Present the decision in their language:
- the task and commercial purpose;
- pilot evidence and limitations;
- affected clients, people and data;
- options, including stop and delay;
- total operating cost;
- governance and accountable owners;
- release boundary;
- measures and review triggers;
- residual risks requiring acceptance.
This gives sceptics something testable and protects the initiative from being sold as inevitability.
Make the implicit decision visible
A firm that continues a pilot has decided to spend more time and expose more data or attention. Record that choice with the same discipline as deployment.
For every current experiment, ask: what question is it answering, when will the evidence be sufficient, who decides and what are the available outcomes? If those answers do not exist, pause the activity long enough to create them.
If you want to assess whether your firm meets the four readiness conditions - and if not, what specifically needs to happen and by when - we run a two-hour transition planning session that ends with a clear go/no-go recommendation and a specific action plan either way. There's also a one-page commitment readiness checklist (free download below) you can work through ahead of a partnership discussion, if that's the conversation you need to have first.



