When an AI Pilot Stalls: Find the Breakdown Before Adding More Training
The pilot has launched, the team has attended a session and usage is disappointing. More training might help. First, find out what happens when someone tries to use the tool on a real task.
Investigate the attempted task.
1
Can they use it?
Access and working conditions
2
Does it help?
Task fit, context and output quality
3
Can work continue?
Capability, ownership and handoffs
Reconstruct an attempt
Ask a participant to walk through a recent attempt using information they are authorized to share. Where did they begin? What output did they receive? What did they need to check or fix? At what point did the old process become easier?
Compare that account with the sponsor’s intended outcome. The pilot may be measuring tool activity while the team is trying to complete work under different quality or time constraints.
Check five possible breakdowns
Access: can the right people use the approved tool in their actual working environment? A blocked account or awkward permission process needs an operational fix.
Task fit: does the tool perform the relevant task well enough on representative cases? If it struggles with the work itself, additional practice may not resolve the issue.
Information: does the task require context that is missing, outdated or inaccessible? Identify the source and who can validate it.
Capability: can people recognize a useful output, revise it and handle exceptions? If this is the gap, targeted practice with feedback may help.
Ownership: does someone have authority and capacity to accept the output, resolve exceptions and change the process? A training session cannot assign that authority.
Work through an illustrative case
A service team uses AI to draft replies. Drafting becomes easier, but supervisors return many replies for correction. The team starts using the tool less often.
Observation reveals that straightforward replies are acceptable. Cases involving a replacement promise need judgment about stock, policy and customer history. Agents cannot tell which promises require approval, and supervisors apply different criteria.
The next test might establish shared escalation rules and review a bounded group of cases. Training can then focus on recognizing those cases. A technical evaluation is still needed if the tool invents information or fails the agreed quality standard. Several causes can coexist.
Borrow a useful habit from practitioners
In his account of AI adoption at Imprint, Will Larson describes using adoption data to investigate what makes tools useful or difficult for people. That is a useful starting posture for a stalled pilot: investigate the experience behind the metric. His account describes one organization’s approach, not a universal benchmark.
Record the observed breakdown, your explanation and what evidence would contradict it. Choose an intervention small enough to examine, with an owner, a baseline, a review point and a stop condition.
For example: if unclear escalation is the main barrier, clearer criteria should reduce returned replies without lowering quality. Compare comparable cases, record supervisor effort and check for other changes that could explain the result. A short test can inform the next decision without proving organization-wide impact.
Gobekli can work with consultants on a focused pilot review and the human context it needs. Bring the original goal, the attempts already made and an example of where the work gets stuck. That is enough to begin scoping something useful.
Try this with one client situation.
Copy these prompts into your working notes. Keep observations, assumptions and open questions separate.