Key Takeaways
- AI pilots should measure whether AI improves specific work outcomes, not just whether employees adopt or use the technology.
- Establishing a baseline before an AI pilot helps organizations determine whether performance, quality, speed or other targeted outcomes improved.
- AI pilot results should account for factors such as quality, errors, rework, human review, cost and risk rather than focusing only on productivity gains.
- Organizations can use AI pilot evidence to decide whether to scale, change or stop an AI implementation based on its impact on the targeted work.
Organizations are experimenting with artificial intelligence (AI) across a growing range of work. However, at some point, we need to move beyond whether employees are using AI and ask what we are learning from these pilots. Organizations need enough evidence to determine whether a pilot should scale, change or stop.
A pilot should help answer a basic business question: Did AI make the work better? Adoption, usage, confidence and satisfaction can be useful, but they answer different questions.
Using AI does not mean AI improved the work. Moreover, improved work does not necessarily mean AI produced a meaningful business result. The evidence needs to match the decision the organization is trying to make.
Start With What You Want to Improve
Before the pilot begins, identify the problem or opportunity AI is expected to address. What work should be different if the pilot succeeds? Is the goal to reduce cycle time, improve quality, increase output, reduce errors, improve service or accomplish something else?
One approach I have used is to start with the existing workflow. Where is the friction? What is taking too long, producing errors, creating rework or otherwise affecting performance? Only then do we ask where AI might help. This gives us something specific to evaluate.
It also helps avoid a common problem: assuming AI will have the same effect across different kinds of work. A 2026 study of 758 knowledge workers found that AI helped people complete some tasks faster and produce higher-quality work. On another task, however, those using AI were 19 percentage points less likely to reach the correct solution.
The practical point is simple. Do not start an AI pilot with a broad objective such as “improve productivity.” Identify the work you want to improve and what better performance would look like.
Know Where You Are Starting
Once you know what should improve, establish the starting point. If the pilot is intended to reduce task completion time, how long does the task take now? If it is intended to improve quality, what does quality look like now? If the goal is to reduce errors or rework, what is the current rate?
This does not mean turning every AI pilot into a research study. Organizations may already have useful information in operational systems, work products, quality reviews or other performance data. The point is much simpler: If you want to know whether something improved, you need something meaningful to compare it with.
In my work on an AI capability initiative, we built measurement into the initiative before it began. We identified the work participants would be expected to perform with AI, examined where friction existed in the current workflow, established baseline evidence and determined what we would measure later. We wanted to know more than whether participants could use AI. We wanted to know whether they applied it in their work, whether performance improved and ultimately whether the evidence supported what to do next.
Consider a learning and development (L&D) team piloting AI to help develop scenario-based assessment items. Before introducing AI, the team could establish how long development typically takes, the quality criteria the items must meet and how much revision or rework is normally required. This gives the team something to compare with what happens during the pilot.
Measure What Actually Changes
During the pilot, measure the performance the pilot was designed to improve.
A 2025 study of more than 5,000 customer-support agents found that access to a generative AI assistant increased productivity by 15% on average. However, the average told only part of the story.
Less experienced and lower-skilled employees experienced the largest gains, while the most experienced and highest-skilled employees saw smaller gains in speed and small declines in quality.
That is why another organization’s results cannot tell you what will happen in yours. Your pilot needs to tell you what changed, for whom, and under what conditions.
It also needs to tell you whether the work actually got better.
Suppose AI reduces the time required to complete a task from three hours to two. That sounds promising. However, what happened to quality? Did errors increase? Was more rework required? Did someone need to spend another 30 minutes reviewing and correcting the AI-assisted work?
Faster does not necessarily mean better. If AI reduces development time by 30% while maintaining or improving quality, that is meaningful evidence. If reviewers then spend considerably more time correcting weak items, the organization has learned something different.
Look at what else changed. Depending on the work, that might include errors, rework, human review, cost, risk or unintended effects on the workflow. You do not need to measure everything that happens during an AI pilot. Measure what will help you make the decision.
Make the Decision
At the end of the pilot, bring the evidence together. What improved? By how much? What stayed the same? What got worse? What additional costs, risks or work were introduced?
If the pilot was intended to influence a business outcome that can reasonably be observed, examine that as well. However, not every pilot needs to demonstrate increased revenue or another distant organizational result. If AI allows employees to complete important work substantially faster while maintaining or improving quality, and the costs and risks are acceptable, that may be enough evidence to make a business decision.
Then make the decision the pilot was intended to support:
- Scale when the evidence shows that AI improved the targeted work enough to justify broader implementation and the consequences are acceptable.
- Change when the opportunity still appears worthwhile, but something needs to be corrected — the workflow, the AI tool, training or support, human review, or another part of the implementation.
- Stop when the expected improvement does not materialize or when the costs, risks or other consequences outweigh the benefits.
An AI pilot should reduce uncertainty. At the end, leaders should know more than whether employees tried the technology or liked using it. They should know what changed, whether it mattered, what the evidence allows them to conclude and what they should do next. The goal is not to prove that people can use AI. It is to provide enough evidence to make a business decision.


