Prove one job before
you automate ten.
The useful AI workflow is the one that leaves less work for your team. Count the checking, the exceptions and the upkeep before you call it a saving.
A fast answer is only one part of the job.
A reply appears in seconds. Someone still checks the facts, fixes the missing detail, sends it to the right place and deals with the unusual case. If that takes longer than doing the job directly, the speed of the draft has not helped the team.
Start by describing an accepted outcome. For an enquiry, that might be a correctly routed record, a draft grounded in the current policy and an explicit flag where information is missing. “Generate a reply” is only one step toward that outcome.
What changed our own approach
In our internal automation work, a routine could keep running without producing useful accepted work. We stopped treating repeated execution as proof of value. The test became what completed correctly and what someone still had to fix.
The business lesson is straightforward: measure the finished handoff. A green run status is not a customer served.
For a first pilot, choose a frequent, bounded task with accessible source information and a person who can judge its output. A rule or a normal form may solve it without AI. Use a model where interpreting varied information is actually part of the work.
One enquiry. Four handoffs.
A fictional service business receives “Can you deliver by Friday?” Change the evidence available and see where the workflow should stop.
Keep the original request and its source.
Match the request to the current facts.
Draft, clarify or flag the exception.
A person owns the customer promise.
Illustrative workflow, not a live customer record. These controls show the decision boundary; they do not send a message.
Keep the boundary easy to explain.
A good candidate has a clear input, an observable output and a manageable exception path. For example: classify enquiries into agreed categories and prepare a sourced reply for review. The initial boundary ends before sending or making a commercial commitment.
Write down who owns the source, who reviews the output and what happens if the source is missing. If you cannot answer those questions, fix the process before asking AI to run it.
| Candidate | A useful first boundary |
|---|---|
| Customer enquiries | Draft and route; a person approves promises and sends. |
| Document intake | Extract fields with source locations; flag missing or conflicting values. |
| Internal reporting | Prepare a summary linked to current records; an owner checks interpretation. |
| Website updates | Make a preview of a bounded change; check the actual result before release. |
Anthropic advises beginning with the simplest design that works. Its evaluation guidance also separates an agent's actions from the outcome they produce. For a business owner, that means checking the destination record or delivered result, not only reading the generated explanation. Building effective agents; Evaluating agent outcomes.
Does this actually save your team time?
Start with the fictional example, then use measurements from your own pilot. The calculation includes the work that comes back to people.
Net human time saved per week
- Before AI
- Review
- Exceptions
- Upkeep
- Human time after AI
- Setup time recovered in
This is a human-time estimate, not a financial return calculation or a promise. It excludes software fees, model costs and elapsed machine runtime. Setup recovery assumes your weekly volume and quality hold. Measure quality separately.
Give the exceptions their own number.
In the starting example, fifty jobs at eight minutes take 400 human minutes a week. After AI, reviewing every item takes 100 minutes, exceptions take another 120 and upkeep takes 30. That leaves 150 minutes saved, or 2.5 hours. Recovering eight hours of setup would take 3.2 weeks at that rate.
Now raise the exception rate in the tool. The attractive draft speed stays the same, but the business result can turn negative. That is the point of the exercise: find the hidden work before expanding it.
Do not compensate for a wrong customer promise by assigning it a few extra minutes. Define hard quality conditions separately. A serious failure can stop the pilot even when the time calculation looks good.
See the calculation
Before = weekly items × current minutes per item.
After = (weekly items × review minutes) + (weekly items × exception percentage ÷ 100 × extra exception minutes) + weekly upkeep.
Net hours saved = (before − after) ÷ 60. Setup recovery in weeks = setup hours ÷ positive weekly hours saved.
If the net saving is zero or negative, there is no setup recovery under those assumptions.
Agree what would make you keep it.
Before the pilot, name the job owner and collect a baseline from real work. Keep a record of the source input, the result, whether it passed, the correction time and any exception. Use representative cases rather than only the easiest ones.
Track first-pass acceptance separately from eventual completion after correction. Divide each by all eligible items, including failed or missing outputs. Otherwise a team that rescues every bad draft can make weak automation look successful.
A short pilot might start with twenty historical examples followed by supervised live work. That is a learning exercise, not statistical proof. Include difficult and out-of-scope requests, and repeat important cases because model outputs can vary.
- Keep it small. Give it the permissions needed for the pilot, with a clear review point.
- Check the destination. Confirm the right record was created or the right output delivered.
- Count all human work. Include checking, repairs, monitoring and training.
- Keep, revise or stop. Expand only when the accepted outcome and the actual workload justify it.
Re-test when the workflow, source policy, tools or model changes. Preserve the previous working process so the team has a way back. The point is useful work completed reliably, not the largest number of automated steps.
A good first pilot gives you a decision, even when that decision is “not yet”.
Run the smallest useful experiment.
Need to choose the AI tool before running the pilot? Use the comparison guide and shared-context template.
Sources and scope
Checked 28 September 2026. The calculator and enquiry example are teaching tools with illustrative starting values. They are not DLVX client results. The operational lesson describes our internal practice without publishing private records.
- Anthropic: building effective agents, selecting a simple workflow and defining success.
- Anthropic: demystifying evaluations for AI agents, evaluating outcomes and the actions behind them.
- OpenAI: evaluation best practices, representative cases and continuing checks.