On this page 11 sections
Ask a business owner where their team’s time goes and you get the org chart. Ask the team and you get a list of small repetitive tasks nobody counts, because none of them individually seems worth mentioning. That gap is where most automatable work hides, and a two-week log finds it faster than a process-mining exercise. The harder half of the method is the four shapes that look automatable and are not - because building one of those is how a first automation project turns a team against the second.
Key takeaways
- A rough two-week log from three or four people beats a workshop, because workshops surface designed steps and the waste is in the undesigned ones.
- Frequency matters more than duration. A three-minute task done forty times a day beats a two-hour task done monthly, and it is far easier to evaluate.
- Four shapes are worth automating; four look automatable and are not. Recognising the second set is what protects the second project.
- Planning for 50-70% automation on a first version is realistic. Anyone promising 95% has not met your edge cases.
- Decide what the recovered hours are for before you start, or the work reads as a cost exercise and cooperation evaporates.
Who this applies to
You suspect there is automatable work in your team and you do not want to spend three months proving it. Also relevant if a previous automation project landed badly and you are deciding whether to try again.
Start with a two-week log, not a workshop
Ask three or four people to keep a rough log for two weeks: what they did, roughly how long, and one word for how it felt - judgement, routine, or chasing.
That third word does most of the work. “Chasing” - waiting for an approval, hunting for a file, asking someone to confirm something - is almost always pure coordination overhead, and it is invisible in every formal process map because it is not a step anybody designed.
This is also why the workshop underperforms. A workshop reconstructs the process as people believe it works, which is the version with the undesigned steps removed. The log catches them because they happened.
You are looking for tasks that are:
- Frequent - at least weekly, ideally daily
- Rule-shaped - a competent new hire could learn them in a week
- Bounded - a clear input, a clear output, a clear “done”
- Currently annoying - because you will need the team’s goodwill
Frequency matters more than duration, for a reason that is not obvious. A three-minute task done forty times a day is a better target than a two-hour task done monthly not because the hours are larger - sometimes they are not - but because you have volume to evaluate against. With forty instances a day you can measure a change in a week. With twelve instances a year you cannot measure anything, ever.
The four shapes worth automating
Most of what surfaces falls into four categories.
Reading and extracting. Invoices, purchase orders, forms, CVs, contracts, inbound emails. Somebody reads unstructured input and types structured output into a system. This is the single most reliable category, and the one traditional automation could never touch because it needed a fixed template. It also has the clearest verification story - see document extraction: what to verify.
Answering the same question. Support tickets, internal IT and HR queries, order status, policy lookups. The answer exists in a document or a record; the work is finding it and phrasing it. Grounding a system in your own content handles this well, provided you have decided what happens when the answer is not there.
Moving data between systems. Rekeying between CRM and finance. Copying an order into a spreadsheet. Often the cheapest to fix and, oddly, the last thing anyone raises, because it feels too obvious to complain about.
Sorting and routing. Triaging tickets, assigning leads, flagging orders for review, prioritising a queue. Small decisions made hundreds of times, where consistency matters more than brilliance.
The four that look automatable and are not
Just as useful to recognise early.
Rare and high-stakes. A decision made twice a year with serious consequences gives you no volume to evaluate against and no tolerance for error. Leave it with a person.
Undocumented judgement. When the rule lives entirely in one experienced person’s head and they cannot articulate it, the first project is not automation, it is getting the rule written down. Sometimes that alone solves the problem - see better decisions, not just faster ones, where this is one of three distinct failure modes and the only one a rule usually fixes.
Genuinely relational. Renewal conversations with your largest accounts, difficult complaints, anything where the point is that a person cared. Automating these saves minutes and costs relationships.
Already broken. If the process itself is wrong, automating it locks in the mistake and makes it harder to change. Fix or delete the process first. This is the outcome worth being pleased about: concluding that three steps should be deleted and no system is needed is a good result, not a failed exercise.
Sizing the prize honestly
For each candidate, estimate three numbers, then do the arithmetic in the open.
- Volume - how many times per week
- Handling time - minutes per instance, measured not guessed
- Realistic automation rate - what fraction can plausibly complete without a human
That third number is where optimism creeps in. For a first version on a well-chosen workflow, planning for 50-70% is reasonable; the remainder escalates. Anyone promising 95% on version one is either describing a very narrow task or has not met your edge cases.
volume × handling time × automation rate gives you hours returned per week. Compare against build cost plus running cost - and note that the running cost is not zero and is not only the API bill, which is the other half of most disappointing business cases. See cost per resolved task for the metric that makes this comparable.
If the payback is longer than a year, it is probably not the right first project. Pick another and come back once the team has a win behind them.
What we do differently because of this
We run the log rather than the workshop, and we insist on measured handling time rather than reported handling time, because the two differ by a consistent direction: people under-report the frequent short tasks and over-report the memorable long ones.
The number we push back on hardest is the automation rate, and it is the number clients most want to hear a large version of. Our position is to quote the business case at the conservative rate and treat anything above it as upside. That makes our proposal look weaker than one quoting 90%, and the alternative is a project that hits its technical target and misses its business case, which is worse for everyone including us.
The habit that has produced the most value and looks least like AI work: the ranked candidate list frequently contains one item whose correct answer is “stop doing this”. We report those, they are free, and they are the reason clients let us run the exercise again.
Where we get it wrong: we have under-scoped the exception path more than once. The 30-40% that escalates needs somewhere to go and someone to handle it, and if that queue is not designed the automation moves work rather than removing it - which the team notices immediately and which poisons the next project.
When this is not worth doing
Teams under about five people doing highly varied work. The log will show that nothing recurs enough, which is a real answer and not worth two weeks to reach.
When you already know the answer. If everyone names the same task unprompted, skip the exercise and size that one.
When the constraint is not hours. If the bottleneck is a decision nobody will make, or a system nobody will replace, returning hours changes nothing.
What to do with the hours
This is the part that decides whether the second project gets approved.
If the hours come back and nothing visibly improves, automation reads as a cost exercise and cooperation evaporates. Decide in advance what the recovered time is for - a backlog that never gets touched, faster response times, or genuinely giving people their Friday afternoon back - and say so before you start.
The teams that get the most out of this work are the ones where the people doing the task were involved in choosing it. They know where the exceptions are. They will tell you, if the conversation is about removing the tedious part of their job rather than the job.
Frequently asked questions
Two weeks feels long. Can we do one?
One week works if the task cycle is daily. If anything important happens weekly or fortnightly, one week will miss it entirely and you will size the wrong thing.
Should the log be detailed?
No. A line per task with a rough duration and one of the three words. Detailed logs do not get filled in, and an accurate log of half the work beats a precise log of nothing.
What if people log strategically?
They will, somewhat, and it matters less than it sounds because you are looking for shape and frequency rather than exact totals. If the concern is serious, the tell is a log with no “chasing” entries at all.
How do we measure handling time properly?
Time ten instances with a stopwatch, or take it from system timestamps where the task lives in a tool. Do not ask people to estimate it, because estimates on short frequent tasks are consistently wrong in the same direction.
Do we need volume data before we start?
You need it before you size, not before you log. Most of it is already in a helpdesk, a CRM or an inbox and takes an afternoon to pull.
Next step
If you have a candidate task and no measured volume behind it, that measurement is the first hour of work and it decides everything after it. The business process automation engagement starts with the ranked candidate list, including the items whose answer is to stop doing them.
Related: Scoping an AI project that actually ships · Cost per resolved task · Why automations fail silently · Business process automation