Every vendor deck contains a percentage. Productivity up by some impressive figure, hours saved per week, payback in weeks. Almost none of it survives contact with a real business, for three reasons: the baseline was never measured, the saved time was never converted into anything, and the cost of getting the tool working was left out entirely.
You can do better than that with a page of arithmetic, and the discipline of doing it will stop you buying at least one thing you were about to buy.
Everything reduces to four numbers, and the hard part is that three of them must be measured before you start.
1. The baseline. How long does the task take today, how often does it happen, and who does it? Measure this for two weeks with a timer, not from memory. Memory inflates unpleasant tasks and shrinks pleasant ones.
2. The full cost. Subscription fees, per-use charges, and the setup time it takes to get from purchase to working. Include the hours spent learning it and the hours spent fixing it in month two. Setup is the cost people consistently forget, and it is often larger than the subscription.
3. The new state. After the change, how long does the task take, including the time spent checking the output? Verification is real work and belongs in this figure, not outside it.
4. What happened to the freed time. This is the one that decides whether any of it counted.
If a tool saves your office manager six hours a month and your office manager is salaried, you have not saved a penny. You have created six hours of capacity. That capacity becomes money in exactly three ways: those hours go into revenue-generating work, you avoid a hire you were about to make, or you reduce paid hours.
If none of those three happen, the honest ROI on the time saving is zero, and the tool must justify itself on quality or speed instead. There is nothing wrong with that. A faster quote turnaround might win work even if nobody's hours change. But say which claim you are making, because they are measured differently.
The same test applies to revenue claims. "More leads" is not a result. More leads that were answered, quoted and closed is a result. If the extra volume simply sat unworked, the tool generated cost.
The calculation is deliberately plain. Over a chosen period, usually a quarter:
Use loaded cost, not salary. Wages plus employer taxes plus benefits, divided by actual worked hours. It is typically well above the headline rate, and using the headline rate flatters every calculation you will ever run.
Be conservative on the benefit side and generous on the cost side. If the answer is still clearly positive, you have something. If it only works when you assume every saved minute becomes billable, you do not.
Long enough to get past the novelty, short enough to stop a bad decision.
One control worth adding: keep one person or one location on the old method for the quarter. It is the only cheap way to tell the tool's effect apart from a busy season.
Sometimes the number comes out negative and that is a good outcome, because you found out in a quarter rather than a year. Cancel, write down why, and move on. The usual reasons are that the task was not frequent enough to matter, that checking the output took as long as doing the work, or that the tool solved a problem the business did not actually have.
The single best predictor of a positive return is task frequency. High-volume, repetitive, low-consequence work pays back. Rare, high-stakes, judgement-heavy work almost never does. If you want a view on what the ongoing spend looks like before you start, our note on AI marketing costs for small businesses lays out the recurring side.
Ninety days is a fair first assessment for a single, well-defined task. Anything shorter is measuring novelty, and anything longer usually means nobody will ever run the numbers. Expect the first month to look worse than your baseline while people learn the tool, and judge on the final six weeks of the quarter.
Use the loaded hourly cost of the person whose time was freed, meaning wages plus employer taxes and benefits divided by hours actually worked. Then apply a test: count only the hours that genuinely went into revenue work, a hire avoided, or reduced paid hours. Hours that simply evaporated into the day are worth nothing.
Setup and training hours, the time spent checking output, and ongoing maintenance when something breaks or a process changes. Verification is the most commonly omitted, and on tasks needing careful review it can consume most of the saving. Integration work and the cost of any parallel process kept running during the trial also belong in the total.
High frequency, repetitive, low consequence if slightly imperfect, and easy for a human to check quickly. Drafting replies, summarising, formatting, classifying and first-pass writing all fit. Rare, high-stakes work requiring judgement and accountability rarely pays back, because the checking effort matches the effort of just doing it properly.
Want this handled for you?
Get a free audit of your website, Google reviews, and local SEO — we’ll show you exactly where you’re losing customers. Delivered in 24 hours, no sales call.
Get my free audit → or book a 15-min callWe help local businesses in Stamford, Greenwich, Norwalk, and Fairfield County implement AI marketing that generates real results.
Get Your Free AI Marketing Audit →