The standard vendor case compares a fully loaded salary against a monthly API bill and presents the difference as savings. It is a compelling slide and the arithmetic is incomplete in ways that matter.
The real calculation has more terms, several of which are negative, and the honest version is still frequently positive. It is just positive for different reasons and by different amounts than the sales deck suggests.
The costs that get left out
Build cost
The obvious omission. Getting a system into production means process discovery, integration work, prompt and behaviour design, testing against real cases, and deployment. For a moderately complex process this is a substantial project, and the integration portion is routinely underestimated because legacy systems resist in ways nobody predicts.
The exception tail
This is the term that most often determines whether a case is real. A system handling eighty percent of cases autonomously leaves twenty percent for humans, and those are the difficult twenty percent.
If the straightforward cases took four minutes and the difficult ones take twenty-five, removing the easy cases removes less cost than the volume suggests. Worse, the remaining work is now uniformly demanding, which affects both throughput and retention.
Model this properly. Take real case data, separate handling time by complexity, and calculate the saving on the cases the system will actually handle rather than on average handling time across all cases.
Oversight cost
Someone reviews sampled output, investigates flagged cases and monitors quality. This is ongoing operational cost, and it does not fall to zero as confidence grows because that is precisely when quality drift becomes hardest to notice.
Maintenance
Models change, source systems change, processes change, edge cases surface. A deployed system requires ongoing attention from someone who understands it. Budgeting zero for this produces systems that degrade quietly over twelve to eighteen months.
Failure cost
Errors have consequences, and at machine speed a systematic error affects many cases before detection. Expected annual cost of error, at realistic error rates and realistic detection lag, belongs in the model even though it is uncomfortable to estimate.
A reasonable rule when reviewing a business case: if the model shows payback in under three months, it has almost certainly omitted the exception tail, the oversight cost, or both. Genuine cases in this category typically show payback across quarters, not weeks, and remain worth doing.
Where the value genuinely is
Having discounted the naive case, the real benefits are worth stating precisely, because they are frequently understated in the very deck that overstates the cost saving.
Capacity without proportional cost
The most economically significant effect in most deployments is not reducing cost at constant volume. It is handling substantially more volume without adding cost. For a growing business, avoiding hires is often worth more than reducing headcount, and it carries none of the organisational damage.
Response time as revenue
In several categories, speed converts directly. Quoting within an hour rather than three days changes win rates measurably. Responding to an enquiry in minutes rather than the following morning changes conversion. This value is genuine and rarely appears in cost-focused models, because it sits in a different column of the P&L.
Coverage outside working hours
A system operating overnight and at weekends provides coverage that would otherwise require shift patterns. The comparison is not against a salary, it is against the cost of staffing hours you currently do not staff at all.
Consistency
Variance between operators, and within an operator across a day, is a real cost that most businesses never quantify. Systems are consistent, which is valuable where consistency matters, and it is worth noting that consistent means consistently wrong when the specification is wrong.
Building an honest model
A defensible case has roughly this shape.
Measure the baseline properly. Volume, handling time separated by case complexity, current error rate, current cost per case including overhead. Most organisations do not have these numbers and are surprised by them.
Estimate the autonomous rate conservatively. Take a real sample of cases, and assess how many could complete without intervention. Halve your instinct. Teams are consistently optimistic here.
Cost the remaining work at its actual complexity rather than at the average.
Include run costs in full. Model usage at realistic volume including retries and failures, infrastructure, oversight time, and a maintenance allowance.
Add the revenue-side effects where you can evidence them: conversion improvement from response time, volume growth accommodated without hiring.
Subtract expected error cost.
What emerges is usually a slower payback and a more durable case than the vendor version, and it survives scrutiny from a finance director, which the vendor version generally does not.
Industry effects worth watching
Two broader dynamics matter for anyone modelling this over several years.
Competitive erosion of the advantage. Early adopters capture genuine advantage. As adoption spreads, the gain moves from advantage to table stakes, and pricing in the category adjusts to reflect lower cost of service. Business cases assuming margin capture indefinitely tend to be wrong; the saving increasingly gets competed away to customers.
Model costs falling. Inference costs have declined substantially and continue to. This makes marginal cases viable over time and argues for building the integration and governance capability now, since that portion does not get cheaper. The expensive part of these systems was never the model.
The distributional question
Worth stating plainly because business cases tend to elide it. If productivity gains are real, someone captures them. They can go to customers as lower prices, to shareholders as margin, to employees as higher-value work and better conditions, or to growth as expanded capacity.
That is a decision, not an outcome. Organisations that make it deliberately generally do better on retention and on the quality of what they build, because people who believe the gains will be taken from them do not help build the thing that takes them.