AI agents in SME operations: where they work, and where they quietly fail
The pitch has changed. Two years ago AI tools were sold as assistants that would help a person work faster. They are now sold as agents that will do the work themselves: handle the inbox, qualify the leads, process the invoices, chase the debtors.
Some of that is real. Enough of it is real that ignoring the category is a mistake. But the failure mode of an agent is different from the failure mode of software, and most growth-stage businesses are evaluating agents with instincts calibrated for software.
Software fails loudly, agents fail quietly
When a conventional system breaks, it usually announces itself. The integration errors, the invoice does not sync, the report returns nothing. Someone notices, because the absence is visible.
An agent that gets something wrong produces a plausible output. It categorises the expense incorrectly, summarises the client call with a detail inverted, or replies to an enquiry with an answer that is almost right. Nothing errors. The work looks done. The mistake enters your records and is discovered later, if at all, by which time several downstream decisions have been made on it.
This is the central operational property of the category, and it determines where agents can safely be deployed. The question is not whether the agent is accurate. It is whether a wrong answer would be caught before it caused damage.
Where they earn their place
Three characteristics show up in the deployments that work in businesses of this size.
The output is checked as part of the normal workflow
Drafting is the strongest category. A proposal draft, a first pass at a client email, a summary of a long thread, a set of file notes from a recording. Someone reads the output before it goes anywhere, because reading it is already part of the job. Errors are caught by a review step that costs nothing extra.
The task is high volume and low consequence per item
Categorising inbound enquiries, tagging documents, extracting fields from supplier invoices for a human to approve, drafting meeting notes. Individually, a mistake is cheap and recoverable. Collectively, the time saved is substantial. The economics work precisely because the downside per item is small.
The underlying process is already defined
An agent applied to a process with an agreed standard performs that standard faster. An agent applied to a process where five people each do it differently learns nothing coherent, because there is no coherent thing to learn. This is the same precondition that governs AI adoption generally, and it is why the tool is rarely the reason a deployment fails.
Automating an undefined process does not produce an automated process. It produces an undefined process running faster and with less visibility.
Where they fail in ways you will not see immediately
Anything that is the system of record
If an agent writes directly into your CRM, practice management system or finance ledger without a review step, errors become facts. The record is now wrong and nothing distinguishes the wrong entries from the right ones. Cleaning this up later is significantly more expensive than the labour the agent saved.
Client-facing communication without review
An agent answering client enquiries unsupervised will eventually commit the business to something it did not intend, or answer confidently on a matter it has no basis to answer. In regulated sectors the exposure is larger than commercial embarrassment. The volume of correct responses does not offset the one that creates a liability.
Judgement dressed as a task
Deciding which debtor to escalate, which candidate to progress, which client complaint is serious. These look like rules and are not. They depend on context the agent does not hold: the client’s history, the commercial relationship, what happened last time. Automating them produces decisions that are defensible on paper and wrong in the room.
Work nobody was doing anyway
A meaningful share of agent deployments automate tasks that were being skipped without consequence. The business now generates summaries nobody reads and reports nobody actions, at a cost. The activity increases and the operation does not improve.
How to evaluate one honestly
Four questions, before any trial. What specifically will this agent do, expressed as a task a person currently performs? Who reviews the output, and is that review already part of their work or an addition to it? What is the cost of a wrong answer that nobody catches for a month? And is the underlying process defined well enough that a new employee could follow it?
If the fourth answer is no, fix that first. The definition work has value whether or not you deploy anything, because undefined manual process is already costing you in a form you are not measuring.
Then run the agent in parallel rather than in production. Have it do the work alongside the person doing it, compare outputs for a few weeks, and measure the disagreement rate. That number tells you whether the agent is ready to be trusted, and where. Businesses that skip this step do not avoid the evaluation. They run it in production, on live data, with clients as the test.