An AI chatbot is measured on how many questions it finished correctly without a person, and somebody inside the business — not the vendor who built it — has to be the one counting. Volume tells a business the widget is visible. Resolution tells it the widget is useful. Melissa Elizondo Reyes leads this area for SB Intelligence at Studio Bananas Group, where the review below is the one the firm runs on its own chatbot; it is described here as at August 2026.

A chatbot that answers confidently and wrongly produces excellent-looking statistics. Sessions rise, response times stay fast, nobody complains inside the tool — and meanwhile a steady share of customers are told something untrue and quietly go elsewhere. A claim about a system is not evidence about that system. The dashboard is a claim.

Why conversation volume is not performance

The metrics that ship pre-built in most chatbot platforms measure activity, not outcome: sessions, messages per session, average response time, and sometimes a thumbs-up rate that a small and unrepresentative fraction of users ever touch.

None of those distinguishes a customer who got what they needed from a customer who gave up. A rising messages per session figure is ambiguous in the worst way: it either means people are engaged, or it means they are asking the same thing four different ways because the first answer was unusable.

Five things to measure instead

  1. Resolution rate. Of the conversations where a user asked something answerable, what share ended with the question actually answered? The headline number, and it almost never appears on a default dashboard, because producing it requires somebody to decide what counts as resolved.
  2. Escalation quality, not escalation rate. Handing off to a human is not a failure; it is often the correct behaviour. What matters is whether the handoffs are the right ones. A bot that escalates everything is a contact form with extra steps. A bot that escalates nothing is answering things it should not be touching.
  3. Wrong-answer rate. The figure nobody measures, because measuring it requires reading transcripts. Also the only one with real downside attached: a wrong answer about pricing, availability, scope or eligibility costs more than a hundred unanswered questions.
  4. Coverage gaps. What are people asking that the bot has no source for? The most commercially useful output a chatbot produces, and the one most owners never look at — a live list of what customers want to know and the website does not say, which is a content plan nobody had to commission.
  5. Task completion, where a task exists. If the bot is meant to book, quote, qualify or route, the measure is how many of those completed, not how many were attempted.

“It sounded right” is not evidence

The most common failure in reviewing a chatbot is reading a handful of conversations, finding the tone pleasant and the answers plausible, and concluding it works. Plausibility is exactly what these systems are good at. Plausibility is not accuracy, and the failure mode is specifically that wrong answers do not look wrong.

The fix is unglamorous: somebody reads transcripts, on a schedule, against the source. Not all of them — a sample, weekly at first, checked against what the business would actually have said. Every wrong answer gets traced to a cause, and in Studio Bananas Group’s own experience building Mango, the cause is almost never the model. It is one of three things:

  1. The knowledge base did not contain the answer, so the system produced its best guess.
  2. The knowledge base contained an old answer, and the bot repeated it faithfully. A chatbot will repeat a business’s worst page with total confidence — which is the argument for being deliberate about what goes into it in the first place.
  3. The bot was never told to refuse. Anything it has not been explicitly instructed to decline — pricing it should not quote, advice it should not give, commitments it cannot make — it will attempt.

Who should own the review, and how often

Someone inside the business. The builder does not close the finding — if the only party reviewing performance is the party that built the thing, the business has a report rather than a check. Studio Bananas Group applies that rule to its own work, which is why the review above is written to be run by a client rather than by us.

A workable rhythm for a small business: read a sample of transcripts weekly for the first month, then monthly. Keep a running list of wrong answers and coverage gaps. Update the knowledge base against that list on a fixed cadence, with a named person responsible.

The named person is the part most deployments are missing, and it is not a technical role. Deciding what the bot may say, what it must refuse, which gaps get closed first and when something is switched off are business decisions with commercial consequences. In most small businesses those decisions land nowhere: the vendor owns the software, an operations manager owns the inbox, and the decisions themselves have no owner at all. That gap is the reason fractional AI leadership exists as a seat — someone accountable for the decisions, part-time, without a permanent hire. What that covers sits on AI and automation.

Ownership also becomes concrete rather than theoretical at exactly this point. A business that cannot export its transcripts and its knowledge base cannot run this review properly and cannot take the work anywhere else, which is the practical edge of who ends up owning the bot.

What to do when the numbers are bad

Bad numbers usually point to one of three fixes, in ascending order of cost:

  1. A knowledge gap — the answer is not there. Cheapest fix, and the coverage-gap list often hands over the work already in priority order.
  2. A refusal gap — the bot is answering things it should decline. Also cheap, and mostly a matter of writing the boundaries down explicitly.
  3. A design mismatch — the bot is being asked to do a job that is not a conversation. If the task is a fixed sequence of steps with a defined outcome, it is a workflow rather than an agent, and no amount of prompt tuning makes the wrong shape fit.

Replacing the model is the last thing to try and the first thing most people suggest.

Where we sit in this

Studio Bananas Group builds and runs chatbots, including its own, so the honest position is that this describes what the firm does rather than what it recommends other people do. The review above is the one it runs. The reason it runs it is that these failure modes were found in its own work, before customers found them.

A business with a bot live and no idea whether it is helping can run the diagnostic in a few hours of transcript reading, and does not need us to do it.

// The list

One useful idea, once a month.

No spam, no drip funnel, no "10x your growth" nonsense. Just one specific, usable note — and you can leave any time. Same promise as the rest of the studio.

By subscribing, you give express consent (CASL) for Studio Bananas Group to email you one update a month. We identify ourselves in every message and you can unsubscribe in one click. See our Privacy Policy.

Share this LinkedIn X / Twitter Email