13 min readBy Tommy Dempsey

The Support Questions You Should Never Let AI Answer

Most advice about AI support argues over auto-send versus draft. There is a category underneath that argument: messages the AI should not answer at all, where a draft is not caution enough and the human should own the reply from the first read. Angry customers, exceptions to policy, and anything touching money. Here is why those three families are different, and the routing that gets them to a person.

A lot of the AI support conversation stops at one question: should the AI send on its own, or write a draft you approve? That is the right question for most of your inbox, but it skips a smaller, sharper category. Some messages should not go to the AI at all, not even as a draft you glance at and approve. They should land on a human from the first read, because the thing that makes them hard is not the answer. It is the judgment, the timing, and who the customer needs to hear from. Three families fall into this bucket: angry customers, exceptions to policy, and anything touching money. This post is about why those three are different, and how to route them to a person before the AI ever gets a turn.

Draft is not the same as never

Draft-and-approve is a good default, and I have written a lot about why. The AI writes, you read, you send, and most of the time that catches the mistakes. But draft mode has a quiet failure mode of its own. A draft primes you. When the AI hands you a fluent, confident reply and all you have to do is click send, you are reviewing under a nudge toward yes. On a routine order-status question that nudge is fine, because the draft is probably right and the stakes are low. On a furious customer threatening to dispute a charge, that same nudge is a trap: you are now editing the AI's framing of a delicate situation instead of deciding from scratch how to handle it.

There is a real line between two postures. One: let the AI take a first pass and I will approve or fix it. The other: this message should reach a human before any AI reply exists, because a person needs to own it from the start. The three families below belong in the second posture. Not draft. Route to a person.

A quick test: if you would be uncomfortable having the AI's draft be the anchor for your reply, do not let the AI draft it. Route it clean, so the human starts from the customer's words and not from a machine's guess at the tone.

Family one: the angry and the emotional

An angry customer is not asking a question. They are having an experience, and the reply they need is shaped by that experience more than by any fact in your knowledge base. The AI can retrieve your return policy perfectly and still make things worse, because the problem was never that the customer did not know the policy. The problem is that they feel wronged, and a crisp, correct, slightly chirpy reply, with a satisfaction check tacked on the bottom, reads as a brush-off. Technically fine, relationally awful, and the reply that ends up screenshotted.

The routing rule for this family is blunt: catch the signal early and hand it to a person. Emotional escalation shows up in words. Furious, disgusted, unacceptable, never again, done with you, ridiculous. A message carrying that vocabulary should skip the AI and route to human review, no matter how answerable the underlying question looks. You are not building a perfect sentiment classifier, just catching the obvious cases, and the obvious cases are most of them.

There is a subtler version that gets missed: the calm complaint. A customer writes, no big deal, just wanted to flag that my order arrived damaged, and moves on. The words are relaxed. The situation is not. A keyword scan will not catch this one, and neither will a naive sentiment score, because the surface is polite. This is why the whole category should be conservative. When a message describes something going wrong, even quietly, a person should read it, because the right response might be a refund or an apology that no policy document told the AI to offer.

Family two: exceptions, where the AI knows the rule but not the context

This is the family people underestimate the most, because on the surface these messages look answerable. The customer is asking about your return window, and you have one documented, so surely the AI can handle it. Sometimes. But a large share of these are not policy questions. They are requests for an exception to the policy, and an exception is a judgment call the AI is not equipped to make.

Here is the shape of it. Your policy is thirty days, and a good customer writes on day thirty-four asking for four days of grace on a single item. The AI knows the policy, so it produces a correct, polite denial. A human reads the same message, sees that this person has ordered fifteen times, and bends the rule, because keeping the relationship is worth more than four days. Same facts, opposite outcome, and the AI's version is the one that loses you a good customer. It does not know this account's history the way you feel it, or that you would rather eat a small loss than earn a bad review from someone with a following. The tell is a request for something slightly outside the documented rule: a late return, a one-time courtesy, a swap you do not normally do, a discount you did not advertise. When the ask is for an exception rather than an answer, route it to a person.

The danger with this family is that the AI sounds most confident exactly when it should not be. A clear policy and a clear question give it high certainty, but certainty about the rule is not the same as knowing whether the rule should apply here. High confidence on an exception request is a reason to be more careful, not less.

Family three: anything that touches money

Money is the family with the least ambiguity and the highest cost of getting it wrong. Refund amounts, compensation, billing disputes, chargebacks, credits, waived fees, anything where the AI is not describing a policy but effectively committing you to move money. Describing how refunds work is fine. Deciding that this customer gets one, and how much, is not.

The distinction is easy to blur, so draw it carefully. There is a difference between the AI saying, our refund policy is thirty days and here is how to start one, and the AI saying, I have refunded your order. The first is information from your docs. The second is a commitment a customer will hold you to. Even when the AI cannot literally push a refund through your payment processor, a sentence promising one creates an obligation, and a customer told they will get forty dollars back and then does not is angrier than before they wrote in.

Chargebacks and disputes deserve their own hard rule. The moment a customer mentions a chargeback, a dispute, their bank, or their card company, the message should route to a human immediately and stay there. These have a clock and a cost, sometimes a legal edge, and an autonomous reply, however polite, is the wrong move every single time. The same goes for any message that reaches for legal language: lawyer, attorney, lawsuit, the letters BBB. A person reads them, and a person decides what happens next.

How the routing actually works

Naming the families is the easy part. Getting the messages to a human reliably is the part that has to work, and there are three tools underneath it. The first is keyword-driven rules. You build a list of words and phrases that force a message straight to human review regardless of what the AI thinks the category is. Chargeback, refund, dispute, lawyer, lawsuit, cancel, furious, unacceptable, and whatever else your own inbox has taught you to watch for. A rule like this fires before the AI drafts anything, so a message carrying one of those words never gets an AI reply in the first place. This is the bluntest instrument and also the most reliable, because it does not depend on the AI understanding the situation. It depends on a word being present, which is easy to check and hard to get wrong.

The second tool is a confidence floor. When the AI generates a reply, it scores how sure it is, and below a floor you set, the message routes to a human instead of being sent, no matter what the category rule said. This catches the cases your keyword list did not anticipate: the weird one-off phrasings and the questions that sit outside your documentation. It is not a guarantee of accuracy, and I would not sell it as one, because an AI can be confidently wrong when its knowledge base has a bad fact in it. But as a net for uncertainty it catches a lot, and it turns I do not know into a routed ticket instead of a guess.

The third tool matters most in live chat, where a visitor is sitting there in real time. When the chat AI decides it should hand off, it routes to a human, and if it offers to connect the customer with a person, that offer itself should trigger the handoff rather than being an empty promise. A bot that says someone will reach out and then does not is worse than a bot that never offered. When no agent is available, the honest move is to say so and capture the customer's email so a person can follow up, rather than faking presence. The handoff should also carry context: what the AI understood and where it got stuck, as an internal note, so the human starts warm instead of cold.

Why routing beats catching it in the draft

If draft mode already puts a human in the loop, why route these families away from the AI entirely? Two reasons. One is the priming problem from the top of this post: a draft anchors you, and on a delicate reply you want to start from the customer's words rather than from the AI's opening move. The other is the rhythm of a review queue. Most drafts are fine, so you fall into a habit of approving, and the angry message and the exception request are exactly the ones that habit waves through. Routing them into a separate human-review lane breaks the rhythm on purpose. This is not an argument against draft mode, which is right for the large, boring middle of your inbox. The three families sit outside that middle, and treating them like it is how they slip through.

Building your own never-answer list

You do not need a committee for this. You need last month's inbox and an honest hour. Here is the exercise.

  1. Pull the messages from the last thirty days where you would have been unhappy if an AI had answered on its own. Not the ones it got wrong, the ones where you are glad a person handled it. Those are your seed set.
  2. Sort them into the three families: emotional, exception, money. The ones that resist sorting are usually a sign the message belongs with a human anyway.
  3. From the emotional and money piles, pull the words that show up again and again. Those become your keyword rules. Chargeback and refund will be there. So will your own product's specific angry vocabulary.
  4. For the exception pile, accept that keywords will not catch these. Lean on the confidence floor plus a habit of routing anything that reads as a request for special treatment. Late, one-time, just this once, as a courtesy. That phrasing is your signal.
  5. Write the list down and put it where your rules live. A never-answer list that exists only in your head is not a rule, it is a hope.

The list is not static. Every month you will find a message that should have been routed and was not, and each one is a new rule. Convert every miss into a rule so the same miss does not happen twice.

What Trigli does here, and what it does not

This is the exact problem I built the routing around. Trigli uses category-based automation rules with four actions: auto-send, draft, human review, and skip. You can force a category or a keyword straight to human review so the AI never drafts it, which is how you implement a never-answer list. Underneath the category rules sits a confidence floor: below it, a message routes to a human regardless of the rule. In chat, the AI escalates when it should hand off, and as a safety net, an offer to connect the customer with a person triggers the handoff on its own rather than sitting there as an unkept promise. When no agent is online, the widget captures the visitor's email so a person can follow up.

Escalations carry context, so the human is not starting from a blank forwarded email. Tickets have priority, assignment, internal notes, and SLA tracking with overdue alerts, which matters here, because a message you correctly routed and then let sit for three days is not better than a wrong AI reply. The customer does not know it was flagged. They only know nobody answered. Notifications go out across in-app, email, browser push, Microsoft Teams, and Slack, so a routed message is something a person actually sees.

What Trigli does not do: it cannot process a refund inside Shopify or your payment processor. The AI can explain how refunds work and route the request to a person, but the money movement is a human step, which for the money family above is arguably the correct boundary anyway. There is no phone or SMS support, and no skill-based routing that auto-assigns tickets to specific people by expertise yet, so routing billing questions to one teammate is a manual setup. If you are a ten-plus person CX org that needs those things, a full helpdesk like Zendesk is the honest recommendation. If you are a small team on Gmail or Outlook that wants the hard messages routed to a person reliably, that is the shape Trigli is built for.

The honest summary

Most of your inbox is answerable, and the AI plus a draft review handles it well. This post is about the part that is not. Angry customers need a human because the reply is about the feeling, not the fact. Exceptions need a human because the AI knows the rule but not the context that decides whether it should bend. Money needs a human because a commitment with a dollar figure is not a call to hand to a machine. For those three, do not draft and hope you catch it. Route it clean, to a person, before the AI ever gets a turn.

The setup is not exotic: a keyword list that forces the obvious cases to human review, a confidence floor for the cases you did not anticipate, a real handoff in chat instead of an empty promise, and the discipline to turn every miss into a new rule. For the neighboring question of when to auto-send versus draft the rest of your inbox, /blog/when-to-let-ai-reply-autonomously-small-teams walks through the categories, and /blog/set-up-ai-customer-support-stop-hallucinations covers the knowledge base and escalation setup in more depth.


Build your never-answer list on a real inbox. The free tier is 50 emails, 25 chats, and 10 tickets a month with no card required. Set the hard families to human review and watch a week of routing.

Related reading

  • What "AI Reads Your Docs" Actually Means: Chunking and Retrieval

    People upload a 60-page policy PDF, watch the AI miss a question that is answered on page 41, and assume the tool is broken. It is not. The AI never read page 41. It reads a handful of retrieved chunks, and that one fact changes how you should write every doc you give it.

  • Why Your AI Support Bot Sounds Wrong (And How to Fix It)

    Most small teams set up an AI support bot, watch it answer questions correctly, and still get complaints that it sounds off. Too stiff. Too eager. Not like us. This post is the honest guide to what actually goes wrong with AI tone, how to audit it yourself in an afternoon, and why the real fix is simpler than any conversation design framework.

  • How to Keep Your AI Support Bot Knowledge Base Current

    If your AI support bot is flagging everything for human review or confidently quoting a policy you changed three months ago, stale docs are the culprit. Here is the scrappy, low-overhead sync routine that founders and small teams actually use to keep their knowledge base current when the product moves fast.

Ready to handle support without the headcount?

Free plan, no credit card. 14-day trial on paid plans. Takes about 5 minutes to set up.