AI Customer Support Setup: Stop Guessing, Flag Unknowns
The fear is real: AI support sounds great until it confidently tells a customer the wrong refund policy. Here is the operational playbook for structuring your knowledge base, automation rules, and review thresholds so your AI drafts on what it knows and flags everything else instead of making things up.
The fear is not irrational. AI support confidently tells a customer the wrong refund window, or invents a policy that does not exist, and now you have a trust problem that is harder to fix than the original ticket. Here is the thing though: most AI support failures are not a model problem. They are a setup problem. And setup problems are solvable.
Why AI support goes wrong (and it is not the model)
Klarna reversed course on AI customer service after leaning in hard. The public story was messy. But the underlying lesson was not that AI support is bad. It was that confident wrong answers erode trust faster than slow answers do. A customer who waits two hours for a human reply is annoyed. A customer who gets a fast, confident, wrong answer about their refund feels deceived.
Most teams I have talked to hit one of three failure modes. First: the AI answers outside its knowledge, pulling from training data or guessing instead of sticking to what you actually told it. Second: auto-send is turned on before anyone has reviewed enough drafts to know if the AI is ready. Third: there is no escalation path, so when the AI is uncertain it either sends anyway or drops the message entirely. All three of these are setup failures, not model failures.
The good news is that all three are fixable before you ever send a reply to a customer. That is what this post is about.
The short version if you are in a hurry
If you are skimming: draft-and-approve is the correct default posture, not auto-send. Your knowledge base should be scoped tightly to your actual policies and nothing else. Below a confidence threshold, the AI routes to human review instead of auto-sending. And flag-and-escalate is not a fallback for when things go wrong. It is the design. Build it in from day one.
The rest of this post is the how. If you want the full playbook, keep reading. If you want to see how Trigli handles this specifically, the setup is at trigli.com.
Step one: build a knowledge base that has edges
A vague knowledge base produces vague answers. And vague answers from an AI are dangerous because they sound confident even when they are not. The fix is to give the AI a knowledge base with clear edges: here is what I know, here is where my knowledge ends.
What to include: your actual return and refund policy with specific numbers (30-day window, not 'we are flexible'), your current pricing page, your shipping carriers and estimated timelines, your FAQ with real answers, and any product-specific docs that customers ask about regularly. A 200-word return policy doc with specific numbers beats a 2,000-word brand story for support accuracy. In my experience, every time.
What to leave out: anything you are not sure is current, anything that changes weekly without a process to update it, and anything that is aspirational rather than operational. If your policy doc says 'we aim to process refunds within 3-5 days' but the actual timeline is 10 days right now, that doc will produce wrong answers. Fix the doc or leave it out.
Trigli takes knowledge base inputs three ways: PDF uploads, paste-in text, and URL imports. The AI is designed to answer from those sources and flag messages when it is uncertain rather than guess. It does not browse the open web. That constraint is a feature, not a limitation. The narrower the knowledge base, the more reliable the answers. You can read more about the broader setup approach at /blog/ai-email-autoresponders-small-business.
Step two: set your automation rules before you touch auto-send
Automation rules are how you tell the AI what to do with each category of incoming message. The four actions are: auto-send (the AI sends the reply without human review), draft (the AI writes a reply and queues it for your approval), human review (the message gets flagged for a person immediately), and skip (the message is ignored by the AI, useful for spam or internal threads).
Week one rule: nothing goes on auto-send. Everything goes to draft or human review. This is not being overly cautious. This is how you build the data you need to know when auto-send is actually safe.
A practical rule set for a small e-commerce team to start with: order status questions on draft, refund requests on human review, shipping questions on draft, billing disputes on human review, and obvious spam on skip. Priority ordering matters here. More specific rules fire first. Your catch-all rule sits at the bottom and catches everything that did not match a more specific category.
Only move a category to auto-send after you have reviewed 20-30 drafts in that category and they were consistently right. Not mostly right. Consistently right. The goal is to shrink your review pile over time, not to skip the review phase entirely. See /blog/automate-customer-support-without-losing-human-touch for more on how to think about what should and should not get automated.
Step three: the confidence threshold is your safety floor
The confidence threshold is the mechanism that stops the worst cases from reaching customers. Here is how it works: the AI scores its own certainty on a reply. Below a set threshold, the message routes to human review instead of auto-sending, regardless of the category rule.
Setting the threshold is a judgment call. Too low and low-quality replies slip through. Too high and you are reviewing everything manually, which defeats the purpose. The right starting point for most small teams is conservative. You can loosen it as you build confidence in the knowledge base and the drafts.
The threshold does not guarantee accuracy. It is a filter that catches the obvious misses. An AI can score high confidence on a reply that is still wrong if the knowledge base has bad information in it. The threshold stops the AI from sending when it is uncertain. It does not fix a bad knowledge base. That is why step one matters.
Step four: design your escalation path before you need it
Escalation is not failure. It is the system working. A message that gets flagged, assigned to a human, and answered correctly is a better outcome than a message that gets auto-sent with a wrong answer. Design the escalation path before you need it, not after something goes wrong.
A good escalation looks like this: the AI flags the message, assigns it to a team member, and adds an internal note explaining what it knew and why it was uncertain. The human is not starting cold. They have context. They know what the AI tried to do and where it got stuck. That is a faster handoff than a raw forwarded email.
SLA tracking on escalated tickets matters here. A ticket that gets flagged for human review and then sits for three days is worse than no automation at all. The customer does not know the AI flagged it. They just know nobody answered. Trigli tracks SLAs with overdue alerts to reduce the chance escalated tickets go unanswered. More on response time targets at /blog/reduce-customer-support-response-time.
The categories most likely to go wrong
Pricing and billing questions: keep these on human review until your pricing docs are locked, current, and you have reviewed at least 20-30 drafts in this category. Pricing changes break AI answers fast.
Refund and return edge cases: the AI knows your policy but it does not know the context of this specific order. A customer who has ordered 15 times and wants an exception is different from a first-time buyer asking the same question. The AI cannot make that call. You can.
Legal or compliance questions: skip or human review, no exceptions. Do not let an AI answer questions about data privacy, liability, or anything that could have legal weight.
Account access or security: human review, always. An AI should never be the one handling password resets, account verification, or anything involving account security.
New product launches: your knowledge base is always behind until you update it. Route new-product questions to draft until the docs are in and you have reviewed the drafts. The AI will try to answer from what it knows, and what it knows is your old product line.
How to review drafts without it becoming a full-time job
Batch review once or twice a day beats real-time monitoring for small teams. Set a time in the morning and one in the afternoon. Review the queue. Approve what is right, edit what is close, reject what is wrong. The goal is 15-20 minutes per session, not a running tab you check every hour.
When you review a draft, look for three things: did it cite the right policy, is the tone right for your brand, and is the specific number or date correct. Tone and number are the two places AI drafts most often go slightly wrong.
When a draft is wrong, update the knowledge base, not just the draft. If the AI got the return window wrong, find the doc that has the wrong number and fix it. Fixing the draft fixes one email. Fixing the source fixes every future email in that category.
Per-message good/bad feedback is a lightweight signal worth using. It takes two seconds per message and the AI gets better at your specific patterns over time. Use it.
A realistic timeline: what week one through week four looks like
Week one: connect your inbox, upload your docs, set everything to draft or human review. Do not touch auto-send. Review every draft that comes through.
Week two: review drafts daily. You will find gaps in your knowledge base. Fill them. You will find categories where the drafts are consistently good. Note them. Still do not touch auto-send.
Week three: pick one or two low-risk categories where the drafts have been consistently right. Order confirmations, basic FAQ answers, shipping timeline questions. Move those to auto-send. Watch them closely.
Week four: check the quality of what auto-sent. If it held up, expand. If something slipped through that should not have, pull that category back to draft and figure out why. This is normal. It is not a failure.
The honest message: it takes a few weeks to tune, not a few minutes. Anyone who tells you AI support is plug-and-play on day one has not run it. The setup described here is what actually works. See /blog/cost-of-support-rep-vs-ai for context on what the economics look like once it is running.
What Trigli does specifically to prevent wrong answers
I built Trigli to solve this exact problem. Here is what it does and how it maps to the steps above.
Answers from your docs. You upload your policies, your FAQs, your pricing page. The AI is designed to answer from those sources and flag messages when it is uncertain rather than guess. It does not browse the open web. When it is uncertain, it routes the message for review rather than sending something it is not confident about. That is the design.
Draft-and-approve is the default mode. Not auto-send. You have to actively choose to move a category to auto-send. The default posture is human in the loop.
Confidence-based routing with a safety floor. Below the threshold, messages route to human review instead of auto-sending, regardless of the category rule. The two systems work together. The category rule says what to do in ideal conditions. The confidence threshold catches the cases that are not ideal.
Reply style scan. Before drafting, Trigli reads your past sent emails to learn your tone. Drafts sound like your team, not a generic support bot. This matters more than most people expect. A technically correct answer that sounds like it came from a robot still damages the relationship.
Honest limitation: Trigli is not magic. A thin knowledge base still produces thin answers. The tool is only as good as the docs you give it. If you upload three bullet points and expect the AI to handle every edge case, it will not. Put in the work on the knowledge base and the rest follows.
What Trigli does not do (so you can plan around it)
No phone or SMS support. Trigli handles email (Gmail and Outlook via OAuth 2.0) and chat. If you need voice or text message support, you will need a different tool for that.
No native Shopify actions. The AI can reference your refund policy and tell a customer how the process works. It cannot process the refund inside Shopify. That step still requires a human.
No per-agent productivity dashboards. If you need reporting on tickets resolved per agent per week, Trigli does not have that yet. You can see ticket status and SLA tracking, but not agent-level performance metrics.
No skill-based agent routing. Trigli routes by ticket category to an action. It does not auto-assign tickets to specific agents based on tagged expertise. If you have a team where billing questions always go to one person and technical questions go to another, you will set that up manually.
No multi-brand routing inside one account. One Trigli account equals one brand. If you run multiple brands, you would need separate accounts.
Honest summary: AI support that flags what it does not know is not a compromise
The teams that get burned are the ones that turn on auto-send on day one and walk away. They come back a week later to find out the AI has been confidently answering questions about a return policy that changed three months ago. That is not an AI problem. That is a setup problem.
Draft-and-approve plus a tight knowledge base is not the cautious option. It is the correct option. It is how you get to 60-70% of repetitive volume handled automatically while keeping the risky stuff in human hands. The goal is not to replace human judgment on hard questions. It is to stop humans from spending time on questions that have a clear, documented answer.
A support rep costs $3-4K a month. AI that handles the repetitive 60-70% and escalates the rest costs $149 a month on the Growth plan. The math works. But only if the setup is right. A badly configured AI tool at $149 a month is not cheap if it is sending wrong answers. The invoice looks fine. The customer trust damage does not show up there.
Trigli has a free tier: 50 emails, 25 chats, and 10 tickets at $0. It is enough to run through week one and two of the timeline above and see whether the setup described here actually works for your volume and your docs. If the fear of wrong answers has been the thing stopping you from trying AI support, the setup described here is how you solve that.
Start with the free tier. Upload your docs, connect your inbox, and review a week of drafts before you commit to anything. If it is not working, no hard feelings.
Related reading
- Draft-first AI email: why "AI writes, you approve" wins the first 90 days
Every AI support tool wants to auto-send on day one because it demos better. For a real inbox, that is backwards. The first 90 days should be draft-first: the AI writes, you approve, and every approval teaches you where the AI is trustworthy and where it is not. Here is what that arc actually looks like week by week.
- Open-Source Customer Support: Self-Host vs SaaS Break-Even
Open-source customer support tools like Chatwoot and Papercups are genuinely good software. But free to download is not free to run. Here is the honest break-even math for a 1-10 person team deciding whether to self-host or just pay for a hosted tool.
- SLA Tracking for Small Teams: A No-Fluff Setup Guide
You have outgrown a shared Gmail inbox but a full Zendesk rollout feels like overkill for a team of three. Here is a practical guide to defining SLA tiers, assigning ticket priority without a dedicated ops person, and making sure nothing gets buried when everyone is wearing five hats.