Draft-first AI email: why "AI writes, you approve" wins the first 90 days
Every AI support tool wants to auto-send on day one because it demos better. For a real inbox, that is backwards. The first 90 days should be draft-first: the AI writes, you approve, and every approval teaches you where the AI is trustworthy and where it is not. Here is what that arc actually looks like week by week.
The first time you connect an AI to your support inbox, you have no idea how good it is. Not really. You have a demo, a marketing page, and a hunch. What you do not have is a single reply it has written to one of your actual customers about one of your actual policies. So the question is not "should I trust this AI." You cannot answer that yet. The question is "how do I find out whether to trust it without a customer paying for my education." Draft-first is the answer, and the first 90 days are when it matters most.
What draft-first actually means
Draft-first is simple. The AI reads an incoming email, writes a reply, and parks it as a draft in your inbox. Nothing goes out. You open the draft, read it, edit if you want, and hit send yourself. The customer gets a reply that a human approved, every single time. The only thing the AI did on its own was write.
The alternative is autonomous send. The AI reads, writes, and sends, all without you seeing it first. That is the mode most tools push you toward, because "the AI just answered your customer for you" is a better sentence to put on a landing page than "the AI wrote a draft and you approved it." One of those sounds like the future. The other sounds like extra work.
But for a new inbox, autonomous-first has the incentives exactly backwards. It sends the replies you have not learned to trust yet, and it hides the ones that went wrong until a customer replies to tell you. Draft-first does the opposite. It shows you every reply before it costs you anything.
Why the first 90 days are different
A support inbox that has been running AI for a year is a different animal from one on day three. After a year you know which question types the AI nails, which ones it fumbles, where your documentation has holes, and how your customers react. On day three you know none of that. You are guessing.
The whole point of the first 90 days is to convert guessing into knowing. Every draft you review is a small, cheap experiment. Did the AI get the return window right? Did it match your tone or sound like a call center? Did it invent a policy you do not have? You find out by reading, not by waiting for a complaint. Ninety days is roughly how long it takes a small inbox to see enough variety, including the weird edge cases that only show up once a month, to actually know where the AI stands.
Turn on full automation during that window and you throw the experiment away. You stop seeing the replies, so you stop learning from them, so you never build the confidence that would let you automate for the right reasons. You just automated on faith. That is the thing to avoid.
The draft is a training tool, not only a safety net
Most writing about draft mode frames it as a seatbelt. Keep the human in the loop so the AI cannot hurt you. That is true, but it undersells the point. The draft is also the fastest feedback loop you will ever get on your own support operation.
When you read a draft and it is perfect, that tells you the AI has good documentation for that question and you can probably automate it later. When you read a draft and you have to fix one specific fact, that tells you your docs are wrong or ambiguous on that fact, and you can go fix the source. When the AI writes a confident, fluent, completely wrong answer, that is the most valuable draft of all, because it shows you a gap you did not know you had before a customer found it.
Autonomous mode deletes this signal. The replies still go out, some of them still wrong, but you never see the pattern. You lose the single best diagnostic you have for improving your knowledge base. The teams that end up with genuinely good AI support are almost always the ones who spent their first months reading drafts and fixing the docs behind them.
A week-by-week plan for the first 90 days
Here is the arc I would actually run. It is not rigid. Move faster if your inbox is simple, slower if it is not. But the shape holds.
Week 1: everything drafts, and you read all of it
Auto-send off. Completely off. The AI drafts every reply and you approve every one before it goes out. Your only job this week is to read carefully and notice your own reactions. When a draft makes you nod, note the question type. When a draft makes you wince, note that too, and note exactly what was wrong: a bad fact, a wrong tone, a policy the AI made up, an over-promise.
You are building two lists without trying to. A list of question types where the AI is consistently right, and a list of places where it is not. Do not automate anything yet. A week is usually not enough volume to trust a pattern, but it is enough to start seeing one.
Weeks 2 to 4: fix the docs, keep drafting
By now you have seen the AI answer the same questions several times. The wince list from week one is mostly a documentation problem in disguise. If the AI keeps getting your refund window wrong, your refund policy is probably vague or missing. Fix the source document, and the drafts improve on their own, because the AI answers from your docs, not from memory.
This is the least glamorous and most important part of the whole 90 days. The quality of AI replies is mostly a function of the quality of the documents behind them. Tight, specific, non-contradictory docs make every AI tool look good. Sparse or conflicting docs make every tool struggle, including the expensive ones. Weeks two through four are when you do that work, and draft mode is what shows you exactly which documents to fix.
Month 2: turn on auto-send for the safe, boring stuff only
Now you have earned the right to automate a little. Pick the question types where you saw the AI right basically every time and where a wrong answer would be low-stakes anyway. Business hours. A link to your pricing page. Standard shipping timeframes. Password reset instructions. The deterministic questions where a competent new hire could answer correctly on day one with just your FAQ in front of them.
Leave everything else in draft mode. Anything involving money, anything emotional, anything where the customer is clearly upset, anything the AI has not seen enough of yet. The rule for month two is to automate the questions you are bored of answering, not the questions you are nervous about. If you feel a flicker of nervousness about auto-sending a category, that flicker is data. Keep it in drafts.
Month 3: review your auto-sends and expand deliberately
A month into partial automation, go back and read what the AI sent on its own. Not the drafts, the actual autonomous sends. Are they still clean? Did any of them make you wince in hindsight? If a category is holding up, you can add an adjacent one. If a category slipped, pull it back into draft mode. There is no shame in demoting a category. It is cheaper than a bad reply.
By the end of month three you should have a real, evidence-based split: a set of categories running autonomously that you have watched perform, and a larger set still in draft mode that either needs more data or will always need a human. That split is yours, earned from your own inbox, not borrowed from a vendor benchmark. That is the entire goal of the 90 days.
What to measure while you do this
You do not need a dashboard for this. A notes file and a bit of honesty will do. Three things are worth tracking.
- Edit rate. Of the drafts you reviewed, how many did you send unchanged versus how many you had to fix? A category you almost never edit is an auto-send candidate. A category you always edit is telling you the docs are wrong or the question needs judgment.
- Wince rate. How often did a draft make you glad you were reviewing? These are the ones that would have been bad autonomous sends. If the wince rate for a category is above zero, it is not ready to automate, no matter how good the average looks.
- Doc fixes triggered. How many documents did you improve because a draft revealed a gap? This number should be high early and fall over time. When it stops falling, your knowledge base is close to complete.
None of these require special tooling. They require you to actually read the drafts for a few weeks, which is the whole point. The measurement is a side effect of doing the review honestly.
The objection: is not draft mode just slower?
Yes, in the narrow sense. Reviewing a draft takes longer than not reviewing it. Let us be honest about the number rather than hand-wave it away. If you get 200 support emails a month and you are reviewing all of them at roughly two minutes each, that is about seven hours of review across a month. In month one, when everything drafts, that is real time.
But two things happen. First, the review time is not wasted, because it is also your documentation work and your calibration work rolled together. You would have had to learn where the AI is weak eventually. Draft mode just front-loads it into a form you can act on. Second, that seven hours shrinks fast. By month two, a chunk of your volume is auto-sending, so you are only reviewing the categories that still need eyes. By month three, review time for a simple inbox can be an hour or two a week. The cost is real and temporary. The alternative, a bad reply going out under your name to a customer you cannot un-send it to, is real and permanent.
There is a reason to prefer draft-first that has nothing to do with safety, too. A reply you approved sounds like you approved it. Customers can feel the difference between a support operation that is paying attention and one that is spraying automated answers. For the first 90 days, while you are also figuring out your own tone and your own boundaries, being in the loop is not a tax. It is how you keep the operation feeling like yours.
Where draft-first is not the right call
I am not going to pretend draft-first is correct for everyone forever. It is not. If you are a large operation with a mature knowledge base and months of data already, sitting in full draft mode is leaving time on the table. Automation you have genuinely earned is good. The argument here is specifically about the first 90 days of a new inbox, when you have not earned it yet.
There are also inboxes where the volume is so high and so repetitive that a short draft period, a couple of weeks rather than three months, is enough to calibrate. A narrow SaaS product with a tight feature set and clean docs gets to trust faster than a service business full of custom situations. Read your own edit rate and let it tell you when to move. The 90 days is a default, not a law.
What is almost never right is skipping the draft period entirely on a brand-new inbox because the demo looked good. The demo was a curated question on curated docs. Your inbox is neither.
How Trigli handles this, and what it does not do
Draft-and-approve is the default mode in Trigli, not an afterthought. The AI drafts inside your Gmail inbox and you approve before anything goes out. Autonomous send rules exist and you can turn them on per category or keyword, but they are opt-in by design. There is also a confidence floor below which the AI never auto-sends regardless of your settings, so a reply the AI is genuinely unsure about lands as a draft or a human ticket instead of going out as a guess. Separately, you control the threshold above that floor that decides when a confident reply is allowed to send on its own, so you are never forced to automate a category before you are ready.
Trigli learns your reply style from your past sent emails, so the drafts you are reviewing in month one already sound roughly like you rather than like a generic bot. That matters more than it sounds, because a draft that already sounds like you is a draft you are more likely to send unedited once you trust it. The chat widget bundled with the email product runs on the same knowledge base and the same routing logic, so what you learn reviewing email drafts carries over.
What Trigli does not do: it is not a full enterprise helpdesk. There is ticketing with priority, assignment, internal notes, and SLA tracking, and there are automation rules, but there are no per-agent productivity dashboards and no skill-based routing that auto-assigns tickets to specific people by expertise. No phone or SMS. No native Shopify app, so if you need the AI to pull order data and process returns inside Shopify, you will want a different tool or a custom integration. If you are a team of ten or more running a real CX org with those requirements, a full helpdesk like Zendesk is the honest recommendation. If you are a small team on Gmail that wants to run a careful, draft-first first 90 days, Trigli fits that shape well.
The short version
On a new inbox, you do not know how good the AI is, and you cannot find out for free with autonomous send. Draft-first turns the first 90 days into a cheap experiment: the AI writes, you approve, and every approval teaches you where the AI is trustworthy and where your docs have holes. Automate the boring, deterministic stuff once you have watched it perform. Keep the money, the judgment, and the emotion in draft mode until the data says otherwise. Earn the automation. Do not assume it.
For a closer look at which specific question types are safe to automate versus which should always wait for a human, the post at /blog/when-to-let-ai-reply-autonomously-small-teams walks through the categories. For the documentation side of this, /blog/build-ai-support-knowledge-base covers how to structure docs the AI can actually answer from.
Run a real draft-first test on your own inbox. The free tier is 50 emails, 25 chats, and 10 tickets a month, no card required. Read the drafts for a week and you will know more about your support operation than any demo could tell you.
Related reading
- AI Support Knowledge Base: Build One That Actually Works
The Klarna reversal, the Jeff Geerling post, the Hacker News pile-on - they all trace back to the same root cause: an AI with no grounded knowledge base confidently making things up. Here is how to fix that before it embarrasses you.
- Klarna AI Support Reversal: What Small Teams Should Learn
Klarna made headlines replacing support agents with AI, then quietly hired humans back. Founders are now asking: does AI customer support actually work, or was it hype? The honest answer is neither. AI handles the right ticket types extremely well and falls apart on the wrong ones. Here is the practical breakdown for small teams deciding how much to trust AI with their inbox right now.
- Autonomous AI Customer Support: When to Send vs. Approve
Klarna built a fully autonomous AI support system, then quietly started rehiring humans. The lesson is not that AI failed - it is that fully autonomous replies fail at the edges. Here is a practical framework for 1-10 person teams: exactly when to let AI send on its own, and when to make it wait for a human to approve first.